September field notes
Eleven measurements from one RTX 5090 desktop, 12 to 21 September 2026, led by a correction of our own record: MiniMax M2.7's 96,000-token leg did not fail, it timed out against a launch flag we had inherited without noticing. With the flag fixed, a direct read on 19 September evaluated 85,763 tokens in 145.48 seconds of prompt time, and through the same agent framework on 21 September the model found all three planted codes at 53,584 tokens in 122.1 seconds and at 103,931 tokens in 251.2 seconds.
A roundup of short measured entries from one desktop. It opens with the correction: our first records called MiniMax M2.7 too slow to use, and the slowness was our own launch line. The same entry keeps what still stands from the first week, including a chat template that starts the model inside a thought block with no way out. The others include a maker-stated 124-billion-parameter model that held a 131,072-token window in 6,941 MiB of graphics memory, a flash model that recalled three sealed codes at 459,911 tokens, and a single trailing zero in a file header that decided whether a model could run at all. Every figure carries a label (measured arithmetic vendor stated) with the meanings defined on our method page, and every measured figure traces to a log or result file in the data package.
What this page is
Each entry is a measurement or a small set of measurements made between 2026-09-12 and 2026-09-21 on one machine: one NVIDIA GeForce RTX 5090 with 32,607 MiB of VRAM, an Intel Core Ultra 9 285K and 188 GiB of RAM, every model served by llama.cpp, one model at a time, one request at a time. Each entry names the model file and quantization, the context window and the day it ran. measured
None of this is a benchmark and none of it ranks models. Several entries exist because something surprised us: a refusal, an empty string, a number that matched its prediction to within two mebibytes. The lead entry is a correction of our own record, a genre this site keeps on purpose.
What we ran, exactly
| Context sweep, 2026-09-12 | llama-server started directly, only the window varied; speaking is the best of two warm ~200-token prose replies, reading one ~21,000-token ledger with a planted code; a 97 GB download ran throughout, so load times are upper bounds. |
|---|---|
| 256K sweep, 2026-09-15 | One long rung per model, a ~230,000-token prompt with three sealed codes at 5, 50 and 95 percent depth, card memory at load and at peak; about 340 GB of downloads ran during it. |
| Wave-two check, 2026-09-13 | Recall of three codes and a tool call at about 48,000 and 96,000 tokens through an agent framework, one run per model per arm. A derived summary ships in data/wave2/; one of its entries records a timeout as a recall miss, and CORRECTION.md beside it annotates that (section 15). |
| MiniMax M2.7 audition, 2026-09-13, and budget test, 2026-09-15 | Header and template reads, two served rungs and the probes, then the server's --reasoning-budget at 0 and 256. Both ran before the launch fix of 19 September, so their wall times are not reprinted. |
| MiniMax M2.7, 2026-09-19 and 21 | Ladders and a restart test directly against the server, then the wave-two check re-run through the same agent framework. |
| MiniMax M3, 2026-09-19 to 20 | The old file's micro-batch sweep, then the correct file, its 262,144 load ladder and a two-machine attempt. |
| Flash-Next context study, 2026-09-20 | Three sealed codes at windows of 262,144, 393,216 and 524,288, cache at f16 and q8_0, plus the rope-flag trap; one pass each, temperature 0. |
| Batch-size sweep, 2026-09-20 to 21 | --batch-size and --ubatch-size swept on ~48,000-token prompts, one run per setting: hand-copied flags at a 131,072 window on 20 September, each model's own start script at its served window on 21 September, so Ling-3.0-flash's two baselines sit on different windows (section 08). The sweep is its own study. |
| Qwen3.6-27B files and the two-box run, 2026-09-13 | A header comparison and a processor-only load; Llama 4 Maverick split across the desktop and an ASUS ROG Flow Z13 laptop over Thunderbolt. |
Card memory is the raw nvidia-smi figure in MiB, never divided by 1000.
MiniMax M2.7, re-opened
The file: Unsloth's UD-IQ4_XS build, four shards,
108,413,781,312 bytes, verified by count and byte size against the
repository tree. The header, read before any load: architecture
minimax-m2, 62 blocks, 48 attention heads with 8
key-value heads, 256 experts with 8 used, native context 196,608
tokens. measured 2026-09-13
What we recorded, and what was wrong
On 13 September, through an agent framework at a 131,072 window,
the 48,000-token leg (53,581 prompt tokens) took 1,682.9 seconds,
about 28 minutes, and the 96,000-token leg returned no answer before
the turn ended at 1,800.4 seconds. We closed the model as too slow. The cause was our
launch line: it carried -ub 128, llama.cpp's micro-batch
size, copied from MiniMax M3's start script, where no reason for it is
recorded. llama.cpp's default is 512.
measured 2026-09-13
On 19 September, at a 131,072 window with an 8-bit cache and
-b 4096 throughout, one 3,658-token prompt read at 32.3
tokens a second at -ub 128 and at 564.0 at
-ub 4096, with every expert in system memory. At 43,909
tokens and -ub 4096, prompt evaluation ran at 640.2 tokens
a second with every expert layer there, 656.5 with the experts of 59 of
the 62 layers there (the placement the installed launch script uses),
and 666.1 at 58 of 62 with 30,695 MiB on the card. Those are
prompt-evaluation rates; the whole requests, each ending in a short
answer, took 72.1, 70.1 and 69.1 seconds. One cold run each; the
batch-size study has the whole ladder.
measured 2026-09-19
At the installed placement, a direct read of a prompt seeded to the same 96,000-token target, 85,763 tokens as the server counted them, was evaluated in 145.48 seconds, 589.5 tokens a second; the whole request, including a 16-token answer, took 147.49 seconds of client wall time. measured 2026-09-19
Then the same agent framework, on 21 September, against the installed server with the fixed setting: the 48,000-token leg took 122.1 seconds at 53,584 prompt tokens and the 96,000-token leg 251.2 seconds at 103,931, with all three codes found on both legs and the tool call passed. measured 2026-09-21
Decode was not what changed: 32-token answers decoded at 9.9 to 10.2 tokens a second after the 3,658-token prompt at micro-batches from 128 to 2,048, and at 9.0 to 10.0 after 43,909 tokens at 4,096; through the framework, 8.0 on 13 September and 8.6 and 7.2 on 21 September, on completions of 160, 180 and 182 tokens. These are short answers, not prose speaking speeds. measured 2026-09-13 to 21
The native 196,608-token window fits on the one card only at a
smaller micro-batch. With every expert in system memory,
-ub 4096 failed to load at that window,
-ub 2048 aborted, and -ub 1024 loaded with
31,560 MiB on the card. Only the 3,658-token prompt was run there
(188.8 tokens a second), so deep reading at 196,608 is unmeasured.
measured 2026-09-19
A warm session survives a real server restart: 85,778 tokens saved and 85,778 restored, 1 token counted as new work, a 2.19-second restore, then a 2.17-second request; the warm-wake study has those clocks and the mechanism. On 21 September the installed launch script repeated the test: the server evaluated 1 prompt token (136.8 ms) and 16 answer tokens, total time 1.97 seconds, stop count 42,898 tokens. measured 2026-09-19 and 2026-09-21
What stands from the first week, dated
The chat template starts the model inside a thought block on every
turn and offers no switch to close it: the default,
enable_thinking: false, thinking: false and
reasoning_effort: "low", each at 65,536 (f16 cache) and
131,072 (8-bit cache), eight calls, and every one produced a reasoning
block. measured 2026-09-13, not
re-run since; every turn of the 21 September re-run also streamed its
reasoning before any answer text.
The server-side budget does not close it either. With
--reasoning-budget 0 the model still wrote 1,734 and
3,003 characters of thought on two prompts, and the second hit its
600-token cap with zero visible characters, stop reason
length; a budget of 256 gave answers and well-formed
tool calls. measured 2026-09-15, before the
launch fix and not re-run; its log ships unchanged.
The first coherence probe of the audition, given a 900-token
output budget, returned finish=length, 900 completion
tokens, 0 characters of content and 4,034 characters
of reasoning: the whole budget went on thinking.
measured 2026-09-13, 65,536 context, f16
cache. That is the finding of our Empty
Answer study, on another vendor's model served locally; it
concerns the output budget, not the launch flag, and was not re-run.
Given 4,096 tokens the model answered coherently (three sentences of
the answer are quoted in the package), and its tool calls were well
formed, parsed correctly from its own XML dialect.
The header arithmetic was always about memory: 248 KiB of cache per token, 2 times 8 key-value heads times 128 dimensions times 2 bytes times 62 blocks, so 131,072 tokens at f16 is 31.0 GiB, too much for a 31.8 GiB card beside the weights; that window needs an 8-bit cache. arithmetic Measured, the 65,536 rung at f16 took 21,232 MiB and the 131,072 rung at 8-bit 22,240 MiB, because both hold about 15.5 GiB of cache. measured 2026-09-13. Our first notes offered this arithmetic as the reason the model was slow (section 15).
MiniMax M3
Two inherited problems. M3's own start
script is where MiniMax M2.7's -ub 128 came from, and no
reason for it is recorded there either. And the file we had,
Unsloth's UD-IQ3_XXS conversion, predates llama.cpp's
support for this model's sparse attention: a tensor inventory finds
zero indexer tensors among its 948, so every M3 number
we had was a dense fallback. measured
2026-09-19
On that old file, at a fixed -b 4096, reading a
3,658-token prompt scaled with the micro-batch: 22.6 tokens a second
at 128, then 67.7, 123.1, 221.3 and 393.8 at 512, 1,024, 2,048 and
4,096. One cold run per rung. measured
2026-09-19
On the correct file, bartowski's Q2_K_L (four shards,
153,086,988,768 bytes, verified), at a 131,072 window,
-ub 2048 and an 8-bit cache: prompt evaluation at 147.5
tokens a second on 3,658 tokens; a 58,307-token recall prompt took
300.2 seconds for the whole request, including 165 generated tokens,
about 194 prompt tokens per second of wall time, which understates
pure reading; decode at 8.9 on an 82-token answer and 9.0 on a forced
128-token generation; all three codes found. At the 196,608 window
(-ub 512): 72.6 on 3,658 tokens, and the same recall took
856.3 seconds including 207 generated tokens (about 68.1 per second of
wall time), 3 of 3 again, so no rise with depth shows at that window.
measured 2026-09-20
A 262,144-token window with the sparse attention engaged fits this
card only at -ub 128: the compute buffer alone asks for
19,619 MiB at -ub 2048, the allocation fails at every
micro-batch from 2,048 down to 256, and at 128 the window loads with
31,525 MiB on the card and reads a 3,658-token prompt at 24.2 tokens a
second. We treat 256K as not viable for this model here.
measured 2026-09-20. A
split across the desktop and the laptop, with bartowski's larger
IQ3_XXS build of the same attention (167.6 GiB), loaded
at -ub 1024 and read the 3,658-token prompt at 17.58
tokens a second; speaking was not measured, because the run was
stopped during the first speaking probe. The
two-box study has that row.
measured 2026-09-20
GLM-4.7-Flash
The fastest speaker in the 12 September sweep. At 131,072 context,
all weights on the card with an f16 cache, Unsloth's
UD-Q4_K_XL build: 223.65 tokens a second speaking
and 4,447.44 reading on a 20,221-token prompt, the planted
code found, 24,814 MiB on the card, loaded in 21.2 seconds.
measured 2026-09-12
At its native ceiling of 202,752 tokens it held 28,702 MiB (28,765 at peak), loaded in 22.2 seconds, spoke at 227.43 on a short turn, and read a 189,931-token prompt at 819.70 tokens a second, finding all three codes; that 29-token three-code answer decoded at 70.48. measured 2026-09-15
Reading falls from 4,447 to 820 between the two rungs, which is why this page gives a prompt size with every reading rate.
On 21 September a larger micro-batch read a ~48,000-token prompt at most 19.5 percent faster (2,594.6 to 3,100.2 tokens a second, at a rung whose visible answer came back empty, the code only in its reasoning) and was not adopted; the batch-size study has the rungs. measured 2026-09-21
Gemma-4-26B
At 131,072 context, all on the card with an f16 cache, 211.07 speaking and 10,593.27 reading on a 21,690-token prompt, the planted code found, 20,646 MiB. measured 2026-09-12
At a full 262,144-token window it read a 230,855-token prompt at 3,799.69 tokens a second and found all three codes; it spoke at 202.64 on a short turn and decoded the 33-token three-code answer at depth at 121.36; 23,446 MiB loaded, 23,513 at peak, loaded in 23.2 seconds. measured 2026-09-15
On 21 September a larger micro-batch read a ~48,000-token prompt
faster, from 8,972.3 tokens a second at the default to a best of
12,063.4 at -ub 2048. At that best reading rung the
17-character code answer decoded at 153.8 tokens a second, against
163.7 at the default; the following rung, -ub 4096,
decoded that same short answer at 148.4 and read slower, 11,811. It
was not adopted; the batch-size study has
the rungs. measured 2026-09-21
Gemma-4-31B
This one fills the card. At 131,072 context: 73.68 speaking, 3,166.30 reading on a 21,690-token prompt, 29,844 MiB on the card with 2,763 MiB free. At 65,536: 73.73 speaking, 3,175.57 reading on the same prompt, 24,660 MiB with 7,947 free. Both rungs found the planted code. measured 2026-09-12
Speed is flat between 64K and 128K; only the room changes. A larger micro-batch measured on 21 September read a ~48,000-token prompt at most 3.6 percent faster (2,687.5 to 2,784.5) and was not adopted. measured 2026-09-21
Ling-3.0-flash
A 124-billion-parameter mixture-of-experts by its maker's count stated (the maker's figure, not our measurement). From the file header: 43 blocks, 512 experts with 8 used and 1 shared, and only 8 of the 43 blocks hold an attention cache, in a compressed form (rank 512 plus 64 rotary dimensions); the other 35 run a linear path with a fixed-size state. measured header read, 2026-09-13
That makes context cheap: 9 KiB of cache per token, 0.56 GiB at 65,536, 1.13 GiB at 131,072 and 2.25 GiB at its native 262,144, 27 times less per token than MiniMax M2.7's 248 KiB, both computed the same way from the files' headers. arithmetic
In the 13 September agent-framework check, at a 131,072 window and the default micro-batch, it held 6,941 MiB, found all three codes on both legs and made its tool calls; the first read reached first content after 506.68 seconds on a 103,807-token prompt (the history was seeded to about 95,913 tokens). A warm turn on that history reached first content after 0.96 seconds; the turn itself took 2.0 seconds. measured 2026-09-13
At its native 262,144, still at the default micro-batch, it loaded in 103.8 seconds holding 8,314 MiB (8,473 at peak), read 230,759 tokens at 206.25 a second, and found all three codes. measured 2026-09-15. The 2.25 GiB is the cache alone, by arithmetic; the 8,314 MiB is the whole model as served.
The batch-size sweep changed its reading. With
-b 4096 -ub 2048, the setting that now ships, it read
727.7 tokens a second on 47,986 tokens and 694.6 on
149,711 at the 262,144 window (21 September, 8,824 MiB at peak),
against 238.2 on 47,992 tokens at the default micro-batch at a 131,072
window (20 September). The windows differ. On one window, 262,144, a
3,007-token read went from 201.5 at -ub 512 to 444.0 at
-ub 2048, about 2.2 times.
measured 2026-09-20 to 21
A faster setting was withdrawn. At -ub 4096 the served
window read 1,173.9 on 48,084 tokens and 1,067.6 on 150,153, but one
3,007-token prompt crashes the server with a CUDA illegal memory
access. The shipped crash log is that one case, at the 262,144 window;
the same prompt passed at 2,048, 1,024 and 512 on that window, and 60
code and prose prompts and 50 ledger prompts ran with zero crashes at
either setting. The crash log and a standalone reproducer ship in the
package; the diagnostic runs are in the
batch-size study.
measured 2026-09-20 to 21
And it has a real thinking off-switch, three in the template:
enable_thinking: false, thinking_option: "off",
or the phrase "detailed thinking off" in a system message; off emits a
closed, empty thought block. measured
template read, 2026-09-13. The opposite of MiniMax M2.7.
Qwen3.8-Flash-Next
At 131,072 context, Unsloth's
UD-Q3_K_XL build with the experts in system memory:
11,366 MiB, 22.50 speaking, 195.84 reading on a
21,740-token prompt, the planted code found, about 2 seconds to load
warm. Its 65,536 preset holds 9,670 MiB and reads 246.9 on a
2,624-token prompt. measured 2026-09-12
At 262,144 it loaded in 52.4 seconds, held 15,124 MiB (15,626 at peak), read the 230,803-token prompt at 165.70 a second, about 23 minutes, and found all three codes; it spoke at 25.33 on a short turn and decoded the 32-token three-code answer at depth at 13.58. measured 2026-09-15
Five days later we pushed it past its trained window. With the rope-scaling flags set, recall held at three of three sealed codes out to 459,911 tokens, 1.75 times the trained 262,144, with time the limit: 19.8 minutes at f16 and 262,144 (the q8_0 twin took 19.4), then 31.2 and 44.6 minutes at q8_0 at 393,216 and 524,288. At 262,144 the q8_0 cache gave the same recall with 22,019 MiB at peak against 25,226 at f16. measured 2026-09-20
These runs show retrieval, not reasoning across the window: the integration question fails against a 5,974-token ledger and against every longer one; the question itself is a 526-token follow-up with the ledger cached.
One trap, measured the same day: ask for 524,288 without the rope flags and the server logs that it caps the slot at 262,144, while the card still pays for the big window: 23,816 MiB, the same as a genuine 524,288 window (23,814), against 15,394 on a plain 262,144 load (that server came up, but its probe then logged a 600-second timeout). measured 2026-09-20. Its batch-size results, which differ between the 131,072 window they were swept at and the 262,144 window it serves, are on the batch-size study.
Qwen3.6-27B
The measurement here is which file could run at all. Two GGUF
distributions bearing the Qwen3.6-27B label exist on this machine. The
one distributed through Ollama's model library declares its rope
sections as [11, 11, 10], three elements, and every
llama.cpp build here refuses it, expected 4, got 3, in
under a second, before any memory is touched. The upstream file
declares [11, 11, 10, 0] and loads. One trailing
zero in a header decided whether this model could run at
all. measured 2026-09-13
The fix was proved without touching the contended card: the
upstream file, Unsloth's UD-Q4_K_XL, was loaded
processor-only with graphics initialisation disabled, and the server
reported model loaded in 6.39 seconds.
measured 2026-09-13. This is an interaction
between a file variant and a loader's array-length requirement, not an
accusation against either project.
Qwen3-32B
The arithmetic success. Its header gives 256 KiB of cache per token
at f16, so at 32,768 context (its served rung; the file's own ceiling
is 40,960) the prediction was 18.81 GiB of weights plus 8.0 GiB of
cache plus 1.04 GiB of compute buffers, 28,518 MiB. The
measurement was 28,520 MiB, two MiB out.
arithmetic confirmed
measured 2026-09-13, Q4_K_M
The same arithmetic says a 131,072 window fits this card only with a 4-bit cache, 27.8 GiB in total; f16 and 8-bit do not fit. That rung was not run. arithmetic
One caution from the wave-two check: with thinking off this model passed, tool call included; with thinking on the same check failed, the tool call not emitted. One run each, one evening; we do not generalise it. measured 2026-09-13
Llama 4 Maverick
Llama 4 Maverick, Unsloth's UD-Q3_K_XL at 167.21 GiB
by its header and never served from one machine here, loaded split
across the desktop and the laptop at 131,072 context on 13 September
and answered at 13.09 tokens a second on a short 51-token reply: a
capacity demonstration on a model deleted the next day, with its row
and header read on the two-box study.
measured 2026-09-13
256,000 tokens, from the headers
What would a quarter-million-token window cost each model? The arithmetic column is the attention cache alone, computed from each file's header by a standalone reader that never touches tensor data (reader and outputs in the package). The measured column is the whole card at the long window on 15 September, one run per model; the context-256k study carries the full table.
| Model | Cache at ~256K, from the header | Label | Measured at the long window, 2026-09-15 | Label |
|---|---|---|---|---|
| Inkling-Small | about 7 GiB | arithmetic | 262,144: 15,744 MiB loaded, 15,814 at peak; read 230,827 prompt tokens at 115.07 t/s; 3 of 3 codes; the 27-token code answer decoded at 8.78 t/s | measured |
| Qwen3.8-Flash-Next | about 6 GiB | arithmetic | 262,144: 15,124 MiB loaded, 15,626 at peak; read 230,803 prompt tokens at 165.70 t/s; 3 of 3 codes | measured |
| Ling-3.0-flash | about 2.25 GiB | arithmetic | 262,144: 8,314 MiB loaded, 8,473 at peak; read 230,759 prompt tokens at 206.25 t/s; 3 of 3 codes | measured |
| Qwen3.5-397B | about 7.5 GiB | arithmetic | 262,144: 30,908 MiB loaded, 31,068 at peak; read 230,803 prompt tokens at 194.51 t/s; 3 of 3 codes; the 32-token code answer decoded at 17.00 t/s with speculative decoding from its own draft head (29 of 29 drafted tokens accepted) | measured |
| Qwen3.8-27B | about 16 GiB | arithmetic | no logged measurement; its field card records a 17 September serve at this window that kept no log | |
| Gemma-4-26B | not computed | 262,144: 23,446 MiB loaded, 23,513 at peak; read 230,855 prompt tokens at 3,799.69 t/s; 3 of 3 codes | measured | |
| GLM-4.7-Flash | cannot reach 262,144; its ceiling is 202,752 | header read | 202,752: 28,702 MiB loaded, 28,765 at peak; read 189,931 prompt tokens at 819.70 t/s; 3 of 3 codes | measured |
The arithmetic rows are estimates.
Qwen3.8-Flash-Next
and Ling-3.0-flash have later figures in sections 08 and 09. MiniMax
M3's answer for this window is in section 04: with its sparse
attention engaged it loads only at -ub 128.
What this does not show
- Nothing here measures model quality, intelligence or benchmark standing.
- Most entries rest on one run each. The wave-two check is one run per model per arm, the sweeps one run per rung, the budget probes one call per cell, and the batch-sweep and MiniMax M3 figures one cold run per setting.
- Header arithmetic is not a run. The 256K arithmetic column is an estimate; the measured column is the evidence.
- Load times are upper bounds. Large downloads competed for disk and page cache during the 12 and 15 September sweeps.
- Reading rates belong to their prompt sizes and windows. Section 05 shows one model at 4,447 and at 820. The batch-sweep figures are ~48,000-token prompts and are not comparable to the 12 September sweep's ~21,000-token prompts.
- Short answers are not speaking speeds. Where a decode rate comes from a planted code or an answer of a few dozen tokens, the page says so.
- The agent-framework speeds are not directly comparable to direct-probe speeds. The 21 September MiniMax M2.7 re-run went through the same framework as the original failure for exactly that reason.
- MiniMax M2.7's thinking and budget findings are dated 13 to 15 September and were not re-run after the launch fix. At its native 196,608 window only a 3,658-token prompt has been run.
- The 459,911-token result is retrieval, not reasoning across the window.
What we got wrong
"MiniMax M2.7 is slow." Our first records called
this model slow, explained it with the model's cache arithmetic, and
closed it on that basis. The speed was our own launch line,
-ub 128, inherited from MiniMax M3's start script with no
recorded reason; the cache arithmetic only ever described the memory
bill. The "why it is slow" reasoning is gone, the wall times that
depended on the flag are withdrawn or dated, and section 03 carries
the re-measurements.
The timeout that looked like a recall miss. Our
gate harness left error: null when a turn ran out of time
without an answer and counted the recall as 0 of 3, so a timeout and a
genuine miss looked identical. MiniMax M2.7's 96,000-token leg was one:
1,800.4 seconds, no answer events, written down as "l128 recall 0/3",
and that record reached this page before we caught it. The same week,
a separate run on another model that produced only reasoning and no
visible text was also scored as recall 0 of 3. On yet another model's
run the harness divided by a missing first-content time and wrote
987.9 tokens a second next to a pass; that run had 326 completion
tokens, so it was not an empty answer. The harness was fixed before
the 21 September re-run: a turn with no answer is now recorded as
UNANSWERED. The wave-two summary in the package keeps its original
values; CORRECTION.md annotates
them.
The 700-to-900 note. In the early hours of 2026-09-13, after the audition had measured an empty answer at a 900-token output budget, a working note of ours advised giving the model "700 to 900 output tokens or the answer comes back empty", quoted here word for word. The advice pointed straight at the measured failure: 900 is exactly the budget that returned zero characters. It was corrected in place the same night with a dated marker. Measure your own floor; the Empty Answer study shows how.
The "~40 t/s" verdict. The budget test's verdict paragraph reports reading at about 40 tokens a second, a rate that does not follow from the log lines in the same record; we do not print it, and the probe log ships unchanged.
The unfinished run. An earlier record of ours, written while the 256K sweep was still running, has the Qwen3.8-Flash-Next run at that window as unfinished: the read had reached 192,512 of about 230,000 tokens when the snapshot was taken. The run completed the same evening, and section 09 prints the completed measurement. When two of our records disagree, the log is the number we print.
Related pages on this site
The Empty Answer study is the parent finding for section 03's budget trap. The batch-size study covers the flag behind this page's biggest correction. The warm-wake study covers restoring a server's state from disk. The two-box study covers the desktop-and-laptop pair behind sections 04 and 12. The context-256k study carries the full 256K table. Labels are defined on the method page.
Sources and artifacts
The data package is a leak-swept selection of the run records this page cites, with derived summaries where no public primary record exists. data/README.md maps every file and redaction and opens with where the package seems to argue with itself; data/NUMBERS.md maps every figure to its file and field. If a number here disagrees with a file in the package, the file is right; tell us and we will fix the page.
Every path on our machines was replaced in full
with <REDACTED_PATH>, keeping only a model file's
own name; addresses, ports, process ids and serving aliases became
placeholders; request and fingerprint identifiers were removed from the
response bodies, which otherwise ship whole. Our session write-ups,
working notes and the per-leg records of both agent-framework checks do
not ship. Model repositories were read at the revisions downloaded on
2026-09-13 (Unsloth's MiniMax-M2.7-GGUF and Qwen3.6-27B-GGUF,
inclusionAI's Ling-3.0-flash) and 2026-09-20 (bartowski's MiniMax-M3
builds).
| 2026-09-12 | Context sweep (GLM-4.7-Flash, both Gemma models, Qwen3.8-Flash-Next) and the Flash-Next install measurements. |
|---|---|
| 2026-09-13 | MiniMax M2.7 audition. Qwen3-32B. The Qwen3.6-27B header finding and processor-only load. The wave-two check. Llama 4 Maverick on the pair. |
| 2026-09-15 | MiniMax M2.7 budget test. The 256K sweep: Gemma-4-26B, GLM-4.7-Flash, Qwen3.8-Flash-Next, Ling-3.0-flash, Inkling-Small and Qwen3.5-397B at their long windows. |
| 2026-09-19 | The MiniMax M2.7 re-launch: the -ub 128 control, the micro-batch and placement ladders, the 196,608-token load ladder, the warm restore. MiniMax M3's micro-batch sweep on the old file. |
| 2026-09-20 | MiniMax M3 on the correct file, its 262,144 load ladder and the two-machine attempt. The Qwen3.8-Flash-Next context study and the rope-flag trap. The batch-size sweep begins. |
| 2026-09-21 | The MiniMax M2.7 gate re-run and the installed script's restore test. The batch sweep concludes: Ling-3.0-flash's faster setting withdrawn; larger micro-batches measured and not adopted for GLM-4.7-Flash and the two Gemma models. |
Every claim here is dated; if today is much later than the rows above, treat version-specific findings as unverified since then.