Graphometer
Field notes · 12 to 21 September 2026 · one RTX 5090 desktop · llama.cpp · eleven entries

September field notes

Eleven measurements from one RTX 5090 desktop, 12 to 21 September 2026, led by a correction of our own record: MiniMax M2.7's 96,000-token leg did not fail, it timed out against a launch flag we had inherited without noticing. With the flag fixed, a direct read on 19 September evaluated 85,763 tokens in 145.48 seconds of prompt time, and through the same agent framework on 21 September the model found all three planted codes at 53,584 tokens in 122.1 seconds and at 103,931 tokens in 251.2 seconds.

A roundup of short measured entries from one desktop. It opens with the correction: our first records called MiniMax M2.7 too slow to use, and the slowness was our own launch line. The same entry keeps what still stands from the first week, including a chat template that starts the model inside a thought block with no way out. The others include a maker-stated 124-billion-parameter model that held a 131,072-token window in 6,941 MiB of graphics memory, a flash model that recalled three sealed codes at 459,911 tokens, and a single trailing zero in a file header that decided whether a model could run at all. Every figure carries a label (measured arithmetic vendor stated) with the meanings defined on our method page, and every measured figure traces to a log or result file in the data package.

Independent measurements. Not affiliated with, endorsed by, or connected to MiniMax, Z.ai, Google, inclusionAI, Alibaba Group, DeepSeek, Thinking Machines Lab, Meta, Ollama, Unsloth, bartowski, NVIDIA, Intel, ASUS, or the llama.cpp project.
01

What this page is

Each entry is a measurement or a small set of measurements made between 2026-09-12 and 2026-09-21 on one machine: one NVIDIA GeForce RTX 5090 with 32,607 MiB of VRAM, an Intel Core Ultra 9 285K and 188 GiB of RAM, every model served by llama.cpp, one model at a time, one request at a time. Each entry names the model file and quantization, the context window and the day it ran. measured

None of this is a benchmark and none of it ranks models. Several entries exist because something surprised us: a refusal, an empty string, a number that matched its prediction to within two mebibytes. The lead entry is a correction of our own record, a genre this site keeps on purpose.

02

What we ran, exactly

Context sweep, 2026-09-12llama-server started directly, only the window varied; speaking is the best of two warm ~200-token prose replies, reading one ~21,000-token ledger with a planted code; a 97 GB download ran throughout, so load times are upper bounds.
256K sweep, 2026-09-15One long rung per model, a ~230,000-token prompt with three sealed codes at 5, 50 and 95 percent depth, card memory at load and at peak; about 340 GB of downloads ran during it.
Wave-two check, 2026-09-13Recall of three codes and a tool call at about 48,000 and 96,000 tokens through an agent framework, one run per model per arm. A derived summary ships in data/wave2/; one of its entries records a timeout as a recall miss, and CORRECTION.md beside it annotates that (section 15).
MiniMax M2.7 audition, 2026-09-13, and budget test, 2026-09-15Header and template reads, two served rungs and the probes, then the server's --reasoning-budget at 0 and 256. Both ran before the launch fix of 19 September, so their wall times are not reprinted.
MiniMax M2.7, 2026-09-19 and 21Ladders and a restart test directly against the server, then the wave-two check re-run through the same agent framework.
MiniMax M3, 2026-09-19 to 20The old file's micro-batch sweep, then the correct file, its 262,144 load ladder and a two-machine attempt.
Flash-Next context study, 2026-09-20Three sealed codes at windows of 262,144, 393,216 and 524,288, cache at f16 and q8_0, plus the rope-flag trap; one pass each, temperature 0.
Batch-size sweep, 2026-09-20 to 21--batch-size and --ubatch-size swept on ~48,000-token prompts, one run per setting: hand-copied flags at a 131,072 window on 20 September, each model's own start script at its served window on 21 September, so Ling-3.0-flash's two baselines sit on different windows (section 08). The sweep is its own study.
Qwen3.6-27B files and the two-box run, 2026-09-13A header comparison and a processor-only load; Llama 4 Maverick split across the desktop and an ASUS ROG Flow Z13 laptop over Thunderbolt.

Card memory is the raw nvidia-smi figure in MiB, never divided by 1000.

03

MiniMax M2.7, re-opened

The file: Unsloth's UD-IQ4_XS build, four shards, 108,413,781,312 bytes, verified by count and byte size against the repository tree. The header, read before any load: architecture minimax-m2, 62 blocks, 48 attention heads with 8 key-value heads, 256 experts with 8 used, native context 196,608 tokens. measured 2026-09-13

What we recorded, and what was wrong

On 13 September, through an agent framework at a 131,072 window, the 48,000-token leg (53,581 prompt tokens) took 1,682.9 seconds, about 28 minutes, and the 96,000-token leg returned no answer before the turn ended at 1,800.4 seconds. We closed the model as too slow. The cause was our launch line: it carried -ub 128, llama.cpp's micro-batch size, copied from MiniMax M3's start script, where no reason for it is recorded. llama.cpp's default is 512. measured 2026-09-13

On 19 September, at a 131,072 window with an 8-bit cache and -b 4096 throughout, one 3,658-token prompt read at 32.3 tokens a second at -ub 128 and at 564.0 at -ub 4096, with every expert in system memory. At 43,909 tokens and -ub 4096, prompt evaluation ran at 640.2 tokens a second with every expert layer there, 656.5 with the experts of 59 of the 62 layers there (the placement the installed launch script uses), and 666.1 at 58 of 62 with 30,695 MiB on the card. Those are prompt-evaluation rates; the whole requests, each ending in a short answer, took 72.1, 70.1 and 69.1 seconds. One cold run each; the batch-size study has the whole ladder. measured 2026-09-19

At the installed placement, a direct read of a prompt seeded to the same 96,000-token target, 85,763 tokens as the server counted them, was evaluated in 145.48 seconds, 589.5 tokens a second; the whole request, including a 16-token answer, took 147.49 seconds of client wall time. measured 2026-09-19

Then the same agent framework, on 21 September, against the installed server with the fixed setting: the 48,000-token leg took 122.1 seconds at 53,584 prompt tokens and the 96,000-token leg 251.2 seconds at 103,931, with all three codes found on both legs and the tool call passed. measured 2026-09-21

Decode was not what changed: 32-token answers decoded at 9.9 to 10.2 tokens a second after the 3,658-token prompt at micro-batches from 128 to 2,048, and at 9.0 to 10.0 after 43,909 tokens at 4,096; through the framework, 8.0 on 13 September and 8.6 and 7.2 on 21 September, on completions of 160, 180 and 182 tokens. These are short answers, not prose speaking speeds. measured 2026-09-13 to 21

The native 196,608-token window fits on the one card only at a smaller micro-batch. With every expert in system memory, -ub 4096 failed to load at that window, -ub 2048 aborted, and -ub 1024 loaded with 31,560 MiB on the card. Only the 3,658-token prompt was run there (188.8 tokens a second), so deep reading at 196,608 is unmeasured. measured 2026-09-19

A warm session survives a real server restart: 85,778 tokens saved and 85,778 restored, 1 token counted as new work, a 2.19-second restore, then a 2.17-second request; the warm-wake study has those clocks and the mechanism. On 21 September the installed launch script repeated the test: the server evaluated 1 prompt token (136.8 ms) and 16 answer tokens, total time 1.97 seconds, stop count 42,898 tokens. measured 2026-09-19 and 2026-09-21

What stands from the first week, dated

The chat template starts the model inside a thought block on every turn and offers no switch to close it: the default, enable_thinking: false, thinking: false and reasoning_effort: "low", each at 65,536 (f16 cache) and 131,072 (8-bit cache), eight calls, and every one produced a reasoning block. measured 2026-09-13, not re-run since; every turn of the 21 September re-run also streamed its reasoning before any answer text.

The server-side budget does not close it either. With --reasoning-budget 0 the model still wrote 1,734 and 3,003 characters of thought on two prompts, and the second hit its 600-token cap with zero visible characters, stop reason length; a budget of 256 gave answers and well-formed tool calls. measured 2026-09-15, before the launch fix and not re-run; its log ships unchanged.

The first coherence probe of the audition, given a 900-token output budget, returned finish=length, 900 completion tokens, 0 characters of content and 4,034 characters of reasoning: the whole budget went on thinking. measured 2026-09-13, 65,536 context, f16 cache. That is the finding of our Empty Answer study, on another vendor's model served locally; it concerns the output budget, not the launch flag, and was not re-run. Given 4,096 tokens the model answered coherently (three sentences of the answer are quoted in the package), and its tool calls were well formed, parsed correctly from its own XML dialect.

The header arithmetic was always about memory: 248 KiB of cache per token, 2 times 8 key-value heads times 128 dimensions times 2 bytes times 62 blocks, so 131,072 tokens at f16 is 31.0 GiB, too much for a 31.8 GiB card beside the weights; that window needs an 8-bit cache. arithmetic Measured, the 65,536 rung at f16 took 21,232 MiB and the 131,072 rung at 8-bit 22,240 MiB, because both hold about 15.5 GiB of cache. measured 2026-09-13. Our first notes offered this arithmetic as the reason the model was slow (section 15).

04

MiniMax M3

Two inherited problems. M3's own start script is where MiniMax M2.7's -ub 128 came from, and no reason for it is recorded there either. And the file we had, Unsloth's UD-IQ3_XXS conversion, predates llama.cpp's support for this model's sparse attention: a tensor inventory finds zero indexer tensors among its 948, so every M3 number we had was a dense fallback. measured 2026-09-19

On that old file, at a fixed -b 4096, reading a 3,658-token prompt scaled with the micro-batch: 22.6 tokens a second at 128, then 67.7, 123.1, 221.3 and 393.8 at 512, 1,024, 2,048 and 4,096. One cold run per rung. measured 2026-09-19

On the correct file, bartowski's Q2_K_L (four shards, 153,086,988,768 bytes, verified), at a 131,072 window, -ub 2048 and an 8-bit cache: prompt evaluation at 147.5 tokens a second on 3,658 tokens; a 58,307-token recall prompt took 300.2 seconds for the whole request, including 165 generated tokens, about 194 prompt tokens per second of wall time, which understates pure reading; decode at 8.9 on an 82-token answer and 9.0 on a forced 128-token generation; all three codes found. At the 196,608 window (-ub 512): 72.6 on 3,658 tokens, and the same recall took 856.3 seconds including 207 generated tokens (about 68.1 per second of wall time), 3 of 3 again, so no rise with depth shows at that window. measured 2026-09-20

A 262,144-token window with the sparse attention engaged fits this card only at -ub 128: the compute buffer alone asks for 19,619 MiB at -ub 2048, the allocation fails at every micro-batch from 2,048 down to 256, and at 128 the window loads with 31,525 MiB on the card and reads a 3,658-token prompt at 24.2 tokens a second. We treat 256K as not viable for this model here. measured 2026-09-20. A split across the desktop and the laptop, with bartowski's larger IQ3_XXS build of the same attention (167.6 GiB), loaded at -ub 1024 and read the 3,658-token prompt at 17.58 tokens a second; speaking was not measured, because the run was stopped during the first speaking probe. The two-box study has that row. measured 2026-09-20

05

GLM-4.7-Flash

The fastest speaker in the 12 September sweep. At 131,072 context, all weights on the card with an f16 cache, Unsloth's UD-Q4_K_XL build: 223.65 tokens a second speaking and 4,447.44 reading on a 20,221-token prompt, the planted code found, 24,814 MiB on the card, loaded in 21.2 seconds. measured 2026-09-12

At its native ceiling of 202,752 tokens it held 28,702 MiB (28,765 at peak), loaded in 22.2 seconds, spoke at 227.43 on a short turn, and read a 189,931-token prompt at 819.70 tokens a second, finding all three codes; that 29-token three-code answer decoded at 70.48. measured 2026-09-15

Reading falls from 4,447 to 820 between the two rungs, which is why this page gives a prompt size with every reading rate.

On 21 September a larger micro-batch read a ~48,000-token prompt at most 19.5 percent faster (2,594.6 to 3,100.2 tokens a second, at a rung whose visible answer came back empty, the code only in its reasoning) and was not adopted; the batch-size study has the rungs. measured 2026-09-21

06

Gemma-4-26B

At 131,072 context, all on the card with an f16 cache, 211.07 speaking and 10,593.27 reading on a 21,690-token prompt, the planted code found, 20,646 MiB. measured 2026-09-12

At a full 262,144-token window it read a 230,855-token prompt at 3,799.69 tokens a second and found all three codes; it spoke at 202.64 on a short turn and decoded the 33-token three-code answer at depth at 121.36; 23,446 MiB loaded, 23,513 at peak, loaded in 23.2 seconds. measured 2026-09-15

On 21 September a larger micro-batch read a ~48,000-token prompt faster, from 8,972.3 tokens a second at the default to a best of 12,063.4 at -ub 2048. At that best reading rung the 17-character code answer decoded at 153.8 tokens a second, against 163.7 at the default; the following rung, -ub 4096, decoded that same short answer at 148.4 and read slower, 11,811. It was not adopted; the batch-size study has the rungs. measured 2026-09-21

07

Gemma-4-31B

This one fills the card. At 131,072 context: 73.68 speaking, 3,166.30 reading on a 21,690-token prompt, 29,844 MiB on the card with 2,763 MiB free. At 65,536: 73.73 speaking, 3,175.57 reading on the same prompt, 24,660 MiB with 7,947 free. Both rungs found the planted code. measured 2026-09-12

Speed is flat between 64K and 128K; only the room changes. A larger micro-batch measured on 21 September read a ~48,000-token prompt at most 3.6 percent faster (2,687.5 to 2,784.5) and was not adopted. measured 2026-09-21

08

Ling-3.0-flash

A 124-billion-parameter mixture-of-experts by its maker's count stated (the maker's figure, not our measurement). From the file header: 43 blocks, 512 experts with 8 used and 1 shared, and only 8 of the 43 blocks hold an attention cache, in a compressed form (rank 512 plus 64 rotary dimensions); the other 35 run a linear path with a fixed-size state. measured header read, 2026-09-13

That makes context cheap: 9 KiB of cache per token, 0.56 GiB at 65,536, 1.13 GiB at 131,072 and 2.25 GiB at its native 262,144, 27 times less per token than MiniMax M2.7's 248 KiB, both computed the same way from the files' headers. arithmetic

In the 13 September agent-framework check, at a 131,072 window and the default micro-batch, it held 6,941 MiB, found all three codes on both legs and made its tool calls; the first read reached first content after 506.68 seconds on a 103,807-token prompt (the history was seeded to about 95,913 tokens). A warm turn on that history reached first content after 0.96 seconds; the turn itself took 2.0 seconds. measured 2026-09-13

At its native 262,144, still at the default micro-batch, it loaded in 103.8 seconds holding 8,314 MiB (8,473 at peak), read 230,759 tokens at 206.25 a second, and found all three codes. measured 2026-09-15. The 2.25 GiB is the cache alone, by arithmetic; the 8,314 MiB is the whole model as served.

The batch-size sweep changed its reading. With -b 4096 -ub 2048, the setting that now ships, it read 727.7 tokens a second on 47,986 tokens and 694.6 on 149,711 at the 262,144 window (21 September, 8,824 MiB at peak), against 238.2 on 47,992 tokens at the default micro-batch at a 131,072 window (20 September). The windows differ. On one window, 262,144, a 3,007-token read went from 201.5 at -ub 512 to 444.0 at -ub 2048, about 2.2 times. measured 2026-09-20 to 21

A faster setting was withdrawn. At -ub 4096 the served window read 1,173.9 on 48,084 tokens and 1,067.6 on 150,153, but one 3,007-token prompt crashes the server with a CUDA illegal memory access. The shipped crash log is that one case, at the 262,144 window; the same prompt passed at 2,048, 1,024 and 512 on that window, and 60 code and prose prompts and 50 ledger prompts ran with zero crashes at either setting. The crash log and a standalone reproducer ship in the package; the diagnostic runs are in the batch-size study. measured 2026-09-20 to 21

And it has a real thinking off-switch, three in the template: enable_thinking: false, thinking_option: "off", or the phrase "detailed thinking off" in a system message; off emits a closed, empty thought block. measured template read, 2026-09-13. The opposite of MiniMax M2.7.

09

Qwen3.8-Flash-Next

At 131,072 context, Unsloth's UD-Q3_K_XL build with the experts in system memory: 11,366 MiB, 22.50 speaking, 195.84 reading on a 21,740-token prompt, the planted code found, about 2 seconds to load warm. Its 65,536 preset holds 9,670 MiB and reads 246.9 on a 2,624-token prompt. measured 2026-09-12

At 262,144 it loaded in 52.4 seconds, held 15,124 MiB (15,626 at peak), read the 230,803-token prompt at 165.70 a second, about 23 minutes, and found all three codes; it spoke at 25.33 on a short turn and decoded the 32-token three-code answer at depth at 13.58. measured 2026-09-15

Five days later we pushed it past its trained window. With the rope-scaling flags set, recall held at three of three sealed codes out to 459,911 tokens, 1.75 times the trained 262,144, with time the limit: 19.8 minutes at f16 and 262,144 (the q8_0 twin took 19.4), then 31.2 and 44.6 minutes at q8_0 at 393,216 and 524,288. At 262,144 the q8_0 cache gave the same recall with 22,019 MiB at peak against 25,226 at f16. measured 2026-09-20

These runs show retrieval, not reasoning across the window: the integration question fails against a 5,974-token ledger and against every longer one; the question itself is a 526-token follow-up with the ledger cached.

One trap, measured the same day: ask for 524,288 without the rope flags and the server logs that it caps the slot at 262,144, while the card still pays for the big window: 23,816 MiB, the same as a genuine 524,288 window (23,814), against 15,394 on a plain 262,144 load (that server came up, but its probe then logged a 600-second timeout). measured 2026-09-20. Its batch-size results, which differ between the 131,072 window they were swept at and the 262,144 window it serves, are on the batch-size study.

10

Qwen3.6-27B

The measurement here is which file could run at all. Two GGUF distributions bearing the Qwen3.6-27B label exist on this machine. The one distributed through Ollama's model library declares its rope sections as [11, 11, 10], three elements, and every llama.cpp build here refuses it, expected 4, got 3, in under a second, before any memory is touched. The upstream file declares [11, 11, 10, 0] and loads. One trailing zero in a header decided whether this model could run at all. measured 2026-09-13

The fix was proved without touching the contended card: the upstream file, Unsloth's UD-Q4_K_XL, was loaded processor-only with graphics initialisation disabled, and the server reported model loaded in 6.39 seconds. measured 2026-09-13. This is an interaction between a file variant and a loader's array-length requirement, not an accusation against either project.

11

Qwen3-32B

The arithmetic success. Its header gives 256 KiB of cache per token at f16, so at 32,768 context (its served rung; the file's own ceiling is 40,960) the prediction was 18.81 GiB of weights plus 8.0 GiB of cache plus 1.04 GiB of compute buffers, 28,518 MiB. The measurement was 28,520 MiB, two MiB out. arithmetic confirmed measured 2026-09-13, Q4_K_M

The same arithmetic says a 131,072 window fits this card only with a 4-bit cache, 27.8 GiB in total; f16 and 8-bit do not fit. That rung was not run. arithmetic

One caution from the wave-two check: with thinking off this model passed, tool call included; with thinking on the same check failed, the tool call not emitted. One run each, one evening; we do not generalise it. measured 2026-09-13

12

Llama 4 Maverick

Llama 4 Maverick, Unsloth's UD-Q3_K_XL at 167.21 GiB by its header and never served from one machine here, loaded split across the desktop and the laptop at 131,072 context on 13 September and answered at 13.09 tokens a second on a short 51-token reply: a capacity demonstration on a model deleted the next day, with its row and header read on the two-box study. measured 2026-09-13

13

256,000 tokens, from the headers

What would a quarter-million-token window cost each model? The arithmetic column is the attention cache alone, computed from each file's header by a standalone reader that never touches tensor data (reader and outputs in the package). The measured column is the whole card at the long window on 15 September, one run per model; the context-256k study carries the full table.

ModelCache at ~256K, from the headerLabelMeasured at the long window, 2026-09-15Label
Inkling-Smallabout 7 GiBarithmetic262,144: 15,744 MiB loaded, 15,814 at peak; read 230,827 prompt tokens at 115.07 t/s; 3 of 3 codes; the 27-token code answer decoded at 8.78 t/smeasured
Qwen3.8-Flash-Nextabout 6 GiBarithmetic262,144: 15,124 MiB loaded, 15,626 at peak; read 230,803 prompt tokens at 165.70 t/s; 3 of 3 codesmeasured
Ling-3.0-flashabout 2.25 GiBarithmetic262,144: 8,314 MiB loaded, 8,473 at peak; read 230,759 prompt tokens at 206.25 t/s; 3 of 3 codesmeasured
Qwen3.5-397Babout 7.5 GiBarithmetic262,144: 30,908 MiB loaded, 31,068 at peak; read 230,803 prompt tokens at 194.51 t/s; 3 of 3 codes; the 32-token code answer decoded at 17.00 t/s with speculative decoding from its own draft head (29 of 29 drafted tokens accepted)measured
Qwen3.8-27Babout 16 GiBarithmeticno logged measurement; its field card records a 17 September serve at this window that kept no log
Gemma-4-26Bnot computed262,144: 23,446 MiB loaded, 23,513 at peak; read 230,855 prompt tokens at 3,799.69 t/s; 3 of 3 codesmeasured
GLM-4.7-Flashcannot reach 262,144; its ceiling is 202,752header read202,752: 28,702 MiB loaded, 28,765 at peak; read 189,931 prompt tokens at 819.70 t/s; 3 of 3 codesmeasured

The arithmetic rows are estimates. Qwen3.8-Flash-Next and Ling-3.0-flash have later figures in sections 08 and 09. MiniMax M3's answer for this window is in section 04: with its sparse attention engaged it loads only at -ub 128.

14

What this does not show

  • Nothing here measures model quality, intelligence or benchmark standing.
  • Most entries rest on one run each. The wave-two check is one run per model per arm, the sweeps one run per rung, the budget probes one call per cell, and the batch-sweep and MiniMax M3 figures one cold run per setting.
  • Header arithmetic is not a run. The 256K arithmetic column is an estimate; the measured column is the evidence.
  • Load times are upper bounds. Large downloads competed for disk and page cache during the 12 and 15 September sweeps.
  • Reading rates belong to their prompt sizes and windows. Section 05 shows one model at 4,447 and at 820. The batch-sweep figures are ~48,000-token prompts and are not comparable to the 12 September sweep's ~21,000-token prompts.
  • Short answers are not speaking speeds. Where a decode rate comes from a planted code or an answer of a few dozen tokens, the page says so.
  • The agent-framework speeds are not directly comparable to direct-probe speeds. The 21 September MiniMax M2.7 re-run went through the same framework as the original failure for exactly that reason.
  • MiniMax M2.7's thinking and budget findings are dated 13 to 15 September and were not re-run after the launch fix. At its native 196,608 window only a 3,658-token prompt has been run.
  • The 459,911-token result is retrieval, not reasoning across the window.
15

What we got wrong

"MiniMax M2.7 is slow." Our first records called this model slow, explained it with the model's cache arithmetic, and closed it on that basis. The speed was our own launch line, -ub 128, inherited from MiniMax M3's start script with no recorded reason; the cache arithmetic only ever described the memory bill. The "why it is slow" reasoning is gone, the wall times that depended on the flag are withdrawn or dated, and section 03 carries the re-measurements.

The timeout that looked like a recall miss. Our gate harness left error: null when a turn ran out of time without an answer and counted the recall as 0 of 3, so a timeout and a genuine miss looked identical. MiniMax M2.7's 96,000-token leg was one: 1,800.4 seconds, no answer events, written down as "l128 recall 0/3", and that record reached this page before we caught it. The same week, a separate run on another model that produced only reasoning and no visible text was also scored as recall 0 of 3. On yet another model's run the harness divided by a missing first-content time and wrote 987.9 tokens a second next to a pass; that run had 326 completion tokens, so it was not an empty answer. The harness was fixed before the 21 September re-run: a turn with no answer is now recorded as UNANSWERED. The wave-two summary in the package keeps its original values; CORRECTION.md annotates them.

The 700-to-900 note. In the early hours of 2026-09-13, after the audition had measured an empty answer at a 900-token output budget, a working note of ours advised giving the model "700 to 900 output tokens or the answer comes back empty", quoted here word for word. The advice pointed straight at the measured failure: 900 is exactly the budget that returned zero characters. It was corrected in place the same night with a dated marker. Measure your own floor; the Empty Answer study shows how.

The "~40 t/s" verdict. The budget test's verdict paragraph reports reading at about 40 tokens a second, a rate that does not follow from the log lines in the same record; we do not print it, and the probe log ships unchanged.

The unfinished run. An earlier record of ours, written while the 256K sweep was still running, has the Qwen3.8-Flash-Next run at that window as unfinished: the read had reached 192,512 of about 230,000 tokens when the snapshot was taken. The run completed the same evening, and section 09 prints the completed measurement. When two of our records disagree, the log is the number we print.

16

Related pages on this site

The Empty Answer study is the parent finding for section 03's budget trap. The batch-size study covers the flag behind this page's biggest correction. The warm-wake study covers restoring a server's state from disk. The two-box study covers the desktop-and-laptop pair behind sections 04 and 12. The context-256k study carries the full 256K table. Labels are defined on the method page.

17

Sources and artifacts

The data package is a leak-swept selection of the run records this page cites, with derived summaries where no public primary record exists. data/README.md maps every file and redaction and opens with where the package seems to argue with itself; data/NUMBERS.md maps every figure to its file and field. If a number here disagrees with a file in the package, the file is right; tell us and we will fix the page.

Every path on our machines was replaced in full with <REDACTED_PATH>, keeping only a model file's own name; addresses, ports, process ids and serving aliases became placeholders; request and fingerprint identifiers were removed from the response bodies, which otherwise ship whole. Our session write-ups, working notes and the per-leg records of both agent-framework checks do not ship. Model repositories were read at the revisions downloaded on 2026-09-13 (Unsloth's MiniMax-M2.7-GGUF and Qwen3.6-27B-GGUF, inclusionAI's Ling-3.0-flash) and 2026-09-20 (bartowski's MiniMax-M3 builds).

2026-09-12Context sweep (GLM-4.7-Flash, both Gemma models, Qwen3.8-Flash-Next) and the Flash-Next install measurements.
2026-09-13MiniMax M2.7 audition. Qwen3-32B. The Qwen3.6-27B header finding and processor-only load. The wave-two check. Llama 4 Maverick on the pair.
2026-09-15MiniMax M2.7 budget test. The 256K sweep: Gemma-4-26B, GLM-4.7-Flash, Qwen3.8-Flash-Next, Ling-3.0-flash, Inkling-Small and Qwen3.5-397B at their long windows.
2026-09-19The MiniMax M2.7 re-launch: the -ub 128 control, the micro-batch and placement ladders, the 196,608-token load ladder, the warm restore. MiniMax M3's micro-batch sweep on the old file.
2026-09-20MiniMax M3 on the correct file, its 262,144 load ladder and the two-machine attempt. The Qwen3.8-Flash-Next context study and the rope-flag trap. The batch-size sweep begins.
2026-09-21The MiniMax M2.7 gate re-run and the installed script's restore test. The batch sweep concludes: Ling-3.0-flash's faster setting withdrawn; larger micro-batches measured and not adopted for GLM-4.7-Flash and the two Gemma models.

Every claim here is dated; if today is much later than the rows above, treat version-specific findings as unverified since then.