# Every figure on the page, and the file and field it came from

Revised 2026-09-26 for the snapshot "The roster, as of 26 September 2026", and again in the fix pass that day. The 15 September version of this file
mapped the first list; its mappings for the figures this page still prints are carried below, and
`ROSTER_22.csv` keeps the first list as it was.

**If a number on the page disagrees with a file in this package, the file is right and the page is wrong.**

Model names here are the public names the page uses; file names in this package use the same names. Folder names
below are relative to this `data/` folder. "Result" means a run's `.result` file, whose lines hold the harness's own
JSON (`prompt_n`, `prefill_tps`, `decode_tps`); "log" means the server's own log, whose `print_timing` lines hold the
server's count of tokens read and generated. Where both exist they agree; the answer lengths quoted on the page
("an 11-token answer") are the log's `eval time = ... / N tokens`.

## Where a file will look like it contradicts the page

1. **Qwen3-235B's 8.18 in `batch-sweep-2026-09-21/Qwen3-235B-A22B-Instruct-2507_2048.result`** is a speaking rate on
   an 11-token answer (the log's `eval time ... / 11 tokens`). The page prints the letters of 26 September instead
   (5.51, 4.68, 4.22) and says why in section 05. Both files are here.
2. **MiniMax M2.7's 9.80 in `minimax-m2.7/2026-09-13_ctx65536_f16.server.log`** (task 261, 900 tokens) is a reply
   made entirely of hidden reasoning, cut off at the length limit before any answer (`finish=length`,
   `content_chars=0` in `minimax-m2.7/2026-09-13_extract.txt`, section 3). The page prints the same request's
   finished reply from the 131,072-window run: 9.73 over 1,522 tokens.
3. **Inkling-Small's 337.9 and Qwen3.8-Flash-Next's 511.7** were read at a 131,072 window (the `ctx=131072` in each
   result's `[load]` line); both models serve 262,144. The page prints them as 131,072 figures and leads each row with
   what exists at the served window: the deep reads at the old batch setting, and 3,000-token reads at the new one.
4. **MiniMax M3's 194**, which our records carried, is not in any file as a rate. It is 58,307 tokens divided by
   the 300.2-second wall time of the whole request (`minimax-m3/2026-09-20_q2kl_needle_58k.json`, `wall_s`,
   `prompt_tokens`). The server's own reading figure for the same request is 208.11 (the server log, task 216).
5. **The Qwen3.8-Flash-Next records hold a failed question the page never mentions.** Each
   `qwen3.8-flash-next-2026-09-20/*.result` ends with `[q3 INTEGRATE] letter=N tally=N task=N`: a question that needs
   facts combined from across the ledger, answered wrongly at every length including the 5,986-token control. The
   page prints only the reading and speaking rates and the three-code retrieval (`[q1 RETRIEVE] 3/3`), and makes no
   claim about reasoning across a window.
6. **Short reads right after a start.** Every 3,000-token read in the batch-size folders was the first completion
   request after the server started (the drivers `bsweep.sh` and `series_probe.sh` send nothing before it). The page
   says so where it prints one as the only figure at a window.
7. **The framework's `t1` speaking column** in `tool-recall-wave1-2026-09-12.tsv` is still not printed, for the
   reason the 15 September version gave: it is a short first turn whose token count leaves out a thinking model's
   reasoning.

## Section 01, the machine and what changed

| page says | file | field |
|---|---|---|
| RTX 5090, 32,607 MiB | `context-sweep-2026-09-12.tsv` | header line 1 |
| Intel Core Ultra 9 285K, 188 GiB of RAM | `install-runs-qwen3.8-flash-next-2026-09-12.tsv` for the RAM figure; the processor is the machine's own specification, stated on the site's method page | |
| Laptop, 128 GB unified memory, direct Thunderbolt cable | `two-box-deepseek-probes-2026-09-13-to-15.txt` | run headings |
| Every start script refuses while another model is up; seven guards each refused | `refusal-guards-2026-09-12.txt` | tests G1 to G7 |
| Twenty-three endpoints | `ROSTER_23.csv` | 23 rows |
| Ornith-1.5-35B-A3B and GLM-5.3-Flash installed on 16 September, MiniMax M2.7 on 21 September; Kimi K3 retired 15 September | our operator records, labelled on the page as such; not measurements. The first runs of the two 16 September models are in `first-runs-2026-09-16/` and M2.7's framework check of 21 September is `gate-records/minimax-m2.7-2026-09-21.json` | |
| Kimi K2.7-Code's weights removed on 17 September | `file-listing-2026-09-26.txt` | section B: 0 model files; the folder last changed 2026-09-17 11:46 |
| MiniMax M3 serves a different file since 21 September | `file-listing-2026-09-26.txt` section A row 19 (the Q2_K_L shards) and section C (its start script); the date is our operator record | |
| Laguna S 2.1 serves 262,144 since 21 September | `file-listing-2026-09-26.txt` section C (`CTX="${1:-262144}"`); `batch-sweep-2026-09-21/Laguna-S-2.1_ctx262144_4096_series.result`, `served: n_ctx_slot = 262144` | |
| Qwen3.6-27B serves a different file since 26 September; our records give it 262,144 from that day, and the runs here passed that window explicitly | `file-listing-2026-09-26.txt`, section A row 10 and section C (`case "${Q36_FILE:-mtp-q5}"`, the cache block dated by the script's own comment); the window itself is set in a settings file that is not published, so the date and the window are our operator record | |

## Section 02, method

| page says | file | field |
|---|---|---|
| 21 September: each model started through its own start script | `batch-sweep-2026-09-21/batch_sweep.sh`, `series_probe.sh`, `series_probe2.sh` | each drives the model's own start script |
| 19 and 20 September: the server started by hand with the same flags | `minimax-m2.7/2026-09-19_*.json`, `config` (micro-batch, placement, window, cache, build) with the probe `m27_probe.py`; the header line of each `minimax-m3/*launch-check.log`; `batch-sweep-2026-09-20/bsweep.sh`; the `[load]` lines in `batch-sweep-2026-09-20/*.result` | |
| temperature 0, a code planted halfway that the answer had to quote exactly | `batch_sweep.sh` (the request: `"temperature": 0.0`, `max_tokens` 900; the ledger with `SEALED REFERENCE` at `n//2`) | |
| llama.cpp's defaults are `-b 2048 -ub 512` | `batch_sweep.sh`, the result line `ubatch=skip (the script's own default)`, and the 20 September `bsweep.sh` header | |
| Paragraph replies: 150 words on 15 September, about 200 tokens on 12 September, better of two warm replies | `long-reads-2026-09-15/CONTEXT_256K_MEASUREMENTS.tsv` header (`decode_tps = best of 2 warm ~200-token replies`) and the request text inside `first-runs-2026-09-16/audition_rung.sh` ("In exactly one paragraph of about 150 words ..."); `context-sweep-2026-09-12.tsv` header (`best of 2 WARM reps on a ~200-token prose generation`) | |
| First runs: thinking switch off, a tool call, about 120,000 tokens with three codes, default batch settings | `first-runs-2026-09-16/first-runs-2026-09-16.tsv` (`tools`, `deep_prompt_tok`, `codes_hit`); the `[cmd]` line of each `first-runs-2026-09-16/*.log` carries no batch flag | |
| Letters: 20,000, 60,000 and 100,000 (Qwen3-235B), 20,000, 100,000 and 230,000 (Qwen3.6-27B) or 20,000, 48,000, 100,000 and 230,000 (Qwen3.8-27B, added in the 26 September update; Qwen3.5-122B-A10B and Ornith-1.5-35B-A3B, added that evening); letters of about 330 to 430 words | `letters-2026-09-26/*.jsonl`, `qwen3.8-27b-2026-09-26/*.jsonl`, `qwen3.5-122b-2026-09-26/*ctx262144_b2048_ub512.jsonl` and `ornith-1.5-35b-2026-09-26/*ctx262144_b2048_ub2048.jsonl`, `depth_target` and `answer_words` (366, 333, 331; 398, 430, 429; 367, 411, 392, 400; 407, 426, 418, 420; 348, 328, 343, 330); the request is `PROSE_Q` in `letters-2026-09-26/prose_probe.py` and `qwen3.8-27b-2026-09-26/probe38.py` (the evening runs used the same probe file under another name) | |
| MiniMax M2.7 had the same request at about 20,000, 48,000 and 190,000 tokens, and GLM-5.3-Flash at about 20,000 to 230,000; their rows say what came back | `minimax-m2.7/2026-09-26_ctx196608_cmoe_ub1024.jsonl` and `glm-5.3-flash-2026-09-26/*ctx262144*.jsonl`, `depth_target` (20000, 48000, 190000; 20000, 48000, 100000, 230000) | |
| Long reads: 150,000 to 231,000 tokens, three codes, 262,144 (202,752 for GLM-4.7-Flash) | `long-reads-2026-09-15/CONTEXT_256K_MEASUREMENTS.tsv`, `n_ctx`, `deep_prompt_tok` (150,475 to 231,065 at 262,144 and 189,931 at 202,752; the table's one row at 131,072 is not used on this page), `codes_hit` | |
| Framework: 48,000 and 96,000 seeded tokens, three codes, a tool call; the true size printed; short answers of 30 to 1,004 tokens; reading is the harness's estimate | `gate-records/*.json`: `l64` and `l128`, `seed_tokens_est`, `A_recall.prompt_tokens`, `A_recall.completion_tokens` (30 to 1,004), `A_recall.prefill_tps_approx` | |
| Gemma-4-26B: 8,972.3 direct at 47,983 against about 6,726.8 through the framework at 53,209 | `batch-sweep-2026-09-21/Gemma-4-26B-A4B_skip.result`; `gate-records/gemma-4-26b-a4b.json`, `l64.A_recall` | |
| Qwen3-235B: 9.3 on a 62-token framework answer at 53,458; 5.51 on a 478-token letter at 20,063 | `gate-records/qwen3-235b-a22b-2507.json`, `l64.A_recall`; `letters-2026-09-26/Qwen3-235B_2048.jsonl`, the first `prose` line | |

## Section 03, the twenty-three rows

Byte counts and shard counts for every row are in `file-listing-2026-09-26.txt`, section A (read 2026-09-26).
Every framework figure is the `l64.A_recall` (48K leg) or `l128.A_recall` (96K leg) block of the model's file in
`gate-records/`: `prompt_tokens`, `prefill_tps_approx`, `decode_tps`, `completion_tokens` and `hits`; the dates are
each record's `started_utc` in local time (UTC minus four hours).

| row | figure on the page | file and field |
|---|---|---|
| 1 | window: 131,072 by the script's fallback; 262,144 loaded on 15 and 21 September | `file-listing-2026-09-26.txt` section C (the split script's `CTX` line); `installed-windows-2026-09-21.txt` section C; the 21 September runs passed 262144 (`batch-sweep-2026-09-21/DeepSeek-V4-Flash-pair_chain.sh`, `SCRIPT_ARGS`), served `n_ctx_slot = 262144` in `DeepSeek-V4-Flash-pair_512_series.result`; the 15 September long-read run passed `--ctx-size 262144` (`long-reads-2026-09-15/CONTEXT_256K_MEASUREMENTS.tsv`, METHOD line and the `DeepSeek-V4-Flash-pair_262144` row); the 14 September framework records show 131,072 served and carry no launch arguments (`gate-records/deepseek-v4-flash-0731-two-machines.json` and `-attempt3.json`, `S.n_ctx`) |
| 1 | 14.73, 12.05, 7.85 on replies of about 400 tokens | same file, `warm_decode_tps` and `warm_gen_n` (400, 394) at 2,998 and 48,024; `..._512_150k.result`, 7.85 on 380 |
| 1 | 161.8 on 2,998, 102.5 on 48,024, 51.4 on 150,103 | the same two results, `prompt_n`, `prefill_tps`; the setting is the header's `ubatch=512 batch=2048` |
| 1 | framework, 2026-09-14 | `gate-records/deepseek-v4-flash-0731-two-machines.json` |
| 1 and 3 | 6.7 and 12.3 times | arithmetic: 690.9 / 102.5 and 632.1 / 51.4 |
| 1 | 19.55 speaking with the 4-bit preset on the pair | `two-box-deepseek-probes-2026-09-13-to-15.txt`, the 2026-09-15 block, `decode 19.55` |
| 2 | 11.77 on a 202-token paragraph | `long-reads-2026-09-15/CONTEXT_256K_MEASUREMENTS.tsv` `decode_tps` 11.772 (best of two); the token count is `timings.predicted_n` of the faster of `Inkling-Small_ctx262144.short.r1.json` and `.r2.json` |
| 2 | 8.78 on 27 tokens and 115.07 on 230,827 | `long-reads-2026-09-15/Inkling-Small_ctx262144.deep.json`, `timings` |
| 2 | 263.6 on 3,033 (2026-09-20) | `batch-sweep-2026-09-20/Inkling-Small_2048_ctx262144_3k.result`; its `[load]` line has `ctx=262144` and `--batch-size 4096 --ubatch-size 2048` |
| 2 | 187.9 on 3,033 (2026-09-21) | `batch-sweep-2026-09-21/Inkling-Small_ctx262144_3k.result`, `served: n_ctx_slot = 262144`; the setting is the start script's default (`file-listing-2026-09-26.txt`, section C) |
| 2 | 337.9 on 48,115 at 131,072 | `batch-sweep-2026-09-20/Inkling-Small_2048.result`, `ctx=131072` in its `[load]` line |
| 2 | framework, 2026-09-14 | `gate-records/inkling-small.json` |
| 2 | the 14 September probes served 131,072; a run that passed no window was served 262,144 on 21 September | `inkling-probes-2026-09-14.md` (run 1, "128K window"), `gate-records/inkling-small.json` (`S.n_ctx` 131072); `installed-windows-2026-09-21.txt` section A |
| 3 | 8-bit: 9.93 on a 209-token reply | `CONTEXT_256K_MEASUREMENTS.tsv` row `DeepSeek-V4-Flash-8bit-desktop_262144`, `decode_tps`; tokens from the faster short reply JSON |
| 3 | 8-bit: 480.0 on 150,324 at 262,144; 9.37 on 106 tokens | `batch-sweep-2026-09-20/DeepSeek-V4-Flash-8bit-desktop_verify-150k.out`; the log `..._8192_150k_ctx262144.log` (`n_ctx_slot = 262144`, `eval time ... / 106 tokens`) |
| 3 | 8-bit: 607.1 on 48,073 at 131,072 | `..._verify-48k.out`, third line; `..._8192_48k.log`, `n_ctx_slot = 131072` |
| 3 | 3-bit: 12.65 on 59 tokens at 2,998; 690.9 on 48,024 | `batch-sweep-2026-09-21/DeepSeek-V4-Flash-3bit-desktop_4096_series.result` and its log (59 tokens) |
| 3 | 3-bit: 632.1 on 150,103; 11.79 on 86 tokens | `..._3bit-desktop_4096_150k.result` and its log |
| 3 | experts of 36 of 43 layers; the 8-bit file is the default preset | `file-listing-2026-09-26.txt`, section C (the `fast` preset's `PLACE=(--n-cpu-moe 36)`, the default `PRESET="${DS_PRESET:-quality}"`); `model-file-headers.md` (`n_layer = 43`) |
| 3 | framework, 2026-09-12 | `gate-records/deepseek-v4-flash-0731-one-machine.json`; `tool-recall-wave1-2026-09-12.tsv` |
| 4 | 16.41 on 11 tokens after 48,029; 932.2 | `batch-sweep-2026-09-21/Qwen3.5-397B-A17B_4096.result` and its log |
| 4 | 224.5 at the defaults | `Qwen3.5-397B-A17B_skip.result` |
| 4 | 12.84 on a 184-token paragraph at 262,144 | `CONTEXT_256K_MEASUREMENTS.tsv` row `Qwen3.5-397B-A17B_262144`; the short reply JSONs |
| 4 | framework, 96K leg, 2026-09-13 | `gate-records/qwen3.5-397b-a17b.json` |
| 5 | 5.51, 4.68, 4.22 on 478, 426, 426 tokens; 741.8, 652.0, 542.1 | `letters-2026-09-26/Qwen3-235B_2048.jsonl`, the `read` and `prose` lines (`prefill_tps`, `decode_tps`, `predicted_n`) |
| 5 | `-b 4096 -ub 2048` | the `cmdline:` line of `letters-2026-09-26/Qwen3-235B_2048.result` |
| 5 | 703.3 on 48,020 (2026-09-21) | `batch-sweep-2026-09-21/Qwen3-235B-A22B-Instruct-2507_2048.result` |
| 5 | with `-ub 4096`: 1,069.0 on 20,063; 5.82 on letters | `letters-2026-09-26/Qwen3-235B_4096.jsonl`; the `env:` lines of its `.result` |
| 5 | framework, 2026-09-12 | `gate-records/qwen3-235b-a22b-2507.json` |
| 6 | (evening of 26 September) window: "262,144, the installed default: a run that passed no window was served it (2026-09-26)" | `installed-starts-2026-09-26.txt`, Qwen3.5-122B-A10B: started 21:07:29 through the installed start script and settings file with no window passed; served `n_ctx` 262144 (`GET /props`) and `n_ctx_slot = 262144`; card 30,848 MiB at load, close to the measured run's 30,898 at the same setting. The installed script (byte-identical to the run's proposed copy, written 17:14) keeps `-b 2048 -ub 512` at 262,144 (listing C line 137). Earlier loads by passing the window: 15 September (`long-reads-2026-09-15/`) and the 26 September runs (`qwen3.5-122b-2026-09-26/`) |
| 6 | `-b 2048 -ub 512` at 262,144, draft head on; 131,072 accepted, at `-b/-ub 4096` | section C, lines 33 (accepted windows) and 137 (batch by window); the runs' `cmdline:` lines (`--batch-size 2048 --ubatch-size 512`, `--spec-type draft-mtp`), the installed script's command line at that window (the staged copy differs only in its settings-file line, an error message and comments) |
| 6 | at 262,144: 24.6, 24.47, 23.84 and 21.64 on letters of 482, 520, 508 and 491 tokens (407 to 426 words) at 20,067, 47,992, 99,872 and 229,090 | `qwen3.5-122b-2026-09-26/Qwen3.5-122B-A10B_ctx262144_b2048_ub512.jsonl`, `prose` lines and the `read` line of each depth; the server log's `eval time` lines agree (24.60, 24.47, 23.84, 21.64) |
| 6 | at 262,144: 505.7 on 20,067, 511.4 on 47,992, 491.6 on 99,872, 448.7 on 229,090 (510.6 s); every code quoted, 3 of 3 at 229,090 | the same file, `read` lines (`prefill_tps`, `prompt_n`, `prompt_ms` 510,586; `code_in_answer`, `codes3_hits` 3); server log 505.74, 511.40, 491.56, 448.68 |
| 6 | row note: `-b 2048 -ub 1024` read 877.8 on 20,067 and 738.9 on 229,090 (310.1 s), peak 31,785 MiB, not adopted | `Qwen3.5-122B-A10B_ctx262144_b2048_ub1024.result` and `.jsonl`; the reason is the start script's comment, line 135 (section C) |
| 6 | row note: 30,898 MiB at load, 31,124 at peak | `Qwen3.5-122B-A10B_ctx262144_b2048_ub512.result` (`VRAM at load`, `peak VRAM during probes`) |
| 6 | row note: at 131,072 with `-b/-ub 4096` the same day, 2,088.3 on 47,992 and 22.82 on a 516-token letter; about four times faster (2,088.3 / 511.4 = 4.08, arithmetic, two settings and two windows) | `Qwen3.5-122B-A10B_ctx131072_b4096_ub4096.jsonl` and `.result` (`n_ctx_slot = 131072`, `--batch-size 4096 --ubatch-size 4096`) |
| 6 | framework cell, 131,072 window | `gate-records/qwen3.5-122b-a10b.json` (`S.n_ctx` 131072) |
| 6 | 32.72 on a 716-token answer after 48,027; 2,101.7 | `batch-sweep-2026-09-21/Qwen3.5-122B-A10B_4096.result` and its log; `reasoning_len` 2,212 characters shows the answer is mostly reasoning |
| 6 | 520.2 at the defaults | `Qwen3.5-122B-A10B_skip.result` |
| 6 | 22.39 on a 188-token paragraph at 262,144 | `CONTEXT_256K_MEASUREMENTS.tsv` and the short reply JSONs |
| 6 | framework, 48K leg, 2026-09-12 | `gate-records/qwen3.5-122b-a10b.json` |
| 7 | (26 September update, as corrected before push 6) window: "262,144 as installed on 26 September (the install record), loaded on 26 September by passing the window; the start script's comment records a 17 September measurement" | no run in these records started the installed script, with or without a window, after the change (every run used the staged copy with an explicit `Q38_CTX`; no log prints the installed script's own note), so the cell takes the install-record form. The install record: the start script as installed (byte-identical to the run's proposed copy, written at 15:17 on 26 September; its own fallback is 65,536, `file-listing-2026-09-26.txt` section C, line 37, so the window is set in settings this package does not publish and which we did not open). The 17 September measurement: the script's own comment (section C, line 67), not a run file. Loaded on 26 September, by passing the window: `qwen3.8-27b-2026-09-26/Qwen3.8-27B_ctx262144_draft-off.result` (`env: Q38_CTX=262144`, `served: n_ctx_slot = 262144`) |
| 7 | at 262,144 on the default file: 8-bit cache, draft head off, `-b 2048 -ub 512`; the draft head stays on for the UD-Q4_K_XL and UD-Q5_K_S files at 262,144, and for 131,072 on the default file | section C, lines 53 to 57 (the file choice), 150 to 158 (at 262,144: the 8-bit cache; the draft head cleared on the default file, kept with an 8-bit draft cache on the others) and 188 (the batch line, no window condition); the run's `cmdline:` line (`--cache-type-k q8_0 --cache-type-v q8_0`, `--batch-size 2048 --ubatch-size 512`, no draft flags), which is the installed script's command line for the default file at 262,144. The two smaller files at 262,144 keep the draft flags: `qwen3.8-27b-2026-09-26/Qwen3.8-27B-UD-Q4_K_XL_ctx262144_draft-on.result` and `...UD-Q5_K_S_ctx262144_draft-on.result` (`n_ctx_slot = 262144`; `--spec-type draft-mtp` with an 8-bit draft cache in each `cmdline:`). 131,072 on the default file with the draft head: the 12 and 21 September records below. Our operator records call these configurations presets and the 131,072 one the coding preset; the start script names none of them, so the page does not |
| 7 | the UD-Q4_K_XL file, 17,559,178,144 bytes, and the UD-Q5_K_S file, 18,665,753,504 bytes, both on disk | `file-listing-2026-09-26.txt` section A, row 7 (a directory listing with `stat` byte counts; the sizes equal the start script's `WANT_SIZE` checks, section C lines 56 and 57) |
| 7 | row note: "the two smaller files are different sets of weights" | they are other quantizations of the model, with other byte counts (section A) |
| 7 | at 262,144, draft head off: 60.7, 55.18, 49.08 and 34.89 on letters of 436, 497, 477 and 504 tokens (367 to 411 words) at 20,067, 48,054, 99,934 and 229,152 tokens | `qwen3.8-27b-2026-09-26/Qwen3.8-27B_ctx262144_draft-off.jsonl`: `kind` `prose` (`decode_tps`, `predicted_n`, `answer_words`) and the `read` line of each depth (`prompt_n`); `draft_n` null throughout |
| 7 | at 262,144: 3,118.2 on 20,067, 2,717.6 on 48,054, 1,958.2 on 99,934, 1,180.7 on 229,152; every planted code quoted, 3 of 3 at 229,152 | the same file, `kind` `read`: `prefill_tps`, `prompt_n`, `code_in_answer`; `codes3_hits` 3 at 229,152 |
| 7 | row note: the card held 30,058 MiB at load and 30,254 at the peak of the reads | the `.result`: `VRAM at load 30058 MiB`; `peak VRAM during probes: 30254 MiB` |
| 7 | row note: the edit-and-reread request, capped at 40 tokens, did not return the code at any depth | the same file, `kind` `edit_reread` (`code_in_answer` false, `predicted_n` 40); `probe38.py` (`max_tokens` 40 for that request) |
| 7 | 86.01 on paragraph replies (2026-09-12), at 131,072 with the draft head on | `context-sweep-2026-09-12.tsv`, Qwen3.8-27B `long-128K`, `decode_tps`; its flags column, `all GPU + gated MTP` |
| 7 | 99.26 on 106 tokens; 2,634.8 on 48,069 at the defaults, at 131,072 with the draft head on | `batch-sweep-2026-09-21/Qwen3.8-27B_skip.result` and its log (`n_ctx_slot = 131072`; `creating MTP draft context`) |
| 7 | `-ub 4096` would not load: 1,568.13 MiB | `batch-sweep-2026-09-21/Qwen3.8-27B_4096.log` |
| 7 | framework, 2026-09-13, 131,072 window | `gate-records/qwen3.8-27b.json` (`S.n_ctx` 131072) |
| 8 | (evening of 26 September) window: "262,144, the installed default: a run that passed no window was served it (2026-09-26)" | `installed-starts-2026-09-26.txt`, Ornith-1.5-35B-A3B: started 21:06:33, no window passed; `n_ctx` 262144 and `n_ctx_slot = 262144`; card 31,256 MiB at load, close to the measured `-ub 2048` run's 31,270. The installed script (written 16:28) uses `-b 2048 -ub 2048` at 262,144 (listing C lines 148, 154); its fallback is 131,072 (line 38), so the window comes from its settings file |
| 8 | `-b 2048 -ub 2048` at 262,144 since 26 September (`-b 2048 -ub 512` before); 131,072 accepted, at `-b/-ub 4096` | section C, lines 50, 148 and 154 (line 148: "since 2026-09-26; it was 2048/512"); the default run's `cmdline:` (`--batch-size 2048 --ubatch-size 2048`, passed explicitly, the installed script's command line at that window) |
| 8 | at 262,144: 98.4, 89.8, 82.61 and 67.73 on letters of 422, 405, 423 and 405 tokens (328 to 348 words) at 20,067, 47,992, 99,872 and 229,090 | `ornith-1.5-35b-2026-09-26/Ornith-1.5-35B_ctx262144_b2048_ub2048.jsonl`, `prose` lines; server log 98.40, 89.80, 82.61, 67.73 |
| 8 | at 262,144: 4,568.0 on 20,067, 4,626.9 on 47,992, 4,219.0 on 99,872, 3,193.3 on 229,090 (71.7 s); every code quoted, 3 of 3 at 229,090 | the same file, `read` lines (`prompt_ms` 71,742; `codes3_hits` 3); server log 4,567.99, 4,626.87, 4,218.96, 3,193.25 |
| 8 | row note: at `-b 2048 -ub 512`, 1,780.4 on 20,067 and 1,508.7 on 229,090 (151.8 s), speaking about the same (92.73 to 68.11) | `Ornith-1.5-35B_ctx262144_b2048_ub512.jsonl` |
| 8 | row note: 31,270 MiB at load, 31,417 at peak | `Ornith-1.5-35B_ctx262144_b2048_ub2048.result` |
| 8 | row note: at 131,072 with `-b/-ub 4096` the same day, 6,114.1 on 47,992 and 89.64 on a 383-token letter | `Ornith-1.5-35B_ctx131072_b4096_ub4096.jsonl` and `.result` (`n_ctx_slot = 131072`) |
| 8 | 89.72 on a 198-token paragraph, thinking off | `first-runs-2026-09-16/first-runs-2026-09-16.tsv` `decode_tps`; `Ornith-1.5-35B_ctx131072.short.json`, `timings.predicted_n` |
| 8 | 73.2 on 83 tokens; 6,283.3 on 48,027 | `batch-sweep-2026-09-21/Ornith-1.5-35B_4096.result` and its log |
| 8 | 1,649.22 on 120,334, 3 of 3 | `first-runs-2026-09-16/Ornith-1.5-35B_ctx131072.deep.json` (`timings`, the three codes in the answer); `first-runs-2026-09-16.tsv`, `codes_hit` |
| 9 | (evening of 26 September) window: "262,144, the installed default: a run that passed no window was served it (2026-09-26); `-b 4096 -ub 1024` there, where `-ub 4096` did not load" | `installed-starts-2026-09-26.txt`, GLM-5.3-Flash: started 21:10:34, no window passed; `n_ctx` 262144 and `n_ctx_slot = 262144`; card 28,714 MiB at load, close to the `-ub 1024` run's 28,834. The installed script (byte-identical to the run's proposed copy, written 21:06) has fallback 131,072 (listing C line 54), accepts 262,144 and 131,072 (line 63) and sets `-ub 1024` above 131,072 (lines 173, 180, 181) |
| 9 | `-ub 4096` did not load at 262,144: "its compute buffer alone asked for 13,281.37 MiB and the allocation failed" | `glm-5.3-flash-2026-09-26/GLM-5.3-Flash_ctx262144_b4096_ub4096_did-not-load.server.log` (`allocating 13281.37 MiB on device 0: cudaMalloc failed: out of memory`; `failed to allocate compute pp buffers`); its `.result` (`DIED on load`) |
| 9 | at `-ub 1024`: speaking 10.22 on 1,200 tokens at 20,065, all reasoning, no letter; 9.98 on 729 tokens, a 351-word letter, at 47,992 | `glm-5.3-flash-2026-09-26/GLM-5.3-Flash_ctx262144_b4096_ub1024.jsonl`, `prose` lines (`decode_tps`, `predicted_n`, `answer_words` 0 and 351, `reasoning_chars` 5,127 and 1,206); `env: GLM53_EFFORT=none` in its `.result`, and `--chat-template-kwargs {"reasoning_effort":"none"}` in its command line |
| 9 | at `-ub 1024`: reads 135.1 on 20,065 and 142.6 on 47,992, the code quoted at both; the only depths at this setting | the same file, `read` lines (`prefill_tps`, `prompt_n`, `code_in_answer`); the run's `depths=20000,48000` in its first line |
| 9 | at `-ub 2048`, not adopted: speaking 10.11 on 1,200 tokens (102 words, cut off) at 20,065; 9.87 on 1,200 (119 words, cut off) at 47,992; 9.56 on 567, a 337-word letter, at 100,003; 9.04 on 1,200 at 230,039, all reasoning, no letter; reads 206.0, 250.3, 243.1, 211.9 (1,085.7 s), every code quoted, 3 of 3 at 230,039 | `GLM-5.3-Flash_ctx262144_b4096_ub2048.jsonl`: the four `prose` lines (`decode_tps`; `predicted_n` 1200, 1200, 567, 1200; `answer_words` 102, 119, 337, 0; `code_in_answer` false, false, true, false; `reasoning_chars` 4,727 and an empty tail on the last), and the `read` lines (`prompt_ms` 1,085,744 on the last; `codes3_hits` 3) |
| 9 | row note: `-ub 2048` peaked at 31,951 MiB; `-ub 1024` held 28,834 at load and 29,261 at peak | the two `.result` files (`VRAM at load`, `peak VRAM during probes`) |
| 9 | row note: several letters ran out of their 1,200-token budget; the 9.04 at 230,039 returned no letter | `predicted_n` 1,200 on four `prose` lines across the two 262,144 runs (the probe's `--max-tokens 1200`); the 9.04 is the `-ub 2048` file's `prose` line at `depth_target` 230000 (`answer_words` 0, `predicted_n` 1200) |
| 9 | row note: the reads at `-ub 1024` reach only 20,065 and 47,992; at 262,144 the figures past that depth (100,003 and 230,039) are the `-ub 2048` run's; the 120,979-token read is the 16 September default-batch run, the 100,003-token read at 131,072 the same day's `-b/-ub 4096` run | `GLM-5.3-Flash_ctx262144_b4096_ub1024.result` (`depths=20000,48000`); the `-ub 2048` file's `read` lines; `first-runs-2026-09-16/GLM-5.3-Flash_ctx131072.deep.json` (120,979 at 75.48); `GLM-5.3-Flash_ctx131072_b4096_ub4096.jsonl` (100,003 at 373.4; `--batch-size 4096 --ubatch-size 4096` in its `.result`) |
| 9 | row note: at 131,072 with `-b/-ub 4096` the same day, 359.4, 404.9 and 373.4 on 20,065, 47,992 and 100,003 | `GLM-5.3-Flash_ctx131072_b4096_ub4096.jsonl` and `.result` (`n_ctx_slot = 131072`) |
| 9 | row note: the crash check: no crash in 40; 36 with an empty reply after using the whole budget (160 and 80 tokens); four with text in both replies, the first a 3,014-token ledger quoting its code | `glm-5.3-flash-2026-09-26/crash-check_ctx262144_ub1024.summary.txt` (`crashes` 0, `empty_answers` 36) and `.jsonl` (`empty` false on `i` 0, 4, 8 and 11; `n1` 160 and `n2` 80 on every empty row; `prompt_n` 3014 and `code_ok` true on `i` 0); `stress_fw.py` (`turn(msgs, 160)`, `turn(msgs, 80)`, `empty = (not a1) or (not a2)`) |
| 9 | 9.58 on a 181-token paragraph | `first-runs-2026-09-16.tsv`; `GLM-5.3-Flash_ctx131072.short.json` |
| 9 | 9.12 on 310 tokens; 425.1 on 48,168 | `batch-sweep-2026-09-21/GLM-5.3-Flash_4096_series.result`, fourth read, and its log |
| 9 | 75.48 on 120,979, 3 of 3 | `GLM-5.3-Flash_ctx131072.deep.json`; `first-runs-2026-09-16.tsv` |
| 10 | 60.85, 46.66, 34.61 on 481, 527, 526 tokens; 3,162.8, 1,971.1, 1,165.6 at 20,067, 99,934, 229,564 | `letters-2026-09-26/Qwen3.6-27B_ctx262144_draft-off.jsonl`, `read` and `prose` lines |
| 10 | `-b 2048 -ub 512`, 262,144 | the `cmdline:` line of `Qwen3.6-27B_ctx262144_draft-off.result` |
| 10 | draft head on at 262,144: "failed to create MTP context" | `letters-2026-09-26/Qwen3.6-27B_ctx262144_draft-on.result` |
| 10 | the cache setting referenced but never defined before 26 September | `file-listing-2026-09-26.txt`, section C, both Qwen3.6-27B blocks |
| 10 | framework, the earlier file, 2026-09-13 | `gate-records/qwen3.6-27b.json` |
| 11 | 23.50 on 32 tokens after 5,986 (the 236.50 on 5,986 left the page in the 26 September update) | `qwen3.8-flash-next-2026-09-20/q8cache_short_control.result`, `[q1 RETRIEVE]` line, and its log |
| 11 | 13.08 on 32 tokens and 197.14 on 229,982 (the 197.14 now in the row note, "a 229,982-token read at the defaults, before the batch change") | `q8cache_ctx262144_deep.result`, `[q1 RETRIEVE]`, and its log |
| 11 | both at the defaults, 8-bit cache, 262,144 | the `[load]` lines (`asked=262144 got=262144 ... kv=q8_0`); `run_hard.sh` passes no batch flag |
| 11 | 45.4 on 3,041 (2026-09-20), "the first request after a start (load 65 s)" | `batch-sweep-2026-09-20/Qwen3.8-Flash-Next_2048_ctx262144_3k.result`, `ctx=262144`, `--batch-size 4096 --ubatch-size 2048`, `load=65s`; the first request of that load. The file records no memory state, so the page gives no cause for it (corrected before push 6; the earlier text inferred one from the 26 September cold start) |
| 11 | 112.3 on 2,999 (2026-09-21; left the page in the 26 September update) | `batch-sweep-2026-09-21/Qwen3.8-Flash-Next_ctx262144_3k.result`, `served: n_ctx_slot = 262144`; the setting is the start script's default (`file-listing-2026-09-26.txt`, section C) |
| 11 | 511.7 on 48,069 at 131,072 (row note since the 26 September update) | `batch-sweep-2026-09-20/Qwen3.8-Flash-Next_2048.result`, `ctx=131072` |
| 11 | 393,216 and 524,288 on request, same setting in the script; the reads at those windows at the defaults, before the change | `file-listing-2026-09-26.txt`, section C (the accepted windows, line 84; the batch line 235 has no window condition); `q8cache_ctx393216_deep.result` and `q8cache_ctx524288_deep.result` (344,981 and 459,911 tokens, 3/3; `run_hard.sh` passes no batch flag) |
| 11 | framework, 96K leg, 2026-09-12 | `gate-records/qwen3.8-flash-next-125b.json` |
| 11 | (26 September update) at 262,144, `-b 4096 -ub 2048`, after the server's first request, with 57.7 to 60.1 GB of the 90 GB file in memory: 517.6 on 3,019 tokens, 660.6 on 48,075, 535.4 on 229,981 (429.6 s) | `qwen3.8-flash-next-2026-09-26/shipped_b4096_ub2048.jsonl`, rows `n` 2, 3 and 4: `prefill_tps`, `prompt_n`, `prompt_ms` (429,586); `resident_gb_before` of row 2, 57.66, to `resident_gb_after` of row 4, 60.05 (printed 60.1, the rounding the other pages of this push use), with the defaults rows at 60.05 to 60.06; the file is 89.99 GB (`size_gb` in `run_fn.out`) |
| 11 | at llama.cpp's defaults 255.7, 240.2 and 205.5 (1,119.3 s) | `default_b2048_ub512.jsonl` in the same folder, rows `n` 2, 3 and 4 (`prompt_ms` 1,119,327) |
| 11 | every planted code quoted, 3 of 3 at 229,981 | both files: `code_ok` true on all eight rows; `codes3_hits` 3 on each row 4 |
| 11 | the window and the two settings | `run_fn.out`: `n_ctx 262144` after each load; `run_fn.sh`: the defaults rung exports `FN_BATCH=2048 FN_UBATCH=512`, the other leaves the start script's own `-b 4096 -ub 2048` (`file-listing-2026-09-26.txt` section C, line 235). The server was started through a copy of the start script that differs from the installed one only in the line naming its settings file and in a comment the installed script gained after the run (checked with `diff`); by the run's own notes that file held the installed settings without the helper process, and we did not open it |
| 11 | row note: "every read was a new prompt" | `fn_reads.py`: a new ledger seed for each read, a planted code at the middle, or three at 5, 50 and 95 percent for the 229,981-token read; reasoning off for these reads |
| 11 | row note: the first start began with none of the file in memory and is the only start recorded with it out of memory; its first request, 3,035 tokens, read at 87.7 while the file's share in memory grew from 43.5 to 57.7 GB, 14.1 GB read in from disk, and the next, 3,019 tokens, at 517.6; the run at the defaults came second and began with 60.1 GB in memory; the file never became fully resident | `shipped_b4096_ub2048.jsonl` row 1 (`prefill_tps` 87.7; `resident_gb_before` 43.54 and `resident_gb_after` 57.66, 14.12 GB) and row 2; `run_fn.out` (0.0 GB resident before the first start); `default_b2048_ub512.jsonl` row 1 (`resident_gb_before` 60.05); no row of either file shows more than 60.06 of the 89.99 GB |
| 11 | row note: the card held 26,795 MiB at load with `-b 4096 -ub 2048` and 21,333 at the defaults | `run_fn.out`, the two `healthy in` lines |
| 12 | 227.43 on a 160-token paragraph at 202,752 | `CONTEXT_256K_MEASUREMENTS.tsv` row `GLM-4.7-Flash_202752`; the short reply JSONs |
| 12 | 129.6 on 429 tokens; 2,594.6 on 47,986 | `batch-sweep-2026-09-21/GLM-4.7-Flash_skip.result` (`reasoning_len` 1,388) and its log |
| 12 | framework, 96K leg, 2026-09-12 | `gate-records/glm-4.7-flash-31b.json` |
| 13 | 5.51 and 61.36 on 20,221 (2026-09-12) | `context-sweep-2026-09-12.tsv`, GLM-4.7 Full `long-128K` |
| 13 | the check failed at the 96K tool call | `gate-records/glm-4.7-full-358b.json`, `verdict` FAIL, `fails` |
| 14 | 202.64 on a 163-token paragraph; 3,799.69 on 230,855 | `CONTEXT_256K_MEASUREMENTS.tsv` row `Gemma-4-26B-A4B_262144`; the short and deep JSONs |
| 14 | 163.68 on 11 tokens; 8,972.3 on 47,983 | `batch-sweep-2026-09-21/Gemma-4-26B-A4B_skip.result` and its log |
| 14 | framework, 2026-09-12 | `gate-records/gemma-4-26b-a4b.json` |
| 15 | 73.68 on paragraph replies (2026-09-12) | `context-sweep-2026-09-12.tsv`, Gemma-4-31B IT QAT `long-128K` |
| 15 | 56.45 on 11 tokens; 2,687.5 on 47,983 | `batch-sweep-2026-09-21/gemma-4-31b-it-qat_skip.result` and its log |
| 15 | framework, 2026-09-12 | `gate-records/gemma-4-31b-it-qat.json` |
| 16 | 29.57 on a 172-token paragraph | `CONTEXT_256K_MEASUREMENTS.tsv` row `Ling-3.0-flash_262144`; the short reply JSONs |
| 16 | 23.25 on 77 tokens and 727.7 on 47,986; 22.6 on 135 and 694.6 on 149,711 | `batch-sweep-2026-09-21/Ling-3.0-flash_2048_series.result`, reads 2 and 3, and its log |
| 16 | framework, 2026-09-13 | `gate-records/ling-3.0-flash.json` |
| 17 | 29.55 on a 249-token paragraph (29.548) | `CONTEXT_256K_MEASUREMENTS.tsv` row `Mistral-Small-4_262144`; the short reply JSONs |
| 17 | 20.97 on 11 tokens; 2,192.3 on 48,697 at 8192; 481.3 at the defaults | `batch-sweep-2026-09-21/Mistral-Small-4_8192.result` and log; `Mistral-Small-4_skip.result` |
| 17 | framework, 2026-09-12 | `gate-records/mistral-small-4-119b.json` |
| 18 | 9.73 on a prose reply of 1,313 characters, 1,522 tokens, after a 70-token prompt | `minimax-m2.7/2026-09-13_ctx131072_q8_0.server.log`, task 261 (`70 tokens`, `1522 tokens`, `9.73 tokens per second`); `2026-09-13_extract.txt` section 3 (`content_chars=1313`, `finish=stop`) and section 1 (`-cmoe`, every expert in system memory) |
| 18 | 9.83 on 32 tokens after 43,909; 656.5 | `minimax-m2.7/2026-09-19_one_ub4096_59_ctx131072.json`, `sizes.49152.cold` (`prompt_n`, `prefill_tps` 656.48, `decode_tps` 9.834, `predicted_n`) |
| 18 | experts of 59 of 62 layers | the same JSON, `config.place` "59"; the layer count in `2026-09-13_extract.txt` section 4 (`n_layer = 62`) |
| 18 | (evening of 26 September) window: "196,608, its native ceiling and the installed default: a run that passed no window was served it (2026-09-26)" | `installed-starts-2026-09-26.txt`, MiniMax M2.7: started 21:09:33, no window passed; `n_ctx` 196608 and `n_ctx_slot = 196608`; card 31,512 MiB at load. The installed script (unchanged since 21 September, fallback 131,072, listing C line 77) sets the one shape at 196,608 itself (lines 108 to 115). "Native ceiling": the script's comment, line 35. Earlier loads by passing the window: 19 September (`minimax-m2.7/2026-09-19_one_ub1024_cmoe_ctx196608.json`) and 26 September (`minimax-m2.7/2026-09-26_ctx196608_cmoe_ub1024.result`) |
| 18 | the one shape that loads at 196,608: every expert in system memory, `-b 4096 -ub 1024`, 8-bit cache | section C lines 35, 108 to 115 (the script sets `NCMOE=cmoe; M27_UBATCH=1024` itself) and 269; the run's `cmdline:` (`--ctx-size 196608`, `--batch-size 4096 --ubatch-size 1024`, `-cmoe`, `--cache-type-k q8_0 --cache-type-v q8_0`) |
| 18 | at 196,608, counting hidden reasoning: 10.01 and 8.96 on 4,096 tokens at 20,039 and 47,933, all reasoning, no letter; 6.05 on 515 tokens, a 275-word letter, at 189,491 | `minimax-m2.7/2026-09-26_ctx196608_cmoe_ub1024.jsonl`, `prose` lines (`decode_tps`, `predicted_n` 4096, 4096, 515; `answer_words` 0, 0, 275; `reasoning_chars` 14,671, 14,785, 787), and each depth's `read` line for the token count |
| 18 | at 196,608: 167.9 on 20,039, 181.2 on 47,933, 155.9 on 189,491 (1,215.4 s); every code quoted, 3 of 3 at 189,491 | the same file, `read` lines (`prefill_tps`, `prompt_n`, `prompt_ms` 1,215,362, `code_in_answer`, `codes3_hits` 3) |
| 18 | row note: 31,709 MiB at load, 31,907 at peak | the `.result` (`VRAM at load`, `peak VRAM during probes`) |
| 18 | row note: 656.5 on 43,909 at 131,072 is 3.6 times the 181.2 on 47,933 at 196,608 | arithmetic, 656.5 / 181.2 = 3.62, across two windows, placements, micro-batches and prompt sizes, all named in the sentence |
| 18 | 131,072 accepted; the 19 September figures at that window (656.5 on 43,909, 9.83 on 32 tokens) measured with experts of 59 of 62 layers in system memory and `-b/-ub 4096` | section C line 91 (accepted windows); `minimax-m2.7/2026-09-19_one_ub4096_59_ctx131072.json` (`config.place` "59", `config.ub` 4096, `config.ctx` 131072). The page does not say which placement a 131,072 start uses now: that depends on settings this package does not publish |
| 18 | the 13 September prose figure (9.73 on 1,522 tokens) measured with every expert in system memory and `-ub 128` | `minimax-m2.7/2026-09-13_extract.txt` section 1 (the launch block: `-cmoe -ub 128`, no `-b`); the window and cache from section 3 and the log's name (`ctx131072_q8_0`) |
| 18 | row note: 45.6 on 6,776 at `-ub 128` (13 September) | `minimax-m2.7/2026-09-13_ctx65536_f16.server.log`, task 129 (`6776 tokens`, `45.58`); the flag in `2026-09-13_extract.txt` section 1 |
| 18 | row note: 100.8 on 3,658 at the default | `minimax-m2.7/2026-09-19_A_ub512_cmoe.json`, `sizes.4096.cold` |
| 18 | row note: with a 900-token limit the request returned no answer | `2026-09-13_extract.txt` section 3 (`finish=length`, `completion_tokens=900`, `content_chars=0`) |
| 18 | row note: `-ub 128` "copied from another model's launcher" | our session notes; stated, not a measurement |
| 18 | framework, 2026-09-21, 3 of 3 and the tool call, PASS, 131,072 window | `gate-records/minimax-m2.7-2026-09-21.json`, `l64`, `l128`, `verdict`, `fails`, `S.n_ctx` 131072 |
| 19 | 8.90 on 82 tokens at 3,680 | `minimax-m3/2026-09-20_q2kl_ctx131072_ub2048.server.log`, task 4 |
| 19 | 208.11 on 58,307; 8.29 on 165 tokens; 3 of 3 | the same log, task 216; `2026-09-20_q2kl_needle_58k.json`, `score` |
| 19 | 186.8 on 3,009 through the model manager (2026-09-21) | `batch-sweep-2026-09-21/MiniMax-M3_model-manager_3k.result` |
| 19 | the file it serves carries the sparse-attention index; the launch check reports it engaged | `file-listing-2026-09-26.txt` (the Q2_K_L shards); `minimax-m3/2026-09-20_q2kl_ctx131072_ub2048.launch-check.log` (`MSA ENGAGED`) |
| 19 | the earlier file lacked the tensors | our session notes; stated, not a measurement |
| 19 | at 262,144 no micro-batch from 512 to 2048 could allocate | `minimax-m3/2026-09-20_q2kl_ctx262144_attempts.log` (buffers of 20,572,169,216, 10,286,650,368 and 5,143,890,944 bytes) |
| 19 | 196,608 loads only at `-ub 512`; 58,307 tokens at 70.16, 3 of 3 | `minimax-m3/2026-09-20_q2kl_ctx196608_attempts-and-launch-check.log` (2048 and 1024 fail); `2026-09-20_q2kl_ctx196608_ub512.server.log`, task 271; `2026-09-20_q2kl_ctx196608_needle_58k.json` |
| 20 | 15.74 and 14.69; 898.8 on 24,013 and 911.6 on 150,158 through the model manager | `batch-sweep-2026-09-21/Laguna-S-2.1_ctx262144_model-manager.result` |
| 20 | 12.73 on 35 tokens and 866.1 on 200,098, `-ub 4096` | `Laguna-S-2.1_ctx262144_4096_series.result`, read 2, and its log |
| 20 | row note: 72.2 at `-ub 128` and 1,434.0 at 8192, on 24,071 at 32,768 | `Laguna-S-2.1_baseline_ub128.result`; `Laguna-S-2.1_8192.result` (`served: n_ctx_slot = 32768`) |
| 21 | 193.65 on 139 tokens; 2,737.7 on 48,090 | `batch-sweep-2026-09-21/Muse-Glimmer-30B_skip.result` (`reasoning_len` 448) and its log |
| 22 | did not answer within the check's timeout, 2026-09-12 | `tool-recall-wave1-2026-09-12.tsv`, its `fails` cell; `gate-records/mistral-medium-3.5-128b.json` |
| 22, 23 | 32,768 | `file-listing-2026-09-26.txt`, section C |

Windows as served (fix pass, 26 September). Every window cell now says what shows it.
`installed-windows-2026-09-21.txt` lists, for each endpoint that has one, a run on 21 September that started the model through
its own start script **without passing a window**, and the window the server then reported: the crash check
(`batch-sweep-2026-09-21/stress_chain.sh`, `stress_model.py`, summary `STRESS.txt`) for rows 2 (262,144), 3 (262,144
on both presets), 4, 5, 9, 17 and 18 (131,072) and 11 (262,144), and rows 6, 8, 9 and 18 as they were before their change of 26 September (131,072); the starts of those four with no window passed on the evening of 26 September are in `installed-starts-2026-09-26.txt` (section E of the windows file); the batch sweep's `_skip` runs
(`batch_sweep.sh` passes no argument; each `_skip.result` and `_skip.log` is in `batch-sweep-2026-09-21/`) for rows 15 and 21 (131,072), 12 (202,752), 14 and 16 (262,144); row 7's shows the 131,072 it served on 21 September, before its change of 26 September (see the row 7 lines above). Row 20's
262,144 is its start script's own fallback, `CTX="${1:-262144}"`, and that script reads no settings file
(`file-listing-2026-09-26.txt` section C); row 19's 131,072 likewise, `CTX="${1:-131072}"`. Rows 1 and 10: every run that loaded 262,144 passed it explicitly, so the cell gives the
script's fallback (131,072) and the dates 262,144 was loaded (`installed-windows-2026-09-21.txt` section C). Row
13's 131,072 with a 5-bit cache is its 12 September framework record's `S.n_ctx` and the 12 September sweep's
`long-128K` row (`kv` q5_1). Rows 22 and 23: the start scripts' default profile and accepted windows (listing
section C). The batch setting printed with each window is the start script's setting at that window (listing
section C); `batch_sweep.sh` records that the settings files do not set the batch sizes.

### Section 03 rows and notes changed in the fix pass (26 September)

| row | figure on the page | file and field |
|---|---|---|
| 1 | 14.76, 12.04 and 7.66 on the short code answers of 67, 69 and 83 tokens | `DeepSeek-V4-Flash-pair_512_series.result` and `..._512_150k.result`: `decode_tps`, `gen_n` |
| 1 | the framework record's verdict is FAIL, on a separate leg | `gate-records/deepseek-v4-flash-0731-two-machines.json`: `verdict`, `fails` |
| 1 and 3 | 6.7 and 12.3 times, each machine at its own setting | 690.9 / 102.5 and 632.1 / 51.4 (arithmetic) |
| 1 and 3 | at `-ub 2048`: the pair 224.9 on 2,998 and 117.0 on 48,024; the desktop (`-b 4096`) 208.1 and 426.5; 3.6 times | `DeepSeek-V4-Flash-pair_2048_series.result`, `DeepSeek-V4-Flash-3bit-desktop_2048_series.result`; the `-b` values in `DeepSeek-V4-Flash-pair_chain.sh` and `series_probe.sh` |
| 1 and 3 | whole requests at 48,024 tokens: 75.2 s alone, 474.2 s on the pair kept, 416.4 s at `-b/-ub 2048` | `wall_s` in `DeepSeek-V4-Flash-3bit-desktop_4096_series.result` (`#2`), `DeepSeek-V4-Flash-pair_512_series.result` and `DeepSeek-V4-Flash-pair_2048_series.result` (`size 48000`) |
| 5 and 10 | the quote-the-code read passed at every depth; the edit-and-reread request, capped at 40 tokens of reply, did not return the code at any depth | `letters-2026-09-26/Qwen3-235B_2048.jsonl` and `Qwen3.6-27B_ctx262144_draft-off.jsonl`: `kind` `read` (`code_in_answer` true) and `edit_reread` (`code_in_answer` false, `predicted_n` 40); the cap is `max_tokens` 40 in `prose_probe.py` |
| 20 | the old script set `-b 512 -ub 128`; it now uses `-b 8192 -ub 4096` above 32,768 | `file-listing-2026-09-26.txt` section C (the saved copy's lines 54 and 55; the current script's lines 68 and 69) |
| hero | Laguna: 72.2 at `-b 512 -ub 128` and 1,434.0 at `-b/-ub 8192` at 32,768; 898.8 on 24,013 at 262,144 and `-b 8192 -ub 4096` | `batch-sweep-2026-09-21/Laguna-S-2.1_baseline_ub128.result` and `Laguna-S-2.1_8192.result` (`n_ctx_slot = 32768`); `Laguna-S-2.1_ctx262144_model-manager.result` (898.8 on 24,013; the series run at `-ub 4096`, `Laguna-S-2.1_ctx262144_4096_series.result`, read 901.9 on the same prompt); the batch lines as row 20 |


## Section 04, what a batch setting did

| page says | file |
|---|---|
| the six rows of the table | `batch-sweep-2026-09-21/<model>_skip.result` and `<model>_<setting>.result` for Qwen3.5-397B-A17B (4096), Qwen3.5-122B-A10B (4096), Mistral-Small-4 (8192), Ornith-1.5-35B (4096), Qwen3-235B-A22B-Instruct-2507 (2048); GLM-5.3-Flash `_skip` (24,008 tokens) and `_4096_series` (read 4, 48,168) |
| every setting that loaded quoted the code | the same results, `needle_in_answer` true |
| speaking moved by 6.2 percent at most, 30.81 to 32.72 | the same results, `decode_tps` (arithmetic: 32.72 / 30.81 = 1.062) |
| at 262,144 the scripts of the Qwen3.5 pair and Mistral Small 4 keep `-b 2048 -ub 512` because the larger buffer does not fit; since 26 September Ornith's has used `-b 2048 -ub 2048` there and GLM-5.3-Flash's `-b 4096 -ub 1024` | `file-listing-2026-09-26.txt` section C, rows 4, 6, 8, 9 and 17 (the `if [ "$CTX" -gt 131072 ]` lines and the comments above them) |
| the models kept at the defaults: on each one's best rung, from a 0.6 percent loss (Muse Glimmer, 2,721.3 against 2,737.7 at 1024) to a 34.5 percent gain (Gemma-4-26B-A4B at 2048, 12,063.4 against 8,972.3, speaking 153.82 against 163.68, 6 percent slower); Qwen3.6-27B's rungs are the earlier Q4 file at 131,072; GLM-4.7-Flash at 2048 returned an empty answer (`needle_in_answer` false, `answer_len` 0, `needle_found` true) | `batch-sweep-2026-09-21/{Qwen3.8-27B,Qwen3.6-27B,GLM-4.7-Flash,Gemma-4-26B-A4B,gemma-4-31b-it-qat,Muse-Glimmer-30B}_{1024,2048,4096}.result` against each `_skip.result`: Gemma-4-26B-A4B at 2048 read 12,063.4 against 8,972.3 (+34 percent) and spoke 153.82 against 163.68; Muse Glimmer's best rung read 2,721.3 against 2,737.7 |
| Qwen3.8-27B's rungs were run at 131,072, its installed window on 21 September | `batch-sweep-2026-09-21/Qwen3.8-27B_skip.result` (`n_ctx_slot = 131072`); the row 7 lines above |

## Section 05, what changed since 15 September

`summary-table-corrections-2026-09-26.csv` carries the table with, for each line, where the old figure stood and
the package files that correct it. The old figures sat in our operator records and session results, which are not
published; the 15 September list's own figures are in `ROSTER_22.csv`.

| line | the correcting file and field |
|---|---|
| Qwen3-235B 8.2 | `Qwen3-235B-A22B-Instruct-2507_2048.result` (8.18) and its log (11 tokens); the letters |
| Qwen3.5-397B 385 | the 385's prompt size and window come from the model's spec card as quoted in our records (August); stated. The September figures are in the row 4 files; 215.5 against 195.2 are `l128.A_recall.prefill_tps_approx` in the two framework records |
| GLM-4.7 Full passed | `gate-records/glm-4.7-full-358b.json`, `verdict` |
| MiniMax M3 194 | see point 4 at the top |
| MiniMax M2.7 45.6 on "4,000" | `minimax-m2.7/2026-09-13_extract.txt` section 2: the probe line `("long(~4k tok)", longp)` and the recorded `(6776 tok)` |
| Laguna 72 (old window, `-b 512 -ub 128`) and 1,434.0 (`-b/-ub 8192`); 898.8 on 24,013 at 262,144 now | `Laguna-S-2.1_baseline_ub128.result` and `Laguna-S-2.1_8192.result` (`n_ctx_slot = 32768`); the old script's batch lines (`file-listing-2026-09-26.txt` section C); `Laguna-S-2.1_ctx262144_model-manager.result` |
| Qwen3.8-Flash-Next: our records' 511.7 and 45.4; the 26 September served-window figures | the 511.7 and 45.4 as row 11 above; `batch-sweep-2026-09-20/Qwen3.8-Flash-Next_2048.result` (`ctx=131072`); the new figures as row 11's 26 September update rows; where our records stood: `summary-table-corrections-2026-09-26.csv` row 7 |
| the pair at 262,144 on 21 September: 14.73 and 12.05 on replies of about 400 tokens; 14.76 and 12.04 on the short code answers | `DeepSeek-V4-Flash-pair_512_series.result`: `warm_decode_tps`, `warm_gen_n`; `decode_tps`, `gen_n` (67, 69) |
| the pair's 15.9 to 17.1 | `ROSTER_22.csv` row 1 and `two-box-deepseek-probes-2026-09-13-to-15.txt` run 5 (13 September, 131,072); the 21 September figures in row 1's files |
| 22 endpoints; ten rows now serving more than 131,072 by default, each shown by a run that passed no window (rows 2, 3, 6, 8, 9, 11, 12, 14, 16, 18; the four of 26 September by `installed-starts-2026-09-26.txt`), Laguna S 2.1 (row 20) serving 262,144 by its start script's own default (`file-listing-2026-09-26.txt` section C, `CTX="${1:-262144}"`; the only no-window run of it in this package, the 21 September crash check at 14:25, served 32,768 before its change), Qwen3.8-27B installed at 262,144 on 26 September (the row 7 lines above), two more that loaded 262,144 when it was passed | `ROSTER_22.csv`; `ROSTER_23.csv`, `window_and_batch_setting` (rows 2, 3, 11, 12, 14, 16, 20; and rows 1, 10); `installed-windows-2026-09-21.txt` |
| every read that asked for a planted code returned it in its answer, except GLM-4.7-Flash at `-ub 2048`; the letter runs' edit-and-reread requests did not return it at any depth | every file named in section 03 with `needle_in_answer`, `needle_found`, `hits`, `score`, `codes_hit` or `code_in_answer`; the exception is `batch-sweep-2026-09-21/GLM-4.7-Flash_2048.result`; the `edit_reread` lines of `letters-2026-09-26/*.jsonl` (`code_in_answer` false, `predicted_n` 40, the reply cap `prose_probe.py` sets for that request) |

## Section 06, on request

| page says | file |
|---|---|
| Qwen3.5-397B and Mistral Small 4 read about 231,000 tokens at 262,144 (15 September, at `-b 2048 -ub 512`, which their scripts keep there; Qwen3.5-122B left this sentence on 26 September, when 262,144 became its installed window) | `CONTEXT_256K_MEASUREMENTS.tsv`, rows `Qwen3.5-397B-A17B_262144` and `Mistral-Small-4_262144` (230,803 and 231,065; 3/3) |
| Qwen3.8-Flash-Next read 344,981 and 459,911 at its two larger windows, at llama.cpp's default batch sizes, before its batch change | `qwen3.8-flash-next-2026-09-20/q8cache_ctx393216_deep.result`, `q8cache_ctx524288_deep.result`; `run_hard.sh` passes no batch flag |
| MiniMax M3 58,307 at 196,608 (`-ub 512`) (MiniMax M2.7 left this sentence on the evening of 26 September, when 196,608 became its installed window) | see row 19 above |
| (left the page on the evening of 26 September, when Ornith-1.5-35B was installed at 262,144 and measured there) Ornith-1.5-35B's script offers 262,144 on a 17 September run whose logs were not kept | `file-listing-2026-09-26.txt`, section C (the `262144 MEASURED 2026-09-17` comment); no run record exists |
| eight rows have prose at depth (1, 5, 6, 7, 8, 9, 10 and 18), and most have a paragraph at a short prompt (row 7 added before push 6; rows 6, 8, 9 and 18 on the evening of 26 September; row 18's is one 275-word letter at 189,491, row 9's a 351-word letter at 47,992 and a 337-word one at 100,003) | row 1: `batch-sweep-2026-09-21/DeepSeek-V4-Flash-pair_512_series.result` (`warm_gen_n` about 400); rows 5 and 10: `letters-2026-09-26/*.jsonl` `prose` lines; row 7: `qwen3.8-27b-2026-09-26/Qwen3.8-27B_ctx262144_draft-off.jsonl` `prose` lines (367 to 411 words); rows 6 and 8: the `prose` lines of `qwen3.5-122b-2026-09-26/Qwen3.5-122B-A10B_ctx262144_b2048_ub512.jsonl` (407 to 426 words) and `ornith-1.5-35b-2026-09-26/Ornith-1.5-35B_ctx262144_b2048_ub2048.jsonl` (328 to 348); the paragraph replies as in section 02 |

## What has no file in this package, stated plainly

1. The installation and retirement dates in section 01, and Qwen3.6-27B's served window: our operator records.
2. That MiniMax M2.7's `-ub 128` was copied from another model's launcher, and that MiniMax M3's earlier file lacked
   the sparse-attention tensors: our session notes.
3. The 385 for Qwen3.5-397B and its prompt size, and the other old figures in section 05: our own records as they
   stood, which are not published; each correcting figure is in the package.
4. The processor model: the machine's specification, stated on the site's method page.
5. Every framework reading figure is the harness's estimate, not a server timing; the page says so.
