# Number map

The TSV’s `decode_tps` is the best of two short-prompt prose replies, not the at-depth code-answer rate. The page uses `deep_decode_tps`, corroborated by each deep response’s `timings.predicted_per_second`. The TSV includes an extra 131,072-window split run; the main table includes only the ten large-window configurations. If a number on the page disagrees with a file in its package, the file is right and the page is wrong.

All main-table results are measured on September 15, 2026, one deep request per configuration. Model names, quantization names, code strings, dates, revision hashes and section numbers are identifiers. Rounding is to one decimal for read seconds/minutes and two decimals for table rates. GPU memory stays MiB.

| Configuration | Printed figures | Primary record and fields |
|---|---|---|
| Gemma-4-26B-A4B, UD-Q4_K_XL | Window 262,144; peak 23,513 MiB; read 230,855 tokens in 60.8 s at 3,799.69 tokens/s; answer 33 tokens at 121.36 tokens/s; 3/3 | `logs256/gemma26_262144.deep.json`: `timings.prompt_n`, `prompt_ms / 1000`, `prompt_per_second`, `predicted_n`, `predicted_per_second`, `choices[0].message.content`; `CONTEXT_256K_MEASUREMENTS.tsv`: row `gemma26_262144`, `n_ctx`, `vram_peak_mib`, `codes_hit` |
| GLM-4.7-Flash, UD-Q4_K_XL | Window 202,752; peak 28,765 MiB; read 189,931 tokens in 231.7 s at 819.70 tokens/s; answer 29 tokens at 70.48 tokens/s; 3/3 | `logs256/glm47flash_202752.deep.json`: `timings.prompt_n`, `prompt_ms / 1000`, `prompt_per_second`, `predicted_n`, `predicted_per_second`, `choices[0].message.content`; `CONTEXT_256K_MEASUREMENTS.tsv`: row `glm47flash_202752`, `n_ctx`, `vram_peak_mib`, `codes_hit` |
| Qwen3.5-122B-A10B, UD-Q4_K_S | Window 262,144; peak 30,794 MiB; read 230,803 tokens in 531.8 s at 433.99 tokens/s; answer 32 tokens at 37.10 tokens/s; 3/3 | `logs256/qwen122_262144.deep.json`: `timings.prompt_n`, `prompt_ms / 1000`, `prompt_per_second`, `predicted_n`, `predicted_per_second`, `choices[0].message.content`; `CONTEXT_256K_MEASUREMENTS.tsv`: row `qwen122_262144`, `n_ctx`, `vram_peak_mib`, `codes_hit` |
| Mistral Small 4, UD-Q4_K_M | Window 262,144; peak 30,356 MiB; read 231,065 tokens in 705.0 s at 327.75 tokens/s; answer 32 tokens at 23.76 tokens/s; 3/3 | `logs256/mistral4_262144.deep.json`: `timings.prompt_n`, `prompt_ms / 1000`, `prompt_per_second`, `predicted_n`, `predicted_per_second`, `choices[0].message.content`; `CONTEXT_256K_MEASUREMENTS.tsv`: row `mistral4_262144`, `n_ctx`, `vram_peak_mib`, `codes_hit` |
| Ling-3.0-flash, Q4_K_M | Window 262,144; peak 8,473 MiB; read 230,759 tokens in 1,118.8 s at 206.25 tokens/s; answer 32 tokens at 26.11 tokens/s; 3/3 | `logs256/ling_262144.deep.json`: `timings.prompt_n`, `prompt_ms / 1000`, `prompt_per_second`, `predicted_n`, `predicted_per_second`, `choices[0].message.content`; `CONTEXT_256K_MEASUREMENTS.tsv`: row `ling_262144`, `n_ctx`, `vram_peak_mib`, `codes_hit` |
| Qwen3.5-397B-A17B, UD-IQ3_XXS | Window 262,144; peak 31,068 MiB; read 230,803 tokens in 1,186.6 s at 194.51 tokens/s; answer 32 tokens at 17.00 tokens/s; 3/3 | `logs256/qwen397_262144.deep.json`: `timings.prompt_n`, `prompt_ms / 1000`, `prompt_per_second`, `predicted_n`, `predicted_per_second`, `choices[0].message.content`; `CONTEXT_256K_MEASUREMENTS.tsv`: row `qwen397_262144`, `n_ctx`, `vram_peak_mib`, `codes_hit` |
| Qwen3.8-Flash-Next, UD-Q3_K_XL | Window 262,144; peak 15,626 MiB; read 230,803 tokens in 1,392.9 s at 165.70 tokens/s; answer 32 tokens at 13.58 tokens/s; 3/3 | `logs256/flashnext_262144.deep.json`: `timings.prompt_n`, `prompt_ms / 1000`, `prompt_per_second`, `predicted_n`, `predicted_per_second`, `choices[0].message.content`; `CONTEXT_256K_MEASUREMENTS.tsv`: row `flashnext_262144`, `n_ctx`, `vram_peak_mib`, `codes_hit` |
| Inkling-Small, UD-Q3_K_XL | Window 262,144; peak 15,814 MiB; read 230,827 tokens in 2,005.9 s at 115.07 tokens/s; answer 27 tokens at 8.78 tokens/s; 3/3 | `logs256/inkling_262144.deep.json`: `timings.prompt_n`, `prompt_ms / 1000`, `prompt_per_second`, `predicted_n`, `predicted_per_second`, `choices[0].message.content`; `CONTEXT_256K_MEASUREMENTS.tsv`: row `inkling_262144`, `n_ctx`, `vram_peak_mib`, `codes_hit` |
| DeepSeek V4 Flash, desktop, UD-Q8_K_XL | Window 262,144; peak 18,306 MiB; read 150,475 tokens in 2,062.2 s at 72.97 tokens/s; answer 25 tokens at 9.14 tokens/s; 3/3 | `logs256/dsv4q8_262144.deep.json`: `timings.prompt_n`, `prompt_ms / 1000`, `prompt_per_second`, `predicted_n`, `predicted_per_second`, `choices[0].message.content`; `CONTEXT_256K_MEASUREMENTS.tsv`: row `dsv4q8_262144`, `n_ctx`, `vram_peak_mib`, `codes_hit` |
| DeepSeek V4 Flash, desktop + laptop, UD-IQ3_XXS | Window 262,144; peak 25,824 MiB; read 230,835 tokens in 6,580.7 s at 35.08 tokens/s; answer 25 tokens at 6.49 tokens/s; 3/3 | `logs256/dsv4split_262144.deep.json`: `timings.prompt_n`, `prompt_ms / 1000`, `prompt_per_second`, `predicted_n`, `predicted_per_second`, `choices[0].message.content`; `CONTEXT_256K_MEASUREMENTS.tsv`: row `dsv4split_262144`, `n_ctx`, `vram_peak_mib`, `codes_hit` |

| Other figure or claim | File / field / derivation |
|---|---|
| Nine models; ten configurations; nine desktop-only and one split | First ten rows of `CONTEXT_256K_MEASUREMENTS.tsv`; two DeepSeek configurations share the model family but differ in quantization and placement; commands in `context256_bodies.sh` and the server logs |
| Read range 150,475 to 231,065; peak range 8,473 to 31,068 MiB; short answers 25 to 33 tokens | Min/max of the ten main rows above, not the extra 131,072 row |
| 60.8 s shortest; 6,580.7 s / 109.7 min longest | `logs256/gemma26_262144.deep.json`: `prompt_ms=60756.249`; `logs256/dsv4split_262144.deep.json`: `prompt_ms=6580695.006`; divide by 1000 or 60000, round to one decimal |
| 262,144 meaning of 256K; 202,752 GLM window | TSV `n_ctx`; `[result]` lines corroborate served capacity; 202,752 is the requested window, confirmed by the result; the launch notes call it the ceiling, but no failed attempt above it is recorded |
| Three codes at about 5%, 50%, 95%; one deep trial each | `deep_recall_probe.py`: `CODES`, `build()` marks by entry index, `QUESTION`; `context256_rung.sh`: one deep request per invocation |
| Every code answer in exact order | Each main-row deep response `choices[0].message.content`, checked against `CODES`; no extra words |
| Sampling temperature 0, nonstreaming | `deep_recall_probe.py`: request object; requests override launch defaults |
| One slot; 24 compute and batch threads; MTP maximum 6 and minimum draft probability 0.75 | `context256_bodies.sh` COMMON and per-model argv; server logs `[cmd]`; drafts in Qwen response `timings.draft_n` and `draft_n_accepted` |
| Quantizations, files, expert-layer counts, RPC split 18,25 and layers 8 to 17, cache and placement options | Server logs `[cmd]`; derived `CONFIGURATIONS.md`; filenames identify quantization. `n-cpu-moe` is an expert-layer count, not an individual-expert count. |
| Flash-Next: 99 expert layers requested on CPU; fit off | `logs256/flashnext_262144.log` `[cmd]`: `--n-cpu-moe 99 --fit off`; no claim that 99 covers every expert layer |
| Most configurations kept expert weights in system RAM | Eight of the ten main commands use `--n-cpu-moe`, `-cmoe`, or CPU expert-tensor overrides; `logs256/*.log` `[cmd]`. Card-use ranges are sampled `nvidia-smi` readings, not total model memory. |
| Build hashes | `build-records.tsv`, derived from original deep response build identifiers before identifier-field removal; no inferred later build |
| 5-second peak sampling; load times upper bounds during downloads; best of two short-prompt replies | `context256_rung.sh`: `sleep 5`, warm decode loop `1 2`; TSV header caveat and `load_s` |
| Cached prefixes present | Deep JSON `timings.cache_n`, retained unchanged; processed token counts use `prompt_n` |
| DeepSeek September 20 (collection date): 262,144 window, 150,324 tokens, 313,172.17 ms, 480.00 tokens/s | `later/v_W_256k_ub8192.log`: initialization and final `prompt eval time` |
| DeepSeek 8192 / 8192; 26,023 MiB at load; all experts CPU-side | `later/v_W_256k_ub8192.result`: `[load]`, `args`, `vram`; load memory is not a request peak |
| DeepSeek 313.2 s / 5.2 min versus 2,062.2 s / 34.4 min | Later prompt ms / 1000 or 60000; earlier `logs256/dsv4q8_262144.deep.json` `timings.prompt_ms=2062226.608` / 1000 or 60000; no causal ratio printed |
| Laguna September 21 (collection date): 262,144; micro-batch 4096; 200,098 tokens; 866.1 prefill tokens/s, not wall ÷ tokens; 233.9 s request wall time; peak 27,995 MiB; string found | `later/Laguna_ctx262144_ub4096.result`: `served`, header, second JSON row `prompt_n`, `prefill_tps`, `wall_s`, `needle_found`; final peak. First request confirms sequence. No prompt-time equivalence claimed for wall time. |
| Ornith-35B: Q6_K, 6 CPU expert layers, MTP, 262,144, refused 512.00 MiB; Ornith-397B: IQ3_XXS-UD, 57 CPU expert layers, MTP, 262,144, refused 1,862.00 MiB | `failures/ornith35_262144.log` and `failures/ornith397_262144.log`: `[cmd]`, allocation failures. The failure logs and audition table have no calendar date; no September 16 date is assigned. |
| GLM Full setup note: 131,072, 93 versus 92 expert layers | `later/GLM-4.7-Full-config-excerpt.txt`: 128K comments and all-expert note; 128 × 1024 = 131,072 (arithmetic). Local configuration note, not a raw measured recall row. |
| MiniMax M3: 262,144, Q2_K_L, q8_0, experts on CPU (cmoe in each test label), micro-batches 2048 / 1024 / 512; buffers 20,572,169,216 / 10,286,650,368 / 5,143,890,944 bytes | `later/MSA_256K.console.log`: headers and corresponding failure lines. September 2026 is collection provenance; no exact timestamp in file. |
| 32 GB card / 188 GiB desktop RAM / 128 GB laptop unified memory; hardware identities | Bench description supplied with the study, not capacity measurements in the logs. The page labels these hardware descriptions. |
| September 15 to 21 study span / September 26 publication | Main response `created` timestamps and dated sweep provenance; later result provenance in README; publication metadata is editorial. |

The title, subtitle, card copy and feed repeat only these scoped counts, ranges and converted times. No later duration is extrapolated to another row. The GLM excerpt includes an original timing figure not repeated on the page; it remains source context, not a new measured result adopted here.

## Added 2026-09-26 (PM): Qwen3.8-Flash-Next at 262,144, two batch settings (fixed after the pre-publish check)

| Figure on the page | File | Field |
|---|---|---|
| 229,981 tokens, 535.4 t/s, 429.6 s (4096 / 2048); 3 of 3 codes | `flashnext-fullwindow/shipped_b4096_ub2048.jsonl`, row 4 | `prompt_n`, `prefill_tps`, `codes3_hits`; 429.6 s = `prompt_ms` 429,586 / 1000 |
| 229,981 tokens, 205.5 t/s, 1,119.3 s (2048 / 512); 3 of 3 codes | `flashnext-fullwindow/default_b2048_ub512.jsonl`, row 4 | same fields; 1,119.3 s = `prompt_ms` 1,119,327 / 1000 |
| 3,035 tokens at 87.7 t/s, the first request on a server started with none of the file in memory; the file's share in memory 43.5 to 57.7 GB during it; the next request, 3,019 tokens at 517.6 | `flashnext-fullwindow/shipped_b4096_ub2048.jsonl`, rows 1 and 2; `flashnext-fullwindow/run.console.log`, first line (0 bytes resident before the start) | `prompt_n`, `prefill_tps`, `resident_gb_before` 43.54, `resident_gb_after` 57.66 |
| About 60 GB of the 90 GB file in memory before both long reads | both JSONL files, row 4; `run.console.log`, first line (three shards, 89,986,353,824 bytes) | `resident_gb_before` 59.48 and 60.06 |
| Placement: 40 expert layers requested on CPU, fit off, 8-bit (`q8_0`) cache, for both rows | `flashnext-fullwindow/launch-flags.txt` (the launch copy's flags, paths omitted) | the run's own notes record the same: ctx 262144, q8_0 KV, `--n-cpu-moe 40` |
| 26,795 and 21,333 MiB at load | `flashnext-fullwindow/run.console.log` (the "healthy in" line of each start) | `nvidia-smi` memory used, read once the server answered |

## Added 2026-09-26 (push 7): Ornith-1.5-35B and Qwen3.5-122B-A10B at 262,144

| Figure on the page | File | Field |
|---|---|---|
| Ornith-1.5-35B, 2048 / 2048: 229,090 tokens, 71.7 s, 3,193.3 tokens/s, 31,270 MiB at load, 3 of 3 codes | `ornith35-fullwindow/orn_256k_b2048_ub2048.result` and `.server.log` | header `VRAM at load 31270 MiB`, `served: n_ctx_slot = 262144`; the JSON row with `depth_target` 230000 and `kind` read: `prompt_n`, `prefill_tps`, `codes3_hits` 3; server log `prompt eval time = 71741.90 ms / 229090 tokens`, / 1000 and rounded |
| Ornith-1.5-35B, 2048 / 512: 229,090 tokens, 151.8 s, 1,508.7 tokens/s, 30,250 MiB at load, 3 of 3 codes | `ornith35-fullwindow/orn_256k_b2048_ub512.result` and `.server.log` | same fields; `151841.86 ms` |
| Qwen3.5-122B-A10B, 2048 / 512: 229,090 tokens, 510.6 s, 448.7 tokens/s, 30,898 MiB at load, 3 of 3 codes | `qwen122-fullwindow/q122_256k_b2048_ub512.result` and `.server.log` | same fields; `510585.60 ms` |
| Qwen3.5-122B-A10B, 2048 / 1024: 229,090 tokens, 310.0 s, 738.9 tokens/s, 31,633 MiB at load, 3 of 3 codes; peak 31,785 MiB over the series | `qwen122-fullwindow/q122_256k_b2048_ub1024.result` and `.server.log` | same fields; `310049.67 ms`; `peak VRAM during probes: 31785 MiB` |
| Placements, cache, draft head, mmap, threads (the four row labels; "same file, placement, cache and day" on each second row) | each `.result`, the `cmdline:` line | `--n-cpu-moe 6` or `38`; no `--cache-type` flag (llama.cpp's default 16-bit cache); `--spec-type draft-mtp` on the Qwen3.5 runs only; `--no-mmap` on the Qwen3.5 runs only; `--threads 24 --threads-batch 24` |
| Qwen3.5-122B-A10B: the launch of its September 15 row; no batch values then, which at this build means 2048 / 512; same build | `logs256/qwen122_262144.log` `[cmd]` against the 26 September `cmdline:`; `build-records.tsv` (c8e03ce); `qwen122-fullwindow/launch-notes.txt` | the flag lists match apart from the written-out batch values; llama.cpp's defaults at c8e03ce from its `common/common.h` lines 443 and 444 (recorded in the notes, not shipped) |
| 230,803 tokens in 531.8 s at 433.99 tokens/s (September 15) | `logs256/qwen122_262144.deep.json` | the main-table row above |
| 31,785 MiB over the 2048 / 1024 series, 822 MiB below the card's total of 32,607 MiB as reported that day | `qwen122-fullwindow/q122_256k_b2048_ub1024.result` (`peak VRAM during probes: 31785 MiB`); the total: `/qwen38-256k/data/records/card-capacity.txt`, published with Qwen3.8-27B at 256K (`same-card total 32607 MiB`, a card reading of 26 September) | arithmetic: 32,607 - 31,785 = 822 |
| Its start script keeps 512 at this window | `qwen122-fullwindow/launch-notes.txt` | the installed script's comment, verbatim (it gives the spare as about 0.8 GB; the page prints the arithmetic above instead) |
| Ornith-1.5-35B: an Ornith-1.5-35B Q6_K file of the same name as the one in the failed launch, 6 expert layers in RAM and default cache, without the draft head; the failure launch requested MTP | `failures/ornith35_262144.log` `[cmd]` (`--spec-type draft-mtp`) against the 26 September `cmdline:` (no `--spec-type`); `ornith35-fullwindow/launch-notes.txt` | model filename `Ornith-1.5-35B-Q6_K.gguf` in both; result rows' `draft_n` null |
| Section 05: these runs wrote the batch out (2048 / 512, or 2048 / 2048), used a 7,200-second timeout and set no CORS flag; the failed launch set no batch, a 3,600-second timeout and a localhost CORS flag; its preflight card reading 1,113 MiB, none before launch on 26 September; the directories differ and neither log records a hash or a file size | `failures/ornith35_262144.log` line 1 (`[preflight] gpu=1113MiB`) and `[cmd]` (`--timeout 3600 --cors-origins localhost`, no `--batch-size` or `--ubatch-size`); each `.result`'s `cmdline:` (`--batch-size 2048 --ubatch-size 512` or `2048`, `--timeout 7200`, no `--cors-origins`) and header (`VRAM at load` only) | the directories are in the unredacted primary records, removed here; no `file size`, hash or checksum line in either log |
| The 2048 / 512 row ran at the start-script copy's own setting then; 2048 / 2048 was passed in for the other row and is what the script now sets at this window | the two `.result` headers (`ORNITH_35B_BATCH=2048` and `ORNITH_35B_UBATCH=2048` on `orn_256k_b2048_ub2048` only); `ornith35-fullwindow/launch-notes.txt` | the installed script's comment, verbatim ("the default is 2048/2048 ... it was 2048/512") |
| Thinking off per request; a fresh ledger at each depth; the deep read after three shorter ones; three codes in the deepest | each `.result`'s JSON rows | `reasoning_chars` 0 in every row; `depth_target` 20000, 48000, 100000, 230000 in that order; `codes3_in_answer` on the deepest read only |

## Added 2026-09-26 (push 7c): MiniMax M2.7 at 196,608

| Figure on the page | File | Field |
|---|---|---|
| MiniMax M2.7, 4096 / 1024: 189,491 tokens, 1,215.4 s, 155.9 tokens/s, 31,709 MiB at load, 3 of 3 codes | `minimax-m27-fullwindow/m27_192k_cmoe_ub1024.result` and `.server.log` | header `VRAM at load 31709 MiB`, `served: n_ctx_slot = 196608`; the JSON row with `depth_target` 190000 and `kind` read: `prompt_n`, `prefill_tps`, `codes3_hits` 3; server log `prompt eval time = 1215361.54 ms / 189491 tokens`, / 1000 and rounded |
| Earlier reads on the same server: 20,039 tokens at 167.9 and 47,933 at 181.2 tokens a second | the same `.result` and `.server.log` | JSON rows with `depth_target` 20000 and 48000 and `kind` read (`prompt_n`, `prefill_tps`); server log `119375.45 ms / 20039 tokens` (167.87) and `264543.67 ms / 47933 tokens` (181.19) |
| Peak 31,907 MiB over the series | the same `.result` | `peak VRAM during probes: 31907 MiB` |
| 196,608, the context length its model file declares (the row and the section 04 intro) | `minimax-m27-fullwindow/launch-notes.txt` | GGUF key `minimax-m2.context_length` = 196608 in the first shard, read for this package (the model file is not shipped); `n_ctx_slot = 196608` in the server log |
| Every expert layer in RAM, -b 4096 -ub 1024, 8-bit cache (the row label) | the `.result`, the `cmdline:` line | `-cmoe`; `--batch-size 4096 --ubatch-size 1024`; `--cache-type-k q8_0 --cache-type-v q8_0`; no `--spec-type` |
| Its start script always uses that shape at this window; 2048 and 4096 failing to load there | the server log's first three lines (the script's own note); the batch-size study's package, `/batch-size/data/runs/2026-09-19_minimax-m2.7/one_ub4096_cmoe_ctx196608.server.log` and `one_ub2048_cmoe_ctx196608.server.log` | cited, not shipped here |
| Thinking left on | the `.result`'s JSON rows | `reasoning_chars` 395, 438 and 480 on the reads; no thinking switch in the requests (`minimax-m27-fullwindow/launch-notes.txt`) |

## Added 2026-09-26 (push 7d): GLM-5.3-Flash at 262,144

| Figure on the page | File | Field |
|---|---|---|
| GLM-5.3-Flash, 4096 / 2048: 230,039 tokens, 1,085.7 s, 211.9 tokens/s, 31,305 MiB at load, 3 of 3 codes | `glm53-fullwindow/glm53_256k_ub2048.result` and `.server.log` | header `VRAM at load 31305 MiB`, `served: n_ctx_slot = 262144`; the JSON row with `depth_target` 230000 and `kind` read: `prompt_n`, `prefill_tps`, `codes3_hits` 3; server log `prompt eval time = 1085743.67 ms / 230039 tokens`, / 1000 and rounded |
| Earlier reads at 2048: 20,065, 47,992 and 100,003 tokens at 206.0, 250.3 and 243.1 tokens a second | the same files | JSON rows `depth_target` 20000, 48000, 100000 (`prompt_n`, `prefill_tps`); server log `97381.06 ms / 20065 tokens`, `191738.81 ms / 47992 tokens`, `411406.12 ms / 100003 tokens`; tokens / seconds, rounded to one decimal |
| Peak 31,951 MiB, 656 MiB below the card's 32,607 MiB total | the same `.result`; the total as in the Qwen3.5-122B-A10B row above | `peak VRAM during probes: 31951 MiB`; arithmetic: 32,607 - 31,951 = 656 |
| GLM-5.3-Flash, 4096 / 1024: 47,992 tokens, 336.4 s, 142.6 tokens/s, 28,834 MiB at load; peak 29,261; 20,065 tokens at 135.1; nothing longer read | `glm53-fullwindow/glm53_256k_ub1024.result` and `.server.log` | header `depths=20000,48000`, `VRAM at load 28834 MiB`, `peak VRAM during probes: 29261 MiB`; server log `336435.32 ms / 47992 tokens`, `148522.88 ms / 20065 tokens` |
| 1024 now its script's setting at this window; 4096 its setting at 131,072 | `glm53-fullwindow/launch-notes.txt` | the installed start script's code: micro-batch default 1024 above 131,072, else 4096; batch default 4096 |
| At 131,072 with 4096: 47,992 tokens at 404.9 tokens a second | `glm53-fullwindow/glm53_128k_default.result` and `.server.log` | `served: n_ctx_slot = 131072`, `--batch-size 4096 --ubatch-size 4096`; JSON row `depth_target` 48000; server log `118517.90 ms / 47992 tokens` |
| 1,048,576, the context length its model file declares | `glm53-fullwindow/launch-notes.txt` | GGUF key `glm5next.context_length` = 1048576 in the first shard, read for this package (the model file is not shipped) |
| First 42 layers' experts in RAM, default 16-bit cache, reasoning effort none (the row labels) | each `.result`, the `cmdline:` line | `--n-cpu-moe 42` (llama.cpp's own description: the MoE weights of the first N layers kept on the CPU); no `--cache-type`; `--chat-template-kwargs {"reasoning_effort":"none"}`; no `--spec-type` |
| A few hundred characters of reasoning in each read's reply | the `.result` JSON rows | `reasoning_chars` 362 to 617 on the reads at 262,144 |
| Section 05: 4096 / 4096 at 262,144 failed to load, a refused 13,281.37 MiB compute buffer | `glm53-fullwindow/glm53_256k_ub4096.result` (env lines, `DIED on load`) and `.server.log` | `allocating 13281.37 MiB on device 0: cudaMalloc failed: out of memory`; `failed to allocate compute pp buffers` |
| Section 04 intro: the later reads run from 47,992 to 230,039 tokens | the section 04 table's Tokens read column | its smallest and largest values |
