256K on a 32 GB card: the cost of reading
Ten configurations retrieved every planted code. Reading took from 60.8 seconds to 6,580.7 seconds (109.7 minutes), with the longest run split across two machines; most configurations kept expert weights in system RAM, and sampled card use ranged from 8,473 to 31,068 MiB.
On September 15, nine models in ten configurations returned all three codes from synthetic ledgers after reading 150,475 to 231,065 tokens. Nine configurations used the desktop alone; one added the laptop. Peak card use ranged from 8,473 to 31,068 MiB. A large served window was usable here, but the waiting time depended heavily on the model and placement.
Measured: one retrieval request per configuration, the files and settings below. Bench specification: NVIDIA RTX 5090 (32 GB) and 188 GiB of desktop RAM; the split also used an ASUS ROG Flow Z13 with 128 GB unified memory. These are hardware descriptions, not measured capacity tests.
Independent measurements. Not affiliated with, endorsed by, or connected to NVIDIA, ASUS, DeepSeek, MiniMax, ggml-org, or the makers of Gemma, GLM, Qwen, Mistral, Ling, Inkling, Ornith or Laguna.
-b 8192 -ub 8192, against 2,062.2 seconds earlier. That is 5.2 versus 34.4 minutes (arithmetic: seconds divided by 60). The later build is not identified in its saved 256K records, so this is a dated before-and-after comparison, not a one-variable experiment.What it is
A served window is the space the server makes available for a request and its answer. It does not say how long reading will take, or how much of that space a particular request fills. Here, 256K means a served window of 262,144 tokens. GLM-4.7-Flash was served at 202,752, the window this sweep requested. The launch notes call that its ceiling. This package has no failed attempt above 202,752.
The test generated a synthetic ledger with three fixed sealed codes near 5, 50 and 95 percent of its entries. Those are approximate depths in the text, not exact token offsets. After reading the ledger, each model was asked to quote the three codes in order. All ten answers contained the exact strings in the requested order.
This measures allocation, reading cost and a narrow retrieval task. The prompts did not fill every token of the served windows. In particular, the desktop DeepSeek Q8 prompt was shorter than the other 262,144-window prompts.
What we ran, exactly
Measured, September 15, 2026. Each configuration ran by itself with --parallel 1 --threads 24 --threads-batch 24. Requests used temperature: 0 and did not stream. The probe checked the served window through /props, warmed the model, ran two short-prompt prose requests, then sent one deep retrieval request. The table below uses that last request only.
The launch lines contain no batch or micro-batch override. “Default batch” here means those flags were absent; the saved logs do not print their effective numeric values. Flash attention was requested in every launch. The Qwen3.5 pair used draft-mtp with maximum draft length 6 and a 0.75 gate. All other main-table launches omitted a draft option. Requests disabled thinking where the launcher supplied that option; Inkling used reasoning_effort: none.
| Configuration | Placement and key settings | Commit |
|---|---|---|
| Gemma-4-26B-A4B | All model layers on card | d3146f2b5 |
| GLM-4.7-Flash | All model layers on card | d3146f2b5 |
| Qwen3.5-122B-A10B | 38 expert layers in RAM; MTP on; no mmap | c8e03ce |
| Mistral Small 4 | 26 expert layers in RAM; no mmap | 5f55650 |
| Ling-3.0-flash | All experts in RAM; fit off; no draft requested | d3146f2b5 |
| Qwen3.5-397B-A17B | 57 expert layers in RAM; MTP on; no mmap | c8e03ce |
| Qwen3.8-Flash-Next | 99 expert layers requested on CPU; fit off | d3146f2b5 |
| Inkling-Small | 42 expert layers in RAM; f16 key and value cache | 946fc11d1 |
| DeepSeek V4 Flash, desktop | All experts in RAM; no repack | 5f55650 |
| DeepSeek V4 Flash, desktop + laptop | RPC split 18,25; expert tensors of layers 8 to 17 in desktop RAM; no repack | d3146f2b5 |
mmap is memory-mapping the file, fit is the runtime’s automatic placement, repack is a weight rewrite, RPC is the link to the second machine, MTP is the model’s built-in draft head, and 0.75 is the minimum draft probability.
The configuration record lists the exact model filenames and complete redacted commands. Quantizations are beside each measured row. Build commits were extracted from the response metadata; the build table preserves that mapping. Different builds and placements make this a deployment comparison, not a controlled ranking of model architectures.
Observed: the September 15 reads
Measured, one deep request per configuration. “Read s” is the server’s prompt-processing time, converted from milliseconds to seconds (arithmetic), excluding model loading and answer generation. Rates are the server’s recorded tokens per second. “Short answer” is the code-only output, 25 to 33 tokens, not prose speed.
| Model / quantization | Served window | Tokens read | Read s | Read tokens/s | Peak card MiB | Answer tokens | Short answer tokens/s | Codes |
|---|---|---|---|---|---|---|---|---|
Gemma-4-26B-A4BUD-Q4_K_XL | 262,144 | 230,855 | 60.8 | 3,799.69 | 23,513 | 33 | 121.36 | 3/3 |
GLM-4.7-FlashUD-Q4_K_XL | 202,752 | 189,931 | 231.7 | 819.70 | 28,765 | 29 | 70.48 | 3/3 |
Qwen3.5-122B-A10BUD-Q4_K_S | 262,144 | 230,803 | 531.8 | 433.99 | 30,794 | 32 | 37.10 | 3/3 |
Mistral Small 4UD-Q4_K_M | 262,144 | 231,065 | 705.0 | 327.75 | 30,356 | 32 | 23.76 | 3/3 |
Ling-3.0-flashQ4_K_M | 262,144 | 230,759 | 1,118.8 | 206.25 | 8,473 | 32 | 26.11 | 3/3 |
Qwen3.5-397B-A17BUD-IQ3_XXS | 262,144 | 230,803 | 1,186.6 | 194.51 | 31,068 | 32 | 17.00 | 3/3 |
Qwen3.8-Flash-NextUD-Q3_K_XL | 262,144 | 230,803 | 1,392.9 | 165.70 | 15,626 | 32 | 13.58 | 3/3 |
Inkling-SmallUD-Q3_K_XL | 262,144 | 230,827 | 2,005.9 | 115.07 | 15,814 | 27 | 8.78 | 3/3 |
DeepSeek V4 Flash, desktopUD-Q8_K_XL | 262,144 | 150,475 | 2,062.2 | 72.97 | 18,306 | 25 | 9.14 | 3/3 |
DeepSeek V4 Flash, desktop + laptopUD-IQ3_XXS | 262,144 | 230,835 | 6,580.7 | 35.08 | 25,824 | 25 | 6.49 | 3/3 |
Peak card use is the largest sampled nvidia-smi memory-used reading while the deep request ran, sampled every 5 seconds and including other card use. It is not model-only allocation or the increase caused by context. The split row reports the desktop card, not the laptop’s unified memory.
Gemma’s 60.8-second read was the shortest. The two-machine DeepSeek read took 6,580.7 seconds, or 109.7 minutes (arithmetic: milliseconds divided by 60,000). It used a different quantization and read more tokens than desktop DeepSeek Q8, so those two rows do not isolate the cost of splitting a model.
Load times during this sweep are upper bounds. The TSV header records large concurrent model downloads competing for disk and page cache. Its load_s values are separate from the reading column. The runs were not performed on an otherwise idle storage system.
The TSV’s decode_tps column is also a different measurement: the best of two earlier, short-prompt prose replies. The at-depth rates above come from deep_decode_tps and the deep response JSON. Small amounts of prefix reuse appear in some deep responses; prompt_n is the server’s processed-token count, and cache_n remains in the files.
The later batch settings change the wait
A micro-batch sets how much of the prompt the runtime processes in one chunk. The later records below are long reads, from 47,992 to 230,039 tokens, with the window still set to 262,144, except MiniMax M2.7’s: it was served at 196,608, the context length its model file declares. They are separate measurements, not estimates applied to the earlier rows.
| Run | Batch / micro-batch | Tokens read | Recorded time | Read tokens/s | Card MiB |
|---|---|---|---|---|---|
| DeepSeek V4 Flash UD-Q8_K_XL, desktop September 20 (collection date) | 8192 / 8192 | 150,324 | 313,172.17 ms prompt processing | 480.00 | 26,023 at load |
| Laguna S 2.1 September 21 (collection date) | not recorded / 4096 | 200,098 | 233.9 s request wall time | 866.1 prefill, not wall ÷ tokens | 27,995 series peak |
| Qwen3.8-Flash-Next UD-Q3_K_XL September 26 40 expert layers on CPU, 8-bit cache | 4096 / 2048 | 229,981 | 429.6 s prompt processing | 535.4 | 26,795 at load |
| Qwen3.8-Flash-Next same file, placement, cache and day | 2048 / 512 (llama.cpp’s default) | 229,981 | 1,119.3 s prompt processing | 205.5 | 21,333 at load |
| Ornith-1.5-35B Q6_K September 26 6 expert layers in RAM, default 16-bit cache, no draft head | 2048 / 2048 | 229,090 | 71.7 s prompt processing | 3,193.3 | 31,270 at load |
| Ornith-1.5-35B same file, placement, cache and day | 2048 / 512 (llama.cpp’s default) | 229,090 | 151.8 s prompt processing | 1,508.7 | 30,250 at load |
| Qwen3.5-122B-A10B UD-Q4_K_S September 26 38 expert layers in RAM, default 16-bit cache, MTP on, no mmap: the September 15 launch | 2048 / 512 (llama.cpp’s default) | 229,090 | 510.6 s prompt processing | 448.7 | 30,898 at load |
| Qwen3.5-122B-A10B same file, placement, cache and day | 2048 / 1024 | 229,090 | 310.0 s prompt processing | 738.9 | 31,633 at load |
| MiniMax M2.7 UD-IQ4_XS September 26, at 196,608 every expert layer in RAM, 8-bit cache | 4096 / 1024 | 189,491 | 1,215.4 s prompt processing | 155.9 | 31,709 at load |
| GLM-5.3-Flash UD-Q3_K_XL September 26 first 42 layers’ experts in RAM, default 16-bit cache, reasoning effort none | 4096 / 2048 | 230,039 | 1,085.7 s prompt processing | 211.9 | 31,305 at load |
| GLM-5.3-Flash same file, placement, cache and day; read to 47,992 tokens only | 4096 / 1024 (now its script’s setting at this window) | 47,992 | 336.4 s prompt processing | 142.6 | 28,834 at load |
DeepSeek’s later log names the same Q8 model file, one slot and the same served window. Its result records all experts in RAM and both batch settings at 8192. The prompt lengths differ, and the later record lacks a build identifier and full launch command. It records a completed answer, but does not retain the three-code answer needed to repeat the earlier retrieval score. No 3/3 result is claimed for this run.
The two Flash-Next rows are one measured comparison, run one after the other on 26 September: the same file, window,
placement and 229,981-token prompt, and both returned 3 of 3 planted codes. That placement is not the one in the 15 September
table above: 40 expert layers requested on CPU instead of 99, fit off, and an 8-bit (q8_0) cache. So 205.5 tokens a
second at the default batch is not a new measurement of the 165.70 row; the placement and the cache changed too. Both long reads
started from the same state, with about 60 GB of the 90 GB model file already in system memory. The first server began with none
of the file in memory, and its first request, 3,035 tokens, read at 87.7 tokens a second while the file's share in memory grew
from 43.5 to 57.7 GB; the next request on that server read 3,019 tokens at 517.6. It is the only start recorded with the file out
of memory, so it shows what this start did, not a rule for other models. The batch-size study has the rest of
this run.
The Ornith-1.5-35B and Qwen3.5-122B-A10B rows are two more runs on September 26 at 262,144, each through a copy of the model’s start script, with thinking switched off per request. At each setting the 229,090-token document, a fresh ledger with three planted codes, followed shorter reads on the same server, and all four answers returned all three codes. Qwen3.5-122B’s 2048 / 512 row keeps the launch of its September 15 row: the same file, build, 38 expert layers in RAM, no mmap and draft-head settings. That launch set no batch values, which at this build means 2048 / 512. It read 229,090 tokens in 510.6 seconds at 448.7 tokens a second, against 230,803 in 531.8 seconds at 433.99 then, with a different probe and document. At 2048 / 1024 it read the same document in 310.0 seconds, but the card peaked at 31,785 MiB over that series, 822 MiB below the card’s total of 32,607 MiB as reported that day (published in the Qwen3.8-27B at 256K package), and its start script keeps 512 at this window. Ornith-1.5-35B has no September 15 row; its failure at this window is in section 05. These rows loaded an Ornith-1.5-35B Q6_K file of the same name as the one in the failed launch, with 6 expert layers in RAM and without the draft head; section 05 lists the other differences. The 2048 / 512 row ran at the start-script copy’s own setting for this window at the time; 2048 / 2048 was passed in for the other row and is what the script now sets at this window.
MiniMax M2.7 has no September 15 row. Its model file declares a 196,608-token context, and on September 26 it ran at
that window through its start script, which there always uses the one shape recorded loading: every expert layer in RAM,
-b 4096 -ub 1024 and an 8-bit cache (the batch-size study records 2048 and 4096 failing to
load at this window). Thinking was left on. After reads of 20,039 and 47,933 tokens on the same server, at 167.9 and 181.2
tokens a second, it read a 189,491-token ledger with three planted codes in 1,215.4 seconds, 155.9 tokens a second, and
returned all three. The card read 31,709 MiB at load and peaked at 31,907 MiB over the series.
GLM-5.3-Flash has no September 15 row either. Its model file declares a 1,048,576-token context; on September 26 it ran
at 262,144 through a copy of its start script, with --n-cpu-moe 42 (the expert weights of its first 42 layers in
RAM), the default 16-bit cache and its reasoning effort set to none, though each read’s reply still carried a few hundred
characters of reasoning. At -b 4096 -ub 4096, its script’s setting at 131,072, it did not load at this window
(section 05). At 2048 it loaded at 31,305 MiB and read 20,065, 47,992 and 100,003 tokens at 206.0, 250.3 and 243.1 tokens a
second, then a 230,039-token ledger with three planted codes in 1,085.7 seconds, 211.9 tokens a second, and returned all
three; the card peaked at 31,951 MiB, 656 MiB below its 32,607 MiB total. At 1024, now its script’s setting at this window,
it loaded at 28,834 MiB and peaked at 29,261 over reads of 20,065 and 47,992 tokens at 135.1 and 142.6 tokens a second;
nothing longer was read at 1024. At 131,072, with 4096, the same probe read 47,992 tokens at 404.9 tokens a second.
Laguna’s record reports the planted string found. Its test used one string and a different prompt harness. The 200,098-token request followed a shorter request in the same series; the peak covers that series. The saved result does not identify the quantization, build, complete placement or sampling settings. It proves this recorded run, not an exact reproducible profile or a before-and-after gain.
The batch-size study covers the broader tuning sweep. There is no new reading time here for a September 15 configuration that was not re-measured at comparable depth.
Where memory stopped the run
Measured failures. Ornith-1.5-35B Q6_K, with 6 expert layers in RAM and MTP enabled, reached draft-context creation at 262,144 and failed on a 512.00 MiB allocation. Ornith-1.5-397B IQ3_XXS-UD, with 57 expert layers in RAM and MTP enabled, failed on a 1,862.00 MiB compute-buffer allocation. The latter log does not establish that the draft context caused the failure. Neither failure log proves a successful draft-off run at that window.
Update, September 26. Runs without the draft head served and read this window with an Ornith-1.5-35B Q6_K file of the same name, the same build, 6 expert layers in RAM and the default cache: 229,090 tokens in 151.8 seconds of prompt processing at 2048 / 512 and in 71.7 seconds at 2048 / 2048, all three codes returned both times (section 04). Against the failure log, these runs omitted the MTP draft head. They also wrote the batch out (2048 / 512, or 2048 / 2048), used a 7,200-second timeout and set no CORS flag; the failed launch set no batch, a 3,600-second timeout and a localhost CORS flag, and its log’s preflight card reading is 1,113 MiB, where the September 26 results record no reading before launch. The filename matches; the directories do not, and neither log records a hash or a file size. This shows that filename serving the window without the draft head. It does not show the failure’s copy serving the window with the head: MTP was not used, and the failure stands as the record of the launch that requested it.
A draft head does not always have to go. Both Qwen3.5 configurations in the main table created their MTP context and used it in the code answer at 262,144. The logs and response draft counters establish that directly.
The GLM-4.7 Full configuration note records a narrower trade: at a 131,072 window, its MTP setup fitted with all 93 expert layers in CPU RAM and failed with 92. This is a recorded setup note, without a raw response or failure log in this package. It is not a 256K retrieval result.
Measured allocation failures, September 2026. MiniMax M3’s sparse-attention test used Q2_K_L, a 262,144 window, q8_0 key/value cache and flash attention on. Each tested header also placed experts on CPU (cmoe in the test label). At micro-batches 2048, 1024 and 512, the log records refused CUDA buffers of 20,572,169,216, 10,286,650,368 and 5,143,890,944 bytes respectively. All three tested settings failed before a retrieval result. These failures do not establish that every possible configuration must fail.
Measured allocation failure, September 26. GLM-5.3-Flash UD-Q3_K_XL at 262,144, with --n-cpu-moe 42,
the default 16-bit cache and -b 4096 -ub 4096, its start script’s setting at 131,072, failed before a retrieval
result: the log records a refused compute-buffer allocation of 13,281.37 MiB. The same launch loaded at 2048 and 1024
(section 04).
What this does not show
- Reasoning across a long document. This is retrieval of three planted strings in one run per configuration. It does not demonstrate reasoning across 230K tokens, handling contradictory evidence, or reliable work over an arbitrary document.
- A full-window generation test. The server exposed the requested capacity; the ledger occupied only part of it. The answers were short, fixed strings. Their rates do not predict long prose or code generation at that depth.
- A reliability estimate. There were no repeated deep trials, confidence intervals or alternate sets of codes. The same three codes appeared in every main-table ledger.
- Total machine memory or a universal card requirement. The table reports sampled card use for these files and placements. Most configurations also used system RAM. It does not measure their total RAM footprint or the minimum card capacity that could serve them.
- An isolated batch effect. The later DeepSeek timing is materially faster, but the saved records do not establish identical builds, prompts, full launch settings or background load. The September 15 downloads also limit interpretation of load times.
What to take from these runs
For a comparable setup, use the model filename, placement and cache settings together with the requested window. Check the served window after loading. Then measure a fresh prompt at the depth you actually need: a successful load and a quick short prompt leave most of the cost untested.
Record batch settings explicitly. The later DeepSeek result shows why a default-batch reading time should not become a permanent expectation for that file. Memory also needs another check after changing the micro-batch or enabling a draft head, as the failed allocations show.
Related pages on this site
The batch-size study follows the prompt-reading change. Slower by default measures a draft-confidence setting. The Qwen3.5 pair field card covers the two Qwen3.5 files, and Method explains the site’s evidence labels.
Sources and artifacts
The data package contains the recorded TSV, response JSON, server logs, probe and launch scripts, allocation failures and later long-window results. The number map traces every figure to a field or a labelled calculation. If a number on this page disagrees with a file in its package, the file is right and the page is wrong.
Paths, addresses, ports, serving aliases and request identifiers have been redacted. Local model answers and timing fields remain. The scripts are archival records with placeholders. Build revisions for the September 15 responses are preserved in a separate derived table. No vendor performance claim is used as a measured result.