Graphometer
Measured study · 19 to 21 September 2026, one model re-measured 26 September · one RTX 5090 desktop (32 GB card, 188 GiB RAM) · llama.cpp at seven commits, one of them built twice · 12 models with experts in system RAM, 6 that fit on the card

llama.cpp's batch size

Prompt reading 2.5 to 7.3 times faster than the default, up to 20 against an inherited -ub 128.

llama.cpp reads a prompt in micro-batches of 512 tokens unless you tell it otherwise, and on this desktop almost nothing told it otherwise: before 20 September, 20 of our 25 start scripts left both batch settings at llama.cpp's defaults, and three of the five that did set them used a micro-batch smaller than the default. We swept --ubatch-size, with --batch-size kept at least as large, across the models we serve. On the twelve whose expert weights sit in system RAM, the first read of a long prompt, mostly 48,000 tokens, became 2.5 to 7.3 times faster than at llama.cpp's default; where a script had inherited -ub 128, reading became 19.9 times faster on a 24,071-token prompt and 17.5 times on a 3,658-token one. Qwen3.8-Flash-Next, re-measured at its full 262,144-token window on 26 September with about 58 to 60 GB of its 90 GB file in memory, kept the gain there (48,075 tokens at 660.6 t/s against 240.2 at the default). On a server started with none of the file in memory, the first request (3,035 tokens) began with 43.5 GB loaded and read at 87.7 t/s, against 517.6 for the next 3,019 tokens (section 04). Inkling-Small was timed on long prompts only at half the window it serves. With a large enough token budget the planted code came back at every rung we checked, and decode on short answers moved less than 7 percent for every shipped value but one (same day, same window). Five of the six models that fit on the card gained 3 to 34.5 percent, the sixth got slower, and all six kept the default. Every model had its own ceiling, set by card memory and by the window it serves, and one setting that read 150,153 tokens correctly and passed 110 other prompts crashed the server on a single 3,007-token input, so we withdrew it. This is one machine, the llama.cpp builds of that week, one request at a time, and reading speed, not the quality of the answers.

Independent measurements; no affiliation with the llama.cpp project, ggml-org, NVIDIA, ASUS, Intel, Alibaba (Qwen), DeepSeek, inclusionAI, Thinking Machines Lab, Poolside, Mistral AI, Z.ai, MiniMax, Google, Meta, Unsloth, or the makers of Ornith.
01

What it is

Before a model writes the first word of a reply, llama.cpp reads the whole prompt, in batches whose size two settings decide. --batch-size (-b) is the logical batch: the most prompt tokens the server hands the model in one step. --ubatch-size (-ub) is the physical micro-batch: the most tokens that go through the model's layers in one pass. At the build most of our models ran, the defaults are 2048 and 512 (common/common.h lines 451 and 452 at commit d3146f2b5; --help calls them the “logical maximum batch size” and the “physical maximum batch size”). A 48,000-token prompt is 94 micro-batches at the default and 12 at 4096, counting the partial last batch (48,000 / 512 = 93.75; arithmetic).

Why it matters when the experts sit in system RAM. A mixture-of-experts model keeps most of its weights in “experts”, of which each token uses only a few. When the file is larger than the graphics card, the experts usually stay in system RAM and the rest goes on the card. During prompt reading, when an operation's weights live in host memory and the batch is large enough, llama.cpp's scheduler hands the operation to the graphics card, which means carrying those weights across for that pass (ggml/src/ggml-backend.cpp lines 966 to 975; the CUDA backend takes any operation with a batch of at least 32 tokens, ggml-cuda.cu lines 5536 to 5539 and 5710, same build). That trip costs about the same for a micro-batch of 512 tokens as for one of 8,192, so a bigger micro-batch pays it fewer times per prompt. This is our reading of the code; we did not instrument the transfer itself.

What it does not touch. Writing the reply (decode) goes one token at a time, so the micro-batch does not enter into it. Our decode figures moved little where before and after share a day and a window, with one exception explained in section 06.

What it costs. Card memory: the card holds a working buffer sized for one micro-batch. On MiniMax M3 split across two machines (20 September), the server asked the card for 20,746,340,352 bytes (19,785.25 MiB) at -ub 4096 whether 36 or 50 of the model's 60 layers stayed on the desktop, and for exactly half, 10,373,736,448 bytes, at -ub 2048; all three requests were refused. The buffer also grows with the window (section 04).

02

What we ran, exactly

Every figure on this page is measured on this machine on the date given, unless it is marked arithmetic.

MachineOne desktop: one NVIDIA GeForce RTX 5090 (32 GB), 188 GiB of system RAM. Every server ran with --threads 24 --threads-batch 24, one request slot (--parallel 1) and flash attention on or auto, one model at a time.
The readA synthetic ledger of numbered, near-identical lines (the generator is in the package), seed 20260920, sized with the server's own tokenizer to within 2 percent of 48,000 tokens (24,000 on Laguna S 2.1's 32,768-token window), with one sealed code at half depth and the question “quote the sealed reference exactly. Answer with the code only.” Temperature 0, max_tokens 900, a freshly loaded server (usually as its first request, with the page cache not recorded; section 04), nothing in the prompt cache. Reading speed is the server's own timings.prompt_per_second; decode is its predicted_per_second.
Pass ruleFrom 21 September, a non-empty answer had to contain the code itself; an empty answer passed for reading speed if the code was in the model's reasoning (section 08). The 20 September harness recorded only whether the code appeared in the answer or the reasoning, and the answer's length: 17 characters, the length of the code, in every passing row.
19 SeptemberMiniMax M2.7 with its own harness: one 3,658-token prompt at -ub 128, 512, 1024, 2048 and 4096 with every expert layer in RAM (131,072 window, q8_0 cache); then 43,909-token reads at 4096 with 62, 61, 60, 59 and 58 of its 62 expert layers in RAM; then its full 196,608 window. That prompt was a repeated paragraph, and this harness did not check the answer.
20 SeptemberFlags copied by hand from each start script, at a 131,072 window: Ling-3.0-flash, Inkling-Small and Qwen3.8-Flash-Next (the default and a candidate; 8192 as well on Inkling and Flash-Next), and DeepSeek V4 Flash with its own harness (16,011-token reads for the first comparisons, then 48,073). The three winners were then loaded at the served 262,144 window and sent one 3,000-token prompt each (sections 04 and 05), and DeepSeek's read 150,324 tokens there. This pass could not start Laguna S 2.1: the hand-copied --n-gpu-layers 999 tried to put all 64,771.15 MiB of its weights on the card.
21 SeptemberThe remaining models, and Ling again, through their own start scripts at each script's served window. The scripts take the batch setting from the environment; a script that had none was first given one whose default is llama.cpp's own, so its baseline did not change, and nothing else was edited. Rungs: the script's own setting, then 2048, 4096 and 8192 (1024, 2048 and 4096 on the models that fit on the card and on Ornith-1.5-35B), stopping at the first rung that failed to load, lost the code, or read slower than the one before. Card memory was sampled every 2 to 5 seconds during each read.
Stress check21 September, the shipped settings of the other eleven models with experts in RAM (twelve runs: DeepSeek V4 Flash's Q8 and IQ3 files separately). Each: a fresh load whose first request was the exact 3,007-token prompt that crashed Ling (section 05), then 40 prompts of real mixed text from llama.cpp's own source tree (C, C++, CUDA, Python, Markdown), sized mostly around one and two micro-batches, max_tokens 32, with a restart and a saved prompt after any crash. Ling had its own check; the six models on the card were not stress-checked.
26 SeptemberQwen3.8-Flash-Next at the 262,144 window it serves, through a copy of its start script (--n-cpu-moe 40, --fit off and the q8_0 cache, as on 20 September): the shipped -b 4096 -ub 2048 on a server started with none of the 90 GB model file in the page cache, then llama.cpp's default on a server started with 60.05 GB of it already there. Four fresh documents per load, of about 3,000, 3,000, 48,000 and 230,000 tokens (one code in each, three in the last), thinking off, with the file's share of the page cache read before and after each request (mincore).
Card memoryRaw nvidia-smi MiB, desktop included. The 20 September runs recorded it once, after load; the 21 September runs recorded the peak during the read, which ran 0 to 465 MiB above the load reading; the 26 September run recorded it at load and after each read.
Scripts beforeOf the 25 start scripts on the machine before 20 September, 20 left both settings at llama.cpp's defaults. Five set them: one at -b 64 -ub 32, two at -ub 128, and two above the default, one of those on only two of its three window settings. (Our first-pass harness's header says four; the recount is in the package.)

The builds

Each model ran on the llama.cpp binary its start script names. The commit comes from each build's own build-info record, and every binary was built before the sweep began.

CommitOriginModels in this study
d3146f2b5ggml-org mainline, build 10919 (a second build of it with RPC enabled)Ling-3.0-flash, Qwen3.8-Flash-Next, GLM-4.7-Flash, Gemma 4 26B-A4B; with RPC: MiniMax M2.7, MiniMax M3, DeepSeek V4 Flash IQ3
946fc11d1llama.cpp pull request #25731, a draft adding the Inkling architectureInkling-Small
04b2b72Poolside's llama.cpp forkLaguna S 2.1
c8e03ceggml-org mainlineQwen3.5-397B-A17B, Qwen3.5-122B-A10B, Qwen3-235B-A22B-Instruct-2507, Qwen3.8-27B, Qwen3.6-27B, Ornith-1.5-35B
5f55650ggml-org mainlineDeepSeek V4 Flash Q8, Mistral Small 4
d94f44eUnsloth's llama.cpp fork, glm5next branchGLM-5.3-Flash
3cb7ffbggml-orgGemma 4 31B, Muse Glimmer 30B
03

Observed: reading before and after

Twelve models keep some or all of their expert weights in system RAM. Eleven of their files are larger than the card (68.2 to 161.9 GB); Ornith-1.5-35B's 29.2 GB file keeps 6 of its 41 layers' experts in RAM. “Was” is the setting before the sweep; “now” is the shipped setting, read at the window given under the table, half the served window on three models (note e). Decode figures are short answers (tokens generated on the second line): they show whether decode moved, not how fast these models speak.

Model · fileRead, tokensWas, t/sNow, t/sShippedCard, MiBDecode, t/sDate
DeepSeek V4 Flash
UD-Q8_K_XL, 161.9 GB e
48,07383.3607.1 (7.3×)-b 8192 -ub 819217,063 → 22,089
after load
9.91 → 9.58
74, 70 tokens
20 Sep
GLM-5.3-Flash
UD-Q3_K_XL, 147.5 GB b
24,008 → 20,33380.7446.0 (5.5×)-b 4096 -ub 409624,352 → 29,872
peak during read
9.22 → 9.35
62, 60 tokens
21 Sep
Mistral Small 4 119B
UD-Q4_K_M, 73.8 GB
48,697481.32,192.3 (4.6×)-b 8192 -ub 819226,882 → 29,756
peak during read
21.01 → 20.97
11, 11 tokens
21 Sep
Qwen3.5-397B-A17B
UD-IQ3_XXS, 149.8 GB
48,029224.5932.2 (4.2×)-b 4096 -ub 409626,195 → 29,360
peak during read
16.39 → 16.41
11, 11 tokens
21 Sep
Qwen3.5-122B-A10B
UD-Q4_K_S, 73.4 GB
48,027520.22,101.7 (4.0×)-b 4096 -ub 409626,911 → 29,619
peak during read
30.81 → 32.72
648, 716 tokens
21 Sep
Ornith-1.5-35B
Q6_K, 29.2 GB
48,0271,860.26,283.3 (3.4×)-b 4096 -ub 409627,304 → 28,691
peak during read
73.57 → 73.20
87, 83 tokens
21 Sep
Ling-3.0-flash
Q4_K_M, 77.8 GB a
47,992 → 47,986238.2727.7 (3.1×)-b 4096 -ub 20487,233 after load;
8,824 peak
27.46 → 23.25
77, 77 tokens
20 & 21 Sep
Ling-3.0-flash, one
day and window a
3,007201.5444.0 (2.2×)-b 4096 -ub 20488,037 → 8,871
peak during read
24.41 → 23.0
104, 64 tokens
21 Sep
Inkling-Small
UD-Q3_K_XL, 119.6 GB e
48,115123.0337.9 (2.7×)-b 4096 -ub 204812,641 → 13,183
after load
12.23 → 11.77
97, 97 tokens
20 Sep
Qwen3.8-Flash-Next
UD-Q3_K_XL, 90.0 GB e
48,069203.2511.7 (2.5×)-b 4096 -ub 204819,444 → 21,394
after load
18.14 → 18.75
99, 104 tokens
20 Sep
Qwen3-235B-A22B-Instruct-2507
Q4_K_M, 142.2 GB
48,020285.8703.3 (2.5×)-b 4096 -ub 204828,317 → 29,421
peak during read
8.27 → 8.18
11, 11 tokens
21 Sep
Laguna S 2.1
Q4_K_M, 68.2 GB c
24,07172.2 at -ub 1281,434.0 (19.9×)-b 8192 -ub 819227,973 → 28,272
peak during read
18.40 → 16.50
12, 12 tokens
21 Sep
MiniMax M2.7
UD-IQ4_XS, 108.4 GB d
3,65832.3 at -ub 128564.0 (17.5×)-b 4096 -ub 409622,530 → 24,359
after load
9.9 → 9.3
32, 32 tokens
19 Sep

Windows: 131,072 tokens, except Ling's 21 September figures (262,144) and Laguna (32,768). Card memory is at load or the peak during the read, as marked; the two are not comparable. The code came back at every rung shown except M2.7's, whose harness did not check (note d). a Ling's first line spans two passes and two windows, 238.2 on 20 September (hand-copied flags, 131,072) and 727.7 on 21 September (its own script, 262,144), each with its own card and decode figures; the second line is one pass at one window. -ub 4096 read 1,140.4 on the 20th and was withdrawn (section 05). b GLM-5.3-Flash's default read was a load's first request; the tuned 20,333-token read was the third on its load, and the fourth, 48,168 tokens, read at 425.1; the card peak covers all four. c Laguna's “was” is its inherited -b 512 -ub 128, not llama.cpp's default (section 06). d M2.7's figures come from its own harness on 19 September, a repeated-paragraph prompt with no answer check; at llama.cpp's default 512 the same prompt read 100.8. e These “now” figures are 48,000-token reads at 131,072, half the 262,144-token window these scripts serve. At 262,144 DeepSeek read 150,324 tokens at 480.0 (below), and Qwen3.8-Flash-Next, re-measured on 26 September with about 58 to 60 GB of its 90 GB file in memory, read 48,075 tokens at 660.6 t/s against 240.2 at the default (section 04). Inkling-Small read only 3,000-token prompts there, at 263.6 and 187.9 t/s; its long reads at that window were not re-measured.

Against llama.cpp's own default, the gain on long reads was 2.5 to 7.3 times: from Qwen3-235B's 2.46 (285.8 to 703.3, with 4096 not shipped, section 13) to DeepSeek V4 Flash's 7.29, whose 48,000-token read went from 576.9 seconds to 79.2. Against an inherited -ub 128, it was 17.5 and 19.9 times (MiniMax M2.7 on a 3,658-token read, Laguna S 2.1 on 24,071 tokens).

At depth, on the same build

One comparison sits at a full window: DeepSeek V4 Flash, the same Q8 file and build (5f55650), every expert in RAM, 24 threads, a 262,144-token window. On 15 September at llama.cpp's default it read 150,475 tokens in 2,062.2 seconds (34.4 minutes, 72.97 t/s) and returned all three planted codes; on 20 September at -b 8192 -ub 8192 it read 150,324 tokens in 313.2 seconds (5.2 minutes, 480.0 t/s) and returned its one code: 6.6 times the prompt-reading rate, on two different synthetic ledgers. More than the batch differed: the earlier request had thinking off, the launches differed in listening address and default sampling flags (both requests set temperature 0), and the 15 September read was a later request on a server whose load took 220 seconds, the 20 September read the first request after a load of about 4 seconds, a difference of the kind that moved an untuned read from 72.3 to 83.3 t/s (section 12). The default's cost at a full window on other models has its own page (section 14).

Models that fit on the card

Six models fit entirely on the card, so no weights cross, and the gains were smaller: 3 to 34.5 percent on five, while the sixth read slower. All six kept llama.cpp's default, for reasons their start scripts record: Gemma 4 26B-A4B already read about 9,000 t/s and decoded slower as the batch grew (on 11-token answers, the kind of figure section 13 corrects); GLM-4.7-Flash and Qwen3.8-27B kept card room (1,371 MiB on Qwen3.8-27B's almost-full card); the rest gained 5.6 percent or less, or read slower.

Model · fileWindowDefault, t/sFastest read, t/sChangeDecode, t/s
Gemma 4 26B-A4B
UD-Q4_K_XL, 17.0 GB
262,1448,972.312,063.4 at 2048+34.5%163.68 → 153.82
11, 11 tokens
GLM-4.7-Flash
UD-Q4_K_XL, 17.5 GB
202,7522,594.63,100.2 at 2048,
answer empty (failed)
+19.5%,
reading only
129.60 → 130.09
429, 900 tokens
Qwen3.6-27B
UD-Q4_K_XL, 17.6 GB
131,0722,966.53,133.3 at 2048+5.6%63.15 → 63.14
182, 166 tokens
Gemma 4 31B
q4_0, 17.7 GB
131,0722,687.52,784.5 at 1024+3.6%56.45 → 55.92
11, 11 tokens
Qwen3.8-27B
UD-Q5_K_XL, 20.2 GB
131,0722,634.82,717.5 at 2048+3.1%99.26 → 94.41
106, 106 tokens
Muse Glimmer 30B
Q4_K_M, 16.8 GB
131,0722,737.7none faster−0.6% at 1024, −7.1% at 2048193.65 → 179.59 at 1024
and 182.69 at 2048;
139, 139, 138 tokens

21 September, 48,000-token reads through each start script. The code came back in the answer at every rung except GLM-4.7-Flash's 2048, which spent all 900 tokens reasoning and returned an empty answer, a failure as an answer though our harness counted it for reading speed (section 08); its 1024 and 4096 rungs answered, at 3,022.0 and 3,070.1. Qwen3.8-27B could not load -ub 4096 (section 04).

04

The ceiling is per model

No single value was safe for every model. The largest setting that loads and finishes a read depends on the model, its placement and its window, and the failure can come at load or in the middle of a read.

  • At load. -ub 8192 failed to load on Qwen3.8-Flash-Next at a 131,072 window on 20 September (a 16,171.13 MiB working buffer refused) and, on 21 September, on GLM-5.3-Flash (15,424.89 MiB), Qwen3.5-397B (3,712.25), Qwen3.5-122B (3,264.25) and Qwen3-235B (2,752.23). Qwen3.8-27B, which fits on the card, failed at -ub 4096 (1,568.13 MiB).
  • Mid-read. Inkling-Small at -ub 8192 loaded at 18,130 MiB, then 8,203 tokens into the read the server tried to allocate 29,242.27 MiB, was refused, and died. The client saw only a dropped connection; that was the out-of-memory signature here.
  • Where 8192 worked. DeepSeek V4 Flash Q8, Laguna S 2.1 at 32,768 tokens, Mistral Small 4, and Ornith-1.5-35B, which read 7,393.7 t/s at 8192 but peaked at 30,087 MiB, so we shipped 4096.

The window moves the ceiling. The same setting took more card at the 262,144-token window these models serve than at the 131,072 we swept: Ling at -ub 4096 8,107 → 10,285 MiB, Inkling-Small at 2048 13,183 → 17,281, DeepSeek at 8192 22,089 → 26,023, and Qwen3.8-Flash-Next at 2048 21,394 → 26,561, all after load, 2,178 to 5,167 MiB more. MiniMax M2.7 with every expert layer in RAM read 43,909 tokens at -ub 4096 with a 131,072 window, but at its full 196,608 window 4096 failed to load (a 2,916.16 MiB buffer refused), 2048 failed too, and only 1024 loaded, at 31,560 MiB; its script now forces 1024 there. Until 26 September, four other scripts (Qwen3.5-397B, Qwen3.5-122B, Mistral Small 4 and Ornith-1.5-35B) returned to llama.cpp's default above 131,072, and the bigger value had not been measured there; Laguna S 2.1 drops from 8192 to 4096 above 32,768.

Update, 26 September. Two of those four were then measured at 262,144, each reading a 229,090-token document whose three planted codes came back at every setting (the 256K page has the rows). Ornith-1.5-35B's prompt processing took 71.7 seconds at -b 2048 -ub 2048 against 151.8 at 2048 / 512. The 151.8-second run used the start-script copy's own setting above 131,072 at the time; 2048 / 2048 was passed in for the other run, and is what the script now sets there. Qwen3.5-122B's prompt processing took 310.0 seconds at -ub 1024 against 510.6 at 512, but the card peaked at 31,785 MiB in that series, 822 MiB below the card's total of 32,607 MiB as reported that day (published in the Qwen3.8-27B at 256K package), and its script keeps 512 there. The Qwen3.5-397B and Mistral Small 4 scripts still return to the default there, unmeasured.

Update, 26 September: MiniMax M2.7. Its installed script was then run at 196,608, where it sets the one shape that loaded: every expert layer in RAM, -b 4096 -ub 1024 and an 8-bit cache. It loaded at 31,709 MiB and peaked at 31,907. Its prompt processing read 20,039 tokens at 167.9 t/s, 47,933 at 181.2 and 189,491 at 155.9, the last in 1,215.4 seconds, and that deepest read, a document with three planted codes, returned all three, with thinking on. At 131,072 its shipped setting had read 43,909 tokens at 656.5 t/s, on another prompt (section 06). A 40-prompt stress check at 196,608, on prompts of 534 to 11,160 tokens, ran without a crash, but all 40 are recorded with an empty answer, every reply running to its token limit, so it checked that the server stayed up, not the answers.

Update, 26 September: GLM-5.3-Flash. Its start script now also offers a 262,144 window. Through a copy of the script there, the setting it keeps at 131,072, -b 4096 -ub 4096, did not load: a 13,281.37 MiB compute buffer was refused. -ub 2048 loaded and read 20,065 to 230,039 tokens at 206.0 to 250.3 t/s, the deepest returning all three planted codes, but peaked at 31,951 MiB, 656 below the card's total; -ub 1024 peaked at 29,261 and read 20,065 and 47,992 tokens at 135.1 and 142.6, and nothing longer. The script ships 1024 at 262,144, for 2,690 MiB more room at the peak and, its comment says, to stay off -ub 2048, the micro-batch that crashes in upstream report #28282 (section 05); our reads at 2048 ran without that error. At 131,072, where it keeps 4096, the same probe read 47,992 tokens at 404.9 t/s, 2.8 times (arithmetic) the shipped 262,144 setting's 142.6. A 40-prompt stress check at 262,144 with -ub 1024 ran without a crash, but 36 of the 40 are recorded with an empty answer, so it mostly checked that the server stayed up.

Short first reads at the served window. On 20 September Inkling-Small and Qwen3.8-Flash-Next each read one 3,000-token prompt at 262,144: Inkling 3,033 tokens at 263.6 t/s (187.9 when its shipped script replayed the crash prompt on 21 September), Flash-Next 3,041 at 45.4 (2,999 at 112.3 on the replay). Short first reads are slower in general: on the seven loads here that read both lengths, a 3,000-token first request ran at 49 to 78 percent of the same load's 48,000-token rate. Inkling's fit that; Flash-Next's, 9 and 22 percent of its 511.7, did not. Its 20 September read spent 57.1 of 67.0 seconds on the first 989 tokens and 9.7 on the next 2,048, which a partly filled micro-batch does not explain. Inkling-Small's long reads at 262,144 were not re-measured.

Flash-Next at its served window, 26 September. The shipped server started with none of the 90 GB model file in the page cache and 43.5 GB of it there after its load. Its first request, 3,035 tokens, read at 87.7 t/s while that grew to 57.7 GB, again with most of the time on the first 983 tokens (28.5 of 34.6 seconds); the next 3,019 tokens on the same server read at 517.6. The default's server started with 60.05 GB already in memory, and its first request read at 253.7, in line with its second at 255.7: a first request is not slow in itself, and this one was slow while the file paged in. The two slow reads above were also first requests after a start, though their page cache was not recorded. With about 58 to 60 GB of the file in memory (it never became fully resident), the shipped -b 4096 -ub 2048 against llama.cpp's default read about 3,000 tokens at 517.6 against 255.7 t/s and 48,075 tokens at 660.6 against 240.2, each read returning its one code, and 229,981 tokens in 429.6 seconds at 535.4 against 1,119.3 seconds at 205.5, with 3 of 3 codes on both: 2.02, 2.75 and 2.61 times, so the gain holds at the window the script serves. The card held 26,795 MiB at load and at most 27,271 after a read, against 21,333 and 21,694; decode was 12.3 to 22.6 t/s on both, on answers of 11 and 33 tokens.

Two files of one model need two values. The DeepSeek V4 Flash IQ3 file, with 7 of its 43 expert layers on the card instead of in RAM, ships -ub 4096 while the Q8 file ships 8192. At the 262,144 window the IQ3 file read 48,024 tokens at 426.5 t/s at 2048 and 690.9 at 4096; at 2048 the same load read 149,898 tokens at 403.8 t/s, and the shipped script at 4096 read 150,103 tokens at 632.1 (21 September, code returned each time).

05

One correct read does not make a setting safe

On 20 September, Ling-3.0-flash at -b 4096 -ub 4096 read 47,992 tokens at 1,140.4 t/s with the code correct, at a 131,072 window. Loaded at its served 262,144 window (10,285 MiB), the server died on its first request:

CUDA error: an illegal memory access was encountered

Our records then called that setting confirmed (section 13). On 21 September, through the model's own script at 262,144 and -ub 4096, five reads on one load (3,005, 2,991, 19,962, 48,084 and 150,153 tokens) were all correct, the three long ones at 1,067.6 to 1,173.9 t/s. Then a fresh load got the exact 3,007-token ledger from the 20 September check as its first request, and died with the same error. At -ub 2048, 1024 and 512 the same prompt passed with the code correct (444.0, 325.7 and 201.5 t/s).

By hand, with the model's exact launch flags, we narrowed it down:

  • Not the window. It crashed at -ub 4096 with windows of 262,144 (the control), 131,072 and 65,536, and at -ub 3072 with 262,144, the only window tried there.
  • Not the length alone. The same kind of ledger at nine other sizes, 1,760 to 4,001 tokens, passed at 4096, as did raw-token prompts of real code at exactly 3,006, 3,007 and 3,008 tokens and five other lengths.
  • Rare. At 4096, 60 prompts of real mixed text and 50 random ledgers of 2,134 to 6,541 tokens, 110 in all, passed without a crash, and so did the same 110 at 2048.

We shipped -b 4096 -ub 2048: 727.7 t/s on a 47,986-token read and 694.6 on 149,711, code correct, 8,824 MiB at peak, and the crash prompt passed through the shipped script. The shipped settings of the other eleven models with experts in RAM then got the same crash prompt first and 40 prompts of real text: twelve runs (DeepSeek's Q8 and IQ3 files separately), 40 of 40 each, no crash, the code returned on every replay. The six models on the card were not stress-checked.

We did not find the cause. The closest public report we found is llama.cpp issue #28282, opened on 2 September and still open when we checked on 26 September: an illegal memory access during a long prompt read on GLM-5.3-Flash at -ub 2048, on the same GPU family, with 512 or 128 as that report's workaround. It is a different model, and we have not shown a shared cause. Our GLM-5.3-Flash at -ub 4096, on its own build, read 3,019 to 48,168 tokens without it and passed its 40 stress prompts through its start script, which sets 4096 (the stress log does not print the micro-batch).

The package includes repro_ling_crash.py, which rebuilds the same request (same generator, same seed) against any running server. The crashes on this page come from the logs, not from that script.

06

Inherited values

Laguna S 2.1. Its start script, written in July, set --batch-size 512 --ubatch-size 128, a micro-batch a quarter of llama.cpp's default, with no reason recorded. Through its own script on 21 September, a 24,071-token read at its 32,768 window took 333.2 seconds at 72.2 t/s. At 2048 it read 698.4 t/s, at 4096 1,070.5, and at 8192 1,434.0, 16.8 seconds, code correct at every rung; the shipped script with no override read 24,013 tokens at 1,406.5. Decode on its 12-token answer fell from 18.40 to 16.50 t/s (15.32 in the shipped-script run), the one shipped value whose decode moved more than 7 percent. The script lets llama.cpp choose the placement (--fit on, 4,096 MiB kept free) and the card total barely moved (27,973 to 28,272 MiB at peak); we suspect the bigger buffer pushed some experts back to RAM and slowed decode, but the logs do not record the placement.

MiniMax M2.7. It was first tried with -cmoe -ub 128, the same pair of flags as the MiniMax M3 start script, where no reason for either is recorded. On 19 September the same 3,658-token prompt, every expert layer in RAM, read 32.3 t/s at 128, 100.8 at llama.cpp's default 512, 184.7 at 1024, 324.3 at 2048 and 564.0 at 4096. At 43,909 tokens, -ub 4096 evaluated the prompt at 640.2 t/s with every expert layer in RAM and 656.5 with 59 of 62, the shipped placement; the whole requests, each ending in a short answer, took 72.1 and 70.1 seconds. We did not run 128 at that depth; at its short-read rate, 43,909 tokens would take about 22.7 minutes (arithmetic). On 21 September the installed script (-b 4096 -ub 4096, 59 of 62 layers' experts in RAM) returned 3 of 3 codes planted in a 43,679-token document, passed its 40 stress prompts, and returned the code on the crash prompt.

07

Knobs we measured and refused

On DeepSeek V4 Flash Q8, 20 September, 16,011-token reads at a 131,072 window:

SettingReading, t/sCard after load, MiB
default micro-batch, f16 cache, every expert in RAM81.617,229
default micro-batch, q8_0 K and V cache82.416,662
-b 4096 -ub 2048, f16 cache260.817,933
-b 4096 -ub 2048, q8_0 K and V cache255.117,465

Short reads: the request asked for a one-word reply and generated 8 tokens; these runs did not check an answer.

  • A quantized cache: no gain, and a known risk. The q8_0 cache read 0.9 percent faster at the default micro-batch and 2.2 percent slower at 2048. Two upstream reports describe garbage output with a quantized K cache on this architecture: llama.cpp #25382, opened and closed as completed on 7 July, and #26423, from a build that already carried the earlier fix, closed as not planned on 5 August. We refused it.
  • Moving experts onto the card: not measured on this file. Our records said it was measured here and gained nothing; it was not (section 13). An upstream report describes this model's output degrading in proportion to how many expert layers run on the card (llama.cpp #25582, opened 12 July on a different machine, closed as not planned on 5 September). We refuse it on that report alone. The IQ3 file we also run keeps 7 of its 43 expert layers on the card and returned its code at every depth we read, 2,998 to 150,103 tokens; one planted code is a narrow check, not a quality test.
08

A thinking model with a small budget returns an empty answer

Our first 48,000-token check on DeepSeek V4 Flash (20 September) set max_tokens to 64. Every rung, the untuned default included, came back empty: the model spent all 64 tokens reasoning. An empty answer after a settings change looks exactly like corruption; here it was the budget. Rerun at 900, the same reads used 70 to 79 tokens and returned the code at the default, 4096 and 8192, and the harness has used 900 since, reading the reasoning field too.

The same thing happened once at 900: GLM-4.7-Flash's -ub 2048 rung spent all 900 tokens reasoning, with the code in its reasoning and an empty answer, while the rungs either side, 1024 and 4096, answered with the code in 433 and 488 tokens. The pass rule in section 02 counts that as a pass for reading speed; as an answer it failed, and the table in section 03 marks it so.

09

What a bigger micro-batch costs after the first read

Everything above reads a document the server had not seen, with nothing in its prompt cache. After a read the server keeps the prompt in its cache; the question is how much of the next request it reads again. For models whose cache cannot be rolled back token by token, or that use sliding-window attention, llama-server keeps restore points 4 + n_ubatch tokens and 4 tokens before the end of every prompt (tools/server/server-context.cpp lines 3465 to 3472 and 3559 to 3565 at d3146f2b5, added by llama.cpp pull request #20288). DeepSeek V4 Flash declares sliding-window attention in its file header (deepseek4.attention.sliding_window = 128).

We measured the cost on one configuration: DeepSeek V4 Flash IQ3 split across the desktop and a laptop (section 10), 262,144 window, 21 September. For each document, three requests: the cold read; the same document with a different last question (about 200 words asked for); and a real next turn, sending back the first request, the reply as returned, and a short new question.

Re-read on the second or third request-ub 512-ub 2048-ub 4096
Swapped last question, 3,000-token document533 tok, 4.4 s2,065 tok, 9.4 s3,015 tok (all of it), 12.4 s
Swapped last question, 48,000-token document533 tok, 10.0 s2,065 tok, 28.2 sno result (section 10)
Real next turn, 3,000-token document245 tok, 3.0 s227 tok, 2.8 s233 tok, 2.9 s
Real next turn, 48,000-token document238 tok, 5.4 s228 tok, 5.0 sno result

At 150,103 tokens (-ub 512 only): swapped question 533 tokens in 20.7 seconds, real next turn 239 tokens in 9.7 seconds.

A normal next turn re-read only the new material, 227 to 245 tokens, at every setting. A changed last message went back to the restore point one micro-batch before the end, so it re-read about one micro-batch: 3.9 times as many tokens at 2048 as at 512 (2,065 against 533), and 2.8 times the wait on the 48,000-token document. If you often edit or regenerate the last message on a model like this, that is the price of the bigger micro-batch; a normal conversation does not pay it. The code applies the same restore points to every model whose cache cannot be rolled back or that uses sliding-window attention; we measured the cost on DeepSeek only.

10

Two machines are the exception

The one setup with experts in RAM where we kept the default: DeepSeek V4 Flash IQ3 split across the desktop and an ASUS ROG Flow Z13 laptop over Thunderbolt with llama.cpp's RPC backend (21 September, 262,144 window; layers 0 to 7 on the desktop's card, 8 to 17 with their experts in the desktop's RAM, 18 to 42 on the laptop).

  • 512 (the default): 161.8 t/s on a 2,998-token read, 102.5 on 48,024 tokens (468.4 seconds).
  • 2048: 224.9 and 117.0 (410.5 seconds), 14 percent faster at 48,000 tokens.
  • 4096: 250.5 on the short read; on the 48,024-token read, after 40,960 tokens, the desktop's server logged the laptop's worker as crashed (“Remote RPC server crashed or returned malformed response”). The laptop's log of that moment is not in our records.

We kept 512: the split read 150,103 tokens in 2,922.9 seconds (48.7 minutes, 51.4 t/s) with the code correct, and decoded 12.05 t/s on a reply of about 400 tokens at 48,000 and 7.85 at 150,000. The IQ3 file on the desktop alone, at its shipped -ub 4096, read the same 150,103 tokens in 237.5 seconds (632.1 t/s); the two-machine study is linked in section 14.

11

How to tune your own

What we would do on a new model, in order. Setup takes a few minutes; then each rung is one long read, and one untuned read on our largest file took about 10 minutes by itself (576.9 seconds). With the stress check, Qwen3-235B took 35 minutes in all and Qwen3.5-397B 34.

  1. Find out where the weights live. On the card, expect less: five of our six such models read 3 to 34.5 percent faster, the sixth slower, and all six kept the default (step 4). With any expert layers in system RAM (-cmoe, --n-cpu-moe, --fit, or a file bigger than your card), this is worth doing.
  2. Build one read you can check. A long prompt (we used 48,000 tokens) with a code planted in the middle and a question asking for it; temperature 0; a document the server has not seen; max_tokens large enough for a thinking model (we used 900). Take timings.prompt_per_second from the response, and look for the code in the answer, not only in the reasoning. Time the second read after a start as well: on the one start we recorded that began with none of Qwen3.8-Flash-Next's file in memory, its first request read at 87.7 t/s while about 14 GB of the file paged in, and the next at 517.6 (section 04). Our harness is in the package.
  3. Sweep upward. Your current setting, then -b 4096 -ub 2048, -b 4096 -ub 4096, -b 8192 -ub 8192. Keep -b at least as large as -ub. Stop at the first rung that fails to load, drops the connection mid-read, loses the code, or reads slower than the one before.
  4. Decide whether the gain is worth its cost. The stop rule finds the fastest rung that works, not the one to ship. We shipped a bigger setting on every model with experts in RAM, where reading got at least twice as fast, and on none of the six on the card, where the gain was 34.5 percent at most and their scripts record the trade (section 03).
  5. Watch the card during the read, not just after load. The working buffer fills while the prompt is read; our peaks ran up to 465 MiB above the load reading. Leave room for anything else that shares the card.
  6. Re-run the winner at the window you actually serve, with a long read. At double the window the same setting took 2,178 to 5,167 MiB more card here, and one model that loaded -ub 4096 at 131,072 could not load it at 196,608. A short read there does not tell you the long-read speed (section 04).
  7. Stress it with real text before you trust it. Forty to sixty prompts of real mixed text sized around one and two micro-batches, plus repetitive input (logs, tables, ledgers), a small max_tokens, and a restart after any crash, keeping the prompt that caused it. Replay any prompt that ever crashed a server. Ling's -ub 4096 passed 110 prompts and a 150,153-token read, and still died on one 3,007-token ledger.
  8. Check that the setting reached the server. In our logs, llama-server's progress lines advance in steps of the logical batch (512 on Laguna's inherited -b 512, 2,048 at the default, 4,096 at -b 4096); at our log level the micro-batch itself was not printed.
  9. Judge decode on a real reply. Most of our decode figures are answers of 11 to 100 tokens. They can show a large change; they are not speaking speeds.
  10. If you often edit or regenerate the last message, each edit re-reads about one micro-batch on models that keep restore points (section 09); a normal next turn does not.
12

What this does not show

  • One machine. One RTX 5090 (32 GB), 188 GiB of RAM, one CPU, and a RAM-to-card link we did not record or measure; the gain depends on how fast weights cross, so a different card or link will move it.
  • Experts in system RAM. The large gains belong to that placement. Models on the card gained 34.5 percent at most, and we shipped none of it.
  • The builds of that week. Seven llama.cpp commits, three of them forks or a pull-request branch. Defaults and kernels change.
  • Reading, not quality. The check is one planted code (three in two runs), and the stress runs looked for crashes, not answers. Nothing here measures whether a model's replies got better or worse.
  • Decode from short answers. Most decode figures come from answers of 11 to 100 tokens (up to 900 on thinking models).
  • Mixed passes. Hand-copied flags at 131,072 on 20 September, each model's own script at its served window on 21 September. Ling's ratio spans both; GLM-5.3-Flash's spans two read lengths; M2.7's comes from a harness with no answer check.
  • Speed at the served window. Inkling-Small was timed on short reads only at the 262,144 it serves; Qwen3.8-Flash-Next was re-measured there on 26 September (section 04).
  • Run-to-run variation. The untuned DeepSeek 48,000-token read measured 72.3 t/s in one run and 83.3 in another with the same flags, after loads of 3 minutes 55 seconds and 3 seconds. We use 83.3, the less flattering baseline; single runs carry noise of that size.
  • First requests after a start. Most sweep reads were the first request after a load, and the sweep did not record the page cache. The 26 September run, which did, kept the gain with about 58 to 60 GB of the 90 GB file in memory: 2.02, 2.75 and 2.61 times at about 3,000, 48,075 and 229,981 tokens (section 04).
  • One request at a time. Nothing reused from the prompt cache, no parallel slots.
  • A stress check is not a proof. Forty prompts per changed setting found nothing; the Ling fault passed 110.
  • Upstream report states are as we found them on 26 September 2026.
13

What we got wrong

This site publishes corrections of its own records, and this study has five.

Two inherited micro-batches nobody had measured

Laguna S 2.1 ran at -ub 128 from July to 21 September, and MiniMax M2.7 was first judged at -cmoe -ub 128, the pair of flags in the MiniMax M3 start script. Measured, those values cost a factor of 19.9 and 17.5 in reading (sections 03 and 06).

“Confirmed” at the served window

After the 20 September pass, three start scripts carried the same comment above their new setting: “Confirmed to LOAD and answer at the full served window (ctx 262144)”, then a card figure. The run cited for Ling's -ub 4096 (10,285 MiB) died on its first request: the comment took its figure from the load line of a result file that also records the crash, and the script shipped 4096 until the afternoon of 21 September. The runs cited for Inkling-Small and Qwen3.8-Flash-Next at 2048 (17,281 and 26,561 MiB) did load and answer, but each was one 3,000-token read; the comments gave reading speeds from 131,072 only, and Flash-Next's 45.4 t/s at 262,144 appeared in neither. Flash-Next's long reads at that window were first measured on 26 September, and both comments were corrected that day (section 04).

An “experts on the card” test that moved no experts

Our records said moving DeepSeek V4 Flash's experts onto the card was measured and bought nothing (255.1 against 260.8 t/s). Those runs used --n-cpu-moe 50 on a 43-layer model (deepseek4.block_count = 43), so every expert stayed in RAM and card use fell by 468 to 567 MiB; they measured the q8_0 cache. The same records called the reports behind both refusals open: #25382 had been closed as completed on 7 July, before any of these runs, and #25582 as not planned on 5 September (section 07).

A speaking figure that was an 11-token answer

We shipped -ub 2048 on Qwen3-235B instead of 4096 partly because an 11-token answer decoded 12 percent slower at 4096 (7.20 against 8.18 t/s). An 11-token answer is not a speaking measurement. The other half of the reason stands: at 4096 the card peaked at 31,053 MiB. 4096 read 944.2 t/s. The same kind of figure is part of why Gemma 4 26B-A4B kept the default (section 03).

Labels in our own harness

The result files of runs with no override say “llama.cpp default 512” even when the script itself set -ub 128 (Laguna) or 4096 (Ling's crash run), and the M2.7 console printed “FIT” for the two window settings that failed to load. The server logs are right and the labels are wrong; the package README lists each case.

14

Related pages on this site

The DeepSeek V4 Flash page covers the model this sweep sped up the most, and the Inkling-Small page a model on a draft llama.cpp branch. The September field notes collect the month's shorter results from this machine; Two boxes studies the desktop and laptop as one host; the cost of a full 262,144-token window at the default has its own page. Slower by default and The wrong flag measure two other llama.cpp defaults on this desktop. Empty Answer is the thinking-budget trap in detail; the method page explains the labels and the data packages.

15

Sources and artifacts

The evidence for this page ships as a data package: the files the runs produced, not a summary. If a number on this page disagrees with a file in the package, the file is right and the page is wrong; tell us and we will fix the page.

  • data/README.md: every folder, what it holds, where the files disagree with each other, and every redaction.
  • data/NUMBERS.md: every figure on this page, with its file and field.
  • The runs (data/runs/): result files, server logs and card-memory samples for every run on this page.
  • The tools (data/tools/): the sweep drivers, the stress and turn probes, the Ling diagnostics, the standalone crash reproducer, and the served-window read harness with its page-cache reader, with private paths removed.
  • Derived tables (data/derived/): the server-log reads on this page in one table, the model files and their headers, the builds, the start-script census, and the source lines quoted above.
  • Read at pinned revisions: llama.cpp source at commit d3146f2b5; llama.cpp pull request #20288 and issues #25382, #26423, #25582 and #28282, checked on 26 September 2026.
2026-09-15DeepSeek V4 Flash Q8 at its full window and the default batch, 150,475 tokens in 34.4 minutes (a separate study; the depth baseline here).
2026-09-19MiniMax M2.7: the micro-batch ladder, the placement ladder at 4096, the 196,608-window check.
2026-09-20First pass with hand-copied flags (Ling, Inkling, Flash-Next; Laguna failed to start); the DeepSeek V4 Flash sweep, empty-answer pass and 150,324-token read; Ling's crashed full-window load; MiniMax M3 split buffers.
2026-09-21The remaining models and Ling through their own scripts; the Ling crash reproduced, diagnosed and fixed; the stress checks; the DeepSeek IQ3 values; the two-machine sweep and turn probes.
2026-09-26Qwen3.8-Flash-Next re-measured at its served window, with the model file's page-cache residency recorded; upstream report states checked; file headers and build records read for this page.