Graphometer
Reference · as of 26 September 2026 · one RTX 5090 desktop, 188 GiB of RAM · llama.cpp · one model at a time

The roster, as of 26 September 2026

Twenty-three local model endpoints on one RTX 5090 desktop, one running at a time; one row also uses a linked laptop. Every reading speed says its prompt size and batch setting, and every speaking speed says what reply it timed.

Start with the fact that makes the rest of the page readable: only one of these models can run at a time. They share one graphics card and one pool of system memory, and the launch policy refuses a second model while one is up. So this is not a fleet. It is a list of what this single desktop can be, one model at a time, with the speed each one produced, the prompt that produced it, and the dated run each figure comes from. Since the list was first compiled on 15 September, three models were added and two removed, and most reading speeds were measured again after we found that llama.cpp's batch settings had never been tuned on this desktop. The largest change was Laguna S 2.1: at its old 32,768-token window it read the same 24,071-token prompt at 72.2 tokens a second with -b 512 -ub 128 and at 1,434.0 with -b/-ub 8192; at the 262,144-token window it now serves, it reads 898.8 on 24,013 tokens at -b 8192 -ub 4096. Several figures in our own records did not survive a check against their logs; section 05 lists them. Two rows have no September figure for the configuration they serve, and say so.

Independent measurements. Not affiliated with, endorsed by, or connected to DeepSeek, Thinking Machines Lab, Alibaba Group, Z.ai, Google, MiniMax, Meta, Mistral AI, inclusionAI, Poolside, Ornith AI, Moonshot AI, NVIDIA, Intel, ASUS, or the llama.cpp project.
01

Twenty-three endpoints on a one-model machine

The machine is one desktop: an NVIDIA GeForce RTX 5090 with 32,607 MiB of graphics memory, an Intel Core Ultra 9 285K and 188 GiB of system RAM. Row 1 also uses a second machine, an ASUS ROG Flow Z13 with 128 GB of unified memory, joined to the desktop by a direct Thunderbolt cable; its cells say so. Every model is served by llama-server from llama.cpp, and every start script refuses to launch while another model is up; the package holds one recorded test in which seven separate guards each refused a launch.

What "endpoint" means here One installed service: one set of serving flags, with the model file or files that service is set up to read, as installed on this machine. Rows 1 and 3 are two separate services for DeepSeek V4 Flash 0731, one split across the desktop and the laptop and one on this desktop alone. Each can serve either of two files as presets: row 3's figures name their file, and row 1's are all from its default file except the one its row note names.

Changes since 15 September, as our operator records have them: Ornith-1.5-35B-A3B and GLM-5.3-Flash were installed on 16 September and MiniMax M2.7 on 21 September. Kimi K3 was retired on 15 September and Kimi K2.7-Code's weights were removed on 17 September. MiniMax M3 serves a different file since 21 September, Laguna S 2.1 a 262,144-token window since the same day, and Qwen3.6-27B a different file since 26 September; our records also give it a 262,144-token window from that day, which the runs on this page loaded by passing it explicitly. Qwen3.8-27B was installed at a 262,144-token window on the afternoon of 26 September; its start script still accepts 131,072. Qwen3.5-122B-A10B and Ornith-1.5-35B-A3B were installed at 262,144 the same afternoon, MiniMax M2.7 at 196,608, its native ceiling, that evening, and GLM-5.3-Flash at 262,144 later that evening; a start of each of these four with no window passed was served that window. Each row's window cell says whether a run that passed no window shows the installed default. A count of twenty-three installed endpoints is not a capacity claim: one runs at a time.

02

What we ran, exactly

The figures come from several exercises between 12 and 26 September 2026. Two methods never share a column: a request sent straight to the server (direct), and the same kind of request made through an agent framework.

The batch-size sweep
19 to 21 September
On 21 September each model was started through its own start script; the 19 and 20 September runs (MiniMax M2.7 and M3, DeepSeek's 8-bit file, Inkling-Small, Qwen3.8-Flash-Next) started the server by hand with the same flags. Each run read a synthetic prompt cold at temperature 0, most often a ledger of about 48,000 tokens with a code word planted halfway that the answer had to quote exactly. Reading is prompt tokens per second. The speaking figure is the short answer's generation rate: a code of about eleven tokens on models that answer directly, longer on thinking models because their hidden reasoning counts. Each row names the batch setting it ran (-b batch, -ub micro-batch; llama.cpp's defaults are -b 2048 -ub 512). The batch-size study has the method.
Paragraph replies
12 and 15 September
A short request for plain prose, the better of two warm replies: one paragraph of about 150 words on 15 September, about 200 tokens of prose on 12 September. These are the page's prose speaking figures for most rows.
First runs of two new models
16 September
For Ornith-1.5-35B-A3B and GLM-5.3-Flash: the same paragraph request with the thinking switch off, a tool call, and a read of about 120,000 tokens with three planted codes, at llama.cpp's default batch settings.
Letters at depth
26 September
For five rows: a ledger read cold at about 20,000, 60,000 and 100,000 tokens (Qwen3-235B), 20,000, 100,000 and 230,000 (Qwen3.6-27B) or 20,000, 48,000, 100,000 and 230,000 (Qwen3.8-27B, Qwen3.5-122B-A10B and Ornith-1.5-35B-A3B), then a letter of about 330 to 430 words written on the same cached ledger: prose at depth, as are row 1's replies of about 400 tokens. MiniMax M2.7 had the same request at about 20,000, 48,000 and 190,000 tokens, and GLM-5.3-Flash at about 20,000 to 230,000; their rows say what came back.
Long reads at a full window
15 to 16 September
Prompts of about 150,000 to 231,000 tokens with three planted codes, at a 262,144-token window (202,752 for GLM-4.7-Flash). Used here only for the window column and a few depth figures; the full table is its own page, what a full window costs.
Through an agent framework
12 to 14 and 21 September
The model driven through an agent framework with 48,000 and 96,000 tokens of seeded history, three codes planted in it, and a tool call made at depth. The framework adds its own prompt, so the server read more than the seeded figure; that true size is printed. Its speaking figures are its short answers (30 to 1,004 tokens, mostly the three codes plus any reasoning) and its reading figure is the harness's own estimate from token counts and timing.

Why there are two speed columns and not one

A model measured directly and through a framework does not give the same number. Gemma-4-26B read at 8,972.3 tokens a second directly at 47,983 tokens and at an estimated 6,726.8 through the framework at 53,209. Qwen3-235B spoke at 9.3 through the framework on a 62-token answer at 53,458 tokens, and at 5.51 directly on a 478-token letter at 20,063. Neither figure is wrong: they answer different questions, and mixing them in one column would be the worst error available on this page. Every speaking figure below also says what it timed, because a short answer and a page of prose are not the same measurement.

03

The twenty-three, with their newest figures

Every speed below is measured, on the dates given, one request at a time. Files and byte counts are read from the files on this machine (a listing taken on 26 September); maker names are the makers' own. "Short answer" means the model's reply to a quote-the-code question, and its token count includes hidden reasoning. Reading speeds after the batch-size sweep name the flag and the depth. Row numbers match ROSTER_23.csv in the data package.

# Model and file Window as served, and the batch setting there Speaking t/s, direct
(what was timed)
Reading t/s, direct
(prompt, setting)
Through an agent framework
(48K leg / 96K leg)
Newest figure
1 DeepSeek V4 Flash 0731, across two machinesDeepSeek. UD-IQ3_XXS, 4 shards, 104,207,848,032 bytes; a UD-Q4_K_XL preset with the maker's draft model 131,072 by the script's fallback; the installed knob may raise it; 262,144 was loaded on 15 and 21 September. -b 2048 -ub 512 At 262,144 (2026-09-21): 14.73, 12.05, 7.85 on replies of about 400 tokens at 2,998, 48,024 and 150,103 tokens; 14.76, 12.04 and 7.66 on the short code answers of 67, 69 and 83 tokens at the same depths; hidden reasoning included At 262,144, -b 2048 -ub 512: 161.8 on 2,998 tokens, 102.5 on 48,024, 51.4 on 150,103 3 of 3 codes and the tool call at both legs; reading about 97.4 on 52,928 tokens and 66.6 on 102,478; speaking 11.7 and 10.0 on 134 and 128-token answers (2026-09-14, 131,072 window). The record's verdict is FAIL, on a separate leg of our own runtime 2026-09-21
2 Inkling-SmallThinking Machines Lab. UD-Q3_K_XL, 4 shards, 119,554,379,840 bytes 262,144, the installed default: a run that passed no window was served it (2026-09-21); the 14 September probes served 131,072. -b 4096 -ub 2048 11.77 on a 202-token paragraph (2026-09-15, this window); 8.78 on a 27-token answer after 230,827 tokens At 262,144: 115.07 on 230,827 tokens at the defaults (2026-09-15); with -b 4096 -ub 2048, short reads only, 263.6 on 3,033 (2026-09-20) and 187.9 on 3,033 (2026-09-21). At 131,072 with that setting: 337.9 on 48,115 (2026-09-20) 3 of 3 and the tool call at both legs; reading about 107.9 on 52,925 and 111.1 on 102,473; speaking 12.9 and 10.9 on 145 and 31-token answers (2026-09-14) 2026-09-21
3 DeepSeek V4 Flash 0731, this desktop aloneDeepSeek. Two presets: 8-bit UD-Q8_K_XL, 5 shards, 161,869,615,520 bytes (the default); 3-bit UD-IQ3_XXS as row 1, experts of 36 of 43 layers in system memory 262,144, the installed default for both presets: runs that passed no window were served it (2026-09-21). 8-bit -b/-ub 8192, 3-bit -b/-ub 4096 8-bit: 9.93 on a 209-token reply (2026-09-15); 9.37 on a 106-token answer after 150,324 tokens. 3-bit: 12.65 on a 59-token answer at 2,998 tokens; 11.79 on 86 tokens after 150,103 8-bit, -b/-ub 8192: 480.0 on 150,324 tokens at 262,144, and 607.1 on 48,073 at 131,072 (2026-09-20). 3-bit, -b/-ub 4096, at 262,144: 690.9 on 48,024 and 632.1 on 150,103 8-bit, at the defaults: 3 of 3 and the tool call at both legs; reading about 75.5 on 52,931 and 73.8 on 102,481; speaking 9.7 and 9.2 on 139 and 119-token answers (2026-09-12) 2026-09-21
4 Qwen3.5-397B-A17BAlibaba Group. UD-IQ3_XXS, 4 shards, 149,840,422,542 bytes 131,072 at -b/-ub 4096; 262,144 on request at -b 2048 -ub 512, where the larger buffer would not fit 16.41 on an 11-token answer after 48,029 tokens; 12.84 on a 184-token paragraph at the 262,144 rung (2026-09-15) 932.2 on 48,029 tokens, -b/-ub 4096; 224.5 at the defaults 96K leg only: 3 of 3 and the tool call; reading about 195.2 on 103,266; speaking 19.1 on a 34-token answer (2026-09-13) 2026-09-21
5 Qwen3-235B-A22B Instruct 2507Alibaba Group. Q4_K_M, one file, 142,154,073,632 bytes, with a 0.6B draft model 131,072 at -b 4096 -ub 2048 5.51, 4.68, 4.22 on letters of about 350 words (478, 426 and 426 tokens) at 20,063, 60,180 and 99,864 tokens 741.8 on 20,063 tokens, 652.0 on 60,180, 542.1 on 99,864, -b 4096 -ub 2048; 703.3 on 48,020 (2026-09-21) 3 of 3 and the tool call at both legs; reading about 241.0 on 53,458 and 215.5 on 115,508; speaking 9.3 and 4.4 on 62-token answers (2026-09-12) 2026-09-26
6 Qwen3.5-122B-A10BAlibaba Group. UD-Q4_K_S, 3 shards, 73,428,564,992 bytes 262,144, the installed default: a run that passed no window was served it (2026-09-26); -b 2048 -ub 512 there, where the larger buffer would not fit; draft head on. The start script also accepts 131,072, at -b/-ub 4096 At 262,144 (2026-09-26): 24.6, 24.47, 23.84 and 21.64 on letters of about 420 words (482, 520, 508 and 491 tokens) at 20,067, 47,992, 99,872 and 229,090 tokens, thinking off. Earlier: 32.72 on a 716-token answer, most of it reasoning, after 48,027 tokens at 131,072 (2026-09-21); 22.39 on a 188-token paragraph at 262,144 (2026-09-15) At 262,144, -b 2048 -ub 512 (2026-09-26): 505.7 on 20,067 tokens, 511.4 on 47,992, 491.6 on 99,872, 448.7 on 229,090 (510.6 s); every planted code quoted, 3 of 3 at 229,090. At 131,072 (2026-09-21): 2,101.7 on 48,027 tokens, -b/-ub 4096; 520.2 at the defaults 48K leg only: 3 of 3 and the tool call; reading about 467.1 on 53,674; speaking 31.3 on a 240-token answer (2026-09-12, 131,072 window) 2026-09-26
7 Qwen3.8-27BAlibaba Group. Dense, with a built-in draft head. UD-Q5_K_XL, 20,218,178,624 bytes, the default; the start script can also serve UD-Q4_K_XL, 17,559,178,144 bytes, and UD-Q5_K_S, 18,665,753,504 bytes, both on disk 262,144 as installed on 26 September (the install record), loaded on 26 September by passing the window; the start script's comment records a 17 September measurement. At 262,144 on the default file: 8-bit cache, draft head off, -b 2048 -ub 512. The draft head stays on for the UD-Q4_K_XL and UD-Q5_K_S files at 262,144, and for 131,072 on the default file At 262,144, draft head off (2026-09-26): 60.7, 55.18, 49.08 and 34.89 on letters of about 400 words (436, 497, 477 and 504 tokens) at 20,067, 48,054, 99,934 and 229,152 tokens. At 131,072 with the draft head on: 86.01 on paragraph replies (2026-09-12, better of two); 99.26 on a 106-token answer after 48,069 tokens (2026-09-21) At 262,144, -b 2048 -ub 512 (2026-09-26): 3,118.2 on 20,067 tokens, 2,717.6 on 48,054, 1,958.2 on 99,934, 1,180.7 on 229,152; every planted code quoted, 3 of 3 at 229,152. At 131,072: 2,634.8 on 48,069 at the defaults, kept; -ub 4096 would not load at that window (2026-09-21) 3 of 3 and the tool call at both legs; reading about 2,376.8 on 54,073 and 1,797.6 on 103,936; speaking 99.6 and 97.7 on 244 and 252-token answers (2026-09-13, 131,072 window) 2026-09-26
8 Ornith-1.5-35B-A3BOrnith AI. Q6_K, one file, 29,208,731,392 bytes. Installed 2026-09-16 262,144, the installed default: a run that passed no window was served it (2026-09-26). At 262,144: -b 2048 -ub 2048 since 26 September (-b 2048 -ub 512 before). The start script also accepts 131,072, at -b/-ub 4096 At 262,144 (2026-09-26): 98.4, 89.8, 82.61 and 67.73 on letters of about 340 words (422, 405, 423 and 405 tokens) at 20,067, 47,992, 99,872 and 229,090 tokens, thinking off. At 131,072: 89.72 on a 198-token paragraph, thinking off (2026-09-16); 73.2 on an 83-token answer after 48,027 tokens (2026-09-21) At 262,144, -b 2048 -ub 2048 (2026-09-26): 4,568.0 on 20,067 tokens, 4,626.9 on 47,992, 4,219.0 on 99,872, 3,193.3 on 229,090 (71.7 s); every planted code quoted, 3 of 3 at 229,090. At 131,072: 6,283.3 on 48,027 tokens, -b/-ub 4096 (2026-09-21); 1,649.22 on 120,334 at the defaults, 3 of 3 codes (2026-09-16) not run through the framework 2026-09-26
9 GLM-5.3-FlashZ.ai. UD-Q3_K_XL, 4 shards, 147,535,921,955 bytes. Installed 2026-09-16 262,144, the installed default: a run that passed no window was served it (2026-09-26); -b 4096 -ub 1024 there, where -ub 4096 did not load. The start script also accepts 131,072, at -b/-ub 4096 At 262,144, -ub 1024 (2026-09-26), reasoning effort set to none and counting the reasoning it still did: 10.22 on 1,200 tokens at 20,065 tokens deep, all reasoning, no letter; 9.98 on 729 tokens, a 351-word letter, at 47,992. At -ub 2048, measured and not adopted: 10.11 on 1,200 tokens (102 words, cut off) at 20,065; 9.87 on 1,200 tokens (119 words, cut off) at 47,992; 9.56 on 567 tokens, a 337-word letter, at 100,003; 9.04 on 1,200 tokens at 230,039, all reasoning, no letter. Earlier, at 131,072: 9.58 on a 181-token paragraph (2026-09-16); 9.12 on a 310-token answer after 48,168 tokens (2026-09-21) At 262,144, -b 4096 -ub 1024 (2026-09-26): 135.1 on 20,065 tokens and 142.6 on 47,992, the only depths measured at this setting, the planted code quoted at both. At -ub 2048, measured and not adopted: 206.0, 250.3, 243.1 and 211.9 on 20,065, 47,992, 100,003 and 230,039 (1,085.7 s), every code quoted, 3 of 3 at 230,039. At 131,072: 425.1 on 48,168 tokens, -b/-ub 4096 (2026-09-21); 75.48 on 120,979 at the defaults, 3 of 3 codes (2026-09-16) not run through the framework 2026-09-26
10 Qwen3.6-27BAlibaba Group. Dense hybrid. Since 2026-09-26 a UD-Q5_K_XL build with a draft head, 20,350,682,240 bytes; the earlier UD-Q4_K_XL, 17,612,564,704 bytes, stays selectable 131,072 by the script's fallback; the installed knob may raise it; 262,144 was loaded on 26 September. -b 2048 -ub 512 At 262,144 (2026-09-26): 60.85, 46.66, 34.61 on letters of about 400 words (481, 527 and 526 tokens) at 20,067, 99,934 and 229,564 tokens At 262,144, -b 2048 -ub 512: 3,162.8 on 20,067 tokens, 1,971.1 on 99,934, 1,165.6 on 229,564 The earlier file at 131,072: 3 of 3 and the tool call at both legs; reading about 2,739.7 on 53,177 and 2,066.0 on 102,680; speaking 62.0 and 55.2 on 291 and 1,004-token answers (2026-09-13) 2026-09-26
11 Qwen3.8-Flash-NextAlibaba Group. UD-Q3_K_XL, 3 shards, 89,986,353,824 bytes 262,144, the installed default: a run that passed no window was served it (2026-09-21); 8-bit cache since 2026-09-20; -b 4096 -ub 2048. 393,216 and 524,288 on request, same setting in the script; the reads at those windows were at the defaults, before the change 23.50 on a 32-token answer after 5,986 tokens; 13.08 on 32 tokens after 229,982 (2026-09-20, this window and cache) At 262,144, -b 4096 -ub 2048 (2026-09-26), after the server's first request, with 57.7 to 60.1 GB of the 90 GB file in memory: 517.6 on 3,019 tokens, 660.6 on 48,075, 535.4 on 229,981 (429.6 s); at llama.cpp's defaults 255.7, 240.2 and 205.5 (1,119.3 s); every planted code quoted, 3 of 3 at 229,981. The 45.4 on 3,041 (2026-09-20) was the first request after a start (load 65 s). 96K leg only: 3 of 3 and the tool call; reading about 175.8 on 103,377; speaking 16.1 on a 301-token answer (2026-09-12) 2026-09-26
12 GLM-4.7-FlashZ.ai. UD-Q4_K_XL, one file, 17,520,169,312 bytes 202,752, its own ceiling and the installed default (a run that passed no window, 2026-09-21); llama.cpp's defaults 227.43 on a 160-token paragraph at this window (2026-09-15); 129.6 on a 429-token answer, most of it reasoning, after 47,986 tokens 2,594.6 on 47,986 tokens at the defaults, kept 96K leg only: 3 of 3 and the tool call; reading about 1,406.6 on 105,183; speaking 99.0 on a 299-token answer (2026-09-12) 2026-09-21
13 GLM-4.7 FullZ.ai. IQ3_XXS, 3 shards, 145,120,517,024 bytes 131,072 with a 5-bit cache, as served on 2026-09-12; llama.cpp's defaults 5.51 on paragraph replies (2026-09-12, better of two, before a draft head was switched on) 61.36 on 20,221 tokens at the defaults (2026-09-12) 3 of 3 codes at both legs; the tool call worked at 48K and failed at 96K, so the check failed; reading about 58.6 on 54,138 and 58.5 on 105,333; speaking 3.8 and 3.1 on 30-token answers (2026-09-12) 2026-09-12
14 Gemma-4-26B-A4BGoogle. UD-Q4_K_XL, one file, 17,010,980,576 bytes 262,144, the installed default (a run that passed no window, 2026-09-21); llama.cpp's defaults 202.64 on a 163-token paragraph at this window (2026-09-15); 163.68 on an 11-token answer after 47,983 tokens 8,972.3 on 47,983 tokens at the defaults, kept; 3,799.69 on 230,855 (2026-09-15) 3 of 3 and the tool call at both legs; reading about 6,726.8 on 53,209 and 5,048.5 on 102,635; speaking 124.1 and 97.3 on 36-token answers (2026-09-12) 2026-09-21
15 Gemma-4-31B IT QATGoogle. Dense, q4_0 quantization-aware build, one file, 17,651,001,568 bytes 131,072; llama.cpp's defaults 73.68 on paragraph replies (2026-09-12, better of two); 56.45 on an 11-token answer after 47,983 tokens 2,687.5 on 47,983 tokens at the defaults, kept 3 of 3 and the tool call at both legs; reading about 2,075.4 on 53,213 and 1,620.5 on 102,640; speaking 47.4 and 41.9 on 36-token answers (2026-09-12) 2026-09-21
16 Ling-3.0-flashinclusionAI. Q4_K_M, 2 shards, 77,804,990,144 bytes, MIT licence 262,144, the installed default (a run that passed no window, 2026-09-21); -b 4096 -ub 2048 29.57 on a 172-token paragraph at this window (2026-09-15); 23.25 on a 77-token answer after 47,986 tokens; 22.6 on 135 tokens after 149,711 727.7 on 47,986 tokens and 694.6 on 149,711, -b 4096 -ub 2048 3 of 3 and the tool call at both legs; reading about 202.6 on 53,718 and 204.9 on 103,807; speaking 24.0 and 23.8 on 173 and 243-token answers (2026-09-13) 2026-09-21
17 Mistral Small 4 119BMistral AI. UD-Q4_K_M, 3 shards, 73,763,180,544 bytes 131,072 at -b/-ub 8192; 262,144 on request at -b 2048 -ub 512, where the larger buffer would not fit 29.55 on a 249-token paragraph at the 262,144 rung (2026-09-15); 20.97 on an 11-token answer after 48,697 tokens 2,192.3 on 48,697 tokens, -b/-ub 8192; 481.3 at the defaults 3 of 3 and the tool call at both legs; reading about 463.3 on 52,820 and 421.6 on 101,684; speaking 23.5 and 21.7 on 40 and 39-token answers (2026-09-12) 2026-09-21
18 MiniMax M2.7MiniMax. UD-IQ4_XS, 4 shards, 108,413,781,312 bytes. Installed 2026-09-21 196,608, its native ceiling and the installed default: a run that passed no window was served it (2026-09-26); there the start script sets the one shape that loads: every expert in system memory, -b 4096 -ub 1024, 8-bit cache. The start script also accepts 131,072. The 19 September figures at that window were measured with experts of 59 of 62 layers in system memory and -b/-ub 4096; the 13 September prose figure was measured with every expert in system memory and -ub 128. At 196,608 (2026-09-26), counting its hidden reasoning (it always thinks): 10.01 and 8.96 on 4,096 tokens at 20,039 and 47,933 tokens deep, all of it reasoning, with no letter returned; 6.05 on 515 tokens, a 275-word letter with its reasoning, at 189,491. At 131,072: 9.73 on a prose reply of 1,313 characters, 1,522 tokens with its hidden reasoning, after a 70-token prompt (2026-09-13, every expert in system memory); 9.83 on a 32-token answer after 43,909 tokens (2026-09-19) At 196,608, -b 4096 -ub 1024, every expert in system memory (2026-09-26): 167.9 on 20,039 tokens, 181.2 on 47,933, 155.9 on 189,491 (1,215.4 s); every planted code quoted, 3 of 3 at 189,491. At 131,072: 656.5 on 43,909 tokens, -b/-ub 4096, experts of 59 of 62 layers in system memory (2026-09-19) 3 of 3 and the tool call at both legs; reading about 529.2 on 53,584 and 460.2 on 103,931; speaking 8.6 and 7.2 on 180 and 182-token answers (2026-09-21, 131,072 window) 2026-09-26
19 MiniMax M3MiniMax. Since 2026-09-21 a Q2_K_L file that carries the model's sparse-attention index, 4 shards, 153,086,988,768 bytes 131,072 at -b 4096 -ub 2048; 196,608 on request at -b 4096 -ub 512 8.90 on an 82-token answer at 3,680 tokens; 8.29 on a 165-token answer after 58,307 tokens (2026-09-20) 208.11 on 58,307 tokens, -ub 2048, 3 of 3 codes (2026-09-20); 186.8 on 3,009 through the model manager (2026-09-21) not run through the framework 2026-09-21
20 Laguna S 2.1Poolside. Q4_K_M, one file, 68,248,759,648 bytes 262,144, the start script's own default since 2026-09-21; -b 8192 -ub 4096 15.74 and 14.69 on short answers after 24,013 and 150,158 tokens; 12.73 on a 35-token answer after 200,098 898.8 on 24,013 tokens and 911.6 on 150,158 through the model manager; 866.1 on 200,098; -ub 4096 not run through the framework 2026-09-21
21 Muse Glimmer 30BMeta. Q4_K_M, one file, 16,756,683,904 bytes, with a block drafter: a companion model that proposes several tokens at once 131,072; llama.cpp's defaults 193.65 on a 139-token answer, most of it reasoning, after 48,090 tokens 2,737.7 on 48,090 tokens at the defaults, kept not run through the framework 2026-09-21
22 Mistral Medium 3.5Mistral AI. Dense, 128B. UD-Q4_K_XL, 3 shards, with a 3B draft model patched to its vocabulary 32,768 in the start script's default profile, at -b 4096 -ub 1024 No September figure for this configuration. It was started for the framework check on 2026-09-12 and did not answer within the check's timeout. A different, smaller file of this model ran split across two machines on 16 September; that run is on the two-machine study, not here. none in September
23 Mistral Large 3 675BMistral AI. UD-IQ1_S, 4 shards 32,768, the largest its start script accepts, at -b 4096 -ub 2048 No September figure. It was not included in any September run. none in September

"Speaking" is generation speed, tokens produced per second; "reading" is prompt processing, tokens of prompt consumed per second, which decides how long you wait before the first word. Every figure's own qualifier (file, window, setting, date) controls. Rows without a date in a cell carry the date in the last column. Every row that shows a planted-code read quoted back all the codes at the depths listed; row 13 is the only failed tool call at depth. The framework's reading figures are its own estimates and its speaking figures are short answers, as section 02 says.

Row notes

1 and 3, DeepSeek V4 Flash 0731 The speaking figures count hidden reasoning; this model thinks by default. On the same file, window (262,144) and prompts, the desktop alone at -b/-ub 4096 read 6.7 times faster than the pair at its kept -b 2048 -ub 512 at 48,024 tokens, and 12.3 times at 150,103 (arithmetic on the table's figures). At the same micro-batch of 2,048, with -b 4096 on the desktop and -b 2048 on the pair, the pair was slightly faster on 2,998 tokens (224.9 against 208.1) and 3.6 times slower at 48,024 (117.0 against 426.5). A whole 48,024-token request took 75.2 seconds alone, 474.2 on the pair at its kept setting and 416.4 at -b/-ub 2048. The two-machine study has the comparison and the 4-bit preset's 19.55 speaking on the pair (2026-09-15, at 131,072). Row 1's framework cell is the pair's attempt 2, whose long legs passed and whose verdict is FAIL on a separate leg; its attempt 3 passed with the short legs only, so no single record holds both.
2, Inkling-Small The framework figures come from a run made while the model was installed under a temporary name; that run failed a separate leg that checks how a model is listed in our own runtime, and its recall and tool results at both depths passed. Its batch setting changed on 20 September; at the served window the only reads with it are of 3,033 tokens, each the first request after the server started, and the 337.9 was read at 131,072.
5, Qwen3-235B Its speaking figures are real prose at depth, and they replace an 8.2 from the batch-size sweep that was an 11-token answer. With -ub 4096 it read faster (1,069.0 on 20,063 tokens) and spoke the same on letters (5.82 against 5.51), at a higher card peak; the served setting is 2048. In the same letter runs the quote-the-code read passed at every depth; the edit-and-reread request, whose reply is capped at 40 tokens, did not return the code at any depth, because the cap ends the letter before the line that quotes it.
6, Qwen3.5-122B-A10B Installed at 262,144 on the afternoon of 26 September, at the batch setting its script already used at that window. There -b 2048 -ub 1024 read faster (877.8 on 20,067 tokens, 738.9 on 229,090 in 310.1 s) but peaked at 31,785 MiB of the card's 32,607, so it was measured and not adopted; the setting adopted held 30,898 MiB at load and 31,124 at the peak of the reads. The same day at 131,072 with -b/-ub 4096 it read about four times faster (2,088.3 on 47,992) and spoke no faster (22.82 on a 516-token letter, against 24.47 at 262,144). The 26 September runs passed their window through a copy of the start script whose flags at that window match the installed ones; the draft head was on throughout.
7, Qwen3.8-27B Installed at 262,144 on the afternoon of 26 September. On this file the draft head does not fit beside a 256K cache, so the start script turns it off at that window; the card held 30,058 MiB at load and 30,254 at the peak of the reads. The draft head stays on for the UD-Q4_K_XL file at 262,144 and for the UD-Q5_K_S file at 262,144 (18,665,753,504 bytes), and for 131,072 on the default file, which is the window of the 12 to 21 September figures; the two smaller files are different sets of weights. The 26 September run passed its window itself, through a copy of the start script whose flags match the installed ones; the 256K field card has the figures for the other configurations. As in rows 5 and 10, the edit-and-reread request, capped at 40 tokens of reply, did not return the code at any depth. At 131,072, -ub 4096 could not allocate a 1,568.13 MiB compute buffer and the server exited, so that window keeps the defaults.
8, Ornith-1.5-35B-A3B Installed at 262,144 on the afternoon of 26 September, with the micro-batch at that window raised from 512 to 2048: at -b 2048 -ub 512 the same reads ran at 1,780.4 on 20,067 tokens and 1,508.7 on 229,090 (151.8 s), with speaking about the same. The card held 31,270 MiB at load and 31,417 at the peak of the reads. The same day at 131,072 with -b/-ub 4096 it read 6,114.1 on 47,992 and spoke at 89.64 on a 383-token letter. The 26 September runs passed their window through a copy of the start script whose flags at that window match the installed ones. It has not been run through the agent framework.
9, GLM-5.3-Flash Installed at 262,144 on the evening of 26 September at -b 4096 -ub 1024. At that window -ub 4096 did not load: its compute buffer alone asked for 13,281.37 MiB and the allocation failed. -ub 2048 loaded and read faster but peaked at 31,951 MiB of the card's 32,607; -ub 1024 held 28,834 MiB at load and 29,261 at the peak of its reads. The reads at -ub 1024 reach only 20,065 and 47,992 tokens. At 262,144, the figures past that depth (100,003 and 230,039) are the -ub 2048 run's. The 120,979-token read is the 16 September default-batch run, and the 100,003-token read at 131,072 is the same day's -b/-ub 4096 run. These runs set the reasoning effort to none; the model still reasoned, and several letters ran out of their 1,200-token budget; the 9.04 at 230,039 returned no letter. The same day at 131,072 with -b/-ub 4096 it read 359.4, 404.9 and 373.4 on 20,065, 47,992 and 100,003 tokens. A 40-prompt crash check at 262,144 and -ub 1024 ran without a crash; in 36 of the 40 a reply came back with no visible text after using its whole budget (160 tokens for the reply, 80 for the follow-up), and in the other four both replies had text, the first of them, a 3,014-token ledger, quoting its planted code. It has not been run through the agent framework.
10, Qwen3.6-27B At 262,144 the draft head is off: a start with it on could not create the draft context (2026-09-26). Its start script only began applying the 8-bit cache that window uses on 26 September; the setting was referenced but never defined before. The framework column is the earlier file at the earlier window. As in row 5, the quote-the-code read passed at every depth and the edit-and-reread request, capped at 40 tokens of reply, did not return the code at any depth.
11, Qwen3.8-Flash-Next Its configuration changed three times in one week: expert placement on 17 September, and the 8-bit cache and the batch setting on 20 September. On 26 September the server was started once at each setting, and every read was a new prompt. The first start began with none of the 90 GB file in memory and is the only start recorded with the file out of memory: its first request, 3,035 tokens, read at 87.7 while the file's share in memory grew from 43.5 to 57.7 GB, 14.1 GB read in from disk, and the next, 3,019 tokens, at 517.6. The run at the defaults came second and began with 60.1 GB in memory; the file never became fully resident. The card held 26,795 MiB at load with -b 4096 -ub 2048 and 21,333 at the defaults. The earlier 511.7 was a 48,069-token read at 131,072 (20 September), and the 197.14 a 229,982-token read at the defaults, before the batch change.
13, GLM-4.7 Full A failed check, not a pass. A built-in draft head was switched on in its start script on 17 September; its speed with that change is recorded only in a same-day table whose raw logs were not kept, so no figure for it is printed. It was not re-measured by the batch-size sweep.
18, MiniMax M2.7 Installed at 196,608, its native ceiling, on the evening of 26 September. That window loads in one shape only, every expert in system memory at -ub 1024, and the start script sets it itself; the card held 31,709 MiB at load and 31,907 at the peak of the reads. At 131,072, with experts of 59 of 62 layers in system memory and -b/-ub 4096, it read 656.5 on 43,909 tokens (19 September), 3.6 times the 181.2 at 196,608 on 47,933. It always thinks: at 20,039 and 47,933 tokens the letter request spent its whole 4,096-token budget on hidden reasoning and returned no letter, so those two speaking figures time reasoning only. The 26 September run went through the installed start script with the window passed. Its first desktop run, on 13 September, used -ub 128, a value that, per our session notes, was copied from another model's launcher, and read 45.6 tokens a second on 6,776 tokens. At llama.cpp's default it read 100.8 on 3,658 tokens, and at 4096 the 656.5 above (2026-09-19). The 13 September prose reply is unaffected by that setting. With a 900-token limit the same request had spent every token on hidden reasoning and returned no answer.
19, MiniMax M3 The file it served until 21 September was an earlier conversion that, per our session notes, lacked the tensors this model's sparse attention needs; the launch check on the file it serves now reports the sparse attention engaged. At 262,144 the new file could not allocate its compute buffer at any micro-batch from 512 to 2048, so it serves 131,072; at 196,608 it loads only at -ub 512, and read 58,307 tokens at 70.16 a second, 3 of 3 (2026-09-20). It has never run through the framework.
20, Laguna S 2.1 Until 21 September its script set -b 512 -ub 128, and at its old 32,768 window it read 72.2 tokens a second on 24,071 tokens; at -b/-ub 8192 it read 1,434.0 on the same prompt. Above 32,768 its script now uses -b 8192 -ub 4096, the value measured at the larger windows.
Not listed Kimi K3 (retired 15 September) and Kimi K2.7-Code (weights removed 17 September) are no longer endpoints; K3's August record is its own field card, and K2.7's one September measurement is on the two-machine study.
04

What a batch setting did, and did not, change

The largest movements in this table are reading speeds on models whose experts sit in system memory. Each read is about 48,000 tokens, the same prompt at both settings, except GLM-5.3-Flash's default-setting read of 24,008 tokens; all six at a 131,072-token window, measured on 21 September through each model's own start script (Qwen3.5-122B, Ornith and GLM-5.3-Flash have served 262,144 since 26 September; rows 6, 8 and 9):

ModelTokens readAt the defaultsAt the setting it now runsSetting, at 131,072
Qwen3.5-397B-A17B48,029224.5932.2-b/-ub 4096
Qwen3.5-122B-A10B48,027520.22,101.7-b/-ub 4096
Mistral Small 4 119B48,697481.32,192.3-b/-ub 8192
Ornith-1.5-35B-A3B48,0271,860.26,283.3-b/-ub 4096
Qwen3-235B-A22B48,020285.8703.3-b 4096 -ub 2048
GLM-5.3-Flash24,008 / 48,16880.7 (on 24,008)425.1 (on 48,168)-b/-ub 4096

All measured, 2026-09-21; in these six, the planted code came back in the answer at every setting that loaded. The speaking figures in those runs moved by 6.2 percent at most, Qwen3.5-122B's 30.81 to 32.72 (arithmetic on the same files). At 262,144 the start scripts of Qwen3.5-397B, Qwen3.5-122B and Mistral Small 4 keep -b 2048 -ub 512: the larger buffer does not fit there; since 26 September Ornith's has used -b 2048 -ub 2048 there and GLM-5.3-Flash's -b 4096 -ub 1024. The models that fit whole on the card kept the defaults: Qwen3.8-27B, GLM-4.7-Flash, the two Gemmas and Muse Glimmer. On each one's best rung, reading moved from a 0.6 percent loss (Muse Glimmer) to a 34.5 percent gain (Gemma-4-26B-A4B at -ub 2048, which spoke 6 percent slower), arithmetic on each model's rungs. Qwen3.8-27B's rungs were run at 131,072, its installed window on 21 September. Qwen3.6-27B kept them too; its rungs were run on its earlier Q4 file at 131,072. GLM-4.7-Flash at -ub 2048 put the code only in its hidden reasoning and returned an empty answer.

Three rows moved for other reasons: MiniMax M2.7 and Laguna S 2.1 had inherited -ub 128, Laguna with -b 512 (row notes 18 and 20), and MiniMax M3 changed file. DeepSeek's two services moved further still, and the two-machine study sets them side by side. How to run the sweep on your own models, and the one value that passed every read and still crashed a server on a specific input, are on the batch-size study.

05

What changed since 15 September

This list was first compiled on 15 September. The later runs corrected figures in our own summary table and in that first list. We publish those corrections the same way we publish everything else.

Our records saidWhat the logs showThis page
Qwen3-235B speaks at 8.2 tokens a second (21 September)That was an 11-token answer. On letters of about 350 words it speaks 5.51, 4.68 and 4.22 at 20,063, 60,180 and 99,864 tokensThe letters (row 5)
Qwen3.5-397B is the fastest-reading large model, at 385By our own records, the 385 was a 2,617-token prompt at a 64K window, measured in August. At the served window the batch-size sweep reads 224.5 at the defaults and 932.2 at -b/-ub 4096 on 48,029 tokens, and Qwen3-235B reads faster than it at depth on the framework's 96K leg (215.5 against 195.2, estimates)The measured figures, with their prompt sizes (row 4)
GLM-4.7 Full passed the framework checkIts record's verdict is a failure: the tool call at 96,000 tokens failedA failed check (row 13); the 15 September list already printed it as partial
MiniMax M3 reads at 194 tokens a second at depth194 is 58,307 tokens divided by the whole request's 300.2 seconds, answer included. The server's own timing is 208.11208.11 (row 19)
MiniMax M2.7 reads 45.6 on a 4,000-token promptThe probe was named for about 4,000 tokens and sent 6,776, at an inherited -ub 1286,776 tokens, the setting, and the later figures (row 18)
Laguna S 2.1 reads at 72 tokens a secondIts script had hard-coded -b 512 -ub 128; at -b/-ub 8192 the same read ran at 1,434.0, both at the old 32,768 window. At the 262,144 window it now serves it reads 898.8 on 24,013 tokens at -b 8192 -ub 4096The window and setting it now serves (row 20)
Qwen3.8-Flash-Next reads at 511.7 with its current setting, and at the window it serves had only a 45.4 on a 3,041-token readThe 511.7 was a 48,069-token read at a 131,072 window. The 45.4 was the first request after a start (load 65 s). At the 262,144 window it serves, with -b 4096 -ub 2048, after the server's first request and with 57.7 to 60.1 GB of the 90 GB file in memory, it reads 517.6 on 3,019 tokens, 660.6 on 48,075 and 535.4 on 229,981, against 255.7, 240.2 and 205.5 at llama.cpp's defaults (26 September)The served-window figures (row 11)
The two-machine DeepSeek speaks at 15.9 to 17.1That band is 13 September at a 131,072 window. At a 262,144 window on 21 September: 14.73 on a 400-token reply at 2,998 tokens and 12.05 at 48,024 (14.76 and 12.04 on the short code answers)The 21 September figures (row 1)
The 15 September list: 22 endpoints, all at 128K, seven listed without speedsThree added, two retired; ten rows now serve more than 131,072 by default, each shown by a run that passed no window (202,752 or 262,144, and 196,608 for MiniMax M2.7); Laguna S 2.1 serves 262,144 by its start script's own default; Qwen3.8-27B was installed at 262,144 on 26 September; two more loaded 262,144 when it was passed; two rows without a September figure23 rows, each with its served window

What held up, in the same breath: every read behind this page's figures that asked for a planted code returned it in its answer, except GLM-4.7-Flash at -ub 2048, whose code stayed in its hidden reasoning (section 04). The letter runs' edit-and-reread requests, capped at 40 tokens of reply, did not return the code at any depth (row notes 5, 7 and 10).

06

What this does not show

  • No measure of quality. Speed, window, recall of a planted code and whether a tool call came back are the only things measured. A model can be the fastest row here and the wrong choice for your work.
  • Short answers are not prose. Where a row's only speaking figure at depth is a short answer, it says so; eight rows have prose at depth (1, 5, 6, 7, 8, 9, 10 and 18), and most have a paragraph at a short prompt. A thinking model's short answer counts its reasoning.
  • Figures from different exercises are not a ranking. Rows were measured on different days, at different depths and settings, and each cell says which. Compare within a column and a date.
  • One desktop. These are honest samples on one machine, not benchmarks, and not comparable with a vendor's published throughput or with a hosted service.
  • Windows above the served one are not claimed here. "On request" means the start script offers a larger window and a September run loaded it: Qwen3.5-397B and Mistral Small 4 read about 231,000 tokens at 262,144 (15 September, at -b 2048 -ub 512, the setting their scripts keep there), Qwen3.8-Flash-Next read 344,981 and 459,911 tokens at its two larger windows (20 September, at llama.cpp's default batch sizes, before its batch change), and MiniMax M3 58,307 at 196,608 (20 September, -ub 512). What a full window costs is on its own page.
  • The one-at-a-time constraint is this machine's. On more hardware several of these would run side by side.
  • Two rows have no figure for what they serve. That is a statement about our September, not about the models.
07

Related pages on this site

08

Sources and artifacts

The data package holds the redacted primary files these figures come from, a few clearly identified extracts, one rebuilt machine-readable table, and the vendor documents cited at pinned revisions. Machine paths, addresses, port numbers, service names and our internal keys for each model are removed; so is a column naming private work. Nothing that carries a measurement is altered, and no figure is recomputed except where a file says it is a derivation.

  • data/README.md: every file, where it came from, what was removed from it, and the places where the package looks like it argues with itself.
  • data/NUMBERS.md: every figure on this page mapped to its file and field.
  • data/ROSTER_23.csv: the table, machine readable, with the file and date behind each figure.
  • The runs of 19 to 26 September: the batch-size sweep's result files and logs, the long reads, the paragraph replies, the letters at depth, the MiniMax records, the new framework record, and the 26 September runs of Qwen3.8-Flash-Next, Qwen3.8-27B, Qwen3.5-122B-A10B and Ornith-1.5-35B-A3B and GLM-5.3-Flash at 262,144 and of MiniMax M2.7 at 196,608, and that evening's starts with no window passed.
  • The runs of 12 to 15 September: the direct sweep, the framework check's ledgers and per-model records, the refusal-guard test, and what the model files say.

If a number on this page disagrees with a file in its package, the file is right and the page is wrong. Tell us at hello@graphometer.ai and the correction goes on the page with its date.

2026-09-15List first compiled: 22 endpoints, runs of 12 to 15 September.
2026-09-26Rebuilt as this snapshot: 23 endpoints, runs of 12 to 26 September, with the corrections in section 05.
2026-09-26, updateQwen3.8-Flash-Next read at its served 262,144 window at both batch settings, and Qwen3.8-27B installed at 262,144 with its figures there: rows 7 and 11, their notes, and sections 01, 02, 04, 05 and 06, with the runs' files in the package.
2026-09-26, eveningQwen3.5-122B-A10B, Ornith-1.5-35B-A3B and GLM-5.3-Flash installed at 262,144, and MiniMax M2.7 at 196,608, with their figures there and a start of each with no window passed: rows 6, 8, 9 and 18, their notes, and sections 01, 02, 04, 05 and 06, with the runs' files in the package.

This page is dated on purpose. It is a snapshot as of 26 September 2026, not a living list; a later roster will be a later page with a later date.