The roster, as of 26 September 2026
Twenty-three local model endpoints on one RTX 5090 desktop, one running at a time; one row also uses a linked laptop. Every reading speed says its prompt size and batch setting, and every speaking speed says what reply it timed.
Start with the fact that makes the rest of the page
readable: only one of these models can run at a time.
They share one graphics card and one pool of system memory, and the
launch policy refuses a second model while one is up. So this is not a
fleet. It is a list of what this single desktop can be, one model at a
time, with the speed each one produced, the prompt that produced it,
and the dated run each figure comes from. Since the list was first
compiled on 15 September, three models were added and two removed, and
most reading speeds were measured again after we found that llama.cpp's
batch settings had never been tuned on this desktop. The largest change
was Laguna S 2.1: at its old 32,768-token window it read the same
24,071-token prompt at 72.2 tokens a second with -b 512 -ub
128 and at 1,434.0 with -b/-ub 8192; at the
262,144-token window it now serves, it reads 898.8 on 24,013 tokens at
-b 8192 -ub 4096. Several figures in our own records did not survive a
check against their logs; section 05 lists them. Two rows have no
September figure for the configuration they serve, and say so.
Twenty-three endpoints on a one-model machine
The machine is one desktop: an NVIDIA GeForce RTX 5090 with 32,607
MiB of graphics memory, an Intel Core Ultra 9 285K and 188 GiB of system
RAM. Row 1 also uses a second machine, an ASUS ROG Flow Z13 with 128 GB
of unified memory, joined to the desktop by a direct Thunderbolt cable;
its cells say so. Every model is served by llama-server
from llama.cpp, and every start script refuses to launch while another
model is up; the package holds one recorded test in which seven separate
guards each refused a launch.
Changes since 15 September, as our operator records have them: Ornith-1.5-35B-A3B and GLM-5.3-Flash were installed on 16 September and MiniMax M2.7 on 21 September. Kimi K3 was retired on 15 September and Kimi K2.7-Code's weights were removed on 17 September. MiniMax M3 serves a different file since 21 September, Laguna S 2.1 a 262,144-token window since the same day, and Qwen3.6-27B a different file since 26 September; our records also give it a 262,144-token window from that day, which the runs on this page loaded by passing it explicitly. Qwen3.8-27B was installed at a 262,144-token window on the afternoon of 26 September; its start script still accepts 131,072. Qwen3.5-122B-A10B and Ornith-1.5-35B-A3B were installed at 262,144 the same afternoon, MiniMax M2.7 at 196,608, its native ceiling, that evening, and GLM-5.3-Flash at 262,144 later that evening; a start of each of these four with no window passed was served that window. Each row's window cell says whether a run that passed no window shows the installed default. A count of twenty-three installed endpoints is not a capacity claim: one runs at a time.
What we ran, exactly
The figures come from several exercises between 12 and 26 September 2026. Two methods never share a column: a request sent straight to the server (direct), and the same kind of request made through an agent framework.
| The batch-size sweep 19 to 21 September |
On 21 September each model was started through its own start
script; the 19 and 20 September runs (MiniMax M2.7 and M3, DeepSeek's
8-bit file, Inkling-Small, Qwen3.8-Flash-Next) started the server by
hand with the same flags. Each run read a synthetic prompt cold at
temperature 0, most often a ledger of about 48,000 tokens with a code
word planted halfway that the answer had to quote exactly. Reading is
prompt tokens per second. The speaking figure is the
short answer's generation rate: a code of about eleven tokens on
models that answer directly, longer on thinking models because their
hidden reasoning counts. Each row names the batch setting it ran
(-b batch, -ub micro-batch; llama.cpp's
defaults are -b 2048 -ub 512). The
batch-size study has the method. |
|---|---|
| Paragraph replies 12 and 15 September |
A short request for plain prose, the better of two warm replies: one paragraph of about 150 words on 15 September, about 200 tokens of prose on 12 September. These are the page's prose speaking figures for most rows. |
| First runs of two new models 16 September |
For Ornith-1.5-35B-A3B and GLM-5.3-Flash: the same paragraph request with the thinking switch off, a tool call, and a read of about 120,000 tokens with three planted codes, at llama.cpp's default batch settings. |
| Letters at depth 26 September |
For five rows: a ledger read cold at about 20,000, 60,000 and 100,000 tokens (Qwen3-235B), 20,000, 100,000 and 230,000 (Qwen3.6-27B) or 20,000, 48,000, 100,000 and 230,000 (Qwen3.8-27B, Qwen3.5-122B-A10B and Ornith-1.5-35B-A3B), then a letter of about 330 to 430 words written on the same cached ledger: prose at depth, as are row 1's replies of about 400 tokens. MiniMax M2.7 had the same request at about 20,000, 48,000 and 190,000 tokens, and GLM-5.3-Flash at about 20,000 to 230,000; their rows say what came back. |
| Long reads at a full window 15 to 16 September |
Prompts of about 150,000 to 231,000 tokens with three planted codes, at a 262,144-token window (202,752 for GLM-4.7-Flash). Used here only for the window column and a few depth figures; the full table is its own page, what a full window costs. |
| Through an agent framework 12 to 14 and 21 September |
The model driven through an agent framework with 48,000 and 96,000 tokens of seeded history, three codes planted in it, and a tool call made at depth. The framework adds its own prompt, so the server read more than the seeded figure; that true size is printed. Its speaking figures are its short answers (30 to 1,004 tokens, mostly the three codes plus any reasoning) and its reading figure is the harness's own estimate from token counts and timing. |
Why there are two speed columns and not one
A model measured directly and through a framework does not give the same number. Gemma-4-26B read at 8,972.3 tokens a second directly at 47,983 tokens and at an estimated 6,726.8 through the framework at 53,209. Qwen3-235B spoke at 9.3 through the framework on a 62-token answer at 53,458 tokens, and at 5.51 directly on a 478-token letter at 20,063. Neither figure is wrong: they answer different questions, and mixing them in one column would be the worst error available on this page. Every speaking figure below also says what it timed, because a short answer and a page of prose are not the same measurement.
The twenty-three, with their newest figures
Every speed below is measured, on the dates given, one request at a time. Files and byte counts are read from the files on this machine (a listing taken on 26 September); maker names are the makers' own. "Short answer" means the model's reply to a quote-the-code question, and its token count includes hidden reasoning. Reading speeds after the batch-size sweep name the flag and the depth. Row numbers match ROSTER_23.csv in the data package.
| # | Model and file | Window as served, and the batch setting there | Speaking t/s, direct (what was timed) |
Reading t/s, direct (prompt, setting) |
Through an agent framework (48K leg / 96K leg) |
Newest figure |
|---|---|---|---|---|---|---|
| 1 | DeepSeek V4 Flash 0731, across two machinesDeepSeek. UD-IQ3_XXS, 4 shards, 104,207,848,032 bytes; a UD-Q4_K_XL preset with the maker's draft model | 131,072 by the script's fallback; the installed knob may raise it; 262,144 was loaded on 15 and 21 September. -b 2048 -ub 512 |
At 262,144 (2026-09-21): 14.73, 12.05, 7.85 on replies of about 400 tokens at 2,998, 48,024 and 150,103 tokens; 14.76, 12.04 and 7.66 on the short code answers of 67, 69 and 83 tokens at the same depths; hidden reasoning included | At 262,144, -b 2048 -ub 512: 161.8 on 2,998 tokens, 102.5 on 48,024, 51.4 on 150,103 |
3 of 3 codes and the tool call at both legs; reading about 97.4 on 52,928 tokens and 66.6 on 102,478; speaking 11.7 and 10.0 on 134 and 128-token answers (2026-09-14, 131,072 window). The record's verdict is FAIL, on a separate leg of our own runtime | 2026-09-21 |
| 2 | Inkling-SmallThinking Machines Lab. UD-Q3_K_XL, 4 shards, 119,554,379,840 bytes | 262,144, the installed default: a run that passed no window was served it (2026-09-21); the 14 September probes served 131,072. -b 4096 -ub 2048 |
11.77 on a 202-token paragraph (2026-09-15, this window); 8.78 on a 27-token answer after 230,827 tokens | At 262,144: 115.07 on 230,827 tokens at the defaults (2026-09-15); with -b 4096 -ub 2048, short reads only, 263.6 on 3,033 (2026-09-20) and 187.9 on 3,033 (2026-09-21). At 131,072 with that setting: 337.9 on 48,115 (2026-09-20) |
3 of 3 and the tool call at both legs; reading about 107.9 on 52,925 and 111.1 on 102,473; speaking 12.9 and 10.9 on 145 and 31-token answers (2026-09-14) | 2026-09-21 |
| 3 | DeepSeek V4 Flash 0731, this desktop aloneDeepSeek. Two presets: 8-bit UD-Q8_K_XL, 5 shards, 161,869,615,520 bytes (the default); 3-bit UD-IQ3_XXS as row 1, experts of 36 of 43 layers in system memory | 262,144, the installed default for both presets: runs that passed no window were served it (2026-09-21). 8-bit -b/-ub 8192, 3-bit -b/-ub 4096 |
8-bit: 9.93 on a 209-token reply (2026-09-15); 9.37 on a 106-token answer after 150,324 tokens. 3-bit: 12.65 on a 59-token answer at 2,998 tokens; 11.79 on 86 tokens after 150,103 | 8-bit, -b/-ub 8192: 480.0 on 150,324 tokens at 262,144, and 607.1 on 48,073 at 131,072 (2026-09-20). 3-bit, -b/-ub 4096, at 262,144: 690.9 on 48,024 and 632.1 on 150,103 |
8-bit, at the defaults: 3 of 3 and the tool call at both legs; reading about 75.5 on 52,931 and 73.8 on 102,481; speaking 9.7 and 9.2 on 139 and 119-token answers (2026-09-12) | 2026-09-21 |
| 4 | Qwen3.5-397B-A17BAlibaba Group. UD-IQ3_XXS, 4 shards, 149,840,422,542 bytes | 131,072 at -b/-ub 4096; 262,144 on request at -b 2048 -ub 512, where the larger buffer would not fit |
16.41 on an 11-token answer after 48,029 tokens; 12.84 on a 184-token paragraph at the 262,144 rung (2026-09-15) | 932.2 on 48,029 tokens, -b/-ub 4096; 224.5 at the defaults |
96K leg only: 3 of 3 and the tool call; reading about 195.2 on 103,266; speaking 19.1 on a 34-token answer (2026-09-13) | 2026-09-21 |
| 5 | Qwen3-235B-A22B Instruct 2507Alibaba Group. Q4_K_M, one file, 142,154,073,632 bytes, with a 0.6B draft model | 131,072 at -b 4096 -ub 2048 |
5.51, 4.68, 4.22 on letters of about 350 words (478, 426 and 426 tokens) at 20,063, 60,180 and 99,864 tokens | 741.8 on 20,063 tokens, 652.0 on 60,180, 542.1 on 99,864, -b 4096 -ub 2048; 703.3 on 48,020 (2026-09-21) |
3 of 3 and the tool call at both legs; reading about 241.0 on 53,458 and 215.5 on 115,508; speaking 9.3 and 4.4 on 62-token answers (2026-09-12) | 2026-09-26 |
| 6 | Qwen3.5-122B-A10BAlibaba Group. UD-Q4_K_S, 3 shards, 73,428,564,992 bytes | 262,144, the installed default: a run that passed no window was served it (2026-09-26); -b 2048 -ub 512 there, where the larger buffer would not fit; draft head on. The start script also accepts 131,072, at -b/-ub 4096 |
At 262,144 (2026-09-26): 24.6, 24.47, 23.84 and 21.64 on letters of about 420 words (482, 520, 508 and 491 tokens) at 20,067, 47,992, 99,872 and 229,090 tokens, thinking off. Earlier: 32.72 on a 716-token answer, most of it reasoning, after 48,027 tokens at 131,072 (2026-09-21); 22.39 on a 188-token paragraph at 262,144 (2026-09-15) | At 262,144, -b 2048 -ub 512 (2026-09-26): 505.7 on 20,067 tokens, 511.4 on 47,992, 491.6 on 99,872, 448.7 on 229,090 (510.6 s); every planted code quoted, 3 of 3 at 229,090. At 131,072 (2026-09-21): 2,101.7 on 48,027 tokens, -b/-ub 4096; 520.2 at the defaults |
48K leg only: 3 of 3 and the tool call; reading about 467.1 on 53,674; speaking 31.3 on a 240-token answer (2026-09-12, 131,072 window) | 2026-09-26 |
| 7 | Qwen3.8-27BAlibaba Group. Dense, with a built-in draft head. UD-Q5_K_XL, 20,218,178,624 bytes, the default; the start script can also serve UD-Q4_K_XL, 17,559,178,144 bytes, and UD-Q5_K_S, 18,665,753,504 bytes, both on disk | 262,144 as installed on 26 September (the install record), loaded on 26 September by passing the window; the start script's comment records a 17 September measurement. At 262,144 on the default file: 8-bit cache, draft head off, -b 2048 -ub 512. The draft head stays on for the UD-Q4_K_XL and UD-Q5_K_S files at 262,144, and for 131,072 on the default file |
At 262,144, draft head off (2026-09-26): 60.7, 55.18, 49.08 and 34.89 on letters of about 400 words (436, 497, 477 and 504 tokens) at 20,067, 48,054, 99,934 and 229,152 tokens. At 131,072 with the draft head on: 86.01 on paragraph replies (2026-09-12, better of two); 99.26 on a 106-token answer after 48,069 tokens (2026-09-21) | At 262,144, -b 2048 -ub 512 (2026-09-26): 3,118.2 on 20,067 tokens, 2,717.6 on 48,054, 1,958.2 on 99,934, 1,180.7 on 229,152; every planted code quoted, 3 of 3 at 229,152. At 131,072: 2,634.8 on 48,069 at the defaults, kept; -ub 4096 would not load at that window (2026-09-21) |
3 of 3 and the tool call at both legs; reading about 2,376.8 on 54,073 and 1,797.6 on 103,936; speaking 99.6 and 97.7 on 244 and 252-token answers (2026-09-13, 131,072 window) | 2026-09-26 |
| 8 | Ornith-1.5-35B-A3BOrnith AI. Q6_K, one file, 29,208,731,392 bytes. Installed 2026-09-16 | 262,144, the installed default: a run that passed no window was served it (2026-09-26). At 262,144: -b 2048 -ub 2048 since 26 September (-b 2048 -ub 512 before). The start script also accepts 131,072, at -b/-ub 4096 |
At 262,144 (2026-09-26): 98.4, 89.8, 82.61 and 67.73 on letters of about 340 words (422, 405, 423 and 405 tokens) at 20,067, 47,992, 99,872 and 229,090 tokens, thinking off. At 131,072: 89.72 on a 198-token paragraph, thinking off (2026-09-16); 73.2 on an 83-token answer after 48,027 tokens (2026-09-21) | At 262,144, -b 2048 -ub 2048 (2026-09-26): 4,568.0 on 20,067 tokens, 4,626.9 on 47,992, 4,219.0 on 99,872, 3,193.3 on 229,090 (71.7 s); every planted code quoted, 3 of 3 at 229,090. At 131,072: 6,283.3 on 48,027 tokens, -b/-ub 4096 (2026-09-21); 1,649.22 on 120,334 at the defaults, 3 of 3 codes (2026-09-16) |
not run through the framework | 2026-09-26 |
| 9 | GLM-5.3-FlashZ.ai. UD-Q3_K_XL, 4 shards, 147,535,921,955 bytes. Installed 2026-09-16 | 262,144, the installed default: a run that passed no window was served it (2026-09-26); -b 4096 -ub 1024 there, where -ub 4096 did not load. The start script also accepts 131,072, at -b/-ub 4096 |
At 262,144, -ub 1024 (2026-09-26), reasoning effort set to none and counting the reasoning it still did: 10.22 on 1,200 tokens at 20,065 tokens deep, all reasoning, no letter; 9.98 on 729 tokens, a 351-word letter, at 47,992. At -ub 2048, measured and not adopted: 10.11 on 1,200 tokens (102 words, cut off) at 20,065; 9.87 on 1,200 tokens (119 words, cut off) at 47,992; 9.56 on 567 tokens, a 337-word letter, at 100,003; 9.04 on 1,200 tokens at 230,039, all reasoning, no letter. Earlier, at 131,072: 9.58 on a 181-token paragraph (2026-09-16); 9.12 on a 310-token answer after 48,168 tokens (2026-09-21) |
At 262,144, -b 4096 -ub 1024 (2026-09-26): 135.1 on 20,065 tokens and 142.6 on 47,992, the only depths measured at this setting, the planted code quoted at both. At -ub 2048, measured and not adopted: 206.0, 250.3, 243.1 and 211.9 on 20,065, 47,992, 100,003 and 230,039 (1,085.7 s), every code quoted, 3 of 3 at 230,039. At 131,072: 425.1 on 48,168 tokens, -b/-ub 4096 (2026-09-21); 75.48 on 120,979 at the defaults, 3 of 3 codes (2026-09-16) |
not run through the framework | 2026-09-26 |
| 10 | Qwen3.6-27BAlibaba Group. Dense hybrid. Since 2026-09-26 a UD-Q5_K_XL build with a draft head, 20,350,682,240 bytes; the earlier UD-Q4_K_XL, 17,612,564,704 bytes, stays selectable | 131,072 by the script's fallback; the installed knob may raise it; 262,144 was loaded on 26 September. -b 2048 -ub 512 |
At 262,144 (2026-09-26): 60.85, 46.66, 34.61 on letters of about 400 words (481, 527 and 526 tokens) at 20,067, 99,934 and 229,564 tokens | At 262,144, -b 2048 -ub 512: 3,162.8 on 20,067 tokens, 1,971.1 on 99,934, 1,165.6 on 229,564 |
The earlier file at 131,072: 3 of 3 and the tool call at both legs; reading about 2,739.7 on 53,177 and 2,066.0 on 102,680; speaking 62.0 and 55.2 on 291 and 1,004-token answers (2026-09-13) | 2026-09-26 |
| 11 | Qwen3.8-Flash-NextAlibaba Group. UD-Q3_K_XL, 3 shards, 89,986,353,824 bytes | 262,144, the installed default: a run that passed no window was served it (2026-09-21); 8-bit cache since 2026-09-20; -b 4096 -ub 2048. 393,216 and 524,288 on request, same setting in the script; the reads at those windows were at the defaults, before the change |
23.50 on a 32-token answer after 5,986 tokens; 13.08 on 32 tokens after 229,982 (2026-09-20, this window and cache) | At 262,144, -b 4096 -ub 2048 (2026-09-26), after the server's first request, with 57.7 to 60.1 GB of the 90 GB file in memory: 517.6 on 3,019 tokens, 660.6 on 48,075, 535.4 on 229,981 (429.6 s); at llama.cpp's defaults 255.7, 240.2 and 205.5 (1,119.3 s); every planted code quoted, 3 of 3 at 229,981. The 45.4 on 3,041 (2026-09-20) was the first request after a start (load 65 s). |
96K leg only: 3 of 3 and the tool call; reading about 175.8 on 103,377; speaking 16.1 on a 301-token answer (2026-09-12) | 2026-09-26 |
| 12 | GLM-4.7-FlashZ.ai. UD-Q4_K_XL, one file, 17,520,169,312 bytes | 202,752, its own ceiling and the installed default (a run that passed no window, 2026-09-21); llama.cpp's defaults | 227.43 on a 160-token paragraph at this window (2026-09-15); 129.6 on a 429-token answer, most of it reasoning, after 47,986 tokens | 2,594.6 on 47,986 tokens at the defaults, kept | 96K leg only: 3 of 3 and the tool call; reading about 1,406.6 on 105,183; speaking 99.0 on a 299-token answer (2026-09-12) | 2026-09-21 |
| 13 | GLM-4.7 FullZ.ai. IQ3_XXS, 3 shards, 145,120,517,024 bytes | 131,072 with a 5-bit cache, as served on 2026-09-12; llama.cpp's defaults | 5.51 on paragraph replies (2026-09-12, better of two, before a draft head was switched on) | 61.36 on 20,221 tokens at the defaults (2026-09-12) | 3 of 3 codes at both legs; the tool call worked at 48K and failed at 96K, so the check failed; reading about 58.6 on 54,138 and 58.5 on 105,333; speaking 3.8 and 3.1 on 30-token answers (2026-09-12) | 2026-09-12 |
| 14 | Gemma-4-26B-A4BGoogle. UD-Q4_K_XL, one file, 17,010,980,576 bytes | 262,144, the installed default (a run that passed no window, 2026-09-21); llama.cpp's defaults | 202.64 on a 163-token paragraph at this window (2026-09-15); 163.68 on an 11-token answer after 47,983 tokens | 8,972.3 on 47,983 tokens at the defaults, kept; 3,799.69 on 230,855 (2026-09-15) | 3 of 3 and the tool call at both legs; reading about 6,726.8 on 53,209 and 5,048.5 on 102,635; speaking 124.1 and 97.3 on 36-token answers (2026-09-12) | 2026-09-21 |
| 15 | Gemma-4-31B IT QATGoogle. Dense, q4_0 quantization-aware build, one file, 17,651,001,568 bytes | 131,072; llama.cpp's defaults | 73.68 on paragraph replies (2026-09-12, better of two); 56.45 on an 11-token answer after 47,983 tokens | 2,687.5 on 47,983 tokens at the defaults, kept | 3 of 3 and the tool call at both legs; reading about 2,075.4 on 53,213 and 1,620.5 on 102,640; speaking 47.4 and 41.9 on 36-token answers (2026-09-12) | 2026-09-21 |
| 16 | Ling-3.0-flashinclusionAI. Q4_K_M, 2 shards, 77,804,990,144 bytes, MIT licence | 262,144, the installed default (a run that passed no window, 2026-09-21); -b 4096 -ub 2048 |
29.57 on a 172-token paragraph at this window (2026-09-15); 23.25 on a 77-token answer after 47,986 tokens; 22.6 on 135 tokens after 149,711 | 727.7 on 47,986 tokens and 694.6 on 149,711, -b 4096 -ub 2048 |
3 of 3 and the tool call at both legs; reading about 202.6 on 53,718 and 204.9 on 103,807; speaking 24.0 and 23.8 on 173 and 243-token answers (2026-09-13) | 2026-09-21 |
| 17 | Mistral Small 4 119BMistral AI. UD-Q4_K_M, 3 shards, 73,763,180,544 bytes | 131,072 at -b/-ub 8192; 262,144 on request at -b 2048 -ub 512, where the larger buffer would not fit |
29.55 on a 249-token paragraph at the 262,144 rung (2026-09-15); 20.97 on an 11-token answer after 48,697 tokens | 2,192.3 on 48,697 tokens, -b/-ub 8192; 481.3 at the defaults |
3 of 3 and the tool call at both legs; reading about 463.3 on 52,820 and 421.6 on 101,684; speaking 23.5 and 21.7 on 40 and 39-token answers (2026-09-12) | 2026-09-21 |
| 18 | MiniMax M2.7MiniMax. UD-IQ4_XS, 4 shards, 108,413,781,312 bytes. Installed 2026-09-21 | 196,608, its native ceiling and the installed default: a run that passed no window was served it (2026-09-26); there the start script sets the one shape that loads: every expert in system memory, -b 4096 -ub 1024, 8-bit cache. The start script also accepts 131,072. The 19 September figures at that window were measured with experts of 59 of 62 layers in system memory and -b/-ub 4096; the 13 September prose figure was measured with every expert in system memory and -ub 128. |
At 196,608 (2026-09-26), counting its hidden reasoning (it always thinks): 10.01 and 8.96 on 4,096 tokens at 20,039 and 47,933 tokens deep, all of it reasoning, with no letter returned; 6.05 on 515 tokens, a 275-word letter with its reasoning, at 189,491. At 131,072: 9.73 on a prose reply of 1,313 characters, 1,522 tokens with its hidden reasoning, after a 70-token prompt (2026-09-13, every expert in system memory); 9.83 on a 32-token answer after 43,909 tokens (2026-09-19) | At 196,608, -b 4096 -ub 1024, every expert in system memory (2026-09-26): 167.9 on 20,039 tokens, 181.2 on 47,933, 155.9 on 189,491 (1,215.4 s); every planted code quoted, 3 of 3 at 189,491. At 131,072: 656.5 on 43,909 tokens, -b/-ub 4096, experts of 59 of 62 layers in system memory (2026-09-19) |
3 of 3 and the tool call at both legs; reading about 529.2 on 53,584 and 460.2 on 103,931; speaking 8.6 and 7.2 on 180 and 182-token answers (2026-09-21, 131,072 window) | 2026-09-26 |
| 19 | MiniMax M3MiniMax. Since 2026-09-21 a Q2_K_L file that carries the model's sparse-attention index, 4 shards, 153,086,988,768 bytes | 131,072 at -b 4096 -ub 2048; 196,608 on request at -b 4096 -ub 512 |
8.90 on an 82-token answer at 3,680 tokens; 8.29 on a 165-token answer after 58,307 tokens (2026-09-20) | 208.11 on 58,307 tokens, -ub 2048, 3 of 3 codes (2026-09-20); 186.8 on 3,009 through the model manager (2026-09-21) |
not run through the framework | 2026-09-21 |
| 20 | Laguna S 2.1Poolside. Q4_K_M, one file, 68,248,759,648 bytes | 262,144, the start script's own default since 2026-09-21; -b 8192 -ub 4096 |
15.74 and 14.69 on short answers after 24,013 and 150,158 tokens; 12.73 on a 35-token answer after 200,098 | 898.8 on 24,013 tokens and 911.6 on 150,158 through the model manager; 866.1 on 200,098; -ub 4096 |
not run through the framework | 2026-09-21 |
| 21 | Muse Glimmer 30BMeta. Q4_K_M, one file, 16,756,683,904 bytes, with a block drafter: a companion model that proposes several tokens at once | 131,072; llama.cpp's defaults | 193.65 on a 139-token answer, most of it reasoning, after 48,090 tokens | 2,737.7 on 48,090 tokens at the defaults, kept | not run through the framework | 2026-09-21 |
| 22 | Mistral Medium 3.5Mistral AI. Dense, 128B. UD-Q4_K_XL, 3 shards, with a 3B draft model patched to its vocabulary | 32,768 in the start script's default profile, at -b 4096 -ub 1024 |
No September figure for this configuration. It was started for the framework check on 2026-09-12 and did not answer within the check's timeout. A different, smaller file of this model ran split across two machines on 16 September; that run is on the two-machine study, not here. | none in September | ||
| 23 | Mistral Large 3 675BMistral AI. UD-IQ1_S, 4 shards | 32,768, the largest its start script accepts, at -b 4096 -ub 2048 |
No September figure. It was not included in any September run. | none in September | ||
"Speaking" is generation speed, tokens produced per second; "reading" is prompt processing, tokens of prompt consumed per second, which decides how long you wait before the first word. Every figure's own qualifier (file, window, setting, date) controls. Rows without a date in a cell carry the date in the last column. Every row that shows a planted-code read quoted back all the codes at the depths listed; row 13 is the only failed tool call at depth. The framework's reading figures are its own estimates and its speaking figures are short answers, as section 02 says.
Row notes
| 1 and 3, DeepSeek V4 Flash 0731 | The speaking figures count hidden reasoning; this model thinks by
default. On the same file, window (262,144) and prompts, the desktop
alone at -b/-ub 4096 read 6.7 times faster than the pair
at its kept -b 2048 -ub 512 at 48,024 tokens, and 12.3
times at 150,103 (arithmetic on the table's figures). At the same
micro-batch of 2,048, with -b 4096 on the desktop and
-b 2048 on the pair, the pair was slightly faster on
2,998 tokens (224.9 against 208.1) and 3.6 times slower at 48,024
(117.0 against 426.5). A whole 48,024-token request took 75.2
seconds alone, 474.2 on the pair at its kept setting and 416.4 at
-b/-ub 2048. The two-machine
study has the comparison and the 4-bit preset's 19.55 speaking
on the pair (2026-09-15, at 131,072). Row 1's framework cell is the
pair's attempt 2, whose long legs passed and whose verdict is FAIL on
a separate leg; its attempt 3 passed with the short legs only, so no
single record holds both. |
|---|---|
| 2, Inkling-Small | The framework figures come from a run made while the model was installed under a temporary name; that run failed a separate leg that checks how a model is listed in our own runtime, and its recall and tool results at both depths passed. Its batch setting changed on 20 September; at the served window the only reads with it are of 3,033 tokens, each the first request after the server started, and the 337.9 was read at 131,072. |
| 5, Qwen3-235B | Its speaking figures are real prose at depth, and they replace an
8.2 from the batch-size sweep that was an 11-token answer. With
-ub 4096 it read faster (1,069.0 on 20,063 tokens) and
spoke the same on letters (5.82 against 5.51), at a higher card peak;
the served setting is 2048. In the same letter runs the quote-the-code
read passed at every depth; the edit-and-reread request, whose reply
is capped at 40 tokens, did not return the code at any depth, because
the cap ends the letter before the line that quotes it. |
| 6, Qwen3.5-122B-A10B | Installed at 262,144 on the afternoon of 26 September, at the
batch setting its script already used at that window. There
-b 2048 -ub 1024 read faster (877.8 on 20,067 tokens,
738.9 on 229,090 in 310.1 s) but peaked at 31,785 MiB of the
card's 32,607, so it was measured and not adopted; the setting
adopted held 30,898 MiB at load and 31,124 at the peak of the
reads. The same day at 131,072 with -b/-ub 4096 it read
about four times faster (2,088.3 on 47,992) and spoke no faster
(22.82 on a 516-token letter, against 24.47 at 262,144). The 26
September runs passed their window through a copy of the start
script whose flags at that window match the installed ones; the
draft head was on throughout. |
| 7, Qwen3.8-27B | Installed at 262,144 on the afternoon of 26 September. On this
file the draft head does not fit beside a 256K cache, so the start
script turns it off at that window; the card held 30,058 MiB at
load and 30,254 at the peak of the reads. The draft head stays on for
the UD-Q4_K_XL file at 262,144 and for the UD-Q5_K_S file at 262,144
(18,665,753,504 bytes), and for 131,072 on the default file, which is
the window of the 12 to 21 September figures; the two smaller files
are different sets of weights. The 26 September run passed its window
itself, through a copy of the start script whose flags match the
installed ones; the 256K field card has
the figures for the other configurations. As in rows 5 and 10, the edit-and-reread request,
capped at 40 tokens of reply, did not return the code at any depth.
At 131,072, -ub 4096 could not allocate a
1,568.13 MiB compute buffer and the server exited, so that
window keeps the defaults. |
| 8, Ornith-1.5-35B-A3B | Installed at 262,144 on the afternoon of 26 September, with the
micro-batch at that window raised from 512 to 2048: at -b 2048
-ub 512 the same reads ran at 1,780.4 on 20,067 tokens and
1,508.7 on 229,090 (151.8 s), with speaking about the same. The card
held 31,270 MiB at load and 31,417 at the peak of the reads. The
same day at 131,072 with -b/-ub 4096 it read 6,114.1 on
47,992 and spoke at 89.64 on a 383-token letter. The 26 September
runs passed their window through a copy of the start script whose
flags at that window match the installed ones. It has not been run
through the agent framework. |
| 9, GLM-5.3-Flash | Installed at 262,144 on the evening of 26 September at -b
4096 -ub 1024. At that window -ub 4096 did not
load: its compute buffer alone asked for 13,281.37 MiB and the
allocation failed. -ub 2048 loaded and read faster but
peaked at 31,951 MiB of the card's 32,607; -ub 1024
held 28,834 MiB at load and 29,261 at the peak of its reads. The
reads at -ub 1024 reach only 20,065 and 47,992 tokens. At
262,144, the figures past that depth (100,003 and 230,039) are the
-ub 2048 run's. The 120,979-token read is the 16 September
default-batch run, and the 100,003-token read at 131,072 is the same
day's -b/-ub 4096 run. These
runs set the reasoning effort to none; the model still reasoned, and
several letters ran out of their 1,200-token budget; the 9.04 at
230,039 returned no letter. The same day at
131,072 with -b/-ub 4096 it read 359.4, 404.9 and 373.4
on 20,065, 47,992 and 100,003 tokens. A 40-prompt crash check at
262,144 and -ub 1024 ran without a crash; in 36 of the
40 a reply came back with no visible text after using its whole
budget (160 tokens for the reply, 80 for the follow-up), and in the
other four both replies had text, the first of them, a 3,014-token
ledger, quoting its planted code. It has not been run through the
agent framework. |
| 10, Qwen3.6-27B | At 262,144 the draft head is off: a start with it on could not create the draft context (2026-09-26). Its start script only began applying the 8-bit cache that window uses on 26 September; the setting was referenced but never defined before. The framework column is the earlier file at the earlier window. As in row 5, the quote-the-code read passed at every depth and the edit-and-reread request, capped at 40 tokens of reply, did not return the code at any depth. |
| 11, Qwen3.8-Flash-Next | Its configuration changed three times in one week: expert placement
on 17 September, and the 8-bit cache and the batch setting on 20
September. On 26 September the server was started once at each
setting, and every read was a new prompt. The first start began with
none of the 90 GB file in memory and is the only start recorded with
the file out of memory: its first request, 3,035 tokens, read at 87.7
while the file's share in memory grew from 43.5 to 57.7 GB, 14.1 GB
read in from disk, and the next, 3,019 tokens, at 517.6. The run at
the defaults came second and began with 60.1 GB in memory; the file
never became fully resident. The card held 26,795
MiB at load with -b 4096 -ub 2048 and 21,333 at the
defaults. The earlier 511.7 was a 48,069-token read at 131,072 (20
September), and the 197.14 a 229,982-token read at the defaults,
before the batch change. |
| 13, GLM-4.7 Full | A failed check, not a pass. A built-in draft head was switched on in its start script on 17 September; its speed with that change is recorded only in a same-day table whose raw logs were not kept, so no figure for it is printed. It was not re-measured by the batch-size sweep. |
| 18, MiniMax M2.7 | Installed at 196,608, its native ceiling, on the evening of 26
September. That window loads in one shape only, every expert in
system memory at -ub 1024, and the start script sets it
itself; the card held 31,709 MiB at load and 31,907 at the peak
of the reads. At 131,072, with experts of 59 of 62 layers in system
memory and -b/-ub 4096, it read 656.5 on 43,909 tokens
(19 September), 3.6 times the 181.2 at 196,608 on 47,933. It always
thinks: at 20,039 and 47,933 tokens the letter request spent its
whole 4,096-token budget on hidden reasoning and returned no letter,
so those two speaking figures time reasoning only. The 26 September
run went through the installed start script with the window passed.
Its first desktop run, on 13 September, used -ub 128,
a value that, per our session notes, was copied from another model's
launcher, and read 45.6 tokens a
second on 6,776 tokens. At llama.cpp's default it read 100.8 on 3,658
tokens, and at 4096 the 656.5 above (2026-09-19). The 13 September
prose reply is unaffected by that setting. With a 900-token limit the
same request had spent every token on hidden reasoning and returned
no answer. |
| 19, MiniMax M3 | The file it served until 21 September was an earlier conversion
that, per our session notes, lacked the tensors this model's sparse
attention needs; the launch check on the file it serves now reports
the sparse attention engaged. At 262,144 the new file could not
allocate its compute buffer at any micro-batch from 512 to 2048, so it
serves 131,072; at 196,608 it loads only at -ub 512, and
read 58,307 tokens at 70.16 a second, 3 of 3 (2026-09-20). It has
never run through the framework. |
| 20, Laguna S 2.1 | Until 21 September its script set -b 512 -ub 128,
and at its old 32,768 window it read 72.2 tokens a second on 24,071
tokens; at -b/-ub 8192 it read 1,434.0 on the same
prompt. Above 32,768 its script now uses -b 8192 -ub
4096, the value measured at the larger windows. |
| Not listed | Kimi K3 (retired 15 September) and Kimi K2.7-Code (weights removed 17 September) are no longer endpoints; K3's August record is its own field card, and K2.7's one September measurement is on the two-machine study. |
What a batch setting did, and did not, change
The largest movements in this table are reading speeds on models whose experts sit in system memory. Each read is about 48,000 tokens, the same prompt at both settings, except GLM-5.3-Flash's default-setting read of 24,008 tokens; all six at a 131,072-token window, measured on 21 September through each model's own start script (Qwen3.5-122B, Ornith and GLM-5.3-Flash have served 262,144 since 26 September; rows 6, 8 and 9):
| Model | Tokens read | At the defaults | At the setting it now runs | Setting, at 131,072 |
|---|---|---|---|---|
| Qwen3.5-397B-A17B | 48,029 | 224.5 | 932.2 | -b/-ub 4096 |
| Qwen3.5-122B-A10B | 48,027 | 520.2 | 2,101.7 | -b/-ub 4096 |
| Mistral Small 4 119B | 48,697 | 481.3 | 2,192.3 | -b/-ub 8192 |
| Ornith-1.5-35B-A3B | 48,027 | 1,860.2 | 6,283.3 | -b/-ub 4096 |
| Qwen3-235B-A22B | 48,020 | 285.8 | 703.3 | -b 4096 -ub 2048 |
| GLM-5.3-Flash | 24,008 / 48,168 | 80.7 (on 24,008) | 425.1 (on 48,168) | -b/-ub 4096 |
All measured, 2026-09-21;
in these six, the planted code came back in the answer at every setting
that loaded. The speaking figures in those runs moved by 6.2 percent at
most, Qwen3.5-122B's 30.81 to 32.72 (arithmetic on the same files). At
262,144 the start scripts of Qwen3.5-397B, Qwen3.5-122B and Mistral
Small 4 keep -b 2048 -ub 512: the larger buffer does not
fit there; since 26 September Ornith's has used -b 2048 -ub
2048 there and GLM-5.3-Flash's -b 4096 -ub 1024. The models that fit whole on the card kept the defaults:
Qwen3.8-27B, GLM-4.7-Flash, the two Gemmas and Muse Glimmer. On each
one's best rung, reading moved from a 0.6 percent loss (Muse Glimmer)
to a 34.5 percent gain (Gemma-4-26B-A4B at -ub 2048, which
spoke 6 percent slower), arithmetic on each model's rungs. Qwen3.8-27B's
rungs were run at 131,072, its installed window on 21 September. Qwen3.6-27B
kept them too; its rungs were run on its earlier Q4 file at 131,072.
GLM-4.7-Flash at -ub 2048 put the code only in its hidden
reasoning and returned an empty answer.
Three rows moved for other reasons: MiniMax M2.7 and Laguna S 2.1 had
inherited -ub 128, Laguna with -b 512 (row notes 18 and 20), and MiniMax M3
changed file. DeepSeek's two services moved further still, and the
two-machine study sets them side by side.
How to run the sweep on your own models, and the one value that passed
every read and still crashed a server on a specific input, are on the
batch-size study.
What changed since 15 September
This list was first compiled on 15 September. The later runs corrected figures in our own summary table and in that first list. We publish those corrections the same way we publish everything else.
| Our records said | What the logs show | This page |
|---|---|---|
| Qwen3-235B speaks at 8.2 tokens a second (21 September) | That was an 11-token answer. On letters of about 350 words it speaks 5.51, 4.68 and 4.22 at 20,063, 60,180 and 99,864 tokens | The letters (row 5) |
| Qwen3.5-397B is the fastest-reading large model, at 385 | By our own records, the 385 was a 2,617-token prompt at a 64K window, measured in August. At the served window the batch-size sweep reads 224.5 at the defaults and 932.2 at -b/-ub 4096 on 48,029 tokens, and Qwen3-235B reads faster than it at depth on the framework's 96K leg (215.5 against 195.2, estimates) | The measured figures, with their prompt sizes (row 4) |
| GLM-4.7 Full passed the framework check | Its record's verdict is a failure: the tool call at 96,000 tokens failed | A failed check (row 13); the 15 September list already printed it as partial |
| MiniMax M3 reads at 194 tokens a second at depth | 194 is 58,307 tokens divided by the whole request's 300.2 seconds, answer included. The server's own timing is 208.11 | 208.11 (row 19) |
| MiniMax M2.7 reads 45.6 on a 4,000-token prompt | The probe was named for about 4,000 tokens and sent 6,776, at an inherited -ub 128 | 6,776 tokens, the setting, and the later figures (row 18) |
| Laguna S 2.1 reads at 72 tokens a second | Its script had hard-coded -b 512 -ub 128; at -b/-ub 8192 the same read ran at 1,434.0, both at the old 32,768 window. At the 262,144 window it now serves it reads 898.8 on 24,013 tokens at -b 8192 -ub 4096 | The window and setting it now serves (row 20) |
| Qwen3.8-Flash-Next reads at 511.7 with its current setting, and at the window it serves had only a 45.4 on a 3,041-token read | The 511.7 was a 48,069-token read at a 131,072 window. The 45.4 was the first request after a start (load 65 s). At the 262,144 window it serves, with -b 4096 -ub 2048, after the server's first request and with 57.7 to 60.1 GB of the 90 GB file in memory, it reads 517.6 on 3,019 tokens, 660.6 on 48,075 and 535.4 on 229,981, against 255.7, 240.2 and 205.5 at llama.cpp's defaults (26 September) | The served-window figures (row 11) |
| The two-machine DeepSeek speaks at 15.9 to 17.1 | That band is 13 September at a 131,072 window. At a 262,144 window on 21 September: 14.73 on a 400-token reply at 2,998 tokens and 12.05 at 48,024 (14.76 and 12.04 on the short code answers) | The 21 September figures (row 1) |
| The 15 September list: 22 endpoints, all at 128K, seven listed without speeds | Three added, two retired; ten rows now serve more than 131,072 by default, each shown by a run that passed no window (202,752 or 262,144, and 196,608 for MiniMax M2.7); Laguna S 2.1 serves 262,144 by its start script's own default; Qwen3.8-27B was installed at 262,144 on 26 September; two more loaded 262,144 when it was passed; two rows without a September figure | 23 rows, each with its served window |
What held up, in the same breath: every read behind this page's
figures that asked for a planted code returned it in its answer, except
GLM-4.7-Flash at -ub 2048, whose code stayed in its hidden
reasoning (section 04). The letter runs' edit-and-reread requests,
capped at 40 tokens of reply, did not return the code at any depth (row
notes 5, 7 and 10).
What this does not show
- No measure of quality. Speed, window, recall of a planted code and whether a tool call came back are the only things measured. A model can be the fastest row here and the wrong choice for your work.
- Short answers are not prose. Where a row's only speaking figure at depth is a short answer, it says so; eight rows have prose at depth (1, 5, 6, 7, 8, 9, 10 and 18), and most have a paragraph at a short prompt. A thinking model's short answer counts its reasoning.
- Figures from different exercises are not a ranking. Rows were measured on different days, at different depths and settings, and each cell says which. Compare within a column and a date.
- One desktop. These are honest samples on one machine, not benchmarks, and not comparable with a vendor's published throughput or with a hosted service.
- Windows above the served one are not claimed here.
"On request" means the start script offers a larger window and a
September run loaded it: Qwen3.5-397B and Mistral Small 4 read about
231,000 tokens at 262,144 (15 September, at
-b 2048 -ub 512, the setting their scripts keep there), Qwen3.8-Flash-Next read 344,981 and 459,911 tokens at its two larger windows (20 September, at llama.cpp's default batch sizes, before its batch change), and MiniMax M3 58,307 at 196,608 (20 September,-ub 512). What a full window costs is on its own page. - The one-at-a-time constraint is this machine's. On more hardware several of these would run side by side.
- Two rows have no figure for what they serve. That is a statement about our September, not about the models.
Related pages on this site
- Method: the labels used on this page.
- The batch-size study: the sweep behind most reading figures here, and how to tune your own.
- Two boxes, one model: rows 1 and 3, and the models run across two machines.
- What a full window costs: the long reads behind the window column.
- DeepSeek V4 Flash, Inkling-Small, Qwen3.8-27B, Qwen3.8-27B at 256K, the Qwen3.5 pair, Mistral Medium 3.5, Muse Glimmer 30B and Kimi K3: field cards for models on or formerly on this list, each with its own dates.
- Which Qwen should I run?: the family guide behind seven of the rows.
- Empty Answer: why a thinking model's short answer can come back empty, which is why the counts here include its reasoning.
Sources and artifacts
The data package holds the redacted primary files these figures come from, a few clearly identified extracts, one rebuilt machine-readable table, and the vendor documents cited at pinned revisions. Machine paths, addresses, port numbers, service names and our internal keys for each model are removed; so is a column naming private work. Nothing that carries a measurement is altered, and no figure is recomputed except where a file says it is a derivation.
- data/README.md: every file, where it came from, what was removed from it, and the places where the package looks like it argues with itself.
- data/NUMBERS.md: every figure on this page mapped to its file and field.
- data/ROSTER_23.csv: the table, machine readable, with the file and date behind each figure.
- The runs of 19 to 26 September: the batch-size sweep's result files and logs, the long reads, the paragraph replies, the letters at depth, the MiniMax records, the new framework record, and the 26 September runs of Qwen3.8-Flash-Next, Qwen3.8-27B, Qwen3.5-122B-A10B and Ornith-1.5-35B-A3B and GLM-5.3-Flash at 262,144 and of MiniMax M2.7 at 196,608, and that evening's starts with no window passed.
- The runs of 12 to 15 September: the direct sweep, the framework check's ledgers and per-model records, the refusal-guard test, and what the model files say.
If a number on this page disagrees with a file in its package, the file is right and the page is wrong. Tell us at hello@graphometer.ai and the correction goes on the page with its date.
| 2026-09-15 | List first compiled: 22 endpoints, runs of 12 to 15 September. |
|---|---|
| 2026-09-26 | Rebuilt as this snapshot: 23 endpoints, runs of 12 to 26 September, with the corrections in section 05. |
| 2026-09-26, update | Qwen3.8-Flash-Next read at its served 262,144 window at both batch settings, and Qwen3.8-27B installed at 262,144 with its figures there: rows 7 and 11, their notes, and sections 01, 02, 04, 05 and 06, with the runs' files in the package. |
| 2026-09-26, evening | Qwen3.5-122B-A10B, Ornith-1.5-35B-A3B and GLM-5.3-Flash installed at 262,144, and MiniMax M2.7 at 196,608, with their figures there and a start of each with no window passed: rows 6, 8, 9 and 18, their notes, and sections 01, 02, 04, 05 and 06, with the runs' files in the package. |
This page is dated on purpose. It is a snapshot as of 26 September 2026, not a living list; a later roster will be a later page with a later date.