Field cards: models, measured
A field card is a measured record of a model actually running on this project's bench: one desktop, dated runs, evidence kept. Every number on a card comes from those recorded runs or from the model files on our disk, and where a fact is a vendor's claim instead, the card says so in place. That makes a field card a different document from a vendor's model card: the model card describes what a model is built to do, and a field card records what one copy of it measurably did here.
Side by side
One row per measured model, ordered small to large by total parameters. Every cell repeats a number its card already publishes, with that card's own qualifiers; nothing in this table is a new measurement, and the rows are not a benchmark of one model against another.
| Model | What it is | Measured decode, tokens per second | Measured |
|---|---|---|---|
| LFM2.5-2.6B | A single network, 2.69B per the card | 489.8 median on the recommended Q4_K_M file | 2026-08-04, updated 2026-08-07 |
| Qwen3.8-27B | Dense, 27B | 72.4 prose, 118.7 structured, gated draft head, thinking off | 2026-08-14, release day |
| Muse Glimmer 30B | Dense, 30B, with a separate block-16 DFlash drafter | 76.22 prose and 76.345 structured with no speculation, 114.86 and 253.19 once the drafter is asked for its own block size | 2026-08-16 |
| Qwen3.5-122B | Mixture of experts, 122B total, about 9.8B active per token | 26 to 38, depending on prompt shape | 2026-08-05 to 08-06 |
| Mistral Medium 3.5 | Dense, 128B class, 125.03B served | about 1 alone, 2 to 5 with the vocabulary-repaired draft | 2026-08-06 |
| Qwen3.5-397B | Mixture of experts, 397B total, about 17.4B active per token | 14 to 20, depending on prompt shape | 2026-08-05 to 08-06 |
| Kimi K3 | Mixture of experts, 2.78T total, about 104B active per token | 0.18 to 0.28 (about 0.2) | 2026-08-01 to 08-05 |
Each speed is the card's own headline number and keeps the card's meaning: one configuration on one machine, a field measurement rather than a benchmark. The LFM2.5-2.6B figure is an end-to-end median on that card's recommended download, a run diagnostic in the card's own words. The Qwen3.8-27B figures were taken with thinking off on release day; the Kimi K3 range is the current quant generation. Where a number needs more qualification than a cell can carry, the card carries it; follow the links for the full conditions.
The cards
- GLM-5.3-Flash on one desktop · measured 16 to 26 September 2026. UD-Q3_K_XL through a pinned llama.cpp fork: at 131,072 tokens on 21 September, the published pair is 80.7 t/s on a 24,008-token first request, with 62 generated tokens at 9.22 t/s, versus 446.0 t/s on a 20,333-token third request, with 60 generated tokens at 9.35 t/s. The fourth tuned request read 48,168 tokens at 425.1 t/s and generated 310 tokens at 9.12 t/s. Prompt lengths differ, and the guarded output fixture passed only in isolation.
- Qwen3.8-27B at 256K · measured 26 September 2026. A 262,144-token window on one RTX 5090, one request at a time: the installed Q5 file without its draft head, or a smaller Q4 file with it. Prose and tool-shaped speeds, 3 of 3 long-read recall, and a measured speed and fidelity trade across three files.
- DeepSeek V4 Flash 0731, on one desktop and on two · measured 12 to 21 September 2026. With its batch settings tuned, one desktop answered a 48,024-token request in 75.2 seconds. The same file split across two machines took 416.4 seconds at its best completed setting and 474.2 at the one it keeps.
- Inkling-Small at 12 tokens a second, from a 32-token prompt to 95,041 · measured 14 to 21 September 2026. Thinking Machines Lab's 276-billion-parameter, 12-billion-active open model (vendor), served text-only on one RTX 5090. Speaking stayed between 11.6 and 12.2 tokens a second from a 32-token prompt to a 95,041-token prompt (measured); a later batch-size change raised reading on a 48,115-token prompt from 123.0 to 337.9 tokens a second, and a full 262,144-token window read 230,827 tokens and returned all three planted codes.
- Muse Glimmer 30B: trained at 16, asked for 3 · measured 2026-08-16. Meta trained its DFlash drafter to draft a block of 16 tokens per pass and prints that number; llama.cpp's generic draft length is 3; the two flags Meta's GGUF card documents leave that default in charge, and on six structured prompts the gap is worth 2.17 times the throughput. Verbatim Apache 2.0, with a separate usage policy beside it.
- Qwen3.8-27B: day zero · released and measured 2026-08-14. One model family with two licenses, and a built-in draft head decoding prose at 72.4 tokens a second gated, 65.9 with no speculation at all, 44.5 with the gate removed.
- Field card: Mistral Medium 3.5 · runs 2026-08-06. A dense 128B-class model made genuinely usable on one desktop: about 1 token per second alone, 2 to 5 once a two-token vocabulary repair let the draft model load.
- Field card: Qwen3.5 (122B and 397B) · runs 2026-08-05 to 08-06. Two models into service on one desktop by the same method, and the trap next to their built-in guessing head: run it without its confidence gate and creative text goes slower than no speculation at all.
- Field card: LFM2.5-2.6B · measured 2026-08-04, updated 2026-08-07. The context window as measured, which file to download, how big an output budget it really needs, and which sampling settings are actually in effect.
- Field card: Kimi K3 · runs 2026-08-01 to 2026-08-05. The complete model, all 93 blocks and all 896 experts, generating at about 0.2 tokens per second: one careful question answered overnight from hardware you own.
What gets a model on this page
Three things, and only these: the model ran on this bench, it was measured while it ran, and the run was recorded with its evidence kept. No entry comes from a spec sheet, a leaderboard, or another site's numbers. What a number must have behind it before it is published here is written down on the method page.