Graphometer
Field cards · published 2026-08-16

Field cards: models, measured

A field card is a measured record of a model actually running on this project's bench: one desktop, dated runs, evidence kept. Every number on a card comes from those recorded runs or from the model files on our disk, and where a fact is a vendor's claim instead, the card says so in place. That makes a field card a different document from a vendor's model card: the model card describes what a model is built to do, and a field card records what one copy of it measurably did here.

Independent measurements. Not affiliated with, endorsed by, or connected to any model maker named on this page.
01

Side by side

One row per measured model, ordered small to large by total parameters. Every cell repeats a number its card already publishes, with that card's own qualifiers; nothing in this table is a new measurement, and the rows are not a benchmark of one model against another.

ModelWhat it isMeasured decode, tokens per secondMeasured
LFM2.5-2.6BA single network, 2.69B per the card489.8 median on the recommended Q4_K_M file2026-08-04, updated 2026-08-07
Qwen3.8-27BDense, 27B72.4 prose, 118.7 structured, gated draft head, thinking off2026-08-14, release day
Muse Glimmer 30BDense, 30B, with a separate block-16 DFlash drafter76.22 prose and 76.345 structured with no speculation, 114.86 and 253.19 once the drafter is asked for its own block size2026-08-16
Qwen3.5-122BMixture of experts, 122B total, about 9.8B active per token26 to 38, depending on prompt shape2026-08-05 to 08-06
Mistral Medium 3.5Dense, 128B class, 125.03B servedabout 1 alone, 2 to 5 with the vocabulary-repaired draft2026-08-06
Qwen3.5-397BMixture of experts, 397B total, about 17.4B active per token14 to 20, depending on prompt shape2026-08-05 to 08-06
Kimi K3Mixture of experts, 2.78T total, about 104B active per token0.18 to 0.28 (about 0.2)2026-08-01 to 08-05

Each speed is the card's own headline number and keeps the card's meaning: one configuration on one machine, a field measurement rather than a benchmark. The LFM2.5-2.6B figure is an end-to-end median on that card's recommended download, a run diagnostic in the card's own words. The Qwen3.8-27B figures were taken with thinking off on release day; the Kimi K3 range is the current quant generation. Where a number needs more qualification than a cell can carry, the card carries it; follow the links for the full conditions.

02

The cards

  • GLM-5.3-Flash on one desktop · measured 16 to 26 September 2026. UD-Q3_K_XL through a pinned llama.cpp fork: at 131,072 tokens on 21 September, the published pair is 80.7 t/s on a 24,008-token first request, with 62 generated tokens at 9.22 t/s, versus 446.0 t/s on a 20,333-token third request, with 60 generated tokens at 9.35 t/s. The fourth tuned request read 48,168 tokens at 425.1 t/s and generated 310 tokens at 9.12 t/s. Prompt lengths differ, and the guarded output fixture passed only in isolation.
  • Qwen3.8-27B at 256K · measured 26 September 2026. A 262,144-token window on one RTX 5090, one request at a time: the installed Q5 file without its draft head, or a smaller Q4 file with it. Prose and tool-shaped speeds, 3 of 3 long-read recall, and a measured speed and fidelity trade across three files.
  • DeepSeek V4 Flash 0731, on one desktop and on two · measured 12 to 21 September 2026. With its batch settings tuned, one desktop answered a 48,024-token request in 75.2 seconds. The same file split across two machines took 416.4 seconds at its best completed setting and 474.2 at the one it keeps.
  • Inkling-Small at 12 tokens a second, from a 32-token prompt to 95,041 · measured 14 to 21 September 2026. Thinking Machines Lab's 276-billion-parameter, 12-billion-active open model (vendor), served text-only on one RTX 5090. Speaking stayed between 11.6 and 12.2 tokens a second from a 32-token prompt to a 95,041-token prompt (measured); a later batch-size change raised reading on a 48,115-token prompt from 123.0 to 337.9 tokens a second, and a full 262,144-token window read 230,827 tokens and returned all three planted codes.
  • Muse Glimmer 30B: trained at 16, asked for 3 · measured 2026-08-16. Meta trained its DFlash drafter to draft a block of 16 tokens per pass and prints that number; llama.cpp's generic draft length is 3; the two flags Meta's GGUF card documents leave that default in charge, and on six structured prompts the gap is worth 2.17 times the throughput. Verbatim Apache 2.0, with a separate usage policy beside it.
  • Qwen3.8-27B: day zero · released and measured 2026-08-14. One model family with two licenses, and a built-in draft head decoding prose at 72.4 tokens a second gated, 65.9 with no speculation at all, 44.5 with the gate removed.
  • Field card: Mistral Medium 3.5 · runs 2026-08-06. A dense 128B-class model made genuinely usable on one desktop: about 1 token per second alone, 2 to 5 once a two-token vocabulary repair let the draft model load.
  • Field card: Qwen3.5 (122B and 397B) · runs 2026-08-05 to 08-06. Two models into service on one desktop by the same method, and the trap next to their built-in guessing head: run it without its confidence gate and creative text goes slower than no speculation at all.
  • Field card: LFM2.5-2.6B · measured 2026-08-04, updated 2026-08-07. The context window as measured, which file to download, how big an output budget it really needs, and which sampling settings are actually in effect.
  • Field card: Kimi K3 · runs 2026-08-01 to 2026-08-05. The complete model, all 93 blocks and all 896 experts, generating at about 0.2 tokens per second: one careful question answered overnight from hardware you own.
03

What gets a model on this page

Three things, and only these: the model ran on this bench, it was measured while it ran, and the run was recorded with its evidence kept. No entry comes from a spec sheet, a leaderboard, or another site's numbers. What a number must have behind it before it is published here is written down on the method page.