Studies and reports
Everything Graphometer has published that measures, verifies, or guides, dated, newest first, each with its finding in the page's own words. The per-model records are the field cards, which have an index of their own.
Kindness doesn't cost
Give AI agents a way to ask, object, pause, decline or step away, and honor it. In our pilot the ordinary work got done just the same, as far as we could measure, and a paragraph of permission went with agents raising the real problem first more often. The small protocol and the paragraph are on the page. What the pilot could and could not show is there too, with its data.
Describing pictures locally: six Qwen vision models on one desktop, on the card and on the CPU
Six Qwen vision models through Ollama on one RTX 5090 desktop. On the card, Qwen3-VL 30B-A3B found all 39 known strings at 1.3 to 3.2 seconds a warm picture. On the CPU, the server ran 4 of 24 threads unless told otherwise, and its first 4K picture took 273 seconds; 16 threads took 111, and a 1,536-pixel cap brought it to 38.5, at a cost on small text.
Silent settings: twelve settings that cost speed or correctness on one desktop without an error, and the one-line check for each
A checklist from our own records: settings a local model server accepted as tested while the speed or the number was wrong anyway, eight with a kept load and six of those with a request that returned text, two failing only at the full window or the draft's load where nobody had tried them, one with no run at its window in the records (the shell check exits 0 and the launch line gains no cache flags), and one whose two misreads are a dated note with its logs not kept. A served window that is not the one you passed; a capped window the card still pays for; a micro-batch that loads at one window and fails at another; an inherited -ub 128; a first read that pages the file in; an idle card; a model file without its sparse-attention index; two start-script bugs; a stale header; a help flag; an Ollama thread default. The symptom, the one-line check, the cost and the record for each.
Bytes per token: Predicting Inkling 975B: 5.2 to 5.7 tokens per second by arithmetic, server-reported 3.43 on short prose
Count the routed expert bytes in the actual quantization, then divide an effective expert throughput by that budget. Recorded short-prompt prose on this desktop implies an effective expert throughput of 31.0 to 34.3 GB/s, an arithmetic proxy rather than a direct RAM measurement. Inkling 975B's arithmetic budget is 6.01 GB of routed expert weights per token; its short-prose sample had a server-reported decode rate of 3.43 tokens per second alone; a different short-prose sample on the pair decoded at 3.13.
llama.cpp's batch size: prompt reading 2.5 to 7.3 times faster than the default, up to 20 against an inherited -ub 128
Swept on one RTX 5090 desktop: on twelve models whose experts sit in system RAM, long prompts read 2.5 to 7.3 times faster than at llama.cpp's default, and raising an inherited -ub 128 made one read 19.9 times faster. Each model had its own ceiling, and one input crashed a setting that had passed everything else.
Six large models on the ASUS ROG Flow Z13 alone
At 48K the laptop read prompts 2.8 to 23 times slower than the desktop. Ling-3.0-flash matched desktop short-answer speaking after a 3.6-minute read. The run used Balanced mode on mains.
256K on a 32 GB card: the cost of reading
Ten configurations of nine models read prompts of 150,000 to 231,000 tokens at full windows and returned every planted code. Reading took from 60.8 seconds to 109.7 minutes; most configurations kept expert weights in system RAM.
Five local coding models, four points apart
Five open-weight models ran four agentic coding tasks through OpenCode and scored 387 to 391 of 400. Every first scored run took all 60 mechanical points; the graded half was noisy and steered by instructions that put most competent work between 30 and 37 of 40. One re-run of one task moved 37 points. With the harness traps behind it.
Two boxes, one model: splitting one llama.cpp model across a desktop and a laptop, measured
A desktop and a laptop joined by one cable, running one large model between them. The pair adds memory. With llama.cpp's batch size raised, the desktop alone answered a 48,024-token request with the same DeepSeek V4 Flash file in 75.2 seconds; the pair took 416.4 at its best completed setting and 474.2 at the one it keeps. The pair still speaks faster on short prompts.
September field notes: eleven measurements on one RTX 5090 desktop
Eleven measurements from one desktop, led by a correction of our own record: MiniMax M2.7's 96,000-token leg did not fail, it timed out against a launch flag we had inherited without noticing. With the flag fixed, it found all three planted codes at 103,931 tokens through the same agent framework in 251.2 seconds. Also MiniMax M3 on a file older than llama.cpp's support for its attention, Ling's crash input, and what Flash-Next held past its window.
The roster, as of 26 September 2026
Twenty-three local model endpoints on one RTX 5090 desktop, one running at a time. Every reading speed with its prompt size and batch setting, every speaking speed with the reply it timed, and the corrections since 15 September.
Seven minutes to 1.7 seconds
Saving a local llama.cpp server's conversation state to disk. After a restart, Qwen3-235B reached first content in 1.65 seconds on a 102,912-token history that took 461.97 seconds to read cold. Cold reading has since got faster; those restart times were not re-run. A separate MiniMax M2.7 test restored 85,778 saved tokens.
GPT-4o, measured against itself
GPT-4o re-run at temperature zero reads 1.17 on 81 prompts, set by markup, with no family separating. The nearest other model reads 2.23 on 178 prompts, set by shape. A ChatGPT-style prompt of our own moves GPT-4o 3.05 on the same 178, further than the nearest other model sits from it. The nearest was Mistral Medium 3.5 running locally; counted text features only, no judge, no lineage claim.
Reasoning effort is a sentence, not a dial: measured on one Qwen3.8-27B template
On the chat template shipped inside the Qwen3.8-27B GGUF we served, the reasoning-effort setting is not a sampler dial: the value you send is looked up in that template and a matching instruction is pasted into your prompt as a system message. xhigh adds 42 tokens, low adds 30, medium adds nothing at all, and leaving the field out renders byte-identical to xhigh.
Release watch: Qwen3.8
Both promised open releases landed and were verified at pinned revisions, and for practitioners the headline is the license: the Max-class weights carry a custom license with revenue thresholds, while the 27B carries a genuine Apache 2.0 file.
The wrong flag: a 675B decode gap traced to a default
Six days of records blamed warmup for a 3.6 versus 6 tokens per second decode gap on Mistral Large 3 675B; a three-arm A/B traced it to llama.cpp's default thread count instead.
Slower by default: llama.cpp's speculation gate, measured
llama.cpp's shipped default made prose generation 19.8% slower than running with no speculation at all, in the same run in which it made structured output 22.5% faster.
Empty Answer: the thinking-budget trap, measured
A reasoning model can spend your entire output budget thinking and hand back nothing at all: HTTP 200, no error, an empty answer, and an eight-fold gap in the budget floor between two specific deployments.
Which Qwen should I run?
We run the family for real on one desktop, from the small daily drivers to the 122B and 397B, so the guidance is what we measured, labeled as measured, next to what the vendor claims, labeled as claims.
Looking for a specific model?
The per-model measured records live on the field cards index, and the free, open instruments live on the tools page.