Graphometer
Index · Published 2026-08-16

Studies and reports

Everything Graphometer has published that measures, verifies, or guides, dated, newest first, each with its finding in the page's own words. The per-model records are the field cards, which have an index of their own.

01
Proposal · Agency Layer Pilot I, 30 August to 3 September 2026 · study paused

Kindness doesn't cost

Give AI agents a way to ask, object, pause, decline or step away, and honor it. In our pilot the ordinary work got done just the same, as far as we could measure, and a paragraph of permission went with agents raising the real problem first more often. The small protocol and the paragraph are on the page. What the pilot could and could not show is there too, with its data.

Read the proposal

02
Measured study · 20 September 2026 · one RTX 5090 desktop

Describing pictures locally: six Qwen vision models on one desktop, on the card and on the CPU

Six Qwen vision models through Ollama on one RTX 5090 desktop. On the card, Qwen3-VL 30B-A3B found all 39 known strings at 1.3 to 3.2 seconds a warm picture. On the CPU, the server ran 4 of 24 threads unless told otherwise, and its first 4K picture took 273 seconds; 16 threads took 111, and a 1,536-pixel cap brought it to 38.5, at a cost on small text.

Read the study

03
Guide · twelve items from dated records, 12 to 26 September 2026 · one RTX 5090 desktop

Silent settings: twelve settings that cost speed or correctness on one desktop without an error, and the one-line check for each

A checklist from our own records: settings a local model server accepted as tested while the speed or the number was wrong anyway, eight with a kept load and six of those with a request that returned text, two failing only at the full window or the draft's load where nobody had tried them, one with no run at its window in the records (the shell check exits 0 and the launch line gains no cache flags), and one whose two misreads are a dated note with its logs not kept. A served window that is not the one you passed; a capped window the card still pays for; a micro-batch that loads at one window and fails at another; an inherited -ub 128; a first read that pages the file in; an idle card; a model file without its sparse-attention index; two start-script bugs; a stale header; a help flag; an Ollama thread default. The symptom, the one-line check, the cost and the record for each.

Read the guide

04
Guide · arithmetic with measured anchors · runs of 15 September 2026

Bytes per token: Predicting Inkling 975B: 5.2 to 5.7 tokens per second by arithmetic, server-reported 3.43 on short prose

Count the routed expert bytes in the actual quantization, then divide an effective expert throughput by that budget. Recorded short-prompt prose on this desktop implies an effective expert throughput of 31.0 to 34.3 GB/s, an arithmetic proxy rather than a direct RAM measurement. Inkling 975B's arithmetic budget is 6.01 GB of routed expert weights per token; its short-prose sample had a server-reported decode rate of 3.43 tokens per second alone; a different short-prose sample on the pair decoded at 3.13.

Read the guide

05
Measured study · 19 to 21 September 2026 · one RTX 5090 desktop

llama.cpp's batch size: prompt reading 2.5 to 7.3 times faster than the default, up to 20 against an inherited -ub 128

Swept on one RTX 5090 desktop: on twelve models whose experts sit in system RAM, long prompts read 2.5 to 7.3 times faster than at llama.cpp's default, and raising an inherited -ub 128 made one read 19.9 times faster. Each model had its own ceiling, and one input crashed a setting that had passed everything else.

Read the study

06
Measured study · 20 to 21 September 2026 · a 128 GB laptop by itself

Six large models on the ASUS ROG Flow Z13 alone

At 48K the laptop read prompts 2.8 to 23 times slower than the desktop. Ling-3.0-flash matched desktop short-answer speaking after a 3.6-minute read. The run used Balanced mode on mains.

Read the study

07
Measured study · 15 to 21 September 2026 · one RTX 5090 desktop, one run split with a laptop

256K on a 32 GB card: the cost of reading

Ten configurations of nine models read prompts of 150,000 to 231,000 tokens at full windows and returned every planted code. Reading took from 60.8 seconds to 109.7 minutes; most configurations kept expert weights in system RAM.

Read the study

08
Measured study · runs 16 September 2026 · OpenCode 1.18.31

Five local coding models, four points apart

Five open-weight models ran four agentic coding tasks through OpenCode and scored 387 to 391 of 400. Every first scored run took all 60 mechanical points; the graded half was noisy and steered by instructions that put most competent work between 30 and 37 of 40. One re-run of one task moved 37 points. With the harness traps behind it.

Read the study

09
Measured study · 13 to 21 September 2026 · one desktop and a 128 GB laptop over Thunderbolt

Two boxes, one model: splitting one llama.cpp model across a desktop and a laptop, measured

A desktop and a laptop joined by one cable, running one large model between them. The pair adds memory. With llama.cpp's batch size raised, the desktop alone answered a 48,024-token request with the same DeepSeek V4 Flash file in 75.2 seconds; the pair took 416.4 at its best completed setting and 474.2 at the one it keeps. The pair still speaks faster on short prompts.

Read the study

10
Field notes · 12 to 21 September 2026 · one RTX 5090 desktop

September field notes: eleven measurements on one RTX 5090 desktop

Eleven measurements from one desktop, led by a correction of our own record: MiniMax M2.7's 96,000-token leg did not fail, it timed out against a launch flag we had inherited without noticing. With the flag fixed, it found all three planted codes at 103,931 tokens through the same agent framework in 251.2 seconds. Also MiniMax M3 on a file older than llama.cpp's support for its attention, Ling's crash input, and what Flash-Next held past its window.

Read the study

11
Guide · measured 2026-09-12 to 09-26 · one RTX 5090 desktop

The roster, as of 26 September 2026

Twenty-three local model endpoints on one RTX 5090 desktop, one running at a time. Every reading speed with its prompt size and batch setting, every speaking speed with the reply it timed, and the corrections since 15 September.

Read the guide

12
Measured study · 13 September 2026 · updated 26 September 2026

Seven minutes to 1.7 seconds

Saving a local llama.cpp server's conversation state to disk. After a restart, Qwen3-235B reached first content in 1.65 seconds on a 102,912-token history that took 461.97 seconds to read cold. Cold reading has since got faster; those restart times were not re-run. A separate MiniMax M2.7 test restored 85,778 saved tokens.

Read the study

13
Measured study · runs 3 to 5 September 2026 · 26 models, one battery

GPT-4o, measured against itself

GPT-4o re-run at temperature zero reads 1.17 on 81 prompts, set by markup, with no family separating. The nearest other model reads 2.23 on 178 prompts, set by shape. A ChatGPT-style prompt of our own moves GPT-4o 3.05 on the same 178, further than the nearest other model sits from it. The nearest was Mistral Medium 3.5 running locally; counted text features only, no judge, no lineage claim.

Read the study

14
Measured study · hosted leg 2026-08-13 · local leg 2026-08-16

Reasoning effort is a sentence, not a dial: measured on one Qwen3.8-27B template

On the chat template shipped inside the Qwen3.8-27B GGUF we served, the reasoning-effort setting is not a sampler dial: the value you send is looked up in that template and a matching instruction is pasted into your prompt as a system message. xhigh adds 42 tokens, low adds 30, medium adds nothing at all, and leaving the field out renders byte-identical to xhigh.

Read the study

15
Release watch · weights verified 2026-08-13 and 2026-08-14 · measured 2026-08-09, 2026-08-13 and 2026-08-14

Release watch: Qwen3.8

Both promised open releases landed and were verified at pinned revisions, and for practitioners the headline is the license: the Max-class weights carry a custom license with revenue thresholds, while the 27B carries a genuine Apache 2.0 file.

Read the release watch

16
Measured study · pre-registered three-arm A/B, 2026-08-11

The wrong flag: a 675B decode gap traced to a default

Six days of records blamed warmup for a 3.6 versus 6 tokens per second decode gap on Mistral Large 3 675B; a three-arm A/B traced it to llama.cpp's default thread count instead.

Read the study

17
Measured study · main run 2026-08-10 · lower-gate follow-up run 2026-08-11

Slower by default: llama.cpp's speculation gate, measured

llama.cpp's shipped default made prose generation 19.8% slower than running with no speculation at all, in the same run in which it made structured output 22.5% faster.

Read the study

18
Incident report · measured 2026-08-09

Empty Answer: the thinking-budget trap, measured

A reasoning model can spend your entire output budget thinking and hand back nothing at all: HTTP 200, no error, an empty answer, and an eight-fold gap in the budget floor between two specific deployments.

Read the report

19
Guide · created 2026-08-09 · Qwen3.8 row refreshed 2026-08-14

Which Qwen should I run?

We run the family for real on one desktop, from the small daily drivers to the 122B and 397B, so the guidance is what we measured, labeled as measured, next to what the vendor claims, labeled as claims.

Read the guide

20

Looking for a specific model?

The per-model measured records live on the field cards index, and the free, open instruments live on the tools page.