Graphometer
Guide · twelve items from dated records, 12 to 26 September 2026 · llama.cpp start scripts and one Ollama default · one RTX 5090 desktop

Silent settings

Twelve settings that cost speed or correctness on one desktop without an error, and the one-line check for each.

A silent setting is one the server accepted as it was tested, while the speed we lived with or the number we published was wrong anyway. For eight of the twelve items below a kept record shows the server loading under the setting (items 02, 03, 05, 06, 07, 08, 11 and 13), and for six of those a request that returned text (05, 06, 07, 08, 11 and 13). Two failed loudly, but only where nobody had tried them: a micro-batch that had loaded at one window is refused, or dies, at the model's full window (item 04), and a draft type that --help lists dies at load with a missing header key (item 12); the silent part was the setting that had passed and the help text that offered it. Item 09 has no run at the 262,144 window in the records; the shell check exits 0 and the launch line gains no cache flags. Item 10's two misreads are a dated note, and the runs' logs were not kept. Between 12 and 26 September 2026 the records on this machine caught twelve of them: in llama.cpp launch flags, in our own start scripts, in one shared library, and in one Ollama default. This page is the list. Each item gives the symptom you would see, or fail to see; the one-line check that catches it; what it cost us; and where the record is. Most of the records live on the pages that measured them, and this page links to those rather than repeating them. Nothing here is a bug report against any project: every item is an interaction between a default, a file and a script, and most of the mistakes were ours.

Independent measurements; no affiliation with the llama.cpp project, ggml-org, Ollama, MiniMax, Z.ai, DeepSeek, Alibaba Group, Poolside, Unsloth, NVIDIA, Intel, or the makers of the other models the package's tables name (Google, Mistral AI, Moonshot AI, Thinking Machines Lab, inclusionAI, Meta, Ornith AI).
01

What it is

A checklist for anyone who keeps more than a few start scripts for local models. The test a server applies to itself, the process came up and the health check answered, is in a kept record for eight of the twelve items (02, 03, 05, 06, 07, 08, 11 and 13), and a request that returned text for six of them (05, 06, 07, 08, 11 and 13). Items 04 and 12 failed that test loudly, at a window or a load nobody had tried, after a setting or a help entry had looked fine. Item 09 has no run at the 262,144 window in the records; the shell check exits 0 and the launch line gains no cache flags. Item 10's two misreads are a dated note, and the runs' logs were not kept. What was wrong showed up later: a 3,658-token read that took 17.5 times longer than the same model managed once one flag changed, a card with 8,311 MiB of its 32,607 unused while a model served, a window that was not the one we thought we served, or a number on a page that a later run could not reproduce.

The machine is one desktop: an RTX 5090 with 32,607 MiB as nvidia-smi reports it, 188 GiB of RAM and a 24-core Core Ultra 9 285K. The servers are llama.cpp's llama-server at the builds each record names, started by our own scripts, and Ollama for the last item. Every figure below is measured on a dated run unless it is marked arithmetic, or marked as a dated note whose run log was not kept; three items carry such notes, and each says so. Where another page on this site owns a figure, this page prints only enough of it to recognise the trap and links to the page and its package. The records that exist nowhere else, script excerpts, source lines, five result files and one header dump, ship in this page's package (section 18).

02

The window you pass is not the window you serve

SymptomA start script accepts a window and serves it when you pass one. What it serves when nothing is passed, the installed default, can be another number, and a figure measured by passing the window says nothing about the default. The server log prints only the window it served, so nothing looks wrong.
CheckStart once with nothing passed and read the served window back: curl -s http://127.0.0.1:PORT/props and look for n_ctx, or grep n_ctx_slot in the server log.
What it cost usOur 15 September list of endpoints put all 22 of them at 128K. By 26 September ten rows served more than 131,072 by default, each shown by a start that passed no window. On the evening of 26 September four installed scripts were started with nothing passed, 21:06 to 21:13, and served 262,144, 262,144, 196,608 and 262,144 (measured 2026-09-26).
RecordThe September roster: section 05 for the 15 September list and the ten rows, section 03 for each row's window cell, which says whether a run that passed no window shows the installed default; its package file installed-starts-2026-09-26.txt for the four starts.
03

A window past the file's training length is capped, and the card still pays for it

SymptomAsk for 524,288 on a file trained to 262,144 without the rope-scaling flags (the settings that stretch a model's position encoding past the length it was trained at). The server comes up. Its log says the slot context exceeds the training context and that it is capping; the card holds the memory of the window you asked for, not the one you got.
CheckAfter the load, compare the served window (n_ctx from /props, or n_ctx_slot in the log) with the one you asked for, and read nvidia-smi against a plain load at the capped window.
What it cost usQwen3.8-Flash-Next, 20 September: asked for 524,288 with no rope flags, it served 262,144 and peaked at 23,816 MiB, the same as a genuine 524,288 window with the flags (23,814) and 8,422 MiB more than a plain 262,144 load (15,394; the difference is arithmetic). The capped server's own record is up in 2 seconds; it was the plain 262,144 load whose probe then logged a 600-second timeout (measured 2026-09-20).
RecordThe September field notes, section 09; its package files flashnext-context/e_512k_norope.log (lines 11 and 12) and e_512k_norope.vram for the cap and the 23,816, b_512k_unlock.vram for the 23,814, and a_256k_base.vram for the 15,394 and the 600-second timeout.
04

A micro-batch that loads at one window fails at another

Symptom-ub 4096 loads and reads fast at 131,072. At the model's full window the same script either dies at load with a refused CUDA buffer, or loads and peaks within a few hundred MiB of the card's total. A script that names one batch setting for every window has usually been checked at one.
CheckLoad once at every window the script offers, and read two things: the compute-buffer allocation at load, and the peak card use during one long read at that window.
What it cost usMiniMax M2.7 read 43,909 tokens at -ub 4096 at 131,072; at its 196,608 window 4096 was refused (a 2,916.16 MiB buffer), 2048 failed too, and only 1024 loaded, at 31,560 MiB. GLM-5.3-Flash at 262,144: -b 4096 -ub 4096, its setting at 131,072, was refused (a 13,281.37 MiB compute buffer); 2048 loaded and read 230,039 tokens but peaked at 31,951 MiB, 656 below the card's total, and the script ships 1024 there. Until 26 September four other scripts returned to llama.cpp's default above 131,072, unmeasured there; that day two of them were measured at 262,144, and Ornith-1.5-35B's script now sets -b 2048 -ub 2048 there while Qwen3.5-122B's keeps 512 for the card's sake. The Qwen3.5-397B and Mistral Small 4 scripts still return to the default above 131,072, unmeasured there (measured 2026-09-19 to 2026-09-26).
RecordThe batch-size study, section 04, and its package.
05

An inherited micro-batch

SymptomA prompt reads many times slower than the same model can: here 17.5 times on a 3,658-token read and 19.9 times on a 24,071-token read (arithmetic on the rates below). At the log level our scripts use, the server logs we kept do not print the batch sizes at load, so nothing in the log says why.
CheckRead the running server's real arguments: tr '\0' ' ' < /proc/PID/cmdline | grep -o -E -- '(-ub|--ubatch-size) [0-9]+'. No match means the build's default, 512 on the builds we ran. A value copied from another script is worth one long read at 2048 and at 4096.
What it cost usLaguna S 2.1 ran --batch-size 512 --ubatch-size 128 from July to 21 September, with no reason recorded: at its 32,768 window a 24,071-token read took 333.2 seconds at 72.2 t/s, and 16.8 seconds at 1,434.0 t/s with -b 8192 -ub 8192, the code correct at both, while decode on the 12-token answer fell from 18.40 to 16.50 t/s. MiniMax M2.7 was first judged at -cmoe -ub 128, the pair of flags in the MiniMax M3 start script: with -b 4096 held, every expert layer in RAM and a 131,072 window, one 3,658-token prompt read at 32.3 t/s at -ub 128 and 564.0 at -ub 4096; and its 96,000-token check, which the record had scored “l128 recall 0/3” when the turn ended at 1,800.4 seconds, had timed out against that flag (measured 2026-09-19 and 2026-09-21).
RecordThe batch-size study, sections 03, 06 and 13; the September field notes, section 03 for the 1,800.4 seconds and the -ub 128 cause, and section 15 with its package file wave2/CORRECTION.md for the score string.
06

The first read after a start pages the file in

SymptomThe first long request after a start reads at a fraction of the model's rate. The second request of the same length reads at the full rate. No message says the file was still coming off the disk.
CheckTime the second request as well as the first. To see it directly, read the file's share of the page cache before the first request (fincore from util-linux, or vmtouch; our record used a small mincore script that ships with the batch-size study).
What it cost usQwen3.8-Flash-Next, 26 September, one start with none of its 90 GB file in memory and 43.5 GB of it there after the load: the first request, 3,035 tokens, read at 87.7 t/s while the file's share grew to 57.7 GB; the next 3,019 tokens on the same server read at 517.6. A server started with 60.05 GB already in memory read its first request at 253.7 and its second at 255.7 (measured 2026-09-26). It is the only start recorded with the file out of memory, so it shows what this start did, not a rule.
RecordThe batch-size study, section 04; the 256K study, section 04.
07

Every expert in system memory leaves the card idle

SymptomA mixture-of-experts model served with -cmoe, every expert layer in system RAM, shows a card with room to spare after load. Nothing suggests moving a layer or two back.
Checknvidia-smi --query-gpu=memory.used,memory.total --format=csv after load, then a short ladder of --n-cpu-moe values with one prompt, reading card use and decode at each step. On this model the first step below the layer count bought nothing and the third did, so one step is not a test.
What it cost usMiniMax M2.7, 19 September (Unsloth's UD-IQ4_XS file, llama.cpp build 10919, a 131,072 window, an 8-bit cache, -ub 4096, one 43,909-token prompt with a 32-token answer, one run per placement): with every expert layer in RAM the card held 24,296 MiB after load, 8,311 below its 32,607 total (arithmetic), and the answer decoded at 9.04 t/s. The first step back, the experts of 61 of its 62 layers in RAM, was slower on both counts (8.97 t/s, and 636.2 t/s reading against 640.2); 60 gained 0.23 t/s; the speed moved at 59 (29,094 MiB, 9.83 t/s) and 58 (30,695 MiB, 10.02 t/s), with reading at 656.5 and 666.1 t/s (measured 2026-09-19; all five rows are below). The script's default at 131,072 is 59; at the 196,608 window it now serves by default it forces every expert into RAM with -ub 1024, the one shape that loads there, and a start with no window passed on 26 September held 31,512 MiB at load. These are 32-token answers, not prose speed; the whole gain is under one token a second; and the only repeat in each file is the same prompt sent again, a cache hit, whose decode moved by up to 0.34 t/s (58 layers: 10.02, then 9.68), more than the 0.19 step from 59 to 58. A 17 September pass that changed the placement in four other scripts kept no per-arm logs, so its gains are not printed here.
RecordThis page's package: records/minimax-m2.7-placement-2026-09-19/, the five result files and placement.tsv; excerpts/minimax-m2.7-window-and-placement.txt for the script's two windows; the September roster's installed-starts-2026-09-26.txt for the 31,512.
Experts in RAMCard after load, MiBSpare of 32,607Reading, t/sDecode, first requestDecode, repeat
all 62 layers (-cmoe)24,2968,311640.29.048.95
61 of 6225,8846,723636.28.979.08
60 of 6227,4835,124644.99.279.48
59 of 62 (the script's 131,072 default)29,0943,513656.59.839.77
58 of 6230,6951,912666.110.029.68

One run per row, 2026-09-19: one 43,909-token prompt, a 32-token answer, then the same prompt again (the repeat, a cache hit). Card memory is read after load, not during the read; the spare is arithmetic against the card's 32,607 MiB. Decode on 32 tokens is not prose speed.

08

Sparse attention that runs dense

SymptomMiniMax M3's sparse attention (MSA) runs as dense attention unless flash attention is on and the cache has one sequence per stream, which means one slot, or a cache that is not unified. The source's own warning says output may be degraded; the only sign is that one warning line at load, which says running DENSE attention.
Checkgrep -E 'running DENSE|MSA requires' server.log after every load; keep --flash-attn on and --parallel 1.
What it cost usThis door has not cost us a number: the harness that measured M3 greps its log for those strings before taking one, and the installed script sets --flash-attn on --parallel 1. The door beside it did. The first M3 file here predated llama.cpp's support for this attention, with zero indexer tensors among its 948, so every M3 number we had before the correct file arrived on 20 September was a dense fallback (measured 2026-09-19).
Recordllama.cpp src/models/minimax-m3.cpp at commit d3146f2b5 (build 10919), lines 228 to 244, excerpted in this page's package and checked against the upstream file at that commit on 26 September; the September field notes, section 04.
09

A cache the script promised and never set

SymptomA start script's comment above its window check says the 262,144 rung sets an 8-bit cache “for you below”. The launch line expands an array named KVARGS. Nothing assigns it. The server starts with the default 16-bit cache and says nothing, because an array that was never assigned expands to nothing: bash 5.2, set -u on, no error.
Checkgrep -n KVARGS start_server.sh: every array the launch line expands needs an assignment line. Or read the cache type back from the server log instead of from the comment.
What it cost usNine days, by the file times, 17 September 16:10 to 26 September, in which the Qwen3.6-27B script offered a 262,144 window that its comment said it could serve and its launch line could not. The error text two lines below the check still said 262,144 was “deliberately not offered”. The assignment was added on 26 September; the copy dated 2026-09-26 22:28 says in its own comment that the rung had “started with an f16 cache that cannot fit”. No run through the script at that window in those nine days is in the records we read.
RecordThis page's package: excerpts/qwen3.6-27b-start-script-kvargs.txt (both versions, with line numbers and file times) and excerpts/kvargs-expansion-check.txt.
10

A settings file wins over your environment

SymptomYou export a variable, a test port or a window, and run the start script. It binds its usual port. Nothing says the file overrode you.
CheckFOO_PORT=8199 bash -c '. settings.env; echo $FOO_PORT' prints what the script will see. Nineteen of the 26 start scripts here source their settings file before any ${VAR:-default} read; the twentieth, MiniMax M2.7, reads only the file's own location first and sources it on the next lines; six have no settings file. In all twenty a plain NAME=value line in the file wins. That is the design, not a fault: the file is the operator's control surface. The trap is expecting the environment to win.
What it cost usA 17 September session note records two verification runs misread as failures for this reason. The runs' logs were not kept, so that is a dated note, not a measurement.
RecordThis page's package: excerpts/settings-file-order.txt, the line numbers for every script and the transcript of the check above.
11

A header that said the draft head did not work

SymptomThe GLM-4.7 Full script's header said --spec-type draft-mtp “DOES NOT WORK on this model here” and put it down to the source. The shared library the script loads implements it: graph_mtp appears 252 times in its strings, and the source tree it was built from carries a 444-line glm4-moe.cpp with the head, against a 284-line older tree with a TODO where the head would be. The note described a different tree from the one the binary came from.
CheckTest the flag on a small window and read the log. When a note blames a build, grep the tree the binary was built from (grep -c graph_mtp src/models/<arch>.cpp), not the tree you remember.
What it cost usFrom 12 September, when this model was serving at 131,072, to 17 September, GLM-4.7 Full served without its draft head. The September roster's row 13 prints its speaking then as 5.51 t/s on paragraph replies, the better of two, on 2026-09-12, at 131,072 with a 5-bit cache and llama.cpp's default batch, before the head was switched on; the script's rung table puts that figure at 131072/92/q5_1, the experts of 92 layers in RAM. The script's own dated note, written when the head was switched on, records 7.02 t/s for the rung 131072/93/q5_1 with the draft head, every expert in RAM (at 92 the draft context did not fit and died with “failed to create MTP context”); 7.41 at 65,536 with --n-cpu-moe 92 and the head on; and one matched pair at 32,768 and --n-cpu-moe 92, 5.23 without the head and 6.54 with it, three runs per arm. The run logs were not kept: a dated note, not a measurement, and no kept log compares the head on and off at the same window and the same placement.
RecordThis page's package: excerpts/glm-4.7-full-start-script-mtp.txt, both headers, the rung table and the library checksums. The September roster, row 13, prints the 5.51 “before a draft head was switched on”.
12

A flag in --help the build cannot load

Symptomllama-server --help lists draft-dspark among the speculation types. Loading DeepSeek V4 Flash's DSpark draft file on that build dies with key not found in model: dflash.attention.sliding_window_pattern.
CheckA name in --help comes from a parser table (common/speculative.cpp, one line per type). Whether the loader can read the file is a separate question, answered by one load.
What it cost usA 17 September note records the failure on the older build here (commit 5f55650); the log was not kept. What the files show: the draft's header carries dflash.attention.sliding_window = 128 and no sliding_window_pattern key; that build's DFlash loader requires the pattern key whenever a window is set (its src/models/dflash.cpp, lines 25 to 27); the newer build (commit d3146f2b5) reads the DSpark draft through a branch of its own that does not. The two builds are not interchangeable for this draft, and the help text of the older one does not say so.
RecordThis page's package: excerpts/deepseek-draft-flag.txt and records/deepseek-draft-header-keys.txt.
13

Ollama's default thread count on the CPU

SymptomA vision model run on the CPU takes minutes on its first picture. When the request sets no thread count, Ollama's server picks one and logs it once, at load; nothing on the request says what it chose.
CheckRead the server's load line, system_info: n_threads = ... (on a standard Linux install, journalctl -u ollama | grep n_threads); then pass num_thread explicitly and cap the picture's long side, and compare the prompt-processing time on one picture.
What it cost usQwen3-VL 30B-A3B through Ollama 0.30.10 (the version the serving process logged when it started on 19 September) on the CPU, 20 September, the same 4K picture first each time, cold, with the model load included, and the card all but empty (1,040 MiB in use): with no thread count set, the server's log reported n_threads = 4 (n_threads_batch = 4) / 24, and the picture took 273.17 seconds of wall time, its 4,154 image-and-prompt tokens processed in 242.93 seconds; 157.31 seconds at 8 threads (123.53 processing); 111.01 at 16 (87.77); 107.88 at 24 (76.57). At 16 threads with the picture capped at 1,536 pixels: 1,370 tokens in 18.17 seconds, 38.52 of wall time (measured 2026-09-20). The answers ran from 174 to 303 tokens, so the wall times include unequal output.
RecordDescribing pictures locally, which publishes with this page and ships the server log excerpt with the thread and version lines; this page's package, records/ollama-threads-2026-09-20.tsv and the log line beside it.
14

The checks, together

Everything above reduces to a few lines, run in this order: before the start, on the script; after the start, on the process, the log and the card; then on the first two reads. Each is a command we ran or a log line our records show; on another build the log wording can differ.

grep -n -E '\[@\]|=\(' start_server.sh                    # every array the launch line expands has an assignment (item 09)
FOO_PORT=8199 bash -c '. settings.env; echo $FOO_PORT'    # what the script will see, not what you exported (item 10)

tr '\0' ' ' < /proc/PID/cmdline                             # the arguments the server really got (item 05)
curl -s http://127.0.0.1:PORT/props                        # the window it serves, n_ctx (items 02, 03)
grep -E 'n_ctx_slot|capping|compute buffer|running DENSE|MSA requires' server.log   # (items 02, 03, 04, 08)
nvidia-smi --query-gpu=memory.used,memory.total --format=csv   # the spare, at load and at the peak of a long read (items 03, 04, 07)

journalctl -u ollama | grep n_threads                       # what Ollama's server chose when the request set nothing (item 13)

# then time the second long request as well as the first (item 06)

Two of the twelve have no command: a note that blames a build (item 11) is checked by testing the flag on the binary you run, and a name in --help (item 12) is checked by one load of the file. Ollama's thread count (item 13) is read from its server's load line, then set.

15

What this does not show

  • One machine, one fortnight, twelve items. This is a list of what our records caught, not a survey; it says nothing about how common any of these is elsewhere.
  • Three items rest on dated notes. The settings-file misreads (item 10), the GLM-4.7 Full speeds (item 11) and the DSpark load failure (item 12) are recorded in a session note or a script note whose run logs were not kept. The page prints them as notes; what the files on disk show for items 11 and 12 is verified and labelled separately.
  • The M2.7 placement rows are one run each. One prompt, 32-token answers; the only second request in each file is the same prompt sent again, a cache hit, whose decode moved by up to 0.34 t/s, more than the 0.19 step between 59 and 58 layers. The ladder is not monotonic (61 was slower than every expert in RAM), the whole gain is under one token a second, and nothing here tests it against run-to-run variation. The card-memory figures are the point, and those are read after load, not at the peak of a read.
  • The Ollama rows are one picture, one run each, with answers of unequal length. The thread count and the version come from the server's own log, which the sibling page ships; our harness recorded neither.
  • The M3 launch-flag fallback was read from source, not observed. No log here shows the warning firing; what we measured was the other door, the old file.
  • Builds move. The source lines are at commit d3146f2b5 (build 10919) and, for the older build, 5f55650; the M3 lines matched the upstream file at that commit on 26 September, and a later commit can change any of them, including whether a log prints a batch size.
  • Nothing here measures output quality beyond the planted codes and planted strings the linked pages describe.
16

What we got wrong

This site publishes corrections of its own records as a genre, and this page is made of them. Three items correct a wrong record of ours that an earlier page had already put right: the served windows on the September roster (its section 05, item 02); the inherited micro-batches on the batch-size study (its section 13, item 05) and the timed-out check on the September field notes (its sections 03 and 15, item 05); the M3 file without its index on the field notes (its section 04, item 08). Three more are traps those pages measured and printed without an earlier wrong record: the capped window (field notes section 09, item 03), the per-window ceilings and the cold first read (batch-size section 04, items 04 and 06). Five are corrected here for the first time: an idle card that a first placement step did not fix (item 07), a start script that promised a cache it could not set for nine days (item 09), a settings order that misled two verification runs (item 10), a header that blamed a build for a head the binary carried all along (item 11), and a note that took a name in --help for a capability (item 12). The twelfth, the thread default (item 13), is measured on the page that publishes with this one. The scripts are ours; so were the notes.

17

Related pages on this site

The batch-size study is the record for three of the twelve (items 04, 05 and 06). The September field notes are the record for item 03 and share items 05 and 08. The September roster is the record for item 02 and supplies the 31,512 MiB start and the 5.51 cell. The 256K study shares the cold first read. The wrong flag is the same kind of finding on a different default, llama.cpp's thread count on a hybrid CPU, and Slower by default traces a shipped default that costs one workload speed. The empty answer is a silent setting of another kind, a thinking budget that returns no text. Describing pictures locally measures the Ollama rows in item 13 in full, and the Qwen3.8-27B at 256K package holds the card-capacity record this page's arithmetic uses.

18

Sources and artifacts

The records that exist nowhere else on this site ship as this page's data package: the files the runs and reads produced, not a summary. If a number on this page disagrees with a file in the package, the file is right and the page is wrong; tell us and we will fix the page. data/README.md maps every file, states its provenance and states the package's redactions plainly; data/NUMBERS.md maps every figure on this page to its file and field, or to the page and package that own it.

2026-09-12 to 09-17GLM-4.7 Full serves at 131,072 without its draft head; its header says the head does not work (item 11). The Qwen3.6-27B script gains a 262,144 rung with a promised cache and no assignment (item 09, from 17 September 16:10).
2026-09-17The GLM-4.7 Full head is tested and switched on; the placement pass over four other scripts runs without keeping its logs (items 07 and 11); the session note on the settings-file order and the DSpark flag (items 10 and 12).
2026-09-19 to 09-21MiniMax M2.7's -ub 128 is found and the batch sweep runs (items 04, 05, 07); the M3 file is replaced (item 08); the Flash-Next rope trap (item 03) and the Ollama thread rows (item 13, 20 September).
2026-09-26Flash-Next's cold first read (item 06); four starts with no window passed (item 02); the GLM-5.3-Flash ceiling at 262,144 (item 04); the Qwen3.6-27B assignment added (item 09); this page's source and header checks.

Every claim on this page is dated. A default, a log line or a loader can change between builds; if today is much later than the dates above, treat the build-specific items as unverified since then.