Graphometer
Field card · measured 2026-09-26 · one NVIDIA RTX 5090 · one request at a time · thinking off

Qwen3.8-27B at 256K

Three files on one RTX 5090.

Measured: a 262,144-token window fits on this 32,607 MiB card. The installed unsloth UD-Q5_K_XL file failed to load MTP beside a 256K q8_0 cache in the two recorded attempts. The smaller UD-Q4_K_XL with MTP was faster than installed Q5 without MTP, especially on tool-shaped replies; this comparison changes both the file and MTP. Q4 showed less agreement with the Q8_0 reference on one public text. Q5_K_S with MTP is the closer same-revision option, with 664 MiB spare at the observed peak. The installed Q5 is revision fe1e2a23; Q4, Q5_K_S and Q8_0 are 4ca72078. This is a choice between measured speed and measured fidelity, on one card and these exact files.

Independent measurements; no affiliation with the Qwen team / Alibaba Group, unsloth, NVIDIA, the llama.cpp project, or Hugging Face.
Update, September 26, 2026

This follows the day-zero field card with a new recorded probe from 26 September 2026. It does not recover the lost 17 September log. A 262,144-token window was served again with a retained log, including a deepest measured read of 229,152 tokens. The full-window comparison covers other long-read measurements. It does not claim a prompt filled every token of that window.

01

What it is

Qwen3.8-27B is the Qwen team's model; unsloth supplies the compressed GGUF files tested here. MTP, or multi-token prediction, is a draft head built into these files. It proposes tokens for the main model to check. It needs additional GPU memory as well as room for the main model's context cache.

Here, 256K means 262,144 tokens of served context, shared by the prompt and answer. The cache stores information used to attend to earlier tokens. At this window all successful runs used q8_0 for both cache keys and values. All serving measurements used llama.cpp, text only.

02

What to set

Measured choices, September 26. For these files and this card, use the installed Q5_K_XL without MTP if you want the closest tested copy to the Q8_0 reference on the one public text scored here. The installed Q5 is revision fe1e2a23; Q4, Q5_K_S and Q8_0 are 4ca72078. Consider Q4_K_XL with MTP for reply speed: it was faster than installed Q5 without MTP, especially on tool-shaped replies, but both the file and MTP change. Q5_K_S with MTP is the closer same-revision copy to Q8_0, with 664 MiB spare at the observed peak. Its small prose-speed differences from Q4 are not a stable ranking.

Measured peak total card use; spare memory is arithmetic
File / settingPeak MiBSpare MiB
Installed Q5_K_XL, 256K, MTP off30,2542,353
Q4_K_XL, 256K, MTP on30,7221,885
Q5_K_S, 256K, MTP on31,943664
Installed Q5_K_XL, 128K, MTP on
default cache (no cache-type override in the recorded command), described as f16 in the launch comment
30,4582,149

Peak is the largest total-card reading sampled during each probe run, including desktop use. Spare = 32,607 minus peak. These are not guaranteed loading margins. Q5_K_S has only 664 MiB spare at the observed peak.

For the installed Q5_K_XL, the measured serving flags were:

llama-server --model Qwen3.8-27B-UD-Q5_K_XL.gguf \
  --jinja --ctx-size 262144 --parallel 1 \
  --batch-size 2048 --ubatch-size 512 \
  --cache-type-k q8_0 --cache-type-v q8_0 \
  --n-gpu-layers 999 --threads 24 --threads-batch 24 \
  --flash-attn on --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0

This lists the model and inference settings from the recorded command; set your own local endpoint. There are no MTP flags on this configuration. For Q4_K_XL or Q5_K_S at the same window, change the model filename and add:

--spec-type draft-mtp --spec-draft-n-max 6 --spec-draft-p-min 0.75 \
--cache-type-k-draft q8_0 --cache-type-v-draft q8_0

For the installed Q5_K_XL at 128K (131,072 tokens), the tested baseline uses MTP with the same draft length and confidence gate and default cache (no cache-type override in the recorded command), described as f16 in the launch comment. Omit the target and draft cache overrides and set --ctx-size 131072. This is the faster baseline to consider when the shorter window is enough, but its default cache, described as f16 in the launch comment, differs from the 256K q8_0 cache, so the comparison does not isolate window size or MTP; the tool-shaped test below is not a coding benchmark.

Keep the confidence gate explicit. The speculation-gate study explains why output shape matters. The batch-size study explains the reading-side setting; this card measured -b 2048 -ub 512 on the successful runs.

03

What we ran, exactly

Machine and dateOne NVIDIA RTX 5090, 32,607 MiB reported card capacity. Serving runs on 2026-09-26, one request at a time.
Serving buildThe launch script identifies llama.cpp mainline c8e03ce, built 2026-08-05, CUDA sm_120. This is script-recorded identity, not a fresh binary hash.
Request settingsThinking off via chat_template_kwargs: {"enable_thinking": false}. Real prose: temperature 0.7, seed 7, cap 900 tokens. Tool-shaped replies: temperature 0, seed 7, cap 1,600. Recall short answers: temperature 0, seed 1, cap 300.
Input and outputA generated ledger at each depth, followed by a real prose letter or a JSON array of ledger entries. These are tool-shaped replies, not actual tool execution. The prose and JSON requests reuse the ledger prefix from the preceding cold read.
RepetitionsOne serving run per configuration; one observation per depth and reply type. No speed confidence intervals.

The server command sets temperature 1.0, but the request overrides it. The retained probe and invocations specify 0.7 for prose. Small differences between sampled replies should not be treated as stable rankings.

The filename alone is not enough

Recorded file identities: the installed UD-Q5_K_XL is the August 14 upload at revision fe1e2a23, 20,218,178,624 bytes. An expected-size list in this package names another UD-Q5_K_XL at 20,876,938,144 bytes (stated). That file was not tested. Do not apply the installed file's memory result to it.

Recorded tested files; GB (bytes/10^9) and GiB (bytes/2^30) are arithmetic, rounded
FileRecorded revisionBytesGB (bytes/10^9)GiB
Installed UD-Q5_K_XLfe1e2a2320,218,178,62420.218.83
UD-Q5_K_S4ca7207818,665,753,50418.717.38
UD-Q4_K_XL4ca7207817,559,178,14417.616.35

Revisions are those named by the retained launch record. The package includes file hashes and the size-verification record. The comparison holds across these actual files, not just their quantization labels.

04

Observed: speaking and reading

Measured tokens per second. Depth columns name approximate ledger depths, not just the short new message. Real prose replies generated 414 to 504 tokens across these runs; tool-shaped replies generated 907 to 1,192. Reading rows time the whole cold prompt. Decode rates exclude the cold read.

Measured: real prose replies, tokens/s
File / setting~20K~48K~100K~230K
Installed Q5_K_XL
256K, MTP off
60.755.249.134.9
Q4_K_XL
256K, MTP on
75.264.854.938.4
Q5_K_S
256K, MTP on
71.564.552.239.9
Installed Q5_K_XL
128K, MTP on
default cache (no cache-type override in the recorded command), described as f16 in the launch comment
72.070.560.8Not run
Measured: tool-shaped replies, tokens/s
File / setting~20K~48K~100K~230K
Installed Q5_K_XL
256K, MTP off
60.754.948.934.8
Q4_K_XL
256K, MTP on
144.7129.9114.787.0
Q5_K_S
256K, MTP on
137.4126.0110.282.8
Installed Q5_K_XL
128K, MTP on
default cache (no cache-type override in the recorded command), described as f16 in the launch comment
141.5136.9127.0Not run
Measured: cold prompt processing, tokens/s
File / setting~20K~48K~100K~230K
Installed Q5_K_XL
256K, MTP off
3,1182,7181,9581,181
Q4_K_XL
256K, MTP on
3,0882,7221,9371,132
Q5_K_S
256K, MTP on
3,0012,6671,9051,124
Installed Q5_K_XL
128K, MTP on
default cache (no cache-type override in the recorded command), described as f16 in the launch comment
2,8962,6081,973Not run

The 128K baseline stops at ~100K. Its default cache, described as f16 in the launch comment with no cache-type override in the recorded command, differs from the 256K runs' q8_0 cache, so these rows do not isolate the effect of window size or MTP alone.

The largest observed speed difference is on the tool-shaped replies in this test. At ~20K, Q4_K_XL with MTP produces real prose at 75.2 tokens/s and tool-shaped replies at 144.7 tokens/s. At ~230K those rates are 38.4 and 87.0 respectively. Neither pair is a single general-purpose speaking speed.

05

The long read and the crash check

Measured: at a 262,144-token window, each of the three successful files read a 229,152-token prompt and returned all 3 sealed codes. The generator placed them at approximately 5%, 50% and 95% of ledger entries. These are positions in the generated document, not exact tokenizer offsets.

Measured: cold read followed by a short code answer
File / settingRead sWhole request sCodes
Installed Q5_K_XL, MTP off194.086196.93 / 3
Q4_K_XL, MTP on202.377205.23 / 3
Q5_K_S, MTP on203.831206.63 / 3

Read seconds = recorded prompt milliseconds ÷ 1,000 (arithmetic conversion). Whole-request time includes the short answer and other overhead. Recall is the dedicated all-codes question, not the later letter or JSON reply.

Measured crash checks: the installed Q5_K_XL without MTP and Q4_K_XL with MTP each completed 40 of 40 prompts with zero crashes and zero empty answers. Each prompt had an initial reply and a follow-up turn. These were shorter mixed-text prompts at the 256K serving setting, not repeated full-window reads. Q5_K_S has no corresponding crash-check result in this package.

06

Why the installed file loses the draft head

Measured failures: the installed Q5_K_XL could not create the MTP context beside its 256K q8_0 target cache. Both recorded MTP-on attempts of the installed Q5 file used --ubatch-size 256 (micro-batch size). The failed requests were:

Failed CUDA allocations during MTP context creation
Draft cacheFailed bufferBytes
f161,024.00 MiB, KV cache1,073,741,824
q8_01,192.27 MiB, compute buffer1,250,189,440

Those are allocation sizes, not measured amounts by which the card was short. The logs show the requested buffers and the out-of-memory failures. They do not establish the smallest extra card capacity that would have made either load succeed.

The smaller Q4_K_XL and Q5_K_S files did load with MTP and q8_0 draft caches with --ubatch-size 512. That is the practical route measured here. These attempts do not prove that every other possible Q5_K_XL configuration fails.

07

What the smaller file changes

Measured fidelity to a reference, not answer quality. Perplexity measures how well a file predicts a fixed text. Across 24 chunks of 4,096 tokens from public llama.cpp documentation, the three scores were too close to distinguish at the reported uncertainty.

Measured perplexity; reported ± terms retained
FilePerplexity
Installed Q5_K_XL2.7203 ± 0.02379
Q5_K_S2.7186 ± 0.02375
Q4_K_XL2.7206 ± 0.02379

A second test compared next-token predictions against unsloth's plain Q8_0 reference at revision 4ca72078, using 12 chunks of 4,096 tokens from the same text. Q5_K_S and Q4_K_XL carry that same recorded revision; the installed Q5_K_XL is the older upload. KL divergence measures how far the full predicted probability distribution moves from the reference; lower is closer.

Measured against Q8_0 on one public text
FileSame top tokenMean KL divergence
Installed Q5_K_XL97.830 ± 0.093%0.004369 ± 0.000153
Q5_K_S97.675 ± 0.096%0.005406 ± 0.000201
Q4_K_XL97.240 ± 0.105%0.007839 ± 0.000345

In plain words, the smaller Q4 file picks the same most likely next token as the 8-bit reference about 97 times in 100 on this text. The installed Q5 file is closer to 98 in 100. A token can be a word or part of one. This is not a percentage of correct answers, and it does not measure the effect of MTP on generated answers.

The reference is Q8_0, not the original full-precision weights. Its recorded size is 29,047,086,048 bytes. The scoring scripts use a separate llama-perplexity executable recorded under build label b10453. All scores use the fixed 414,554-byte corpus identified by hash in the package. The logs' ± terms are reproduced as reported, without treating them as speed intervals or answer-quality guarantees.

08

What this does not show

  • One card, one request at a time, one serving run per configuration. Concurrency, other GPUs and other builds were not tested.
  • A served 262,144-token window and a successful 229,152-token read. Neither code recall nor a short crash check establishes reliable reasoning over every part of a full window.
  • Speed is not quality of answers. The ledger letter and JSON array are controlled output shapes, not a representative coding or tool-use evaluation.
  • Fidelity was measured on one public technical text against one quantized reference. Different upload revisions also limit attribution to quantization alone.
  • The larger Q5_K_XL file named in the expected-size list has no serving result here. A matching filename is not evidence of matching memory use.
10

Sources and artifacts

The data package contains the recorded probe rows, server logs, failed allocations, memory samples, crash-check records, scoring logs and redacted script excerpts. The number map identifies every printed figure. File revisions and build identity are the ones recorded in those scripts; the package does not contain an independent binary-hash audit.

The fidelity corpus is identified by byte count and SHA-256; its text and the reference logits are not included. The scoring results are inspectable, but the package alone is not a complete rerun kit for that test.

If a number on this page disagrees with a file in its package, the file is right and the page is wrong; tell us and we will fix the page.