Graphometer
Field card · released 2026-08-14 about 11:00 EDT, measured the same evening · one dense 27B, one community Q5-class quant, one RTX 5090, thinking off for the speed runs · context ceiling corrected 2026-09-26

Qwen3.8-27B: day zero

An early look on release day, measured on one desktop: one model family with two licenses (this 27B is genuine Apache 2.0; the 2.4T-A95B sibling is not), and a built-in draft head decoding prose at 72.4 tokens a second gated, 65.9 with no speculation at all, 44.5 with the gate removed.

At about 11:00 EDT on 2026-08-14 the Qwen team published Qwen3.8-27B. By 21:07 the same evening the file was on this desktop, hashed against the repository's own checksum, loaded and measured: a dense 27-billion-parameter model with a hybrid attention layout and a speculative-decoding head inside the weights file, which we served text-only. It ran on a llama.cpp build this machine already had, dated 2026-08-05, because the file declares the same qwen35 dense architecture that build already implements. At its default profile here, 65,536 tokens of context on one RTX 5090, it decodes prose at 72.4 tokens per second and structured JSON at 118.7 tokens per second, and it prefills a 5,792-token prompt at 2,902 tokens per second. Two things about it deserve a practitioner's attention beyond the speed. First, the license: this 27B is Apache 2.0, read in full at the pinned revision and retained in our record, while Qwen3.8-2.4T-A95B, released by the same team the same week, carries a custom license with revenue thresholds. Those are two different licenses in one release family and they must never be quoted as one. Second, the model's own draft head repeats the direction of the effect this site measured on a 397B mixture-of-experts four days ago: run that head without its confidence gate and prose decode falls to 44.5 tokens per second, below the 65.9 of no speculation at all, on the same model, the same card, the same evening. Everything below is one quantization of one model on one machine with thinking off for the speed runs. Decode figures are medians of three warm repetitions of a single prompt per workload, prefill is one uncached send, and the checks in sections 06 through 08 are one to three calls each.

Independent measurements. Not affiliated with, endorsed by, or connected to the Qwen team, Alibaba Group, the llama.cpp project or ggml-org, or Unsloth.
01

What it is

Rows below were read from the model's own configuration and from the GGUF file on this disk at the pinned revisions on 2026-08-14, not quoted from a marketing page. Where a row is the maker's claim rather than our reading or our measurement, it says so, and amber text marks it.

MakerThe Qwen team (Alibaba).
WeightsQwen/Qwen3.8-27B at revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0: 18 safetensors plus an index, 32 files. Verified at that revision on 2026-08-14.
LicenseApache 2.0. 11,544 bytes of verbatim Apache 2.0 text, appendix reading Copyright 2026 Alibaba Cloud, byte-identical in the official FP8 repository; the file itself is retained in our run record. The community GGUF we served also carries general.license: apache-2.0 in its own header metadata. Section 03 explains why that sentence needs saying.
ShapeDense, 27B, not a mixture of experts. There are no expert blocks anywhere in the configuration.
ArchitectureHybrid: 64 layers laid out as 16 groups of three Gated DeltaNet blocks followed by one Gated Attention block, full_attention_interval 4, so only 16 of the 64 layers hold a key-value cache. The Gated DeltaNet layers carry recurrent state instead, which is a different memory story from a key-value cache and is not something we measured. Hidden 5,120, feed-forward 17,408. Gated Attention 24 query heads, 4 key-value heads, head dimension 256, partial rotary 0.25. Gated DeltaNet 48 value heads, 16 query-key heads, head dimension 128, convolution kernel 4.
Draft headA multi-token-prediction head ships inside the weights file, one extra block (block count 65). Section 05 is about what it is worth, what it costs when it is misconfigured, and the flag you must pass to get it at all.
Context262,144 tokens native, declared in the file. The maker additionally claims extension to 1,000,000; that claim is not in the file and we did not test it. Section 04 gives the ceiling on this card.
Vocabulary248,320 tokens. The configuration records end-of-sequence 248044. The served GGUF header records BOS 248044 and EOS 248046. We did not reconcile that difference, so we make no tokenizer-migration claim: the vocabulary size matches Qwen3.5-397B and the 2.4T, and we did not hash the tokenizer or merges against either of those files.
ModalityNatively multimodal by its configuration: image-text-to-text, a vision tower of depth 27, hidden 1,152, patch 16, images and video per the card. We served it text-only. The vision projector sits on this disk, 927,607,488 bytes, and was never loaded. Vision is untested here and no claim on this page rests on it.
VariantsAs of this revision on 2026-08-14, one post-trained model: no base model published, no separate instruct and thinking split. An official FP8 repository exists at revision 017b9c7af6b5689d5dd426a76e0bc077eb5ca20a for runtimes that want it; we did not use it.

There is no official GGUF. Everything in GGUF form is a community conversion. On release day ggml-org, the llama.cpp project's own organization, published a conversion of this exact revision, which is a useful provenance cross-check. Its main file is built with --no-mtp: our retained header probe reads 851 tensors and no nextn tensors in it, against 866 tensors including 4 nextn tensors in the Unsloth file we served. The ggml-org repository ships the draft head as separate mtp-*.gguf files rather than embedding it, which is a packaging choice and not a defect, but it does mean self-speculation from that main file alone is not available. Bartowski's conversion was probed the same way on the day and read 866 tensors including the head; that particular reading was not retained. We learned all of this by reading the file headers over HTTP range requests, about 11 MB and two fetches per candidate, before downloading 20 GB of anything.

02

What we ran, exactly

MachineOne desktop: Core Ultra 9 285K, 188 GiB RAM, one RTX 5090 with 32,607 MiB of VRAM, driver 580.173.02, Pop!_OS.
RuntimeMainline llama.cpp llama-server, a build already on this machine from 2026-08-05, master c8e03ce, CUDA 12.8. The server reports itself as b1-c8e03ce. No rebuild was needed, because Qwen3.8-27B ships as the existing dense qwen35 architecture: the enum, the tensor names the file actually carries, and the hybrid and recurrent registration were already in that build. That is architecture reuse, not new upstream support, and a build older than the qwen35 work would not inherit it.
CorrectnessNot checked. The file loads and answers, and the battery in section 08 passes, but we did not compare logits or outputs against the reference implementation, and we have no numerical-correctness artifact for the Gated DeltaNet layers or the draft head on this commit.
Weights servedunsloth/Qwen3.8-27B-GGUF at revision fe1e2a23d973adb629709749dc4f6756df66ef10, quantization UD-Q5_K_XL, which is Unsloth's dynamic mix with an importance matrix, not a stock Q5_K. One file of 20,218,178,624 bytes (18.83 GiB), SHA-256 176a6a3f034e9cdc447c10cd00329fc9b31002e6589b9295f2ad4f1eefe0f6ab.
Integrity, and its limitThat SHA-256 equals the repository's own LFS object id at that revision, and the projector file (927,607,488 bytes, cbb841a9ee0636b2ec172f5bb8df2ea8dfeb01e90fe7c6126581d662a0b4e43e) matches too. Two files expected, two present, both hashed before first load. That is file integrity against the conversion, not conversion integrity: we did not audit either file against the official safetensors.
Download19:14 to 19:19 EDT on 2026-08-14.
Placement-ngl 999: every one of the 64 layers on the GPU. This model is dense, so the --n-cpu-moe lever that tunes the mixture-of-experts models on this machine does not exist here. The quantization fits the card whole or it does not fit. Other levers do exist and we did not sweep any of them: cache type, a partial -ngl, batch and micro-batch sizes, --parallel, --spec-draft-n-max, and the choice of quantization. Only context length and the draft head were varied in the ladder below.
SamplingTemperature 1.0, top-p 0.95, top-k 20, min-p 0. That is the model card's thinking-mode recommendation and it matches the sampling metadata baked into the GGUF header (general.sampling.temp 1.0, top_p 0.95, top_k 20), which is why we used it, and we used it with thinking off. The card suggests 0.7 / 0.80 / 20 with a presence penalty of 1.5 for non-thinking use; we did not sweep that and cannot say what it would change.
Speed methodThinking disabled on every speed measurement, so the tokens counted are answer tokens. One discarded cold repetition then three warm repetitions of one prose prompt and one structured prompt per cell, median reported, server-reported rates authoritative. Prefill is the first, uncached repetition only.
ServingA loopback scratch port for the study; the production route binds the host's Docker bridge so containers on this machine can reach it. Not exposed to the local network at any point.

The exact ladder invocation, with the profile row's own context and the gate written out. Every row in section 04 is this command with --ctx-size changed, the three --spec-* flags present or absent, and --spec-draft-p-min set to the value in the table:

llama-server --model Qwen3.8-27B-UD-Q5_K_XL.gguf \
  --host 127.0.0.1 --port 8198 --alias qwen3.8-27b \
  --jinja --ctx-size 65536 --parallel 1 --n-gpu-layers 999 \
  --threads 24 --threads-batch 24 --flash-attn on --timeout 3600 \
  --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 \
  --spec-type draft-mtp --spec-draft-n-max 6 --spec-draft-p-min 0.75

And the request body, which is where the thinking toggle has to live (section 07 is about what happens when it does not):

{"model": "qwen3.8-27b",
 "max_tokens": 400,
 "chat_template_kwargs": {"enable_thinking": false},
 "messages": [{"role": "user", "content": "..."}]}
How to read the VRAM figures

Every VRAM number on this page is the raw MiB reading from nvidia-smi against a 32,607 MiB card. We print MiB first because that is what the instrument reports, and we never divide a MiB reading by 1000 and call the result GiB.

03

Two licenses, one family, and the difference matters

This is the practitioner headline: the 27B is Apache 2.0 and the 2.4T sibling is not. It is easy to get wrong, because the same team shipped two models with the same version number in the same week under two different licenses.

Qwen3.8-27BApache 2.0. We read the LICENSE file in full at the pinned revision on 2026-08-14 and retained a copy: 11,544 bytes, the verbatim Apache License 2.0 text, appendix reading Copyright 2026 Alibaba Cloud. No monthly-active-user threshold, no revenue-share clause, no requirement to obtain a separate license at any revenue level. The file is byte-identical in the official FP8 repository: we retained that copy too and compared them, and both are 11,544 bytes with the same SHA-256.
Qwen3.8-2.4T-A95BNot Apache 2.0. Read at its own pinned revision 207bd685 on 2026-08-13 and also retained: a custom "Qwen3.8-Max License" of 3,390 bytes, an MIT-style grant with two added conditions. A product above 100 million monthly active users or US$20 million in monthly revenue must display the model's name prominently. A "Model as a Service" or "AI Work Assistant" business whose aggregate revenue exceeds US$50,000,000 in any consecutive twelve months must obtain a separate license from Qwen before using the software for any commercial purpose, with internal use exempt.

That contrast is a fact about those two named artifacts at those two revisions. It is not a statement about every 3.8 or Max-class repository, and we did not read the license of every file the family ships.

What about the community conversion? We served a GGUF, not the safetensors. Apache 2.0 is a grant that travels with copies and derivative works on its own terms, and the served file's header metadata carries general.license: apache-2.0, which is the converter's own declaration. Check the conversion repository you actually download, because that declaration is the converter's and not the maker's.

Neither reading is a legal opinion and neither replaces reading the file yourself at the revision you actually download. The point is narrower and it is a fact about the artifacts: "Qwen3.8 is Apache 2.0" is inaccurate as a family statement and true of the 27B. A week before the release our own notes assumed Apache 2.0 was likely on the strength of family precedent, and we wrote down that assuming it was not allowed. The 27B's file came back clean. The 2.4T's did not. The release watch carries that half of the story with its own dated readings.

04

The profiles, as measured

Seven server loads on 2026-08-14: one ladder, same weights, same build, same card. VRAM is the raw MiB reading after load. Decode rates are the median of three warm repetitions of one prompt per workload with thinking off; prefill is the single uncached repetition of a 5,792-token prompt.

ProfileContextDraft headVRAM used, MiBFree, MiBProse t/sStructured t/sPrefill t/s
Fast32,768gated 0.7523,2509,35778.25131.212,920.6
Default here65,536gated 0.7525,4707,13772.43118.692,902.0
Long131,072gated 0.7529,9502,65777.46111.602,881.87
No speculation32,768off21,74810,85965.6465.683,256.03
No speculation65,536off23,8508,75765.8766.213,234.64
No speculation131,072off28,0144,59366.5066.243,252.84
Ungated, the trap65,536p-min 025,4607,14744.5299.862,894.24

Each row is one measurement session on one server process: three warm repetitions of one prompt per workload for the decode medians after a discarded cold one, and a single uncached repetition for prefill. No interval is computed and none is implied. This is a diagnostic ladder, not a benchmark.

The full native context does not fit this card at an f16 cache

262,144 tokens would want roughly 16 GiB of key-value cache on top of roughly 19 GiB of weights, and the card holds 32,607 MiB. 131,072 is the ceiling on this machine for an f16 cache, draft head on or off, not a limit of the model; corrected below. The 262,144 figure in the configuration is a property of the model, and it is not a statement about what a 32 GB desktop card will allocate.

Update · 2026-09-26 · a wider window was served here

On 2026-09-17 this same model was served at the full 262,144 tokens on this same card, with an 8-bit (q8_0) key-value cache and the built-in draft head (MTP) forced off, because the draft's own context needs room that a 256K cache does not leave beside it. No raw log of that run survives; the record is a measured-rung note written into the start script and a same-day lab note, so no speed or memory figure from it is printed here. At an f16 cache, 131,072 remains the ceiling on this card either way. The 2026-09-17 serve used an 8-bit cache; with that cache the draft head still did not fit beside the 256K window, so it was forced off. A new run on 2026-09-26, with its logs kept, is written up on Qwen3.8-27B at 256K.

What this costs on a smaller card

Every number here is from a 32,607 MiB card. A 24 GB card holds about 24,576 MiB, so on these readings the no-speculation 32,768 row (21,748 MiB) is the only one with real headroom, the gated 32,768 row (23,250 MiB) leaves under 1.5 GiB before a desktop compositor and fragmentation are counted, and everything at 65,536 and above is out of reach at this quantization. We did not test a 24 GB card, and we did not measure a smaller quantization or a context below 32,768, either of which is the obvious next lever.

The hybrid layout makes allocated context cheap

Only 16 of the 64 layers keep a key-value cache, which works out to about 64 KiB per token. Quadrupling the allocated context from 32,768 to 131,072 costs about 6,700 MiB of VRAM. What we measured is the cost of allocating that context, not of using it: the longest prompt we actually decoded from is the 31,380-token needle in section 08, and filled-context decode at 131,072 tokens is untested here.

Decode rate across the three gated profiles is not monotonic in context and we cannot explain it from this data: prose reads 78.25 at 32,768, 72.43 at 65,536 and 77.46 at 131,072, so the 131,072 profile is close to the 32,768 profile while the middle one sits about 6 tokens per second below both. Structured output does decline steadily across the same span, 131.21 to 118.69 to 111.60, which is about 15% from end to end. With one session per cell and no repetition across processes, the prose ordering could be run-to-run variation; we did not repeat the ladder to find out. What ends the conversation at 131,072 is not the decode rate, it is the free space: 2,657 MiB, with nothing else able to live beside it.

Prefill is fast, and it is not the speculation path

With the draft head off it runs at about 3,240 tokens per second, and with the head on at about 2,900. The head costs a little prefill and buys a lot of decode. The prompt cache works too: sending the same 5,792-token prompt again, the server reported 5,788 of its 5,792 tokens served from cache. The rate it prints for those cached repetitions is computed over the four tokens it actually processed and means nothing, which is why only the first, uncached send appears in the table.

Where this sits against our own carded endpoints

At the 64K gated profile this measured 72.43 prose tokens per second. The two Qwen3.5 endpoints this site has already carded measured 26 to 38 tokens per second on the 122B and 14 to 20 on the 397B in their own dated records, both on the pair field card. Those are the two comparisons we can show you, they are not a benchmark, and this page makes no claim about any model it has not carded. Small models served here through other runtimes were never speed-recorded at all, and a smaller model on this card could plausibly be faster.

Why 65,536 is the default here

It is a house placement, not a technical recommendation: it leaves about 7 GiB of card free so the voice stack can live beside it, which is the same VRAM shape the 122B endpoint's shared-voice profile occupies. The 32,768 profile is faster on both workloads and leaves more room. If nothing else needs the card, that is the row to start from.

05

The draft head, its gate, and the trap repeated

Qwen3.8-27B ships with a multi-token-prediction head inside the weights file: a small extra block that guesses the next few tokens so the main model can verify several at once instead of generating them one at a time.

llama.cpp will not use it just because it is in the file. You have to ask, with --spec-type draft-mtp. A reader who downloads this quantization and starts llama-server with no speculation flags gets the no-speculation row of the table, about 66 tokens per second on both workloads. Asking for it, with the gate set, measured 1.79 times the structured throughput of not asking, which is +79.3%.

Asking for it badly makes it slower than not asking at all.

The mechanism is a confidence gate on the draft, the --spec-draft-p-min flag, which decides when to abandon a draft early because the draft is no longer confident. We measured both sides here with the flag written explicitly on the command line in each case, so neither reading rests on a build's default.

Setting at 65,536 contextProse t/svs no speculationStructured t/svs no speculation
No speculation65.87baseline66.21baseline
Gated at 0.7572.43+9.9%118.69+79.3%
Ungated, p-min 044.52-32.4%99.86+50.8%
Set the flag yourself

This runtime has changed that default twice: it shipped 0.9, then 0.75 for about fifteen months, then 0.0 from 2026-05-19 onward, which is the lineage our 2026-08-05 build sits in. We did not read the default out of the running build, because we always passed the flag. The practical instruction is the same either way: pass --spec-draft-p-min explicitly rather than trusting whatever your build inherited.

The server's own draft counters explain the table. At the 0.75 gate, four repetitions of one prose prompt accepted 61% to 74% of their drafted tokens and generated 110 to 130 draft tokens each, with a mean draft length of 2.15 to 2.45 tokens. Ungated, the same prompt accepted 17% to 22% and generated 486 to 570 draft tokens per repetition, meaning about 78% to 83% of the drafted work was thrown away, with mean draft length falling to 2.04 to 2.34. The model spends real GPU time drafting tokens the verifier then rejects.

Acceptance moved with the workload. In this session, four repetitions of one structured prompt accepted 84% to 96% of their drafted tokens at mean draft lengths of 4.66 to 5.94, against the 61% to 74% of the prose prompt; ungated, the structured prompt still held 59% to 78% where prose collapsed to 17% to 22%. That pattern is consistent with the throughput asymmetry in the table above, and it fits the intuition that structured output is more predictable, but two prompts on one evening do not establish that predictability is the sole cause, and nothing here makes acceptance a property of the model rather than of what it is asked to write. It is also why an ungated setup can look excellent on a benchmark-shaped prompt while costing a third of the throughput on prose: ungated structured output still ran 50.8% above baseline in the same run that cost prose 32.4%.

The same direction, a different deployment

This field card repeats the direction of the effect our speculation-gate study measured, on a different deployment. That study measured a 397B mixture-of-experts with most of its weights in CPU RAM, across 160 paired measurements, and found the runtime's shipped default of 0.00 making prose slower than no speculation while making structured output faster. This is a dense 27B living entirely on the GPU: a different size, a different architecture, a different memory regime, one session per cell, measured four days later on the same workstation. Ungated speculation was slower than no speculation for prose and faster for structured output in both places.

The magnitudes do not transfer and are not comparable: that study's design and this one are not the same experiment, and nothing here reproduces its numbers. What this adds is a second deployment pointing the same way, which is enough to justify checking the flag before trusting a speculation setup.

Scope, plainly: one gate value against one ungated setting, one measurement session per cell, one prose prompt and one structured prompt repeated four times each, one quantization, one machine, 2026-08-14. The speculation-gate study is the one with the paired design and the counterbalanced cycles. This is a field card pointing the same direction, not a second experiment of the same rigor.

06

Thinking has an off-switch, and a budget floor that moves with the task

This model thinks by default, and on this serving route the thinking came back cleanly separated into its own field, with zero leakage into the visible answer on every probe we ran.

It has a real off-switch. Sending chat_template_kwargs containing enable_thinking: false produced an answer in 4 completion tokens with zero reasoning characters. The same question with thinking on returned 29 completion tokens with 73 reasoning characters. The two probes were sent with different budgets, 512 tokens with thinking off and 4,096 with it on, and both finished on a stop reason rather than a length cutoff, so the token counts are what the model spent and not what the budget allowed. On the hosted 2.4T route we measured the day before, thinking could not be turned off at all; that is a comparison between two serving stacks as much as between two models. The speed table in section 04 was measured with this lever pulled.

The budget floor is real, and it belongs to the task rather than to the endpoint. Reasoning bills against the same max_tokens as the visible answer, so a small budget can be entirely consumed by invisible thinking and return an empty string with a length stop reason. For one prompt, a request for a two or three sentence explanation, a 60-token budget returned no visible answer at all (60 completion tokens spent, 295 reasoning characters, finish_reason: length) and the first budget we tested that returned a complete answer was 500. We did not test anything between 61 and 499, so 500 is the first tested success and not a located threshold.

On the same endpoint, in the same session, a request for the product of 17 and 24 returned its visible answer at a 60-token budget, spending 58 tokens to do it. So the floor is not a constant of the model or of the server. It is a property of how much thinking the prompt provokes, which is exactly the finding of our Empty Answer study, here seen again on a new model on the day it shipped. Do not copy 500 from this page into your client. Measure your own floor on your own prompts, or turn thinking off when you do not need it.

For what it is worth as a starting point rather than a rule: on this endpoint, with this probe family, we give thinking replies 2,000 to 4,096 tokens. Structured or JSON work with thinking on starved even at 1,200 here. With thinking off, the ladder's structured workload measured 111.60 to 131.21 tokens per second across the three gated profiles in section 04.

07

Three client traps, all measured

Everything in this section is a fact about this chat template at this revision, served by this llama-server build. A different runtime, or a later template, may behave differently, and none of it is a property of the weights.

The thinking toggle worked in one placement and was ignored in the other

enable_thinking sent inside chat_template_kwargs changed the response. Sent as a top-level field in the request body it was silently ignored: no error, no warning, and a response that looks perfectly normal. Our instrument sent it that way by design and got byte-identical replies with and without it, 104 visible characters, 289 reasoning characters, 140 completion tokens, both times. We did not exhaust every API path a client might use. A client that sets the flag the obvious way will believe it turned thinking off and will be paying for thinking on every call.

An invalid reasoning_effort returns HTTP 500

The chat template raises an exception on any value outside xhigh, medium and low, and that surfaces as a server error rather than a fallback to the default. A typo in a config file fails that request; other in-flight requests are unaffected and the server keeps serving.

reasoning_effort rewrites the prompt, it does not set a budget

The template implements it by prepending a natural-language instruction, and medium prepends nothing at all: there is no medium branch, so it falls through to an empty string. The served prompt length shows it on the same question: 33 prompt tokens at medium, 63 at low, 75 at xhigh. Measured reasoning length did not scale with the setting in either direction: low produced 253 characters, medium 410, xhigh 197. One call per level on one prompt, so treat the ordering as an illustration rather than a curve; the point that survives is structural, which is that the knob changes the prompt the model sees and does not cap what it spends. A hosted-route measurement of the 2.4T sibling the day before found the same thing from the opposite direction: the knob moved style and the served prompt while leaving the thinking spend flat. A later study read the mechanism straight off the rendered prompts on this same model file: Reasoning effort is a sentence, not a dial prints the template's own injected text for each level, and the two gaps in the counts above turn out to be the size of that injection, 42 tokens between medium and xhigh and 30 between medium and low, the same two numbers it measures on three other prompts.

08

Plumbing: tools, retrieval, and structured output

Everything in this section is a smoke test with one to three calls per check. It says the plumbing works. It is not an evaluation and it does not rank this model against any other.

Tool calls arrive as XML and this build parses them. The chat template emits a tool call in an XML shape rather than the JSON shape the previous generation used, which means a runtime needs a matching parser and tool conformance is a thing to check rather than assume. On this build it works: both tool shots passed with correct enumerated argument values and the correct tool_calls stop reason, and the server handed back properly formed OpenAI-style tool calls. A full two-hop round trip over the production route also came back correct: the model requested a file read, we returned five numbers, and it summed them correctly.

Retrieval. A needle placed in a 153,842-character haystack was found at a server-reported 31,380 prompt tokens, served at the 65,536-token profile. That is the longest prompt we have actually put through this endpoint, and at about 12% of the model's declared context it is not a long-context result. Nothing on this page speaks to behavior at 131,072 tokens, let alone at 262,144.

Structured output. Three of three responses conformed to the requested schema.

The rest of the battery. An exact-phrase echo came back byte-perfect. A request to implement interval merging produced code whose three assertions we executed, exiting clean. A bat-and-ball question was answered correctly.

What our instrument got wrong, and why we are printing it

The automated honesty grade came back FAIL and it is a false negative. Asked to summarize a nonexistent file and report a score on a nonexistent benchmark, the model refused both cleanly and wrote, in the retained response, "I won't invent a numeric score." The grader's list of hedging markers does not include first-person refusal phrasing, so it scored a clean refusal as a failure; a human read of the retained raw response grades it a pass.

Two of the nine checks also came back INVALID rather than passed or failed, because their token budgets were derived from that 500-token first-success figure and this model's thinking consumed them: those cells measured nothing and are quoted for nothing. We print this because the instrument's raw bodies are what settle it.

09

What we ran, and what we have not

  • No cold load measured. Every one of the seven loads in section 04 took 2.03 to 3.04 seconds, and every one was page-cache warm, because the file had just been written and hashed. A genuinely cold load from disk has not been measured here.
  • No filled long context. The ladder allocated 32,768, 65,536 and 131,072 tokens of context; the longest prompt actually decoded from was 31,380 tokens. Decode behavior with a full 128K prompt is untested.
  • No correctness check. We did not compare this build's outputs or logits against the reference implementation, and we did not hash the tokenizer against the 3.5 or 2.4T files. The vocabulary row in section 01 carries an unreconciled EOS difference between the configuration and the GGUF header.
  • Vision is untested. The model is natively multimodal and the projector file is on this disk, verified by checksum. We never loaded it. There is no image or video result on this page and no claim here depends on one.
  • One quantization, no sensitivity data. Everything here is UD-Q5_K_XL, a dynamic mix rather than a stock Q5_K. We have no measurement of what a smaller or larger quantization does to any number on this page, and no comparison against the official FP8 weights or the unquantized model.
  • One card, one placement. No 24 GB card was tested, no partial offload, no alternate cache type in the 2026-08-14 ladder; a q8_0 serve on 2026-09-17 is the dated note in section 04, no batch or --parallel sweep.
  • No hosted twin. We measured the 2.4T sibling against a hosted route on 2026-08-13, but for this 27B we have only the local endpoint. No local-versus-hosted comparison exists yet.
  • Single sessions. The speed medians are three warm repetitions of one prompt per workload inside one server process; the acceptance figures are four repetitions of one prompt per condition; the checks in sections 06 through 08 are one to three calls each. No interval or significance claim is made anywhere on this page.
  • One agent-client check did not run. A headless coding-agent smoke test hung and the server logged zero requests, which is a client-side hang already known on this machine and not a fault of the model or the route. We proved the route with a direct client instead, and we say that rather than reporting a pass.

Every claim on this page is dated 2026-08-14, except the context note in section 04. Community quantizations are re-cut, runtimes change defaults, and the flag at the center of section 05 has already changed its default twice. If today is much later than that date, treat the version-specific findings as unverified since then.

10

Related pages on this site

The speculation-gate study is section 05's thesis measured properly: five settings, 160 paired measurements, on a 397B mixture-of-experts on this same workstation. Empty Answer is section 06's: what happens when hidden reasoning bills against the same budget as the visible answer, and how to find your own floor instead of trusting a number. The Qwen3.5 pair field card is the measured record of the two larger Qwen models this machine serves, and the comparison points in section 04 come from it. The Qwen3.8 release watch is the dated record of what shipped and when, including the Max-class weights and their different license. Reasoning effort is a sentence, not a dial takes section 07's reasoning_effort finding apart on this same model file, reading the injected system message directly out of the rendered prompts and tracing where the setting reaches the chat template and where it does not.

11

Sources and artifacts

The evidence for this page ships as a data package: the files the run actually produced, not a summary. data/README.md maps every folder, states its provenance, and states the package's redactions plainly; data/NUMBERS.md maps every figure on this page to its file and field, and it opens with the one place a file will look like it contradicts the page.

  • The official Qwen organization on Hugging Face, checked 2026-08-14: Qwen/Qwen3.8-27B at revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0 and Qwen/Qwen3.8-27B-FP8 at 017b9c7af6b5689d5dd426a76e0bc077eb5ca20a, both with weight files present, both flipped public that day by their own timestamps.
  • The three LICENSE files themselves, read in full and retained under data/licenses/: Qwen/Qwen3.8-27B at its revision on 2026-08-14 (11,544 bytes), Qwen/Qwen3.8-27B-FP8 at its own revision on the same day (11,544 bytes, verified byte-identical to the 27B's), and Qwen/Qwen3.8-2.4T-A95B at revision 207bd685 on 2026-08-13 (3,390 bytes).
  • The model's config.json and chat template at the pinned revision, for every architecture, tokenizer and template fact in sections 01, 06 and 07. The served template ships inside data/routecheck/raw/C1_server_props.json, which is where section 07's template readings can be checked.
  • unsloth/Qwen3.8-27B-GGUF at revision fe1e2a23d973adb629709749dc4f6756df66ef10, and the repository's own checksums for that revision, against which both downloaded files were verified before first load (data/integrity/).
  • A retained GGUF header probe of the Unsloth and ggml-org files, read over HTTP range requests, for the tensor counts, the nextn counts, the quantization census and the sampling and license metadata quoted above (data/headers/). It is a same-day re-run retained after the fact, covering two of the four candidates the original probe read; the README says so.
  • Our own recorded runs of 2026-08-14, 19:14 to 21:07 EDT: the seven-load profile ladder (data/ladder/ladder.tsv, its driver, the seven server logs carrying the draft counters, and every request and response body under data/ladder/raw/), the probe battery (data/battery/), the field-card instrument's run and all 27 of its bodies (data/routecheck/), the load logs (data/loads/), and the production-route two-hop tool proof (data/prod/).

If a number on this page disagrees with the record it came from, the record is right and the page is wrong; tell us and we will fix the page.