Reasoning effort is a sentence, not a dial
On the chat template shipped inside the Qwen3.8-27B GGUF we served, the xhigh / medium / low setting is a system sentence added to your prompt: xhigh adds 42 tokens, low adds 30, medium adds none at all, and leaving the field out is byte-identical to xhigh. Read directly off the rendered prompts on a local 27B, three days after a hosted 2.4T-A95B reported the same three prompt-token counts with the mechanism invisible, plus the llama.cpp build where the top-level field started reaching the template.
Most model APIs now take a reasoning_effort field, and it
presents itself as a dial: high, medium, low, turn it up for hard problems and down for
cheap ones. On the Qwen3.8 template we served it is not a dial. It is a sentence. The
value you send is looked up in the chat template that ships inside the model file, and
the template pastes a matching instruction into your prompt as a system message before
the model ever sees your question. Setting
xhigh adds 42 tokens of system message telling the model to think
carefully, validate its assumptions and consider alternatives. Setting low
adds 30 tokens telling it to keep its thinking brief. Setting medium adds
nothing at all: there is no medium branch in the template, so medium is the absence of
an instruction rather than a middle amount of one. And omitting the field entirely
renders a prompt that is byte-identical to xhigh, because the template's
default is xhigh, so there is no way to turn the setting off by leaving it
out. Two of those facts fight each other in plain English, and it is worth holding both
at once before reading on: the value named medium gets you no instruction,
and sending no value at all gets you the maximum one. We directly measured the mechanism
once, on a local 27B. An earlier hosted run observed matching prompt-token counts on the
same message, but could not inspect prompt construction: on 2026-08-13, on the hosted
2.4-trillion-parameter Qwen3.8-2.4T-A95B, we could see only that the server reported
three different prompt-token counts, 112 and 70 and 100, for three sends of a
byte-identical message, and we wrote down that we did not know why. On 2026-08-16, on
the 27-billion-parameter Qwen3.8-27B running locally, we asked the server to render the
prompt and read it. The rendered prompts tokenize to 112, 70 and 100. The third finding
belongs to the runtime rather than the model: on the older llama.cpp build we tested,
the top-level OpenAI-style reasoning_effort field did not reach the
template at all, so every value rendered the default prompt, and on the newer build it
reaches the template and an unrecognized value returns HTTP 500 while thinking is on.
The run log attributes that change to pull request #26941, reported merged on 2026-08-14
and first shipped in release b10434; we measured
b10290 and b10453, one build on each side
of it.
Summary
The reasoning_effort field, on the Qwen3.8 build we served, is
implemented in the chat template that ships inside the model file, not in the sampler
and not in the serving runtime. The template maps the value to a string it calls
reasoning_instructions and injects that string as a system message ahead
of your own message. That single implementation detail explains everything below.
Everything in this section describes a server that applies the GGUF's own chat
template. On llama.cpp that means --jinja with no
--chat-template or --chat-template-file override, which is the
configuration both of our builds ran and which section 03 states in full. A server that
substitutes a built-in template renders a different prompt, and none of these numbers
describe it.
- The three levels inject three different amounts of text, and the amounts
are fixed. On the served template,
xhighinjects a system block of 42 tokens,lowinjects 30, andmediuminjects zero. Those deltas were identical on all three test prompts, whose lengths differ, so the injection is a constant added to your prompt rather than anything proportional to it. - Medium is not a middle setting. It is the absence of a setting.
The template has an
xhighbranch and alowbranch and nomediumbranch, somediumfalls through with the instruction string empty and the rendered prompt carries no system message at all. - Omitting the field is identical to sending
xhigh. The template's default isxhigh, so the rendered prompts for "field absent" and forxhighwere byte-identical, and the generations they produced were byte-identical too. There is no "unset" state. - What moved most consistently was answer style, not thinking
volume. On the local 27B,
lowproduced the least thinking on all three prompts (62% to 73% of whatxhighspent on the same prompt), butmediumwas erratic: 85% and 97% ofxhighon two prompts and 3.3 timesxhighon the third.xhighgave the tersest answer on every prompt andmediumthe longest, by factors of roughly 2 to 3. These are three prompts at one observation each (section 04), so read them as three readings, not as a rate. The hosted 2.4T leg had shown the same style direction three days earlier. - On the hosted 2.4T-A95B, thinking spend did not order by
effort. One call per level: 184 reasoning tokens at
xhigh, 197 atmedium, 168 atlow. A 17% band with the nominal maximum-effort setting in the middle of it, from three calls split across two providers. That is not a measurement of a small effect; it is a failure to see a large one. Section 07 carries the cost figures and the provider confound that comes with them. - The same three integers appeared on both legs. The hosted study recorded server prompt-token counts of 112 / 70 / 100 for xhigh / medium / low on a marble-probability prompt and wrote "mechanism unknown". Our locally rendered prompts for the same three levels on the same message tokenize to 112 / 70 / 100. Two different models on two different serving stacks: the hosted counts are an unexplained observation that is consistent with the injection we can read locally, and nothing more. This is a convergence of two independent readings, not a replication, and section 07 says exactly what it does and does not license.
- On llama.cpp, whether the field reaches the template at all depends on
your build. On the older build we tested, b10290, the
OpenAI-compatibility layer handled the value
noneand left every other value unhandled, under a source comment saying those values are model-specific and not yet handled. The effect on the wire is that every level rendered the default prompt:xhigh,medium,low, absent and even the nonsense valueultraall rendered the same 112-token prompt. On the newer build, b10453, the field reaches the template, the three levels render three different prompts, andultrareturns HTTP 500. Those are the two builds we ran. The run log attributes the change to pull request #26941, reported merged on 2026-08-14 and first shipped in b10434; section 08 says which parts of that we measured and which parts we are repeating. - There are two paths into the template, and they disagree about one
word. Passing effort inside
chat_template_kwargsalways worked, on both builds, because it is generic passthrough. Passing it at top level is the path the merge added. The wordnoneis intercepted by the server's OpenAI-compatibility handler on the top-level path (it turns thinking off, HTTP 200) and passed straight through to the template on the kwargs path, where it is not a valid effort and raises (HTTP 500). Same word, two paths, opposite outcomes, on both builds.
Everything on this page is one model family, two deployments, two builds, on the dates stated. The local leg is a single observation per cell; section 04 explains why, and why we think it is still worth reading.
What the setting actually is
The clearest way to state this is to show the prompt the server built. This is the
local 27B's own render of the same user message under three effort values, retrieved
from the server's /apply-template endpoint, on a server started with
--jinja and no template override, so the template being rendered here is
the one inside the GGUF.
/apply-template is not the endpoint you generate through, so we checked
it against the one you do. On the four marble conditions we ran both ways, the
prompt-token count the server reported for a real /v1/chat/completions
generation is the same as the token count of the retained render: 112 for
xhigh, 70 for medium, 100 for low, 112 for the
field omitted. The two endpoints also share their prompt-construction call in the source
of the build we ran, but the matching counts are the part we can hand you as a
record.
reasoning_effort: "xhigh" renders
<|im_start|>system Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer.<|im_end|> <|im_start|>user A bag holds 3 red, 4 blue and 5 green marbles. ...<|im_end|> <|im_start|>assistant <think>
reasoning_effort: "low" renders
<|im_start|>system Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration.<|im_end|> <|im_start|>user A bag holds 3 red, 4 blue and 5 green marbles. ...<|im_end|> <|im_start|>assistant <think>
reasoning_effort: "medium" renders
<|im_start|>user A bag holds 3 red, 4 blue and 5 green marbles. ...<|im_end|> <|im_start|>assistant <think>
That is the whole finding in three blocks. There is no system message in the medium
render. Not a shorter instruction: none. And the render with the field omitted
entirely is byte-identical to the xhigh render, because the
template line that resolves the value reads, in the shipped jinja,
reasoning_effort|default('xhigh').
The template also normalizes high to xhigh before the
lookup, and validates: any value that is not one of xhigh,
medium or low after normalization calls
raise_exception. That validation is what produces the HTTP 500 in section
08, and it sits inside the template's thinking branch, which is why it disappears when
thinking is off (section 09).
One detail worth stating plainly because it changes how you read the numbers: the 42
and 30 token figures are the size of the entire injected system block,
the <|im_start|>system and <|im_end|> role markers
included, not of the English sentence alone. They are the difference between what your
prompt costs at that setting and what it costs at medium, which is the
number a practitioner is actually billed for.
Environment
Local leg, 2026-08-16
| Model | Qwen3.8-27B, unsloth/Qwen3.8-27B-GGUF, UD-Q5_K_XL quantization, already on disk from a prior run, SHA-256 verified 176a6a3f… |
|---|---|
| Chat template | The one embedded in that GGUF (9,993 characters, carrying Unsloth's "Unsloth fixes" marker). This matters: the high to xhigh normalization is an Unsloth addition and is absent from the ggml-org conversion's template. |
| Runtime | llama.cpp release b10453, commit 3cb7ffb1a1f612d5e4a46244ae5a3c77ad934a70, published 2026-08-16, built fresh for this study in a scratch tree. The production build on this machine was read and never modified. |
| Control runtime | llama.cpp b10290, commit c8e03ce, the pre-merge build, run read-only on a second scratch port with zero GPU layers for the render comparison in section 08. |
| Serving flags | --jinja, ctx 32,768, --n-gpu-layers 999, --flash-attn on, --parallel 1, loopback only, no speculative decoding of any kind. No --chat-template and no --chat-template-file: nothing overrode the template inside the GGUF. The control build ran the same flags with --n-gpu-layers 0 and ctx 2,048. |
| Template actually applied | The GGUF's own jinja, on both builds, and each build's server log says so. An invalid effort produces a Minja traceback quoting line 64 of the served template, raise_exception('Unexpected reasoning effort ...'). That validator exists only in this template; llama.cpp's built-in templates carry nothing like it. Both logs also print chat template supports preserving reasoning at startup, which is capability detection over the template that was loaded. |
| Thinking | On, by the template's own default. No reasoning-related server flag was passed, and no enable_thinking key appears in any of the 36 retained request bodies; the template takes its thinking branch when enable_thinking is undefined. Every render on this page ends at <think> except the three thinking-off ones in section 09. |
| Hardware | One RTX 5090, 32,607 MiB. Load used 21,856 MiB. Decode ran at roughly 65 tokens per second throughout; the study counts tokens, not speed. |
| Per call | max_tokens: 2000, temperature: 0, seed: 42, effort sent as the top-level reasoning_effort field. |
Speculative decoding was deliberately excluded: llama.cpp issue #27122 reports reproducible CUDA lockups with the multi-token-prediction path on this model class on recent builds. The server log confirms the draft head was loaded but unused.
Those two flag rows are load-bearing for everything this page says about the
template. The finding is that a specific string inside this GGUF's jinja gets pasted
into your prompt. If your server does not apply that jinja, because it substitutes a
built-in template or was started without --jinja, the mechanism described
here is not what your server is doing, and the token counts below are not your token
counts.
Hosted leg, 2026-08-13
| Model | Qwen3.8-2.4T-A95B, a mixture-of-experts model, served through OpenRouter |
|---|---|
| Providers | Not pinned. The three effort calls landed on DeepInfra (xhigh, medium) and Together (low). DeepInfra listed quantization fp4; Together listed none. This is a confound and is treated as one in section 07. |
| Request shape | OpenRouter's reasoning: {"effort": ...} object, temperature: 0, max_tokens: 2000, one call per level |
| Template actually applied | Unknown. We did not see either provider's serving configuration and could not. No template claim on this page extends to the hosted leg. |
| What was visible | Server-reported token counts, cost, and the returned text. Nothing about prompt construction. |
The two legs are different models on different serving stacks. Nothing below treats one as a check on the other in the sense of a replication.
The three request shapes
One setting, three body shapes, and the study touched all three. The local leg sent the field at the top level of an OpenAI-shaped body, which is the path section 08's build boundary is about:
{"model": "...", "messages": [ ... ], "reasoning_effort": "low",
"temperature": 0, "seed": 42, "max_tokens": 2000}
llama.cpp also accepts the same value inside its own passthrough object, which is the path that worked on both of our builds:
{"model": "...", "messages": [ ... ], "chat_template_kwargs": {"reasoning_effort": "low"}}
The hosted leg went through OpenRouter, which takes its own object:
{"model": "...", "messages": [ ... ], "reasoning": {"effort": "low"}}
We did not test whether any of these servers accepts the others' shapes, so treat the shapes as tied to the stacks we sent them to. If you are debugging a field that appears to do nothing, the shape your stack expects is the first thing to check and the cheapest.
The experiment, and the limitation that came out of it
The local run was pre-registered: the question, the grid, the expectations and the honesty rules were written and saved before the first model call, and the source-level reading of what the llama.cpp merge changed was done before the run so the prediction could be scored rather than retrofitted. The run log carries all of it, including the expectations that were confirmed and the deviations that were not planned.
The grid: three prompts by four effort levels by three repetitions, 36 calls.
- The marble prompt is the hosted study's own prompt, carried over verbatim so that one cell of the local grid is the same message the hosted leg sent.
- Two further prompts came from our standing probe set: a word problem with a checkable integer answer, and a short prose question with no single right answer.
- The four levels are
xhigh,medium,low, and absent, the last meaning the field is omitted from the request body entirely.
Separately, 16 rendered prompts were retrieved and retained: the
four levels on the top-level path and three of them on the
chat_template_kwargs path, since there is no kwargs spelling of an omitted
field, plus the special values high, none and
ultra on both paths, plus three thinking-off renders. Those renders are the evidence in sections 02, 05, 08 and 09, and
none of them required a single generated token.
At temperature: 0 with a fixed seed the server was fully
deterministic. All three repetitions of all twelve cells returned byte-identical
visible text and byte-identical reasoning text: 12 of 12 cells. The
three repetitions therefore measured the server's determinism, not the model's
variance, and the effective observation count per cell is one. Every
generation number on this page is a single observation and is labeled as one. Nothing
here has an interval, and no language on this page implies a distribution. Getting a
spread would need a temperature above zero or varied seeds; that run has not
happened.
The renders are not affected by this. A rendered prompt is a deterministic function of the request, and section 05's injection sizes are exact counts of text, not samples of behavior.
All 36 calls returned HTTP 200 with a stop reason of
stop. Nothing was truncated: the largest completion was 651 tokens against
a 2,000-token cap. No reasoning text leaked into the visible answer in any cell. Zero
retries, zero errors.
Observed: the injection is a fixed-size block
Prompt tokens for the same four levels, on three prompts of different lengths. The counts come from the server's own tokenizer applied to the rendered prompt.
| Prompt | xhigh | field absent | low | medium | xhigh minus medium | low minus medium |
|---|---|---|---|---|---|---|
| Marble probability | 112 | 112 | 100 | 70 | +42 | +30 |
| Word problem | 111 | 111 | 99 | 69 | +42 | +30 |
| Prose question | 99 | 99 | 87 | 57 | +42 | +30 |
Every cell is the prompt-token count the server reported for the generations. The marble row is also confirmed independently by the retained renders, tokenized by the server itself (section 02); the other two prompts were measured through generations only and were not rendered separately.
Three readings of the same two constants. The injected block does not scale with
your prompt: it is 42 tokens for xhigh, 30 for low, and 0 for
medium, added on top of whatever you sent. In characters, the same three
renders of the marble prompt measure 566, 495 and 329, so xhigh adds 237
characters of system message and low adds 166.
The xhigh and absent columns are not merely equal in count. On the
marble prompt the two rendered prompts are byte-identical strings, and the two
generations they produced are byte-identical too: same reasoning text, same visible
answer, same token counts, on all three prompts. If you are choosing between "send
xhigh" and "leave it out", you are choosing between two spellings of the same
request.
Observed: what it does to thinking, and to answers
Single observation per cell. Reasoning tokens are counted exactly, by sending the returned reasoning text back to the server's own tokenizer, rather than trusting a reported field.
| Prompt | Effort | Reasoning tokens | Answer characters | Answer shape |
|---|---|---|---|---|
| Marble probability | xhigh | 399 | 284 | bare LaTeX, boxed result |
| field absent | 399 | 284 | identical to xhigh | |
| medium | 339 | 797 | markdown headers, a table | |
| low | 283 | 711 | markdown headers, a table | |
| Word problem | xhigh | 126 | 134 | two sentences |
| field absent | 126 | 134 | identical to xhigh | |
| medium | 122 | 345 | numbered steps in bold | |
| low | 78 | 148 | three short lines | |
| Prose question | xhigh | 123 | 346 | three sentences |
| field absent | 123 | 346 | identical to xhigh | |
| medium | 405 | 686 | one long paragraph | |
| low | 90 | 391 | two long sentences |
Every row is one observation. The three repetitions of each cell were byte-identical (section 04).
On thinking volume, low behaves and medium does
not. low spent less than xhigh on every prompt, at
71%, 62% and 73% of the xhigh figure. medium spent 85% on one
prompt, 97% on another, and 329% on the third: on the prose question,
the level with no injected instruction at all produced 3.3 times the thinking of the
level that explicitly asks for careful thought. That is not a middle setting
misbehaving. It is what you would expect if medium simply removes an
instruction and lets the model do whatever it does unprompted.
On answer style, the direction held on all three prompts.
xhigh produced the shortest visible answer in each of the three, and
medium the longest in each, by factors of 2.8, 2.6 and 2.0. Three prompts,
one observation each. The direction is the same one the hosted 2.4T leg showed three
days earlier: there, xhigh returned 340 characters of terse mathematics
while medium returned 720 characters of markdown with headers.
There is a readable mechanism for that, and it is in the injected text itself: the
xhigh sentence ends by asking for "correctness, consistency, and clarity in
the final answer", and the low sentence asks the model to move "directly to
the conclusion without unnecessary elaboration". Both of those are instructions about
the answer as much as about the thinking. The setting is named for reasoning effort and
reads, in practice, as a house-style instruction. We are reading the injected sentence
here, not instrumenting the model, and the distinction matters: this is an
interpretation of a measured text, offered as one.
Correctness: the two prompts with a checkable answer were right in all eight of their cells (the marble probability came back 3/11 with 60 favourable outcomes and 220 total; the word problem came back 27). The third prompt is open prose and has no key, so we make no correctness claim about it.
Observed: the hosted 2.4T leg, and what converges
Three days before the local run, the same marble prompt went to the hosted Qwen3.8-2.4T-A95B through OpenRouter, one call per level, same temperature and same output cap. What the hosted study could see was this:
| Effort | Provider | Prompt tokens | Reasoning tokens | Answer characters | Completion tokens | Cost |
|---|---|---|---|---|---|---|
| xhigh | DeepInfra | 112 | 184 | 340 | 356 | $0.00236 |
| medium | DeepInfra | 70 | 197 | 720 | 582 | $0.003632 |
| low | Together | 100 | 168 | 527 | 505 | $0.00341 |
Single observation per level, 2026-08-13. Cost is the server's own reported figure, not a computation.
The thinking-volume dial did not appear. Reasoning spend across the three levels was 168 to 197 tokens, a 17% band, with the nominal maximum-effort setting sitting in the middle of it. The study's pre-registered expectation had been that spend would order xhigh above medium above low. It did not.
The cost ordering inverted the knob's apparent intent: medium cost
the most, xhigh the least, because the visible answer grew as the nominal effort fell
and output tokens are what you pay for. State that with its confound attached: the
low call landed on Together, whose listed prices were higher than
DeepInfra's, so the medium-versus-low ordering is not clean. The xhigh-versus-medium
comparison is clean, both being DeepInfra calls, and on that pair the lowest nominal
effort setting of the two cost 54% more.
And the three prompt-token counts had no explanation. The message bytes were identical across the three calls, verified from the retained request bodies, yet the server reported 112, 70 and 100 prompt tokens. The study recorded them and wrote, in the finding itself, "Mechanism unknown; recorded as evidence the effort level rewrites the served prompt."
Our locally rendered prompts, for the same three levels on the same message, tokenize to 112, 70 and 100.
These are two different models on two different serving stacks: a 2.4-trillion-parameter mixture-of-experts model served by two commercial GPU clouds through an aggregator, and a 27-billion-parameter dense model in a community-quantized GGUF served by llama.cpp on a desktop. We did not read the hosted providers' prompt construction and cannot. This is a convergence of two independent readings, not a replication. What it supports is that the hosted numbers are consistent with the same template mechanism we can read locally, which is a good deal more than the hosted study could say on its own and a good deal less than proof of what DeepInfra's server does. The two models are the same family and, on this evidence, tokenize this prompt identically, which is the premise the convergence rests on and is itself an inference from the matching integers rather than something we verified independently.
The style direction converged too: terse at
xhigh, long and formatted at medium, on both legs. The
thinking-volume result did not fully converge. Hosted spend was flat across all three
levels; locally, low came in below xhigh on each of the three
prompts while medium was erratic. Two observations of one cell each
disagreeing about a
small effect is exactly what you would expect whether or not the effect is real, and we
draw nothing from it.
The version boundary: llama.cpp b10290 and b10453
The mechanism above lives in the chat template, so it applies wherever that template is the one being rendered. Whether the field ever reaches the template is a property of your runtime, and on llama.cpp that changed recently.
Be precise about which half of this we measured. Our renders show different behavior
on b10290 and on b10453, and both builds
were started with the same flags, --jinja and no template override. What we
are repeating rather than measuring is the release mapping: the run log reports that
pull request #26941 merged on 2026-08-14 and first shipped in release
b10434, read from the project's own records. We did not build
b10434, and we tested nothing between b10435 and b10453. Confirm that mapping against
your deployed build rather than taking the number from this page.
The two builds do differ in the way that pull request describes. On
b10290 the OpenAI-compatibility layer handles the value
none and leaves every other value unhandled, under a source comment saying
in full that other reasoning_effort values are model-specific and not yet
handled. That reads as an incomplete compatibility stub rather than a defect: on that
build there was no Qwen effort vocabulary on that field for the server to forward. The
user-visible consequence is the same either way, and it is what the table shows. On
b10453 the value is forwarded into the template's variable
context.
We measured both ends rather than taking the diff's word for it. The same model file,
the same prompt, the same flags and the same render conditions were sent to both builds,
the pre-merge one run read-only on a separate scratch port with zero GPU layers.
Every cell below is an /apply-template render, tokenized
by that build's own tokenizer. Exactly three cells differ:
| Value sent | Path | Pre-merge b10290 | Post-merge b10453 | |
|---|---|---|---|---|
| xhigh | top-level field | 112 tok | 112 tok | unchanged by coincidence: read the note below |
| medium | top-level field | 112 tok | 70 tok | changed |
| low | top-level field | 112 tok | 100 tok | changed |
| field absent | top-level field | 112 tok | 112 tok | |
| xhigh / medium / low | chat_template_kwargs | 112 / 70 / 100 | 112 / 70 / 100 | unchanged: this path always worked |
| high | either path | 112 tok | 112 tok | normalized to xhigh by this template |
| none | top-level field | 72 tok | 72 tok | handled pre-merge too |
| none | chat_template_kwargs | HTTP 500 | HTTP 500 | not a valid template value |
| ultra | top-level field | 112 tok | HTTP 500 | changed: HTTP 200 with the default prompt became HTTP 500 |
| ultra | chat_template_kwargs | HTTP 500 | HTTP 500 |
Both columns are renders of the same prompt through the same model
file, retrieved from each build's /apply-template endpoint and tokenized by
that build's own tokenizer.
Read the xhigh row carefully, because it is the trap in this table. On
the pre-merge build xhigh rendered 112 tokens, the same as the post-merge
build, and that is a coincidence rather than support. Pre-merge, every
top-level value rendered 112 tokens, because the field never reached the template and
the template fell back to its xhigh default. A user who sent
xhigh to that build got the right prompt for the wrong reason. A user who
sent low got the xhigh prompt and no error. A user who sent a
typo got the xhigh prompt and no error.
Which endpoint produced these cells. Every cell in the table is a
render from /apply-template on the build named in its column. On the
post-merge build we also sent the invalid value through
/v1/chat/completions itself, both at the top level and inside
chat_template_kwargs, and both returned HTTP 500 there too, so that result
is not an artifact of the render endpoint. We did not run the pre-merge build through
/v1/chat/completions: its column is a render comparison only.
Upstream issue #27023, "Misc. bug: reasoning_effort seems broken", has been open since 2026-08-13, and a webui proposal (#27118, 2026-08-15) asks for effort and thinking budget to become two separate settings. The boundary above is consistent with confusion from both directions: before the merge the field did nothing, and after it the field does something and rejects values that used to pass. We make no claim about what any particular reporter experienced.
If you run llama.cpp with this GGUF's template applied, and thinking is on, and
your client sends a top-level reasoning_effort that the template does not
recognize, then crossing this build boundary turns a value that had no effect into an
HTTP 500. All three conditions are load-bearing: with thinking off the same request
returns HTTP 200 (section 09), and with a built-in template applied instead it is a
different question that this page did not measure. Anything a client library sends by
default, including vocabulary borrowed from other vendors' APIs, is worth auditing
before that upgrade.
Special values, two paths, and one escape hatch
Three behaviors are worth separating out, because each of them will bite somebody.
none means two different things depending on how you send
it. Sent as the top-level field, it is intercepted by the server's own
compatibility handler before the template is reached: thinking is turned off, the render
carries a closed empty thinking block, and a generation returns an answer with zero
reasoning tokens. That render is the medium render plus that block and
nothing else, which is where its 72 tokens come from: 340 characters against medium's
329, 72 tokens against medium's 70. The two extra tokens are the closed block, not an
instruction. Sent inside chat_template_kwargs, it goes straight into
the template, which does not list none among its valid values, and the
template raises: HTTP 500. Same word, two paths, opposite outcomes, on both builds we
tested.
high is normalized, but that normalization belongs to the
quantizer's template, not to the standard. The template in the file we served
maps high to xhigh before validating, so high
renders the xhigh prompt on both paths and both builds. That normalization
line is an Unsloth addition. The ggml-org conversion of the same model carries a
template without it. We did not serve the ggml-org file, so we make no measurement claim
about it, but on a plain reading of that template text a high sent to it
would fail validation the way ultra does. Treat this as a prediction to
test, not a result.
Invalid values fail loudly, and the failure is conditional on thinking being on. An unrecognized value returns HTTP 500 with the template's own traceback, which names the supported set in plain text:
Error: Jinja Exception: Unexpected reasoning effort ultra. Supported types are xhigh (default), medium, and low.
That is an unusually good error message, and it is worth knowing it can appear at
all. But the validation lives inside the template's thinking branch, so it is
unreachable when thinking is off. We tested that: with thinking turned off, the invalid
value ultra returned HTTP 200 and a render byte-identical to the
thinking-off render without it. On this model the thinking switch belongs inside
chat_template_kwargs, as {"enable_thinking": false}; a
top-level enable_thinking field was measured as ignored on this endpoint
two days earlier. One retention note, because this page asks you to check its records:
the three thinking-off entries retain the rendered prompt but not the request body, so
the render is the artifact and the key used is the run log's account of it.
The consequence is worth stating in both directions. If your integration disables thinking, an invalid effort value passes silently. If somebody later turns thinking back on, the same request starts returning HTTP 500.
We tested one invalid value, ultra. The template's validation rejects
anything outside xhigh, medium and low after
normalization, so on a plain reading of that source any other unrecognized value takes
the same branch. That is a read of the template, labeled as such, not a second
measurement. The practical point is that clients send what they send, and a value
another vendor's API accepts is not thereby a value this template accepts. The only way
to know is to send it at the template and runtime you actually deploy and look at what
comes back.
What this does not show
- Every generation figure here is a single observation. Temperature 0 with a fixed seed collapsed three repetitions into one, in 12 cells out of 12. No intervals, no significance, no distribution. Differences of a few percent between cells mean nothing on this page; the ones we lean on are large (a 3.3-times ratio, a 2-to-3-times answer length ratio, an exact zero in the medium render).
- One template, applied one way. Both builds ran
--jinjawith no template override, so every template result here describes a server that applies the GGUF's own jinja; we did not measure what happens with a built-in template substituted for it. And the jinja we read is the Unsloth conversion's. We did not serve the ggml-org conversion, whose template differs at least in thehighnormalization (section 09), and we never read the 2.4T model's template at all. - Two deployments, not a survey. One hosted 2.4T-A95B through one aggregator on 2026-08-13, split across two providers without pinning, and one local 27B in one community quantization on one desktop on 2026-08-16. Other quantizations ship other templates. Other runtimes handle the field their own way.
- The hosted mechanism was not read, only inferred. We do not know how DeepInfra or Together construct the prompt. The match of three integers is consistent with the template mechanism and is not proof of it.
- Three prompts. Two with checkable answers, one prose. The prompts are short. Whether the levels separate more on genuinely hard problems, where a reasoning instruction has more to bite on, is unmeasured and is the obvious next run.
- We did not instrument the model. Section 06's reading of why answer style moves more than thinking volume is an interpretation of the injected sentence's wording, not a mechanistic result.
- No comparison of models and no quality ranking. Nothing here says one model reasons better than another, and the two legs are not scored against each other.
- Other stacks were not examined. We did not test vLLM, TensorRT-LLM, text-generation-inference, or the vendors' own first-party endpoints, and nothing here speaks to how they handle this field.
- Dates matter more than usual on this page. The llama.cpp behavior changed two days before we measured it, and an open upstream issue and an open proposal both touch this surface. Treat every version-specific statement as unverified after 2026-08-16.
If you send reasoning_effort to Qwen3.8
Seven narrow things, each scoped to what is above, and all of them assuming a server that applies this GGUF's own chat template.
- Know that you are writing a sentence into your prompt. On this model the field is not a sampler control. It prepends a system message. If you already send your own system message about how the model should work, you are now sending two, and the template's one comes first.
- Send
mediumwhen you want no instruction at all. This is the counterintuitive one.mediumis the only value that leaves your prompt alone. It is also the cheapest of the three on input tokens, by 42 tokens againstxhigh. What it is not is a middle amount of effort. - Do not treat omitting the field as neutral. It renders as
xhigh, byte for byte. If your intent is "do not touch my prompt", omitting the field does the opposite of what you want andmediumis the value you are looking for. lowbought less thinking on our three prompts, and a different answer with it. In these three deterministic local prompts,lowused 62% to 73% as many reasoning tokens asxhigh. Test that pattern on your own prompts before relying on it. It also made answers longer and more heavily formatted thanxhighdid, which may be the opposite of what a trim-the-thinking setting is being reached for.- To turn thinking off, send
noneat the top level, and never putnonein the kwargs. This is the most useful single fact on the page for anyone watching a bill. Top-levelnoneis intercepted by the server before the template: thinking off, HTTP 200, zero reasoning tokens. The same word insidechat_template_kwargsreaches the template, which does not list it as a valid effort, and returns HTTP 500. This model also has a real thinking switch of its own,chat_template_kwargs {"enable_thinking": false}, which we measured on this endpoint two days before this study. - On llama.cpp, know which side of the build boundary you are on.
On the older build a top-level
reasoning_effortother thannonenever reached the template and every request rendered the default. On the newer one the value takes effect, and an unrecognized value returns HTTP 500 while thinking is on. Before upgrading across that boundary, four things are worth auditing: what your clients actually put in that field, including values borrowed from other vendors' vocabularies; whether thinking is on, since the HTTP 500 needs it; whether your server applies the GGUF's template or a built-in one; and which of the request shapes in section 03 your code sends. A setting that has been quietly doing nothing in your stack starts doing something the day you rebuild. - Pick one path and know its quirks. The top-level field is the
standard-shaped one and handles
noneas thinking-off. Thechat_template_kwargspath worked on both builds and passesnonethrough to a template error. If you need one code path that behaves the same on old and new llama.cpp builds,chat_template_kwargsis that path, with the caveat that it is a llama.cpp-specific request shape.
This page cannot tell you what vLLM, text-generation-inference or a vendor's own
endpoint does with this field, because we did not test them. It can tell you how to
find out in two requests. Send the same message twice, once with the field at
xhigh and once at medium, and compare the server's reported
prompt_tokens. If the two counts differ, something in your stack is
rewriting your prompt from that field, and the difference is the size of what it
injects. If they are identical, the field is not reaching a template that does
anything with it, and whatever behavior you are attributing to it is coming from
somewhere else. On llama.cpp you can go one step further and read the prompt itself
out of /apply-template.
And one thing to measure if the levels matter to you: whether
they separate on your prompts. On ours they mostly changed the shape of the answer. Send
your real prompts at xhigh and at medium, count reasoning
tokens and read the two answers side by side. It is a cheap test and it will tell you
which of the two effects you are actually buying.
Artifacts
The complete evidence for this page ships as a data package: the files the two runs actually produced, not a summary. If a number on this page disagrees with a file in the package, the file is right and the page is wrong; tell us and we will fix the page.
- data/README.md: what is in each folder, the record schema, and the provenance of every file.
- data/NUMBERS.md: every figure printed on this page, with the file and the field it came from.
- The pre-registration and run log (data/local/RUN_LOG.md) for the local study, written and saved before the first model call, including the source-level reading of the llama.cpp change made before the run, the expectations it predicted, and the deviations logged as they happened, including the determinism finding that reduced the effective observation count to one per cell.
- The 16 rendered prompts, verbatim, every condition in sections
02, 05, 08 and 09 (
data/local/raw/render_*.json). These are the primary evidence: the rendered text is the finding. - The 36 per-call records from the grid
(
data/local/raw/grid_*.json), each carrying the full request and response body, the server's token counts and timings, the exact reasoning-token count, and the wall time, plus the summarized grid (data/local/effort_grid.tsv). - The cross-build comparison
(data/local/merge_ab_comparison.txt)
against the pre-merge llama.cpp build, plus that build's server log and the 13
pre-merge renders (
data/local/raw_premerge/). - The hosted study's effort records from 2026-08-13
(
data/hosted/raw/effort_*.json): the three request bodies proving the messages were byte-identical, and the three response bodies carrying the provider name, the token counts and the server-reported cost. - The release verification (data/local/release_verification.txt) taken at the end of the local run: GPU memory, ports, and process state.
| 2026-08-13 | Hosted battery on Qwen3.8-2.4T-A95B through OpenRouter, including the three-level effort sweep on the marble prompt. One call per level; providers recorded per call and not pinned. The three unexplained prompt-token counts were recorded that day as "mechanism unknown". |
|---|---|
| 2026-08-14 | llama.cpp pull request #26941 merged, reported as first carried in release b10434: the top-level reasoning_effort field begins reaching the chat template. This row is the project's record, not our measurement. |
| 2026-08-16 | Local study on Qwen3.8-27B: pre-registration written before the first call, 36 generations and 16 renders on the post-merge build b10453, then the same renders on the pre-merge build b10290 with zero GPU layers. No unexpected operational failures or exclusions; the planned invalid-value probes returned the documented HTTP 500 results. GPU released and verified clean. |
Every claim on this page is dated. The llama.cpp behavior described in section 08 changed on 2026-08-14 and was measured on 2026-08-16; if today is much later, treat it as unverified since that date.
Related pages on this site
Qwen3.8-27B: day zero is the release-day field card for
the local model measured here, on the same desktop.
Release watch: Qwen3.8 tracks the family's releases
and licenses, including the split between the Apache-licensed 27B and its
differently-licensed 2.4T sibling.
Muse Glimmer 30B is the same class of knob on another
vendor's model, measured the same day on the same desktop: there the lever is called
reasoning_strength, the template interpolates the value into a system line
without validating it, and an undefined value was accepted without complaint.
Empty Answer is the same editorial thesis in the
neighbouring domain: what a reasoning model does with your output budget when the
thinking does not fit. The Qwen hub page carries downloadable
serving profiles for the local models on this machine.