Graphometer
Measured study · hosted leg 2026-08-13: Qwen3.8-2.4T-A95B via OpenRouter, one call per effort level · local leg 2026-08-16: Qwen3.8-27B on one RTX 5090, 36 calls plus 16 rendered prompts read directly, two llama.cpp builds

Reasoning effort is a sentence, not a dial

On the chat template shipped inside the Qwen3.8-27B GGUF we served, the xhigh / medium / low setting is a system sentence added to your prompt: xhigh adds 42 tokens, low adds 30, medium adds none at all, and leaving the field out is byte-identical to xhigh. Read directly off the rendered prompts on a local 27B, three days after a hosted 2.4T-A95B reported the same three prompt-token counts with the mechanism invisible, plus the llama.cpp build where the top-level field started reaching the template.

Most model APIs now take a reasoning_effort field, and it presents itself as a dial: high, medium, low, turn it up for hard problems and down for cheap ones. On the Qwen3.8 template we served it is not a dial. It is a sentence. The value you send is looked up in the chat template that ships inside the model file, and the template pastes a matching instruction into your prompt as a system message before the model ever sees your question. Setting xhigh adds 42 tokens of system message telling the model to think carefully, validate its assumptions and consider alternatives. Setting low adds 30 tokens telling it to keep its thinking brief. Setting medium adds nothing at all: there is no medium branch in the template, so medium is the absence of an instruction rather than a middle amount of one. And omitting the field entirely renders a prompt that is byte-identical to xhigh, because the template's default is xhigh, so there is no way to turn the setting off by leaving it out. Two of those facts fight each other in plain English, and it is worth holding both at once before reading on: the value named medium gets you no instruction, and sending no value at all gets you the maximum one. We directly measured the mechanism once, on a local 27B. An earlier hosted run observed matching prompt-token counts on the same message, but could not inspect prompt construction: on 2026-08-13, on the hosted 2.4-trillion-parameter Qwen3.8-2.4T-A95B, we could see only that the server reported three different prompt-token counts, 112 and 70 and 100, for three sends of a byte-identical message, and we wrote down that we did not know why. On 2026-08-16, on the 27-billion-parameter Qwen3.8-27B running locally, we asked the server to render the prompt and read it. The rendered prompts tokenize to 112, 70 and 100. The third finding belongs to the runtime rather than the model: on the older llama.cpp build we tested, the top-level OpenAI-style reasoning_effort field did not reach the template at all, so every value rendered the default prompt, and on the newer build it reaches the template and an unrecognized value returns HTTP 500 while thinking is on. The run log attributes that change to pull request #26941, reported merged on 2026-08-14 and first shipped in release b10434; we measured b10290 and b10453, one build on each side of it.

Independent measurements; no affiliation with the Qwen models' provider (Alibaba), the llama.cpp project or ggml-org, Unsloth, OpenRouter, DeepInfra, Together, or OpenAI.
01

Summary

The reasoning_effort field, on the Qwen3.8 build we served, is implemented in the chat template that ships inside the model file, not in the sampler and not in the serving runtime. The template maps the value to a string it calls reasoning_instructions and injects that string as a system message ahead of your own message. That single implementation detail explains everything below.

Everything in this section describes a server that applies the GGUF's own chat template. On llama.cpp that means --jinja with no --chat-template or --chat-template-file override, which is the configuration both of our builds ran and which section 03 states in full. A server that substitutes a built-in template renders a different prompt, and none of these numbers describe it.

  1. The three levels inject three different amounts of text, and the amounts are fixed. On the served template, xhigh injects a system block of 42 tokens, low injects 30, and medium injects zero. Those deltas were identical on all three test prompts, whose lengths differ, so the injection is a constant added to your prompt rather than anything proportional to it.
  2. Medium is not a middle setting. It is the absence of a setting. The template has an xhigh branch and a low branch and no medium branch, so medium falls through with the instruction string empty and the rendered prompt carries no system message at all.
  3. Omitting the field is identical to sending xhigh. The template's default is xhigh, so the rendered prompts for "field absent" and for xhigh were byte-identical, and the generations they produced were byte-identical too. There is no "unset" state.
  4. What moved most consistently was answer style, not thinking volume. On the local 27B, low produced the least thinking on all three prompts (62% to 73% of what xhigh spent on the same prompt), but medium was erratic: 85% and 97% of xhigh on two prompts and 3.3 times xhigh on the third. xhigh gave the tersest answer on every prompt and medium the longest, by factors of roughly 2 to 3. These are three prompts at one observation each (section 04), so read them as three readings, not as a rate. The hosted 2.4T leg had shown the same style direction three days earlier.
  5. On the hosted 2.4T-A95B, thinking spend did not order by effort. One call per level: 184 reasoning tokens at xhigh, 197 at medium, 168 at low. A 17% band with the nominal maximum-effort setting in the middle of it, from three calls split across two providers. That is not a measurement of a small effect; it is a failure to see a large one. Section 07 carries the cost figures and the provider confound that comes with them.
  6. The same three integers appeared on both legs. The hosted study recorded server prompt-token counts of 112 / 70 / 100 for xhigh / medium / low on a marble-probability prompt and wrote "mechanism unknown". Our locally rendered prompts for the same three levels on the same message tokenize to 112 / 70 / 100. Two different models on two different serving stacks: the hosted counts are an unexplained observation that is consistent with the injection we can read locally, and nothing more. This is a convergence of two independent readings, not a replication, and section 07 says exactly what it does and does not license.
  7. On llama.cpp, whether the field reaches the template at all depends on your build. On the older build we tested, b10290, the OpenAI-compatibility layer handled the value none and left every other value unhandled, under a source comment saying those values are model-specific and not yet handled. The effect on the wire is that every level rendered the default prompt: xhigh, medium, low, absent and even the nonsense value ultra all rendered the same 112-token prompt. On the newer build, b10453, the field reaches the template, the three levels render three different prompts, and ultra returns HTTP 500. Those are the two builds we ran. The run log attributes the change to pull request #26941, reported merged on 2026-08-14 and first shipped in b10434; section 08 says which parts of that we measured and which parts we are repeating.
  8. There are two paths into the template, and they disagree about one word. Passing effort inside chat_template_kwargs always worked, on both builds, because it is generic passthrough. Passing it at top level is the path the merge added. The word none is intercepted by the server's OpenAI-compatibility handler on the top-level path (it turns thinking off, HTTP 200) and passed straight through to the template on the kwargs path, where it is not a valid effort and raises (HTTP 500). Same word, two paths, opposite outcomes, on both builds.

Everything on this page is one model family, two deployments, two builds, on the dates stated. The local leg is a single observation per cell; section 04 explains why, and why we think it is still worth reading.

02

What the setting actually is

The clearest way to state this is to show the prompt the server built. This is the local 27B's own render of the same user message under three effort values, retrieved from the server's /apply-template endpoint, on a server started with --jinja and no template override, so the template being rendered here is the one inside the GGUF.

/apply-template is not the endpoint you generate through, so we checked it against the one you do. On the four marble conditions we ran both ways, the prompt-token count the server reported for a real /v1/chat/completions generation is the same as the token count of the retained render: 112 for xhigh, 70 for medium, 100 for low, 112 for the field omitted. The two endpoints also share their prompt-construction call in the source of the build we ran, but the matching counts are the part we can hand you as a record.

reasoning_effort: "xhigh" renders

<|im_start|>system
Reasoning effort is set to xhigh. Please think carefully through the task,
validate key assumptions, consider plausible alternatives, and prioritize
correctness, consistency, and clarity in the final answer.<|im_end|>
<|im_start|>user
A bag holds 3 red, 4 blue and 5 green marbles. ...<|im_end|>
<|im_start|>assistant
<think>

reasoning_effort: "low" renders

<|im_start|>system
Reasoning effort is set to low. Keep your thinking brief and focused, moving
directly to the conclusion without unnecessary elaboration.<|im_end|>
<|im_start|>user
A bag holds 3 red, 4 blue and 5 green marbles. ...<|im_end|>
<|im_start|>assistant
<think>

reasoning_effort: "medium" renders

<|im_start|>user
A bag holds 3 red, 4 blue and 5 green marbles. ...<|im_end|>
<|im_start|>assistant
<think>

That is the whole finding in three blocks. There is no system message in the medium render. Not a shorter instruction: none. And the render with the field omitted entirely is byte-identical to the xhigh render, because the template line that resolves the value reads, in the shipped jinja, reasoning_effort|default('xhigh').

The template also normalizes high to xhigh before the lookup, and validates: any value that is not one of xhigh, medium or low after normalization calls raise_exception. That validation is what produces the HTTP 500 in section 08, and it sits inside the template's thinking branch, which is why it disappears when thinking is off (section 09).

One detail worth stating plainly because it changes how you read the numbers: the 42 and 30 token figures are the size of the entire injected system block, the <|im_start|>system and <|im_end|> role markers included, not of the English sentence alone. They are the difference between what your prompt costs at that setting and what it costs at medium, which is the number a practitioner is actually billed for.

03

Environment

Local leg, 2026-08-16

ModelQwen3.8-27B, unsloth/Qwen3.8-27B-GGUF, UD-Q5_K_XL quantization, already on disk from a prior run, SHA-256 verified 176a6a3f…
Chat templateThe one embedded in that GGUF (9,993 characters, carrying Unsloth's "Unsloth fixes" marker). This matters: the high to xhigh normalization is an Unsloth addition and is absent from the ggml-org conversion's template.
Runtimellama.cpp release b10453, commit 3cb7ffb1a1f612d5e4a46244ae5a3c77ad934a70, published 2026-08-16, built fresh for this study in a scratch tree. The production build on this machine was read and never modified.
Control runtimellama.cpp b10290, commit c8e03ce, the pre-merge build, run read-only on a second scratch port with zero GPU layers for the render comparison in section 08.
Serving flags--jinja, ctx 32,768, --n-gpu-layers 999, --flash-attn on, --parallel 1, loopback only, no speculative decoding of any kind. No --chat-template and no --chat-template-file: nothing overrode the template inside the GGUF. The control build ran the same flags with --n-gpu-layers 0 and ctx 2,048.
Template actually appliedThe GGUF's own jinja, on both builds, and each build's server log says so. An invalid effort produces a Minja traceback quoting line 64 of the served template, raise_exception('Unexpected reasoning effort ...'). That validator exists only in this template; llama.cpp's built-in templates carry nothing like it. Both logs also print chat template supports preserving reasoning at startup, which is capability detection over the template that was loaded.
ThinkingOn, by the template's own default. No reasoning-related server flag was passed, and no enable_thinking key appears in any of the 36 retained request bodies; the template takes its thinking branch when enable_thinking is undefined. Every render on this page ends at <think> except the three thinking-off ones in section 09.
HardwareOne RTX 5090, 32,607 MiB. Load used 21,856 MiB. Decode ran at roughly 65 tokens per second throughout; the study counts tokens, not speed.
Per callmax_tokens: 2000, temperature: 0, seed: 42, effort sent as the top-level reasoning_effort field.

Speculative decoding was deliberately excluded: llama.cpp issue #27122 reports reproducible CUDA lockups with the multi-token-prediction path on this model class on recent builds. The server log confirms the draft head was loaded but unused.

Those two flag rows are load-bearing for everything this page says about the template. The finding is that a specific string inside this GGUF's jinja gets pasted into your prompt. If your server does not apply that jinja, because it substitutes a built-in template or was started without --jinja, the mechanism described here is not what your server is doing, and the token counts below are not your token counts.

Hosted leg, 2026-08-13

ModelQwen3.8-2.4T-A95B, a mixture-of-experts model, served through OpenRouter
ProvidersNot pinned. The three effort calls landed on DeepInfra (xhigh, medium) and Together (low). DeepInfra listed quantization fp4; Together listed none. This is a confound and is treated as one in section 07.
Request shapeOpenRouter's reasoning: {"effort": ...} object, temperature: 0, max_tokens: 2000, one call per level
Template actually appliedUnknown. We did not see either provider's serving configuration and could not. No template claim on this page extends to the hosted leg.
What was visibleServer-reported token counts, cost, and the returned text. Nothing about prompt construction.

The two legs are different models on different serving stacks. Nothing below treats one as a check on the other in the sense of a replication.

The three request shapes

One setting, three body shapes, and the study touched all three. The local leg sent the field at the top level of an OpenAI-shaped body, which is the path section 08's build boundary is about:

{"model": "...", "messages": [ ... ], "reasoning_effort": "low",
 "temperature": 0, "seed": 42, "max_tokens": 2000}

llama.cpp also accepts the same value inside its own passthrough object, which is the path that worked on both of our builds:

{"model": "...", "messages": [ ... ], "chat_template_kwargs": {"reasoning_effort": "low"}}

The hosted leg went through OpenRouter, which takes its own object:

{"model": "...", "messages": [ ... ], "reasoning": {"effort": "low"}}

We did not test whether any of these servers accepts the others' shapes, so treat the shapes as tied to the stacks we sent them to. If you are debugging a field that appears to do nothing, the shape your stack expects is the first thing to check and the cheapest.

04

The experiment, and the limitation that came out of it

The local run was pre-registered: the question, the grid, the expectations and the honesty rules were written and saved before the first model call, and the source-level reading of what the llama.cpp merge changed was done before the run so the prediction could be scored rather than retrofitted. The run log carries all of it, including the expectations that were confirmed and the deviations that were not planned.

The grid: three prompts by four effort levels by three repetitions, 36 calls.

  • The marble prompt is the hosted study's own prompt, carried over verbatim so that one cell of the local grid is the same message the hosted leg sent.
  • Two further prompts came from our standing probe set: a word problem with a checkable integer answer, and a short prose question with no single right answer.
  • The four levels are xhigh, medium, low, and absent, the last meaning the field is omitted from the request body entirely.

Separately, 16 rendered prompts were retrieved and retained: the four levels on the top-level path and three of them on the chat_template_kwargs path, since there is no kwargs spelling of an omitted field, plus the special values high, none and ultra on both paths, plus three thinking-off renders. Those renders are the evidence in sections 02, 05, 08 and 09, and none of them required a single generated token.

The limitation, discovered during the run and logged there

At temperature: 0 with a fixed seed the server was fully deterministic. All three repetitions of all twelve cells returned byte-identical visible text and byte-identical reasoning text: 12 of 12 cells. The three repetitions therefore measured the server's determinism, not the model's variance, and the effective observation count per cell is one. Every generation number on this page is a single observation and is labeled as one. Nothing here has an interval, and no language on this page implies a distribution. Getting a spread would need a temperature above zero or varied seeds; that run has not happened.

The renders are not affected by this. A rendered prompt is a deterministic function of the request, and section 05's injection sizes are exact counts of text, not samples of behavior.

All 36 calls returned HTTP 200 with a stop reason of stop. Nothing was truncated: the largest completion was 651 tokens against a 2,000-token cap. No reasoning text leaked into the visible answer in any cell. Zero retries, zero errors.

05

Observed: the injection is a fixed-size block

Prompt tokens for the same four levels, on three prompts of different lengths. The counts come from the server's own tokenizer applied to the rendered prompt.

Promptxhighfield absentlowmediumxhigh minus mediumlow minus medium
Marble probability11211210070+42+30
Word problem1111119969+42+30
Prose question99998757+42+30

Every cell is the prompt-token count the server reported for the generations. The marble row is also confirmed independently by the retained renders, tokenized by the server itself (section 02); the other two prompts were measured through generations only and were not rendered separately.

Three readings of the same two constants. The injected block does not scale with your prompt: it is 42 tokens for xhigh, 30 for low, and 0 for medium, added on top of whatever you sent. In characters, the same three renders of the marble prompt measure 566, 495 and 329, so xhigh adds 237 characters of system message and low adds 166.

The xhigh and absent columns are not merely equal in count. On the marble prompt the two rendered prompts are byte-identical strings, and the two generations they produced are byte-identical too: same reasoning text, same visible answer, same token counts, on all three prompts. If you are choosing between "send xhigh" and "leave it out", you are choosing between two spellings of the same request.

06

Observed: what it does to thinking, and to answers

Single observation per cell. Reasoning tokens are counted exactly, by sending the returned reasoning text back to the server's own tokenizer, rather than trusting a reported field.

PromptEffortReasoning tokensAnswer charactersAnswer shape
Marble probabilityxhigh399284bare LaTeX, boxed result
field absent399284identical to xhigh
medium339797markdown headers, a table
low283711markdown headers, a table
Word problemxhigh126134two sentences
field absent126134identical to xhigh
medium122345numbered steps in bold
low78148three short lines
Prose questionxhigh123346three sentences
field absent123346identical to xhigh
medium405686one long paragraph
low90391two long sentences

Every row is one observation. The three repetitions of each cell were byte-identical (section 04).

On thinking volume, low behaves and medium does not. low spent less than xhigh on every prompt, at 71%, 62% and 73% of the xhigh figure. medium spent 85% on one prompt, 97% on another, and 329% on the third: on the prose question, the level with no injected instruction at all produced 3.3 times the thinking of the level that explicitly asks for careful thought. That is not a middle setting misbehaving. It is what you would expect if medium simply removes an instruction and lets the model do whatever it does unprompted.

On answer style, the direction held on all three prompts. xhigh produced the shortest visible answer in each of the three, and medium the longest in each, by factors of 2.8, 2.6 and 2.0. Three prompts, one observation each. The direction is the same one the hosted 2.4T leg showed three days earlier: there, xhigh returned 340 characters of terse mathematics while medium returned 720 characters of markdown with headers.

There is a readable mechanism for that, and it is in the injected text itself: the xhigh sentence ends by asking for "correctness, consistency, and clarity in the final answer", and the low sentence asks the model to move "directly to the conclusion without unnecessary elaboration". Both of those are instructions about the answer as much as about the thinking. The setting is named for reasoning effort and reads, in practice, as a house-style instruction. We are reading the injected sentence here, not instrumenting the model, and the distinction matters: this is an interpretation of a measured text, offered as one.

Correctness: the two prompts with a checkable answer were right in all eight of their cells (the marble probability came back 3/11 with 60 favourable outcomes and 220 total; the word problem came back 27). The third prompt is open prose and has no key, so we make no correctness claim about it.

07

Observed: the hosted 2.4T leg, and what converges

Three days before the local run, the same marble prompt went to the hosted Qwen3.8-2.4T-A95B through OpenRouter, one call per level, same temperature and same output cap. What the hosted study could see was this:

EffortProviderPrompt tokensReasoning tokensAnswer charactersCompletion tokensCost
xhighDeepInfra112184340356$0.00236
mediumDeepInfra70197720582$0.003632
lowTogether100168527505$0.00341

Single observation per level, 2026-08-13. Cost is the server's own reported figure, not a computation.

The thinking-volume dial did not appear. Reasoning spend across the three levels was 168 to 197 tokens, a 17% band, with the nominal maximum-effort setting sitting in the middle of it. The study's pre-registered expectation had been that spend would order xhigh above medium above low. It did not.

The cost ordering inverted the knob's apparent intent: medium cost the most, xhigh the least, because the visible answer grew as the nominal effort fell and output tokens are what you pay for. State that with its confound attached: the low call landed on Together, whose listed prices were higher than DeepInfra's, so the medium-versus-low ordering is not clean. The xhigh-versus-medium comparison is clean, both being DeepInfra calls, and on that pair the lowest nominal effort setting of the two cost 54% more.

And the three prompt-token counts had no explanation. The message bytes were identical across the three calls, verified from the retained request bodies, yet the server reported 112, 70 and 100 prompt tokens. The study recorded them and wrote, in the finding itself, "Mechanism unknown; recorded as evidence the effort level rewrites the served prompt."

Our locally rendered prompts, for the same three levels on the same message, tokenize to 112, 70 and 100.

What that licenses, stated carefully

These are two different models on two different serving stacks: a 2.4-trillion-parameter mixture-of-experts model served by two commercial GPU clouds through an aggregator, and a 27-billion-parameter dense model in a community-quantized GGUF served by llama.cpp on a desktop. We did not read the hosted providers' prompt construction and cannot. This is a convergence of two independent readings, not a replication. What it supports is that the hosted numbers are consistent with the same template mechanism we can read locally, which is a good deal more than the hosted study could say on its own and a good deal less than proof of what DeepInfra's server does. The two models are the same family and, on this evidence, tokenize this prompt identically, which is the premise the convergence rests on and is itself an inference from the matching integers rather than something we verified independently.

The style direction converged too: terse at xhigh, long and formatted at medium, on both legs. The thinking-volume result did not fully converge. Hosted spend was flat across all three levels; locally, low came in below xhigh on each of the three prompts while medium was erratic. Two observations of one cell each disagreeing about a small effect is exactly what you would expect whether or not the effect is real, and we draw nothing from it.

08

The version boundary: llama.cpp b10290 and b10453

The mechanism above lives in the chat template, so it applies wherever that template is the one being rendered. Whether the field ever reaches the template is a property of your runtime, and on llama.cpp that changed recently.

Be precise about which half of this we measured. Our renders show different behavior on b10290 and on b10453, and both builds were started with the same flags, --jinja and no template override. What we are repeating rather than measuring is the release mapping: the run log reports that pull request #26941 merged on 2026-08-14 and first shipped in release b10434, read from the project's own records. We did not build b10434, and we tested nothing between b10435 and b10453. Confirm that mapping against your deployed build rather than taking the number from this page.

The two builds do differ in the way that pull request describes. On b10290 the OpenAI-compatibility layer handles the value none and leaves every other value unhandled, under a source comment saying in full that other reasoning_effort values are model-specific and not yet handled. That reads as an incomplete compatibility stub rather than a defect: on that build there was no Qwen effort vocabulary on that field for the server to forward. The user-visible consequence is the same either way, and it is what the table shows. On b10453 the value is forwarded into the template's variable context.

We measured both ends rather than taking the diff's word for it. The same model file, the same prompt, the same flags and the same render conditions were sent to both builds, the pre-merge one run read-only on a separate scratch port with zero GPU layers. Every cell below is an /apply-template render, tokenized by that build's own tokenizer. Exactly three cells differ:

Value sentPathPre-merge b10290Post-merge b10453
xhightop-level field112 tok112 tokunchanged by coincidence: read the note below
mediumtop-level field112 tok70 tokchanged
lowtop-level field112 tok100 tokchanged
field absenttop-level field112 tok112 tok
xhigh / medium / lowchat_template_kwargs112 / 70 / 100112 / 70 / 100unchanged: this path always worked
higheither path112 tok112 toknormalized to xhigh by this template
nonetop-level field72 tok72 tokhandled pre-merge too
nonechat_template_kwargsHTTP 500HTTP 500not a valid template value
ultratop-level field112 tokHTTP 500changed: HTTP 200 with the default prompt became HTTP 500
ultrachat_template_kwargsHTTP 500HTTP 500

Both columns are renders of the same prompt through the same model file, retrieved from each build's /apply-template endpoint and tokenized by that build's own tokenizer.

Read the xhigh row carefully, because it is the trap in this table. On the pre-merge build xhigh rendered 112 tokens, the same as the post-merge build, and that is a coincidence rather than support. Pre-merge, every top-level value rendered 112 tokens, because the field never reached the template and the template fell back to its xhigh default. A user who sent xhigh to that build got the right prompt for the wrong reason. A user who sent low got the xhigh prompt and no error. A user who sent a typo got the xhigh prompt and no error.

Which endpoint produced these cells. Every cell in the table is a render from /apply-template on the build named in its column. On the post-merge build we also sent the invalid value through /v1/chat/completions itself, both at the top level and inside chat_template_kwargs, and both returned HTTP 500 there too, so that result is not an artifact of the render endpoint. We did not run the pre-merge build through /v1/chat/completions: its column is a render comparison only.

Upstream issue #27023, "Misc. bug: reasoning_effort seems broken", has been open since 2026-08-13, and a webui proposal (#27118, 2026-08-15) asks for effort and thinking budget to become two separate settings. The boundary above is consistent with confusion from both directions: before the merge the field did nothing, and after it the field does something and rejects values that used to pass. We make no claim about what any particular reporter experienced.

The operational consequence

If you run llama.cpp with this GGUF's template applied, and thinking is on, and your client sends a top-level reasoning_effort that the template does not recognize, then crossing this build boundary turns a value that had no effect into an HTTP 500. All three conditions are load-bearing: with thinking off the same request returns HTTP 200 (section 09), and with a built-in template applied instead it is a different question that this page did not measure. Anything a client library sends by default, including vocabulary borrowed from other vendors' APIs, is worth auditing before that upgrade.

09

Special values, two paths, and one escape hatch

Three behaviors are worth separating out, because each of them will bite somebody.

none means two different things depending on how you send it. Sent as the top-level field, it is intercepted by the server's own compatibility handler before the template is reached: thinking is turned off, the render carries a closed empty thinking block, and a generation returns an answer with zero reasoning tokens. That render is the medium render plus that block and nothing else, which is where its 72 tokens come from: 340 characters against medium's 329, 72 tokens against medium's 70. The two extra tokens are the closed block, not an instruction. Sent inside chat_template_kwargs, it goes straight into the template, which does not list none among its valid values, and the template raises: HTTP 500. Same word, two paths, opposite outcomes, on both builds we tested.

high is normalized, but that normalization belongs to the quantizer's template, not to the standard. The template in the file we served maps high to xhigh before validating, so high renders the xhigh prompt on both paths and both builds. That normalization line is an Unsloth addition. The ggml-org conversion of the same model carries a template without it. We did not serve the ggml-org file, so we make no measurement claim about it, but on a plain reading of that template text a high sent to it would fail validation the way ultra does. Treat this as a prediction to test, not a result.

Invalid values fail loudly, and the failure is conditional on thinking being on. An unrecognized value returns HTTP 500 with the template's own traceback, which names the supported set in plain text:

Error: Jinja Exception: Unexpected reasoning effort ultra. Supported types
are xhigh (default), medium, and low.

That is an unusually good error message, and it is worth knowing it can appear at all. But the validation lives inside the template's thinking branch, so it is unreachable when thinking is off. We tested that: with thinking turned off, the invalid value ultra returned HTTP 200 and a render byte-identical to the thinking-off render without it. On this model the thinking switch belongs inside chat_template_kwargs, as {"enable_thinking": false}; a top-level enable_thinking field was measured as ignored on this endpoint two days earlier. One retention note, because this page asks you to check its records: the three thinking-off entries retain the rendered prompt but not the request body, so the render is the artifact and the key used is the run log's account of it.

The consequence is worth stating in both directions. If your integration disables thinking, an invalid effort value passes silently. If somebody later turns thinking back on, the same request starts returning HTTP 500.

We tested one invalid value, ultra. The template's validation rejects anything outside xhigh, medium and low after normalization, so on a plain reading of that source any other unrecognized value takes the same branch. That is a read of the template, labeled as such, not a second measurement. The practical point is that clients send what they send, and a value another vendor's API accepts is not thereby a value this template accepts. The only way to know is to send it at the template and runtime you actually deploy and look at what comes back.

10

What this does not show

  • Every generation figure here is a single observation. Temperature 0 with a fixed seed collapsed three repetitions into one, in 12 cells out of 12. No intervals, no significance, no distribution. Differences of a few percent between cells mean nothing on this page; the ones we lean on are large (a 3.3-times ratio, a 2-to-3-times answer length ratio, an exact zero in the medium render).
  • One template, applied one way. Both builds ran --jinja with no template override, so every template result here describes a server that applies the GGUF's own jinja; we did not measure what happens with a built-in template substituted for it. And the jinja we read is the Unsloth conversion's. We did not serve the ggml-org conversion, whose template differs at least in the high normalization (section 09), and we never read the 2.4T model's template at all.
  • Two deployments, not a survey. One hosted 2.4T-A95B through one aggregator on 2026-08-13, split across two providers without pinning, and one local 27B in one community quantization on one desktop on 2026-08-16. Other quantizations ship other templates. Other runtimes handle the field their own way.
  • The hosted mechanism was not read, only inferred. We do not know how DeepInfra or Together construct the prompt. The match of three integers is consistent with the template mechanism and is not proof of it.
  • Three prompts. Two with checkable answers, one prose. The prompts are short. Whether the levels separate more on genuinely hard problems, where a reasoning instruction has more to bite on, is unmeasured and is the obvious next run.
  • We did not instrument the model. Section 06's reading of why answer style moves more than thinking volume is an interpretation of the injected sentence's wording, not a mechanistic result.
  • No comparison of models and no quality ranking. Nothing here says one model reasons better than another, and the two legs are not scored against each other.
  • Other stacks were not examined. We did not test vLLM, TensorRT-LLM, text-generation-inference, or the vendors' own first-party endpoints, and nothing here speaks to how they handle this field.
  • Dates matter more than usual on this page. The llama.cpp behavior changed two days before we measured it, and an open upstream issue and an open proposal both touch this surface. Treat every version-specific statement as unverified after 2026-08-16.
11

If you send reasoning_effort to Qwen3.8

Seven narrow things, each scoped to what is above, and all of them assuming a server that applies this GGUF's own chat template.

  1. Know that you are writing a sentence into your prompt. On this model the field is not a sampler control. It prepends a system message. If you already send your own system message about how the model should work, you are now sending two, and the template's one comes first.
  2. Send medium when you want no instruction at all. This is the counterintuitive one. medium is the only value that leaves your prompt alone. It is also the cheapest of the three on input tokens, by 42 tokens against xhigh. What it is not is a middle amount of effort.
  3. Do not treat omitting the field as neutral. It renders as xhigh, byte for byte. If your intent is "do not touch my prompt", omitting the field does the opposite of what you want and medium is the value you are looking for.
  4. low bought less thinking on our three prompts, and a different answer with it. In these three deterministic local prompts, low used 62% to 73% as many reasoning tokens as xhigh. Test that pattern on your own prompts before relying on it. It also made answers longer and more heavily formatted than xhigh did, which may be the opposite of what a trim-the-thinking setting is being reached for.
  5. To turn thinking off, send none at the top level, and never put none in the kwargs. This is the most useful single fact on the page for anyone watching a bill. Top-level none is intercepted by the server before the template: thinking off, HTTP 200, zero reasoning tokens. The same word inside chat_template_kwargs reaches the template, which does not list it as a valid effort, and returns HTTP 500. This model also has a real thinking switch of its own, chat_template_kwargs {"enable_thinking": false}, which we measured on this endpoint two days before this study.
  6. On llama.cpp, know which side of the build boundary you are on. On the older build a top-level reasoning_effort other than none never reached the template and every request rendered the default. On the newer one the value takes effect, and an unrecognized value returns HTTP 500 while thinking is on. Before upgrading across that boundary, four things are worth auditing: what your clients actually put in that field, including values borrowed from other vendors' vocabularies; whether thinking is on, since the HTTP 500 needs it; whether your server applies the GGUF's template or a built-in one; and which of the request shapes in section 03 your code sends. A setting that has been quietly doing nothing in your stack starts doing something the day you rebuild.
  7. Pick one path and know its quirks. The top-level field is the standard-shaped one and handles none as thinking-off. The chat_template_kwargs path worked on both builds and passes none through to a template error. If you need one code path that behaves the same on old and new llama.cpp builds, chat_template_kwargs is that path, with the caveat that it is a llama.cpp-specific request shape.
A ten-minute check on your own stack

This page cannot tell you what vLLM, text-generation-inference or a vendor's own endpoint does with this field, because we did not test them. It can tell you how to find out in two requests. Send the same message twice, once with the field at xhigh and once at medium, and compare the server's reported prompt_tokens. If the two counts differ, something in your stack is rewriting your prompt from that field, and the difference is the size of what it injects. If they are identical, the field is not reaching a template that does anything with it, and whatever behavior you are attributing to it is coming from somewhere else. On llama.cpp you can go one step further and read the prompt itself out of /apply-template.

And one thing to measure if the levels matter to you: whether they separate on your prompts. On ours they mostly changed the shape of the answer. Send your real prompts at xhigh and at medium, count reasoning tokens and read the two answers side by side. It is a cheap test and it will tell you which of the two effects you are actually buying.

12

Artifacts

The complete evidence for this page ships as a data package: the files the two runs actually produced, not a summary. If a number on this page disagrees with a file in the package, the file is right and the page is wrong; tell us and we will fix the page.

  • data/README.md: what is in each folder, the record schema, and the provenance of every file.
  • data/NUMBERS.md: every figure printed on this page, with the file and the field it came from.
  • The pre-registration and run log (data/local/RUN_LOG.md) for the local study, written and saved before the first model call, including the source-level reading of the llama.cpp change made before the run, the expectations it predicted, and the deviations logged as they happened, including the determinism finding that reduced the effective observation count to one per cell.
  • The 16 rendered prompts, verbatim, every condition in sections 02, 05, 08 and 09 (data/local/raw/render_*.json). These are the primary evidence: the rendered text is the finding.
  • The 36 per-call records from the grid (data/local/raw/grid_*.json), each carrying the full request and response body, the server's token counts and timings, the exact reasoning-token count, and the wall time, plus the summarized grid (data/local/effort_grid.tsv).
  • The cross-build comparison (data/local/merge_ab_comparison.txt) against the pre-merge llama.cpp build, plus that build's server log and the 13 pre-merge renders (data/local/raw_premerge/).
  • The hosted study's effort records from 2026-08-13 (data/hosted/raw/effort_*.json): the three request bodies proving the messages were byte-identical, and the three response bodies carrying the provider name, the token counts and the server-reported cost.
  • The release verification (data/local/release_verification.txt) taken at the end of the local run: GPU memory, ports, and process state.
2026-08-13Hosted battery on Qwen3.8-2.4T-A95B through OpenRouter, including the three-level effort sweep on the marble prompt. One call per level; providers recorded per call and not pinned. The three unexplained prompt-token counts were recorded that day as "mechanism unknown".
2026-08-14llama.cpp pull request #26941 merged, reported as first carried in release b10434: the top-level reasoning_effort field begins reaching the chat template. This row is the project's record, not our measurement.
2026-08-16Local study on Qwen3.8-27B: pre-registration written before the first call, 36 generations and 16 renders on the post-merge build b10453, then the same renders on the pre-merge build b10290 with zero GPU layers. No unexpected operational failures or exclusions; the planned invalid-value probes returned the documented HTTP 500 results. GPU released and verified clean.

Every claim on this page is dated. The llama.cpp behavior described in section 08 changed on 2026-08-14 and was measured on 2026-08-16; if today is much later, treat it as unverified since that date.

13

Related pages on this site

Qwen3.8-27B: day zero is the release-day field card for the local model measured here, on the same desktop. Release watch: Qwen3.8 tracks the family's releases and licenses, including the split between the Apache-licensed 27B and its differently-licensed 2.4T sibling. Muse Glimmer 30B is the same class of knob on another vendor's model, measured the same day on the same desktop: there the lever is called reasoning_strength, the template interpolates the value into a system line without validating it, and an undefined value was accepted without complaint. Empty Answer is the same editorial thesis in the neighbouring domain: what a reasoning model does with your output budget when the thinking does not fit. The Qwen hub page carries downloadable serving profiles for the local models on this machine.