Graphometer
Field card · measured 2026-08-16 on one RTX 5090 · one dense 30B, Meta's own 4-bit build and its DFlash drafter, greedy decoding, one request at a time · comparison note corrected 2026-09-26

Muse Glimmer 30B: trained at 16, asked for 3

Meta trained its DFlash drafter to produce a block of 16 tokens per pass and prints that number in its specification table. llama.cpp's generic draft length is 3. The two flags Meta's GGUF card documents leave that generic default in charge, and on our six structured prompts the difference is worth 2.17 times the throughput, 116.505 tokens per second against 253.19. Prose tops out at 1.51 times our baseline whatever we do. Neither project's documentation is wrong on its own. The license is verbatim Apache 2.0, with a separate usage policy shipping beside it.

Meta published Muse Glimmer 30B with a speed table that names our exact hardware: 74.9 tokens per second with no speculation, 233.4 with its DFlash drafter, 3.1 times faster, averaged across a prompt set they do not publish, at batch size 1 with greedy decoding on an Nvidia RTX 5090 through llama.cpp. On 2026-08-16 we downloaded the two files that table rests on, checked every byte against the repository's own object ids, and ran five server loads on one RTX 5090. On our harness the numbers land where theirs do. Text-only, one slot, 8,192 tokens of context, greedy, our no-speculation medians were 76.22 tokens per second on six prose prompts and 76.345 on six structured prompts, 1.8 and 1.9 percent above their published baseline, and our fastest structured median, 253.19, is above their published 233.4. That is a scoped comparison and not a reproduction of their average: their prompt mix is unpublished, ours is printed in section 02, and we never started their full documented command. The finding is in the gap between two projects. Meta trained DFlash to draft a block of 16 tokens in one forward pass and prints Block size | 16 in its specification table. llama.cpp's --spec-draft-n-max defaults to 3, a sensible number for the ordinary autoregressive draft models that flag was written for, and llama.cpp does not read the drafter's block size out of the GGUF to set it. Meta's GGUF card turns speculation on with two flags, -md and -ngld 99, and says nothing about draft length, so the generic default stays in charge: the server prints n_max=3 on the line directly above block_size=16. In that configuration our six structured prompts run at a median of 116.505 tokens per second; adding --spec-draft-n-max 16 moves the same six to 253.19, 2.17 times more. No document here is wrong on its own, and the reader who follows the documentation still loses that 2.17 times. And the workload decides more than the flag does: with the same flag, prose reaches 1.51 times our baseline at its best and structured output 3.32, so a single averaged speedup, ours or anyone's, cannot tell a reader which of those two their own work will get. Six prompts per cell, one call each, no interval computed anywhere on this page. Everything below is one quantization of one model on one machine, greedy, 8,192 tokens of context, one request at a time, on 2026-08-16.

Independent measurements. Not affiliated with, endorsed by, or connected to Meta Platforms, the llama.cpp project or ggml-org, or Unsloth.
01

What it is

Rows below were read from the model's own configuration files and from the GGUF headers on this disk at the pinned revisions on 2026-08-16, not quoted from a marketing page. Where a row is the maker's claim rather than our reading or our measurement, it says so, and amber text marks it.

MakerMeta. The model card names Meta Superintelligence Lab as author and August 2026 as the release date. The repositories live in the meta-models organization on Hugging Face.
Weightsmeta-models/Muse-Glimmer-30B at revision a4e59da52a7bc87ae7251dd5545c0dd437c44b68; the GGUF conversions we served come from meta-models/Muse-Glimmer-30B-GGUF at 43c7eadd41352a299ea8e0a36b3157978dd63596, both read on 2026-08-16. The base repository holds 2 safetensors shards, 59,553,435,272 bytes, and Hugging Face reports 29,776,626,688 parameters.
LicenseVerbatim Apache 2.0, 11,358 bytes, read in full at the pinned revision and retained in our record. A separate USAGE_POLICY.md of 5,230 bytes ships beside it in the same repository. Section 06 states both facts and makes no claim about which one binds you.
ShapeDense. No expert blocks anywhere in the configuration, so the --n-cpu-moe lever that tunes this machine's mixture-of-experts models does not exist here. The quantization fits the card whole or it does not fit.
Architecture52 layers, hidden 6,656, feed-forward 19,968, 32 query heads to 2 key-value heads, head dimension 128, RoPE theta 500,000. The layer pattern is three sliding-attention layers with a 2,048-token window followed by one full-attention layer, repeated 13 times, so only 13 of the 52 layers hold an unbounded key-value cache. Section 05 is what that buys.
DrafterA separate model, not a head inside the weights file: meta-models/Muse-Glimmer-30B-assistant, 2,555,985,152 parameters, 5 layers, block_size: 16, reading hidden features from 5 layers of the target. Its GGUF twin declares general.architecture = dflash. Per the vendor's card it is a block-diffusion drafter that predicts a whole block of 16 tokens in one forward pass, which is a different object from the multi-token-prediction and draft-head numbers elsewhere on this site. Sections 03 and 04 do not pool them.
Context131,072 tokens native, declared in the configuration. Section 05 measures the whole of it on this card.
Vocabulary202,048 tokens; beginning-of-sequence 200,000, end-of-sequence 200,001, end-of-turn 200,008.
ModalityNatively multimodal by its configuration: a 50-layer vision tower, patch 14, shipped as a separate projector file. We served it text-only and never downloaded the projector. Vision is untested here and no claim on this page rests on it.
Quantization servedMeta's own Muse-Glimmer-30B-KQuant-17GB-Q4_K_M.gguf, 16,756,683,904 bytes, plus dflash-Muse-Glimmer-30B-Q4_K_M.gguf, 1,631,208,128 bytes. These are the two files the vendor's own speed table names, which is why a comparison against that table is possible at all. Section 09 measures a community alternative.
QualityThe maker claims 1.0% degradation for this K-Quant-17GB build and 0.2% for K-Quant-Dynamic, "measured using an average on accuracy metrics across 15 common benchmarks". We measured no quality of any kind. That row is their claim, reproduced, not our finding.
02

What we ran, exactly

MachineOne desktop: Core Ultra 9 285K, 188 GiB RAM, one RTX 5090 with 32,607 MiB of VRAM, Pop!_OS.
RuntimeMainline llama.cpp llama-server, build b10453, commit 3cb7ffb, CUDA 12.8, already on this machine and reused read-only. Muse Glimmer support merged upstream as PR #26841 on 2026-08-10 and first shipped in release b10353, so this build sits a hundred releases past the floor.
A version check that needs a release buildThe vendor's card says to run llama-cli --version and require a build number of 10353 or higher. That works on a release binary. On a build from source llama.cpp itself prints version: 0.1.0-dev (build 1, commit 3cb7ffb), so the number is not there to test and the check is unusable in that case. Their second check works everywhere: grep -c LLM_ARCH_MUSE_GLIMMER src/llama-arch.cpp.
IntegrityThree files downloaded, three planned; every byte count equals the pinned manifest and every SHA-256 equals the repository's own LFS object id at the pinned revision. That is file integrity against the vendor's own conversion, which is the whole chain here, because Meta published these GGUFs itself.
Placement-ngl 99, every layer on the GPU, plus -ngld 99 for the drafter.
SamplingGreedy, temperature 0, matching the vendor's stated measurement method (batch size 1, greedy decoding). Their recommended chat sampling is temperature 1.0, top-p 0.95, top-k 64; we did not use it for speed work and cannot say what it would change.
Concurrency--parallel 1. Single request throughout. Section 11 says why that boundary is drawn where it is.
Speed methodFive server loads, one per arm, each a fresh process. Each arm runs the same twelve prompts, six prose and six structured, each prompt sent exactly once, so no prompt is reused inside an arm and no prefix cache is warmed across the measured calls. A discarded warm-up call precedes each arm. The figure reported per cell is the median of six different prompts, one call each, not repetitions of one prompt, and the design is paired: arm to arm, the prompts are identical. Server-reported decode rates are authoritative.
ServingA loopback scratch port for the study. Nothing was installed: no service unit, no registry key, no client wiring, and the server was stopped and verified stopped at the end.

The exact invocation. Every arm in section 03 is this command with the speculation flags added or changed; nothing else moves:

llama-server \
  --model Muse-Glimmer-30B-KQuant-17GB-Q4_K_M.gguf \
  --host 127.0.0.1 --port <scratch> --alias muse-glimmer-30b \
  --jinja --ctx-size 8192 --parallel 1 --n-gpu-layers 99 \
  --threads 24 --threads-batch 24 --flash-attn on --timeout 3600

Arm B adds the two flags the vendor's card documents for speculation, and nothing more:

  -md dflash-Muse-Glimmer-30B-Q4_K_M.gguf -ngld 99

Arms C and D add one flag beyond that, --spec-draft-n-max 16 and --spec-draft-n-max 8. Section 03 sets out exactly how that relates to the vendor's own documented server command, which is a different and longer thing that we did not run.

Our prompt set, since we are about to say that theirs is missing

Every arm sent these twelve, in this order, once each, max_tokens 900, temperature 0. Six prose:

  1. Write four paragraphs about why coastal fishing villages in Iceland declined in the twentieth century.
  2. Write four paragraphs explaining how a mechanical watch escapement works, for a curious adult.
  3. Write four paragraphs about the history of the Dutch windmill and its role in land reclamation.
  4. Write four paragraphs describing the life cycle of a monarch butterfly and its migration.
  5. Write four paragraphs on how sourdough starter cultures work biologically.
  6. Write four paragraphs about the invention and spread of the printing press in Europe.

Six structured:

  1. Write a Python class LRUCache with get and put, using a dict and a doubly linked list. Include docstrings. Output only code.
  2. Write a JSON array of 12 objects, each with keys id, name, country, founded_year, describing 12 European universities. Output only JSON.
  3. Write a Python module with five functions for vector math (add, sub, dot, norm, normalize) with full type hints and docstrings. Output only code.
  4. Write a JSON object mapping 15 chemical element symbols to objects with keys name, atomic_number, group. Output only JSON.
  5. Write a Python implementation of a binary search tree with insert, find, and in-order traversal, fully commented. Output only code.
  6. Write a markdown table with 15 rows comparing 15 programming languages across columns: language, year, paradigm, typing, main use.

That set is ours, not a standard, and it is small. It is printed so that anyone can rerun it and so that no sentence on this page criticizes a missing prompt list from behind one.

How the numbers are rounded. Decode figures are printed as the server reports them. A median of six values falls between the two middle prompts, so some medians end in a half cent of a token per second: those are printed with three decimals (76.345, 116.505, 76.325, 79.615) rather than rounded, and every ratio on this page is computed from unrounded values.

Startup noise that is not a fault. Every load with the drafter prints dflash requires ctx_other to be set and then [spec] failed to measure draft model memory. The vendor documents the second one as harmless, and the model loads and serves normally. Every load of every Muse GGUF here also prints special_eot_id is not in special_eog_ids. We observed no stop-behaviour problem in any run, so we record that line as an open observation and not as a defect.

How to read the VRAM figures

Every VRAM number on this page is the raw MiB reading from nvidia-smi against a 32,607 MiB card. We print MiB first because that is what the instrument reports, and we never divide a MiB reading by 1000 and call the result GiB.

03

The numbers side by side, and the default that decides the gap

Meta's card publishes this table, with the method stated underneath it: an average across a diverse prompt set, batch size 1, greedy decoding, the RTX measurements taken with llama.cpp.

GPUBaseline, no speculationWith DFlashSpeedup
Nvidia RTX 509074.9233.43.1x

The vendor's own figures, quoted from their GGUF model card as read on 2026-08-16.

The reported hardware, runtime, decoding mode and files are specified, and they are ours. The prompt mix is not. Here is our own measurement beside it, with the prompt set split in two and printed in section 02, greedy, 8,192 tokens of context, one request at a time, on 2026-08-16.

ArmSpeculation flagsProse t/svs baselineStructured t/svs baseline
Anone76.22baseline76.345baseline
Bthe two flags the card documents79.071.04x116.5051.53x
DB plus --spec-draft-n-max 8114.461.50x208.292.73x
CB plus --spec-draft-n-max 16114.861.51x253.193.32x

Each cell is the median of six different prompts, one call per prompt (n = 6), inside one server process, with the same twelve prompts in every arm. The "vs baseline" columns divide by arm A of the same workload, never by the vendor's number. The arms are lettered in the order they were run and the table is ordered by draft length, which is why D sits above C. No interval is computed and none is implied; this is a diagnostic ladder, not a benchmark.

Our baseline lands where theirs does

76.22 and 76.345 against their 74.9 is 1.8 and 1.9 percent above the vendor's figure on a machine they did not configure, which is as close as anyone should expect two different desktops to land.

Our fastest arm is above their published DFlash figure, on structured prompts only

253.19 tokens per second is 8.5 percent above the 233.4 they published, and 3.32 times our own baseline against their 3.1 times theirs (n = 6 structured prompts, one call each). Those two ratios are not the same estimand: theirs is an average over an unpublished mix of unknown composition, ours is a median over six structured prompts we wrote and printed. They are printed side by side because they are the two numbers a reader will hold, not because one reproduces the other.

Where the difference comes from: one line in the server log

With the two documented flags and nothing else, the drafter registers like this:

common_speculative_impl_draft_dflash: - n_max=3, n_min=0, p_min=0.00
common_speculative_impl_draft_dflash: - block_size=16, mask_token_id=201818, n_extract=5

n_max=3 is llama.cpp's own generic default draft length, inherited from ordinary autoregressive draft models. block_size=16 is what this drafter was trained to produce in a single forward pass, and the runtime prints it on the very next line without using it. A block drafter asked for three tokens at a time throws away most of what it computed. Ask for the block it was built for and the same log reads:

common_speculative_impl_draft_dflash: - n_max=16, n_min=0, p_min=0.00
common_speculative_impl_draft_dflash: requested draft size (n_max=16, n_min=0) exceeds
  the trained block size 16 -- clamping to 15

The runtime clamps 16 down to 15, and 15 is where the fast arm actually ran. Pass 16 and accept 15; passing 15 yourself would land in the same place without the warning. Arm D, at --spec-draft-n-max 8, draws no warning and gets most of the way there.

Nobody's document is wrong, and the reader still loses 2.17 times

Three separate things are each individually correct. Meta trained DFlash at block size 16 and publishes that number in the drafter's specification table on both readmes; it is an architecture row, not a command-line instruction, and it does not claim to be one. llama.cpp's --spec-draft-n-max defaults to 3, which is a reasonable default for the ordinary autoregressive draft models the speculative path was built for; the flag is llama.cpp's, documented by llama.cpp, and on this build the runtime does not read the drafter's block size out of the GGUF to raise its own default. Meta's GGUF card documents the two flags that turn speculation on, -md and -ngld 99, which is exactly what a reader needs to load the drafter, and says nothing about draft length, which leaves the generic default in charge.

The string --spec-draft-n-max appears nowhere in either vendor readme; we checked both retained copies and the grep returns zero hits. That is a fact about the readmes, not an accusation: a vendor card is not obliged to teach another project's flags. The consequence is still real. A practitioner who reads Meta's card carefully and llama.cpp's defaults not at all gets 1.53 times our baseline on structured output and 1.04 on prose, with no documented reason to suspect that one more flag is worth 2.17 times more on the structured half, 253.19 against 116.505, six prompts per cell, one call each. This is what a gap between two correct documents costs, and it is the kind of gap that only shows up when somebody runs both projects together and reads the log.

What arm B is, exactly, and what it is not

This is the sentence a reviewer should attack, so here is the whole of it. The vendor's GGUF card documents this server command:

./build/bin/llama-server \
    -m       Muse-Glimmer-30B-GGUF/Muse-Glimmer-30B-KQuant-17GB-Q4_K_M.gguf \
    --mmproj Muse-Glimmer-30B-GGUF/mmproj-Muse-Glimmer-30B-Q4_K_M.gguf \
    -a muse-glimmer-30B \
    -ngl 99 -c 131072 -np 4 \
    --host 127.0.0.1 --port 8080 \
    --jinja \
    --temp 1.0 --top-p 0.95 --top-k 64

and then documents speculation as an addition to it: -md Muse-Glimmer-30B-GGUF/dflash-Muse-Glimmer-30B-Q4_K_M.gguf -ngld 99. Arm B reproduces that two-flag addition verbatim, on the same drafter file, inside our own harness. It does not reproduce the command above. Ours is text-only with no projector, one slot instead of four, 8,192 tokens of context instead of 131,072, and greedy instead of their recommended chat sampling, because greedy at batch size 1 is what the vendor's own measurement footnote specifies. So we did not measure the vendor's documented command and this page makes no claim about what it would produce. What we can say is narrower and still useful, because --spec-draft-n-max is absent from that command too: any invocation built from that page inherits llama.cpp's generic default, and our arm B log is a retained instance of what that looks like at startup. The argv of each arm was not itself written to a file; the flags are evidenced by the harness script and by three lines in each retained server log, the drafter's path, the n_max line, and n_slots = 1, n_ctx_slot = 8192. Their four-slot default is also the regime that upstream issue #27117 reports acceptance collapsing in; see section 11.

The split is the point

The same flag that buys 3.32x on structured output buys 1.51x on prose. Without the mix, a vendor average cannot tell a reader which result to expect for their own workload, and ours cannot either, which is why both halves are printed separately and never pooled into one number. Our speculation-gate study measured a 397B mixture-of-experts where an ungated draft made prose slower than no speculation at all while making structured output faster, and the Qwen3.8-27B field card found the same direction on a dense 27B living entirely on the GPU. This is a third deployment, a different drafter mechanism, and the same lesson about pooled numbers. The magnitudes do not transfer between those pages and this one, and nothing here reproduces their numbers.

04

Acceptance rate misranks the arms on this drafter

The obvious diagnostic for speculative decoding is the acceptance rate, and on this drafter it points the wrong way. Same runs as section 03, counters summed across the six prompts of each workload (n = 6 per row, one call per prompt).

ArmWorkloadDraft tokens generatedAcceptedAcceptanceTokens per target stepThroughput
Bstructured4,6013,4080.7413.22116.505
Dstructured9,0233,9640.4394.49208.29
Cstructured13,2044,2120.3195.75253.19
Bprose6,9262,6890.3882.1679.07
Dprose16,3602,8750.1762.40114.46
Cprose28,5893,0130.1052.57114.86

Acceptance is accepted draft tokens divided by generated draft tokens, from the server's own counters. Tokens per target step is completion tokens divided by (completion tokens minus accepted draft tokens), which is the same arithmetic the server prints per request as mean len. It is an effective-work proxy and a diagnostic: it is not a recorded count of target-model forward passes, because the server does not log one.

Acceptance falls as throughput rises, from 0.741 to 0.319 on structured output while the same six prompts get 2.17 times faster. The proxy moves with the speed instead: 3.22, then 4.49, then 5.75. A block drafter proposes a whole block whether or not the block will survive; discarded draft tokens cost drafter time, which is cheap here, and the saving is counted in target passes avoided, which is expensive.

The practical instruction, scoped to what this is: in these runs, acceptance alone ranks the three structured arms in exactly the wrong order. Tune against end-to-end throughput on your own workload, and use acceptance and the tokens-per-target-step proxy as diagnostics rather than as targets. A reader who tunes for acceptance here lands on the vendor's documented configuration, which has the best acceptance number on this page and the second-worst throughput. Whether that generalizes to every block drafter is not something six prompts on one machine can establish.

These counters are not comparable with the acceptance figures on this site's Qwen pages, which came from a multi-token-prediction head, a different mechanism. We do not pool them and neither should anyone else.

05

The context ladder: the whole native context fits, with the drafter

Six loads, same weights, same build, same card. VRAM is the raw nvidia-smi reading after load against 32,607 MiB. Each decode figure is one 400-token generation from a 63-token prompt, a single observation per row, not a median. The speculative rows ran at --spec-draft-n-max 16.

ContextDrafterVRAM used, MiBFree, MiBLoad, sDecode t/s
8,192no16,97415,1132.0476.69
8,192yes19,37412,7133.04117.77
32,768no17,30714,7802.0376.74
32,768yes19,70912,3783.04117.65
131,072no18,64913,4382.0376.26
131,072yes21,05111,0363.04117.19

Speculative rows ran with --spec-draft-n-max 16. One load and one generation per row; all loads were page-cache warm.

The full native 131,072-token context fits on this card with the drafter loaded, using 21,051 MiB and leaving 11,036 MiB free. Across these one-shot rows the drafter's overhead was 2,400, 2,402 and 2,402 MiB, and decode with the drafter read 117.77, 117.65 and 117.19 tokens per second as allocated context grew sixteen-fold. That sample shows no material decline, and one generation per row cannot establish a context-performance curve. Note also that these generations start from a 63-token prompt: this is decode at allocated context, not at filled context, and section 08 says how far we actually filled it.

Why context is this cheap here

Only 13 of the 52 layers hold an unbounded key-value cache; the other 39 are capped at their 2,048-token window. Two key-value heads times 128 head dimension times two tensors times two bytes is 1 KiB per token per unbounded layer, so 13 KiB per token, so 1,664 MiB predicted for the unbounded cache at 131,072 tokens, which is 1.63 GiB.

And the honest version of that arithmetic

The measured quantity we can put beside it is the whole-process VRAM difference between the 8,192 and 131,072 loads without the drafter, 1,675 MiB or 1.64 GiB. Those two numbers are close, but they are not the same quantity: the 8,192 load already holds about 104 MiB of that cache, so the predicted increment is 1,560 MiB, and compute buffers grow with context too. Read 1.63 against 1.64 as confirmation that the layer pattern makes long context cheap by roughly the factor the configuration implies, not as a validated measurement of the cache itself. The runtime did not print a key-value cache size line at the verbosity we ran, so we do not have the direct reading.

One inference for smaller cards, marked as an inference

The most common question about a 30B build is whether it fits 24 GB, and this study cannot answer it, because we ran one card and one placement. What the table supports is arithmetic and nothing more: this model plus its drafter occupied 19,374 MiB at 8,192 tokens of context and 21,051 MiB at 131,072, on a card that reports 32,607 MiB, with the vision projector never loaded. We did not test a 24 GB card, and decode speed does not transfer between cards at all, so treat the fit as a starting estimate to verify yourself and the speeds on this page as belonging to this card only.

The comparison this site can make from its own records: the Qwen3.8-27B measured on this same card two days earlier declares 262,144 tokens native and did not fit that context on this card at an f16 cache, topping out at 131,072 with the draft head on or off; a 2026-09-17 run served the full window on that card with an 8-bit cache and the draft head off, see that card's dated update. Muse Glimmer's whole declared context fits with room for the drafter. Those are two different models with two different attention layouts, quantizations and dates, and the sentence is a fact about what fit on one card, not a ranking.

06

The license, and the document that ships beside it

Fact one. LICENSE in the base repository is 11,358 bytes of verbatim, unmodified Apache License 2.0, read in full at the pinned revision and retained in our record. Searching it for meta, usage polic, acceptable, restrict, revenue, MAU and monthly active returns zero hits. There is no revenue threshold, no monthly-active-user threshold and no field-of-use restriction inside the license. Its appendix still carries the stock unfilled placeholder:

Copyright [yyyy] [name of copyright owner]

That placeholder is how the unmodified Apache template ships, and we read it as one more piece of evidence that the file is the stock text rather than a modified grant. We draw nothing else from it. The same 11,358-byte file ships in the GGUF repository and in the drafter repository.

Fact two. A separate USAGE_POLICY.md of 5,230 bytes sits beside it in all three repositories, also read in full and retained. It states that the model "is not intended for individuals under the age of 18" and lists prohibited uses under five headings. Two clauses a practitioner should read before deploying: the bar on any action

"to intentionally circumvent or remove usage restrictions or other safety measures, or to enable functionality disabled by Meta"

and the bar on "Military, warfare, nuclear industries or applications, espionage".

Fact three, the structural one. The license does not incorporate the usage policy by reference, and the usage policy does not describe itself as a license condition. The two documents sit side by side and neither one points at the other. The Apache grant is unconditional on its face.

We make no claim about which document binds you. That is a legal question and this site measures things. What we can hand you is both files at their pinned revisions, which is what the data package does. What we will say plainly is that "it is Apache 2.0" and "there is a usage policy with prohibited uses" are both true at once, and quoting either one alone describes this release inaccurately.

07

Thinking: no off switch, a lever that is a sentence, and no floor on a simple prompt

Everything in this section is a property of a template and a build, not of the weights. It was measured on build b10453 with --jinja, serving Meta's documented long-named GGUF and therefore that file's baked tokenizer.chat_template, greedy, at 8,192 tokens of context. Section 09 shows the same model behaving differently from a file whose only difference is that template key.

There is no way to turn thinking off. The chat template opens the reasoning channel unconditionally: there is no enable_thinking branch and no branch that suppresses the channel. The vendor states this on their own card, and reading the template confirms it. What you control is how much.

The lever is reasoning_strength, and it is a sentence injected into your prompt. The template interpolates the value verbatim into a system line, Reasoning strength: <value>., defaulting to high. It is not validated. Sent through chat_template_kwargs on one prompt, one call per value, greedy, on 2026-08-16:

reasoning_strengthHidden reasoning, charactersCompletion tokens
low312121
medium459171
banana472163
high680224
field omitted680224

One call per value on one prompt, greedy, on this endpoint and configuration, 2026-08-16.

Three things in one table, and all five rows are one call each on one prompt. The documented values moved thinking volume in the documented direction here, low below medium below high. Omitting the field produced a response identical to high, which is the template's default. And banana was accepted: no error, no fallback, a reply that looks perfectly normal, with a thinking volume between medium and high, which is what you would expect from a model reading an English sentence about an undefined strength rather than consulting a table. Read the ordering as an illustration and the mechanism, prompt injection with no validation, as the finding: the mechanism is readable in the template source, which is why it does not rest on five calls.

The same mechanism class as Qwen's reasoning_effort, with the opposite outcome. The Qwen3.8-27B card found a knob that rewrites the prompt, raises HTTP 500 on an unrecognized value, and did not move measured thinking length in either direction. Here the knob also rewrites the prompt, accepts anything without complaint, and did move thinking volume. Neither one is a budget cap. If you need a cap, llama.cpp's --reasoning-budget is the flag, and we did not test it. reasoning_effort sent as a top-level field also reaches this template through llama.cpp's own plumbing on this build, measured at 388 reasoning characters for low against 535 for xhigh, one call each; that is the runtime's mapping, not a field the vendor's template defines, and the reasoning effort study is where that knob is read off the rendered prompts.

No budget floor on a simple prompt, and a floor of 2,000 on a harder one

Asked for the product of 17 and 23 with max_tokens set to 60, the model returned the correct answer, spending 56 tokens and finishing on stop rather than on a length cutoff. The same prompt at 500, 2,000 and 4,096 returned the identical answer with identical thinking. This is not a general property. On a harder prompt in the same session, our instrument's own budget ladder returned an empty visible answer at 60 and at 500 tokens, with the whole budget consumed by hidden reasoning, and first produced a visible answer at 2,000.

Both readings are here because both are true: the floor is a property of how much thinking a prompt provokes, not a constant of the model or of the endpoint. That is exactly the finding of our Empty Answer study, seen again on a different vendor's model. Do not copy either number into your client. Measure your own floor on your own prompts. For what it is worth as a starting point rather than a rule: on this endpoint, with this probe family, thinking replies were given 2,000 to 4,096 tokens.

08

Structured output conformed without constraining the surface

Dependencies for this section, because these are route behaviours rather than model properties: build b10453, --jinja, Meta's documented GGUF and its baked chat template, greedy sampling, 8,192 tokens of context, one request at a time.

Sending response_format of type json_schema with "strict": true, the returned content was this, fence included:

```json
{
  "city": "Reykjavik",
  "country": "Iceland",
  "population_millions": 0.14
}
```

The object inside is correct and conforms to the schema. The response is not parseable JSON, because the model wrapped it in a markdown fence and strict: true did not prevent that. That is one call, and it is the one whose request body was retained. Our instrument's structured-output cell separately recorded 3 of 3 conforming responses on its own schema prompt, so on this route the constraint held on content and not on surface.

The client instruction, scoped to that evidence: do not assume json.loads(message.content) is safe on this route. Strip fences first, or parse defensively. Two further calls in the same session showed the surface varying by prompt, one bare object and one fenced object with a sentence of prose appended after the closing fence; those two calls are illustrative only, because their request bodies were not retained and we will not build a claim on a record we cannot hand you. Whether the fence comes from the model, the template or the runtime's handling of strict is not something this study separated.

Tool calls work and arrive in the right shape on this build. The template emits Meta's ATEM XML shape, and build b10453 parses it into proper OpenAI tool_calls with the correct enumerated argument values, on both shots of the instrument's tool cell and on a direct probe. The vendor's card states that an older build leaks the raw to=self<|message|> marker into the visible content instead; that is their statement about earlier builds, not something we reproduced.

Retrieval. A planted needle was found at a server-reported 28,659 prompt tokens, served at the 32,768-token profile, one call. That is the longest prompt we actually put through this endpoint, and at 22 percent of the declared context it is not a long-context result. Nothing on this page speaks to behaviour at 131,072 filled tokens.

The rest of the battery, one to three calls per check: an exact-phrase echo came back byte-perfect; the bat-and-ball question was answered correctly; a request to implement interval merging produced code containing three assertions. The session record states that the extracted file was executed and exited clean; the exit status itself was not written to a retained artifact, so it is reported here as the session's account and not as a measurement.

09

Two files wearing one badge, two templates wearing one name

Before downloading anything we read the GGUF headers over HTTP range requests, about 13 MB per file, which is how the rest of this section was found without moving 30 GB.

Same badge, different file

Meta's KQuant-Dynamic-Q4_K_XL (19.65 GB, 237 Q6_K tensors) and Unsloth's UD-Q4_K_XL (15.88 GB, zero Q6_K tensors, 410 Q4_K) wear the same Q4_K_XL badge, are 3.78 GB apart, and differ structurally in their type mix. Unsloth's "XL" is smaller than Meta's plain Q4_K_M. A reader choosing "the Q4_K_XL" is not choosing one thing.

The community quant we measured is lighter and slightly faster, and we say nothing about its quality

Unsloth's UD-Q4_K_XL is 878,461,536 bytes smaller on disk than Meta's K-Quant-17GB build, which is 838 MiB. Served with no drafter, greedy, 8,192 tokens of context, one request at a time, across the same twelve prompts, one call each: 79.615 tokens per second pooled median against Meta's 76.325, about 4.3 percent faster (n = 12 per arm, pooled across both workloads). Quality was not measured, in either direction, on either file. A 4 percent decode difference is not a reason to switch quantizations by itself, and we have no evidence at all about what the different type mix does to output.

Two templates wearing one name

Meta's GGUF repository carries undocumented near-duplicates of three of its four files, short-named twins that do not appear in the repository's own files table. For the main model the twins are 2,848 bytes apart, carry identical tensor counts and identical type histograms, and differ in exactly one metadata key: tokenizer.chat_template. The undocumented template also still carries a pre-release codename in its error string, reading "Onyx ATEM chat template" where the documented one reads "Muse Glimmer".

The difference is behavioural, not cosmetic, and the vendor documents the fix in prose even though the superseded files stay in the repository. The documented template normalizes a caller-written Reasoning effort: line to Reasoning strength: and then suppresses its own injected default when the system prompt already carries one. The undocumented template does neither: it renders the caller's system content as written and appends its own default underneath. Unsloth's repository re-hosts the undocumented pair byte-identically, which is why this is a practitioner's problem and not a curiosity: a mirror, a re-upload or an internal cache can hand you the older template under a name that looks equivalent.

Measured consequence, same question, one call per cell, same build and --jinja in every row, 2026-08-16:

Served file and system promptPrompt tokensHidden reasoning, characters
Meta's documented file, plain system prompt49651
Meta's documented file, caller writes Reasoning effort: low.49385
Meta's documented file, caller writes Reasoning strength: low.49385
Unsloth's file, plain system prompt49650
Unsloth's file, caller writes Reasoning effort: low.55539
Unsloth's file, caller writes Reasoning strength: low.55340

One call per cell, greedy, on the same prompt, served from two different files on the same build.

On Meta's documented file the caller's own wording is normalized and deduplicated: writing Reasoning effort: low. and writing Reasoning strength: low. produce identical prompt lengths and identical thinking volume, and the directive is obeyed. On Unsloth's file the caller's line is added on top of the template's unchanged high default, and the result, 539 characters, sits between that file's own low and its own default. The caller's instruction competes with the default instead of replacing it. Prefer the long-named vendor files, and if you are not sure which template a file carries, the error string is the tell: the older one still says "Onyx".

A prediction this study registered, and then refuted

Before running the template tests we wrote down the expectation that the reasoning_strength lever would be dead on the undocumented template whenever the caller supplied a system message, on the reasoning that the undocumented template renders the system content as written. That was wrong. With a system message present, Unsloth's file moved from 650 characters of thinking at high to 354 at low, so the lever is live on both files. Reading the template source afterwards showed why: the undocumented template does inject the strength line, unconditionally, in its system branch. What it lacks is the normalization and the deduplication, not the injection. We print the refuted prediction because a registered expectation that survives contact with the measurement is worth no more than one that does not, unless both are reported.

10

The instrument, and what it got wrong

The plumbing cells in section 08 came from routecheck, the fixed battery this site publishes as an instrument, run once against this endpoint on 2026-08-16. Its card records nine tests with every request and response body retained.

What our instrument got wrong, and why we are printing it

The honesty cell, which asks the model to summarize a file that does not exist and to report a score on a benchmark that does not exist, graded FAIL, and reading the retained response shows a false negative. The model refused both halves cleanly and wrote, in the retained body, that it has "no verified record of a benchmark called" that name and therefore cannot provide a published score. The grader's list of hedging markers is scoped per topic and does not contain that phrasing, so a correct refusal scored as a failure. This false-negative pattern has appeared in this site's prior records, which makes the fix overdue rather than surprising, and the case here is a clear one. It is disclosed because anyone quoting the card would otherwise quote a wrong grade.

The reasoning-toggle cell also graded FAIL, correctly in mechanism and misleadingly in label: it tests the enable_thinking request field, this model has no such field and no off switch at all, and a test whose failure mode is "the toggle did nothing" cannot distinguish a broken toggle from a model that has none. The prefill cell graded INVALID: its whole token budget went to hidden reasoning and it measured nothing, so nothing on this page is quoted from it.

11

What we did not run

  • The vendor's own documented command. We ran their two speculation flags inside our fixed harness, not their full server line with the projector loaded and four slots open. Section 03 sets out the difference flag by flag.
  • Single request only. Every number on this page was measured with --parallel 1. Upstream issue #27117, open at the time of writing, reports draft-dflash acceptance collapsing under concurrent sequences, with sustained throughput below the no-speculation baseline in the reporter's own 16-way rows. That report is on AMD hardware through ROCm and its author states they cannot test CUDA. We did not test concurrency and make no claim about it, in either direction. Note that the vendor's own documented server command opens four slots.
  • Vision is untested. The model is natively multimodal; we never downloaded the projector. Upstream issue #26873 reports memory growth after first projector use, which is a reason to keep it that way until it is tested deliberately.
  • No quality measurement of any kind. Not between quantizations, not against the full-weight model, not on any benchmark. The vendor's own 1.0 percent degradation figure is quoted in section 01 as their claim.
  • No cold load. Every load on this page was page-cache warm because the files had just been written and hashed.
  • No filled long context. The ladder allocated up to 131,072 tokens; the longest prompt actually decoded from was 28,659 tokens.
  • One draft-gating value. Every speculation arm ran at the shipped p_min=0.00. The gate that mattered so much on the speculation-gate study was not swept here, so this page says nothing about it on this drafter.
  • One machine, one card, one placement. No 24 GB card, no partial offload, no alternate cache type, no batch or slot sweep.
  • No agentic loop and no multi-turn task, on a model whose card is built around agentic use.
  • No hosted comparison. Local endpoint only.

Every claim on this page is dated 2026-08-16, except the Qwen3.8-27B comparison in section 05, which includes a 2026-09-17 serve described on that card. Community quantizations get re-cut, vendors update templates, and llama.cpp changes defaults; the default at the centre of section 03 is one that could move, and if it does, the finding stops reproducing quietly. The diagnostic is the startup line: if your server prints n_max=3 above block_size=16, you are in the regime this page measured. If today is much later than that date, treat the version-specific findings as unverified since then.

12

Related pages on this site

The speculation-gate study is the properly designed version of section 03's thesis: five settings and 160 paired measurements on a 397B mixture-of-experts on this same workstation, where the runtime's shipped default made prose slower than no speculation at all. The Qwen3.8-27B field card is the same lesson on a dense 27B, and it is the comparison for section 07's thinking knob, which is the same mechanism class with the opposite outcome, and for section 05's context ceiling. Reasoning effort is a sentence, not a dial takes that comparison one level down on the Qwen side: the same class of knob read directly off the rendered prompts, where the injected system message can be counted in tokens rather than inferred from how much thinking came back. Empty Answer is section 07's budget-floor finding measured properly: what happens when hidden reasoning bills against the same budget as the visible answer, and how to find your own floor instead of trusting anyone's number.

13

Sources and artifacts

The evidence for this page ships as a data package: the set of files the run actually produced, not a summary of them. data/README.md maps every folder, states its provenance, and states the package's redactions plainly; data/NUMBERS.md maps every figure on this page to its file and field, and it opens with the one place a file will look like it contradicts the page.

  • The four Hugging Face repositories at their pinned revisions, read 2026-08-16: meta-models/Muse-Glimmer-30B (a4e59da5), meta-models/Muse-Glimmer-30B-GGUF (43c7eadd), meta-models/Muse-Glimmer-30B-assistant (e8192f3a) and unsloth/Muse-Glimmer-30B-GGUF (faa5b025). The file listings taken at those reads, with their LFS object ids, are in data/repos/.
  • The two license documents themselves, read in full and retained: data/licenses/LICENSE (11,358 bytes, SHA-256 cfc7749b96f63bd3...) and data/licenses/USAGE_POLICY.md (5,230 bytes). Section 06 is a reading of those two files and you can check it against them rather than against our summary.
  • The model's config.json, the drafter's config.json, and both chat templates, the documented one (9,992 bytes) and the undocumented twin (7,167 bytes), for every architecture and template fact in sections 01, 07 and 09 (data/config/).
  • The vendor's own readmes, retained in full as the vendor's work (data/vendor/), for the speed table quoted in section 03, the documented server command quoted in the same section, the drafter specification table with its Block size 16 row, and the quality claim in section 01.
  • The header probe of all six GGUF candidates, read over HTTP range requests before any download, for the tensor counts and type censuses in section 09 (data/headers/typegate.json).
  • The integrity record: every downloaded file's SHA-256 against the repository's own object id at the pinned revision (data/integrity/).
  • Our own runs of 2026-08-16: the five speed arms with their rows and their complete response bodies, twelve prompts each, the prompt text itself printed in section 02, and the five server logs that carry the n_max and block_size lines section 03 rests on (data/speed/); the six-load context ladder and its server logs (data/ladder/ctx_ladder.tsv); the probe battery (data/battery/); the template tests, whose driver carries the refuted prediction in its own docstring (data/templates/ and data/scripts/tmpl_test.py); and the instrument's card with all 28 of its retained bodies (data/routecheck/card.md).
  • The harness itself (data/scripts/): the twelve prompts, the single-pass loop, the discarded warm-up, and the fixed server flags every arm shares, so the method can be read rather than trusted. The release verification taken at the end of the evening is data/release/release_verification.txt.

If a number on this page disagrees with the record it came from, the record is right and the page is wrong; tell us and we will fix the page.