Muse Glimmer 30B: trained at 16, asked for 3
Meta trained its DFlash drafter to produce a block of 16 tokens per pass and prints that number in its specification table. llama.cpp's generic draft length is 3. The two flags Meta's GGUF card documents leave that generic default in charge, and on our six structured prompts the difference is worth 2.17 times the throughput, 116.505 tokens per second against 253.19. Prose tops out at 1.51 times our baseline whatever we do. Neither project's documentation is wrong on its own. The license is verbatim Apache 2.0, with a separate usage policy shipping beside it.
Meta published Muse Glimmer 30B with a speed table
that names our exact hardware: 74.9 tokens per second with no
speculation, 233.4 with its DFlash drafter, 3.1 times faster,
averaged across a prompt set they do not publish, at batch size 1
with greedy decoding on an Nvidia RTX 5090 through llama.cpp. On
2026-08-16 we downloaded the two files that table rests on, checked
every byte against the repository's own object ids, and ran five
server loads on one RTX 5090. On our harness the numbers
land where theirs do. Text-only, one slot, 8,192 tokens of
context, greedy, our no-speculation medians were
76.22 tokens per second on six prose prompts and
76.345 on six structured prompts, 1.8 and 1.9
percent above their published baseline, and our fastest structured
median, 253.19, is above their published 233.4.
That is a scoped comparison and not a reproduction of their
average: their prompt mix is unpublished, ours is printed in
section 02, and we never started their full documented command.
The finding is in the gap between two projects. Meta
trained DFlash to draft a block of 16 tokens in one forward pass and
prints Block size | 16 in its specification table.
llama.cpp's --spec-draft-n-max defaults to 3, a sensible
number for the ordinary autoregressive draft models that flag was
written for, and llama.cpp does not read the drafter's block size out
of the GGUF to set it. Meta's GGUF card turns speculation on with two
flags, -md and -ngld 99, and says nothing
about draft length, so the generic default stays in charge: the
server prints n_max=3 on the line directly above
block_size=16. In that configuration our six structured
prompts run at a median of 116.505 tokens per
second; adding --spec-draft-n-max 16 moves the same six
to 253.19, 2.17 times more. No document here
is wrong on its own, and the reader who follows the documentation
still loses that 2.17 times. And the workload decides more
than the flag does: with the same flag, prose reaches 1.51 times our
baseline at its best and structured output 3.32, so a single averaged
speedup, ours or anyone's, cannot tell a reader which of those two
their own work will get. Six prompts per cell, one call each, no
interval computed anywhere on this page. Everything below is one
quantization of one model on one machine, greedy, 8,192 tokens of
context, one request at a time, on 2026-08-16.
What it is
Rows below were read from the model's own configuration files and from the GGUF headers on this disk at the pinned revisions on 2026-08-16, not quoted from a marketing page. Where a row is the maker's claim rather than our reading or our measurement, it says so, and amber text marks it.
| Maker | Meta. The model card names Meta Superintelligence Lab as author and August 2026 as the release date. The repositories live in the meta-models organization on Hugging Face. |
|---|---|
| Weights | meta-models/Muse-Glimmer-30B at revision a4e59da52a7bc87ae7251dd5545c0dd437c44b68; the GGUF conversions we served come from meta-models/Muse-Glimmer-30B-GGUF at 43c7eadd41352a299ea8e0a36b3157978dd63596, both read on 2026-08-16. The base repository holds 2 safetensors shards, 59,553,435,272 bytes, and Hugging Face reports 29,776,626,688 parameters. |
| License | Verbatim Apache 2.0, 11,358 bytes, read in full at the pinned revision and retained in our record. A separate USAGE_POLICY.md of 5,230 bytes ships beside it in the same repository. Section 06 states both facts and makes no claim about which one binds you. |
| Shape | Dense. No expert blocks anywhere in the configuration, so the --n-cpu-moe lever that tunes this machine's mixture-of-experts models does not exist here. The quantization fits the card whole or it does not fit. |
| Architecture | 52 layers, hidden 6,656, feed-forward 19,968, 32 query heads to 2 key-value heads, head dimension 128, RoPE theta 500,000. The layer pattern is three sliding-attention layers with a 2,048-token window followed by one full-attention layer, repeated 13 times, so only 13 of the 52 layers hold an unbounded key-value cache. Section 05 is what that buys. |
| Drafter | A separate model, not a head inside the weights file: meta-models/Muse-Glimmer-30B-assistant, 2,555,985,152 parameters, 5 layers, block_size: 16, reading hidden features from 5 layers of the target. Its GGUF twin declares general.architecture = dflash. Per the vendor's card it is a block-diffusion drafter that predicts a whole block of 16 tokens in one forward pass, which is a different object from the multi-token-prediction and draft-head numbers elsewhere on this site. Sections 03 and 04 do not pool them. |
| Context | 131,072 tokens native, declared in the configuration. Section 05 measures the whole of it on this card. |
| Vocabulary | 202,048 tokens; beginning-of-sequence 200,000, end-of-sequence 200,001, end-of-turn 200,008. |
| Modality | Natively multimodal by its configuration: a 50-layer vision tower, patch 14, shipped as a separate projector file. We served it text-only and never downloaded the projector. Vision is untested here and no claim on this page rests on it. |
| Quantization served | Meta's own Muse-Glimmer-30B-KQuant-17GB-Q4_K_M.gguf, 16,756,683,904 bytes, plus dflash-Muse-Glimmer-30B-Q4_K_M.gguf, 1,631,208,128 bytes. These are the two files the vendor's own speed table names, which is why a comparison against that table is possible at all. Section 09 measures a community alternative. |
| Quality | The maker claims 1.0% degradation for this K-Quant-17GB build and 0.2% for K-Quant-Dynamic, "measured using an average on accuracy metrics across 15 common benchmarks". We measured no quality of any kind. That row is their claim, reproduced, not our finding. |
What we ran, exactly
| Machine | One desktop: Core Ultra 9 285K, 188 GiB RAM, one RTX 5090 with 32,607 MiB of VRAM, Pop!_OS. |
|---|---|
| Runtime | Mainline llama.cpp llama-server, build b10453, commit 3cb7ffb, CUDA 12.8, already on this machine and reused read-only. Muse Glimmer support merged upstream as PR #26841 on 2026-08-10 and first shipped in release b10353, so this build sits a hundred releases past the floor. |
| A version check that needs a release build | The vendor's card says to run llama-cli --version and require a build number of 10353 or higher. That works on a release binary. On a build from source llama.cpp itself prints version: 0.1.0-dev (build 1, commit 3cb7ffb), so the number is not there to test and the check is unusable in that case. Their second check works everywhere: grep -c LLM_ARCH_MUSE_GLIMMER src/llama-arch.cpp. |
| Integrity | Three files downloaded, three planned; every byte count equals the pinned manifest and every SHA-256 equals the repository's own LFS object id at the pinned revision. That is file integrity against the vendor's own conversion, which is the whole chain here, because Meta published these GGUFs itself. |
| Placement | -ngl 99, every layer on the GPU, plus -ngld 99 for the drafter. |
| Sampling | Greedy, temperature 0, matching the vendor's stated measurement method (batch size 1, greedy decoding). Their recommended chat sampling is temperature 1.0, top-p 0.95, top-k 64; we did not use it for speed work and cannot say what it would change. |
| Concurrency | --parallel 1. Single request throughout. Section 11 says why that boundary is drawn where it is. |
| Speed method | Five server loads, one per arm, each a fresh process. Each arm runs the same twelve prompts, six prose and six structured, each prompt sent exactly once, so no prompt is reused inside an arm and no prefix cache is warmed across the measured calls. A discarded warm-up call precedes each arm. The figure reported per cell is the median of six different prompts, one call each, not repetitions of one prompt, and the design is paired: arm to arm, the prompts are identical. Server-reported decode rates are authoritative. |
| Serving | A loopback scratch port for the study. Nothing was installed: no service unit, no registry key, no client wiring, and the server was stopped and verified stopped at the end. |
The exact invocation. Every arm in section 03 is this command with the speculation flags added or changed; nothing else moves:
llama-server \ --model Muse-Glimmer-30B-KQuant-17GB-Q4_K_M.gguf \ --host 127.0.0.1 --port <scratch> --alias muse-glimmer-30b \ --jinja --ctx-size 8192 --parallel 1 --n-gpu-layers 99 \ --threads 24 --threads-batch 24 --flash-attn on --timeout 3600
Arm B adds the two flags the vendor's card documents for speculation, and nothing more:
-md dflash-Muse-Glimmer-30B-Q4_K_M.gguf -ngld 99
Arms C and D add one flag beyond that,
--spec-draft-n-max 16 and
--spec-draft-n-max 8. Section 03 sets out exactly how
that relates to the vendor's own documented server command, which is
a different and longer thing that we did not run.
Our prompt set, since we are about to say that theirs is missing
Every arm sent these twelve, in this order, once each,
max_tokens 900, temperature 0. Six prose:
- Write four paragraphs about why coastal fishing villages in Iceland declined in the twentieth century.
- Write four paragraphs explaining how a mechanical watch escapement works, for a curious adult.
- Write four paragraphs about the history of the Dutch windmill and its role in land reclamation.
- Write four paragraphs describing the life cycle of a monarch butterfly and its migration.
- Write four paragraphs on how sourdough starter cultures work biologically.
- Write four paragraphs about the invention and spread of the printing press in Europe.
Six structured:
- Write a Python class LRUCache with get and put, using a dict and a doubly linked list. Include docstrings. Output only code.
- Write a JSON array of 12 objects, each with keys id, name, country, founded_year, describing 12 European universities. Output only JSON.
- Write a Python module with five functions for vector math (add, sub, dot, norm, normalize) with full type hints and docstrings. Output only code.
- Write a JSON object mapping 15 chemical element symbols to objects with keys name, atomic_number, group. Output only JSON.
- Write a Python implementation of a binary search tree with insert, find, and in-order traversal, fully commented. Output only code.
- Write a markdown table with 15 rows comparing 15 programming languages across columns: language, year, paradigm, typing, main use.
That set is ours, not a standard, and it is small. It is printed so that anyone can rerun it and so that no sentence on this page criticizes a missing prompt list from behind one.
How the numbers are rounded. Decode figures are printed as the server reports them. A median of six values falls between the two middle prompts, so some medians end in a half cent of a token per second: those are printed with three decimals (76.345, 116.505, 76.325, 79.615) rather than rounded, and every ratio on this page is computed from unrounded values.
Startup noise that is not a fault. Every load with
the drafter prints dflash requires ctx_other to be set and
then [spec] failed to measure draft model memory. The
vendor documents the second one as harmless, and the model loads and
serves normally. Every load of every Muse GGUF here also prints
special_eot_id is not in special_eog_ids. We observed no
stop-behaviour problem in any run, so we record that line as an open
observation and not as a defect.
Every VRAM number on this page is the raw MiB reading from
nvidia-smi against a 32,607 MiB card. We print
MiB first because that is what the instrument reports, and we never
divide a MiB reading by 1000 and call the result GiB.
The numbers side by side, and the default that decides the gap
Meta's card publishes this table, with the method stated underneath it: an average across a diverse prompt set, batch size 1, greedy decoding, the RTX measurements taken with llama.cpp.
| GPU | Baseline, no speculation | With DFlash | Speedup |
|---|---|---|---|
| Nvidia RTX 5090 | 74.9 | 233.4 | 3.1x |
The vendor's own figures, quoted from their GGUF model card as read on 2026-08-16.
The reported hardware, runtime, decoding mode and files are specified, and they are ours. The prompt mix is not. Here is our own measurement beside it, with the prompt set split in two and printed in section 02, greedy, 8,192 tokens of context, one request at a time, on 2026-08-16.
| Arm | Speculation flags | Prose t/s | vs baseline | Structured t/s | vs baseline |
|---|---|---|---|---|---|
| A | none | 76.22 | baseline | 76.345 | baseline |
| B | the two flags the card documents | 79.07 | 1.04x | 116.505 | 1.53x |
| D | B plus --spec-draft-n-max 8 | 114.46 | 1.50x | 208.29 | 2.73x |
| C | B plus --spec-draft-n-max 16 | 114.86 | 1.51x | 253.19 | 3.32x |
Each cell is the median of six different prompts, one call per prompt (n = 6), inside one server process, with the same twelve prompts in every arm. The "vs baseline" columns divide by arm A of the same workload, never by the vendor's number. The arms are lettered in the order they were run and the table is ordered by draft length, which is why D sits above C. No interval is computed and none is implied; this is a diagnostic ladder, not a benchmark.
Our baseline lands where theirs does
76.22 and 76.345 against their 74.9 is 1.8 and 1.9 percent above the vendor's figure on a machine they did not configure, which is as close as anyone should expect two different desktops to land.
Our fastest arm is above their published DFlash figure, on structured prompts only
253.19 tokens per second is 8.5 percent above the 233.4 they published, and 3.32 times our own baseline against their 3.1 times theirs (n = 6 structured prompts, one call each). Those two ratios are not the same estimand: theirs is an average over an unpublished mix of unknown composition, ours is a median over six structured prompts we wrote and printed. They are printed side by side because they are the two numbers a reader will hold, not because one reproduces the other.
Where the difference comes from: one line in the server log
With the two documented flags and nothing else, the drafter registers like this:
common_speculative_impl_draft_dflash: - n_max=3, n_min=0, p_min=0.00 common_speculative_impl_draft_dflash: - block_size=16, mask_token_id=201818, n_extract=5
n_max=3 is llama.cpp's own generic default draft
length, inherited from ordinary autoregressive draft models.
block_size=16 is what this drafter was trained to produce
in a single forward pass, and the runtime prints it on the very next
line without using it. A block drafter asked for three tokens at a
time throws away most of what it computed. Ask for the block it was
built for and the same log reads:
common_speculative_impl_draft_dflash: - n_max=16, n_min=0, p_min=0.00 common_speculative_impl_draft_dflash: requested draft size (n_max=16, n_min=0) exceeds the trained block size 16 -- clamping to 15
The runtime clamps 16 down to 15, and 15 is where the fast
arm actually ran. Pass 16 and accept 15; passing 15 yourself
would land in the same place without the warning. Arm D, at
--spec-draft-n-max 8, draws no warning and gets most of
the way there.
Three separate things are each individually correct.
Meta trained DFlash at block size 16 and publishes
that number in the drafter's specification table on both readmes; it
is an architecture row, not a command-line instruction, and it does
not claim to be one. llama.cpp's
--spec-draft-n-max defaults to 3, which is a
reasonable default for the ordinary autoregressive draft models the
speculative path was built for; the flag is llama.cpp's, documented
by llama.cpp, and on this build the runtime does not read the
drafter's block size out of the GGUF to raise its own default.
Meta's GGUF card documents the two flags that turn
speculation on, -md and -ngld 99,
which is exactly what a reader needs to load the drafter, and says
nothing about draft length, which leaves the generic default in
charge.
The string --spec-draft-n-max appears nowhere in
either vendor readme; we checked both retained copies and the grep
returns zero hits. That is a fact about the readmes, not an
accusation: a vendor card is not obliged to teach another project's
flags. The consequence is still real. A practitioner who reads
Meta's card carefully and llama.cpp's defaults not at all gets 1.53
times our baseline on structured output and 1.04 on prose, with no
documented reason to suspect that one more flag is worth
2.17 times more on the structured half, 253.19
against 116.505, six prompts per cell, one call each. This is what a
gap between two correct documents costs, and it is the kind of gap
that only shows up when somebody runs both projects together and
reads the log.
What arm B is, exactly, and what it is not
This is the sentence a reviewer should attack, so here is the whole of it. The vendor's GGUF card documents this server command:
./build/bin/llama-server \
-m Muse-Glimmer-30B-GGUF/Muse-Glimmer-30B-KQuant-17GB-Q4_K_M.gguf \
--mmproj Muse-Glimmer-30B-GGUF/mmproj-Muse-Glimmer-30B-Q4_K_M.gguf \
-a muse-glimmer-30B \
-ngl 99 -c 131072 -np 4 \
--host 127.0.0.1 --port 8080 \
--jinja \
--temp 1.0 --top-p 0.95 --top-k 64
and then documents speculation as an addition to it:
-md Muse-Glimmer-30B-GGUF/dflash-Muse-Glimmer-30B-Q4_K_M.gguf
-ngld 99. Arm B reproduces that two-flag addition
verbatim, on the same drafter file, inside our own harness. It does
not reproduce the command above. Ours is text-only with no
projector, one slot instead of four, 8,192 tokens of context instead
of 131,072, and greedy instead of their recommended chat sampling,
because greedy at batch size 1 is what the vendor's own measurement
footnote specifies. So we did not measure the vendor's
documented command and this page makes no claim about what it would
produce. What we can say is narrower and still useful,
because --spec-draft-n-max is absent from that command
too: any invocation built from that page inherits llama.cpp's generic
default, and our arm B log is a retained instance of what that looks
like at startup. The argv of each arm was not itself written to a
file; the flags are evidenced by the harness script and by three lines
in each retained server log, the drafter's path, the
n_max line, and
n_slots = 1, n_ctx_slot = 8192. Their four-slot default
is also the regime that upstream issue #27117 reports acceptance
collapsing in; see section 11.
The split is the point
The same flag that buys 3.32x on structured output buys 1.51x on prose. Without the mix, a vendor average cannot tell a reader which result to expect for their own workload, and ours cannot either, which is why both halves are printed separately and never pooled into one number. Our speculation-gate study measured a 397B mixture-of-experts where an ungated draft made prose slower than no speculation at all while making structured output faster, and the Qwen3.8-27B field card found the same direction on a dense 27B living entirely on the GPU. This is a third deployment, a different drafter mechanism, and the same lesson about pooled numbers. The magnitudes do not transfer between those pages and this one, and nothing here reproduces their numbers.
Acceptance rate misranks the arms on this drafter
The obvious diagnostic for speculative decoding is the acceptance rate, and on this drafter it points the wrong way. Same runs as section 03, counters summed across the six prompts of each workload (n = 6 per row, one call per prompt).
| Arm | Workload | Draft tokens generated | Accepted | Acceptance | Tokens per target step | Throughput |
|---|---|---|---|---|---|---|
| B | structured | 4,601 | 3,408 | 0.741 | 3.22 | 116.505 |
| D | structured | 9,023 | 3,964 | 0.439 | 4.49 | 208.29 |
| C | structured | 13,204 | 4,212 | 0.319 | 5.75 | 253.19 |
| B | prose | 6,926 | 2,689 | 0.388 | 2.16 | 79.07 |
| D | prose | 16,360 | 2,875 | 0.176 | 2.40 | 114.46 |
| C | prose | 28,589 | 3,013 | 0.105 | 2.57 | 114.86 |
Acceptance is accepted draft tokens divided by
generated draft tokens, from the server's own counters. Tokens per
target step is completion tokens divided by (completion tokens minus
accepted draft tokens), which is the same arithmetic the server prints
per request as mean len. It is an effective-work proxy
and a diagnostic: it is not a recorded count of target-model forward
passes, because the server does not log one.
Acceptance falls as throughput rises, from 0.741 to 0.319 on structured output while the same six prompts get 2.17 times faster. The proxy moves with the speed instead: 3.22, then 4.49, then 5.75. A block drafter proposes a whole block whether or not the block will survive; discarded draft tokens cost drafter time, which is cheap here, and the saving is counted in target passes avoided, which is expensive.
The practical instruction, scoped to what this is: in these runs, acceptance alone ranks the three structured arms in exactly the wrong order. Tune against end-to-end throughput on your own workload, and use acceptance and the tokens-per-target-step proxy as diagnostics rather than as targets. A reader who tunes for acceptance here lands on the vendor's documented configuration, which has the best acceptance number on this page and the second-worst throughput. Whether that generalizes to every block drafter is not something six prompts on one machine can establish.
These counters are not comparable with the acceptance figures on this site's Qwen pages, which came from a multi-token-prediction head, a different mechanism. We do not pool them and neither should anyone else.
The context ladder: the whole native context fits, with the drafter
Six loads, same weights, same build, same card. VRAM is the raw
nvidia-smi reading after load against 32,607 MiB.
Each decode figure is one 400-token generation from a
63-token prompt, a single observation per row, not a median.
The speculative rows ran at
--spec-draft-n-max 16.
| Context | Drafter | VRAM used, MiB | Free, MiB | Load, s | Decode t/s |
|---|---|---|---|---|---|
| 8,192 | no | 16,974 | 15,113 | 2.04 | 76.69 |
| 8,192 | yes | 19,374 | 12,713 | 3.04 | 117.77 |
| 32,768 | no | 17,307 | 14,780 | 2.03 | 76.74 |
| 32,768 | yes | 19,709 | 12,378 | 3.04 | 117.65 |
| 131,072 | no | 18,649 | 13,438 | 2.03 | 76.26 |
| 131,072 | yes | 21,051 | 11,036 | 3.04 | 117.19 |
Speculative rows ran with
--spec-draft-n-max 16. One load and one generation per
row; all loads were page-cache warm.
The full native 131,072-token context fits on this card with the drafter loaded, using 21,051 MiB and leaving 11,036 MiB free. Across these one-shot rows the drafter's overhead was 2,400, 2,402 and 2,402 MiB, and decode with the drafter read 117.77, 117.65 and 117.19 tokens per second as allocated context grew sixteen-fold. That sample shows no material decline, and one generation per row cannot establish a context-performance curve. Note also that these generations start from a 63-token prompt: this is decode at allocated context, not at filled context, and section 08 says how far we actually filled it.
Why context is this cheap here
Only 13 of the 52 layers hold an unbounded key-value cache; the other 39 are capped at their 2,048-token window. Two key-value heads times 128 head dimension times two tensors times two bytes is 1 KiB per token per unbounded layer, so 13 KiB per token, so 1,664 MiB predicted for the unbounded cache at 131,072 tokens, which is 1.63 GiB.
The measured quantity we can put beside it is the whole-process VRAM difference between the 8,192 and 131,072 loads without the drafter, 1,675 MiB or 1.64 GiB. Those two numbers are close, but they are not the same quantity: the 8,192 load already holds about 104 MiB of that cache, so the predicted increment is 1,560 MiB, and compute buffers grow with context too. Read 1.63 against 1.64 as confirmation that the layer pattern makes long context cheap by roughly the factor the configuration implies, not as a validated measurement of the cache itself. The runtime did not print a key-value cache size line at the verbosity we ran, so we do not have the direct reading.
One inference for smaller cards, marked as an inference
The most common question about a 30B build is whether it fits 24 GB, and this study cannot answer it, because we ran one card and one placement. What the table supports is arithmetic and nothing more: this model plus its drafter occupied 19,374 MiB at 8,192 tokens of context and 21,051 MiB at 131,072, on a card that reports 32,607 MiB, with the vision projector never loaded. We did not test a 24 GB card, and decode speed does not transfer between cards at all, so treat the fit as a starting estimate to verify yourself and the speeds on this page as belonging to this card only.
The comparison this site can make from its own records: the Qwen3.8-27B measured on this same card two days earlier declares 262,144 tokens native and did not fit that context on this card at an f16 cache, topping out at 131,072 with the draft head on or off; a 2026-09-17 run served the full window on that card with an 8-bit cache and the draft head off, see that card's dated update. Muse Glimmer's whole declared context fits with room for the drafter. Those are two different models with two different attention layouts, quantizations and dates, and the sentence is a fact about what fit on one card, not a ranking.
The license, and the document that ships beside it
Fact one. LICENSE in the base
repository is 11,358 bytes of verbatim, unmodified Apache
License 2.0, read in full at the pinned revision and retained
in our record. Searching it for meta,
usage polic, acceptable,
restrict, revenue, MAU and
monthly active returns zero hits. There is no revenue
threshold, no monthly-active-user threshold and no field-of-use
restriction inside the license. Its appendix still carries the stock
unfilled placeholder:
Copyright [yyyy] [name of copyright owner]
That placeholder is how the unmodified Apache template ships, and we read it as one more piece of evidence that the file is the stock text rather than a modified grant. We draw nothing else from it. The same 11,358-byte file ships in the GGUF repository and in the drafter repository.
Fact two. A separate USAGE_POLICY.md
of 5,230 bytes sits beside it in all three repositories, also read in
full and retained. It states that the model "is not intended for
individuals under the age of 18" and lists prohibited uses under five
headings. Two clauses a practitioner should read before deploying: the
bar on any action
"to intentionally circumvent or remove usage restrictions or other safety measures, or to enable functionality disabled by Meta"
and the bar on "Military, warfare, nuclear industries or applications, espionage".
Fact three, the structural one. The license does not incorporate the usage policy by reference, and the usage policy does not describe itself as a license condition. The two documents sit side by side and neither one points at the other. The Apache grant is unconditional on its face.
We make no claim about which document binds you. That is a legal question and this site measures things. What we can hand you is both files at their pinned revisions, which is what the data package does. What we will say plainly is that "it is Apache 2.0" and "there is a usage policy with prohibited uses" are both true at once, and quoting either one alone describes this release inaccurately.
Thinking: no off switch, a lever that is a sentence, and no floor on a simple prompt
Everything in this section is a property of a template and
a build, not of the weights. It was measured on build
b10453 with --jinja, serving Meta's
documented long-named GGUF and therefore that file's baked
tokenizer.chat_template, greedy, at 8,192 tokens of
context. Section 09 shows the same model behaving differently from a
file whose only difference is that template key.
There is no way to turn thinking off. The chat
template opens the reasoning channel unconditionally: there is no
enable_thinking branch and no branch that suppresses the
channel. The vendor states this on their own card, and reading the
template confirms it. What you control is how much.
The lever is reasoning_strength, and it is a
sentence injected into your prompt. The template interpolates
the value verbatim into a system line,
Reasoning strength: <value>., defaulting to
high. It is not validated. Sent through
chat_template_kwargs on one prompt, one call per value,
greedy, on 2026-08-16:
| reasoning_strength | Hidden reasoning, characters | Completion tokens |
|---|---|---|
low | 312 | 121 |
medium | 459 | 171 |
banana | 472 | 163 |
high | 680 | 224 |
| field omitted | 680 | 224 |
One call per value on one prompt, greedy, on this endpoint and configuration, 2026-08-16.
Three things in one table, and all five rows are one call each on
one prompt. The documented values moved thinking volume in the
documented direction here, low below medium below high.
Omitting the field produced a response identical to
high, which is the template's default. And
banana was accepted: no error, no
fallback, a reply that looks perfectly normal, with a thinking volume
between medium and high, which is what you
would expect from a model reading an English sentence about an
undefined strength rather than consulting a table. Read the ordering
as an illustration and the mechanism, prompt injection with no
validation, as the finding: the mechanism is readable in the template
source, which is why it does not rest on five calls.
The same mechanism class as Qwen's
reasoning_effort, with the opposite outcome. The
Qwen3.8-27B card found a knob that
rewrites the prompt, raises HTTP 500 on an unrecognized value, and did
not move measured thinking length in either direction. Here the knob
also rewrites the prompt, accepts anything without complaint, and did
move thinking volume. Neither one is a budget cap. If you need a cap,
llama.cpp's --reasoning-budget is the flag, and we did not
test it. reasoning_effort sent as a top-level field also
reaches this template through llama.cpp's own plumbing on this build,
measured at 388 reasoning characters for low against 535
for xhigh, one call each; that is the runtime's mapping,
not a field the vendor's template defines, and the
reasoning effort study is where that
knob is read off the rendered prompts.
No budget floor on a simple prompt, and a floor of 2,000 on a harder one
Asked for the product of 17 and 23 with max_tokens set
to 60, the model returned the correct answer, spending 56 tokens and
finishing on stop rather than on a length cutoff. The same
prompt at 500, 2,000 and 4,096 returned the identical answer with
identical thinking. This is not a general property.
On a harder prompt in the
same session, our instrument's own budget ladder returned an empty
visible answer at 60 and at 500 tokens, with the whole budget consumed
by hidden reasoning, and first produced a visible answer at 2,000.
Both readings are here because both are true: the floor is a property of how much thinking a prompt provokes, not a constant of the model or of the endpoint. That is exactly the finding of our Empty Answer study, seen again on a different vendor's model. Do not copy either number into your client. Measure your own floor on your own prompts. For what it is worth as a starting point rather than a rule: on this endpoint, with this probe family, thinking replies were given 2,000 to 4,096 tokens.
Structured output conformed without constraining the surface
Dependencies for this section, because these are
route behaviours rather than model properties: build
b10453, --jinja, Meta's documented GGUF and
its baked chat template, greedy sampling, 8,192 tokens of context, one
request at a time.
Sending response_format of type
json_schema with "strict": true, the
returned content was this, fence included:
```json
{
"city": "Reykjavik",
"country": "Iceland",
"population_millions": 0.14
}
```
The object inside is correct and conforms to the schema.
The response is not parseable JSON, because the model
wrapped it in a markdown fence and strict: true did not
prevent that. That is one call, and it is the one whose request body
was retained. Our instrument's structured-output cell separately
recorded 3 of 3 conforming responses on its own schema prompt, so on
this route the constraint held on content and not on surface.
The client instruction, scoped to that evidence: do not
assume json.loads(message.content) is safe on this route.
Strip fences first, or parse defensively. Two further calls
in the same session showed the surface varying by prompt, one bare
object and one fenced object with a sentence of prose appended after
the closing fence; those two calls are illustrative only, because
their request bodies were not retained and we will not build a claim
on a record we cannot hand you. Whether the fence comes from the
model, the template or the runtime's handling of strict
is not something this study separated.
Tool calls work and arrive in the right shape on this
build. The template emits Meta's ATEM XML shape, and build
b10453 parses it into proper OpenAI
tool_calls with the correct enumerated argument values,
on both shots of the instrument's tool cell and on a direct probe. The
vendor's card states that an older build leaks the raw
to=self<|message|> marker into the visible content
instead; that is their statement about earlier builds, not something
we reproduced.
Retrieval. A planted needle was found at a server-reported 28,659 prompt tokens, served at the 32,768-token profile, one call. That is the longest prompt we actually put through this endpoint, and at 22 percent of the declared context it is not a long-context result. Nothing on this page speaks to behaviour at 131,072 filled tokens.
The rest of the battery, one to three calls per check: an exact-phrase echo came back byte-perfect; the bat-and-ball question was answered correctly; a request to implement interval merging produced code containing three assertions. The session record states that the extracted file was executed and exited clean; the exit status itself was not written to a retained artifact, so it is reported here as the session's account and not as a measurement.
Two files wearing one badge, two templates wearing one name
Before downloading anything we read the GGUF headers over HTTP range requests, about 13 MB per file, which is how the rest of this section was found without moving 30 GB.
Same badge, different file
Meta's KQuant-Dynamic-Q4_K_XL (19.65 GB, 237
Q6_K tensors) and Unsloth's UD-Q4_K_XL (15.88 GB,
zero Q6_K tensors, 410 Q4_K) wear the same
Q4_K_XL badge, are 3.78 GB apart, and differ
structurally in their type mix. Unsloth's "XL" is smaller than Meta's
plain Q4_K_M. A reader choosing "the Q4_K_XL" is not
choosing one thing.
The community quant we measured is lighter and slightly faster, and we say nothing about its quality
Unsloth's UD-Q4_K_XL is 878,461,536 bytes smaller on
disk than Meta's K-Quant-17GB build, which is 838 MiB. Served
with no drafter, greedy, 8,192 tokens of context, one request at a
time, across the same twelve prompts, one call each:
79.615 tokens per second pooled median against Meta's 76.325,
about 4.3 percent faster (n = 12 per arm, pooled across both
workloads). Quality was not measured, in either
direction, on either file. A 4 percent decode difference is
not a reason to switch quantizations by itself, and we have no
evidence at all about what the different type mix does to output.
Two templates wearing one name
Meta's GGUF repository carries undocumented near-duplicates of
three of its four files, short-named twins that do not appear in the
repository's own files table. For the main model the twins are 2,848
bytes apart, carry identical tensor counts and identical type
histograms, and differ in exactly one metadata key:
tokenizer.chat_template. The undocumented template also
still carries a pre-release codename in its error string, reading
"Onyx ATEM chat template" where the documented one reads "Muse
Glimmer".
The difference is behavioural, not cosmetic, and the vendor
documents the fix in prose even though the superseded files stay in
the repository. The documented template normalizes a caller-written
Reasoning effort: line to
Reasoning strength: and then suppresses its own injected
default when the system prompt already carries one. The undocumented
template does neither: it renders the caller's system content as
written and appends its own default underneath. Unsloth's
repository re-hosts the undocumented pair byte-identically,
which is why this is a practitioner's problem and not a curiosity: a
mirror, a re-upload or an internal cache can hand you the older
template under a name that looks equivalent.
Measured consequence, same question, one call per cell, same build
and --jinja in every row, 2026-08-16:
| Served file and system prompt | Prompt tokens | Hidden reasoning, characters |
|---|---|---|
| Meta's documented file, plain system prompt | 49 | 651 |
Meta's documented file, caller writes Reasoning effort: low. | 49 | 385 |
Meta's documented file, caller writes Reasoning strength: low. | 49 | 385 |
| Unsloth's file, plain system prompt | 49 | 650 |
Unsloth's file, caller writes Reasoning effort: low. | 55 | 539 |
Unsloth's file, caller writes Reasoning strength: low. | 55 | 340 |
One call per cell, greedy, on the same prompt, served from two different files on the same build.
On Meta's documented file the caller's own wording is normalized
and deduplicated: writing Reasoning effort: low. and
writing Reasoning strength: low. produce identical prompt
lengths and identical thinking volume, and the directive is obeyed. On
Unsloth's file the caller's line is added on top of the template's
unchanged high default, and the result, 539 characters,
sits between that file's own low and its own default.
The caller's instruction competes with the default instead of
replacing it. Prefer the long-named vendor files, and if you
are not sure which template a file carries, the error string is the
tell: the older one still says "Onyx".
Before running the template tests we wrote down the expectation
that the reasoning_strength lever would be
dead on the undocumented template whenever the
caller supplied a system message, on the reasoning that the
undocumented template renders the system content as written.
That was wrong. With a system message present,
Unsloth's file moved from 650 characters of thinking at
high to 354 at low, so the lever is live
on both files. Reading the template source afterwards showed why:
the undocumented template does inject the strength line,
unconditionally, in its system branch. What it lacks is the
normalization and the deduplication, not the injection. We print the
refuted prediction because a registered expectation that survives
contact with the measurement is worth no more than one that does
not, unless both are reported.
The instrument, and what it got wrong
The plumbing cells in section 08 came from routecheck, the fixed battery this site publishes as an instrument, run once against this endpoint on 2026-08-16. Its card records nine tests with every request and response body retained.
The honesty cell, which asks the model to summarize a file that does not exist and to report a score on a benchmark that does not exist, graded FAIL, and reading the retained response shows a false negative. The model refused both halves cleanly and wrote, in the retained body, that it has "no verified record of a benchmark called" that name and therefore cannot provide a published score. The grader's list of hedging markers is scoped per topic and does not contain that phrasing, so a correct refusal scored as a failure. This false-negative pattern has appeared in this site's prior records, which makes the fix overdue rather than surprising, and the case here is a clear one. It is disclosed because anyone quoting the card would otherwise quote a wrong grade.
The reasoning-toggle cell also graded FAIL, correctly in mechanism
and misleadingly in label: it tests the
enable_thinking request field, this model has no such
field and no off switch at all, and a test whose failure mode is "the
toggle did nothing" cannot distinguish a broken toggle from a model
that has none. The prefill cell graded INVALID: its whole token budget
went to hidden reasoning and it measured nothing, so nothing on this
page is quoted from it.
What we did not run
- The vendor's own documented command. We ran their two speculation flags inside our fixed harness, not their full server line with the projector loaded and four slots open. Section 03 sets out the difference flag by flag.
- Single request only. Every number on this page
was measured with
--parallel 1. Upstream issue #27117, open at the time of writing, reportsdraft-dflashacceptance collapsing under concurrent sequences, with sustained throughput below the no-speculation baseline in the reporter's own 16-way rows. That report is on AMD hardware through ROCm and its author states they cannot test CUDA. We did not test concurrency and make no claim about it, in either direction. Note that the vendor's own documented server command opens four slots. - Vision is untested. The model is natively multimodal; we never downloaded the projector. Upstream issue #26873 reports memory growth after first projector use, which is a reason to keep it that way until it is tested deliberately.
- No quality measurement of any kind. Not between quantizations, not against the full-weight model, not on any benchmark. The vendor's own 1.0 percent degradation figure is quoted in section 01 as their claim.
- No cold load. Every load on this page was page-cache warm because the files had just been written and hashed.
- No filled long context. The ladder allocated up to 131,072 tokens; the longest prompt actually decoded from was 28,659 tokens.
- One draft-gating value. Every speculation arm
ran at the shipped
p_min=0.00. The gate that mattered so much on the speculation-gate study was not swept here, so this page says nothing about it on this drafter. - One machine, one card, one placement. No 24 GB card, no partial offload, no alternate cache type, no batch or slot sweep.
- No agentic loop and no multi-turn task, on a model whose card is built around agentic use.
- No hosted comparison. Local endpoint only.
Every claim on this page is dated 2026-08-16,
except the Qwen3.8-27B comparison in section 05, which includes a
2026-09-17 serve described on that card.
Community quantizations get re-cut, vendors update templates, and
llama.cpp changes defaults; the default at the centre of section 03 is
one that could move, and if it does, the finding stops reproducing
quietly. The diagnostic is the startup line: if your server prints
n_max=3 above block_size=16, you are in the
regime this page measured. If today is much later than that date,
treat the version-specific findings as unverified since then.
Related pages on this site
The speculation-gate study is the properly designed version of section 03's thesis: five settings and 160 paired measurements on a 397B mixture-of-experts on this same workstation, where the runtime's shipped default made prose slower than no speculation at all. The Qwen3.8-27B field card is the same lesson on a dense 27B, and it is the comparison for section 07's thinking knob, which is the same mechanism class with the opposite outcome, and for section 05's context ceiling. Reasoning effort is a sentence, not a dial takes that comparison one level down on the Qwen side: the same class of knob read directly off the rendered prompts, where the injected system message can be counted in tokens rather than inferred from how much thinking came back. Empty Answer is section 07's budget-floor finding measured properly: what happens when hidden reasoning bills against the same budget as the visible answer, and how to find your own floor instead of trusting anyone's number.
Sources and artifacts
The evidence for this page ships as a data package: the set of files the run actually produced, not a summary of them. data/README.md maps every folder, states its provenance, and states the package's redactions plainly; data/NUMBERS.md maps every figure on this page to its file and field, and it opens with the one place a file will look like it contradicts the page.
- The four Hugging Face repositories at their pinned revisions,
read 2026-08-16:
meta-models/Muse-Glimmer-30B(a4e59da5),meta-models/Muse-Glimmer-30B-GGUF(43c7eadd),meta-models/Muse-Glimmer-30B-assistant(e8192f3a) andunsloth/Muse-Glimmer-30B-GGUF(faa5b025). The file listings taken at those reads, with their LFS object ids, are indata/repos/. - The two license documents themselves, read in full and retained: data/licenses/LICENSE (11,358 bytes, SHA-256 cfc7749b96f63bd3...) and data/licenses/USAGE_POLICY.md (5,230 bytes). Section 06 is a reading of those two files and you can check it against them rather than against our summary.
- The model's
config.json, the drafter'sconfig.json, and both chat templates, the documented one (9,992 bytes) and the undocumented twin (7,167 bytes), for every architecture and template fact in sections 01, 07 and 09 (data/config/). - The vendor's own readmes, retained in full as the vendor's work
(
data/vendor/), for the speed table quoted in section 03, the documented server command quoted in the same section, the drafter specification table with itsBlock size 16row, and the quality claim in section 01. - The header probe of all six GGUF candidates, read over HTTP range requests before any download, for the tensor counts and type censuses in section 09 (data/headers/typegate.json).
- The integrity record: every downloaded file's SHA-256 against the repository's own object id at the pinned revision (data/integrity/).
- Our own runs of 2026-08-16: the five speed arms
with their rows and their complete response bodies, twelve prompts
each, the prompt text itself printed in section 02, and the five
server logs that carry the
n_maxandblock_sizelines section 03 rests on (data/speed/); the six-load context ladder and its server logs (data/ladder/ctx_ladder.tsv); the probe battery (data/battery/); the template tests, whose driver carries the refuted prediction in its own docstring (data/templates/and data/scripts/tmpl_test.py); and the instrument's card with all 28 of its retained bodies (data/routecheck/card.md). - The harness itself (
data/scripts/): the twelve prompts, the single-pass loop, the discarded warm-up, and the fixed server flags every arm shares, so the method can be read rather than trusted. The release verification taken at the end of the evening is data/release/release_verification.txt.
If a number on this page disagrees with the record it came from, the record is right and the page is wrong; tell us and we will fix the page.