GLM-5.3-Flash on one desktop
At 262,144 tokens, the installed micro-batch setting was measured through 47,992 tokens at 142.6 t/s. A guarded fixture passed only in isolation.
Measured: Unsloth's four-shard UD-Q3_K_XL file runs through a pinned GLM-specific llama.cpp fork. The installed launcher and the 21 and 26 September runs place 42 blocks' experts in system memory. The 16 September -cmoe launch passed no explicit expert-block count; its finished 181-token answer decoded at 9.58 tokens per second. On 21 September, the published comparison was a first request at 24,008 tokens and 80.7 tokens per second, with 62 generated tokens at 9.22 tokens per second, against a third request on the tuned load at 20,333 tokens and 446.0 tokens per second, with 60 generated tokens at 9.35 tokens per second. The fourth request on that tuned load read 48,168 tokens at 425.1 tokens per second and generated 310 tokens at 9.12 tokens per second. The prompt lengths still differ, so the 5.5 times ratio is descriptive. On 26 September, -ub 1024 at 262,144, the setting that ships, was measured only through 47,992 tokens at 142.6 tokens per second. A separate -ub 2048 run read 230,039 tokens and returned all three planted codes; that is not the installed setting. -ub 4096 failed to load when a 13,281.37 MiB allocation was refused. On 17 September, one synthetic high-effort request put more reasoning and a literal </think> in the visible answer. An isolated guard passed later that day, but the installed launcher does not carry it.
The audition table calls 20.5 seconds a first-word time. Its client was non-streaming and measured the whole HTTP request. This page calls it request wall time. There is no measured time to first token for that answer.
What it is
GLM-5.3-Flash is a Z.ai mixture-of-experts model. The tested GGUF is Unsloth UD-Q3_K_XL in four shards. The header names architecture glm5next, 46 blocks, 288 experts with 8 selected, and a native context length of 1,048,576 tokens. Those are file metadata, not measurements of useful context.
The four recorded shard sizes sum to 147,535,921,955 bytes, or 137.404 GiB by division by 230 (arithmetic). The first shard is metadata only. Its header names Unsloth as the quantizer and MIT as the license.
What to set
At 131,072 tokens: start with -b 4096 -ub 4096. That profile read 48,168 tokens at 425.1 tokens per second and peaked at 29,872 MiB. -ub 8192 failed to load.
At 262,144 tokens: use -b 4096 -ub 1024 for the measured headroom choice. It read 47,992 tokens at 142.6 tokens per second and peaked at 29,261 MiB. Only -ub 2048 was read beyond 47,992 tokens: it reached 230,039 tokens at 211.9 tokens per second, returned 3 of 3 codes, and peaked at 31,951 MiB. -ub 4096 failed to load at this window when a 13,281.37 MiB compute allocation failed.
The batch-size study explains the risk around the -ub 2048 shape and upstream llama.cpp issue 28282. The full-window study carries the cross-model comparison. This card does not generalize one setting across windows.
llama-server --ctx-size 262144 --parallel 1 \ --batch-size 4096 --ubatch-size 1024 \ --n-gpu-layers 999 --n-cpu-moe 42 \ --threads 24 --threads-batch 24 --flash-attn auto
The command shows the measured serving shape without machine paths, bind addresses or operator-specific configuration.
What we ran, exactly
| Model file | GLM-5.3-Flash, Unsloth UD-Q3_K_XL, four GGUF shards, 147,535,921,955 bytes total. The header reports file type 12. |
|---|---|
| Machine | Intel Core Ultra 9 285K, 188 GiB RAM, one NVIDIA GeForce RTX 5090 with 32,607 MiB reported card capacity. |
| Runtime | A GLM-specific, non-mainline llama.cpp fork pinned at d94f44e79aa219d8057e8de21f95360a187ebf41. |
| 16 September placement | --n-gpu-layers 999 -cmoe, with no batch or micro-batch flags passed, so llama.cpp used its defaults. The 75.48 tokens per second read and 9.58 tokens per second answer are from this launch. |
| 21 and 26 September placement | --n-gpu-layers 999 --n-cpu-moe 42, memory mapped weights, one slot and one request at a time. Batch settings are scoped to each measured row below. |
| Windows | 131,072 tokens on 16 and 21 September; 131,072 and 262,144 on 26 September. |
| Audition request | Temperature 0, low reasoning effort, 400-token cap. The answer stopped normally after 181 generated tokens. |
| Long-window requests | Code reads used temperature 0, seed 1 and a 1,024-token cap. Letters used temperature 0.7, seed 7 and a 1,200-token cap. The server was launched with reasoning effort none; the returned records still contain reasoning. |
| Repetitions | One run per serving configuration. No confidence intervals. |
Observed
| Run and request | Prompt evaluated | Prompt processing | Generated | Decode | Visible words | What came back |
|---|---|---|---|---|---|---|
| 16 Sep, 131,072, short answer | 29 of 38 tokens, 9 cached | 18.89 t/s | 181 | 9.58 t/s | not counted | A finished one-paragraph explanation. Whole-request wall time was 20.5 s. |
| 16 Sep, 131,072, deep read | 120,979 tokens | 75.48 t/s | 58 | 8.71 t/s | 3 codes | A finished answer with all 3 planted codes. |
26 Sep, 131,072, -ub 4096, code at 48,000 depth | 47,992 tokens | 404.9 t/s | 98 | 9.84 t/s | 1 | The planted code. |
26 Sep, 131,072, -ub 4096, letter on the cached 20,000-token document | 4,148 tokens | 312.1 t/s | 1,200 | 10.09 t/s | 0 | No visible words. |
26 Sep, 131,072, -ub 4096, letter on the cached 48,000-token document | 4,148 tokens | 294.8 t/s | 1,200 | 9.85 t/s | 0 | No visible words. |
26 Sep, 131,072, -ub 4096, letter on the cached 100,000-token document | 4,148 tokens | 260.9 t/s | 1,200 | 9.51 t/s | 0 | No visible words. |
26 Sep, 262,144, -ub 1024, code at 48,000 depth | 47,992 tokens | 142.6 t/s | 125 | 10.02 t/s | 1 | The planted code. |
26 Sep, 262,144, -ub 1024, letter on the cached 20,000-token document | 1,080 tokens | 97.0 t/s | 1,200 | 10.22 t/s | 0 | No visible words. |
26 Sep, 262,144, -ub 1024, letter on the cached 48,000-token document | 1,080 tokens | 95.3 t/s | 729 | 9.98 t/s | 351 | A finished letter. |
26 Sep, 262,144, -ub 2048, code at 230,000 depth | 230,039 tokens | 211.9 t/s | 231 | 8.88 t/s | 3 | All 3 codes. Prompt processing took 1,085.744 s; whole-request wall time was 1,112.9 s. |
26 Sep, 262,144, -ub 2048, letter on the cached 230,000-token document | 2,088 tokens | 131.6 t/s | 1,200 | 9.04 t/s | 0 | No visible words. |
Prompt processing is the server's prompt-eval timing. Each letter reused the cached document from the paired read at that target depth. The paired reads evaluated 20,065, 47,992, 100,003, or 230,039 tokens. On a letter row, Prompt evaluated is only the uncached tail, and the prompt-processing rate is the rate for that tail. Decode includes reasoning tokens. Whole-request wall time includes prompt processing, generation and client overhead.
This table is not every 26 September request. On -ub 2048, the letters at target depths 20,000, 48,000, and 100,000 returned 102, 119, and 337 visible words; only the 230,000-depth letter returned none. The 102-word and 119-word replies reached the 1,200-token cap; the 337-word reply finished. Across the displayed 26 September rows, decode stayed near 10 tokens per second, while reading moved sharply with window and micro-batch. A short visible answer can hide far more generation. At 230,039 prompt tokens, the visible answer was three codes but the server generated 231 tokens, including 617 reasoning characters. A separate 40-prompt check on the shipped 262,144, -ub 1024 setting recorded 36 empty answers and 0 crashes.
The reading-side lever
The 21 September sweep used a 131,072-token window and one request at a time. Every successful read returned the planted code in the visible answer. The default-setting and tuned rows used different prompt lengths, so their rate ratio is descriptive, not a same-prompt A/B.
| Batch and micro-batch | Request | Prompt tokens | Read t/s | Generated | Decode t/s | Peak MiB |
|---|---|---|---|---|---|---|
| llama.cpp defaults, 2048 / 512 | first | 24,008 | 80.7 | 62 | 9.22 | 24,352 |
| 4096 / 2048 | recorded long read | 48,168 | 252.0 | 9.21 | 26,763 | |
| 4096 / 4096 | third | 20,333 | 446.0 | 60 | 9.35 | 29,872 |
| 4096 / 4096 | fourth | 48,168 | 425.1 | 310 | 9.12 | 29,872 |
| 8192 / 8192 | Did not load |
The published comparison, 446.0 divided by 80.7, differs by 5.5 times after rounding to one decimal place. The prompt lengths differ, so this is descriptive rather than a same-prompt comparison. The longer 48,168-token result at 425.1 tokens per second was the fourth request on the tuned load. Its 310 generated tokens decoded at 9.12 tokens per second; the default first request generated 62 tokens at 9.22 tokens per second.
The visible-output boundary
Measured failure, 17 September: one synthetic high-effort streaming request generated a correct three-sentence answer, then more reasoning, a literal </think>, and a second final answer in the visible field. It finished normally after 689 generated tokens. The raw generation contained a second assistant control token before the restarted reasoning. Independent parsing of the saved wire bytes found the closing marker in server delta.content. That rules out the test client as the source of the extra visible text. It does not prove whether the checkpoint, quantization or runtime caused the generation.
Measured isolated guard: the same synthetic smoke, a tool call and its tool result all passed when the server used --special and each request stopped on four strings:
<|assistant|> <|endoftext|> <|user|> <|observation|>
The smoke returned one clean 254-character answer; its stop matched <|assistant|>. The tool cycle finished with the expected call and a visible result. This was one isolated fixture. Its own comparison record says the exact profile passed and was not production-qualified.
The installed launcher is different. On 26 September it set neither --special nor request stop strings. It set reasoning_effort, with high as the code fallback, plus temperature 1.0 and top-p 0.95. It selected -b 4096 -ub 4096 at windows up to 131,072 and -b 4096 -ub 1024 above that. The launcher accepts 131,072 and 262,144, but its code fallback is 131,072 even though its comment calls 262,144 the new default. It sources an optional operator configuration first, so the deployed value can differ. The record does not establish whether the installed profile still reproduces the 17 September failure.
One coding snapshot
The local coding trial gives this speed a little context. Across four single-run tasks, GLM-5.3-Flash scored 387 of 400 in 6,058 seconds of task wall time. Each task received 60 of 60 mechanical points; the four blind judgment scores were 34, 39, 37 and 37 of 40. A separate one-question judgment probe scored 39 of 40 in 423 seconds from question to answer.
Those are measured scores from that exercise, not a general coding rank. The coding page owns the tasks, rubric, other models and limits. This card does not duplicate them.
What we got wrong
Our audition record labelled a non-streaming request's 20.5-second wall time as first_word_s. The client started a clock, waited for the complete JSON response, and stopped it after parsing. It never observed the first token. The 9.58 tokens-per-second decode rate remains a server timing for the finished 181-token answer. Only the latency label changes.
We also avoid calling the isolated output guard a deployed fix. The exact guarded profile passed its three-part test, but the installed launcher does not set that profile.
What this does not show
- No full native window. The header states 1,048,576 tokens. The deepest measured prompt here is 230,039 tokens on the separate
-ub 2048run, not on the installed-ub 1024setting. - Recall is not reasoning. Returning three planted codes once does not establish reliable reasoning across the window.
- No stable quality ranking. The coding scores are one run per task, and the judgment probe is one question.
- No general output verdict. The failure and guarded pass each use one synthetic fixture on one quantization and one pinned fork. Separately, a 40-prompt check on the shipped 262,144,
-ub 1024setting recorded 36 empty answers and 0 crashes. That is evidence of frequent empty output in that check, not a quality verdict. - No measured RAM floor. The launcher carries a 150 GB availability check. The package records that figure as inferred from another deployment, not measured for this model. This page makes no claim that 150 GB is necessary or sufficient for this model.
- No concurrency result. Every serving measurement was one request at a time on one desktop.
Sources and artifacts
The data package contains redacted extracts from the audition, the batch sweep, the 26 September long-window rows, the shipped-profile stress check, the output diagnostic and combined guard, the installed launch profile, the GGUF header, and the coding totals. The number map identifies every printed figure and derivation.
Machine paths, bind addresses, request identifiers, fingerprints, service names and internal harness names are removed or replaced. The raw private coding prompts and hosted grader text are not included. The public coding-trial package holds that exercise's releasable records.
If a number on this page disagrees with a file in its package, the file is right and the page is wrong. Tell us at hello@graphometer.ai and the correction goes on the page with its date.