Seven minutes to 1.7 seconds
Saving a local llama.cpp server's conversation state to disk. On 13 September, after a restart, Qwen3-235B reached first content in 1.65 seconds on a history that took 461.97 seconds to read cold. Cold reading has since got faster; these restart times were not re-run.
Measured, 13 September 2026. In one restart cycle, Qwen3-235B reached first content in 1.65 seconds on a 102,912-token history that had taken 461.97 seconds to read cold. That 1.65 seconds is the first request once the server was back; loading the model is a separate cost, and this page prints it. GLM-4.7-Flash went from 74.73 seconds to 0.67 seconds on 105,181 tokens, with first content 15.3 seconds after the stop command. After the DeepSeek V4 Flash restart no restore was performed, and the first request re-read the whole history. These original results are three configurations, one request at a time, on one machine, on one date.
-b 4096 -ub 2048, against 285.8 at the earlier default. One run at each setting. The same probe at -b 4096 -ub 4096 read at 944.2 tokens a second and was not adopted. The 703.3 result is the launch default, not the fastest rung. The cold wait is smaller at that measured depth; no new roughly 100K cold time was found, so the headline remains the dated 13 September result. A later MiniMax M2.7 test restored its slot and completed the next request in 2.17 seconds, after a separate 2.19-second restore. That is total request time, not time to first content.
What it is
A long conversation is expensive to read again after a local model server restarts. The server has lost the attention state it built while reading earlier turns, so the next request can begin with minutes of waiting.
llama-server can write a slot's state to a file and read it back, and a small program of our own decides when to save it and asks for the restore after a restart. The test recorded a cold turn, a warm turn, a server stop and start, then another turn with the same history. Time to first content is the measure in the original three rows. The later MiniMax row times the entire completion and says so. On the GLM runs the first content is a reasoning event, not the answer.
What we ran, exactly
| Machine | Measured. One NVIDIA GeForce RTX 5090, an Intel Core Ultra 9 285K, and 188 GiB of RAM. |
|---|---|
| GLM-4.7-Flash | Measured. A 131,072-token served window, one slot, and 105,181 prompt tokens in the cold leg. From our run records: Z.ai GLM-4.7-Flash, UD-Q4_K_XL. |
| Qwen3-235B | Measured. A 131,072-token served window, one slot, and 102,912 prompt tokens in the cold leg. From our run records: Alibaba Group Qwen3-235B Instruct 2507, Q4_K_M with a 0.6B draft model (a small model that proposes tokens for the large one to verify). |
| DeepSeek V4 Flash | Measured. A 131,072-token served window, one slot, and 102,480 prompt tokens in the cold leg. From our run records: DeepSeek V4 Flash 0731, Q8 body. |
| What each 13 September run did | Measured. One seeded history with a 96,000-token target, one tool call and a three-code recall check per run, through an agent framework. The framework is not named. |
| MiniMax M2.7, 19 September 2026 campaign | Measured configuration. Unsloth UD-IQ4_XS, a 131,072-token window, one slot, --n-cpu-moe 59 -b 4096 -ub 4096 --no-repack --flash-attn on, q8_0 key and value caches, and 24 threads. The probe sent the same synthetic text before and after a real server stop and start, directly to llama-server. It used the server's own slot-save and slot-restore endpoints. |
| Runtime revision | The records do not identify a llama.cpp build revision, so this page does not assign one. |
| Sampling | The 13 September records do not state sampling settings. The MiniMax probe used temperature 0, non-streaming completion, and a 16-token output limit. |
The original three model-file descriptions come from our run records, not from the result files. MiniMax's file and flags are recorded in the shipped launch excerpt and server log. For the original rows, timings, prompt sizes, window, and slot count come from the redacted JSON records. MiniMax's measurements come from its console and server logs.
Observed
Each row is one restart cycle. These are not medians.
| Configuration | History | Cold request | After restart | Label |
|---|---|---|---|---|
| GLM-4.7-Flash, UD-Q4_K_XL | 105,181 tokens | 74.73 s | 0.67 s | measured first content, 2026-09-13 |
| Qwen3-235B, Q4_K_M | 102,912 tokens | 461.97 s | 1.65 s | measured first content, 2026-09-13 |
| DeepSeek V4 Flash, Q8 body | 102,480 tokens | 1,396.6 s | 797.57 s, history re-read | measured first content; failure, 2026-09-13 |
| MiniMax M2.7, UD-IQ4_XS | 85,763 tokens | 147.49 s, full request | 2.17 s, full request after restore | measured 16-token completion; 19 September campaign, console has no clock time |
For the 13 September rows, time is from the request to first content. MiniMax is a non-streaming request through completion of 16 generated tokens; it is not a first-content measurement or an end-to-end restart time. On the GLM runs the first content is a reasoning event, not the answer. The DeepSeek row is a failure, not a partial success: its post-restart request re-read the whole history, 102,846 tokens. The next request in the same record reached first content in 2.86 seconds with no tokens re-read.
Measured, 13 September 2026. The GLM-4.7-Flash cycle in that table timed the whole stop and start. First content arrived 15.3 seconds after the stop command, model load included. Reading the saved state back and replaying the last prompt took 1.03 seconds, and the request produced first content in 0.67 seconds. The saved state was 5,722,437,888 bytes, about 5.7 GB.
Measured, 13 September 2026. End to end, the Qwen3-235B restart took 230.5 seconds from the stop command, 222.1 of them loading the model; the 1.65 seconds is the first request once the server was back.
Measured timings, 19 September 2026 campaign; the console has no clock time. MiniMax M2.7 saved and restored 85,778 tokens. The saved state was 11,573,855,516 bytes; saving took 3.05 seconds client wall time. The repeated prompt then required 1 token of new prompt work. Its client-side restore call took 2.19 seconds (the server reported 2,171.784 ms); the following full request took 2.17 seconds. The saved count includes generated tokens, so it is not the prompt length. These records show a working restore through llama-server's own slot endpoints, tested directly; they do not time first content or the complete stop-to-answer interval.
-b 4096 -ub 2048, in a 131,072-token window, one request at a time. This is prompt-reading throughput on one synthetic prose prompt that asks for a planted reference, not the 13 September conversation and not source code. One run at each setting. The same probe at -b 4096 -ub 4096 read at 944.2 tokens a second and was not adopted. The 703.3 result is the launch default, not the fastest rung. No new roughly 100K cold duration is substituted for 461.97 seconds. On the GLM batch trial's 202,752-token window, larger micro-batches read a 47,986-token prompt at 3,022.0 to 3,100.2 tokens a second against 2,594.6 at the default. The ubatch 2048 run returned an empty answer. The launch script still uses the default. None of this replaces the 74.73-second cold result.
An earlier GLM-4.7-Flash cycle that morning ran the same sequence and reached 0.68 seconds after the restart from a 74.79-second cold read, with no prompt tokens re-read and the same 1.03-second restore and replay. Its record reports a saved state of 1,687,552 bytes, about 1.7 MB, against 5.7 GB in the cycle above. The records do not explain the difference, so the page reports the later cycle and ships both files.
How the warm wake works
llama-server can write a slot's state to a file and read it back, using its slot-save flag and its slot endpoints. A small program of our own sits between the client and the server: when a turn goes quiet it asks the server to save the slot; after a restart it asks for the restore, replays the last prompt, and keeps the tool list in a stable order. The program is ours and is not yet published. The saving and restoring is llama.cpp's.
- The server reads a long prompt and builds the attention state used by later turns.
- After the conversation goes quiet, our program asks the server to write that slot state to a file.
- After a server restart, the saved slot is read back and the last prompt is sent again, so the server's state matches the conversation.
- The next request can reuse the restored prefix when the model architecture and the server allow it.
Measured, 13 September 2026. The Qwen3-235B record reports a 10,553,464,396-byte slot file and a 4.95-second restore and replay. Its next request re-read 0 prompt tokens and reached first content in 1.65 seconds.
Measured, 13 September 2026. Saving is not free. In one record the turn that triggered a save waited 20.78 seconds for first content, with no prompt tokens re-read; that turn is in glm-4.7-flash-warm-wake.json. This is why the save waits for a quiet turn instead of running after every request.
This saves prompt state, not model weights. Loading the model still takes time. The 15.3-second GLM cycle includes a 12-second model load, and the Qwen3-235B server needed 222.1 seconds to come back.
The client can throw the cache away
The prompt prefix has to stay byte-stable. On these runs the client reshuffled its tool list between requests. The tools sit near the start of the prompt, so changing their order changed the prefix and discarded the reusable cache.
The fix was simple: sort the tool list on every request. Stated: our diagnosis, from the fix working. No separate record ships for it. The wider lesson is not specific to one client. If a local model server depends on prefix reuse, keep stable instructions and tool definitions in a stable order.
Where it does not work
After the DeepSeek V4 Flash restart no restore was performed: the record's restore field is empty, and the records do not say why. The first request re-read the whole history, 102,846 tokens, and reached first content after 797.57 seconds. Measured failure, 13 September 2026. The next request in the same record reached first content in 2.86 seconds with no tokens re-read, so on that run the re-read was paid once, not on every turn.
Our operator record from the next morning says the restored state was not reused on the sliding-window and recurrent designs we checked: DeepSeek V4 Flash, Qwen3.5-122B, Qwen3.5-397B, Qwen3.8-Flash-Next, Gemma 4 26B and Gemma 4 31B. For those the program replays the history at server start instead. Stated: our operator record, 2026-09-13. No result file ships for them. On the model we measured, the replay happened on the request itself, which is the 797.57 seconds above.
Check the model file before testing reuse
Inspect the GGUF metadata for a key ending in attention.sliding_window. Our Qwen3-235B Q4_K_M and MiniMax M2.7 UD-IQ4_XS files have no such key; both reused saved state in the tests above. DeepSeek V4 Flash's Q8 header has deepseek4.attention.sliding_window = 128. Measured headers, read 26 September 2026. The absence of the key is a useful sign for these files, not a guarantee for every architecture or build. A sliding-window key warns that restored state may need a full re-read. This page's DeepSeek cycle has restart.restore = null, so it does not measure reuse after a successful restore.
Run python3 read_attention.py /path/to/model.gguf with the metadata-only reader supplied here, using the first shard for a split file. It scans the metadata without loading tensors or starting a model. The captured header results show the keys found. Then test the same prompt after save, restart and restore: the next request's prompt-eval count decides whether state was reused; a large restored-token count alone does not.
This is an interaction between the model's attention design and what a restored slot can be reused for. It is not a claim that any vendor caused the failure.
What we got wrong
Our first notes grouped DeepSeek V4 Flash with the models expected to reuse a restored slot. The next measured run showed otherwise. No restore was performed, the request re-read the history, and the result was marked FAIL.
That correction matters because a saved file on disk is not by itself a warm next turn. The user-visible request is the check that decides whether the method worked.
What this does not show
- One restart cycle per configuration, two for GLM-4.7-Flash. None of it is a distribution: there are no medians or confidence intervals, and the two GLM cycles are not merged.
- The study does not measure answer quality. The original rows measure first content; MiniMax measures the complete short request. Both check whether the history was re-read.
- The original records omit build and sampling settings. The new MiniMax probe records sampling settings but no build revision.
- The successful restores do not prove the method works for every full-attention model, runtime revision, prompt template, or storage device.
- The failure does not prove that every sliding-window or recurrent model must fail on every later build.
- The client-side prefix finding is one diagnosed cache loss on these runs. It is advice to test byte stability, not a benchmark of other clients.
Related pages on this site
The Method page defines Graphometer's evidence labels and the rule that files outrank prose. The wrong flag is another measured llama.cpp study on the same workstation. The model field cards place these runs beside the exact files and configurations used elsewhere on this bench.
Sources and artifacts
The evidence package keeps the four original result files: three configurations, two cycles for GLM-4.7-Flash. It adds MiniMax's console and server records and redacted probe and launch excerpts, plus the Qwen and GLM batch records and measurement instrument. Operational names, addresses, ports, private tool names, and internal paths were removed. Measurement fields were retained.
- MiniMax restore and request timings, server log and probe excerpt: the 2.19-second restore and subsequent 2.17-second complete request are different intervals. 96,000 is the seed target. The server counted 85,763 prompt tokens. 147.49 seconds is the client wall for the whole completion. 589.5 tokens a second is server prompt eval over 145.48 seconds, not 85,763 divided by 147.49. The server logged 145,476.38 ms; 145.48 seconds is that value divided by 1,000 and rounded.
- Qwen earlier default and larger batch: the 48,020-token prose reading update, plus the unadopted larger rung.
- data/README.md: package contents, provenance, redactions, and the disagreements readers should know first.
- data/NUMBERS.md: every figure on this page mapped to a file and field.
- GLM-4.7-Flash full restart result: the 74.73 to 0.67 second cycle, 15.3 seconds from the stop command to first content, model load included.
- Qwen3-235B result: the 461.97 to 1.65 second cycle, and 230.5 seconds from the stop command.
- DeepSeek V4 Flash result: the measured failure.
- GLM-4.7-Flash first cycle: the 74.79 to 0.68 second cycle, with the 1.7 MB saved state the records do not explain.
If a number on the page disagrees with a file in its package, the file is right and the page is wrong.
The saving and restoring are llama.cpp's own: llama-server's slot-save flag and its slot endpoints. The program that decides when to save, asks for the restore and replays the last prompt is ours, and is not yet published. No upstream source revision is cited because the records do not identify one.