DeepSeek V4 Flash 0731, on one desktop and on two
With its batch settings tuned, one desktop answered a 48,024-token request in 75.2 seconds. The same file split across two machines took 416.4 seconds at its best completed setting and 474.2 at the one it keeps. Measured 21 September 2026.
The practical choice for long prompts is now the desktop alone, with experts split between system RAM and the card. The IQ3 desktop at -b 4096 -ub 4096 read the 48,024-token prompt at 690.9 tokens a second; whole-request time was 75.2 seconds, against 474.2 seconds on the IQ3 two-machine split at retained -b 2048 -ub 512 measured, 21 September. Reading means prompt tokens processed per second. Speaking means generated tokens per second, including hidden reasoning on this model. These are tested configurations on one RTX 5090 desktop, alone or paired with an ASUS ROG Flow Z13, not an isolated test of machine count.
Our earlier recommendation led with the two-machine 4-bit body and its draft model. Tuning the amount of prompt text processed in each batch changed the whole-turn result. The desktop now finishes the measured long-prompt requests sooner. The September 12 to 15 measurements remain below as history; the 4-bit split with a draft was not re-tuned in the new sweep.
What it is
DeepSeek V4 Flash 0731 is a mixture-of-experts language model. The maker's model card, read on Hugging Face at repository revision 7872f01b, says the release ships with a speculative decoding module attached vendor. Our 3-bit GGUF header reports 284.33 billion model parameters, 256 experts, 6 experts used per token, and 1 shared expert measured. The maker's model card at that revision states no active-parameter figure, so this page prints none, and the header does not carry the expert width an active total would need.
| Maker and release | DeepSeek; 31 July 2026 vendor. The date is the creation date of the maker's repository, read at revision 7872f01b. |
|---|---|
| License record | MIT. The LICENSE file was read in full at repository revision 7872f01b and ships in the data package vendor. The community GGUF header also records mit measured. |
| Base shape | 43 blocks, 284.33 billion parameters, 256 experts, 6 used per token plus 1 shared measured. |
| Served window | 131,072 tokens in the earlier September 12 to 15 V4 runs and the September 20 Q8 runs at 48,073 prompt tokens; 262,144 in the September 15 full-window sweep, the September 20 Q8 run at 150,324 prompt tokens and the September 21 IQ3 runs, one request at a time measured. The header records a trained 1,048,576-token window. This card uses 128K and 256K for the two served windows, not prompt occupancy. |
| Runtime | llama.cpp build 10919 at source commit d3146f2b5 for the two-machine sweep measured. |
What we ran, exactly
The bench is one NVIDIA GeForce RTX 5090 with 32,607 MiB of VRAM, an Intel Core Ultra 9 285K and 188 GiB of RAM. The second machine is an ASUS ROG Flow Z13 with 128 GB of unified memory, joined directly by Thunderbolt. One model process ran at a time. The two-machine runs used one host process on the desktop and a worker on the laptop, built from the same source commit measured.
| 8-bit, desktop alone | Q8 body, 131,072-token window, one slot. Loaded in 213.1 seconds and used 16,860 MiB on the card on 12 September 2026 measured. The speed record comes from the tool-and-recall check. |
|---|---|
| 3-bit, two machines | Unsloth UD-IQ3_XXS, 4 shards totaling 104.2 GB. Placement: 8 whole layers on the card, 10 layers with attention on the card and experts in desktop RAM, then 25 whole layers on the laptop. Healthy in 163 seconds and 24,907 MiB on the card on 13 September 2026 measured. |
| 4-bit, no draft | Unsloth UD-Q4_K_XL, 5 shards totaling 155.1 GB. Placement: 5 whole layers on the card, 14 expert sets in desktop RAM, and 24 whole layers on the laptop. Served at 131,072 tokens on 13 September 2026 measured. |
| 4-bit, installed preset | The same 4-bit body plus DeepSeek's draft model. Placement: 2 whole layers on the card, 13 expert sets in desktop RAM, and 28 whole layers on the laptop. Healthy in 225 seconds, 27,430 MiB on the card, and 91 GB used on the laptop on 15 September 2026 measured. |
On the two-machine runs the common launch shape was a 131,072-token window, one parallel slot, flash attention selected automatically, 24 processing threads, Jinja chat formatting, DeepSeek reasoning formatting, and one request at a time. The server was launched with temperature 1.0, top-p 1.0, top-k 0 and min-p 0.05. The direct speed requests explicitly set temperature 0.7 measured. The 8-bit desktop-alone run came from a different launcher on 12 September 2026: its record gives the served window and the registered window, both 131,072 tokens, and no launch record for it is in this evidence, so nothing above is claimed for it beyond the window. The draft sweep asked for at most 8 draft tokens; the installed 4-bit preset reported at most 6, with minimum 0 and probability floor 0.50 measured.
The practical result now
On the desktop, use the UD-IQ3_XXS body with --n-cpu-moe 36 -b 4096 -ub 4096 --no-repack for the measured fast configuration. The larger UD-Q8_K_XL body used -cmoe -b 8192 -ub 8192 --no-repack: all experts in system RAM. The names “fast” and “quality” identify those presets; they do not report an answer-quality comparison.
These September 20 and 21 probes ran one request at a time, temperature zero, with a synthetic ledger and a sealed code at its midpoint. The code requests allowed 900 generated tokens, including thinking. Both desktop presets used 24 processing and batch threads, automatic flash attention, Jinja formatting and DeepSeek reasoning formatting measured configuration. The new logs do not print a build commit, so the older build identity below is not assigned to them.
| Configuration and date | Window | Prompt tokens | Reading, t/s | Speaking, t/s | Label |
|---|---|---|---|---|---|
| Q8 desktop, recorded extra arguments -cmoe only, 20 September | 131,072 | 48,073 | 83.3 | 9.91 | measured |
| Q8 desktop, b / ub 4096, 20 September | 131,072 | 48,073 | 424.4 | 9.71 | measured |
| Q8 desktop, b / ub 8192, 20 September | 131,072 | 48,073 | 607.1 | 9.58 | measured |
| Q8 desktop, b / ub 8192, 20 September | 262,144 | 150,324 | 480.0 | 9.37 | measured |
| IQ3 desktop, b / ub 4096, 21 September | 262,144 | 2,998 | 539.6 | 12.65 | measured |
| IQ3 desktop, b / ub 4096, 21 September | 262,144 | 48,024 | 690.9 | 12.16 | measured |
| IQ3 desktop, b / ub 4096, 21 September | 262,144 | 150,103 | 632.1 | 11.79 | measured |
The visible code answer was 17 characters. Q8 reasoning text was 216 to 332 characters; timed generation was 70 to 106 tokens on Q8 and 59 to 86 tokens on desktop IQ3. Every speaking rate counts hidden reasoning, not just the visible answer measured. The slow Q8 row records only -cmoe as extra arguments; its progress interval supports logical batch 2,048, but the micro-batch is not recorded. The Q8 date is the 20 September session date, not a calendar timestamp inside its server logs.
The Q8 batch change accelerated reading without a speaking gain. At the 131,072 window, card use at load rose from 17,063 to 22,089 MiB: 5,026 MiB more arithmetic difference. At the 262,144 window the Q8 run loaded at 26,023 MiB. The IQ3 desktop peaked at 28,930 MiB on the deep read measured. Each displayed code probe returned the exact code. The IQ3 preset also completed 40 of 40 mixed public-text stress prompts without a crash on 21 September measured; that is a crash check, not an answer-quality score. Its longest prompt was 9,135 tokens, not the 48,024 or 150,103 tokens of the code requests measured.
Keep the batch settings specific to the body. The measurements above validate -ub 4096 on IQ3 and -ub 8192 on Q8. Keep --no-repack as used in these runs; removing it was not part of this successful comparison.
Two other Q8 settings did not earn a change
The first sweep combined --n-cpu-moe 50 with --cache-type-k q8_0 --cache-type-v q8_0. On a 16,011-token prompt it read 82.41 versus 81.64 t/s at default batch, and 255.12 versus 260.82 t/s with -b 4096 -ub 2048 measured, 20 September. This was a combined configuration test, not an isolated gain from either flag; the model has only 43 blocks, so a CPU-expert setting of 50 is not evidence that experts moved onto the card. The probe allowed only 8 generated tokens and did not verify an answer. We kept the Q8 cache and placement unchanged.
Upstream reports cover quantized K-cache corruption and garbled output with CUDA experts. Read on 26 September 2026, the first is closed; the follow-up cache report and the expert report are closed as not planned. Those statuses differ from the tuning notes' “open” description. Closure does not establish that these local builds were fixed. The IQ3 results above are their own measured checks.
The two-machine split, re-tuned
The September 21 split used IQ3 at a 262,144-token window, no draft, with 8 whole layers on the card, 10 expert sets in desktop RAM and 25 whole layers on the laptop. At -b 2048 -ub 512 it read 161.8 t/s on 2,998 tokens and 102.5 t/s on 48,024; timed generation on those short code answers (67 and 69 tokens, hidden reasoning included) ran at 14.76 and 12.04 t/s. Raising -ub to 2048 gave 224.9 and 117.0 t/s reading measured, a 14.1% gain at the longer depth arithmetic.
At -b 4096 -ub 4096, the 48K request failed after the server logged 40,960 tokens processed and then “Remote RPC server crashed” measured failure. The same 48,024-token request had already completed with the exact code in 416.4 seconds at -b 2048 -ub 2048 measured whole-request time. We kept -ub 512 anyway as a precaution after the laptop GPU hang at 4096, not because 2048 failed. The desktop's 75.2-second whole request remains shorter than the split's best completed 416.4-second result at this depth. At 150,103 tokens it took 2,922.9 seconds to read and timed generation, including hidden reasoning, ran at 7.66 t/s measured. That read alone is 48.7 minutes arithmetic. The older 15.9 to 17.1 t/s split band below came from short probes at a 131,072-token window, not this larger-window run.
The 15 September full-window sweep
At a 262,144-token window, the desktop UD-Q8_K_XL body read 150,475 tokens in 2,062,226.61 ms at 72.97 t/s with no explicit batch flags on 15 September measured, or 2,062.2 seconds rounded arithmetic, milliseconds divided by 1,000. The 20 September session's Q8 run at the same window read 150,324 tokens in 313,172.17 ms at 480.00 t/s with -b 8192 -ub 8192 measured, or 313.2 seconds rounded arithmetic. Both launch records name the same binary path. The older response identifies build b1-5f55650; the later record omits a build identity, so an unchanged binary cannot be established. The prompt lengths and probes also differ; this is not an isolated one-variable comparison.
The 15 September UD-IQ3_XXS split at a 262,144-token window read 230,835 tokens in 6,580,695.01 ms at 35.08 t/s measured, or 109.7 minutes rounded arithmetic, milliseconds divided by 60,000. It returned a 25-token answer at 6.49 t/s, found 3 of 3 codes and peaked at 25,824 MiB on the card measured. This is generated-token throughput, including any hidden reasoning, not a prose rate. At a 131,072-token window, the same split read 120,305 tokens at 58.70 t/s measured. Both split records identify build 10919 / d3146f2b5, one slot, no explicit batch flags, 24 processing and batch threads, and the 8 / 10 / 25 placement described above.
What the next turn costs
A real follow-up appends a reply and a new message to the cached history. At the split's 48K depth it processed just 238 tokens in 5.4 seconds at ub 512, or 228 in 5.0 seconds at ub 2048. Swapping the last question in the existing prompt instead went back roughly one micro-batch: 533 tokens in 10.0 seconds versus 2,065 in 28.2 seconds measured, 21 September. An edited or regenerated last turn can pay that replay cost; it is not the cost of every ordinary follow-up.
Observed, 12 to 15 September
Speaking is generated-token throughput. Reading is prompt throughput. This model thinks by default, so speaking counts every generated token, the hidden reasoning included: the short probes behind 19.55, 17.06 and 10.59 tokens a second each spent their whole 200-token budget thinking and returned no visible reply. Each row below is a recorded run, not an average across configurations.
| Body and placement | Speaking, t/s | Reading, t/s | Prompt, tokens | Label |
|---|---|---|---|---|
| 8-bit, desktop alone, 12 September | 9.2 to 9.7 | 73.8 to 75.5 | 48K and 96K legs | measured |
| 3-bit, two machines, no draft, 13 September | 15.9 to 17.1 | 184.6, 187.9 fresh | 21, 7,841 and 8,210 | measured |
| 4-bit, two machines, no draft, 13 September | 10.6 to 13.0 | 145.4, 173.0 fresh | 21, 7,841 and 8,210 | measured |
| 4-bit, two machines, installed draft preset, 15 September | 19.55, 19.59 | 125.1 | 21 and 7,841 | measured |
| 3-bit, two machines, at 48,000 tokens of context, 14 September | 11.7 | 97.4 | 52,928 | measured |
| 3-bit, two machines, at 96,000 tokens of context, 14 September | 10.0 | 66.6 | 102,478 | measured |
A fresh reading is a newly written prompt the server has never seen, so nothing is answered from cache; a planted code is a short code word hidden at a known depth in that prompt, which the model is then asked to repeat. The last two rows were taken through an agent framework rather than by a direct probe, and they show what a practitioner should expect at depth: on this configuration both speaking and reading fall as the history grows. The 8-bit-alone and 3-bit-split rows use different files, so their difference mixes quantization with placement. The 4-bit rows share a body but differ in placement and draft configuration. The fresh 4-bit no-draft probe missed its planted code; its speed remains a real timing, but it is not a clean recall result.
On the short 8-bit turn on 12 September 2026, the first visible word arrived after 60.5 seconds and the reply then ran at 15.0 tokens a second measured. That is faster than the 9.2 to 9.7 band in the table because it is a short turn on a small prompt, while the band was measured at 48,000 and 96,000 tokens of context, where speaking slows. The 60.5-second delay includes the time it took to read the agent framework's own system prompt. It is not a measurement of how long the model thought.
The earlier draft-model comparison
A draft model is a small companion model that guesses the next few tokens. The large model then checks the guesses in one pass and keeps the ones it agrees with. Accepted guesses cost almost nothing extra, so the share of guesses accepted is what turns into speed.
The installed 4-bit preset loaded DeepSeek's own draft model natively. Its runtime record reports a maximum draft length of 6, a minimum of 0, and a probability floor of 0.50 measured. The remaining draft settings are recorded in the package README.
In the earlier sweep on 13 September 2026, where the launcher allowed up to 8 draft tokens measured, the draft's guesses were accepted at 61.4%, 62.8%, and 63.9% on the 3-bit file, and 68.0%, 84.6%, and 69.4% on the 4-bit file measured. Each set is the three probe turns of one run, in order: the short turn, the 8,000-token depth probe, and the fresh-prompt probe. Mean accepted runs ranged from 2.82 to 3.75 tokens measured.
At the 8,000-token depth probe on 13 September 2026, the best 4-bit configuration with the draft reached 19.25 tokens a second against 12.78 for the best without it, a difference of 51% arithmetic. On the 3-bit file the same comparison is 18.12 against 16.15, a difference of 12% arithmetic. The underlying speeds are measured; the percentages are arithmetic on them.
The placement changed as well as the draft, so these are comparisons of two configurations, not an isolated measurement of the draft. Without the draft the 4-bit file ran 5 whole layers on the card, 14 expert sets in desktop RAM and 24 whole layers on the laptop; with the draft it ran 2, 13 and 28. The 3-bit file ran 8, 10 and 25 without the draft and 4, 14 and 25 with it measured. In both cases whole layers moved off the card to make room for the draft model. Reading fell in both configurations too, from 184.6 to 148.8 tokens a second on the 3-bit file and from 145.4 to 121.9 on the 4-bit file measured, and that loss carries the same caution.
The 4-bit file stores more information per weight than the 3-bit file. We did not measure answer quality, so "better weights" here means a higher-bit file, not a quality verdict. The 15 September preset measured speaking near 19.6 tokens a second, including hidden reasoning, with the 4-bit body measured. It gave up reading speed to do that. The 26 September update above gives the current choice for the measured long-prompt code requests; the 4-bit draft preset was not re-tuned.
Context, recall, and tools
The 3-bit run's eight cache buffers sum to 892.25 MiB at a 131,072-token window measured. Doubling that layout, which is what a 256K window would be, gives about 1,785 MiB arithmetic, not run. That was an untested memory estimate on 15 September. The new runs above served 262,144 tokens; their cache layout is not inferred from this doubling.
An earlier summary of ours called the 128K attention cache about 1 GB. This page prints the exact sum of the eight recorded buffers instead of the rounded figure. The change is precision, not a correction of a wrong number.
Tool calls and recall at 48,000 and 96,000 tokens of context were checked through an agent framework on 14 September 2026. At the first depth a seeded history of about 48,000 tokens became a 52,928-token prompt once the framework's own scaffolding was counted; the model found 3 of 3 planted codes and wrote a file through a tool at depth measured. At the second depth the prompt was 102,478 tokens, again 3 of 3, again a tool used at depth measured. First visible content arrived after 543.65 seconds at the first depth and 1,537.75 seconds at the second measured; both times include the framework seeding the history and reading its own system prompt, so neither is model time alone. The same records give the speaking and reading rates at those depths, 11.7 and 97.4 tokens a second at 48,000 and 10.0 and 66.6 at 96,000 measured.
The long-context check took three attempts measured. The first failed its time budgets. The second holds the 48K and 96K legs, the recall and the tool use, and failed on a separate leg. The third reran the short legs and passed. Across attempts 2 and 3 every component passed; no single recorded attempt contains both.
The installed 4-bit preset passed only the short tool-and-recall legs, on 15 September 2026. Its direct tool-schema call took 5.8 seconds measured. The 48K and 96K legs were not rerun on that preset.
Thinking is stated, not verified here
Our operator record for both of our DeepSeek V4 Flash endpoints, the desktop-alone one and the two-machine one, says the model thinks by default and accepts a per-request switch to turn thinking off stated. The loaded chat template reported thinking enabled measured, and the 200-token short speed probes of 12 to 15 September spent their whole token budget on reasoning. The new code probes stopped early and returned a visible answer; their timed generation also counts hidden reasoning. The timed probes packaged here do not establish the off-switch on DeepSeek V4 Flash 0731. We therefore do not label the switch measured.
Cold wake, then and now
The old cold-wake times below remain measurements of their dated configurations. With the tuned desktop IQ3 body on 21 September, the server read 150,103 tokens in 237,469.00 ms and the complete code-answer request took 244.8 seconds, against 2,933.8 seconds on the split at ub 512 measured. The desktop read time is 237.469 seconds arithmetic, milliseconds divided by 1,000. At 48,024 tokens the corresponding whole requests took 75.2 and 474.2 seconds measured. These are direct requests to an already loaded server. They are not replacement first-visible-content measurements for the old framework history and do not include loading the weights.
On the desktop alone on 13 September 2026, the first visible content after a 102,480-token history took 1,396.6 seconds, with the prompt read at about 73.4 tokens a second measured. That is 23.3 minutes arithmetic. The same record found 3 of 3 planted codes measured.
On the pair, the nearest measured equivalent is the deep leg above: a 102,478-token history reached first visible content in 1,537.75 seconds, with the prompt read at 66.6 tokens a second measured. Reading speed falls with depth on this configuration, so the 184.6 tokens a second in the table, measured on a 7,841-token prompt, must not be extended to a 100,000-token history. We did not measure the pair's reading rate at that depth with a direct probe.
The saved-slot experiment did not make this architecture warm after a restart. The saved slot was not put back after the restart: the run record's restore field is empty, and the next request read the history again, 102,846 prompt tokens processed and first content after 797.57 seconds against 1,396.6 seconds cold measured. The result record marks that attempt FAIL. Its GGUF records attention.sliding_window = 128 measured header, the sliding-window mechanism behind the full re-read after a restored slot described in the warm-wake study.
DeepSeek V4.1 Flash, measured and parked
We also measured the 4.1 version on the same desktop through a community port of llama.cpp on 14 September 2026. That port keeps a small expert cache on the graphics card and a larger second tier of expert weights in system memory, and streams the rest from a local drive as the model needs them. The tier sizes below are the size of that second cache in system memory. The run worked, but it did not complete a tool call through an agent framework with thinking off or on measured failure. A structured tool call sent straight to the server did succeed with thinking off; with thinking on it did not. We parked it.
| Check | Result | Served window | Label |
|---|---|---|---|
| File used | MXFP4 engram GGUF, 11 shards, 475G reported by the download log | n/a | measured |
| Load to healthy | 143 to 207 seconds across 5 runs | 4,096 to 131,072 | measured |
| Card use at 128K | 29,615 and 29,584 MiB in runs D and E | 131,072 | measured |
| Cold speaking, runs A to D | 6.18 to 6.86 tokens a second | 4,096 to 131,072 | measured |
| Warm speaking, 110 GiB tier | 10.14 to 10.82 tokens a second | 4,096 to 131,072 | measured |
| Speaking at about 100K | 5.44 tokens a second | 131,072 | measured |
| Reading repetitive test text | 22.9 to 39.0 tokens a second | 4,096 to 131,072 | measured |
| Reading real prose | 12.5 tokens a second on 38,198 prompt tokens | 131,072 | measured |
| Early recall, runs A and B | 3 misses at about 2,000 and 3,500 tokens | 4,096 | measured failure |
| Later recall, runs C to E | Exact at 2K, 8K, 30K, 32K, and 100K | 32,768 and 131,072 | measured |
| Tool calls through an agent framework | Not completed, thinking off or on | 131,072 | measured failure |
Runs A to E swept 4,096, 32,768 and 131,072-token windows and second-tier caches of 72, 110 and 125 GiB, so the rows above are not all one configuration. Run A, with the 72 GiB tier, missed the planted code at about 2,000 and 3,500 tokens; its warm speaking rate was 9.25 tokens a second. Run B, with 110 GiB, found the code at 2,000 and missed it at 3,500. Runs C, D, and E then recalled correctly at the listed depths. These are three recorded misses across the first two runs measured.
A separate real-prose probe took 3,064.2 seconds to read 38,198 prompt tokens at 12.5 tokens a second, then recalled its planted code measured. The prose was about 120,000 characters of our own prose documentation, with a code planted at the midpoint. On this V4.1 port, the later batch-setting attempt failed to create a model context. It remains a failed V4.1 attempt; the successful September 20 and 21 batch results above concern V4 Flash 0731.
What this does not show
- No answer-quality result. The page measures speed, memory, recall, and tool plumbing. It does not show that 4-bit answers are better than 3-bit answers.
- No controlled machine-count comparison. The earlier desktop and split rows changed quantization as well as placement. The new IQ3 rows share weights but compare separately tuned configurations and different machines, not machine count alone.
- No full-window occupancy test. The new window is 262,144 tokens; the deepest new IQ3 prompt is 150,103 tokens. The old doubled-cache figure remains arithmetic.
- No validated quantized-cache recommendation. The Q8 combined trial timed reading without an answer check. It does not establish correctness of a quantized attention cache.
- No measured thinking off-switch. The switch is stated in our operator record and remains unprobed here.
- No long-context result for the installed 4-bit preset. Its framework check covered the short legs only.
- No general claim about two-machine serving. This card covers one model family. The wider comparison belongs on the two-box study.
What we got wrong
Our own summary rounded the 3-bit split to 17 tokens a second. The raw run contains 15.89, 16.15, and 17.06 measured. This page prints 15.9 to 17.1 measured.
Our ledger gave the 4-bit no-draft placement as 5, 10, and 28 layers. The raw run says 5 whole layers on the card, 14 expert sets in desktop RAM, and 24 whole layers on the laptop measured. This page uses the raw run.
Our own notes disagreed with each other on active parameters, and one of them carried a figure from a secondary source. We print no active-parameter figure at all: the maker's model card at the pinned revision states none, and the GGUF header records expert counts without the expert width an active total would need measured.
Related pages on this site
The two-box study carries the wider comparison across the models tried on the desktop and laptop. The warm-wake study records the successful restores and this model's unsuccessful attempt. The method page defines the evidence labels used here.
Sources and artifacts
The evidence package under data/ holds the records these numbers come from. Local paths, machine names, network addresses and internal service names are replaced in every file; nothing measured is changed. data/README.md maps every file, lists the redactions with counts, and opens with the places where the package looks as if it argues with the page. data/NUMBERS.md maps each displayed figure to a file and a field.
- The DeepSeek V4 Flash 0731 model card and LICENSE file, read in full at repository revision
7872f01band kept indata/vendor/. - The community GGUF header emitted by llama.cpp, for the base parameter count, expert layout, license metadata, cache layout, served window, runtime build, and model identity.
- The recorded direct runs for the 3-bit and 4-bit bodies, the installed 4-bit preset, the tool-and-recall JSON records, the cold-wake JSON record, and the V4.1 port results.
- The probe and launch scripts that produced these figures.
- llama.cpp build 10919 at source commit
d3146f2b5for the two-machine sweep.
If a number on this page disagrees with a file in its package, the file is right and the page is wrong.