Two boxes, one model
A desktop and a laptop joined by one cable, running one large model between them. The pair adds memory. With llama.cpp's batch size raised, the desktop alone answered a 48,024-token request with the same DeepSeek V4 Flash file in 75.2 seconds; the pair took 416.4 at its best completed setting and 474.2 at the one it keeps. The pair still speaks faster on short prompts.
- If the model fits on one machine, run it there, and
raise llama.cpp's micro-batch first
(
--ubatch-size, the number of prompt tokens pushed through the model in one step, with--batch-sizeat least as large). Keep the fastest value that loads and answers correctly at your real window; the batch-size study has the method. On this desktop at-b/-ub 4096the 3-bit DeepSeek V4 Flash file answered a 48,024-token request in 75.2 seconds, 69.5 of them reading the prompt. The pair took 474.2 seconds at its kept-b 2048 -ub 512, 468.4 of them reading, and 416.4 at-b/-ub 2048, its best completed setting. - Use the second machine for memory: for a model that does not fit on one machine, or a dense model much larger than the card (section 05). Expect it to cost reading speed, more the longer the prompt.
- On the pair, a larger micro-batch helped until it
broke.
-ub 2048finished the same request faster; at 4,096 the laptop's side failed 40,960 tokens into the read. We keep 512 as a precaution after that failure, not because 2,048 failed. - Build both ends from the same llama.cpp commit, or a mismatch can answer in fluent nonsense (section 08).
llama.cpp can serve one model from two computers: the server runs on the desktop, and a worker program on a second machine lends it that machine's memory. Between 13 and 20 September 2026 we ran ten models that way, with the desktop's RTX 5090 and an ASUS ROG Flow Z13 with 128 GB of unified memory on the other end of a direct Thunderbolt cable. On 21 September we ran the pair's best case, DeepSeek V4 Flash at 3-bit, against the same file on the desktop alone: same day, same llama.cpp commit, same window, same prompts. Reading is how fast a server takes a prompt in; speaking is how fast it writes the answer; both are llama.cpp's own timings. The pair always added memory: a 283.03 GiB file, larger than either machine's memory, loaded and answered on it. It did not beat a well-set desktop at reading: every comparison re-run after the batch change reads faster on the desktop alone. The one case that had pointed the other way, MiniMax M2.7, was measured against a desktop run left at a quarter of llama.cpp's default micro-batch; at the default the desktop read faster, on a shorter prompt (3,658 tokens against the pair's 7,844). GLM-4.7 Full and the control were not re-run after the batch change. Every speed here is measured on these machines, one request at a time; the one maker's figure is labelled vendor.
Summary
Each finding is scoped to the runs that produced it; the dates are in section 03.
- For DeepSeek V4 Flash, one machine now beats two on long
prompts. Same 3-bit file, same build, same 262,144-token
window and identical prompts, 21 September: the desktop alone at
-b/-ub 4096read 690.9 tokens a second at 48,024 tokens and 632.1 at 150,103; the pair at its kept-b 2048 -ub 512read 102.5 and 51.4. Whole requests took 75.2 seconds against 474.2 at 48,024 tokens (416.4 on the pair at-b/-ub 2048), and 244.8 seconds against 2,933.8 at 150,103, of which reading took 237.5 seconds and 2,922.9, or 48.7 minutes (section 04). - The pair's speaking edge is a short-prompt edge. On the same day and prompts the pair spoke at 14.76 tokens a second against 12.65 at 2,998 tokens, 12.04 against 12.16 at 48,024 and 7.66 against 11.79 at 150,103, and its reading fell from 226.8 to 30.7 tokens a second across one long prompt.
- The pair always adds memory; it did not add reading
speed. A 400.71-billion-parameter model and a
1.03-trillion-parameter model ran on the pair, and the larger file,
283.03 GiB, is bigger than either machine's memory. Against the
desktop at its tuned setting, the pair read slower in every
comparison re-run after the batch change. MiniMax M2.7, which our
records had as faster on the pair, was compared with a desktop run
at
-ub 128; at the default the desktop read faster, on a shorter prompt (section 05). GLM-4.7 Full and the control were not re-run. - A dense model is the pair's best case for speaking. Mistral Medium 3.5, which does not fit the card, spoke at 5.54 tokens a second on the pair on a 441-token prose reply, against 1.97 on its desktop record, a different file from August; it read about ten times slower (section 05).
- A large micro-batch can break the pair. At
-ub 4096the laptop's side failed 40,960 tokens into a 48,024-token read and the request died.-ub 2048finished it, reading 14 percent faster, in 416.4 seconds against 474.2; the pair is kept at 512 as a precaution after the 4,096 failure (sections 04 and 08). - The path works, and it is not free. A 29.94-billion-parameter control model split half and half spoke at 63.8 to 68.0 tokens a second on answers of 54, 10 and 8 tokens, against 223.65 alone on replies of about 200 tokens; its reading figures are on different prompt lengths and windows, so neither pair measures what the cable costs (section 05). Two traps cost more time than any tuning (section 08).
What it is
llama.cpp has an RPC backend. One machine runs
llama-server and holds the model; another machine runs
ggml-rpc-server, and the first machine treats it as a
device, in the same list as its graphics card and its processor.
Layers are assigned to devices at load time. Nothing is pipelined
across requests and nothing is sharded inside a layer: when a token
needs a layer that lives on the second machine, the work and its data
cross the cable and come back.
That makes three places a layer can live on this bench, and the
launcher expresses them as three numbers: whole layers on the
graphics card, layers whose attention stays on the
card while their experts sit in the desktop's system
memory, and whole layers on the second
machine. Every run on the pair is written as those three
numbers in that order, for example 8 / 10 / 25. A whole
layer on the laptop includes its attention: the load logs put the
attention cache for those layers on the laptop, so the laptop does
that part of the attention work.
Every model placed with that middle number is a mixture-of-experts model: most of its weight sits in expert tensors that only some tokens use. That is why the knob exists at all, and it is why the placement question has an answer worth measuring.
| Desktop | One NVIDIA GeForce RTX 5090, 32,086 MiB as llama.cpp reports it (32,607 MiB as nvidia-smi reports it); Intel Core Ultra 9 285K, 16 compute threads of 24 in the control run and 24 in the others; 192,591 MiB of system memory, which is 188 GiB. CUDA 12.8.93, sm_120. measured |
|---|---|
| Second machine | An ASUS ROG Flow Z13 with 128 GB of unified memory: Ryzen AI Max+ 395, Radeon 8060S Graphics, served through Vulkan, which reports uma: 1, fp16: 1, bf16: 0, fp4: 0. measured |
| Memory it offered | 110,592 MiB total, 109,602 MiB free, as the server's own device probe reported it at launch. This is a device report, not a memory benchmark. measured |
| Link | A direct Thunderbolt cable between the two machines, host to host, no switch and no router in between. Measured on 2026-09-15; figures in section 03. measured |
| Model under the headline | DeepSeek V4 Flash 0731: 284 billion parameters, read from the model file's header (measured); the maker's repository page lists 304B because a draft module ships with it (vendor, read 2026-09-15). The maker's model card at the revision we retained states no parameter count at all, and we print no figure for how many parameters are active per token. |
What we ran, exactly
The builds
Both ends of the pair were built from one llama.cpp source
commit, d3146f2b5 (build 10919, GNU 13.3.0,
Linux x86_64), with the RPC backend on, as the server's own first
log line records on 13 September: a CUDA build on the desktop, a
Vulkan build of the worker, ggml-rpc-server, on the
laptop. The later pair runs and the 21 September runs of the 3-bit
DeepSeek file alone used a desktop build of the same commit, read
from the build's own source directory because their logs do not
print it. Other desktop-alone runs used each model's own serving
build: 5f55650a7 for the 8-bit DeepSeek runs
of 15 and 20 September and the 13 September MiniMax run,
c8e03ce81 for the 21 September Qwen runs.
The link, measured
On 2026-09-15, with no model server running: 0.667 ms average round trip over 30 pings, and one TCP stream carrying 2,092 to 2,120 MB/s from desktop to laptop and 1,107 to 1,111 MB/s back. That is a plain socket test, a lower bound rather than a tuned benchmark; the table, the method and its scripts are in the data package. measured
The launch shape
One llama-server on the desktop, one worker process
on the laptop, one request at a time. The 13 to 14 September runs
served a 131,072-token window, except the control, which served
8,192; the 15 and 21 September runs served 262,144 where the rows say
so. Every run on the pair from 13 to 21 September left llama.cpp's
batch sizes at their defaults, -b 2048 -ub 512, except
the 21 September rungs that raised them.
The probes
- 13 to 15 September: a short probe (a 21-to-55-token question, up to 200 tokens of answer), a depth probe (filler padded to about 8,000 tokens with a code word planted in the middle) and, on some runs, a fresh 8,000-token prompt with a different code.
- 15 to 21 September: a synthetic ledger with a sealed code planted halfway, read cold, answered with the code. On 21 September the desktop-alone and pair runs used identical prompts, 2,998, 48,024 and 150,103 tokens long, and on the pair each cold read was followed by the same ledger with its last question changed, then a real follow-up turn.
Speeds are llama.cpp's own timings fields, per request. Speaking figures count every token the model generates, hidden reasoning included: DeepSeek V4 Flash thinks by default, and a short answer can be mostly reasoning. Where a figure comes from a short answer, the table says how long the answer was. A whole-request time is the probe's wall clock from sending the request to the finished answer; a reading time is the server's own prompt-evaluation time inside it.
Dates
| 2026-09-13 | Link brought up, both ends built from one commit, control run, then the first sweep on the pair: DeepSeek's 8-bit, 3-bit and 4-bit files, the draft model, the trillion-parameter model, GLM-4.7 Full, Qwen3.5-397B and the 400B. |
|---|---|
| 2026-09-14 | Three tool-and-recall attempts on the 3-bit pair; Qwen3-235B and MiniMax M2.7 on the pair. |
| 2026-09-15 to 16 | The 4-bit preset on the installed service; the link measured; long reads at a 262,144-token window; Mistral Medium 3.5 on the pair on the 16th. |
| 2026-09-17 | The 3-bit DeepSeek file measured alone on the desktop for the first time. |
| 2026-09-19 to 21 | MiniMax M3 on the pair early on the 20th. The micro-batch sweep on the desktop alone, then, on the night of the 21st, the pair re-tuned against the 3-bit file alone. |
The same file on one machine and on two
On 21 September we ran the 3-bit DeepSeek V4 Flash file
(UD-IQ3_XXS, 104.2 GB) both ways: the same llama.cpp binary, a
262,144-token window and identical prompts. On the desktop alone the
file ran with every layer's attention on the card and the experts of
36 of its 43 layers in system memory. On the pair it ran at
8 / 10 / 25. All figures measured,
2026-09-21, one request at a time, the sealed code quoted back
correctly on every read that finished.
| Setup | Reading, 2,998 tokens | Reading, 48,024 tokens | Reading, 150,103 tokens | Speaking, 2,998 | Speaking, 48,024 | Speaking, 150,103 |
|---|---|---|---|---|---|---|
Desktop alone, -b/-ub 4096 (the setting it now runs at; see the note for the 150,103-token run) |
539.6 | 690.9 | 632.1 | 12.65 | 12.16 | 11.79 |
Desktop alone, -b 4096 -ub 2048 |
208.1 | 426.5 | 403.8 (149,898 tokens) | 12.62 | 12.29 | 11.80 |
The pair, -b 2048 -ub 512 (kept) |
161.8 | 102.5 | 51.4 | 14.76 | 12.04 | 7.66 |
The pair, -b/-ub 2048 |
224.9 | 117.0 | not run | 14.61 | 11.70 | not run |
The pair, -b/-ub 4096 |
250.5 | failed after 40,960 tokens | not run | 14.39 | none | not run |
Speaking figures are short answers: 59 to 86
tokens generated, most of it hidden reasoning before a one-line code.
On the pair, a longer reply of about 400 tokens on the same documents
spoke within 3 percent of the short answer at every depth (14.73,
12.05 and 7.85); we did not time a longer reply on the desktop alone.
The desktop run at -ub 2048 read a different
150,000-token ledger, 149,898 tokens long. The desktop's 150,103-token
run records ubatch=skip: the probe passed no batch
setting, and the server was started on the fast preset, whose default
is -b/-ub 4096; its card reading at load, 28,570 MiB,
matches the explicit 4,096 run (28,490) and not the 2,048 one
(27,092). The package README lists this as its contradiction 10.
| Time to the answer, seconds | 48,024 tokens: reading | 48,024: whole request | 150,103 tokens: reading | 150,103: whole request |
|---|---|---|---|---|
Desktop alone, -b/-ub 4096 | 69.5 | 75.2 | 237.5 | 244.8 |
The pair, -b/-ub 2048 (best completed) | 410.5 | 416.4 | not run | not run |
The pair, -b 2048 -ub 512 (kept) | 468.4 | 474.2 | 2,922.9 | 2,933.8 |
Reading is the server's own prompt-evaluation time; the whole request adds the answer and the probe's round trip. 2,922.9 seconds of reading is 48.7 minutes, and 237.5 is 4.0. measured
Reading. The desktop alone read faster at every
length, each machine at its own setting (-b/-ub 4096
alone, -b 2048 -ub 512 on the pair): 3.3 times the pair's
speed at 2,998 tokens, 6.7 times at 48,024 and 12.3 times at 150,103.
At the same micro-batch of 2,048, with -b 4096 on the
desktop and -b 2048 on the pair, the pair was slightly
faster on the short prompt, 224.9 against 208.1, and 3.6 times slower
at 48,024 tokens, 117.0 against 426.5. The gap grows with depth, and
the server's own progress lines on the 150,103-token read show it
directly:
| Stretch of the 150,103-token prompt | Desktop alone, -ub 4096, tokens a second | The pair, -ub 512, tokens a second |
|---|---|---|
| tokens 0 to 4,096 | 690.7 | 226.8 |
| tokens 40,960 to 49,152 | 673.7 | 75.2 |
| tokens 90,112 to 98,304 | 614.6 | 43.3 |
| tokens 139,264 to 143,360 | 573.7 | 30.7 |
Tokens read per second over each stretch, worked out from the timestamps in each server's progress lines for the same prompt by the script in the package. measured
Speaking. The pair still spoke faster on the short prompt, 14.76 against 12.65, about a sixth. By 48,024 tokens the two were level (12.04 against 12.16), and at 150,103 the desktop alone was faster (11.79 against 7.66). Our reading of the pattern, which no run here instruments: the laptop holds 25 whole layers, attention included, and attention work grows with the prompt, so the laptop's share of every token grows with depth.
The micro-batch on the pair. Raising it from 512 to 2,048 bought 39 percent on the 2,998-token read and 14 percent at 48,024 tokens, where the whole request took 416.4 seconds against 474.2. At 4,096 the 48,024-token read stopped: the server's last progress line was at 40,960 tokens, then it logged that the remote worker had crashed or returned a malformed response, and the request failed. Our notes from the laptop's system log record its graphics driver resetting the device and the worker restarting itself; that log was not kept, so the desktop's log is the only file behind this paragraph. The worker was serving again three minutes later for a 150,103-token read at 512, which completed without a fault. We keep 512 as a precaution after that failure; 2,048 did not fail.
What a follow-up turn costs on the pair. A real next turn, with the model's reply sent back and a short new question added, re-read only the new messages at either setting. A changed last message, here the same document with its last question swapped, went back one micro-batch, so a larger micro-batch costs more there; we measured the swapped question only.
| The pair, 48,024 tokens deep | -ub 512 | -ub 2048 |
|---|---|---|
| A real next turn: tokens re-read, time | 238, 5.4 s | 228, 5.0 s |
| The last question changed: tokens re-read, time | 533, 10.0 s | 2,065, 28.2 s |
At 150,103 tokens deep and -ub 512
the same two cases re-read 239 tokens in 9.7 seconds and 533 tokens
in 20.7 seconds. measured
How the picture changed
- 13 September, the comparison our records first made. The 3-bit file on the pair, at a 131,072-token window and default batch sizes, spoke at 15.9 to 17.1 tokens a second and read 184.6 to 187.9 on prompts of about 8,000 tokens. It was set against the 8-bit file alone in a 12 September framework check: speaking 9.2 to 9.7, and reading of about 75.5 and 73.8, the framework's own estimates, on counted prompts of 52,931 and 102,481 tokens. Different files, depths, days and methods. measured
- 17 September, the first desktop run of the same 3-bit file. It spoke at 13.5, the mean of three 400-token replies to a short prompt, at 262,144 with the experts of 36 layers in system memory. That comes from the day's own results table; its raw logs were not kept. measured
- 21 September, the reversal: the tables above. The 15 September reads at a full window are on what a full window costs, and the 8-bit file's batch runs of 20 September on the DeepSeek V4 Flash card.
The laptop by itself. Run alone, with the same
3-bit file at a 131,072-token window and -ub 512, the
laptop read the 48,024-token prompt at 56.7 tokens a second, slower
than the pair. Six models on the laptop alone have
their own page.
The other models
One row per model, on the pair at llama.cpp's default batch sizes, one request at a time. Where a model has a desktop-alone figure, the run that produced it is named, because it was always a different night and often a different setting. All speeds measured; parameter counts and sizes are read from each file's own header. The full cells, with every probe, are in the data package.
| Model and file | On the pair | On the desktop alone | What we take from it |
|---|---|---|---|
| GLM-4.7-Flash, the control, 29.94 B; half the layers each side, 8,192-token window | Speaking 63.80, 67.97 and 63.86 on answers of 54, 10 and 8 tokens; reading 1,125.28 on a 1,940-token prompt | Speaking 223.65 on replies of about 200 tokens; reading 4,447.44 on 20,221 tokens at a 131,072 window (2026-09-12) | The path works. The speaking figures are different reply lengths and the reading figures different prompt lengths and windows, so they are not a measured cost of the cable. Not re-run after the batch change. |
| Qwen3-235B-A22B 2507, Q4_K_M, 235.09 B; 6 / 36 / 52 | Speaking 5.97 on a 49-token answer, 5.54 on a 10-token answer at 8K depth; reading 97.6 on 7,844 tokens | 2026-09-21, 48,020 tokens: reading 285.8 at the defaults, 703.3 at -b 4096 -ub 2048 (the setting it serves) and 944.2 at -b/-ub 4096 |
Slower to read on the pair, even against the desktop's default micro-batch on a prompt six times longer. Speaking not comparable on answers of 10 and 49 tokens. |
| MiniMax M2.7, UD-IQ4_XS, 228.69 B; 6 / 16 / 40 | Speaking 9.66 on a 401-token answer, 8.98 on 128 tokens at 8K depth; reading 70.4 on 7,844 tokens | 2026-09-13: speaking 9.8 to 10.2 across eight replies of 64 to 900 tokens (the 900 all hidden reasoning), reading 45.6 on 6,776 tokens at -ub 128, a 64K window, f16 cache. 2026-09-19, 131,072 window, q8_0 cache: reading 100.8 on 3,658 tokens at the default -ub 512, 656.5 on 43,909 at -ub 4096 |
Same speaking speed. The faster reading on the pair was an artifact of the desktop run's -ub 128, a quarter of the default, inherited from another model's launcher. At the default the desktop read faster, on a shorter prompt: 3,658 tokens against 7,844. |
| GLM-4.7 Full, IQ3_XXS, 358.34 B; 4 / 34 / 55, q8 cache | Speaking 5.66 short, 5.59 at 8K depth; reading 51.1 at depth | Speaking 5.51, reading 61.36 on 20,221 tokens at 131,072 with a q5_1 cache (2026-09-12) | Slower reading on the pair; the cache precision differs too. Not re-run after the batch change. |
| Qwen3.5-397B-A17B, UD-IQ3_XXS, 402.94 B; 6 / 25 / 30 | Speaking 12.61 on a 10-token answer at 8K depth, 15.31 on a 108-token answer; reading 98.7 and 108.2 | 2026-09-21, 131,072 window: reading 224.5 on 48,029 tokens at the defaults, 932.2 at -b/-ub 4096; speaking 16.39 on an 11-token answer |
Slower to read on the pair even against the desktop's default on a prompt six times longer. |
| Mistral Medium 3.5, dense, Q3_K_S, 54.4 GB, with a 3B draft model patched to its vocabulary; 16 layers on the card, 72 on the laptop; 2026-09-16 | Prose speaking 5.54 on a 441-token reply to a 1,453-token prompt (3.81 with no draft at 8 / 80 layers); 4.71 on a 311-token reply to an 8,437-token request, 1,377 of its tokens already cached and 7,060 read; 2.82 after a 32,509-token read that took 1,027.1 seconds at 31.65. Reading 43.8 on the 1,453-token prompt | Its field card, 2026-08-06, a different and larger file (UD-Q4_K_XL) at a 32K window: prose 1.97 with the draft, reading 435 | The one model here that spoke clearly faster on the pair, across different files and months: on the desktop much of a dense model this size sits in system memory, on the pair in the laptop's graphics memory. Reading fell about ten times. |
| MiniMax M3, IQ3_XXS with its sparse-attention tensors, 167.6 GiB; 4 / 32 / 24; 2026-09-20 | At -ub 4096 and 2048 it could not allocate its compute buffers at either of two layer splits. At -ub 1024 it loaded and read a 3,658-token prompt at 17.58; speaking not measured |
A smaller file with the same attention (Q2_K_L, 142.6 GiB) read the same 3,658 tokens at 147.5 with -ub 2048 and spoke 8.90 on an 82-token answer (2026-09-20) |
The compute buffer grew with the micro-batch, not the layer split. At the one setting that fit, the pair read at an eighth of the desktop's speed, on a larger file, since deleted. |
| Llama 4 Maverick, UD-Q3_K_XL, 400.71 B, 167.21 GiB | Speaking 12.02 to 13.18; reading 147.4 and 148.5 | Never served from one machine here | A capacity demonstration; files deleted 2026-09-14. |
| Kimi K2.7-Code, IQ1_M, 1.03 T, 283.03 GiB | Cold first turn 0.67 speaking; warm 8K turn 5.17 speaking and 26.10 reading | Not measured alone | The second capacity demonstration; weights removed 2026-09-17. |
Ten models loaded on the pair, including the control. Against the same file alone, DeepSeek V4 Flash's speaking edge holds only on short prompts (section 04); Mistral Medium 3.5 spoke faster than its desktop record, on a different file; four models read more slowly on the pair and spoke at about the same pace or slower; MiniMax M3 loaded only at a small micro-batch; two are capacity demonstrations; the control is a proof of path.
The placement rule, revised
Moving layers onto the second machine speeds up speaking only when it lets the graphics card hold more whole layers, and only while the prompt is short. As the prompt grows, the second machine's share of the attention work grows with it, and the gain shrinks, then reverses. Reading on the pair was slower than the desktop alone at the desktop's tuned setting in every case we re-measured.
The mechanism we believe is simple, and worth stating as belief rather than measurement. Every token drags some bytes through some memory: bytes on the card are fastest, bytes in the desktop's system memory are slower, and bytes on the other end of a cable that carries about two gigabytes a second (section 03) are slower again. Moving expert weight off the desktop helps when the freed card space fills with layers the card can serve at full speed. But a whole layer on the laptop brings its attention with it, and attention work grows with the prompt, so the laptop's slower share of every token grows too. That fits DeepSeek's speaking going from 17 percent ahead at 2,998 tokens to 35 percent behind at 150,103, and the pair's reading falling from 226.8 to 30.7 tokens a second across one long prompt. The dense Mistral Medium 3.5 fits it from the other side. We did not instrument any of it: no byte counter, link utilisation or per-layer timing.
Reading has a plainer cause too. For models whose experts sit in system memory, a larger micro-batch makes the desktop's reading several times faster (the batch-size study explains why); on the pair it bought 14 percent before the laptop's side failed. The 13 September comparisons ran the desktop at the default setting or below it.
The draft model, the installed preset and the checks
DeepSeek's draft model, on the card, 2026-09-13, raised speaking at 8,000 tokens of depth by 12 percent on the 3-bit file (16.15 to 18.12) and 51 percent on the 4-bit file (12.78 to 19.25), and cost reading on both (184.6 to 148.8; 145.4 to 121.9); placement moved as well. It was not re-measured at 262,144 or with a larger micro-batch. The 4-bit file with the draft, as a preset of the installed service on 2026-09-15 at 131,072, spoke 19.55 and 19.59 and read 125.1 on 7,841 tokens; its check covered the short legs only. measured
A separate check through an agent framework, 2026-09-14, planted three codes in 48,000 and 96,000 tokens of history and required a tool call at depth. On the 3-bit pair the long legs passed, 3 of 3 codes and a tool call at 52,928 and 102,478 tokens with the first word after 543.65 and 1,537.75 seconds, only inside an attempt that failed on an unrelated leg; a later attempt passed and covers the short legs only. The tables, the attempts and the load figures are in the data package, and the DeepSeek V4 Flash card carries the model-level record.
Two traps behind the advice
Each is an interaction between a reasonable default and a two-machine setup the default did not anticipate, and each is cheap to avoid once seen. The figures are from the campaign's logs.
1. Build both ends from one commit, or get confident gibberish
The first split run of the campaign loaded cleanly, was ready to
serve after 414 seconds, and returned HTTP 500 on the
first request: the laptop's worker had been built from a different
commit than the desktop's server. The second time was quieter. With a
server built from a llama.cpp pull-request branch and a worker built
from mainline, a different model loaded, stayed up, and produced
fluent nonsense: the same word repeated. The branch adds one ggml
operation, shifting every operation identifier after it
(GGML_OP_COUNT 101 to 102, RPC protocol 0 to 1), and the
host never asks the worker what it can do. Build both ends
from the same source commit and record it. If split output
is fluent and wrong, suspect the build before the model.
2. A micro-batch the laptop cannot carry
At -b/-ub 4096 the pair read a 2,998-token prompt
normally, then failed 40,960 tokens into a 48,024-token read (section
04). The short read gave no warning. On a pair, test any
larger micro-batch at your longest real prompt, not a short
one.
Two more traps from the first nights, a watchdog that stopped a healthy server by probing a busy worker's port and a worker that started before the laptop's graphics driver and served from its processor for hours, are written up in the data package with their dates.
What we got wrong
This site publishes corrections of its own records as a genre. Where our records and the logs disagree, the log's number is printed.
| Our records said | What the records show | This page now says |
|---|---|---|
| The pair made DeepSeek V4 Flash faster than the desktop alone | That comparison set a 3-bit file on the pair against an 8-bit file alone, with default batch sizes on both. The same 3-bit file alone at its tuned setting read 3.3 to 12.3 times faster than the pair at its kept one, finished a 48,024-token request in 75.2 seconds against 474.2, and spoke as fast at 48,024 tokens and faster at 150,103 | One machine is the better choice for long prompts; the pair's edge is short-prompt speaking (sections 04 and 06) |
| MiniMax M2.7 read faster on the pair (70.4 against 45.6) | The desktop run used -ub 128, a quarter of llama.cpp's default. At the default the desktop read 100.8, on a 3,658-token prompt against the pair's 7,844 |
At the default micro-batch the desktop alone read faster, on a shorter prompt; the apparent gain came from the desktop run's -ub 128 |
| M2.7's desktop reading was on "a 4,000-token prompt", speaking "9.9 to 10.2" | The probe was named for about 4,000 tokens; the server counted 6,776. The run's eight timings span 9.80 to 10.19 | 6,776 tokens, and 9.8 to 10.2 |
| Qwen3.5-397B reads at 385 on the desktop alone, carried from an earlier run | The 385 was a 2,617-token prompt at a 64K window in August. Measured on 21 September: 224.5 on 48,029 tokens at the default batch sizes, 932.2 at -b/-ub 4096 |
The measured figures; the 385 is retired |
| Every run served a 131,072-token window | The control run served 8,192 | The control's row says so |
| No run showed the laptop's graphics chip doing attention work | The load logs put the attention cache for the laptop's layers on the laptop, so it computes those layers' attention | It does; how much time that takes is not measured |
| The pair speaks the 3-bit file at 15.9 to 17.1 | True at a 131,072 window on 13 September. Other nights gave 14.8 to 15.4 on similar short answers, including 15.2 to 15.4 at the same window | The band, dated (section 04); the other nights are in the data package |
Our internal ledger of the first runs also rounded several numbers or quoted the best probe without saying so; the nine rows where it and the logs disagree are in the data package. The pattern is worth more than any single row: the numbers that drifted all drifted in the flattering direction, for the pair. A band became its top end, three attempts became one pass, and a comparison across two files and two default settings became a speed-up.
What this does not show
- One same-file comparison, one model. Section 04 compares one DeepSeek file on one machine and on two, on one day, each machine at its own tuned or kept setting. The other models' rows cross nights, builds, windows and prompt lengths, and each row says which.
- Nothing here measures quality. The only correctness checks are planted codes quoted back and tool calls that returned.
- Windows up to 262,144 tokens. The longest prompt on the pair was 230,835 tokens, the longest on the desktop alone in these records 150,475. We have no data past that.
- One failure at
-ub 4096, not retried; the laptop's side of it rests on notes, not a kept log. - DeepSeek's speaking is measured on short answers, which 400-token replies matched within 3 percent on the pair; on the desktop alone we did not time a long reply.
- Three files on this page are no longer here. Llama 4 Maverick was deleted on 2026-09-14, Kimi K2.7-Code's weights on 2026-09-17 and the MiniMax M3 file used on the pair by 2026-09-21. Their rows record that the files loaded, not a recommendation.
- One cable, one pair of machines, runs from 13 to 21 September. A different second machine, link or llama.cpp build could move all of it.
Sources and artifacts
The evidence ships as a data package: the files the runs actually produced, not a summary. If a number on this page disagrees with a file in the package, the file is right and the page is wrong; tell us and we will fix the page. data/README.md maps every file, states its provenance and the redactions, and opens with the places where the package looks like it argues with itself; data/NUMBERS.md maps every figure on this page to its file and field; data/moved-from-the-page-2026-09-26.md holds the tables and write-ups that left this page on 26 September, each with its sources.
- 15 to 21 September: server logs, result files and memory samples for the pair and the 3-bit file alone, the long reads at a 262,144-token window, the 8-bit file's micro-batch runs, the MiniMax and Qwen desktop runs, and the probe scripts.
- 13 to 15 September: the link measurement with its scripts, one log per model on the pair, the loader's header lines, the tool-and-recall records, the installed service's run record, the placement comment block and the maker's model card at a pinned revision.
Machine names, network addresses, ports, absolute paths, service names, our internal keys and the agent framework's tool names are replaced with marked placeholders or public names, file by file, as the package README lists; nothing that carries a measurement is altered. Withheld: the internal ledger that section 09 quotes, the session notes, and the desktop-only records of 12 September, which carry private labels.
Every claim on this page is dated. llama.cpp's RPC backend is under active development; if today is much later than 21 September 2026, treat everything here as unverified since then.
Related pages on this site
The DeepSeek V4 Flash field card is the model-level record for the model in sections 04 and 07. The batch-size study is the desktop-alone half of this story across the whole roster. What a full window costs has the long reads of 15 September. The laptop-alone study runs six models on the second machine by itself. The September roster lists what else this desktop serves. The wrong flag is the same editorial habit applied to an earlier record: a default nobody had checked, measured and corrected.