Six large models on the ASUS ROG Flow Z13 alone
At 48K, the laptop read prompts 2.8 to 23 times slower than the desktop. Ling-3.0-flash was the practical exception: desktop-level short-answer speaking at 48K, with a 3.6-minute read.
Measured, 2026-09-20 to 2026-09-21. A 128 GB unified-memory laptop can hold models that do not fit on a 32 GB graphics card. Holding them is not the same as serving them well. We ran six local models on the Z13 by itself, over loopback, then compared the 48K measurements with the same sealed-code probe family on a desktop with one RTX 5090 and 188 GiB of RAM. All six returned the code. The laptop's prompt reading slowed with depth. Only Ling remained a practical alternative at 48K, and even Ling took 216.2 seconds to read the prompt.
The answer
Measured, 2026-09-20 to 2026-09-21. At about 48K prompt tokens, the laptop read at 36.9 to 222.0 tokens per second. The selected desktop comparisons read at 511.7 to 2,594.6 tokens per second. Dividing each desktop rate by its laptop counterpart gives a range of 2.8 to 23 times slower on the laptop. Flash-Next supplies the low end at 2.76 times, rounded to 2.8. The laptop's own best Flash-Next setting, -ub 2048 at 213.0 tokens per second, narrows that low end to 2.4 times. GLM-4.7-Flash supplies the high end at 22.70 times, rounded to 23.
The laptop did not win on deep speaking. These measurements used a sealed-code answer, so every decode figure below is labelled speaking on a short answer. Ling was effectively level at 48K, 23.79 tokens per second on the laptop and 23.25 on the desktop. Qwen3-235B was also close, but the two machines used different quantizations. The laptop's clearest speaking win was shallow Ling: 34.84 tokens per second after a 3K prompt versus 23.22 on the desktop.
-ub 512, and no multi-token prediction (MTP) draft. It read the 47,986-token prompt at 222.0 tokens per second, returned the exact code, spoke the short answer at 23.79 tokens per second, and peaked at 72,179 MiB GTT and 76,002 MiB RAM used.What we ran, exactly
Measured configuration, 2026-09-21. The machine and runtime facts below describe the recorded test configuration, not a general specification for either machine class.
| Laptop | ASUS ROG Flow Z13 GZ302EA, AMD Ryzen AI Max+ 395, Radeon 8060S, 128 GB unified memory with about 119 GiB usable. |
|---|---|
| Laptop runtime | llama.cpp build 10919 at revision d3146f2b5, Vulkan, -ngl 999 -fa on -t 16 --parallel 1, one model and one request at a time, bound to loopback. |
| Power | Balanced mode, on mains. We changed neither the power mode nor the machine configuration. |
| Desktop | One NVIDIA GeForce RTX 5090 (32,607 MiB), 188 GiB system RAM, and a 24-core Intel Core Ultra 9 285K. CUDA llama.cpp kept model experts in system RAM where required. |
| Probe | A synthetic ledger placed a sealed code at its midpoint. Temperature 0; one cold request at about 3K and 48K. A warm prose follow-up was recorded for five models and is absent from the Qwen3-235B log. This page uses only the cold sealed-code decode figures for cross-machine speaking. |
| Dates | Laptop runs and five desktop comparisons: 2026-09-21. Flash-Next desktop comparison: 2026-09-20. All used the same probe family. |
Laptop files were Qwen3-235B-A22B-Instruct-2507 UD-Q3_K_XL, Qwen3.8-Flash-Next UD-Q3_K_XL, Ling-3.0-flash Q4_K_M, DeepSeek V4 Flash 0731 UD-IQ3_XXS, GLM-4.7-Flash UD-Q4_K_XL, and Qwen3.5-122B-A10B UD-Q4_K_S. The desktop used the same files except Qwen3-235B, which used Q4_K_M plus a Qwen3-0.6B Q8_0 draft. All laptop comparison rows used -ub 512; windows were 128K for both Qwen models and the main DeepSeek run, 198K for GLM, and 256K for Ling and Flash-Next. In the desktop table, Ling used 256K and -ub 2048; Flash-Next used 128K and -ub 2048 on the separate Flash-Next server binary, not build 10919; Qwen3.5 used 128K and -ub 4096; GLM used 198K at that model's own default micro-batch, whose number the log does not record; DeepSeek used 256K and -ub 4096; and Qwen3-235B used 128K and -ub 2048 with the draft already disclosed. A separate load allocated a 256K DeepSeek slot and answered a 3K prompt, peaking at 98,482 MiB GTT.
Observed at 48K
| Model | Laptop read | Desktop read | Slower | Laptop speak | Desktop speak |
|---|---|---|---|---|---|
| Ling-3.0-flash | 222.0 t/s | 727.7 t/s | 3.28x | 23.79 t/s | 23.25 t/s |
| Qwen3.8-Flash-Next | 185.4 t/s | 511.7 t/s | 2.76x | 14.39 t/s | 18.75 t/s |
| Qwen3.5-122B-A10B | 159.3 t/s | 2,101.7 t/s | 13.19x | 19.07 t/s | 32.72 t/s |
| GLM-4.7-Flash | 114.3 t/s | 2,594.6 t/s | 22.70x | 10.70 t/s | 129.60 t/s |
| DeepSeek V4 Flash | 56.7 t/s | 690.9 t/s | 12.19x | 11.03 t/s | 12.16 t/s |
| Qwen3-235B | 36.9 t/s | 703.3 t/s | 19.06x | 8.42 t/s | 8.18 t/s |
Measured rates from one cold sealed-code request per configuration. Laptop prompts ranged from 47,985 to 48,052 tokens; desktop prompts ranged from 47,986 to 48,069. Speaking columns are speaking on a short answer, including reasoning tokens where the runtime counted them. The “slower” column is arithmetic: desktop read rate divided by laptop read rate. The ratios divide the one-decimal result fields; server-log rates give 22.71 for GLM and 19.08 for Qwen3-235B. Desktop configurations were Ling at 256K and -ub 2048; Flash-Next at 128K and -ub 2048 on the separate Flash-Next server binary, not build 10919; Qwen3.5 at 128K and -ub 4096; GLM at 198K with its own default micro-batch, whose number the log does not record; DeepSeek at 256K and -ub 4096; and Qwen3-235B at 128K and -ub 2048 with the disclosed draft. Laptop Flash-Next used -ub 512; its fastest tested laptop setting, -ub 2048, reached 213.0 tokens per second.
The cold-read wait
| Model | 48K read time | Peak GTT | Practical result |
|---|---|---|---|
| Ling-3.0-flash | 3.6 min | 72,179 MiB | practical |
| Qwen3.8-Flash-Next | 4.3 min | 61,948 MiB | works, slower |
| Qwen3.5-122B-A10B | 5.0 min | 68,875 MiB | desktop preferred |
| GLM-4.7-Flash | 7.0 min | 24,455 MiB | desktop preferred |
| DeepSeek V4 Flash | 14.1 min | 97,477 MiB | desktop preferred |
| Qwen3-235B | 21.7 min | 103,358 MiB | desktop preferred |
Measured read times for the first five rows, rounded from seconds in the result files. Qwen3-235B did not record a separate read duration, so its 21.7 minutes is arithmetic: 48,027 prompt tokens divided by 36.9 tokens per second. GTT is the laptop's sampled GPU-visible memory counter, not power use.
Flash-Next is the credible fallback if Ling is not the desired model. It fit with a 256K configured window and finished the 48K read in 4.3 minutes, but its short-answer speaking rate at that depth remained below the desktop's. Qwen3.5-122B also finished its read in about five minutes, yet gave up much more speaking speed. GLM used the least GTT and still suffered the largest relative reading gap. DeepSeek and Qwen3-235B came closest to their desktop short-answer speaking rates, but waiting 14.1 or 21.7 minutes for a cold 48K read makes that parity less useful.
Measured trend. Every laptop model read more slowly at 48K than at 3K. The declines ranged from 23% for Flash-Next and Qwen3.5-122B to 79% for GLM. The Qwen3-235B server log shows the cumulative rate continuing to fall after about 30K. That pattern is consistent with attention work becoming dominant, but the run did not instrument attention separately, so the cause is an inference rather than a direct measurement.
Why Ling is the exception
Measured. Ling combines a usable read time with speaking that stayed level with the desktop at the tested depth. Its 47,986-token read took 216.2 seconds, or 3.6 minutes, before the short answer. At 103,887 tokens it read at 149.0 tokens per second in 697.4 seconds and spoke the short answer at 15.09 tokens per second. The desktop was not measured at 104K; its deeper recorded point was 149,711 tokens, with reading at 694.6 tokens per second and short-answer speaking at 22.60 tokens per second.
Measured. The laptop also survived a targeted stability check. A fresh Ling server first answered the recorded 3,007-token crash replay and stayed alive. It then completed 40 of 40 public-source prompts with zero crashes and one server start. Those 40 prompts ranged from 339 to 10,930 tokens. This was not a 48K stability test. It reduces one specific crash concern; it does not establish long-window reliability.
Measured. The MTP draft was the wrong trade at depth. At -ub 2048, it raised speaking on a short answer after 3K from 34.65 to 41.55 tokens per second, then lowered the 48K figure from 23.80 to 16.30. We left it off. Raising -ub from 512 to 2048 also did not improve Ling's 48K read: 222.0 versus 217.7 tokens per second.
What we got wrong
Measured correction. An early note treated the warm follow-up as a normal next turn. It was not. The probe replaced the last question on the same ledger. With Ling, that edit caused 535 tokens to be read again at -ub 512 and 2,067 at -ub 2048. At -ub 2048 the regenerated answer included the sealed code; at -ub 512 it did not. A normal appended turn parts at the end of the old prompt and can use a different restore point. This page therefore uses the result only as advice about editing or regenerating the last prompt, not ordinary conversation.
What this does not show
- No Performance-mode run was made. Balanced mode stayed on throughout, so the study does not measure the laptop's fastest possible power profile.
- No battery life, wall power, temperature, fan noise, sustained thermal throttling, or simultaneous-request capacity was measured.
- The cross-machine speaking task was a short sealed-code answer, not prose. The study does not claim the same ratios for long generation.
- Exact-code retrieval is a correctness check, not a model-quality evaluation. There was one main cold request per depth and configuration, not a statistical sample.
- Only Ling was probed beyond 48K, at about 104K. A configured 128K, 198K, or 256K window is capacity, not proof of speed or stability at its limit.
- The Qwen3-235B 48K prompts, 48,020 tokens on the desktop and 48,027 on the laptop, were longer than the 40,960-token trained context reported in the desktop log.
- The machines used different backends, context windows, micro-batches, batch sizes, cache settings, thread counts, and some model-specific flags. Qwen3-235B also used different quantization and drafting on the desktop, and Flash-Next used a separate desktop server binary. These are practical deployment comparisons, not isolated hardware benchmarks.
Related pages on this site
The two-machine study measures what changes when the laptop helps the desktop instead of serving alone. The provisional batch-size study shows why prompt-reading settings matter on the desktop comparison machine.
Sources and artifacts
The data package contains the redacted result files, memory samples, server logs, probes, launch chains, and the six desktop comparison records. NUMBERS.md maps every displayed figure to a file and field. The runtime was llama.cpp at revision d3146f2b5. If a number on this page disagrees with a file in its package, the file is right and the page is wrong.