Graphometer
Guide · 26 September 2026 · arithmetic with recorded runs from 15 September (desktop local calendar) · one RTX 5090 desktop; one laptop comparison

Bytes per token

Predicting Inkling 975B: 5.2 to 5.7 tokens per second by arithmetic, server-reported 3.43 on short prose.

A giant mixture-of-experts model can use only a small part of its weights for each new token and still speak slowly. For the Inkling 975B UD-IQ1_S files here, that part includes 6.01 GB of routed expert weights per token (arithmetic). Transferring the effective rate from smaller models gives 5.2 to 5.7 tokens per second (arithmetic), before allowing for different kernels, attention, or disk misses. The recorded short-prose reply reached 3.43 tokens per second (server-reported). This is a useful budget before a download, not a promise of speed.

Independent work. No affiliation with the makers identified in the footer.
01

What it is

A mixture-of-experts model chooses a few feed-forward networks, called routed experts, at each expert layer. The file contains all the choices. A token uses only the selected ones, plus shared experts, attention and the rest of the network. Total file size answers a capacity question. The bytes used at each step help answer a speed question.

For a configuration that keeps routed experts in system RAM, start with their stored byte count. Use the actual quantization of each tensor, the named array of weights in a GGUF file. A label such as IQ1_S does not mean every tensor costs exactly one bit per weight.

B = sum over CPU expert layers of
    (bytes of the routed expert banks x selected experts / total experts)

R = B_anchor x measured tokens per second
predicted tokens per second = R / B_target

All three expressions are arithmetic. B is an expert-weight budget per logical output token. R is an effective expert throughput inferred from a complete decode, not a direct measurement of RAM traffic. We found no direct RAM-bandwidth record in the evidence used here. If instead you divide bytes per token by bytes per second, the result is seconds per token.

The estimate assumes the selected expert weights have to be supplied for each step. It ignores weight reuse in cache and treats quantized file bytes as a proxy for runtime traffic. Computation, synchronization, activations and attention can all consume time. Speculative decoding can change how many output tokens one weight pass buys. This is an approximate budget, not a proven hardware upper bound.

02

Count the bytes at their actual placement

We read the GGUF metadata and tensor directories afresh on 26 September, stopping before the weight payloads. The package lists every shard, its exact file size and header hash, and every tensor and its contribution. No weights were loaded. Decimal GB means bytes divided by 1,000,000,000 (unit conversion).

Model filesFile GB
(arithmetic)
Routed expert selection
(measured header)
RAM expert GB/token
(arithmetic)
Inkling-Small UD-Q3_K_XL119.556 of 2562.655
Qwen3.5-397B-A17B UD-IQ3_XXS149.8410 of 5122.563
DeepSeek V4 Flash 0731 UD-Q8_K_XL161.876 of 2563.449
Inkling 975B UD-IQ1_S270.166 of 2566.006

File sizes are measured filesystem byte counts, converted and rounded here. Header values describe these files. Qwen's RAM count covers blocks 0 to 56 only; the other rows cover every routed expert bank. The extra Qwen draft block is not counted as a normal output-token pass. The DeepSeek Q8 file name is the product name; its routed expert banks are MXFP4 in this header.

For Inkling 975B, the routed banks contain 256,276,168,704 bytes. Multiply by 6 / 256 to get 6,006,472,704 bytes per token (arithmetic). The first two of its 66 blocks are dense, so this sum contains the expert banks in blocks 2 to 65. It uses each bank's own storage type, including its quantization scales.

Other weights still matter. Here is the target's remaining inventory and how it enters this particular calculation:

Inkling 975B categoryStored bytes
(header arithmetic)
Treatment
Shared experts5,322,866,688Both shared experts run; configured on GPU
Attention and short convolution weights6,091,720,704Configured on GPU; cache traffic extra
Dense feed-forward weights662,962,176Configured on GPU
Routing, norms and other weights407,511,304Excluded from expert-only denominator
Output head and output norm694,763,520Configured on GPU
Input embedding694,738,944One row is 3,456 bytes, not the whole table

These categories are an inventory, not a measured traffic trace. The RAM denominator counts only routed expert matrices. It excludes the tiny input lookup as well as work on the GPU. Adding all those GPU bytes to the RAM count would mix two different memory systems. If your placement puts shared experts, attention or the output head on the CPU, account for those separately before using this estimate.

03

What we ran, exactly

The anchor runs were recorded on this site's RTX 5090 desktop on 15 September 2026 in the desktop's local calendar; the stored created values run from 15 September 23:52 UTC through 16 September 02:52 UTC. They served a 262,144-token window but used a short water-pump prompt, not a full window. Each requested a paragraph with a 320-token cap and thinking disabled; the server was started at temperature 1.0; the water-pump request set temperature 0. Each returned visible prose and stopped before the cap. One request ran at a time, with 24 compute threads and 24 batch threads.

Recorded setupllama.cpp commitPlacement and draft
Inkling-Small946fc11d1--n-cpu-moe 42; f16 cache; no speculative draft
Qwen3.5-397Bc8e03ce--n-cpu-moe 57; --no-mmap; draft-mtp, maximum 6, gate 0.75
DeepSeek V4 Flash Q85f55650-cmoe, all routed experts in system RAM; --no-repack; no speculative draft

All three launches left batch-size flags unset. The full anchor flags and request construction ship with the page. The sweep's contemporary record notes concurrent download activity. These are historical observations, not a controlled bandwidth experiment.

Inkling 975B used a 131,072-token window, f16 cache, flash attention, one slot and 24 desktop threads. Its launcher points to the draft Inkling implementation whose retained build record is 10897 / 946fc11d1. Run A used --n-cpu-moe 66. Run A is the paging placement: the desktop log warns that CPU tensor overrides were used with mmap enabled, and the run ledger marks that placement as paging. The 3.43 tokens per second figure is the server's decode rate for that placement, not a measurement of resident RAM traffic. Run B used --n-cpu-moe 42 and moved blocks 42 to 65's routed experts to an ASUS ROG Flow Z13 through remote procedure call (RPC), with the worker requested at 8 threads. Attention and shared experts stayed on the desktop card. The probe used temperature 1.0 and reasoning effort none.

Neither target launch set -b or -ub. That branch's defaults were 2048 and 512, respectively (recorded source). This matters to prompt reading. A tuned reading result cannot replace these decode observations; see the batch-size study.

04

Observed: prose anchors, not tiny answers

These speeds are the server's decode rate, timings.predicted_per_second, not request wall time. On Inkling build 946fc11d1, the rate is (completed tokens minus one) / eval seconds. The Qwen and DeepSeek predicted_per_second values divide the completed token count itself by eval seconds. We retain these different server conventions in the arithmetic. The rows show individual repeated results, not confidence intervals. Word counts split the recorded visible text on whitespace.

Short-prompt outputDecode t/s
(server-reported, 15 Sep, desktop local calendar)
B x rate, GB/s
(arithmetic)
Inkling-Small: finished 174-word paragraph, 202 generated tokens each11.77; 11.6931.26; 31.04
Qwen3.5-397B: finished 161-word paragraph, 184 generated tokens each12.76; 12.8432.70; 32.92
DeepSeek V4 Flash Q8: second finished 166-word paragraph, 209 generated tokens9.9334.26

DeepSeek's first prose reply was much slower: 2.07 t/s (measured), also a finished 166-word paragraph of 209 tokens. We retain it in the package. The second request reused a prompt prefix. We use that second decode as the warm operating estimate, not the first request or an average that hides the difference. These records do not isolate why the first reply was slow.

Qwen used speculative decoding and recorded 90 accepted draft tokens out of 106 proposed in each reply (measured counters). Its product of bytes and output rate is therefore only an effective proxy. It lands between the two non-speculative models; removing Qwen leaves the selected range unchanged. Neither the proxy nor this agreement measures physical RAM bandwidth.

We use 31.0 to 34.3 GB/s (arithmetic) as the observed operating envelope of that proxy. It is configuration-dependent. The batch sweep's Qwen result, for example, reports 16.41 t/s after a long read but returned just a 17-character code word (measured, 21 September). That is not the prose anchor used here.

05

Apply the budget to Inkling 975B

Expert budget: 6,006,472,704 bytes/token
Effective rate from selected prose runs: 31.044 to 34.263 GB/s
Predicted decode: 5.168 to 5.704 tokens/s
Rounded estimate: 5.2 to 5.7 tokens/s (arithmetic)

This transfers an effective rate between different models and builds. The estimate was reconstructed on 26 September using already recorded runs; it is not a prospectively tested prediction. The target's measurements on 15 September in the run ledger's local calendar were:

Reply requestedDesktop: tokens; t/s; eval ms
(server-reported)
Pair: tokens; t/s; eval ms
(server-reported)
Short harbour-town prose71; 3.43; 20424.3580; 3.13; 25201.99
Weather tool call18; 4.98; 3414.7518; 3.26; 5222.13
Planted code word, 12 tokens, from a 1,950-token prompt12; 4.65; 2367.2312; 3.34; 3293.87
Planted code word, 12 tokens, from a 7,641-token promptNot run12; 3.35; 3279.71

On build 946fc11d1, the server's printed tokens per second is (completed tokens minus one) / eval seconds. Raw eval milliseconds are shown beside every desktop and pair rate. Dividing the completed token count itself by eval seconds gives a different rate; the desktop tool call then falls inside the arithmetic prediction band. The short-prose result remains below the estimate.

All returned visible prose, a tool call or the requested code word, with no recorded reasoning text. The prose replies stayed below the 120-token cap; the tool and recall requests had caps of 80 and 40. The saved prose stdout contains only a prefix, so we do not assign a full-reply word count. The server logs record an untruncated stop.

The desktop's server-reported range is 3.43 to 4.98 t/s, but its top value is an 18-token tool call. The pair's complete server-reported range is 3.13 to 3.35 t/s. These ranges use the Inkling timing convention above. The most relevant prose comparison is 3.43 versus 3.13, on different sampled replies to the same short prompt. None of these target runs is a sustained essay-speed measurement.

The arithmetic gets the scale of the problem: a low single-digit speaking rate is plausible even with very low-bit files. It does not reproduce the short-prose speed. The model leaves out costs that the recorded decode includes, and the runs do not apportion the gap among them.

06

Why the second machine did not add speaking speed

The split changes where expert work runs, not how many expert bytes one token needs. From the same headers, the pair puts 3,647,766,528 bytes per token on the desktop's routed-expert path and 2,358,706,176 on the laptop's path (arithmetic). Their sum is still the same expert budget.

Approximate time for one token on the pair:
  desktop expert bytes / desktop effective rate
+ laptop expert bytes / laptop effective rate
+ dependent transfers, attention and other work

Those layer stages depend on the preceding activations. Their times cannot be replaced by a single division by the sum of both machines' memory bandwidths. In this placement, remote experts feed results back into work that remains on the desktop. More memory capacity can help a model fit, while remote work and communication can still make a token take longer.

The recorded slowdown is consistent with that added cost, but this run did not time RPC transfers or measure the laptop's effective memory rate. It does not prove that cable latency alone caused the slowdown, or that every split would lose. It shows that this particular placement lost on each matched prompt type. Two boxes covers the broader capacity and speed trade.

07

Before you download a model

  1. Get an exact tensor inventory for the quantization. Use a published header dump or header bytes for every shard. A parameter count and an advertised average bit width cannot tell you which tensors received higher precision.
  2. Choose a plausible placement. Count selected routed expert banks in RAM; account separately for shared experts, attention and the output head if those also stay on the CPU. Leave room for the cache and working buffers on the card.
  3. Calibrate on a recorded prose run from your own machine. Match the build, thread count, --n-cpu-moe, quantization family and draft settings as closely as possible. Keep first-request and warm figures separate.
  4. Divide effective bytes per second by target bytes per token. Treat the answer as a budget. If even that looks too slow, extra model capacity does not solve the speaking problem. If it looks acceptable, a real run is still needed.

The shipped header reader takes existing GGUF paths and writes metadata and tensor descriptors only. The calculator recomputes this page from its shipped header dumps and response records. Neither script starts a model.

08

What this does not show

  • A portable RAM rating. The effective rate depends on the build, thread count, placement, quantization kernels, cache state and draft branch. The Inkling implementation here was a draft llama.cpp branch. Qwen's speculative draft is a different issue and is separately disclosed.
  • Long-context speaking or prompt reading. Serving a large window does not mean the anchor replies filled it. Attention and cache work grow differently from routed expert bytes. Batch settings change prompt reading; this decode calculation does not predict reading time.
  • A measured total memory-traffic count. Headers describe storage. Repacking, dequantization, caches and speculation change runtime traffic. The tensor inventory documents what was included and excluded.
  • A controlled two-machine mechanism test. There was one run per target placement, sampled short replies, and no breakdown of computation, paging or RPC time.
  • Model quality or other giant models' speed. Nothing here ranks answers or predicts an unverified trillion-parameter file from its name.
10

Sources and artifacts

The evidence ships as a data package: recorded responses and log excerpts, fresh header-reader output, and the arithmetic that connects them. If a number on this page disagrees with a file in its package, the file is right and the page is wrong.

Source excerpts document the Inkling architecture and expert-offload pattern at the retained build revision 946fc11d1. Anchor build observations record c8e03ce and 5f55650 as well. Header hashes identify the bytes read, not the unread weight payloads. No model weights are distributed here.