Graphometer
Field card · measured 14 to 21 September 2026; updated 26 September · one RTX 5090 desktop · llama.cpp on an unmerged branch · 128K window and a 262,144-token full-window sweep · text only

Inkling-Small at 12 tokens a second, from a 32-token prompt to 95,041

Across the September 14 direct probes, Inkling-Small spoke at 11.6 to 12.2 tokens a second on one RTX 5090 desktop. Batch tuning later raised 48,115-token reading from 123.0 to 337.9 t/s; 95K was not remeasured.

Independent measurements. Not affiliated with, endorsed by, or connected to Thinking Machines Lab, Unsloth, NVIDIA, Intel, ASUS, Hugging Face, or the llama.cpp project.

This is one quantization of one model on one machine, measured on 14 and 15 September 2026, one request at a time. The runtime used an open, unmerged draft branch; its 15 September status snapshot is retained below. With every expert in system memory, reply speed stayed flat across the September 14 direct probes from 32 to 95,041 prompt tokens while reading speed rose with prompt size. At the default batch settings on 14 September, a 95,041-token history read at 185.5 tokens a second, and the planted code came back exactly measured.

Update, 26 September 2026

Reading is faster with -b 4096 -ub 2048. On 20 September, the same 48,115-token prompt read at 337.9 t/s, against 123.0 t/s at default batch settings measured. The older direct-probe speaking band still stands as a dated result. We have no tuned reading measurement at 95,041 tokens to replace its 185.5 t/s.

01

What it is

The maker describes Inkling-Small as a 276-billion-parameter mixture-of-experts model with 12 billion parameters active for each token vendor. Its repository says it accepts text, images, and audio and returns text vendor. We served text only.

FileInkling-Small, Unsloth UD-Q3_K_XL, 4 shards, 119,554,379,840 bytes measured
Shape42 blocks: 2 dense and 40 mixture-of-experts. The header reports 256 experts, 6 used and 2 shared measured
Attention32 query heads, 8 key-value heads, keys and values 128 wide, a 4,096-wide embedding, and a 512-token sliding window on 5 of every 6 blocks measured
Full-cache layers7 of 42 blocks carry the full attention cache measured
Served hereThe 14 September probes served 131,072 tokens. Later runs served 262,144. F16 cache, one request at a time, text only measured
PublishedOn Hugging Face on 27 July 2026, the date the public repository was created, with its last update on 31 July 2026 vendor
LicenceApache 2.0 by the repository's metadata at revision 8cc5877b. No separate licence file ships in the repository tree, so we did not read one in full vendor

The file, shape, attention and cache-layer rows were read from the file header and the chat template on 14 September 2026. Served windows come from the dated server logs below. Both repository facts come from the maker's repository metadata at that revision, read on 15 September 2026, not from a licence document we retained.

02

What we ran, exactly

The GGUF was checked byte for byte against the file tree of the public unsloth/Inkling-Small-GGUF repository before the run measured. The runtime was llama.cpp pull request 25731 at commit 946fc11d1, build 10897, based on df750f76b, with CUDA 12.8.93 for the RTX 5090 measured. PR #25731 was open, marked draft, and not merged when read on 15 September 2026 vendor.

CardNVIDIA GeForce RTX 5090, 32,086 MiB as llama.cpp reports it measured
MachineIntel Core Ultra 9 285K, 24 threads, and 188 GiB of system memory measured
Every expert in system memoryThe mixture-of-experts weights stay in system memory rather than on the card, so the card holds attention, the dense blocks and the cache measured
The September 14 configurationThe launch settings kept on the desktop for everyday use: every expert in system memory, a 131,072-token window, f16 cache, the template's default effort measured

The branch notes set two boundaries on how it may be driven: one request at a time, and flash attention on vendor. They also set one hardware limit: the branch's new attention operation is available only on CPUs and NVIDIA GPUs vendor. Every measurement here stayed inside all three. The single-machine run kept every expert in system memory. A second run put 5 expert blocks on the graphics card. The two-machine run placed routed experts on the second machine while attention stayed on the card measured.

Scope

The 14 September probes used one Unsloth UD-Q3_K_XL file, f16 attention cache, a 131,072-token served window, one request at a time, and one RTX 5090 desktop. The original 14 and 15 September direct-probe programs set temperature 1.0; temperature zero applies to the 20 September batch requests and the 21 September replay. Section 03 gives the later window and the separate 15 September full-window sweep, whose requests also set temperature zero measured.

03

Reading update, then the original probes

The batch comparison used Unsloth UD-Q3_K_XL on the same desktop, every expert block in system RAM (--n-cpu-moe 42 on a model with 40 expert blocks), f16 key and value caches, flash attention on, 24 processing and batch threads, one request at a time, and temperature zero. The synthetic ledger asked for a sealed code with a 900-token output cap. The new logs do not print a build commit, so the September 14 commit is not assigned to this run.

Date / settingServed windowPrompt tokensReadingCard at load
20 September, default b 2048 / ub 512131,07248,115123.0 t/s12,641 MiB
20 September, b 4096 / ub 2048131,07248,115337.9 t/s13,183 MiB
20 September, direct binary, b 4096 / ub 2048262,1443,033263.6 t/s17,281 MiB
21 September, start script, ubatch not overridden262,1443,033187.9 t/s16,935 MiB at health; 17,273 MiB peak

measured At the 262,144-token window, the 20 September direct binary launch read at 263.6 t/s and generated 104 tokens at 12.48 t/s. The 21 September start-script replay read at 187.9 t/s (187.88 in the log) and generated 104 tokens at 10.59 t/s. Both used 3,033-token prompts; we do not know why the script replay was slower. These are not 48K or 95K reading measurements at that window. The paired 48K card cost is 542 MiB, computed as 13,183 minus 12,641 arithmetic.

-b 8192 -ub 8192 failed mid-read at the smaller 131,072-token window: the log reports a CUDA allocation failure and the request lost its connection measured failure, 20 September. On 21 September the start script, whose defaults that afternoon were -b 4096 -ub 2048, served a 262,144-token window and finished 40 of 40 public-text prompts without a crash measured. The harness record says ubatch: skip and does not print the batch flags; the flags are inferred from the script defaults. Prompt lengths were 352 to 5,885 tokens, drawn for a 512-token physical batch measured configuration. This does not show that a 48K or 262K prompt is stable at those flags.

The 15 September full-window sweep

The same UD-Q3_K_XL file ran at a 262,144-token served window with --n-cpu-moe 42, f16 cache, no explicit batch flags, one request at a time and temperature-zero requests. Build 10897 / 946fc11d1 loaded at 15,744 MiB and peaked at 15,814 MiB on the card measured, 15 September. It read 230,827 tokens in 2,005,919.14 ms at 115.07 t/s measured, or 2,005.9 seconds rounded arithmetic, milliseconds divided by 1,000. It returned 3 of 3 codes planted at approximately 5 / 50 / 95 percent of the ledger's entries, in a 27-token answer at 8.78 t/s; its two warm, roughly 200-token replies ran at 11.69 to 11.77 t/s measured. The earlier flat speaking band applies to the September 14 direct probes through 95,041 tokens, not this deeper request.

The September 14 direct probes

Across our direct probes speaking stayed between 11.6 and 12.2 tokens a second from a 32-token prompt to a 95,041-token prompt measured. Reading looks different. Small prompts do not fill a batch, so they read slowly and unevenly: prompts of 32 to 75 tokens read at 23 to 34 tokens a second across six probes; from about 2,000 tokens the rate climbs with prompt size.

PromptProbeReadingSpeakingRecall
32 tokensshort prose, effort none29.0 t/s11.91 t/snot tested
32 tokensshort prose, repeat33.8 t/s12.09 t/snot tested
34 tokensshort prose, thinking on32.1 t/s11.84 t/snot tested
73 tokenstool call, effort none22.9 t/s12.09 t/snot tested
75 tokenstool call, thinking on26.3 t/s12.00 t/snot tested
75 tokenstool call, template default27.1 t/s12.19 t/snot tested
1,950 tokensplanted code118.7 t/s11.66 t/sexact
7,641 tokensplanted code158.2 t/s11.57 t/sexact
30,441 tokensplanted code182.1 t/s11.63 t/sexact
95,041 tokensplanted code185.5 t/s11.62 t/sexact

measured All ten direct probes on the single-machine, every-expert-in-system-memory placement. Recall used one planted code at each tested depth. This is not a multi-needle suite.

The 95,041-token read took about 8.5 minutes arithmetic, recorded milliseconds divided by 60,000. The first run became ready to serve (healthy) in 40 seconds from a cold page cache measured. The card held 12,416 MiB after load and 12,425 MiB after the probes, leaving about 19,600 MiB of its 32,086 MiB unused arithmetic difference. About 3,584 MiB of that is the attention cache for the 7 blocks that carry it at a 131,072-token window, plus about 70 MiB for the 35 blocks on a 512-token sliding window, computed from the file header arithmetic. The window pattern is why the cache is small.

Through the agent framework, on the same server, decode fell from about 13.5 tokens a second at a 3,300-token prompt to 12.8 at 52,925 tokens and 10.7 at 102,473 tokens measured, so the flatness we measured is a property of these direct probes, not of every request shape.

Two faster placements

Putting 5 expert blocks on the card raised speaking speed to 13.0 to 13.7 tokens a second. That run held 26,870 MiB after load and 26,954 MiB after the probes, and became ready to serve in 6 seconds with the page cache already warm measured.

The 14 September launch kept every expert in system memory. Its smoke run at 23:22 became ready to serve in 14 seconds and spoke at 13.3 to 14.2 tokens a second, and that run's own note recorded 12.2 GB on the card measured. A second launch of the same configuration later that evening recorded 12,430 MiB at its health check, and the gate record in this package reports 12,410 MiB before and after its own run measured. Three readings, two launches, and each figure keeps the unit its record used. The page cache state differed from the first run, so the page does not treat the speed gap as a placement result.

04

The effort dial is prompt text

The chat template maps seven names to six values: none 0.0, minimal 0.1, low 0.2, medium 0.7, high 0.9, xhigh and max 0.99. It also accepts any number from 0.0 to 0.99. The default is 0.9, the value the name high carries measured. The template emits the choice as a system sentence. It is guidance to the model, not a spending limit.

Live direct requests honoured both none and medium. Medium produced 961 characters of reasoning in a separate channel and then a clean answer measured. In the recorded framework turns the first visible content arrived after 39.05 s, 11.1 s and 35.49 s measured. Most of that was prompt processing: the 39.05-second turn spent 38.4 seconds reading a 3,328-token prompt, and the installed turn that took 35.49 seconds produced no reasoning events at all measured. At these prompt sizes the default effort is not what the reader is waiting for.

The related reasoning-effort study shows why a prompt sentence should not be mistaken for a hard budget.

05

Tools and recall through a framework

Three direct probes produced a well-formed weather tool call: with thinking off, at medium effort, and at the template's high default measured.

The 14 September launch then passed every short leg of a tool-and-recall check through an agent framework. It emitted a direct tool call in 5.6 seconds, wrote and read back a test file, accepted a 118,000-token working window, and completed a turn in which the framework swapped the active model and back measured. An earlier run of the check, made under a temporary registration, failed one step unrelated to the model: a roster check rejected the temporary entry. That record stands as a failure, even though its model legs completed.

Seeded contextPromptFirst contentReadingRecall
48,00052,925 tokens490.72 s108 t/s3 of 3
96,000102,473 tokens922.55 s111 t/s3 of 3

measured Earlier framework run, template-default effort. A tool was also used at depth in both seeded legs. The two reading rates are the server's own prompt-eval timings for those requests, 108.05 and 111.31 tokens a second. The direct 95,041-token probe read at 185.5 tokens a second, so the framework path was visibly slower. These are different request shapes, not an estimate of framework overhead alone.

06

What we got wrong: the two-machine worker

The second machine is an ASUS ROG Flow Z13 with 128 GB of unified memory, joined to the desktop by a direct Thunderbolt cable. With a worker built from the same branch commit, the split placement returned coherent text, recalled the code, and formed the tool call. It spoke at 7.0 to 7.7 tokens a second and read at 48 to 52, against about 12 speaking and 119 to 158 reading on one machine at the comparable tested depths measured. For a file that already fits one machine, the second machine bought nothing. Its use is capacity, for a file that does not fit, and never speed.

The first split attempt used a worker built from a different source commit. It returned repeated, unusable output. The draft branch adds one operation, moving the operation count from 101 to 102 and the protocol patch level from 0 to 1 measured. The worker reported that it supported every operation, so nothing caught the mismatch and it ran the wrong operation ids. The practical rule is exact: build both sides from the same source commit.

07

The 975-billion sibling

The larger sibling is Inkling, the 975-billion-parameter model in the same family, with 41 billion parameters active for each token vendor. We recorded it on 15 September 2026, between 00:13 and 00:30. It used a 270 GB UD-IQ1_S file with 66 blocks, served at 131,072 tokens measured. Alone on the desktop, with a file larger than the machine's 188 GiB of system memory so that part of it pages from an NVMe drive, it became ready to serve in 96 seconds, used 22,887 MiB on the card, spoke at 3.4 to 5.0 tokens a second, and read a 1,950-token prompt at 20.05 tokens a second measured.

On two machines it became ready to serve in 272 seconds, used 22,829 MiB on the card and 95 GB on the second machine, spoke at 3.13 to 3.35 tokens a second, and read at 21.38 to 26.40 tokens a second across the 1,950-token and 7,641-token probes measured. The 26.40 reading figure belongs to the pair. It does not belong to the single-machine run. The pair did not improve speaking speed.

08

What this does not show

  • No stable runtime claim. The original measurements used a draft branch; PR #25731 was still an unmerged draft when checked on 26 September 2026. Its behaviour and flags may change.
  • No concurrent serving. The branch required one request at a time, and every run stayed at one request.
  • No image or audio result. The maker says the model accepts both vendor. We tested text only.
  • The larger-window run has a measured depth. Doubling the window doubles the cache the 7 full-cache blocks need: 7 blocks, 8 key-value heads, keys and values 128 wide, two of them, at two bytes each, over 262,144 tokens is about 7,168 MiB, which would put the card near 16,000 MiB of its 32,086 arithmetic. That was the original memory estimate. The 15 September full-window sweep read 230,827 tokens at a 262,144-token served window and measured 15,744 MiB loaded and 15,814 MiB peak measured. The later tuned-window checks used only 3,033 prompt tokens. The estimate is not their measured card use.
  • No broad recall claim. Recall was exact for one planted code at each of four tested depths, at prompts of 1,950, 7,641, 30,441 and 95,041 tokens measured. That is narrower than a multi-needle evaluation.
  • No speed claim beyond these files and placements. The small and large siblings used different quantizations, shapes, and placements.
09

Related pages on this site

Reasoning effort is a sentence, not a dial examines the same mechanism class on a different model. The model field-card index places this result beside other bodies measured on the same bench. The method page defines the evidence labels used here.

10

Sources and artifacts

The evidence for this page is the run record package, mapped file by file. data/README.md maps the records, states their provenance and lists what was removed from them. data/NUMBERS.md maps every figure to a file and field.

  • data/runs/MEASUREMENTS.md holds the measurement tables for every direct probe, placement, branch build, and sibling result used here.
  • data/logs/ holds the server, probe, tune and smoke logs the figures were read from, and data/probes/ the programs that produced them, prompts included.
  • data/model/chat_template.txt carries the file header dump and the complete effort mapping.
  • data/gates/installed-check.json carries the installed short-leg pass. data/gates/earlier-check.json carries the earlier failed verdict and the 48,000 and 96,000-token legs measured.
  • The public model repository was read at revision 8cc5877b and llama.cpp pull request 25731 was read on 15 September 2026 vendor. The repository was created on 27 July 2026 and last updated on 31 July 2026, and its metadata carries the Apache 2.0 licence; no licence file ships in its tree. data/repository/ keeps the maker's model card at that revision, the pull request's recorded state, and the public file listings and byte totals for both files.

If a number on this page disagrees with a file in its package, the file is right and the page is wrong.