Describing pictures locally
On the card, Qwen3-VL 30B-A3B found all 39 known strings at 1.3 to 3.2 seconds a warm picture. On the CPU, its first 4K picture took 273 seconds at Ollama's default 4 threads, 111 at 16, and 38.5 at 16 with a 1,536-pixel cap.
A picture describer is a vision model that writes down what a picture shows, so that a person or a text-only model can use it. On 20 September 2026 we timed six Qwen vision models doing that job through Ollama 0.30.10 on one desktop, with the describer settings we use: a plain description of what is visible, including any text in the picture, a reply of at most 2,048 tokens, and 150 seconds per picture. The test set is seven pictures holding 39 strings we knew were there, one run per configuration. On the RTX 5090, Qwen3-VL 30B-A3B, Qwen3-VL 8B and Qwen3.6 35B-A3B found all 39, at 1.3 to 3.8 seconds a picture once loaded. The CPU is where a describer can go when a large model holds the card, and there most of the wait is the model reading the picture. With no thread count in the request, the server Ollama started reported 4 threads out of 24, and Qwen3-VL 30B-A3B took 273 seconds on the first 4K picture, 243 of them reading it. At 16 threads the same picture took 111 seconds; at 16 threads with every picture capped at 1,536 pixels on its longest side, 38.5, and that run found all 39 strings. With 26,396 MiB of the card already in use by another program, Ollama's default request put 11 percent of the model on the card and took 207 seconds on that first picture, past the 150-second limit, while a CPU-only request with 16 threads, an 8,192-token window and the cap took 39.4 and added nothing to the card; its launch also used a larger batch and turned memory mapping off (section 06). The cap has a cost the 39 strings did not show: on a 4K screenshot of small text, the 30B, on the card, misread 4 of 24 codes at 1,536 pixels and none uncapped, and every miss was a wrong code, not a gap. Section 08 shows what else a string count cannot see. One desktop, one run per configuration, one Ollama version: this measures speed and string recall, not the quality of a description.
What it is
A vision model reads a picture as tokens, much as it reads words. On this desktop a 3840 x 2160 picture came to 4,154 prompt tokens with the Qwen3-VL models, against 1,974 for a 1200 x 1600 poster. Reading those tokens is prompt processing; writing the description is generation. On the graphics card both are quick. On the CPU, reading a large picture is most of the time, and two settings decide how long it takes: how many CPU threads the server uses, and how big the picture is when it arrives.
The card is fast, but the models here took 4.3 to 23 GB of memory as Ollama reported it (section 02), and a card that already holds a large model may not have that to spare. The CPU path uses system memory instead and leaves the card alone. This page measures both, with the settings that decide the CPU's speed, and checks what the fast settings cost.
What we ran, exactly
Every figure on this page is measured on this machine on 20 September 2026 unless it is marked arithmetic. The screenshot's counts in section 07 were made on 26 September from the replies recorded on 20 September, and the logo file's pixels were read on 26 September (section 08).
| Machine | One desktop: an NVIDIA GeForce RTX 5090 (32 GB card), an Intel Core Ultra 9 285K (24 cores: 8 performance and 16 efficiency, no hyperthreading, as The wrong flag records), 188 GiB of RAM. Runs from 08:19 to 09:16 local time, one at a time. |
|---|---|
| Server | Ollama 0.30.10: the version the serving process logged when it started on 19 September. The same process served every request here; our harness did not record the version itself. For each model Ollama started a llama-server process (llama.cpp's server, as Ollama ships it) with a 32,768-token window unless the request set one, and one request at a time. Every Qwen3-VL launch also asked for at least 1,024 image tokens per picture (--image-min-tokens 1024); the Qwen3.5 4B, Qwen3.5 9B and Qwen3.6 launches did not carry that flag. The log does not give that server's build. Its launch lines are in the package. |
| The request | The describer prompt we use, sent with each picture. In our words: it asks for a description limited to what is plainly visible (what the picture shows, how it is laid out, its colours, its mood and style, and any text that can be read), with no guessing at things that are absent, and it allows a question to be answered afterwards. We sent no question. Options: num_predict 2048 (the longest reply allowed), keep_alive 60 seconds, no sampling options, so each model used its own defaults. Qwen3.5 and Qwen3.6 were sent think: false and returned no reasoning text. Where a table says so, the request also set num_gpu 0 (CPU only), num_thread or num_ctx. |
| The limit | The describer settings we use allow 150 seconds per picture. The harness recorded whether each picture finished inside it; it did not cut any request short. |
| Time | “Total” is the harness's wall clock around one request, from sending the picture to the end of the reply. Load, reading and writing are Ollama's own durations from its reply (load_duration, prompt_eval_duration, eval_duration); reading includes turning the picture into tokens. |
| First picture and warm | Before each run the harness told Ollama to unload the model, so the first picture of every run includes starting the server and loading the model; the pictures after it are “warm”. The first picture is always the 3840 x 2160 nebula photograph (or, in section 07, the screenshot). The operating system's file cache was not recorded: the first card load of the 30B that morning took 36.97 seconds, and its later card loads 4.56 to 4.86. |
| Scoring | score.py (in the package) lowercases each reply, drops £, $ and commas, collapses spaces, and checks whether each expected string appears in it. The 39 strings are 34 pieces of text drawn into the pictures and 5 plain subject words: nebula, stars, earth, house, sun. |
| Overlaps | Downloads of four other models ran from 08:35 to 08:40, during the Qwen3.6 and Qwen3-VL 4B card runs and the first two minutes of the 4B's full-size CPU run. Two runs against hosted services finished during the default-thread CPU run. |
The models
Ollama's library tags, each a Q4_K_M file (the load log's “Q4_K - Medium”). “Memory” is what ollama ps reported with the model on the card at the 32,768-token window.
| Ollama tag | Model | Parameters | Memory |
|---|---|---|---|
| qwen3-vl:30b-a3b-instruct | Qwen3-VL 30B-A3B, mixture of experts; the log calls its model type 30B.A3B (about 3B active per token by that name, not a figure from this page) | 30.53 B | 22 GB |
| qwen3.6:35b | Qwen3.6 35B-A3B, mixture of experts, thinking off | 35.51 B | 23 GB |
| qwen3-vl:8b-instruct | Qwen3-VL 8B | 8.19 B | 10 GB |
| qwen3-vl:4b-instruct | Qwen3-VL 4B | 4.02 B | 7.9 GB |
| qwen3.5:9b | Qwen3.5 9B, thinking off | 9.20 B | 6.7 GB |
| qwen3.5:4b | Qwen3.5 4B, thinking off | 4.33 B | 4.3 GB |
The pictures
| Picture | What it is | Pixels | Strings |
|---|---|---|---|
| 1 nebula | Hubble's Orion Nebula photograph (credit in section 12) | 3840 x 2160 | 2 |
| 2 earth | Earth photographed from the International Space Station (credit in section 12) | 3840 x 2160 | 1 |
| 3 poster | A made-up event poster: title, date, time, hall, address, price, phone number, contact name, printer, job number | 1200 x 1600 | 10 |
| 4 invoice | A made-up bicycle-repair invoice with small print: five line items, total, terms | 1240 x 1754 | 13 |
| 5 chart | A bar chart, rainfall for January to June with each value printed on its bar | 1400 x 900 | 9 |
| 6 cottage | A drawn cottage with sun, grass and caption, stored on its side with a rotation tag that says to turn it upright for display | 1200 x 900 | 3 |
| 7 logo | The COSMIC wordmark (a mark of System76): dark gray letters with a coloured bar under the O, on an opaque light gray field, with thin coloured curves in two corners. A third-party mark, described here and not shipped. Our notes called it a white logo on a transparent background; its file has no transparent pixels (section 08) | 3840 x 2160 | 1 |
| 8 screenshot | Section 07 only: a made-up 4K screenshot, 24 codes in text 28, 20, 16 and 13 pixels high | 3840 x 2160 | 24 codes |
Pictures 3 to 6 and 8 were drawn by our build scripts, which are in the package with the strings each picture is expected to yield.
Observed: six models on the card
All seven pictures at full size, each model on the card at the 32,768-token window. The server reported 4 threads for each of these runs; the request set none.
| Model | Strings | First picture, s (load) | Warm pictures, s | Card, MiB before → after the first picture |
|---|---|---|---|---|
| Qwen3-VL 30B-A3B | 39 / 39 | 41.69 (36.97) | 1.31 to 3.16 | 1,077 → 24,755 |
| Qwen3.6 35B-A3B | 39 / 39 | 39.60 (36.11) | 1.82 to 3.47 | 1,011 → 25,764 |
| Qwen3-VL 8B | 39 / 39 | 8.85 (5.16) | 1.81 to 3.82 | 24,208 → 13,348 |
| Qwen3-VL 4B | 38 / 39 | 8.74 (3.86) | 1.15 to 2.61 | 1,001 → 10,656 |
| Qwen3.5 9B | 37 / 39 | 9.12 (4.73) | 2.14 to 4.49 | 1,139 → 9,537 |
| Qwen3.5 4B | 30 / 39 | 6.63 (4.19) | 3.02 to 9.35 | 946 → 6,732 |
Card memory is nvidia-smi's reading in MiB, desktop included.
The 8B run started with the 30B still loaded from the run before it; Ollama evicted the 30B
first (its log, 08:57:36), so that 24,208 is not an empty card. Downloads ran during the
Qwen3.6 and 4B runs (section 02).
Three models read everything. Qwen3-VL 30B-A3B, Qwen3.6 35B-A3B and Qwen3-VL 8B found all 39 strings, with warm pictures between 1.31 and 3.82 seconds. The first picture is where the model load shows: 36.97 and 36.11 seconds for the two large models, loaded for the first time that morning, and 3.86 to 5.16 for the others.
The misses say different things. Qwen3-VL 4B's one miss was “house” on the sideways cottage, which it called “an arrow pointing leftward”. Qwen3.5 9B's two misses were correct readings in other words: it wrote “£2/day” and “12 months” where the invoice says “£2 per day” and “twelve-month”. Qwen3.5 4B missed four poster details (the address, the price, the phone number and the contact name) and five of the six chart values, and three of its replies ran to 1,562, 2,048 and 2,048 tokens (section 08).
A smaller window saves card memory. A picture plus the
longest reply our settings allow fits in 8,192 tokens (4,158 + 2,048 =
6,206; arithmetic, with the largest prompt recorded here). At
num_ctx 8192, ollama ps reported 19 GB for
the 30B instead of 22, and 4.2 GB for the 4B instead of 7.9. The
card read 22,562 MiB with the 30B loaded (1,150 before) and
7,343 with the 4B (1,148 before), on runs of two pictures (the nebula and
the invoice, 4 threads) that found all 15 of their strings. A later 30B run at 8,192 with 16 threads found all
39 strings at 1.23 to 2.95 seconds a warm picture.
For reference, two hosted services. On the same
morning the same seven pictures went, with the describer settings we use,
to two hosted services. Google's found 39 of 39. Mistral AI's found 22 of
39 with one model (2 of the invoice's 13) and 39 of 39 with Mistral
Medium. Every hosted reply ends with a line that repeats the picture's
file name, and we count the description only: score.py,
which reads that line too, gives the first Mistral run 23, its extra hit
the word “earth” in the file name, not in the description.
The result files name the provider, not the model version, and the
Medium name is the run's own label. We publish only these counts, not
the replies.
On the CPU, the thread count sets the wait
Qwen3-VL 30B-A3B with num_gpu 0, the 32,768-token
window, full-size pictures, one run per thread count. The first picture
is the 3840 x 2160 nebula (4,154 prompt tokens), including the
model load; the second is the poster (1,974 tokens).
| Threads, as the server reported | Nebula: total, s | Load | Reading | Writing | Poster: total, s | Reading | Writing |
|---|---|---|---|---|---|---|---|
| 4 of 24 (request set none) | 273.17 | 12.60 | 242.93 | 174 tokens, 17.36 s | not run | ||
| 8 of 24 | 157.31 | 12.62 | 123.53 | 303 tokens, 20.89 s | 60.30 | 40.14 | 364 tokens, 19.74 s |
| 16 of 24 | 111.01 | 13.13 | 87.77 | 193 tokens, 9.85 s | 41.87 | 29.57 | 278 tokens, 11.90 s |
| 24 of 24 | 107.88 | 14.56 | 76.57 | 208 tokens, 16.47 s | 45.77 | 26.55 | 235 tokens, 18.82 s |
Amber: over the 150-second limit. The 4-thread run was stopped after its first picture, so it left no reply text and no string count; its figures are from the harness's console log and the server's own timing lines. The 8, 16 and 24-thread runs found all 12 strings in their two pictures.
Where the 4 comes from. When the request set no thread
count, Ollama's launch line for the server carried no thread flag, and the
server reported n_threads = 4 (n_threads_batch = 4) / 24.
With num_thread 8, 16 or 24 the launch line carried
-t with that number and the server reported it back. Four
is also what llama.cpp's own default chose on this CPU in
The wrong flag, where we traced it through
llama.cpp's code; we did not read the code of the server Ollama ships.
The card runs that set no thread count were launched with the same
4.
Reading and writing peak in different places. Reading the nebula ran at 17.1, 33.6, 47.3 and 54.3 tokens a second with 4, 8, 16 and 24 threads (arithmetic, tokens over reading time), so it kept improving all the way to 24. Writing did not: 10.0, 14.5, 19.6 and 12.6 tokens a second on the nebula's reply, and 18.4, 23.4 and 12.5 on the poster's with 8, 16 and 24 (arithmetic). Over a whole picture, 24 threads were 3.1 seconds faster than 16 on the nebula, and 16 was 3.9 seconds faster on the poster (arithmetic). This CPU has 8 performance and 16 efficiency cores; we did not record which cores the threads ran on.
The smaller model reads faster, but not enough. Qwen3-VL 4B at 16 threads, full-size pictures: 95.69 seconds on the nebula (71.48 reading), 23.07 to 80.80 seconds on the warm pictures (the two slowest, 80.80 and 80.66, were the other 4K pictures), and 38 of 39 strings. Model downloads were running during its first two minutes.
On the CPU, the picture size sets the rest
The capped sets were made by one script (in the package): it applies each picture's rotation tag, then shrinks any picture whose longest side is over the cap. Prompt tokens as Ollama counted them for the Qwen3-VL models, instruction included:
| Picture | Full size | Capped at 1,536 | Capped at 1,024 |
|---|---|---|---|
| nebula, earth, logo (3840 x 2160) | 4,154 | 1,370 | 1,106 |
| poster (1200 x 1600) | 1,974 | 1,802 | 1,110 |
| invoice (1240 x 1754) | 2,219 | 1,706 | 1,127 |
| chart (1400 x 900) | 1,306 | 1,306 (under the cap) | 1,114 |
| cottage (1200 x 900 stored) | 1,138 | 1,138 (turned upright, not shrunk) | 1,110 |
At the 1,024 cap every picture came to 1,106 to 1,127 tokens; that run's
Qwen3-VL 4B launch carried --image-min-tokens 1024, which suggests a floor there (our
reading of the flag). The
Qwen3.5 and Qwen3.6 models counted 4 more tokens per picture.
Then the same CPU settings as section 04 at 16 threads, 32,768-token window, all seven pictures unless stated:
| Model | Cap | Nebula: total, s (load, reading) | Warm pictures, s | Strings |
|---|---|---|---|---|
| Qwen3-VL 30B-A3B | none | 111.01 (13.13, 87.77) | 41.87 (poster only) | 12 / 12, two pictures |
| Qwen3-VL 30B-A3B | 1,536 | 38.52 (12.30, 18.17) | 22.64 to 38.93 | 39 / 39 |
| Qwen3-VL 4B | none | 95.69 (5.50, 71.48) | 23.07 to 80.80 | 38 / 39 |
| Qwen3-VL 4B | 1,536 | 29.11 (6.73, 14.30) | 19.83 to 39.36 | 39 / 39 |
| Qwen3-VL 4B | 1,024 | 27.20 (5.24, 10.95) | 18.12 to 32.64 | 39 / 39 |
The cap took the 4K picture's reading from 87.77 seconds to
18.17 on the 30B, and the whole first picture from 111.01 to
38.52. With the cap, the 30B's warm pictures took 22.64 to 38.93
seconds, close to the 4B's 19.83 to 39.36 at the same cap (174.27 against
172.17 seconds for the six warm pictures together; arithmetic). The log
calls the 30B's model type 30B.A3B, a mixture of experts that by that name
uses about 3B parameters per token (the name's figure, not a measurement
here); on this CPU it cost little more time than the 4B while needing 22 GB of memory to the 4B's 8.0, as
ollama ps reported them.
The 4B's one miss. Its only full-size miss was “house”, on the cottage. At the 1,536 cap the cottage was turned upright and not shrunk (1,138 prompt tokens both ways), and that reply contains “house”. With one run each, that is all the evidence.
On the card the cap barely matters. There the 30B read the 4K screenshot in 1.85 seconds uncapped and 0.49 at the 1,536 cap (section 07).
With the card mostly full
Four runs started while another program already held 26,390 to
26,405 MiB of the card (nvidia-smi, desktop included);
Ollama's log put the card's free memory at 5.1 GiB. What that
program was, and whether it was working during these runs, is not in
our records. Two pictures each: the nebula first, including the load,
then the poster.
| Model | Request | Where it ran (ollama ps) | Threads | Window | Batch | Cap | Nebula: total, s (reading) | Poster: total, s (reading) | Card, MiB before → after load |
|---|---|---|---|---|---|---|---|---|---|
| Qwen3-VL 30B-A3B | nothing set | 89% CPU, 11% card | 4 | 32,768 | 512 | none | 206.87 (180.68) | 74.91 (54.52) | 26,396 → 29,453 |
| Qwen3-VL 30B-A3B | num_thread 16, num_ctx 8192 | 88% CPU, 12% card | 16 | 8,192 | 512 | none | 85.45 (65.20) | 31.02 (21.29) | 26,390 → 29,408 |
| Qwen3-VL 30B-A3B | num_gpu 0, num_thread 16, num_ctx 8192 | 100% CPU | 16 | 8,192 | 1,024 | 1,536 | 39.39 (18.52) | 42.68 (26.14) | 26,405 → 26,391 |
| Qwen3-VL 4B | num_thread 16, num_ctx 8192 | 52% CPU, 48% card | 16 | 8,192 | 512 | none | 56.27 (45.20) | 18.56 (13.19) | 26,391 → 29,816 |
Threads as the server reported them; batch is the launch line's -b and
-ub. The CPU-only launch also carried --no-mmap (memory mapping off); the
other three did not. All four runs found all 12 strings in their two pictures. The capped run's pictures were smaller: 1,370 and 1,802 prompt tokens, against 4,154 and
1,974 for the others.
The default request was the slowest thing we measured with a
full card. Ollama fitted part of the model into the free memory,
kept the picture encoder off the card (its launch line said
--no-mmproj-offload, and the encoder's log line says it ran
on the CPU), and ran everything left on the CPU with 4 threads. The
first picture took 206.87 seconds.
Setting 16 threads and an 8,192-token window, two changes at once, brought a similar split (88 percent on the CPU) to 85.45 seconds on the nebula and 31.02 on the poster.
CPU only, with 16 threads, the 8,192 window and the cap, the nebula took 39.39 seconds, on a picture capped to a third of the tokens (1,370 against 4,154; arithmetic), and the poster 42.68, slower than the split's 31.02. Against the default request its launch also changed the batch (1,024 against 512) and turned memory mapping off, so the fall from 206.87 to 39.39 seconds is all of those changes together, not the threads and the cap alone. It is the only one of the four that took no card memory: the split runs raised the card's reading by 3,057, 3,018 and 3,425 MiB (arithmetic, after load minus before), memory the program already on the card could no longer use.
The CPU path took about as long with the card full. The capped CPU settings with the card empty (section 05; the two launches differ only in the window, 32,768 against 8,192) took 38.52 seconds on the nebula and 38.93 on the poster, against 39.39 and 42.68 here, one run each.
What the 1,536-pixel cap costs on small text
The seven pictures have fairly large text, so they did not show what
shrinking does to small print. The screenshot does: 24 codes such as
F12-zephyr-6823, six in each of four text sizes, on a
3840 x 2160 page. Qwen3-VL 30B-A3B on the card, 4 threads,
32,768-token window, one run per cap, each the first picture after a
load (8.56 to 9.88 seconds including a 4.56 to 4.84-second load).
score.py has no expected strings for this picture, so we counted the
codes on 26 September with its matching rule (score_hires.py
in the package).
| Cap | Prompt tokens | 28 px | 20 px | 16 px | 13 px | Codes found |
|---|---|---|---|---|---|---|
| none (3840) | 4,154 | 6 | 6 | 6 | 6 | 24 / 24 |
| 2,560 | 3,674 | 6 | 6 | 6 | 6 | 24 / 24 |
| 2,048 | 2,378 | 6 | 6 | 6 | 5 | 23 / 24 |
| 1,536 | 1,370 | 5 | 6 | 6 | 3 | 20 / 24 |
Every missing code was replaced by a wrong one. At
2,048 the model wrote F12-zephyr-6829 for
F12-zephyr-6823. At 1,536 it wrote
D14-kettle-B104 for D14-kettle-8104 (28 px
text), and, in the 13 px lines, F12-pebble-4823,
E28-ribbon-7319 and H40-pebble-8359 for
F12-zephyr-6823, E26-ribbon-7519 and
H20-pebble-8359. A reader of those replies would have no
way to tell the four wrong codes from the twenty right ones. At the
1,536 cap, text drawn 13 pixels high becomes 5.2 pixels high, and at
2,048 about 6.9 (arithmetic).
This ran on the card. On the CPU the model receives the same pixels, so we expect the same kind of misreading there, but we did not measure it (transferred).
What 39 strings show, and what they do not
The count shows whether a model copied specific text and named obvious subjects, and it separated the models: three found everything, and Qwen3.5 4B dropped nine strings. It is also a narrow instrument. We read the cottage replies of all eleven seven-picture runs, the chart replies of five, three logo replies, and Qwen3.5 4B's long replies; this is what the count could not see.
- A picture on its side. The cottage is stored lying on its side with a tag that says to turn it upright. All six models on the card described it as stored: a green band down the right-hand side and the caption running vertically, and the 4B called the house an arrow. Five of them scored 3 of 3. When the capped sets sent it upright, the 30B and the 4B described sky at the top and green ground at the bottom. As far as these replies show, the picture reached the models without its rotation applied.
- Wrong statements beside right numbers. On the card, the 30B listed all six chart values correctly and then wrote that “the bars increase in height from January to April”; February's 35 is lower than January's 42. The 8B called the bars “arranged in ascending order from left to right, except for the peak in April”. Both scored 9 of 9.
- Replies that run away. Qwen3.5 4B's replies to the Earth photograph, the poster and the chart ran to 2,048, 1,562 and 2,048 tokens. Within its first 1,400 characters each became one long unpunctuated stream, ending in lists of loosely related words (winemaking terms, music genres, military ranks). The Earth reply still scored 1 of 1.
- Right readings in other words. “£2/day” and “12 months” counted as misses (section 03), and so did a full-size CPU reply that called the house “a building or cottage”.
- Answers nobody asked for. Some replies, most of them the 4B's, wrote a question of their own after the description and answered it, though we sent none; the 30B did it in three replies.
- Wrong text that looks right. A misread code (section 07) scores as a miss, but in a reply it reads as a fact.
- Our own notes, wrong. Our notes and the file's name call picture 7 a white logo on a transparent background. Its pixels, read on 26 September, are all opaque: a light gray field, dark gray letters with a coloured bar under the O, and thin coloured curves in the top-right and bottom-left corners. The card replies we read (the 30B's, the 8B's and Qwen3.6's) describe that picture: a dark wordmark, a TM mark (which we did not check in the pixels) and coloured curves in those corners. “COSMIC” matched either way, so the count could not have told which account was right, and the set has no transparency test.
We did not grade the replies for invented details beyond these examples. All the local replies are in the package, with 13 prompt passages and one word removed (section 12).
What to run
From one desktop's measurements; your card, CPU and pictures will move every number. Each line says what it rests on.
- If the card has room, use it. Qwen3-VL 30B-A3B found all 39 strings at 1.31 to 3.16 seconds a warm picture in 22 GB, as Ollama reported it (19 GB at an 8,192-token window). Qwen3-VL 8B found all 39 at 1.81 to 3.82 seconds in 10 GB. Qwen3.5 4B did worst here: 30 of 39, and three runaway replies.
- If a large model holds the card, go to the CPU on
purpose. Set
num_gpu0. Left to itself, Ollama fitted part of the model into the card's last free memory, using 3,018 to 3,425 MiB of it in our runs (arithmetic), and still ran the picture encoder and most of the model on the CPU (section 06). - Set the thread count. With no
num_thread, the server here ran 4 threads of 24. On this CPU, 16 and 24 threads were within 4 seconds of each other on both pictures we timed (arithmetic; 24 read faster, 16 wrote faster), 16 leaves more of the CPU for everything else, and 8 missed the 150-second limit on a 4K picture. Check your own server: on every load, Ollama 0.30.10's log printedsystem_info: n_threads = …(on a standard Linux install,journalctl -u ollama | grep n_threads). - Shrink large pictures before sending them, and turn them upright first. A 1,536-pixel cap cut a 4K picture from 4,154 tokens to 1,370, and the 30B's first picture from 111.01 seconds to 38.52 at 16 threads, with all 39 strings found. Apply the rotation tag yourself: as far as the replies show, nothing on the way did.
- Keep more pixels for small text. The same cap misread 4 of 24 codes on a 4K screenshot, where 2,560 found all 24 (one run each, on the card). For screenshots and fine print, cap at 2,560 or send them to the card.
- Set
num_ctx8192 for single pictures. A 4K picture and the longest reply we allow need 6,206 tokens (arithmetic), and at the smaller windowollama psput the 30B at 19 GB on the card instead of 22. Raise it if you send several pictures or a long question at once. - Budget the wait. With 16 threads and the 1,536 cap,
the 30B took 22.64 to 38.93 seconds a warm picture on this CPU, plus
12.30 seconds of model load on the first, at the 32,768-token window;
ollama psput its memory at 22 GB there and 19 GB at 8,192. The 4B at a 1,024 cap took 18.12 to 32.64 seconds in 8.0 GB and found all 39 strings in our set, with the small-text risk of section 07.
What this does not show
- One run per configuration. No repeats, so we cannot say how much a second run would move any figure; a difference of a few seconds or one string may be noise.
- A small test. Seven pictures and 39 strings, plus one screenshot with 24 codes. Not a benchmark, and not a grade of description quality: section 08 lists what the count misses, and we did not score invented details.
- One machine and one CPU. The thread results belong to a Core Ultra 9 285K with 8 performance and 16 efficiency cores. The best thread count on another CPU may differ.
- One Ollama version, one quantization. Ollama 0.30.10, whose bundled server build is not in its log, and each model's Q4_K_M library file. What the server chose when a request said nothing (the 32,768-token window, 4 threads, the split on a full card) may differ in other versions and setups.
- First pictures include loads. The file cache was not recorded, so how much of each first-picture time was disk is unknown; the first card load of the 30B took 36.97 seconds and the later ones 4.56 to 4.86.
- The full card was a test of memory, not of shared work. We do not know what held the card or whether it was busy. A large model whose experts sit in system RAM also uses the CPU, and a describer on the CPU beside it would compete for threads and memory bandwidth; we did not measure that.
- Few pictures per CPU and full-card run. The thread runs and the full-card runs used two pictures each, the 4-thread run one.
- Overlaps. Model downloads ran during three runs, and two hosted-service runs during the 4-thread CPU run (section 02).
- The capped sets changed two things. They shrank the large pictures and turned the cottage upright, so the string counts at a cap mix both effects; for the cottage only the rotation changed (section 05).
- No transparency test. Picture 7 was meant to be one; its file has no transparent pixels (section 08).
- Sampling. The request set no sampling options, so each reply is one sample at the model's own defaults.
- The hosted counts. The result files record the provider, not the model version; they are counts only.
Related pages on this site
The wrong flag traced a 4-thread default on this CPU to llama.cpp's thread heuristic, on a 675B model whose decode it roughly halved. llama.cpp's batch size measures the reading settings of the models this desktop serves with llama.cpp; Ollama launched the describers here with batches of 1,024 or 512 tokens, and we did not vary them. Slower by default is another default measured on this desktop, and the method page explains the labels and the data packages.
Sources and artifacts
The evidence for this page ships as a data package: the files the runs produced, not a summary. If a number on this page disagrees with a file in the package, the file is right and the page is wrong; tell us and we will fix the page.
- data/README.md: every file, where the files disagree with each other, and every redaction.
- data/NUMBERS.md: every figure on this page, with its file and field.
- The runs (
data/results/): one result file per run with every reply, timing and card reading, and the console log of the 4-thread run. Where a small model repeated our private prompt word for word, 13 passages in five files read[prompt wording removed], and one word in one reply reads[one word removed]; data/README.md lists both. - The server's log
(
data/service-log/): Ollama's version line and, for these runs, each launch command, thread report, file type, memory and timing line. - The pictures and the scripts: the test pictures
except the third-party logo, the scripts that drew and shrank them, the
harness with the prompt wording removed,
score.py,score_hires.py, andtables.py, which rebuilds the derived tables indata/derived/. - Picture credits. Nebula: NASA, ESA, M. Robberto (Space Telescope Science Institute/ESA) and the Hubble Space Telescope Orion Treasury Project Team (ESA/Hubble image heic0601a, licensed CC BY 4.0). The JPEG we ship is a 3840 x 2160 derivative taken from a desktop wallpaper set, not the release file. Earth: photograph ISS064-E-29444, image courtesy of the Earth Science and Remote Sensing Unit, NASA Johnson Space Center. Their use here does not imply endorsement by NASA or ESA.
| 2026-09-20 | All runs, 08:19 to 09:16 local time: six models on the card, the thread ladder, the capped CPU runs, the screenshot at four caps, the 8,192-token window, and the four full-card runs. |
|---|---|
| 2026-09-26 | Scores reproduced from the recorded replies; the screenshot's codes counted; the server's log read for the version, launch lines and thread reports; the logo file's pixels read; picture credits checked against the ESA/Hubble and NASA image pages. |