Graphometer
Family guide · measured rows dated · vendor rows labeled

Which Qwen should I run?

Qwen is an unusually wide open model family: real Apache-licensed weights from under one billion parameters to a 397B mixture-of-experts, plus API-only frontier tiers above that. That width is the point of this page. We run the family for real on one desktop, from the small daily drivers to the 122B and 397B, so the guidance here is what we measured, labeled as measured, next to what the vendor claims, labeled as claims. Find your hardware in the chooser, install your first one with two commands once Ollama is on your machine, take the serving profiles we actually run, and know when the honest answer is "use the API instead."

Independent guidance. Not affiliated with, endorsed by, or connected to the Qwen team or Alibaba Group.
01

The family on one screen

Four generations matter today. The one-line version: 3.5 and 3.6 are the open ones you can run; 3.7 is API-only; 3.8 is now open too, in two very different sizes and under two different licenses.

Qwen3.5 (open)The current big open generation: mixture-of-experts models topping out at the 397B and the 122B we serve here, Apache 2.0, with built-in speculative-decoding heads. Our measured deep-dive is the Qwen3.5 pair field card.
Qwen3.6 (open)The newest open weights: a 27B and a 35B mixture-of-experts, Apache 2.0, files verified on the official organization. Our daily drivers below come from this generation and its predecessors.
Qwen3.7 (API only)Max, Plus and Flash tiers on the official API and aggregators; the vendor states they are proprietary. No 3.7 weights have ever shipped.
Qwen3.8 (open, in two sizes)The Max is live as an API, and both promised open releases have now shipped: a 2.4-trillion-parameter Max-class model on 2026-08-12 under a custom license with revenue thresholds, and a 27B on 2026-08-14 under Apache 2.0. Two models, one version number, two different licenses, so check the file in the repository you actually download. Verified status, revisions and the fake-repository warning live on our Qwen3.8 release watch; the 27B's measured record, taken on the day it shipped, is the Qwen3.8-27B field card.

Generation status checked 2026-08-09; the Qwen3.8 row refreshed 2026-08-14. The release-watch page carries the dated table and updates as things actually ship.

02

The chooser: what your machine can actually run

Rows marked measured come from recorded runs on our machine (a single RTX 5090 with 32 GB of VRAM and 188 GiB of system RAM; dates given). Rows marked arithmetic are file-size reasoning we consider sound but have not run. Rows marked vendor are the maker's sizing. Speeds are warm generation, one configuration each, not benchmarks.

Your hardwareWhat we would runSpeed seenBasis
Laptop or desktop, no real GPUThe sub-1B and 2B-class tags as utilities and pipeline testers; anything bigger belongs to the APIsee basistransferred: the 0.6B measured 857 t/s on our 32 GB GPU (July 2026); CPU-only speed untested here, but this size class runs acceptably on CPUs generally (vendor and community consensus, labeled as such)
8 to 12 GB VRAMThe 4B and 9B-class tags at 4-bituntested herevendor sizing; we have not run these sizes and say so
16 GB VRAMA 14B-class at 4-bit comfortably; the 27B at 4-bit with partial CPU offloaduntested herearithmetic from file sizes
24 GB VRAMThe Qwen3.6 27B at 4-bit (a 17 GB download)see basisarithmetic; measured only on our 32 GB card (72 tokens per second, July 2026), where it never had to squeeze
32 GB VRAMThe sweet spot we live on: the 35B mixture-of-experts as the fast generalist, the 27B, and a 30B-class vision model, all fully on GPU227 / 72 / 291 t/smeasured, July 2026, Ollama serving (qwen3.6:35b / qwen3.6:27b / the 30B-class vision tag, in that order)
32 GB VRAM + roughly 96 GB or more RAMThe Qwen3.5 122B: every layer on GPU, expert blocks walked to CPU, gated speculation on26 to 38 t/smeasured, 2026-08-05/06, llama.cpp serving on a 188 GiB machine; full story on the field card
32 GB VRAM + roughly 176 GB or more RAMThe Qwen3.5 397B, same technique, more patience14 to 20 t/smeasured, 2026-08-05/06, same 188 GiB machine; low end is thinking replies
None of the above fitsHosted Qwen: the same open weights served by API providers, or the official frontier tiersn/asection 06

The RAM rule behind those floors: at the quantizations we run, the 122B's weights are a 73 GB file and the 397B's are 150 GB, and the technique copies weights into RAM, so the weights plus operating-system headroom must fit. We measured both models only on a 188 GiB machine. With less RAM than the weights, the model streams from disk instead, which we did not measure here and which is dramatically slower. Quantization tier changes everything; ours are named on the field card.

Behind that table are three questions, and they are worth asking in order, because a yes to one does not imply a yes to the next. Does it load? Does it generate at a speed you can work at? Does it still do that with enough context to be useful for your actual task? Most sizing advice answers only the first and lets the reader assume the rest. The 397B on this machine loads at the placements we use, generates comfortably, and has never been fed a prompt longer than 4,367 tokens: three different answers about one model.

Total parameters and active parameters are different numbers Every name in this family carries two: the 122B is a 122.1 billion-parameter model of which 9.8 billion compute for any given token, and the 397B is 396.3 billion of which 17.4 billion do. Those are header sums, read out of the files on our disk on 2026-08-08 by adding up tensor shapes (899 tensor headers in one file and 1,118 in the other) rather than figures quoted from a model card, and they are the served text stacks: each file also carries a built-in guessing head, 2.5 and 6.6 billion parameters, which the field card counts separately. The architecture is why: the smaller model holds 256 routed experts and fires 8 of them per token plus one always-on shared expert; the larger holds 512 and fires 10, plus its shared expert. Under five percent of the big model works on any one token. The total decides what you must store; with the weights in memory, the active count is what decides how fast a token is. That is the single most useful thing to know before buying hardware, and getting it backwards is the most common sizing mistake we see: an "A17B" label does not mean the model needs 17 billion parameters' worth of memory. At the quantizations we run, what has to fit is 73 GB for the 122B and 150 GB for the 397B, and our method copies those bytes into system RAM rather than reading them from disk. The active count is why a 397-billion-parameter model answers at conversational speed at all.

One more distinction the chooser above compresses. Loads, generates, worth studying, and worth recommending are four different findings, and a page that reports only the first is telling you a quarter of the story. Here is where we actually got to on each model we have run, and where we deliberately stopped. "Not established" means we did not do the work, not that the answer is no.

Model, as we ran itLoadsGeneratesWorth studyingWorth recommending
Qwen3.5-122B, 4-bit, 38 expert blocks on CPU Yes: 24.6 GiB of a 32 GB card, 20 to 24 seconds warm Yes: 26 to 29 t/s prose, about 38 structured Yes: a full recorded battery, 2026-08-06 Yes, with an output-token floor in the thousands: our recorded 700-token caps came back empty
Qwen3.5-397B, 3-bit tier, 57 expert blocks on CPU Yes: 23.5 GiB of the same card, 160 to 215 seconds Yes: 19.5 t/s structured, 14 to 17 on real thinking replies Yes: the same recorded battery Yes for depth, if you can wait; the hard puzzle we ran failed at a 4,096-token cap and completed in 7,608 tokens under an 8,192-token cap
The same 397B, one placement step further onto the GPU (53 blocks on CPU) No: every weight loaded, then it died minutes in, unable to allocate a 1,862 MiB working buffer n/a n/a No, and this is why "it fits" is not the same question as "it loads"
Qwen3.6-27B, 4-bit, fully on GPU (Ollama, July 2026) Yes Yes: 72 t/s Not established: a battery of ours ran on it, then the thinking-budget trap invalidated most of the card it produced Yes, as a first Qwen (section 03)
Qwen3.6-35B mixture-of-experts, fully on GPU (Ollama, July 2026) Yes Yes: 227 t/s Not established Yes, as the fast generalist at this card size
A 397B-class model on a machine with less RAM than its weights Not attempted, predicted yes: the weights can be read from disk instead of copied into memory Predicted no at a usable speed: the profile copies 150 GB of weights into RAM, and with less RAM than that the model streams from the drive on every token. The nearest thing we have measured is a different 744B model on this machine whose weights exceed its RAM: 4 to 6 tokens per second Not established No, on that arithmetic, and we say predicted because we did not run it

Rows one to three are recorded runs on this machine, 2026-08-05 to 08-06. Rows four and five are July 2026 Ollama-era measurements at a different serving stack and carry that date. The last row is arithmetic and is labeled as a prediction throughout; the 744B comparison it leans on is a recorded run of a different model family, quoted for the memory regime only.

03

Your first Qwen, without being an infrastructure person

The short path is Ollama. One hardware gate first: the model below is a 17 GB download that wants roughly 20 GB or more of VRAM to run fully on GPU. If your card is smaller, pick a smaller tag from the chooser above, or read section 06. Then install Ollama from the official installer for your platform, and:

ollama pull qwen3.6:27b
ollama run qwen3.6:27b

Two commands, and that is genuinely it for a first conversation. Two things will save you the hours they cost us:

First, the empty-answer trap. Modern Qwens think before answering, and the thinking spends the same output budget as the visible reply. If your client caps output around 500 or 1,000 tokens, a hard question can consume the whole cap thinking and hand back nothing, which looks exactly like a broken install. It is not broken. Start with the output limit at 4,096: 2,000 to 4,096 is a starting range that worked for our jobs, not a universal safe number, because the floor depends on the task. In one recorded pair of measurements, the same hosted model answered one prompt from a 256-token budget and had produced nothing visible at 12,000 tokens on another the same day, a spread of more than 46 times (Empty Answer carries the method for measuring your own floor). We measured this same behavior all the way up at the API frontier (the release watch has the table).

Second, sampling. For thinking-family Qwens use temperature 0.6, top-k 20, top-p 0.95. The values are in the first profile below so you do not have to remember them.

You graduate from Ollama to llama.cpp's llama-server when you want a model bigger than your VRAM: the technique that puts a 122B on a 32 GB card is a llama.cpp flag set, and the second profile below is its shape.

04

Known-good profiles, downloadable

Three small plain-text files, the settings we actually run, commented so you know why each value is what it is.

  • ollama-qwen-thinking.txt: the parameter set for thinking-family Qwens under Ollama, including the output-budget floor that avoids the empty-answer trap.
  • llamacpp-qwen35-moe.txt: the llama-server shape for big Qwen3.5 mixture-of-experts models on one GPU, including the expert-offload knob and the speculation confidence gate our field card measured (ungated speculation ran slower than none on creative text, measured on the 122B; we keep the same gate on the 397B on that evidence; the gate is one flag, off by default in the builds we used). The gate now has its own measured study on the 397B itself, two recorded runs with stated intervals and a downloadable evidence package: Slower by default.
  • opencode-local-qwen.json: a provider block that points the OpenCode coding agent at your local Qwen, either serving route.

The profiles encode our recorded configurations generalized to standard ports and paths. They are starting points, not guarantees; your quantization, context size, and card will move the numbers.

05

The big ones, on one desktop

The reason this page can speak about 122B and 397B models in the first person: both serve on the single 32 GB card described above, using expert offload and the models' built-in speculation heads with the confidence gate set. The 122B runs at 26 to 38 tokens per second depending on prompt shape, the 397B at 14 to 20, and the complete measured story, artifact hashes to placement arithmetic to the gate trap, is the Qwen3.5 pair field card. The Qwen3.8 weights have since landed, and the 27B needed none of this technique: it is dense, it fits the card whole, and the Qwen3.8-27B field card is the measured record from the day it shipped. The Max-class half is a different problem and the release watch page says where it stands.

06

Local or hosted, honestly

Hosted Qwen is startlingly cheap. As of 2026-08-09, aggregator pricing put the same 397B we run locally at roughly fifty to sixty cents per million input tokens and about three and a half dollars per million output; a heavy personal month costs single-digit dollars. The official API's flash tier is cheaper still, with a million-token context. Local never wins that arithmetic.

What local buys instead: your prompts and documents never leave the machine; it works with the network down; nobody deprecates, throttles, or silently revises your model; and you learn more about how these systems actually behave than any API will teach you. Our rule of thumb: privacy-sensitive or high-volume-and-predictable work runs local; frontier-quality one-offs and anything needing a million-token context run hosted. A same-task local-versus-hosted study on identical weights is in progress and will get its own page.

Prices observed 2026-08-09 and they change; check the provider's live page before planning around them.

07

What this page does not know

  • We have not run the 4B-to-14B-class sizes on small GPUs; those rows are sizing arithmetic and vendor claims, labeled as such.
  • Our speed figures are one machine, one quantization tier, one serving configuration each, at the dates given. They are honest samples, not benchmarks.
  • The July 2026 daily-driver numbers predate our current llama.cpp serving work; a fresh measured pass on the 27B and 35B is planned and this page will update with dated entries.
  • Vision, audio, and long-context behavior are not covered here; the field card states exactly what long-context ground the big pair has actually been tested on.
08

Updates and sources

2026-08-09Page created. Generation status, pricing, and tag names checked that day; chooser speeds carry their own dates.
2026-08-11The speculation confidence gate was measured on the 397B itself, in two recorded runs with stated intervals; the full study and its data package live on their own page.
2026-08-14Both Qwen3.8 open releases have now shipped, so the generation row in section 01 is refreshed: the Max-class weights on 2026-08-12 under a custom license with revenue thresholds, and a 27B on 2026-08-14 under Apache 2.0, measured here the same evening on the card described in section 02. Its record is the Qwen3.8-27B field card. Nothing else on this page was re-checked on that date.

If today is later than the newest row above, treat the generation-status and pricing lines as unverified since that date; the release watch carries the living Qwen3.8 record.

  • The Qwen3.5 pair field card: dated recorded runs behind every 122B/397B number here.
  • The Qwen3.8 release watch: verified generation status and the API measurements.
  • July 2026 daily-driver speeds: our recorded Ollama benchmark pass (project records that can be produced if questioned).
  • Official Qwen organization pages for licenses and file sizes, checked 2026-08-09; aggregator pricing pages for section 06, observed 2026-08-09.