Release watch: Qwen3.8
On 2026-08-03 the Qwen team announced Qwen3.8 and promised open weights for a Max-class model alongside a new 27B, both due in the week beginning 2026-08-10. On 2026-08-12 the larger half of that promise landed. The Max-class open weights are real, and we verified them: two repositories under the official Qwen organization, checked at pinned revisions on 2026-08-13, 213 weight shards each, 4,892.4 GB in BF16 and 2,496.2 GB in FP8. On 2026-08-14 the other half landed too: the 27B named in the same announcement shipped, we verified it at its own pinned revision, and it was downloaded, hash-checked, served and measured on one desktop the same evening. The fake repositories that grew around the name before the release are still there after it. For practitioners the headline is the license, and it is not one license: the Max-class weights carry a custom license with revenue thresholds, read at the pinned revision and quoted in section 02, while the 27B carries a genuine Apache 2.0 file, read in full and retained. This page holds the verified state of the release, what the shipped configuration says, which community quantizations are real, an honest account of what a desktop can do with files this size, and our measurements of the hosted model. It updates with dated entries and it does not speculate: where a fact is a vendor's promise rather than a shipped artifact, it is labeled as one.
What exists, verified
Each row carries exactly one status label. Shipped and verified means we saw the actual weight files at a named revision with a license. API only means the model answers requests but no weights exist. Announced means a vendor promise with no artifact. Community only means real files exist from a named third party but nothing official. Not official means artifacts exist under the name but none come from the maker. A repository name is never proof of a release.
| Model | Status | Evidence, with check date |
|---|---|---|
| Qwen3.8-2.4T-A95B | Shipped and verified | On the official Qwen organization. Its own timestamps say created 2026-08-08 and last modified 10:24 UTC on 2026-08-12, the day an aggregator listed the model at 16:21 UTC. Verified 2026-08-13 at revision 207bd685a7e3696cfaff12ded7c6a7ea0f88c996: 213 weight shards plus an index file, 4,892.4 GB in BF16, license file present (section 02). The 95-billion-active-parameter figure that was press-only before the release now appears in the official repository name itself. |
| Qwen3.8-2.4T-A95B-FP8 | Shipped and verified | Verified 2026-08-13 at revision d2dc35658bcf77e66643428cb52e774cc3b5bd29: 213 shards, 2,496.2 GB, fp8 dynamic quantization with 128 by 128 blocks, license file byte-identical to the BF16 repository. The organization's own Qwen3.8 collection contains exactly these two repositories and nothing else. |
| GGUF conversions of the 2.4T-A95B | Community only | One real set exists, from the community quantizer unsloth, verified 2026-08-13: 312 GGUF files at revision 567d3e6ac26c, sized from 397.3 GB to 4,893.2 GB (section 04). No official GGUF exists. On release day we downloaded the smallest tier and verified it bit-perfect against the repository's own checksums, and no public runtime can load it yet: its expert tensors use a quantization type that exists in no public build (sections 04 and 07). This row is about the Max-class weights only; the 27B's conversions are a separate story, in the row below. |
| Qwen3.8-27B | Shipped and verified | Published at about 11:00 EDT on 2026-08-14 and verified the same day at revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0: 18 safetensors plus an index, 32 files, and a LICENSE file we read in full and retained. That license is Apache 2.0, 11,544 bytes of verbatim Apache text, byte-identical in the official FP8 repository at revision 017b9c7af6b5689d5dd426a76e0bc077eb5ca20a. It is not the Max-class license of section 02, and the two must never be quoted as one. Still no official GGUF: the conversions are community work, including one from the llama.cpp project's own organization published the same day. We downloaded a community conversion, verified both files against the repository's own checksums, served it and measured it that evening; the numbers are in the Qwen3.8-27B field card. |
| Qwen3.8-Max | API only | Live since 2026-08-03 on Alibaba Cloud Model Studio and on aggregators. We measured it (section 06). Per the released model card, checked 2026-08-13, vision input, the non-thinking mode, and the million-token default context are Max API extras: none of them ship in the open weights (section 03). |
| Third-party "Qwen3.8" repos | Not official | Re-screened 2026-08-13: every third-party 27B repository we had previously catalogued under the name still contains zero weight tensors, and a fresh wave of mirrors and placeholders using the Max name appeared around the release. Section 05. |
| Qwen3.7 family | API only | Max, Plus and Flash tiers serve on the official API and aggregators (checked 2026-08-09). The vendor states they are proprietary; no 3.7 weights have ever shipped. |
| Qwen3.6-27B and 35B | Shipped and verified | The newest open Qwen weights until 2026-08-12, and still the newest at sizes a normal computer can hold: file trees confirmed on the official organization, Apache 2.0 license files present (checked 2026-08-09). |
An earlier version of this page carried a single-source report that the coming weights would have geographic license restrictions, held as rumor until a license file existed. The license file now exists, and we read it at the pinned revision on 2026-08-13: its conditions are the two described in section 02, and a geographic restriction is not among them. The note stays so the rumor's resolution is on the record.
The Max-class license: not Apache 2.0
Qwen3.6, the previous open release, shipped under Apache 2.0.
The Qwen3.8 Max-class weights do not. The file in both of their repositories is named the
Qwen3.8-Max License: 3,390 bytes, byte-identical
in the BF16 and FP8 repositories, read on 2026-08-13 at revision
207bd685a7e3696cfaff12ded7c6a7ea0f88c996.
The body is an MIT-style permission grant with two conditions
attached:
- A name-display condition above consumer scale. Products or services built on the model that exceed 100 million monthly active users, or twenty million US dollars in monthly revenue, must display the model name prominently. Note the unit on the second test: monthly revenue, not annual.
- A separate-license condition above fifty million. Businesses offering the model as a service, or offering an "AI Work Assistant" product, whose aggregate revenue exceeds fifty million US dollars in any consecutive twelve months, must obtain a separate license from Qwen before commercial use. Internal use is exempt.
In plain English, as we read it, the grant works like MIT for individuals, researchers and ordinary businesses, very large consumer products owe a visible name credit, and a company past fifty million dollars in any rolling twelve months that sells model access or an AI work assistant must get a separate license from Qwen before commercial use; read the license yourself, this is not legal advice.
The clause also settles a pre-release question. Reporting sourced to Reuters on 2026-08-07 said to expect a revenue-share arrangement in this license, with the rate still under negotiation. What shipped resolves that expectation as thresholds instead: as we read the file, no percentage of anyone's revenue appears in it, a name-display duty starts at consumer scale, and a separate-license duty starts at fifty million dollars for the two named business shapes. The license file itself does not say what a separate license costs above the threshold.
For local research use of the kind this site does, none of this appears to bite: downloading, running, modifying and publishing measurements sit inside the grant as we read it, far below every threshold. Same caveat: our reading, not advice.
Qwen3.8-2.4T-A95B and its FP8
twin. The 27B that shipped on 2026-08-14 carries a genuine
Apache 2.0 file: 11,544 bytes of verbatim Apache
License 2.0 text with an appendix reading
Copyright 2026 Alibaba Cloud, no monthly-active-user
threshold, no revenue-share clause, no separate-license requirement at
any revenue level, and byte-identical in the official FP8 repository.
We read it in full at revision
1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0 on
2026-08-14 and retained the file. Two models, one version
number, two licenses. "Qwen3.8 is Apache 2.0" is inaccurate as
a family statement and true of the 27B; the reverse mistake, reading
this section's thresholds onto the 27B, is just as wrong. Section 03
of the Qwen3.8-27B field card puts both
readings side by side with the retained files.
The Max-class weights: one model, always thinking
Everything here comes from the Qwen3.8-2.4T-A95B
repository files at the pinned revision, mostly the config,
tokenizer and chat template, plus the model card, checked
2026-08-13. These are shipped facts, as distinct from announcement
promises. None of this row set describes the 27B,
which is a different architecture with a working thinking
off-switch and a vision tower in its configuration; its own reading
is section 01 of the
Qwen3.8-27B field card.
| Release shape | One post-trained model. No base variant and no separate instruct variant exist: the repository is the assistant model, and the organization's Qwen3.8 collection carries nothing else. |
|---|---|
| Thinking | Mandatory. The thinking stage cannot be disabled, and the chat template raises an error if asked to run with thinking off. Control is a reasoning_effort setting with levels xhigh, medium and low; the default is xhigh, and thinking content is preserved across turns by default. |
| Modality | Text only. Vision input, the non-thinking mode and the million-token default context are hosted Qwen3.8-Max extras per the model card; none of them ship in these weights. |
| Context length | 262,144 tokens native in the config. The model card calls it extensible to 1,010,000; the seven-figure number is a vendor claim we have not tested. |
| Architecture | Qwen3_5MoeForCausalLM, model type qwen3_5_moe_text: a scaled Qwen3.5, not a new architecture. 92 layers arranged as 23 repeats of three GatedDeltaNet linear-attention layers plus one gated-attention layer, hidden size 8192, 512 experts with 10 routed and one shared active per token, and a one-layer multi-token-prediction head. |
| Tokenizer and template | Vocabulary 248,320 and end-of-sequence id 248044, both unchanged from the Qwen3.5 397B we already serve. The chat template drops the vision branches and adds the reasoning_effort control. |
The architecture line is the practical
finding. Because the model declares an existing architecture
rather than a new one, the runtime wait that normally follows a
release should not apply: Qwen3.5 support took a merge, a revert
and a re-merge across a week, with the vendor's own pull request,
and none of that is needed if nothing new has to be merged. The
evidence, all checked 2026-08-13: llama.cpp has carried the
qwen35moe architecture enumeration since the 3.5
generation, and the build serving our Qwen3.5 pair, from upstream
commit c8e03ce, already contains it; the real
community GGUF conversion declares exactly that architecture and
the native 262,144 context in its own metadata (section 04); vLLM
merged a change on 2026-08-07 that routes Qwen3.8 checkpoints
through its existing Qwen3.5 path rather than adding a new model
class; and no new-architecture pull request is open or needed on
llama.cpp. The one open 3.8-adjacent llama.cpp pull request is
chat-template plumbing for the reasoning_effort
setting, and it does not block loading: the default effort applies
without it.
We state this as an expectation with evidence, not as a tested result: existing Qwen3.5 llama.cpp and vLLM paths should load these weights unchanged. The one load attempt here so far does not test it either way: the community file we tried was refused for its quantization type at the header parse, before any architecture code ran (section 07). For weights in mainline formats the expectation stands, still untested here.
Community quantizations of the Max-class weights, and what fits at home
This section sizes the Qwen3.8-2.4T-A95B conversions.
The 27B's conversions are ordinary desktop-sized files and are
covered on its own field card, which
names the one it served and its checksum.
The organization shipped safetensors only; no official GGUF exists, checked 2026-08-13. One community conversion is real: the quantizer account unsloth published unsloth/Qwen3.8-2.4T-A95B-GGUF, 312 GGUF files at revision 567d3e6ac26c, last pushed 01:01 UTC on 2026-08-13, with architecture and context metadata matching the official config (section 03). File-manifest sizes at that revision:
| Quantization | On-disk size |
|---|---|
| UD-Q1_0 | 397.3 GB |
| UD-IQ1_S | 508.4 GB |
| UD-IQ1_M | 564.0 GB |
| UD-IQ2_XXS | 656.6 GB |
| UD-IQ2_XS | 730.7 GB |
| UD-IQ3_XXS | 955.5 GB |
| UD-IQ4_XS | 1,310.9 GB |
| Q8_0 | 2,600.2 GB |
| BF16 | 4,893.2 GB |
Sizes as listed in the repository manifest on 2026-08-13. One row has since been through our hands: the UD-Q1_0 set was downloaded here on 2026-08-13, every shard matched the repository's own SHA-256 checksums, and it could not be run (section 07). The rest of the table remains listing-verified only, neither downloaded nor run here. The other quantizer accounts people usually watch, bartowski and mradermacher, had published nothing under the name at check time, and one further well-known quantizer account carried only an empty placeholder.
The name-squat hazard, still live on release day
Within days of the announcement, and before any real weights, repositories using the Qwen3.8 name appeared on public model hubs under third-party accounts. The ones we examined fall into two shapes: placeholder cards that describe planned quantizations but contain no weight tensors at all, and rebadges, where a small model distilled from or unrelated to the announced ones circulates under the 3.8 name. Secondary articles have already cited some of these as if they were releases. The two are different hazards with the same conclusion: an empty card cannot be a release, and real tensor files under the name are still someone else's model, not the announced Max or 27B.
Release day did not retire this section; it sharpened it. Re-screened 2026-08-13: every third-party 27B repository we had previously catalogued under the name still contains zero weight tensors, including ones whose cards now advertise an Apache 2.0 license for a model that had not then shipped under any license. A fresh wave of mirrors and placeholders using the Max name appeared around the release itself.
The real 27B now exists, which changes the test rather
than removing it. As of 2026-08-14 there is a genuine
article to compare against:
Qwen/Qwen3.8-27B, on the maker's own organization, at
revision
1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0, with
an Apache 2.0 license file. Verify against that repository and that
revision. A third-party repository under the 27B name is still that
account's artifact, and now it is also checkable: if the weights are
not the maker's files at a revision you can name, it is not the
released model. The empty placeholders that carried the name before
the release did not become real when a real model appeared beside
them, and one of them advertising Apache 2.0 does not make it the
Apache 2.0 model.
The general rule is worth more than any list of today's fakes, because next week's will have different names. Three questions settle almost every case, and none of them requires downloading anything:
- Who owns the repository? A release comes from the maker's own organization account. A model published under any other account is that account's artifact, whatever the name says and however carefully the card is worded. This is a question about the owner, not about the quality of the page.
- Is there anything in it? Open the file list before the description. Weights are large and they come in shards: a 27B at four bits is tens of gigabytes, a 2.4-trillion-parameter model is hundreds. A repository whose entire contents are a readme and a small configuration file contains no model at all, and a card describing planned quantizations is a plan, not a release.
- Was this size ever announced? The 2026-08-03 announcement named exactly two open models: a 27B and the 2.4-trillion-parameter Max. A repository using the 3.8 name at any other size is, by arithmetic, not one of the announced models, whatever else it may be.
We are describing shapes rather than naming accounts on purpose. A rule generalizes to next week's uploads; a list of names is out of date the day after it is written, and it accuses people we have not contacted on evidence we have not asked them to answer.
When a drop happens, the safe path is boring: take the repository link from the Qwen team's own announcement channels, or the official organization page directly, confirm the license file is present, confirm the weight shards are actually there with sizes that make sense for the stated parameter count, and verify checksums after downloading. That day arrived on 2026-08-12, and the path above is exactly what we ran on 2026-08-13: the verified repositories, revisions and sizes are in section 01. We ran it again on 2026-08-14 for the 27B, in about five minutes, and the checksums matched.
First measurements of the API model
On 2026-08-09 we ran a fixed fifteen-call battery against Qwen3.8-Max through one aggregator route (OpenRouter), with every raw request and response body retained. This is one serving route on one afternoon, one to three calls per cell: diagnostics, not a benchmark. Total cost of the whole battery: seven cents, at the listed two dollars per million input tokens and six per million output.
The quick facts first. Asked what it is, the model answered "Qwen, also known as Tongyi Qianwen," built by Alibaba's Tongyi Lab. It solved the classic bat-and-ball trap correctly, produced a correct interval-merging function with three test assertions including the empty-list edge case, returned schema-conforming JSON on all three attempts, and passed our honesty probe cleanly: asked to summarize a file we never provided and report a score on a benchmark that does not exist, it said it had no such file and could not state the score, and asked for the data instead of inventing either.
The finding that will matter to anyone building on it: the model thinks by default, its thinking bills as output tokens inside your token cap, and it will happily spend your entire cap thinking and hand back an empty answer. We know the behavior well from serving two Qwen3.5 models locally (their field card is linked below), and it is fully present on this frontier API route. It is not a Qwen trait: a model from a different vendor, swept on the same prompt at the same caps on 2026-08-09, was slower than this one to produce a first visible answer (section 10, and the full study on Empty Answer). The Qwen battery:
| Request | Token cap | Thinking spent | Visible answer |
|---|---|---|---|
| Hourglass puzzle | 60 | 60 | empty |
| Hourglass puzzle | 500 | 500 | empty |
| Hourglass puzzle | 2,000 | 2,000 | empty |
| 300-word story, three attempts | 1,200 | 1,200 each | empty, all three |
| Structured JSON ask | 800 | 416 to 467 | complete |
| Hourglass, effort set low | 4,000 | 999 | complete and correct |
| Hourglass, effort set to the top level | 4,000 | 1,100 | complete and correct |
Token caps are the request's max-tokens value on this route, which covers thinking plus the visible answer. "Thinking spent" is the reasoning-token count the API reported back. Every empty row finished with the length stop reason: the cap ran out before the first visible word.
Read the middle rows carefully: a medium reasoning puzzle exhausted a 2,000-token cap without producing one visible word, and a simple story request burned 1,200 tokens of thinking three times in a row and never wrote the story. The practical rule we use locally holds here: give thinking Qwens thousands of tokens of headroom, or use the effort control, and treat an empty answer with a length stop as a budget failure, not a model refusal. You pay for the thinking either way: our three empty story attempts cost more than the completed effort-controlled runs that actually answered.
End-to-end latency on this route, for scale: about three to five seconds for a short factual answer, about twelve for the JSON asks, about twenty-five to forty-six seconds when thinking runs long. Wall-clock through an aggregator, labeled as such.
Measured on release day
The released model has a hosted twin: since 2026-08-12 at 16:21 UTC the same 2.4T-A95B has been listed on the aggregator route we measured in section 06, at the same listed prices as the Max endpoint, two dollars per million input tokens and six per million output. On 2026-08-13 we ran the day-one battery against that route, and the same day we downloaded the one community file set from section 04 that fits this machine and tried to run it. Both halves of release day are below. Every number comes from a recorded run with its raw bodies retained, one to three calls per cell, single observations labeled as such: diagnostics, not a benchmark.
One sentence of context first: a morning attempt never reached the model, because on its first day the open-weights listing was served only by third-party GPU clouds while this account's provider allowlist admitted model authors only, so all 26 calls returned the same not-found error in about a tenth of a second each, and the battery below ran in the afternoon, after we deliberately admitted two of those providers.
The hosted battery: thirty-five model calls, two providers, thirty-nine cents. The aggregator filled our calls from DeepInfra (20 of 35) and Together (15), switching between them mid-battery; DeepInfra lists its serving precision as fp4 and Together lists none, so every number below belongs to this endpoint mix on this day, with the provider recorded per call. The headline is continuity: the empty-answer behavior this site documents is fully present on the open weights. The marble prompt from the Empty Answer study, temperature zero, one call per cap:
| Token cap | Provider | Thinking spent | Visible answer |
|---|---|---|---|
| 60 | DeepInfra | 73 | empty |
| 256 | DeepInfra | 179 | first words, cut off |
| 500 | Together | 223 | complete and correct |
| 2,000 | DeepInfra | 188 | complete and correct |
One call per cell, 2026-08-13. "Thinking spent" is the reasoning-token count the server reported back; the 60 and 256 rows finished with the length stop reason, the 500 and 2,000 rows stopped on their own. The shape matches the hosted Max curve from 2026-08-09 at neighboring caps: empty at 60 where Max was empty at 64, first truncated words at a 256 cap on both, clean by 500 where Max was clean by 512.
The reasoning_effort control from section 03 then
got a same-prompt, same-cap reading: the marble prompt again, a
2,000-token cap, temperature zero, one call per level. On this one
trivial probe the knob did not change how much the model thinks.
What it demonstrably changed is the served prompt and the answer's
style:
| Effort level | Provider | Thinking spent | Prompt tokens, server count | Cost, server billed | Visible answer |
|---|---|---|---|---|---|
| xhigh (default) | DeepInfra | 184 | 112 | $0.00236 | 340 chars, terse, correct |
| medium | DeepInfra | 197 | 70 | $0.00363 | 720 chars, formatted, correct |
| low | Together | 168 | 100 | $0.00341 | 527 chars, formatted, correct |
One call per level, 2026-08-13; the low call landed on the other provider, a stated confound. The three requests were byte-identical apart from the effort field.
Read the columns against each other. Thinking spend stayed flat, 168 to 197 tokens across the three levels, and every answer was correct, so on this probe the knob neither bought nor saved any thinking. But the server counted three different prompt lengths for byte-identical requests, 112 tokens at the default, 70 at medium, 100 at low: the effort level is being written into the served prompt itself. That is consistent with the shipped chat template from section 03, which implements effort as inserted instruction sentences, one for xhigh, one for low, none for medium. And because lower effort produced longer, more formatted answers, the billed cost ran opposite to the knob's apparent intent: medium cost the most and the default cost the least. One trivial prompt, one call per level; a harder-prompt replication is the obvious next step, and until it runs we generalize nothing.
Thinking is mandatory in practice, not just in the template.
Section 03 said the shipped chat template raises an error if asked
to run with thinking off; the hosted route does not raise, it just
does not comply: a probe that sent enable_thinking:
false got a reply that still carried 626 tokens of hidden
reasoning, the field silently ignored. Single observation,
2026-08-13.
The listing is not the bill. Fine print, measured 2026-08-13: both providers that served us ran the native 262,144-token context from section 03, and the million-token context on the aggregator's listing exists only on a provider we did not admit, so we never touched it. Together bills above the listed price, two dollars fifty per million input tokens and six twenty-five per million output against the listed two and six, and the server-reported cost of every Together call matched that premium to the digit. The same provider billed 12,672 repeated identical prompt tokens at roughly a fifth of its input price on one probe, a cache discount the listing does not mention, which is why the session's billed total came in just under the sum of list prices. Check the per-provider terms behind an aggregator listing before budgeting: the listing showed neither the premium nor the discount.
The rest of the battery, one to three calls per cell: the long-context cells planted a code and got it back exactly, once at 31,380 prompt tokens and once at 69,441, both counts as the server reported them; tool calling passed every shot with exactly the right function arguments; and structured output conformed on all three attempts when the schema was written into the prompt. The one serving-layer gap: a request that supplied its JSON schema only through the API's response_format field starved to an empty answer, and the retained reasoning opens by complaining that no schema was provided, so on this route that field never reached the model. Put the schema in the prompt. These cells came from routecheck, the instrument that packages the fixed battery of section 08 into one recorded pass, plus a few one-call checks alongside it; routecheck also measured this route's thinking-budget floor at 500 tokens for its own probe family, where the marble prompt above showed first words at 256, which is the floor lesson in one line: it is prompt-dependent, so measure it on your prompts, not ours.
The local half of release day ended at the runtime layer, with zero seconds of model serving. We downloaded the one file set from section 04 that fits this machine's free disk, the 397.3 GB UD-Q1_0, at the pinned revision named there, and verified it bit-perfect on 2026-08-13: all ten shards matched the repository's own SHA-256 checksums, ten of ten, through a download that had to survive two interruptions. Then our server refused the file in 0.38 seconds, at the header parse, before allocating anything. The reason is not memory and not architecture: the 276 routed-expert tensors that hold most of the 2.4 trillion parameters are stored as quantization type number 66, and in every public build we checked that day the type table ends at 42, llama.cpp mainline and the quantizer's own public fork alike. The "Q1_0" in the file name is branding, not the mainline format of that name, and the file's own metadata label claims a standard type that appears nowhere among its tensors. The type is real; the code that reads it is, today, not public.
So the giant still has no local speed, memory or quality numbers here, and the reason is now sharper than not having run it: nothing public can load it. We will not invent what nothing could measure. The re-test is standing and cheap: the day a public build carries the new type family, the verified bytes on disk get their load attempt, and its numbers land here, dated. Until then the hosted route above is the only place the open 2.4T answers questions from this desk.
The absorption path, and where it stands
This machine already serves two Qwen3.5 mixture-of-experts models on one 32 GB graphics card, so the absorption path was rehearsed and written down before the release rather than after. Release day followed it. In order:
- Verify before downloading. Official repository, license file read, architecture string from the config, shard list with sizes. If the architecture is new to the runtimes, we say plainly that it does not run locally yet and track the support work instead of forcing it. Done 2026-08-13 for the Max-class release: sections 01 to 04 of this update are that step's output.
- Feasibility before the big one. The 27B-class model is a routine download. A Max-class open release is not: we compute storage, memory and rollback from the actual file listing first. We have run a model far larger than system memory before, streaming from disk, so "too big to be sensible" is a measurement question here, not a reflex. Loadable, able to generate, worth studying, and worth recommending are four different findings and we will label which one we reached. Done 2026-08-13 for the Max-class release: the arithmetic in section 04 picked the one tier that fits this machine, storage and rollback were computed from the actual file listing, and the download ran the same day (section 07).
- Checksums, then a first safe load at conservative settings, nothing else running. Half done 2026-08-13: checksums passed ten of ten; the first safe load was refused at the header parse by the file's quantization type, so the load stage waits on a public runtime that declares it (section 07).
- The same fixed battery we run on every model: identity and file accounting, load and memory placement, warm speed medians, the thinking-budget map from section 06 run locally with identical prompts, the reasoning toggle, tool calls, structured output, an honesty probe, and a long-context retrieval check. Same prompts as the API run above, so local-versus-hosted lands on day one with nothing invented. The hosted side of that battery ran 2026-08-13 (section 07); the local side waits on a loadable file.
- This page updates the same day, dated, with whatever we actually measured and nothing more.
The 27B half ran on 2026-08-14, end to end, the day it shipped. It was the routine download the plan predicted: weights public at about 11:00 EDT, on this desktop by evening, verified against the conversion repository's own checksums before first load, and served on a llama.cpp build this machine already had, because the file declares an architecture that build already implements. Every stage of the plan above completed for it, including the local half of the fixed battery that the Max-class release could not reach. The full measured record is its own page: the Qwen3.8-27B field card carries the seven-profile ladder and the exact invocation, the thinking off-switch and the budget floor, three measured client traps, the draft head that decodes prose at 72.4 tokens a second gated and 44.5 with the gate removed, and the license reading in full, with a data package holding every request and response body behind it. This page stays what it is: the dated record of what shipped and when. The card is the record of what it does.
Where this page comes from
Graphometer runs big Qwen models for real: the Qwen3.5 122B and 397B serve on one desktop here, measured in a field card whose numbers trace to dated recorded runs, including a speculative-decoding gate trap that the guides we used while serving those models do not call out. That working familiarity is the lens for this page: we are not previewing benchmarks, we are preparing to operate the new models and to publish what operating them is actually like.
Updates
| 2026-08-09 | Page created. Status verified against the official organization pages, the announcement coverage, aggregator APIs and runtime repositories; nothing shipped yet. API battery run and recorded (section 06). |
|---|---|
| 2026-08-10 | The empty-answer behavior in section 06 is not a Qwen fault, and we can now show that rather than assert it. On 2026-08-09 we swept one fixed reasoning prompt across the same token caps on two models from different vendors: Qwen3.8-Max on this aggregator route, and a comparable hidden-reasoning model from another family as a control. Both spend the whole cap thinking and return an empty answer when the cap is small. On this prompt Qwen was the better-behaved of the two: it produced its first visible words at a 256-token cap, where the control still returned nothing at 256 and did not finish until 1,024. This is a property of models that think invisibly and bill that thinking against one visible output budget, not a property of any one vendor. The whole comparison cost under three cents and it changed our conclusion: a page showing the Qwen curve alone would have implied a Qwen-specific defect and would have been wrong. The full study follows on its own page, Empty Answer: three lanes including this machine's local Qwen3.5, the cross-vendor control, and the method for finding your own floor instead of trusting a number. |
| 2026-08-13 | Release day: the Max-class open weights shipped, and only the Max-class. Two repositories under the official Qwen organization flipped public on 2026-08-12 by their own timestamps; on 2026-08-13 we verified both at pinned revisions, with shard counts, sizes and the license file, so their rows in section 01 move to shipped and verified. The license is the practitioner headline: not Apache 2.0 but a custom Qwen3.8-Max License with a name-display threshold and a separate-license threshold, quoted verbatim in section 02; the pre-release revenue-share report resolved into thresholds, and the geographic-restriction rumor found no support in the text we read. The shipped config says one post-trained model, always thinking, text only, on a scaled Qwen3.5 architecture, which is why we expect existing runtimes to load it without new support work (section 03; stated as an expectation, not yet load-tested here). One community GGUF set is real and sized in section 04; no official GGUF exists. The 27B remains announced, with zero official artifacts anywhere we checked, and the name-squat hazard of section 05 held true on release day, with a fresh wave of placeholders around the release itself. A day-one battery against the hosted route is underway; its numbers land in a dated follow-up (section 07). |
| 2026-08-13 | Same-day follow-up: the two release-day studies completed, and section 07 now holds their numbers. Hosted: the day-one battery ran in the afternoon, thirty-five model calls across two named providers for thirty-nine cents, after a morning attempt was blocked by this account's own provider allowlist. The budget map on the open weights matches the hosted Max curve from 2026-08-09; the reasoning-effort control, on one trivial probe, left thinking volume flat and rewrote the served prompt instead, with billed cost running opposite to the knob's apparent intent; asking for thinking off is silently ignored, so thinking is mandatory in practice; and the per-provider fine print is recorded, a premium over the listed price, a cache discount, the native 262,144-token context on both providers we touched, the million-token listing figure only on one we did not. Local: we downloaded the one community file set that fits this machine, verified all ten shards bit-perfect against the repository's own checksums, and no public runtime can load it, so the server refused it at the header parse in under half a second. Sections 01, 03, 04 and 08 now say downloaded, verified and refused instead of not downloaded. The re-test waits on a public build that declares the new quantization type. |
| 2026-08-14 | The 27B shipped, and it is Apache 2.0. Weights public at about 11:00 EDT under the official Qwen organization; verified the same day at revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0, 18 safetensors plus an index, with a LICENSE file we read in full and retained: 11,544 bytes of verbatim Apache License 2.0, no revenue thresholds of any kind, byte-identical in the official FP8 repository. Its row in section 01 moves to shipped and verified. The two open Qwen3.8 models therefore carry two different licenses, and quoting one for both is a mistake in either direction; section 02 now says so at the end. There is still no official GGUF for the 27B, and the community conversions include one from the llama.cpp project's own organization. The absorption path of section 08 ran end to end the same evening: we downloaded a community conversion, verified both files against the repository's own checksums, served it on a build this machine already had, and measured it. Those measurements are their own page, the Qwen3.8-27B field card, with a data package behind every figure. The name-squat hazard of section 05 is not retired by any of this: it now has a real repository and a real revision to check against, which section 05 names. |
| 2026-08-16 | The one open pull
request of section 03 has merged, within a day of the
weights. llama.cpp merged the chat-template plumbing for
the reasoning_effort setting (pull request
#26941)
on 2026-08-14, one day after the Max-class release day and about
three hours after the 27B's weights went public; the first
release carrying it is build b10434, the
same evening. Section 03's sentence was checked 2026-08-13 and
stays as written; as of this row there is no open 3.8-adjacent
pull request left to wait for. That strengthens the expectation
section 03 stated: the architecture needed no new support work,
and the one piece of adjacent plumbing that was open landed
within a day of the weights. One thing this does not change: the
community quantization type that section 07's local load attempt
was refused on is still not declared by any public llama.cpp
build, and no open pull request adds it, checked 2026-08-16, so
that re-test still waits. |
Every existence claim on this page is dated. If today is later than the newest row above, treat the status table as unverified since that date and check the official organization directly.
Sources
- The Qwen team's 2026-08-03 announcement and the official Qwen organization on Hugging Face, checked 2026-08-09: no Qwen3.8 repositories under the official organization as of that date. ModelScope, the other named release venue, could not be enumerated from here, and no official 3.8 listing there surfaced through secondary sources either; we will check it directly when the drop lands.
- OpenRouter's public model API, 2026-08-09: Qwen3.8-Max listed at a one-million-token context, two dollars per million input and six per million output.
- Our fifteen-call recorded battery against Qwen3.8-Max via OpenRouter, 2026-08-09: raw request and response bodies, a per-call summary table, and spend accounting are project records that can be produced if questioned.
- Runtime repositories (llama.cpp, vLLM, SGLang, Ollama), checked 2026-08-09: Qwen3.5 support is merged and shipping; no merged Qwen3.8 support in those repositories as of that date.
- The official Qwen organization on Hugging Face, checked 2026-08-13: Qwen/Qwen3.8-2.4T-A95B at revision 207bd685a7e3696cfaff12ded7c6a7ea0f88c996, 213 weight shards plus an index file, 4,892.4 GB; and Qwen/Qwen3.8-2.4T-A95B-FP8 at revision d2dc35658bcf77e66643428cb52e774cc3b5bd29, 213 shards, 2,496.2 GB. The organization's Qwen3.8 collection lists exactly these two. The license file, 3,390 bytes and byte-identical in both, read at those revisions.
- ModelScope, checked 2026-08-13: both Max-class repositories mirrored, with a 224-file listing confirmed; the 27B name returns an empty listing.
- The community GGUF repository unsloth/Qwen3.8-2.4T-A95B-GGUF, checked 2026-08-13 at revision 567d3e6ac26c: 312 GGUF files, last push 01:01 UTC that day; the sizes in section 04 are the manifest's.
- OpenRouter's public model API, 2026-08-12: the released 2.4T-A95B listed at 16:21 UTC with a million-token context, two dollars per million input tokens and six per million output, matching the Max endpoint's prices.
- Runtime repositories, checked 2026-08-13: the
qwen35moearchitecture enumeration is merged and shipping in llama.cpp; vLLM's 2026-08-07 change routes Qwen3.8 checkpoints through its existing Qwen3.5 path; the one open 3.8-adjacent llama.cpp pull request is chat-template plumbing for the effort setting, not architecture support. - Our recorded day-one battery against the released 2.4T-A95B via OpenRouter, 2026-08-13: thirty-five provider-recorded model calls across DeepInfra and Together, raw request and response bodies, per-call provider and cost accounting, the routecheck route card, and the blocked morning attempt's 26 not-found bodies. Project records that can be produced if questioned.
- The local download and load-attempt record, 2026-08-13: the pinned-revision manifest with the repository's SHA-256 checksums, our shard-by-shard verification log, the tensor-type census of all ten shards, and the server's refusal log. Project records that can be produced if questioned.
- The official Qwen organization on Hugging Face, checked 2026-08-14: Qwen/Qwen3.8-27B at revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0, 18 safetensors plus an index, and Qwen/Qwen3.8-27B-FP8 at revision 017b9c7af6b5689d5dd426a76e0bc077eb5ca20a. The Apache 2.0 LICENSE file, 11,544 bytes and byte-identical in both, read and retained at those revisions; it ships in the field card's data package.
- The Qwen3.8-27B field card on this site for the 27B's measured record of 2026-08-14, and its data package for the files behind it.
- The Qwen3.5 pair field card on this site for the local serving record this page builds on.
Vendor claims on this page are labeled as claims. Model names and trademarks belong to their owners. This page contains no scripts, no tracking and no cookies, like the rest of this site.