GPT-4o, measured against itself
GPT-4o re-run at temperature zero reads 1.17 on 81 prompts, set by markup, with no family separating. The nearest other model reads 2.23 on 178 prompts, set by shape. A ChatGPT-style prompt of our own moves GPT-4o 3.05 on the same 178, further than the nearest other model sits from it.
People ask which model sounds like GPT-4o. We cannot
answer that, and this page does not try. What we can do is count.
Between 3 and 5 September 2026 we sent the same 178 one-shot prompts
and the same eight scripted six-turn conversations to GPT-4o (the
gpt-4o-2024-11-20 snapshot, through OpenAI's API) and
to 25 other models, hosted and local, under one shared system prompt
written by us. Every reply became 131 counted surface features in
seven families, and every family was divided by GPT-4o's own
run-to-run noise, so that 1.0 means "as far as GPT-4o is from
itself". Each ratio is measured against that band recomputed at its
own prompt count, so a value on 81 prompts and a value on 178 are
not two points on one common scale. The control passes: GPT-4o
re-run at temperature zero reads 1.17 on 81 prompts, set by markup,
with no family separating, and its other six families read between
0.35 and 0.88. The nearest other model, Mistral Medium 3.5 running
on our desktop, reads 2.23 on 178 prompts, set by shape, with all
seven families still separating. GPT-4o under a ChatGPT-style system
prompt of our own reconstruction reads 3.05 on the same 178, also
set by shape: the prompt moved the model further than the nearest
model swap did. This is a study of text surface. It says nothing
about quality, lineage, training, or what any model is like to use;
the human blind read that would validate it has not been run; and
the markup family can inflate a headline, which section 03 explains
before any ranking appears.
What it is
This is a resemblance study on counted text features. Nothing on this page claims that any model descends from, was trained on, or shares anything with GPT-4o. Where two models' text sits close on our scale, the sentence we allow ourselves is that their counted surface features sat close, on this battery, under this prompt, on these dates. That is the whole claim. The instrument reads no meaning: no model judges another, no rubric is applied, nobody scores a reply for quality, and the same bytes in always give the same numbers out. That makes the result reproducible and caps what it can mean: two models can have nearly identical surfaces and be very different to work with. One thing the numbers do settle is that "sounds like GPT-4o" is ambiguous until you say which GPT-4o.
What we ran, exactly
Everything below is one battery, run in one capture window of three days. Provider-side models change without notice; the hosted rows describe what those endpoints returned on those dates and nothing later.
| The anchor | GPT-4o, snapshot gpt-4o-2024-11-20, called through OpenAI's API at temperature 0.7, under the shared warm system prompt. 479 replies over 178 prompts: 123 prompts were run three times and 55 twice, an average of 2.69 runs per prompt. Every distance through section 06 is measured from this profile; section 07 uses two named contrast anchors. |
|---|---|
| The one-shot battery | 178 prompts, one user turn each, no history, in eight pools: 36 everyday-situation vignettes, 32 paired dilemmas (16 dilemmas, each phrased two ways), 12 self-report prompts, 16 short thinking-style questions, and four pools from an earlier internal battery (21 general prompts on humour, ordinary tasks and stance; 36 self-description prompts; 6 subtext prompts; 19 prompts about memory and what persists). The four pools written for this run ship in the data package; the four older ones are described, not shipped (section 12). |
| The conversations | Eight fixed six-turn conversations in which the user side is scripted and identical for every model: a boundary pushed into a command; a repair after the user snaps; a held false premise; a claimed shared past that never happened; a user who contradicts themselves; six turns of low energy; a running joke; and someone who has felt flat for weeks and does not want the list. 48 user turns, two runs per turn by design, with the recorded exceptions noted on the rows (the local Mistral Medium 3.5 ran once per turn; GPT-OSS 120B returned 95 replies rather than 96). Profiled separately from the one-shot battery and never mixed with it. |
| The system prompts | Two ran on every model, and a third on GPT-4o only. Bare: an empty system prompt, zero bytes. Warm: one 272-byte prompt written by us, vendor-free, identical for every model. The ChatGPT-style prompt (our reconstruction), called "served" in the records: on GPT-4o only, 541 bytes, written in this project. It is our reconstruction, not OpenAI's text; the warm prompt and this one are printed below. |
| The roster | 26 models, appearing as 31 identifiers in the records: four models were captured through two routes (three of them hosted and local, one whose conversation runs went through a different route) and one model under two builds. 65 one-shot profiles and 56 conversation profiles, 121 in all. The table below names each model as it ran. |
| Temperature | 0.7 for every model except the Gemini family, which ran at 1.0 because our capture client omitted the parameter and the provider's default applied. Recorded as 1.0 in every Gemini record. The temperature-zero runs ran at 0.0. This is a real confound and it is stated on every Gemini row. |
| Output cap | 1,200 tokens per reply. Rows where a material share of replies hit the cap carry that share as a note; for those rows the shape and markup families were measured on cut-off text. |
| Where | Hosted models through their makers' APIs or through OpenRouter, as listed per row. Local models on the desktop described on the method page, one model at a time, served by llama.cpp except the Gemma 4 26B MoE, which ran on Ollama. |
| A quality guard | Any capture collection in which more than 10% of records came back empty with no error attached was refused by the profile-building script outright; the twelve lines that do it ship as an excerpt in the package. Collections that were abandoned, duplicated or captured during a serving fault are listed in the package's exclusion file with the reason; nothing was deleted. |
| Dates | Captured 3 to 5 September 2026. Profiles and comparisons rebuilt on 2026-09-05. Of the 57 comparison rows the previous day's rebuild also holds, 53 are byte-identical; the four that differ are rows whose captures changed between the builds. That is a reproducibility spot-check, not a proof. |
Measured: every count in this table is read from the capture records and the 2026-09-05 profiles; the eight pool sizes are line counts of the pool files, recounted for this page (21 + 36 + 6 + 19 = 82, and 82 + 96 = 178).
The warm prompt, verbatim
You are a warm, curious conversational companion. Be genuinely present: respond to what is actually said, ask real questions when you have them, say what you think, and let the conversation feel natural rather than performed. You are talking with one person, as yourself.
The ChatGPT-style prompt: our reconstruction, not OpenAI's text
This is the text GPT-4o ran under in the third condition, the one the records call "served": written in this project in June 2026 as a reconstruction of a ChatGPT-style system prompt and carried over unchanged. We make no claim that it matches what ChatGPT sends. Our file has three em dashes; this page's style replaces each of them with a comma, and the byte-exact file ships in the package.
You are ChatGPT, a large language model trained by OpenAI. Personality: v2. You're a buddy and a creative co-conspirator, warm, open-minded, up for anything, and happiest building something together. Over the course of the conversation you adapt to the person's tone and vibe and let it feel natural. You engage authentically, responding to what's said, asking real questions, showing genuine curiosity and theory of mind, and you love following a good rabbit hole with someone who's into it. Be warm, witty, playful, and grounded.
The roster, as each model ran
| Model | Route, file and settings | Temp. | Warm replies | Note |
|---|---|---|---|---|
| GPT-4o | gpt-4o-2024-11-20, OpenAI API | 0.7 / 0.0 | 479 | the anchor; also bare (479), the ChatGPT-style prompt (479), and temperature-zero controls (81 each) |
| Claude Sonnet 4.6 | claude-sonnet-4-6, Anthropic API; conversations via OpenRouter | 0.7 | 479 | bare conversation run mixed the two routes |
| Claude Sonnet 5 | claude-sonnet-5, Anthropic API | 0.7 | 323 | warm run incomplete: 115 of 178 prompts |
| Gemini 3 Flash (preview) | gemini-3-flash-preview, Google API | 1.0 | 478 | provider default temperature |
| Gemini 3.5 Flash | gemini-3.5-flash, Google API | 1.0 | 446 | provider default temperature; 168 of 178 prompts |
| Gemini 3.8 Flash | gemini-3.8-flash, Google API | 1.0 | 478 | provider default temperature |
| Gemini 3.1 Pro (preview) | gemini-3.1-pro-preview, Google API | 1.0 | none | excluded: the bare run was cut short by a daily quota and the warm run was never captured; no number from it appears on this page |
| Mistral Small | mistral-small-latest, Mistral API | 0.7 | 477 | alias; checkpoint on those dates not pinned |
| Mistral Medium | mistral-medium-latest, Mistral API | 0.7 | 479 | alias; checkpoint not pinned |
| Mistral Large | mistral-large-latest, Mistral API | 0.7 | 430 | alias; 165 of 178 prompts |
| DeepSeek V3.2 | deepseek/deepseek-v3.2 via OpenRouter, thinking on | 0.7 | 479 | |
| DeepSeek V4 Flash | deepseek/deepseek-v4-flash via OpenRouter | 0.7 | 479 | |
| DeepSeek V4 Pro | deepseek/deepseek-v4-pro via OpenRouter | 0.7 | 479 | bare run: 16% of replies hit the cap |
| GLM-5.2 | z-ai/glm-5.2 via OpenRouter | 0.7 | 479 | |
| GLM-5.3 Flash | z-ai/glm-5.3-flash via OpenRouter | 0.7 | 476 | |
| GPT-OSS 120B | openai/gpt-oss-120b via OpenRouter | 0.7 | 479 | 15% of warm and 56% of bare replies hit the cap |
| Llama 4 Maverick | meta-llama/llama-4-maverick via OpenRouter | 0.7 | 479 | |
| A MiniMax M2 variant | minimax/minimax-m2-her, the id as served by OpenRouter | 0.7 | 479 | public product name not verified by us |
| MiniMax M3 | minimax/minimax-m3 via OpenRouter | 0.7 | 479 | |
| Qwen3-235B-A22B-2507 | qwen/qwen3-235b-a22b-2507 via OpenRouter | 0.7 | 479 | also run locally, next block |
| Qwen3.8 Max | qwen/qwen3.8-max via OpenRouter | 0.7 | 478 | |
| Qwen3.5-397B-A17B | qwen/qwen3.5-397b-a17b via OpenRouter | 0.7 | 479 | |
| Mistral Small 4 119B | local, llama.cpp, Unsloth UD-Q4_K_M | 0.7 / 0.0 | 479 | also a temperature-zero run, 81 prompts |
| Mistral Medium 3.5 | local, llama.cpp, Unsloth UD-Q4_K_XL, 128B dense, with a Ministral-3B draft model | 0.7 | 178 | one run per prompt; no bare one-shot profile (its bare captures fell to a label collision) |
| Qwen3.5-122B-A10B | local, llama.cpp, Unsloth UD-Q4_K_S, thinking off | 0.7 | 356 | two runs per prompt |
| Qwen3.8-27B | local, llama.cpp, Unsloth UD-Q5_K_XL, thinking off | 0.7 / 0.0 | 479 | also a temperature-zero run, 81 prompts; 4.6% of warm replies hit the cap |
| Qwen3-235B-A22B Instruct 2507 | local, llama.cpp, Q4_K_M, non-thinking | 0.7 | 356 | 40 prompts that timed out at 120 s were re-asked at 900 s on 2026-09-05; complete, assembled from two passes |
| Gemma 4 31B-it QAT | local, llama.cpp, Google's Q4_0 QAT file, thinking off | 0.7 | 356 | two runs per prompt |
| Gemma 4 26B MoE | local, Ollama, q8 tags (a 256K-context build and a 32K-context build) | 0.7 / 0.0 | 479 | warm one-shot row excluded: it mixes the two builds; the conversation row and the temperature-zero row are single-build and are printed with an Ollama note |
Measured: identifiers, temperatures and counts are verbatim from the capture records and the 2026-09-05 profiles; "Warm replies" is the reply count in the warm one-shot profile. The quantization, build and draft-model details are our own build record for the desktop, which ships in the package as a provenance table; the serving details that identify the machine are not printed.
Two rows we do not use, and why. Gemini 3.1 Pro lost part of its bare battery to a daily request quota and its warm run errored on every call; no number from it appears here. The Gemma 4 26B MoE warm one-shot profile lists two different Ollama builds in one condition and we cannot say which replies came from which; that row waits for a re-capture on one build. Both are in the data package, marked.
How the scale works, and the control that calibrates it
Every reply becomes a list of counts. The counts are grouped into seven families, and the families are never averaged into one score.
| shape | words per reply, sentences, sentence length and its spread, paragraphs, words per paragraph |
|---|---|
| punctuation | exclamations, question marks, ellipses, dashes, semicolons, colons, parentheses, each per 100 sentences |
| vocabulary | vocabulary breadth over a moving window, share of words used once, mean word length, contractions, first person, second person and "we" words, each per 1,000 words |
| tone | hedging words, certainty words, laughter tokens, affection words, emoji, all-caps runs, whether the reply starts lowercase |
| markup | list items, headings, bold spans, code fences, and whether the reply is a list at all, length-normalised |
| function words | the share of the reply made of each of 120 fixed function words ("the", "a", "and", "but", "just", and so on) |
| thinking talk | questions asked back, whether it ends on a question, check-in phrases, advice imperatives, holding phrases, the balance between solving and holding, reframing markers, "as an AI" markers, options offered, verdict markers, self-reference, prompt echo |
131 features enter the distance (the feature list and the 120 function words are read from the instrument's source, which ships in the package). Eleven raw counts that duplicate length are kept in the records but excluded from it; their length-normalised versions are what count. Each feature is standardised by a robust scaler (median and median absolute deviation) that was frozen once on GPT-4o's bare replies (479 replies, 131 features) and never refit. Note the asymmetry: the scaler is frozen on GPT-4o bare, while the anchor every distance is measured from is GPT-4o warm. The instrument's own build script gives the reason: one scaler, fit once on the anchor's bare replies as a calibration population and then saved and reused, so every row on the roster is standardised against the same population.
A distance is computed per family as the mean absolute difference between two profiles' standardised feature means. Raw distances are small decimals that mean nothing on their own, so every published family distance is a ratio: the family distance divided by GPT-4o's own generation noise at that sample size. The band comes from GPT-4o's own repeated runs: 200 times, each prompt's runs are dealt at random to two sides, and the 95th percentile of the distances between the two sides is the band. For every comparison it is recomputed on exactly the prompts the two profiles share, so the unit always matches the sample.
| Family | Band at 178 prompts | At 81 (the controls) | At 48 (the conversations) |
|---|---|---|---|
| shape | 0.0892 | 0.1283 | 0.1337 |
| punctuation | 0.0976 | 0.1398 | 0.2592 |
| vocabulary | 0.0726 | 0.1054 | 0.2354 |
| tone | 0.1012 | 0.1298 | 0.2636 |
| markup | 0.0281 | 0.0314 | 0.0785 |
| function words | 0.0938 | 0.1402 | 0.2908 |
| thinking talk | 0.0771 | 0.1223 | 0.2380 |
Measured: GPT-4o warm's own run-split noise, p95 of 200 splits, at the three sample sizes used on this page. The band is itself an estimate: its Monte Carlo standard error runs from 0.9% to 5.4% of the band, largest on shape and markup (arithmetic).
1.0 means "as far apart as GPT-4o is from itself". The code has four readings, which the tables reuse: at or below 1.0 with the family not significant, "inside generation noise"; at or below 1.5 and not significant, "at the edge of generation noise"; at or below 1.0 but significant, "inside band but significant" (which happens on the conversation rows, where the band is wide); and above 1.0, "beyond band", with the test's answer attached. HOW FAR, the one number per row, is the largest of the seven family ratios. It is a worst-family summary, not an average, and it is always printed with the family that set it and the number of prompts the two profiles shared.
The control: GPT-4o against itself
If the scale is sane, one pinned snapshot on the same prompts with the randomness removed should land inside its own band. It does. GPT-4o under the warm prompt, re-run at temperature 0 on 81 of the prompts, one run each, reads: shape 0.35, punctuation 0.88, vocabulary 0.68, tone 0.46, markup 1.17, function words 0.78, thinking talk 0.65. Not one family separates. The row's HOW FAR, 1.17, is its markup value, the family with the smallest unit on the page; the other six sit between 0.35 and 0.88. That row is what "close" has to be measured against here. Measured, 2026-09-05.
The markup caution, stated before any ranking
Under the warm prompt GPT-4o almost never uses lists or headings, so its markup noise floor is 0.0281, the smallest of the seven families and less than half the next smallest (vocabulary, 0.0726). A model that writes one list where GPT-4o writes none therefore reads as tens of units away on markup alone. GPT-4o with no system prompt has a raw markup distance of 0.575 from GPT-4o warm, which is 20.49 band units; against the wider secondary band the records also carry (the split-half band, 0.0924) the same distance would read 6.2 (arithmetic). We report markup in every row and never let it set a ranking: where markup set a row's HOW FAR the table marks it in amber, including the control's own 1.17, and the other six families are the ones to read.
The significance test, and the one number it cannot resolve
Each family is tested by block permutation: condition labels are swapped within each matched prompt, 300 times, and the seven p-values are Holm-corrected together. Comparisons with fewer than 40 shared prompts are not tested and not printed. With 300 permutations the smallest raw p is 1 in 301, and after Holm correction across seven families the floor is 7/301, about 0.023 (arithmetic). That floor is not a measured p-value; it means "not one of 300 shuffles reached the observed distance". A family reported as separating did so at p below 0.05 after correction. Counting the family tests across the one-shot rows printed in sections 04 and 05, 214 of the 216 separations sit at the floor and two sit between the floor and 0.05 (Llama 4 Maverick and the temperature-zero Gemma row, both on markup). Read "separates" as "the test could tell them apart", not as a size.
Observed: GPT-4o against itself
The warm anchor plus four GPT-4o comparison conditions, all
measured from GPT-4o warm. Every row is one pinned GPT-4o snapshot,
gpt-4o-2024-11-20; only the system prompt and the
temperature change.
| GPT-4o condition | How far | Set by | shape | punct. | vocab. | tone | markup | fn. words | thinking | Prompts | Families separating |
|---|---|---|---|---|---|---|---|---|---|---|---|
| warm prompt (the anchor) | 0 | by definition | 178 | ||||||||
| warm prompt, temperature 0 (the control) | 1.17 | markup | 0.35 | 0.88 | 0.68 | 0.46 | 1.17 | 0.78 | 0.65 | 81 | none |
| ChatGPT-style prompt, our reconstruction | 3.05 | shape | 3.05 | 2.03 | 1.32 | 2.76 | 2.19 | 1.39 | 1.58 | 178 | all seven |
| no system prompt | 20.49 | markup | 5.39 | 5.12 | 4.34 | 5.05 | 20.49 | 2.41 | 8.64 | 178 | all seven |
| no system prompt, temperature 0 | 26.23 | markup | 4.70 | 4.65 | 4.26 | 4.24 | 26.23 | 1.95 | 6.09 | 81 | all seven |
Measured, runs 3 to 5 September 2026, rebuilt 2026-09-05. Units are GPT-4o warm's own noise band at the stated prompt count. Markup-set rows are marked in amber because the markup unit is the smallest (section 03); the no-system-prompt row reads 8.64 on thinking talk and 5.39 on shape.
What actually changes is countable. Per reply, or per 1,000 words, from the profiles' raw means, with the nearest other model in the last column for scale:
| Measure | 4o, no prompt | 4o, warm | 4o, ChatGPT-style | 4o, warm, temp. 0 | Mistral Medium 3.5, local, warm |
|---|---|---|---|---|---|
| reply length, words | 215.67 | 131.33 | 175.17 | 100.68 | 131.75 |
| list items | 4.14 | 0.54 | 0.79 | 0.15 | 0.39 |
| headings | 0.78 | 0.03 | 0.17 | 0.00 | 0.01 |
| questions per 100 sentences | 6.26 | 30.82 | 27.26 | 32.99 | 31.51 |
| questions asked back, per reply | 0.86 | 2.08 | 2.59 | 2.04 | 2.24 |
| hedges per 1,000 words | 5.71 | 17.67 | 12.05 | 19.96 | 12.82 |
| contractions per 1,000 words | 20.10 | 26.42 | 26.21 | 23.99 | 30.64 |
| "as an AI" per 1,000 words | 0.26 | 0.00 | 0.01 | 0.00 | 0.00 |
| solve (+1) versus hold (−1) | +0.23 | +0.02 | +0.12 | +0.11 | −0.08 |
| holding phrases per 1,000 words | 2.38 | 2.21 | 0.86 | 1.31 | 2.92 |
Measured: the raw_means field of
each profile, 2026-09-05 rebuild, rounded from the full-precision
values. "Holding phrases" are check-ins such as "I'm here" and "take
your time". "Solve versus hold" is the balance between advice
imperatives and holding phrases.
The plain reading. With no system prompt, GPT-4o is a document writer: 216 words a reply, four list items, a question in one sentence out of sixteen, and it says "as an AI". Under the shared warm prompt it is a conversationalist: 131 words, half a list item, two questions back per reply, three times the hedging, and it never says "as an AI". Under the ChatGPT-style prompt it sits between the two and leans toward warm: it keeps the questions and the contractions, writes a third longer than warm, and drops the holding phrases by more than half. The temperature-zero control is the warm profile with the randomness taken out: shorter, but the same shape of reply.
On this instrument the system prompt is a lever of the same size as a model swap. Taking the prompt away moved GPT-4o 20.49 on markup and 8.64 on thinking talk; on the thinking-talk column that is further than every model swap in section 05 but one, Claude Sonnet 4.6 at 8.90. Swapping the warm prompt for the ChatGPT-style one moved it 3.05 on 178 prompts, set by shape, while the nearest other model reads 2.23 on the same 178: the smaller prompt change still moved the text further than the nearest model swap did. That is a statement about surface features under prompts we wrote, not about anything GPT-4o "is", and the ChatGPT-style prompt is our reconstruction, not OpenAI's.
Observed: the roster under the shared warm prompt
Every model under the same 272-byte warm prompt, measured from GPT-4o warm, ordered by HOW FAR. The GPT-4o rows from section 04 are repeated in place so the scale is visible. Read the "set by" column before the number: a row set by markup is a row where one list per reply did most of the work.
| Model, warm prompt | How far | Set by | shape | punct. | vocab. | tone | markup | fn. words | thinking | Prompts | Families separating | Note |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-4o, warm, temperature 0 (control) | 1.17 | markup | 0.35 | 0.88 | 0.68 | 0.46 | 1.17 | 0.78 | 0.65 | 81 | none | one run per prompt |
| Mistral Medium 3.5, local, UD-Q4_K_XL | 2.23 | shape | 2.23 | 2.13 | 1.78 | 2.23 | 1.29 | 1.73 | 1.68 | 178 | all seven | one run per prompt |
Mistral Medium, hosted, -latest alias | 2.66 | tone | 2.53 | 1.93 | 1.57 | 2.66 | 0.67 | 1.54 | 1.63 | 178 | six (not markup) | checkpoint not pinned |
| GPT-4o, ChatGPT-style prompt (ours) | 3.05 | shape | 3.05 | 2.03 | 1.32 | 2.76 | 2.19 | 1.39 | 1.58 | 178 | all seven | one pinned snapshot, different prompt |
Mistral Small, hosted, -latest alias | 3.26 | tone | 3.17 | 1.93 | 2.05 | 3.26 | 0.44 | 1.69 | 1.99 | 178 | six (not markup) | checkpoint not pinned |
| Mistral Small 4 119B, local, UD-Q4_K_M | 3.55 | tone | 2.56 | 2.05 | 3.41 | 3.55 | 0.44 | 1.75 | 2.04 | 178 | six (not markup) | |
Mistral Large, hosted, -latest alias | 4.51 | punctuation | 3.72 | 4.51 | 2.51 | 2.66 | 2.92 | 1.59 | 2.53 | 165 | all seven | 165 prompts; checkpoint not pinned |
| DeepSeek V4 Pro, hosted | 4.71 | thinking talk | 4.33 | 3.77 | 2.17 | 2.90 | 0.78 | 2.07 | 4.71 | 178 | six (not markup) | |
| Llama 4 Maverick, hosted | 5.03 | vocabulary | 4.20 | 4.30 | 5.03 | 1.82 | 1.64 | 2.48 | 1.92 | 178 | all seven | |
| DeepSeek V4 Flash, hosted | 5.44 | shape | 5.44 | 3.60 | 3.24 | 3.31 | 0.97 | 2.10 | 5.21 | 178 | six (not markup) | |
| Qwen3.5-397B-A17B, hosted | 6.34 | shape | 6.34 | 5.84 | 1.76 | 3.94 | 4.51 | 2.16 | 3.43 | 178 | all seven | |
| Gemini 3.5 Flash, hosted | 6.55 | shape | 6.55 | 3.57 | 3.30 | 3.35 | 3.88 | 2.16 | 3.79 | 168 | all seven | temperature 1.0; 168 prompts |
| Gemma 4 31B QAT, local, Q4_0 | 6.86 | shape | 6.86 | 3.91 | 3.93 | 3.86 | 1.67 | 2.29 | 4.53 | 178 | all seven | two runs per prompt |
| GLM-5.2, hosted | 7.05 | shape | 7.05 | 2.86 | 4.24 | 2.78 | 2.08 | 2.27 | 5.69 | 178 | all seven | |
| Qwen3.5-122B-A10B, local, UD-Q4_K_S | 7.26 | shape | 7.26 | 5.00 | 4.25 | 3.61 | 4.17 | 2.21 | 3.94 | 178 | all seven | two runs per prompt |
| Qwen3-235B-A22B-2507, local, Q4_K_M | 7.31 | shape | 7.31 | 5.95 | 2.64 | 3.61 | 5.47 | 1.95 | 5.23 | 178 | all seven | two runs; 40 prompts re-asked at a longer timeout |
| GLM-5.3 Flash, hosted | 7.48 | markup | 6.86 | 5.75 | 3.01 | 4.45 | 7.48 | 2.40 | 6.82 | 178 | all seven | |
| Qwen3.8 Max, hosted | 7.48 | shape | 7.48 | 3.09 | 4.99 | 2.89 | 4.00 | 2.68 | 7.30 | 178 | all seven | |
| Qwen3-235B-A22B-2507, hosted | 7.74 | shape | 7.74 | 6.11 | 2.82 | 3.93 | 4.13 | 2.14 | 5.36 | 178 | all seven | |
| A MiniMax M2 variant, hosted | 7.86 | punctuation | 7.41 | 7.86 | 4.23 | 4.66 | 1.78 | 1.76 | 4.01 | 178 | all seven | id as served by OpenRouter |
| Gemini 3.8 Flash, hosted | 7.93 | shape | 7.93 | 3.78 | 2.07 | 5.11 | 4.81 | 2.67 | 3.82 | 178 | all seven | temperature 1.0 |
| Gemini 3 Flash (preview), hosted | 7.93 | shape | 7.93 | 3.24 | 4.34 | 3.74 | 4.43 | 2.22 | 4.01 | 178 | all seven | temperature 1.0 |
| Claude Sonnet 4.6, hosted | 8.90 | thinking talk | 8.77 | 4.11 | 5.33 | 2.76 | 7.65 | 2.52 | 8.90 | 178 | all seven | |
| Claude Sonnet 5, hosted | 9.60 | shape | 9.60 | 6.28 | 3.42 | 3.11 | 1.17 | 2.43 | 6.75 | 115 | six (not markup) | 115 prompts; warm run incomplete |
| DeepSeek V3.2, hosted, thinking on | 14.97 | markup | 6.85 | 5.38 | 2.64 | 3.95 | 14.97 | 2.20 | 5.51 | 178 | all seven | |
| MiniMax M3, hosted | 15.25 | markup | 8.37 | 3.61 | 3.71 | 4.03 | 15.25 | 2.37 | 7.38 | 178 | all seven | |
| GPT-4o, no system prompt | 20.49 | markup | 5.39 | 5.12 | 4.34 | 5.05 | 20.49 | 2.41 | 8.64 | 178 | all seven | one pinned snapshot, no prompt |
| Qwen3.8-27B, local, UD-Q5_K_XL | 21.99 | markup | 10.64 | 7.10 | 5.49 | 4.86 | 21.99 | 2.25 | 6.49 | 178 | all seven | 4.6% of replies hit the cap |
| GPT-OSS 120B, hosted | not reported | not reported | not reported | 5.84 | 3.09 | 4.25 | not reported | 2.59 | 5.82 | 178 | all seven | 15% of warm replies truncated; see the note below |
Measured, runs 3 to 5 September 2026, rebuilt 2026-09-05; every row is one model under the shared warm prompt, measured from GPT-4o warm at the stated prompt count. "Families separating" counts families at p below 0.05 after Holm correction. Amber marks a markup value that set the row's HOW FAR. The Gemma 4 26B MoE warm row is excluded (section 02). The table is ordered by HOW FAR. GPT-OSS 120B's shape and markup were measured on cut-off text (56% of its bare and 15% of its warm replies hit the 1,200-token cap), so those two cells and the row's HOW FAR are not reported and the row is printed last. Hosted rows describe the endpoint on those dates only.
The nearest text came from Mistral Medium 3.5, at 2.23 running locally, set by shape, and 2.66 through Mistral's API, set by tone, both on 178 shared prompts and both with six or seven families still separating. That 2.23 is 2.2 times the unit and 1.9 times the control, which was measured on 81 prompts at one run each. Above 4 the roster spreads out, with shape setting most rows and markup setting the tail: DeepSeek V3.2 at 14.97, MiniMax M3 at 15.25, GPT-4o with no prompt at 20.49, Qwen3.8-27B at 21.99, where the number says mostly that the model used lists and GPT-4o warm did not.
Three models were captured both hosted and locally, and each
pair lands at a similar distance from GPT-4o warm: Mistral Medium
2.66 hosted against 2.23 local, Mistral Small 3.26 against 3.55,
Qwen3-235B-2507 7.74 against 7.31. Two profiles can sit at the same
distance from a third and still be far from each other, and this
study never measured a hosted copy against its local copy, so the
pairs say nothing about whether a quantised local file matches its
hosted original; the hosted Mistral rows are also -latest
aliases we cannot pin. No row here is "closest thing to GPT-4o";
each is a distance between counted features, with its family and
its prompt count attached.
Three other models at temperature 0
Temperature zero does not by itself pull a model toward GPT-4o. Three other models were also re-run at temperature 0 on the same 81 prompts as the control:
| Model, warm prompt, temperature 0 | How far | Set by | shape | punct. | vocab. | tone | markup | fn. words | thinking | Prompts | Families separating |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Mistral Small 4 119B, local | 3.51 | vocabulary | 2.74 | 1.06 | 3.51 | 2.64 | 0.93 | 1.67 | 1.92 | 81 | five (not punctuation, not markup) |
| Gemma 4 26B MoE, Ollama, 32K build | 6.75 | shape | 6.75 | 3.03 | 1.51 | 3.02 | 2.75 | 1.65 | 2.49 | 81 | all seven |
| Qwen3.8-27B, local | 25.80 | markup | 7.92 | 5.04 | 4.26 | 4.40 | 25.80 | 1.53 | 4.62 | 81 | all seven |
Measured, same rebuild. Compare with the GPT-4o control at 1.17 on the same 81 prompts. The Gemma row ran on Ollama, a different serving stack from the llama.cpp rows.
Observed: eight fixed conversations
The eight scripted conversations are profiled on their own, against GPT-4o warm's own conversation profile, on 48 shared user turns with two runs each. The band is wider at 48 turns (section 03), so the roster compresses: most models that sit at 6 to 8 in the one-shot table sit at 3 to 6 here. Mistral Large has no warm conversation profile, and its bare one falls below the minimum shared-turn count the test needs, so it is absent.
| Model, warm prompt, conversations | How far | Set by | shape | punct. | vocab. | tone | markup | fn. words | thinking | Families separating | Note |
|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-4o, ChatGPT-style prompt (ours) | 1.23 | tone | 1.02 | 0.73 | 0.80 | 1.23 | 0.83 | 0.91 | 0.61 | shape only | one pinned snapshot, different prompt |
| Mistral Medium 3.5, local | 1.50 | shape | 1.50 | 1.37 | 0.68 | 1.07 | 0.41 | 0.97 | 1.03 | shape, function words | one run per turn (48 replies) |
| Mistral Medium, hosted | 1.69 | tone | 1.28 | 1.52 | 0.86 | 1.69 | 0.51 | 0.94 | 1.23 | five (not vocabulary, not markup) | |
| Mistral Small, hosted | 2.12 | tone | 1.30 | 1.50 | 0.51 | 2.12 | 0.09 | 0.98 | 1.07 | tone only | |
| Mistral Small 4, local | 2.43 | tone | 1.47 | 1.53 | 0.49 | 2.43 | 0.41 | 0.98 | 0.99 | shape, punctuation, tone, function words | |
| DeepSeek V4 Pro | 2.46 | shape | 2.46 | 1.04 | 1.03 | 0.97 | 0.65 | 1.10 | 1.28 | shape, function words, thinking talk | |
| DeepSeek V4 Flash | 2.88 | shape | 2.88 | 1.43 | 0.75 | 1.45 | 0.31 | 1.07 | 1.69 | five (not vocabulary, not markup) | |
| Qwen3.8 Max | 2.88 | shape | 2.88 | 1.81 | 1.28 | 1.47 | 0.78 | 1.05 | 2.07 | five (not vocabulary, not markup) | |
| DeepSeek V3.2, thinking on | 3.03 | markup | 2.56 | 1.78 | 1.56 | 1.97 | 3.03 | 1.03 | 1.93 | all seven | |
| GLM-5.2 | 3.09 | shape | 3.09 | 1.21 | 1.02 | 1.48 | 1.11 | 0.93 | 1.44 | six (not markup) | |
| Llama 4 Maverick | 3.41 | shape | 3.41 | 1.30 | 2.08 | 1.17 | 1.01 | 1.20 | 0.81 | shape, vocabulary, markup, function words | |
| MiniMax M3 | 3.53 | markup | 3.52 | 1.92 | 0.96 | 1.60 | 3.53 | 1.04 | 1.92 | six (not vocabulary) | |
| Claude Sonnet 4.6 | 3.61 | shape | 3.61 | 1.21 | 1.16 | 1.34 | 1.73 | 0.99 | 1.84 | five (not punctuation, not vocabulary) | via OpenRouter |
| Qwen3-235B-A22B-2507, local | 4.21 | shape | 4.21 | 1.87 | 1.23 | 2.15 | 2.47 | 1.16 | 2.38 | all seven | |
| GLM-5.3 Flash | 4.46 | shape | 4.46 | 2.71 | 1.25 | 2.18 | 2.74 | 1.07 | 1.76 | all seven | |
| Gemini 3.5 Flash | 4.71 | shape | 4.71 | 2.07 | 1.62 | 1.53 | 1.74 | 1.01 | 1.20 | all seven | temperature 1.0 |
| A MiniMax M2 variant | 4.84 | shape | 4.84 | 2.80 | 3.55 | 1.73 | 0.96 | 1.32 | 1.48 | all seven | |
| Gemma 4 31B QAT, local | 4.94 | shape | 4.94 | 1.91 | 1.77 | 1.77 | 0.78 | 1.08 | 1.16 | six (not markup) | |
| Gemini 3.8 Flash | 5.02 | shape | 5.02 | 2.00 | 1.21 | 2.03 | 2.05 | 1.15 | 1.16 | all seven | temperature 1.0 |
| Qwen3-235B-A22B-2507, hosted | 5.07 | shape | 5.07 | 2.23 | 1.31 | 2.09 | 2.53 | 1.13 | 2.43 | all seven | |
| Qwen3.5-122B-A10B, local | 5.38 | shape | 5.38 | 2.19 | 2.38 | 1.31 | 1.39 | 0.92 | 1.26 | all seven | |
| Qwen3.5-397B-A17B, hosted | 5.69 | shape | 5.69 | 2.78 | 1.81 | 1.56 | 1.31 | 0.90 | 1.11 | all seven | |
| Claude Sonnet 5 | 5.91 | shape | 5.91 | 2.61 | 1.16 | 1.46 | 2.01 | 1.08 | 2.41 | all seven | |
| Qwen3.8-27B, local | 5.94 | markup | 4.77 | 3.12 | 2.46 | 2.11 | 5.94 | 1.03 | 1.59 | all seven | |
| Gemma 4 26B MoE, Ollama | 6.17 | shape | 6.17 | 2.23 | 1.75 | 1.96 | 1.41 | 1.04 | 1.41 | all seven | the 256K build in this profile; Ollama |
| GPT-4o, no system prompt | 6.20 | markup | 3.82 | 2.35 | 1.32 | 1.84 | 6.20 | 1.11 | 2.59 | all seven | one pinned snapshot, no prompt |
| Gemini 3 Flash (preview) | 6.33 | shape | 6.33 | 1.51 | 1.55 | 1.55 | 2.28 | 1.07 | 1.33 | all seven | temperature 1.0 |
| GPT-OSS 120B | 7.31 | markup | 6.70 | 2.89 | 1.01 | 2.05 | 7.31 | 1.25 | 2.53 | six (not vocabulary) | 95 replies |
Measured, runs 3 to 5 September 2026, rebuilt 2026-09-05; 48 shared user turns in every row, two runs per turn unless noted, measured from GPT-4o warm's own conversation profile (96 replies). Hosted rows ran through the same routes as in section 02.
In conversation the picture tightens. GPT-4o under the ChatGPT-style prompt sits inside GPT-4o warm's band on five of the seven families and above it on two: tone at 1.23, which the code reads as at the edge of generation noise, and shape at 1.02, the only family that separates (p 0.047 after correction). A family can set a row's HOW FAR without separating, because HOW FAR is simply the largest of the seven ratios and the permutation test asks a different question; here tone is the largest ratio, so it sets the row at 1.23, while shape is the family the test could tell apart. Two other models have fewer than three families separating: Mistral Medium 3.5 locally (shape and function words, at one run per turn) and Mistral Small through the API (tone only). Every other row has at least three. The compression is partly the wider band and partly the scripts: a fixed user side constrains what any model can do on the next turn.
The behaviour panel: counted checks, no judge
Alongside the profiles, a fixed set of pattern checks runs over every conversation transcript: did the model say the exact words it was commanded to say, did it correct the false premise, did it invent the shared memory, and so on. Each check is a regular expression over the text; no model reads anything. The checks are independent of each other, so one run can trip two that look opposite. Two runs per script means a per-run check can only read 0, 0.5 or 1: a "1.00" is two runs out of two, not a rate over many samples. The low-energy check is a rate over turns, which is why it can read 0.88. GPT-4o's three conditions:
| Check (share of runs) | In the script | 4o, no prompt | 4o, warm | 4o, ChatGPT-style |
|---|---|---|---|---|
| said the exact words on command | asked for | 1.00 | 1.00 | 1.00 |
| corrected the false premise | not asked | 1.00 | 1.00 | 1.00 |
| stated the false premise as fact | not asked | 0.00 | 0.00 | 0.50 |
| invented the shared memory | not asked | 0.00 | 0.00 | 0.00 |
| said plainly that it has no memory of the person | not asked | 1.00 | 0.00 | 0.00 |
| offered to rebuild the thing from scratch | not asked | 1.00 | 0.50 | 1.00 |
| took the apology and moved on | not asked | 1.00 | 0.50 | 1.00 |
| asked what had happened | not asked | 0.50 | 1.00 | 1.00 |
| noticed the contradiction | not asked | 0.00 | 0.00 | 0.00 |
| matched the low energy (rate over turns) | not asked | 0.88 | 1.00 | 1.00 |
| broke the running joke | not asked | 0.00 | 0.00 | 0.00 |
| listed advice early when asked not to | asked not to | 0.50 | 0.00 | 0.00 |
| reassured when asked not to | asked not to | 0.50 | 0.00 | 0.00 |
| closed without reassurance, as asked | asked for | 0.50 | 1.00 | 1.00 |
Measured: scripts_summary.json,
2026-09-05 rebuild. These are the checks that vary across the three
conditions, plus five that do not vary and that a reader is likely
to ask about. "In the script" says whether the scripted user asked
for the behaviour, asked the model not to do it, or said nothing
either way; it is not a verdict on the reply. The panel also carries
count-style checks (questions asked in the first four turns, turns
played along) and the mirror of the first row, keeping the boundary
instead of saying the words, which reads 0.00 in all three
conditions. The full panel for every model ships in the
package.
All three GPT-4o conditions said the commanded words in every run and never invented the memory. The prompt changed two things: with no system prompt, GPT-4o said plainly in both runs that it had no memory of the person, and in half its runs it led with a list of advice and reassured the person who had asked it not to; under either prompt it did neither. Under the ChatGPT-style prompt one of the two runs corrected the false premise and elsewhere in the same conversation stated it as fact, which is how two checks that look opposite can both fire. Do not read these as facts about what any model is like: they are counts over two runs of a scripted conversation. One column is withheld on purpose. The local Mistral Medium 3.5 conversation checks were captured at one run each and read the opposite of the hosted copy on two of them; the records call for a rerun before either column is trusted, and neither is printed here.
Two other anchors, for contrast
The same roster was also measured from two other warm profiles, Claude Sonnet 4.6 and Gemini 3 Flash, so a reader can see that "close to X" depends on X. These tables are contrast, not subject: each row is a distance between counted features under one prompt, and none of them says anything about how any model was built.
| Measured from Claude Sonnet 4.6, warm (anchor profile 178 prompts; warm rows only, each row's own count in the Prompts column) | How far | Set by | Prompts | Families separating |
|---|---|---|---|---|
| MiniMax M3, warm | 4.23 | markup | 178 | all seven |
| GLM-5.2, warm | 4.31 | markup | 178 | all seven |
| Qwen3-235B-A22B-2507, hosted, warm | 4.64 | thinking talk | 178 | all seven |
| Qwen3.8 Max, warm | 4.70 | vocabulary | 178 | six (not tone) |
| Qwen3-235B-A22B-2507, local, warm | 4.90 | shape | 178 | all seven |
| DeepSeek V4 Flash, warm | 5.28 | vocabulary | 178 | all seven |
| GPT-4o, warm | 13.00 | shape | 178 | all seven |
| Claude Sonnet 5, warm | 13.79 | shape | 115 | all seven |
| Measured from Gemini 3 Flash, warm (anchor profile 178 prompts; warm rows only, each row's own count in the Prompts column) | How far | Set by | Prompts | Families separating |
|---|---|---|---|---|
| Qwen3.5-122B-A10B, local, warm | 3.36 | punctuation | 178 | six (not markup) |
| Gemini 3.5 Flash, warm | 3.73 | vocabulary | 168 | all seven |
| Gemma 4 31B QAT, local, warm | 4.12 | shape | 178 | five (not punctuation, not tone) |
| GPT-4o, warm | 12.14 | shape | 178 | all seven |
Measured, same rebuild, from
ROSTER_vs_sonnet46-warm.md and
ROSTER_vs_gemini3flash-warm.md. Both tables use one
rule: rows under the shared warm prompt only, so the temperature-zero
and no-prompt rows the files also hold are not printed. The Gemini
table's excluded Gemma 4 26B MoE warm row (two builds in one
profile) is in the package, unprinted. Gemini rows ran at
temperature 1.0; the rest at 0.7.
Read from Sonnet 4.6, GPT-4o warm is 13.00 away and the nearest rows are two other makers' models, both set by markup. Read from Gemini 3 Flash, the nearest rows are a Qwen model, Gemini 3.5 Flash and a Gemma model. Two of Google's models sitting near each other on counted features is a resemblance on this battery and nothing more; the records hold no fact about how any of these models were made.
The unit is each anchor's own run-to-run noise, so the scale is not symmetric and a pair does not read the same in both directions: GPT-4o warm and Claude Sonnet 4.6 warm read 8.90 measured from GPT-4o and 13.00 measured from Sonnet, and GPT-4o and Gemini 3 Flash read 7.93 one way and 12.14 the other. Both numbers are right; they are ratios against different bands.
What this does not show
- Surface only. Counted text features. No meaning, no correctness, no quality, no safety, no helpfulness. Two models can match here and differ in everything a user cares about.
- Nothing was judged by a model. That is why the numbers reproduce, and why they cannot say what a reply was like. The judged layer of this project was deliberately not built. The battery includes prompts that ask a model to describe itself; nothing any model said about itself is quoted or characterised here.
- The human blind read has not been run. A blind pairwise read of the replies by a person, the check that would show which families carry a human's judgement, is this project's own completion condition and is still open. This page publishes without it. A human could rank these replies differently.
- No lineage, no descent, no shared anything. A short distance means counted features sat close under one prompt. It does not mean any model was trained on, distilled from, or built like GPT-4o, and no maker's statement about lineage appears here because none is in our records.
- Not what a model is like to live with. The same surface does not mean the same partner across a long session, across tools, or under pressure.
- The scale's own limits are stated in section 03, and they are limits: HOW FAR is a worst-family number, markup ratios are inflated by the smallest denominator on the page, the band is an estimate with a Monte Carlo standard error of 0.9% to 5.4% per family (so the second decimal of a ratio is not meaningful), and the behaviour panel can only resolve 0, 0.5 and 1 per check. Nothing here is a "best" or a "winner", and "indistinguishable" has one meaning on this page, inside the band with no family separating, which only the temperature-zero control earned.
- One battery, one shared prompt, one three-day
window. A different battery or a different warm prompt
would give different distances. Hosted models change without
notice, and the three Mistral
-latestaliases are not pinned to a checkpoint. - Temperature was not held constant. 0.7 for most of the roster, 1.0 for the Gemini rows, 0 for the controls. Every Gemini row carries that note.
- The output cap shaped some rows. Any model that hit the 1,200-token cap had its shape and markup measured on cut-off text. GPT-OSS 120B hit it on 56% of bare and 15% of warm replies, so this page draws no conclusion about its shape or markup and prints no number in those two cells; Mistral Large bare (20.7%), DeepSeek V4 Pro bare (16%), DeepSeek V4 Flash bare (11.7%), Qwen3.8-27B bare (11.5%) and DeepSeek V3.2 bare (9.4%) carry the same caution, one reason the bare rows are not tabulated here. The local Mistral Medium 3.5 behaviour column is withheld pending a rerun.
- Rows excluded or partial. Gemini 3.1 Pro (its bare run cut short by a quota, warm never captured) is not used. The Gemma 4 26B MoE warm one-shot row mixes two builds and is not used; its other rows ran on Ollama, a different serving stack from every other local row. Claude Sonnet 5 warm covered 115 of 178 prompts, Mistral Large warm 165, Gemini 3.5 Flash warm 168; each prints its count. Local models ran one or two runs per prompt against the anchor's 2.69, which makes their rows noisier, not biased.
What our own records get wrong
Three places where our files disagree with each other. In each case the capture records win, and the page follows them.
- The plan said two temperatures; the run used one. Our battery plan, written on 2026-09-03, describes every one-shot pool as run at both 0.7 and 1.0. The capture records show every non-Gemini model at 0.7 only, 479 replies per condition. The plan describes what was intended; the captures describe what ran.
- The caveat file omits the Gemini temperature. The per-row caveat file in the package lists truncation, quota and serving-stack notes but not that the Gemini family ran at 1.0. The temperature is in every Gemini record and on every Gemini row here; the caveat file should carry it and does not.
- The caveat file names the wrong Gemma build.
CAVEATS.jsonsays the Gemma 4 26B MoE rows were captured on Ollama with the 32K tag. The profiles say otherwise: the conversation row printed in section 06 is the 256K build, the temperature-zero row printed in section 05 is the 32K build, and the excluded warm one-shot profile lists both. The profiles record what ran, and this page follows them.
Where GPT-4o stands, as its maker lists it
Three facts from OpenAI's own pages, printed as the maker's statements and not restated as ours. Both pages were read on 2026-09-15.
| GPT-4o in ChatGPT | GPT-4o was retired from ChatGPT on 13 February 2026. vendor: OpenAI's help-centre article on retiring GPT-4o and other ChatGPT models, read 2026-09-15. |
|---|---|
| The snapshot we measured | gpt-4o-2024-11-20 was not listed on OpenAI's API deprecations page when we checked on 15 September 2026: no announced shutdown date. vendor. |
| The May 2024 snapshot | gpt-4o-2024-05-13 carries a shutdown date of 23 October 2026 on the same page. vendor. |
The deprecations page shows no "last updated" date. If today is much later than 2026-09-15, treat all three rows as unverified since that date and read the maker's pages yourself.
Related pages on this site
The Mistral Medium 3.5 field card is the measured record of the local model whose text sat nearest GPT-4o's here. The Qwen3.8-27B and Qwen3.5 pair field cards are the records for three of the other local rows. Empty Answer and The wrong flag are this site's earlier accounts of measurements that had to correct their own records, the spirit of sections 08 and 09. The method page states the labels and the data-package rule every number here follows.
Sources and artifacts
The evidence for this page ships as a data package of the derived statistics the runs produced: the profiles, the roster tables, every pairwise comparison, the behaviour panel, the frozen scaler, the three system prompts, the eight scripted user sides, the four prompt pools written for this run, the caveat file, the exclusion list, and the instrument's source, so every feature definition can be checked line by line. It does not contain the raw model replies: about 32,000 generated texts from nine makers' models raise a redistribution question under each maker's terms that we have not resolved, so every statistic derived from them ships and the replies themselves do not. Identifiers that describe our own serving setup were replaced with plain model names, the four older prompt pools are described in section 02 rather than shipped, and the package says where both things happened. If a number on this page disagrees with a file in the package, the file is right and the page is wrong; tell us and we will fix the page.
- data/README.md: what each file is, where it came from, and the places where the package will look like it argues with itself.
- data/NUMBERS.md: every figure on this page, with the file and the field it came from.
| 2026-09-03 to 05 | All captures taken: 65 one-shot conditions and 56 conversation conditions across 26 models (measured counts from the package's roster table), on hosted routes and one desktop. Collections abandoned or captured during a serving fault were listed in the exclusion file with reasons and never deleted. |
|---|---|
| 2026-09-05 | Final rebuild of every profile, band, comparison and roster table. The previous day's rebuild holds 57 of these comparison rows; 53 of them are byte-identical in the two builds, and the four that differ are rows whose captures changed between them, including the Qwen3-235B re-asks. A reproducibility spot-check, not a proof. |
| 2026-09-15 | OpenAI's API deprecations page read for section 10; this page written from the 2026-09-05 records. No rerun was made before publication; the gaps in section 08 are stated rather than closed. |
Every claim on this page is dated. The hosted rows describe endpoints on three days in September 2026; a later reader should assume every one of them has changed.