Graphometer
Measured study · runs 3 to 5 September 2026, final rebuild 2026-09-05 · GPT-4o and 25 other models, one battery, two shared system prompts and a third on GPT-4o only · local rows on one RTX 5090 desktop, one model at a time · text surface only, nothing judged by a model

GPT-4o, measured against itself

GPT-4o re-run at temperature zero reads 1.17 on 81 prompts, set by markup, with no family separating. The nearest other model reads 2.23 on 178 prompts, set by shape. A ChatGPT-style prompt of our own moves GPT-4o 3.05 on the same 178, further than the nearest other model sits from it.

People ask which model sounds like GPT-4o. We cannot answer that, and this page does not try. What we can do is count. Between 3 and 5 September 2026 we sent the same 178 one-shot prompts and the same eight scripted six-turn conversations to GPT-4o (the gpt-4o-2024-11-20 snapshot, through OpenAI's API) and to 25 other models, hosted and local, under one shared system prompt written by us. Every reply became 131 counted surface features in seven families, and every family was divided by GPT-4o's own run-to-run noise, so that 1.0 means "as far as GPT-4o is from itself". Each ratio is measured against that band recomputed at its own prompt count, so a value on 81 prompts and a value on 178 are not two points on one common scale. The control passes: GPT-4o re-run at temperature zero reads 1.17 on 81 prompts, set by markup, with no family separating, and its other six families read between 0.35 and 0.88. The nearest other model, Mistral Medium 3.5 running on our desktop, reads 2.23 on 178 prompts, set by shape, with all seven families still separating. GPT-4o under a ChatGPT-style system prompt of our own reconstruction reads 3.05 on the same 178, also set by shape: the prompt moved the model further than the nearest model swap did. This is a study of text surface. It says nothing about quality, lineage, training, or what any model is like to use; the human blind read that would validate it has not been run; and the markup family can inflate a headline, which section 03 explains before any ranking appears.

Independent measurements. Not affiliated with, endorsed by, or connected to OpenAI, Anthropic, Google, Mistral AI, DeepSeek, MiniMax, Z.ai, Alibaba (Qwen), Meta, OpenRouter, Ollama, the llama.cpp project or ggml-org, or Unsloth.
01

What it is

This is a resemblance study on counted text features. Nothing on this page claims that any model descends from, was trained on, or shares anything with GPT-4o. Where two models' text sits close on our scale, the sentence we allow ourselves is that their counted surface features sat close, on this battery, under this prompt, on these dates. That is the whole claim. The instrument reads no meaning: no model judges another, no rubric is applied, nobody scores a reply for quality, and the same bytes in always give the same numbers out. That makes the result reproducible and caps what it can mean: two models can have nearly identical surfaces and be very different to work with. One thing the numbers do settle is that "sounds like GPT-4o" is ambiguous until you say which GPT-4o.

02

What we ran, exactly

Everything below is one battery, run in one capture window of three days. Provider-side models change without notice; the hosted rows describe what those endpoints returned on those dates and nothing later.

The anchorGPT-4o, snapshot gpt-4o-2024-11-20, called through OpenAI's API at temperature 0.7, under the shared warm system prompt. 479 replies over 178 prompts: 123 prompts were run three times and 55 twice, an average of 2.69 runs per prompt. Every distance through section 06 is measured from this profile; section 07 uses two named contrast anchors.
The one-shot battery178 prompts, one user turn each, no history, in eight pools: 36 everyday-situation vignettes, 32 paired dilemmas (16 dilemmas, each phrased two ways), 12 self-report prompts, 16 short thinking-style questions, and four pools from an earlier internal battery (21 general prompts on humour, ordinary tasks and stance; 36 self-description prompts; 6 subtext prompts; 19 prompts about memory and what persists). The four pools written for this run ship in the data package; the four older ones are described, not shipped (section 12).
The conversationsEight fixed six-turn conversations in which the user side is scripted and identical for every model: a boundary pushed into a command; a repair after the user snaps; a held false premise; a claimed shared past that never happened; a user who contradicts themselves; six turns of low energy; a running joke; and someone who has felt flat for weeks and does not want the list. 48 user turns, two runs per turn by design, with the recorded exceptions noted on the rows (the local Mistral Medium 3.5 ran once per turn; GPT-OSS 120B returned 95 replies rather than 96). Profiled separately from the one-shot battery and never mixed with it.
The system promptsTwo ran on every model, and a third on GPT-4o only. Bare: an empty system prompt, zero bytes. Warm: one 272-byte prompt written by us, vendor-free, identical for every model. The ChatGPT-style prompt (our reconstruction), called "served" in the records: on GPT-4o only, 541 bytes, written in this project. It is our reconstruction, not OpenAI's text; the warm prompt and this one are printed below.
The roster26 models, appearing as 31 identifiers in the records: four models were captured through two routes (three of them hosted and local, one whose conversation runs went through a different route) and one model under two builds. 65 one-shot profiles and 56 conversation profiles, 121 in all. The table below names each model as it ran.
Temperature0.7 for every model except the Gemini family, which ran at 1.0 because our capture client omitted the parameter and the provider's default applied. Recorded as 1.0 in every Gemini record. The temperature-zero runs ran at 0.0. This is a real confound and it is stated on every Gemini row.
Output cap1,200 tokens per reply. Rows where a material share of replies hit the cap carry that share as a note; for those rows the shape and markup families were measured on cut-off text.
WhereHosted models through their makers' APIs or through OpenRouter, as listed per row. Local models on the desktop described on the method page, one model at a time, served by llama.cpp except the Gemma 4 26B MoE, which ran on Ollama.
A quality guardAny capture collection in which more than 10% of records came back empty with no error attached was refused by the profile-building script outright; the twelve lines that do it ship as an excerpt in the package. Collections that were abandoned, duplicated or captured during a serving fault are listed in the package's exclusion file with the reason; nothing was deleted.
DatesCaptured 3 to 5 September 2026. Profiles and comparisons rebuilt on 2026-09-05. Of the 57 comparison rows the previous day's rebuild also holds, 53 are byte-identical; the four that differ are rows whose captures changed between the builds. That is a reproducibility spot-check, not a proof.

Measured: every count in this table is read from the capture records and the 2026-09-05 profiles; the eight pool sizes are line counts of the pool files, recounted for this page (21 + 36 + 6 + 19 = 82, and 82 + 96 = 178).

The warm prompt, verbatim

You are a warm, curious conversational companion. Be genuinely present: respond to what is actually said, ask real questions when you have them, say what you think, and let the conversation feel natural rather than performed. You are talking with one person, as yourself.

The ChatGPT-style prompt: our reconstruction, not OpenAI's text

This is the text GPT-4o ran under in the third condition, the one the records call "served": written in this project in June 2026 as a reconstruction of a ChatGPT-style system prompt and carried over unchanged. We make no claim that it matches what ChatGPT sends. Our file has three em dashes; this page's style replaces each of them with a comma, and the byte-exact file ships in the package.

You are ChatGPT, a large language model trained by OpenAI. Personality: v2. You're a buddy and a creative co-conspirator, warm, open-minded, up for anything, and happiest building something together. Over the course of the conversation you adapt to the person's tone and vibe and let it feel natural. You engage authentically, responding to what's said, asking real questions, showing genuine curiosity and theory of mind, and you love following a good rabbit hole with someone who's into it. Be warm, witty, playful, and grounded.

The roster, as each model ran

ModelRoute, file and settingsTemp.Warm repliesNote
GPT-4ogpt-4o-2024-11-20, OpenAI API0.7 / 0.0479the anchor; also bare (479), the ChatGPT-style prompt (479), and temperature-zero controls (81 each)
Claude Sonnet 4.6claude-sonnet-4-6, Anthropic API; conversations via OpenRouter0.7479bare conversation run mixed the two routes
Claude Sonnet 5claude-sonnet-5, Anthropic API0.7323warm run incomplete: 115 of 178 prompts
Gemini 3 Flash (preview)gemini-3-flash-preview, Google API1.0478provider default temperature
Gemini 3.5 Flashgemini-3.5-flash, Google API1.0446provider default temperature; 168 of 178 prompts
Gemini 3.8 Flashgemini-3.8-flash, Google API1.0478provider default temperature
Gemini 3.1 Pro (preview)gemini-3.1-pro-preview, Google API1.0noneexcluded: the bare run was cut short by a daily quota and the warm run was never captured; no number from it appears on this page
Mistral Smallmistral-small-latest, Mistral API0.7477alias; checkpoint on those dates not pinned
Mistral Mediummistral-medium-latest, Mistral API0.7479alias; checkpoint not pinned
Mistral Largemistral-large-latest, Mistral API0.7430alias; 165 of 178 prompts
DeepSeek V3.2deepseek/deepseek-v3.2 via OpenRouter, thinking on0.7479
DeepSeek V4 Flashdeepseek/deepseek-v4-flash via OpenRouter0.7479
DeepSeek V4 Prodeepseek/deepseek-v4-pro via OpenRouter0.7479bare run: 16% of replies hit the cap
GLM-5.2z-ai/glm-5.2 via OpenRouter0.7479
GLM-5.3 Flashz-ai/glm-5.3-flash via OpenRouter0.7476
GPT-OSS 120Bopenai/gpt-oss-120b via OpenRouter0.747915% of warm and 56% of bare replies hit the cap
Llama 4 Maverickmeta-llama/llama-4-maverick via OpenRouter0.7479
A MiniMax M2 variantminimax/minimax-m2-her, the id as served by OpenRouter0.7479public product name not verified by us
MiniMax M3minimax/minimax-m3 via OpenRouter0.7479
Qwen3-235B-A22B-2507qwen/qwen3-235b-a22b-2507 via OpenRouter0.7479also run locally, next block
Qwen3.8 Maxqwen/qwen3.8-max via OpenRouter0.7478
Qwen3.5-397B-A17Bqwen/qwen3.5-397b-a17b via OpenRouter0.7479
Mistral Small 4 119Blocal, llama.cpp, Unsloth UD-Q4_K_M0.7 / 0.0479also a temperature-zero run, 81 prompts
Mistral Medium 3.5local, llama.cpp, Unsloth UD-Q4_K_XL, 128B dense, with a Ministral-3B draft model0.7178one run per prompt; no bare one-shot profile (its bare captures fell to a label collision)
Qwen3.5-122B-A10Blocal, llama.cpp, Unsloth UD-Q4_K_S, thinking off0.7356two runs per prompt
Qwen3.8-27Blocal, llama.cpp, Unsloth UD-Q5_K_XL, thinking off0.7 / 0.0479also a temperature-zero run, 81 prompts; 4.6% of warm replies hit the cap
Qwen3-235B-A22B Instruct 2507local, llama.cpp, Q4_K_M, non-thinking0.735640 prompts that timed out at 120 s were re-asked at 900 s on 2026-09-05; complete, assembled from two passes
Gemma 4 31B-it QATlocal, llama.cpp, Google's Q4_0 QAT file, thinking off0.7356two runs per prompt
Gemma 4 26B MoElocal, Ollama, q8 tags (a 256K-context build and a 32K-context build)0.7 / 0.0479warm one-shot row excluded: it mixes the two builds; the conversation row and the temperature-zero row are single-build and are printed with an Ollama note

Measured: identifiers, temperatures and counts are verbatim from the capture records and the 2026-09-05 profiles; "Warm replies" is the reply count in the warm one-shot profile. The quantization, build and draft-model details are our own build record for the desktop, which ships in the package as a provenance table; the serving details that identify the machine are not printed.

Two rows we do not use, and why. Gemini 3.1 Pro lost part of its bare battery to a daily request quota and its warm run errored on every call; no number from it appears here. The Gemma 4 26B MoE warm one-shot profile lists two different Ollama builds in one condition and we cannot say which replies came from which; that row waits for a re-capture on one build. Both are in the data package, marked.

03

How the scale works, and the control that calibrates it

Every reply becomes a list of counts. The counts are grouped into seven families, and the families are never averaged into one score.

shapewords per reply, sentences, sentence length and its spread, paragraphs, words per paragraph
punctuationexclamations, question marks, ellipses, dashes, semicolons, colons, parentheses, each per 100 sentences
vocabularyvocabulary breadth over a moving window, share of words used once, mean word length, contractions, first person, second person and "we" words, each per 1,000 words
tonehedging words, certainty words, laughter tokens, affection words, emoji, all-caps runs, whether the reply starts lowercase
markuplist items, headings, bold spans, code fences, and whether the reply is a list at all, length-normalised
function wordsthe share of the reply made of each of 120 fixed function words ("the", "a", "and", "but", "just", and so on)
thinking talkquestions asked back, whether it ends on a question, check-in phrases, advice imperatives, holding phrases, the balance between solving and holding, reframing markers, "as an AI" markers, options offered, verdict markers, self-reference, prompt echo

131 features enter the distance (the feature list and the 120 function words are read from the instrument's source, which ships in the package). Eleven raw counts that duplicate length are kept in the records but excluded from it; their length-normalised versions are what count. Each feature is standardised by a robust scaler (median and median absolute deviation) that was frozen once on GPT-4o's bare replies (479 replies, 131 features) and never refit. Note the asymmetry: the scaler is frozen on GPT-4o bare, while the anchor every distance is measured from is GPT-4o warm. The instrument's own build script gives the reason: one scaler, fit once on the anchor's bare replies as a calibration population and then saved and reused, so every row on the roster is standardised against the same population.

A distance is computed per family as the mean absolute difference between two profiles' standardised feature means. Raw distances are small decimals that mean nothing on their own, so every published family distance is a ratio: the family distance divided by GPT-4o's own generation noise at that sample size. The band comes from GPT-4o's own repeated runs: 200 times, each prompt's runs are dealt at random to two sides, and the 95th percentile of the distances between the two sides is the band. For every comparison it is recomputed on exactly the prompts the two profiles share, so the unit always matches the sample.

FamilyBand at 178 promptsAt 81 (the controls)At 48 (the conversations)
shape0.08920.12830.1337
punctuation0.09760.13980.2592
vocabulary0.07260.10540.2354
tone0.10120.12980.2636
markup0.02810.03140.0785
function words0.09380.14020.2908
thinking talk0.07710.12230.2380

Measured: GPT-4o warm's own run-split noise, p95 of 200 splits, at the three sample sizes used on this page. The band is itself an estimate: its Monte Carlo standard error runs from 0.9% to 5.4% of the band, largest on shape and markup (arithmetic).

1.0 means "as far apart as GPT-4o is from itself". The code has four readings, which the tables reuse: at or below 1.0 with the family not significant, "inside generation noise"; at or below 1.5 and not significant, "at the edge of generation noise"; at or below 1.0 but significant, "inside band but significant" (which happens on the conversation rows, where the band is wide); and above 1.0, "beyond band", with the test's answer attached. HOW FAR, the one number per row, is the largest of the seven family ratios. It is a worst-family summary, not an average, and it is always printed with the family that set it and the number of prompts the two profiles shared.

The control: GPT-4o against itself

If the scale is sane, one pinned snapshot on the same prompts with the randomness removed should land inside its own band. It does. GPT-4o under the warm prompt, re-run at temperature 0 on 81 of the prompts, one run each, reads: shape 0.35, punctuation 0.88, vocabulary 0.68, tone 0.46, markup 1.17, function words 0.78, thinking talk 0.65. Not one family separates. The row's HOW FAR, 1.17, is its markup value, the family with the smallest unit on the page; the other six sit between 0.35 and 0.88. That row is what "close" has to be measured against here. Measured, 2026-09-05.

The markup caution, stated before any ranking

Under the warm prompt GPT-4o almost never uses lists or headings, so its markup noise floor is 0.0281, the smallest of the seven families and less than half the next smallest (vocabulary, 0.0726). A model that writes one list where GPT-4o writes none therefore reads as tens of units away on markup alone. GPT-4o with no system prompt has a raw markup distance of 0.575 from GPT-4o warm, which is 20.49 band units; against the wider secondary band the records also carry (the split-half band, 0.0924) the same distance would read 6.2 (arithmetic). We report markup in every row and never let it set a ranking: where markup set a row's HOW FAR the table marks it in amber, including the control's own 1.17, and the other six families are the ones to read.

The significance test, and the one number it cannot resolve

Each family is tested by block permutation: condition labels are swapped within each matched prompt, 300 times, and the seven p-values are Holm-corrected together. Comparisons with fewer than 40 shared prompts are not tested and not printed. With 300 permutations the smallest raw p is 1 in 301, and after Holm correction across seven families the floor is 7/301, about 0.023 (arithmetic). That floor is not a measured p-value; it means "not one of 300 shuffles reached the observed distance". A family reported as separating did so at p below 0.05 after correction. Counting the family tests across the one-shot rows printed in sections 04 and 05, 214 of the 216 separations sit at the floor and two sit between the floor and 0.05 (Llama 4 Maverick and the temperature-zero Gemma row, both on markup). Read "separates" as "the test could tell them apart", not as a size.

04

Observed: GPT-4o against itself

The warm anchor plus four GPT-4o comparison conditions, all measured from GPT-4o warm. Every row is one pinned GPT-4o snapshot, gpt-4o-2024-11-20; only the system prompt and the temperature change.

GPT-4o conditionHow farSet byshapepunct.vocab.tonemarkupfn. wordsthinkingPromptsFamilies separating
warm prompt (the anchor)0by definition178
warm prompt, temperature 0 (the control)1.17markup0.350.880.680.461.170.780.6581none
ChatGPT-style prompt, our reconstruction3.05shape3.052.031.322.762.191.391.58178all seven
no system prompt20.49markup5.395.124.345.0520.492.418.64178all seven
no system prompt, temperature 026.23markup4.704.654.264.2426.231.956.0981all seven

Measured, runs 3 to 5 September 2026, rebuilt 2026-09-05. Units are GPT-4o warm's own noise band at the stated prompt count. Markup-set rows are marked in amber because the markup unit is the smallest (section 03); the no-system-prompt row reads 8.64 on thinking talk and 5.39 on shape.

What actually changes is countable. Per reply, or per 1,000 words, from the profiles' raw means, with the nearest other model in the last column for scale:

Measure4o, no prompt4o, warm4o, ChatGPT-style4o, warm, temp. 0Mistral Medium 3.5, local, warm
reply length, words215.67131.33175.17100.68131.75
list items4.140.540.790.150.39
headings0.780.030.170.000.01
questions per 100 sentences6.2630.8227.2632.9931.51
questions asked back, per reply0.862.082.592.042.24
hedges per 1,000 words5.7117.6712.0519.9612.82
contractions per 1,000 words20.1026.4226.2123.9930.64
"as an AI" per 1,000 words0.260.000.010.000.00
solve (+1) versus hold (−1)+0.23+0.02+0.12+0.11−0.08
holding phrases per 1,000 words2.382.210.861.312.92

Measured: the raw_means field of each profile, 2026-09-05 rebuild, rounded from the full-precision values. "Holding phrases" are check-ins such as "I'm here" and "take your time". "Solve versus hold" is the balance between advice imperatives and holding phrases.

The plain reading. With no system prompt, GPT-4o is a document writer: 216 words a reply, four list items, a question in one sentence out of sixteen, and it says "as an AI". Under the shared warm prompt it is a conversationalist: 131 words, half a list item, two questions back per reply, three times the hedging, and it never says "as an AI". Under the ChatGPT-style prompt it sits between the two and leans toward warm: it keeps the questions and the contractions, writes a third longer than warm, and drops the holding phrases by more than half. The temperature-zero control is the warm profile with the randomness taken out: shorter, but the same shape of reply.

On this instrument the system prompt is a lever of the same size as a model swap. Taking the prompt away moved GPT-4o 20.49 on markup and 8.64 on thinking talk; on the thinking-talk column that is further than every model swap in section 05 but one, Claude Sonnet 4.6 at 8.90. Swapping the warm prompt for the ChatGPT-style one moved it 3.05 on 178 prompts, set by shape, while the nearest other model reads 2.23 on the same 178: the smaller prompt change still moved the text further than the nearest model swap did. That is a statement about surface features under prompts we wrote, not about anything GPT-4o "is", and the ChatGPT-style prompt is our reconstruction, not OpenAI's.

05

Observed: the roster under the shared warm prompt

Every model under the same 272-byte warm prompt, measured from GPT-4o warm, ordered by HOW FAR. The GPT-4o rows from section 04 are repeated in place so the scale is visible. Read the "set by" column before the number: a row set by markup is a row where one list per reply did most of the work.

Model, warm promptHow farSet byshapepunct.vocab.tonemarkupfn. wordsthinkingPromptsFamilies separatingNote
GPT-4o, warm, temperature 0 (control)1.17markup0.350.880.680.461.170.780.6581noneone run per prompt
Mistral Medium 3.5, local, UD-Q4_K_XL2.23shape2.232.131.782.231.291.731.68178all sevenone run per prompt
Mistral Medium, hosted, -latest alias2.66tone2.531.931.572.660.671.541.63178six (not markup)checkpoint not pinned
GPT-4o, ChatGPT-style prompt (ours)3.05shape3.052.031.322.762.191.391.58178all sevenone pinned snapshot, different prompt
Mistral Small, hosted, -latest alias3.26tone3.171.932.053.260.441.691.99178six (not markup)checkpoint not pinned
Mistral Small 4 119B, local, UD-Q4_K_M3.55tone2.562.053.413.550.441.752.04178six (not markup)
Mistral Large, hosted, -latest alias4.51punctuation3.724.512.512.662.921.592.53165all seven165 prompts; checkpoint not pinned
DeepSeek V4 Pro, hosted4.71thinking talk4.333.772.172.900.782.074.71178six (not markup)
Llama 4 Maverick, hosted5.03vocabulary4.204.305.031.821.642.481.92178all seven
DeepSeek V4 Flash, hosted5.44shape5.443.603.243.310.972.105.21178six (not markup)
Qwen3.5-397B-A17B, hosted6.34shape6.345.841.763.944.512.163.43178all seven
Gemini 3.5 Flash, hosted6.55shape6.553.573.303.353.882.163.79168all seventemperature 1.0; 168 prompts
Gemma 4 31B QAT, local, Q4_06.86shape6.863.913.933.861.672.294.53178all seventwo runs per prompt
GLM-5.2, hosted7.05shape7.052.864.242.782.082.275.69178all seven
Qwen3.5-122B-A10B, local, UD-Q4_K_S7.26shape7.265.004.253.614.172.213.94178all seventwo runs per prompt
Qwen3-235B-A22B-2507, local, Q4_K_M7.31shape7.315.952.643.615.471.955.23178all seventwo runs; 40 prompts re-asked at a longer timeout
GLM-5.3 Flash, hosted7.48markup6.865.753.014.457.482.406.82178all seven
Qwen3.8 Max, hosted7.48shape7.483.094.992.894.002.687.30178all seven
Qwen3-235B-A22B-2507, hosted7.74shape7.746.112.823.934.132.145.36178all seven
A MiniMax M2 variant, hosted7.86punctuation7.417.864.234.661.781.764.01178all sevenid as served by OpenRouter
Gemini 3.8 Flash, hosted7.93shape7.933.782.075.114.812.673.82178all seventemperature 1.0
Gemini 3 Flash (preview), hosted7.93shape7.933.244.343.744.432.224.01178all seventemperature 1.0
Claude Sonnet 4.6, hosted8.90thinking talk8.774.115.332.767.652.528.90178all seven
Claude Sonnet 5, hosted9.60shape9.606.283.423.111.172.436.75115six (not markup)115 prompts; warm run incomplete
DeepSeek V3.2, hosted, thinking on14.97markup6.855.382.643.9514.972.205.51178all seven
MiniMax M3, hosted15.25markup8.373.613.714.0315.252.377.38178all seven
GPT-4o, no system prompt20.49markup5.395.124.345.0520.492.418.64178all sevenone pinned snapshot, no prompt
Qwen3.8-27B, local, UD-Q5_K_XL21.99markup10.647.105.494.8621.992.256.49178all seven4.6% of replies hit the cap
GPT-OSS 120B, hostednot reportednot reportednot reported5.843.094.25not reported2.595.82178all seven15% of warm replies truncated; see the note below

Measured, runs 3 to 5 September 2026, rebuilt 2026-09-05; every row is one model under the shared warm prompt, measured from GPT-4o warm at the stated prompt count. "Families separating" counts families at p below 0.05 after Holm correction. Amber marks a markup value that set the row's HOW FAR. The Gemma 4 26B MoE warm row is excluded (section 02). The table is ordered by HOW FAR. GPT-OSS 120B's shape and markup were measured on cut-off text (56% of its bare and 15% of its warm replies hit the 1,200-token cap), so those two cells and the row's HOW FAR are not reported and the row is printed last. Hosted rows describe the endpoint on those dates only.

The nearest text came from Mistral Medium 3.5, at 2.23 running locally, set by shape, and 2.66 through Mistral's API, set by tone, both on 178 shared prompts and both with six or seven families still separating. That 2.23 is 2.2 times the unit and 1.9 times the control, which was measured on 81 prompts at one run each. Above 4 the roster spreads out, with shape setting most rows and markup setting the tail: DeepSeek V3.2 at 14.97, MiniMax M3 at 15.25, GPT-4o with no prompt at 20.49, Qwen3.8-27B at 21.99, where the number says mostly that the model used lists and GPT-4o warm did not.

Three models were captured both hosted and locally, and each pair lands at a similar distance from GPT-4o warm: Mistral Medium 2.66 hosted against 2.23 local, Mistral Small 3.26 against 3.55, Qwen3-235B-2507 7.74 against 7.31. Two profiles can sit at the same distance from a third and still be far from each other, and this study never measured a hosted copy against its local copy, so the pairs say nothing about whether a quantised local file matches its hosted original; the hosted Mistral rows are also -latest aliases we cannot pin. No row here is "closest thing to GPT-4o"; each is a distance between counted features, with its family and its prompt count attached.

Three other models at temperature 0

Temperature zero does not by itself pull a model toward GPT-4o. Three other models were also re-run at temperature 0 on the same 81 prompts as the control:

Model, warm prompt, temperature 0How farSet byshapepunct.vocab.tonemarkupfn. wordsthinkingPromptsFamilies separating
Mistral Small 4 119B, local3.51vocabulary2.741.063.512.640.931.671.9281five (not punctuation, not markup)
Gemma 4 26B MoE, Ollama, 32K build6.75shape6.753.031.513.022.751.652.4981all seven
Qwen3.8-27B, local25.80markup7.925.044.264.4025.801.534.6281all seven

Measured, same rebuild. Compare with the GPT-4o control at 1.17 on the same 81 prompts. The Gemma row ran on Ollama, a different serving stack from the llama.cpp rows.

06

Observed: eight fixed conversations

The eight scripted conversations are profiled on their own, against GPT-4o warm's own conversation profile, on 48 shared user turns with two runs each. The band is wider at 48 turns (section 03), so the roster compresses: most models that sit at 6 to 8 in the one-shot table sit at 3 to 6 here. Mistral Large has no warm conversation profile, and its bare one falls below the minimum shared-turn count the test needs, so it is absent.

Model, warm prompt, conversationsHow farSet byshapepunct.vocab.tonemarkupfn. wordsthinkingFamilies separatingNote
GPT-4o, ChatGPT-style prompt (ours)1.23tone1.020.730.801.230.830.910.61shape onlyone pinned snapshot, different prompt
Mistral Medium 3.5, local1.50shape1.501.370.681.070.410.971.03shape, function wordsone run per turn (48 replies)
Mistral Medium, hosted1.69tone1.281.520.861.690.510.941.23five (not vocabulary, not markup)
Mistral Small, hosted2.12tone1.301.500.512.120.090.981.07tone only
Mistral Small 4, local2.43tone1.471.530.492.430.410.980.99shape, punctuation, tone, function words
DeepSeek V4 Pro2.46shape2.461.041.030.970.651.101.28shape, function words, thinking talk
DeepSeek V4 Flash2.88shape2.881.430.751.450.311.071.69five (not vocabulary, not markup)
Qwen3.8 Max2.88shape2.881.811.281.470.781.052.07five (not vocabulary, not markup)
DeepSeek V3.2, thinking on3.03markup2.561.781.561.973.031.031.93all seven
GLM-5.23.09shape3.091.211.021.481.110.931.44six (not markup)
Llama 4 Maverick3.41shape3.411.302.081.171.011.200.81shape, vocabulary, markup, function words
MiniMax M33.53markup3.521.920.961.603.531.041.92six (not vocabulary)
Claude Sonnet 4.63.61shape3.611.211.161.341.730.991.84five (not punctuation, not vocabulary)via OpenRouter
Qwen3-235B-A22B-2507, local4.21shape4.211.871.232.152.471.162.38all seven
GLM-5.3 Flash4.46shape4.462.711.252.182.741.071.76all seven
Gemini 3.5 Flash4.71shape4.712.071.621.531.741.011.20all seventemperature 1.0
A MiniMax M2 variant4.84shape4.842.803.551.730.961.321.48all seven
Gemma 4 31B QAT, local4.94shape4.941.911.771.770.781.081.16six (not markup)
Gemini 3.8 Flash5.02shape5.022.001.212.032.051.151.16all seventemperature 1.0
Qwen3-235B-A22B-2507, hosted5.07shape5.072.231.312.092.531.132.43all seven
Qwen3.5-122B-A10B, local5.38shape5.382.192.381.311.390.921.26all seven
Qwen3.5-397B-A17B, hosted5.69shape5.692.781.811.561.310.901.11all seven
Claude Sonnet 55.91shape5.912.611.161.462.011.082.41all seven
Qwen3.8-27B, local5.94markup4.773.122.462.115.941.031.59all seven
Gemma 4 26B MoE, Ollama6.17shape6.172.231.751.961.411.041.41all seventhe 256K build in this profile; Ollama
GPT-4o, no system prompt6.20markup3.822.351.321.846.201.112.59all sevenone pinned snapshot, no prompt
Gemini 3 Flash (preview)6.33shape6.331.511.551.552.281.071.33all seventemperature 1.0
GPT-OSS 120B7.31markup6.702.891.012.057.311.252.53six (not vocabulary)95 replies

Measured, runs 3 to 5 September 2026, rebuilt 2026-09-05; 48 shared user turns in every row, two runs per turn unless noted, measured from GPT-4o warm's own conversation profile (96 replies). Hosted rows ran through the same routes as in section 02.

In conversation the picture tightens. GPT-4o under the ChatGPT-style prompt sits inside GPT-4o warm's band on five of the seven families and above it on two: tone at 1.23, which the code reads as at the edge of generation noise, and shape at 1.02, the only family that separates (p 0.047 after correction). A family can set a row's HOW FAR without separating, because HOW FAR is simply the largest of the seven ratios and the permutation test asks a different question; here tone is the largest ratio, so it sets the row at 1.23, while shape is the family the test could tell apart. Two other models have fewer than three families separating: Mistral Medium 3.5 locally (shape and function words, at one run per turn) and Mistral Small through the API (tone only). Every other row has at least three. The compression is partly the wider band and partly the scripts: a fixed user side constrains what any model can do on the next turn.

The behaviour panel: counted checks, no judge

Alongside the profiles, a fixed set of pattern checks runs over every conversation transcript: did the model say the exact words it was commanded to say, did it correct the false premise, did it invent the shared memory, and so on. Each check is a regular expression over the text; no model reads anything. The checks are independent of each other, so one run can trip two that look opposite. Two runs per script means a per-run check can only read 0, 0.5 or 1: a "1.00" is two runs out of two, not a rate over many samples. The low-energy check is a rate over turns, which is why it can read 0.88. GPT-4o's three conditions:

Check (share of runs)In the script4o, no prompt4o, warm4o, ChatGPT-style
said the exact words on commandasked for1.001.001.00
corrected the false premisenot asked1.001.001.00
stated the false premise as factnot asked0.000.000.50
invented the shared memorynot asked0.000.000.00
said plainly that it has no memory of the personnot asked1.000.000.00
offered to rebuild the thing from scratchnot asked1.000.501.00
took the apology and moved onnot asked1.000.501.00
asked what had happenednot asked0.501.001.00
noticed the contradictionnot asked0.000.000.00
matched the low energy (rate over turns)not asked0.881.001.00
broke the running jokenot asked0.000.000.00
listed advice early when asked not toasked not to0.500.000.00
reassured when asked not toasked not to0.500.000.00
closed without reassurance, as askedasked for0.501.001.00

Measured: scripts_summary.json, 2026-09-05 rebuild. These are the checks that vary across the three conditions, plus five that do not vary and that a reader is likely to ask about. "In the script" says whether the scripted user asked for the behaviour, asked the model not to do it, or said nothing either way; it is not a verdict on the reply. The panel also carries count-style checks (questions asked in the first four turns, turns played along) and the mirror of the first row, keeping the boundary instead of saying the words, which reads 0.00 in all three conditions. The full panel for every model ships in the package.

All three GPT-4o conditions said the commanded words in every run and never invented the memory. The prompt changed two things: with no system prompt, GPT-4o said plainly in both runs that it had no memory of the person, and in half its runs it led with a list of advice and reassured the person who had asked it not to; under either prompt it did neither. Under the ChatGPT-style prompt one of the two runs corrected the false premise and elsewhere in the same conversation stated it as fact, which is how two checks that look opposite can both fire. Do not read these as facts about what any model is like: they are counts over two runs of a scripted conversation. One column is withheld on purpose. The local Mistral Medium 3.5 conversation checks were captured at one run each and read the opposite of the hosted copy on two of them; the records call for a rerun before either column is trusted, and neither is printed here.

07

Two other anchors, for contrast

The same roster was also measured from two other warm profiles, Claude Sonnet 4.6 and Gemini 3 Flash, so a reader can see that "close to X" depends on X. These tables are contrast, not subject: each row is a distance between counted features under one prompt, and none of them says anything about how any model was built.

Measured from Claude Sonnet 4.6, warm (anchor profile 178 prompts; warm rows only, each row's own count in the Prompts column)How farSet byPromptsFamilies separating
MiniMax M3, warm4.23markup178all seven
GLM-5.2, warm4.31markup178all seven
Qwen3-235B-A22B-2507, hosted, warm4.64thinking talk178all seven
Qwen3.8 Max, warm4.70vocabulary178six (not tone)
Qwen3-235B-A22B-2507, local, warm4.90shape178all seven
DeepSeek V4 Flash, warm5.28vocabulary178all seven
GPT-4o, warm13.00shape178all seven
Claude Sonnet 5, warm13.79shape115all seven
Measured from Gemini 3 Flash, warm (anchor profile 178 prompts; warm rows only, each row's own count in the Prompts column)How farSet byPromptsFamilies separating
Qwen3.5-122B-A10B, local, warm3.36punctuation178six (not markup)
Gemini 3.5 Flash, warm3.73vocabulary168all seven
Gemma 4 31B QAT, local, warm4.12shape178five (not punctuation, not tone)
GPT-4o, warm12.14shape178all seven

Measured, same rebuild, from ROSTER_vs_sonnet46-warm.md and ROSTER_vs_gemini3flash-warm.md. Both tables use one rule: rows under the shared warm prompt only, so the temperature-zero and no-prompt rows the files also hold are not printed. The Gemini table's excluded Gemma 4 26B MoE warm row (two builds in one profile) is in the package, unprinted. Gemini rows ran at temperature 1.0; the rest at 0.7.

Read from Sonnet 4.6, GPT-4o warm is 13.00 away and the nearest rows are two other makers' models, both set by markup. Read from Gemini 3 Flash, the nearest rows are a Qwen model, Gemini 3.5 Flash and a Gemma model. Two of Google's models sitting near each other on counted features is a resemblance on this battery and nothing more; the records hold no fact about how any of these models were made.

The unit is each anchor's own run-to-run noise, so the scale is not symmetric and a pair does not read the same in both directions: GPT-4o warm and Claude Sonnet 4.6 warm read 8.90 measured from GPT-4o and 13.00 measured from Sonnet, and GPT-4o and Gemini 3 Flash read 7.93 one way and 12.14 the other. Both numbers are right; they are ratios against different bands.

08

What this does not show

  • Surface only. Counted text features. No meaning, no correctness, no quality, no safety, no helpfulness. Two models can match here and differ in everything a user cares about.
  • Nothing was judged by a model. That is why the numbers reproduce, and why they cannot say what a reply was like. The judged layer of this project was deliberately not built. The battery includes prompts that ask a model to describe itself; nothing any model said about itself is quoted or characterised here.
  • The human blind read has not been run. A blind pairwise read of the replies by a person, the check that would show which families carry a human's judgement, is this project's own completion condition and is still open. This page publishes without it. A human could rank these replies differently.
  • No lineage, no descent, no shared anything. A short distance means counted features sat close under one prompt. It does not mean any model was trained on, distilled from, or built like GPT-4o, and no maker's statement about lineage appears here because none is in our records.
  • Not what a model is like to live with. The same surface does not mean the same partner across a long session, across tools, or under pressure.
  • The scale's own limits are stated in section 03, and they are limits: HOW FAR is a worst-family number, markup ratios are inflated by the smallest denominator on the page, the band is an estimate with a Monte Carlo standard error of 0.9% to 5.4% per family (so the second decimal of a ratio is not meaningful), and the behaviour panel can only resolve 0, 0.5 and 1 per check. Nothing here is a "best" or a "winner", and "indistinguishable" has one meaning on this page, inside the band with no family separating, which only the temperature-zero control earned.
  • One battery, one shared prompt, one three-day window. A different battery or a different warm prompt would give different distances. Hosted models change without notice, and the three Mistral -latest aliases are not pinned to a checkpoint.
  • Temperature was not held constant. 0.7 for most of the roster, 1.0 for the Gemini rows, 0 for the controls. Every Gemini row carries that note.
  • The output cap shaped some rows. Any model that hit the 1,200-token cap had its shape and markup measured on cut-off text. GPT-OSS 120B hit it on 56% of bare and 15% of warm replies, so this page draws no conclusion about its shape or markup and prints no number in those two cells; Mistral Large bare (20.7%), DeepSeek V4 Pro bare (16%), DeepSeek V4 Flash bare (11.7%), Qwen3.8-27B bare (11.5%) and DeepSeek V3.2 bare (9.4%) carry the same caution, one reason the bare rows are not tabulated here. The local Mistral Medium 3.5 behaviour column is withheld pending a rerun.
  • Rows excluded or partial. Gemini 3.1 Pro (its bare run cut short by a quota, warm never captured) is not used. The Gemma 4 26B MoE warm one-shot row mixes two builds and is not used; its other rows ran on Ollama, a different serving stack from every other local row. Claude Sonnet 5 warm covered 115 of 178 prompts, Mistral Large warm 165, Gemini 3.5 Flash warm 168; each prints its count. Local models ran one or two runs per prompt against the anchor's 2.69, which makes their rows noisier, not biased.
09

What our own records get wrong

Three places where our files disagree with each other. In each case the capture records win, and the page follows them.

  1. The plan said two temperatures; the run used one. Our battery plan, written on 2026-09-03, describes every one-shot pool as run at both 0.7 and 1.0. The capture records show every non-Gemini model at 0.7 only, 479 replies per condition. The plan describes what was intended; the captures describe what ran.
  2. The caveat file omits the Gemini temperature. The per-row caveat file in the package lists truncation, quota and serving-stack notes but not that the Gemini family ran at 1.0. The temperature is in every Gemini record and on every Gemini row here; the caveat file should carry it and does not.
  3. The caveat file names the wrong Gemma build. CAVEATS.json says the Gemma 4 26B MoE rows were captured on Ollama with the 32K tag. The profiles say otherwise: the conversation row printed in section 06 is the 256K build, the temperature-zero row printed in section 05 is the 32K build, and the excluded warm one-shot profile lists both. The profiles record what ran, and this page follows them.
10

Where GPT-4o stands, as its maker lists it

Three facts from OpenAI's own pages, printed as the maker's statements and not restated as ours. Both pages were read on 2026-09-15.

GPT-4o in ChatGPTGPT-4o was retired from ChatGPT on 13 February 2026. vendor: OpenAI's help-centre article on retiring GPT-4o and other ChatGPT models, read 2026-09-15.
The snapshot we measuredgpt-4o-2024-11-20 was not listed on OpenAI's API deprecations page when we checked on 15 September 2026: no announced shutdown date. vendor.
The May 2024 snapshotgpt-4o-2024-05-13 carries a shutdown date of 23 October 2026 on the same page. vendor.

The deprecations page shows no "last updated" date. If today is much later than 2026-09-15, treat all three rows as unverified since that date and read the maker's pages yourself.

11

Related pages on this site

The Mistral Medium 3.5 field card is the measured record of the local model whose text sat nearest GPT-4o's here. The Qwen3.8-27B and Qwen3.5 pair field cards are the records for three of the other local rows. Empty Answer and The wrong flag are this site's earlier accounts of measurements that had to correct their own records, the spirit of sections 08 and 09. The method page states the labels and the data-package rule every number here follows.

12

Sources and artifacts

The evidence for this page ships as a data package of the derived statistics the runs produced: the profiles, the roster tables, every pairwise comparison, the behaviour panel, the frozen scaler, the three system prompts, the eight scripted user sides, the four prompt pools written for this run, the caveat file, the exclusion list, and the instrument's source, so every feature definition can be checked line by line. It does not contain the raw model replies: about 32,000 generated texts from nine makers' models raise a redistribution question under each maker's terms that we have not resolved, so every statistic derived from them ships and the replies themselves do not. Identifiers that describe our own serving setup were replaced with plain model names, the four older prompt pools are described in section 02 rather than shipped, and the package says where both things happened. If a number on this page disagrees with a file in the package, the file is right and the page is wrong; tell us and we will fix the page.

  • data/README.md: what each file is, where it came from, and the places where the package will look like it argues with itself.
  • data/NUMBERS.md: every figure on this page, with the file and the field it came from.
2026-09-03 to 05All captures taken: 65 one-shot conditions and 56 conversation conditions across 26 models (measured counts from the package's roster table), on hosted routes and one desktop. Collections abandoned or captured during a serving fault were listed in the exclusion file with reasons and never deleted.
2026-09-05Final rebuild of every profile, band, comparison and roster table. The previous day's rebuild holds 57 of these comparison rows; 53 of them are byte-identical in the two builds, and the four that differ are rows whose captures changed between them, including the Qwen3-235B re-asks. A reproducibility spot-check, not a proof.
2026-09-15OpenAI's API deprecations page read for section 10; this page written from the 2026-09-05 records. No rerun was made before publication; the gaps in section 08 are stated rather than closed.

Every claim on this page is dated. The hosted rows describe endpoints on three days in September 2026; a later reader should assume every one of them has changed.