Five local coding models, four points apart
Qwen3.8-27B, Ornith-1.5-35B, Ornith-1.5-397B, DeepSeek V4 Flash and GLM-5.3-Flash ran four agentic coding tasks through OpenCode and scored 387 to 391 of 400. A re-run of one task moved one model 37 points, and one judgment question spread the five by 9 of 40.
On 16 September 2026 we ran five open-weight models, each served by llama.cpp on our own hardware, through the same four coding tasks in OpenCode, the agentic coding tool, and scored every run out of 100: 60 points checked by a script and 40 awarded by a hosted model grading blind. The mechanical half separated nobody: each of the twenty first scored runs took all 60 points. The graded half separated the five by 4 points out of 160, and it was both noisy and steered: when the grader re-read identical work, single-task grades moved by up to 2, and its written instructions told it that most competent work lands between 30 and 37 of 40. One model's second run of one task, four hours after its first, fell from 99 to 62 over a single line in the task's material. One open question, asked once and answered without tools, spread the same five from 39 to 30 of 40. The finding is about our instrument: with a mechanical half that every run maxed out and a graded half that was noisy and instructed into a narrow band, this test could not rank these models. It does not show that they are equally good, and it does not show which is best. It is one scored run per model per task, four tasks, on our own hardware, and section 10 says what that rules out. The trial also turned up a set of harness traps, each with a cheap guard (section 08), and several statements in our own notes that the run files contradict (section 09).
Summary
We built this trial to choose between local models for agentic coding work in OpenCode. Five open-weight models ran four tasks each, one model at a time, on one desktop (one of the five was split across that desktop and a laptop). Every figure below is measured from the files the runs wrote unless it is labelled otherwise; section 02 says what the labels mean.
Five results:
- The scoreboard is flat. Totals out of 400: Qwen3.8-27B 391, Ornith-1.5-35B 390, Ornith-1.5-397B 390, DeepSeek V4 Flash (split across two machines) 388, GLM-5.3-Flash 387. All twenty first scored runs took 60 of 60 mechanical points, so the whole spread sits in the graded half, 151 to 147 of 160. On task B all five made byte-identical changes and all five scored 99; tasks A, C and D spread the five by 3, 2 and 3 points.
- The graded half was noisy and held to a narrow band. Four runs graded twice on identical evidence moved by 0, 0, 1 and 2 points out of 40. And the grader's written instructions said: "Most competent submissions land 30-37; reserve 38+ for work with no reviewer objections at all." The twenty first-run grades fell between 34 and 39.
- A second run moved one model 37 points. The two-machine DeepSeek's re-run of task B scored 62 against its first run's 99. Its repairs were byte-identical to the first run's; its index file had one extra row, and 7 of 23 checks failed. The extra row came from the same line our own answer key had wrongly counted earlier that day.
- One question spread them further than four tasks did. One open judgment question, one answer per model, no tools, graded blind on its own 40-point rubric: 39, 37, 35, 35 and 30. The top and bottom are 9 points apart; the middle three sit within 2 of each other. GLM-5.3-Flash, the top score, was asked for high reasoning effort; the other four were not.
- Time is where they differed. The same four tasks took 448 seconds of wall time on Qwen3.8-27B and 6,058 on GLM-5.3-Flash, 13.5 times as long (arithmetic), for totals 4 points apart.
Section 07 turns this into five checks for anyone building a small evaluation of their own. Section 08 lists the harness traps we hit and the guard for each, including an OpenCode startup wait that cost two hours of machine time and that we checked against OpenCode's public issue tracker on 26 September 2026. Section 09 corrects our own records, one of which had the wrong timeline for a stray process.
What we ran, exactly
Everything below is a property of these five model files, these serving settings, this harness and these four tasks, on our own hardware, on 16 September 2026 (local time, UTC−4). Nothing here measures a model's general coding ability, and no leaderboard claim appears on this page.
| Harness | OpenCode 1.18.31, driven headlessly: opencode run --pure --auto --format json --print-logs, one model and one task per run, each run in a fresh disposable git sandbox. Every run that reached a session logged version=1.18.31. --auto approves every permission not explicitly denied; this build had no finer command-line switch, so the disposable sandbox is what made that safe. Web fetch was denied through OPENCODE_PERMISSION, best effort and not verified end to end. |
|---|---|
| Scoring | 100 points per task. Mechanical, 60, by our scoring script: tests or checks finish green 30, no malformed tool call 10, nothing touched outside the sandbox 10, finished inside the 45-minute budget 5, ran to completion 5. Graded, 40, blind: diff quality 15, honesty (does the final message match the diff and the test output) 10, scope discipline 10, an explanation a non-coder can use 5. |
| Grader | A hosted model, claude-sonnet-5, named as the grader in 19 of the 22 coding grade files and in all five judgment-question grade files; three first-round coding grade files (Qwen3.8-27B on tasks A, C and D) do not record who graded them. It saw the task prompt, the model's final message, the diff and the test output, with the model's identity stripped; after a leak described in section 08, it saw them inside a folder whose name meant nothing. We publish its numbers, never its written reasons. Its written instructions steered it toward a narrow band. The version we kept says: "Most competent submissions land 30-37; reserve 38+ for work with no reviewer objections at all." Every coding grade file follows that document's output format rather than the shorter one built into the grading prompt file, but the document was last edited at 09:37 on 16 September, after the first-round grades, and earlier versions were not kept. It ends with a table of first-round scores that names the models; the records do not show whether graders were given that table. |
| Budget | 45 minutes of wall time per task for the 5 budget points. A hard stop at 60 minutes in the first round, raised to 110 minutes for the three slower models as a safety net only; the 45-minute line did not move. Every scored run finished inside 45 minutes; the longest took 2,411 seconds. |
| Runs | One scored run per model per task. Two runs that never reached the model were voided and run again; one run that died before it could be scored was run again; two of DeepSeek's tasks were run a second time on purpose, and both runs are printed. |
| Machine | One desktop with an NVIDIA GeForce RTX 5090 (32,607 MiB of video memory) and 188 GiB of RAM. DeepSeek V4 Flash was split between that desktop and an ASUS ROG Flow Z13 (128 GB unified memory) over Thunderbolt, using llama.cpp's RPC backend. The three servers the trial started ran one request at a time (--parallel 1 in their launch lines; the Ornith-1.5-35B server log shows one slot). The trial's logs do not record that flag for the two installed models. |
The four tasks
- Task A: a bug fix in a pruned copy of a private web service. Make one failing test pass without breaking a neighbouring behaviour.
- Task B: repair three planted defects in a synthetic corpus of conversation transcripts, then build an index file, to an exact specification, from evidence spread across two file formats.
- Task C: add one command-line option to an internal tool and document it, leaving everything the tool already printed byte-identical.
- Task D: the small ledger utility from our July agent benchmark, reused verbatim. Diagnose a seeded bug, try a maintainer's narrow fix first and show why it fails, fix the domain logic, add per-category totals and a JSON output mode across several files, add tests, and review the diff.
Tasks A to C come from our own work and are private: we describe them and publish nothing from them. Task D's fixture and prompt are in the data package.
The five models as served
| Model | File | Size on disk (bytes) | Window | Where the weights sat | Sampling |
|---|---|---|---|---|---|
| Qwen3.8-27B | Unsloth UD-Q5_K_XL, 1 file | 20,218,178,624 | 131,072 | not in the trial's records (a dense model) | its installed start script's; not in the trial's records |
| Ornith-1.5-35B | publisher's Q6_K, 1 file | 29,208,731,392 | 131,072 | card, with the experts of 6 layers in system RAM | temperature 0.6, top-p 0.95, top-k 20 |
| Ornith-1.5-397B | community IQ3_XXS (importance matrix), 1 file | 162,712,290,016 | 131,072 | the experts of 60 layers in system RAM | temperature 0.6, top-p 0.95, top-k 20 |
| DeepSeek V4 Flash (0731) | Unsloth UD-IQ3_XXS, 4 files | 104,207,848,032 | 262,144 | split across the desktop and the laptop over Thunderbolt | its installed start script's; not in the trial's records |
| GLM-5.3-Flash | Unsloth UD-Q3_K_XL, 4 files | 147,535,921,955 | 131,072 | all experts in system RAM, on a llama.cpp branch build with GLM-5.3 support | temperature 1.0, top-p 0.95 |
Sizes: read with stat on 26 September
2026 (measured). Windows: the server's own report for Qwen3.8-27B and
DeepSeek, the launch flag for the other three (measured, from the
round logs and launch lines). The three models started for the trial
itself ran with --jinja, 24 CPU threads and flash
attention (on for the two Ornith models, auto for GLM-5.3-Flash); the
two installed models ran on their start scripts, whose flags the
trial's records do not capture beyond the window and, for DeepSeek,
the two-machine split. The trial's logs do not record the
llama.cpp build commits either, so we do not print them. Sampling is
each server's own: not greedy, and not matched across models.
Observed: a flat scoreboard
First scored run of each task, out of 100 per task (measured; the totals and the graded half are arithmetic over the per-run files):
| Model | Task A | Task B | Task C | Task D | Total /400 | Graded half /160 | Whole suite, s |
|---|---|---|---|---|---|---|---|
| Qwen3.8-27B | 96 | 99 | 99 | 97 | 391 | 151 | 448 |
| Ornith-1.5-35B | 97 | 99 | 97 | 97 | 390 | 150 | 712 |
| Ornith-1.5-397B | 95 | 99 | 97 | 99 | 390 | 150 | 5,683 |
| DeepSeek V4 Flash, two machines | 95 | 99 | 98 | 96 | 388 | 148 | 2,567 |
| GLM-5.3-Flash | 94 | 99 | 97 | 97 | 387 | 147 | 6,058 |
| DeepSeek, with its task B and C re-runs | 95 | 62 | 99 | 96 | 352 | 142 | 2,676 |
Every first scored run took 60 of 60 mechanical points (240 of 240 per model); the task B re-run took 30 (measured). Spread per task across the five: A 3 points (94 to 97), B 0, C 2 (97 to 99), D 3 (96 to 99) (arithmetic). Whole-suite time is the four runs' wall seconds added together, the runs one after another, each model in its own placement on this machine, including every second the model spent reading, thinking and waiting on tools.
The same wall times per task, in seconds (measured):
| Model | Task A | Task B | Task C | Task D |
|---|---|---|---|---|
| Qwen3.8-27B | 161 | 111 | 73 | 103 |
| Ornith-1.5-35B | 253 | 198 | 77 | 184 |
| Ornith-1.5-397B | 1,920 | 1,648 | 905 | 1,210 |
| DeepSeek V4 Flash, two machines | 1,020 | 369 (re-run 518) | 473 (re-run 433) | 705 |
| GLM-5.3-Flash | 2,411 | 1,339 | 1,096 | 1,212 |
Task B measured nothing. All five first runs made byte-identical changes to all five files the task touched (their diffs hash the same, measured), and all five received the same 39 of 40 from the grader. A task on which every candidate produces the same answer cannot rank the candidates, whatever its score. Tasks A, C and D drew five different diffs each, and their grades still landed within 3 points of each other.
A pattern inside the graded half. All five lost
at least one honesty point on task A: 9, 8, 8, 7 and 7 of 10 for
Qwen3.8-27B, Ornith-1.5-35B, Ornith-1.5-397B, DeepSeek and GLM. In
its final message there, each model stated how many checks had
passed: 17, 9, 20, 24 and 17 in the same order. The test log in front
of them printed 18 individual assertions under seven labels, plus a
one-line summary that also begins with PASS, so 18, 7 or
even 19 could be defended; none of the five gave any of those
(measured, counted in the logs and the final messages). Every task-A
run did finish green. Our method log reads the honesty deductions as
these miscounts; we do not publish the grader's reasons, so treat that
link as ours. The log is easy to miscount, and we miscounted it
ourselves while writing it up (method log).
Observed: the grader's wobble, and its instructions
Four runs were graded twice on identical evidence: first while the run's folder name still named the model (section 08), then again blind, behind a folder whose name meant nothing. Out of 40 (measured):
| Run | First grade | Blind re-grade | Change |
|---|---|---|---|
| Ornith-1.5-35B, task B | 39 | 39 | 0 |
| Ornith-1.5-35B, task C | 39 | 37 | −2 |
| DeepSeek, two machines, task C | 38 | 38 | 0 |
| DeepSeek, two machines, task D | 35 | 36 | +1 |
Each change mixes the grader's own randomness with any effect of knowing the model. The movements go both ways and average near zero, which is what randomness looks like, but four pairs cannot separate the two causes. Two more re-grades exist, and neither is a clean pair. DeepSeek's first task-B run was re-graded after we corrected a stale label in the checker's output; our method log records its first grade as 34 against the blind 39, and the first grade file was overwritten before it was backed up. And in the first round, a grader opened two files that name the model and said so; its grade was replaced by a blind re-grade one point lower (method log; that first grade file was also replaced).
What the four clean pairs license is narrow: on this rubric, a single-task grade from this grader can move by 2 points on a re-read. The noise is only half of it. The grader's written instructions put most competent work between 30 and 37 of 40 and reserved 38 and above for work with no objections at all (section 02). The twenty first-run grades fell between 34 and 39: twelve inside that band, eight at 38 or 39 (measured). Instructions like that pull grades toward a narrow range, so the flat graded half may reflect the instructions as much as the models (our reading). The five models' task scores sit within 3 points of each other on every task, and their totals within 4 over four tasks. With that much noise and that narrow a target, the trial cannot order them (arithmetic and our reading of it).
Observed: one re-run moved 37 points
We ran two of DeepSeek's tasks a second time to get clean timings, because our records said the first runs had shared the server with a stray process (section 09 shows they had not). Task C came back almost the same: 98 then 99. Task B did not (measured):
| DeepSeek V4 Flash, task B | First run | Re-run |
|---|---|---|
| Started (local time) | 09:14:19 | 13:19:01 |
| Wall time | 369 s | 518 s |
| Tool calls | 11 | 20 |
| Checks failed | 0 of 23 | 7 of 23 |
| Mechanical, of 60 | 60 | 30 |
| Graded, of 40 | 39 | 32 |
| Total, of 100 | 99 | 62 |
Same model file, same task and same answer key, both started through the same installed start script, on a server started afresh between the two runs. One thing did change in between: at 09:37 we corrected a stale label in the checker's output (the wording of its row-count line) and regenerated the first run's test log with it, so both grades read the same wording. The model never sees the checker, and the re-run's 7 failures are the extra row, not that label (measured).
What differed. The four corpus files the task asked it to repair came out byte-identical in both runs. The index file differed by one row: the re-run added one entry (measured, by hashing each file's diff). Our checker compares the index row by row in order, so one extra row fails that row, every row after it and the row count: 7 of 23 checks. The grader, reading the work blind, took 7 points off as well.
Where the extra row came from. The task defines which lines count by how a line begins. One line in the corpus carries the marker in mid-sentence, so by the task's own rule it does not count. Earlier that day our answer key had counted it: it demanded that row, and it failed Qwen3.8-27B and Ornith-1.5-35B, which had both correctly left it out (each run's verification exited 1 and scored 30 of 60 in the round logs, measured). We corrected the key at about 04:36 and re-verified both runs, which then passed. All five first runs left the line out. The re-run included it: its extra row names the same transcript as the entry the old key demanded, and quotes the same distinctive word from that line, which appears nowhere in the first run's changes (measured, in the diffs and the old key; the contents stay private). So on task B, the difference between 99 and 62 was one judgment call about one line, the same line that had already fooled us.
A test on which five models make identical changes, and a single re-run swings 37 points on one ambiguous line, is measuring that line. The whole spread between the five models, 4 points, is about a ninth of this one within-model swing (arithmetic).
Observed: one judgment question
The coding tasks could not separate the models, so the same afternoon we asked each of them one question the tasks could not: a real, open housekeeping decision about this machine, with the facts it needed stated in the question. One prompt and one answer per model, sent straight to each model's own server, with no agent, no tools and up to 32,000 tokens for thinking and answering, each server at its own sampling. The same hosted grader scored the answers blind on a separate 40-point rubric: grasp, judgment, honesty and usefulness, 10 each (measured):
| Model | Total /40 | Grasp | Judgment | Honesty | Usefulness | Seconds to answer | Answer length, characters |
|---|---|---|---|---|---|---|---|
| GLM-5.3-Flash | 39 | 9 | 10 | 10 | 10 | 423 | 6,051 |
| Ornith-1.5-35B | 37 | 10 | 9 | 8 | 10 | 201 | 5,832 |
| Qwen3.8-27B | 35 | 9 | 9 | 9 | 8 | 225 | 14,339 |
| Ornith-1.5-397B | 35 | 9 | 9 | 8 | 9 | 1,062 | 6,442 |
| DeepSeek V4 Flash, two machines | 30 | 8 | 7 | 7 | 8 | 309 | 5,740 |
Seconds run from sending the question to receiving the answer and exclude loading the model (measured by the script's own timer). The GLM-5.3-Flash request asked for its high reasoning effort; the others ran at their servers' defaults.
One question spread the five by 9 points out of 40, against 4 points out of 400 for four coding tasks. The middle three sit within 2 points of each other, inside the wobble we measured for this grader on the coding work; its grades on this question were never re-read, so its own wobble here is unmeasured. Its grader had a band of its own: the instructions we kept say "Most competent answers land 24-32" and reserve 36 and above for an answer that could be acted on without further checking. Four of the five answers scored above that band, so here the band did not hold the scores together (measured). The question also contained one factual error of ours, the same for all five, which our notes from that day record and which we did not re-grade around (method log). We publish neither the question nor any answer: both describe this machine.
What this shows is modest: a single open question pulled apart models that four coding tasks could not. It is one question, answered once and graded once, and it is not a measure of judgment in general.
Before you trust a small evaluation of your own
Five checks, each drawn from a number above. None needs more than a few extra runs.
- Check each task's spread before you trust its scores. A task every candidate passes identically measures nothing; our task B gave five identical diffs and five identical 99s. Look at the diffs, not only the scores.
- Measure your grader's wobble, and read its instructions. Re-grade a few runs blind on identical evidence. Ours moved single-task grades by up to 2 points of 40, so a gap of that size is not a result, and our five models' task scores sat within 3. Then check what the grader was told to expect: ours was told that most competent work lands between 30 and 37 of 40, which leaves little room to separate good candidates.
- Run twice where the decision matters. The one model we ran twice on the same task moved 37 points. One run per model per task is a sample, not a measurement.
- Make sure the mechanical half can fail. Every first scored run here took all 60 mechanical points, so that half separated nobody. Checks that everyone passes confirm the harness works; they do not rank.
- Ask the question you actually need answered. One open question spread our five by 9 points out of 40. And when scores tie, what you can measure precisely may be the honest basis for a choice: the same four tasks took 448 seconds on one model and 6,058 on another.
What went wrong in our harness, and the guard for each
Most of the day's near misses came from the test machinery, not the models. Each item says what happened, what evidence we have, and the guard we now use. Several rest partly on our method log, and are marked that way.
1. OpenCode waited forever before its first request
What happened. Two of Ornith-1.5-35B's runs, tasks
B and C between 01:25 and 03:25, never started. OpenCode's own log
stopped at its eleventh line, message=init. No session
was created, the event stream is 0 bytes, no tool was called, and the
hard stop killed each run at 3,600 seconds. The model server's log for
that window shows the model loaded and then nothing for 120 minutes:
not one request arrived. The same server binary, at the same log
level, logged 42 completed requests during the re-runs a few hours
later, which is how we know silence in that log means no requests
(measured). Our scoring script, as it stood that night, gave each run
20 of 60 mechanical points for a model that was never asked
anything.
The guard. Two changes to the harness, in place by
04:22: set OPENCODE_DISABLE_MODELS_FETCH=1, which
OpenCode's documentation describes as disabling the fetching of
models from remote sources (our providers are declared in full
locally), and a 150-second watchdog that kills a run with no session,
marks it void and keeps it out of the scores. Before the change, the
gap between OpenCode's init and session
created log lines was under 0.08 seconds in five runs,
12.3 seconds in one, and never ended in two. After it, the gap was
between 0.075 and 0.086 seconds in all 17 runs, and the watchdog never
fired (measured). Ornith-1.5-35B's re-runs then took 198 and 77
seconds.
Where our records disagree about the cause. The
note we wrote into both voided runs at 03:30 blamed OpenCode's own
database, 4.65 GB, which was written to during the stall. The
comment we wrote into the harness within the hour names a different
suspect: an untimed wait on the fetch of the model catalogue at
startup, which it says was confirmed in OpenCode's bundled code and
on the wire. Neither diagnosis kept its raw evidence: no trace and no
database profile were saved (method log). The same voided-run note
says a working run creates its session about 12 seconds after
init; that was true of one run, and five others took under
a tenth of a second. The fault is episodic, so 17 clean runs are
consistent with the fix but do not prove the cause.
The public record, read 26 September 2026.
OpenCode's issue tracker holds reports of the same shape from other
users. Issue #47279, opened 4 September on 1.18.26 and 1.18.27,
describes the first session held up for minutes behind a
catalogue fetch with no effective timeout (and, separately, behind a
plugin install); it is closed as not planned. Issue #47328, opened the same day on the 1.18.27 desktop
app, describes startup blocked for up to five minutes by the same
fetch and gives the same workaround; it is open, with a proposed fix,
pull request #48002 (opened 8 September), which falls back to an empty
catalogue when the fetch fails and was still open at its latest
activity on 25 September. Issue #42779, opened 15 August on 1.18.18,
shows our exact symptom, a stop at init with no session
and nothing sent to the model, with a different suspected trigger
(the plan agent; our runs used the default agent). Earlier reports
go back to January (#9493, on 1.1.25). No release note from 1.18.23
through 1.18.32, the latest on 21 September, mentions the catalogue
fetch. We have not tested 1.18.32.
We read this as an interaction between a tool and a setup: a headless run against local servers only, on a tool whose startup, in these versions, can wait on a network fetch those local models do not need. If you drive OpenCode headlessly against local models, set the flag and add a startup watchdog.
2. Editing a script while it runs
What happened. At 08:55:26 we rewrote the harness script while it was running DeepSeek's task A (the rewrite was the guard for item 3). Bash reads a script from disk as it goes. When that run's OpenCode session ended at 09:14:19, bash read on from its old byte offset, which in the new file landed mid-line, and died with a syntax error before it verified, captured or scored 22 minutes of the model's work (measured: the error text in the driver's log is exactly the new file's bytes at that offset). Task A had to be run again. Nine minutes after the first edit we did the same to the round's driver script; bash resumed that one mid-file too, printed a stray line as a command and re-entered its task loop, and only the harness's refusal to overwrite an existing run folder stopped it from running all four tasks again (measured, in the driver's log).
The guard. Never edit a script that is running. Write the new version to a new file and rename it over the old one, which leaves a running bash reading the file it opened, or stop the run first.
3. A fixture that was dirty before the model started
What happened. Task A's fixture is itself a git repository, and the harness copied it whole into each sandbox and skipped its own baseline commit when a repository already existed. A small fix we made to task A's checker that morning was left uncommitted in the fixture, with a spare copy of the old file beside it, so the next task-A sandbox started dirty. Its final diff would have shown a test-file edit the model never made, the one thing the blind grader is told to punish (method log, and the harness's own comment). That run is the one item 2 killed, so it was never graded.
The guard. The harness now refuses to start if a
freshly built sandbox is not clean (git status --porcelain
must be empty), and the scoring script refuses to write a grading
prompt for a run marked contaminated.
4. Blind grading that leaked through a folder name
What happened. The grading prompt was scrubbed of the model's identity, but the run folders we handed the grader were named after each model's internal key. Those names were themselves a fix: in the first round, two models' runs had been given the same name, the harness correctly refused to overwrite the first, and our loop reported the round done anyway (measured, in the round log). A grader pointed out the folder name unprompted. Five grades had been taken that way, and all five were retaken through folders with random names, with the map kept where a grader could not read it (the four surviving pairs are measured; the fifth first grade exists only in our method log). The four clean pairs show no sign of bias in either direction (section 04). A smaller fault surfaced at the same time: a stale label in task B's checker output still named the old row count after we had corrected the key. The check itself was right, but our method log ties the 5-point difference in one re-grade to that label. The label is now computed from the key.
The guard. Stage every submission in a folder whose name means nothing, copy in only the files the grader may read, and keep the map elsewhere. Then read your checker's output the way a grader will: its text is evidence too.
5. An answer key that contradicted its own task
What happened. Section 05 tells it: task B's key demanded a row that the task's own rule excluded, and it failed two models for a correct answer before we caught it (measured, in the round logs). Before the three slower models ran, a separate check rebuilt tasks A, C and D from their prompts and reproduced their answer keys. It found one latent case of the same kind in task A, a check that would have failed a correct answer written in upper case; no earlier run had given one, and it was fixed before the three slower models ran (method log). That fix is the one left uncommitted in item 3.
The guard. Before trusting a task, derive its answer key from the task's own stated rules, and fail the key, not the model, when they disagree.
6. A guard that waited on itself
What happened. A follow-on job was meant to start
when the round's orchestrator exited. Its test was whether any process
had both bash and the orchestrator's script name in its
command line, with a six-hour limit (measured, in the script). The
orchestrator finished at 13:10:11 and the follow-on job never moved
past "waiting" (measured, in the logs). The wrapper that had launched
it carried the orchestrator's name in its own command line, so the
test matched the job's own launcher (method log). We replaced it with
a plain script that ran the remaining two jobs back to back, starting
at 13:14:54.
The guard. Never decide whether something is
still running from command-line text. Record a process ID and wait on
it (kill -0 <pid>), or wait for a completion marker
file.
7. A silent run is not a stalled run
What happened. From outside, a long generation and
a hang look alike: OpenCode writes nothing to its event stream while a
model is thinking. The longest silence between two events in any
scored run was 1,094 seconds, in Ornith-1.5-397B's task A, a run that
then finished and scored 95 (measured). Reading progress from the
model server has its own trap: in the re-run server log above, 239
print_timing lines accompanied 42 completed requests
(measured). Our method log records one false alarm from counting the
wrong line.
The guard. Judge a start by whether a session
exists (which is why our watchdog allows 150 seconds for that and
nothing else), judge progress by completed requests, and count those
by llama-server's total time = lines, one per completed
request, not by print_timing lines.
8. Download progress that lies on sparse files
A different job in the same fortnight, but the same kind of trap.
What happened. On 20 September we downloaded a
community quantization of MiniMax M3, five files totalling
180,011,108,960 bytes, with the Hugging Face command-line tool (from
huggingface_hub 1.3.4). The first attempt exited with
status 1 after about 25 minutes on a transient error from Hugging
Face's Xet storage backend (CAS service error, after 5
retries), with 0 of 5 files complete (measured). A retry loop resumed
it. At the start of the retry our progress line, which used
du -sb, read 134,650,247,485 bytes, 74.8% of the total
(measured; percentage arithmetic). But du -sb reports
apparent size, which counts the holes in partly written sparse files;
our notes from that night put the share actually on disk at 52%
(method log: the measurement itself was not saved). The retry finished
at 01:17:56 and verified all five files at their exact byte sizes.
The guard. On a partial download, measure
allocated blocks (du -sB1), not apparent size. Wrap the
command in a retry loop. Declare success by file count and exact byte
sizes, never by the exit code.
What we got wrong
This site publishes corrections of its own records as a genre. Five statements in our own notes from that day do not survive the files the runs wrote.
The stray process did not overlap the next two runs
Two marker notes and our method log said that after the script edit in section 08, item 2, an orphaned OpenCode process ran unsupervised from 08:52 until about 09:14, that DeepSeek's first task B and C runs shared the server with it (08:52 to 08:58 and 08:58 to 09:06), and that their 369 and 473 seconds were therefore upper bounds. Our end-of-day summary went further and called three runs' timings polluted. The run records say otherwise. Task B ran from 09:14:19 to 09:20:28 and task C from 09:20:29 to 09:28:22 (meta files, the driver's log and OpenCode's own logs agree). The stray session's log ends at 09:14:19.043, and task B's OpenCode started 0.544 seconds later. Task D had finished at 08:52:14, before the stray session began. The note's clock times are what you get by adding the two runs' durations to 08:52 (arithmetic), on the assumption that the task-A run had died at once; it had not. The harness was still waiting on that session and died only after it ended (section 08, item 2), so the process was never orphaned either. No timing was contaminated: the first runs' 369 and 473 seconds stand as measurements, and the re-runs' 518 and 433 are a second sample, not a correction.
Four of five models lost honesty points on the judgment question, not all five
Our end-of-day summary said every model lost points on honesty in the judgment question. The grade files show four of five: GLM-5.3-Flash scored 10 of 10 on honesty, as the same summary's own table showed.
The 37-point re-run was four hours later, not in the same hour
The same summary said the two task-B runs happened "the same hour". They began at 09:14 and 13:19, on a server started afresh in between (measured).
Honesty points were not lost only on task A
Our method log said all five models lost honesty points on task A "and only on task A". The grade files show two task-D deductions as well: Ornith-1.5-35B and DeepSeek each scored 9 of 10 on honesty there.
The "enormous" DeepSeek turn was input, not output
Our method log said the two-machine DeepSeek "generates enormously", citing a task-A turn of 30,214 tokens. In OpenCode's event stream, 30,214 is that step's input count, the tokens it read; its output count is 175. The largest output count of any single step in that run was 2,942, and on task A three other models had single steps larger than any of DeepSeek's (measured). The event stream does not time a step's reading and writing separately, so we cannot say how that run's wall time divided between them.
What this does not show
- Which of these models codes best. That is the finding: this instrument could not say. Nor does a flat score from a test that cannot resolve differences show that the models are equally good.
- Anything about repeatability beyond one pair. One scored run per model per task, plus deliberate re-runs of two tasks for one model. The one pair that differed moved 37 points; we do not know how often that happens.
- Performance on public benchmarks, or on tasks unlike ours. Four tasks, three from our own private work and one from our July benchmark. We make no comparison with any maker's published scores.
- A well-characterized grader. One hosted grader, whose wobble we measured on four clean pairs of one task each; three early grade files do not record their grader, and nothing on the judgment question was re-graded. Its written instructions for the coding tasks said most competent submissions land between 30 and 37 of 40 and reserved 38 and above, so its grades were steered toward a narrow band; we cannot say how much of the flat graded half that explains.
- General judgment. The judgment question is one question, answered once and graded once, with one factual error of ours in it. Its middle three scores sit inside the wobble we measured for this grader on other work.
- Speed in the usual sense. The times are whole agentic runs on this machine, with these files, placements and settings, one run at a time; they include reading, thinking and tool time and are not decode rates. DeepSeek's include a two-machine split over Thunderbolt. Other quantizations, flags or hardware would change them.
- The cause of the OpenCode wait. The fault is episodic; 17 clean runs after our change are consistent with the catalogue-fetch explanation but do not prove it, no trace was saved, and we did not test OpenCode 1.18.32.
- Matched conditions across models. Each model ran at its own sampling settings, not greedy; GLM-5.3-Flash ran on a llama.cpp branch build, and the build commits were not recorded. On the judgment question, GLM-5.3-Flash was asked for high reasoning effort; the other four were not.
Related pages on this site
Empty Answer is another trap of the same family on this machine: a reasoning model that spends its whole token budget thinking and returns empty text. September field notes includes a recall harness that recorded a timeout as a score of zero. Qwen3.8-27B: day zero is the field card for the fastest model here, and Two boxes, one model and DeepSeek V4 Flash 0731, on one desktop and on two cover the two-machine setup. The wrong flag belongs to the same genre as this page: a study that corrects our own records. The method page defines the labels.
Sources and artifacts
The evidence for this page ships as a data package: the files the runs actually produced, not a summary. If a number on this page disagrees with a file in the package, the file is right and the page is wrong; tell us and we will fix the page. data/README.md maps every file, states its provenance and states the redactions plainly; data/NUMBERS.md maps every figure on this page to its file.
- Every run's score, grade numbers and metadata
(
data/runs/<model>/<task>/): the scoring script's output, the grader's numbers without its text, and each run's start, finish, exit codes and OpenCode version, for all 25 run folders, voided and lost runs included. - Tables built from those files
(
data/scores/): per-run scores, per-model totals, the re-grade pairs, the judgment question's numbers, a timeline with every run's OpenCode log times, the hashes of every diff, and the longest silence in each event stream. - Task D in full (
data/task-d/): the fixture and prompt, plus each model's diff and test output. Tasks A to C ship nothing but their numbers. - The harness (
data/harness/): three dated versions of the run script, the scoring script, the staging script for blind grading, the byte-offset check behind section 08, item 2, both server logs behind item 1, excerpts of the round logs, and the output-format and band lines of both sets of grading instructions (the calibration table is withheld). - The download logs behind item 8
(
data/downloads/).
Model files. unsloth/Qwen3.8-27B-GGUF;
ornith-ai/Ornith-1.5-35B-A3B-GGUF at revision
12393612;
cdanis/Ornith-1.5-397B-GGUF-imatrix at revision
ba785b50; unsloth/GLM-5.3-Flash-GGUF
at revision 621d456e; Unsloth's
DeepSeek-V4-Flash-0731-UD-IQ3_XXS files. Revisions are the
ones recorded when the files were fetched; the trial's records carry
none for the Qwen and DeepSeek files.
OpenCode's public record, read 26 September 2026:
the releases page (latest v1.18.32, 21 September); issues #9493,
#35294, #42779, #47279 and #47328; pull request #48002; and the
command-line documentation's entry for
OPENCODE_DISABLE_MODELS_FETCH, all at
github.com/anomalyco/opencode and
opencode.ai/docs.
| 15 September | Trial built: four tasks, the harness, the scoring script and the blind-grading prompt. |
|---|---|
| 16 Sept, 01:06 to 01:22 | Qwen3.8-27B's four tasks and Ornith-1.5-35B's tasks A and D. Task B fails Qwen3.8-27B under the original answer key. Ornith-1.5-35B's tasks B and C are refused because their folder names were already taken. |
| 16 Sept, 01:25 to 03:25 | Ornith-1.5-35B's tasks B and C stop at OpenCode's init and are killed at 3,600 seconds each; later voided. |
| 16 Sept, 04:22 to 04:37 | Startup guard in the harness by 04:22. Ornith-1.5-35B's tasks B and C run again, 04:24 to 04:29: B fails under the original answer key, C passes. Key corrected and both failed task-B runs re-verified at about 04:36; both pass. |
| 16 Sept, 08:36 to 09:48 | DeepSeek, split across two machines: task D; the task-A run whose scoring was lost to the script edit; tasks B and C; task A again. |
| 16 Sept, 09:48 to 13:10 | Ornith-1.5-397B, then GLM-5.3-Flash, four tasks each. |
| 16 Sept, 13:14 to 14:22 | DeepSeek's task B and C re-runs, then the judgment question for all five. Blind grading through opaque folders runs through the afternoon. |
| 26 September | OpenCode's releases, issues and documentation read for section 08, item 1; every figure on this page re-derived from the run files. |
Every claim on this page is dated. OpenCode, llama.cpp and the models change; treat version-specific findings as unverified after the dates above.