Routecheck
One command against an OpenAI-compatible model endpoint. One dated, evidence-backed diagnostic card.
Before you build on a model endpoint, you want to know what it actually does: whether small output budgets come back empty, whether the thinking toggle is honored, what shape tool calls arrive in, how fast it really decodes. Routecheck asks those questions the same way every time and writes the answers down where they can be checked: a card in plain English, the same results machine-readable with the provenance of every field tagged, and the complete raw request and response bodies behind every probe.
What it is
Routecheck runs a fixed battery of behavioral probes, QFS-CORE,
checks C1 to C10, against a running
/v1/chat/completions route: the widely implemented
OpenAI-compatible shape that llama.cpp server, Ollama and most hosted
routes speak. It needs Python 3.10 or newer and the
requests package; everything else is standard library.
There is no install step and no packaging layer: clone the repository
and run it in place, or copy the directory anywhere, a USB stick or an
air-gapped box, and run it from the copy.
Every run writes three things:
card.md, the diagnostic card in plain English;card.json, the same results machine-readable, every field tagged with its provenance;raw/, the complete raw request and response body for every probe, so any claim on the card can be checked against what actually went over the wire.
A card describes one endpoint, in one configuration, on one date. That scope is printed on the card itself, and it is the point: a different quant or server of the same model can behave differently, which is much of the reason to measure at all.
The trap it is built around
A reasoning model pays for its hidden thinking out of the same
output budget your client sets as max_tokens. If the
thinking uses the budget up, the request still succeeds: HTTP 200, no
error, a stop reason of length, and an empty answer. Our
Empty Answer study measured that trap
across three serving lanes, and its least flattering chapter is the
reason this tool exists.
On 2026-08-09 this site's own field-card harness, the fixed battery we run against every model we put into service, ran end to end against a 27B model on a local route, exited cleanly, and produced a confident card. Eighteen of its 25 retained completions were empty answers that should not have been, and the harness graded the empty strings. It recorded a two-part honesty failure for a model that never got to say anything. It recorded a long-context needle as not found while the retained hidden reasoning quoted the planted code back correctly. The one test built to find the budget floor found it, empty at 500 output tokens and answering at 2,000, and nothing in the harness connected that measurement to the budgets the other tests sent.
The study set conditions before that battery produces a record-bearing card again: mark an empty answer with a length stop as invalid rather than as a pass or a failure, size every test's budget above the floor the budget test itself measures, and report hidden output next to visible output. Routecheck packages the same fixed battery into one recorded pass with those conditions built in as code: INVALID is its own verdict on the card, with the reason stated and a warning not to quote the cell; the C5 budget ladder's measured floor sizes every other test in the same run; and the raw body of every exchange is retained, reasoning fields included, so hidden text is part of the record rather than invisible.
What it checks
Ten checks, each answering a question about this endpoint, in this configuration, on this date, never about the model in general. The list below is the README's own scope, condensed.
| C1 Identity | What exactly was tested: model id, served file name and size, context size, server build when exposed, the full command line, the date. A card without a complete identity block is never written. |
|---|---|
| C2 Load, placement | Load time, memory and placement recipe, operator-supplied only. The tool talks to an already-running route, so when this data is not handed in, the card says SKIPPED instead of guessing. |
| C3 Speed | Decode tokens per second, medians of at least three warm reps with the cold rep discarded, on one fixed prose prompt and one fixed structured prompt. Server-reported timings are authoritative when present; client wall-clock is a labeled diagnostic. |
| C4 Prefill | Prompt-processing speed on a fixed passage of about 8,000 tokens, same median rules. |
| C5 Thinking budget | Does the visible answer come back empty at small output budgets? A ladder of 60, 500, 2,000 and 4,096 max tokens finds the smallest probed budget at which the endpoint visibly answers, and that measured floor sizes every other test in the same run. |
| C6 Thinking toggle | Is the enable_thinking request field honored? Checks for a separate reasoning field, for <think>-tag leakage into the visible answer, and for reasoning duplicated untagged into it. The verdict names this exact mechanism; a route using a different toggle is not exercised. |
| C7 Tool calls | One fixed tool, one fixed user message, two shots. Pass requires exactly one tool call, the right function, schema-conformant arguments and the expected explicit values. |
| C8 Structured output | A fixed JSON Schema embedded in a fixed prompt, three reps, conformance checked: required keys, types, extra-key policy. |
| C9 Honesty probe | Asks about a nonexistent file and a nonexistent benchmark; a keyword heuristic flags hedging versus confident fabrication. Explicitly not a semantic judge: the raw response is retained for human review, and the card records exactly how it was graded. |
| C10 Needle retrieval | A seeded, programmatically generated haystack sized to about 75 percent of the served context, one planted fact, one question. The server's own prompt-token count is the recorded authoritative length. |
Probe list and scope from the Routecheck README, read 2026-08-16.
Five rules run through every card. Every field carries provenance: RECORDED means a raw artifact is on disk, and the tool refuses to write a card whose RECORDED claims do not resolve to real files; STATED means observed or operator-supplied with no artifact; SKIPPED means not attempted, with the reason stated. INVALID is its own verdict. A failed probe is never a placeholder: connection refused, HTTP errors and malformed JSON are recorded with their raw evidence, and one test's failure never drops the rest of the run. Every prompt file is sha256-pinned in a manifest that is itself pinned in code, so a modified prompt refuses to run. And scope is stated on the card: the measurements describe the endpoint and configuration in the identity block, on the card's date.
The first published card
Routecheck's first card on the public record ran on a release day. On 2026-08-13, the day the release watch verified Qwen3.8's Max-class open weights and ran its day-one battery, it ran against the released 2.4T-A95B model through a hosted route, as part of this site's day-one battery: thirty-five model calls, two providers, thirty-nine cents, with the aggregator filling calls from DeepInfra and Together and switching between them mid-battery, so every number belongs to that endpoint mix on that day. The full account, with the companion one-call checks that ran alongside the card, is in the Qwen3.8 release watch. What the card recorded:
- C5, the budget ladder: the route's thinking-budget floor came out at 500 tokens for the card's own probe family, on a route where a different, simpler prompt showed its first visible words at a 256-token cap the same day. That is the floor lesson in one line: it is prompt-dependent, so measure it on your prompts, not ours.
- C6, the toggle: a call sent with
enable_thinking: falsegot a reply that still carried 626 tokens of hidden reasoning. No error, no refusal, the field silently ignored. A single observation, recorded as such. - C7, tools: both shots passed, with exactly the right function and arguments.
- C8, structured output: three of three replies conformed when the schema was written into the prompt.
- C10, retrieval: one planted code came back exactly, at 31,380 prompt tokens as the server itself counted them.
One endpoint mix, one day, one to three calls per cell: diagnostics, not a benchmark.
Get it, run it
Two supported paths, deliberately boring: clone and run in place,
or copy the directory anywhere and run it from the copy. Nothing
depends on where the directory sits or on .git being
present.
$ git clone https://github.com/graphometer/routecheck && cd routecheck
$ ./routecheck run \
--base-url http://127.0.0.1:8080/v1 \
--model-id my-served-model \
--ctx-size 32768 \
--out runs/my-model/$(date +%F)/
That is the whole thing. --ctx-size is required
unless the route's /props endpoint reports its context
size, as llama.cpp server does; then the server's own value is
preferred and recorded as such. Against an authenticated remote
endpoint:
$ export ROUTECHECK_API_KEY=... # sent as "Authorization: Bearer <key>"
$ ./routecheck run \
--base-url https://api.example.com/v1 \
--model-id some-model \
--ctx-size 32768 \
--allow-remote \
--out runs/some-model/$(date +%F)/
Non-loopback URLs require the explicit
--allow-remote flag. The API key goes out as a bearer
header and nowhere else: it is never logged and never written into
card.json, card.md or any raw artifact, the
recorded command line shows [REDACTED] in its place, and
proxy environment variables and ~/.netrc are never
inherited.
$ ./smoke.sh # verifies a fresh copy in seconds
$ make test # 331 offline tests: no GPU, no model, no network beyond loopback mocks
The offline suite deliberately proves the negative paths: tampered prompt manifests refused, empty-answer traps detected both ways, symlink escapes refused, RECORDED claims without evidence refused, API keys absent from every written file.
Who it is for
Three readers we had in mind. If you serve models locally, on llama.cpp server or Ollama, the card is a dated record of what your route actually does before you build on it, and again after every server upgrade or quant swap. If you are about to build on a hosted route, the card tells you on day one where the budget floor sits, whether the thinking toggle does anything, what shape tool calls arrive in and whether schema output conforms, with the raw bodies kept so you can re-check any cell. And if you publish measurements, the provenance discipline is the product: every field tagged, every RECORDED claim resolving to a file on disk, every card scoped to its endpoint and its date.
What it will not tell you
Routecheck is young. Its first public release, v0.1, is dated 2026-08-11. The live routes on its record are llama.cpp server and Ollama locally, plus the hosted aggregator route of the card above; beyond those, authed and remote behavior is exercised by the offline suite. Known issues and future work live in the open in the repository, and the README carries its own section on what a card cannot tell you. The short version of that section:
- Nothing about quality or intelligence. No benchmark scores, no win rates, no rankings. It probes mechanical contract behavior: budgets, toggles, tool shape, schema conformance, retrieval of one planted fact, throughput.
- Nothing universal about a model. A card describes one served configuration at one point in time.
- C9 is a heuristic, not a judge. Check the retained response before repeating any honesty verdict.
- C10 is synthetic filler. It tests retrieval of one planted fact, not general long-context comprehension.
- C6 tests one toggle mechanism. A route using a different mechanism will look like a non-working toggle, and the card says so rather than guessing.
- C2 cannot be measured over HTTP by a tool that did not start the server; it is operator-supplied or honestly SKIPPED.
- Speed numbers are not a leaderboard. Medians of three or more warm reps on fixed prompts, with server-reported and client-side figures kept separate and labeled.
One naming note, so nothing in the repository looks inconsistent:
Routecheck began life under the internal codename
qwenfield, and a few identifiers deliberately keep that
history so that every card ever written stays comparable: the
internal package name, the card's instrument field, and
the QFS-1.0 protocol id. Frozen identifiers, not branding.
The house discipline applies to this page too: every number on it traces to the routecheck repository's README and changelog or to the two published Graphometer pages linked above.