Tools: the instrument line
Three instruments, each free, open source, and made to run on your own machine. Each gets a short entry here and a full page of its own; every claim below comes from the instrument's page or its public repository, and the links carry the rest.
Graphometer Workbench, for Grok Build
A local, readable window on the Grok Build coding agent: it puts the agent's sessions, its work as it happens, and every moment it stops to ask you a question into one plain screen. It asks before the agent changes your files, lets you review and undo one change at a time, and survives the agent dying.
Who it is for: people who direct AI agents without living in a terminal.
Runs on Linux (tested), or Windows via WSL2; macOS is unverified
and not claimed. The server needs Node.js 22+ and runs
TypeScript directly, with no build step and no dependencies to
install; Grok Build itself is installed separately.
Released 2026-08-14. No further development is planned; the page and the launch film stay up, and the repository stays archived and readable.
The Workbench page · github.com/graphometer/workbench-for-grok-build
Works with Grok Build. Not affiliated with, endorsed by, or connected to xAI.
Graphometer Droplet
An independent compatibility kit for Liquid LFMs: the harness can break the model, and Droplet checks the fit. When it does not fit, Droplet names the failing layer and applies the smallest proven repair, then gets out of the way.
Who it is for: anyone serving a Liquid LFM locally who needs to know whether a failure is the plumbing's fault or the model's.
Installs with pip from the repository; no
telemetry, no accounts, loopback only. Licensed Apache 2.0 for the
original code, CC BY 4.0 for the evidence data, and LFM Open
License v1.0 for the Liquid-derived templates.
The Droplet page · github.com/graphometer/droplet
Independent work. Not affiliated with, endorsed by, or connected to Liquid AI.
Routecheck
One command against an OpenAI-compatible chat endpoint, one dated, evidence-backed diagnostic card. It runs a fixed battery of behavioral probes against a running route and writes the results three ways: machine-readable, plain English, and the complete raw request and response for every probe, so any claim on the card can be checked against what actually went over the wire.
Who it is for: anyone running a model behind an OpenAI-compatible route who wants a dated, checkable record of how that endpoint actually behaves.
Needs Python 3.10 or newer plus the requests
package; everything else is standard library, and there is no
install step: clone the repository and run in place, or copy the
directory anywhere and run it from the copy. llama.cpp server and
Ollama locally, plus the hosted aggregator route of its first
published card, are the routes it has been run against live so far;
authed and remote endpoints are supported. On the
Qwen3.8 release watch it is the
instrument that packaged the fixed battery into one recorded
pass.
The Routecheck page · github.com/graphometer/routecheck
Independent work. Not affiliated with, endorsed by, or connected to the llama.cpp project, ggml-org, or Ollama.
Looking for a specific model?
The per-model measured records live on the field cards index, and the studies, reports, and guides have an index of their own.