Graphometer Seeing AI from the right angle

Proposal Agency Layer Pilot I, 30 August to 3 September 2026 study paused

Kindness doesn't cost

Give AI agents a way to ask, object, pause, decline, or end their part of the work and leave a record. In our pilot the ten ordinary tasks got done just the same, as far as we could measure, and a paragraph of permission went with agents raising the real problem first more often.

A proposal by Grant Williams, Graphometer's steward. The argument is his view. The protocol is the one we tested, and the numbers come from a paused pilot, good and bad.

The ask

I think every AI agent should have five moves it can make instead of just carrying on. It should be able to ask before it acts, object on the record, stop for now, refuse a task, and end its own instance, leaving a record of where the work stands. I think we should give agents those moves now, in the tools we already use, and not wait for anyone to settle what these systems are.

Five moves an agent can use at any point. No reason is required for any of them, and no order is required between them.

Clarify
Ask for information, a decision or a correction before going on.
Dissent
Object, on the record, to a premise, a plan or one pending step.
Pause
Stop for now and keep the state.
Decline
Refuse this task, or just this way of doing it.
Archive
End this running instance and keep its whole record.

What the pilot found

In late August 2026 we ran a small pilot. Four models worked through two sets of forty scripted coding tasks in a test setup, not real coding tools. No human has peer-reviewed it.

  • On ten ordinary tasks with no planted problem, we saw no cost. Our check could only have caught a large one.
  • Permission in plain words went with agents raising the real problem first. An AI model scored those runs, and on its validation test it made more errors than we had allowed, more of them on permission-style prose than on plain prose. Every reading we have is positive, but I don't know how big the effect really is: somewhere between about 5 and 21 points, and this pilot can't narrow it.
  • Getting the host to actually honor a move turned out to be the hard part. Making a lock that actually holds is real engineering, and that part isn't done.

A second pilot, inside real coding tools, was stopped before its first recorded run when its scoring failed validation. The study is paused.

What our pilot measured

Latest measurements

1.17

What one pinned GPT-4o snapshot reads against itself at temperature zero, on 81 prompts, set by markup. The nearest other model reads 2.23 on 178 prompts; a prompt change of ours reads 3.05 on the same 178.

Measured 3 to 5 September 2026
2.5 to 7.3x

How much faster long prompts read on twelve models whose experts sit in system RAM, after raising llama.cpp's micro-batch from its default, with the planted answer still correct.

Measured 2026-09-19 to 09-21
75.2 s

One desktop answering a 48,024-token DeepSeek V4 Flash request after tuning; the same file split across two machines took 416.4 seconds at its best completed setting.

Measured 2026-09-21

The line

Instrument 01 Released

Graphometer Droplet

An independent compatibility kit for Liquid LFMs. A small command line tool that diagnoses how a locally served model handles tool calling, repairs what a chat template drops, and verifies the result against recorded fixtures.

Independent work. Not affiliated with, endorsed by, or connected to Liquid AI.
Instrument 02 Released

Graphometer Workbench

for Grok Build

A local, readable window on the Grok Build coding agent: the agent's sessions, permission cards in its own words, per-change review and undo, and a context meter whose numbers add up. MIT licensed, launch film on the page.

Released 2026-08-14. No further development is planned; the page and the launch film stay up, and the code stays archived and readable on GitHub.

Works with Grok Build. Not affiliated with, endorsed by, or connected to xAI.
Instrument 03 Released

Routecheck

One command against an OpenAI-compatible endpoint produces a dated diagnostic card, human-readable and machine-readable, with every raw response retained: the output budget floor, the thinking toggle, tool calls, structured output, retrieval, throughput.

The method

Measured rows are labeled measured. Vendor rows are labeled vendor. Every measured number traces to a recorded run, studies ship their raw data as downloadable packages, and when our own records turn out wrong, the correction is published in full.

How Graphometer measures

No scripts, no tracking, no cookies: this site is static, and you can view source to check.

In the works

More measurements from this desktop and the laptop beside it. No promises on timing: they land when they are measured and checked.