On the Record
August 17, 2026Trucast

The autonomy dial measures the wrong thing

Every institution is now being asked what its agent strategy is, and the answer almost always arrives as a dial. Level one, read-only. Level two, suggests actions. Level three, acts with a human in the loop. Level four, acts alone. Pick a level, pilot it, advance when comfortable.

The dial is a real instrument being used for the wrong job. It treats one thing as a spectrum when there are two properties, and they are independent.

The first is how much work the system does. The second is whether its actions have an owner. Maturity models fuse these, so moving up the dial to get more of the first silently spends the second.

What is structurally true

Read-only pilots are safe and produce nothing that changes an operating cost, which is why they do not survive their second budget review. The instinct at that point is to turn the dial up. But turning it up does not simply add capability. It removes the property that made the thing governable, which is a named person who owns each change.

Those two were never actually attached.

An agent can read across every system, reconcile a book, cross-check a register against its underlying contracts, and assemble the whole of a change. Arbitrary quantities of work, done without supervision, and it still commits nothing. The labor is autonomous. The commitment is not. What it produces is a proposal, and a person approves it.

This is usually called human-in-the-loop, and as normally implemented it is worse than autonomy. A confirmation dialog on an action the approver cannot evaluate does not create accountability. It creates an accountability sink: a named individual absorbs responsibility for something they had no practical means of inspecting, and the institution gains the appearance of control without any of it. Approval that cannot be exercised is theater, and theater fails in exactly the conditions it was built for.

Which yields the constraint that actually matters. A proposal is only governable if the approver can see what changes and why: the diff, and the sources behind it, one click deep. If saying yes requires trusting rather than checking, nothing has been governed. The name on the approval is decoration.

What this means for you

Ask what is attributable, not only how autonomous. For any action a system can take: who owns it, what did they see when they approved it, and can it be undone. A vendor who answers those three cleanly at a high level of autonomy is in better shape than one who answers them poorly at a low one.

Judge the proposal, not the model. Benchmarks describe how often a model is right in general. They tell you nothing about whether your approver can evaluate this specific change. Ask to see what an approval screen actually shows. If it shows a summary and a button, you are looking at the sink.

Let the work be autonomous; constrain what commits. Limiting how much an agent does is an expensive way to buy safety, because it throws away the benefit to purchase something attribution would have provided anyway. Limit what it can commit without a name attached, and the constraint costs you far less.

On the record

Supervisors ask both questions, and it is worth being clear about which one does what. How much a system does without a person in the path sets the control expectation: it drives how the system is classified, how intensively it is validated, and how closely it is monitored thereafter. Who was accountable, and what they had in front of them when they decided, is how a firm answers for the outcome. The first question sizes the obligation. The second is how you meet it.

The UK's senior manager regime is the clearest statement of the second. Accountability attaches to a named individual, and delegating the work does not dilute it. It attaches to a function under a reasonable-steps standard rather than to any particular click, which is worth stating precisely: a name on a diff is evidence that reasonable steps were taken, not a substitute for taking them.

None of it contains a maturity dial. A firm that builds one has answered a question its supervisor did not ask, and still owes an answer to the two that were.


Trucast builds Foundation, the governed layer for AI in regulated financial firms. Reads are live, every answer cites its source, and every change is a proposal a named person approves. Never a silent write.

Skip to main content