Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Approval boundaries and exact-action binding

Three projects put a human decision in front of a consequential change. All three separate two things that are easy to merge:

  • What the approval interface displays to the human, which may come from the model, from a page rendered seconds ago, or from a cache.
  • What the mutation boundary independently verifies when the change is applied, against the current authoritative state, under a lock.

The second must never depend on the first being true.

The three designs

I · template (opt-in)III · ETF Research (mandatory)V · Tauros
What is decideda ticket's new prioritycommit an evaluation, shortlist, assignapprove, reject or request changes on an invoice
Shown to the humanthe ticket's current priority as the model passed itthe engine's decision as the model reported itthe invoice's current revision and its payload, rendered by the server
How the decision is boundHMAC-signed token: action, resource, actor, request, expected_choice, payload digest, nonce, expirysigned token binding the choice, the displayed engine decision, the payload (note), actor and requestthe approver submits revision_id and payload_hash with the decision
Re-verified at mutationnonce consumed; row locked; state re-derived and compared with expected_choice; apply_policyrow locked; evaluation recomputed and compared with the displayed decision; hard constraints re-checked; nonce consumedinvoice row locked; revision must be current; hash must be that revision's; policy: human approver who owns the agent
Written atomicallymutation + audit rowmutation + nonce + history row with policy versionsapproval + state transition + event; unique index on approvals.revision_id
ChaptersApprovals, lab 09Approvals, A5lesson 8, exact-payload approval

What "exact" buys you

An approval that says "yes" without naming its object can be applied to something the human never saw. All three designs prevent that. They differ in what they name:

  • The template and ETF Research bind the premise, meaning the state the human was told about, and refuse if reality differs. A model that misreports the current state cannot get a change applied. At worst it gets the human's approval refused.
  • Tauros binds the object, an immutable revision identified by id and content hash. Because revisions never change, "approve this hash" can only ever mean one payload. A change after submission creates a new revision, which needs a new decision.

Binding an immutable, content-addressed object is the stronger pattern when the thing approved can be represented that way. Binding a premise is what you can do when the approved thing is a transition on mutable state.

What the boundary does not verify

Exactness has limits, and the book notes them where they occur:

  • In the template and ETF Research the human may be shown a model-supplied premise. The approval card is honest about what the human chose, not about what was true. The recomputation covers the gap.
  • In Tauros the revision id and hash come back from the review form. The server verifies that they are current and consistent. It does not separately compare them with what the page rendered to that human (lesson 8 note).
  • No project verifies that the human understood what they approved. Clear displays, rationale requirements for overrides (Parts I and III) and reasons for rejection (Part V) are the mitigations.
  • No project implements separation of duties. The approver may be the person who configured or owns the proposing agent.
  • The signer is trusted. In the template and ETF Research the approval token is an HMAC signed by the agent runtime with HITL_APPROVAL_SECRET, the same secret the MCP server verifies with. The model never sees it, so a prompt injection cannot mint a token. A compromise of the runtime or its environment could, though. The recomputation would still refuse a stale or impossible change, but a forged token for a permitted change would be accepted as a human decision. In Tauros the decision is an authenticated human session calling the action directly, so the equivalent trusted component is the web application and its session handling. See Approvals — the trust model.

Checklist for your own approval boundary

  1. Name the approved object exactly: an id plus a content hash, or a signed premise.
  2. Make the approval single-use (a nonce, or a unique index per object).
  3. At mutation: lock, re-derive, compare, apply policy, then write the change and the audit record in one transaction.
  4. Treat everything displayed to the human as a claim, and verify it.
  5. Make the refusal path explicit and visible: "nothing was decided".
  6. Test each refusal, and test the database-level backstop by bypassing the application, as Tauros lesson 9 does with raw SQL.
  7. Write down who can produce a valid approval: which process holds the signing key or session, and what an attacker inside that process could approve.