Approval boundaries and exact-action binding
Three projects put a human decision in front of a consequential change. All three separate two things that are easy to merge:
- What the approval interface displays to the human, which may come from the model, from a page rendered seconds ago, or from a cache.
- What the mutation boundary independently verifies when the change is applied, against the current authoritative state, under a lock.
The second must never depend on the first being true.
The three designs
| I · template (opt-in) | III · ETF Research (mandatory) | V · Tauros | |
|---|---|---|---|
| What is decided | a ticket's new priority | commit an evaluation, shortlist, assign | approve, reject or request changes on an invoice |
| Shown to the human | the ticket's current priority as the model passed it | the engine's decision as the model reported it | the invoice's current revision and its payload, rendered by the server |
| How the decision is bound | HMAC-signed token: action, resource, actor, request, expected_choice, payload digest, nonce, expiry | signed token binding the choice, the displayed engine decision, the payload (note), actor and request | the approver submits revision_id and payload_hash with the decision |
| Re-verified at mutation | nonce consumed; row locked; state re-derived and compared with expected_choice; apply_policy | row locked; evaluation recomputed and compared with the displayed decision; hard constraints re-checked; nonce consumed | invoice row locked; revision must be current; hash must be that revision's; policy: human approver who owns the agent |
| Written atomically | mutation + audit row | mutation + nonce + history row with policy versions | approval + state transition + event; unique index on approvals.revision_id |
| Chapters | Approvals, lab 09 | Approvals, A5 | lesson 8, exact-payload approval |
What "exact" buys you
An approval that says "yes" without naming its object can be applied to something the human never saw. All three designs prevent that. They differ in what they name:
- The template and ETF Research bind the premise, meaning the state the human was told about, and refuse if reality differs. A model that misreports the current state cannot get a change applied. At worst it gets the human's approval refused.
- Tauros binds the object, an immutable revision identified by id and content hash. Because revisions never change, "approve this hash" can only ever mean one payload. A change after submission creates a new revision, which needs a new decision.
Binding an immutable, content-addressed object is the stronger pattern when the thing approved can be represented that way. Binding a premise is what you can do when the approved thing is a transition on mutable state.
What the boundary does not verify
Exactness has limits, and the book notes them where they occur:
- In the template and ETF Research the human may be shown a model-supplied premise. The approval card is honest about what the human chose, not about what was true. The recomputation covers the gap.
- In Tauros the revision id and hash come back from the review form. The server verifies that they are current and consistent. It does not separately compare them with what the page rendered to that human (lesson 8 note).
- No project verifies that the human understood what they approved. Clear displays, rationale requirements for overrides (Parts I and III) and reasons for rejection (Part V) are the mitigations.
- No project implements separation of duties. The approver may be the person who configured or owns the proposing agent.
- The signer is trusted. In the template and ETF Research the approval token
is an HMAC signed by the agent runtime with
HITL_APPROVAL_SECRET, the same secret the MCP server verifies with. The model never sees it, so a prompt injection cannot mint a token. A compromise of the runtime or its environment could, though. The recomputation would still refuse a stale or impossible change, but a forged token for a permitted change would be accepted as a human decision. In Tauros the decision is an authenticated human session calling the action directly, so the equivalent trusted component is the web application and its session handling. See Approvals — the trust model.
Checklist for your own approval boundary
- Name the approved object exactly: an id plus a content hash, or a signed premise.
- Make the approval single-use (a nonce, or a unique index per object).
- At mutation: lock, re-derive, compare, apply policy, then write the change and the audit record in one transaction.
- Treat everything displayed to the human as a claim, and verify it.
- Make the refusal path explicit and visible: "nothing was decided".
- Test each refusal, and test the database-level backstop by bypassing the application, as Tauros lesson 9 does with raw SQL.
- Write down who can produce a valid approval: which process holds the signing key or session, and what an attacker inside that process could approve.