Design exercise: an agent-proposed payment, end to end
Important
This is a design exercise. The integrated system it describes does not exist. At the pinned revisions no project signs transactions, moves value or settles payments, and no two projects are connected at runtime. See What is integrated, and what is only designed.
The scenario
An operations agent prepares invoices for a small company. When a customer invoice is approved and the customer pays in stablecoins, the company pays a contractor's share onward to the contractor's registered receiving address. You must design how an AI agent can prepare that onward payment while:
- no model ever holds a key or a secret;
- a human decides, and the decision binds one exact payment;
- a crash, a retry or a replay cannot pay twice;
- an auditor can later reconstruct who intended what, who approved it, under which rules, and what actually settled.
Building blocks you may reuse
| Need | Pattern | Taught in |
|---|---|---|
| Bounded, authenticated agent with traced tool calls | gateway, MCP capability boundary, evaluation | Part I |
| A run that survives restarts | runs vs checkpoints; resume; replay window | Part II |
| A deterministic rule deciding whether a payment is allowed at all | policy as versioned data; hard constraints; recomputation at mutation | Part III |
| Keys that never reach the model | key hierarchy, least-capability tools, operator custody | Part IV |
| Financial intent, ownership, exact approval, idempotency | immutable revisions, payload hash, state machine, one decision per revision | Part V |
Tasks
Work through these in order. Write your answers as an architecture decision record with a diagram.
- Intent. Define a
PayoutInstruction: its fields, its content hash, and what makes it immutable. Which existing Tauros concepts does it reuse (destinations, revisions, idempotency keys)? Which fields must be impossible for an agent to set? - Capability. List the MCP tools the agent gets. Show that none of them can approve, sign or change a destination. What does the agent's tool list not contain, and which test pins it?
- Policy. Write the deterministic rules that decide whether an instruction may even be proposed: amount limits, an allow-listed destination, matching an approved invoice. Where do they live, how are they versioned, and what does the audit row record?
- Human authority. Design the approval. What exactly does the human see? Which of those values could a model influence? What does the approval bind (id plus hash, or premise)? Add a four-eyes rule. Who may not approve?
- Signing boundary. Arktos does not sign today. Design the signing
capability as a separate service, or an extension of Arktos, that:
- accepts only instructions approved upstream, verified by signature or by a query back to the system of record, never just "because the caller says so";
- re-checks the destination and amount against its own policy;
- is idempotent on the instruction hash;
- exposes no signing tool to any agent. Compare with Arktos's challenge on safe transaction signing.
- Durability. The agent's run crashes after calling "submit". Walk through what a Sophos-style resume replays, and show why your design still pays at most once. Name the idempotency identity at each hop.
- Settlement. Payment is not done when it is submitted. Define the settlement event, how it is observed, how it is reconciled with the instruction, and what happens when it never arrives (Tauros's eventual-consistency concept).
- Custody. For each key in your design, state who holds it, who can decrypt it, and what an operator could do alone. Is your design custodial, and from whose point of view?
- Evidence. List the tests, including adversarial ones, that would convince you each guarantee holds. Mark which need real concurrency.
Review questions
- If the model were fully compromised by a prompt injection, what is the worst thing it could cause? Which component stops the next step?
- If the approval UI showed the wrong amount, would a wrong payment be signed?
- If the agent runtime (not just the model) were compromised, which credentials would the attacker hold, and what could they approve or submit with them?
- If the signer's database were restored from yesterday's backup, could an instruction be paid twice?
- Which guarantees are enforced by a type, a constraint, a policy, a signature or a person? Is any guarantee enforced only by a prompt?
There is no published solution. A good answer is one where every guarantee names the component that enforces it, and none of those components is the model.