Idempotency, replay, concurrency and recovery
Four failure situations look alike from far away. In all four, an action may take effect more than once, or not at all, when it should take effect exactly once:
| Situation | Cause | Typical defence |
|---|---|---|
| Retry | a client did not see the response and sends the request again | idempotency key; natural uniqueness |
| Replay | a runtime re-executes a step after a crash or resume | replay-safe tools; external deduplication |
| Concurrency | two requests act on the same object at the same time | row locks; unique constraints; conditional updates |
| Recovery | a process died; its work is half done | durable status, explicit resume, reconciliation |
Checkpointing, nonces, idempotency keys and locks each defend against one of these. None defends against all four.
What each project does
| Project | Retry | Replay | Concurrency | Recovery |
|---|---|---|---|---|
| I · template | an approval token's nonce is single-use (primary key) | n/a: no durable runs | row lock at mutation; nonce conflict enforced by the database (no automated concurrency test) | the transaction rolls back on any failure, nonce included |
| II · Sophos | a new message is refused (409 SESSION_IN_PROGRESS) while a run is active in that conversation; nothing deduplicates a retried message | resume re-executes the interrupted step; all tool calls of that turn re-run | one active run per conversation | running runs become interrupted on the next database access; resume from the last checkpoint |
| III · ETF Research | single-use nonce (ON CONFLICT (nonce) DO NOTHING, then refused) | n/a | row lock, then recompute | one transaction for mutation, nonce and history |
| IV · Arktos | wallet names are unique per owner; addresses are derived once, recorded, and returned unchanged afterwards | n/a | single instance per SQLCipher database | backup and the master-key recovery trap (C8) |
| V · Tauros | idempotency key unique per agent; same key and same payload returns the original, a different payload is refused | n/a | invoice row lock on every command; unique index on approvals.revision_id | a state machine whose transitions are atomic |
Three observations
Checkpointing is not exactly-once. Sophos makes this its central lesson. Durable execution lets a workflow continue. It cannot know whether a tool's external effect landed before the crash. A replay-safe tool must be idempotent on its own terms, for example by accepting a caller-supplied key, which is what Tauros does for proposals (R7, Tauros lesson 7). If Sophos were given a Tauros proposal tool, the idempotency key would have to be derived from something stable across replays, such as the run and tool-call ids, and not generated freshly inside the replayed step.
Locks keep the application consistent; constraints keep the database consistent. Tauros states this directly, and lesson 9 proves the second half with a raw-SQL insert that the unique index refuses. The same split appears in the template (row lock plus a nonce primary key) and in ETF Research.
Tests that look concurrent may not be. Several projects' concurrency tests run inside a single database connection or transaction (Tauros's SQL sandbox), or are documented as untested (the template's nonce conflict under real concurrency). They prove the rules: a second decision is refused and a replay returns the original. They do not prove lock contention between real connections. Note which kind of evidence you have.
Design rule of thumb
For every state-changing tool an agent can call, write down:
- its idempotency identity: what makes two calls "the same";
- what a replay after a crash will send, and whether it has the same identity;
- the lock or constraint that serialises concurrent calls;
- the durable status that tells a restarted process what is unfinished;
- the test that proves each, and whether it uses real concurrency.