Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Idempotency, replay, concurrency and recovery

Four failure situations look alike from far away. In all four, an action may take effect more than once, or not at all, when it should take effect exactly once:

SituationCauseTypical defence
Retrya client did not see the response and sends the request againidempotency key; natural uniqueness
Replaya runtime re-executes a step after a crash or resumereplay-safe tools; external deduplication
Concurrencytwo requests act on the same object at the same timerow locks; unique constraints; conditional updates
Recoverya process died; its work is half donedurable status, explicit resume, reconciliation

Checkpointing, nonces, idempotency keys and locks each defend against one of these. None defends against all four.

What each project does

ProjectRetryReplayConcurrencyRecovery
I · templatean approval token's nonce is single-use (primary key)n/a: no durable runsrow lock at mutation; nonce conflict enforced by the database (no automated concurrency test)the transaction rolls back on any failure, nonce included
II · Sophosa new message is refused (409 SESSION_IN_PROGRESS) while a run is active in that conversation; nothing deduplicates a retried messageresume re-executes the interrupted step; all tool calls of that turn re-runone active run per conversationrunning runs become interrupted on the next database access; resume from the last checkpoint
III · ETF Researchsingle-use nonce (ON CONFLICT (nonce) DO NOTHING, then refused)n/arow lock, then recomputeone transaction for mutation, nonce and history
IV · Arktoswallet names are unique per owner; addresses are derived once, recorded, and returned unchanged afterwardsn/asingle instance per SQLCipher databasebackup and the master-key recovery trap (C8)
V · Taurosidempotency key unique per agent; same key and same payload returns the original, a different payload is refusedn/ainvoice row lock on every command; unique index on approvals.revision_ida state machine whose transitions are atomic

Three observations

Checkpointing is not exactly-once. Sophos makes this its central lesson. Durable execution lets a workflow continue. It cannot know whether a tool's external effect landed before the crash. A replay-safe tool must be idempotent on its own terms, for example by accepting a caller-supplied key, which is what Tauros does for proposals (R7, Tauros lesson 7). If Sophos were given a Tauros proposal tool, the idempotency key would have to be derived from something stable across replays, such as the run and tool-call ids, and not generated freshly inside the replayed step.

Locks keep the application consistent; constraints keep the database consistent. Tauros states this directly, and lesson 9 proves the second half with a raw-SQL insert that the unique index refuses. The same split appears in the template (row lock plus a nonce primary key) and in ETF Research.

Tests that look concurrent may not be. Several projects' concurrency tests run inside a single database connection or transaction (Tauros's SQL sandbox), or are documented as untested (the template's nonce conflict under real concurrency). They prove the rules: a second decision is refused and a replay returns the original. They do not prove lock contention between real connections. Note which kind of evidence you have.

Design rule of thumb

For every state-changing tool an agent can call, write down:

  1. its idempotency identity: what makes two calls "the same";
  2. what a replay after a crash will send, and whether it has the same identity;
  3. the lock or constraint that serialises concurrent calls;
  4. the durable status that tells a restarted process what is unfinished;
  5. the test that proves each, and whether it uses real concurrency.