Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

What this book teaches

Agentic systems stop being demos once their output starts to change something: a ticket's priority, a recorded investment decision, a wallet, an invoice that someone will pay. From that point on, getting the right answer is not enough. You also have to know who was allowed to cause each change, on exactly what, and whether you can prove it afterwards.

This book teaches the engineering judgment for that situation, through five runnable systems. Each one isolates one engineering problem:

PartEngineering questionReference implementation
IHow is a production agent engineered around an untrusted probabilistic component?simple-agent-template
IIOnce you have an agent, how does it become durable software that survives time, state, crashes and restarts?sophos-agent
IIIHow do you put deterministic policy, evidence and human authority around probabilistic reasoning?etf-research-agent
IVHow can probabilistic software request cryptographic capabilities without becoming the custodian of cryptographic authority?arktos-wallet
VHow can agents do operational financial work while humans keep the authority?tauros-revenue
VIWhat do the five answers have in common, and where do they differ?the book's own synthesis

The idea that runs through every part

The model is treated as a probabilistic component: useful for interpreting intent, choosing among tools, synthesising evidence and explaining. It is unreliable as a source of truth, and it has no place as a holder of authority. So each project draws the same line in a different domain:

Use the model for decisions that benefit from interpretation. Use ordinary code for decisions that can be specified deterministically.

Drawing that line puts most of the work in the software around the model:

  • Capability boundaries. What can the model cause at all? (MCP tool lists, four wallet tools, eight reviewed Ash actions.)
  • Identity. Who is the caller, and who decides that? (A gateway, an API key resolved out of band, distinct human and agent actors.)
  • Authority. Who may make this particular change? (Policies, signed approvals, a human approver who must name the exact revision.)
  • Durability. What survives a crash, and what happens on a retry? (Checkpoints, run records, idempotency keys, nonces.)
  • Evidence and audit. Can the decision be explained and reconstructed later? (Evidence contracts, append-only histories, policy versions.)

What you should be able to do afterwards

  • Point at the exact line of code that stops a model from authorising its own action, in each of the five systems, and explain why a prompt could not bypass it.
  • Tell apart what an approval interface displays and what the mutation boundary independently verifies, and design the second so it never relies on the first.
  • Separate runtime state, domain state and audit history, and say which one answers "what happened?".
  • Design a state-changing tool so that a retry, a replay after a crash or two concurrent requests cannot apply it twice.
  • Give an agent a capability, such as deriving an address, without giving it the secret behind it, and say who still holds custody.
  • Decide which guarantees belong in a database constraint, a type system, a policy engine, a protocol boundary or a human decision.

How the book is grounded

The projects are runnable, and their lessons are written against their code and tests. Wherever the book's editors found a lesson's statement broader than the pinned code supports, the chapter carries a labelled Book edition note. The upstream text is left as it is. These notes are listed, with their evidence, in the project's content audit. Implemented behaviour, behaviour demonstrated by a test, documented limitations and future exercises are kept distinct throughout.