Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Part I — How is a production agent engineered around an untrusted model?

Reference implementation: cognokratos/simple-agent-template, branch main, pinned in Source revisions.

Stack: Python (NVIDIA NeMo Agent Toolkit 1.9, NeMo Guardrails 0.21), Rust (gateway, MCP server), Keycloak, PostgreSQL, OpenTelemetry → MLflow, assistant-ui.

Part I is the foundation for the rest of the book. It starts from one observation: agentic AI is software engineering around a probabilistic decision-making component. Then, stage by stage, it asks why each next architectural component is needed.

What the system is

A customer-support agent over a small ticket database. A user signs in through Keycloak. A Rust gateway acting as a backend-for-frontend handles OIDC with PKCE, opaque sessions and CSRF, and mints the identity headers downstream services trust. The agent runs a bounded ReAct loop in the NeMo Agent Toolkit. Its tools live in a separate Rust MCP server, the only component besides PostgreSQL on the data network. The MCP server authenticates the agent with a service key and grants it exactly the tools on the agent's include: list. Input and output guardrails surround the loop. Every request produces one trace, and four evaluation suites measure behaviour.

The sample application is read-only by default: its two tools, search_tickets and get_ticket, only read. A state-changing action, changing a ticket's priority, can be enabled as an opt-in. It then requires a signed human approval. The agent runtime signs it with a secret the model never sees, and the MCP server verifies it at the point of mutation: it consumes a single-use nonce, locks the row, re-derives the current state, checks policy, and applies the change with an append-only audit record in one transaction.

What it does not do

The template's Limitations chapter is part of the curriculum. Among other things, it lists that there is no per-user data authorisation, no secrets manager, no migrations, no access control on the trace store and no dedicated guard model. Some guarantees are enforced by the database but not covered by an automated test, such as nonce conflicts under real concurrency. Read Limitations before you adapt the template.

How this part is organised

  1. The learning path. Stages 0–10, from the model as a probabilistic component to the composed production architecture. Each stage names the failure it addresses.
  2. Follow one request. One request through every trust boundary, with the code at each hop.
  3. Concepts. Eight concept chapters and the anti-patterns, which explain why each component exists.
  4. Labs. Ten hands-on labs: run, extend, break, evaluate, trace, add a mutation and an approval, then build your own domain agent.
  5. Challenges. Competency exercises without solutions.
  6. Reference. The precise how: architecture, security, approvals, guardrails, evaluation, observability, configuration, test scenarios, extension and limitations.

Running the labs

The full stack runs about ten services; a Docker VM with roughly 8 GB of memory is the practical minimum, and PII masking adds more. You also need make with bash, a host Python 3, Node 22 or newer for the UI checks, and Rust 1.88 or newer for the labs that change the MCP server (03 and 08). The model endpoint is any OpenAI-compatible API. The default is a local Ollama with qwen3:8b, which handles single-tool questions reliably but struggles with multi-ticket reasoning, as the labs themselves show. See Setting up each track.

What later parts take from here

Later partBuilds on
II · Sophosstages 0–3: model, loop, tool calling, MCP
III · ETF Researchthe whole path; its code is derived from this template
IV · Arktosstage 3 (capability boundaries), stage 8 (identity)
V · Taurosthe distinction between a model's capability and authority