Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Learning path

From cognokratos/simple-agent-template · docs/LEARNING-PATH.md · pinned revision c66ce19d7b0c

A progressive curriculum for experienced software engineers who want to build production AI agents. It assumes you know HTTP, APIs, databases, authentication, Docker, distributed systems, testing and observability. It teaches what is different when one component of your system is a probabilistic decision maker.

Agentic AI is software engineering around a probabilistic decision-making component.

The LLM is an untrusted decision maker. Security and authorization must be enforced deterministically outside the model.

Each stage answers one question: why did we need the next architectural component? Each points to a concept page (the why), a lab (the hands-on), and the reference manual (the precise how).

How long things take

InYou canRead
30 secondsSay what this project isREADME
5 minutesName the major components and why each existsThis page's summary table, ARCHITECTURE.md
30 minutesExplain how one request becomes tool calls and an answerFollow one request
A few hoursAdd a tool, run evaluations, read tracesLabs 01–07
The full pathReplace the sample domain with your own production agentLabs 08–10, CHALLENGES.md

The path at a glance

StageConceptComponent it addsFailure it addresses
0LLM as a probabilistic component—Treating model output as a return value
1Agent and agent loopNAT ReAct runtime, loop boundsOne model call cannot act on the world
2Tool callingTyped tools, native tool callingParsing prose into actions; unbounded actions
3MCP and capability boundariesRust MCP server, include: listAn agent that can do whatever its credentials can
4GroundingTools over the system of recordConfident answers from model memory
5Guardrails and untrusted dataNeMo input/output rails, deterministic patternsHostile input; leaked secrets and PII
6EvaluationMLflow suites, deterministic scorers"It worked when I tried it"
7ObservabilityOpenTelemetry → MLflow tracesNot knowing what the agent actually did
8Identity and trust boundariesKeycloak, gateway, service credentials, segmented networksThe model or the browser choosing who the user is
9Human-in-the-loop mutationSigned approvals, point-of-mutation policy, auditThe model authorizing its own actions
10Production architectureAll of the above, composed—

The whole path rests on one split:

Use the model for decisions that benefit from interpretation; use ordinary code for decisions that can be specified deterministically.

Probabilistic: the modelDeterministic: the code around it
interpreting intentauthentication
choosing tools where the request is ambiguousauthorization
synthesizing information from tool resultsvalidation of inputs and tool arguments
handling natural-language ambiguitybusiness rules
recommendations and explanationssorting and ranking when the rule is explicit
state transitions
approval verification
persistence constraints
audit records
the allowed capability surface
evaluation assertions
network boundaries

Lab 04 shows what happens when this line is drawn in the wrong place. Asked "which ticket should we handle first?", the default model, in one set of runs, picked a high ticket over an urgent one, working from grounded data. If "highest priority, then oldest" is the rule, it can be written as an ORDER BY, and handing that decision to the model adds risk without adding value. Let the model explain the ranking; let code compute it.


Stage 0: The LLM as a probabilistic component

Concept. A model call returns a sample: fluent, often right, sometimes wrong, never guaranteed. It has no access to your systems and no memory between calls.

Why it matters. Every later component exists because of this. You can't unit-test the model into correctness, and you can't let it hold authority.

In this repo. llms.primary in agent/config.yml: any OpenAI-compatible endpoint, temperature: 0.0. While writing these labs, the same prompt, on the same model at temperature 0, with the same configuration digest, behaved consistently within one agent build and differently after a rebuild (concept 3). Temperature 0 narrows variation. It does not make behaviour a stable property of your configuration.

Failure it addresses. Treating model output as a return value you can trust.

Lab. 04 — Break the agent Read. Concept 1, CONFIGURATION.md — model endpoint

Stage 1: Agent and agent loop

Why we need it. A single model call can only produce text. To answer "what's the history of TKT-1001?" something has to fetch data, show it to the model and let it continue. That something is an agent runtime: a loop that asks the model for its next action, executes it, and feeds back the result.

In this repo. NAT's ReAct graph, wrapped by streaming_react_agent in register.py. The loop is bounded by max_tool_calls, max_history, retries and timeouts.

Failure it addresses. Models that can't act. Bounding the loop addresses its own new failure: runaway loops.

Lab. 01 — Run the agent Read. Concept 1, Follow one request

Stage 2: Tool calling

Why we need it. The loop needs a contract for "actions". Tools are typed functions the model can request. Native tool calling returns those requests as structured data rather than prose to be parsed.

In this repo. use_native_tool_calling: true. search_tickets and get_ticket with JSON Schemas derived from Rust structs. Descriptions tuned in tool_overrides.

Failure it addresses. Brittle text parsing, and actions with no contract. It also introduces a new risk: tool descriptions are prompts, and tool granularity determines how many round trips (and failure points) a task needs.

Lab. 02 — Understand tool calling Read. Concept 2

Stage 3: MCP and capability boundaries

Why we need it. If tools are functions inside the agent process, the agent process needs every credential every tool needs. Moving tools behind a protocol (MCP) into a separate service makes the tool list the capability boundary. The agent can do exactly what the tools implement, and nothing else.

In this repo. The Rust MCP server (mcp-server/src/main.rs), service-key authenticated, the only service on data_net. The agent's include: list grants the tools. A granted tool that doesn't exist stops the agent from starting.

Failure it addresses. An agent with database credentials, and a run_sql tool. See ANTI-PATTERNS.md.

Lab. 03 — Add an MCP tool Read. Concept 2, ARCHITECTURE.md — network segmentation

Stage 4: Grounding in authoritative systems

Why we need it. With tools available, the model can still answer from memory. Grounding makes the system of record the only source of domain facts, and makes answers checkable against what the tools returned.

In this repo. The system prompt's "use the tools for every factual statement"; PostgreSQL CHECK constraints; the grounding evaluation suite.

Failure it addresses. Hallucinated authoritative state. It also shows the limit: grounded inputs don't guarantee correct reasoning (lab 04, experiment 3).

Lab. 04 — Break the agent (experiments 3–4) Read. Concept 3

Stage 5: Guardrails and untrusted data

Why we need it. Users can send hostile input, and tool results can carry hostile text. Answers can contain secrets or PII. Guardrails filter what goes in and what comes out, using an LLM classifier backed by deterministic patterns.

In this repo. NeMo Guardrails via text_guardrails: length bound, critical patterns, LLM self-check, anchored allow templates, regex output blocking, buffered Presidio masking.

Failure it addresses. Direct prompt injection, harmful requests, credential and PII disclosure. It does not address indirect injection or authorization. Guardrails are not authorization.

Lab. 07 — Experiment with guardrails, 04 — Break the agent (experiment 1) Read. Concept 4, GUARDRAILS.md

Stage 6: Evaluation

Why we need it. Every previous stage changes model behaviour: prompts, tool descriptions, rails. Without measurement, you can't tell whether a change helped or hurt.

In this repo. Four suites (tools, guardrails, grounding, injection), source-controlled datasets, deterministic scorers, gate metrics, latency percentiles and provenance, all in MLflow.

Failure it addresses. Shipping on anecdotes. It also teaches its own lesson: a scorer only catches what it asserts.

Lab. 05 — Evaluate the agent Read. Concept 5, EVALUATION.md

Stage 7: Observability and traces

Why we need it. An evaluation tells you that a case failed. Only a trace tells you why: which tools, which arguments, which rail decision, where the time went.

In this repo. NAT and Guardrails spans joined into one trace per request (trace_context.py), readable question and answer, header redaction, opt-in user attribution, exported over OTLP to MLflow.

Failure it addresses. Debugging a runtime-chosen code path from scattered logs, and accidentally turning the trace store into a credential or PII store.

Lab. 06 — Debug with traces Read. Concept 6, OBSERVABILITY.md

Stage 8: Authentication, identity and trust boundaries

Why we need it. Once the agent serves real users, it must know who is asking, and neither the browser nor the model may decide that. Every internal hop must prove who is calling, because services that trust identity headers must only accept them from the component that minted them.

In this repo. Keycloak OIDC + PKCE in the Rust gateway; opaque sessions; CSRF; schema re-serialisation; gateway-minted x-authenticated-* headers; service credentials on gateway → NAT and NAT → MCP; seven segmented networks.

Failure it addresses. Spoofed identity, model-chosen identity, and "it's on a private network" as authentication.

Lab. 01 — Run the agent (Break it) Read. Concept 7, SECURITY.md, ARCHITECTURE.md

Stage 9: Human-in-the-loop and controlled mutations

Why we need it. Reading is recoverable. Writing is not. If the model can change state, an injected instruction or a reasoning error becomes a real change with an audit record that looks intentional. The model may propose. A human decides. Deterministic code verifies and applies.

In this repo. Optional and off by default: approval.py (proposal → human prompt → signed token), interaction_guard.py (only the prompted user, only an offered choice), mutation.rs (nonce, row lock, re-derived state, apply_policy, apply + append-only audit in one transaction).

Failure it addresses. The model authorizing its own actions; replayed, stale or altered approvals; editable audit trails.

Lab. 08 — Add a state-changing action, 09 — Add human approval Read. Concept 8, APPROVALS.md

Stage 10: Production agent architecture

Why it looks like this. Put the stages together and every component has a specific job, with the model contained in the middle:

  • the gateway decides who the user is;
  • the agent runtime runs a bounded loop and decides nothing about authority;
  • guardrails filter text at the edges of the loop;
  • MCP tools define everything the model can cause;
  • the database is the authority on state and enforces its own invariants;
  • approvals turn model proposals into human-authorized, policy-checked, audited changes;
  • traces record what happened, and evaluations measure how well.

See the full diagram, with trust boundaries, in ARCHITECTURE.md, and the end-to-end request in Follow one request.

ConceptImplementation in this repo
Agent runtimeNeMo Agent Toolkit (NAT) 1.9
Tool interoperabilityMCP (streamable HTTP)
Capability implementationRust MCP server (rmcp, sqlx)
GuardrailsNeMo Guardrails 0.21 + application-level deterministic layers
TracingOpenTelemetry → OpenTelemetry Collector
Experiment tracking, trace storeMLflow
AuthenticationKeycloak (OIDC), Rust gateway (BFF)
PersistencePostgreSQL
UIassistant-ui on Next.js

These are implementation choices. The concepts carry over to other runtimes, protocols and stores. What doesn't carry over automatically is the discipline: deterministic boundaries around a probabilistic core.

Before production, read LIMITATIONS.md. It lists what this template does not do (per-user data authorization, secrets management, migrations, trace-store access control, a dedicated guard model, among others).

Lab. 10 — Build your own domain agent Then. CHALLENGES.md, ANTI-PATTERNS.md