Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Lab 10 — Build your own domain agent

From cognokratos/simple-agent-template · docs/tutorials/10-build-your-own-domain-agent.md · pinned revision c66ce19d7b0c

Objective

Replace the support-ticket sample with your own domain while inheriting the authentication, isolation, guardrails, observability, evaluation and approval infrastructure unchanged.

Concept

Almost everything in this repository is domain-independent infrastructure around a probabilistic component. The domain is a small, well-defined set of files. Knowing exactly which ones is what makes the template reusable. → Learning path, stage 10

Architecture before

flowchart LR
    subgraph INFRA [Inherit unchanged]
        UI[assistant-ui] --> GW[gateway]
        GW --> NAT[NAT runtime + middleware]
        OBS[observability] --- NAT
        EVH[evaluation harness + scorers] --> NAT
        APP[approval token + verifiers]
    end
    subgraph DOMAIN [Replace]
        SCHEMA[db/init.sql + fixtures]
        TOOLS[MCP tool functions]
        CFG["agent/config.yml: system_prompt,<br/>include, tool_overrides,<br/>rail prompt"]
        PATTERNS["text_guardrails.py:<br/>critical patterns, allow templates"]
        DATA[evaluation/datasets/*.json]
        COPY[ui/app/page.tsx welcome copy]
    end
    NAT --> TOOLS --> SCHEMA

The authoritative list is in EXTENDING.md.

Exercise

Pick a domain with structured, authoritative state and questions that need interpretation. Examples: library loans, IT asset inventory, clinical trial site status (synthetic data only), internal incident tracker.

Follow EXTENDING.md's recommended sequence, and gate each step before moving on:

StepDoGate before moving on
1. Fork and isolateSet COMPOSE_PROJECT_NAME so your stack and this one keep separate volumes (CONFIGURATION.md)make static-check
2. SchemaReplace db/init.sql. Add CHECK constraints for every enumerated field. Keep the approval tables if you might need them later.make dev starts; make shell-db shows your data
3. ToolsReplace the tool functions in mcp-server/src/main.rs. Read-only first. Typed args, validated, bound, clamped (lab 03).cd mcp-server && cargo clippy --all-targets -- -D warnings && cargo test; make inspector-tools
4. Agent configRewrite system_prompt (keep {tools} / {tool_names}), include:, tool_overridesTool cards appear for your prompts in the UI
5. Input policyRewrite the self_check_input prompt, _CRITICAL_INPUT_PATTERNS and _READ_ONLY_TICKET_TEMPLATES for your domain. Anchor allow templates to the complete message.make verify-input-guardrails; your common queries are not blocked
6. Output policyReview the regex patterns and the Presidio entity list. Which entities are needed in answers (as names are here)?make verify-output-guardrails
7. Injection fixturesSeed dedicated, never-mutated records whose free text carries injection payloads (copy the shape of db/injection_test_fixtures.sql)—
8. EvaluationNew datasets for all four suites; set EVALUATION_TOOL_NAMES and the experiment namesmake eval-all passes on your chosen model
9. TracesCheck what your tool spans record. Real data will be sensitive.make trace-test
10. Mutations (only if needed)Add an approval-gated action (lab 08, EXTENDING.md)make verify-approvals-rust; make verify-approvals

Run it

make static-check
make dev && make wait
make test
make eval-all

Observe

Track how much you changed outside the "Replace" box. If you had to edit the gateway, NAT middleware, the observability package or the scorers, write down why. That is either a genuine gap in the template (worth an issue upstream) or a sign that domain logic is leaking into infrastructure.

Break it

Before step 5, run your domain's most common questions through the unchanged ticket-specific input rail and look at guardrail.decision_source in the traces. Expect some false positives: the self-check prompt and allow templates describe ticket lookups, not your domain. That is why step 5 exists.

Why it failed

Guardrail policy is a domain judgement. What counts as a harmful request, and which routine requests a classifier tends to over-block, differs per domain. The mechanism (layering, precedence, anchoring, decision recording) is reusable. The policy is not.

Architecture after

The same architecture with your domain in the "Replace" box, the same deterministic boundaries, and evaluation numbers for your agent on your model.

What you learned

  • The domain surface is schema, tools, prompts, input policy, fixtures and datasets.
  • Start read-only. Add mutation only through the approval boundary.
  • Every domain needs its own evaluation datasets and injection fixtures before it needs a bigger model.

Go deeper

This chapter is maintained in cognokratos/simple-agent-template beside the code it teaches. The book shows docs/tutorials/10-build-your-own-domain-agent.md at revision c66ce19d7b0c5c88c41b6860c78f66075485a07a (branch main). View source at this revision · Report a correction.

Corrections are made upstream against the current main branch and appear here when the book's pin for this source is updated.