Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Applied decision engineering

From cognokratos/etf-research-agent · docs/applied/README.md · pinned revision 493a67a721ef

Lessons, labs and case studies for engineers who already understand how a secure production agent is built and want to see what it takes to put one inside a consequential decision. The front door is APPLIED-LEARNING-PATH.md.

If you…Go to
are new to production agentsthe template's learning path first
already understand the template's architectureAPPLIED-LEARNING-PATH.md
want one decision end to endFollow one decision
want the real failuresCase studies
want to test yourselfChallenges
need implementation detailthe reference documents below

Learn

LessonStageYou will
01 — Policy is a programA1Change policy without touching code, then break it and see which validator notices
02 — Uncertainty is policyA2Predict the renormalised score of an incomplete record, and the false claim each naïve design would make
03 — Design evidence for the modelA3Build the three evidence contracts the IEAC-LSE incident went through
04 — Model the domain before the agentA4See why ranking and mutation need different identity semantics
05 — Recommendation, authority and consentA5Walk the trust matrix, and trace a model lying about the engine to a human
06 — Decisions that survive policy changeA6Commit a decision, move the mandate, and read two policy generations apart
07 — Evaluate the system, not just the modelA6Classify metrics by what they can prove, using the repository's own incidents
08 — Adversarial domain dataA3, A5Poison issuer text and separate "model compromised" from "authority compromised"

Also: Follow one decision · Case studies · Challenges

Reference

The lessons explain why and guide experiments. These documents are the canonical description of what the system does, and the lessons link into them rather than repeating them:

DocumentFor
ARCHITECTURE.mdDesign decisions and rejected alternatives
APPROVALS.mdThe approval boundary, token claims, transactional order
SECURITY.mdEach control and how to check it
EVALUATION.mdSuites and scoring methodology
EVALUATION_ANALYSIS.mdThe measured figures and what they mean
VERIFICATION.mdWhich command proves which control
LIMITATIONS.mdKnown gaps
DEMO.mdPrompts to type and what should happen

Ground rules for the labs

  • Run labs that modify tracked files only from a clean worktree. Check with git status --short, and commit or stash your own work first. The documented restore commands (git checkout -- data/, git checkout -- mcp-server/, …) deliberately discard the lab's local edits, and they discard any other uncommitted changes in the same paths along with them.
  • Most labs need no cluster. make rules-explain ETF=<etf_id> prints one fund's evaluation and component_evidence from the shipped engine; FACTS='<json>' overrides scored fields in memory. It never writes anything.
  • Labs that edit data/ say so and say how to undo it. The undo is always git checkout -- data/ evaluation/results/deterministic-etf-baseline.json. make rules-test regenerates that baseline from whatever policy is on disk, so an experiment leaves it modified until you restore it.
  • The running MCP server reads data/ at boot, mounted read-only. A policy or profile edit takes effect after docker compose restart mcp-server agent — no image rebuild. Restore and restart again when you are done.
  • Labs that mutate state use funds the approval-boundary suite resets (VJPN-LSE, VHYL-LSE, …), so make verify-approvals puts them back. audit_events is append-only by trigger; its rows stay, which is the point.
  • Model behaviour varies. Anything quoted from qwen3:8b says so, with its date and build. Your results may differ, and finding out is part of the exercise.