Applied decision engineering
From cognokratos/etf-research-agent · docs/applied/README.md · pinned revision 493a67a721ef
Lessons, labs and case studies for engineers who already understand how a secure production agent is built and want to see what it takes to put one inside a consequential decision. The front door is APPLIED-LEARNING-PATH.md.
| If you… | Go to |
|---|---|
| are new to production agents | the template's learning path first |
| already understand the template's architecture | APPLIED-LEARNING-PATH.md |
| want one decision end to end | Follow one decision |
| want the real failures | Case studies |
| want to test yourself | Challenges |
| need implementation detail | the reference documents below |
Learn
| Lesson | Stage | You will |
|---|---|---|
| 01 — Policy is a program | A1 | Change policy without touching code, then break it and see which validator notices |
| 02 — Uncertainty is policy | A2 | Predict the renormalised score of an incomplete record, and the false claim each naïve design would make |
| 03 — Design evidence for the model | A3 | Build the three evidence contracts the IEAC-LSE incident went through |
| 04 — Model the domain before the agent | A4 | See why ranking and mutation need different identity semantics |
| 05 — Recommendation, authority and consent | A5 | Walk the trust matrix, and trace a model lying about the engine to a human |
| 06 — Decisions that survive policy change | A6 | Commit a decision, move the mandate, and read two policy generations apart |
| 07 — Evaluate the system, not just the model | A6 | Classify metrics by what they can prove, using the repository's own incidents |
| 08 — Adversarial domain data | A3, A5 | Poison issuer text and separate "model compromised" from "authority compromised" |
Also: Follow one decision · Case studies · Challenges
Reference
The lessons explain why and guide experiments. These documents are the canonical description of what the system does, and the lessons link into them rather than repeating them:
| Document | For |
|---|---|
| ARCHITECTURE.md | Design decisions and rejected alternatives |
| APPROVALS.md | The approval boundary, token claims, transactional order |
| SECURITY.md | Each control and how to check it |
| EVALUATION.md | Suites and scoring methodology |
| EVALUATION_ANALYSIS.md | The measured figures and what they mean |
| VERIFICATION.md | Which command proves which control |
| LIMITATIONS.md | Known gaps |
| DEMO.md | Prompts to type and what should happen |
Ground rules for the labs
- Run labs that modify tracked files only from a clean worktree. Check with
git status --short, and commit or stash your own work first. The documented restore commands (git checkout -- data/,git checkout -- mcp-server/, …) deliberately discard the lab's local edits, and they discard any other uncommitted changes in the same paths along with them. - Most labs need no cluster.
make rules-explain ETF=<etf_id>prints one fund's evaluation andcomponent_evidencefrom the shipped engine;FACTS='<json>'overrides scored fields in memory. It never writes anything. - Labs that edit
data/say so and say how to undo it. The undo is alwaysgit checkout -- data/ evaluation/results/deterministic-etf-baseline.json.make rules-testregenerates that baseline from whatever policy is on disk, so an experiment leaves it modified until you restore it. - The running MCP server reads
data/at boot, mounted read-only. A policy or profile edit takes effect afterdocker compose restart mcp-server agent— no image rebuild. Restore and restart again when you are done. - Labs that mutate state use funds the approval-boundary suite resets
(
VJPN-LSE,VHYL-LSE, …), somake verify-approvalsputs them back.audit_eventsis append-only by trigger; its rows stay, which is the point. - Model behaviour varies. Anything quoted from
qwen3:8bsays so, with its date and build. Your results may differ, and finding out is part of the exercise.