Part III — How are consequential decisions governed around a model?
Reference implementation: cognokratos/etf-research-agent, branch main, pinned in Source revisions.
Stack: Rust (versioned rules engine, MCP server, gateway), Python (NeMo Agent Toolkit), PostgreSQL, Keycloak. Built from the Part I template.
Part I ends with a production agent that can propose a change and have a human approve it. Part III moves into a domain where decisions have consequences and asks what else must be engineered around the agent: who owns the decision, what evidence supports it, what incomplete data means, how a human overrides it safely, and whether the decision can still be explained after the policy changes.
The domain is ETF research. It is used as a consequential domain to engineer in, not as financial content. Nothing in this part is an investment recommendation.
What the system is, precisely
- Policy is data.
data/rules_spec.jsonanddata/investor_profile.jsondefine the decision. Generic Rust code inmcp-server/src/rules.rsinterprets them. The rules specification is validated at boot; an invalid policy stops the service instead of silently rescoring. - The engine is authoritative.
rules::evaluatescores each fund against the mandate, handles missing data by explicit renormalisation and caps, and producesrules_decision, versioned byrules_versionandprofile_version. - The model is advisory. It receives an evidence contract, explains, and may recommend. A recommendation may be equal to the engine's or more conservative, never more optimistic, and it never becomes the decision.
- A human consents. Every state change (commit an evaluation, shortlist, assign) pauses for a human through a signed approval. Unlike Part I, where approvals are opt-in, they are mandatory here.
- The backend re-derives. At the point of mutation the MCP server locks the row, recomputes the evaluation, refuses the token if the decision the approval card displayed differs from the recomputation, re-checks hard constraints, and applies the change, consumes the nonce and appends the history row in one transaction.
The fund data is a dated snapshot (data_as_of values in data/etfs.json).
There is no market-data feed, no brokerage connection and no trade
execution anywhere in the code.
Display versus verification
The approval card shows "the engine's decision" as the model reported it. The approval layer does not fetch it. The signed token binds that displayed premise, and the backend refuses the token unless it equals its own recomputation. So a model that misreports the engine can mislead the card but cannot cause a wrong decision to be recorded. Approvals explains this in its section on the displayed premise. Part VI generalises it in Approval boundaries and exact-action binding.
How this part is organised
- The applied learning path, stages A1–A6 and the diagram the whole part elaborates.
- Lessons A1–A4 (policy, uncertainty, evidence, domain identity). Almost
all of their labs need only
cargoandpython3. - Follow one decision from fund facts to an audit row, with the owner of every step.
- Lessons A5–A6 and the adversarial-data lesson, which need a running stack and a model for the live experiments.
- Case studies and Challenges.
- Reference: architecture, approvals and limitations.
Book edition notes in this part
The lessons describe the model's recommendation as persisted "only inside an
approved record". At the pinned revision, read-only evaluations also write an
ETF_EVALUATED history row that can include a model-supplied recommendation.
It is still advisory and never becomes the decision. Notes in
A5,
the walkthrough
and the adversarial-data lesson
mark the places where this matters.
Running the labs
The deterministic labs need Rust 1.85 or newer and Python 3. The first
cargo test downloads and compiles dependencies. The live labs need the Part I
style Compose stack and a model endpoint. One caution: make rules-test
rewrites the committed deterministic baseline file, so run it in a scratch
checkout. See
Setting up each track.
Prerequisite. The Part I learning path, especially human-in-the-loop and Approvals.