Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Applied learning path: from agent to decision system

From cognokratos/etf-research-agent · docs/APPLIED-LEARNING-PATH.md · pinned revision 493a67a721ef

simple-agent-template teaches how to build a production AI agent.

etf-research-agent teaches how to turn that agent into a governed decision system.

Here the interesting question is no longer "how does an agent call a tool?" It is: who owns the decision, what evidence supports it, what happens when the data is incomplete, how does a human override it safely, and can we still explain that decision after the policy changes?

A production agent is only part of the system. In a consequential domain, authority, policy, evidence, uncertainty, human consent, auditability and evaluation have to be engineered explicitly around it.

Before you start

This path assumes you have completed, or could teach, the template's learning path: the model as a probabilistic component, agent loops, tool calling, MCP as a capability boundary, grounding, guardrails, evaluation, tracing, identity, signed approvals. None of that is re-taught here. Where a lesson depends on one of those concepts it links to it, and the link is the prerequisite.

If you have not done the template path, start there. If you only want to run the ETF application, the README and DEMO.md are the right entry points, not this page.

The domain is ETF research, and it is used as a consequential domain to engineer in, not as financial content. The system ends at decision support: no brokerage, no orders, no forecast. A score is a policy and fit evaluation of a dated data snapshot against a written mandate. Nothing in this curriculum is an investment recommendation, and every example is about software and system design.

The path at a glance

StageApplied questionMain ideaFailure it preventsLesson
A1What belongs in deterministic policy?Policy as executable, versioned dataThe model, or a code edit nobody reviewed as policy, deciding the outcome01 — Policy is a program
A2What does incomplete evidence mean?Missing data, renormalisation and capsA score that claims a measurement nobody made02 — Uncertainty is policy
A3What evidence must the model receive?Evidence contracts for explanationCorrect decisions explained wrongly03 — Design evidence for the model
A4What exactly is the domain entity?Fund identity versus listing identityDouble counting in rankings; mutations on a guessed record04 — Model the domain before the agent
A5Who recommends, decides and authorizes?Rules versus model versus humanThe model's opinion, or the model's claim, becoming the decision05 — Recommendation, authority and consent
A6How do we prove and preserve decisions?Evaluation, provenance and policy-versioned auditGreen dashboards over wrong answers; audit rows nobody can interpret06 — Decisions that survive policy change, 07 — Evaluate the system

Cross-cutting, linked from the stages that need it:

The diagram the whole path elaborates:

flowchart TD
    F["Verified fund facts<br/>data/etfs.json → PostgreSQL"] --> E
    P["Policy<br/>data/rules_spec.json"] --> E
    M["Mandate<br/>data/investor_profile.json"] --> E
    E["Deterministic engine<br/>rules::evaluate — authoritative"] --> RD["rules_decision<br/>+ component_evidence"]
    RD --> L["LLM: explains, may recommend<br/>advisory only"]
    L --> H["Human: confirms or overrides<br/>with a rationale"]
    H --> T["Signed approval<br/>binds the premise the human was shown"]
    T --> B["Backend at the point of mutation<br/>re-derives, re-checks, refuses or applies"]
    B --> A["Mutation + append-only audit<br/>with rules_version and profile_version"]
    RD -. "recomputed, never trusted" .-> B

Authority enters at the top as data, is computed once by code, passes through the model without being owned by it, is consented to by a person, and is re-derived by the backend before anything persists. Every lesson is about one edge of that graph.

How long things take

InYou canRead
10 minutesSay how this repository differs from the template, and where authority livesThis page, the walkthrough's summary table
45 minutesTrace one decision from fund facts to an audit rowFollow one decision
An afternoonDo the labs in A1–A4; almost all of them need no cluster and no modelLessons 01–04 with make rules-explain
A dayRun the live experiments in A5–A6 and lesson 08A running cluster and a model; README
Open-endedPort the architecture to another consequential domainChallenges

Every deterministic lab runs with python3 and cargo only. make rules-explain prints one fund's evaluation exactly as the engine and the read models produce it, and accepts in-memory overrides of the scored facts, so most experiments never touch data/ at all.


Stage A1: Policy is a program

Question. If a decision can be specified, where should the specification live?

Main idea. As reviewable, versioned, validated data, interpreted by generic code. rules.rs defines how policy is interpreted; rules_spec.json and investor_profile.json determine what the decision is.

In this repository. data/rules_spec.json, data/investor_profile.json, RulesSpec::validate and rules::evaluate in mcp-server/src/rules.rs.

Failure it prevents. A threshold buried in code that nobody reviews as policy; an invalid policy that silently rescored the universe instead of refusing to boot.

Experiment. Change a cost band and watch five funds move with every test still green; break the specification and see which validator catches each break.

Prerequisite. Template Stage 0 and the split it ends on, and overusing agents for deterministic workflows.

Reference. ARCHITECTURE.md — policy is data, not code

Stage A2: Uncertainty is policy

Question. Grounded data can still be incomplete. What does "unknown" mean for a decision?

Main idea. Absence is neither zero nor nothing. Absent weight leaves the denominator, the absence is published with the weight it removed, and caps exist because renormalisation flatters exactly the records it helps.

In this repository. missing_data_policy and decision_caps in the specification; the Normalization and ComponentBreakdown types in rules.rs. AGGH-XETRA is the shipped example: 84 renormalised, held at research.

Failure it prevents. "We measured it and it was bad" said about a fund nobody measured; scores that look comparable and are not.

Experiment. Remove metrics from a fund in memory, predict the denominator and the factor, then compare with what the engine returns — and with what the two naïve designs would have claimed.

Prerequisite. Template Stage 4 — grounding.

Reference. ARCHITECTURE.md — missing data is a policy

Stage A3: Design evidence for the model

Question. The backend's answer is right. What must the model receive to explain it right?

Main idea. Tool design is information architecture for a probabilistic consumer. A relationship the contract does not make explicit is a relationship the model will eventually drop or invert.

In this repository. component_evidence in mcp-server/src/domain.rs and RESEARCH_CONTEXT_REQUIRED_ELEMENTS in mcp-server/src/server.rs.

Failure it prevents. The IEAC-LSE regression: the explanation stopped naming "bond", then named it with the direction inverted. Both with the correct decision.

Experiment. Build three progressively richer payloads from the engine's real output and see what each one makes it possible to explain.

Prerequisite. Template concept 2 — when agent problems are API-design problems.

Reference. EVALUATION_ANALYSIS.md — the explanation lost its facts

Stage A4: Model the domain before the agent

Question. What is the thing being decided about?

Main idea. Fund identity (ISIN) and listing identity (ticker + venue) are different entities, and the right one depends on the operation. Aggregation may collapse identities; mutation must resolve one canonical resource.

In this repository. fund_identity() in domain.rs, collapse_listings and resolve in server.rs, the cross-listing check in scripts/validate_etf_fixtures.py.

Failure it prevents. One fund taking two slots in a top five; a mutation landing on a listing the system guessed.

Experiment. Rank United States shortlist candidates with and without listing collapse; resolve VUSA and see why a ranking may group it and a mutation may not.

Prerequisite. Template Stage 3 — capability boundaries.

Reference. ARCHITECTURE.md — listing identity versus fund identity

Question. The engine computes, the model recommends, the human chooses. Which of those is the decision, and who can change it?

Main idea. rules_decision, llm_recommendation, the human's choice and the persisted final_decision are four separate values with four separate owners. The backend re-derives the authoritative one at the point of mutation and never trusts anybody's claim about it — including a claim the human approved.

In this repository. rules::reconcile_decision, blocking_hard_constraint, compare_recommendation; the commit path in server.rs; the approval prompt in approval.py.

Failure it prevents. The model holding the default in the conservative direction; a model misreporting the engine to get a promotion past a human.

Experiment. Walk the trust matrix through reconcile_decision; trace the observed case where the model told a human that a rejected fund's engine decision was shortlist.

Prerequisite. Template Stage 9 — human-in-the-loop mutation. Token mechanics are not re-taught.

Reference. ARCHITECTURE.md — who holds the default decision, APPROVALS.md

Then: 08 — adversarial domain data, which asks the same authority question about text instead of actors.

Stage A6: Prove and preserve decisions

Question. How do you show the system decides correctly, and how does a decision stay interpretable after the policy moves?

Main idea. Two halves. Preservation: a historical decision is only interpretable with the policy and mandate generation it was made under, so the audit record carries both and never merges generations. Proof: separate deterministic claims from model-dependent ones from presentation signals, and give each metric a maturity — gate, diagnostic, experimental — that matches what it can actually prove.

In this repository. decided_rules_version / decided_profile_version, policy_generations in domain.rs, audit_events in db/init.sql; the five suites in evaluation/scorers.py and the deterministic baseline emitted by make rules-test.

Failure it prevents. An assignment event presenting a v1 score under a v2 version; a 100× expense-ratio error shipping under green gates.

Experiment. Commit a decision, change the mandate, assign the fund, and read the two generations apart. Then take the TER, IEAC and policy-0.4 incidents and say, for each, what the evaluation proved and what it did not.

Prerequisite. Template Stage 6 — evaluation and concept 3 — current state versus history.

Reference. ARCHITECTURE.md — audit events are one snapshot, EVALUATION.md, EVALUATION_ANALYSIS.md


What this path deliberately does not teach

TopicWhere it is taught
Agent loops, ReAct, native tool callingTemplate stages 1–2
MCP, the tool list as the capability boundaryTemplate stage 3
Grounding basics; grounded is not correctTemplate concept 3
Input/output rails, data-plane injection basicsTemplate concept 4
Deterministic scorers, evaluation methodologyTemplate concept 5
Tracing and the trace pipelineTemplate concept 6, OBSERVABILITY.md
Gateway, OIDC, service credentials, networksTemplate concept 7, SECURITY.md
Approval tokens, nonces, the interaction guardTemplate concept 8, APPROVALS.md
The network path of one requestTemplate request walkthrough

How to read the claims in these lessons

The same rule as the reference documentation:

Implementation and executable verification are authoritative. Reference documentation is the canonical description. These lessons explain why and guide experiments.

If a lesson and the code disagree, the code is right and the lesson is a bug. make docs-check keeps the links, anchors and make targets honest; it cannot check that prose is true.

Two kinds of statement appear, and they are never mixed:

  • Guaranteed — a property of deterministic code, asserted by make rules-test, make etf-check, make verify-approvals or make verify-approvals-rust. A failure is a defect.
  • Observed — a behaviour of qwen3:8b, on a named build and date, in a stated number of runs. It can differ on your machine, your model or tomorrow, and finding out is part of the exercise. Observations are never promoted into architecture guarantees.

Numbers quoted from the engine (scores, weights, factors) come from the shipped fixtures — rules 1.1.0, profile 1.0.0 — and make rules-explain reproduces every one of them.