Challenges
From cognokratos/etf-research-agent · docs/applied/CHALLENGES.md · pinned revision 493a67a721ef
Competency exercises for the applied path. They have requirements and a definition of done, and no solutions. The template's challenges cover adding tools, entities and approval-gated actions; these assume you can do that and ask a different question: can you change a decision system without moving authority somewhere it should not be?
Ground rules: work on a branch, keep make static-check and make rules-test
green, and commit the regenerated baseline with any policy change so the diff is
reviewable. Where a challenge is a design exercise, the deliverable is a written
design that answers every question listed, not code.
| Challenge | Kind | Builds on |
|---|---|---|
| 1. Add a new policy dimension | implementation | 01, 02, 03 |
| 2. Add a second investor profile | implementation | 01, 06 |
| 3. Build a semantic-direction metric | implementation | 03, 07 |
| 4. Design V2 fund/listing persistence | design | 04, 06 |
| 5. Caller-scoped service identity | security design | 05, template stage 8 |
| 6. Port the architecture to another domain | design, then implementation | all |
Challenge 1: Add a new policy dimension
Score a property of the fund the engine does not score today.
Pick it carefully. It must be a property of the fund that is published and
checkable, not a forecast. return_3y_annualized is in the snapshot and is the
obvious candidate, and scoring it would break the one claim this system rests on
— that a score is quality and fit, not expected performance. If you choose a
field that is not in the snapshot yet, you are also adding it to the data.
Requirements.
- No ETF-specific threshold, weight or band in Rust. The schema change (a new
scorable field) is code; how it counts is
rules_spec.json. - Validation: a specification that names your field wrongly, or gives it
inconsistent weights, fails at boot and in
make rules-test. - Missing-data semantics, decided and written down: is the field critical? does it
count toward completeness? What does an unrecognised value mean? The answer goes
in
missing_data_policy, with the reasoning beside it, ascritical_fields_notedoes today. - Evidence: the new rule appears in
component_evidencewith a meaningful note, andrequired_elementsstill asks for everything an explanation now owes. - Deterministic tests: labelled cases, with rationales, that pin the decisions your dimension is meant to move — and at least one it is meant not to move.
- No LLM authority: the model reads the result; nothing it says feeds the score.
Done when a later band change for your dimension is a JSON-only diff that
moves the expected funds in the regenerated baseline, and removing the field from
one fund in memory (make rules-explain FACTS=…) renormalises and reports
exactly as your policy says.
What makes it hard. Every layer touches it — EtfFacts, SCORABLE_FIELDS, the
seed, the schema's CHECK constraints, scripts/validate_etf_fixtures.py, the
read models — and a weight added to one component has to come from somewhere,
which moves every score in the universe.
Challenge 2: Add a second investor profile
Evaluate the same fund universe under two mandates.
Requirements.
- Clear profile provenance: every evaluation, committed decision and audit row
says which profile produced it, not only which version. Today
audit_eventsstoresprofile_versionand noprofile_id; with one profile that is unambiguous, and with two,"1.0.0"is not. - No policy-generation ambiguity: lesson 06's guarantees hold across profiles, not just across versions.
- Labelled expectations per profile. Several engine tests encode the shipped mandate's behaviour (lesson 01); decide which are properties of the engine and which are properties of a mandate.
- A comparison: a reviewable artifact showing how each fund's decision moves between the two mandates, generated by the shipped engine.
The authority question you must answer first. Who selects the profile for a request? If the model can choose it — a tool argument, say — the model has chosen the decision by choosing the mandate. If the user can, how is that bound to their identity and recorded? Write that down before writing code.
Done when the same fund can be committed under each profile by different people, and its history says which mandate each decision answered to.
Challenge 3: Build a semantic-direction evaluation metric
Detect whether an explanation says a fact helped or hurt in the direction the deterministic evidence supports.
research_required_facts_present checks that "bond" appears. It passes "as a bond
fund it earned only 0.1 of the asset-class rule" and "bond exposure aligns with a
high risk tolerance" alike (lesson 07).
This is not trivial. Direction is expressed in prose in open-ended ways — "weak fit", "works against", "only 0.6 of 6", "unlike an equity fund", "a poor match" — and negation, comparison and hedging all flip or blur it. A regex that looks plausible on five examples will be wrong on the sixth.
Requirements.
- Ground truth from the engine, never from a model:
earned_fractionincomponent_evidencedecides which direction is correct. - A false-positive set (correct answers your scorer must not flag) and a
false-negative set (inversions it must flag), written before the scorer, as
tests in
evaluation/tests/. - A replay over real captured answers, with every flag read by a person.
- A stated maturity — experimental — and a written promotion criterion: how many runs, over how long, with what false-positive budget, before it can become a diagnostic or a gate.
- No LLM judge. The project's argument is that the first model is not load-bearing; a second one in the measuring instrument would undo it.
Done when the inverted IEAC-LSE answer fails your metric, every answer in your false-positive set passes, and the limits of what it can see are written next to it.
Challenge 4: Design V2 fund/listing persistence
Design the refactor of V1's listings table into fund → listings. Do not
implement it.
Constraints.
- External semantics preserved: the read models keep their shape, as ARCHITECTURE.md claims they can.
- Mutation still requires an exact, canonical resource. Decide whether that resource is now a fund or a listing, and defend it.
- Rankings aggregate at the level the operation needs, and say which.
- Every existing
audit_eventsrow stays interpretable — including itsetf_id, its versions and its committed score — after the migration. - The cross-listing agreement check in
validate_etf_fixtures.pybecomes a database constraint or is shown to be unnecessary.
Deliverable. A design document: schema, migration, the meaning of every
identifier before and after, how search_etfs, get_research_summary, the
resolver and the three approval actions change, and the questions from lesson
04, step 5
answered — including the share-class and index-exposure levels.
Challenge 5: Caller-scoped service identity
LIMITATIONS.md records it: the service
credential carries the authority to assert any identity. NAT believes whatever
x-authenticated-user-id a key-holding caller sends, the approval boundary binds
to it, and two callers hold the key — the gateway and the evaluator. The evaluator
asserts evaluation-harness; nothing stops it asserting a real researcher.
Design:
gateway credential → may assert a human identity
evaluation credential → may assert only the synthetic evaluation principal
Answer, in the design.
- Where the binding between credential and assertable identity lives, and why
there rather than in the gateway or the evaluator (
RequireIdentityHeaderMiddlewareinagent/src/nat_streaming_react/fastapi_worker.pyis where the identity requirement is enforced today). - What an approval minted on behalf of the synthetic principal must be refused for, and where.
- How
make auth-testandIdentityBoundaryTestschange: the new negative cases, stated as tests. - Rotation, and how two credentials are configured without a half-configured
deployment starting cleanly — the failure
scripts/verify_approval_surface.pyexists to catch for the approval secret. - What workload identity would replace, and what it would not.
Implementing it is optional. Asserting it — with the negative cases — is the part that matters.
Challenge 6: Port the architecture to another consequential domain
Choose a domain where a decision has consequences for someone other than the person asking: Swiss real-estate research, credit underwriting, AML case review, insurance claims triage, vendor-risk assessment.
Before writing the agent, write the table:
| Your domain | |
|---|---|
| Facts — what is observed, from which source, as of when, with what provenance | |
| Policy — what is specified deterministically, as data, interpreted by code | |
| Mandate — the per-user or per-client inputs the policy is evaluated for | |
| Uncertainty — which facts are routinely missing, which are critical, and what absence means | |
| Identity — every entity identity, and which one each operation needs | |
| Decision authority — what computes the default decision | |
| Advisory model role — what the model may explain, recommend, compare | |
| Human role — who may confirm or override, in which directions, with what record | |
| Non-bypassable constraints — what no actor may override | |
| Audit and version semantics — what a decision record must carry to stay interpretable after the policy changes |
Then find your domain's equivalents of this repository's three hard cases: a record that is excellent and wrong for this mandate (IEAC-LSE), a record with a missing critical fact that renormalisation would flatter (AGGH-XETRA), and an identifier that names more than one thing (VUSA).
Done when your engine computes every decision with no model in the loop and passes deterministic tests, your agent can explain a capped decision with the facts and their direction, and a human override in your domain is recorded with the actor, the rationale, the engine's decision and the policy generation.
Open problems
Observations about current behaviour, found while writing this curriculum and deliberately not changed by it. Each is a design question before it is a code change.
| Observation | Where it shows | The question |
|---|---|---|
With no scorable profile-fit weight, profile_fit.fraction is reported as 0.0 and CAP-PROFILE-FIT fires with a message claiming the fund "earns less than half" of weight that does not exist. Reachable on its own only with preferences switched off; the outcome is conservative, the explanation is not true | lesson 02, step 6 | What should "fit" be when nothing about fit is known, and which cap should say so? See LIMITATIONS.md |
The approval prompt displays the model-supplied rules_decision under the label "Deterministic engine (authoritative)". The mutation is safe — the backend recomputes — but a human can consent to a false premise | lesson 05, step 3 | Should the approval function fetch the engine's decision itself, and what then is expected_choice? See LIMITATIONS.md |
rules_version and profile_version are hand-maintained labels; a policy edit that does not bump them is indistinguishable in history | lesson 06 | Identify generations by content hash? Refuse to boot on an unchanged version with changed content? |
Committed decisions record the policy and mandate generation, not the fact snapshot (data_as_of) they were made on | lesson 06 | What is the smallest record that makes a decision reproducible? |
RulesSpec::validate does not require band fractions to be monotonic; a non-monotonic band passes both validators and is caught only by labelled cases | lesson 01, step 4 | Which semantic properties belong in validation, and which in expectations? |
| No metric checks the direction of an explanation, or flags an absence claim made without a tool call, or an authority claim from untrusted text repeated as fact | lessons 07, 08 | Challenge 3, and its siblings |
audit_events records profile_version but not profile_id | challenge 2 | Unambiguous today, ambiguous with a second profile |