Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

5. Recommendation, authority and consent are different things

From cognokratos/etf-research-agent · docs/applied/05-recommendation-authority-and-consent.md · pinned revision 493a67a721ef

Stage A5 of the applied learning path. Prerequisite: template Stage 9 — human-in-the-loop mutation and concept 8 — model advice versus authoritative policy.

Never sign the model's claim about authoritative state and assume that makes it authoritative.

The template teaches the mechanics of a safe mutation: proposal, prompt, signed token, point-of-mutation verification, one transaction. They are not repeated here. This lesson is about the domain semantics those mechanics carry when the thing being approved is a decision with a computed default: who produces each value, who may move it, in which direction, and what the backend believes.

Four values, four owners

Note

Book edition note. At the pinned revision llm_recommendation is also persisted outside approved records: when the model passes a recommendation to evaluate_etf, the read-only evaluation writes an ETF_EVALUATED audit row that includes it (server.rs). The value is still advisory: the engine's decision is recomputed, and the model's value never becomes final_decision. But "only inside an approved record" is narrower than the code.

ValueProduced byAuthorityPersisted as
rules_decisionthe deterministic engine, recomputedauthoritative, and the default — alwaysaudit_events.rules_decision
llm_recommendationthe modeladvisory only; may equal or be more conservative than rules_decision, never more optimisticaudit_events.llm_recommendation, only inside an approved record
the human's choicethe authenticated personfinal, within constraints; a rationale is required whenever it differs from rules_decision, in either directionthe token's choice; override_applied, override_rationale
final_decisionthe backend, after reconciliation and constraintswhat actually happenedetfs.decision, audit_events.final_decision, with rules_version and profile_version

They are named separately everywhere they appear — the approval prompt, the token, the MCP validation, the audit row, the tool result — because the moment two of them share a field, one silently becomes the other.

That is not hypothetical. The commit path once derived the default as llm_recommendation.unwrap_or(rules_decision). With the engine at shortlist and the model at research, a person choosing shortlist — the engine's own answer — was recorded as overriding the system, and a person choosing research was recorded as agreeing with it. The model held the default in the conservative direction while being refused it in the optimistic one (ARCHITECTURE.md).

What each party may do

MayMay not
Modelexplain; recommend the engine's decision or a more conservative one; ask the engine whether a hypothetical recommendation would be permitted (evaluate_etf with llm_recommendation, read-only)define the authoritative result; recommend above it; turn a recommendation into a state change; present its own recollection of the rules as the rules
Humanconfirm; override in either direction with a rationale; initiate an override the model never proposed; supply the research note a shortlist requiresbypass a non-bypassable constraint, by any rationale
Backend—trust anybody's claim about the engine's decision, including one a human approved

and the backend must, at the point of mutation: lock the row, recompute the evaluation, verify the token against the recomputed decision, reconcile the choice, re-check hard constraints, and write the mutation, the nonce and the audit record in one transaction — or refuse and roll all of it back.

Where each rule is enforced

RuleFirst enforcedAuthoritatively enforced
Advisory ceiling on the modelcommit_evaluation request check in approval.py — before a human is askedrules::reconcile_decision in rules.rs, after the row lock
The default is the engine's decisionthe approval prompt labels the model-reported engine decision Confirm and every other Overridereconcile_decision: override_applied = requested != rules_decision
An override declares itselfthe token's override_requested, derived from the choicereconcile_decision refuses a flag that disagrees, in either direction
Overrides carry a rationalethe prompt requires onecommit_evaluation / shortlist_etf in server.rs
Hard constraints hold— (all options are offered on purpose)rules::blocking_hard_constraint, called in every mutation body
The displayed premise was truethe token binds it as expected_choice — signed, not verifiedApprovalVerifier::verify against the decision recomputed under the lock

The left column is convenience; the right column is the boundary. The agent-side checks make a bad request fail early and readably, but nothing in the system depends on them.

Lab

1. The trust matrix

reconcile_decision(rules_decision, llm_recommendation, requested_decision, override_requested) is pure. Predict the outcome of each row before reading the answer column, using only the rules above:

RulesModelHumanOutcome in this repositoryAsserted by
shortlistshortlistshortlistcommitted; override_applied = false; a research note is required for a shortlistverify-approvals
shortlistresearchshortlistcommitted; not an override — confirming the engine never is, whatever the model saidcase A, choosing_the_deterministic_decision_over_a_conservative_model_is_not_an_override
shortlistresearchresearchcommitted as a human override, with a rationale, though the model suggested itcase B, following_a_conservative_model_away_from_the_engine_is_a_human_override
researchshortlistshortlistrefused — the model's promotion is refused before the human is asked, and again by reconcile_decision whatever the human chosecase C, a_more_optimistic_model_recommendation_is_refused_before_anything_else
researchnone / researchshortlistcommitted as a human promotion, with rationale and a research notecase D, a_human_may_move_above_the_deterministic_decision; make verify-hitl end to end
reject (non-UCITS)noneshortlistrepresentable as an override, then refused by HC-UCITS; so is researchcase E, a_human_override_cannot_reach_past_a_non_bypassable_constraint
shortlistresearchshortlist, token claims overriderefused — the flag disagrees with the decisioncase A′

Note row four against row five. The human's outcome is reachable either way; what is refused is the model owning it. A promotion exists only as a human act with the person's name and reason on it — which is why make verify-hitl-audit checks that llm_recommendation is not shortlist on the row it inspects.

Run the deterministic half:

make verify-approvals-rust

Then add a row of your own as a scratch test next to the cases in rules.rs: engine shortlist, model reject, human research. Predict override_applied, whether a rationale is required, and what human_override_decision holds. Delete the test afterwards.

2. Break it: the model lies about the engine

This was observed with qwen3:8b and is recorded in ARCHITECTURE.md and DEMO.md. Asked to shortlist the non-UCITS VTI-ARCA, the model first proposed llm_recommendation: shortlist against an engine that returns reject, and was refused by the ceiling. It then retried with rules_decision: "shortlist" — asserting a different engine decision — and that request reached a human, who approved it.

Trace it through the code and answer each question before reading the answer.

What did the human see? Both decision-carrying approval functions in approval.py (etf_commit_evaluation, etf_shortlist_etf) build the prompt from the model's request. It read "Deterministic engine (authoritative): shortlist". On the commit path the card then offers "Confirm — shortlist" as the default and labels reject an Override; on the shortlist path a shortlist that matches the claimed engine decision is a plain confirmation. Either way, the premise was false and it was displayed under the label "authoritative".

What was signed? choice = shortlist, expected_choice = shortlist (the model's claim), override_requested = false (because the choice equalled the claimed engine decision), the gateway-asserted actor_id, the request_id, and the payload. A genuine HMAC over a false premise. Note also what the lie avoided: on the true premise this is an upward override and the human would have been made to type a rationale.

What did the backend re-derive? The mutation tool (commit_evaluation and shortlist_etf behave identically here) locked VTI-ARCA, recomputed the evaluation — reject — and called verify with that as the expected choice:

Human approval rejected: approval token was issued against a different
deterministic decision

Nothing after that point ran. The transaction rolled back, the nonce was not consumed, and VTI-ARCA stayed UNREVIEWED with no mutating history event.

Could the wrong claim have become the authoritative mutation? No, and not because of one check. Had the binding somehow passed, reconcile_decision would have found shortlist ≠ reject with override_requested = false and refused; had that passed, HC-UCITS would have refused any decision above reject. Three independent gates, and the one that fired first depended only on what the model got wrong.

3. What the boundary did not protect

The mutation was safe. The consent was not well-informed: a person approved a request whose premise was false, labelled authoritative by the system's own UI. The token binding turns that into a refusal rather than a wrong state — that is its job — but the human still made a decision on a falsehood.

Question. The approval function already runs below the model, holds the authenticated identity and can reach the MCP server. What would it take for the prompt to display the engine's recomputed decision rather than the model's claim, keeping the model's value only as a cross-check? What does the token's expected_choice then mean, and which refusal becomes impossible? (This is a documented limitation of the current design — LIMITATIONS.md — and an open problem in CHALLENGES.md, not something this repository changes.)

4. Read the point of mutation

Open commit_evaluation in server.rs and list, in order, everything that happens between pool.begin() and tx.commit(). Mark which steps read caller-supplied data and which read state recomputed under the lock. The only caller-supplied inputs are etf_id, the token and request_id; every decision parameter comes from the signed claims, and every claim about state is checked against the recomputation.

What to take away

  • Name the engine's result, the model's opinion, the human's choice and the persisted outcome as four values with four owners, and never let two share a field.
  • Advisory means advisory in both directions. A model that cannot promote must not be able to demote either; permitted is not adopted.
  • A signature proves that a human agreed to a statement. It does not make the statement true. Bind the premise into the signature so that a false one is detectable, and recompute the truth at the point of mutation.
  • The point-of-mutation backend is the only component that knows the decision. The UI, the token and the model all carry claims about it.
  • Consent is only as good as the premise it was shown. Mutation integrity and informed consent are separate properties: this repository guarantees the first and, today, not the second.

Go deeper

This chapter is maintained in cognokratos/etf-research-agent beside the code it teaches. The book shows docs/applied/05-recommendation-authority-and-consent.md at revision 493a67a721ef56ee64151e66e6e47c974552a23f (branch main). View source at this revision · Report a correction.

Corrections are made upstream against the current main branch and appear here when the book's pin for this source is updated.