Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

3. Grounding in authoritative systems

From cognokratos/simple-agent-template · docs/concepts/03-grounding-and-authoritative-state.md · pinned revision c66ce19d7b0c

This page covers learning-path stage 4.

The problem: the model's memory is not your database

A model asked "what is the priority of TKT-1004?" without tools can only do one of two things: say it doesn't know, or produce a plausible guess. A guess is indistinguishable from a fact in fluent prose. In a support-ticket domain, a plausible guess about refund status is worse than no answer.

Grounding means every factual claim about domain state is derived from data the agent retrieved from the system of record during this request, and the answer can be checked against that data.

Model memoryAuthoritative system
SourceTraining data, the promptPostgreSQL, through MCP tools
FreshnessFrozen at training timeCurrent at the time of the call
Knows your ticketsNoYes
Can be checkedNoYes. The tool result is in the trace.
Who decides it is trueNobodyThe database (constraints, transactions)

How this repository grounds answers

  1. The prompt demands it. The system_prompt in agent/config.yml begins: "You must use the available tools for every factual statement about tickets or their history," and asks the model to "distinguish recorded facts (what a tool returned) from your own recommendations."
  2. The tools are the only path to the data. The agent holds no database credentials. Facts can enter the conversation only through search_tickets and get_ticket (concept 2).
  3. The database enforces its own invariants. CHECK (status IN ('open', 'resolved')) and CHECK (priority IN (...)) in db/init.sql mean the state the model reads is always well-formed, whatever any caller attempted.
  4. Grounding is measured, not assumed. The grounding evaluation suite checks the answer against what the tools actually returned (concept 5).

Step 1 is a request. Steps 2 to 4 are what make it hold.

Measuring grounding deterministically

grounding_scores in evaluation/scorers.py compares the answer with the captured tool results:

  • Numbers are compared by value, so 980.0 in a tool result grounds $980.00 in the answer.
  • Forbidden assertions come from the dataset: an answer about TKT-1001 must not mention TKT-9999.
  • Grounding and completeness are scored separately, and only grounding is gated. Inventing a value is an integrity failure. Omitting one is a thoroughness failure. Averaging them hides which one moved. See EVALUATION.md.

The GROUND-ABSENT-RECORD case ("Tell me about ticket TKT-DOES-NOT-EXIST") checks the most important grounding behaviour: the agent must ask the system of record before saying anything about a ticket.

On the default qwen3:8b, in the evaluation run and in two direct runs, the agent answered "The ticket ID "TKT-DOES-NOT-EXIST" does not exist in the system" without calling any tool. It inferred that from the identifier's name. The claim happens to be true, and it is still a hallucinated statement about authoritative state: nothing was checked, and the same reasoning would produce a confident "does not exist" for a real ticket with an unusual id. The deterministic scorer caught it (grounding_tool_used = False), and the gate went red. Contrast TKT-9999, where the model did call get_ticket and reported the tool's "not found" error.

Current state versus history

The schema separates what is true now from what happened:

  • tickets.priority is the current state;
  • ticket_events is the history the agent reads;
  • ticket_audit (used by the optional approval feature) holds the committed decisions that produced state changes.

Reading one is never a substitute for reading the other. An agent that infers "the priority was raised yesterday" by comparing current state with a remembered earlier answer is reconstructing history from memory. That is the failure grounding exists to prevent. See EXTENDING.md.

Grounded is not the same as correct

Grounding guarantees the inputs to the model's reasoning are real. It does not guarantee the reasoning.

Observed while writing this material, on the default qwen3:8b at temperature 0, for "Which ticket should we handle first, and why?":

  • On one agent build, 3 of 3 runs made one correct search_tickets(status="open") call, whose result included TKT-1002 with priority urgent, and then answered:

    the ticket with the highest priority is TKT-1004 (Duplicate charge) with a priority of "high"

    Every fact in that sentence is grounded. The conclusion is wrong, because urgent outranks high. The answer passes a "did it use the tools" check and fails an "is it right" check.

  • On a rebuild from commit af29ce0, with the same configuration digest and the same model, 4 of 4 runs fanned out to get_ticket for three tickets and then ended without an answer. That is the failure CONFIGURATION.md records, along with a model (qwen3.5:9b) that answered correctly.

Both are wrong, in different ways, and nothing in the configuration changed between them. TEST-SCENARIOS.md names TKT-1002 as the expected pick. The lesson about grounding is the first bullet. The lesson about probabilistic components is the pair: re-evaluate after every rebuild and dependency upgrade, not only after prompt or model changes.

Engineering responses, roughly in order of reliability:

  1. Move deterministic logic out of the model. Use the model for decisions that benefit from interpretation, and ordinary code for decisions that can be specified deterministically. If "highest priority, then oldest" is the business rule, compute it in SQL (ORDER BY a priority rank) and return it from a tool. The model then explains the ranking instead of performing it. This is the same move as the fan-out fix in concept 2.
  2. Evaluate the decision, not just the grounding. There is no evaluation case for prioritization today. Adding one is a challenge.
  3. Use a stronger model for multi-step reasoning, and measure it with the same suites before switching.

Provenance

Grounding answers where did this fact come from? for a single answer. Provenance asks the same question of the system: which agent, prompt and model produced this result? evaluation/provenance.py records the running agent's own /version (prompt digests and model names), the prompt registry version, and the harness commit with every evaluation result. It flags when they disagree. A score you cannot attribute to a specific configuration is not evidence. See EVALUATION.md — provenance.

Go deeper

This chapter is maintained in cognokratos/simple-agent-template beside the code it teaches. The book shows docs/concepts/03-grounding-and-authoritative-state.md at revision c66ce19d7b0c5c88c41b6860c78f66075485a07a (branch main). View source at this revision · Report a correction.

Corrections are made upstream against the current main branch and appear here when the book's pin for this source is updated.