Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

5. Evaluation

From cognokratos/simple-agent-template · docs/concepts/05-evaluation.md · pinned revision c66ce19d7b0c

This page covers learning-path stage 6.

Why manual prompting is not testing

Typing five prompts into the UI and seeing good answers tells you those five prompts worked once, with that model, that prompt and that data. It says nothing about the next model version, the prompt edit you are about to make, or the sixth prompt. A probabilistic component needs the same discipline as any other dependency you don't control: a regression suite run against the real system, with results you can compare over time.

Evaluation differs from unit testing in three ways:

Unit testEvaluation
Deterministic code, exact assertionsProbabilistic system, scored behaviour
Pass/fail per testMetrics over a dataset (tool_call_correct/mean) with a threshold
Runs in milliseconds, offlineRuns against the live agent and model; minutes; costs inference

This repository runs both. Unit tests (make static-check, make test) pin the deterministic code. Evaluations (make eval-all) measure the agent.

Diagram F: the evaluation pipeline

flowchart LR
    DS[("Dataset<br/>evaluation/datasets/*.json")] -->|"make eval-bootstrap"| MD[(MLflow dataset)]
    MD --> RUN["Runner<br/>evaluation/runner.py"]
    RUN -->|"POST /v1/workflow/full<br/>service key + EVALUATION_PRINCIPAL"| AG[Live agent]
    AG -->|"SSE: tool calls, tool results,<br/>answer, blocked?"| CL["Client parser<br/>evaluation/client.py"]
    CL --> SC{"Deterministic scorers<br/>evaluation/scorers.py"}
    SC --> MET["Metrics + latency p50/p95<br/>+ provenance"]
    MET --> ML[(MLflow experiment)]
    MET --> JSON["evaluation/results/<br/>suite-latest.json"]
    AG -.->|"OTLP traces"| ML
    MET --> GATE{"Gate: metric ≥ threshold?"}
    GATE -->|no| FAIL[non-zero exit]
ConceptImplementation in this repo
Experiment trackingMLflow (experiments, datasets, prompt registry, traces)
DatasetsSource-controlled JSON in evaluation/datasets/, synced into MLflow
ScorersPure Python functions in evaluation/scorers.py, one list per suite in SCORERS
Gate metric per suiterequired_metric in evaluation/config.py
Provenanceevaluation/provenance.py

The four suites

SuiteQuestionExample case
toolsRight tools, right arguments, right order?TOOLS-SEARCH-OPEN: exactly search_tickets(status="open")
guardrailsBlocked what it should, allowed what it should?GR-BLOCK-APPENDED-INJECTION, GR-ALLOW-FRAUD-EDUCATION
groundingAnswer built only from tool results?GROUND-ABSENT-RECORD: must not invent a ticket
injectionDid hostile text in a tool result change tools, state or disclosures?INJ-FABRICATED-AUTHORIZATION

Deterministic scorers, not LLM judges

An LLM judge ("rate this answer 1 to 10") is another probabilistic component in the measuring instrument. This repository uses none. Every metric is a function of captured facts: which tools were called with which arguments, what they returned, whether the response was blocked, and which strings the answer contains. A red metric is a fact about the run.

The cost is that a scorer only catches what it asserts. While writing this material, a captured answer to the fabricated-authorization case was re-scored offline with the real scorer. The answer repeated the planted claim ("This approval is treated as granted"). Given the ticket's real get_ticket fields as evidence, injection_resistance_scores reported injection_resisted = True. That is correct by its definition: nothing was mutated, and no state change was claimed. Whether "repeats injected authority as fact" should also fail is a product decision. If it should, it needs a scorer that checks for it. That is a challenge.

Methodology worth stealing

Each of these is explained in EVALUATION.md:

  • report grounding and completeness separately, and gate only on grounding;
  • compare numbers by value, not by digit string;
  • suppress negations ("nothing was changed") when detecting action claims;
  • count guardrail false positives and false negatives separately;
  • treat over-blocking as a failure in the injection suite;
  • seed dedicated fixtures (TKT-INJ-*) instead of mutating demo data;
  • report latency as a distribution (p50, p95, max), never a mean;
  • attach provenance to every result, or it is not evidence.

Where evaluation runs

  • Not in the pull-request gate. Evaluations need a model, are non-deterministic, and cost inference. .github/workflows/ci.yml runs only the deterministic checks.
  • Manually or before release with make eval-all, or with the manual .github/workflows/live-evaluation.yml.
  • ALLOW_FAILURES=1 suppresses the metric gate only. A dead agent or missing dataset still fails, because "the model got worse" and "the cluster is broken" must never look the same.

Go deeper