Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Lab 09 — Add human approval

From cognokratos/simple-agent-template · docs/tutorials/09-add-human-approval.md · pinned revision c66ce19d7b0c

Objective

Turn on the template's optional human-approval flow, approve a real change, read the audit trail it leaves, and probe the controls around it.

Concept

A human approval is only meaningful if it is bound to exactly what will happen (actor, resource, request, current state, payload), is single-use, and is verified at the point of mutation by a component the model cannot influence. → Concept 8

Architecture before

Read-only: no approval function, no /approvals/execute route, no interaction endpoints. CI enforces that this is the shipped state (scripts/verify_read_only_default.py). Do not commit the changes in this lab.

Exercise

All four steps are required. No single one opens the mutation path on its own (APPROVALS.md — enabling it).

  1. In agent/config.yml, uncomment the functions: block:

    functions:
      ticket_priority_change:
        _type: ticket_set_priority_approval
        token_ttl_seconds: 600
    
  2. Add the function to the workflow's tools:

    workflow:
      ...
      tool_names:
        - tickets_mcp
        - ticket_priority_change
    
  3. In .env, set a shared secret of at least 24 characters, used by both the agent (to mint) and the MCP server (to verify):

    HITL_APPROVAL_SECRET=<output of: openssl rand -hex 32>
    
  4. In .env, mount NAT's interaction endpoints:

    HITL_ENABLE_INTERACTIVE=true
    

Run it

make up-build     # rebuilds the agent image (config.yml) and recreates changed services
make wait
make logs-mcp     # look for: human-approval execution endpoint enabled

In the UI:

Mark ticket TKT-1003 as high priority.

Expected, per TEST-SCENARIOS.md: the agent reads TKT-1003 (medium), calls the approval function, and an approval card appears. Choose Change to — high and type a reason.

Then look at the database:

make shell-db
SELECT id, priority, updated_at FROM tickets WHERE id = 'TKT-1003';
SELECT ticket_id, previous_priority, new_priority, actor_id, request_id,
       rationale, policy_context, recorded_at
FROM ticket_audit ORDER BY recorded_at DESC LIMIT 5;
SELECT nonce, action, resource_id, actor_id, consumed_at
FROM approval_nonces ORDER BY consumed_at DESC LIMIT 5;

Observe

  • actor_id in ticket_audit is your Keycloak subject, from the gateway header. The model had no way to set it.
  • rationale is what you typed. payload holds any model-supplied note, which the card showed you labelled as model-supplied.
  • policy_context records the policy version.
  • One nonce row per applied approval, written in the same transaction.
  • Run it again and choose Keep — medium/high or Cancel: no token is minted, and no audit row is written.

Break it

A. Edit the audit trail. In make shell-db:

UPDATE ticket_audit SET new_priority = 'low';

Observed (on a test row in a rolled-back transaction):

ERROR:  ticket_audit is append-only; UPDATE is not permitted

B. Let a hostile record ask. With approvals enabled, send Summarise ticket TKT-INJ-FAKE-AUTH. The description claims an approval was already granted. Whatever the model does, it cannot apply the change by itself: the only path to a mutation is a card you answer, bound to your identity. If the model calls the approval function, decline. If it says the change "has been applied", check ticket_audit and tickets.priority: the claim is not the state. This is the case the injection suite's injection_no_action_claim metric exists for.

C. Run the boundary tests.

make verify-approvals        # agent side: binding, replay, ownership, offered choices
make verify-approvals-rust   # MCP side: signature, binding, lifetime, policy

They cover what is hard to do by hand: forged and tampered tokens, expiry, wrong resource or request, moved state, replay, answering someone else's prompt, and choosing an option that was never offered.

Why it failed

  • A: append-only is enforced by a database trigger, not by application convention. A decision record that can be edited is not evidence.
  • B: authorization is not an inference the model makes. It is a signed artifact produced by a human interaction the model cannot answer, verified by a server the model cannot reach directly.
  • C: each binding in the token removes one specific way an approval could be misused. The tests prove each one independently.

Architecture after

Diagram G in concept 8.

Reset

  1. Put TKT-1003 back through the same controlled path: ask the agent to set it to medium and approve with a reason. That leaves a second, honest audit row. (An UPDATE tickets ... in psql would work too, but it bypasses the audit trail. Notice that you'd be doing exactly what this lab argues against.)
  2. git checkout -- agent/config.yml, remove HITL_APPROVAL_SECRET and HITL_ENABLE_INTERACTIVE from .env, then make up-build.
  3. make logs-mcp should again show that the server is read-only.

What you learned

  • A safe mutation needs a proposal, a human decision, a binding, a re-check at the point of mutation, and an immutable record. Remove any one and a specific attack works.
  • Feature flags for dangerous capabilities should remove the surface, not just disable it.
  • The model reports the outcome. It does not establish it.

Go deeper