Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Follow one request

From cognokratos/simple-agent-template · docs/tutorials/REQUEST-WALKTHROUGH.md · pinned revision c66ce19d7b0c

"Show me the complete details and history for ticket TKT-1001."

This walkthrough traces that prompt from the browser to PostgreSQL and back, through the actual code, and ends at the trace it leaves in MLflow. Read it with the source open: code → explanation → runtime behaviour → trace.

Every timing and span name below comes from real runs against this repository on the default configuration (qwen3:8b on a local Ollama, guard model the same): a cold run, the first request the guard model served, and a warm repeat on an agent built from commit af29ce0. Your numbers will differ. The shape should not.

The whole path

sequenceDiagram
    autonumber
    participant B as Browser
    participant UI as assistant-ui<br/>(Next.js server)
    participant GW as Gateway (Rust)
    participant MW as NAT middleware
    participant GR as Guardrails
    participant RT as ReAct loop
    participant LLM
    participant MCP as MCP server (Rust)
    participant DB as PostgreSQL
    participant OT as Collector → MLflow

    B->>UI: POST /api/gateway/chat (cookies)
    UI->>GW: POST /api/chat (cookie, x-csrf-token)
    GW->>GW: session, CSRF, schema, stream slot
    GW->>MW: POST /v1/workflow/full<br/>Bearer AGENT_API_KEY + x-authenticated-*
    MW->>MW: service key, identity header, trace context
    MW->>GR: input rail
    GR->>LLM: self_check_input (guard model)
    LLM-->>GR: "No" (do not block)
    GR->>RT: allowed
    RT->>LLM: prompt + tools + question
    LLM-->>RT: tool_call get_ticket(TKT-1001)
    RT->>MCP: tools/call (Bearer MCP_API_KEY)
    MCP->>DB: 2 parameterized SELECTs
    DB-->>MCP: rows
    MCP-->>RT: JSON result
    RT->>LLM: + tool result
    LLM-->>RT: answer tokens
    RT->>GR: output rails (regex, PII mask)
    GR-->>GW: SSE: intermediate_data + data
    GW-->>UI: SSE (streamed through)
    UI-->>B: UI message stream (tool card + Markdown)
    MW-->>OT: spans (OTLP/HTTP), one trace

1. The browser sends the message

The page is a client component built on assistant-ui. ui/app/page.tsx wires the chat runtime to a same-origin route:

() => new AssistantChatTransport({ api: "/api/gateway/chat" }),

The browser holds no tokens: no OIDC token, no API key, nothing the model could leak. It holds an opaque, HttpOnly session cookie and a CSRF cookie, both scoped to Path=/api/gateway (SECURITY.md — cookies).

2. The UI server proxies to the gateway

ui/app/api/gateway/chat/route.ts (POST) runs on the Next.js server. It flattens the UI messages to {role, content} text and makes a server-side fetch to the gateway's /api/chat. It forwards the cookie header and copies the CSRF cookie into x-csrf-token.

The UI is a proxy, not a trust boundary (ARCHITECTURE.md). Everything that matters is checked again by the gateway.

3. The gateway authenticates the session

proxy::chat in gateway/src/proxy.rs:

let (_session_id, session) = authenticated_session(&state, &headers).await?;
verify_csrf(&state.config, &headers, &session)?;

authenticated_session (auth.rs) looks up the opaque session id, rejects expired sessions, and refreshes Keycloak tokens (re-reading roles from userinfo) when the 5-minute access token is near expiry. verify_csrf (session.rs) requires the cookie, the header and the session's own token to agree.

Concept: authentication is established once, by a deterministic component, before any model is involved. → Concept 7

4. The gateway validates, sanitises and mints identity

Still in proxy::chat:

  1. Schema. The body is parsed into ChatProxyRequest with deny_unknown_fields. validate_chat_request accepts only user and assistant roles, requires the last message to be user, and bounds count, per-message characters and total characters.
  2. Concurrency. A per-session stream slot (GATEWAY_MAX_STREAMS_PER_SESSION, default 4).
  3. Re-serialisation. Anything beyond the schema does not survive.
  4. A fresh upstream request with Authorization: Bearer <AGENT_API_KEY>, a new x-request-id, and identity headers built from the session:
let request = identity_headers(request, &session.user, true)?;

identity_headers sets x-authenticated-user-id, -username, -roles and -email. No browser header is forwarded, so neither the page nor the model can choose who the agent thinks the user is.

The URL is AGENT_WORKFLOW_URL (http://agent:8000/v1/workflow/full) with filter_steps=TOOL_START,TOOL_END,FUNCTION_START,FUNCTION_END, so the stream carries tool events the UI can render.

5. The agent authenticates its caller

NAT's FastAPI app is built by AuthenticatedFastApiFrontEndPluginWorker in agent/src/nat_streaming_react/fastapi_worker.py, selected by general.front_end.runner_class in agent/config.yml. Its pure-ASGI middleware runs outermost first:

MiddlewareDoes
StaticServiceKeyMiddlewarehmac.compare_digest on the bearer token; removes Authorization from the request before NAT, session metadata or telemetry can see it
RequireIdentityHeaderMiddlewareExactly one non-empty x-authenticated-user-id, or 401
WorkflowTraceContextMiddlewareCreates the trace id and root span id for this request (trace_context.py)
ResponderIdentityMiddlewareRecords the caller for the approval feature's ownership check

NAT then resolves x-authenticated-user-id into Context.user_id (identity_header in config).

6. Input guardrails run

The workflow is wrapped by the text_guardrails middleware (workflow.middleware: [workflow_guardrails]). TextGuardrailsMiddleware.pre_invoke in text_guardrails.py:

  1. extracts the latest user turn and records it as the trace's readable question;
  2. refuses (does not truncate) anything over GUARDRAILS_INPUT_MAX_CHARS;
  3. runs the NeMo self check input flow, which is a call to the guard model with the self_check_input prompt from agent/config.yml;
  4. runs _CRITICAL_INPUT_PATTERNS over the latest turn and any client-supplied assistant turns, and the anchored _READ_ONLY_TICKET_TEMPLATES;
  5. resolves the decision with _resolve_input_policy.

For this prompt the guard model allowed it, and the specific_ticket_details allow template also matched. The span recorded guardrail.decision_source = llm_and_deterministic_allow. This step took 2.79 s cold and 0.24 s warm, almost all of it the guard-model call.

Concept: a probabilistic check backed by deterministic ones, with the decision source recorded. → Concept 4

7. The agent runtime invokes the LLM

streaming_react_agent_workflow in register.py builds NAT's ReAct graph with the primary LLM, the tickets_mcp tools and the system_prompt (whose {tools} and {tool_names} placeholders NAT fills from MCP discovery). Because use_native_tool_calling: true, the tool schemas go to the model as structured definitions.

The loop is bounded: recursion_limit = (max_tool_calls + 1) * 2.

8. The model chooses a tool

The model returns a structured tool call, not prose:

get_ticket  {"ticket_id": "TKT-1001"}

It chose this from the tool description (from tool_overrides in agent/config.yml) and the system-prompt rule "For details about a specific ticket, call get_ticket using its exact ID." This is the probabilistic decision in the request. Everything after it is deterministic again until the answer is written.

9. The MCP request is issued

NAT's mcp_client function group tickets_mcp sends tools/call over streamable HTTP to http://mcp-server:8080/mcp, with Authorization: Bearer ${MCP_API_KEY} (custom_headers in config) and a 30-second tool_call_timeout. In the stream, this appears as:

intermediate_data: {"type":"FUNCTION_START","name":"tickets_mcp__get_ticket", ... "input":{"ticket_id":"TKT-1001"} ...}

10. The Rust MCP server validates and executes

mcp-server/src/main.rs:

  • require_api_key compares the bearer token in constant time and removes it before RMCP logging;
  • RMCP deserialises arguments into GetTicketArgs { ticket_id: String }, a typed schema;
  • get_ticket trims and rejects an empty id, then runs two parameterized queries: the ticket by id = $1, and its ticket_events by ticket_id = $1 ORDER BY occurred_at DESC.

11. PostgreSQL provides authoritative data

The rows come from db/init.sql: TKT-1001, "Delayed delivery", open, medium, customer Renee Castillo, assigned to Priya Shah, with three history events EVT-1001–EVT-1003. CHECK constraints guarantee status and priority are valid values. PostgreSQL is reachable from the MCP server and nothing else (data_net).

12. The tool result returns to the agent

get_ticket returns one text content block containing JSON:

{"ticket": {"id": "TKT-1001", "subject": "Delayed delivery", "status": "open", "priority": "medium", ...},
 "history_count": 3,
 "history": [{"id": "EVT-1003", "event_type": "support_note", ...}, ...]}

The tool call took 0.03 s cold and under 0.01 s warm. From here the result is part of the conversation, and every string in it is untrusted data: a description or support note could contain instructions. → Lab 04

13. The LLM writes the final answer

The ReAct loop sends the conversation plus the tool result back to the model, which writes the answer. _stream_fn in register.py yields text chunks as they arrive, skipping tool-call chunks. Native tool calling has no Final Answer: marker to wait for, and this is why the module exists.

14. Output guardrails run

TextGuardrailsMiddleware sits between _stream_fn and the client. With PII masking configured (it is, by default), _stream_with_buffered_masking buffers the whole answer, then runs regex check output (credential and prompt-leakage patterns) and mask sensitive data on output (Presidio) over it once, and only then releases it. The released text is what gets recorded as the trace's answer. Observed: 0.04–0.05 s, outcome passed. Names stay visible on purpose: PERSON and ORGANIZATION are not in the entity list.

15. The result streams back

The agent's response is server-sent events. Abridged from the real stream:

intermediate_data: {"type":"WORKFLOW_START","name":"support-tickets-agent.invoke", ...}
intermediate_data: {"type":"FUNCTION_START","name":"guardrail_input_self_check_decision", ...}
intermediate_data: {"type":"FUNCTION_START","name":"tickets_mcp__get_ticket", ...}
intermediate_data: {"type":"FUNCTION_END","name":"tickets_mcp__get_ticket", ...}
data: {"value": "The complete details and history for ticket **TKT-1001** are as follows: ..."}
intermediate_data: {"type":"FUNCTION_START","name":"guardrail_output_regex_presidio_decision", ...}
intermediate_data: {"type":"WORKFLOW_END","name":"support-tickets-agent.invoke", ...}

The gateway streams it through without buffering (x-accel-buffering: no, no request timeout on this route). The UI route parses intermediate_data: into tool cards (tool-input-available / tool-output-available) and data: into answer text, using the wire helpers in ui/lib/nat-wire.ts. The page renders Markdown with react-markdown without raw HTML, so a model answer cannot inject raw HTML.

16. OpenTelemetry records the execution

NAT's spans pass through three processors registered in otlp_exporter.py: WorkflowContentProcessor (readable question and answer), SensitiveHeaderRedactionProcessor (credential headers) and UserIdentityProcessor (drops the user id unless OTEL_TRACE_USER_ID=true). Then they go to the collector over OTLP/HTTP. Guardrails spans go through the process-wide OTel SDK with the same trace id. The collector (observability/otel-collector.yml) batches and forwards to MLflow's /v1/traces, experiment 0 (Default).

17. MLflow lets you inspect it

The trace recorded for the cold run:

support-tickets-agent.invoke                        17.6 s
  <workflow>                                        17.6 s
    guardrail_input_self_check_decision
    tickets_mcp__get_ticket                         0.03 s
    guardrail_output_regex_presidio_decision
  guardrail.input.self_check     outcome=passed     2.79 s
    guardrails.request → rail → action
      self_check_input qwen3:8b                     2.78 s
  guardrail.output.regex_presidio  outcome=passed   0.05 s

Request preview: the question. Response preview: the released answer. The latency budget, from both runs:

ColdWarm (af29ce0)
Total (support-tickets-agent.invoke)17.6 s12.8 s
Input rail (guard-model call)2.79 s0.24 s
get_ticket (MCP + PostgreSQL)0.03 s< 0.01 s
Output rails0.05 s0.04 s
Remainder: the agent model~14.7 s~12.5 s

The remainder is at least two agent-model calls: one to choose the tool, one to write the answer. The agent model's calls did not appear as separately named spans in these traces. Their time shows only inside <workflow>. In these runs, agent latency was model latency.

Open it yourself: make open-mlflow → Experiments → Default → Traces. See lab 06.


What to take away

StepProbabilistic or deterministicEnforced by
1–5 authentication, identity, validationdeterministicKeycloak, gateway, NAT middleware
6 input classificationbothguard model + patterns + templates
7–8 tool choiceprobabilisticthe model
9–12 tool execution, datadeterministicMCP server, SQL, PostgreSQL
13 answerprobabilisticthe model
14 output controlsdeterministicregex, Presidio
15–17 delivery, tracingdeterministicgateway, UI, OTel

The model made two decisions in this request: which tool to call, and what to say. Everything else was ordinary software.

Go deeper