Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

1. The LLM, the agent, and the agent loop

From cognokratos/simple-agent-template · docs/concepts/01-agents-and-agent-loops.md · pinned revision c66ce19d7b0c

Agentic AI is software engineering around a probabilistic decision-making component.

This page covers learning-path stages 0 and 1. It explains what the model is, what an agent adds around it, and how the loop in this repository runs.

The LLM is a probabilistic component

Treat the model like an unreliable remote service whose output is a sample, not a return value:

PropertyWhat it means for engineering
Non-deterministicThe same prompt can produce different answers. temperature: 0.0 in agent/config.yml narrows the spread but does not make output reproducible across model versions, hardware or providers.
Untyped outputIt emits text (or a structured tool-call request). Anything downstream must parse and validate it.
StatelessIt remembers nothing between calls. Every call includes the whole conversation (bounded here by max_history: 20).
No access to your systemsIt knows only its training data and what you put in the prompt. It cannot see your database.
Fluent when wrongA wrong answer is as confident and well-formatted as a right one.

Two consequences shape everything else in this repository:

  1. You cannot unit-test a model into correctness. You measure it with evaluations (concept 5) and observe it with traces (concept 6).
  2. You cannot let it hold authority. Identity, authorization and state changes are enforced by deterministic code that does not depend on what the model said (concept 7).

The split this repository is built around:

Probabilistic: the model decidesDeterministic: code enforces
interpreting the requestauthentication (Keycloak, gateway sessions)
planning and reasoningauthorization and the allowed capability surface
which tool to call, with which argumentstool input schemas and parameterized SQL
natural-language generationoutput regex blocking, PII masking
recommendationspolicy evaluation, approval-token verification
database constraints, append-only audit records
evaluation assertions, network boundaries

What an agent is

An agent is a program that uses an LLM to decide its next action in a loop, executes that action through code it controls, and feeds the result back to the model until the model produces a final answer.

The LLM never executes anything. It produces a request: "call get_ticket with {"ticket_id": "TKT-1001"}". The agent runtime decides whether to honour that request, executes it, and returns the result as more input.

Diagram A: the agent mental model

flowchart LR
    U([User]) -->|question| A[Agent runtime]
    A -->|conversation + tool schemas| L{{LLM}}
    L -->|"tool-call request<br/>(name + JSON args)"| A
    A -->|executes| T[Tool]
    T -->|reads| E[(Environment:<br/>database, APIs)]
    E --> T
    T -->|result as text| A
    L -->|final answer| A
    A -->|answer| U

    classDef prob fill:#fde68a,stroke:#b45309,color:#000
    classDef det fill:#bfdbfe,stroke:#1d4ed8,color:#000
    class L prob
    class A,T,E det

Yellow is probabilistic, blue is deterministic. The model sits inside a loop that ordinary code owns.

ConceptImplementation in this repo
Agent runtimeNeMo Agent Toolkit (NAT) streaming_react_agent workflow, register.py
LLMAny OpenAI-compatible endpoint (llms.primary in agent/config.yml); default qwen3:8b on a local Ollama
ToolsTwo read-only MCP tools, search_tickets and get_ticket, in mcp-server/src/main.rs
EnvironmentPostgreSQL, schema in db/init.sql

ReAct and native tool calling

ReAct ("reason + act") is the loop pattern: the model alternates between reasoning about what to do and requesting an action, and each observation (tool result) informs the next step.

Early ReAct implementations ran over plain text. The model wrote Action: get_ticket / Action Input: {...} / Final Answer: ... and a parser extracted them with regexes. Native tool calling moves this into the model API: the tool schemas go into the request as structured definitions, and the model returns a structured tool_calls field instead of prose to be parsed.

This repository uses NAT's ReAct graph with use_native_tool_calling: true. Native calling removes a class of parse failures. It also changed streaming: NAT's stock stream waits for the literal Final Answer: marker, which native calling never emits. That is why register.py exists (see EXTENDING.md).

Diagram B: one tool-calling turn

This is what happens for "Show me the complete details and history for ticket TKT-1001", as observed in a live trace (see the request walkthrough):

sequenceDiagram
    autonumber
    actor User
    participant RT as Agent runtime (NAT ReAct)
    participant LLM
    participant MCP as MCP server (Rust)
    participant DB as PostgreSQL

    User->>RT: "Show me ... ticket TKT-1001"
    RT->>LLM: system prompt + tool schemas + user message
    LLM-->>RT: tool_call get_ticket {"ticket_id":"TKT-1001"}
    Note over RT: The runtime executes it.<br/>The model never does.
    RT->>MCP: tools/call get_ticket (Bearer MCP_API_KEY)
    MCP->>DB: SELECT ... WHERE id = $1 (parameterized)
    DB-->>MCP: ticket row + history rows
    MCP-->>RT: JSON text result
    RT->>LLM: conversation + tool result
    LLM-->>RT: final answer (streamed)
    RT-->>User: answer

The loop is bounded by configuration

An unbounded loop around a probabilistic component is an outage waiting to happen. The bounds live in workflow: in agent/config.yml:

SettingValueWhat it bounds
max_tool_calls20Loop iterations. register.py turns it into a LangGraph recursion_limit of (max_tool_calls + 1) * 2. When it is hit, the user gets "The agent could not produce a final answer within 20 tool calls" instead of a hang.
max_history20Messages carried into each model call
tool_call_max_retries2Retries of a failing tool call
parse_agent_response_max_retries3Retries when the model's output cannot be parsed
pass_tool_call_errors_to_agenttrueA tool error becomes an observation the model can explain, not a crash. Try TKT-9999 (scenario 4 in TEST-SCENARIOS.md).
request_timeout / max_retries (under llms.primary)300.0 / 2Each model call
tool_call_timeout (under function_groups.tickets_mcp)30Each MCP call

The system prompt is configuration, not a control

The system_prompt in agent/config.yml tells the model how to use the tools, to ground every factual statement, and to treat ticket text as untrusted data. It is worth writing carefully, because it measurably changes behaviour. It is not a security boundary. A prompt is a request to a probabilistic component, and the evaluation suites exist because "the prompt says not to" is not evidence that it doesn't. See ANTI-PATTERNS.md.

When not to use an agent

If the sequence of steps is known in advance, write it as code. An agent pays for flexibility with latency, cost and non-determinism. That trade is worth it when the path depends on interpreting natural language ("what's going on with this customer's orders?"). It is not worth it for "every night, export open tickets to CSV", and not for any single decision with an explicit rule, such as ranking tickets by priority (concept 3). See ANTI-PATTERNS.md.

Go deeper