Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Part II — How does an agent become durable software?

Reference implementation: cognokratos/sophos-agent (Σοφός), branch main, pinned in Source revisions.

Stack: TypeScript, SvelteKit (adapter-node), LangGraph.js, SQLite, MCP (Memory and Fetch servers), Ollama.

Part I builds an agent. Part II asks what happens to it over time: across many runs, a dropped connection, a model server that goes away, a kill -9 in the middle of a tool call. Its claim is that the runtime is part of the agent's architecture. What the agent remembers, whether a task finishes and where its data goes are all decided by the runtime, not by the model.

What the system is, precisely

One Node process (SvelteKit) contains the UI, the HTTP API and an explicit single-agent LangGraph StateGraph with two nodes, agent and tools. The model runs in Ollama. Tools come from two MCP servers, Memory and Fetch, started as child processes. Conversations, runs and LangGraph checkpoints live in one SQLite file, written with synchronous durability. The Memory server keeps its own knowledge graph in memory.jsonl. Live output reaches the browser over Server-Sent Events, with Last-Event-ID reconnection served from an in-memory buffer.

The system is small on purpose: every runtime concern in this part can be pointed at in a few hundred lines of code and observed on your machine.

What it is not

  • Not a multi-agent system. There is one graph with two nodes and no orchestration of multiple agents.
  • No exactly-once side effects. Checkpointing makes the workflow resumable. All tool calls of one assistant turn run inside one graph step, so a crash inside that step re-runs every call of the turn on resume. Lesson R7 is devoted to this.
  • No approvals, guardrails, OpenTelemetry tracing, evaluation framework or run cancellation. These are on the roadmap. Where the lessons discuss them, they are labelled as challenges or future design.

Four words that must not be conflated

TermOwnerLifetime
Conversationthe application (conversations table)until the user deletes it
ThreadLangGraph (thread_id)the checkpoint history of one conversation
Runthe application (runs table: running, completed, failed, interrupted)one user message's execution
CheckpointLangGraphone completed graph step

Lesson R2 is about why the application needed runs even though the framework already had threads and checkpoints.

How this part is organised

  1. The runtime learning path, including the lab environment that every lesson uses.
  2. Follow one run through HTTP, LangGraph, SQLite, MCP and SSE.
  3. Lessons R1–R8, each with a lab that predicts and then observes runtime behaviour.
  4. Case studies and Challenges: design cancellation, durable approval, replay-safe tools or a worker split.
  5. Reference: the architecture chapters the lessons cite.

Running the labs

Node 24 with Corepack pnpm, Ollama with qwen3 pulled, npx and uv for the MCP servers, and curl, jq, sqlite3, lsof, pgrep/pkill. Docker Compose is needed only for part of R8. The labs run a production build against disposable state under data/lab/ and never touch your real conversations. The Fetch labs need internet access. See Setting up each track.

Prerequisite. Part I stages 0–3. The lessons link to the specific Part I chapters they build on.