Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

3a) Durable Execution & Persistence

From cognokratos/sophos-agent · docs/architecture/3a-durable-execution-persistence.md · pinned revision 8d9fe52182d8

Learn: 02 — Model execution state, 03 — Design durability boundaries, 06 — Failure, restart and resume, run lifecycle walkthrough.

The agent's state lives in one SQLite file and survives restarts and crashes. LangGraph's checkpointer stores the execution state; a few application tables store metadata. Markdown is an export, not a database.

SvelteKit  (/api/chat, /api/conversations)
    │
    ▼
LangGraph  StateGraph: START → agent ⇄ tools → END
    │            │
    │            └── MCP tools (Memory, Fetch), resolved when the tools node runs
    ▼
SqliteSaver  (official @langchain/langgraph-checkpoint-sqlite)
    │
    ▼
data/db/sophos.db

Terminology

TermWhat it isWhere it lives
ConversationThe chat the user sees in the sidebar: title, timestamps, message count.conversations table
ThreadLangGraph's persistent execution identity. Its state is the message list. thread_id = conversation id.checkpoints / writes tables (owned by SqliteSaver)
RunOne execution of the graph on a thread: from a new user message (or a resume) until the graph stops.runs table (status, error, timestamps)
CheckpointA snapshot of the thread's state, saved after every step (super-step) of a run.checkpoints table
Conversation  (conversations.id)
     │
     └── thread_id  (same value)
            │
            ├── Run 1  (runs row: completed)
            │     └── checkpoints: input → agent → tools → agent
            │
            ├── Run 2  (runs row: failed / interrupted)
            │     └── checkpoints: input → agent → tools ✗   ← resume continues here
            │
            └── Run N

LangGraph itself has no "run" record in its checkpoints (their metadata is source, step, parents); the runs table is how this application gives runs an identity and an outcome.

Storage layout

One file, DATABASE_PATH (default data/db/sophos.db, plus SQLite's -wal/-shm files), one better-sqlite3 connection:

TableOwnerContents
conversationsapp (persistence/conversations.ts)id, title, created_at, updated_at, message_count
runsappid, conversation_id, status, error_code, error_message, started_at, finished_at
checkpointsSqliteSaverserialized state per checkpoint (messages, including tool calls/results)
writesSqliteSaverpending writes of steps that have not completed

Why one file: the checkpointer needs a better-sqlite3 connection anyway, so the application tables reuse it. One file is one thing to back up, inspect (sqlite3 data/db/sophos.db .tables) or delete. Table names do not collide, and the application never reads the checkpoint tables directly; it goes through LangGraph (getState, getStateHistory).

Messages are stored once, in the thread. The conversations table holds only metadata; message_count is denormalized after each run so the sidebar does not deserialize every thread.

Schema setup is automatic. openDatabase() (persistence/database.ts) applies numbered migrations and records progress in PRAGMA user_version; SqliteSaver creates its tables on first use. Nothing needs to be run by hand.

A run, step by step

POST /api/chat {message, conversation?}
  │  validate id · create conversation (if new) · 409 if a run is active
  ▼
runs row: running ─────────────────────────────── (SQLite)
  │
  ▼
graph.stream({messages: [HumanMessage]}, {thread_id, durability: 'sync', recursionLimit})
  │   LangGraph loads the thread's last checkpoint, appends the message,
  │   and saves a checkpoint after every step
  ├── agent: system prompt + thread messages → Ollama → AIMessage
  ├── tools: ToolNode → MCP → ToolMessage(s)
  └── … until the model answers without tool calls
  │
  │   tokens / trace events ──► SSE subscribers (GET /api/chat)
  ▼
runs row: completed | failed, conversations.message_count updated
  │
  ▼
SSE `end` (or `chat-error`) → the UI reloads GET /api/conversations/:id
  • The caller sends only the new message. History comes from the checkpoint; nothing is rebuilt from files.
  • The system prompt is not stored in the thread. The agent node prepends config/system.md on every model call, so editing it affects existing conversations.
  • durability: 'sync' writes each checkpoint before the next step starts, so a crash loses at most the step in progress.

Restart, failure and resume

Note

Book edition note. "A tool that already ran is not called again" is true for completed graph steps. All tool calls of one assistant turn run inside a single tools step, so if the process dies part-way through that step, a resume re-runs every tool call of the turn, including ones that already took effect. Lesson R7 · Side effects and idempotency covers this failure window; checkpointing does not make side effects exactly-once.

SituationWhat is persistedruns.statusWhat happens next
Run finishedfinal checkpoint, next = []completednext message starts a new run
Error in a node (e.g. Ollama down)checkpoints up to the failing step, next = [node]failedPOST {resume: true} retries from that step, or send a new message
Process crash / restart mid-runevery checkpoint completed before the crashrunning → interrupted on the new process's first database accessif next is non-empty, POST {resume: true} continues
Step limit reachedcheckpoints so farfailed (RECURSION_LIMIT)resume (with a fresh limit) or send a new message

A resume calls graph.stream(null, {thread_id}): null input means "continue from the last checkpoint". Completed steps are not repeated (a tool that already ran is not called again). GET /api/conversations/:id reports resumable: true when the thread has pending steps and no run is active; the UI then shows a Resume button.

Resumable is a property of the thread, not of runs.status: a conversation is resumable exactly when its latest checkpoint has a non-empty next and no run is active. interrupted only says the process that owned the run died before recording an outcome. Two crash windows produce an interrupted run that has no work of its own to resume:

  • after LangGraph wrote the final checkpoint but before finishRun() recorded completed: the answer is in the thread and next = [];
  • after startRun() inserted the run row but before LangGraph wrote the input checkpoint: the new message itself was never recorded (send it again), and next is whatever the thread had before this run.

Graceful shutdown. Sophos has no explicit run-drain or cancellation policy. On SIGTERM/SIGINT, adapter-node waits for open HTTP connections (including SSE streams) up to SHUTDOWN_TIMEOUT, then the shutdown hook closes MCP and SQLite. An in-flight run's outcome depends on timing:

TimingTypical result
The run finishes before the hook runs (e.g. an open SSE stream held the server open)completed
The hook closes MCP / SQLite while the run is still executingthe run can fail against the closed dependency and be recorded failed (one observed case: The database connection is not open)
The process exits before the run's outcome is recorded (e.g. SIGKILL after a container grace period)the row stays running and is recovered as interrupted by the next process

In every case, whether the thread can be resumed is decided by next, as above.

These are the same states a future human-in-the-loop interrupt needs (interrupted → resume with a decision), so approval can be added on top of this model without changing persistence.

Execution limits

AGENT_RECURSION_LIMIT (default 25) is passed as LangGraph's recursionLimit: the maximum number of steps per run. Each node execution is one step, so every agent → tools round trip costs two; 25 allows about 12 model calls. In this graph tool iterations are always one fewer than model iterations, so one limit covers both. Exceeding it raises GraphRecursionError, recorded as RECURSION_LIMIT.

What is durable, what is in memory

Durable (SQLite)In memory only (lost on restart, by design)
conversations and their metadataactiveRuns: the run currently executing per conversation
every run record; its outcome once recordedits buffered SSE events (for Last-Event-ID replay) and subscribers
all messages, tool calls and tool resultsthe compiled graph and MCP connections (rebuilt on demand)
checkpoints of every step

Run outcomes (completed, failed) are durable once the process records them; a run whose process died first stays running until the next process recovers it as interrupted. After a restart a conversation is fully usable: its messages load from the thread, and it is resumable if its thread has a pending step.

Markdown and JSONL: exports only

GET /api/conversations/:id/export generates a Markdown transcript (messages, tool calls, tool results) from the thread; ?format=jsonl lists the thread's checkpoints, one JSON line each, oldest first. Both are generated on request and never read back: deleting an export cannot lose conversation state.

Legacy data/chat/

Earlier versions stored conversations as data/chat/{id}/0001.user.md … trace.jsonl and rebuilt the history from those files. That directory is no longer read or written. Existing files are left in place (nothing deletes them) but do not appear in the UI; there is no importer. Delete the directory when you no longer need it.

Code map

FileResponsibility
src/lib/agent/graph.tsthe StateGraph (agent/tools nodes, routing), threadConfig()
src/lib/agent/index.tsruntime: builds the graph once, runAgent(), getThreadState()
src/lib/agent/runs.tsstarts runs, records outcomes, active-run SSE registry
src/lib/agent/persistence/database.tsopens SQLite, migrations, createCheckpointer()
src/lib/agent/persistence/conversations.tsconversations and runs tables
src/lib/agent/messages.tsLangChain messages → MessageDto (the only UI-facing shape)
src/lib/agent/export.tsMarkdown / checkpoint JSONL exports

This chapter is maintained in cognokratos/sophos-agent beside the code it teaches. The book shows docs/architecture/3a-durable-execution-persistence.md at revision 8d9fe52182d8441454916ec8a6ab13c0773228e2 (branch main). View source at this revision · Report a correction.

Corrections are made upstream against the current main branch and appear here when the book's pin for this source is updated.