3a) Durable Execution & Persistence
From cognokratos/sophos-agent · docs/architecture/3a-durable-execution-persistence.md · pinned revision 8d9fe52182d8
Learn: 02 — Model execution state, 03 — Design durability boundaries, 06 — Failure, restart and resume, run lifecycle walkthrough.
The agent's state lives in one SQLite file and survives restarts and crashes. LangGraph's checkpointer stores the execution state; a few application tables store metadata. Markdown is an export, not a database.
SvelteKit (/api/chat, /api/conversations)
│
▼
LangGraph StateGraph: START → agent ⇄ tools → END
│ │
│ └── MCP tools (Memory, Fetch), resolved when the tools node runs
▼
SqliteSaver (official @langchain/langgraph-checkpoint-sqlite)
│
▼
data/db/sophos.db
Terminology
| Term | What it is | Where it lives |
|---|---|---|
| Conversation | The chat the user sees in the sidebar: title, timestamps, message count. | conversations table |
| Thread | LangGraph's persistent execution identity. Its state is the message list. thread_id = conversation id. | checkpoints / writes tables (owned by SqliteSaver) |
| Run | One execution of the graph on a thread: from a new user message (or a resume) until the graph stops. | runs table (status, error, timestamps) |
| Checkpoint | A snapshot of the thread's state, saved after every step (super-step) of a run. | checkpoints table |
Conversation (conversations.id)
│
└── thread_id (same value)
│
├── Run 1 (runs row: completed)
│ └── checkpoints: input → agent → tools → agent
│
├── Run 2 (runs row: failed / interrupted)
│ └── checkpoints: input → agent → tools ✗ ← resume continues here
│
└── Run N
LangGraph itself has no "run" record in its checkpoints (their metadata is source, step, parents); the runs table is how this application gives runs an identity and an outcome.
Storage layout
One file, DATABASE_PATH (default data/db/sophos.db, plus SQLite's -wal/-shm files), one better-sqlite3 connection:
| Table | Owner | Contents |
|---|---|---|
conversations | app (persistence/conversations.ts) | id, title, created_at, updated_at, message_count |
runs | app | id, conversation_id, status, error_code, error_message, started_at, finished_at |
checkpoints | SqliteSaver | serialized state per checkpoint (messages, including tool calls/results) |
writes | SqliteSaver | pending writes of steps that have not completed |
Why one file: the checkpointer needs a better-sqlite3 connection anyway, so the application tables reuse it. One file is one thing to back up, inspect (sqlite3 data/db/sophos.db .tables) or delete. Table names do not collide, and the application never reads the checkpoint tables directly; it goes through LangGraph (getState, getStateHistory).
Messages are stored once, in the thread. The conversations table holds only metadata; message_count is denormalized after each run so the sidebar does not deserialize every thread.
Schema setup is automatic. openDatabase() (persistence/database.ts) applies numbered migrations and records progress in PRAGMA user_version; SqliteSaver creates its tables on first use. Nothing needs to be run by hand.
A run, step by step
POST /api/chat {message, conversation?}
│ validate id · create conversation (if new) · 409 if a run is active
▼
runs row: running ─────────────────────────────── (SQLite)
│
▼
graph.stream({messages: [HumanMessage]}, {thread_id, durability: 'sync', recursionLimit})
│ LangGraph loads the thread's last checkpoint, appends the message,
│ and saves a checkpoint after every step
├── agent: system prompt + thread messages → Ollama → AIMessage
├── tools: ToolNode → MCP → ToolMessage(s)
└── … until the model answers without tool calls
│
│ tokens / trace events ──► SSE subscribers (GET /api/chat)
▼
runs row: completed | failed, conversations.message_count updated
│
▼
SSE `end` (or `chat-error`) → the UI reloads GET /api/conversations/:id
- The caller sends only the new message. History comes from the checkpoint; nothing is rebuilt from files.
- The system prompt is not stored in the thread. The agent node prepends
config/system.mdon every model call, so editing it affects existing conversations. durability: 'sync'writes each checkpoint before the next step starts, so a crash loses at most the step in progress.
Restart, failure and resume
Note
Book edition note. "A tool that already ran is not called again" is true for completed graph steps. All tool calls of one assistant turn run inside a single
toolsstep, so if the process dies part-way through that step, a resume re-runs every tool call of the turn, including ones that already took effect. Lesson R7 · Side effects and idempotency covers this failure window; checkpointing does not make side effects exactly-once.
| Situation | What is persisted | runs.status | What happens next |
|---|---|---|---|
| Run finished | final checkpoint, next = [] | completed | next message starts a new run |
| Error in a node (e.g. Ollama down) | checkpoints up to the failing step, next = [node] | failed | POST {resume: true} retries from that step, or send a new message |
| Process crash / restart mid-run | every checkpoint completed before the crash | running → interrupted on the new process's first database access | if next is non-empty, POST {resume: true} continues |
| Step limit reached | checkpoints so far | failed (RECURSION_LIMIT) | resume (with a fresh limit) or send a new message |
A resume calls graph.stream(null, {thread_id}): null input means "continue from the last checkpoint". Completed steps are not repeated (a tool that already ran is not called again). GET /api/conversations/:id reports resumable: true when the thread has pending steps and no run is active; the UI then shows a Resume button.
Resumable is a property of the thread, not of runs.status: a conversation is resumable exactly when its latest checkpoint has a non-empty next and no run is active. interrupted only says the process that owned the run died before recording an outcome. Two crash windows produce an interrupted run that has no work of its own to resume:
- after LangGraph wrote the final checkpoint but before
finishRun()recordedcompleted: the answer is in the thread andnext = []; - after
startRun()inserted the run row but before LangGraph wrote the input checkpoint: the new message itself was never recorded (send it again), andnextis whatever the thread had before this run.
Graceful shutdown. Sophos has no explicit run-drain or cancellation policy. On SIGTERM/SIGINT, adapter-node waits for open HTTP connections (including SSE streams) up to SHUTDOWN_TIMEOUT, then the shutdown hook closes MCP and SQLite. An in-flight run's outcome depends on timing:
| Timing | Typical result |
|---|---|
| The run finishes before the hook runs (e.g. an open SSE stream held the server open) | completed |
| The hook closes MCP / SQLite while the run is still executing | the run can fail against the closed dependency and be recorded failed (one observed case: The database connection is not open) |
| The process exits before the run's outcome is recorded (e.g. SIGKILL after a container grace period) | the row stays running and is recovered as interrupted by the next process |
In every case, whether the thread can be resumed is decided by next, as above.
These are the same states a future human-in-the-loop interrupt needs (interrupted → resume with a decision), so approval can be added on top of this model without changing persistence.
Execution limits
AGENT_RECURSION_LIMIT (default 25) is passed as LangGraph's recursionLimit: the maximum number of steps per run. Each node execution is one step, so every agent → tools round trip costs two; 25 allows about 12 model calls. In this graph tool iterations are always one fewer than model iterations, so one limit covers both. Exceeding it raises GraphRecursionError, recorded as RECURSION_LIMIT.
What is durable, what is in memory
| Durable (SQLite) | In memory only (lost on restart, by design) |
|---|---|
| conversations and their metadata | activeRuns: the run currently executing per conversation |
| every run record; its outcome once recorded | its buffered SSE events (for Last-Event-ID replay) and subscribers |
| all messages, tool calls and tool results | the compiled graph and MCP connections (rebuilt on demand) |
| checkpoints of every step |
Run outcomes (completed, failed) are durable once the process records them; a run whose process died first stays running until the next process recovers it as interrupted. After a restart a conversation is fully usable: its messages load from the thread, and it is resumable if its thread has a pending step.
Markdown and JSONL: exports only
GET /api/conversations/:id/export generates a Markdown transcript (messages, tool calls, tool results) from the thread; ?format=jsonl lists the thread's checkpoints, one JSON line each, oldest first. Both are generated on request and never read back: deleting an export cannot lose conversation state.
Legacy data/chat/
Earlier versions stored conversations as data/chat/{id}/0001.user.md … trace.jsonl and rebuilt the history from those files. That directory is no longer read or written. Existing files are left in place (nothing deletes them) but do not appear in the UI; there is no importer. Delete the directory when you no longer need it.
Code map
| File | Responsibility |
|---|---|
src/lib/agent/graph.ts | the StateGraph (agent/tools nodes, routing), threadConfig() |
src/lib/agent/index.ts | runtime: builds the graph once, runAgent(), getThreadState() |
src/lib/agent/runs.ts | starts runs, records outcomes, active-run SSE registry |
src/lib/agent/persistence/database.ts | opens SQLite, migrations, createCheckpointer() |
src/lib/agent/persistence/conversations.ts | conversations and runs tables |
src/lib/agent/messages.ts | LangChain messages → MessageDto (the only UI-facing shape) |
src/lib/agent/export.ts | Markdown / checkpoint JSONL exports |