Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

03 — Design durability boundaries

From cognokratos/sophos-agent · docs/runtime/03-design-durability-boundaries.md · pinned revision 8d9fe52182d8

Persist semantics, not everything.

Stage R3 · What must survive? · Learning path · Previous: 02 · Next: 04

Objective

Name every piece of state in the Sophos runtime, say whether it is durable or ephemeral and why, and predict what a graceful stop and a hard kill each leave behind.

Why it matters

Every durable byte is a commitment: it must be migrated, backed up, secured, deleted on request and kept consistent with everything else. Every ephemeral byte is a bet that losing it is acceptable. The boundary between the two is an architectural decision, and in agent systems it is often made by default: by whatever the framework happens to persist.

Mental model

flowchart TB
    RT["Sophos runtime"] --> D["Durable<br/>(survives restart)"]
    RT --> E["Ephemeral<br/>(dies with the process, by design)"]
    D --> SQL["data/db/sophos.db<br/>conversations · runs<br/>checkpoints · writes"]
    D --> KG["data/memory/memory.jsonl<br/>Memory MCP knowledge graph"]
    D --> CFG["config/ · infra/app/config/<br/>mcp.json · system.md"]
    E --> AR["activeRuns<br/>one ActiveRun per conversation"]
    E --> BUF["ActiveRun.events<br/>SSE replay buffer"]
    E --> SUB["ActiveRun.subscribers<br/>+ SSE ping intervals"]
    E --> G["graph · mcpReady<br/>compiled graph, discovery promise"]
    E --> MCP["MCPAdapter<br/>stdio children · HTTP sessions"]
    E --> H["database handle<br/>(the file is durable, the connection is not)"]

The rule Sophos applies:

DurableEphemeral
Facts about what happened and what is left to do: messages, tool calls and results, run outcomes, pending steps, learned knowledgeMachinery for doing it right now: connections, child processes, subscribers, buffers, compiled code
Rebuilding it is impossible or would change the meaningRebuilding it is cheap and deterministic, or losing it is harmless

Where it lives in Sophos

StateName in codeStored inOwner
Conversation metadataconversations tableSQLiteSophos
Run outcomesruns tableSQLiteSophos
Messages, tool calls, tool resultsmessages channel of the thread statecheckpoints (SQLite)LangGraph (SqliteSaver)
Pending task outputs and task errorspending writeswrites (SQLite)LangGraph
Long-term knowledgeentities, relations, observationsmemory.jsonlMemory MCP server
Active run per conversationactiveRuns (src/lib/agent/runs.ts)RAMSophos
SSE replay bufferActiveRun.eventsRAMSophos
Connected browsersActiveRun.subscribersRAMSophos (route handler)
Compiled graph, discovery stategraph, mcpReady (src/lib/agent/index.ts)RAMSophos
MCP connections and child processesMCPAdapter in MCPClientServiceRAM / OS processesthe adapter
System promptconfig/system.mdfile, read on every model call, never stored in the threadyou

The comment at the top of src/lib/agent/runs.ts states the split; 3a has the reference table.

Experiment

Classify first

Fill in the table before you run anything. For each item, decide whether it should survive a restart and write down why in one line.

ItemShould survive restart?Why?
assistant message
tool result (e.g. a fetched page)
active HTTP connection (SSE stream)
checkpoint
pending write of a failed task
stdio pipe to the Memory server
run status
SSE replay buffer
the "one active run per conversation" flag
compiled graph
long-term memory entity
the system prompt as sent to the model

Graceful stop

Use the lab environment. Create a conversation with a completed tool call:

C=$(curl -s -X POST $B/api/chat -H 'content-type: application/json' \
  -d '{"message":"Fetch https://example.com and tell me in one sentence what it says."}' | jq -r .conversation)
curl -sN "$B/api/chat?conversation=$C" | tail -2
ls data/lab/db

Predict: after Ctrl+C, which files remain in data/lab/db? Is the fetched page still somewhere on disk?

Stop the lab server with Ctrl+C, then:

ls data/lab/db
sqlite3 $DB "SELECT count(*) FROM checkpoints WHERE thread_id = '$C'"

Start it again and reload the conversation (curl -s $B/api/conversations/$C | jq '.messages[] | {role, name}', or the UI).

Observed: while the server ran, data/lab/db held sophos.db, sophos.db-wal and sophos.db-shm; after the graceful stop, only sophos.db (closing the connection checkpointed the WAL). After restart, every message reloaded, including the tool result: the fetched page content is now part of your durable state.

Hard kill

Start a long run and kill the process (not Ctrl+C) while it streams:

C=$(curl -s -X POST $B/api/chat -H 'content-type: application/json' \
  -d '{"message":"Write a detailed 500-word essay about SQLite WAL mode."}' | jq -r .conversation)
sleep 3
kill -9 $(lsof -nP -tiTCP:5174 -sTCP:LISTEN)
ls data/lab/db
sqlite3 $DB "SELECT status FROM runs WHERE conversation_id = '$C'"

Predict: is the run running, interrupted or failed right now? What happens to it when the server starts?

Observed: running, with the -wal file still on disk. Start the lab server, call curl -s $B/api/conversations > /dev/null, and query again: interrupted. SQLite replays its WAL on open; Sophos replays nothing: recoverInterruptedRuns() simply relabels runs that can no longer be running. The essay tokens the stream had shown are gone. Here the thread was resumable (next: ["agent"]), because the kill landed mid-step; interrupted alone doesn't promise that (lesson 06).

Why the system behaves this way

  • One file for everything durable that Sophos owns. App tables reuse the checkpointer's better-sqlite3 connection: one thing to back up, inspect or delete (5) Trade-off Decisions).
  • Messages are stored once, in the thread. The conversations table holds only metadata and a denormalized message_count, so the sidebar never deserializes a checkpoint.
  • durability: 'sync' (src/lib/agent/index.ts) makes each checkpoint durable before the next step starts. SQLite's WAL makes each write atomic.
  • Nothing ephemeral is worth restoring. A connection to a process that no longer exists, a subscriber whose socket is closed, a buffer of tokens whose meaning is already in the checkpoint: restoring them would be wrong, not just wasteful.

What this does NOT guarantee

  • Durable is not safe. The SQLite file and memory.jsonl are unencrypted, and tool results (whole fetched pages) are stored in the checkpoints (9) Security Posture).
  • Nothing is pruned. Every checkpoint of every run is kept; the database only grows.
  • Two durable stores, no transaction across them. A run can write to memory.jsonl and then fail before its checkpoint lands. See lesson 07.
  • Configuration is durable but not versioned with the state. Editing config/system.md changes how every existing conversation continues.

Questions

Answer both with examples from Sophos:

  1. What would go wrong if we persisted every transient transport detail? Think about restoring an ActiveRun with subscribers that no longer exist, an SSE buffer of tokens next to the checkpointed message they spell out, or an MCP session id for a server process that died.
  2. What would go wrong if we persisted too little? Think about the pre-SQLite design, which stored user and assistant text but not tool calls, tool results or run status.

Takeaway

Persist semantics, not everything.

Durable state in Sophos answers two questions: what happened? and what is left to do? Everything else is rebuilt or let go.

Go deeper

This chapter is maintained in cognokratos/sophos-agent beside the code it teaches. The book shows docs/runtime/03-design-durability-boundaries.md at revision 8d9fe52182d8441454916ec8a6ab13c0773228e2 (branch main). View source at this revision · Report a correction.

Corrections are made upstream against the current main branch and appear here when the book's pin for this source is updated.