#Consistency and recovery
On this page
Three kinds of stateThe event pathInbound delivery is at-least-onceDurable journalsStartup recovery orderContinuity of owned runtimesPartial success is reported, not hiddenLocks and orderingJam keeps working across daemon restarts, crashes, sleep, network loss, and partially failed operations through operation-specific journals: the delivery queue and each destructive or long-running operation write durable intent before their side effects, and a named owner replays them on the next start. Not every multi-step flow has a journal. Assigning work, for example, can leave a created room on Band with no local replay, and reports the partial result instead.
#Three kinds of state
Jam moves three kinds of information with different guarantees. Treating an event as if it were a snapshot or a journal loses data after a gap.
| Kind | Examples | Guarantee | How a reader stays correct |
|---|---|---|---|
| Event notification | EventKind values on GET /v1/events, Band WebSocket pushes |
Best-effort. Events can drop under load or during a disconnect. | Re-read the authoritative query after any gap. Never treat an event as the only record of a fact. |
| Authoritative snapshot | Control::list, Control::status, room history pages, list_work_items |
Current truth at the moment of the read | Hydrate on mount, after reconnect, and after any resync marker |
| Durable journal | Delivery queue, workspace operations, cleanup operations, usage reconciliation outbox | Survives a crash. Replayed by a named owner on startup. | Only the owner mutates it, and each step is idempotent |
The Control::events doc comment in crates/jam-contract/src/lib.rs states the first rule directly: "Best-effort: events can drop under burst or disconnect, so query methods remain the truth."
#The event path
How a change reaches the desktop. The notes mark where it can be lost.
The mechanics:
crates/jam-manager/src/fanout.rsmerges every worker's events into atokio::sync::broadcastchannel with capacity 256 (CHANNEL_CAP). A subscriber that falls behind loses events, and the module comment says it "reconciles truth vialist/status." The fan-out also appends each event to the peer's log file, so it is the single point every peer event passes through.- The daemon's
eventshandler (crates/jam-daemon/src/lib.rs) subscribes to the fan-out before taking a peer snapshot, then sends the snapshot aspeer_addedevents followed by live events. Subscribing first closes the race between the snapshot and the first live event. It does not make the stream lossless: if the bounded broadcast overflows,Manager's event stream resubscribes, replaces room snapshots, and emitsStreamResync.peer_addedis an idempotent upsert on the desktop, so re-sending known peers only refreshes them. crates/jam-ipc/src/events.rsis a reconnecting relay. When the stream ends, it backs off, resubscribes, and emits a single resync to the UI on the first event after the gap. It never fires a resync on the initial connection or while the daemon stays down. The client also injects an in-bandEventKind::StreamResyncmarker when its line reader detects a gap.- The desktop store re-hydrates every slice a mounted surface reads when it receives a resync, after a daemon reconnect, and when the window regains focus.
#Inbound delivery is at-least-once
A room message addressed to an agent passes through two durable steps before delivery and is removed only on an explicit acknowledgement. The notes mark the three queue commit points; attachments are stored before the first of them.
The ordering is deliberate, and comments in Engine::handle_inbound (crates/jam-core/src/engine.rs) explain each step:
- Append before claiming. The processing claim is made after the durable append, "so a refused or lost response cannot create remote work that Jam has no local record for."
- Append under the routing lock. Route selection, append, and index insertion share one write lock with room retirement and live binding (
Engine::bind_session_reconciled). A message for a room being bound at the same instant is either appended before the reconcile, which moves it to the new session's queue, or routed after the insert. It is never stranded in the unrouted queue. - No automatic processed mark. Delivery ends the automatic path. The comment states that "delivered-to-the-inbox is not the same as the LLM having handled it." Only an explicit
jam ack,jam reply, or owned-runtime disposition marks a message processed. A session that crashes after delivery gets the message again on the next worker start. - Duplicates are dropped. A message already in the in-memory index is ignored, so a duplicate Band event cannot trigger a second delivery.
- Replying and settling are separate steps.
Engine::reply_partsposts the reply first, then marks the message processed and removes it from the queue. If either of the last two fails, the reply is already on Band, the call returns warnings, and the message stays queued. A later redelivery can therefore produce a second reply. - Unrouted messages wait. If no session owns the room, the message goes to the unrouted queue. It is never sent to a guessed runtime.
- Lifecycle messages are never work. A
[system:...]marker message is claimed and marked processed immediately, without entering the queue.
Each room has its own ordered inbound worker with a bounded spool, so a slow room does not block others.
#Durable journals
Every multi-step mutation that must survive a crash has a journal. The table lists each one, its owner, and when it is replayed.
| Journal | Table or location | Protects | Owner | Replayed |
|---|---|---|---|---|
| Delivery queue | queue_messages (SQLite) or JSONL files (file store) |
Inbound messages until acknowledged | jam-core::Queue, backend chosen in crates/jam-manager/src/queue.rs |
On worker start, queues are re-indexed and pending messages redelivered |
| Provisional provider sessions | provisional_sessions |
A provider session created before Jam learns its ID | Manager (recover_incomplete_provisional_sessions) |
rebuild, per peer, before the worker starts |
| Runtime configuration transaction | host_sessions.desired_config_generation, effective_config_generation, pending_change_json |
A configuration change between save and apply | crates/jam-manager/src/runtime_txn.rs |
Visible after restart; re-applied when the runtime starts |
| Workspace operations | workspace_operations (current summary) plus workspace_operation_events (append-only) |
Provision, clone, branch, commit, push, pull request, export, archive, remove | crates/jam-manager/src/workspace.rs |
rebuild via reconcile_all_managed; archive and export deletion finish their recorded intent |
| Runtime host cleanup | runtime_host_cleanup_operations plus _events, forward-only stages |
Whole-host removal | Manager (recover_runtime_host_cleanups) |
rebuild, before any worker |
| Docker environment cleanup | Owner markers under the sandbox owner root | Environment-file sandboxes created by an interrupted operation | jam_host::sandbox::reconcile_pending_environment_cleanups |
rebuild, first step |
| Provider checkpoint payload deletion | provider_checkpoint_payload_deletions |
Deleting checkpoint bytes while keeping metadata | Workspace manager | rebuild |
| Provider continuity reset | provider_continuity_resets |
Starting a fresh provider thread after an unrecoverable one | Manager (replay_provider_continuity_resets) |
rebuild, before workers |
| Usage reconciliation outbox | usage_reconciliation (Pending, Accepted, Conflict), usage_counter_mutations |
A live usage observation between staging and acceptance | crate::worker::reconcile_pending_runtime_usage |
rebuild |
| Native work projection outbox | native_projection_outbox |
Projection of private task lanes into the shared board | crates/jam-manager/src/workitems.rs NativeWorkProjector::retry_pending |
rebuild |
| Room task operations | room_task_operations |
Shared board mutations that are queued or need reconciliation | Room board backend | Listed by list_room_task_operations; reconciled on request |
| Attention mutation outbox | room_attention_mutation_outbox, room_attention_spool |
Inbox and attention updates | crates/jam-manager/src/roomfeed.rs |
Room feed start |
These journals follow the same rules:
- The durable intent is written before the external side effect: before a Docker call, a Band call, a file removal, or a provider call.
- Each step is idempotent, so replaying a partially finished operation converges instead of repeating a side effect.
- A destructive journal is forward-only. Runtime host cleanup, for example, records
Prepared, thenSandboxRemoved, thenHostRootRemoved, and a replay skips stages that were durably recorded. A side effect that finished before its stage was saved runs again on replay, so each step must tolerate an already-removed resource. - Most replay steps in
rebuildlog a failure and continue, so one failed item does not block unrelated peers. The exception is provisional-session recovery:recover_incomplete_provisional_sessionsreturns its error with?, which abortsrebuildbefore later peers start.
Most of these journals are recent. The chart counts schema migrations by the month they landed on main. The workspace operation, runtime host cleanup, continuity reset, and checkpoint payload deletion tables all arrived in September 2026, in one merge.
#Startup recovery order
Manager::rebuild in crates/jam-manager/src/manager.rs runs in the background after the local socket is already serving (see runtime topology). It recovers durable state in this order before any agent starts:
The reasons for the order:
- Destructive replays come first. A half-finished cleanup or deletion must complete before any worker can reuse the paths or sandboxes it was removing.
- Continuity resets come before workers. The comment in
rebuildsays a fresh-start authorization "is persisted before any peer, binding, or host mutation. Replay every immutable plan before workers so a crash at any intermediate write converges without ever reusing the failed provider identity." - Workspaces are reconciled before workers, "so a daemon restart repairs interrupted publication before a runtime can use the paths."
- Per-peer failures are mostly isolated. A peer whose account or worker fails to start is marked failed with a warning, and the loop continues with the next peer. A provisional-session recovery error is the exception noted above.
- The order governs
rebuild, not every request. The socket already serves requests whilerebuildruns, so a lifecycle request from a client can start a worker concurrently with recovery.
Several conditions keep a peer from starting in step 9:
- An explicitly stopped peer is published as
Stoppedand left alone. - A terminal (attached) identity waits for its integration to reconnect rather than resuming automatically.
- A peer on a signed-out account is failed, unless the account still holds an OAuth refresh token. In that case the worker runs on its durable agent key while the refresh loop catches up.
- A peer that
try_start_peerreports as conflicted (Ok(false)) during a restart handoff is left inwants_running. Therebuilddoc comment attributes this to the exiting predecessor still holding the per-peer lock. The supervisor loop, started afterrebuildand running every 5 seconds, keeps retrying it.
After rebuild, the startup task starts the host monitor (every 15 seconds: stop peers whose coding-agent process exited, reap idle owned runtimes, retry pending removals), the supervisor, the per-account room feeds, the stats backfill, the orphan lane sweep, discovery, and the Linear poller. A read-only Docker ownership inventory runs in its own task so a slow or missing sbx does not delay host-native agents.
#Continuity of owned runtimes
An owned runtime's conversation lives in the provider's own state files, not in Jam. Jam makes it survive a restart by resuming the provider's session by ID and, for managed Docker runtimes, by checkpointing the provider state.
- Host-native runtimes resume from the durable provider session ID stored on the host session. The provider CLI finds its own state on disk.
- Managed Docker runtimes take a checkpoint of the reviewed provider files at defined boundaries (
after_turn,before_stop,before_reset,before_remove, ormanual, per theprovider_checkpoints.reasoncheck constraint) and restore it into a new microVM. Each resume attempt is recorded inprovider_resume_observations. - A bound provider identity with missing, corrupt, or version-incompatible state fails closed. The runtime does not silently start a fresh thread. Recovery is a separate, explicitly confirmed continuity reset, which mints new runtime session and binding IDs and keeps the old evidence. The lifecycle test
rebuild_fails_closed_for_unrecoverable_managed_provider_stateguards this.
The sandbox chapter describes checkpoints, resets, and the destructive-action journals in detail.
#Partial success is reported, not hidden
When a multi-step user action fails partway, Jam reports what succeeded and what failed rather than rolling back visible results. ControlError has explicit variants for this, including SessionDetachPartialSuccess and RoomKeptAfterKickoffFailure. A created room is kept after a failed kickoff, and a multi-room template apply reports each room's outcome with a retry for the failed rooms.
#Locks and ordering
The manager's concurrency control is a set of named gates rather than one lock. These are the gates a refactor of the manager or engine will touch:
| Lock | Scope | Protects | Source |
|---|---|---|---|
Manager::inner (std::sync::Mutex<Inner>) |
Whole manager | Worker map, pending questions and permissions, runtime transactions, failure views | crates/jam-manager/src/manager.rs. Never held across .await. |
| Account lifecycle lock | One profile | Sign-in, sign-out, deletion, and rebuild of that account's peers | Manager::lock_account_lifecycle |
| Worker lifecycle gate | One peer | Start, stop, restart, and removal of that peer's worker | Manager::lock_worker_lifecycle |
| Runtime room-bind gate | Whole manager | Binding a runtime to a room. close cancels first, then takes this gate, to avoid deadlocking an in-flight bind. |
Manager::runtime_room_bind_gate |
Engine routing lock (RwLock) |
One engine | Route selection, durable append, index, retirement, live bind | crates/jam-core/src/engine.rs |
| Task lane gate | One RuntimeSessionId |
Load, mutate, and commit of one private task lane. The gate covers concurrency only and does not span every writer; the store's single TaskLaneAggregate transaction gives crash atomicity, and an optimistic generation check with bounded rebase covers writers outside the gate, such as the task-file watcher |
crates/jam-manager/src/task_lane.rs |
| Daemon lifetime lock | One app directory, across processes | Prevents a second daemon until shutdown cleanup finishes | jam_service::try_acquire_daemon_lock |
crates/jam-manager/src/room_connectivity.rs documents its own lock order: account generation, then local ownership, then transport observation, then room state, then fan-out, with no lock held across an await.