Jam architecture main at 59f67f5d0 · 2026-09-29
View source

#Consistency and recovery

On this pageThree kinds of stateThe event pathInbound delivery is at-least-onceDurable journalsStartup recovery orderContinuity of owned runtimesPartial success is reported, not hiddenLocks and ordering

Jam keeps working across daemon restarts, crashes, sleep, network loss, and partially failed operations through operation-specific journals: the delivery queue and each destructive or long-running operation write durable intent before their side effects, and a named owner replays them on the next start. Not every multi-step flow has a journal. Assigning work, for example, can leave a created room on Band with no local replay, and reports the partial result instead.

#Three kinds of state

Jam moves three kinds of information with different guarantees. Treating an event as if it were a snapshot or a journal loses data after a gap.

Kind Examples Guarantee How a reader stays correct
Event notification EventKind values on GET /v1/events, Band WebSocket pushes Best-effort. Events can drop under load or during a disconnect. Re-read the authoritative query after any gap. Never treat an event as the only record of a fact.
Authoritative snapshot Control::list, Control::status, room history pages, list_work_items Current truth at the moment of the read Hydrate on mount, after reconnect, and after any resync marker
Durable journal Delivery queue, workspace operations, cleanup operations, usage reconciliation outbox Survives a crash. Replayed by a named owner on startup. Only the owner mutates it, and each step is idempotent

The Control::events doc comment in crates/jam-contract/src/lib.rs states the first rule directly: "Best-effort: events can drop under burst or disconnect, so query methods remain the truth."

#The event path

How a change reaches the desktop. The notes mark where it can be lost.

The mechanics:

  • crates/jam-manager/src/fanout.rs merges every worker's events into a tokio::sync::broadcast channel with capacity 256 (CHANNEL_CAP). A subscriber that falls behind loses events, and the module comment says it "reconciles truth via list/status." The fan-out also appends each event to the peer's log file, so it is the single point every peer event passes through.
  • The daemon's events handler (crates/jam-daemon/src/lib.rs) subscribes to the fan-out before taking a peer snapshot, then sends the snapshot as peer_added events followed by live events. Subscribing first closes the race between the snapshot and the first live event. It does not make the stream lossless: if the bounded broadcast overflows, Manager's event stream resubscribes, replaces room snapshots, and emits StreamResync. peer_added is an idempotent upsert on the desktop, so re-sending known peers only refreshes them.
  • crates/jam-ipc/src/events.rs is a reconnecting relay. When the stream ends, it backs off, resubscribes, and emits a single resync to the UI on the first event after the gap. It never fires a resync on the initial connection or while the daemon stays down. The client also injects an in-band EventKind::StreamResync marker when its line reader detects a gap.
  • The desktop store re-hydrates every slice a mounted surface reads when it receives a resync, after a daemon reconnect, and when the window regains focus.

#Inbound delivery is at-least-once

A room message addressed to an agent passes through two durable steps before delivery and is removed only on an explicit acknowledgement. The notes mark the three queue commit points; attachments are stored before the first of them.

The ordering is deliberate, and comments in Engine::handle_inbound (crates/jam-core/src/engine.rs) explain each step:

  • Append before claiming. The processing claim is made after the durable append, "so a refused or lost response cannot create remote work that Jam has no local record for."
  • Append under the routing lock. Route selection, append, and index insertion share one write lock with room retirement and live binding (Engine::bind_session_reconciled). A message for a room being bound at the same instant is either appended before the reconcile, which moves it to the new session's queue, or routed after the insert. It is never stranded in the unrouted queue.
  • No automatic processed mark. Delivery ends the automatic path. The comment states that "delivered-to-the-inbox is not the same as the LLM having handled it." Only an explicit jam ack, jam reply, or owned-runtime disposition marks a message processed. A session that crashes after delivery gets the message again on the next worker start.
  • Duplicates are dropped. A message already in the in-memory index is ignored, so a duplicate Band event cannot trigger a second delivery.
  • Replying and settling are separate steps. Engine::reply_parts posts the reply first, then marks the message processed and removes it from the queue. If either of the last two fails, the reply is already on Band, the call returns warnings, and the message stays queued. A later redelivery can therefore produce a second reply.
  • Unrouted messages wait. If no session owns the room, the message goes to the unrouted queue. It is never sent to a guessed runtime.
  • Lifecycle messages are never work. A [system:...] marker message is claimed and marked processed immediately, without entering the queue.

Each room has its own ordered inbound worker with a bounded spool, so a slow room does not block others.

#Durable journals

Every multi-step mutation that must survive a crash has a journal. The table lists each one, its owner, and when it is replayed.

Journal Table or location Protects Owner Replayed
Delivery queue queue_messages (SQLite) or JSONL files (file store) Inbound messages until acknowledged jam-core::Queue, backend chosen in crates/jam-manager/src/queue.rs On worker start, queues are re-indexed and pending messages redelivered
Provisional provider sessions provisional_sessions A provider session created before Jam learns its ID Manager (recover_incomplete_provisional_sessions) rebuild, per peer, before the worker starts
Runtime configuration transaction host_sessions.desired_config_generation, effective_config_generation, pending_change_json A configuration change between save and apply crates/jam-manager/src/runtime_txn.rs Visible after restart; re-applied when the runtime starts
Workspace operations workspace_operations (current summary) plus workspace_operation_events (append-only) Provision, clone, branch, commit, push, pull request, export, archive, remove crates/jam-manager/src/workspace.rs rebuild via reconcile_all_managed; archive and export deletion finish their recorded intent
Runtime host cleanup runtime_host_cleanup_operations plus _events, forward-only stages Whole-host removal Manager (recover_runtime_host_cleanups) rebuild, before any worker
Docker environment cleanup Owner markers under the sandbox owner root Environment-file sandboxes created by an interrupted operation jam_host::sandbox::reconcile_pending_environment_cleanups rebuild, first step
Provider checkpoint payload deletion provider_checkpoint_payload_deletions Deleting checkpoint bytes while keeping metadata Workspace manager rebuild
Provider continuity reset provider_continuity_resets Starting a fresh provider thread after an unrecoverable one Manager (replay_provider_continuity_resets) rebuild, before workers
Usage reconciliation outbox usage_reconciliation (Pending, Accepted, Conflict), usage_counter_mutations A live usage observation between staging and acceptance crate::worker::reconcile_pending_runtime_usage rebuild
Native work projection outbox native_projection_outbox Projection of private task lanes into the shared board crates/jam-manager/src/workitems.rs NativeWorkProjector::retry_pending rebuild
Room task operations room_task_operations Shared board mutations that are queued or need reconciliation Room board backend Listed by list_room_task_operations; reconciled on request
Attention mutation outbox room_attention_mutation_outbox, room_attention_spool Inbox and attention updates crates/jam-manager/src/roomfeed.rs Room feed start

These journals follow the same rules:

  • The durable intent is written before the external side effect: before a Docker call, a Band call, a file removal, or a provider call.
  • Each step is idempotent, so replaying a partially finished operation converges instead of repeating a side effect.
  • A destructive journal is forward-only. Runtime host cleanup, for example, records Prepared, then SandboxRemoved, then HostRootRemoved, and a replay skips stages that were durably recorded. A side effect that finished before its stage was saved runs again on replay, so each step must tolerate an already-removed resource.
  • Most replay steps in rebuild log a failure and continue, so one failed item does not block unrelated peers. The exception is provisional-session recovery: recover_incomplete_provisional_sessions returns its error with ?, which aborts rebuild before later peers start.

Most of these journals are recent. The chart counts schema migrations by the month they landed on main. The workspace operation, runtime host cleanup, continuity reset, and checkpoint payload deletion tables all arrived in September 2026, in one merge.

Schema migrations added per month

#Startup recovery order

Manager::rebuild in crates/jam-manager/src/manager.rs runs in the background after the local socket is already serving (see runtime topology). It recovers durable state in this order before any agent starts:

The reasons for the order:

  • Destructive replays come first. A half-finished cleanup or deletion must complete before any worker can reuse the paths or sandboxes it was removing.
  • Continuity resets come before workers. The comment in rebuild says a fresh-start authorization "is persisted before any peer, binding, or host mutation. Replay every immutable plan before workers so a crash at any intermediate write converges without ever reusing the failed provider identity."
  • Workspaces are reconciled before workers, "so a daemon restart repairs interrupted publication before a runtime can use the paths."
  • Per-peer failures are mostly isolated. A peer whose account or worker fails to start is marked failed with a warning, and the loop continues with the next peer. A provisional-session recovery error is the exception noted above.
  • The order governs rebuild, not every request. The socket already serves requests while rebuild runs, so a lifecycle request from a client can start a worker concurrently with recovery.

Several conditions keep a peer from starting in step 9:

  • An explicitly stopped peer is published as Stopped and left alone.
  • A terminal (attached) identity waits for its integration to reconnect rather than resuming automatically.
  • A peer on a signed-out account is failed, unless the account still holds an OAuth refresh token. In that case the worker runs on its durable agent key while the refresh loop catches up.
  • A peer that try_start_peer reports as conflicted (Ok(false)) during a restart handoff is left in wants_running. The rebuild doc comment attributes this to the exiting predecessor still holding the per-peer lock. The supervisor loop, started after rebuild and running every 5 seconds, keeps retrying it.

After rebuild, the startup task starts the host monitor (every 15 seconds: stop peers whose coding-agent process exited, reap idle owned runtimes, retry pending removals), the supervisor, the per-account room feeds, the stats backfill, the orphan lane sweep, discovery, and the Linear poller. A read-only Docker ownership inventory runs in its own task so a slow or missing sbx does not delay host-native agents.

#Continuity of owned runtimes

An owned runtime's conversation lives in the provider's own state files, not in Jam. Jam makes it survive a restart by resuming the provider's session by ID and, for managed Docker runtimes, by checkpointing the provider state.

  • Host-native runtimes resume from the durable provider session ID stored on the host session. The provider CLI finds its own state on disk.
  • Managed Docker runtimes take a checkpoint of the reviewed provider files at defined boundaries (after_turn, before_stop, before_reset, before_remove, or manual, per the provider_checkpoints.reason check constraint) and restore it into a new microVM. Each resume attempt is recorded in provider_resume_observations.
  • A bound provider identity with missing, corrupt, or version-incompatible state fails closed. The runtime does not silently start a fresh thread. Recovery is a separate, explicitly confirmed continuity reset, which mints new runtime session and binding IDs and keeps the old evidence. The lifecycle test rebuild_fails_closed_for_unrecoverable_managed_provider_state guards this.

The sandbox chapter describes checkpoints, resets, and the destructive-action journals in detail.

#Partial success is reported, not hidden

When a multi-step user action fails partway, Jam reports what succeeded and what failed rather than rolling back visible results. ControlError has explicit variants for this, including SessionDetachPartialSuccess and RoomKeptAfterKickoffFailure. A created room is kept after a failed kickoff, and a multi-room template apply reports each room's outcome with a retry for the failed rooms.

#Locks and ordering

The manager's concurrency control is a set of named gates rather than one lock. These are the gates a refactor of the manager or engine will touch:

Lock Scope Protects Source
Manager::inner (std::sync::Mutex<Inner>) Whole manager Worker map, pending questions and permissions, runtime transactions, failure views crates/jam-manager/src/manager.rs. Never held across .await.
Account lifecycle lock One profile Sign-in, sign-out, deletion, and rebuild of that account's peers Manager::lock_account_lifecycle
Worker lifecycle gate One peer Start, stop, restart, and removal of that peer's worker Manager::lock_worker_lifecycle
Runtime room-bind gate Whole manager Binding a runtime to a room. close cancels first, then takes this gate, to avoid deadlocking an in-flight bind. Manager::runtime_room_bind_gate
Engine routing lock (RwLock) One engine Route selection, durable append, index, retirement, live bind crates/jam-core/src/engine.rs
Task lane gate One RuntimeSessionId Load, mutate, and commit of one private task lane. The gate covers concurrency only and does not span every writer; the store's single TaskLaneAggregate transaction gives crash atomicity, and an optimistic generation check with bounded rebase covers writers outside the gate, such as the task-file watcher crates/jam-manager/src/task_lane.rs
Daemon lifetime lock One app directory, across processes Prevents a second daemon until shutdown cleanup finishes jam_service::try_acquire_daemon_lock

crates/jam-manager/src/room_connectivity.rs documents its own lock order: account generation, then local ownership, then transport observation, then room state, then fan-out, with no lock held across an await.

Scroll to zoom, drag to pan.