Jam architecture main at 59f67f5d0 · 2026-09-29
View source

#Execution and lifecycle

On this pageWhere it livesHow it worksThe objects and how many of each existPlacement: Shared and DedicatedCreating an owned agentBinding a room: materialization, dormancy, and wakeThe worker and its supervisor taskRun intent and the supervisor loopThe host monitor loop and idle reapingStopping and restarting at three scopesConfiguration transactions: desired, effective, pendingStartup: Manager::rebuildShutdown: Manager::close and the process deadlinesThe Manager struct and the Inner mutexExperiment gating and knobsState and ownershipContractsProvided: Control methodsProvided: eventsConsumed: the host factory and the Host traitConsumed: BandInvariantsFailure and recoveryExtension pointsAdding a harness that the lifecycle can runAdding a lifecycle operationRefactor notes

This subsystem decides when a coding agent runs, where it runs, and under which configuration. It keeps one durable record per local agent identity (the Peer), one routing row per room (the HostSession), and, for Docker-backed runtimes, durable ownership records for the execution boundary (the runtime host) and each room's runtime context. It starts and supervises one Band worker per peer, materializes room-bound runtimes from a parked template when the agent joins a room, stops idle runtimes and wakes them on the next message, and applies configuration changes without claiming a change is live before a runtime accepted it.

It does not implement any provider protocol (that is jam-host, covered in Providers and attached agents), does not route or queue individual messages (see Messaging and room state), and does not own Docker sandbox, workspace, or checkpoint mechanics (see Sandbox, workspaces, and continuity). It never infers host-wide scope from a room, never routes a room to a guessed runtime, and never lets the desktop own a process.

#Where it lives

Area Crate and module Key types and functions
Identity record crates/jam-domain/src/peer.rs Peer, IdentityClass (Terminal, Owned), PeerState, Peer::id, Peer::routable_sessions, Peer::resolved_identity_class
Room binding row and runtime config crates/jam-domain/src/host_session.rs (8,330 lines) HostSession, HostSession::is_parked, is_owned_runtime, is_owned_codex_app_server and siblings, has_outstanding_config_change, HostRuntime, HostTransport, HostRuntimeConfig, HostRuntimeConfig::resolved_host_placement, resolved_sandbox_capability_cell, the transport_catalog! macro
Identifiers crates/jam-domain/src/ids.rs PeerKey, IdentityId, ChatId, SessionId, RuntimeTemplateId, RuntimeSessionId, RuntimeHostId, RuntimeBindingId, uuid_v4_from_bytes
Runtime host records crates/jam-domain/src/runtime_workspace.rs RuntimeHostPlacement (SharedAgentHost, DedicatedSessionHost), RuntimeHostRecord, RuntimeHostState, RuntimeHostSessionRecord, RuntimeHostSessionState, RuntimeSessionRoomBinding, ProviderCheckpointReason
Configuration transition record crates/jam-domain/src/host_runtime_catalog.rs PendingChange, PendingChangeStatus, PendingChangeAction, ProvisionalSessionRecord and its stage and cleanup enums
Manager and its state crates/jam-manager/src/manager.rs (58,145 lines) Manager, Inner, Manager::lock, ensure_peer, start_and_wait, try_start_peer, rebuild, supervise, spawn_supervisor, spawn_host_monitor, check_host_processes, reap_idle_runtimes, wake_dormant_room, auto_bind_owned_runtime_room, bind_session_live_under_gate, restart_peer, restart_session, detach_session, control_runtime_host, save_runtime_template, close
Worker crates/jam-manager/src/worker.rs (10,609 lines) Worker, LiveHost, Reaped, spawn_worker, supervise (the per-peer task), Worker::bind_live, reap_session, retire_session, reinsert_session, prepare_host_until_shutdown, managed_runtime_context_for_session, ensure_managed_runtime_ownership, owned_runtime_lazy_enabled, STOP_GRACE
Runtime host allocation crates/jam-manager/src/runtime_host.rs RuntimeHostAllocator::allocate, RuntimeHostCompatibility::from_runtime, fingerprint, incompatibility_reason
Configuration transactions crates/jam-manager/src/runtime_txn.rs RuntimeTxnState, RuntimeConfigGeneration, PendingReason, AckOutcome, ActiveRuntimeBinding, ProvisionalSession
Composition root bins/jam/src/jamd.rs main, build_host, FORCE_EXIT_GRACE, the new_host closure, CodexAppServerRegistry, CopilotSdkRegistry construction
Provider seam used here crates/jam-host/src/lib.rs Host trait (prepare, teardown, teardown_for, runtime_pid, runtime_transport_active, idle_reap_policy, readiness_warning), IdleReapPolicy, RuntimeHostContext
Process-local Shared registries crates/jam-host/src/codex/shared_dispatch.rs, codex/app_server.rs, copilot/sdk.rs SharedHostRegistry, CodexAppServerRegistry::build, CopilotSdkRegistry::build
Store tables crates/jam-store/src/sqlite.rs, crates/jam-store/src/migrations/0045_runtime_host_workspace_foundation.sql peers, host_sessions (38 columns), provisional_sessions, runtime_hosts, runtime_host_sessions, runtime_session_room_bindings, settings

The dependency direction matches the intended layering at crate level (jam-domain to jam-contract to jam-manager to jam-daemon, composed in bins/jam), but not at module level: jam-manager depends on jam-host with the runtime-providers feature and names concrete adapter modules. See Refactor notes.

The components that drive an agent's lifecycle, and who calls whom.

#How it works

#The objects and how many of each exist

Six objects participate in running one agent, and most lifecycle bugs come from treating two of them as one. Domain and authority names every identifier. Here is how the lifecycle uses them.

A peer is the local record of one Band agent under one account profile. Peer is keyed by PeerKey, the string profile/scope returned by Peer::id. It holds the Band agent_id, the agent API key (a SecretKey, stored in the OS secret store by SqliteStore::save_peer), the adapter name in host, the operator briefing, and a list of host sessions. identity_class records whether the identity is an interactive terminal window (Terminal) or a daemon-owned runtime (Owned). The field is stamped at onboard because jam detach can clear every session row, so rows alone cannot tell an idle terminal identity from an owned agent. Legacy records without the stamp resolve through Peer::resolved_identity_class: any Owned row means Owned, otherwise Terminal.

A host session is one row in Peer::host_sessions. Each row binds at most one room. Its id is a SessionId, an operator-facing alias such as default or default-<room>, and it is never an identity key. Its runtime field is a HostRuntimeConfig that says how the agent runs: mode (HostRuntime::AttachInbox for a user-owned process, Owned for a jamd-owned process), transport (AttachInbox, Acp, CodexAppServer, CopilotSdk, ClaudeCodeCli, Opencode, the legacy Pty marker, or Unknown), spawn settings, sandbox settings, placement, workspace source, network intent, and the two immutable UUIDs template_id and session_id.

A runtime template is a host session whose room is empty. HostSession::is_parked is literally self.room.as_str().is_empty(). The worker builds no routing surface for a parked row (Peer::routable_sessions filters it out), so a template can never receive a message. For an owned agent the parked row is the configuration that room membership copies. For an attached agent a parked row holds a waiting window's PID and mailbox until the window is bound to a room.

A room-bound session is a host session with a room. For owned runtimes it is materialized from the template: owned_runtime_room_bind_opts in manager.rs copies the template's settings, names the new row <template id>-<room id>, sets runtime_template_id to the template's operator-facing id, and clears provider attribution. Manager::upsert_host_session then copies the template's immutable RuntimeTemplateId UUID into the new row before apply_runtime_opts mints a fresh RuntimeSessionId. The comment on that copy states why: without it every room would become a different template and could never share a compatible host.

A runtime session (RuntimeSessionId) is the UUID in HostRuntimeConfig::session_id. It names one local runtime context and stays stable across renames and restarts. apply_runtime_opts mints it and RuntimeTemplateId with getrandom when an owned runtime has none (mint_runtime_session_id, mint_runtime_template_id).

A runtime host (RuntimeHostId) is one concrete execution boundary: one provider process and, for Docker, one microVM. Runtime host records exist only for Docker-sandboxed owned sessions. managed_runtime_context in worker.rs returns None when !session.is_owned_runtime() || !session.runtime.spawn.sandbox.enabled, and that function is the only production caller of ensure_managed_runtime_ownership, which calls RuntimeHostAllocator::allocate and writes the room binding. A host-native owned Codex session therefore has a RuntimeSessionId but no RuntimeHostRecord, no runtime_host_sessions row, and no RuntimeSessionRoomBinding.

A room binding (RuntimeBindingId, record RuntimeSessionRoomBinding) ties one runtime session to one room for one agent, with an active flag. Room removal sets active = false instead of deleting the row. That inactive row is a tombstone: an authoritative re-add of the same room (auto_bind_owned_runtime_room with authoritative_readd) looks it up with runtime_room_binding_including_inactive, reuses the exact runtime session, and passes the saved provider session id so the provider resumes the same conversation.

The records that exist for one agent, and how many of each.

The cardinalities come from code and schema:

  • One room per host session, one session per room within a peer. Manager::upsert_host_session refuses a bind whose room another session owns unless force (--steal) is set, and refuses a second session on the same attached PID unless allow_pid_share is set.
  • One runtime host per runtime session. runtime_host_sessions.runtime_session_id is the primary key in migration 0045_runtime_host_workspace_foundation.sql.
  • At most one live Shared host per agent. The partial unique index one_live_shared_runtime_host_per_agent covers placement = 'shared_agent_host' AND state NOT IN ('archived', 'removed').
  • At most one active binding per agent and room. The partial unique index one_active_runtime_binding_per_agent_room covers active = 1.
  • A process is not a session. A Shared runtime host serves several runtime sessions from one provider process, and an attached window's PID can back sessions in several rooms when the reconcile path sets allow_pid_share.

#Placement: Shared and Dedicated

RuntimeHostPlacement has two values. SharedAgentHost reuses one compatible runtime host for every room of one agent identity. DedicatedSessionHost gives one runtime session its own host. The Rust Default is SharedAgentHost, but RuntimeHostPlacement::from_existing(None) resolves to DedicatedSessionHost, because records written before placement existed ran one microVM per session and must not be silently merged. HostRuntimeConfig::resolved_host_placement is the one accessor that applies that rule.

RuntimeHostAllocator::allocate in runtime_host.rs makes the decision while holding the allocator's allocation_gate mutex. That gate belongs to one allocator instance, and the only production caller, ensure_managed_runtime_ownership, constructs a new RuntimeHostAllocator on every call, so the gate does not serialize separate production allocations. The store's primary key and partial unique indexes prevent duplicate owners. The live-bind path also holds the room-bind gate, but the spawn path (spawn_worker through managed_runtime_context_for_session) does not. The allocator decides in this order:

  1. If the runtime session already has a membership, it reuses that host after checking compatibility (ensure_session_compatible). This is what makes a retry after a crash converge on the same record.
  2. For Dedicated, the host id is the runtime session UUID itself (RuntimeHostId::new(request.runtime_session_id.as_str())), so a crash between host insert and membership attach heals the same record.
  3. For Shared, it looks for the agent's live Shared host. If one exists it must pass ensure_shared_compatible. Otherwise it mints a random UUID and inserts. If the insert loses a race on the unique index, it reloads the winner and converges when compatible.

Compatibility is an opaque SHA-256 digest (RuntimeHostCompatibility::fingerprint, domain tag jam.runtime-host-compatibility.v1) over the host-wide inputs: runtime mode, transport, command, arguments, sorted environment allowlist, auth mode, runtime version, the sandbox-enabled flag, sandbox agent, kits, static MCP servers, reference folders, Docker configuration source, GitHub access, Codex control channel, network profile, developer destinations, and the local platform bridge flag. Session-local values (working directory, workspace identity, thread settings) are excluded so rooms can differ in those and still share a host. incompatibility_reason also compares agent UUID, template UUID, placement, provider, sandbox name, and terminal state. The allocator tests cover the rules by name, for example shared_placement_reuses_one_compatible_host_for_two_sessions, dedicated_placement_is_stable_and_isolated_per_session, incompatible_live_shared_host_is_rejected_without_second_owner, and same_template_with_different_network_profile_cannot_share_a_host.

Durable allocation is only half of Shared placement. The other half is process-local: jamd creates one CodexAppServerRegistry and one CopilotSdkRegistry, each wrapping a SharedHostRegistry keyed by RuntimeHostId with weak entries. When build_host builds a Codex or Copilot facade whose managed context says SharedAgentHost, the registry returns the existing live runner for that host UUID or initializes one, and fails closed if the new facade's host-level inputs differ. Each room still gets its own facade, provider thread, working directory, and event stream.

Which transports may use Shared is declared in the transport catalog's capability cells (CODEX_SANDBOX_CAPABILITY_CELLS, COPILOT_SANDBOX_CAPABILITY_CELLS, CLAUDE_SANDBOX_CAPABILITY_CELLS, CURSOR_SANDBOX_CAPABILITY_CELLS in host_session.rs). Each of those lists one Shared cell, for a Jam-managed configuration with a Managed workspace. HostRuntimeConfig::resolved_sandbox_capability_cell refuses any combination without a cell.

#Creating an owned agent

Manager::ensure_peer is the single create-or-start operation. The desktop's createLocalAgent in apps/desktop/src/App.tsx calls it through ensureOnboardedAgent with runtime_mode: "owned", an empty room, and the draft's transport and sandbox settings (draftToEnsureOpts in apps/desktop/src/lib/runtimeDraft.ts). The call reaches the daemon at POST /v1/ensure. The client allows it ENSURE_REQUEST_TIMEOUT (60 seconds); the daemon applies its default 45-second unary timeout because /v1/ensure is not in RUNTIME_CHILD_ROUTES.

ensure_peer does the following:

  1. Validates network intent against the operator gate, derives or normalizes the scope to lowercase (reusing a legacy cased record through lookup_peer_ci), and reaps any worker whose supervisor already exited (reap_finished_workers).
  2. If a worker is live, it merges the request into the existing record (upsert_ensure_host_session) and either returns, persists a changed parked row, or live-binds a changed room-bound row. create_only refuses here so a new-agent form can never mutate an existing identity.
  3. Otherwise it inserts the peer key into Inner::provisioning, which makes a concurrent ensure for the same key return ManagerError::Provisioning.
  4. For a stored peer it clears the persisted stop intent and calls start_and_wait. For a new peer it calls provision, which registers the Band agent through the Human API, persists an identity-create recovery marker (identity-create-recovery:v1: setting) before any local write, saves the peer, and clears the marker. A failure after registration deletes the Band agent (rollback_registered_identity).
  5. start_and_wait takes the per-peer worker lifecycle gate, runs GitHub preflights and managed-workspace preparation, calls spawn_worker, registers the worker, and waits up to READINESS_TIMEOUT (35 seconds) for PeerState::Connected or Degraded on the fan-out. On success it adds the key to wants_running. On Failed, a closed channel, or timeout it removes and stops the worker and returns StartupFailed or StartupTimeout. A brand-new peer whose start fails is deprovisioned.
  6. On success for a new peer, the manager publishes EventKind::PeerAdded and an analytics event.

A new owned agent normally has only a parked template, so spawn_worker builds an engine and Band subscription but no host and no provider child. The agent is online on Band and runs nothing until it joins a room.

What happens when the desktop creates an owned agent, and which writes become durable.

#Binding a room: materialization, dormancy, and wake

Room membership is Band's decision. The desktop adds an agent to a room with ipc.addRoomParticipant (a Human API call through Manager::add_room_participant), which does not bind anything locally. Band then sends room_added on the agent's WebSocket. The engine turns it into EventKind::ChatAdded, the worker's supervise task maps it to a RoomMembershipUpdate (room_membership_update in worker.rs), and the manager's room_membership_sink spawns handle_room_membership_update. After the CLI-pairing disposition, that calls auto_bind_owned_runtime_room and then the external-window reconcile.

The same binder also runs level-triggered. On every worker spawn, the first Connected or Degraded state fires the one-shot on_ready hook, whose on_ready_reconcile lists the agent's Band rooms once and calls auto_bind_owned_runtime_rooms. This is how a daemon restart binds rooms joined while it was down. AGENTS.md states the requirement as "Startup/reconcile must bind current rooms, not just future events."

auto_bind_owned_runtime_room runs under runtime_room_bind_gate, re-resolves the peer from the store, and skips rooms that are retiring or known inaccessible unless the event is an authoritative re-add. It picks an anchor: the exact inactive binding's parked session when one exists, otherwise owned_runtime_anchor, which prefers the first owned parked template. Then it chooses between two outcomes:

  • Dormant registration. When the lazy lifecycle is on (the default, JAM_OWNED_RUNTIME_LAZY unset or not 0) and the unrouted queue holds no message for that room (room_has_backlog), register_dormant_runtime_room persists the room-bound row and updates the worker's peer snapshot but builds no host. It then re-checks the backlog under the still-held gate, so a message that arrived during registration triggers an immediate bind.
  • Live bind. Otherwise it calls bind_session_live_under_gate.

A dormant room is woken by its first message. The worker's supervise calls the on_inbound hook for every inbound message. on_inbound_wake checks is_dormant_owned_runtime_room, reserves a per-room slot in Inner::waking_rooms (an RAII WakeGuard clears it even on panic, so a burst coalesces to one wake), and spawns wake_dormant_room. That function takes the room-bind gate, re-resolves the peer, confirms the room is dormant or crashed (session_is_dormant, session_runtime_crashed), retires a crashed runtime's stale host first, and then calls bind_session_live_under_gate with RuntimeStartTrigger::Message and OnlineCause::WakeResume. The live bind reconciles the unrouted queue, so the buffered message is re-homed to the session and redelivered.

bind_session_live_under_gate is the single path that turns a room-bound row into a live runtime. It is also used by explicit attach, invite, restart_session, host restart, configuration replacement, and wake. It is about 700 lines long and has five phases:

  1. Plan in memory. Take the per-peer worker lifecycle gate. Resolve which session the bind lands on (bind_session_id) before mutating anything. Compute the new host-session set with upsert_host_session. Nothing is persisted yet.
  2. Resolve ownership. Prepare legacy host-native workspaces. For Docker-sandboxed owned sessions, managed_runtime_context_for_session allocates or reuses the runtime host and membership, ensures the managed workspace, and upserts the active room binding. These records are durable before any provider starts. If this is a fresh allocation and a later step fails, rollback_fresh_managed_runtime_failure removes only the freshly created ownership.
  3. Journal. Build the RuntimeHostContext (room roster, tools, brokers, checkpoint broker). If the session is already live on the same id, tear it down first so two children never own the same provider state. Write a ProvisionalSessionRecord with stage Spawning to provisional_sessions before calling the host factory.
  4. Prepare. Call Deps::new_host (the build_host closure in jamd) and prepare_host_and_initial_repositories_until_shutdown, which races Host::prepare against the manager's root_cancel. prepare is where the adapter spawns the provider process or creates or wakes the Docker microVM. If a managed workspace has initial repositories and preparation fails, the room row is persisted as a non-routable recovery handle instead of being discarded.
  5. Commit. Inside with_worker_runtime_settlement (the worker's settlement gate, then Inner), with no .await between the checks and the writes: confirm the worker is still registered, save_peer, start the work watcher, and call Worker::bind_live, which retires any displaced host's event ingress, inserts the new LiveHost, and binds the engine session while the durable queues are re-homed under the engine's registry write lock. Then commit_room_rebind clears any retirement fence, the provisional record is deleted, startup event ingress is activated, and backlog is redelivered.

A commit failure tears down the prepared host. If teardown also fails, the host and its provisional record are pushed onto Manager::pending_host_cleanups, which the host monitor's retry_pending_removals and Manager::close retry.

A room-bound owned session has no single state field. Its state is the combination of the persisted row (present, parked, or removed), the worker's LiveHost entry for its SessionId (present or absent), the host's runtime_pid and runtime_transport_active, the durable provisional record, the membership state on Docker hosts, and the retirement fences in Inner. session_is_dormant and session_runtime_crashed in manager.rs derive two of these states on demand; SessionPresence in crates/jam-domain/src/status.rs is the projection clients see (Live, Disconnected, Dormant, Unknown).

The states one room's owned runtime moves through from materialization to removal.

Transition Detail
Start as Dormant Dormant registration, or a session restored as dormant at worker spawn
Start as Preparing Room added with a backlog, or eager mode
Dormant to Preparing The first inbound message calls wake_dormant_room
Preparing to Live prepare succeeds, the worker commits, and the route is bound
Preparing to Dormant prepare or commit failed, and the host was torn down
Preparing to RecoveryHandle An initial repository clone failed
Preparing to Blocked The daemon crashed with a provisional record outstanding
Live to Dormant Idle reap, host stop, or room restart teardown
Live to Crashed The child exited while its PID is still recorded
Crashed to Preparing The next inbound message retires the stale host
Live to Replacing A template apply or restart_session
Replacing to Live The replacement bind committed, or it failed and the old configuration was re-bound
Live or Dormant to Retiring The room was removed, or Band returned 404
Retiring to Live Teardown failed; the child is re-tracked and retried later
Retiring to Retired The row is removed and the Docker binding marked inactive
Retired to Preparing An authoritative re-add resumes the same session

Inferred: the state names are this chapter's labels, not code identifiers. Blocked is the peer-level refusal in rebuild when recover_incomplete_provisional_sessions finds a record it cannot delete; it blocks every room of that peer, not only the one being prepared. RecoveryHandle applies only to Managed workspaces with initial repositories. The Retired to Preparing edge exists only for Docker sessions, because only they keep an inactive RuntimeSessionRoomBinding.

What happens between Band adding the agent to a room and the provider child receiving the first message.

#The worker and its supervisor task

spawn_worker in worker.rs builds everything one peer needs to talk to Band and to its runtimes, and spawns one task, supervise, that runs the engine and forwards its events.

Before building anything it reconciles the durable queue directory to current routing (QueueBackend::reconcile), so a message queued while a room was unowned moves into the session that now owns it. For each routable session it skips rooms fenced as retired, reads the session queue, and, under the lazy lifecycle, skips owned sessions with an empty backlog. Those rooms restore dormant, which is what keeps a daemon restart from spawning one process per room for an agent in many rooms. For every other session it builds a RuntimeHostContext, attaches the managed context and checkpoint broker, enriches it with the room roster from Band (a 404 marks the room for retirement instead of failing the peer), calls Deps::new_host, prepares it, creates an engine Session, and redelivers backlog. A preparation failure tears down every host already prepared in this spawn (cleanup_prepared_startup_hosts) and fails the spawn, except that an initial-repository failure on a Managed workspace leaves that one session dormant for recovery.

The engine is constructed with the peer's Band identity and the account's persisted user_id (a /me/profile call is only the fallback for old accounts). The worker's cancel token is a child of the manager's root_cancel, so manager shutdown cancels every worker.

The supervise task selects over four sources:

  • Engine events. It rewrites PeerState::Connected to Degraded when the parked-readiness warning or any live host's readiness_warning is set, feeds room membership changes to the manager, fires the on_ready reconcile once per spawn, calls on_inbound for every inbound message, and publishes everything else on the fan-out.
  • Runtime lease updates, a watch channel that tells the loop which runtime event leases are active so buffered assistant output and working state follow the current binding.
  • Runtime events from owned adapters, through a bounded channel of 64, projected into room messages, activity, usage, and turn dispositions.
  • The engine run task. When it ends, the supervisor publishes the terminal state: Stopped for a clean return, Failed for an error, and Failed plus an error event for a panic. A panic in one worker never affects another peer (worker_panic_is_isolated_and_rolled_back).

The manager registers a worker with insert_registered_worker, which inserts it into Inner::workers only if the manager is not closed and the key is free. A losing duplicate has its owned hosts torn down. Startup ingresses are activated only after registration, so a buffered runtime event can never reach a room its worker never owned.

PeerState has six values (crates/jam-domain/src/peer.rs). Three writers produce them: the engine maps its Band connection state through map_conn_state in crates/jam-core/src/engine.rs (Connecting to Starting, Up to Connected, Reconnecting to Reconnecting, Down and Superseded to Stopped); the worker's supervise task rewrites Connected to Degraded when a readiness warning exists and publishes the terminal state when the engine task ends; and the manager publishes Failed through fail_peer, Stopped on a kept stop or a skipped explicitly stopped peer at boot, and Degraded when a live-bound host reports a warning. Manager::status_of then clamps any state of a peer with no registered worker to Stopped, except Failed.

The PeerState transitions a consumer can observe, and what causes each.

Stopped is reached by stop_peer, a host-monitor stop, or a clean engine exit. Failed comes from an engine error, fail_peer, or a worker panic. The supervisor retries a Failed peer after a backoff.

Inferred: the edges are drawn from the publish sites named above. There is no transition table, so any writer can publish any value; for example fail_peer can publish Failed for a peer that has no worker at all, and a restart publishes the new worker's Starting without an intervening Stopped when the old worker is removed quietly. Degraded to Connected only happens through a reconnect, because nothing clears the degraded state while the socket stays up. The readiness-timeout edge assumes the cancelled engine returns Ok, so the worker's supervise task publishes Stopped; start_and_wait itself publishes nothing on timeout, and if the supervisor task is aborted after STOP_GRACE no terminal state is published.

#Run intent and the supervisor loop

Two facts decide whether a peer should be running:

  • Inner::wants_running, an in-memory set. start_and_wait adds a peer after it reaches readiness, rebuild adds each owned peer before trying to start it, and every stop removes it.
  • A persisted stop intent, the setting peer.lifecycle.stopped.v1.<profile>/<scope>, written by stop_peer with keep = true before any teardown and cleared by an explicit ensure or restart. Inner::explicitly_stopped mirrors it for this daemon run.

Manager::spawn_supervisor runs Manager::supervise every 5 seconds (interval set in jamd). Each pass reaps finished workers and then, for each stored peer that is wanted but has no worker, checks account eligibility and sign-in (a stored OAuth refresh token counts as signed in), confirms with a Band roster read (3-second SUPERVISOR_PRESENCE_TIMEOUT) that the agent still exists, and calls try_start_peer. Failures back off per peer from SUPERVISOR_RETRY_MIN (10 seconds) doubling to SUPERVISOR_RETRY_MAX (60 seconds), and a successful start is on probation for one delay period so a runtime that exits immediately is not relaunched every tick (autonomous_supervisor_retries_back_off_and_cap, supervisor_does_not_respawn_a_crashing_worker_on_every_tick). A confirmed-absent Band agent removes the intent. The supervisor never resurrects an explicitly stopped peer (supervisor_does_not_resurrect_a_stopped_peer), and explicit start and restart commands bypass the backoff.

#The host monitor loop and idle reaping

Manager::spawn_host_monitor runs every 15 seconds (interval set in jamd) and calls three things in order:

  1. check_host_processes observes attached windows. It distinguishes a process that exited (dead_process, the only verdict allowed to remove a session row) from a session that lapsed (no heartbeat, which is announced but never tears down a binding, so a sleeping laptop keeps its rooms), and also detects owned children that crashed while their PID was still recorded (crashed_owned_sessions). Exited rows are removed and the peer is saved. If no session remains, the worker is stopped with the record kept and a non-explicit intent. Otherwise the peer is restarted so the dead rooms leave the routing registry. A peer whose sessions are all expired pull leases is stopped the same way. The loop uses try_lock on both gates so one busy peer does not stall the sweep.
  2. reap_idle_runtimes stops idle owned children. It is a no-op when JAM_OWNED_RUNTIME_LAZY=0. Under the room-bind gate it snapshots every live runtime session, then for each one, inside the worker settlement gate, skips it if the adapter's idle_reap_policy is KeepAlive, if the process is already dead (that is a crash, not idleness), if the engine reports it mid-turn, or if it was active within the threshold. The threshold is JAM_OWNED_RUNTIME_IDLE_SECS, else the setting runtime.defaults.idle_secs, else 600 seconds. Eligible sessions are detached with Worker::reap_session (host removed, engine route removed, work watcher aborted) and then torn down outside the lock. A hard teardown failure re-inserts the still-live child with Worker::reinsert_session so it is never orphaned (reaper_teardown_failure_re_tracks_the_child_instead_of_orphaning_it). A HostError::RuntimeStopped finalization failure leaves the session dormant and wakeable (reaper_finalization_failure_keeps_the_stopped_child_dormant_and_wakeable). Only the owned Claude Code adapter returns KeepAlive today (crates/jam-host/src/claudecode/owned/mod.rs).
  3. retry_pending_removals first retries pending_host_cleanups, then restores persisted pending removals and retries room removals whose teardown or durable write failed.

#Stopping and restarting at three scopes

The lifecycle has three separate scopes, and a refactor must not merge them.

Peer scope. stop_peer removes run intent, tears down every owned host (stop_all_worker_hosts_under_worker_gate), stops the worker, and either keeps the record (keep = true, publishing PeerState::Stopped) or deprovisions the Band agent and deletes the record (publishing PeerRemoved). restart_peer persists "not stopped", removes and stops the worker, and starts it again with RuntimeStartTrigger::Restart. Because the lazy spawn restores owned rooms dormant, a peer restart stops every room's child and restarts only rooms with backlog. A peer restart is a Band reconnect plus a restart of every room on that peer, including every room on a Shared host.

Session scope. restart_session restarts one room's owned runtime. If the peer has no worker it first restarts the peer (so it doubles as recovery). It then detaches and tears down the exact session with teardown_live_owned_session and checkpoint reason BeforeStop, clears stale provider attribution, and live-binds the same row. Teardown comes before prepare because the old and new adapters would own the same provider process or sandbox name. detach_session removes one row. When the peer has other sessions and the departing one is owned, it uses the same exact-session teardown instead of a peer restart, because a peer restart on Shared placement would bounce sibling rooms (shared_session_restart_and_unbind_never_restart_the_sibling_runtime). Room-scoped reset_runtime_sandbox refuses Shared placement outright (validate_session_sandbox_reset_scope).

Host scope. control_runtime_host with RuntimeHostLifecycleAction::Stop, Restart, or ResetSandbox { discard_changes } is the only operation allowed to stop or reset a Shared host. Its scope comes from the durable membership table, never from the clicked row or the live worker registry. It takes neither the room-bind gate nor a worker lifecycle gate for its validation and teardown phases: each member teardown goes through teardown_live_owned_session (worker settlement gate only), and each re-bind takes the room-bind gate through bind_session_live. control_runtime_host_inner:

  1. Verifies the host belongs to this agent and is not Archived or Removed.
  2. Resolves every membership to exactly one owned peer session, or to an inactive dormant binding owned by this agent. Any mismatch aborts before anything stops. Restart and reset also require the session's placement to match the host's; Stop does not, so a Shared-to-Dedicated edit cannot strand a host that must be parked.
  3. Records which members were live from the worker registry.
  4. For reset, builds and validates the sandbox reset plan and runs a fresh workspace inventory for every member before teardown.
  5. Tears down each live member with the action's checkpoint reason, persists each as Dormant, marks previously dormant members Dormant too, and sets the host Parked. A teardown failure persists the members already stopped and returns the error.
  6. For restart and reset, re-binds only the members that were live, then sets the host Active (or Parked if none were live).

The result lists every affected room (RuntimeHostLifecycleResult). Tests: shared_host_stop_enumerates_every_room_and_persists_all_dormant, shared_host_restart_restarts_all_live_members_without_touching_dedicated_peer, shared_host_reset_quiesces_every_member_then_destroys_once_with_exact_reason.

#Configuration transactions: desired, effective, pending

An owned runtime's configuration has three durable fields on its HostSession, described in Domain and authority: desired_config_generation (what the operator asked for), effective_config_generation (what a live runtime accepted), and pending_change (the change in flight, including a failure reason and the generation recovery should treat as authoritative). HostSession::has_outstanding_config_change is the one comparison.

save_runtime_template (route POST /v1/saveRuntimeTemplate, 15-minute daemon budget) applies a template edit:

  1. Under the room-bind gate, it validates the edit, promotes legacy room sessions that have no template into one family when they match exactly, writes the new template with desired = max(family desired) + 1, and, depending on RuntimeTemplateApplyMode, either leaves room sessions alone (NewSessionsOnly, the default) or marks them for apply (ApplyAndRestart).
  2. Under ApplyAndRestart, dormant rooms receive the desired runtime immediately, with effective = desired, because they have no running configuration to preserve. The saved parked template row itself is written with effective_config_generation = 0.
  3. Live rooms keep their current runtime and receive a pending_change with status Requested. The peer is saved and the gate is released.
  4. For each live room not mid-turn, apply_pending_runtime_configuration_at_generation re-takes the gate, records Applying, tears down the old runtime, and live-binds a candidate built by desired_runtime. On success it sets effective = desired and clears pending_change. On failure it re-binds the old configuration and records PendingChangeStatus::Failed with recovery_target set to the old effective generation.
  5. Rooms that were mid-turn are applied later: the worker's turn-boundary callback (on_runtime_turn_boundary) re-drives the pending generation it observed.

When the edit changes the host compatibility fingerprint of a Docker host, the save instead recreates that runtime host once through control_runtime_host_inner with a reset and a replacement fingerprint (host_compatibility_template_change_recreates_shared_host_once_and_preserves_sessions).

RuntimeTxnState in runtime_txn.rs is a pure, lock-free, in-memory coordinator stored in Inner::runtime_txns. It keeps per-template desired and effective generations, per-binding candidate leases, and an in-memory provisional journal, so that a late acknowledgment from a superseded save is AckOutcome::Rejected and cannot become the effective state. The manager seeds it from the durable generations on each save (restore_template) and replacement (seed_binding), so the durable HostSession fields remain the authority across restarts. Its unit tests (late_acknowledgment_from_a_superseded_save_cannot_commit, failed_replacement_keeps_the_old_configuration_effective, stale_replacement_cannot_take_the_live_binding) describe the contract.

#Startup: Manager::rebuild

jamd calls Manager::rebuild once in its startup task, after an optional lane migration, and starts the host monitor and supervisor only afterward. rebuild replays durable recovery before it starts any worker, in this order: Docker environment cleanup, runtime-host cleanup, provider-checkpoint payload deletion, provider continuity resets, managed workspace reconciliation, pending usage observations, experiment account reconciliation, and pending native work projections. It then backfills account identity for every signed-in eligible account.

For each stored peer, under the account lifecycle gate, it:

  1. Re-reads the peer and runs recover_incomplete_provisional_sessions. A provisional record with no provider session id and no accepted turn is deleted. Any other record is marked Orphaned/Failed, the matching room gets a PendingChange with status Failed and the reason "Jam restarted while this runtime was being prepared", and the peer is failed and skipped (rebuild_refuses_a_second_runtime_when_provisional_cleanup_is_unresolved).
  2. Skips peers with a persisted stop intent (publishing Stopped) and terminal identities, which wait for their integration to reconnect (rebuild_starts_owned_peers_but_never_restores_terminal_identities).
  3. Adds the peer to wants_running, checks eligibility and sign-in, and calls try_start_peer. A failure fails the peer with a warning that Jam will retry automatically; the supervisor does the retry.

Managed provider continuity is validated during the first activation. A bound provider identity with missing or incompatible recovery state fails closed before the host factory runs (rebuild_fails_closed_for_unrecoverable_managed_provider_state), and rebuild_passes_persisted_exact_checkpoint_to_fresh_runtime_host asserts the exact handoff. Consistency and recovery covers the full startup order.

#Shutdown: Manager::close and the process deadlines

Shutdown is bounded by three nested deadlines. They must stay ordered, because a backstop that fires first can kill a daemon in the middle of writing the checkpoint that exact resume depends on.

Deadline Value Owner What it bounds
Worker stop grace 5 s STOP_GRACE in worker.rs Waiting for the supervisor task, then each work watcher, after Worker::stop cancels them. The supervisor is aborted on expiry.
Codex runtime finalization 30 s RUNTIME_FINALIZATION_TIMEOUT in crates/jam-host/src/codex/app_server.rs Quiescing the provider, the 3-second child TEARDOWN_TIMEOUT, parking Docker, clearing PID publication, and one final checkpoint.
Process backstop 45 s FORCE_EXIT_GRACE in bins/jam/src/jamd.rs Everything after the shutdown token fires. Then std::process::exit(0).
Desktop release wait 45 s + 1.5 s RELEASE_WAIT from DAEMON_FORCE_EXIT in apps/desktop/src-tauri/src/daemon.rs How long the desktop waits for a stopping daemon before spawning a new one. release_wait_outlasts_the_daemon_force_exit guards it.

Other adapters use their own teardown bounds: Copilot SDK 10 seconds, ACP 8 seconds, host-native Claude Code 8 seconds, and sandboxed Claude Code DOCKER_SANDBOX_LIFECYCLE_TIMEOUT_SECS + 15 = 315 seconds.

The shutdown edge is one CancellationToken in jamd, cancelled by Ctrl-C or POST /v1/shutdown. The same edge starts the 45-second backstop and a cleanup task that joins the startup task, analytics shutdown, discovery withdrawal, and Manager::close. The local server drains in parallel. Manager::close:

  1. Sets Inner::closed (so no new worker can register) and clears task-access reservations.
  2. Cancels root_cancel before taking any gate. The code comment explains the order: an in-flight bind can hold the room-bind gate while its prepare waits on this token, so taking the gate first would deadlock (manager_close_cancels_an_inflight_live_host_prepare).
  3. Drains attached working-state publishers, takes the room-bind gate, adds every peer with a lifecycle gate to the stop set, closes CLI-only receivers, cancels Copilot bridge liveness tasks, and shuts down plan file watchers.
  4. Stops every peer concurrently with join_all. Each peer takes its lifecycle gate and loops on remove_and_stop_under_worker_gate until it succeeds: owned hosts are torn down first (each adapter runs its final checkpoint), then the worker is stopped. The loop has no retry limit; the 45-second backstop is its only bound (manager_close_finalizes_independent_runtime_hosts_concurrently).
  5. Retries pending_host_cleanups.

After the cleanup task returns, jamd drains abandoned Codex probes and runtime probe tasks, then returns from main, which drops _daemon_lifetime_lock and _instance_lock. The lifetime lock (jam_service::try_acquire_daemon_lock, an exclusive lock on locks/daemon.lock in the config directory) is held from before the socket is bound until after cleanup, so the desktop or a service manager cannot start a second jamd against the same store while the socket is already gone but checkpoints are still being written.

The order in which a daemon shuts down, with its durable points.

#The Manager struct and the Inner mutex

Manager is one struct with 104 fields. About 56 of them are their own std::sync::Mutex, tokio::sync::Mutex, or semaphore, each guarding one concern (artifact fetches, discovery, contacts cache, receivers, account lifecycle gates, and so on). The lifecycle state lives in one field, inner: Mutex<Inner>, a std::sync::Mutex reached only through Manager::lock, which recovers from poisoning with unwrap_or_else(|e| e.into_inner()).

Inner has 29 fields. The lifecycle-relevant ones are:

Field Meaning
workers: HashMap<PeerKey, Worker> The registry of running workers. The only owner of each Worker.
provisioning: HashSet<PeerKey> Peers inside ensure_peer's create path.
wants_running, explicitly_stopped, supervisor_retries Run intent and supervisor backoff, memory only.
runtime_txns: RuntimeTxnState In-memory configuration transaction coordinator.
binding_generations, next_binding_generation, pending_removals, undurable_removals, retirement_revisions, inaccessible_rooms Per-room binding generation fences that let a stale removal lose to a newer binding.
waking_rooms One in-flight wake per (peer, room).
reconciling, reconcile_pending, wait_reconcile_due Coalescing and debounce for room reconciliation.
runtime_failures, runtime_permissions, pending_questions, environment_approvals, pending_task_access Per-session live state owned by other subsystems but kept under the same lock.
oauth_sessions, account_auth_states, account_*_tasks Account state owned by the accounts subsystem.
closed Set once by close.

The rule is that no code holds the Inner guard across an .await. Code takes the lock, copies what it needs (usually an Arc<Engine> or a peer clone), drops the guard, and then awaits. Worker::reap_session returns the detached host precisely so the caller can call teardown().await outside the lock. The commit phase of a live bind is the clearest example: host.prepare(...).await is the last await, and everything under the lock after it is synchronous. Long waits use separate async gates instead of Inner:

  • runtime_room_bind_gate (tokio::sync::Mutex<()>, manager-wide) serializes room materialization, wake, removal, detach, restart, the re-bind phase of host control, template save and apply, reaping, and shutdown. manager.rs takes it at 23 sites: 22 blocking lock().await calls and one try_lock in check_host_processes.
  • worker_lifecycle_gates (one tokio::sync::Mutex per PeerKey) serializes worker teardown and startup for one peer so a replacement cannot overwrite the only tracked owner of a child. It is taken at 13 sites: 11 through lock_worker_lifecycle and 2 through try_lock_worker_lifecycle in check_host_processes.
  • account_lifecycle_gates (one per profile) serializes sign-in, sign-out, deletion, and the per-peer steps of rebuild and the supervisor.
  • The worker's runtime_settlement_gate serializes a terminal reply or acknowledgment with host retirement or replacement. with_worker_runtime_settlement takes it, then Inner, and retries if the worker was replaced in between.

The order in which the lifecycle gates are acquired.

Inferred: the order is reconstructed from the call sites of bind_session_live, restart_peer, stop_peer_with_intent, close, rebuild, and supervise, where each outer gate is taken before the inner one. No type or test enforces a global order, and not every path takes every gate: rebuild and the supervisor take the account gate and then the worker gate without the room gate, and ensure_peer's create path takes only the worker gate. close takes the room gate before per-peer lifecycle gates, matching the order above.

#Experiment gating and knobs

The backend lifecycle is shipped and ungated. Its desktop surfaces are partly behind experiments registered in crates/jam-manager/src/experiments.rs, all disabled by default:

  • runtime-setup shows the Runtime setup card that edits or disables the reusable runtime template.
  • docker-sandboxes shows Docker Sandbox creation and runtime controls. Existing sandboxed agents keep their recovery and cleanup controls when it is off.
  • sessions-section shows the Sessions view of execution processes and room bindings.

JAM_OWNED_RUNTIME_LAZY=0 and JAM_OWNED_RUNTIME_IDLE_SECS are environment knobs, not experiments. The PTY transport is legacy: build_host rejects owned PTY sessions with a message telling the operator to replace the template.

#State and ownership

Aggregate Key and scope Authority Mutation coordinator Durable representation Projections Freshness or version fence Recovery Deletion authority
Peer record PeerKey = profile/scope, one per Band agent per profile Manager Inner::provisioning for create, worker lifecycle gate for start and stop peers row plus the agent key in SecretStore PeerStatus, PeerAdded, PeerUpdated, PeerRemoved Lowercase scope normalization; create_only refusal Identity-create recovery marker replays a half-finished registration stop_peer with keep = false deprovisions the Band agent and deletes the record
Host session rows, including templates (profile, scope, id) Manager runtime_room_bind_gate, then worker lifecycle gate host_sessions (38 columns). save_peer deletes and reinserts every row of the peer in one transaction Sessions list, agent detail, HostSessionView Room binding generation per (peer, room) in Inner; retirement state in room-retirement: settings rebuild restores routing; startup reconcile binds current rooms detach_session, room removal, disable_runtime_template for parked rows only
Runtime template identity RuntimeTemplateId UUID plus the operator-facing runtime_template_id string Manager Minted once in apply_runtime_opts; copied by upsert_host_session Inside host_sessions (runtime_template_uuid, runtime_template_id) and runtime_hosts.template_id Template editor Part of host compatibility Not applicable Deleted with its template row
Runtime session identity RuntimeSessionId UUID Manager Minted in apply_runtime_opts; re-minted only for unrecoverable legacy host-native workspaces host_sessions.runtime_session_id; PK of runtime_host_sessions Logs, sessions list, recovery panels Immutable once bound Restored from the inactive binding on re-add Continuity reset mints new ones; never rewritten in place
Runtime host and membership (Docker only) RuntimeHostId; membership keyed by RuntimeSessionId Manager (RuntimeHostAllocator) Store primary key and partial unique indexes (the per-instance allocation_gate does not span calls); host controls take the room-bind gate only for re-binds runtime_hosts, runtime_host_sessions Runtime sessions table grouped by host Compatibility fingerprint; RuntimeHostState; RuntimeHostSessionState Allocation retry converges on the same record Whole-host cleanup only (see Sandbox, workspaces, and continuity)
Room binding (Docker only) RuntimeBindingId; at most one active per (agent, room) Manager ensure_managed_runtime_ownership during bind; room removal runtime_session_room_bindings with active flag Continuity evidence, host control scope Partial unique index on active = 1 Inactive row restores exact session and provider thread on re-add Never deleted by room removal; retained for host cleanup
Worker and live hosts PeerKey; live hosts by SessionId Manager Inner for registry changes; worker settlement gate for host swaps Memory only PeerStatus.running, session presence (Live, Dormant, Disconnected) Arc<Engine> identity used to detect a completed concurrent restart Rebuilt by rebuild, supervisor, wake remove_and_stop_under_worker_gate
Peer connection state PeerKey Engine and supervisor task Fan-out publish Memory only (Fanout states map) EventKind::PeerState; status_of clamps a non-running peer to Stopped unless Failed None Re-derived on reconnect Not applicable
Run intent PeerKey Operator and manager Inner wants_running in memory; stop intent in settings Stopped rows in the desktop Stop intent written before teardown rebuild re-derives wants_running for owned identities Cleared by explicit ensure or restart
Configuration generations One HostSession row Operator for desired, provider for effective runtime_room_bind_gate; RuntimeTxnState in Inner desired_config_generation, effective_config_generation, pending_change_json Settings surfaces and status Generations; AckOutcome::Rejected for stale acknowledgments pending_change survives restart with Failed and recovery_target Cleared on successful apply
Provisional provider session (profile, scope, provider identity, provider session id) Manager Written before new_host, deleted after commit provisional_sessions Failed PendingChange on the room creation_stage, cleanup_state rebuild deletes empty ones and blocks the peer for the rest Only after cleanup is confirmed

#Contracts

#Provided: Control methods

These methods on jam_contract::Control form the lifecycle surface. control-api.md lists their routes, Tauri commands, and CLI use.

Method Route Guarantee enforced in code
ensure_peer POST /v1/ensure Idempotent for a live peer. create_only never mutates an existing identity. A failed first start deprovisions the new Band agent. Waits at most 35 s for readiness.
adopt_agent POST /v1/adopt Refuses an existing local record at the same scope. Saves the record even if the worker fails to start.
stop_peer, stop_all POST /v1/stop, /v1/stopAll keep = true persists stop intent before teardown. Reports teardown failures as a conflict after stopping the worker.
restart_peer POST /v1/restart Keeps identity. Returns the current status without restarting again if another restart completed while waiting for the gate.
restart_session POST /v1/restartSession Restarts exactly one owned room runtime. Recreates a missing worker first. Refuses parked and attached rows.
detach_session, detach POST /v1/detachSession, /v1/detach Keeps the peer and Band agent. Returns SessionDetachPartialSuccess with a recovery action when the row was removed but a follow-up stop or restart failed.
control_runtime_host POST /v1/controlRuntimeHost Scope from durable membership. All-or-nothing validation before any stop. Reports every affected room.
save_runtime_template, disable_runtime_template POST /v1/saveRuntimeTemplate, /v1/disableRuntimeTemplate Desired generation always advances. Live rooms change only after a replacement commits. Disable accepts parked rows only.
reset_runtime_sandbox POST /v1/resetRuntimeSandbox Refuses Shared placement.
attach_session, invite_session POST /v1/attach, /v1/invite Live-bind into a running worker without a respawn when only the room changes.
list_sessions, status POST /v1/listSessions, /v1/status Pure reads with presence enrichment.

Timeouts are shared between client and daemon through jam_wire. The daemon's default unary timeout is 45 seconds (DEFAULT_UNARY_TIMEOUT_SECS) and a timeout returns HTTP 504 with RequestFailure::Timeout. Routes in jam_wire::RUNTIME_CHILD_ROUTES (restart, restart session, reset sandbox, attach, template save, host control, provider continuity fresh start, and others) receive RUNTIME_CHILD_TIMEOUT_SECS (645 seconds) on the daemon and 15 seconds more on the client (RUNTIME_PROBE_TIMEOUT in jam-client), so the daemon's typed timeout arrives before the client gives up. Template save receives 15 minutes on the daemon and 15 minutes plus 15 seconds on the client. A lifecycle route that newly starts a child must be added to that list, or the daemon will cancel it at 45 seconds.

#Provided: events

The subsystem publishes EventKind::PeerAdded, PeerRemoved, PeerState, PeerUpdated, Warning, and Error through Fanout. Delivery is best effort: the stream can drop events, and clients re-hydrate with status and list_sessions. See events.md. PeerState is also published for pseudo-peers named human/<profile> by the human room feed's reconnect callback, so a consumer cannot assume every PeerState key names a stored peer.

#Consumed: the host factory and the Host trait

Deps::new_host has type NewHost = Arc<dyn Fn(&Peer, &HostSession, &Path, RuntimeHostContext) -> Result<Arc<dyn Host>, String>>. jamd implements it with build_host. The manager wraps every result in QuestionGate. The lifecycle relies on these Host guarantees:

  • prepare is idempotent and may be dropped. prepare_host_until_shutdown drops it on root_cancel and then calls teardown; its doc comment states that owned adapters must reap any child or Docker command they started when prepare is dropped.
  • teardown and teardown_for(reason) are idempotent. teardown_for carries the ProviderCheckpointReason (BeforeStop, BeforeReset, and others) for the final checkpoint; the default body calls teardown. Codex defaults to BeforeStop, which is what reaper and shutdown paths get because they call plain teardown.
  • HostError::RuntimeStopped means the process stopped but finalization failed. The manager then keeps the session dormant instead of re-tracking a dead handle.
  • runtime_pid reports a live child PID and may retain the last PID after a crash so the manager can tell a crash from a reap. runtime_transport_active reports protocol liveness for SDK-owned children whose PID is hidden, and overrides PID checks when present.
  • idle_reap_policy returns Allow by default; KeepAlive exempts a runtime from the reaper.
  • readiness_warning degrades the peer to Degraded.

#Consumed: Band

Room membership comes only from Band room_added and room_removed events and Band room lists. A room is retired locally only on a room-scoped Band read that returns exactly 404 (prune_inaccessible_bound_rooms); absence from a list is never the deletion signal.

#Invariants

Rule Enforced by Known exceptions
A parked session never receives traffic. HostSession::is_parked, Peer::routable_sessions in spawn_worker None found
One room per host session and one session per room within a peer. Manager::upsert_host_session conflict check force (--steal) evicts the other session deliberately
Room-bound owned sessions carry the template's RuntimeTemplateId and their own RuntimeSessionId. Copy in upsert_host_session before apply_runtime_opts mints IDs; desired_runtime preserves the child's IDs An explicit first move of a compatibility session into a Managed Docker family adopts the template UUID (desired_runtime, entering_managed_docker)
One runtime host per runtime session; one live Shared host per agent; one active binding per agent and room. Primary key and partial unique indexes in migration 0045; RuntimeHostAllocator None found
Host-wide scope comes from durable membership, never from a room or the live registry. control_runtime_host_inner; validate_session_sandbox_reset_scope; tests shared_host_stop_enumerates_every_room_and_persists_all_dormant, shared_session_restart_and_unbind_never_restart_the_sibling_runtime None found
Room-scoped restart and unbind never use restart_peer on a peer with siblings. restart_session and detach_session use teardown_live_owned_session detach_session of an attached (not owned) session with siblings still restarts the peer
A live child is never orphaned by a failed teardown. Worker::reinsert_session; teardown_live_owned_session re-tracks when the PID is still alive; pending_host_cleanups A RuntimeStopped error, or a dead PID, retires the host instead
Provider-session creation is journaled before any child can start. upsert_provisional_session before new_host in bind_session_live_under_gate; recover_incomplete_provisional_sessions in rebuild spawn_worker does not journal the sessions it prepares during a fresh spawn
Nothing is persisted as bound unless the worker that received the binding is still registered. Commit inside with_worker_runtime_settlement, no await between check and save_peer None found
The Inner guard is never held across .await. std::sync::MutexGuard is not Send, so any Send future that holds it fails to compile (Control is Send + Sync through async_trait); clippy::await_holding_lock under CI's -D warnings Re-locking Inner while already holding it deadlocks without any await, and nothing but review prevents that; the comment in adopt_agent records one near miss
An explicitly stopped peer is not resurrected. Stop intent persisted before teardown; supervise checks wants_running; supervisor_does_not_resurrect_a_stopped_peer, stop_keep_survives_daemon_rebuild_until_explicit_restart None found
A terminal identity is never auto-started at boot. peer_should_resume_automatically; rebuild_starts_owned_peers_but_never_restores_terminal_identities None found
A configuration change is not reported live before a runtime accepted it. effective_config_generation moves only after a successful replacement bind; RuntimeTxnState::acknowledge rejects stale generations Under ApplyAndRestart, dormant rooms set effective equal to desired immediately, because no runtime can hold the old value. The saved parked template is written with effective 0
Shutdown deadlines nest: 5 s worker stop and 30 s Codex finalization inside the 45 s backstop, inside the desktop's 46.5 s wait. Constants in worker.rs, app_server.rs, jamd.rs, daemon.rs; release_wait_outlasts_the_daemon_force_exit Sandboxed Claude Code teardown is bounded at 315 s, and close retries without a limit; see Refactor notes
Only one jamd owns a config directory, including during shutdown. Two file locks held until main returns: DaemonInstanceLock (jamd.lock, acquired first in jamd.rs) and jam_service::try_acquire_daemon_lock (locks/daemon.lock, also probed by the desktop). DaemonInstanceLock::acquire also waits up to 3 s (LEGACY_LOCK_WAIT) for legacy per-peer lock files held by an older daemon None found

#Failure and recovery

Crash during agent creation. The identity-create recovery marker is written after Band registration and before save_peer. reconcile_identity_create_recovery runs at the next provision for the same scope and resolves the orphan. A blank agent id from Band is refused before any local write, and the agent is deleted by the id recovered from its key.

Crash during a room bind. Durable ownership records written before the child (host, membership, workspace, binding) survive and are reused by the next bind (killed_daemon_retries_one_canonical_managed_sandbox_without_duplicate_ownership in bins/jam/tests/jamd_managed_restart.rs). The provisional record makes rebuild refuse to start the peer automatically and records a Failed pending change explaining why, because an earlier child might still exist.

Startup failure or timeout. start_and_wait removes and stops the worker on Failed or after 35 seconds and returns StartupFailed or StartupTimeout. For rebuild and the supervisor, the peer stays wanted and is retried with backoff.

Worker crash or panic. The supervisor task publishes Failed. The next supervisor pass reaps the finished worker and restarts it after the backoff (supervisor_reaps_a_worker_that_died_after_connecting_and_restarts_it).

Provider child crash. The room keeps its row. The host monitor detects the crash (crashed_owned_sessions), announces the room offline, and publishes a peer update. The next inbound message wakes the room, and wake_dormant_room retires the stale host before re-binding.

Idle reap or teardown failure. A hard failure re-tracks the live child and retries on the next sweep. A stopped-but-unfinalized child stays dormant and wakeable.

Bind commit failure. The prepared host is torn down. Failing that, it and its provisional record are queued in pending_host_cleanups, retried by the host monitor and by close.

Room removal failure. A failed child teardown or durable write leaves the room in pending_removals (or undurable_removals), and retry_pending_removals retries the same binding generation. A newer binding generation makes the stale removal a no-op.

Configuration replacement failure. The old configuration is re-bound and the room records PendingChangeStatus::Failed with recovery_target set to the old effective generation. If restoring the old runtime also fails, both errors are returned.

Host-wide partial failure. control_runtime_host_inner persists every member it already stopped as Dormant, publishes the peer, and returns the error. On a restart that fails partway, the host is marked Active if any member restarted, else Parked.

Shutdown during a bind. root_cancel drops the in-flight prepare, the adapter reaps what it started, and the next daemon reuses the same ownership graph (manager_close_cancels_an_inflight_live_host_prepare, graceful_shutdown_cancels_inflight_managed_provision_and_retries_one_owner in bins/jam/tests/jamd_managed_restart.rs).

Daemon restart race. The daemon instance and lifetime locks prevent two daemons from sharing a config directory, and the desktop waits longer than the backstop before spawning a replacement.

#Extension points

#Adding a harness that the lifecycle can run

The lifecycle is provider-neutral once a Host exists, but a new transport still has to be registered in several exhaustive matches before the manager will create, persist, fingerprint, and build it. Extension seams lists the full cross-crate inventory from the OpenCode case study. The execution-specific steps are:

  1. Domain. Add a HostTransport variant with its serde name, add the string to the hand-written impl Deserialize for HostTransport in host_session.rs, add a transport_catalog! row (label, command, provider, auth modes, settings, and sandbox capability cells), add an is_owned_<name> predicate on HostSession, and, for Docker support, add a SandboxRuntimeProfile and extend the match in HostRuntimeConfig::resolved_sandbox_capability_cell. HostTransport::Unknown appears 15 times in host_session.rs.
  2. Manager. Extend the exhaustive matches in apply_runtime_opts, validate_owned_runtime_configuration, default_owned_runtime_provider, clear_unsupported_thread_settings_for_transport, runtime_transport_name, and the string parser parse_runtime_transport in manager.rs; transport_tag in runtime_host.rs, which feeds the compatibility fingerprint; and agent_runtime_configuration_for in analytics.rs. If the runtime enforces room authority at spawn, add it to validate_runtime_room_authority in worker.rs, which currently branches only on ClaudeCodeCli.
  3. Store. Extend transport_str and parse_transport in crates/jam-store/src/sqlite.rs. The persist_enum fallback that preserves unknown strings is guarded by undecodable_enum_survives_a_load_and_save_round_trip.
  4. Composition. Add a branch to build_host in bins/jam/src/jamd.rs in the correct position. The chain is ordered: PTY rejection, owned Claude Code, owned ACP or OpenCode, owned Codex, owned Copilot SDK, attached Copilot by provider, attached Claude Code by peer.host, then Generic. An omitted or misplaced branch falls through silently to an attached adapter. If the harness supports Shared placement, construct a registry around SharedHostRegistry in main and pass it into build_host.
  5. Adapter. Implement Host with the lifecycle guarantees in Contracts: cancellation-safe prepare, idempotent teardown and teardown_for, runtime_pid or runtime_transport_active so crash detection and dormancy work, and idle_reap_policy if the runtime must not be reaped. A harness that reports neither PID nor transport liveness is always classified as dormant by session_is_dormant once it has no host, and its crashes are invisible to the host monitor.
  6. Work capture. If the harness has a native task file, add an arm to the new_work_source closure in jamd.

#Adding a lifecycle operation

A new Control method is a full vertical change, described in AGENTS.md under "Add a Tauri command". Two execution-specific additions apply:

  • If the operation starts or replaces a provider child, add its route to jam_wire::RUNTIME_CHILD_ROUTES. Both jam-daemon::unary_timeout_for_route and jam_client::Client::runtime_operation_timeout read that list.
  • If it acts on a runtime host, extend RuntimeHostLifecycleAction in jam-contract (three variants today) and handle it in runtime_host_lifecycle_classes and the 18 references in control_runtime_host and control_runtime_host_inner, plus the three in bins/jam/src/main.rs and the desktop's controlRuntimeHost in App.tsx.

A new path that makes a room live should call bind_session_live_under_gate rather than build hosts itself, so it inherits journaling, ownership allocation, the commit fence, and cleanup. A new path that stops one room should use teardown_live_owned_session.

#Refactor notes

The regions of manager.rs, split at its impl blocks and section comments, sized by production lines. Two of the four largest regions are impl Manager blocks with no section comment, so the file's own structure does not say what they hold.

Regions of manager.rs by production lines

Files by size and change rate over the 90 days before the snapshot. manager.rs sits alone in the top right: it is both the largest file and the one changed most often.

Churn against size

File and function size. manager.rs is 58,145 lines with three inherent impl Manager blocks containing 266, 489, and 20 methods, plus a 254-method impl Control for Manager that mostly delegates. Inline #[cfg(test)] modules add thousands more lines. worker.rs is 10,609 lines. The lifecycle functions are individually large: bind_session_live_under_gate is about 700 lines, save_runtime_template about 520, control_runtime_host_inner about 480, spawn_worker about 420, supervise (worker) about 360, rebuild about 265, and ensure_peer about 250. The integration suite crates/jam-manager/tests/lifecycle.rs is 68,007 lines with 900 tests; it is the main safety net for any restructuring.

One lock for many subsystems. Inner holds lifecycle state together with permissions, questions, task access, OAuth sessions, and account background tasks. Manager has 104 fields, about 56 of them independent locks. The lifecycle is therefore coupled to unrelated subsystems through one mutex and one struct, and the no-reentry rule (never call a method that locks Inner while holding it) applies across all of them.

Two representations of one room binding. For Docker sessions, a room binding exists both as a HostSession row (on the peer, rewritten wholesale by save_peer) and as runtime_host_sessions plus runtime_session_room_bindings rows. They are written at different points of the bind, and room removal must update both (handle_room_removed_if_generation marks the membership Dormant, deactivates the binding, then removes the row). Host-native owned sessions have only the HostSession row. A refactor that unifies them must preserve the inactive-binding tombstone that makes authoritative re-add resume the exact provider thread.

Wholesale peer writes. save_peer_with_connectivity is called at 49 non-test sites in manager.rs, and each call deletes and reinserts every host session row for the peer. worker.peer is also overwritten at 7 sites to keep the worker's snapshot aligned. Any path that saves a stale peer clone can overwrite a concurrent change to another row; the gates are what prevent that today.

Placement defaults disagree. The Rust Default for RuntimeHostPlacement is SharedAgentHost, from_existing(None) is DedicatedSessionHost, the desktop create panel defaults to shared_agent_host (LocalAgentCreatePanel.tsx), the desktop edit path maps a missing value to dedicated_session_host (runtimeDraft.ts), and the CLI (sandbox_runtime_intent in bins/jam/src/main.rs) chooses Shared only for Codex and refuses --placement shared for any other transport, although the domain capability cells declare Shared for Copilot, Claude Code, and Cursor ACP.

RuntimeTxnState duplicates durable state. The durable HostSession generations and pending_change are the authority across restarts. RuntimeTxnState is in memory, seeded from them only on template save and replacement, and its provisional journal overlaps the durable provisional_sessions table. Its module comment still says durable persistence "belongs to the store schema, which is U1's unit", but that table now exists and is written by the bind path. PendingChangeStatus::Applied is declared and rendered by the CLI but never written by the manager; a successful apply clears pending_change instead.

Concrete adapter knowledge in the manager. manager.rs and worker.rs name jam_host::claudecode::owned 30 times, jam_host::copilot::sdk 11 times, jam_host::codex::app_server 5 times, and jam_host::opencode several times. Deps::new_host has five call sites, four in manager.rs and one in spawn_worker. Three of the manager's (teardown_hosts, legacy_parked_session_readiness_warning, prompt_unbound_rooms) build a throwaway host with RuntimeHostContext::empty() only to call teardown, read a readiness warning, or push a prompt. That is a boundary exception to "providers translate only at the jam-host edge".

Hard-coded network profile at allocation. ensure_managed_runtime_ownership computes the allocation fingerprint with RuntimeNetworkProfile::Developer and a comment deferring the resolved profile to a later phase, while save_runtime_template and apply_pending_runtime_configuration_at_generation compute the fingerprint with resolved_requested_network_profile. Inferred: for a Production-profile Docker session whose host was recreated with a Production fingerprint, the next bind's allocation request would carry a Developer fingerprint and fail ensure_session_compatible. No test found exercises that sequence.

Shutdown bounds are not uniformly enforced. close loops on remove_and_stop_under_worker_gate without a retry limit, and a peer's owned hosts are torn down one after another inside that peer's task. Inferred: a peer with two or more live Dedicated Codex rooms can need more than one 30-second finalization window, and a sandboxed Claude Code room can need up to 315 seconds, so the 45-second backstop can end the process before every final checkpoint is written. The concurrency test covers independent peers, not multiple hosts under one peer.

Stale comments about a per-peer file lock. The doc comments on rebuild (above recover_incomplete_provisional_sessions), try_start_peer, and Manager::supervise, the comment above spawn_supervisor in jamd.rs, and the comment above FORCE_EXIT_GRACE in jamd.rs describe a per-peer flock held by an exiting predecessor. Commit 533221ff2 removed crates/jam-manager/src/lock.rs and PeerLock, and added DaemonInstanceLock (jamd.lock) with a bounded 3-second wait for legacy per-peer lock files that an older daemon may still hold. The jam_service daemon lifetime lock predates that commit. No current daemon holds per-peer locks, and try_start_peer now returns Ok(false) only when the manager is closed or the worker already exists. The test name create_only_maps_a_cross_manager_peer_lock_to_the_typed_local_conflict also refers to that lock. The conflict it asserts now comes from the second manager's ensure_peer finding the first manager's stored peer record through get_peer and returning already_provisioned_for_create_only.

Order-dependent host dispatch. build_host chooses an adapter through an ordered if chain over is_owned_* predicates, then session.provider, then peer.host. Attached delivery depends on fields other than the transport, so a registry keyed only by (HostRuntime, HostTransport) would not cover it.

Reused peer-state vocabulary. The human room feed publishes PeerState::Connected for the pseudo-peer human/<profile>. Consumers that assume a PeerState event names a stored peer will mis-handle it.

Scroll to zoom, drag to pan.