#Execution and lifecycle
On this page
Where it livesHow it worksThe objects and how many of each existPlacement: Shared and DedicatedCreating an owned agentBinding a room: materialization, dormancy, and wakeThe worker and its supervisor taskRun intent and the supervisor loopThe host monitor loop and idle reapingStopping and restarting at three scopesConfiguration transactions: desired, effective, pendingStartup: Manager::rebuildShutdown: Manager::close and the process deadlinesThe Manager struct and the Inner mutexExperiment gating and knobsState and ownershipContractsProvided: Control methodsProvided: eventsConsumed: the host factory and the Host traitConsumed: BandInvariantsFailure and recoveryExtension pointsAdding a harness that the lifecycle can runAdding a lifecycle operationRefactor notesThis subsystem decides when a coding agent runs, where it runs, and under which configuration. It keeps one durable record per local agent identity (the Peer), one routing row per room (the HostSession), and, for Docker-backed runtimes, durable ownership records for the execution boundary (the runtime host) and each room's runtime context. It starts and supervises one Band worker per peer, materializes room-bound runtimes from a parked template when the agent joins a room, stops idle runtimes and wakes them on the next message, and applies configuration changes without claiming a change is live before a runtime accepted it.
It does not implement any provider protocol (that is jam-host, covered in Providers and attached agents), does not route or queue individual messages (see Messaging and room state), and does not own Docker sandbox, workspace, or checkpoint mechanics (see Sandbox, workspaces, and continuity). It never infers host-wide scope from a room, never routes a room to a guessed runtime, and never lets the desktop own a process.
#Where it lives
| Area | Crate and module | Key types and functions |
|---|---|---|
| Identity record | crates/jam-domain/src/peer.rs |
Peer, IdentityClass (Terminal, Owned), PeerState, Peer::id, Peer::routable_sessions, Peer::resolved_identity_class |
| Room binding row and runtime config | crates/jam-domain/src/host_session.rs (8,330 lines) |
HostSession, HostSession::is_parked, is_owned_runtime, is_owned_codex_app_server and siblings, has_outstanding_config_change, HostRuntime, HostTransport, HostRuntimeConfig, HostRuntimeConfig::resolved_host_placement, resolved_sandbox_capability_cell, the transport_catalog! macro |
| Identifiers | crates/jam-domain/src/ids.rs |
PeerKey, IdentityId, ChatId, SessionId, RuntimeTemplateId, RuntimeSessionId, RuntimeHostId, RuntimeBindingId, uuid_v4_from_bytes |
| Runtime host records | crates/jam-domain/src/runtime_workspace.rs |
RuntimeHostPlacement (SharedAgentHost, DedicatedSessionHost), RuntimeHostRecord, RuntimeHostState, RuntimeHostSessionRecord, RuntimeHostSessionState, RuntimeSessionRoomBinding, ProviderCheckpointReason |
| Configuration transition record | crates/jam-domain/src/host_runtime_catalog.rs |
PendingChange, PendingChangeStatus, PendingChangeAction, ProvisionalSessionRecord and its stage and cleanup enums |
| Manager and its state | crates/jam-manager/src/manager.rs (58,145 lines) |
Manager, Inner, Manager::lock, ensure_peer, start_and_wait, try_start_peer, rebuild, supervise, spawn_supervisor, spawn_host_monitor, check_host_processes, reap_idle_runtimes, wake_dormant_room, auto_bind_owned_runtime_room, bind_session_live_under_gate, restart_peer, restart_session, detach_session, control_runtime_host, save_runtime_template, close |
| Worker | crates/jam-manager/src/worker.rs (10,609 lines) |
Worker, LiveHost, Reaped, spawn_worker, supervise (the per-peer task), Worker::bind_live, reap_session, retire_session, reinsert_session, prepare_host_until_shutdown, managed_runtime_context_for_session, ensure_managed_runtime_ownership, owned_runtime_lazy_enabled, STOP_GRACE |
| Runtime host allocation | crates/jam-manager/src/runtime_host.rs |
RuntimeHostAllocator::allocate, RuntimeHostCompatibility::from_runtime, fingerprint, incompatibility_reason |
| Configuration transactions | crates/jam-manager/src/runtime_txn.rs |
RuntimeTxnState, RuntimeConfigGeneration, PendingReason, AckOutcome, ActiveRuntimeBinding, ProvisionalSession |
| Composition root | bins/jam/src/jamd.rs |
main, build_host, FORCE_EXIT_GRACE, the new_host closure, CodexAppServerRegistry, CopilotSdkRegistry construction |
| Provider seam used here | crates/jam-host/src/lib.rs |
Host trait (prepare, teardown, teardown_for, runtime_pid, runtime_transport_active, idle_reap_policy, readiness_warning), IdleReapPolicy, RuntimeHostContext |
| Process-local Shared registries | crates/jam-host/src/codex/shared_dispatch.rs, codex/app_server.rs, copilot/sdk.rs |
SharedHostRegistry, CodexAppServerRegistry::build, CopilotSdkRegistry::build |
| Store tables | crates/jam-store/src/sqlite.rs, crates/jam-store/src/migrations/0045_runtime_host_workspace_foundation.sql |
peers, host_sessions (38 columns), provisional_sessions, runtime_hosts, runtime_host_sessions, runtime_session_room_bindings, settings |
The dependency direction matches the intended layering at crate level (jam-domain to jam-contract to jam-manager to jam-daemon, composed in bins/jam), but not at module level: jam-manager depends on jam-host with the runtime-providers feature and names concrete adapter modules. See Refactor notes.
The components that drive an agent's lifecycle, and who calls whom.
#How it works
#The objects and how many of each exist
Six objects participate in running one agent, and most lifecycle bugs come from treating two of them as one. Domain and authority names every identifier. Here is how the lifecycle uses them.
A peer is the local record of one Band agent under one account profile. Peer is keyed by PeerKey, the string profile/scope returned by Peer::id. It holds the Band agent_id, the agent API key (a SecretKey, stored in the OS secret store by SqliteStore::save_peer), the adapter name in host, the operator briefing, and a list of host sessions. identity_class records whether the identity is an interactive terminal window (Terminal) or a daemon-owned runtime (Owned). The field is stamped at onboard because jam detach can clear every session row, so rows alone cannot tell an idle terminal identity from an owned agent. Legacy records without the stamp resolve through Peer::resolved_identity_class: any Owned row means Owned, otherwise Terminal.
A host session is one row in Peer::host_sessions. Each row binds at most one room. Its id is a SessionId, an operator-facing alias such as default or default-<room>, and it is never an identity key. Its runtime field is a HostRuntimeConfig that says how the agent runs: mode (HostRuntime::AttachInbox for a user-owned process, Owned for a jamd-owned process), transport (AttachInbox, Acp, CodexAppServer, CopilotSdk, ClaudeCodeCli, Opencode, the legacy Pty marker, or Unknown), spawn settings, sandbox settings, placement, workspace source, network intent, and the two immutable UUIDs template_id and session_id.
A runtime template is a host session whose room is empty. HostSession::is_parked is literally self.room.as_str().is_empty(). The worker builds no routing surface for a parked row (Peer::routable_sessions filters it out), so a template can never receive a message. For an owned agent the parked row is the configuration that room membership copies. For an attached agent a parked row holds a waiting window's PID and mailbox until the window is bound to a room.
A room-bound session is a host session with a room. For owned runtimes it is materialized from the template: owned_runtime_room_bind_opts in manager.rs copies the template's settings, names the new row <template id>-<room id>, sets runtime_template_id to the template's operator-facing id, and clears provider attribution. Manager::upsert_host_session then copies the template's immutable RuntimeTemplateId UUID into the new row before apply_runtime_opts mints a fresh RuntimeSessionId. The comment on that copy states why: without it every room would become a different template and could never share a compatible host.
A runtime session (RuntimeSessionId) is the UUID in HostRuntimeConfig::session_id. It names one local runtime context and stays stable across renames and restarts. apply_runtime_opts mints it and RuntimeTemplateId with getrandom when an owned runtime has none (mint_runtime_session_id, mint_runtime_template_id).
A runtime host (RuntimeHostId) is one concrete execution boundary: one provider process and, for Docker, one microVM. Runtime host records exist only for Docker-sandboxed owned sessions. managed_runtime_context in worker.rs returns None when !session.is_owned_runtime() || !session.runtime.spawn.sandbox.enabled, and that function is the only production caller of ensure_managed_runtime_ownership, which calls RuntimeHostAllocator::allocate and writes the room binding. A host-native owned Codex session therefore has a RuntimeSessionId but no RuntimeHostRecord, no runtime_host_sessions row, and no RuntimeSessionRoomBinding.
A room binding (RuntimeBindingId, record RuntimeSessionRoomBinding) ties one runtime session to one room for one agent, with an active flag. Room removal sets active = false instead of deleting the row. That inactive row is a tombstone: an authoritative re-add of the same room (auto_bind_owned_runtime_room with authoritative_readd) looks it up with runtime_room_binding_including_inactive, reuses the exact runtime session, and passes the saved provider session id so the provider resumes the same conversation.
The records that exist for one agent, and how many of each.
The cardinalities come from code and schema:
- One room per host session, one session per room within a peer.
Manager::upsert_host_sessionrefuses a bind whose room another session owns unlessforce(--steal) is set, and refuses a second session on the same attached PID unlessallow_pid_shareis set. - One runtime host per runtime session.
runtime_host_sessions.runtime_session_idis the primary key in migration0045_runtime_host_workspace_foundation.sql. - At most one live Shared host per agent. The partial unique index
one_live_shared_runtime_host_per_agentcoversplacement = 'shared_agent_host' AND state NOT IN ('archived', 'removed'). - At most one active binding per agent and room. The partial unique index
one_active_runtime_binding_per_agent_roomcoversactive = 1. - A process is not a session. A Shared runtime host serves several runtime sessions from one provider process, and an attached window's PID can back sessions in several rooms when the reconcile path sets
allow_pid_share.
#Placement: Shared and Dedicated
RuntimeHostPlacement has two values. SharedAgentHost reuses one compatible runtime host for every room of one agent identity. DedicatedSessionHost gives one runtime session its own host. The Rust Default is SharedAgentHost, but RuntimeHostPlacement::from_existing(None) resolves to DedicatedSessionHost, because records written before placement existed ran one microVM per session and must not be silently merged. HostRuntimeConfig::resolved_host_placement is the one accessor that applies that rule.
RuntimeHostAllocator::allocate in runtime_host.rs makes the decision while holding the allocator's allocation_gate mutex. That gate belongs to one allocator instance, and the only production caller, ensure_managed_runtime_ownership, constructs a new RuntimeHostAllocator on every call, so the gate does not serialize separate production allocations. The store's primary key and partial unique indexes prevent duplicate owners. The live-bind path also holds the room-bind gate, but the spawn path (spawn_worker through managed_runtime_context_for_session) does not. The allocator decides in this order:
- If the runtime session already has a membership, it reuses that host after checking compatibility (
ensure_session_compatible). This is what makes a retry after a crash converge on the same record. - For Dedicated, the host id is the runtime session UUID itself (
RuntimeHostId::new(request.runtime_session_id.as_str())), so a crash between host insert and membership attach heals the same record. - For Shared, it looks for the agent's live Shared host. If one exists it must pass
ensure_shared_compatible. Otherwise it mints a random UUID and inserts. If the insert loses a race on the unique index, it reloads the winner and converges when compatible.
Compatibility is an opaque SHA-256 digest (RuntimeHostCompatibility::fingerprint, domain tag jam.runtime-host-compatibility.v1) over the host-wide inputs: runtime mode, transport, command, arguments, sorted environment allowlist, auth mode, runtime version, the sandbox-enabled flag, sandbox agent, kits, static MCP servers, reference folders, Docker configuration source, GitHub access, Codex control channel, network profile, developer destinations, and the local platform bridge flag. Session-local values (working directory, workspace identity, thread settings) are excluded so rooms can differ in those and still share a host. incompatibility_reason also compares agent UUID, template UUID, placement, provider, sandbox name, and terminal state. The allocator tests cover the rules by name, for example shared_placement_reuses_one_compatible_host_for_two_sessions, dedicated_placement_is_stable_and_isolated_per_session, incompatible_live_shared_host_is_rejected_without_second_owner, and same_template_with_different_network_profile_cannot_share_a_host.
Durable allocation is only half of Shared placement. The other half is process-local: jamd creates one CodexAppServerRegistry and one CopilotSdkRegistry, each wrapping a SharedHostRegistry keyed by RuntimeHostId with weak entries. When build_host builds a Codex or Copilot facade whose managed context says SharedAgentHost, the registry returns the existing live runner for that host UUID or initializes one, and fails closed if the new facade's host-level inputs differ. Each room still gets its own facade, provider thread, working directory, and event stream.
Which transports may use Shared is declared in the transport catalog's capability cells (CODEX_SANDBOX_CAPABILITY_CELLS, COPILOT_SANDBOX_CAPABILITY_CELLS, CLAUDE_SANDBOX_CAPABILITY_CELLS, CURSOR_SANDBOX_CAPABILITY_CELLS in host_session.rs). Each of those lists one Shared cell, for a Jam-managed configuration with a Managed workspace. HostRuntimeConfig::resolved_sandbox_capability_cell refuses any combination without a cell.
#Creating an owned agent
Manager::ensure_peer is the single create-or-start operation. The desktop's createLocalAgent in apps/desktop/src/App.tsx calls it through ensureOnboardedAgent with runtime_mode: "owned", an empty room, and the draft's transport and sandbox settings (draftToEnsureOpts in apps/desktop/src/lib/runtimeDraft.ts). The call reaches the daemon at POST /v1/ensure. The client allows it ENSURE_REQUEST_TIMEOUT (60 seconds); the daemon applies its default 45-second unary timeout because /v1/ensure is not in RUNTIME_CHILD_ROUTES.
ensure_peer does the following:
- Validates network intent against the operator gate, derives or normalizes the scope to lowercase (reusing a legacy cased record through
lookup_peer_ci), and reaps any worker whose supervisor already exited (reap_finished_workers). - If a worker is live, it merges the request into the existing record (
upsert_ensure_host_session) and either returns, persists a changed parked row, or live-binds a changed room-bound row.create_onlyrefuses here so a new-agent form can never mutate an existing identity. - Otherwise it inserts the peer key into
Inner::provisioning, which makes a concurrent ensure for the same key returnManagerError::Provisioning. - For a stored peer it clears the persisted stop intent and calls
start_and_wait. For a new peer it callsprovision, which registers the Band agent through the Human API, persists an identity-create recovery marker (identity-create-recovery:v1:setting) before any local write, saves the peer, and clears the marker. A failure after registration deletes the Band agent (rollback_registered_identity). start_and_waittakes the per-peer worker lifecycle gate, runs GitHub preflights and managed-workspace preparation, callsspawn_worker, registers the worker, and waits up toREADINESS_TIMEOUT(35 seconds) forPeerState::ConnectedorDegradedon the fan-out. On success it adds the key towants_running. OnFailed, a closed channel, or timeout it removes and stops the worker and returnsStartupFailedorStartupTimeout. A brand-new peer whose start fails is deprovisioned.- On success for a new peer, the manager publishes
EventKind::PeerAddedand an analytics event.
A new owned agent normally has only a parked template, so spawn_worker builds an engine and Band subscription but no host and no provider child. The agent is online on Band and runs nothing until it joins a room.
What happens when the desktop creates an owned agent, and which writes become durable.
#Binding a room: materialization, dormancy, and wake
Room membership is Band's decision. The desktop adds an agent to a room with ipc.addRoomParticipant (a Human API call through Manager::add_room_participant), which does not bind anything locally. Band then sends room_added on the agent's WebSocket. The engine turns it into EventKind::ChatAdded, the worker's supervise task maps it to a RoomMembershipUpdate (room_membership_update in worker.rs), and the manager's room_membership_sink spawns handle_room_membership_update. After the CLI-pairing disposition, that calls auto_bind_owned_runtime_room and then the external-window reconcile.
The same binder also runs level-triggered. On every worker spawn, the first Connected or Degraded state fires the one-shot on_ready hook, whose on_ready_reconcile lists the agent's Band rooms once and calls auto_bind_owned_runtime_rooms. This is how a daemon restart binds rooms joined while it was down. AGENTS.md states the requirement as "Startup/reconcile must bind current rooms, not just future events."
auto_bind_owned_runtime_room runs under runtime_room_bind_gate, re-resolves the peer from the store, and skips rooms that are retiring or known inaccessible unless the event is an authoritative re-add. It picks an anchor: the exact inactive binding's parked session when one exists, otherwise owned_runtime_anchor, which prefers the first owned parked template. Then it chooses between two outcomes:
- Dormant registration. When the lazy lifecycle is on (the default,
JAM_OWNED_RUNTIME_LAZYunset or not0) and the unrouted queue holds no message for that room (room_has_backlog),register_dormant_runtime_roompersists the room-bound row and updates the worker's peer snapshot but builds no host. It then re-checks the backlog under the still-held gate, so a message that arrived during registration triggers an immediate bind. - Live bind. Otherwise it calls
bind_session_live_under_gate.
A dormant room is woken by its first message. The worker's supervise calls the on_inbound hook for every inbound message. on_inbound_wake checks is_dormant_owned_runtime_room, reserves a per-room slot in Inner::waking_rooms (an RAII WakeGuard clears it even on panic, so a burst coalesces to one wake), and spawns wake_dormant_room. That function takes the room-bind gate, re-resolves the peer, confirms the room is dormant or crashed (session_is_dormant, session_runtime_crashed), retires a crashed runtime's stale host first, and then calls bind_session_live_under_gate with RuntimeStartTrigger::Message and OnlineCause::WakeResume. The live bind reconciles the unrouted queue, so the buffered message is re-homed to the session and redelivered.
bind_session_live_under_gate is the single path that turns a room-bound row into a live runtime. It is also used by explicit attach, invite, restart_session, host restart, configuration replacement, and wake. It is about 700 lines long and has five phases:
- Plan in memory. Take the per-peer worker lifecycle gate. Resolve which session the bind lands on (
bind_session_id) before mutating anything. Compute the new host-session set withupsert_host_session. Nothing is persisted yet. - Resolve ownership. Prepare legacy host-native workspaces. For Docker-sandboxed owned sessions,
managed_runtime_context_for_sessionallocates or reuses the runtime host and membership, ensures the managed workspace, and upserts the active room binding. These records are durable before any provider starts. If this is a fresh allocation and a later step fails,rollback_fresh_managed_runtime_failureremoves only the freshly created ownership. - Journal. Build the
RuntimeHostContext(room roster, tools, brokers, checkpoint broker). If the session is already live on the same id, tear it down first so two children never own the same provider state. Write aProvisionalSessionRecordwith stageSpawningtoprovisional_sessionsbefore calling the host factory. - Prepare. Call
Deps::new_host(thebuild_hostclosure injamd) andprepare_host_and_initial_repositories_until_shutdown, which racesHost::prepareagainst the manager'sroot_cancel.prepareis where the adapter spawns the provider process or creates or wakes the Docker microVM. If a managed workspace has initial repositories and preparation fails, the room row is persisted as a non-routable recovery handle instead of being discarded. - Commit. Inside
with_worker_runtime_settlement(the worker's settlement gate, thenInner), with no.awaitbetween the checks and the writes: confirm the worker is still registered,save_peer, start the work watcher, and callWorker::bind_live, which retires any displaced host's event ingress, inserts the newLiveHost, and binds the engine session while the durable queues are re-homed under the engine's registry write lock. Thencommit_room_rebindclears any retirement fence, the provisional record is deleted, startup event ingress is activated, and backlog is redelivered.
A commit failure tears down the prepared host. If teardown also fails, the host and its provisional record are pushed onto Manager::pending_host_cleanups, which the host monitor's retry_pending_removals and Manager::close retry.
A room-bound owned session has no single state field. Its state is the combination of the persisted row (present, parked, or removed), the worker's LiveHost entry for its SessionId (present or absent), the host's runtime_pid and runtime_transport_active, the durable provisional record, the membership state on Docker hosts, and the retirement fences in Inner. session_is_dormant and session_runtime_crashed in manager.rs derive two of these states on demand; SessionPresence in crates/jam-domain/src/status.rs is the projection clients see (Live, Disconnected, Dormant, Unknown).
The states one room's owned runtime moves through from materialization to removal.
| Transition | Detail |
|---|---|
| Start as Dormant | Dormant registration, or a session restored as dormant at worker spawn |
| Start as Preparing | Room added with a backlog, or eager mode |
| Dormant to Preparing | The first inbound message calls wake_dormant_room |
| Preparing to Live | prepare succeeds, the worker commits, and the route is bound |
| Preparing to Dormant | prepare or commit failed, and the host was torn down |
| Preparing to RecoveryHandle | An initial repository clone failed |
| Preparing to Blocked | The daemon crashed with a provisional record outstanding |
| Live to Dormant | Idle reap, host stop, or room restart teardown |
| Live to Crashed | The child exited while its PID is still recorded |
| Crashed to Preparing | The next inbound message retires the stale host |
| Live to Replacing | A template apply or restart_session |
| Replacing to Live | The replacement bind committed, or it failed and the old configuration was re-bound |
| Live or Dormant to Retiring | The room was removed, or Band returned 404 |
| Retiring to Live | Teardown failed; the child is re-tracked and retried later |
| Retiring to Retired | The row is removed and the Docker binding marked inactive |
| Retired to Preparing | An authoritative re-add resumes the same session |
Inferred: the state names are this chapter's labels, not code identifiers. Blocked is the peer-level refusal in rebuild when recover_incomplete_provisional_sessions finds a record it cannot delete; it blocks every room of that peer, not only the one being prepared. RecoveryHandle applies only to Managed workspaces with initial repositories. The Retired to Preparing edge exists only for Docker sessions, because only they keep an inactive RuntimeSessionRoomBinding.
What happens between Band adding the agent to a room and the provider child receiving the first message.
#The worker and its supervisor task
spawn_worker in worker.rs builds everything one peer needs to talk to Band and to its runtimes, and spawns one task, supervise, that runs the engine and forwards its events.
Before building anything it reconciles the durable queue directory to current routing (QueueBackend::reconcile), so a message queued while a room was unowned moves into the session that now owns it. For each routable session it skips rooms fenced as retired, reads the session queue, and, under the lazy lifecycle, skips owned sessions with an empty backlog. Those rooms restore dormant, which is what keeps a daemon restart from spawning one process per room for an agent in many rooms. For every other session it builds a RuntimeHostContext, attaches the managed context and checkpoint broker, enriches it with the room roster from Band (a 404 marks the room for retirement instead of failing the peer), calls Deps::new_host, prepares it, creates an engine Session, and redelivers backlog. A preparation failure tears down every host already prepared in this spawn (cleanup_prepared_startup_hosts) and fails the spawn, except that an initial-repository failure on a Managed workspace leaves that one session dormant for recovery.
The engine is constructed with the peer's Band identity and the account's persisted user_id (a /me/profile call is only the fallback for old accounts). The worker's cancel token is a child of the manager's root_cancel, so manager shutdown cancels every worker.
The supervise task selects over four sources:
- Engine events. It rewrites
PeerState::ConnectedtoDegradedwhen the parked-readiness warning or any live host'sreadiness_warningis set, feeds room membership changes to the manager, fires theon_readyreconcile once per spawn, callson_inboundfor every inbound message, and publishes everything else on the fan-out. - Runtime lease updates, a watch channel that tells the loop which runtime event leases are active so buffered assistant output and working state follow the current binding.
- Runtime events from owned adapters, through a bounded channel of 64, projected into room messages, activity, usage, and turn dispositions.
- The engine run task. When it ends, the supervisor publishes the terminal state:
Stoppedfor a clean return,Failedfor an error, andFailedplus an error event for a panic. A panic in one worker never affects another peer (worker_panic_is_isolated_and_rolled_back).
The manager registers a worker with insert_registered_worker, which inserts it into Inner::workers only if the manager is not closed and the key is free. A losing duplicate has its owned hosts torn down. Startup ingresses are activated only after registration, so a buffered runtime event can never reach a room its worker never owned.
PeerState has six values (crates/jam-domain/src/peer.rs). Three writers produce them: the engine maps its Band connection state through map_conn_state in crates/jam-core/src/engine.rs (Connecting to Starting, Up to Connected, Reconnecting to Reconnecting, Down and Superseded to Stopped); the worker's supervise task rewrites Connected to Degraded when a readiness warning exists and publishes the terminal state when the engine task ends; and the manager publishes Failed through fail_peer, Stopped on a kept stop or a skipped explicitly stopped peer at boot, and Degraded when a live-bound host reports a warning. Manager::status_of then clamps any state of a peer with no registered worker to Stopped, except Failed.
The PeerState transitions a consumer can observe, and what causes each.
Stopped is reached by stop_peer, a host-monitor stop, or a clean engine exit. Failed comes from an engine error, fail_peer, or a worker panic. The supervisor retries a Failed peer after a backoff.
Inferred: the edges are drawn from the publish sites named above. There is no transition table, so any writer can publish any value; for example fail_peer can publish Failed for a peer that has no worker at all, and a restart publishes the new worker's Starting without an intervening Stopped when the old worker is removed quietly. Degraded to Connected only happens through a reconnect, because nothing clears the degraded state while the socket stays up. The readiness-timeout edge assumes the cancelled engine returns Ok, so the worker's supervise task publishes Stopped; start_and_wait itself publishes nothing on timeout, and if the supervisor task is aborted after STOP_GRACE no terminal state is published.
#Run intent and the supervisor loop
Two facts decide whether a peer should be running:
Inner::wants_running, an in-memory set.start_and_waitadds a peer after it reaches readiness,rebuildadds each owned peer before trying to start it, and every stop removes it.- A persisted stop intent, the setting
peer.lifecycle.stopped.v1.<profile>/<scope>, written bystop_peerwithkeep = truebefore any teardown and cleared by an explicit ensure or restart.Inner::explicitly_stoppedmirrors it for this daemon run.
Manager::spawn_supervisor runs Manager::supervise every 5 seconds (interval set in jamd). Each pass reaps finished workers and then, for each stored peer that is wanted but has no worker, checks account eligibility and sign-in (a stored OAuth refresh token counts as signed in), confirms with a Band roster read (3-second SUPERVISOR_PRESENCE_TIMEOUT) that the agent still exists, and calls try_start_peer. Failures back off per peer from SUPERVISOR_RETRY_MIN (10 seconds) doubling to SUPERVISOR_RETRY_MAX (60 seconds), and a successful start is on probation for one delay period so a runtime that exits immediately is not relaunched every tick (autonomous_supervisor_retries_back_off_and_cap, supervisor_does_not_respawn_a_crashing_worker_on_every_tick). A confirmed-absent Band agent removes the intent. The supervisor never resurrects an explicitly stopped peer (supervisor_does_not_resurrect_a_stopped_peer), and explicit start and restart commands bypass the backoff.
#The host monitor loop and idle reaping
Manager::spawn_host_monitor runs every 15 seconds (interval set in jamd) and calls three things in order:
check_host_processesobserves attached windows. It distinguishes a process that exited (dead_process, the only verdict allowed to remove a session row) from a session that lapsed (no heartbeat, which is announced but never tears down a binding, so a sleeping laptop keeps its rooms), and also detects owned children that crashed while their PID was still recorded (crashed_owned_sessions). Exited rows are removed and the peer is saved. If no session remains, the worker is stopped with the record kept and a non-explicit intent. Otherwise the peer is restarted so the dead rooms leave the routing registry. A peer whose sessions are all expired pull leases is stopped the same way. The loop usestry_lockon both gates so one busy peer does not stall the sweep.reap_idle_runtimesstops idle owned children. It is a no-op whenJAM_OWNED_RUNTIME_LAZY=0. Under the room-bind gate it snapshots every live runtime session, then for each one, inside the worker settlement gate, skips it if the adapter'sidle_reap_policyisKeepAlive, if the process is already dead (that is a crash, not idleness), if the engine reports it mid-turn, or if it was active within the threshold. The threshold isJAM_OWNED_RUNTIME_IDLE_SECS, else the settingruntime.defaults.idle_secs, else 600 seconds. Eligible sessions are detached withWorker::reap_session(host removed, engine route removed, work watcher aborted) and then torn down outside the lock. A hard teardown failure re-inserts the still-live child withWorker::reinsert_sessionso it is never orphaned (reaper_teardown_failure_re_tracks_the_child_instead_of_orphaning_it). AHostError::RuntimeStoppedfinalization failure leaves the session dormant and wakeable (reaper_finalization_failure_keeps_the_stopped_child_dormant_and_wakeable). Only the owned Claude Code adapter returnsKeepAlivetoday (crates/jam-host/src/claudecode/owned/mod.rs).retry_pending_removalsfirst retriespending_host_cleanups, then restores persisted pending removals and retries room removals whose teardown or durable write failed.
#Stopping and restarting at three scopes
The lifecycle has three separate scopes, and a refactor must not merge them.
Peer scope. stop_peer removes run intent, tears down every owned host (stop_all_worker_hosts_under_worker_gate), stops the worker, and either keeps the record (keep = true, publishing PeerState::Stopped) or deprovisions the Band agent and deletes the record (publishing PeerRemoved). restart_peer persists "not stopped", removes and stops the worker, and starts it again with RuntimeStartTrigger::Restart. Because the lazy spawn restores owned rooms dormant, a peer restart stops every room's child and restarts only rooms with backlog. A peer restart is a Band reconnect plus a restart of every room on that peer, including every room on a Shared host.
Session scope. restart_session restarts one room's owned runtime. If the peer has no worker it first restarts the peer (so it doubles as recovery). It then detaches and tears down the exact session with teardown_live_owned_session and checkpoint reason BeforeStop, clears stale provider attribution, and live-binds the same row. Teardown comes before prepare because the old and new adapters would own the same provider process or sandbox name. detach_session removes one row. When the peer has other sessions and the departing one is owned, it uses the same exact-session teardown instead of a peer restart, because a peer restart on Shared placement would bounce sibling rooms (shared_session_restart_and_unbind_never_restart_the_sibling_runtime). Room-scoped reset_runtime_sandbox refuses Shared placement outright (validate_session_sandbox_reset_scope).
Host scope. control_runtime_host with RuntimeHostLifecycleAction::Stop, Restart, or ResetSandbox { discard_changes } is the only operation allowed to stop or reset a Shared host. Its scope comes from the durable membership table, never from the clicked row or the live worker registry. It takes neither the room-bind gate nor a worker lifecycle gate for its validation and teardown phases: each member teardown goes through teardown_live_owned_session (worker settlement gate only), and each re-bind takes the room-bind gate through bind_session_live. control_runtime_host_inner:
- Verifies the host belongs to this agent and is not
ArchivedorRemoved. - Resolves every membership to exactly one owned peer session, or to an inactive dormant binding owned by this agent. Any mismatch aborts before anything stops. Restart and reset also require the session's placement to match the host's; Stop does not, so a Shared-to-Dedicated edit cannot strand a host that must be parked.
- Records which members were live from the worker registry.
- For reset, builds and validates the sandbox reset plan and runs a fresh workspace inventory for every member before teardown.
- Tears down each live member with the action's checkpoint reason, persists each as
Dormant, marks previously dormant membersDormanttoo, and sets the hostParked. A teardown failure persists the members already stopped and returns the error. - For restart and reset, re-binds only the members that were live, then sets the host
Active(orParkedif none were live).
The result lists every affected room (RuntimeHostLifecycleResult). Tests: shared_host_stop_enumerates_every_room_and_persists_all_dormant, shared_host_restart_restarts_all_live_members_without_touching_dedicated_peer, shared_host_reset_quiesces_every_member_then_destroys_once_with_exact_reason.
#Configuration transactions: desired, effective, pending
An owned runtime's configuration has three durable fields on its HostSession, described in Domain and authority: desired_config_generation (what the operator asked for), effective_config_generation (what a live runtime accepted), and pending_change (the change in flight, including a failure reason and the generation recovery should treat as authoritative). HostSession::has_outstanding_config_change is the one comparison.
save_runtime_template (route POST /v1/saveRuntimeTemplate, 15-minute daemon budget) applies a template edit:
- Under the room-bind gate, it validates the edit, promotes legacy room sessions that have no template into one family when they match exactly, writes the new template with
desired = max(family desired) + 1, and, depending onRuntimeTemplateApplyMode, either leaves room sessions alone (NewSessionsOnly, the default) or marks them for apply (ApplyAndRestart). - Under
ApplyAndRestart, dormant rooms receive the desired runtime immediately, witheffective = desired, because they have no running configuration to preserve. The saved parked template row itself is written witheffective_config_generation = 0. - Live rooms keep their current runtime and receive a
pending_changewith statusRequested. The peer is saved and the gate is released. - For each live room not mid-turn,
apply_pending_runtime_configuration_at_generationre-takes the gate, recordsApplying, tears down the old runtime, and live-binds a candidate built bydesired_runtime. On success it setseffective = desiredand clearspending_change. On failure it re-binds the old configuration and recordsPendingChangeStatus::Failedwithrecovery_targetset to the old effective generation. - Rooms that were mid-turn are applied later: the worker's turn-boundary callback (
on_runtime_turn_boundary) re-drives the pending generation it observed.
When the edit changes the host compatibility fingerprint of a Docker host, the save instead recreates that runtime host once through control_runtime_host_inner with a reset and a replacement fingerprint (host_compatibility_template_change_recreates_shared_host_once_and_preserves_sessions).
RuntimeTxnState in runtime_txn.rs is a pure, lock-free, in-memory coordinator stored in Inner::runtime_txns. It keeps per-template desired and effective generations, per-binding candidate leases, and an in-memory provisional journal, so that a late acknowledgment from a superseded save is AckOutcome::Rejected and cannot become the effective state. The manager seeds it from the durable generations on each save (restore_template) and replacement (seed_binding), so the durable HostSession fields remain the authority across restarts. Its unit tests (late_acknowledgment_from_a_superseded_save_cannot_commit, failed_replacement_keeps_the_old_configuration_effective, stale_replacement_cannot_take_the_live_binding) describe the contract.
#Startup: Manager::rebuild
jamd calls Manager::rebuild once in its startup task, after an optional lane migration, and starts the host monitor and supervisor only afterward. rebuild replays durable recovery before it starts any worker, in this order: Docker environment cleanup, runtime-host cleanup, provider-checkpoint payload deletion, provider continuity resets, managed workspace reconciliation, pending usage observations, experiment account reconciliation, and pending native work projections. It then backfills account identity for every signed-in eligible account.
For each stored peer, under the account lifecycle gate, it:
- Re-reads the peer and runs
recover_incomplete_provisional_sessions. A provisional record with no provider session id and no accepted turn is deleted. Any other record is markedOrphaned/Failed, the matching room gets aPendingChangewith statusFailedand the reason "Jam restarted while this runtime was being prepared", and the peer is failed and skipped (rebuild_refuses_a_second_runtime_when_provisional_cleanup_is_unresolved). - Skips peers with a persisted stop intent (publishing
Stopped) and terminal identities, which wait for their integration to reconnect (rebuild_starts_owned_peers_but_never_restores_terminal_identities). - Adds the peer to
wants_running, checks eligibility and sign-in, and callstry_start_peer. A failure fails the peer with a warning that Jam will retry automatically; the supervisor does the retry.
Managed provider continuity is validated during the first activation. A bound provider identity with missing or incompatible recovery state fails closed before the host factory runs (rebuild_fails_closed_for_unrecoverable_managed_provider_state), and rebuild_passes_persisted_exact_checkpoint_to_fresh_runtime_host asserts the exact handoff. Consistency and recovery covers the full startup order.
#Shutdown: Manager::close and the process deadlines
Shutdown is bounded by three nested deadlines. They must stay ordered, because a backstop that fires first can kill a daemon in the middle of writing the checkpoint that exact resume depends on.
| Deadline | Value | Owner | What it bounds |
|---|---|---|---|
| Worker stop grace | 5 s | STOP_GRACE in worker.rs |
Waiting for the supervisor task, then each work watcher, after Worker::stop cancels them. The supervisor is aborted on expiry. |
| Codex runtime finalization | 30 s | RUNTIME_FINALIZATION_TIMEOUT in crates/jam-host/src/codex/app_server.rs |
Quiescing the provider, the 3-second child TEARDOWN_TIMEOUT, parking Docker, clearing PID publication, and one final checkpoint. |
| Process backstop | 45 s | FORCE_EXIT_GRACE in bins/jam/src/jamd.rs |
Everything after the shutdown token fires. Then std::process::exit(0). |
| Desktop release wait | 45 s + 1.5 s | RELEASE_WAIT from DAEMON_FORCE_EXIT in apps/desktop/src-tauri/src/daemon.rs |
How long the desktop waits for a stopping daemon before spawning a new one. release_wait_outlasts_the_daemon_force_exit guards it. |
Other adapters use their own teardown bounds: Copilot SDK 10 seconds, ACP 8 seconds, host-native Claude Code 8 seconds, and sandboxed Claude Code DOCKER_SANDBOX_LIFECYCLE_TIMEOUT_SECS + 15 = 315 seconds.
The shutdown edge is one CancellationToken in jamd, cancelled by Ctrl-C or POST /v1/shutdown. The same edge starts the 45-second backstop and a cleanup task that joins the startup task, analytics shutdown, discovery withdrawal, and Manager::close. The local server drains in parallel. Manager::close:
- Sets
Inner::closed(so no new worker can register) and clears task-access reservations. - Cancels
root_cancelbefore taking any gate. The code comment explains the order: an in-flight bind can hold the room-bind gate while itspreparewaits on this token, so taking the gate first would deadlock (manager_close_cancels_an_inflight_live_host_prepare). - Drains attached working-state publishers, takes the room-bind gate, adds every peer with a lifecycle gate to the stop set, closes CLI-only receivers, cancels Copilot bridge liveness tasks, and shuts down plan file watchers.
- Stops every peer concurrently with
join_all. Each peer takes its lifecycle gate and loops onremove_and_stop_under_worker_gateuntil it succeeds: owned hosts are torn down first (each adapter runs its final checkpoint), then the worker is stopped. The loop has no retry limit; the 45-second backstop is its only bound (manager_close_finalizes_independent_runtime_hosts_concurrently). - Retries
pending_host_cleanups.
After the cleanup task returns, jamd drains abandoned Codex probes and runtime probe tasks, then returns from main, which drops _daemon_lifetime_lock and _instance_lock. The lifetime lock (jam_service::try_acquire_daemon_lock, an exclusive lock on locks/daemon.lock in the config directory) is held from before the socket is bound until after cleanup, so the desktop or a service manager cannot start a second jamd against the same store while the socket is already gone but checkpoints are still being written.
The order in which a daemon shuts down, with its durable points.
#The Manager struct and the Inner mutex
Manager is one struct with 104 fields. About 56 of them are their own std::sync::Mutex, tokio::sync::Mutex, or semaphore, each guarding one concern (artifact fetches, discovery, contacts cache, receivers, account lifecycle gates, and so on). The lifecycle state lives in one field, inner: Mutex<Inner>, a std::sync::Mutex reached only through Manager::lock, which recovers from poisoning with unwrap_or_else(|e| e.into_inner()).
Inner has 29 fields. The lifecycle-relevant ones are:
| Field | Meaning |
|---|---|
workers: HashMap<PeerKey, Worker> |
The registry of running workers. The only owner of each Worker. |
provisioning: HashSet<PeerKey> |
Peers inside ensure_peer's create path. |
wants_running, explicitly_stopped, supervisor_retries |
Run intent and supervisor backoff, memory only. |
runtime_txns: RuntimeTxnState |
In-memory configuration transaction coordinator. |
binding_generations, next_binding_generation, pending_removals, undurable_removals, retirement_revisions, inaccessible_rooms |
Per-room binding generation fences that let a stale removal lose to a newer binding. |
waking_rooms |
One in-flight wake per (peer, room). |
reconciling, reconcile_pending, wait_reconcile_due |
Coalescing and debounce for room reconciliation. |
runtime_failures, runtime_permissions, pending_questions, environment_approvals, pending_task_access |
Per-session live state owned by other subsystems but kept under the same lock. |
oauth_sessions, account_auth_states, account_*_tasks |
Account state owned by the accounts subsystem. |
closed |
Set once by close. |
The rule is that no code holds the Inner guard across an .await. Code takes the lock, copies what it needs (usually an Arc<Engine> or a peer clone), drops the guard, and then awaits. Worker::reap_session returns the detached host precisely so the caller can call teardown().await outside the lock. The commit phase of a live bind is the clearest example: host.prepare(...).await is the last await, and everything under the lock after it is synchronous. Long waits use separate async gates instead of Inner:
runtime_room_bind_gate(tokio::sync::Mutex<()>, manager-wide) serializes room materialization, wake, removal, detach, restart, the re-bind phase of host control, template save and apply, reaping, and shutdown.manager.rstakes it at 23 sites: 22 blockinglock().awaitcalls and onetry_lockincheck_host_processes.worker_lifecycle_gates(onetokio::sync::MutexperPeerKey) serializes worker teardown and startup for one peer so a replacement cannot overwrite the only tracked owner of a child. It is taken at 13 sites: 11 throughlock_worker_lifecycleand 2 throughtry_lock_worker_lifecycleincheck_host_processes.account_lifecycle_gates(one per profile) serializes sign-in, sign-out, deletion, and the per-peer steps ofrebuildand the supervisor.- The worker's
runtime_settlement_gateserializes a terminal reply or acknowledgment with host retirement or replacement.with_worker_runtime_settlementtakes it, thenInner, and retries if the worker was replaced in between.
The order in which the lifecycle gates are acquired.
Inferred: the order is reconstructed from the call sites of bind_session_live, restart_peer, stop_peer_with_intent, close, rebuild, and supervise, where each outer gate is taken before the inner one. No type or test enforces a global order, and not every path takes every gate: rebuild and the supervisor take the account gate and then the worker gate without the room gate, and ensure_peer's create path takes only the worker gate. close takes the room gate before per-peer lifecycle gates, matching the order above.
#Experiment gating and knobs
The backend lifecycle is shipped and ungated. Its desktop surfaces are partly behind experiments registered in crates/jam-manager/src/experiments.rs, all disabled by default:
runtime-setupshows the Runtime setup card that edits or disables the reusable runtime template.docker-sandboxesshows Docker Sandbox creation and runtime controls. Existing sandboxed agents keep their recovery and cleanup controls when it is off.sessions-sectionshows the Sessions view of execution processes and room bindings.
JAM_OWNED_RUNTIME_LAZY=0 and JAM_OWNED_RUNTIME_IDLE_SECS are environment knobs, not experiments. The PTY transport is legacy: build_host rejects owned PTY sessions with a message telling the operator to replace the template.
#State and ownership
| Aggregate | Key and scope | Authority | Mutation coordinator | Durable representation | Projections | Freshness or version fence | Recovery | Deletion authority |
|---|---|---|---|---|---|---|---|---|
| Peer record | PeerKey = profile/scope, one per Band agent per profile |
Manager | Inner::provisioning for create, worker lifecycle gate for start and stop |
peers row plus the agent key in SecretStore |
PeerStatus, PeerAdded, PeerUpdated, PeerRemoved |
Lowercase scope normalization; create_only refusal |
Identity-create recovery marker replays a half-finished registration | stop_peer with keep = false deprovisions the Band agent and deletes the record |
| Host session rows, including templates | (profile, scope, id) |
Manager | runtime_room_bind_gate, then worker lifecycle gate |
host_sessions (38 columns). save_peer deletes and reinserts every row of the peer in one transaction |
Sessions list, agent detail, HostSessionView |
Room binding generation per (peer, room) in Inner; retirement state in room-retirement: settings |
rebuild restores routing; startup reconcile binds current rooms |
detach_session, room removal, disable_runtime_template for parked rows only |
| Runtime template identity | RuntimeTemplateId UUID plus the operator-facing runtime_template_id string |
Manager | Minted once in apply_runtime_opts; copied by upsert_host_session |
Inside host_sessions (runtime_template_uuid, runtime_template_id) and runtime_hosts.template_id |
Template editor | Part of host compatibility | Not applicable | Deleted with its template row |
| Runtime session identity | RuntimeSessionId UUID |
Manager | Minted in apply_runtime_opts; re-minted only for unrecoverable legacy host-native workspaces |
host_sessions.runtime_session_id; PK of runtime_host_sessions |
Logs, sessions list, recovery panels | Immutable once bound | Restored from the inactive binding on re-add | Continuity reset mints new ones; never rewritten in place |
| Runtime host and membership (Docker only) | RuntimeHostId; membership keyed by RuntimeSessionId |
Manager (RuntimeHostAllocator) |
Store primary key and partial unique indexes (the per-instance allocation_gate does not span calls); host controls take the room-bind gate only for re-binds |
runtime_hosts, runtime_host_sessions |
Runtime sessions table grouped by host | Compatibility fingerprint; RuntimeHostState; RuntimeHostSessionState |
Allocation retry converges on the same record | Whole-host cleanup only (see Sandbox, workspaces, and continuity) |
| Room binding (Docker only) | RuntimeBindingId; at most one active per (agent, room) |
Manager | ensure_managed_runtime_ownership during bind; room removal |
runtime_session_room_bindings with active flag |
Continuity evidence, host control scope | Partial unique index on active = 1 |
Inactive row restores exact session and provider thread on re-add | Never deleted by room removal; retained for host cleanup |
| Worker and live hosts | PeerKey; live hosts by SessionId |
Manager | Inner for registry changes; worker settlement gate for host swaps |
Memory only | PeerStatus.running, session presence (Live, Dormant, Disconnected) |
Arc<Engine> identity used to detect a completed concurrent restart |
Rebuilt by rebuild, supervisor, wake |
remove_and_stop_under_worker_gate |
| Peer connection state | PeerKey |
Engine and supervisor task | Fan-out publish | Memory only (Fanout states map) |
EventKind::PeerState; status_of clamps a non-running peer to Stopped unless Failed |
None | Re-derived on reconnect | Not applicable |
| Run intent | PeerKey |
Operator and manager | Inner |
wants_running in memory; stop intent in settings |
Stopped rows in the desktop | Stop intent written before teardown | rebuild re-derives wants_running for owned identities |
Cleared by explicit ensure or restart |
| Configuration generations | One HostSession row |
Operator for desired, provider for effective | runtime_room_bind_gate; RuntimeTxnState in Inner |
desired_config_generation, effective_config_generation, pending_change_json |
Settings surfaces and status | Generations; AckOutcome::Rejected for stale acknowledgments |
pending_change survives restart with Failed and recovery_target |
Cleared on successful apply |
| Provisional provider session | (profile, scope, provider identity, provider session id) |
Manager | Written before new_host, deleted after commit |
provisional_sessions |
Failed PendingChange on the room |
creation_stage, cleanup_state |
rebuild deletes empty ones and blocks the peer for the rest |
Only after cleanup is confirmed |
#Contracts
#Provided: Control methods
These methods on jam_contract::Control form the lifecycle surface. control-api.md lists their routes, Tauri commands, and CLI use.
| Method | Route | Guarantee enforced in code |
|---|---|---|
ensure_peer |
POST /v1/ensure |
Idempotent for a live peer. create_only never mutates an existing identity. A failed first start deprovisions the new Band agent. Waits at most 35 s for readiness. |
adopt_agent |
POST /v1/adopt |
Refuses an existing local record at the same scope. Saves the record even if the worker fails to start. |
stop_peer, stop_all |
POST /v1/stop, /v1/stopAll |
keep = true persists stop intent before teardown. Reports teardown failures as a conflict after stopping the worker. |
restart_peer |
POST /v1/restart |
Keeps identity. Returns the current status without restarting again if another restart completed while waiting for the gate. |
restart_session |
POST /v1/restartSession |
Restarts exactly one owned room runtime. Recreates a missing worker first. Refuses parked and attached rows. |
detach_session, detach |
POST /v1/detachSession, /v1/detach |
Keeps the peer and Band agent. Returns SessionDetachPartialSuccess with a recovery action when the row was removed but a follow-up stop or restart failed. |
control_runtime_host |
POST /v1/controlRuntimeHost |
Scope from durable membership. All-or-nothing validation before any stop. Reports every affected room. |
save_runtime_template, disable_runtime_template |
POST /v1/saveRuntimeTemplate, /v1/disableRuntimeTemplate |
Desired generation always advances. Live rooms change only after a replacement commits. Disable accepts parked rows only. |
reset_runtime_sandbox |
POST /v1/resetRuntimeSandbox |
Refuses Shared placement. |
attach_session, invite_session |
POST /v1/attach, /v1/invite |
Live-bind into a running worker without a respawn when only the room changes. |
list_sessions, status |
POST /v1/listSessions, /v1/status |
Pure reads with presence enrichment. |
Timeouts are shared between client and daemon through jam_wire. The daemon's default unary timeout is 45 seconds (DEFAULT_UNARY_TIMEOUT_SECS) and a timeout returns HTTP 504 with RequestFailure::Timeout. Routes in jam_wire::RUNTIME_CHILD_ROUTES (restart, restart session, reset sandbox, attach, template save, host control, provider continuity fresh start, and others) receive RUNTIME_CHILD_TIMEOUT_SECS (645 seconds) on the daemon and 15 seconds more on the client (RUNTIME_PROBE_TIMEOUT in jam-client), so the daemon's typed timeout arrives before the client gives up. Template save receives 15 minutes on the daemon and 15 minutes plus 15 seconds on the client. A lifecycle route that newly starts a child must be added to that list, or the daemon will cancel it at 45 seconds.
#Provided: events
The subsystem publishes EventKind::PeerAdded, PeerRemoved, PeerState, PeerUpdated, Warning, and Error through Fanout. Delivery is best effort: the stream can drop events, and clients re-hydrate with status and list_sessions. See events.md. PeerState is also published for pseudo-peers named human/<profile> by the human room feed's reconnect callback, so a consumer cannot assume every PeerState key names a stored peer.
#Consumed: the host factory and the Host trait
Deps::new_host has type NewHost = Arc<dyn Fn(&Peer, &HostSession, &Path, RuntimeHostContext) -> Result<Arc<dyn Host>, String>>. jamd implements it with build_host. The manager wraps every result in QuestionGate. The lifecycle relies on these Host guarantees:
prepareis idempotent and may be dropped.prepare_host_until_shutdowndrops it onroot_canceland then callsteardown; its doc comment states that owned adapters must reap any child or Docker command they started whenprepareis dropped.teardownandteardown_for(reason)are idempotent.teardown_forcarries theProviderCheckpointReason(BeforeStop,BeforeReset, and others) for the final checkpoint; the default body callsteardown. Codex defaults toBeforeStop, which is what reaper and shutdown paths get because they call plainteardown.HostError::RuntimeStoppedmeans the process stopped but finalization failed. The manager then keeps the session dormant instead of re-tracking a dead handle.runtime_pidreports a live child PID and may retain the last PID after a crash so the manager can tell a crash from a reap.runtime_transport_activereports protocol liveness for SDK-owned children whose PID is hidden, and overrides PID checks when present.idle_reap_policyreturnsAllowby default;KeepAliveexempts a runtime from the reaper.readiness_warningdegrades the peer toDegraded.
#Consumed: Band
Room membership comes only from Band room_added and room_removed events and Band room lists. A room is retired locally only on a room-scoped Band read that returns exactly 404 (prune_inaccessible_bound_rooms); absence from a list is never the deletion signal.
#Invariants
| Rule | Enforced by | Known exceptions |
|---|---|---|
| A parked session never receives traffic. | HostSession::is_parked, Peer::routable_sessions in spawn_worker |
None found |
| One room per host session and one session per room within a peer. | Manager::upsert_host_session conflict check |
force (--steal) evicts the other session deliberately |
Room-bound owned sessions carry the template's RuntimeTemplateId and their own RuntimeSessionId. |
Copy in upsert_host_session before apply_runtime_opts mints IDs; desired_runtime preserves the child's IDs |
An explicit first move of a compatibility session into a Managed Docker family adopts the template UUID (desired_runtime, entering_managed_docker) |
| One runtime host per runtime session; one live Shared host per agent; one active binding per agent and room. | Primary key and partial unique indexes in migration 0045; RuntimeHostAllocator |
None found |
| Host-wide scope comes from durable membership, never from a room or the live registry. | control_runtime_host_inner; validate_session_sandbox_reset_scope; tests shared_host_stop_enumerates_every_room_and_persists_all_dormant, shared_session_restart_and_unbind_never_restart_the_sibling_runtime |
None found |
Room-scoped restart and unbind never use restart_peer on a peer with siblings. |
restart_session and detach_session use teardown_live_owned_session |
detach_session of an attached (not owned) session with siblings still restarts the peer |
| A live child is never orphaned by a failed teardown. | Worker::reinsert_session; teardown_live_owned_session re-tracks when the PID is still alive; pending_host_cleanups |
A RuntimeStopped error, or a dead PID, retires the host instead |
| Provider-session creation is journaled before any child can start. | upsert_provisional_session before new_host in bind_session_live_under_gate; recover_incomplete_provisional_sessions in rebuild |
spawn_worker does not journal the sessions it prepares during a fresh spawn |
| Nothing is persisted as bound unless the worker that received the binding is still registered. | Commit inside with_worker_runtime_settlement, no await between check and save_peer |
None found |
The Inner guard is never held across .await. |
std::sync::MutexGuard is not Send, so any Send future that holds it fails to compile (Control is Send + Sync through async_trait); clippy::await_holding_lock under CI's -D warnings |
Re-locking Inner while already holding it deadlocks without any await, and nothing but review prevents that; the comment in adopt_agent records one near miss |
| An explicitly stopped peer is not resurrected. | Stop intent persisted before teardown; supervise checks wants_running; supervisor_does_not_resurrect_a_stopped_peer, stop_keep_survives_daemon_rebuild_until_explicit_restart |
None found |
| A terminal identity is never auto-started at boot. | peer_should_resume_automatically; rebuild_starts_owned_peers_but_never_restores_terminal_identities |
None found |
| A configuration change is not reported live before a runtime accepted it. | effective_config_generation moves only after a successful replacement bind; RuntimeTxnState::acknowledge rejects stale generations |
Under ApplyAndRestart, dormant rooms set effective equal to desired immediately, because no runtime can hold the old value. The saved parked template is written with effective 0 |
| Shutdown deadlines nest: 5 s worker stop and 30 s Codex finalization inside the 45 s backstop, inside the desktop's 46.5 s wait. | Constants in worker.rs, app_server.rs, jamd.rs, daemon.rs; release_wait_outlasts_the_daemon_force_exit |
Sandboxed Claude Code teardown is bounded at 315 s, and close retries without a limit; see Refactor notes |
Only one jamd owns a config directory, including during shutdown. |
Two file locks held until main returns: DaemonInstanceLock (jamd.lock, acquired first in jamd.rs) and jam_service::try_acquire_daemon_lock (locks/daemon.lock, also probed by the desktop). DaemonInstanceLock::acquire also waits up to 3 s (LEGACY_LOCK_WAIT) for legacy per-peer lock files held by an older daemon |
None found |
#Failure and recovery
Crash during agent creation. The identity-create recovery marker is written after Band registration and before save_peer. reconcile_identity_create_recovery runs at the next provision for the same scope and resolves the orphan. A blank agent id from Band is refused before any local write, and the agent is deleted by the id recovered from its key.
Crash during a room bind. Durable ownership records written before the child (host, membership, workspace, binding) survive and are reused by the next bind (killed_daemon_retries_one_canonical_managed_sandbox_without_duplicate_ownership in bins/jam/tests/jamd_managed_restart.rs). The provisional record makes rebuild refuse to start the peer automatically and records a Failed pending change explaining why, because an earlier child might still exist.
Startup failure or timeout. start_and_wait removes and stops the worker on Failed or after 35 seconds and returns StartupFailed or StartupTimeout. For rebuild and the supervisor, the peer stays wanted and is retried with backoff.
Worker crash or panic. The supervisor task publishes Failed. The next supervisor pass reaps the finished worker and restarts it after the backoff (supervisor_reaps_a_worker_that_died_after_connecting_and_restarts_it).
Provider child crash. The room keeps its row. The host monitor detects the crash (crashed_owned_sessions), announces the room offline, and publishes a peer update. The next inbound message wakes the room, and wake_dormant_room retires the stale host before re-binding.
Idle reap or teardown failure. A hard failure re-tracks the live child and retries on the next sweep. A stopped-but-unfinalized child stays dormant and wakeable.
Bind commit failure. The prepared host is torn down. Failing that, it and its provisional record are queued in pending_host_cleanups, retried by the host monitor and by close.
Room removal failure. A failed child teardown or durable write leaves the room in pending_removals (or undurable_removals), and retry_pending_removals retries the same binding generation. A newer binding generation makes the stale removal a no-op.
Configuration replacement failure. The old configuration is re-bound and the room records PendingChangeStatus::Failed with recovery_target set to the old effective generation. If restoring the old runtime also fails, both errors are returned.
Host-wide partial failure. control_runtime_host_inner persists every member it already stopped as Dormant, publishes the peer, and returns the error. On a restart that fails partway, the host is marked Active if any member restarted, else Parked.
Shutdown during a bind. root_cancel drops the in-flight prepare, the adapter reaps what it started, and the next daemon reuses the same ownership graph (manager_close_cancels_an_inflight_live_host_prepare, graceful_shutdown_cancels_inflight_managed_provision_and_retries_one_owner in bins/jam/tests/jamd_managed_restart.rs).
Daemon restart race. The daemon instance and lifetime locks prevent two daemons from sharing a config directory, and the desktop waits longer than the backstop before spawning a replacement.
#Extension points
#Adding a harness that the lifecycle can run
The lifecycle is provider-neutral once a Host exists, but a new transport still has to be registered in several exhaustive matches before the manager will create, persist, fingerprint, and build it. Extension seams lists the full cross-crate inventory from the OpenCode case study. The execution-specific steps are:
- Domain. Add a
HostTransportvariant with its serde name, add the string to the hand-writtenimpl Deserialize for HostTransportinhost_session.rs, add atransport_catalog!row (label, command, provider, auth modes, settings, and sandbox capability cells), add anis_owned_<name>predicate onHostSession, and, for Docker support, add aSandboxRuntimeProfileand extend the match inHostRuntimeConfig::resolved_sandbox_capability_cell.HostTransport::Unknownappears 15 times inhost_session.rs. - Manager. Extend the exhaustive matches in
apply_runtime_opts,validate_owned_runtime_configuration,default_owned_runtime_provider,clear_unsupported_thread_settings_for_transport,runtime_transport_name, and the string parserparse_runtime_transportinmanager.rs;transport_taginruntime_host.rs, which feeds the compatibility fingerprint; andagent_runtime_configuration_forinanalytics.rs. If the runtime enforces room authority at spawn, add it tovalidate_runtime_room_authorityinworker.rs, which currently branches only onClaudeCodeCli. - Store. Extend
transport_strandparse_transportincrates/jam-store/src/sqlite.rs. Thepersist_enumfallback that preserves unknown strings is guarded byundecodable_enum_survives_a_load_and_save_round_trip. - Composition. Add a branch to
build_hostinbins/jam/src/jamd.rsin the correct position. The chain is ordered: PTY rejection, owned Claude Code, owned ACP or OpenCode, owned Codex, owned Copilot SDK, attached Copilot by provider, attached Claude Code bypeer.host, thenGeneric. An omitted or misplaced branch falls through silently to an attached adapter. If the harness supports Shared placement, construct a registry aroundSharedHostRegistryinmainand pass it intobuild_host. - Adapter. Implement
Hostwith the lifecycle guarantees in Contracts: cancellation-safeprepare, idempotentteardownandteardown_for,runtime_pidorruntime_transport_activeso crash detection and dormancy work, andidle_reap_policyif the runtime must not be reaped. A harness that reports neither PID nor transport liveness is always classified as dormant bysession_is_dormantonce it has no host, and its crashes are invisible to the host monitor. - Work capture. If the harness has a native task file, add an arm to the
new_work_sourceclosure injamd.
#Adding a lifecycle operation
A new Control method is a full vertical change, described in AGENTS.md under "Add a Tauri command". Two execution-specific additions apply:
- If the operation starts or replaces a provider child, add its route to
jam_wire::RUNTIME_CHILD_ROUTES. Bothjam-daemon::unary_timeout_for_routeandjam_client::Client::runtime_operation_timeoutread that list. - If it acts on a runtime host, extend
RuntimeHostLifecycleActioninjam-contract(three variants today) and handle it inruntime_host_lifecycle_classesand the 18 references incontrol_runtime_hostandcontrol_runtime_host_inner, plus the three inbins/jam/src/main.rsand the desktop'scontrolRuntimeHostinApp.tsx.
A new path that makes a room live should call bind_session_live_under_gate rather than build hosts itself, so it inherits journaling, ownership allocation, the commit fence, and cleanup. A new path that stops one room should use teardown_live_owned_session.
#Refactor notes
The regions of manager.rs, split at its impl blocks and section comments, sized by production lines. Two of the four largest regions are impl Manager blocks with no section comment, so the file's own structure does not say what they hold.
Files by size and change rate over the 90 days before the snapshot. manager.rs sits alone in the top right: it is both the largest file and the one changed most often.
File and function size. manager.rs is 58,145 lines with three inherent impl Manager blocks containing 266, 489, and 20 methods, plus a 254-method impl Control for Manager that mostly delegates. Inline #[cfg(test)] modules add thousands more lines. worker.rs is 10,609 lines. The lifecycle functions are individually large: bind_session_live_under_gate is about 700 lines, save_runtime_template about 520, control_runtime_host_inner about 480, spawn_worker about 420, supervise (worker) about 360, rebuild about 265, and ensure_peer about 250. The integration suite crates/jam-manager/tests/lifecycle.rs is 68,007 lines with 900 tests; it is the main safety net for any restructuring.
One lock for many subsystems. Inner holds lifecycle state together with permissions, questions, task access, OAuth sessions, and account background tasks. Manager has 104 fields, about 56 of them independent locks. The lifecycle is therefore coupled to unrelated subsystems through one mutex and one struct, and the no-reentry rule (never call a method that locks Inner while holding it) applies across all of them.
Two representations of one room binding. For Docker sessions, a room binding exists both as a HostSession row (on the peer, rewritten wholesale by save_peer) and as runtime_host_sessions plus runtime_session_room_bindings rows. They are written at different points of the bind, and room removal must update both (handle_room_removed_if_generation marks the membership Dormant, deactivates the binding, then removes the row). Host-native owned sessions have only the HostSession row. A refactor that unifies them must preserve the inactive-binding tombstone that makes authoritative re-add resume the exact provider thread.
Wholesale peer writes. save_peer_with_connectivity is called at 49 non-test sites in manager.rs, and each call deletes and reinserts every host session row for the peer. worker.peer is also overwritten at 7 sites to keep the worker's snapshot aligned. Any path that saves a stale peer clone can overwrite a concurrent change to another row; the gates are what prevent that today.
Placement defaults disagree. The Rust Default for RuntimeHostPlacement is SharedAgentHost, from_existing(None) is DedicatedSessionHost, the desktop create panel defaults to shared_agent_host (LocalAgentCreatePanel.tsx), the desktop edit path maps a missing value to dedicated_session_host (runtimeDraft.ts), and the CLI (sandbox_runtime_intent in bins/jam/src/main.rs) chooses Shared only for Codex and refuses --placement shared for any other transport, although the domain capability cells declare Shared for Copilot, Claude Code, and Cursor ACP.
RuntimeTxnState duplicates durable state. The durable HostSession generations and pending_change are the authority across restarts. RuntimeTxnState is in memory, seeded from them only on template save and replacement, and its provisional journal overlaps the durable provisional_sessions table. Its module comment still says durable persistence "belongs to the store schema, which is U1's unit", but that table now exists and is written by the bind path. PendingChangeStatus::Applied is declared and rendered by the CLI but never written by the manager; a successful apply clears pending_change instead.
Concrete adapter knowledge in the manager. manager.rs and worker.rs name jam_host::claudecode::owned 30 times, jam_host::copilot::sdk 11 times, jam_host::codex::app_server 5 times, and jam_host::opencode several times. Deps::new_host has five call sites, four in manager.rs and one in spawn_worker. Three of the manager's (teardown_hosts, legacy_parked_session_readiness_warning, prompt_unbound_rooms) build a throwaway host with RuntimeHostContext::empty() only to call teardown, read a readiness warning, or push a prompt. That is a boundary exception to "providers translate only at the jam-host edge".
Hard-coded network profile at allocation. ensure_managed_runtime_ownership computes the allocation fingerprint with RuntimeNetworkProfile::Developer and a comment deferring the resolved profile to a later phase, while save_runtime_template and apply_pending_runtime_configuration_at_generation compute the fingerprint with resolved_requested_network_profile. Inferred: for a Production-profile Docker session whose host was recreated with a Production fingerprint, the next bind's allocation request would carry a Developer fingerprint and fail ensure_session_compatible. No test found exercises that sequence.
Shutdown bounds are not uniformly enforced. close loops on remove_and_stop_under_worker_gate without a retry limit, and a peer's owned hosts are torn down one after another inside that peer's task. Inferred: a peer with two or more live Dedicated Codex rooms can need more than one 30-second finalization window, and a sandboxed Claude Code room can need up to 315 seconds, so the 45-second backstop can end the process before every final checkpoint is written. The concurrency test covers independent peers, not multiple hosts under one peer.
Stale comments about a per-peer file lock. The doc comments on rebuild (above recover_incomplete_provisional_sessions), try_start_peer, and Manager::supervise, the comment above spawn_supervisor in jamd.rs, and the comment above FORCE_EXIT_GRACE in jamd.rs describe a per-peer flock held by an exiting predecessor. Commit 533221ff2 removed crates/jam-manager/src/lock.rs and PeerLock, and added DaemonInstanceLock (jamd.lock) with a bounded 3-second wait for legacy per-peer lock files that an older daemon may still hold. The jam_service daemon lifetime lock predates that commit. No current daemon holds per-peer locks, and try_start_peer now returns Ok(false) only when the manager is closed or the worker already exists. The test name create_only_maps_a_cross_manager_peer_lock_to_the_typed_local_conflict also refers to that lock. The conflict it asserts now comes from the second manager's ensure_peer finding the first manager's stored peer record through get_peer and returning already_provisioned_for_create_only.
Order-dependent host dispatch. build_host chooses an adapter through an ordered if chain over is_owned_* predicates, then session.provider, then peer.host. Attached delivery depends on fields other than the transport, so a registry keyed only by (HostRuntime, HostTransport) would not cover it.
Reused peer-state vocabulary. The human room feed publishes PeerState::Connected for the pseudo-peer human/<profile>. Consumers that assume a PeerState event names a stored peer will mis-handle it.