Jam architecture main at 59f67f5d0 · 2026-09-29
View source

#Sandbox, workspaces, and continuity

On this pageWhere it livesHow it worksExecution boundaries and the capability matrixThe ownership model: plan, marker, and nonceDeployment view of a Shared runtime hostCreating a managed Codex runtime through its first turnEnvironment delivery and secretsNetwork policy: exact rules, never adoptionRepository operations inside the guestDocker environment filesThe local-platform bridgeCheckpoint on stop and exact resume on restartContinuity reset: explicit fresh startArchive, export, restore, and removalRuntime-host cleanupWorkspace operation state machineStartup orderState and ownershipContractsControl methodsDestructive actionsHost-side ports and adapter seamsSemantic guarantees enforced by codeInvariantsFailure and recoveryExtension pointsAdding a new sandboxed harnessAdding a new workspace operation kindAdding a destructive actionRefactor notes

This subsystem runs an owned coding runtime inside a Docker Sandboxes microVM instead of on the host, gives each room session a Jam-owned workspace directory on the host, and preserves the provider's conversation state across process, microVM, and daemon restarts so that a restarted runtime resumes the exact provider thread it had before. It also owns the destructive end of that lifecycle: archiving a workspace into an immutable export, deleting exports and old checkpoint payloads, and removing a whole runtime host. It does not own the provider protocol (each adapter in jam-host speaks Codex app-server, Copilot SDK, Claude Code, or ACP itself), does not own room routing or Band connectivity (the daemon keeps the Band socket and credentials; the guest never opens a Band connection), and does not decide which runtime host a session belongs to (that is the allocator described in Execution and lifecycle).

#Where it lives

The code splits along the crate boundary. jam-host owns everything that invokes the sbx CLI or touches a guest. jam-manager owns durable ownership records, the host filesystem layout under the Jam app directory, the operation journals, and the recovery decisions. jam-domain holds the shared record types, and jam-store persists them.

Area Location Key types and functions
sbx CLI boundary crates/jam-host/src/sandbox.rs (about 11,100 lines of production code before its test module, 320 top-level functions) SandboxPlan, SandboxOwnerKey, SandboxRemovalTarget, PreparedExec, SharedSandboxLease, RuntimeEnvFile, HostOwnerMarker, ensure_created_with_owner_compatibility, park_owned, reset_owned, remove_verified_owned_target, network_policy_readiness, github_readiness_with_repositories_classified, clone_repository and siblings, read_owner_inventory, sandbox_inventory
Docker environment files crates/jam-host/src/sandbox/environment_operation.rs EnvironmentOperationLease, reserve, run_reserved, EnvironmentApprovalProvenance (PTY-backed same-child approval via a jamd supervisor helper)
Docker MCP task service crates/jam-host/src/sandbox/task_service.rs Receipt, prepare, load, digest, cleanup
Guest-to-host bridge crates/jam-host/src/local_bridge.rs, crates/jam-host/src/local_bridge_process.rs LocalBridgeGrant, LocalBridgeRelay, LocalBridgeProcess::start
Provider-state selection crates/jam-host/src/codex/provider_state.rs, copilot/provider_state.rs, claudecode/owned/provider_state.rs, acp_provider_state.rs, shared provider_state_fs.rs stage_quiesced_checkpoint, activate_quiesced_checkpoint per provider
Continuity seam types crates/jam-host/src/lib.rs RuntimeHostContext, ManagedRuntimeContext, ManagedRuntimeCheckpoint, ManagedRuntimeHostCheckpoint, ManagedRuntimeOrphanRecovery, RuntimeCheckpointBroker, RuntimeCheckpointRequest, RuntimeCheckpointCoordinator (an alias of codex::shared_dispatch::SharedTurnCoordinator), ProviderResumeObserver
Workspace service crates/jam-manager/src/workspace.rs (about 2,850 production lines) plus workspace/manifest.rs, workspace/repository.rs, workspace/export.rs, workspace/checkpoint.rs WorkspaceManager, ManagedRuntimePaths, managed_runtime_paths, RuntimeHostManifest, inspect_workspace_repositories, authorize_workspace_removal
Checkpoint broker crates/jam-manager/src/checkpoint_broker.rs runtime_checkpoint_broker, runtime_checkpoint_coordinator, runtime_resume_observer, attach_runtime_checkpoint_broker, persisted_provider_session_id
Context composition crates/jam-manager/src/worker.rs managed_runtime_context_for_session, existing_managed_runtime_storage_context_for_session, ensure_managed_runtime_ownership
Startup inventory crates/jam-manager/src/sandbox_reconciliation.rs SandboxInventoryProvider, SandboxReconciliationDisposition, classify_sandbox_inventory
Kit pin policy crates/jam-manager/src/managed_sandbox_kit.rs and JAM_MANAGED_WORKSPACE_KIT_REF in sandbox.rs apply_new_managed_codex_defaults, apply_managed_codex_probe_defaults
Host-native allocation (separate) crates/jam-manager/src/owned_workspace.rs allocate, recover, marker_matches
Manager entry points crates/jam-manager/src/manager.rs control_runtime_host, reset_runtime_sandbox, archive_managed_workspace, delete_managed_workspace_export, recover_managed_workspace, inspect_runtime_host_cleanup, cleanup_runtime_host, recover_runtime_host_cleanups, inspect_provider_continuity_recovery, start_fresh_provider_continuity, replay_provider_continuity_resets, start_managed_workspace_*, sandbox_reconciliation_report, rebuild
Side-effect ports crates/jam-manager/src/lib.rs RuntimeSandboxResetter, RuntimeSandboxRemover, RuntimeSandboxGitHubReadinessProbe, RuntimeSandboxNetworkPolicyProbe, RuntimeSandboxRepositoryOperator, DockerSandboxSupportProbe and their HostRuntimeSandbox* production implementations
Domain records crates/jam-domain/src/runtime_workspace.rs RuntimeHostRecord, RuntimeHostSessionRecord, ManagedWorkspaceRecord, WorkspaceOperationRecord/Event, WorkspaceExportRecord, RuntimeHostCleanupOperation, ProviderCheckpointRecord, ProviderResumeObservationRecord, ProviderContinuityResetRecord, ProviderCheckpointPayloadDeletion
Persistence crates/jam-store/src/lib.rs trait RuntimeWorkspaceRepo (45 methods), backends in sqlite.rs and file.rs, migrations 0045 to 0070 workspace_operation_event, runtime_host_cleanup_event transition validators

Two words recur and mean different things. A managed workspace (ManagedWorkspaceRecord, workspace.rs) is the Docker-era, UUID-owned directory under <app>/runtime-hosts/<host>/workspace-root/sessions/<session>. A host-native managed workspace (jam_domain::ManagedWorkspace inside HostRuntimeConfig, owned_workspace.rs) is an older, unrelated allocation under ~/Jam/<account>/<agent>/<session> for runtimes that are not sandboxed. The two share a name but no code. The rest of this chapter means the Docker-era kind unless it says host-native.

#How it works

#Execution boundaries and the capability matrix

Every owned runtime runs either host-native or inside a Docker Sandbox. Host-native is the default and the daemon must work when sbx is absent (probe_support in sandbox.rs returns a classified SupportProbe rather than failing startup). A sandboxed runtime is described by three independent choices, resolved per runtime template:

  • Configuration source (SandboxConfigurationSource): JamManaged, where Jam chooses kits, MCP registrations, mounts, and network policy, or DockerEnvironment { files }, where Docker merges the operator's .sbxenv files and Jam appends only a private ownership overlay (prepare_environment_overlay).
  • Workspace source (SandboxWorkspaceSource): Managed (Jam allocates the directory; no host path or Git repository needed), PrivateHostClone { host_path } (sbx create --clone on a primary Git checkout), or DirectHost { host_path } (mount and edit in place). resolve_workspace_source in sandbox.rs maps each to the (workspace, clone_workspace) pair that create_args uses.
  • Placement (RuntimeHostPlacement): SharedAgentHost (one microVM per agent identity serving several room sessions) or DedicatedSessionHost (one microVM per session). RuntimeHostPlacement::from_existing defaults an absent value to Dedicated.

Which combinations are legal is not decided here. It is the static sandbox_cells column of the transport catalog macro in crates/jam-domain/src/host_session.rs. Codex, Copilot, Claude Code, and Cursor ACP each publish seven cells; exactly one of the seven is Shared, and it is always JamManaged plus Managed. OpenCode publishes three Dedicated JamManaged cells. Every other transport publishes none. HostRuntimeConfig::resolved_sandbox_capability_cell checks a runtime against this table, and a matching cell proves structural support only. Credentials, Docker version, files, and policy are separate readiness gates.

The desktop controls for creating and configuring Docker runtimes are gated by the docker-sandboxes experiment (DOCKER_SANDBOXES_KEY in crates/jam-manager/src/experiments.rs, read through useExperimentPreference in apps/desktop/src/App.tsx). The experiment is display-only: its registered description states that existing sandboxed agents keep recovery and cleanup controls while it is off, and the CLI and daemon APIs are unchanged. Everything in this chapter below the desktop is shipped behavior.

#The ownership model: plan, marker, and nonce

Jam never identifies a sandbox by its name alone. A SandboxPlan is the exact creation intent: owner, derived name, agent image or kit, primary and additional workspaces, read-only reference folders, kits, static MCP servers, configuration source, and the clone flag. Its owner is a SandboxOwnerKey:

  • RuntimeHost { agent_id, runtime_host_id } is the canonical form and the only one used for a Shared host.
  • RuntimeSession { agent_id, runtime_session_id } is the compatibility shape written by the older per-session path. classify_owner_plan_match accepts it as DedicatedSessionCompatibility only when the stored session UUID equals the requested host UUID. That equality holds because the allocator names a Dedicated host after its session (allocate_dedicated in runtime_host.rs).

derive_name builds the sandbox name as a sanitized prefix plus the first 16 hex characters of SHA-256 over the agent UUID and the stable owner UUID, capped at 63 characters. Room IDs, binding aliases, and display names never enter it.

When Jam creates a sandbox it writes two pieces of evidence (write_owner_with_nonce):

  1. A random nonce, piped over stdin into /home/agent/.jam-sandbox-owner.json inside the guest with umask 077.
  2. A HostOwnerMarker JSON file at <app>/cfg/sandbox-owners/<name>.json holding the full SandboxPlan, the same nonce, the Docker MCP registration digest, and an optional task-service receipt.

The host marker is the authoritative plan. The guest can read its own nonce but can never rewrite the plan that Jam compares on reuse. Every later reuse runs verify_owner_with_mcp_digest (plan match, MCP digest match, and live guest nonce read through owner_marker_read_args). Every removal runs verify_owner_for_cleanup, which checks only plan and nonce. The difference is deliberate, per the comment in that function: cleanup must stay possible for markers that predate MCP digests or whose registrations were later removed.

owner_root derives the marker directory from the store's artifacts root (<app>/cfg/artifacts), so markers live at <app>/cfg/sandbox-owners. An empty artifacts root (probes with an empty context) falls back to a per-user temp directory.

#Deployment view of a Shared runtime host

Where each byte and process of one Shared Codex runtime host lives, and which mounts connect them:

The paths come from managed_runtime_paths in workspace.rs, which derives every directory from validated UUIDs and never from user input: runtime_host_root = <app>/runtime-hosts/<host>, workspace_root = .../workspace-root, session_workspace = .../workspace-root/sessions/<session>, provider_state_root = .../provider-state/<provider>, recovery_root = .../recovery, operation_journal_root = .../operation-journal, and manifest_path = .../manifest.json. The sandbox plan for a Managed runtime uses workspace_root as its primary workspace and adds provider_state_root through SandboxPlan::with_additional_workspace (sandbox_plan in codex/app_server.rs), so repository bytes and provider continuity state are separate mounts. Docker mounts host workspaces at their host paths inside the guest, so cwd arithmetic is the same on both sides.

Several facts about this picture matter for a refactor:

  • One process, many sessions. On a Shared host, CodexAppServerRegistry::build in codex/app_server.rs converges every room facade on one SharedSdkRuntimeHost keyed by runtime-host UUID, and fails closed with shared_runtime_host_incompatible if a second facade arrives with different host-level inputs (SharedRuntimeHostCompatibility). Each room keeps its own provider thread and passes its own session workspace as the per-request cwd. The app-server process itself runs with --workdir /home/agent (GUEST_PROVIDER_HOST_WORKDIR), because the comment on exec_prefix_at records that an inherited mounted cwd can fail getcwd(2) on Docker Desktop when a later room attaches.
  • Live SQLite stays in the guest. Codex's rollout files live in the mounted codex-home, but its SQLite database lives at /home/agent/.jam-provider-state/sqlite-home on the guest's own disk (MANAGED_GUEST_SQLITE_HOME in codex/process_control.rs). It is copied to a host directory only while Codex is quiesced (checkpoint_guest_directory in sandbox.rs) and restored by the fixed guest wrapper script (MANAGED_RUNTIME_WRAPPER_SCRIPT) before exec.
  • Exports are outside the host root. export_root in workspace/export.rs places exports under <app>/workspace-exports/<export_uuid>, so removing a whole runtime host cannot destroy its recovery copies.
  • The lease counts sessions, the microVM parks last. SharedSandboxLease (process-local, SHARED_SANDBOX_LEASES) counts live room claims per (agent, host, sandbox name). release_with parks the microVM (park_owned_unleased) only when the count reaches zero, and a failed final park stays retryable. The Copilot, ACP, and Claude Code adapters acquire this lease directly. Codex uses its own SharedSdkLifecycle::active_sessions count in SharedSdkRuntimeHost::release and parks from the shared supervisor.

#Creating a managed Codex runtime through its first turn

What happens, and where durable state is committed, between a room binding a new Managed/Shared Codex session and that session completing its first turn.

The order inside PreparedExec::prepare_without_runtime_env_with_tasks is the security-sensitive part and is centralized for every adapter: optional task-service composition, environment overlay, ensure_created_with_overlay_and_approval, cancellation check, reference-folder readiness, optional Docker environment control-route probe, network policy reconciliation (Shared hosts go through reconcile_shared_managed_network_policy_claim, which unions every live session's claim), and only then a PreparedExec value. Any failure after creation calls cleanup_pre_provider_failure, which parks a Jam-managed sandbox or runs the environment rollback. configure_runtime_env then refuses to run unless the caller presents the exact plan and policy that preparation used, which prevents a spawn from quietly using different inputs than the ones reconciled.

discover_stock_codex_version in codex/app_server.rs runs the version probe through that same PreparedExec and then hands it to the spawn (prepared_handoff), so preparation happens once. The comment on prepare_managed_provider_mount records why provider directories must exist before this step: Docker captures the mounted tree when it creates the microVM, and a directory added after the version probe created the sandbox would be invisible to the guest.

Room delivery waits for more than the process. prepare_host_and_initial_repositories_until_shutdown in worker.rs runs the manager's materialize_initial_repository_hints after Host::prepare and before the session becomes routable. Each optional template repository becomes an ordinary Clone workspace operation through the same execute_repository_clone path that later operator and model clones use. A failure tears the prepared host down rather than leaving a live but unrouted microVM.

#Environment delivery and secrets

Ambient host environment never crosses into the guest. configure_runtime_env splits the guest environment into two groups by provenance:

  • Allowlisted, non-secret host configuration is written to a short-lived mode-0600 file (RuntimeEnvFile::create) and passed as sbx exec --env-file. PreparedExec::release_env_file deletes it once the long-lived exec process has started, and Drop removes it on every other path.
  • Jam-generated provider-home paths go as explicit --env KEY=VALUE arguments. The module comment explains why: released Docker Sandboxes v0.35 accepts --env-file without delivering its values to the guest.

validate_runtime_env_pairs rejects any name that is_band_custody_env_key matches (BAND_AGENT_KEY, BAND_USER_KEY, or any _-suffixed form), even when the value is a proxy sentinel, and rejects any secret-like name unless its value is a recognized proxy sentinel. Provider credentials are Docker service secrets (sbx secret), substituted by Docker's egress proxy. Jam only inspects the non-secret inventory: global_openai_service_credential, global_github_service_credential, require_claude_docker_credential, and require_cursor_docker_credential parse sbx secret ls output without retaining secret text, and select only the global row.

#Network policy: exact rules, never adoption

Jam requests the reviewed Developer destinations (DEVELOPER_NETWORK_TARGETS), any per-agent HTTPS destinations, and exact localhost targets for the local route and control probes. reconcile_managed_network_policy_inner compares that set with a private journal at <owner_root>/<name>.network-policy (ManagedNetworkPolicyManifest), whose entries are one of:

  • PendingAdd { baseline_rule_ids }: intent recorded before sbx policy allow.
  • Owned { rule_id }: Jam created exactly this rule and may remove only it.
  • CoveredExternally: Docker already allowed the target; Jam must never remove the covering rule.
  • PendingRemove { rule_id }: intent recorded before sbx policy rm.

ensure_managed_network_target checks effective policy first. A target that is already allowed becomes CoveredExternally, except that an exact Docker-environment control target that must be Jam-owned fails instead. Organization-governance denial fails. An existing sandbox-local rule for the target that is not effective also fails, so the recorded baseline is always empty on this path. Otherwise it persists PendingAdd with that baseline, runs the allow, and claims ownership only when exactly one new rule ID appears. Zero new rules with an allowed effective result becomes CoveredExternally. Two or more is an error. Recovery of a PendingAdd applies the same "exactly one beyond baseline" rule. The read-only readiness check is a separate path, network_policy_readiness, which accepts at most 64 targets (MAX_NETWORK_POLICY_TARGETS), runs at most eight checks concurrently, and returns Unavailable when the CLI lacks policy check network.

#Repository operations inside the guest

Managed workspaces start empty. Repositories arrive through six guest operations in sandbox.rs: clone_repository, create_repository_branch, commit_repository, push_repository, create_or_update_pull_request, and inspect_pull_request_checks. Each takes a request type whose constructor validates and canonicalizes input (SandboxRepositoryCloneRequest::parse normalizes owner/repo to a credential-free coordinate and validates a top-level destination; validate_repository_relative_path and validate_git_ref bound the others). Each builds fixed argv with no shell and runs it through sbx exec.

The manager side (start_managed_workspace_clone and siblings in manager.rs) persists a Running WorkspaceOperationRecord of the matching kind through WorkspaceManager::begin_repository_operation, returns the operation immediately, and runs the guest work on a spawned task that races root_cancel. complete_repository_* records a bounded outcome class. Callers poll managed_workspace_operation. The store enforces one Running operation per workspace (see Invariants), so a clone, a commit, an export, and an archive can never interleave on one session workspace.

The model reaches the same operations through the exact-session runtime tool registry in crates/jam-host/src/runtime_workspace_tools.rs (WorkspaceClone, WorkspacePush, and so on), which closes over the authenticated target and runtime-session UUID and exposes no selector. The Providers and attached agents chapter covers that registry.

Two read-only guest checks share the boundary. github_readiness_with_repositories_classified runs git --version, gh --version, and gh auth status --hostname github.com, then up to 16 per-repository gh api checks, eight at a time, each with a 4-second deadline, all inside one SBX_INTERACTIVE_TIMEOUT aggregate (12 seconds, from DOCKER_SANDBOX_SUPPORT_TIMEOUT_SECS). Both checks verify ownership before running anything.

#Docker environment files

When the configuration source is DockerEnvironment, creation goes through sbx env create <files...> <overlay> instead of sbx create. Docker Sandboxes refuses to wait for plan approval when stdin is a pipe, so environment_operation.rs runs that child under a real PTY in a separate jamd supervisor process (--jam-environment-operation-supervisor), which talks to the daemon over a nonce-authenticated private Unix socket and writes an atomic terminal receipt only after the child's process group is gone. The local operator's decision comes from the EnvironmentApprovalBroker in RuntimeHostContext.

Before running the child, ensure_created_with_owner_compatibility persists a PendingEnvironmentCleanup record at <owner_root>/environment-cleanups/<name>.json holding the plan, the owner nonce, the overlay's SHA-256, SHA-256 fingerprints of the source files (audit only; contents are never copied), the supervisor lease, and a phase: Reserved, SupervisorTerminal, EnvironmentCreated, OwnerWritten. Rollback (cleanup_pending_environment_at) waits for the terminal receipt and then runs sbx env rm --force --skip-host-commands <overlay> against the immutable overlay only, never the caller's files. rebuild replays these records first.

#The local-platform bridge

local_bridge.rs is a narrow guest-to-host HTTP and WebSocket relay to one operator-configured loopback origin, for one runtime session. start_local_platform_bridge in codex/app_server.rs refuses unless the runtime is sandboxed and has a managed runtime context. LocalBridgeProcess::start issues a LocalBridgeGrant (hashed capability, at most MAX_BRIDGE_GRANT_TTL_MILLIS of 15 minutes, renewed every five minutes by BRIDGE_RENEW_INTERVAL), launches the public-kit helper through stock Node over sbx exec -i, and serves frames over the helper's stdio. The limits are constants in local_bridge.rs and local_bridge_process.rs: 16 concurrent requests, 120 requests per minute, 4 MiB request bodies, 8 MiB responses, 16 WebSockets, 256 KiB chunks, and 32 MiB per minute per direction. They are per runtime session. Only the Codex adapter starts the bridge today (local_bridge_process is referenced from codex/app_server.rs and lib.rs only).

#Checkpoint on stop and exact resume on restart

Continuity rests on one rule, enforced in managed_runtime_context in worker.rs: a session that has ever bound a provider identity must resume that exact identity from a verified checkpoint, or fail before the host factory runs. There is no fallback to a fresh thread.

The coordination has three parts:

  • The coordinator. runtime_checkpoint_coordinator returns one RuntimeCheckpointCoordinator per runtime-host UUID from a process-wide weak registry. It is a Tokio RwLock: each turn holds a shared permit (enter_turn) and a checkpoint takes the exclusive side (with_all_sessions_quiesced), so no sibling room mutates the provider home while bytes are copied, and a queued checkpoint is not starved.
  • The broker. runtime_checkpoint_broker is the closure the adapter calls. It validates the host, peer, and membership against the store, computes the complete canonical session list for the host (canonical_host_sessions), asks WorkspaceManager::prepare_provider_checkpoint for a fresh staging directory, runs the adapter's capture closure, and publishes with complete_provider_checkpoint. The adapter chooses which files to copy. The manager chooses the checkpoint UUID, the session mapping, and whether publication succeeds.
  • Publication. complete_provider_checkpoint in workspace/checkpoint.rs scans the staged tree with bounds (100,000 files, 20 GiB, depth 64), enforces private permissions, writes an immutable complete.json manifest with a metadata checksum, and only then inserts the ProviderCheckpointRecord. A crash between the file and the row is repaired by an exact retry (test completion_retry_repairs_marker_to_store_crash_window).

How a Codex session's state survives a stop, and how the next launch proves it resumed the same thread.

Checkpoints are taken for several reasons (ProviderCheckpointReason). AfterTurn is set by checkpoint_completed_turn_if_due in the Codex loop when a provider turn completed. BeforeStop is the default shutdown reason (shutdown_checkpoint_reason). BeforeReset is passed through Host::teardown_for by host reset, room-scoped sandbox reset, and rebind. The manager's teardown_live_owned_session is the single place that passes a reason to an adapter. BeforeRemove exists in the enum and the store decoder but no production code constructs it, and Manual appears only in a test in workspace.rs.

On the restart side, inspect_provider_checkpoint_activation returns one of four results, and managed_runtime_context turns each into a context or a refusal:

Inspection result Session has a bound provider identity Outcome
Ready(activation) Yes or no checkpoint set with selected_provider_session_id; the checkpoint must also capture every live sibling's identity (checkpoint_captures_all_provider_sessions)
SessionNotCaptured No, and placement is Shared host_checkpoint set: restore the shared home, but start a new conversation for this room
SessionNotCaptured Otherwise Refuse: session absent from the latest host checkpoint
Missing No Fresh launch
Missing Yes, Shared Cursor ACP orphan_recovery set: the adapter may reuse live host-backed state only after it proves it fenced an orphaned microVM (SharedSandboxLease::orphan_was_fenced)
Missing Yes, anything else Refuse: bound identity but no canonical checkpoint
ProviderVersionMismatch Any Refuse

When a checkpoint is selected, attach_runtime_checkpoint_broker also installs a ProviderResumeObserver. The Codex adapter refuses to start an exact resume without it (from_session returns "managed Codex exact resume is missing its durable outcome observer"). The observer writes one immutable ProviderResumeObservationRecord per attempt. The Codex loop reports Succeeded only when ThreadStarted carries the expected thread ID, and Failed on a stream error during a required exact resume. If recording either outcome fails, the adapter fails the turn rather than emitting provider output.

#Continuity reset: explicit fresh start

When exact continuity cannot be honored, the operator can start fresh, but only through a two-step, evidence-bound path. inspect_provider_continuity_recovery classifies the host as healthy or as one of ProviderContinuityFailure::{Missing, Corrupt, SessionNotCaptured, ProviderVersionMismatch, ResumeFailed} (the last only when the latest observation for the current checkpoint and session failed), and computes evidence_checksum over the host, checkpoint, failure, sessions, and selected resume observations (provider_continuity_evidence_checksum, schema jam-provider-continuity-evidence-v2). can_start_fresh is true only for a failure while no provider runtime is running (provider_continuity_can_start_fresh).

start_fresh_provider_continuity requires confirm_start_fresh, reinspects, and rejects the call if the checksum changed. It then inserts an immutable ProviderContinuityResetRecord that maps every source runtime-session UUID and binding UUID to newly minted replacements, stops the peer, applies the plan (apply_provider_continuity_reset), and restarts the peer. Application rebinds rooms to the replacement sessions, marks source memberships Dormant and source workspaces Archived, and marks the source host Archived as its final write. That last write is the commit marker: replay on startup (replay_provider_continuity_resets, called from rebuild) treats an Archived source host as complete. Nothing is deleted; the source host, workspaces, checkpoints, and identities remain as evidence until a separate whole-host cleanup.

#Archive, export, restore, and removal

Destruction proceeds in scoped steps, each with its own authority. The session-scoped steps are driven by archive_managed_workspace and delete_managed_workspace_export in manager.rs and workspace/export.rs. The host-scoped step is cleanup_runtime_host.

Archive (one session). The manager tears down the live session with BeforeStop if it is live, sets its membership Dormant, and calls WorkspaceManager::archive_managed_workspace with a new UUID that serves as both the archive operation ID and the export ID. The workspace manager runs a fresh inspect_for_removal (bounded repository inventory, authorize_workspace_removal), sets the workspace RemovalBlocked and stops if blocked, otherwise creates an export, then persists a Running Archive operation whose result_class encodes the retention and force intent (archive_running_class), and finally prunes older exports for KeepLatest, removes the live session directory, marks the workspace Archived, and marks the operation Succeeded. Shared siblings are untouched.

Export (immutable). export_managed_workspace_inner copies the session workspace and the latest compatible provider checkpoint generation into <app>/workspace-exports/.<uuid>.partial, writes a checksummed manifest, and renames it into place before recording the WorkspaceExportRecord. An exact retry verifies an existing export instead of overwriting it (test exact_export_retry_repairs_database_publication_after_filesystem_commit).

Restore (non-destructive). restore_managed_workspace_export_inner requires the export's immutable owner to equal the requesting session's workspace, host, and session, refuses if the destination exists, restores the checkpoint first, then restores the workspace through a private staging directory and rename. A restore of an archived workspace returns an Archived or Removed host to Parked. It records a Reconcile operation with result class restored_export.

Export deletion (one export). Allowed only when the workspace is Archived and the caller passes confirm_permanent. The export UUID is carried in the Running Remove operation's result class (recovery_export_delete_running_<uuid>) so replay knows exactly what to delete.

Whole-host cleanup is described with its state machine below.

#Runtime-host cleanup

inspect_runtime_host_cleanup (read-only) combines WorkspaceManager::inspect_runtime_host_cleanup with live worker state and the Docker ownership inventory. It blocks unless the host is Parked (or Archived as a continuity-reset source), every member is Dormant, every Managed member's workspace is Archived with at least one verified export and no running operation, no member is live in a worker, and the sandbox reconciliation shows exactly one Healthy stopped sandbox or a confirmed-absent one. Path-backed members (Private clone, Direct) need no export because Jam does not own their bytes.

cleanup_runtime_host requires confirm_permanent, rejects more than one cleanup operation per host, resumes an existing incomplete one, and otherwise reruns the full inspection before persisting a Prepared RuntimeHostCleanupOperation. finish_runtime_host_cleanup then advances stage by stage.

This state machine answers which side effect happens at each stage and what a crash at each point replays.

Each stage write goes through upsert_runtime_host_cleanup_operation, which appends an event row atomically and rejects any transition that moves backward, changes immutable fields, or follows Succeeded (runtime_host_cleanup_event in jam-store/src/lib.rs). remove_runtime_host_root itself refuses unless the operation is already at SandboxRemoved or later, and treats an absent root as an idempotent retry. rebuild calls recover_runtime_host_cleanups before ordinary workspace reconciliation. Retained exports are never removed by this path. RuntimeHostCleanupOperation records planned_runtime_host_bytes, removed_runtime_host_bytes, and retained_export_bytes separately so capacity reports distinguish them.

The checkpoint payload deletion uses the same pattern with stages Prepared, PayloadRemoved, Succeeded (ProviderCheckpointPayloadDeletionStage). delete_provider_checkpoint_payload in workspace/checkpoint.rs requires manual_delete_allowed from the retention inventory, which excludes the newest checkpoint, any checkpoint referenced by a workspace export or a continuity reset, missing payloads, and corrupt evidence. The immutable ProviderCheckpointRecord metadata is never deleted.

#Workspace operation state machine

This state machine answers which transitions of a WorkspaceOperationRecord the store accepts, and what startup does with an operation left Running.

workspace_operation_event in jam-store/src/lib.rs is the transition validator for both backends. An identical upsert appends nothing. Any change to ID, workspace, host, session, kind, or start time is rejected. Only the six arrows out of Pending and Running are legal, which makes Succeeded, Failed, and Cancelled terminal. Every accepted transition appends a WorkspaceOperationEvent with the next sequence number in the same transaction.

What happens to a Running operation after a crash depends on its kind:

Kind Who creates it On restart
Provision ensure_managed_inner Replayed by reconcile_managed_inner
Reconcile reconcile_managed_inner, restore Replayed by reconcile_managed_inner
Inventory inspect_for_removal_inner Never Running: inserted as Succeeded together with its evidence via complete_workspace_repository_inventory
Clone, Branch, Commit, Push, PullRequest, Checks begin_repository_operation Marked Failed with interrupted; partial bytes preserved; never replayed
Export export_managed_workspace_inner Not replayed automatically; an exact retry of the same export UUID finishes or verifies it
Archive archive_managed_workspace_inner Finished by resume_running_archive_for_session from reconcile_all_managed and retry_recovery
Remove (export deletion) delete_managed_workspace_export_inner Finished by resume_running_export_delete_for_session
Reset none The variant exists in WorkspaceOperationKind, the SQL CHECK, and the generated TypeScript binding, but no production code constructs it

The workspace record has its own lifecycle (WorkspaceLifecycleState). Production code writes Planned, Provisioning, Ready, Recovering, ProvisionFailed, NeedsReconciliation, RemovalBlocked, and Archived. Active, Parked, and Removed are read in matches! guards but never assigned by production code.

#Startup order

Manager::rebuild runs the replay steps of this subsystem before starting any worker, and each step logs and continues on failure so one bad record cannot block unrelated peers:

  1. jam_host::sandbox::reconcile_pending_environment_cleanups (Docker environment rollback records).
  2. recover_runtime_host_cleanups (forward-only host cleanup).
  3. WorkspaceManager::recover_provider_checkpoint_payload_deletions.
  4. replay_provider_continuity_resets.
  5. WorkspaceManager::reconcile_all_managed (archive and export-delete replay, provision and reconcile replay, interrupted repository operations).
  6. Usage reconciliation and the rest of rebuild, then workers.

The read-only Docker inventory runs after rebuild, as a detached task in bins/jam/src/jamd.rs, so a missing or slow Docker cannot delay host-native startup. sandbox_reconciliation_report snapshots sbx ls and the owner markers through the injected SandboxInventoryProvider (HostSandboxInventory in jamd.rs) and classifies each into a SandboxReconciliationDisposition: Healthy, PersistedWithoutOwnerMarker, OwnerWithoutDocker, OrphanOwnerMarker, DedicatedSessionCompatibilityVerificationRequired, or OwnershipMismatch. Foreign sandboxes are counted but never listed. No disposition authorizes a mutation. The cleanup preflight reads the same report on demand.

Orphaned Shared microVMs are handled at the lease, not at startup. The process-local lease registry does not survive SIGKILL, so the first acquire_shared_sandbox_lease for a host in a new daemon runs fence_orphaned_shared_sandbox: a running sandbox with no live lease is ownership-verified and stopped before any provider child starts.

#State and ownership

Aggregate Key and scope Authority Mutation coordinated by Durable representation Projections Fence or version Recovery Deletion authority
Runtime host RuntimeHostId, one per agent for live Shared, one per session for Dedicated Store row; allocator decides membership RuntimeHostAllocator (runtime_host.rs), WorkspaceManager, control_runtime_host_inner, cleanup, continuity reset runtime_hosts; partial unique index one_live_shared_runtime_host_per_agent jam sessions host UUID, desktop host grouping compatibility_fingerprint, changed only through replace_runtime_host_compatibility with an expected value reconcile_managed; compatibility transition file in operation-journal cleanup_runtime_host only
Host membership runtime_session_id (primary key) to host Store Allocator attach, host control, archive, continuity reset runtime_host_sessions with workspace_kind Host-wide scope for stop, restart, reset, cleanup State Attached, Dormant, Failed Read at every host-wide action; never inferred from the live worker registry Not deleted
Managed workspace WorkspaceId, one per runtime session Store row plus manifest.json WorkspaceManager managed_workspaces; directory under host root Recovery status, repository panel root_version; manifest schema 3 with metadata checksum reconcile_all_managed, retry_recovery Archive removes live bytes only
Workspace operation WorkspaceOperationId Store WorkspaceManager and repository task workspace_operations plus append-only workspace_operation_events managed_workspace_operation, recovery journal Forward-only validator; one Running per workspace Per-kind rules above Never deleted
Repository inventory WorkspaceId Store, written with its Inventory operation inspect_for_removal_inner workspace_repository_inventories inspect_managed_workspace_repositories Replaced atomically with a Succeeded Inventory row Recomputed at every destructive boundary Replaced, not deleted
Workspace export WorkspaceExportId Filesystem manifest plus store row Export, archive, prune, delete <app>/workspace-exports/<uuid>; workspace_exports Retention inventory, 30-day advisory Immutable; exact retry idempotent Exact retry; export-delete replay delete_managed_workspace_export or KeepLatest prune during archive
Provider checkpoint ProviderCheckpointId per host complete.json plus store row Checkpoint broker under coordinator exclusive permit recovery/provider-checkpoints/<provider>/<uuid>; provider_checkpoints, provider_checkpoint_sessions Continuity inspection, retention Immutable; provider version; metadata checksum Exact retry repairs marker-to-row window Payload bytes only, through delete_provider_checkpoint_payload; metadata never
Resume observation ProviderResumeObservationId Store runtime_resume_observer provider_resume_observations Continuity inspection and checksum Latest per checkpoint and session wins; timestamps forced monotonic None needed Never deleted
Continuity reset RuntimeHostOperationId Store start_fresh_provider_continuity provider_continuity_resets Reset result Evidence checksum at confirm time replay_provider_continuity_resets Never deleted
Host cleanup operation RuntimeHostOperationId, at most one per host Store cleanup_runtime_host, startup replay runtime_host_cleanup_operations plus events Cleanup result Forward-only stage recover_runtime_host_cleanups Never deleted
Sandbox ownership Sandbox name Host marker file plus guest nonce sandbox.rs create, reuse, remove <app>/cfg/sandbox-owners/<name>.json; guest /home/agent/.jam-sandbox-owner.json Startup inventory Plan equality (with reviewed transitions) and nonce reconcile_pending_cleanup, environment cleanup replay Removed with the sandbox by remove_owned_host_metadata
Network rule ownership Sandbox name plus target Private manifest reconcile_managed_network_policy_inner <owner_root>/<name>.network-policy Readiness reports Baseline rule IDs; exactly-one-new claim Pending entries resolved on next reconciliation Only Owned rule IDs
Shared lease (agent, host, sandbox name) Process memory SharedSandboxLease None (lost on crash; orphan fenced at next acquire) None Plan and owner root must match fence_orphaned_shared_sandbox Final release parks
Host-native allocation Room-session key per peer Setting row plus directory marker prepare_managed_workspaces in manager.rs, owned_workspace.rs managed-workspace-allocation.v1.* setting; .jam-workspace.json marker HostRuntimeConfig.managed_workspace Account instance plus allocation nonce recover; mismatch becomes Orphaned Never removed automatically

#Contracts

#Control methods

The daemon exposes 30 Control methods for this subsystem. All have default trait bodies and none is a streaming route; environment_approvals is a GET, and inspect_managed_workspace_repositories has no route of its own. See the generated Control API crosswalk for routes, Tauri commands, and CLI use. They group as follows:

  • Host lifecycle: control_runtime_host (Stop, Restart, ResetSandbox), reset_runtime_sandbox (room-scoped; refuses Shared through validate_session_sandbox_reset_scope).
  • Read-only inspection: inspect_managed_workspace_removal, inspect_managed_workspace_repositories (no route of its own; its default body in jam-contract calls inspect_managed_workspace_removal with force_discard = false and converts the result), inspect_managed_workspace_github_readiness, inspect_managed_workspace_github_repository_readiness, inspect_managed_runtime_network_readiness, inspect_managed_workspace_recovery, inspect_runtime_host_cleanup, inspect_managed_runtime_capacity, inspect_provider_continuity_recovery, runtime_effective_policy, docker_sandbox_support, managed_runtime_support.
  • Asynchronous repository mutations: start_managed_workspace_clone, _branch, _commit, _push, _pull_request, _checks, polled with managed_workspace_operation.
  • Recovery and destruction: recover_managed_workspace (RetryReconcile, Export, RestoreLatest), archive_managed_workspace, delete_managed_workspace_export, cleanup_runtime_host, start_fresh_provider_continuity, delete_provider_checkpoint_payload.
  • Docker environment approvals: environment_approvals, decide_environment_approval. Credential setup: configure_docker_github_global.

#Destructive actions

Each destructive action has its own confirmation, its own fresh evidence at the mutation boundary, and its own journal. None of them grants authority to another.

Action Scope Authority required Fresh evidence at mutation Journal and replay
Archive managed workspace One runtime session Owning agent target; force_discard flag for blocked work inspect_for_removal rescans the workspace inside the archive call; export created and verified before any live byte is removed Archive WorkspaceOperationRecord with intent in result_class; replayed on startup and explicit retry
Delete recovery export One export of an archived session confirm_permanent; workspace must be Archived Export manifest and owner reverified Remove WorkspaceOperationRecord carrying the export UUID; replayed on startup
Prune older exports (KeepLatest) Older exports of one session Part of the archive request Each export inspected before deletion Inside the Archive operation
Delete checkpoint payload One non-latest, unreferenced checkpoint confirm_permanent Retention inventory rerun; manual_delete_allowed required ProviderCheckpointPayloadDeletion, forward-only; replayed on startup
Start fresh continuity Whole runtime host confirm_start_fresh plus the inspection's evidence_checksum Inspection rerun; checksum must match Immutable ProviderContinuityResetRecord; replayed on startup; deletes nothing
Host reset sandbox Whole runtime host discard_changes for Private clone dirty work Fresh repository inventory per Managed member; Private-clone preservation preflight; guest nonce Managed operation log only; no durable stage journal
Room-scoped sandbox reset One Dedicated session discard_changes Same as host reset Managed operation log only
Whole-host cleanup Whole runtime host confirm_permanent Full preflight rerun before persisting Prepared; nonce verified at removal RuntimeHostCleanupOperation, forward-only; replayed on startup
Retire vacated sandbox Host whose members all left sandboxing Configuration change Nonce-verified removal None; host set Parked after removal

#Host-side ports and adapter seams

The manager never calls sbx directly for these operations. It goes through injectable ports in crates/jam-manager/src/lib.rs, whose production implementations delegate to jam_host::sandbox: RuntimeSandboxResetter (validate_private_clone_reset, reset_owned), RuntimeSandboxRemover (remove_verified_owned_target_with_environment_approval or remove_verified_owned_target), RuntimeSandboxGitHubReadinessProbe, RuntimeSandboxNetworkPolicyProbe, RuntimeSandboxRepositoryOperator, DockerSandboxSupportProbe, and SandboxInventoryProvider. Lifecycle tests inject recorders and never touch Docker.

Adapters receive continuity only through RuntimeHostContext fields: managed_runtime, checkpoints, checkpoint_activation, checkpoint_coordinator, provider_resume_observer, and environment_approvals. The adapter may report outcomes; it never selects a host, checkpoint, runtime session, or provider thread.

#Semantic guarantees enforced by code

  • Ordering. Every destructive or journaled mutation persists its intent before the side effect: Running operations, Prepared cleanup and payload-deletion rows, PendingAdd and PendingRemove network entries, PendingEnvironmentCleanup, and the continuity reset record.
  • Idempotency. Exact retries under the same UUID are accepted and change nothing (insert_workspace_export, insert_provider_checkpoint, insert_provider_continuity_reset, identical operation upserts). A changed record under an existing UUID fails.
  • Timeouts. Read-only sbx version, diagnose, ls use 12 seconds; sbx create uses 600 seconds (SBX_CREATE_TIMEOUT, after field evidence of a 298-second image pull); other lifecycle commands use 300 seconds (DOCKER_SANDBOX_LIFECYCLE_TIMEOUT_SECS), selected by sbx_command_timeout. The Codex adapter bounds sandbox preparation at 630 seconds (SANDBOX_PREPARE_TIMEOUT), child shutdown at 3 seconds, and runtime finalization at 30 seconds. run_network_policy_check uses 15 seconds (NETWORK_POLICY_CHECK_TIMEOUT).
  • Retries. The Codex adapter's park_sandbox retries park_owned every second until it succeeds (CLEANUP_RETRY_DELAY). Nothing else in this subsystem retries automatically; recovery is replay on startup or an explicit operator action.
  • Delivery. Room traffic is not queued behind guest work: repository operations run on spawned tasks and do not block room ingestion.

#Invariants

One live Shared runtime host per agent. Enforced by the partial unique index one_live_shared_runtime_host_per_agent in migration 0045 and by allocate_shared, which reloads and converges on the winner after an insert conflict. Exception: none; Archived and Removed hosts are excluded from the index.

A sandbox is removed only after its host plan and live guest nonce agree. Enforced by verify_owner_for_cleanup inside reset_owned and remove_verified_owned_target, and by load_owned_plan_for_removal, which selects the stored plan by name and immutable owner rather than rebuilding it from a mutable template. Known exception: remove_verified_owned_target removes the host marker without a nonce check when Docker already reports the sandbox absent. Tests: pre_mcp_owner_marker_remains_nonce_verified_and_cleanup_eligible, host_owner_metadata_is_private_and_carries_the_authoritative_plan.

Reuse requires the exact stored plan. Enforced by classify_owner_plan_match. Known exceptions, each reviewed and narrow: the legacy managed-workspace kit digest to the current digest (LEGACY_JAM_MANAGED_WORKSPACE_KIT_REF), the legacy copilot agent template to COPILOT_SANDBOX_KIT_REF, and the session-owned to host-owned Dedicated shape. Tests: legacy_managed_kit_owner_marker_accepts_only_the_reviewed_digest_transition, dedicated_owner_compatibility_never_accepts_reverse_or_cross_agent_changes.

At most one Running operation per managed workspace. Enforced by both store backends in upsert_workspace_operation (SQLite counts state = 'running' inside the transaction; FileStore checks under runtime_ops), and checked again by begin_repository_operation, inspect_for_removal_inner, and export deletion. Exception: an existing Running row may be updated.

Workspace operation transitions are forward-only and journaled atomically. Enforced by workspace_operation_event in jam-store/src/lib.rs, shared by both backends, and exercised by testkit::runtime_workspace_foundation ("a terminal operation cannot be rewritten as a different terminal outcome"). Legacy rows without events get one backfilled event on first read (legacy_operation_summaries_are_durably_backfilled_on_journal_read in file.rs).

Runtime-host cleanup never moves a stage backward and never removes the host root before the sandbox. A side effect whose stage was not yet recorded can run again on replay, so removal treats an already-absent sandbox or root as done. Enforced by runtime_host_cleanup_event (stage monotonic, immutable fields frozen) and by the SandboxRemoved precondition in remove_runtime_host_root. Tests: runtime_host_cleanup_treats_verified_absent_sandbox_as_already_removed, runtime_host_cleanup_preflight_maps_complete_shared_scope_and_blocks_live_members, runtime_host_cleanup_preflight_requires_preservation_for_every_shared_member.

A bound provider identity resumes exactly or fails before the host is built. Enforced by managed_runtime_context in worker.rs and the observer check in codex/app_server.rs Config::from_session. Exception: Shared Cursor ACP orphan recovery, which requires adapter-proven orphan fencing. Tests (in crates/jam-manager/tests/lifecycle.rs): rebuild_passes_persisted_exact_checkpoint_to_fresh_runtime_host, rebuild_fails_closed_for_unrecoverable_managed_provider_state, exact_resume_failure_is_durable_and_verified_success_supersedes_it.

Checkpoints are captured with every sibling quiesced. Enforced by with_all_sessions_quiesced in the broker and adapters, and by enter_turn permits in the Codex, Copilot, ACP, and Claude Code turn loops. Test: managed_runtime_wiring_installs_one_checkpoint_broker asserts sibling contexts share one coordinator instance.

Archive publishes a verified export before removing any live byte. Enforced by the order in archive_managed_workspace_inner and finish_workspace_archive. Tests: archive_reinspects_exports_and_removes_only_the_exact_workspace, startup_reconciliation_finishes_a_running_archive_instead_of_recreating_it.

Repository operations interrupted by a restart are never replayed. Enforced in reconcile_managed_inner. Test: interrupted_repository_clone_is_failed_without_replay_and_partial_bytes_survive.

Band custody credentials never enter the guest. Enforced by validate_runtime_env_pairs (is_band_custody_env_key) for every environment pair. Test: sandbox_env_rejects_band_custody_keys_even_when_names_lack_secret_markers.

Jam removes only network rules it proved it created. Enforced by the ManagedNetworkPolicyEntry journal and the exactly-one-new-rule claim. Test: pending_network_add_claims_only_one_rule_beyond_its_recorded_baseline.

The managed-workspace kit is pinned once. JAM_MANAGED_WORKSPACE_KIT_REF in sandbox.rs is the only source constant; managed_sandbox_kit.rs attaches it only to new Managed Docker Codex templates and probes and removes only that digest on an explicit transition away. Test: reviewed_pin_is_exact_and_applies_only_to_managed_sandboxed_codex. The same digest string is also written out in AGENTS.md and docs/docker-sandbox-runtimes.md.

Live Docker tests run only on the physical host. Every host-like ignored test registers in jam-test-support::LIVE_GATES and requires JAM_LIVE_PHYSICAL_HOST=1 plus its own opt-in; crates/jam-test-support/tests/live_fixture_contract.rs discovers unregistered ones. The decisive gates for this subsystem are managed_codex_guest_cannot_escape_its_owned_isolation_boundary and managed_codex_organization_network_denial_is_authoritative in crates/jam-host/tests/sandbox_probe.rs, live_stock_codex_shared_and_dedicated_topology and live_real_codex_dedicated_docker_resumes_after_recreation in codex/app_server.rs, managed_private_and_public_two_repository_golden_path_is_exact_bounded_and_cleans_up in crates/jam-manager/tests/managed_github_live.rs, and packaged_daemon_materializes_ordinary_room_through_stock_docker_and_resumes in bins/jam/tests/jamd_managed_restart.rs.

#Failure and recovery

Daemon crash during provisioning. A Running Provision operation and a Provisioning host remain. reconcile_all_managed replays it through reconcile_managed_inner, which republishes the layout and manifest or marks the workspace NeedsReconciliation if the manifest is corrupt, preserving it rather than overwriting. Test: reconciliation_replays_a_running_provision_with_no_published_layout, corrupt_manifest_is_preserved_and_marked_for_manual_reconciliation.

Crash between sbx create and the owner marker. If the host marker write fails, write_owner_with_nonce removes the sandbox. If that removal also fails, it writes a pending-cleanup record into the process temp directory (pending_cleanup_root), which reconcile_pending_cleanup consumes on the next ensure_created for that plan, removing the sandbox only if the guest nonce still matches. For Docker environments, the PendingEnvironmentCleanup phase record drives receipt-gated rollback.

Crash with a Shared microVM still running. The in-memory lease is gone. The first acquire in the new process runs fence_orphaned_shared_sandbox, verifies ownership, and stops the microVM before any provider child starts. orphan_was_fenced lets the Cursor ACP adapter authorize its reviewed crash-recovery path.

Crash before a checkpoint publishes. A staged generation without complete.json is reported as Incomplete and is never selected (test prepared_but_unpublished_generation_is_truthfully_incomplete). A crash after complete.json but before the row insert is repaired by an exact retry. If no checkpoint exists for a session with a bound identity, the next start fails closed and the operator uses continuity inspection and start-fresh.

Checkpoint on shutdown fails. The Codex adapter records a finalization failure (record_finalization_failure), and ensure_started refuses to restart that host until teardown is repeated, so a restart cannot run over unsaved state. The daemon's process backstop is 45 seconds (FORCE_EXIT_GRACE in jamd.rs), longer than the 30-second runtime finalization.

Exact resume fails. The observer records Failed. Continuity inspection then reports ResumeFailed for that checkpoint and session. A later verified success supersedes the failure without deleting it.

Archive interrupted. Before the Archive row is persisted, nothing destructive has happened; the export may exist and is reusable only by exact retry. After it is persisted, startup or recover_managed_workspace(RetryReconcile) finishes it from the recorded retention and force intent, reverifying the export first. Test: public_recovery_retry_finishes_archive_after_live_bytes_were_removed.

Export deletion interrupted. Startup finishes it from the export UUID in result_class. A payload found without its canonical record fails closed. Test: startup_replays_export_delete_after_filesystem_commit.

Host cleanup interrupted. Startup resumes from the last stage. A Prepared operation whose sandbox was already absent at preflight must reconfirm absence or stops with a conflict. One failed operation stays incomplete and does not block other peers.

Continuity reset interrupted. The immutable reset record was written first. replay_provider_continuity_resets applies it again, matching either source or replacement session UUIDs, and treats an Archived source host as complete. Test: rebuild_replays_persisted_continuity_reset_before_starting_workers in crates/jam-manager/tests/lifecycle.rs.

Repository operation interrupted. Shutdown cancels it through root_cancel (recorded Cancelled); a crash leaves it Running, and startup marks it Failed with interrupted. Partial bytes stay for the repository panel to show. Initial template repositories that fail leave the room session as a dormant recovery handle rather than failing the whole peer.

Network policy mutation interrupted. PendingAdd and PendingRemove entries are resolved on the next reconciliation using the recorded baseline, or refused if more than one candidate rule exists.

Docker missing or slow. Support probes and inventory return classified unavailability within 12 seconds; the startup inventory runs detached; host-native runtimes are unaffected. The cleanup preflight treats an unavailable inventory as a blocker rather than as absence.

Timeout from the desktop. Read-only checks use the desktop's latest-result-wins controller with a finite UI deadline; the daemon-side work continues to its own bound and cleans up. Destructive mutations are not given UI deadlines.

#Extension points

#Adding a new sandboxed harness

Today a provider becomes Docker-capable by touching all of the following. Counts exclude test modules unless stated.

  1. Capability cells. Add a SandboxRuntimeProfile variant and a *_SANDBOX_CAPABILITY_CELLS constant in crates/jam-domain/src/host_session.rs, and point the transport's sandbox_cells: row at it. Offer a Shared cell only if the adapter can route several provider threads through one process.
  2. Sandbox agent name. Add a match arm to sandbox_agent_for_session in crates/jam-manager/src/manager.rs. It maps HostTransport to the Docker agent image (codex, copilot, claude, opencode, or the reviewed ACP profile's agent) and is called from host reset, room reset, and runtime_host_sandbox_plan. sandbox_agent_for_session returns "copilot" while the Copilot adapter's own SANDBOX_AGENT_TEMPLATE is COPILOT_SANDBOX_KIT_REF; SandboxPlan::from_parts rewrites the legacy name.
  3. Adapter plan construction. Write a sandbox_plan(cfg) function in the adapter that calls SandboxPlan::from_runtime_host_workspace_source when managed_runtime is present (adding provider_state_root with with_additional_workspace) and SandboxPlan::from_spawn otherwise. Four adapters each have their own copy today: codex/app_server.rs, copilot/sdk.rs, claudecode/owned/runtime.rs, and acp.rs (two, one for Cursor and one for OpenCode).
  4. Preparation and spawn. Call one of the PreparedExec::prepare* constructors, then configure_runtime_env, then invocation or invocation_in_workspace, then release_env_file after spawn, and park on every failure path. Adapters use different constructors today (prepare_without_runtime_env_with_tasks in Codex, prepare_with_runtime_tasks in Copilot, Claude Code, and ACP, plain prepare in the bridge and one Copilot probe path).
  5. Shared placement. Acquire acquire_shared_sandbox_lease per room (Copilot, ACP, Claude Code do) or build an equivalent host registry like CodexAppServerRegistry. Hold checkpoint_coordinator.enter_turn() during each turn.
  6. Continuity. Add a provider_state.rs for the provider that implements staging (stage_quiesced_checkpoint) and activation (activate_quiesced_checkpoint), reusing provider_state_fs.rs for bounded traversal and copying. Call the broker with a RuntimeCheckpointRequest at every quiesced boundary. Four production call sites exist today, one each in Codex, Copilot, Claude Code, and ACP. The other constructions in copilot/sdk.rs and runtime_probe_workspace.rs are inside their test modules. Report exact resume outcomes through observe_provider_resume.
  7. Credentials. If Jam-managed mode needs a Docker service secret, add a require_*_docker_credential function beside the existing four in sandbox.rs.
  8. Host compatibility. If the new transport has host-wide settings, make sure RuntimeHostCompatibility::from_runtime in runtime_host.rs covers them, or two incompatible sessions can share a host. Add its tag to transport_tag.
  9. Transport-specific branches in the manager. Outside test modules, manager.rs has about 46 HostTransport::* match-arm lines and 8 equality comparisons (93 lines reference a variant), runtime_host.rs has 8 arms, all in transport_tag, and worker.rs has one comparison and no match arms. Not all concern sandboxing, but each must be reviewed.
  10. Tests. Extend the capability parity test sandbox_capability_cells_match_runtime_validation in manager.rs (module runtime_probe_workspace_tests), which loops over (HostTransport, SandboxRuntimeProfile) pairs, and add a live gate registered in LIVE_GATES.

#Adding a new workspace operation kind

Add the variant to WorkspaceOperationKind in jam-domain, add a migration that rebuilds the workspace_operations table with the new CHECK value (see 0053_workspace_branch_operations.sql for the pattern), decide its restart rule in reconcile_managed_inner (replay, mark interrupted, or finish), add begin_/complete_ methods if it uses the repository slot, and run just bindings so the TypeScript union updates.

#Adding a destructive action

Follow the forward-only pattern: a typed stage enum in jam-domain, a store validator in jam-store/src/lib.rs shared by both backends, persist-before-mutate in the manager, a replay step in rebuild before workers start, a separate confirmation flag on the Control method, and a desktop cancellation regression. Do not reuse another action's confirmation.

#Refactor notes

Oversized files. crates/jam-host/src/sandbox.rs is 16,535 lines; its production portion (before mod tests) is about 11,100 lines and 320 top-level functions spanning at least eight concerns: sbx process execution and output parsing, ownership markers, lifecycle (create, park, reset, remove), environment files, network policy, credentials, GitHub readiness, and repository operations. crates/jam-manager/src/manager.rs is 58,145 lines and holds every Control entry point listed above, plus host control, continuity inspection and reset, and archive orchestration. crates/jam-host/src/codex/app_server.rs is 23,045 lines and interleaves protocol handling, shared-host lifecycle, sandbox preparation, provider-state activation, checkpointing, and the bridge.

Stringly typed errors across the boundary. About 230 function signatures in sandbox.rs (232 lines matching -> Result<_, String>, all before the test module) return a String error. The manager classifies them into bounded result classes by wrapping, and in at least one place the Codex adapter classifies by matching exact error prefixes (managed_version_prepare_failure checks starts_with("Docker MCP registration ") and a fixed list of messages). Changing a message in sandbox.rs can silently change a user-visible classification.

Intent encoded in result_class strings. A running Archive stores retention and force as archive_running_keep_latest_force and similar; a running export deletion stores its target as recovery_export_delete_running_<uuid>. Replay parses these strings (parse_archive_running_class, parse_export_delete_running_class). The column is documented as a bounded redacted classification, so this is a second meaning. The archive also relies on the archive operation UUID equaling the export UUID.

Unused vocabulary. WorkspaceOperationKind::Reset, WorkspaceOperationState::Pending, ProviderCheckpointReason::BeforeRemove, and WorkspaceLifecycleState::{Active, Parked, Removed} are persisted, validated, or exported to TypeScript but never constructed by production code. RuntimeHostState::Active is written by host restart; the runtime host's Removed state is written only by cleanup.

WorkspaceManager is constructed at each call site. Outside test modules there are 19 WorkspaceManager::new( calls in manager.rs, 2 in worker.rs, and 2 in checkpoint_broker.rs. It is stateless apart from app_dir and the store handle, so this is cheap, but there is no single owner through which workspace calls pass.

Duplicated plan construction. Each adapter re-derives its SandboxPlan from ManagedRuntimeContext (step 3 of the harness list), and the manager derives it again in runtime_host_sandbox_plan and runtime_sandbox_reset_plan. They must produce the same name and mounts; runtime_host_sandbox_plan checks plan.name != host.sandbox_name as a guard. The Codex adapter's sandbox_plan hard-codes SandboxWorkspaceSource::Managed whenever a managed runtime context exists, while Config::from_session validates against the session's real source and the ACP, Copilot, and Claude Code adapters pass managed.workspace_source. managed_runtime_context in worker.rs builds a managed context for every sandboxed owned session, including Private clone and Direct sessions (with primary_workspace set to None), and Codex publishes Dedicated Private and Direct cells. Inferred: for such a Codex session, sandbox_plan would call resolve_workspace_source(Managed, None), which returns an error, and its .expect("managed sandbox plan was validated while building Codex config") would panic. Validation in from_session passes because it uses the real source. No Codex test builds a managed context with a path-backed source, so this path may be untested rather than unreachable; confirm before refactoring.

Two shared-host mechanisms. Codex uses SharedSdkRuntimeHost with its own active_sessions counter and parks from its supervisor; Copilot, ACP, and Claude Code use SharedSandboxLease. Both converge sessions on one microVM with a last-release park, but they are separate code paths with separate counters.

Process-global registries. SHARED_SANDBOX_LEASES, the checkpoint coordinator registry in runtime_checkpoint_coordinator, MCP_VALIDATION_CACHE, and NETWORK_POLICY_INVENTORY_CAPABILITY_CACHE are static state in jam-host and jam-manager. Two managers in one process (as in some integration tests) share them. Inferred: a refactor that makes these instance-owned must preserve orphan fencing, which depends on the lease registry being empty in a fresh process.

Layering exceptions. jam-manager/src references jam_host::sandbox directly in 232 places, test modules included (145 in manager.rs, 72 in lib.rs, 10 in runtime_host.rs, 4 in worker.rs, 1 in managed_sandbox_kit.rs), and runtime_host.rs calls jam_host::sandbox::derive_name to compute canonical sandbox names, so the allocator depends on the host crate's naming function. SandboxPlan is a jam-host type that crosses into the manager's port signatures. task_service.rs derives the daemon app root by walking two parents up from owner_root and requires the literal directory names cfg and sandbox-owners, which couples it to the store's layout.

Two unrelated "managed workspace" implementations. owned_workspace.rs plus the settings-table allocation journal in manager.rs (managed-workspace-allocation.v1.*) serve host-native runtimes under ~/Jam; workspace.rs serves Docker runtimes under <app>/runtime-hosts. The comment in prepare_managed_workspaces records a past bug where the host-native backfill reminted Docker session UUIDs, which is the kind of collision the shared name invites.

Pending cleanup in the temp directory. pending_cleanup_root stores the Jam-managed create-rollback record in std::env::temp_dir(), not under the app directory. Inferred: a temp-directory purge between the failed rollback and the next ensure_created loses that record, leaving an orphaned sandbox that only the startup inventory would report as OrphanOwnerMarker or not at all (if the host marker was also removed).

Retention is advisory. Both export retention and checkpoint payload retention compute 30-day automatic_cleanup_candidate flags, but no code deletes automatically. A refactor that adds automatic deletion must keep the newest-generation and referenced-evidence exclusions.

Stale comment. The module comment at the top of sandbox.rs says the module ensures "a per-session microVM". Shared placement makes that false: the canonical owner is the runtime host, and one microVM serves several sessions.

Documentation drift on a timeout. docs/docker-sandbox-runtimes.md and AGENTS.md describe the network policy checks as having "a four-second host-side deadline". The code uses NETWORK_POLICY_CHECK_TIMEOUT = 15s, and the unit test network_policy_check_timeout_covers_a_slow_docker_control_plane asserts it is at least 12 seconds. The four-second figure matches GITHUB_REPOSITORY_READINESS_TIMEOUT, not the network checks.

Scroll to zoom, drag to pan.