Jam architecture main at 59f67f5d0 · 2026-09-29
View source

#Future directions

On this pageRecommendationCriteriaStrategiesState machine inventoryA. Lifecycle rulesWhat the code does todayOptionsRecommendationB. Durable operations and recoveryWhat the code does todayOptionsRecommendationC. Harness boundaryWhat the code does todayOptionsRecommendationD. Owned runtime kernelWhat the code does todayOptionsRecommendationE. Manager structure and concurrencyWhat the code does todayOptionsRecommendationF. Daemon API surfaceWhat the code does todayOptionsRecommendationG. Events and client stateWhat the code does todayOptionsRecommendationH. Runtime ownership graphWhat the code does todayOptionsRecommendationDirections not includedPlanning handoff

The recommended overall direction is Harness platform: make harnesses and lifecycles declarative, split the manager by the state each subsystem owns, and give every owned runtime one ownership graph. The strongest alternative is Harden in place, which fixes the same defects with fewer migrations and less extension payoff.

The chapter covers eight areas. Each area states what the code does today, lists the options including keeping the current design, and scores them on the same six criteria. The Direction picker in the interactive report (just architecture-site, then open target/architecture-site/explore/directions.html) lets a reader change the weights, pick one option per area, and see the combined profile, dependency conflicts, and delivery waves.

Facts cite main at the atlas snapshot. Counts are production code only unless a line says otherwise.

#Recommendation

  • Selected direction: Harness platform.
  • Strongest alternative: Harden in place.
  • Trade-off: The platform strategy changes how transports are identified (C2) and migrates host-native records into the ownership graph (H1). Harden in place avoids both migrations and still removes every silent harness step and every unchecked lifecycle, but the next harness still spans several crates.
  • Evidence floor: The 25-file OpenCode change, 11 of 18 lifecycles with no store or type-level transition check, the binding stored twice, and 49 wholesale peer writes.
  • Reversal trigger: If the first registered harness (OpenCode under C2) still needs edits outside its module beyond registration and desktop copy, fall back to Harden in place for areas C and H.

#Criteria

Every option is scored from 1 to 5 on each criterion, where 5 is always better. The weighted score uses the default weights below; the picker lets a reader change them.

Criterion Default weight Meaning
Adding a harness or feature 3 How much less work, and how few silent steps, it takes to add a coding-agent harness or a cross-cutting feature afterwards.
Simpler system 3 Fewer parallel implementations, fewer concepts, and less code to read after the change lands.
Invariant safety 3 How much harder it becomes to break crash recovery, ownership, authority, or at-least-once delivery.
Delivery risk 2 Chance of regressions and data-migration hazards while the change is in flight. 5 means low risk.
Ships in slices 2 Can land as small pull requests while main keeps moving at about 2,900 commits every 90 days.
Effort 1 Engineering time to reach the end state. 5 means cheap.

Scores are judgments grounded in the evidence each area cites. They are for comparing options with each other, not for measuring a project.

#Strategies

A strategy is one option per area. The five strategies below span the range from no change to a full re-architecture.

Strategy A B C D E F G H Mean weighted score
Keep the current design A0 B0 C0 D0 E0 F0 G0 H0 2.78
Harden in place A1 B1 C1 D1 E0 F3 G1 H2 3.46
Harness platform (recommended) A2 B1 C2 D1 E1 F1 G1 H1 3.79
Re-architecture A4 B2 C2 D2 E2 F2 G2 H1 3.30
Protocol-first challenger A1 B1 C3 D0 E1 F1 G1 H2 3.41
  • Keep the current design. No structural change. Shows the cost of doing nothing.
  • Harden in place. Close the unchecked lifecycles, the duplicated facts, and the duplicated binding without changing the daemon's shape.
  • Harness platform. Make harnesses and lifecycles declarative, split the manager by subsystem, and unify the ownership graph.
  • Re-architecture. The most ambitious option in every area. Listed to show its combined risk, not as a recommendation.
  • Protocol-first challenger. Converge owned runtimes on ACP and keep the rest modest. Tests whether one protocol can carry the product.

How the areas depend on each other. An arrow means an option in the later area requires an option in the earlier one.

#State machine inventory

Every lifecycle the atlas catalogs, and where its transitions are checked today. "Store validator" means both store backends reject an illegal transition. "Owning type" means one type holds the state and every transition goes through its methods. "Guard at some sites" means specific callers check before writing. "None" means any caller can write any state.

Lifecycle Type States Stored in Transitions checked by Unused vocabulary
Peer connection PeerState 6 Memory, published as an event None. No transition check. 56 production references in 5 files publish or match it.
Account credential AccountAuthState 3 Memory in the manager Guard at some sites. refresh_rejection_can_require_auth permits only an absent or Active state to become AuthRequired. Other writes are unchecked. The durable signal is whether a refresh token exists.
Runtime host RuntimeHostState 9 runtime_hosts.state Guard at some sites. The store checks immutable ownership and allows one live Shared host per agent (the one_live_shared_runtime_host_per_agent index), then overwrites state. is_terminal in runtime_host.rs keeps allocation from reusing an archived or removed host. 18 production upsert_runtime_host calls in 4 files.
Runtime session membership RuntimeHostSessionState 3 runtime_host_sessions.state Guard at some sites. attach_runtime_session overwrites state on conflict. control_runtime_host_inner refuses host control when a member is not Attached or Dormant. Failed is never constructed.
Docker managed workspace WorkspaceLifecycleState 11 managed_workspaces.state Guard at some sites. The store checks immutable ownership only and overwrites state. ensure_workspace_is_recoverable in workspace.rs rejects Archived and Removed on the reconciliation and repository paths. Active, Parked, and Removed are matched but never assigned.
Host-native workspace allocation ManagedWorkspaceState 4 JSON inside host_sessions Guard at some sites. prepare_managed_workspaces recovers only Requested and Provisioning allocations and skips Orphaned ones. The stored JSON has no transition check.
Workspace operation WorkspaceOperationState 5 workspace_operations plus an event table Store validator. workspace_operation_event in jam-store/src/lib.rs accepts six arrows and appends an event. Pending is accepted but never written. WorkspaceOperationKind::Reset is never constructed.
Runtime host cleanup RuntimeHostCleanupStage 4 runtime_host_cleanup_operations plus an event table Store validator. runtime_host_cleanup_event rejects backward stages and changed immutable fields.
Checkpoint payload deletion ProviderCheckpointPayloadDeletionStage 3 provider_checkpoint_payload_deletions Store validator. validate_provider_checkpoint_payload_deletion_transition, a second hand-written forward-only validator.
Provider resume observation ProviderResumeObservationOutcome 2 provider_resume_observations, one row per attempt Store validator. Rows are immutable; a changed insert is rejected.
Private lane task WorkStatus 4 work_items None. apply_update in task_engine.rs sets the requested status directly.
Shared board task WorkStatus per assignee plus overall 4 room_tasks or Band None. The rollup is computed. No transition validator was found.
Board operation dispatch RoomTaskOperationState 5 room_task_operations Store validator. validate_room_task_operation_transition in jam-store/src/lib.rs.
Runtime permission RuntimePermissionStatus 2 Memory in the manager Guard at some sites. The decision path refuses an already-resolved request.
Usage reconciliation UsageReconciliationState 3 usage_reconciliation Owning type. stage, accept, and conflict methods on UsageReconciliationRecord.
Runtime event ingress RuntimeEventIngressState 5 Memory in jam-host Owning type. Owned by RuntimeEventIngress; every transition is in one file.
Configuration change PendingChangeStatus plus generations 4 Columns and JSON on host_sessions Guard at some sites. Generation fences in runtime_txn.rs, duplicated in memory by RuntimeTxnState. Applied is never written; a successful apply clears the change instead.
Lane state LaneState 4 Declared only None. No production code uses it. It is re-exported from jam-domain. Lane state is implied by table membership. The whole enum.

#A. Lifecycle rules

Where do the rules for each state machine live, and who enforces them?

#What the code does today

  • The inventory below covers 18 lifecycles. Five have their transitions checked in the store and two by their owning type. Seven rely on guards at some call sites, and four have no check at all. The state machine inventory lists each one.
  • The store's forward-only validators for runtime-host cleanup and checkpoint payload deletion are two hand-written copies of the same rule.
  • Seven declared states or kinds are never written by production code, and LaneState has no production use.
  • WorkspaceLifecycleState shares eight state names with RuntimeHostState. The three it never assigns (Active, Parked, Removed) are process states that do not apply to a directory.
  • The atlas's own state machine diagrams are hand-drawn from write sites and marked Inferred where the code has no table to read.

#Options

Option Extension Simplicity Safety Low risk Slices Effort Weighted
A0 Keep per-machine rules 1 1 2 5 5 5 2.6
A1 Prune, then declare a table per machine 2 3 4 4 5 4 3.5
A2 One lifecycle declaration generates the table, the check, and the diagram (recommended) 3 4 5 4 4 3 3.9
A3 Typestate records 2 2 4 2 2 1 2.4
A4 Event-sourced lifecycles 3 2 4 1 2 1 2.4

#A0. Keep per-machine rules (current design)

Each lifecycle keeps its own validator, guard, or no check. Guards are added when a bug appears.

What happens if nothing changes

  • Every new lifecycle chooses its own enforcement style. Illegal transitions on the eleven machines without a store or type-level check stay possible and are found in review or production.
  • Unused states keep appearing in SQL CHECK constraints and generated TypeScript, where readers take them as real.

Scores: Free today; the cost is paid in each future bug.

#A1. Prune, then declare a table per machine

Delete unused states, give every durable lifecycle a const transition table beside its enum, and check it in the store the way workspace operations already are.

What changes

  • Remove LaneState, RuntimeHostSessionState::Failed, the unassigned WorkspaceLifecycleState variants, PendingChangeStatus::Applied, WorkspaceOperationState::Pending, and WorkspaceOperationKind::Reset, with migrations for the SQL CHECK constraints.
  • Add an explicit table beside each enum and call it from every store upsert that writes that enum.
  • Add a test per machine that walks every write site's transition through the table.

What it fixes

  • Closes the unchecked transitions on runtime hosts, sessions, and workspaces.
  • Removes states that the store, TypeScript, and the atlas advertise but nothing writes.

Costs and risks

  • Removing a variant from a persisted enum needs a migration and an old-row decode plan; Failed sessions or Pending operations may exist in old stores.

  • Wrong if: A removed state turns out to be written by a path the production-only count missed, such as a test seam that ships in a bundle.

  • First step: Delete LaneState and WorkspaceOperationKind::Reset, which have no production references, and add a table for RuntimeHostState.

  • Sequencing: First: foundations other choices build on.

Scores: Each machine is an independent pull request, so risk and slicing score high.

#A2. One lifecycle declaration generates the table, the check, and the diagram

A lifecycle! macro in jam-domain declares states, legal arrows, and terminal states once. It generates can_transition, the store check, and the Mermaid diagram the atlas renders.

What changes

  • Includes everything in A1.
  • One macro produces the enum, the transition table, a validate_transition used by both store backends, and an exported Mermaid source.
  • The forward-only stage validators for cleanup and payload deletion become two declarations of the same macro.
  • The atlas's state-machine chapter reads the generated diagrams instead of hand-drawn ones.

What it fixes

  • Every durable lifecycle is enforced the same way, including the eleven that today rely on guards or nothing.
  • The atlas's lifecycle diagrams can no longer drift from code.

Costs and risks

  • The table must come from every current write site, not from the atlas diagram, which shows main paths only.
  • Macros that emit serde and specta derives can slow compile times and confuse rust-analyzer if written carelessly.

Declaration shape. The arrows below are the atlas's main paths; the real table must be derived from every write site.

lifecycle! {
    pub enum RuntimeHostState {
        Planned      -> Provisioning,
        Provisioning -> Ready | Failed,
        Ready        -> Active | Parked | Recovering,
        Active       -> Parked,
        Parked       -> Active | Recovering | Archived | Removed,
        Recovering   -> Ready | Failed,
        Archived     -> Removed,
        terminal Failed, Removed,
    }
}

// Generated:
//   RuntimeHostState::can_transition(from, to) -> bool
//   RuntimeHostState::TRANSITIONS: &[(Self, Self)]
//   RuntimeHostState::mermaid() -> &'static str
  • Wrong if: More than a few write sites need to bypass the table, which would mean the states model more than one concern and should be split instead.
  • First step: Declare RuntimeHostCleanupStage and ProviderCheckpointPayloadDeletionStage with the macro and delete their two hand-written validators.
  • Sequencing: First: foundations other choices build on.

Scores: Highest safety in the area; the macro is the only new abstraction.

#A3. Typestate records

Encode each state as a type parameter, for example RuntimeHost<Parked>, so an illegal transition does not compile.

What changes

  • Record types become generic over their state; transitions are methods that consume one type and return another.

What it fixes

  • Illegal transitions inside one function fail to compile.

Costs and risks

  • Records are loaded from SQLite with an arbitrary stored state, so every load still needs a runtime check and a match that returns one of several types.

  • Most transitions happen across await points and between separate calls, where the type is erased into an enum again.

  • First step: Prototype on RuntimeEventIngressState, the one machine owned entirely in memory by one type.

  • Sequencing: Next: the main restructuring.

Scores: Strong guarantees where they matter least: most of Jam's machines are persisted and written by several callers.

#A4. Event-sourced lifecycles

Every lifecycle becomes an append-only event log with a reducer, as workspace operations and runtime-host cleanup already half are.

What changes

  • Current-state rows become projections of events.
  • Recovery replays events instead of reading a summary row.

What it fixes

  • Full history for every lifecycle.
  • One recovery model for all machines.

Costs and risks

  • Migrating live rows to event histories for hosts, sessions, and workspaces.

  • Projection rebuilds at startup add to a recovery path that already runs ten ordered steps.

  • Works best with B2, G2.

  • Wrong if: Nobody needs the history of a lifecycle beyond the operation journals that already keep one.

  • First step: Only after A2: make one lifecycle, runtime hosts, a projection of its events, and measure startup time.

  • Sequencing: Later: needs earlier waves or a proven prototype.

Scores: History is valuable for operations, not for every lifecycle.

#Recommendation

  • Selected: A2, One lifecycle declaration generates the table, the check, and the diagram.
  • Strongest alternative: A1, Prune, then declare a table per machine.
  • Trade-off: A macro adds one shared abstraction every lifecycle depends on. A1 gets most of the safety without it but leaves each table hand-written in its own style.
  • Evidence floor: Eleven lifecycles with no store or type-level transition check, four of them with no check at all, and two duplicated forward-only validators.
  • Reversal trigger: If deriving a table from the current write sites of RuntimeHostState or PeerState needs more than a few exceptions, stop at A1 and keep hand-written tables.

#B. Durable operations and recovery

How do multi-step operations record intent and resume after a crash?

#What the code does today

  • The durable journals table lists 12 journals. Each has its own storage (usually a table; configuration changes use host_sessions columns and Docker environment cleanup uses owner markers on disk), its own validation, and its own replay step.
  • Manager::rebuild calls about ten recovery steps in a hand-written order. Adding a journal means finding the right line in that function.
  • A running archive or export deletion stores its retention, force flag, or target inside the result_class string, which replay parses back (parse_archive_running_class, parse_export_delete_running_class).
  • The Docker create-rollback record lives in std::env::temp_dir(), outside the app directory, where an OS temp purge can remove it.
  • Pending questions and permissions exist only in memory, while their inbox projections are durable.

#Options

Option Extension Simplicity Safety Low risk Slices Effort Weighted
B0 Keep one table and replay function per operation 1 1 3 5 5 5 2.9
B1 Typed operation journal library with a replay registry (recommended) 4 4 4 4 4 3 3.9
B2 One operations table for every journal 4 4 3 2 2 2 3.1
B3 Embedded durable workflow engine 4 2 3 1 1 1 2.3

#B0. Keep one table and replay function per operation (current design)

Each operation keeps its own schema, validator, and replay step, placed by hand in rebuild.

What happens if nothing changes

  • Each new destructive action copies the pattern: a stage enum, a validator, a replay step, and a manual choice of where it runs in startup.
  • Intent that does not fit a column keeps being packed into result_class strings.

Scores: Safe today because each journal is tested; the cost is repetition.

#B1. Typed operation journal library with a replay registry

One library supplies the journal record, typed intent, stage ordering, event append, and idempotent step runner. Each operation registers a replayer with a declared recovery phase; rebuild runs the registry.

What changes

  • A generic journal type with a typed intent payload replaces result_class parsing.
  • Stage tables come from the lifecycle declarations in area A.
  • Replayers declare a phase such as Destructive, Continuity, Workspace, or Usage; rebuild runs phases in order instead of hand-placed calls.
  • Tables stay per operation, so no data migration is needed.
  • The create-rollback record moves under the app directory.

What it fixes

  • A new destructive action is a stage declaration, an intent type, and a replayer, with its startup position declared rather than chosen by editing rebuild.
  • Replay no longer depends on string formats.

Costs and risks

  • Touches every crash-recovery path. The lifecycle and restart suites must pass unchanged at each step.

  • Requires A1 or A2 or A4.

  • Wrong if: The journals differ in ways a shared runner cannot express, such as needing different transaction boundaries per stage.

  • First step: Move runtime-host cleanup and checkpoint payload deletion, the two forward-only journals, onto the library together.

  • Sequencing: Next: the main restructuring.

Scores: Keeps tables and replay semantics, so risk stays moderate.

#B2. One operations table for every journal

Replace the per-operation tables with one operations table and one operation_events table keyed by kind.

What changes

  • Includes the library from B1.
  • Migrate in-flight rows of every journal into the shared tables.

What it fixes

  • One place to inspect every operation in flight.
  • One schema for new operations.

Costs and risks

  • Migrating destructive journals while an upgrade may interrupt one mid-flight.

  • One table mixes retention rules: workspace operations are kept, usage outbox rows are pruned.

  • Requires A1 or A2 or A4.

  • Wrong if: Operators never need a cross-operation view; per-table queries already serve the desktop.

  • First step: Only after B1 has run in production for a release.

  • Sequencing: Later: needs earlier waves or a proven prototype.

Scores: A small simplicity gain over B1 for a real migration hazard.

#B3. Embedded durable workflow engine

Run multi-step operations as durable workflows, in the style of Temporal or Restate, with automatic replay of each step.

What changes

  • Operations become workflow functions; the engine persists step results and replays them.

What it fixes

  • Replay and retry are handled by one engine.

Costs and risks

  • Workflow code must be deterministic across releases. A workflow in flight during an upgrade must replay against new code.

  • Adds an engine and its storage to a local daemon whose destructive journals record between two and four stages each.

  • Sequencing: Later: needs earlier waves or a proven prototype.

Scores: Included as a challenger: it shows that versioning in-flight work across releases is the hard constraint.

#Recommendation

  • Selected: B1, Typed operation journal library with a replay registry.
  • Strongest alternative: B0, Keep one table and replay function per operation.
  • Trade-off: B1 moves every journal onto one library, which is a change to crash-recovery code, the least forgiving code in Jam. B0 costs nothing now but repeats the pattern for every new destructive action.
  • Evidence floor: 12 journals with per-journal validators and replay, and operation intent stored in parsed strings.
  • Reversal trigger: If moving the two forward-only journals onto the library changes any replay test in lifecycle.rs, keep per-journal code and only add the replay registry.

#C. Harness boundary

What does adding a coding-agent harness touch, and where do the facts about each harness live?

#What the code does today

  • Adding OpenCode changed 25 files (1,987 inserted and 94 deleted lines, tests and generated bindings included) across seven Rust packages, the desktop, and docs, although it reused the existing ACP adapter.
  • The transport wire names are hand-written in five places (serde, store, manager parse, manager display, CLI display) with different alias sets. Serde decodes an unknown name to Unknown, the store decodes leniently with a one-time warning, and only the manager parser returns an error.
  • build_host in jamd.rs is an ordered if chain; putting a new branch in the wrong position silently gives an owned session mailbox delivery.
  • 188 production HostTransport:: references sit outside jam-host, 95 of them in manager.rs. jam-manager also names concrete adapter modules: jam_host::claudecode:: 31 times, copilot:: 12, opencode:: 10, codex:: 5.
  • The transport_catalog! macro already works as a manifest, and the desktop derives its picker and settings from it.

#Options

Option Extension Simplicity Safety Low risk Slices Effort Weighted
C0 Keep the enum, the catalog, and the checklist 1 1 3 5 5 5 2.9
C1 One codec and catalog-owned rules 3 3 4 4 5 4 3.7
C2 Compile-time harness registry (recommended) 5 4 4 3 3 2 3.8
C3 ACP as the only owned protocol 5 4 2 1 2 2 2.9

#C0. Keep the enum, the catalog, and the checklist (current design)

A new harness follows the 12-step checklist in Extension seams, three steps of which fail silently.

What happens if nothing changes

  • Each harness repeats the OpenCode change shape: the same facts registered in about a dozen places.
  • The three silent steps (dispatch position, native task source, name decoding) keep failing only at runtime.

Scores: Works, at about 25 files per harness.

#C1. One codec and catalog-owned rules

Keep the enum, but make the catalog own every per-transport fact: one name codec with aliases, the default provider, thread-setting rules, and the adapter key build_host dispatches on.

What changes

  • One TransportCodec generated from the catalog replaces the five name maps, and unknown names become typed errors instead of Unknown.
  • default_owned_runtime_provider, clear_unsupported_thread_settings_for_transport, and validate_owned_runtime_configuration read catalog fields.
  • build_host looks up an adapter key from the catalog row instead of testing predicates in order.
  • Desktop copy (providerCardCopy, transportLabel) moves into catalog fields.

What it fixes

  • Removes all three silent checklist steps.
  • Cuts the manager's transport matching roughly to probes and capability checks.

Costs and risks

  • The store's legacy aliases (app-server, codex, copilot) must keep decoding old rows.

  • Wrong if: The per-transport rules in the manager turn out to depend on session state that a static catalog row cannot express.

  • First step: Generate the name codec from the catalog and delete transport_str, TRANSPORTS, and parse_runtime_transport's table.

  • Sequencing: First: foundations other choices build on.

Scores: Most of the safety gain for little risk; the harness still spans several crates.

#C2. Compile-time harness registry

Each harness is one module or crate that implements a Harness trait bundling its catalog row, adapter constructor, probe, model listing, native task source, settings validation, and usage mapping. jamd registers the list.

What changes

  • Includes C1.
  • HostTransport becomes a validated identifier owned by the registry; stored strings stay the same.
  • build_host, work_provider, the probe paths, and the model listing ask the registry.
  • The manager asks capability questions ("does this runtime support compaction?") instead of matching transports.
  • Desktop copy and icons come from registry fields through the existing catalog command.

What it fixes

  • A new harness is its own module plus one registration line.
  • Moves the 95 transport references in manager.rs behind capability calls.

Costs and risks

  • Generated TypeScript loses the exhaustive HostTransport union that currently forces desktop copy for each variant; the registry must supply that copy instead.
  • Analytics enums still need a privacy review per harness under docs/analytics.md.
  • Works best with D1, D2, F1.
  • Wrong if: Attached harnesses cannot fit the same trait, because attached delivery depends on peer.host and session.provider, not only on the transport.
  • First step: Build the registry with OpenCode as its only member, keep the enum for the rest, and measure what still leaks.
  • Sequencing: Next: the main restructuring.

Scores: The largest extension payoff among options that keep every adapter.

#C3. ACP as the only owned protocol

Run every owned harness through the Agent Client Protocol adapter and retire the bespoke Codex, Copilot, and Claude Code adapters.

What changes

  • AcpHost becomes the one owned adapter; each harness is an ACP command line plus a profile.
  • The Codex, Copilot, and owned Claude Code modules leave jam-host: about 29,400 production lines and 29,100 test lines.

What it fixes

  • A new harness with an ACP adapter needs configuration, not code.

Costs and risks

  • The capability matrix shows generic ACP with no live usage, no model list, and questions only for Cursor. Codex dynamic tools, exact provider resume, and Copilot usage would be lost or need ACP extensions.

  • Places a third-party adapter process between Jam and the stock CLI, which the "stock provider executables" rule does not cover.

  • Conflicts with D2.

  • First step: Run the packaged real-agent journeys against an ACP adapter for one harness and list every capability that fails.

  • Sequencing: Later: needs earlier waves or a proven prototype.

Scores: Included as a challenger: the product's capability matrix is per provider, so a single protocol costs features.

#Recommendation

  • Selected: C2, Compile-time harness registry.
  • Strongest alternative: C1, One codec and catalog-owned rules.
  • Trade-off: C2 turns HostTransport from a closed enum into registered identifiers, which touches persistence and wire decoding. C1 keeps the enum and gets about half the benefit.
  • Evidence floor: The OpenCode change set, the five hand-written name maps with inconsistent unknown-name handling, and the ordered dispatch chain.
  • Reversal trigger: If moving OpenCode into a registered harness still needs edits outside its module beyond one registration line and desktop copy, stop at C1.

#D. Owned runtime kernel

How much of an owned adapter is shared code, and how much is protocol translation?

#What the code does today

  • Each of the four owned adapters defines its own EventContext. Codex, Copilot, and Claude Code keep separate copies of active-message tracking and disposition staging, kept in agreement by parallel tests.
  • Instruction assembly is repeated per adapter; the ## Operator briefing header is a literal string in three files.
  • Provider-neutral code lives inside the Codex module: SharedTurnCoordinator, SharedHostRegistry, and the task tool names in codex::task_mcp are used by Copilot, ACP, and Claude Code.
  • Shared placement has two mechanisms: Codex's SharedSdkRuntimeHost and the SharedSandboxLease the other adapters use.
  • The adapters are large. Production lines, excluding tests: codex/app_server.rs 10,034, copilot/sdk.rs 4,818, acp.rs 4,534, claudecode/owned/runtime.rs 4,172. OpenCode is a mode of AcpHost behind 26 production references to is_opencode, opencode_sandbox, and opencode_uses_default_command_pair.

#Options

Option Extension Simplicity Safety Low risk Slices Effort Weighted
D0 Keep one self-contained adapter per protocol 1 1 3 5 5 5 2.9
D1 Extract the shared kernel pieces (recommended) 3 4 4 4 5 3 3.9
D2 Generic runtime engine with narrow provider drivers 5 4 4 2 2 1 3.4

#D0. Keep one self-contained adapter per protocol (current design)

Each adapter owns its turn loop, staging, instructions, sandbox preparation, and checkpointing.

What happens if nothing changes

  • Settlement and staging rules stay shared by convention; a fix in one adapter must be repeated in the others.
  • New adapters copy the largest existing one.

Scores: Each adapter is tested on its own, so safety holds today.

#D1. Extract the shared kernel pieces

Move what is already duplicated into shared jam-host modules: a turn ledger for active messages and disposition staging, an instruction builder, one shared-host registry and lease, and the checkpoint coordinator.

What changes

  • A shared turn ledger replaces the three staging implementations and gives ACP explicit staging.
  • One instruction builder with provider sections replaces four assembly paths.
  • SharedSdkRuntimeHost and SharedSandboxLease converge on one mechanism.
  • Neutral types move out of codex:: into their own modules.

What it fixes

  • Settlement rules live in one place and one test suite.
  • Adapters shrink toward protocol translation.

Costs and risks

  • Merging the two shared-host mechanisms touches Shared placement for Codex, the one path the credential-backed Shared live gate covers.
  • Wrong if: The staging rules differ between providers for a product reason rather than by accident.
  • First step: Move SharedTurnCoordinator, SharedHostRegistry, and the task tool names out of codex:: with no behavior change.
  • Sequencing: First: foundations other choices build on.

Scores: Each extraction is small and behavior-preserving.

#D2. Generic runtime engine with narrow provider drivers

One OwnedRuntime engine owns the turn loop, settlement, compaction policy, checkpointing, and sandbox lifecycle, and calls a narrow ProviderDriver trait for start, send, interrupt, and event mapping.

What changes

  • Includes D1.
  • Each adapter becomes a driver of a few thousand lines.
  • Retry ladders and compaction fallbacks move into the engine with per-driver parameters.

What it fixes

  • A new owned protocol is a driver, not a copy of a 10,000-line adapter.
  • One implementation of every turn-level invariant.

Costs and risks

  • Every owned provider's turn loop changes at once. Only the real-provider journeys and live gates prove it, and some consume subscription quota.

  • Provider differences, such as Copilot's idle during an outstanding question, must fit the driver interface.

  • Requires C1 or C2.

  • Conflicts with C3.

  • Wrong if: D1 shows that the remaining turn loops share little structure.

  • First step: Only after D1: port OpenCode's ACP driver onto the engine first, then Copilot.

  • Sequencing: Later: needs earlier waves or a proven prototype.

Scores: Large payoff, but it moves the most provider-specific code in the repository.

#Recommendation

  • Selected: D1, Extract the shared kernel pieces.
  • Strongest alternative: D2, Generic runtime engine with narrow provider drivers.
  • Trade-off: D2 yields the smallest adapters but moves every turn loop at once. D1 extracts only what is already duplicated and shows how much variation remains.
  • Evidence floor: Four EventContext copies, three staging implementations, and two shared-host mechanisms.
  • Reversal trigger: If after D1 the four adapters still share most of their turn-loop structure, schedule D2; if they diverge, stop.

#E. Manager structure and concurrency

Who owns each piece of manager state, and what serializes changes to it?

#What the code does today

  • manager.rs has 51,465 production lines. Manager has 104 fields, about 56 of them independent locks.
  • The Inner mutex (29 fields) holds lifecycle state together with permissions, questions, task access, OAuth sessions, and account tasks, so one no-reentry rule couples all of them.
  • runtime_room_bind_gate serializes room materialization, wake, removal, restart, template apply, reaping, and shutdown across every agent. manager.rs takes it at 23 sites.
  • save_peer_with_connectivity has 49 call sites, and each rewrites every host-session row for the peer. A stale peer clone can overwrite a concurrent change.
  • WorkspaceManager::new is constructed at 23 production sites, so workspace calls have no single owner.

#Options

Option Extension Simplicity Safety Low risk Slices Effort Weighted
E0 Keep one manager with shared locks 1 1 2 5 5 5 2.6
E1 Subsystem services with their own state (recommended) 3 4 4 3 4 2 3.5
E2 One actor per agent 3 3 4 1 2 1 2.6

#E0. Keep one manager with shared locks (current design)

State stays in Manager and Inner; new features add fields and gates to the same struct.

What happens if nothing changes

  • manager.rs was the most-changed file in the last 90 days, so its merge conflicts and review load keep growing.
  • Every new subsystem inherits the no-reentry rule of a lock it does not need.

Scores: Wholesale peer writes are a live hazard, which keeps safety at 2.

#E1. Subsystem services with their own state

Split the manager into services that each own their state and lock (accounts, permissions, questions, runtime lifecycle, workspaces, usage). Manager becomes a facade that implements Control by delegation. Host-session writes become per row.

What changes

  • Permissions and questions leave Inner first; they already have their own brokers.
  • One WorkspaceService owns the workspace manager instead of 23 constructions.
  • A per-row host-session upsert replaces save_peer_with_connectivity's wholesale rewrite.
  • RuntimeTxnState reads the durable generations instead of duplicating them.

What it fixes

  • Unrelated subsystems no longer share a lock or a file.
  • Removes the stale-clone overwrite hazard.

Costs and risks

  • Some code may rely on reading two subsystems' state under one lock; each split must find and preserve that ordering.

  • lifecycle.rs, the 68,000-line safety net, calls the manager directly and moves with it.

  • Works best with G1, G2.

  • First step: Move the permission broker's state out of Inner behind its existing methods.

  • Sequencing: Next: the main restructuring.

Scores: Each service can move in its own pull request behind the facade.

#E2. One actor per agent

Each agent's lifecycle, host sessions, bindings, and workers belong to a task with a mailbox. Commands for one agent are serialized in its mailbox; different agents proceed in parallel.

What changes

  • Replaces runtime_room_bind_gate and the per-peer lifecycle gates with per-agent mailboxes.
  • Shared runtime hosts, which span rooms of one agent, fit inside one actor.

What it fixes

  • One slow agent's bind or teardown no longer blocks every other agent.
  • "Never hold the lock across .await" stops being a rule for lifecycle code, because the actor owns its state.

Costs and risks

  • Changes the concurrency model of the daemon. Cross-agent operations, such as a room with two local agents, need a protocol between actors.

  • Shutdown and restart ordering must be rebuilt on the actor model.

  • Requires H1 or H2.

  • Works best with G2.

  • Wrong if: Measurements show the global gate is not a latency problem for real multi-agent use.

  • First step: Measure how long runtime_room_bind_gate is held during Docker materialization with several agents before choosing this.

  • Sequencing: Later: needs earlier waves or a proven prototype.

Scores: Included as a challenger: it tests whether the global gate is the real bottleneck.

#Recommendation

  • Selected: E1, Subsystem services with their own state.
  • Strongest alternative: E2, One actor per agent.
  • Trade-off: E2 removes the global gate by giving each agent its own serialized owner, but it changes the concurrency model of the whole daemon. E1 keeps the model and separates state by subsystem.
  • Evidence floor: One mutex shared by unrelated subsystems, 49 wholesale peer writes, and one gate for every agent's room lifecycle.
  • Reversal trigger: If splitting permissions and questions out of Inner exposes ordering that depends on the shared lock, keep them together and split along another boundary first.

#F. Daemon API surface

What does one new daemon operation cost, and what guarantees that every layer has it?

#What the code does today

  • A daemon operation crosses nine layers: Control, manager, wire types and route, daemon handler, client, jam-ipc, Tauri command, two isolation allowlists, and generated bindings.
  • The Control API crosswalk lists 255 methods, 253 daemon routes, and 250 desktop commands.
  • 243 of the 255 Control methods have default bodies, and 241 of those bodies return a "not supported" error, so a missing implementation compiles and fails at runtime.
  • The client (10 seconds) and daemon (45 seconds) disagree on the default unary deadline, and route exceptions are kept in two lists.

#Options

Option Extension Simplicity Safety Low risk Slices Effort Weighted
F0 Keep hand-written layers with parity tests 1 1 3 5 5 5 2.9
F1 One declaration generates every layer (recommended) 4 4 4 4 4 3 3.9
F2 One generic RPC endpoint 4 5 3 3 3 3 3.6
F3 Split Control into required capability traits 2 3 4 4 4 4 3.4

#F0. Keep hand-written layers with parity tests (current design)

Each operation is written in nine places; tests check bindings and allowlists.

What happens if nothing changes

  • The Control-to-route-to-client part has no compile-time check; the generated crosswalk is the only way to find gaps.

Scores: Parity tests make it safe, not cheap.

#F1. One declaration generates every layer

Declare each operation once (name, request, response, route, deadline class, streaming, desktop exposure) and generate the trait method, wire types, route, handler, client method, jam-ipc method, Tauri command, and both allowlists with a just recipe, as just bindings does today.

What changes

  • A checked-in declaration file and a generator run by just bindings.
  • Deadlines come from the declaration, so the client and daemon share one list.
  • Implementations are required methods on the manager side; defaults disappear.

What it fixes

  • A new operation is one declaration plus its manager implementation.
  • Client and daemon deadlines cannot disagree.

Costs and risks

  • Generated code across six crates must stay readable and reviewable.

  • Streaming routes, held waits, and wire-version gates need declaration fields, not special cases.

  • Works best with C2.

  • Wrong if: More than a handful of operations need hand-written exceptions in the generated layers.

  • First step: Generate the client and jam-ipc methods for one section of the crosswalk and compare them with the hand-written ones.

  • Sequencing: Next: the main restructuring.

Scores: Can migrate one section of the API at a time.

#F2. One generic RPC endpoint

Replace 253 routes with one POST /v1/rpc carrying a tagged ControlRequest enum; the client, daemon, and jam-ipc become generic dispatch.

What changes

  • One route, one handler, one client method; per-method metadata for deadlines and admission.
  • Tauri keeps typed commands or uses one command with a method allowlist.

What it fixes

  • Removes most per-operation code in four layers.

Costs and risks

  • Per-route logging, admission, timeouts, and the high-water tracker now key on the method name inside the body.

  • Old CLIs and desktops speak routes; the daemon must serve both during the wire-version transition.

  • First step: Serve the generic endpoint beside the existing routes for read-only methods.

  • Sequencing: Next: the main restructuring.

Scores: Simplest end state; loses the per-route handles operations already rely on.

#F3. Split Control into required capability traits

Break Control into traits such as RuntimeControl, RoomControl, and UsageControl with no default bodies, so a missing implementation fails to compile.

What changes

  • Control becomes a supertrait of capability traits.
  • Test fakes implement only the traits they need.

What it fixes

  • The 241 default bodies that fail at runtime with "not supported" become compile errors.

Costs and risks

  • Every fake Control in tests must implement or stub its traits explicitly.

  • First step: Split off the usage methods, which have one real implementation and few fakes.

  • Sequencing: First: foundations other choices build on.

Scores: Cheap safety; leaves the boilerplate.

#Recommendation

  • Selected: F1, One declaration generates every layer.
  • Strongest alternative: F3, Split Control into required capability traits.
  • Trade-off: F1 adds a code generator to the build. F3 is plain Rust and catches missing implementations, but leaves the nine hand-written layers.
  • Evidence floor: Nine layers per operation, 241 default bodies that fail at runtime, and two deadline lists.
  • Reversal trigger: If the generator cannot express held-open streams and per-route admission without special cases, keep hand-written streaming routes and generate only unary ones.

#G. Events and client state

How do clients learn about changes, and what do they re-read after missing one?

#What the code does today

  • All peers and human feeds share one broadcast channel of 256 events (CHANNEL_CAP in fanout.rs). A burst lags every slow subscriber at once, and each must re-hydrate after StreamResync.
  • There are 59 event kinds; the desktop handles 57.
  • The desktop store daemon.ts is 7,554 lines with a 57-case reducer, 27 hydrators, and 63 module-level bindings outside Zustand, many of them generation fences and replay buffers that decide which response is newest.
  • EventKind::ContactsChanged is also used as an internal refresh signal between the engine and the worker.

#Options

Option Extension Simplicity Safety Low risk Slices Effort Weighted
G0 Keep one channel and hand-written hydration 1 1 3 5 5 5 2.9
G1 Topic channels with slice revisions (recommended) 3 3 4 4 4 3 3.5
G2 Revisioned slice sync with a generic client cache 4 5 4 2 3 1 3.6

#G0. Keep one channel and hand-written hydration (current design)

Events stay best-effort hints on one channel; the desktop keeps per-slice hydrators and fences.

What happens if nothing changes

  • Every new slice adds a hydrator, reducer cases, and fences to daemon.ts.
  • Busy rooms cause resyncs for unrelated views.

Scores: Correct because clients re-hydrate; costly to extend.

#G1. Topic channels with slice revisions

Partition fan-out by topic (peer, room, usage, runtime) and stamp each event with the revision of the slice it changed, so a client re-reads only the slice it lagged on.

What changes

  • Fanout keeps one channel per topic with its own capacity.
  • Events carry (slice, revision); StreamResync names the slices to re-read.
  • Internal signals such as ContactsChanged move off the client event stream.

What it fixes

  • A burst in one room no longer resyncs every view.
  • Revisions replace most hand-written fence variables.

Costs and risks

  • The wire format changes, so the desktop and CLI must accept both shapes during rollout.

  • First step: Split presence and room-message events onto their own topics and measure resync counts.

  • Sequencing: Next: the main restructuring.

Scores: Keeps the events-are-hints model and narrows its cost.

#G2. Revisioned slice sync with a generic client cache

The daemon owns versioned state slices. A client subscribes with its last revision and receives deltas or a snapshot. The desktop store becomes a generic cache keyed by query, invalidated by slice tags.

What changes

  • Includes G1.
  • Daemon services publish slice snapshots and deltas.
  • Hydrators, reducer cases, and fences in daemon.ts give way to one cache layer.

What it fixes

  • A new desktop surface declares the slices it reads instead of writing sync code.
  • Reconnect cost is proportional to what changed.

Costs and risks

  • Every slice needs a daemon-side owner that can produce a revision, which depends on E1 or E2.

  • Large rewrite of the desktop's data layer and its IPC budget tests.

  • Requires E1 or E2.

  • First step: Only after G1: move one slice, usage reports, which already has a polling cache, onto revisioned sync.

  • Sequencing: Later: needs earlier waves or a proven prototype.

Scores: Simplest desktop at the end, with the largest rewrite.

#Recommendation

  • Selected: G1, Topic channels with slice revisions.
  • Strongest alternative: G2, Revisioned slice sync with a generic client cache.
  • Trade-off: G2 removes most hand-written desktop sync code but needs every daemon slice to carry a revision. G1 bounds the blast radius of a lag without changing the client model.
  • Evidence floor: One shared channel, and freshness fences hand-written as module-level bindings.
  • Reversal trigger: If slice revisions are cheap to add while doing E1, move to G2 for the slices that already have them.

#H. Runtime ownership graph

Is there one model of agent, runtime session, runtime host, workspace, and room binding for every owned runtime?

#What the code does today

  • There are two unrelated managed-workspace implementations: owned_workspace.rs with ManagedWorkspaceState for host-native runtimes, and workspace.rs with WorkspaceLifecycleState for Docker. A past backfill bug reminted Docker session UUIDs because of the shared name.
  • A Docker room binding is stored twice, as a HostSession row and as runtime_host_sessions plus runtime_session_room_bindings rows, written at different points of the bind. Host-native sessions have only the HostSession row.
  • The sandbox plan is derived by each adapter and again by the manager, which compares the two names as a guard.
  • Placement defaults disagree in five places: the Rust default, from_existing, the desktop create panel, the desktop edit path, and the CLI.

#Options

Option Extension Simplicity Safety Low risk Slices Effort Weighted
H0 Keep separate host-native and Docker models 1 1 2 5 5 5 2.6
H1 One ownership graph for every owned runtime (recommended) 4 5 5 2 3 2 3.9
H2 One binding table for all owned sessions 2 3 4 3 4 3 3.1

#H0. Keep separate host-native and Docker models (current design)

Host-native runtimes keep their HostSession-only model; Docker keeps runtime hosts and bindings beside it.

What happens if nothing changes

  • Every lifecycle change must update both binding representations for Docker and remember the difference for host-native.
  • The "One representation" section of AGENTS.md names this pattern as the root cause of recurring defects.

Scores: The duplicated binding is a known defect source.

#H1. One ownership graph for every owned runtime

Every owned runtime, host-native or Docker, has a runtime host (with a boundary kind of host or microVM), runtime sessions, a workspace record, and room bindings. HostSession rows stop carrying bindings.

What changes

  • Host-native runtimes get runtime-host records with a host boundary kind.
  • One workspace lifecycle serves both boundaries; owned_workspace.rs folds into the workspace service.
  • One sandbox-plan function, called by the manager and passed to the adapter.
  • One placement default read from the catalog.

What it fixes

  • One place to look up what an agent is running and where.
  • Removes the duplicated binding and the second workspace implementation.

Costs and risks

  • Migrates every existing host-native agent's records at upgrade.

  • The inactive-binding tombstone that lets a re-added room resume its exact provider thread must survive the move.

  • Works best with E2.

  • Wrong if: Host-native runtimes need none of what a runtime host records, making the records empty ceremony.

  • First step: Do H2 first: move bindings for all owned sessions onto runtime_session_room_bindings.

  • Sequencing: Later: needs earlier waves or a proven prototype.

Scores: The only option that removes both duplications.

#H2. One binding table for all owned sessions

Use runtime_session_room_bindings as the only room binding for every owned session. Keep the two workspace implementations for now.

What changes

  • Host-native owned sessions gain binding rows.
  • HostSession room fields become derived or are removed.

What it fixes

  • Removes the binding that is written twice.
  • Makes the tombstone rule uniform.

Costs and risks

  • A migration of binding data for every owned session.

  • First step: Write binding rows for host-native owned sessions alongside the existing fields, then switch reads.

  • Sequencing: First: foundations other choices build on.

Scores: The smallest change that removes a known defect source.

#Recommendation

  • Selected: H1, One ownership graph for every owned runtime.
  • Strongest alternative: H2, One binding table for all owned sessions.
  • Trade-off: H1 gives host-native runtimes the same host, session, and workspace records as Docker, which is a data migration for every existing agent. H2 fixes only the duplicated binding.
  • Evidence floor: Two representations of one binding, two workspace implementations, and a recorded UUID collision.
  • Reversal trigger: If host-native runtimes cannot adopt runtime-host records without changing their on-disk workspace layout under ~/Jam, stop at H2.

#Directions not included

  • Split jam-manager into more crates first. Moving code into crates without changing who owns which state keeps the shared lock and the wholesale writes. E1 splits along ownership, and crate boundaries can follow it.
  • Replace the encrypted SQLite store. No finding in the atlas traces a defect to the store engine. The problems are in what the store accepts, which area A addresses.
  • Move routing or queues onto Band. Delivery depends on a local durable queue before Band is told a message is processing. Moving it off the device breaks at-least-once delivery when the network drops.

#Planning handoff

After a direction is chosen, implementation planning under docs/engineering/planning.md must decide:

  • Which strategy, or which per-area picks, the team selects, recorded with the weights used.
  • For the selected lifecycle option: the transition table for each machine, derived from every production write site rather than from the atlas diagrams.
  • For the selected harness option: whether stored transport strings and the generated TypeScript union keep their current shape.
  • For any option that migrates data (A1 pruning, B2, H1, H2): the upgrade and downgrade behavior for stores written by the previous release.
  • Which existing tests are the acceptance bar for each step, starting with crates/jam-manager/tests/lifecycle.rs, the restart tests in bins/jam/tests/jamd_managed_restart.rs, and the packaged real-agent journeys.
Scroll to zoom, drag to pan.