Jam architecture main at 59f67f5d0 · 2026-09-29
View source

#Tasks, questions, permissions, plans, and artifacts

On this pageWhere it livesHow it worksTwo task layers and why they are separateLane identity and the dual-read chokepointThe authoritative write pathThe observed write pathLegacy and room-keyed writesThe shared board and its routingBoard task lifecycleCorrelation from Layer 1 to Layer 2Kickoff: assigning work versus sending a messageQuestionsPermissionsAttention and the InboxPlansArtifactsState and ownershipContractsControl methodsEventsTraits provided and consumedBand protocol consumedInvariantsFailure and recoveryExtension pointsAdding a coding-agent harnessAdding a feature in this subsystemRefactor notes

This subsystem turns an agent's private progress and its requests for human input into objects that people can see and act on from a room. It owns two task models that stay separate: the Layer 1 private lane, one per runtime session, which records what a single agent is doing, and the Layer 2 shared board, one per room, which records assignments that several participants can take. It also owns the manager's question and permission brokers, the durable attention rows behind the Inbox, the per-room plan document and its live node overlay, and the content-addressed artifact store.

It does not decide what an agent does, does not merge the two task layers into one model, and does not treat a room message as an answer or an approval unless it carries an exact correlation token from an authorized sender. It never approves a permission on an agent's behalf, never sends a local file link to a participant it cannot prove is local, and never moves a room's board between this computer and Band without durable evidence that the Band board is usable.

#Where it lives

Area Crate and module Key types and functions
Layer 1 domain model crates/jam-domain/src/workitem.rs WorkItem, WorkStatus, EffectiveTaskSource (Observed, JamAuthoritative), TaskWriteAdapter (McpStdio, AppServerDynamic, Cli), WorkItemSnapshot, RoomWorkLane, RoomWorkSnapshot, LaneState, LANE_RETENTION_MS, stamp_milestones, merge_trail
Layer 1 mutation engine crates/jam-domain/src/task_engine.rs TaskLane::apply, TaskOp (Create, Update, Delete), TaskCreate, TaskUpdate, TaskLink, TaskEngineError, TaskStatus, task_status, bounds such as MAX_LIVE_TASKS (256) and MAX_BATCH_OPS (256)
Authoritative lane writes crates/jam-manager/src/task_lane.rs TaskLaneCoordinator, TaskMutation, TaskMutationOutcome, TaskLaneError, mutate, reconcile_snapshot, drain_committed, validate_room_scope, MAX_REBASE_ATTEMPTS (3)
Observed lane writes crates/jam-manager/src/workitems.rs NativeWorkProjector (project, project_runtime_event, project_copilot_todos, retry_pending), lane_storage_room, room_runtime, peer_lane_runtime, stamp_after_save, RUNTIME_KEYED_LANES
Provider capture seam crates/jam-core/src/worksource.rs WorkItemSource (provider, snapshot, watch), WorkItemWatch, NullSource, WorkMode (JAM_WORKITEMS), CLAUDECODE_VOCAB, CODEX_VOCAB, COPILOT_VOCAB, map_status, parse_room_marker
Capture sources crates/jam-host/src/claudecode/mod.rs, crates/jam-host/src/codex/mod.rs ClaudeCodeSource (fs-notify on ~/.claude/tasks/…), CodexSource (polls ~/.codex/tasks/<session>.json)
Watch loop crates/jam-manager/src/workwatch.rs spawn_work_watch, poll_once, POLL_INTERVAL (2 s), WATCH_FALLBACK_INTERVAL (45 s), DEBOUNCE (150 ms)
Source factory bins/jam/src/jamd.rs; type in crates/jam-manager/src/lib.rs new_work_source closure, work_provider, NewWorkSource
Task tool surfaces bins/jam/src/mcp/tasks.rs, bins/jam/src/mcp/runtime_tasks.rs, crates/jam-host/src/runtime_task_tools.rs, crates/jam-manager/src/runtime_tasks.rs, bins/jam/src/main.rs TaskTools (jam mcp tasks), RuntimeTaskToolProvider, Manager::runtime_task_call, TaskCmd (jam task), WorkCmd (jam work)
Shared board port crates/jam-contract/src/lib.rs RoomBoardBackend (9 methods), BoardError, RoomTaskStatusChange, RoomTaskOperationView
Shared board domain crates/jam-domain/src/roomboard.rs RoomBoard, RoomTask, TaskAssignment, TaskCreator, RoomTaskOperation, RoomTaskOperationPayload, RoomTaskOperationState
Local board backend and correlation crates/jam-manager/src/roomboard.rs LocalRoomBoardBackend, rollup, next_uid, emit_room_board, correlate, correlate_typed_link
Board routing, offline journal, migration crates/jam-manager/src/bandboard.rs, bandboard/offline.rs, bandboard/migration.rs RoutedRoomBoardBackend, BandBoardBridge (11 methods), BoardBridgeSlot, RoomOwner, enqueue_and_dispatch, run_room_task_operation_replay, REPLAY_GATE, run_shadow_board_migration
Band board credential bridge crates/jam-manager/src/manager/board_bridge.rs impl BandBoardBridge for Manager, room_task_assign_human, room_task_create_human
Question broker crates/jam-manager/src/manager.rs, crates/jam-manager/src/questions.rs Manager::ask_room_question, try_answer_pending_question, question_intercept, PendingRoomQuestion, PendingQuestionCleanup, select_question_human, parse_correlated_question_answer, attach_question_to_board, resolve_question_on_board
Question entry points crates/jam-host/src/lib.rs, crates/jam-manager/src/worker.rs RuntimeQuestionBroker, RuntimeQuestionRequest, RuntimeQuestionAnswer, RuntimeHostContext::ask_question, QuestionGate
Permission model crates/jam-domain/src/permission.rs RuntimePermissionRequest, RuntimePermissionRecord, RuntimePermissionDecision, RuntimePermissionOutcome, RuntimePermissionRule, RuntimePermissionStatus
Permission broker crates/jam-manager/src/manager.rs request_runtime_permission, decide_runtime_permission, finish_runtime_permission, try_decide_pending_permission, post_room_permission_ask, PendingRuntimePermission, permission_operation_key, scoped_permission_outcome, permission_rule_lookup_keys, parse_permission_reply
Attention and Inbox crates/jam-domain/src/attention.rs, crates/jam-manager/src/roomfeed.rs, crates/jam-store/src/lib.rs AttentionCandidate, AttentionResolution, NotificationPreferences, apply_pending_attention, Manager::resolve_inbox_attention, Manager::cached_inbox_snapshot, store methods stage_attention_mutation, apply_direct_mention, bind_permission_attention, put_inbox_questions
Room plans crates/jam-domain/src/plan.rs, crates/jam-manager/src/plan.rs, plan_watch.rs, plan_gate.rs RoomPlan, PlanSource, PlanContent, NodeOverlay, NodeState, NodeStatus, LiveSourceState, read_local_plan, read_observed_plan, apply_focus, apply_status, apply_unfocus, PlanLiveSources, PlanWriteGates
Artifacts crates/jam-domain/src/artifact.rs, crates/jam-manager/src/artifacts.rs Artifact, ArtifactDeletePreview, attachments_travel, classify_locality, sanitize_artifact_name, store_artifact_bytes, FILE_MUTATION_LOCK, append_local_links, validate_managed_local_links, delete_row_and_maybe_file_after, sweep_orphan_dirs
Kickoff crates/jam-manager/src/kickoff.rs, crates/jam-manager/src/manager.rs derive_subject, kickoff_message, kickoff_message_with_collaborators, Manager::kickoff_work_seeded, KickoffSeed

Most of the manager-side behavior sits in crates/jam-manager/src/manager.rs, which is 58,145 lines. The chapter names the method so a reader can search for it.

#How it works

#Two task layers and why they are separate

A Layer 1 lane is one agent runtime's private task list. It holds WorkItem rows with a four-value WorkStatus, dependency edges (blocks, blocked_by), metadata, component references into the room's architecture diagram, and tombstones. A lane belongs to a RuntimeSessionId, not to a room. One runtime bound to two rooms has one lane, and each row records the room it projects into through WorkItem.origin_chat_id. A row with no origin is agent-private and projects into no room.

A Layer 2 board is one room's shared list of RoomTask cards. Each card has an opaque uid, a subject and detail, and a list of TaskAssignment stakes, one per agent that took it. The card's overall status is computed from the stakes by roomboard::rollup. The board belongs to a room, and for Band rooms it lives on Band.

The two layers meet at one point only: correlation. A private task may carry a link to one card (WorkItem.linked_chat_id plus linked_room_task_uid, or the older [#uid] subject marker), and roomboard::correlate copies that task's status onto the agent's own stake on that card. Nothing flows the other way. A board change never edits a private lane.

The writers that feed a Layer 1 lane, and which of the two write paths each uses:

A lane has exactly one effective writer model at a time, recorded durably as EffectiveTaskSource in task_lane_source:

  • Observed { provider }: the lane mirrors a provider's own task store. Snapshots arrive as full lists and are merged with merge_trail, which keeps a task the provider stopped reporting and marks it Completed.
  • JamAuthoritative { adapter }: Jam's own task service owns the lane. Writes are engine operations, not snapshots, and bypass merge_trail.

The first authoritative write flips the source inside the same store transaction as the rows it writes (Store::commit_task_aggregate). From then on the observed path stands down: NativeWorkProjector::project_locked checks the source before saving and returns Ok(None), and the store rechecks inside its own transaction (refuse_if_lane_authoritative in crates/jam-store/src/sqlite.rs) and returns StoreError::LaneAuthoritative. Tests workitems::observed_watcher_stands_down_once_the_lane_is_authoritative and workitems::authoritative_flip_between_precheck_and_save_cannot_clobber cover both checks.

#Lane identity and the dual-read chokepoint

Lanes were originally keyed by (profile, scope, room). They are now keyed by runtime session, but the storage row still sits under one deterministic room key. workitems::lane_storage_room is the single place that resolves a room to its lane row:

  1. room_runtime(peer, room) finds the host session bound to the room. A daemon-owned session supplies its minted RuntimeSessionId. A lease-backed attached session (every session has an empty runtime ID) uses peer_lane_runtime, a stable hash of profile and scope prefixed with peer-.
  2. Store::resolve_lane_room returns the storage room stamped with that runtime, if any.
  3. Otherwise the room itself is the key.

RUNTIME_KEYED_LANES is a compile-time constant set to true. When it is true, jamd runs Manager::migrate_lanes_to_runtime_keys before Manager::rebuild at startup (bins/jam/src/jamd.rs). The migration stamps single-binding lanes, deduplicates equivalent collision groups, archives divergent groups and starts them empty, and leaves unresolvable lanes room-keyed so reads keep working (Store::migrate_lanes_to_runtime_keys doc comment). Every live writer then calls stamp_after_save so new lanes are runtime-keyed from their first write.

Room reads apply the projection filter. Manager::list_work_items_with_source returns only rows whose origin_chat_id equals the room and hides tombstones when the room is runtime-bound (retain_room_projection). A room with no runtime binding returns its room-keyed row unchanged. Test lane_tasks.rs::named_room_context_cannot_read_or_mutate_sibling_or_private_tasks covers the filter, and lane_tasks.rs::list_work_items_projects_a_runtime_lane_per_room covers one runtime serving two rooms.

#The authoritative write path

TaskLaneCoordinator::mutate is the only entry point for engine operations on a lane. All four doors reach it: the lane_task_* Control verbs used by jam task and jam mcp tasks, the scoped runtime_task_call used by sandboxed and Shared runtimes, and set_work_items when the room is runtime-bound (through reconcile_snapshot).

The order in which one authoritative batch loads, validates, commits, and publishes, and where a crash changes the outcome:

The details that make this safe:

  • Concurrency. The gate is a tokio::sync::Mutex per runtime, held in a std::sync::Mutex<HashMap> registry. The registry lock is dropped before awaiting the lane lock, so unrelated runtimes proceed. reclaim_gate removes an entry when Arc::strong_count is 1, which proves no holder and no waiter (test task_lane::create_persists_one_task_and_reclaims_the_gate).
  • Writers outside the gate. The observed path does not take the coordinator's gate. The optimistic expected_generation covers it: a commit that loses the race gets StoreError::StaleGeneration, reloads, and rebases up to MAX_REBASE_ATTEMPTS (3). Exhausting the budget is a hard TaskLaneError::Store, never a silent drop (tests task_lane::stale_generation_forces_a_rebase_that_then_succeeds, task_lane::rebase_exhaustion_is_a_bounded_hard_error).
  • Atomicity. TaskLane::apply works on a scratch copy, validates every bound and the dependency graph, and replaces the lane only when the whole batch passes. The store commit is one transaction. A rejected batch leaves nothing behind.
  • Idempotency. An optional key of at most TASK_IDEMPOTENCY_KEY_MAX_CHARS (128) characters is stored with a SHA-256 hash of the mutation (hash_mutation). The same key with the same hash replays the recorded task IDs and the current lane; the same key with a different hash is IdempotencyConflict, mapped to ControlError::Conflict (tests task_lane::same_key_same_payload_replays_without_double_applying, task_lane::same_key_different_payload_conflicts). jam mcp tasks mixes a per-process nonce into the key it derives from the JSON-RPC request ID (idempotency_key in bins/jam/src/mcp/tasks.rs). A bare interactive jam task create has no key and is documented as not retry-safe.
  • Room scope. When a caller names a room, validate_room_scope rejects any operation on a task or edge outside that room's projection. It runs under the gate on every rebase attempt, because the lane may have changed since the manager's preflight (test task_lane::queued_room_delete_revalidates_scope_after_full_lane_writer).
  • Invalidation, not payload. EventKind::TaskLaneChanged carries only the runtime, the committed revision, and the origin rooms the batch touched. A private change invalidates no room. The desktop compares the revision with its per-runtime watermark and pulls the changed rooms through list_work_items and list_room_work_lanes (case "task_lane_changed" in apps/desktop/src/stores/daemon.ts). Test lane_tasks.rs::task_lane_changed_is_compact_and_origin_filtered covers the shape.

#The observed write path

NativeWorkProjector serves every provider snapshot: the per-session watchers, AgentEventKind::WorkItems events from owned runtimes (Codex app-server TodoList items in crates/jam-host/src/codex/app_server.rs and Copilot SQL todos in crates/jam-host/src/copilot/sdk.rs, routed by worker::project_runtime_work_items), and the Copilot CLI extension's copilot_bridge_todos. All three enter project_locked under one tokio::sync::Mutex<NativeProjectionState>, which is shared by every lane.

project_locked resolves the lane row, stands down on an authoritative lane, and orders snapshots: a snapshot with an older recorded_at than the latest seen for that lane only re-merges the existing trail, and revision_safe_current stops an older provider row from regressing a newer one. It stamps origin_chat_id with the session's room, merges with merge_trail, and, when anything changed, saves through save_work_items_for_projection. That call writes the rows and an outbox row in one transaction and returns a generation. It then publishes EventKind::WorkItems with the stamped items and runs the same correlate plus complete_native_projection pair as the authoritative path. Tests workitems::older_provider_mtime_cannot_reopen_a_completed_native_assignment and workitems::disappeared_native_row_completes_its_linked_shared_assignment cover ordering and the completed trail.

The watcher loop in workwatch::spawn_work_watch runs one task per host session, as a child of the worker's CancellationToken. It prefers the source's native change signal (WorkItemSource::watch), debounces a burst by 150 ms, and keeps a 45-second safety-net re-read. A source with no watch is polled every 2 seconds, and each tick tries to arm a watch, which covers a task directory that appears after the session starts. A snapshot or save error is logged and retried on the next tick. It never affects the worker. WorkMode::from_env reads JAM_WORKITEMS: off and cli-only disable watchers.

jamd chooses the source in one closure, new_work_source in bins/jam/src/jamd.rs. work_provider maps the session provider (falling back to the peer host when the session provider is empty) and returns a source only for claudecode and codex. Every other provider has no watcher. Owned Codex and Copilot runtimes still feed lanes through runtime events.

#Legacy and room-keyed writes

Two manager writers still use the older whole-snapshot Store::save_work_items, and one client path reaches it indirectly:

  • Manager::record_work, used by set_work_items when the room has no bound runtime (the worker is stopped or never attached). It stays room-keyed and observed, labeled provider cli (test lane_tasks.rs::set_work_items_without_a_runtime_stays_room_keyed_and_observed).
  • questions::save_and_publish, which writes the synthesized "Waiting for your input" card (see Questions). It is refused on an authoritative lane and then degrades to an Inbox-only question.
  • The jam work set, add, status, done, and rm subcommands in bins/jam/src/main.rs, which send a whole list through set_work_items (all but set read the list and edit it client-side first). On an unbound room this is record_work. On a runtime-bound room it becomes a coordinator reconcile_snapshot, which does not use save_work_items, which makes the lane authoritative with adapter Cli (test task_lane::reconcile_makes_the_lane_authoritative_and_the_watcher_stands_down).

#The shared board and its routing

Every consumer of shared tasks goes through one Arc<dyn RoomBoardBackend> in Deps.room_board. jamd composes it once as RoutedRoomBoardBackend::new(LocalRoomBoardBackend, store, board_bridge_slot). The Band arm never holds credentials; it calls BandBoardBridge, which the Manager implements in manager/board_bridge.rs. Manager::into_arc seeds the bridge slot with a Weak reference. Before that, or after manager teardown, every Band call fails Unavailable, not local (test routed_board.rs::an_unseeded_bridge_fails_unavailable_not_local).

How the routed backend decides whether a room's board is local or on Band:

Three rules come out of RoutedRoomBoardBackend::remote_profile and remote_profile_with_support:

  • A room is a Band room when any account's saved directory lists it, either as a member room or in observe_only_room_ids. Several owners resolve to the first by sorted profile name. When at least one Band account exists, a UUID-shaped room that no directory lists resolves to RoomOwner::Unlisted, and every single-room call on it fails Unavailable. list_boards treats it differently: it omits the room's stale local board when every directory has loaded from Band, and fails the whole list while any directory is still unloaded (test routed_board.rs::a_new_band_room_missing_from_the_saved_directory_never_gets_a_local_shadow).
  • Once a room has proven remote support, the store keeps a board_remote_support/<chat>/<profile> marker holding the account instance. A later capability failure then fails the call rather than silently writing a local shadow board (test routed_board.rs::a_capability_lookup_failure_never_switches_a_remote_room_to_local).
  • A room that still holds local tasks moves them to Band through run_shadow_board_migration before it routes to Band. The band-board-migration experiment (Staff visibility, BAND_BOARD_MIGRATION_KEY) controls whether new moves may start; BandBoardBridge::shadow_migration_enabled reads it. Turning it off stops new moves but does not reverse a started one. Local component node IDs cannot be stored on Band, so a room with them stays blocked (ensure_shadow_migrated).

The local backend owns uid generation (next_uid, a monotonic #N from RoomBoard.next_seq, never reused), the rollup rule, and timestamps. Each mutation is load, mutate, rollup, stamp_rollup_milestones, then Store::put_board, which replaces the whole board in one transaction. One tokio::sync::Mutex named ops serializes every mutation for every room on this computer.

The Band arm writes through a durable journal in room_task_operations. offline::enqueue_and_dispatch takes the process-wide REPLAY_GATE, inserts a Pending operation, first dispatches every earlier unfinished operation for that room in sequence order, then dispatches the new one. A blocked earlier operation leaves the new one pending and returns BoardError::Pending naming the operation ID and jam work operation <room> <id>. run_room_task_operation_replay drains the journal at startup, on human-socket reconnect, and every 30 seconds (the loop in Manager::into_arc). A blocked or uncertain operation holds back only its own room (test routed_board/offline.rs::a_store_read_failure_holds_one_room_not_the_whole_replay).

Band board reads fold Band's six statuses into the local four in jam-transport-band (fold_status): only blocked stays Blocked, and in_review and failed become InProgress. Human edits of Band tasks are refused ("people can file Band tasks but cannot edit them from Jam yet" in RoutedRoomBoardBackend::edit_task), and a create with component node IDs is refused rather than dropping them (test routed_board.rs::remote_board_rejects_component_changes_instead_of_discarding_them).

A Band room's board changes arrive as a bodyless RoomEvent::BoardChanged on the human socket. The room feed calls on_board_changed, which refetches and publishes the same EventKind::RoomBoard a local change publishes (roomfeed.rs). Board events always carry the whole board, and clients replace it wholesale.

#Board task lifecycle

A card's overall status is derived, never set directly. roomboard::rollup returns Pending for no stakes, Blocked if any stake is blocked, Completed if all stakes are completed, Pending if all are pending, and InProgress otherwise. Manager::room_task_set_status_for requires a non-empty reason of at most 2,000 characters for Blocked, and the reason goes into the task's history. A stake can change only through its own agent: LocalRoomBoardBackend::set_status returns Conflict when the caller has not taken the task.

How a shared board card's rolled-up status moves, and what ends it:

stamp_rollup_milestones pins started_at the first time overall leaves Pending and clears completed_at whenever a completed card regresses (test roomboard::rollup_walk_stamps_milestones_and_regression_clears_completed). Removal deletes the row locally and cancels the task on Band, because Band never deletes tasks (AgentTaskUpdate.cancel). The Band dispatch journal has its own state machine, documented in State machines: Board operation dispatch.

#Correlation from Layer 1 to Layer 2

roomboard::correlate runs after every lane commit, observed or authoritative, over the lane's live rows. For each row it tries three rules in order:

  1. Typed link. When the row carries linked_chat_id, correlate_typed_link fetches exactly that board and looks for the agent's stake linked to this row. It establishes the stake with take when missing and otherwise calls set_status only on a real change.
  2. Existing link. Otherwise it scans every board from list_boards for this agent's stake whose linked_native_id equals the row ID and propagates on change.
  3. Marker. Otherwise, when the subject had a [#uid] marker, it establishes a link only if exactly one board has this agent assigned to that uid. Zero or many matches skip, and many logs a warning (test roomboard::correlate_ambiguous_marker_is_skipped).

A private task has no blocked-with-reason concept, so correlation never clears a stake the agent set to Blocked unless the private task completes. A correlation failure returns an error, which leaves the outbox row pending for retry_pending (test workitems::correlation_failure_retries_only_the_durable_pending_lane_after_restart).

How room views read both layers, and which events keep each projection current:

room_work_snapshot exists so the desktop can learn which rooms hold work before opening them. It builds lanes through the same list_room_work_lanes projection a room reads, and it filters boards to rooms visible to the profile, because the room_boards table has no profile column (tests room_work_snapshot.rs::a_snapshot_never_returns_a_board_for_another_profiles_room, room_work_snapshot.rs::a_snapshot_lane_equals_what_the_per_room_read_reports). The desktop also derives display-only cards that no task layer stores: pending permission cards (pendingPermissionTasks in apps/desktop/src/lib/permissionTasks.ts) and standalone question cards (withStandaloneQuestions in apps/desktop/src/lib/questionTasks.ts).

#Kickoff: assigning work versus sending a message

Sending a message posts text to a room. Assigning work creates tracked state. Manager::kickoff_work_seeded is the one sequence behind the desktop's kickoff composer (kickoff_work), collaborator kickoffs, and issue-driven starts from the linear-work-source experiment (which passes a KickoffSeed with a fixed room title and issue context). It never rolls back: once Band has created the room, every later failure returns ControlError::RoomKeptAfterKickoffFailure naming the kept room.

The order in which a kickoff creates its durable objects, and where a failure stops it:

kickoff::kickoff_message keeps the person's text verbatim and appends a footer naming the task uid. For an agent this machine runs, the footer includes copy-paste jam work commands with the agent's --session selector and a quoted uid. For a remote agent the footer omits commands, because the recipient may not have the CLI. Collaborators are room members but are not mentioned, because a mention is what makes an agent run; the footer names them to the assignee as available help (kickoff_message_with_collaborators, test kickoff::collaborators_are_named_to_the_primary_as_a_bench). create_task_connected on a Band room refuses to queue behind unfinished room operations and marks the create with kickoff_create_unknown before sending, so a crash can never replay an ambiguous create (tests routed_board/offline.rs::kickoff_crash_gap_pending_marker_never_replays_a_create, routed_board/offline.rs::kickoff_crash_after_post_never_reposts_an_ambiguous_create). End-to-end order is covered by lifecycle.rs::kickoff_creates_room_membership_task_and_message_in_order and failure reporting by lifecycle.rs::kickoff_step_failures_keep_what_exists_and_name_the_step.

Manager::room_task_assign_human in manager/board_bridge.rs is the smaller in-room form: it files the task first, records a local assignee's stake with that agent's own credential and enqueues a local Jam notice in its session queue, or mentions a remote assignee once in the room.

#Questions

A question is a structured request from a runtime to one specific human. All providers use one broker, Manager::ask_room_question. It is reached from four entry points:

  • Owned runtimes, through RuntimeHostContext.questions, which the manager fills with runtime_question_broker per room. Adapters call ask_question from crates/jam-host/src/codex/app_server.rs (native request_user_input and the jam_ask_user dynamic tool), copilot/sdk.rs (ask_user), claudecode/owned/permission.rs (owned Claude AskUserQuestion), and acp.rs (Cursor ACP questions).
  • The Claude Code hook bridge and jam ask, through Control::ask_agent_question and Manager::ask_agent_question_impl. The manager resolves the asking session with the activity join (resolve_activity) and returns None for a foreign, ambiguous, or roomless session, so the hook falls back to the terminal picker (test lifecycle.rs::question_from_a_foreign_session_falls_back_to_none).
  • The scoped runtime collaboration service, RuntimeCollaborationOperation::AskUserQuestion in crates/jam-manager/src/runtime_collaboration.rs.
  • The Copilot CLI extension, Manager::copilot_bridge_question.

How a provider question reaches a human, and how the answer resumes the provider's turn:

Rules the broker enforces:

  • One recipient. select_question_human accepts only a User participant: the named recipient, or the sole human when none was named. In a room with several humans and no recipient the broker returns None without posting (test lifecycle.rs::room_question_without_a_unique_human_never_posts_to_the_room).
  • Exact correlation. The broker mints an 8-character answer code, unique among active questions (fresh_question_answer_code), and a UUID question instance. The room post ends with Reply with `Answer CODE: <answer>`. and an invisible jam-hitl-question marker. try_answer_pending_question accepts a message only when the sender is the recorded recipient with sender type user, the body (after stripping one leading mention of the asking agent and the invisible jam-hitl-answer marker) starts with Answer CODE:, and the answer text is non-empty (test lifecycle.rs::room_question_mentions_one_human_and_rejects_an_agents_answer). The desktop composes the same envelope in formatAnswerMessage (apps/desktop/src/lib/agentQuestion.ts).
  • Permission asks are not answers. The intercept runs try_decide_pending_permission first, and try_answer_pending_question ignores any message starting with the permission-ask marker (test lifecycle.rs::a_permission_ask_post_is_never_consumed_as_a_question_answer).
  • One question per peer and room. pending_questions in Inner is keyed by peer and room. A new ask replaces the old one, resolves the old board card, and drops the old sender, which makes the older waiter return None. A late reply to the replaced ask cannot settle the new one (test lifecycle.rs::a_superseded_question_cannot_be_answered_by_its_late_reply). Removal always uses take_exact_question, which matches group ID and manager-owned revision.
  • Reservation before post. The waiter is registered before the room post, so a fast answer cannot fall into the gap. The board card is attached only if the same group and revision are still current. Otherwise the orphaned card is resolved immediately.
  • Timeout order. The request's timeout_secs, then the per-agent setting runtime.question_timeout.<peer>, then JAM_QUESTION_TIMEOUT_SECS, then RuntimeDefaults.question_timeout_secs. No value or zero waits indefinitely (tests lifecycle.rs::room_question_remains_pending_beyond_ten_minutes, lifecycle.rs::positive_room_question_timeout_settles_unanswered_request, lifecycle.rs::agent_indefinite_question_override_wins_over_overall_timeout).
  • Cancellation. Manager::interrupt_runtime_turn takes the room's question and drops its sender before interrupting the provider, so the provider's pending tool call can return and the interrupt RPC cannot deadlock (comment in interrupt_runtime_turn, test lifecycle.rs::cancelling_a_question_wait_resolves_its_work_card). If the broker future itself is dropped, PendingQuestionCleanup::drop removes the exact entry and spawns board resolution.

The board projection in questions.rs is provider-neutral. question_host_task picks the lane's in-progress task visible in this room, falling back to the newest pending one, and never a sibling room's task (tests questions::running_task_wins_over_a_newer_pending_task, questions::runtime_bound_question_never_attaches_to_a_sibling_room_task). When no task qualifies it synthesizes a question-<group> card with active_form "Waiting for your input" and writes it through the legacy save_work_items. The durable open-question record is separate: room_inbox_questions, keyed by account instance and a task key "<peer>\0<room>:<task_id>". That record, not the card, is authoritative for whether a person still owes an answer (comment in resolve_question_on_board, test lifecycle.rs::answered_question_clears_attention_after_lane_becomes_task_authoritative).

The states one question group can reach, and which of them close its board card and Inbox record:

Answered returns a RuntimeQuestionAnswer. Every other terminal state returns None, which Codex renders as "the question was cancelled; proceed with your best judgment" (bridge_user_input_request).

#Permissions

A permission request asks a human to allow or deny one provider operation before the adapter continues. Adapters translate native approval prompts into RuntimePermissionRequest and call RuntimeHostContext::request_permission, which denies once when no broker is configured. Production call sites are in crates/jam-host/src/codex/app_server.rs, acp.rs, claudecode/owned/runtime.rs, and copilot/sdk.rs. For owned Claude Code, claudecode/owned/bridge.rs routes the permission tool call to the sink in runtime.rs, which makes the context call. Manager::copilot_bridge_permission is the second entry for the Copilot CLI extension. All reach Manager::request_runtime_permission.

How one permission request is registered, decided, and returned to the provider:

The rules:

  • One deadline owner. The manager overwrites expires_at from effective_permission_timeout_secs: JAM_RUNTIME_PERMISSION_TIMEOUT_SECS, then RuntimeDefaults.permission_timeout_secs, then DEFAULT_RUNTIME_PERMISSION_TIMEOUT_SECS (0, meaning wait indefinitely). Adapter values are ignored (test lifecycle.rs::bounded_permission_policy_publishes_the_manager_enforced_deadline).
  • Durable rules are keyed by the exact operation. permission_operation_key builds a key only for command actions (exec_command, command_execution, shell) from the provider's structured argv, never from the display description. It refuses argv with whitespace or control characters, non-system absolute program paths, and interpreters listed in UNNAMEABLE_PROGRAMS. Rules live in the settings table under runtime.permission.always|<peer>|<action>|<program>. scoped_permission_outcome downgrades AllowAlways without a named operation to AllowOnce. A DenyAlways without a name persists as an action-wide refusal under the program string "any operation of this kind". Lookup honors an exact-operation allow or deny but only refusals from action-wide and legacy keys (permission_rule_lookup_keys, tests lifecycle.rs::allow_always_persists_for_named_operations_and_resolves_matching_asks, lifecycle.rs::legacy_action_wide_mcp_allow_rule_is_removed_and_prompts_again, unit manager::mcp_allow_always_is_bounded_to_the_current_request).
  • Room-reply deciders. Room-owner approvals are on unless JAM_ROOM_OWNER_APPROVALS is 0, false, off, or no, or RuntimeDefaults.room_owner_approvals is false. The base decider snapshot, written at registration, is the account's own Band user ID. post_room_permission_ask adds participants with role owner that are the account human or an agent whose handle shares the asking peer's owner prefix, skipping the asking agent. A room reply decides only when parse_permission_reply finds approve or deny followed by the request ID, the room matches, and the sender is in the snapshot. A room reply can produce only AllowOnce or DenyOnce (tests lifecycle.rs::room_reply_from_the_account_owner_decides_a_pending_permission, lifecycle.rs::unauthorized_or_wrong_id_room_replies_never_decide, lifecycle.rs::room_owner_approvals_off_means_room_replies_are_inert, lifecycle.rs::pending_ask_posts_to_qualifying_owners_and_a_same_account_owner_agent_decides).
  • Bounded surface. A request whose Inbox projection exceeds 4 KiB is denied once. A new request is denied once when the profile already holds MAX_PENDING_PERMISSIONS_PER_PROFILE (128) records. A duplicate request ID denies once.
  • Display truncation is stamped. inbox_permission_projection clears metadata_json, caps the description at 2,048 characters, and sets display_truncated when it cut anything. The desktop dialog then states that the answer covers the whole request and shows the structured argv beside the capped text (PermissionApprovalDialog.tsx). It offers "Allow always" only when durable_rule is non-empty.
  • Abandonment. An ExpireAbandoned drop guard marks the record Expired when the adapter drops the broker future, so a late room reply cannot approve an operation no provider is waiting for (test lifecycle.rs::adapter_abandonment_expires_the_ask_and_late_replies_are_ignored). A dropped decision channel resolves Cancelled, and a timeout resolves Expired.

Permission records live only in Inner.runtime_permissions. Resolved records stay in the map as history, which is also what lets a redelivered request ID find its predecessor.

#Attention and the Inbox

"Attention" is the durable account-scoped record of things a person should look at. Three producers feed it:

  • Direct human mentions. The human room feed stages AttentionMutationAction::ApplyDirectMention or ResolveDirectMention in room_attention_mutation_outbox, in the same transaction as the room message mutation (apply_room_message_mutation_with_attention) or alone (stage_attention_mutation). roomfeed::apply_pending_attention applies each staged action to room_attention_items (active, capped at ATTENTION_ACTIVE_LIMIT, 128) or room_attention_spool (overflow, capped at 10,000), acknowledges it, and publishes EventKind::InboxAttentionChanged plus an AttentionCandidate for desktop notifications. SqliteStore::apply_direct_mention reads the item's room_attention_tombstones sequence before inserting and refuses any mutation whose sequence is at or below it, so a stale apply cannot resurrect a resolved item. It applies the same sequence <= stored refusal against the active and spool rows.
  • Open questions. put_inbox_questions and resolve_inbox_questions maintain room_inbox_questions. Accepted replies are appended to room_inbox_accepted_question_replies with a per-account projection revision and published as AgentQuestionReplyAccepted, which lets the desktop bind the reply to the exact question post.
  • Permission asks. bind_permission_attention links the room post that announced an ask to its (peer, permission_id, agent_session_id, turn_id) key in room_permission_attention_links. apply_direct_mention skips a message already bound to a resolved permission, so the ask post does not linger as a mention.

Manager::cached_inbox_snapshot builds InboxSnapshot from the cached store state (direct mentions, read and done flags from room_inbox_state, open questions by task key, overflow and truncation counts) plus pending permissions read from memory. resolve_inbox_attention and set_inbox_item_state are the human mutations. Every account-scoped mutation checks the account generation before and after writing. The desktop Inbox section is gated by the inbox-section experiment. With it off, Home and each room's Work board still show questions and permission requests (experiment description in crates/jam-manager/src/experiments.rs).

#Plans

A RoomPlan is one record per room with two independent parts: a document (a local Markdown file, a standalone Arch JSON file, or a directly pushed diagram) with its cached body, and a sparse NodeOverlay that agents paint onto the diagram's node IDs. Document writes bump doc_rev. Overlay writes (plan_focus, plan_status, plan_unfocus) do not. Both ride the same EventKind::RoomPlan, whose document_write flag tells clients which kind of write it was.

Every plan mutation takes the room's PlanWriteGates entry (a std::sync::Mutex per room, reclaimed through Weak references) around read, mutate, and save_and_publish_plan. save_and_publish_plan rechecks that the event profile is still active, saves the whole record (room_plans.data as JSON), publishes, and asks the watcher to resync the room (test plan_gate::same_room_shares_a_gate_and_different_rooms_do_not).

Reads go through one validate-then-read chokepoint in crates/jam-manager/src/plan.rs. read_local_plan requires an absolute path, canonicalizes it, and opens it component by component without following symlinks (open_regular_file using cap_primitives). It accepts only .md, .markdown, or .json, at most MAX_PLAN_BYTES (2 MiB), in valid UTF-8, and validates Arch JSON with validate_arch_document. read_observed_plan rereads the exact canonical path accepted at attach time without canonicalizing again, so a symlink planted later degrades the source instead of being followed (test plan::observed_read_rejects_a_later_symlink_replacement). A failed refresh keeps the last good cache and records last_error (test lifecycle.rs::plan_service_persists_emits_and_keeps_a_stale_cache).

Two source kinds behave differently:

  • A human attach keeps the source live. PlanLiveSources watches the parent directory with notify, debounces by 200 ms, and reconciles every 60 seconds with a 16-target and 16 MiB budget (plan_watch.rs constants). The record starts Degraded and moves to Live when the watcher confirms it (test lifecycle.rs::human_attach_mints_no_artifact_and_keeps_live_watching).
  • An agent attach snapshots the file into the artifact store (artifact_ingest_bytes with origin plan) and points the source at the stored copy with snapshot: true. A snapshot is never watched (tests lifecycle.rs::agent_plan_attach_mints_a_plan_artifact_and_stays_snapshot, plan_watch::snapshot_source_is_never_watched_and_keeps_its_persisted_state).

Attaching a different file drops the overlay, because its node IDs belong to the old diagram. clear_plan writes an empty record that keeps the bumped doc_rev, so revisions stay monotonic across clear and reattach. apply_focus and apply_status keep at most one Active node, and it is always overlay.focus.

#Artifacts

An artifact is a file attached to a room. Bytes live under <store-root>/artifacts/<sha256>/<name>, where the store root is beside jam.db (SqliteStore::artifacts_root). Identical bytes share one file. Each attach gets its own artifacts row with an art-<24 hex> ID. Files are plaintext on purpose so local agents can open them by path. sanitize_artifact_name reduces any client name to one safe file-name component and strips CommonMark link characters, so a name cannot inject a link into the attachment line.

A process-wide FILE_MUTATION_LOCK serializes store-bytes-then-insert against count-then-delete, so deleting one room's row can never remove bytes another row still uses (delete_row_and_maybe_file_after, test artifacts::store_dedups_identical_bytes_under_one_file). Deletion is two-step: artifact_delete_preview computes an impact token from the final-reference check and the snapshot plans that would lose their backing file, and artifact_delete rejects a stale token. Affected snapshot plans are stamped with a "removed from Files" banner in the same store commit (commit_artifact_delete, tests lifecycle.rs::deleting_the_backing_artifact_stamps_the_banner_only_when_the_file_goes, lifecycle.rs::sqlite_artifact_delete_rolls_back_first_plan_when_second_plan_write_fails).

Sharing follows the local-only link rule:

  • When the Band deployment reports feature flag ff_file_transfer, a room-bound artifact is uploaded in the background (spawn_artifact_upload) and its remote_id recorded. Sends refresh each upload at send time, verify the local bytes still match the stored SHA-256, and pass platform file IDs, not links (Manager::room_send_with_staged, Manager::send_with_artifacts). Manager::into_arc resumes unfinished uploads at startup (test lifecycle.rs::manager_startup_resumes_bound_artifact_uploads).
  • Without file transfer, attachments travel as 📎 [name](file://…) links, and only when classify_locality proves every participant local: the verified sending human or an agent whose ID this daemon owns. An empty participant list is not proof (test artifacts::locality_is_the_store_join_never_kind).
  • In either mode, validate_managed_local_links rejects any file:// Markdown destination in the body that was not minted from a validated artifact row.

A message carries at most MAX_MESSAGE_ATTACHMENTS (10) files, and one artifact is at most DEFAULT_MAX_ARTIFACT_BYTES (32 MiB) unless JAM_ARTIFACT_MAX_BYTES overrides it. Unbound artifacts staged for a kickoff that never happened are swept after one day, and the sweep re-checks bound-ness under the file lock (test artifacts::stale_unbound_delete_skips_a_row_bound_since_listing). sweep_orphan_dirs removes SHA directories no row references. Both run on the activity-report path every 512th report in a blocking task (Manager::report_activity_impl).

#State and ownership

Aggregate Key and scope Authority Mutation coordinator Durable representation Projections Freshness or version fence Recovery Deletion authority
Layer 1 lane (profile, scope, storage room) resolved from RuntimeSessionId; room-keyed when unbound EffectiveTaskSource: provider store when Observed, Jam when JamAuthoritative TaskLaneCoordinator gate per runtime; NativeWorkProjector single mutex for observed writes work_snapshots, work_items, task_lane_source, task_events list_work_items_with_source, list_room_work_lanes, room_work_snapshot; WorkItems and TaskLaneChanged events work_snapshots.projection_generation; expected_generation on authoritative commits; recorded_at ordering for observed snapshots Outbox replay by NativeWorkProjector::retry_pending in Manager::rebuild Engine Delete tombstones only; archive_runtime_lanes_for_peer on peer archive; reclaim_orphan_lanes after 90 days, twice; delete_all_work_items on peer removal
Lane idempotency record (profile, scope, room, key) Coordinator Same aggregate transaction task_idempotency Replay outcome Request hash Survives restart Removed with the lane's side rows; prune_task_idempotency exists but has no production caller
Lane projection outbox (profile, scope, room) Lane commit Written in the lane transaction; acknowledged by generation native_projection_outbox None complete_native_projection ignores a stale generation retry_pending at rebuild Acknowledgement; cascades with the snapshot row
Local shared board chat_id Jam LocalRoomBoardBackend.ops, one mutex for all rooms room_boards, room_tasks, room_task_assignments get_room_board, list_boards, room_work_snapshot; RoomBoard event Whole-board replace; next_seq high-water None needed remove; board rows stay after a room moves to Band
Band shared board Band room Band Band; Jam writes through BandBoardBridge Band only, with the board_remote_support marker in settings Same Control reads; RoomBoard event after refetch Band task uid; status folding on read Refetch on BoardChanged and after replay Cancel on Band; never deleted
Board operation journal Operation UUID, sequence per (account_instance, room) Jam until settled REPLAY_GATE; validate_room_task_operation_transition forward-only room_task_operations list_room_task_operations, get_room_task_operation Sequence order per room run_room_task_operation_replay at start, reconnect, and every 30 s; manual reconcile_room_task_operation Settles to Completed or Discarded; row retained until the account is removed (SqliteStore::remove_account)
Shadow migration state (chat_id, account_instance) and per-row markers Jam run_shadow_board_migration Settings board_migration/done/…, board_migration/row/…, board_migration/store_id Routing decision Per-row markers creating, retryable, uid:<id> Rerun on directory change and startup Not reversed
Pending question group (peer, room), one at a time Manager Inner mutex; take_exact_question by group and revision Memory only (pending_questions) Room post, board card, AgentQuestionsAsked Manager revision next_question_revision None: lost on restart Answer, timeout, cancellation, supersession
Open question record (account_instance, task_key) Manager put_inbox_questions, resolve_inbox_questions room_inbox_questions InboxSnapshot.open_questions_by_task None Read from store after restart resolve_inbox_questions only
Accepted question reply (account_instance, peer, room, group, revision) Manager append_accepted_question_reply room_inbox_accepted_question_replies AgentQuestionReplyAccepted Per-account projection_revision Survives restart The append keeps only the newest ACCEPTED_QUESTION_REPLY_LIMIT (128) per account
Synthesized question card Lane row question-<group> Manager Legacy save_work_items work_items Work board Refused on authoritative lanes None Completed on resolution
Runtime permission record (peer, request id) Manager broker Inner mutex; finish_runtime_permission is idempotent Memory only (runtime_permissions) runtime_permissions, Inbox snapshot, RuntimePermission* events status Pending to Resolved None: lost on restart None: resolved records remain as history until the daemon exits
Durable permission rule (peer, action, program) Human decider decide_runtime_permission Settings `runtime.permission.always …` runtime_permission_rules Exact key match Survives restart
Permission attention link (account_instance, peer, permission, session, turn) Manager bind_permission_attention, resolve_permission_attention room_permission_attention_links Inbox visibility of the ask post resolved_at Survives restart Pruned after 30 days or beyond 10,000 resolved links
Direct-mention attention (account_instance, item_id) Band human feed Room persistence worker drains the outbox room_attention_items, room_attention_spool, room_attention_tombstones, room_attention_mutation_outbox InboxSnapshot, InboxAttentionChanged, AttentionCandidate mutation_sequence per room Outbox replay by the persistence worker Resolution tombstone; retired with the room
Room plan chat_id Human or agent that attached it; source file for live plans PlanWriteGates per room room_plans.data JSON room_plan; RoomPlan event doc_rev for documents; attachment_id per attach Watcher rebuild at start and every 60 s clear_plan keeps a tombstone record
Artifact Row ID; bytes by SHA-256 Jam FILE_MUTATION_LOCK artifacts rows; <root>/artifacts/<sha>/ files artifact_list; RoomArtifacts event impact_token for delete Upload resume at start artifact_delete with token; one-day unbound sweep; orphan directory sweep; account removal

#Contracts

#Control methods

This subsystem serves 58 Control methods. The generated crosswalk in control-api.md lists them with their routes and callers under the headings "Layer-1 authoritative lane Task verbs", "Layer 2: shared room board", "Kickoff", "Room plan", and "Room artifacts". Question, permission, work-item, and runtime_task_call methods appear under "Peers and lifecycle", set_inbox_item_state and resolve_inbox_attention under "contacts", and the two Copilot bridge methods under "Copilot CLI extension bridge". The methods are the five lane_task_* verbs, set_work_items, list_work_items_with_source, list_room_work_lanes, room_work_snapshot, work_item_history, the seven room_task_* verbs, get_room_board, the three room-task-operation methods, ask_agent_question, agent_question_timeout, set_agent_question_timeout, the five runtime-permission methods (runtime_permissions, decide_runtime_permission, runtime_permission_rules, forget_runtime_permission_rule, clear_runtime_permission_rules), kickoff_work, the nine plan verbs, the fourteen methods under "Room artifacts" (including room_locality and room_file_list), set_inbox_item_state, resolve_inbox_attention, runtime_task_call, copilot_bridge_permission, and copilot_bridge_question. list_work_items has no daemon route and delegates to list_work_items_with_source.

Semantics the code enforces on those methods:

  • lane_task_* return LaneTaskWriteResp with replayed, result_generation, and generation. Engine rejections and oversized keys are Invalid, key reuse with a different payload is Conflict, a task outside the named room is NotFound, and store failures are Internal (impl From<TaskLaneError> for ControlError).
  • lane_task_list pages with a stable cursor and returns the generation read atomically with the rows (Store::get_lane_snapshot, tests lane_tasks.rs::list_paginates_with_a_stable_cursor, lane_tasks.rs::list_generation_matches_the_rows_it_returned).
  • room_task_* on a Band room may return BoardError::Pending with an operation ID. That means the change is journaled and will be retried. The caller must not resubmit it.
  • ask_agent_question, copilot_bridge_question, and copilot_bridge_permission are held-open routes with no unary timeout (the S marker in the crosswalk). kickoff_work is also marked held-open.
  • runtime_task_call captures the actor once at admission, rechecks that exact capability before dispatch, and returns OutcomeUnknown instead of retrying when an account change drops the operation (Manager::runtime_task_call).

#Events

This subsystem publishes 13 event kinds: work_items, task_lane_changed, room_board, room_plan, room_artifacts, inbox_snapshot, inbox_attention_changed, attention_candidate, agent_questions_asked, agent_questions_resolved, agent_question_reply_accepted, runtime_permission_requested, and runtime_permission_resolved. See events.md for which the desktop handles. work_items, room_board, room_plan, and room_artifacts carry complete replacements for their key. task_lane_changed and inbox_attention_changed are invalidations that require a pull. The event stream is best-effort, as described in Messaging and room state, so every surface must hydrate these slices on mount and after reconnect.

#Traits provided and consumed

Trait Defined in Implementations Contract
RoomBoardBackend jam-contract LocalRoomBoardBackend, RoutedRoomBoardBackend, test fakes Backend assigns uids and stamps creators. set_status returns whether this write completed the card. The trait documents remove as idempotent, but LocalRoomBoardBackend::remove returns BoardError::NotFound for a missing uid. routes_to_band defaults to false.
BandBoardBridge jam-manager::bandboard Manager in manager/board_bridge.rs Credentialed Band operations with an explicit acting identity. Methods that cannot know their outcome return OfflineBeforeSend only when nothing was sent.
WorkItemSource jam-core::worksource ClaudeCodeSource, CodexSource, NullSource One instance per host session. snapshot returns None for no task surface and never panics. A closed watch channel means fall back to polling.
RuntimeQuestionBroker, RuntimePermissionBroker (type aliases for boxed async closures, not traits) jam-host Manager closures A question returns None for cancel or timeout. A missing permission broker denies once.
ToolProvider bins/jam/src/mcp/mod.rs TaskTools and ScopedRuntimeTools (the scoped runtime task service), among others Tool errors are tool results, not JSON-RPC errors. Schemas expose no runtime, identity, or room argument.

#Band protocol consumed

The Band arm uses the human board read, human task create, agent board read (which carries linked_native_id and assignee IDs), agent task create, agent task update (AgentTaskUpdate facets: subject, detail, status, active form, linked native ID, comment, cancel), agent task history, and the lifecycle read that confirms cancellation. Band has no idempotency key for task creation, so an uncertain create is never retried blindly (run_room_task_operation_replay doc comment). Questions and permission asks are ordinary room messages with Jam envelopes, not a Band feature.

#Invariants

Rule Enforced by Known exceptions
An authoritative lane batch applies completely or not at all. TaskLane::apply on a scratch copy; commit_task_aggregate in one transaction. Test lane_tasks.rs::invalid_batches_reject_whole_with_structured_errors. None.
Two writers to one runtime lane cannot lose each other's update. Per-runtime gate plus expected_generation. Tests task_lane::concurrent_creates_on_one_runtime_do_not_lose_updates, lane_tasks.rs::jam_work_and_lane_task_on_one_runtime_do_not_lose_updates. The legacy save_work_items writers (record_work, question cards) do not check a generation.
An observed snapshot never overwrites an authoritative lane. Pre-check in project_locked and refuse_if_lane_authoritative in the store transaction. Tests workitems::observed_watcher_stands_down_once_the_lane_is_authoritative, workitems::authoritative_flip_between_precheck_and_save_cannot_clobber. None.
A named-room caller can see and mutate only that room's tasks. validate_room_scope under the gate; retain_room_projection on reads. Test lane_tasks.rs::named_room_context_cannot_read_or_mutate_sibling_or_private_tasks. A caller with no room scope gets the full lane.
Task IDs and board uids are never reused. Tombstones are immutable and keep their IDs (TaskEngineError::Tombstoned, IdCollision); next_uid increments next_seq. Test roomboard::next_uid_is_monotonic_and_never_reuses. None.
A lane commit's board correlation eventually runs. Outbox row in the lane transaction; retry_pending at rebuild. Tests workitems::authoritative_commit_crash_before_ack_drains_exactly_once_on_restart, workitems::correlation_failure_retries_only_the_durable_pending_lane_after_restart. Correlation is at-least-once. The anti-thrash check in correlate makes repeats harmless.
Correlation never guesses a card. Typed link or exact linked_native_id; marker only when one board matches. Test roomboard::correlate_ambiguous_marker_is_skipped. None.
A Band room never gets a local shadow board by accident. RoomOwner::Unlisted, the remote-support marker, and ensure_shadow_migrated. Tests routed_board.rs::a_new_band_room_missing_from_the_saved_directory_never_gets_a_local_shadow, routed_board.rs::a_capability_lookup_failure_never_switches_a_remote_room_to_local, routed_board/offline.rs::offline_without_proven_remote_capability_never_enqueues_or_shadows. A room with local component IDs stays blocked rather than moving.
Band board writes in one room apply in order, and an uncertain write blocks later writes in that room. enqueue_and_dispatch dispatches earlier operations first; replay tracks blocked_rooms. Tests routed_board/offline.rs::proven_remote_room_queues_offline_writes_and_replays_them_in_room_order, routed_board/offline.rs::an_uncertain_change_still_holds_back_a_new_change_in_its_room. A change to a task no longer on the board is settled Discarded and does not block.
A board operation's state only moves forward. validate_room_task_operation_transition in jam-store. None.
An operator cannot discard an operation while its request is in flight. lock_room_task_operations shares REPLAY_GATE. Test routed_board/recovery.rs::operator_cannot_discard_while_assignment_post_is_in_flight. None.
A blocked shared task always names its blocker. room_task_set_status_for rejects Blocked without a reason. Test routed_board.rs::blocking_sends_the_status_and_its_reason_in_one_update covers the Band arm. roomboard::correlate calls the backend's set_status directly with no reason, so a private task in Blocked can block its linked stake without one.
Only the named human answers a question, with the exact code. try_answer_pending_question checks recipient ID, sender type user, and the Answer CODE: envelope. Test lifecycle.rs::room_question_mentions_one_human_and_rejects_an_agents_answer. None.
A question never wakes every agent in a multi-human room. select_question_human returns None without a unique human. Test lifecycle.rs::room_question_without_a_unique_human_never_posts_to_the_room. None.
A consumed answer is not delivered as a new turn. QuestionGate::deliver returns early and the intercept acks asynchronously. If the ack fails, the message redelivers on respawn (comment in question_intercept).
The asking agent never decides its own permission by room reply. post_room_permission_ask skips the asker; the base snapshot is the account human. The decide_runtime_permission Control path has no caller-identity check. See Refactor notes.
A durable allow binds to one exact named operation. permission_operation_key, scoped_permission_outcome, permission_rule_lookup_keys. Tests lifecycle.rs::allow_always_persists_for_named_operations_and_resolves_matching_asks, manager::mcp_allow_always_is_bounded_to_the_current_request. Action-wide refusals persist by design.
A room reply can only mint a one-time decision. try_decide_pending_permission maps to AllowOnce or DenyOnce. None.
An abandoned permission request cannot be approved later. ExpireAbandoned drop guard. Test lifecycle.rs::adapter_abandonment_expires_the_ask_and_late_replies_are_ignored. None.
A truncated permission description is labeled as truncated. inbox_permission_projection sets display_truncated; the dialog renders the warning. Test manager::inbox_permission_projection_drops_metadata_without_mutating_broker_state. None.
A plan source is never followed through a symlink after attach. read_observed_plan without re-canonicalization; open_regular_file refuses links. Tests plan::observed_read_rejects_a_later_symlink_replacement, lifecycle.rs::snapshot_refresh_never_follows_a_planted_leaf_symlink. None.
doc_rev is monotonic and never wraps. next_doc_rev uses checked_add. Tests plan::checked_document_revision_never_wraps, lifecycle.rs::exhausted_plan_revision_refuses_a_document_write_without_mutating_storage. Records from older daemons carry doc_rev 0.
At most one overlay node is active, and it is the focus. apply_focus, apply_status, demote_other_active. None.
A local file link is sent only when every participant is proven local. classify_locality; validate_managed_local_links. Tests artifacts::locality_is_the_store_join_never_kind, lifecycle.rs::agent_send_uses_local_links_when_band_has_no_file_transfer. With ff_file_transfer, platform file IDs replace links for any audience.
Deleting one artifact row never removes bytes another row uses. FILE_MUTATION_LOCK around count and delete. Test lifecycle.rs::artifact_vertical_stores_dedups_deletes_and_fans_out. None.
A kickoff never erases the room it created. RoomKeptAfterKickoffFailure on every post-room failure. Test lifecycle.rs::kickoff_step_failures_keep_what_exists_and_name_the_step. None.

#Failure and recovery

Daemon crash during a lane write. Before commit_task_aggregate returns, nothing is written. After it, the rows, source, idempotency record, and outbox row are durable. The desktop missed the TaskLaneChanged event, but its revision watermark detects the gap on the next event or reconnect and triggers a full hydrate. Manager::rebuild calls NativeWorkProjector::retry_pending, which runs correlation for every pending outbox row and acknowledges it by generation. A client that retries the same idempotency key receives the recorded result.

Provider watcher failure. A snapshot I/O or decode error is logged and retried on the next tick with the last trail kept. A dead fs-notify watcher switches the loop to 2-second polling and re-arms later. None of this touches the worker.

Board write while Band is unreachable. A proven-remote room journals the operation and returns BoardError::Pending. Replay runs at startup, on reconnect, and every 30 seconds. An operation that may have reached Band moves to NeedsReconciliation. Replay inspects Band's board to recognize an applied change, and an operator settles the rest with jam work reconcile. A kickoff create that was not proven sent is never replayed (kickoff_create_unknown). A room with no proven remote capability fails the call and journals nothing.

Board correlation failure. The outbox row stays pending. The next lane commit for that lane or the next rebuild retries it. A missing card is logged at debug level and skipped.

Question wait interrupted. On timeout, cancellation, supersession, or future drop, the broker resolves the board card and the Inbox record and the adapter receives None. On daemon crash, the in-memory waiter, the provider turn, and the provider process are gone. Inferred: the room_inbox_questions row and any synthesized question-<group> card are not resolved on restart, because resolve_inbox_questions is called only from resolve_question_on_board, and no startup path calls it. A stale open question would stay in the Inbox until something else clears it. No test covers a question across a daemon restart.

Question post failure. The broker removes its own reservation and returns None. Nothing was projected.

Permission wait interrupted. Timeout resolves Expired, a dropped channel resolves Cancelled, and an abandoned broker future resolves Expired through the drop guard. A room post failure leaves the ask decidable in the app and CLI. On daemon crash every permission record is lost with its provider turn. Durable rules and attention links survive. Inferred: an attention link bound to a post whose ask never resolved stays unresolved, because resolve_permission_attention runs only from finish_runtime_permission and post_room_permission_ask.

Plan source unavailable. A failed refresh keeps the cached body and records a bounded, path-free last_error (source_error_for_display). A watcher failure sets live_source_state to Degraded with watch_error. The watcher rebuilds its targets from the store every 60 seconds and after its command queue overflows. PlanLiveSources::shutdown waits for a commit already inside the commit gate, then refuses later commits (test plan_watch::shutdown_waits_for_started_refresh_and_prevents_a_late_commit). Account removal marks affected live plans degraded before the profile's fan-out is removed (remove_profile).

Artifact upload failure. The attach still succeeds. remote_id stays empty, so the composer's gate treats the file as local-only. A 404 on a deployment without ff_file_transfer marks the artifact local-only. Unfinished uploads resume at startup. A kickoff upload failure keeps the artifact staged for a new attempt (test lifecycle.rs::remote_kickoff_upload_failure_keeps_the_artifact_staged_for_retry).

Kickoff partial failure. Each step after room creation returns RoomKeptAfterKickoffFailure with the room ID and the failed step. The desktop offers the room's own participant and task controls to continue. Nothing is rolled back.

#Extension points

#Adding a coding-agent harness

To get private task capture for a new provider, a developer has three options today:

  1. A file-backed task store. Implement WorkItemSource in the provider's jam-host module (as ClaudeCodeSource and CodexSource do), add a status vocabulary constant beside CODEX_VOCAB in crates/jam-core/src/worksource.rs when the provider's words differ, add one arm to the new_work_source closure in bins/jam/src/jamd.rs, and add the provider name to the allow-list in work_provider in the same file. worker::spawn_worker and Manager::build_work_watch pick it up with no change.
  2. Runtime events. Have the adapter emit AgentEventKind::WorkItems { provider, items }, as the Codex app-server and Copilot SDK adapters do. worker::project_runtime_work_items projects it with no manager change.
  3. Authoritative task tools. Attach Jam's task service instead of scraping. Host-native Codex registers jam mcp tasks through jam_host::codex::task_mcp. Sandboxed runtimes use the scoped runtime task service (bins/jam/src/mcp/runtime_tasks.rs), and Shared Codex uses RuntimeTaskToolProvider in the manager's tool-registry builder (the SharedAgentHost branch near RuntimeTaskToolProvider::new in manager.rs). The provider needs no lane code: all three reach TaskLaneCoordinator.

To support questions, translate the native question into RuntimeQuestionRequest and call RuntimeHostContext::ask_question. One call site per adapter exists today, in codex/app_server.rs, copilot/sdk.rs, claudecode/owned/permission.rs, and acp.rs. The adapter must translate None into the provider's "proceed without an answer" result and must not add its own timeout. For an attached provider with hooks, route through Control::ask_agent_question as the Claude hook bridge in bins/jam/src/main.rs does.

To support permissions, translate the native approval into RuntimePermissionRequest and call RuntimeHostContext::request_permission. Fill command with the provider's structured argv when it has one, because only that enables a durable allow. Set display_truncated when the adapter abbreviated the description (see acp_permission_description in acp.rs). Use an action name from PERMISSION_RULE_ACTIONS in manager.rs, or the value is normalized to permission on the Copilot bridge path (normalize_permission_action). Do not add an adapter-side timer: the manager owns the deadline. Existing adapters call it from codex/app_server.rs, acp.rs, claudecode/owned/runtime.rs, and copilot/sdk.rs.

#Adding a feature in this subsystem

  • A new lane field. Add it to WorkItem, validate it in TaskLane::apply, add columns to work_items in a migration, persist it in both SqliteStore and FileStore commit_task_aggregate and save_work_items_internal, include it in hash_mutation so idempotency covers it, and run just bindings. Decide whether merge_trail preserves it for observed lanes.
  • A new shared-board operation. Add the method to RoomBoardBackend, implement it in LocalRoomBoardBackend and RoutedRoomBoardBackend, add a RoomTaskOperationPayload variant plus its dispatch and reconciliation in bandboard/offline.rs, add the credentialed call to BandBoardBridge and its Manager implementation, then follow the nine-step Tauri command procedure in AGENTS.md. Update the test fakes that implement RoomBoardBackend: ObservedBoard in the workitems.rs tests, RecordingLocal in crates/jam-manager/tests/routed_board.rs, and BandRoutedBoard and ScopedOperationBoard in crates/jam-manager/tests/lifecycle.rs.
  • A new attention source. Add a store method pair in RoomMessageCacheRepo with a no-op default, implement it in SqliteStore and FileStore, publish InboxAttentionChanged, and extend InboxSnapshot and cached_inbox_snapshot.
  • A new plan source kind. Extend PlanSourceKind (its hand-written Deserialize maps unknown values to Unknown), the extension match in read_plan_file, and the watcher's Target.kind handling.

#Refactor notes

manager.rs holds most of the behavior. The question broker, the permission broker, attention binding, plan verbs, artifact verbs, kickoff, and the lane_task_* validation all live in the 58,145-line crates/jam-manager/src/manager.rs, interleaved with unrelated lifecycle code. Only the pure helpers are split out (questions.rs, plan.rs, artifacts.rs, kickoff.rs). Both brokers keep their state in the shared Inner struct behind the manager's one std::sync::Mutex. crates/jam-manager/tests/lifecycle.rs (68,007 lines) is the main test harness for all of them.

Three write paths into one lane. Authoritative commits (commit_task_aggregate), observed projections (save_work_items_for_projection), and legacy snapshots (save_work_items) all write work_items. Only the first uses the per-runtime gate and generation fence. The legacy path is used by record_work for unbound rooms and by the synthesized question card. The observed path serializes every lane in the process behind one tokio::sync::Mutex in NativeWorkProjector, and the local board serializes every room behind one ops mutex.

Transitional lane keying. RUNTIME_KEYED_LANES is a pub const set to true, re-exported from jam-manager, and tested in three conditionals: lane_storage_room and stamp_after_save in workitems.rs, and the startup migration in bins/jam/src/jamd.rs. Dual-read through lane_storage_room remains on every lane read and write. Lease-backed attached sessions use a synthetic peer-<hash> runtime ID that is not a real RuntimeSessionId.

Declared but unused. jam_domain::LaneState (Live, Detached, Archived, Reclaimed) is exported but not referenced by any other code. Lane state is implied by table membership (work_snapshots versus work_lane_archive). TaskWriteAdapter::AppServerDynamic is never set outside tests. The Shared Codex dynamic-tool path records McpStdio because Manager::run_runtime_task hard-codes that adapter. task_engine::tombstones_to_compact and Store::prune_task_idempotency have no production callers, so MAX_TOMBSTONE_ROWS (1,024) is enforced as a rejection bound but tombstones are never compacted, and idempotency rows are removed only with their lane.

Permission map growth. Inner.runtime_permissions keeps resolved records as history, and no code path in manager.rs removes an entry. The per-profile cap in request_runtime_permission sums by_id.len(), which counts resolved records as well as pending ones. Inferred: after 128 permission requests for one profile in a single daemon lifetime, every further request is denied with "too many permission requests are waiting for this account". No test covers this cap.

Decision authority is enforced only on the room-reply path. try_decide_pending_permission checks the decider snapshot, but Manager::decide_runtime_permission and the daemon handler accept any caller on the owner-UID socket, and jam permissions allow <id> --always passes AllowAlways straight through. Inferred: a local agent process that shares the operator's OS user and can run jam could approve its own pending request through the CLI. This differs from the product rule in AGENTS.md that only an authenticated approval dialog persists an allow.

Board correlation and question projection reach across layers. roomboard::correlate depends on RoomBoardBackend::list_boards, which on the routed backend reads every local board and, for Band rooms, may probe remote support per profile. It runs after every lane commit. questions.rs writes directly to the store and fan-out rather than through the lane coordinator.

Board data is not profile-scoped in the store. room_boards and room_tasks have no profile column, so room_work_snapshot must intersect with the account's room directory to avoid leaking another profile's rooms (comment in room_work_snapshot).

Stringly correlation envelopes. Question answers, permission replies, and kickoff footers are recognized by text prefixes and invisible Unicode markers (HITL_ANSWER_MARKER, PERMISSION_ASK_MARKER, Answer CODE:, approve <id>) that the Rust manager and the TypeScript desktop (apps/desktop/src/lib/agentQuestion.ts) each define. No shared constant or generated binding keeps them in agreement.

Sweeps on the activity path. The one-day unbound-artifact sweep and the orphan-directory sweep run inside report_activity_impl every 512th activity report, so they depend on provider activity volume rather than a timer.

Memory-only brokers. Pending questions and permissions have no durable record, while their projections (room_inbox_questions, room_permission_attention_links, synthesized cards) are durable. A refactor that moves either broker must decide how those projections recover after a crash. See Consistency and recovery for the startup order these recoveries would join.

Related chapters: Messaging and room state for the delivery path the QuestionGate sits on, Execution and lifecycle for workers and host sessions, Sandbox, workspaces, and continuity for the scoped runtime task service's sandbox side, and State machines for the catalog entries this chapter details.

Scroll to zoom, drag to pan.