Skip to content

Dispatch internals

How the daemon decides what to run, and what the worker's prompt looks like. The operator-facing lifecycle is in Lifecycle.

How dispatch works

  • Poller (daemon/lib/shuttle/poller.ex) owns the tick. It walks each configured felt store with one projected listing (shuttle ls --json --has-field shuttle --json-field …), and considers a fiber eligible iff it carries a shuttle: block owned by this host (shuttle.host matches), felt status is active, and it isn't already running/claimed (see eligible?/2 in poller.ex). A listing that fails or times out degrades that store to its last-known rows for the tick; there is no broader fallback listing. A shuttle CLI too old for the projection flags shows up at boot as a contract skew.
  • Host identity is shuttle's. The Poller takes SHUTTLE_HOST (trimmed) or asks shuttle host --json once at boot and freezes the answer; a shuttle CLI that cannot answer stops the daemon from booting. The host-file and hostname chain lives only in the Go CLI (internal/shuttlecli/host.go).
  • Eligibility is pure; the filesystem is the dispatch action's business. eligible?/2 reads fiber frontmatter and in-memory runtime maps and nothing else. Whether a project_dir exists is decided inside do_dispatch_fiber/3 — for a fiber that has passed every pure gate and is about to have a worker spawned into that directory. The reason is TCC: a project_dir in a macOS file provider (iCloud Drive, ~/Library/CloudStorage) answers every touch with an "access data from other apps" prompt that cannot be granted to a launchd-run daemon, so the cost of a touch is a dialog on someone's screen, not a syscall. A fiber the poller merely looks at each tick — parked, closed, or refused — is never touched, however long it sits there.
  • A missing project_dir refuses the dispatch. Present-but-missing means the checkout lives on another machine, so the fiber is refused with {:project_dir_missing, dir} rather than having its worker silently downgraded to a felt store as its cwd. A force-dispatch skips the check like every other non-force gate. The refusal is recorded as a dispatch failure, so the fiber appears in the snapshot's blocked list with its reason instead of vanishing from the board, and a successful dispatch clears the row. It also rides the preflight cooldown (@preflight_cooldown_ms, 5 minutes): a directory that is absent — or that this daemon is denied, which is the same File.dir? answer — will not appear between two ticks, so it is retried once per window rather than once per tick.
  • Workers are NOT excluded from sharing a checkout. Several workers may run in one project_dir at once, and shuttle says nothing about it. Coordinating concurrent work in a shared checkout is the operator's call, not the dispatcher's.
  • Configured stores come from SHUTTLE_STORES (comma-separated env var) → persisted ~/.config/shuttle/stores.json. There is no implicit default store. POST /api/v1/felt-stores rewrites the persisted file. Inside the daemon, Shuttle.FeltStores maps a fiber to the store that holds it (store_for_fiber, resolve_fiber → %{store, fiber_id, path, uid}).
  • Picker projects are a separate list — SHUTTLE_PROJECTS → persisted ~/.config/shuttle/projects.json (Shuttle.Projects) — and answer a different question: which checkouts a human can file INTO from the Stash/Capture forms. Kept out of the poll list on purpose, so polling never walks TCC-protected paths. Served at origins.<host>.projects of GET /api/v1/felt-stores. The file is hand-editable, and the forms add to it too. Both forms show the destination as host then project (ui/src/forms/ProjectPicker.tsx — HostPicker and ProjectPicker, both plain <select>s, sized to match the agent and effort selects beside them). The host list is the origins of GET /api/v1/felt-stores, derived in projectModel.ts (deriveHosts) alongside the projects so the two can't disagree; it defaults to the LOCAL origin, and the project list is projectsForHost of the selection — one host's projects, no host suffix on the rows.

The "add" row is a <select> option, not a custom combobox: a floating list portalled into <body> sits outside the React root and never sees React 18's delegated clicks. The option is safe because it is a sentinel, not a state: interpretProjectChange maps its value to {kind: 'add'}, so onChange runs the add flow and restores the previous selection (controlled value untouched, plus a direct write back to the DOM node), and it sits alone in a leading <optgroup> so it never reads as a project. The placeholder shown when nothing is selected is disabled, so type-ahead and arrows cannot land there either.

The project select's FIRST option is "Add a new project…", over POST /api/v1/projects ({"path": …, "origin": …} — runs felt -C <path> init, which creates <path>/.felt when absent, then appends the path; idempotent). With the host already settled, that has exactly two shapes: on the local host with a dialog, POST /api/v1/choose-folder raises the host's own (Shuttle.FolderPicker — Finder via osascript, else zenity, else kdialog) and answers {ok: true, path}, {ok: false, cancelled: true} on a dismissal, or 501 when the host has none; it blocks until the human answers (bounded at five minutes). On any remote host (whose dialog would open on a desktop nobody is at) or a local daemon with no dialog, the add row instead asks for the absolute path on that host and posts it straight to /api/v1/projects, showing the owning daemon's own 400 ("not a directory: …") inline. The UI picks between the two from the native_folder_picker flag on each origin of GET /api/v1/felt-stores. Both endpoints are owner-routed. - Dispatcher (daemon/lib/shuttle/dispatcher.ex) resolves the agent, spawns the <leaf>-<uid>-shuttle tmux session. That is the only worker name, and Shuttle.ULID.from_tmux/1 is its only parser: a fiber without a ULID id: is refused with {:uid_missing, message}, a preflight refusal like the others (the blocked row, the dispatch API's 422, the preflight cooldown). - Standing roles — shuttle.kind: standing with a cron schedule:. Nothing due is stored: an armed (active, untempered) role is due when a cron occurrence has passed since it was last serviced — the later of shuttle.runtime.dispatched_at, handed_off_at and the fiber's creation (StandingRoles.standing_role_due?/1). Manual dispatch is ad-hoc (adhoc-<ms> run id). Worker exit closes the role into Awaiting review, and shuttle accept or resume re-arms it and stamps handed_off_at in the same write, so the just-served occurrence never fires again. - A finished run is finished — there is no reopen. When a oneshot's worker is gone, Dispatcher.check_resume_intent/2 decides between resuming its transcript and starting fresh. A handed_off_at newer than dispatched_at (the worker's own shuttle handoff) means fresh. No newer handoff means the session died without handing off, and then its transcript's age decides: last written within the warm window (45 minutes, config :shuttle, :resume_warm_window_s) → resume session_uuid; older, or not on this host → fresh, with the prompt naming the cut-off session and its transcript path. A transcript last written before dispatched_at is not that dispatch's session — a codex/pi launch stamps dispatched_at at once but its own id only when the scrape backfills it, so the marker can still name the predecessor — and that dispatch goes plain fresh. A resume replays the whole transcript into the model, which is cheap only while the harness's prompt cache still holds it, so a worker killed by a host outage hours ago costs less as a fresh worker reading ## Status — whatever the transcript's size. The lookup is one resolve through Shuttle.Transcript.path/2 and one stat, taken only on that no-handoff branch. A surface: app conversation is exempt from the transcript check: it keeps its identity in the Codex App Server, so it resumes whenever it did not hand off. Pinned and standing roles start fresh on the autonomous loop.

resume_mode overrides the autonomous rule: "previous" (the board's Resume) resumes unconditionally; "fresh" (New session) never resumes but still names a cut-off session; "continue" (Shuttle.Delivery, a message to a fiber with no live worker) applies the same no-handoff rule to every kind of fiber, so a message after a clean handoff or to a cold session launches fresh carrying it.

A clean handoff therefore is the end of that conversation: the next worker lands on the rewritten ## Status, and resume/reopen on a closed or awaiting fiber re-arm the document for a fresh dispatch rather than reattaching. The only reattach window is while the run is live (shuttle attach); a worker that wants a human's word before it ends stays alive at the checkpoint instead of handing off (the pinned-role contract). This is the contract, not a gap.

tmux server ownership (macOS)

  • The daemon never roots the tmux server on macOS. A worker's tmux session (Dispatcher.spawn_tmux/4) needs a tmux server to attach to, and if none exists, tmux new-session forks one as a child of whatever invoked it. Under launchd that invoker is the daemon's own beam executable, and macOS TCC charges every file access in a process tree to the tree's responsible process — not the process that actually opened the file, the executable the tree descends from. A daemon-forked server makes beam.smp the responsible process for every worker, every shell the worker opens, and every tool the worker runs, so each one raises its own "wants to access data from other apps" prompt, and the daemon binary has no way to hold the grant those prompts ask for (launchd-run processes don't keep TCC grants across restarts the way a terminal-launched one does). The plist template header documents the same fact for the daemon's own working directory; this is the same rule applied to everything the daemon spawns downstream of it.
  • The fix is to have the user's own terminal own the fork. The daemon already remote-controls kitty (Shuttle.Kitty, kitty @ launch) to open worker windows; a --type=background launch runs a command as kitty's child with no window at all. Before a dispatch or capture that finds no tmux server present, the daemon asks kitty to start one this way — an anchor session (shuttle-anchor, deliberately not -shuttle-suffixed, so nothing that scans session names for workers picks it up) that holds the server alive with nothing running in it. The fork chain is then kitty → tmux, and kitty is what TCC charges — a normal, terminal-launched process that can hold its own grants.
  • If kitty is unreachable, dispatch is refused, never silently daemon-forked. No live tmux server and no kitty remote-control socket means the daemon has no way to start one without becoming its ancestor, so it declines with {:tmux_server_unavailable, message} instead. The refusal rides the same preflight-cooldown and blocked-row machinery as every other dispatch refusal (Poller.record_dispatch_failure, the snapshot's blocked list, the 422 shape on /api/v1/dispatch and /api/v1/capture) — a human sees it on the board rather than a worker quietly inheriting a bad responsible process.
  • Attribution: how a running server is told apart from a daemon-forked one. The kernel is asked, rather than the argv guessed at. launchctl procinfo <pid> names the responsible process outright but requires root (This subcommand requires root privileges: procinfo), so it is not usable at runtime; launchctl print pid/<pid> needs no privileges and prints the process's resource coalition, whose name is the launchd label or app bundle that rooted the tree — the same attribution TCC charges file access to. So: the server pid comes from tmux display-message -p '#{pid}', and its resource-coalition name decides the origin — io.shuttle.daemon (the daemon's own launchd label) is daemon_born, any other name is user_born and is reported verbatim, an unparsable answer is unknown, and no server at all is absent. Only daemon_born is a defect. The parsing lives in internal/shuttlecli/tmux_origin.go (the resource coalition is selected by name: launchctl print emits a jetsam coalition block with the same name key, and only the resource one is TCC's). shuttle status and shuttle doctor surface the classification so a daemon-born server reads as a one-line remedy: restart it from a terminal.
  • A present server is hardened, not just accepted. tmux's exit-empty makes a server exit as soon as it holds no sessions, and the window between tmux ls answering "present" and the dispatcher's new-session is real: a human closing their last session in that window would leave new-session to fork a fresh, daemon-rooted server. So on darwin a present server also gets tmux set-option -s exit-empty off — idempotent, and unlike new-session it never forks a server of its own (with none running it just fails to connect, verified), though it is only run when a server is present or was just started.
  • Non-darwin hosts are unaffected. Every remote in this fleet is Linux, where TCC doesn't exist and a daemon-forked tmux server was never a problem; the check above is gated on os_type and is a no-op everywhere but macOS.

Dispatch prompt structure

Prompt rendering lives in Shuttle.Dispatcher.

Launch messages open with You are a Shuttle worker. Activate the felt and shuttle skills. and otherwise carry only this dispatch's facts:

  • Fiber: and Felt store:; kind, surface, and headless mode.
  • Mode: resume and Sync and re-read the fiber before continuing. on a resumed session, which may have slept for any length of time.
  • Standing run ID and scheduled/ad-hoc mode when applicable.
  • On fresh launches, Previous session: <uuid> (<harness>) when there was one. When the previous session ended without a handoff and is not resumed, the line says so and gives its transcript path.
  • Collaboration: — the assigned role and collaborator (named when the roster has exactly one of each) and the shared role store, when the fiber carries a roster.
  • From User: followed by the exact user message, when nonblank.

Syncing the store, reading the fiber and its ## Status, what to do with a predecessor's transcript, and reading role and collaborator fibers are static instructions, so they live in the shuttle skill (its Survey step and references/collaboration.md) rather than in every prompt.

Capture launches point to the shuttle skill's references/capture.md and carry JSON install metadata and the exact claim endpoint/body. That reference owns the open → install → claim → activate ordering and surface-specific claim behavior. Worker loop, headless handling, and handoff instructions live only in the skill. No decorative rules, constitution snapshots, or duplicated exit contracts are inlined. Previous transcripts support understanding; they do not supply missing instructions. Fresh and resumed workers read the current constitution and Status.

Deploy prompt changes together with the corresponding skill generation: the worker must be able to load every reference named by its launch message before that daemon build starts dispatching.