Lifecycle¶
This page covers the span from "the daemon sees an armed fiber" to "a human accepts the result."
The worker loop¶
The daemon starts one worker per eligible fiber. A terminal worker is a tmux session named
<slug-leaf>-<uid>-shuttle, keyed by the fiber's intrinsic id: (a ULID),
running the agent CLI in shuttle.project_dir. A fiber without an id has no
worker name, so the daemon refuses to dispatch it and the board shows it
blocked; felt backfill-ids gives every fiber in a store one, and every felt
write stamps one on a fiber that lacks it. shuttle session-name <fiber>
prints the name. The daemon composes a deliberately thin
prompt: the fiber id, the felt store path, an exit contract, and an optional
per-dispatch "From User" directive.
A Codex app worker (shuttle.surface: app) runs in the managed App Server and
belongs to the project matching shuttle.project_dir. Shuttle persists its
conversation identity before starting a turn. The conversation stays owned
while idle, survives a Shuttle restart, and can receive replies from another
app client. Existing blocks without surface use the terminal surface.
It does not paste the constitution into the prompt. The worker reads the fiber fresh from disk, so it picks up an edit you make mid-session instead of freezing a stale snapshot.
From there the worker:
- Surveys — reads the constitution, the previous
## Statushandoff, thegit login and around the project directory, and any sub-fibers. - Works — picks the highest-value slice itself. The constitution describes what "done" looks like; the worker sequences the steps.
- Writes back — rewrites
outcome:, rewrites## Status, corrects the spec if the session sharpened it, files findings as sub-fibers, commits. - Exits — with exactly one Shuttle lifecycle verb: to continue, runs
shuttle handoff <fiber>as its final action; to stop, runsshuttle close <fiber>and does nothing else.
To continue, app workers run env -u TMUX shuttle -C <felt-store> handoff
<fiber>, then finish the turn; to stop, they run
shuttle -C <felt-store> close <fiber> and finish the turn without a handoff
call. The daemon releases ownership once
that turn is idle. A normal final reply without a handoff or a close keeps the
conversation available for the human. Stop interrupts an app turn and releases
ownership; Resume uses that same conversation, while New session explicitly
starts another one. Unknown connection state never authorizes a duplicate.
Workers should exit earlier than feels natural. A clean handoff at half a
context window beats pushing through a compaction. ## Status plus the
constitution recovers most of a warm world-model on the next dispatch.
Exit semantics¶
The worker asks three questions, in order. The answer sets status, and
status decides what happens next.
1. Is the desired state realized? Run shuttle close <fiber> and exit.
The card lands in Awaiting review. A human accepts it or resumes it.
Substantive work needs independent fresh-eyes review before the session that
produced it closes — a subagent reviewer in-session, or leave the fiber active
and let the next dispatch review it cold. Edits to the fiber's own surfaces
(spec, ## Status, outcome, report.html) count as handoff, not work
product, and never block a close.
2. Blocked on something only a human can supply? Run
shuttle close <fiber>, and lead the outcome with Blocked: … so the card
reads as a question.
3. More work, not blocked? Leave the fiber active and just hand off. The
daemon starts a fresh worker next tick, and it lands on your ## Status.
Closing parks the work¶
shuttle obeys the vocabulary literally. Know which word does what.
| You say | The worker does | Next |
|---|---|---|
| "hand off" | case 3 — status stays active, then handoff |
Daemon redispatches |
| "close it out" | Run shuttle close <fiber>; it sets status: closed — never handoff |
Card waits for you; no new worker |
Closing puts the work back on a human's desk. It claims nothing about completion. A worker should never upgrade a close-out into a continuation because the work looks unfinished — unfinished is often exactly why you want it back.
Tempered — human only¶
tempered carries the verdict, and a worker never sets it to true.
tempered |
Meaning |
|---|---|
| absent | Awaiting review |
true |
Accepted |
false |
Discarded — mooted or superseded |
Workers also never uninstall their own shuttle: block. Closing and
uninstalling are separate decisions, and the block stays as historical record.
Resume vs fresh¶
Workers start fresh by default. Resuming a transcript happens in three cases:
you press Resume on the board; a oneshot died dirty and its transcript is still
warm; or you send a message to a fiber whose worker died dirty while its
transcript is still warm. A death counts as dirty when handed_off_at is
missing or older than dispatched_at, which is why the handoff verb matters.
A transcript is warm when it was written in the last 45 minutes, while the
model's prompt cache still holds it; resuming a colder one would replay it
uncached, so the worker starts fresh instead and its prompt names the cut-off
session and its transcript. A Codex app conversation (surface: app) that
died dirty resumes whatever its transcript's age.
New session on the board always starts fresh. Scheduled runs, ad-hoc standing runs, daemon recovery, and orphan adoption start fresh too.
The board¶
The daemon serves a board at http://127.0.0.1:4000/: one page, three full-page
views behind a hotkey row.
| Key | View | What it answers |
|---|---|---|
1 |
Desk | What needs doing, and what is running right now — the kanban |
2 |
Chronicle | What a stretch of weeks was about, under a strip of cycle bands |
3 |
Board | What the work produced — every file a worker sent, rendered on a canvas |
The board holds no state of its own. It views the same fibers the daemon polls, plus tmux liveness and the host-local ledgers.
The board covers all three in depth — column and horizon rules, the snooze gesture, Attach, how Chronicle fetches its record, and how to build the bundle.
Dispatch eligibility¶
A fiber dispatches if and only if all of these hold. The daemon evaluates them
in this order (eligible?/2 and dispatch_gates_pass?/3 in
daemon/lib/shuttle/poller.ex).
- It lives in a felt store the daemon polls.
- It carries a
shuttle:block. That block alone defines "shuttle-managed"; no tag predicate exists. shuttle.hostequals this daemon's own host id.- felt-native
statusisactive. - No worker is already running or claimed for it.
- The resume-loop circuit breaker is closed.
- The boot quarantine is released.
shuttle.project_dir, when the block declares one, exists on this host. A block without one still dispatches, with the felt store as the worker's directory; felt's arming verbs refuse to arm such a block.
depends_on is not on this list. It is an ordering annotation for the board
("this is filed after that") and carries no dispatch meaning.
Configured stores come from SHUTTLE_STORES (comma-separated) or the persisted
registry at ~/.config/shuttle/stores.json. shuttle assumes no default store.
The circuit breaker (7) exists because a worker that dies on startup would
otherwise be relaunched forever. Five consecutive worker deaths, each under 90
seconds, pause autonomous dispatch for that fiber for ten minutes and surface it
as blocked. A healthy run or a force-dispatch clears it.
Boot quarantine¶
A daemon restart is not dispatch authority. On every start, the daemon
parks every candidate it has never observed running into pending_launch and
dispatches nothing fresh. Work it did observe alive under its own uptime —
adopted at boot, or dispatched since — resumes normally. That counts as
continuation, not a fresh launch. A due standing role also passes through:
its cron occurrence is a fixed-time "go" the human already gave, bounded to one
run, so a restart that straddles 09:00 does not eat the run. (Under CLI/daemon
contract skew it holds like everything else.)
The guard exists because a daemon that restarts repeatedly (an overloaded machine crash-looping, for example) would otherwise treat every restart as license to relaunch every armed, workerless fiber it can see — a burst of redundant, token-burning launches on each crash.
Release is manual — no timeout — with one narrow exception below.
shuttle daemon release
A human force-dispatch bypasses the quarantine without clearing it.
The one automatic exit: a proven fast bounce¶
Every restart someone asked for holds: a deploy, make restart, make stop, a
supervisor restart (systemctl --user restart, launchctl kickstart -k). Those
all stop the daemon with SIGTERM, and a SIGTERM'd daemon touches a stop marker,
$SHUTTLE_DATA_DIR/heartbeat.stopped, first thing on the way down. make stop,
shuttle daemon stop, shuttle daemon install and bin/shuttle-deploy touch
it themselves before they signal — in the data_dir that shuttle host
--json reports, so they apply the daemon's own trim and ~ rule — and so does
re-running bin/shuttle-launch, whose tmux kill-session SIGHUPs the daemon
with no shutdown at all. The next boot sees the marker and holds.
The exception is a daemon killed hard — a kernel CPU-rlimit SIGKILL on a
capped cluster login node, say — and respawned seconds later. Its workers keep
running (tmux owns them) and nothing went stale, yet a hold there would stop all
new work until someone noticed. A hard kill runs no shutdown code and touches no
marker. Only such a host needs the exception, so it is off unless the host opts
in, in ~/.config/shuttle/host.json (or $SHUTTLE_HOST_CONFIG_FILE):
{"class": "shared-multi-user", "quarantine_auto_release": true}
On any other host a SIGKILL is an out-of-memory kill or a person's kill -9,
and holding is the right answer. With the key absent or anything but true,
every boot holds and logs automatic release is off for this host.
On every host the daemon rewrites $SHUTTLE_DATA_DIR/heartbeat.json (default
~/.shuttle/heartbeat.json) every 10 seconds, and at once when the hold is
released: the time of the write, when this incarnation booted, its host id, the
machine's node name and the daemon's OS pid, whether it was still held, the
workers it has live, and a short ring of recent boot times. On an opted-in host
the next boot reads the file before adoption and judges it after adoption. The
hold lifts by itself only when all of these hold:
- the heartbeat was written by this host id on this machine (a
$HOMEshared across login nodes can hold another node's heartbeat), by a different daemon process (a restart inside the running daemon is not a hard kill); - no stop marker is as new as that incarnation's boot;
- that incarnation had been released: an unreleased hold survives hard kills, so a daemon that booted held after an outage or a deploy stays held;
- the heartbeat is less than 60 seconds old, and newer than the machine's own boot;
- every worker it recorded is live now, established by this daemon's own tmux adoption rather than by trusting the file — and none of them is an app worker, whose liveness adoption cannot observe;
- the previous incarnation ran at least 90 seconds, and the daemon has booted at most 3 times in the last 10 minutes, counting this boot.
Anything else holds: a graceful stop, a stale heartbeat (a real outage), a crash loop, a missing or malformed file, a boot whose tmux scan failed, or a CLI/daemon contract skew — which has no release path at all, automatic or manual. The daemon log records the verdict and its reason either way.
CLI verbs¶
felt owns fiber content and generic frontmatter. It preserves shuttle: as
opaque data; it does not validate the block or write Shuttle lifecycle state.
shuttle owns the block's schema, resolved reads, host/fleet configuration,
and lifecycle. Its local writes validate before touching disk; status, snapshot,
and dispatch can also ask the local daemon. The
CLI reference lists the commands.
shuttle check # validate Shuttle blocks in a store
shuttle ls --has-field shuttle # fiber listing with resolved Shuttle data
shuttle show <fiber> # read one fiber with its resolved facet
shuttle status [<fiber>] # offline eligibility and dispatch state
shuttle snapshot # local daemon snapshot
shuttle dispatch <fiber> [--ad-hoc] # ask the daemon to dispatch
shuttle daemon start [--force] # start the local Mix release
shuttle daemon stop # stop it and record an intentional stop
shuttle daemon status # state JSON; exit 2 when the daemon is down
shuttle daemon release # clear the boot quarantine
shuttle daemon reset <remote> # reset a remote's circuit breaker
shuttle daemon install # install the per-user keep-alive
shuttle daemon uninstall # remove the per-user keep-alive
shuttle doctor # diagnose host, listener, and daemon contract
shuttle version # daemon version, or local Mix release version
The daemon uses felt for fiber content and generic writes, and shuttle for
resolved reads and Shuttle-owned operations. shuttle daemon status,
release, and reset contact the local daemon over HTTP; start and stop
manage the local release, while install and uninstall manage the
keep-alive.
The daemon also speaks HTTP under /api/v1, in four groups: a write plane
(dispatch, transition, kill, lifecycle, felt-edit, …), a read
plane (fibers, agents, felt-stores, …), a temporal read plane the
board's time views live on (activity, sessions, commits,
sent-files/all, each with a /composite sibling that fans in every host a
hub aggregates — the fleet), and
operator routes (state, version, the manual gate releases). The API
reference tabulates them.
curl -s http://127.0.0.1:4000/api/v1/agents | jq
When to uninstall¶
Closing a fiber leaves its block in place, and that is deliberate. Uninstall earns its keep in four cases:
- Mistake recovery — wrong slug, immediate undo.
- Full rebuild — though
shuttle reshapenow covers most of the "change the kind or schedule" case in place, without a rebuild. - Archiving — a closed fiber's card leaves the board entirely.
- Handing ownership to a different dispatcher.
Never uninstall to end a worker session. Use the ordinary exit — shuttle handoff to continue, shuttle close to stop.
Diagnosing a missing card¶
Most "my card isn't showing" reduces to "no block installed yet." Check in this
order: is the fiber in a store the daemon polls, does shuttle status show
a block, is status: active, does shuttle.host match, and is the quarantine
released?
For multi-host setups, remote fibers reach the board over an SSH tunnel from the owning daemon — never via git sync. A git mirror may replicate a remote fiber's files locally. That replication is incidental; treat any behaviour that depends on it as a bug. If a remote card is missing, debug the tunnel, not the git state.
Remotes come from one config file
~/.config/shuttle/remotes.json lists every remote daemon: its name, its
local forwarded port, and how to reach it. The CLI and the daemon both read
it at runtime. Manage it with shuttle remotes list|add|rm|path. A
single-machine setup needs no such file — see Configuring
remotes.