Architecture¶
One repo, one checkout, four artifacts: the felt CLI and shuttle CLI
(Go), the shuttle daemon (Elixir/OTP), and the board UI (TypeScript,
ui/). felt is a lean fiber library and notes CLI; shuttle is the orchestration
layer built on felt. The daemon is the production dispatcher.
Architecture stance¶
- Two Go CLIs, one-way dependency.
feltowns fiber storage and generic frontmatter.shuttleowns the network and orchestration, and reads and writes fibers through felt's Go library. felt has no dependency on shuttle or messaging packages; a test enforces that import boundary. - One repo, one checkout, four artifacts.
cmd/feltandcmd/shuttlebuild the two CLIs, while the same source tree builds the daemon release and board UI. Shuttle needs no separate browser application or dispatcher. - The
shuttle:block is opaque to felt. felt preserves it as ordinary frontmatter. shuttle owns the schema, validation, resolution, and lifecycle fields. The daemon reads and writes fiber content through felt and shells shuttle for Shuttle-owned behavior. Continuation state (session_uuid/dispatched_at/handed_off_at) lives in the block. - Resolved views belong to shuttle.
shuttle lsandshuttle showuse felt's fiber readers and add Shuttle's resolved facet; felt's JSON output contains the plain frontmatter without aresolvedsub-key.
The Elixir/OTP daemon is the production dispatcher. Dispatch, the
per-worker watcher, and the :4000 API are where OTP earns its keep. A Go
rewrite collapsing everything into a single binary is a deferred, must-earn-
itself idea.
Continuity across dispatches lives in frontmatter and git, not an event log.
The daemon detects clean worker exits via the shuttle.runtime.handed_off_at
frontmatter field. The editorial chain lives in the constitution body's
## Status block plus the git log of the fiber.
Stores and views¶
Fiber documents synchronize through ordinary Git. Workers use felt sync
before substantive work, edit locally, and publish committed changes with
felt sync --push. A worker resolves relevant conflicts with the work's
context. Roles and collaborators under roles/<role>/<collaborator>/ have
intrinsic IDs and no host owner. A task's optional collaboration map records
readable collaborator slugs by role; it does not select an execution model.
Synchronizing a constitution does not change its shuttle.host dispatch gate.
Host-addressed board requests still use the daemon's routing for that host's
current content, artifacts, and execution controls.
A store is a .felt/ directory and everything under it — one namespace,
one git repo. A view (or substore) is a project whose .felt is a symlink
into a subdirectory of a larger store: this repo's .felt points into
~/loom/.felt/ai-futures/felt. The loom is the store; the project is a lens on
it. A lens narrows what is listed, never what can be found or reached.
An id is resolved against the local store first, then — on a miss — against
the enclosing store, which resolves it by its own scope and suffix rules from
this view's position out there. A hit is reported as external and the command
runs where the fiber lives, appending (in <root>) so a cross-store write is
never silent. The one thing that never crosses is the basename rescue (the
"your path went stale but the slug is unique" guess): the external probe sits
above it, so felt never guesses at a fiber and then deletes or moves it.
The verb contract, in one line each:
felt lslists the view. Every flag filters this store's own listing. Fast, local, always — in a substore a filtered ls closes with a line namingfelt findand the store it would search.felt findsearches the store. Local hits first, the rest of the enclosing store under a separator naming it, each by its full id there, capped with an exact remainder count (--limit).-jmerges both halves into one array, each entry naming itsstore.- An id reaches anywhere.
show,edit,rm,nest,tree, and everyshuttleverb acts on the fiber the id names, in the store that holds it.
The shuttle: block¶
The shuttle: block is opaque frontmatter to felt. felt preserves it like
any other unknown field and does not validate or resolve it. The shuttle CLI
owns its typed facet and lifecycle writes; the Elixir daemon reads fiber content
with felt, then uses shuttle ls/shuttle show when it needs resolved
Shuttle data and shuttle for Shuttle-owned operations. shuttle contract
checks the CLI/daemon contract at boot, and a test keeps felt's import graph
free of shuttle and messaging packages.
Execution surfaces¶
The agent registry determines the harness, model, and reasoning effort.
shuttle.surface chooses how a Codex conversation runs: cli uses a tmux
worker; app uses the local Codex App Server. An existing block without a
surface retains CLI execution. New Codex selections in the board default to
app execution, with CLI available explicitly. Other harnesses use CLI.
The App Server owns the conversation and its turns. Shuttle stores the conversation identity and task assignment durably on the owning host. Capture creates an unassigned conversation that can claim a fiber; an ordinary dispatch assigns the conversation immediately. Both use the same execution backend. A completed turn can wait for a reply without releasing its task or starting another conversation. An explicit stop or completed handoff releases ownership.
App transport errors preserve uncertainty instead of becoming worker-death signals. Resume targets the recorded conversation and reports failures; creating a fresh conversation is an explicit operation. CLI process liveness continues to come from tmux. API responses distinguish these surfaces instead of presenting an app conversation as a terminal session.
The app surface requires the managed App Server's Unix control socket. A desktop application's private stdio App Server is a separate runtime: it may share persisted rollouts through the same Codex home, but it does not share loaded-turn ownership and is not a safe fallback. Shuttle refuses app dispatch when the managed socket is absent; choose a host with a reachable managed App Server or use the CLI surface. Codex documents the managed daemon lifecycle in its App Server daemon reference. On macOS, desktop attachment to the managed daemon also depends on the desktop transport-selection gate; current failures caused by desktop-injected config overrides are tracked upstream in openai/codex#41014.
Phone access uses the host's configured Codex remote access. App Server creation and a direct link into the ChatGPT app are separate capabilities: session identity does not imply a working phone URL. Shuttle must not invent a cloud-task URL for a local conversation.
Trust boundaries¶
The daemon's control plane has no authentication: anything that reaches its listener can read and write every registered fiber, read transcripts, and launch or kill workers as the user running it. The stance this follows from is that the listener is the boundary, so the listener has to match the trust of the host it runs on rather than the daemon carrying its own authentication layer.
A host declares its trust class — single-user, shared-multi-user, or
exposed — in ~/.config/shuttle/host.json, set with shuttle host class
and read with shuttle host --json, which also reports the host id. The
daemon reads host.json with the same rules, but takes its host id from
shuttle at boot. The class is a declared fact the operator asserts, not
something the daemon infers from the host; shuttle doctor checks that the
assertion still holds against reality (logged-in users, the socket directory's
mode, every fleet-owned listening socket).
The class determines where the daemon binds. single-user listens on
tcp://127.0.0.1:4000 (SHUTTLE_PORT). shared-multi-user and exposed
listen on the Unix socket $SHUTTLE_DATA_DIR/sock/daemon.sock (default
~/.shuttle/sock/daemon.sock), inside a 0700
directory the daemon creates and verifies before binding — a directory found
with looser permissions is a refusal to bind, not a downgrade. host.json's
listen key or SHUTTLE_LISTEN overrides either default, and the CLI
resolves the same address to reach the daemon it is driving. A Unix socket
changes who can even open a connection: a caller has to traverse the
filesystem to the socket path, so only sshd running as the host's owner (an
SSH tunnel's far end) or the local CLI can reach it, which is a permission
check the kernel enforces rather than one the daemon has to implement. A host
whose only inbound is SSH tunnels can run entirely off the socket.
An unprivileged userspace tailscaled cannot target a Unix socket with
tailscale serve, and the macOS system tailscaled cannot reach a filesystem
socket at all. A shared host using either arrangement needs a loopback TCP
listener. An exposed host instead places a front proxy before the daemon; that
proxy dials the daemon's Unix socket.
On shared TCP listeners, PeerGatePlug runs before static assets or
request-body parsing and admits only a peer whose uid from /proc/net/tcp or
/proc/net/tcp6 matches the daemon's effective uid. An unresolved or foreign
uid receives a 403 response. The plug also refuses exposed TCP requests that
reach it, while daemon boot refuses to create an exposed TCP listener.
The gate identifies the last local process, not the original client. A relay
running as the daemon's owner can pass co-tenant traffic with that owner's uid;
examples include userspace Tailscale SOCKS/HTTP proxies, ssh -D/-L, socat,
and code-server or Jupyter /proxy/ routes. Operators must not run relays as
themselves. Root has no separate admission exception; it can already inspect
or control the daemon process.
The daemon refuses defaults.https_proxy on shared and exposed hosts because
the local Tailscale HTTP proxy is an unauthenticated loopback gateway to the
whole tailnet.
For https:// remotes, defaults.tailscale_socket asks the daemon to open a
private, owner-only Unix socket per remote and dial the remote through
tailscaled's LocalAPI. On shared/exposed hosts using the default Unix listener,
bridge sockets sit beside daemon.sock under the same 0700 directory guard.
The bridge verifies the remote TLS certificate and hostname before it relays
traffic; it never exposes a loopback proxy to co-tenants.
When the private socket is configured, HTTPS requests fail closed if their
bridge is missing or unavailable; they never fall back to direct or proxy
routing. The two transport defaults are mutually exclusive. Tailnet discovery
uses the same socket: it reads the status from the LocalAPI and dials each probe
through it with the same certificate and hostname checks.
Every connection gets peer facts in conn.assigns.peer: transport (unix or
tcp), TCP uid when /proc resolves it, whether forwarding headers are
present, and a tailscale_login header. The uid gate consumes those facts on
shared and exposed TCP listeners; single-user TCP and Unix connections bypass
that gate. Unix reachability is bounded by the socket directory's filesystem
permissions. The tailscale_login value is retained on TCP only after the uid
gate admits the peer, but the header remains an assertion rather than an
independent credential. A shared TCP listener refuses to boot when
/proc/net/tcp is unreadable; an exposed TCP listener is refused regardless
of /proc. Use the class's Unix socket or declare the host single-user when
loopback is private to its operator.
The fleet¶
Every daemon finds the others itself. Shuttle.TailnetPeers reads the tailnet
status, keeps the peers owned by this node's own Tailscale user, probes each
online one at https://<magicdns-name>/api/v1/version, and names every
Shuttle daemon that answers by the host id it reports there. Fibers route by
shuttle.host, so a peer is named by its host id, never by its MagicDNS label.
A node shared in from another tailnet is never probed or trusted; a peer that
reports this daemon's own id, reports no id, or shares an id with another peer
is dropped. Phones and other nodes without a daemon drop out because they do
not answer. The round runs shortly after boot, off the boot path, and every
minute after that. A peer that answered once survives ten minutes of failed
probes, so a daemon restarting for a deploy goes stale on the board instead of
vanishing from it.
~/.config/shuttle/remotes.json holds the exceptions: hosts outside the
tailnet (reached through an ssh tunnel port), non-default URLs, ssh names for
the recovery cascade, and deploy metadata. Shuttle.Remotes.resolve/2 merges
the two sources. A configured entry wins over a discovered peer with the same
name or https authority, "enabled": false suppresses a discovered host, and
defaults.discover: false turns discovery off. The resolved fleet feeds the
registries, Shuttle.OriginRouter, messaging, and the dial bridges. A newly
discovered peer reaches them without a restart, because the discovery
generation is part of Shuttle.Remotes.config_token/0.
The Go CLI does not probe the tailnet. shuttle remotes list and every verb
that routes by host name read the local daemon's discovered peers from
/api/v1/version and merge them with the file through resolveRemotes, the Go
mirror of the Elixir resolver. daemon/test/fixtures/tailnet_peers/ holds the
shared cases. When Tailscale is absent, stopped, or failing, the daemon uses
the file alone and records why, and shuttle doctor reports the reason.
Platform story¶
Linux and macOS are both supported for single-host use. One host runs the
shuttle CLI, daemon, board, and workers on either OS. The keep-alive differs —
a launchd LaunchAgent on macOS, a systemd user unit on Linux (shuttle daemon
install picks the branch from uname -s) — and so does the log path, but the
CLI, daemon, and bundle are built from this repo.
Either platform can be the fleet's hub. shuttle tunnels install
writes the hub's autossh jobs as launchd LaunchAgents on macOS and systemd
--user units on Linux, picking the branch the way shuttle daemon install does.
A Linux host with no systemd user session (an HPC login node usually has none)
gets a refusal naming --write-only, not units nothing would start. Either way
the tunnel is installed on the hub, not on the remote. One asymmetry remains:
the daemon's recovery cascade bounces a stalled tunnel with launchctl
kickstart, so on a Linux hub a quiet remote skips the bounce and advances to
the cascade's ssh check.
kitty attach is terminal lock-in, not platform lock-in. Attach opens the
worker's tmux session in kitty via kitty's remote-control CLI, and kitty runs on
Linux. What is mac-specific is only the osascript call that raises the kitty
window, and that is already a no-op off macOS (activate/1 in
daemon/lib/shuttle/kitty.ex). A non-kitty user gets nothing on either OS; shuttle attach <fiber> always works.
Windows is unsupported.
Explicit Codex endpoint¶
SHUTTLE_CODEX_SOCKET selects the same Unix WebSocket endpoint for app dispatch
and native session messaging. Set it in the Shuttle daemon's environment when
using an explicitly shared Desktop backend. Without it, both clients use
$CODEX_HOME/app-server-control/app-server-control.sock, with CODEX_HOME
defaulting to ~/.codex. An empty override uses that default. An unavailable
explicit endpoint is an error; clients do not switch to another conversation
owner or start a replacement backend.