Skip to content

Architecture

One repo, one checkout, four artifacts: the felt CLI and shuttle CLI (Go), the shuttle daemon (Elixir/OTP), and the board UI (TypeScript, ui/). felt is a lean fiber library and notes CLI; shuttle is the orchestration layer built on felt. The daemon is the production dispatcher.

Architecture stance

  • Two Go CLIs, one-way dependency. felt owns fiber storage and generic frontmatter. shuttle owns the network and orchestration, and reads and writes fibers through felt's Go library. felt has no dependency on shuttle or messaging packages; a test enforces that import boundary.
  • One repo, one checkout, four artifacts. cmd/felt and cmd/shuttle build the two CLIs, while the same source tree builds the daemon release and board UI. Shuttle needs no separate browser application or dispatcher.
  • The shuttle: block is opaque to felt. felt preserves it as ordinary frontmatter. shuttle owns the schema, validation, resolution, and lifecycle fields. The daemon reads and writes fiber content through felt and shells shuttle for Shuttle-owned behavior. Continuation state (session_uuid / dispatched_at / handed_off_at) lives in the block.
  • Resolved views belong to shuttle. shuttle ls and shuttle show use felt's fiber readers and add Shuttle's resolved facet; felt's JSON output contains the plain frontmatter without a resolved sub-key.

The Elixir/OTP daemon is the production dispatcher. Dispatch, the per-worker watcher, and the :4000 API are where OTP earns its keep. A Go rewrite collapsing everything into a single binary is a deferred, must-earn- itself idea.

Continuity across dispatches lives in frontmatter and git, not an event log. The daemon detects clean worker exits via the shuttle.runtime.handed_off_at frontmatter field. The editorial chain lives in the constitution body's ## Status block plus the git log of the fiber.

Stores and views

Fiber documents synchronize through ordinary Git. Workers use felt sync before substantive work, edit locally, and publish committed changes with felt sync --push. A worker resolves relevant conflicts with the work's context. Roles and collaborators under roles/<role>/<collaborator>/ have intrinsic IDs and no host owner. A task's optional collaboration map records readable collaborator slugs by role; it does not select an execution model. Synchronizing a constitution does not change its shuttle.host dispatch gate. Host-addressed board requests still use the daemon's routing for that host's current content, artifacts, and execution controls.

A store is a .felt/ directory and everything under it — one namespace, one git repo. A view (or substore) is a project whose .felt is a symlink into a subdirectory of a larger store: this repo's .felt points into ~/loom/.felt/ai-futures/felt. The loom is the store; the project is a lens on it. A lens narrows what is listed, never what can be found or reached.

An id is resolved against the local store first, then — on a miss — against the enclosing store, which resolves it by its own scope and suffix rules from this view's position out there. A hit is reported as external and the command runs where the fiber lives, appending (in <root>) so a cross-store write is never silent. The one thing that never crosses is the basename rescue (the "your path went stale but the slug is unique" guess): the external probe sits above it, so felt never guesses at a fiber and then deletes or moves it.

The verb contract, in one line each:

  • felt ls lists the view. Every flag filters this store's own listing. Fast, local, always — in a substore a filtered ls closes with a line naming felt find and the store it would search.
  • felt find searches the store. Local hits first, the rest of the enclosing store under a separator naming it, each by its full id there, capped with an exact remainder count (--limit). -j merges both halves into one array, each entry naming its store.
  • An id reaches anywhere. show, edit, rm, nest, tree, and every shuttle verb acts on the fiber the id names, in the store that holds it.

The shuttle: block

The shuttle: block is opaque frontmatter to felt. felt preserves it like any other unknown field and does not validate or resolve it. The shuttle CLI owns its typed facet and lifecycle writes; the Elixir daemon reads fiber content with felt, then uses shuttle ls/shuttle show when it needs resolved Shuttle data and shuttle for Shuttle-owned operations. shuttle contract checks the CLI/daemon contract at boot, and a test keeps felt's import graph free of shuttle and messaging packages.

Execution surfaces

The agent registry determines the harness, model, and reasoning effort. shuttle.surface chooses how a Codex conversation runs: cli uses a tmux worker; app uses the local Codex App Server. An existing block without a surface retains CLI execution. New Codex selections in the board default to app execution, with CLI available explicitly. Other harnesses use CLI.

The App Server owns the conversation and its turns. Shuttle stores the conversation identity and task assignment durably on the owning host. Capture creates an unassigned conversation that can claim a fiber; an ordinary dispatch assigns the conversation immediately. Both use the same execution backend. A completed turn can wait for a reply without releasing its task or starting another conversation. An explicit stop or completed handoff releases ownership.

App transport errors preserve uncertainty instead of becoming worker-death signals. Resume targets the recorded conversation and reports failures; creating a fresh conversation is an explicit operation. CLI process liveness continues to come from tmux. API responses distinguish these surfaces instead of presenting an app conversation as a terminal session.

The app surface requires the managed App Server's Unix control socket. A desktop application's private stdio App Server is a separate runtime: it may share persisted rollouts through the same Codex home, but it does not share loaded-turn ownership and is not a safe fallback. Shuttle refuses app dispatch when the managed socket is absent; choose a host with a reachable managed App Server or use the CLI surface. Codex documents the managed daemon lifecycle in its App Server daemon reference. On macOS, desktop attachment to the managed daemon also depends on the desktop transport-selection gate; current failures caused by desktop-injected config overrides are tracked upstream in openai/codex#41014.

Phone access uses the host's configured Codex remote access. App Server creation and a direct link into the ChatGPT app are separate capabilities: session identity does not imply a working phone URL. Shuttle must not invent a cloud-task URL for a local conversation.

Trust boundaries

The daemon's control plane has no authentication: anything that reaches its listener can read and write every registered fiber, read transcripts, and launch or kill workers as the user running it. The stance this follows from is that the listener is the boundary, so the listener has to match the trust of the host it runs on rather than the daemon carrying its own authentication layer.

A host declares its trust class — single-user, shared-multi-user, or exposed — in ~/.config/shuttle/host.json, set with shuttle host class and read with shuttle host --json, which also reports the host id. The daemon reads host.json with the same rules, but takes its host id from shuttle at boot. The class is a declared fact the operator asserts, not something the daemon infers from the host; shuttle doctor checks that the assertion still holds against reality (logged-in users, the socket directory's mode, every fleet-owned listening socket).

The class determines where the daemon binds. single-user listens on tcp://127.0.0.1:4000 (SHUTTLE_PORT). shared-multi-user and exposed listen on the Unix socket $SHUTTLE_DATA_DIR/sock/daemon.sock (default ~/.shuttle/sock/daemon.sock), inside a 0700 directory the daemon creates and verifies before binding — a directory found with looser permissions is a refusal to bind, not a downgrade. host.json's listen key or SHUTTLE_LISTEN overrides either default, and the CLI resolves the same address to reach the daemon it is driving. A Unix socket changes who can even open a connection: a caller has to traverse the filesystem to the socket path, so only sshd running as the host's owner (an SSH tunnel's far end) or the local CLI can reach it, which is a permission check the kernel enforces rather than one the daemon has to implement. A host whose only inbound is SSH tunnels can run entirely off the socket. An unprivileged userspace tailscaled cannot target a Unix socket with tailscale serve, and the macOS system tailscaled cannot reach a filesystem socket at all. A shared host using either arrangement needs a loopback TCP listener. An exposed host instead places a front proxy before the daemon; that proxy dials the daemon's Unix socket.

On shared TCP listeners, PeerGatePlug runs before static assets or request-body parsing and admits only a peer whose uid from /proc/net/tcp or /proc/net/tcp6 matches the daemon's effective uid. An unresolved or foreign uid receives a 403 response. The plug also refuses exposed TCP requests that reach it, while daemon boot refuses to create an exposed TCP listener.

The gate identifies the last local process, not the original client. A relay running as the daemon's owner can pass co-tenant traffic with that owner's uid; examples include userspace Tailscale SOCKS/HTTP proxies, ssh -D/-L, socat, and code-server or Jupyter /proxy/ routes. Operators must not run relays as themselves. Root has no separate admission exception; it can already inspect or control the daemon process.

The daemon refuses defaults.https_proxy on shared and exposed hosts because the local Tailscale HTTP proxy is an unauthenticated loopback gateway to the whole tailnet.

For https:// remotes, defaults.tailscale_socket asks the daemon to open a private, owner-only Unix socket per remote and dial the remote through tailscaled's LocalAPI. On shared/exposed hosts using the default Unix listener, bridge sockets sit beside daemon.sock under the same 0700 directory guard. The bridge verifies the remote TLS certificate and hostname before it relays traffic; it never exposes a loopback proxy to co-tenants. When the private socket is configured, HTTPS requests fail closed if their bridge is missing or unavailable; they never fall back to direct or proxy routing. The two transport defaults are mutually exclusive. Tailnet discovery uses the same socket: it reads the status from the LocalAPI and dials each probe through it with the same certificate and hostname checks.

Every connection gets peer facts in conn.assigns.peer: transport (unix or tcp), TCP uid when /proc resolves it, whether forwarding headers are present, and a tailscale_login header. The uid gate consumes those facts on shared and exposed TCP listeners; single-user TCP and Unix connections bypass that gate. Unix reachability is bounded by the socket directory's filesystem permissions. The tailscale_login value is retained on TCP only after the uid gate admits the peer, but the header remains an assertion rather than an independent credential. A shared TCP listener refuses to boot when /proc/net/tcp is unreadable; an exposed TCP listener is refused regardless of /proc. Use the class's Unix socket or declare the host single-user when loopback is private to its operator.

The fleet

Every daemon finds the others itself. Shuttle.TailnetPeers reads the tailnet status, keeps the peers owned by this node's own Tailscale user, probes each online one at https://<magicdns-name>/api/v1/version, and names every Shuttle daemon that answers by the host id it reports there. Fibers route by shuttle.host, so a peer is named by its host id, never by its MagicDNS label. A node shared in from another tailnet is never probed or trusted; a peer that reports this daemon's own id, reports no id, or shares an id with another peer is dropped. Phones and other nodes without a daemon drop out because they do not answer. The round runs shortly after boot, off the boot path, and every minute after that. A peer that answered once survives ten minutes of failed probes, so a daemon restarting for a deploy goes stale on the board instead of vanishing from it.

~/.config/shuttle/remotes.json holds the exceptions: hosts outside the tailnet (reached through an ssh tunnel port), non-default URLs, ssh names for the recovery cascade, and deploy metadata. Shuttle.Remotes.resolve/2 merges the two sources. A configured entry wins over a discovered peer with the same name or https authority, "enabled": false suppresses a discovered host, and defaults.discover: false turns discovery off. The resolved fleet feeds the registries, Shuttle.OriginRouter, messaging, and the dial bridges. A newly discovered peer reaches them without a restart, because the discovery generation is part of Shuttle.Remotes.config_token/0.

The Go CLI does not probe the tailnet. shuttle remotes list and every verb that routes by host name read the local daemon's discovered peers from /api/v1/version and merge them with the file through resolveRemotes, the Go mirror of the Elixir resolver. daemon/test/fixtures/tailnet_peers/ holds the shared cases. When Tailscale is absent, stopped, or failing, the daemon uses the file alone and records why, and shuttle doctor reports the reason.

Platform story

Linux and macOS are both supported for single-host use. One host runs the shuttle CLI, daemon, board, and workers on either OS. The keep-alive differs — a launchd LaunchAgent on macOS, a systemd user unit on Linux (shuttle daemon install picks the branch from uname -s) — and so does the log path, but the CLI, daemon, and bundle are built from this repo.

Either platform can be the fleet's hub. shuttle tunnels install writes the hub's autossh jobs as launchd LaunchAgents on macOS and systemd --user units on Linux, picking the branch the way shuttle daemon install does. A Linux host with no systemd user session (an HPC login node usually has none) gets a refusal naming --write-only, not units nothing would start. Either way the tunnel is installed on the hub, not on the remote. One asymmetry remains: the daemon's recovery cascade bounces a stalled tunnel with launchctl kickstart, so on a Linux hub a quiet remote skips the bounce and advances to the cascade's ssh check.

kitty attach is terminal lock-in, not platform lock-in. Attach opens the worker's tmux session in kitty via kitty's remote-control CLI, and kitty runs on Linux. What is mac-specific is only the osascript call that raises the kitty window, and that is already a no-op off macOS (activate/1 in daemon/lib/shuttle/kitty.ex). A non-kitty user gets nothing on either OS; shuttle attach <fiber> always works.

Windows is unsupported.

Explicit Codex endpoint

SHUTTLE_CODEX_SOCKET selects the same Unix WebSocket endpoint for app dispatch and native session messaging. Set it in the Shuttle daemon's environment when using an explicitly shared Desktop backend. Without it, both clients use $CODEX_HOME/app-server-control/app-server-control.sock, with CODEX_HOME defaulting to ~/.codex. An empty override uses that default. An unavailable explicit endpoint is an error; clients do not switch to another conversation owner or start a replacement backend.