anvilsign in

collin/anvil · 3f59c98c

Design doc for dial-out remote CI runners

Collin Richards · 2026-08-24 07:40 UTC · 3f59c98cd084d60c2114a1988e1d36d63922c26c · parent a68a1ed4 · browse files

modifiedREADME.md+2 −0
⋯ 36 unchanged lines
3737 ## Docs
3838
3939 - [CI artifacts](docs/ci-artifacts.md)
40+- [Remote runners](docs/remote-runners.md) — dial-out runners so CI executes
41+ off-box; design, not yet built
4042 - [Single sign-on](docs/oidc.md) — OIDC sign-in, alongside passwords
4143 - [Repository secrets](docs/secrets.md) — encrypted to your ssh keys in the
4244 browser; anvil stores ciphertext it cannot open
⋯ 1 unchanged line
addeddocs/remote-runners.md+297 −0
1+# Remote runners
2+
3+Status: **design** (2026-08-24). Supersedes the in-process CI executor in
4+`crates/anvil-ci`.
5+
6+Today anvil *is* the runner: `run_worker` (`anvil-ci/src/lib.rs:73`) drains the
7+queue in-process and `execute` (`:387`) creates the job container directly on
8+the local Docker socket. The crate header states the model outright — "anvil is
9+the broker: it is the only Docker client."
10+
11+That works, and it pins CI to whichever machine anvil runs on. anvil runs on
12+hagrid, a droplet small enough that a release build OOMs it (see
13+[DEPLOY.md](../DEPLOY.md) §3). So CI has to run somewhere else.
14+
15+A **runner** is a separate binary that dials out to anvil, claims a job, runs it
16+in a sandboxed container on its own Docker daemon, and posts the result back.
17+anvild stops executing anything.
18+
19+```
20+ hagrid (droplet, 1 GB) build host (Mac mini M2)
21+ +------------------------+ +---------------------------+
22+ | anvild | | anvil-worker (launchd) |
23+ | queue, tree, vault | <--------- | long-poll claim |
24+ | artifacts, webhook | https | POST result |
25+ | | ---------> | | |
26+ | NO docker socket | | v |
27+ +------------------------+ | [ job container ] |
28+ | cap-drop, no socket |
29+ +---------------------------+
30+```
31+
32+## Why dial-out
33+
34+The alternative was forwarding the build host's Docker socket to hagrid and
35+pointing anvil at it. Dial-out wins on four counts, in order of importance:
36+
37+1. **No inbound path from the public VPS into the home LAN.** A forwarded
38+ Docker socket is unauthenticated root on the machine that owns it. Punching
39+ that from an internet-facing droplet into a home network makes hagrid's
40+ compromise the build host's compromise.
41+2. **NAT and sleep/wake are free.** The build host is a desktop machine behind
42+ a residential connection. A dialing client reconnects; a dialed-into server
43+ needs a tunnel supervised on both ends.
44+3. **The failure mode is legible.** A sleeping Mac becomes "no runner
45+ available," not a Docker API error surfacing as `[runner error]` in a job
46+ log.
47+4. **Runners become plural.** Once anvil addresses runners rather than a
48+ socket, a second one is configuration, not architecture — which is how the
49+ arm64/amd64 split below gets solved properly.
50+
51+## What this buys hagrid
52+
53+anvild stops needing Docker at all. `deploy/run.sh` drops
54+`-v /var/run/docker.sock:/var/run/docker.sock` and `--group-add`, which deletes
55+the warning in [DEPLOY.md](../DEPLOY.md) §4 about the container holding
56+root-equivalent control of the host. The internet-facing process stops being a
57+host-escape vector.
58+
59+This is contingent on agent sessions staying off — see [Scope](#scope).
60+
61+## What moves
62+
63+| Stays in `anvild` | Moves to `anvil-worker` |
64+| ------------------------------------------ | ---------------------------- |
65+| queue, `enqueue`/`requeue_interrupted` | `execute` |
66+| repo + tree resolution, `build_tar` | `collect_artifacts` |
67+| `parse_pipeline`, image allowlist check | `store_artifact` |
68+| the secret vault (`app.vault.take`) | `download_tar` |
69+| step-script assembly, `single_quote` | `parse_meta_tar` |
70+| artifact storage, the swap, `gc_artifacts` | `docker.rs` |
71+| `mask_secrets` (applied on receipt) | |
72+| the deploy webhook | |
73+
74+`run_worker` becomes `run_dispatcher`: same queue drain, but instead of calling
75+`execute` it parks the job until a runner claims it.
76+
77+Script assembly stays server-side deliberately. The runner then never parses
78+`.anvil/ci.yml` and holds no pipeline model — it receives an image, a script,
79+sandbox caps, and a tar. That keeps the wire format stable as the pipeline
80+schema grows.
81+
82+### The crate name
83+
84+The binary is **`anvil-worker`**, not `anvil-runner`, because
85+`anvil-runner:latest` is already the *image* CI jobs and agent sessions run in
86+(see [agent-sessions.md](agent-sessions.md)). Prose says "runner" for the
87+concept; the crate avoids the collision.
88+
89+## The protocol
90+
91+Plain HTTP against the existing axum server, under the `/-/` system namespace,
92+so it inherits Caddy's TLS and needs no new listener.
93+
94+| Endpoint | Method | Purpose |
95+| -------------------------------------------- | ------ | -------------------------------- |
96+| `/-/runner/claim` | POST | long-poll; 204 on timeout |
97+| `/-/runner/jobs/{run_id}/checkout.tar` | GET | the uploaded checkout |
98+| `/-/runner/jobs/{run_id}/heartbeat` | POST | extend the lease |
99+| `/-/runner/jobs/{run_id}/artifacts/{name}` | POST | one artifact tar |
100+| `/-/runner/jobs/{run_id}/result` | POST | exit code + log + artifact meta |
101+
102+The claim response carries everything `execute` takes as arguments today:
103+
104+```json
105+{
106+ "run_id": 42,
107+ "image": "rust:1.95-bookworm",
108+ "platform": "linux/amd64",
109+ "script": "set -e\n...",
110+ "secrets": [{"name": "CARGO_TOKEN", "value": "..."}],
111+ "artifacts": [{"name": "docs", "path": "target/doc", "browse": true}],
112+ "sandbox": {
113+ "memory_mb": 2048, "cpus": 2.0, "pids_limit": 512,
114+ "timeout_secs": 1800, "network": true, "run_as": ""
115+ }
116+}
117+```
118+
119+The checkout is a separate `GET` rather than a base64 field, so a large tree
120+doesn't inflate a JSON body by a third. Secrets ride in the claim body over
121+TLS, never on a separately-fetchable URL.
122+
123+Artifacts upload individually before `result`, for the same reason.
124+
125+### Logs
126+
127+The runner posts the whole log with `result`, which is exactly current
128+behaviour: `append_log` is called only twice in `process` — once on the
129+secrets-failure exit (`:161`) and once when the run ends (`:233`). The log
130+accumulates in memory and lands in one write. The run page is not live today
131+and does not become less live.
132+
133+Live logs are a genuine follow-up, and a remote runner makes them *easier* to
134+justify (there is now a producer that could stream). Out of scope here.
135+
136+### Auth
137+
138+A shared secret in `[ci] runner_token`, sent as `X-Anvil-Runner-Token`,
139+constant-time compared.
140+
141+This matches the existing `deploy_secret` pattern rather than inventing a
142+credential type. It is deliberate: API tokens are read-only and Bearer-only on
143+GET/HEAD (see [untrusted-mode.md](untrusted-mode.md)), a runner must POST, and
144+the write scope is still on [TODO.md](../TODO.md). Per-runner DB-backed tokens
145+with `last_used_at` are the right end state; a single-tenant forge with one
146+runner does not need them to start.
147+
148+### Leases
149+
150+Held **in memory** on the server: `run_id -> (runner_name, expires_at)`, bumped
151+by `heartbeat`, swept periodically. An expired lease returns the run to
152+`queued`.
153+
154+In-memory rather than columns on `CiRun` because Toasty migrations do not exist
155+yet — DEPLOY.md §Operations and TODO.md both flag that schema changes don't
156+auto-apply to the live database. Adding `claimed_by`/`lease_expires_at` to
157+`CiRun` (`models.rs:65`) would need a manual migration on hagrid.
158+
159+It also costs nothing: anvild is the only dispatcher, so a lease has no reason
160+to outlive it, and the anvild-crash case is already handled —
161+`requeue_interrupted` (`ci.rs:245`) re-queues everything left `running` at
162+startup. The sweep covers the new case, a runner that dies mid-job.
163+
164+## Architecture
165+
166+The build host is arm64; hagrid is x86_64. Two separate concerns:
167+
168+**What the shipped image is.** Already solved: `deploy/build.sh:33` cross-
169+compiles with zigbuild and builds `--platform linux/amd64`. Nothing to do,
170+though note the `Dockerfile`'s `RUN apt-get …` does execute amd64 binaries
171+under emulation. Since `anvild` is a static musl binary,
172+`gcr.io/distroless/static:nonroot` would make that build pure `COPY` and
173+genuinely emulation-free. Nice-to-have.
174+
175+**What jobs run on.** On an M2, `rust:1.95-bookworm` resolves to arm64, so
176+`cargo test` tests an architecture you don't ship. Fix: a `platform:` key in
177+`.anvil/ci.yml`, plumbed to `CreateContainerOptions.platform`
178+(bollard 0.18 `container.rs:107`) and `CreateImageOptions.platform`
179+(`image.rs:72`) — `ensure_image` currently leaves both at `Default::default()`.
180+Enable Docker Desktop's Rosetta option; it is far faster than QEMU for amd64
181+Linux binaries.
182+
183+Per-pipeline you then choose: native arm64 for lint and unit tests, amd64 under
184+Rosetta for anything arch-sensitive.
185+
186+Once runners are plural, `platform` also becomes a scheduling hint — a job
187+declaring `linux/amd64` prefers a runner that is natively amd64. Not needed
188+with one runner; the field is forward-compatible with it.
189+
190+## Isolation on macOS
191+
192+Docker on macOS is a Linux VM (LinuxKit under Docker Desktop, Lima under
193+Colima/OrbStack) and every container runs inside it, so the sandbox is enforced
194+by the same kernel primitives as on Linux: `--cap-drop=ALL` and
195+`no-new-privileges` are capabilities and prctl, the pids/memory/cpu caps are
196+cgroups v2.
197+
198+It is a boundary *better* than hagrid's. `untrusted-mode.md` §1 notes that a
199+kernel or runc escape defeats the sandbox; on the Mac that escape reaches a
200+disposable Linux VM, not the host.
201+
202+Two things to get right:
203+
204+- **`Docker::connect_with_socket_defaults()` (`docker.rs:15`) will not find the
205+ socket.** It hardcodes `/var/run/docker.sock`; Docker Desktop only creates
206+ that symlink when "Allow the default Docker socket to be used" is ticked, and
207+ the real path is `~/.docker/run/docker.sock` (Colima differs again). The
208+ runner uses `Docker::connect_with_defaults()`, which honours `DOCKER_HOST`.
209+- **Run the runner natively under launchd, not in a container.** Containerizing
210+ it means mounting the socket into it, rebuilding the root-equivalent hole
211+ this change removes from hagrid.
212+
213+Size the VM's RAM deliberately: per-job `memory_mb` is carved out of a fixed
214+allocation. The dispatcher runs one job at a time, so this is slack rather than
215+a constraint.
216+
217+## Two processes on the build host
218+
219+Image building **cannot be a CI job**. Job containers get no Docker socket by
220+design (`execute`'s doc comment at `:369`, and `untrusted-mode.md` §1), and
221+that invariant is the whole broker model. So the build host runs two things at
222+two trust levels:
223+
224+1. **`anvil-worker`** — claims jobs, runs them sandboxed, no Docker access
225+ *inside* the job container.
226+2. **a deploy agent** — fired on green CI, runs *outside* any sandbox with full
227+ Docker access, does `docker build --platform linux/amd64` and `docker push`
228+ to `registry.vibe.richardscollin.com`, then triggers hagrid to pull.
229+
230+Keeping these separate is what preserves the broker model. Folding the second
231+into the first would give pipeline authors a path to the daemon.
232+
233+The deploy agent is out of scope for M1 — the existing `deploy_webhook` still
234+works, with the receiver moved to the build host and reached over Tailscale.
235+M2 folds it into the same dial-out channel as a privileged "publish" job kind,
236+authorized server-side by the `is_deploy_target` check that already scopes CD
237+to exactly one repository (`config.rs:357`). That removes the last inbound
238+requirement.
239+
240+## Scope
241+
242+**Agent sessions stay local-socket and stay out.** `anvil-agent` calls
243+`anvil_ci::docker::connect()` in five places (`lib.rs:141`,
244+`supervisor.rs:55,278,414,439`) and `docker.rs` is explicitly shared plumbing,
245+so the crate split touches them — but a session is interactive (tmux attach,
246+exec streaming, resize, `pump_transcript`), which is a far harder protocol than
247+fire-and-forget CI. They are `enabled = false` and absent from
248+`deploy/anvil.toml` entirely, so nothing in production regresses.
249+
250+`ensure_image` and `connect` therefore need a home both crates can reach. They
251+move to a thin `anvil-docker` crate rather than being duplicated —
252+`ensure_image`'s local-fallback pull logic is subtle enough to be worth having
253+once.
254+
255+The corollary: turning agent sessions on for hagrid later means either putting
256+the socket back, or moving sessions onto the runner protocol too.
257+
258+## Registry cleanup
259+
260+`ensure_image`'s tolerance of a failed pull exists because `anvil-runner:latest`
261+"exists in no registry" (`docker.rs:22-24`) and is built straight into the local
262+daemon store. With `registry.vibe.richardscollin.com` up, push the image there
263+and the fallback stops being load-bearing.
264+
265+It should stay in the code regardless — a runner on a fresh machine wants a
266+clear failure when the pull fails and nothing is cached — but the comment
267+explaining *why* it exists needs rewriting.
268+
269+## Milestones
270+
271+**M1 — the split.** `anvil-docker` and `anvil-worker` crates, the five
272+endpoints, in-memory leases, `[ci] runner_token`, `connect_with_defaults`,
273+`run_worker` → `run_dispatcher`, socket mount dropped from `deploy/run.sh`.
274+Deploys keep using the existing webhook.
275+
276+**M2 — platform.** `platform:` in `.anvil/ci.yml`, plumbed through both bollard
277+options. Runner advertises its native platform at claim time.
278+
279+**M3 — publish jobs.** The deploy agent folds into the dial-out channel,
280+authorized by `is_deploy_target`. No inbound path to the build host remains.
281+
282+## Not yet done
283+
284+- **Live logs.** Newly worth doing, still not done. The result POST is a single
285+ write; streaming needs chunked append with offsets and a UI that tolerates
286+ gaps.
287+- **Per-runner credentials.** One shared secret means one revocation for all
288+ runners, and no `last_used_at`. Wants the API-token write scope first.
289+- **Runner labels.** `platform` is the only scheduling dimension in M2. Tags
290+ ("has-postgres", "big-memory") are the obvious next axis and are not designed.
291+- **Secrets on the runner host.** They now cross the network and sit in
292+ plaintext in a container on a machine anvil does not own. This needs a
293+ paragraph in `untrusted-mode.md` §1 — the exposure is no longer bounded by
294+ hagrid.
295+- **Concurrency.** The dispatcher still hands out one job at a time, inherited
296+ from `run_worker`. Multiple runners make that the binding constraint rather
297+ than a sensible default.