anvilsign in

collin/anvil

RenderedSource

1# Remote runners
2
3Status: **design** (2026-08-24). Supersedes the in-process CI executor in
4`crates/anvil-ci`.
5
6Today anvil *is* the runner: `run_worker` (`anvil-ci/src/lib.rs:73`) drains the
7queue in-process and `execute` (`:387`) creates the job container directly on
8the local Docker socket. The crate header states the model outright — "anvil is
9the broker: it is the only Docker client."
10
11That works, and it pins CI to whichever machine anvil runs on. anvil runs on
12hagrid, a droplet small enough that a release build OOMs it (see
13[DEPLOY.md](../DEPLOY.md) §3). So CI has to run somewhere else.
14
15A **runner** is a separate binary that dials out to anvil, claims a job, runs it
16in a sandboxed container on its own Docker daemon, and posts the result back.
17anvild stops executing anything.
18
19```
20 hagrid (droplet, 1 GB) build host (Mac mini M2)
21 +------------------------+ +---------------------------+
22 | anvild | | anvil-worker (launchd) |
23 | queue, tree, vault | <--------- | long-poll claim |
24 | artifacts, webhook | https | POST result |
25 | | ---------> | | |
26 | NO docker socket | | v |
27 +------------------------+ | [ job container ] |
28 | cap-drop, no socket |
29 +---------------------------+
30```
31
32## Why dial-out
33
34The alternative was forwarding the build host's Docker socket to hagrid and
35pointing anvil at it. Dial-out wins on four counts, in order of importance:
36
371. **No inbound path from the public VPS into the home LAN.** A forwarded
38 Docker socket is unauthenticated root on the machine that owns it. Punching
39 that from an internet-facing droplet into a home network makes hagrid's
40 compromise the build host's compromise.
412. **NAT and sleep/wake are free.** The build host is a desktop machine behind
42 a residential connection. A dialing client reconnects; a dialed-into server
43 needs a tunnel supervised on both ends.
443. **The failure mode is legible.** A sleeping Mac becomes "no runner
45 available," not a Docker API error surfacing as `[runner error]` in a job
46 log.
474. **Runners become plural.** Once anvil addresses runners rather than a
48 socket, a second one is configuration, not architecture — which is how the
49 arm64/amd64 split below gets solved properly.
50
51## What this buys hagrid
52
53anvild stops needing Docker at all. `deploy/run.sh` drops
54`-v /var/run/docker.sock:/var/run/docker.sock` and `--group-add`, which deletes
55the warning in [DEPLOY.md](../DEPLOY.md) §4 about the container holding
56root-equivalent control of the host. The internet-facing process stops being a
57host-escape vector.
58
59This is contingent on agent sessions staying off — see [Scope](#scope).
60
61## What moves
62
63| Stays in `anvild` | Moves to `anvil-worker` |
64| ------------------------------------------ | ---------------------------- |
65| queue, `enqueue`/`requeue_interrupted` | `execute` |
66| repo + tree resolution, `build_tar` | `collect_artifacts` |
67| `parse_pipeline`, image allowlist check | `store_artifact` |
68| the secret vault (`app.vault.take`) | `download_tar` |
69| step-script assembly, `single_quote` | `parse_meta_tar` |
70| artifact storage, the swap, `gc_artifacts` | `docker.rs` |
71| `mask_secrets` (applied on receipt) | |
72| the deploy webhook | |
73
74`run_worker` becomes `run_dispatcher`: same queue drain, but instead of calling
75`execute` it parks the job until a runner claims it.
76
77Script assembly stays server-side deliberately. The runner then never parses
78`.anvil/ci.yml` and holds no pipeline model — it receives an image, a script,
79sandbox caps, and a tar. That keeps the wire format stable as the pipeline
80schema grows.
81
82### The crate name
83
84The binary is **`anvil-worker`**, not `anvil-runner`, because
85`anvil-runner:latest` is already the *image* CI jobs and agent sessions run in
86(see [agent-sessions.md](agent-sessions.md)). Prose says "runner" for the
87concept; the crate avoids the collision.
88
89## The protocol
90
91Plain HTTP against the existing axum server, under the `/-/` system namespace,
92so it inherits Caddy's TLS and needs no new listener.
93
94| Endpoint | Method | Purpose |
95| -------------------------------------------- | ------ | -------------------------------- |
96| `/-/runner/claim` | POST | long-poll; 204 on timeout |
97| `/-/runner/jobs/{run_id}/checkout.tar` | GET | the uploaded checkout |
98| `/-/runner/jobs/{run_id}/heartbeat` | POST | extend the lease |
99| `/-/runner/jobs/{run_id}/artifacts/{name}` | POST | one artifact tar |
100| `/-/runner/jobs/{run_id}/result` | POST | exit code + log + artifact meta |
101
102The claim response carries everything `execute` takes as arguments today:
103
104```json
105{
106 "run_id": 42,
107 "image": "rust:1.95-bookworm",
108 "platform": "linux/amd64",
109 "script": "set -e\n...",
110 "secrets": [{"name": "CARGO_TOKEN", "value": "..."}],
111 "artifacts": [{"name": "docs", "path": "target/doc", "browse": true}],
112 "sandbox": {
113 "memory_mb": 2048, "cpus": 2.0, "pids_limit": 512,
114 "timeout_secs": 1800, "network": true, "run_as": ""
115 }
116}
117```
118
119The checkout is a separate `GET` rather than a base64 field, so a large tree
120doesn't inflate a JSON body by a third. Secrets ride in the claim body over
121TLS, never on a separately-fetchable URL.
122
123Artifacts upload individually before `result`, for the same reason.
124
125### Logs
126
127The runner posts the whole log with `result`, which is exactly current
128behaviour: `append_log` is called only twice in `process` — once on the
129secrets-failure exit (`:161`) and once when the run ends (`:233`). The log
130accumulates in memory and lands in one write. The run page is not live today
131and does not become less live.
132
133Live logs are a genuine follow-up, and a remote runner makes them *easier* to
134justify (there is now a producer that could stream). Out of scope here.
135
136### Auth
137
138A shared secret in `[ci] runner_token`, sent as `X-Anvil-Runner-Token`,
139constant-time compared.
140
141This matches the existing `deploy_secret` pattern rather than inventing a
142credential type. It is deliberate: API tokens are read-only and Bearer-only on
143GET/HEAD (see [untrusted-mode.md](untrusted-mode.md)), a runner must POST, and
144the write scope is still on [TODO.md](../TODO.md). Per-runner DB-backed tokens
145with `last_used_at` are the right end state; a single-tenant forge with one
146runner does not need them to start.
147
148### Leases
149
150Held **in memory** on the server: `run_id -> (runner_name, expires_at)`, bumped
151by `heartbeat`, swept periodically. An expired lease returns the run to
152`queued`.
153
154In-memory rather than columns on `CiRun` because Toasty migrations do not exist
155yet — DEPLOY.md §Operations and TODO.md both flag that schema changes don't
156auto-apply to the live database. Adding `claimed_by`/`lease_expires_at` to
157`CiRun` (`models.rs:65`) would need a manual migration on hagrid.
158
159It also costs nothing: anvild is the only dispatcher, so a lease has no reason
160to outlive it, and the anvild-crash case is already handled —
161`requeue_interrupted` (`ci.rs:245`) re-queues everything left `running` at
162startup. The sweep covers the new case, a runner that dies mid-job.
163
164## Architecture
165
166The build host is arm64; hagrid is x86_64. Two separate concerns:
167
168**What the shipped image is.** Already solved: `deploy/build.sh:33` cross-
169compiles with zigbuild and builds `--platform linux/amd64`. Nothing to do,
170though note the `Dockerfile`'s `RUN apt-get …` does execute amd64 binaries
171under emulation. Since `anvild` is a static musl binary,
172`gcr.io/distroless/static:nonroot` would make that build pure `COPY` and
173genuinely emulation-free. Nice-to-have.
174
175**What jobs run on.** On an M2, `rust:1.95-bookworm` resolves to arm64, so
176`cargo test` tests an architecture you don't ship. Fix: a `platform:` key in
177`.anvil/ci.yml`, plumbed to `CreateContainerOptions.platform`
178(bollard 0.18 `container.rs:107`) and `CreateImageOptions.platform`
179(`image.rs:72`) — `ensure_image` currently leaves both at `Default::default()`.
180Enable Docker Desktop's Rosetta option; it is far faster than QEMU for amd64
181Linux binaries.
182
183Per-pipeline you then choose: native arm64 for lint and unit tests, amd64 under
184Rosetta for anything arch-sensitive.
185
186Once runners are plural, `platform` also becomes a scheduling hint — a job
187declaring `linux/amd64` prefers a runner that is natively amd64. Not needed
188with one runner; the field is forward-compatible with it.
189
190## Isolation on macOS
191
192Docker on macOS is a Linux VM (LinuxKit under Docker Desktop, Lima under
193Colima/OrbStack) and every container runs inside it, so the sandbox is enforced
194by the same kernel primitives as on Linux: `--cap-drop=ALL` and
195`no-new-privileges` are capabilities and prctl, the pids/memory/cpu caps are
196cgroups v2.
197
198It is a boundary *better* than hagrid's. `untrusted-mode.md` §1 notes that a
199kernel or runc escape defeats the sandbox; on the Mac that escape reaches a
200disposable Linux VM, not the host.
201
202Two things to get right:
203
204- **`Docker::connect_with_socket_defaults()` (`docker.rs:15`) will not find the
205 socket.** It hardcodes `/var/run/docker.sock`; Docker Desktop only creates
206 that symlink when "Allow the default Docker socket to be used" is ticked, and
207 the real path is `~/.docker/run/docker.sock` (Colima differs again). The
208 runner uses `Docker::connect_with_defaults()`, which honours `DOCKER_HOST`.
209- **Run the runner natively under launchd, not in a container.** Containerizing
210 it means mounting the socket into it, rebuilding the root-equivalent hole
211 this change removes from hagrid.
212
213Size the VM's RAM deliberately: per-job `memory_mb` is carved out of a fixed
214allocation. The dispatcher runs one job at a time, so this is slack rather than
215a constraint.
216
217## Two processes on the build host
218
219Image building **cannot be a CI job**. Job containers get no Docker socket by
220design (`execute`'s doc comment at `:369`, and `untrusted-mode.md` §1), and
221that invariant is the whole broker model. So the build host runs two things at
222two trust levels:
223
2241. **`anvil-worker`** — claims jobs, runs them sandboxed, no Docker access
225 *inside* the job container.
2262. **a deploy agent** — fired on green CI, runs *outside* any sandbox with full
227 Docker access, does `docker build --platform linux/amd64` and `docker push`
228 to `registry.vibe.richardscollin.com`, then triggers hagrid to pull.
229
230Keeping these separate is what preserves the broker model. Folding the second
231into the first would give pipeline authors a path to the daemon.
232
233The deploy agent is out of scope for M1 — the existing `deploy_webhook` still
234works, with the receiver moved to the build host and reached over Tailscale.
235M2 folds it into the same dial-out channel as a privileged "publish" job kind,
236authorized server-side by the `is_deploy_target` check that already scopes CD
237to exactly one repository (`config.rs:357`). That removes the last inbound
238requirement.
239
240## Scope
241
242**Agent sessions stay local-socket and stay out.** `anvil-agent` calls
243`anvil_ci::docker::connect()` in five places (`lib.rs:141`,
244`supervisor.rs:55,278,414,439`) and `docker.rs` is explicitly shared plumbing,
245so the crate split touches them — but a session is interactive (tmux attach,
246exec streaming, resize, `pump_transcript`), which is a far harder protocol than
247fire-and-forget CI. They are `enabled = false` and absent from
248`deploy/anvil.toml` entirely, so nothing in production regresses.
249
250`ensure_image` and `connect` therefore need a home both crates can reach. They
251move to a thin `anvil-docker` crate rather than being duplicated —
252`ensure_image`'s local-fallback pull logic is subtle enough to be worth having
253once.
254
255The corollary: turning agent sessions on for hagrid later means either putting
256the socket back, or moving sessions onto the runner protocol too.
257
258## Registry cleanup
259
260`ensure_image`'s tolerance of a failed pull exists because `anvil-runner:latest`
261"exists in no registry" (`docker.rs:22-24`) and is built straight into the local
262daemon store. With `registry.vibe.richardscollin.com` up, push the image there
263and the fallback stops being load-bearing.
264
265It should stay in the code regardless — a runner on a fresh machine wants a
266clear failure when the pull fails and nothing is cached — but the comment
267explaining *why* it exists needs rewriting.
268
269## Milestones
270
271**M1 — the split.** `anvil-docker` and `anvil-worker` crates, the five
272endpoints, in-memory leases, `[ci] runner_token`, `connect_with_defaults`,
273`run_worker` → `run_dispatcher`, socket mount dropped from `deploy/run.sh`.
274Deploys keep using the existing webhook.
275
276**M2 — platform.** `platform:` in `.anvil/ci.yml`, plumbed through both bollard
277options. Runner advertises its native platform at claim time.
278
279**M3 — publish jobs.** The deploy agent folds into the dial-out channel,
280authorized by `is_deploy_target`. No inbound path to the build host remains.
281
282## Not yet done
283
284- **Live logs.** Newly worth doing, still not done. The result POST is a single
285 write; streaming needs chunked append with offsets and a UI that tolerates
286 gaps.
287- **Per-runner credentials.** One shared secret means one revocation for all
288 runners, and no `last_used_at`. Wants the API-token write scope first.
289- **Runner labels.** `platform` is the only scheduling dimension in M2. Tags
290 ("has-postgres", "big-memory") are the obvious next axis and are not designed.
291- **Secrets on the runner host.** They now cross the network and sit in
292 plaintext in a container on a machine anvil does not own. This needs a
293 paragraph in `untrusted-mode.md` §1 — the exposure is no longer bounded by
294 hagrid.
295- **Concurrency.** The dispatcher still hands out one job at a time, inherited
296 from `run_worker`. Multiple runners make that the binding constraint rather
297 than a sensible default.