anvilsign in

collin/anvil

RenderedSource

1# Remote runners
2
3Status: **M1 implemented** (2026-08-24). Supersedes the in-process CI
4executor that used to live in `crates/anvil-ci`.
5
6anvil used to *be* the runner: `run_worker` drained the queue in-process and
7`execute` created the job container on the local Docker socket. That works, and
8it pins CI to whichever machine anvil runs on — which is hagrid, a droplet
9small enough that a release build OOMs it (see [DEPLOY.md](../DEPLOY.md) §3).
10So CI had to run somewhere else.
11
12A **runner** is a separate binary that dials out to anvil, claims a job, runs it
13in a sandboxed container on its own Docker daemon, and posts the result back.
14anvild stops executing anything.
15
16```
17 hagrid (droplet, 1 GB) build host (Mac mini M2)
18 +------------------------+ +---------------------------+
19 | anvild | | anvil-worker (launchd) |
20 | queue, tree, vault | <--------- | long-poll claim |
21 | artifacts, webhook | https | POST result |
22 | | ---------> | | |
23 | NO docker socket | | v |
24 +------------------------+ | [ job container ] |
25 | cap-drop, no socket |
26 +---------------------------+
27```
28
29## Why dial-out
30
31The alternative was forwarding the build host's Docker socket to hagrid and
32pointing anvil at it. Dial-out wins on four counts, in order of importance:
33
341. **No inbound path from the public VPS into the home LAN.** A forwarded
35 Docker socket is unauthenticated root on the machine that owns it. Punching
36 that from an internet-facing droplet into a home network makes hagrid's
37 compromise the build host's compromise.
382. **NAT and sleep/wake are free.** The build host is a desktop machine behind
39 a residential connection. A dialing client reconnects; a dialed-into server
40 needs a tunnel supervised on both ends.
413. **The failure mode is legible.** A sleeping Mac becomes "no runner
42 available," not a Docker API error surfacing as `[runner error]` in a job
43 log.
444. **Runners become plural.** Once anvil addresses runners rather than a
45 socket, a second one is configuration, not architecture — which is how the
46 arm64/amd64 split below gets solved properly.
47
48## What this buys hagrid
49
50anvild stops needing Docker at all. `deploy/run.sh` drops
51`-v /var/run/docker.sock:/var/run/docker.sock` and `--group-add`, which deletes
52the warning in [DEPLOY.md](../DEPLOY.md) §4 about the container holding
53root-equivalent control of the host. The internet-facing process stops being a
54host-escape vector.
55
56This is contingent on agent sessions staying off — see [Scope](#scope).
57
58## What moves
59
60| Stays in `anvild` | Moved to `anvil-worker` |
61| ------------------------------------------ | ---------------------------- |
62| queue, `enqueue`/`requeue_interrupted` | `execute` |
63| repo + tree resolution, `build_tar` | `collect_artifacts` |
64| `parse_pipeline`, image allowlist check | `download_tar` |
65| the secret vault (`app.vault.take`) | `parse_meta_tar` |
66| step-script assembly, `single_quote` | the `ArtifactSink` trait |
67| `store_artifact`, the swap, `gc_artifacts` | |
68| `mask_secrets` (applied on receipt) | |
69| the deploy webhook | |
70
71`store_artifact` deliberately stayed behind. The runner uploads the raw tar it
72pulled out of the container and anvil decides what to do with it, so on-disk
73layout — and `browse`, which turns a tarball into a servable directory tree —
74never becomes a runner's call. The upload response reports the stored size, so
75the per-run artifact budget is still charged what actually landed.
76
77`run_worker` becomes `run_dispatcher`: same queue drain, but instead of calling
78`execute` it parks the job until a runner claims it.
79
80Script assembly stays server-side deliberately. The runner then never parses
81`.anvil/ci.yml` and holds no pipeline model — it receives an image, a script,
82sandbox caps, and a tar. That keeps the wire format stable as the pipeline
83schema grows.
84
85### The crates
86
87| Crate | Holds |
88| --------------- | ------------------------------------------------------ |
89| `anvil-job` | the wire format, and nothing else — serde only |
90| `anvil-docker` | `connect`/`ensure_image`, shared with `anvil-agent` |
91| `anvil-worker` | the runner binary: claim loop, client, executor |
92
93`anvil-job` exists so the runner does not link `anvil-core` — and therefore
94toasty, SQLite, gix and the rest of the forge — just to learn the shape of a
95job. `anvil-docker` exists so `anvil-agent` need not depend on the runner.
96
97The binary is **`anvil-worker`**, not `anvil-runner`, because
98`anvil-runner:latest` is already the *image* CI jobs and agent sessions run in
99(see [agent-sessions.md](agent-sessions.md)). Prose says "runner" for the
100concept; the crate avoids the collision.
101
102## The protocol
103
104Plain HTTP against the existing axum server, under the `/-/` system namespace,
105so it inherits Caddy's TLS and needs no new listener.
106
107| Endpoint | Method | Purpose |
108| -------------------------------------------- | ------ | -------------------------------- |
109| `/-/runner/claim` | POST | long-poll; 204 on timeout |
110| `/-/runner/jobs/{run_id}/checkout.tar` | GET | the uploaded checkout |
111| `/-/runner/jobs/{run_id}/heartbeat` | POST | extend the lease |
112| `/-/runner/jobs/{run_id}/artifacts/{name}` | POST | one artifact tar |
113| `/-/runner/jobs/{run_id}/result` | POST | exit code + log + artifact meta |
114
115The claim response carries everything `execute` takes as arguments today:
116
117```json
118{
119 "run_id": 42,
120 "image": "rust:1.95-bookworm",
121 "platform": "linux/amd64",
122 "script": "set -e\n...",
123 "secrets": [{"name": "CARGO_TOKEN", "value": "..."}],
124 "artifacts": [{"name": "docs", "path": "target/doc", "browse": true}],
125 "sandbox": {
126 "memory_mb": 2048, "cpus": 2.0, "pids_limit": 512,
127 "timeout_secs": 1800, "network": true, "run_as": ""
128 }
129}
130```
131
132The checkout is a separate `GET` rather than a base64 field, so a large tree
133doesn't inflate a JSON body by a third. Secrets ride in the claim body over
134TLS, never on a separately-fetchable URL.
135
136Artifacts upload individually before `result`, for the same reason.
137
138### Logs
139
140The runner posts the whole log with `result`, which is exactly what the
141in-process runner did: `append_log` was only ever called when a run ended (and
142on the secrets-failure exit). The log accumulated in memory and landed in one
143write. The run page was not live before and is no less live now — a header
144naming the runner is now written at claim time, so a `running` run at least
145shows something.
146
147Live logs are a genuine follow-up, and a remote runner makes them *easier* to
148justify (there is now a producer that could stream). Out of scope here.
149
150### Auth
151
152A shared secret in `[ci] runner_token`, sent as `X-Anvil-Runner-Token`,
153constant-time compared. The runner's self-asserted name rides alongside in
154`X-Anvil-Runner-Name` — it labels runs and keys leases, and is explicitly not
155a credential: everyone holding the token is one principal.
156
157The four per-job endpoints additionally require the caller to hold that run's
158lease, so a valid token gets you *a* job rather than everyone else's. A lease
159mismatch answers 409, not 403: the caller is a legitimate runner whose claim
160simply expired.
161
162This matches the existing `deploy_secret` pattern rather than inventing a
163credential type. It is deliberate: API tokens are read-only and Bearer-only on
164GET/HEAD (see [untrusted-mode.md](untrusted-mode.md)), a runner must POST, and
165the write scope is still on [TODO.md](../TODO.md). Per-runner DB-backed tokens
166with `last_used_at` are the right end state; a single-tenant forge with one
167runner does not need them to start.
168
169### Leases
170
171Held **in memory** on the server (`anvil_core::jobs::Dispatch`):
172`run_id -> (runner_name, expires_at, secrets)`, bumped by `heartbeat` every 30s
173against a 120s TTL, swept every 30s. An expired lease returns the run to
174`queued`.
175
176A job's secrets are stashed on its lease rather than re-read from the vault
177when the result lands. `Vault::take` fails once the repository's unlock TTL
178lapses, and a job can easily outlive an unlock — re-reading would mean a long
179run silently skips log masking, which is exactly the run whose log is most
180likely to contain something. The values are already in this process's vault,
181so this is not new exposure.
182
183The honest gap: an expired lease can double-run a job whose runner is alive but
184unreachable. The container keeps going and the requeued run may be claimed
185elsewhere. That is what a lease without fencing buys; CI steps are assumed
186idempotent.
187
188In-memory rather than columns on `CiRun` because Toasty migrations do not exist
189yet — DEPLOY.md §Operations and TODO.md both flag that schema changes don't
190auto-apply to the live database. Adding `claimed_by`/`lease_expires_at` to
191`CiRun` (`models.rs:65`) would need a manual migration on hagrid.
192
193It also costs nothing: anvild is the only dispatcher, so a lease has no reason
194to outlive it, and the anvild-crash case is already handled —
195`requeue_interrupted` (`ci.rs:245`) re-queues everything left `running` at
196startup. The sweep covers the new case, a runner that dies mid-job.
197
198## Architecture
199
200The build host is arm64; hagrid is x86_64. Two separate concerns:
201
202**What the shipped image is.** Already solved: `deploy/build.sh:33` cross-
203compiles with zigbuild and builds `--platform linux/amd64`. Nothing to do,
204though note the `Dockerfile`'s `RUN apt-get …` does execute amd64 binaries
205under emulation. Since `anvild` is a static musl binary,
206`gcr.io/distroless/static:nonroot` would make that build pure `COPY` and
207genuinely emulation-free. Nice-to-have.
208
209**What jobs run on.** On an M2, `rust:1.95-bookworm` resolves to arm64, so
210`cargo test` tests an architecture you don't ship. Fix: a `platform:` key in
211`.anvil/ci.yml`, plumbed to `CreateContainerOptions.platform`
212(bollard 0.18 `container.rs:107`) and `CreateImageOptions.platform`
213(`image.rs:72`) — `ensure_image` currently leaves both at `Default::default()`.
214Enable Docker Desktop's Rosetta option; it is far faster than QEMU for amd64
215Linux binaries.
216
217Per-pipeline you then choose: native arm64 for lint and unit tests, amd64 under
218Rosetta for anything arch-sensitive.
219
220Once runners are plural, `platform` also becomes a scheduling hint — a job
221declaring `linux/amd64` prefers a runner that is natively amd64. Not needed
222with one runner; the field is forward-compatible with it.
223
224## Isolation on macOS
225
226Docker on macOS is a Linux VM (LinuxKit under Docker Desktop, Lima under
227Colima/OrbStack) and every container runs inside it, so the sandbox is enforced
228by the same kernel primitives as on Linux: `--cap-drop=ALL` and
229`no-new-privileges` are capabilities and prctl, the pids/memory/cpu caps are
230cgroups v2.
231
232It is a boundary *better* than hagrid's. `untrusted-mode.md` §1 notes that a
233kernel or runc escape defeats the sandbox; on the Mac that escape reaches a
234disposable Linux VM, not the host.
235
236Two things to get right:
237
238- **`Docker::connect_with_socket_defaults()` (`docker.rs:15`) will not find the
239 socket.** It hardcodes `/var/run/docker.sock`; Docker Desktop only creates
240 that symlink when "Allow the default Docker socket to be used" is ticked, and
241 the real path is `~/.docker/run/docker.sock` (Colima differs again). The
242 runner uses `Docker::connect_with_defaults()`, which honours `DOCKER_HOST`.
243- **Run the runner natively under launchd, not in a container.** Containerizing
244 it means mounting the socket into it, rebuilding the root-equivalent hole
245 this change removes from hagrid.
246
247Size the VM's RAM deliberately: per-job `memory_mb` is carved out of a fixed
248allocation. The dispatcher runs one job at a time, so this is slack rather than
249a constraint.
250
251## Two processes on the build host
252
253Image building **cannot be a CI job**. Job containers get no Docker socket by
254design (`execute`'s doc comment at `:369`, and `untrusted-mode.md` §1), and
255that invariant is the whole broker model. So the build host runs two things at
256two trust levels:
257
2581. **`anvil-worker`** — claims jobs, runs them sandboxed, no Docker access
259 *inside* the job container.
2602. **a deploy agent** — fired on green CI, runs *outside* any sandbox with full
261 Docker access, does `docker build --platform linux/amd64` and `docker push`
262 to `registry.vibe.richardscollin.com`, then triggers hagrid to pull.
263
264Keeping these separate is what preserves the broker model. Folding the second
265into the first would give pipeline authors a path to the daemon.
266
267The deploy agent is out of scope for M1 — the existing `deploy_webhook` still
268works, with the receiver moved to the build host and reached over Tailscale.
269M2 folds it into the same dial-out channel as a privileged "publish" job kind,
270authorized server-side by the `is_deploy_target` check that already scopes CD
271to exactly one repository (`config.rs:357`). That removes the last inbound
272requirement.
273
274## Scope
275
276**Agent sessions stay local-socket and stay out.** `anvil-agent` calls
277`anvil_ci::docker::connect()` in five places (`lib.rs:141`,
278`supervisor.rs:55,278,414,439`) and `docker.rs` is explicitly shared plumbing,
279so the crate split touches them — but a session is interactive (tmux attach,
280exec streaming, resize, `pump_transcript`), which is a far harder protocol than
281fire-and-forget CI. They are `enabled = false` and absent from
282`deploy/anvil.toml` entirely, so nothing in production regresses.
283
284`ensure_image` and `connect` therefore need a home both crates can reach. They
285move to a thin `anvil-docker` crate rather than being duplicated —
286`ensure_image`'s local-fallback pull logic is subtle enough to be worth having
287once.
288
289The corollary: turning agent sessions on for hagrid later means either putting
290the socket back, or moving sessions onto the runner protocol too.
291
292## Registry cleanup
293
294`ensure_image`'s tolerance of a failed pull exists because `anvil-runner:latest`
295"exists in no registry" (`docker.rs:22-24`) and is built straight into the local
296daemon store. With `registry.vibe.richardscollin.com` up, push the image there
297and the fallback stops being load-bearing.
298
299It should stay in the code regardless — a runner on a fresh machine wants a
300clear failure when the pull fails and nothing is cached — but the comment
301explaining *why* it exists needs rewriting.
302
303## Milestones
304
305**M1 — the split. Done.** `anvil-job`, `anvil-docker` and `anvil-worker`
306crates, the five endpoints, in-memory leases, `[ci] runner_token`,
307`connect_with_defaults`, `run_worker` → `run_dispatcher`, socket mount dropped
308from `deploy/run.sh`. Deploys keep using the existing webhook.
309
310**M2 — platform.** The plumbing is already live: `JobSpec.platform` reaches
311`CreateContainerOptions` and `CreateImageOptions`, and a runner advertises its
312native platform when it claims. What is missing is a *source* — a `platform:`
313key in `.anvil/ci.yml` (and probably a `[ci] platform` default), plus routing a
314job to a runner that has that architecture natively.
315
316**M3 — publish jobs.** The deploy agent folds into the dial-out channel,
317authorized by `is_deploy_target`. No inbound path to the build host remains.
318
319## Not yet done
320
321- **Live logs.** Newly worth doing, still not done. The result POST is a single
322 write; streaming needs chunked append with offsets and a UI that tolerates
323 gaps.
324- **Per-runner credentials.** One shared secret means one revocation for all
325 runners, and no `last_used_at`. Wants the API-token write scope first.
326- **Runner labels.** `platform` is the only scheduling dimension in M2. Tags
327 ("has-postgres", "big-memory") are the obvious next axis and are not designed.
328- **Secrets on the runner host.** They now cross the network and sit in
329 plaintext in a container on a machine anvil does not own. This needs a
330 paragraph in `untrusted-mode.md` §1 — the exposure is no longer bounded by
331 hagrid.
332- **Concurrency.** Nothing bounds how many jobs are in flight beyond how many
333 runners exist, and nothing stops one runner claiming repeatedly. The
334 in-process runner's "one job at a time" was a property of the loop, and it is
335 gone; a `max_concurrent` equivalent for CI does not exist.
336- **`ensure_image`'s local fallback vs `platform`.** A cached image of the
337 wrong architecture satisfies the fallback, since it inspects presence and not
338 arch. Only bites an offline runner asked to cross-build.