anvilsign in

collin/anvil

RenderedSource

1# Remote runners
2
3Status: **M1 + M2 implemented** (2026-08-24). Supersedes the in-process CI
4executor that used to live in `crates/anvil-ci`.
5
6anvil used to *be* the runner: `run_worker` drained the queue in-process and
7`execute` created the job container on the local Docker socket. That works, and
8it pins CI to whichever machine anvil runs on — which is hagrid, a droplet
9small enough that a release build OOMs it (see [DEPLOY.md](../DEPLOY.md) §3).
10So CI had to run somewhere else.
11
12A **runner** is a separate binary that dials out to anvil, claims a job, runs it
13in a sandboxed container on its own Docker daemon, and posts the result back.
14anvild stops executing anything.
15
16```
17 hagrid (droplet, 1 GB) build host (Mac mini M2)
18 +------------------------+ +---------------------------+
19 | anvild | | anvil-worker (launchd) |
20 | queue, tree, vault | <--------- | long-poll claim |
21 | artifacts, webhook | https | POST result |
22 | | ---------> | | |
23 | NO docker socket | | v |
24 +------------------------+ | [ job container ] |
25 | cap-drop, no socket |
26 +---------------------------+
27```
28
29## Why dial-out
30
31The alternative was forwarding the build host's Docker socket to hagrid and
32pointing anvil at it. Dial-out wins on four counts, in order of importance:
33
341. **No inbound path from the public VPS into the home LAN.** A forwarded
35 Docker socket is unauthenticated root on the machine that owns it. Punching
36 that from an internet-facing droplet into a home network makes hagrid's
37 compromise the build host's compromise.
382. **NAT and sleep/wake are free.** The build host is a desktop machine behind
39 a residential connection. A dialing client reconnects; a dialed-into server
40 needs a tunnel supervised on both ends.
413. **The failure mode is legible.** A sleeping Mac becomes "no runner
42 available," not a Docker API error surfacing as `[runner error]` in a job
43 log.
444. **Runners become plural.** Once anvil addresses runners rather than a
45 socket, a second one is configuration, not architecture — which is how the
46 arm64/amd64 split below gets solved properly.
47
48## What this buys hagrid
49
50anvild stops needing Docker at all. `compose.yaml` carries no
51`/var/run/docker.sock` mount and no `group_add`, which deletes
52the warning in [DEPLOY.md](../DEPLOY.md) §4 about the container holding
53root-equivalent control of the host. The internet-facing process stops being a
54host-escape vector.
55
56This is contingent on agent sessions staying off — see [Scope](#scope).
57
58## What moves
59
60| Stays in `anvild` | Moved to `anvil-worker` |
61| ------------------------------------------ | ---------------------------- |
62| queue, `enqueue`/`requeue_interrupted` | `execute` |
63| repo + tree resolution, `build_tar` | `collect_artifacts` |
64| `parse_pipeline`, image allowlist check | `download_tar` |
65| the secret vault (`app.vault.take`) | `parse_meta_tar` |
66| step-script assembly, `single_quote` | the `ArtifactSink` trait |
67| `store_artifact`, the swap, `gc_artifacts` | |
68| `mask_secrets` (applied on receipt) | |
69| the deploy webhook | |
70
71`store_artifact` deliberately stayed behind. The runner uploads the raw tar it
72pulled out of the container and anvil decides what to do with it, so on-disk
73layout — and `browse`, which turns a tarball into a servable directory tree —
74never becomes a runner's call. The upload response reports the stored size, so
75the per-run artifact budget is still charged what actually landed.
76
77`run_worker` becomes `run_dispatcher`: same queue drain, but instead of calling
78`execute` it parks the job until a runner claims it.
79
80Script assembly stays server-side deliberately. The runner then never parses
81`.anvil/ci.yml` and holds no pipeline model — it receives an image, a script,
82sandbox caps, and a tar. That keeps the wire format stable as the pipeline
83schema grows.
84
85### The crates
86
87| Crate | Holds |
88| --------------- | ------------------------------------------------------ |
89| `anvil-job` | the wire format, and nothing else — serde only |
90| `anvil-docker` | `connect`/`ensure_image`, shared with `anvil-agent` |
91| `anvil-worker` | the runner binary: claim loop, client, executor |
92
93`anvil-job` exists so the runner does not link `anvil-core` — and therefore
94toasty, SQLite, gix and the rest of the forge — just to learn the shape of a
95job. `anvil-docker` exists so `anvil-agent` need not depend on the runner.
96
97The binary is **`anvil-worker`**, not `anvil-runner`, because
98`anvil-runner:latest` is already the *image* CI jobs and agent sessions run in
99(see [agent-sessions.md](agent-sessions.md)). Prose says "runner" for the
100concept; the crate avoids the collision.
101
102## The protocol
103
104Plain HTTP against the existing axum server, under the `/-/` system namespace,
105so it inherits Caddy's TLS and needs no new listener.
106
107| Endpoint | Method | Purpose |
108| -------------------------------------------- | ------ | -------------------------------- |
109| `/-/runner/claim` | POST | long-poll; 204 on timeout |
110| `/-/runner/jobs/{run_id}/checkout.tar` | GET | the uploaded checkout |
111| `/-/runner/jobs/{run_id}/heartbeat` | POST | extend the lease |
112| `/-/runner/jobs/{run_id}/artifacts/{name}` | POST | one artifact tar |
113| `/-/runner/jobs/{run_id}/result` | POST | exit code + log + artifact meta |
114
115The claim response carries everything `execute` takes as arguments today:
116
117```json
118{
119 "run_id": 42,
120 "image": "rust:1.95-bookworm",
121 "platform": "linux/amd64",
122 "script": "set -e\n...",
123 "secrets": [{"name": "CARGO_TOKEN", "value": "..."}],
124 "artifacts": [{"name": "docs", "path": "target/doc", "browse": true}],
125 "sandbox": {
126 "memory_mb": 2048, "cpus": 2.0, "pids_limit": 512,
127 "timeout_secs": 1800, "network": true, "run_as": ""
128 }
129}
130```
131
132The checkout is a separate `GET` rather than a base64 field, so a large tree
133doesn't inflate a JSON body by a third. Secrets ride in the claim body over
134TLS, never on a separately-fetchable URL.
135
136Artifacts upload individually before `result`, for the same reason.
137
138### Logs
139
140The runner posts the whole log with `result`, which is exactly what the
141in-process runner did: `append_log` was only ever called when a run ended (and
142on the secrets-failure exit). The log accumulated in memory and landed in one
143write. The run page was not live before and is no less live now — a header
144naming the runner is now written at claim time, so a `running` run at least
145shows something.
146
147Live logs are a genuine follow-up, and a remote runner makes them *easier* to
148justify (there is now a producer that could stream). Out of scope here.
149
150### Auth
151
152A shared secret in `[ci] runner_token`, sent as `X-Anvil-Runner-Token`,
153constant-time compared. The runner's self-asserted name rides alongside in
154`X-Anvil-Runner-Name` — it labels runs and keys leases, and is explicitly not
155a credential: everyone holding the token is one principal.
156
157The four per-job endpoints additionally require the caller to hold that run's
158lease, so a valid token gets you *a* job rather than everyone else's. A lease
159mismatch answers 409, not 403: the caller is a legitimate runner whose claim
160simply expired.
161
162This matches the existing `deploy_secret` pattern rather than inventing a
163credential type. It is deliberate: API tokens are read-only and Bearer-only on
164GET/HEAD (see [untrusted-mode.md](untrusted-mode.md)), a runner must POST, and
165the write scope is still on [TODO.md](../TODO.md). Per-runner DB-backed tokens
166with `last_used_at` are the right end state; a single-tenant forge with one
167runner does not need them to start.
168
169### Leases
170
171Held **in memory** on the server (`anvil_core::jobs::Dispatch`):
172`run_id -> (runner_name, expires_at, secrets)`, bumped by `heartbeat` every 30s
173against a 120s TTL, swept every 30s. An expired lease returns the run to
174`queued`.
175
176A job's secrets are stashed on its lease rather than re-read from the vault
177when the result lands. `Vault::take` fails once the repository's unlock TTL
178lapses, and a job can easily outlive an unlock — re-reading would mean a long
179run silently skips log masking, which is exactly the run whose log is most
180likely to contain something. The values are already in this process's vault,
181so this is not new exposure.
182
183The honest gap: an expired lease can double-run a job whose runner is alive but
184unreachable. The container keeps going and the requeued run may be claimed
185elsewhere. That is what a lease without fencing buys; CI steps are assumed
186idempotent.
187
188In-memory rather than columns on `CiRun` because Toasty migrations do not exist
189yet — DEPLOY.md §Operations and TODO.md both flag that schema changes don't
190auto-apply to the live database. Adding `claimed_by`/`lease_expires_at` to
191`CiRun` (`models.rs:65`) would need a manual migration on hagrid.
192
193It also costs nothing: anvild is the only dispatcher, so a lease has no reason
194to outlive it, and the anvild-crash case is already handled —
195`requeue_interrupted` (`ci.rs:245`) re-queues everything left `running` at
196startup. The sweep covers the new case, a runner that dies mid-job.
197
198## Architecture
199
200The build host is arm64; hagrid is x86_64. Two separate concerns:
201
202**What the shipped image is.** Already solved: `deploy/build.sh` cross-compiles
203with zigbuild and `compose.yaml` pins `platforms: [linux/amd64]`. Nothing to do,
204though note the `Dockerfile`'s `RUN apt-get …` does execute amd64 binaries
205under emulation. Since `anvild` is a static musl binary,
206`gcr.io/distroless/static:nonroot` would make that build pure `COPY` and
207genuinely emulation-free. Nice-to-have.
208
209**What jobs run on.** On an M2, `rust:1.95-bookworm` resolves to arm64, so
210`cargo test` tests an architecture you don't ship. That is what M2 fixes.
211
212A job's platform comes from `platform:` in `.anvil/ci.yml`, falling back to
213`[ci] platform`, falling back to the claiming runner's native architecture —
214so an instance that sets neither behaves exactly as it did before. It reaches
215`CreateContainerOptions.platform` and `CreateImageOptions.platform`, and the
216value is validated as `os/arch[/variant]` at parse time (a bare `amd64` would
217otherwise reach Docker as an *operating system* named amd64).
218
219Per-pipeline you then choose: native arm64 for lint and unit tests, amd64 under
220Rosetta for anything arch-sensitive. Turn Rosetta on in Docker Desktop; it is
221far faster than QEMU for amd64 Linux binaries.
222
223### Routing
224
225`platform` is also the scheduling dimension. Every claim and heartbeat records
226the runner's advertised architecture in `Dispatch` (`jobs.rs`), which expires
227after `RUNNER_TTL` — 5 minutes, longer than both the 55s claim poll and the 30s
228heartbeat, so an idle runner and a runner mid-build both stay visible. A
229claiming runner is then offered, in queue order:
230
2311. runs that name its platform, or name none at all;
2322. then runs whose platform *no currently connected runner* is native to.
233
234Tier 2 is what keeps one arm64 Mac usable as the only runner for pipelines that
235declare `linux/amd64`: nobody can run them natively, so it emulates them rather
236than leaving them queued forever. Add an amd64 runner and the Mac stops taking
237those jobs the moment the new runner's first claim registers it — no
238configuration, which is the property the dial-out model was chosen for.
239
240The corollary worth stating plainly: `platform:` is not a promise of native
241execution. It is a promise about *what the job runs*, which the runner enforces
242by inspecting the image it ended up with (`anvil-docker::check_platform`) and
243failing the job if the architecture is not the one asked for. Where it runs is
244a scheduling preference. A run's log header says which it got:
245
246```
247platform: linux/amd64 (emulated on linux/arm64)
248```
249
250## Isolation on macOS
251
252Docker on macOS is a Linux VM (LinuxKit under Docker Desktop, Lima under
253Colima/OrbStack) and every container runs inside it, so the sandbox is enforced
254by the same kernel primitives as on Linux: `--cap-drop=ALL` and
255`no-new-privileges` are capabilities and prctl, the pids/memory/cpu caps are
256cgroups v2.
257
258It is a boundary *better* than hagrid's. `untrusted-mode.md` §1 notes that a
259kernel or runc escape defeats the sandbox; on the Mac that escape reaches a
260disposable Linux VM, not the host.
261
262Two things to get right:
263
264- **`Docker::connect_with_socket_defaults()` (`docker.rs:15`) will not find the
265 socket.** It hardcodes `/var/run/docker.sock`; Docker Desktop only creates
266 that symlink when "Allow the default Docker socket to be used" is ticked, and
267 the real path is `~/.docker/run/docker.sock` (Colima differs again). The
268 runner uses `Docker::connect_with_defaults()`, which honours `DOCKER_HOST`.
269- **Run the runner natively under launchd, not in a container.** Containerizing
270 it means mounting the socket into it, rebuilding the root-equivalent hole
271 this change removes from hagrid.
272
273Size the VM's RAM deliberately: per-job `memory_mb` is carved out of a fixed
274allocation. The dispatcher runs one job at a time, so this is slack rather than
275a constraint.
276
277### Running one on a Mac mini
278
279```sh
280# On the Mac, in a checkout of anvil:
281cargo build --release -p anvil-worker
282sudo cp target/release/anvil-worker /usr/local/bin/
283
284# Docker Desktop → Settings → General:
285# ✓ Use Rosetta for x86_64/amd64 emulation on Apple Silicon
286# Without it, a linux/amd64 job runs under QEMU — correct, and much slower.
287
288cp deploy/worker/com.anvil.worker.plist ~/Library/LaunchAgents/
289# Fill in --url, --name, ANVIL_RUNNER_TOKEN and DOCKER_HOST, then:
290chmod 600 ~/Library/LaunchAgents/com.anvil.worker.plist
291launchctl load -w ~/Library/LaunchAgents/com.anvil.worker.plist
292tail -f /tmp/anvil-worker.log # "anvil-worker macmini (linux/arm64) → …"
293```
294
295The forge side needs `[ci] runner_token` set to the same secret; until it is,
296every claim gets a 503 saying so and queued runs sit.
297
298The plist is a **LaunchAgent**, not a LaunchDaemon, because Docker Desktop's
299socket only exists inside the logged-in user's session — a root daemon starts
300before Docker and never finds it. The consequence to know about: the runner is
301only up while that user is logged in, and a Mac that sleeps stops claiming.
302That is the failure mode the dial-out design chose (a sleeping Mac reads as "no
303runner available"), and after `RUNNER_TTL` its architecture stops counting as
304present, so anything routed to it falls back to another runner.
305
306`--name` matters: the default reads `$HOSTNAME`, which launchd does not set, so
307an unnamed runner is called `runner`.
308
309## Two processes on the build host
310
311Image building **cannot be a CI job**. Job containers get no Docker socket by
312design (`execute`'s doc comment at `:369`, and `untrusted-mode.md` §1), and
313that invariant is the whole broker model. So the build host runs two things at
314two trust levels:
315
3161. **`anvil-worker`** — claims jobs, runs them sandboxed, no Docker access
317 *inside* the job container.
3182. **a deploy agent** — fired on green CI, runs *outside* any sandbox with full
319 Docker access, does `docker build --platform linux/amd64` and `docker push`
320 to `registry.vibe.richardscollin.com`, then triggers hagrid to pull.
321
322Keeping these separate is what preserves the broker model. Folding the second
323into the first would give pipeline authors a path to the daemon.
324
325The deploy agent is out of scope for M1 — the existing `deploy_webhook` still
326works, with the receiver moved to the build host and reached over Tailscale.
327M2 folds it into the same dial-out channel as a privileged "publish" job kind,
328authorized server-side by the `is_deploy_target` check that already scopes CD
329to exactly one repository (`config.rs:357`). That removes the last inbound
330requirement.
331
332## Scope
333
334**Agent sessions stay local-socket and stay out.** `anvil-agent` calls
335`anvil_ci::docker::connect()` in five places (`lib.rs:141`,
336`supervisor.rs:55,278,414,439`) and `docker.rs` is explicitly shared plumbing,
337so the crate split touches them — but a session is interactive (tmux attach,
338exec streaming, resize, `pump_transcript`), which is a far harder protocol than
339fire-and-forget CI. They are `enabled = false` and absent from
340`deploy/anvil.toml` entirely, so nothing in production regresses.
341
342`ensure_image` and `connect` therefore need a home both crates can reach. They
343move to a thin `anvil-docker` crate rather than being duplicated —
344`ensure_image`'s local-fallback pull logic is subtle enough to be worth having
345once.
346
347The corollary: turning agent sessions on for hagrid later means either putting
348the socket back, or moving sessions onto the runner protocol too.
349
350## Registry cleanup
351
352`ensure_image`'s tolerance of a failed pull exists because `anvil-runner:latest`
353"exists in no registry" (`docker.rs:22-24`) and is built straight into the local
354daemon store. With `registry.vibe.richardscollin.com` up, push the image there
355and the fallback stops being load-bearing.
356
357It should stay in the code regardless — a runner on a fresh machine wants a
358clear failure when the pull fails and nothing is cached — but the comment
359explaining *why* it exists needs rewriting.
360
361## Milestones
362
363**M1 — the split. Done.** `anvil-job`, `anvil-docker` and `anvil-worker`
364crates, the five endpoints, in-memory leases, `[ci] runner_token`,
365`connect_with_defaults`, `run_worker` → `run_dispatcher`, and no socket mount in
366the deployed container. Deploys keep using the existing webhook.
367
368**M2 — platform. Done.** `platform:` in `.anvil/ci.yml`, a `[ci] platform`
369default, `CiConfig::resolve_platform`, the two-tier routing above (backed by a
370runner registry in `Dispatch`), the emulation note in the run header, and an
371architecture check on the image the runner actually got.
372
373**M3 — publish jobs.** The deploy agent folds into the dial-out channel,
374authorized by `is_deploy_target`. No inbound path to the build host remains.
375
376## Not yet done
377
378- **Live logs.** Newly worth doing, still not done. The result POST is a single
379 write; streaming needs chunked append with offsets and a UI that tolerates
380 gaps.
381- **Per-runner credentials.** One shared secret means one revocation for all
382 runners, and no `last_used_at`. Wants the API-token write scope first.
383- **Runner labels.** `platform` is the only scheduling dimension in M2. Tags
384 ("has-postgres", "big-memory") are the obvious next axis and are not designed.
385- **A runners page.** `Dispatch::runners()` now knows every runner connected in
386 the last five minutes and what it is. Nothing renders it, so "is my Mac
387 actually claiming?" is still answered by reading logs.
388- **Routing is per-claim, not per-queue.** A runner that can take nothing sleeps
389 until the next wake; it does not reserve the job it declined. With two runners
390 and a job only one can run natively, the other simply keeps polling — correct,
391 but it means a queue can look busy while a runner looks idle.
392- **Concurrency.** Nothing bounds how many jobs are in flight beyond how many
393 runners exist, and nothing stops one runner claiming repeatedly. The
394 in-process runner's "one job at a time" was a property of the loop, and it is
395 gone; a `max_concurrent` equivalent for CI does not exist.
396- **Concurrency, again.** Platform routing makes a second runner useful, which
397 makes the missing `max_concurrent` more pressing rather than less.