anvilsign in

collin/anvil

RenderedSource

1# Remote runners
2
3Status: **M1 + M2 implemented** (2026-08-24). Supersedes the in-process CI
4executor that used to live in `crates/anvil-ci`.
5
6anvil used to *be* the runner: `run_worker` drained the queue in-process and
7`execute` created the job container on the local Docker socket. That works, and
8it pins CI to whichever machine anvil runs on — which is hagrid, a droplet
9small enough that a release build OOMs it (see [DEPLOY.md](../DEPLOY.md) §3).
10So CI had to run somewhere else.
11
12A **runner** is a separate binary that dials out to anvil, claims a job, runs it
13in a sandboxed container on its own Docker daemon, and posts the result back.
14anvild stops executing anything.
15
16```
17 hagrid (droplet, 1 GB) build host (Mac mini M2)
18 +------------------------+ +---------------------------+
19 | anvild | | anvil-worker (launchd) |
20 | queue, tree, vault | <--------- | long-poll claim |
21 | artifacts, webhook | https | POST result |
22 | | ---------> | | |
23 | NO docker socket | | v |
24 +------------------------+ | [ job container ] |
25 | cap-drop, no socket |
26 +---------------------------+
27```
28
29## Why dial-out
30
31The alternative was forwarding the build host's Docker socket to hagrid and
32pointing anvil at it. Dial-out wins on four counts, in order of importance:
33
341. **No inbound path from the public VPS into the home LAN.** A forwarded
35 Docker socket is unauthenticated root on the machine that owns it. Punching
36 that from an internet-facing droplet into a home network makes hagrid's
37 compromise the build host's compromise.
382. **NAT and sleep/wake are free.** The build host is a desktop machine behind
39 a residential connection. A dialing client reconnects; a dialed-into server
40 needs a tunnel supervised on both ends.
413. **The failure mode is legible.** A sleeping Mac becomes "no runner
42 available," not a Docker API error surfacing as `[runner error]` in a job
43 log.
444. **Runners become plural.** Once anvil addresses runners rather than a
45 socket, a second one is configuration, not architecture — which is how the
46 arm64/amd64 split below gets solved properly.
47
48## What this buys hagrid
49
50anvild stops needing Docker at all. `compose.yaml` carries no
51`/var/run/docker.sock` mount and no `group_add`, which deletes
52the warning in [DEPLOY.md](../DEPLOY.md) §4 about the container holding
53root-equivalent control of the host. The internet-facing process stops being a
54host-escape vector.
55
56This is contingent on agent sessions staying off — see [Scope](#scope).
57
58## What moves
59
60| Stays in `anvild` | Moved to `anvil-worker` |
61| ------------------------------------------ | ---------------------------- |
62| queue, `enqueue`/`requeue_interrupted` | `execute` |
63| repo + tree resolution, `build_tar` | `collect_artifacts` |
64| `parse_pipeline`, image allowlist check | `download_tar` |
65| the secret vault (`app.vault.take`) | `parse_meta_tar` |
66| step-script assembly, `single_quote` | the `ArtifactSink` trait |
67| `store_artifact`, the swap, `gc_artifacts` | |
68| `mask_secrets` (applied on receipt) | |
69| the deploy webhook | |
70
71`store_artifact` deliberately stayed behind. The runner uploads the raw tar it
72pulled out of the container and anvil decides what to do with it, so on-disk
73layout — and `browse`, which turns a tarball into a servable directory tree —
74never becomes a runner's call. The upload response reports the stored size, so
75the per-run artifact budget is still charged what actually landed.
76
77`run_worker` becomes `run_dispatcher`: same queue drain, but instead of calling
78`execute` it parks the job until a runner claims it.
79
80Script assembly stays server-side deliberately. The runner then never parses
81`.anvil/ci.toml` and holds no pipeline model — it receives an image, a script,
82sandbox caps, and a tar. That keeps the wire format stable as the pipeline
83schema grows.
84
85### The crates
86
87| Crate | Holds |
88| --------------- | ------------------------------------------------------ |
89| `anvil-job` | the wire format, and nothing else — serde only |
90| `anvil-docker` | `connect`/`ensure_image`, shared with `anvil-agent` |
91| `anvil-worker` | the runner binary: claim loop, client, executor |
92
93`anvil-job` exists so the runner does not link `anvil-core` — and therefore
94toasty, SQLite, gix and the rest of the forge — just to learn the shape of a
95job. `anvil-docker` exists so `anvil-agent` need not depend on the runner.
96
97The binary is **`anvil-worker`**, not `anvil-runner`, because
98`anvil-runner:latest` is already the *image* CI jobs and agent sessions run in
99(see [agent-sessions.md](agent-sessions.md)). Prose says "runner" for the
100concept; the crate avoids the collision.
101
102## The protocol
103
104Plain HTTP against the existing axum server, under the `/-/` system namespace,
105so it inherits Caddy's TLS and needs no new listener.
106
107| Endpoint | Method | Purpose |
108| -------------------------------------------- | ------ | -------------------------------- |
109| `/-/runner/claim` | POST | long-poll; 204 on timeout |
110| `/-/runner/jobs/{run_id}/checkout.tar` | GET | the uploaded checkout |
111| `/-/runner/jobs/{run_id}/heartbeat` | POST | extend the lease |
112| `/-/runner/jobs/{run_id}/artifacts/{name}` | POST | one artifact tar |
113| `/-/runner/jobs/{run_id}/result` | POST | exit code + log + artifact meta |
114
115The claim response carries everything `execute` takes as arguments today:
116
117```json
118{
119 "run_id": 42,
120 "image": "anvil-runner:rust",
121 "platform": "linux/amd64",
122 "script": "set -e\n...",
123 "secrets": [{"name": "CARGO_TOKEN", "value": "..."}],
124 "artifacts": [{"name": "docs", "path": "target/doc", "browse": true}],
125 "sandbox": {
126 "memory_mb": 2048, "cpus": 2.0, "pids_limit": 512,
127 "timeout_secs": 1800, "network": true, "run_as": ""
128 }
129}
130```
131
132The checkout is a separate `GET` rather than a base64 field, so a large tree
133doesn't inflate a JSON body by a third. Secrets ride in the claim body over
134TLS, never on a separately-fetchable URL.
135
136Artifacts upload individually before `result`, for the same reason.
137
138### Logs
139
140The runner posts the whole log with `result`, which is exactly what the
141in-process runner did: `append_log` was only ever called when a run ended (and
142on the secrets-failure exit). The log accumulated in memory and landed in one
143write. The run page was not live before and is no less live now — a header
144naming the runner is now written at claim time, so a `running` run at least
145shows something.
146
147Live logs are a genuine follow-up, and a remote runner makes them *easier* to
148justify (there is now a producer that could stream). Out of scope here.
149
150### Auth
151
152A shared secret in `[ci] runner_token`, sent as `X-Anvil-Runner-Token`,
153constant-time compared. The runner's self-asserted name rides alongside in
154`X-Anvil-Runner-Name` — it labels runs and keys leases, and is explicitly not
155a credential: everyone holding the token is one principal.
156
157The four per-job endpoints additionally require the caller to hold that run's
158lease, so a valid token gets you *a* job rather than everyone else's. A lease
159mismatch answers 409, not 403: the caller is a legitimate runner whose claim
160simply expired.
161
162This matches the existing `deploy_secret` pattern rather than inventing a
163credential type. It is deliberate: API tokens are read-only and Bearer-only on
164GET/HEAD (see [untrusted-mode.md](untrusted-mode.md)), a runner must POST, and
165the write scope is still on [TODO.md](../TODO.md). Per-runner DB-backed tokens
166with `last_used_at` are the right end state; a single-tenant forge with one
167runner does not need them to start.
168
169### Leases
170
171Held **in memory** on the server (`anvil_core::jobs::Dispatch`):
172`run_id -> (runner_name, expires_at, secrets)`, bumped by `heartbeat` every 30s
173against a 120s TTL, swept every 30s. An expired lease returns the run to
174`queued`.
175
176A job's secrets are stashed on its lease rather than re-read from the vault
177when the result lands. `Vault::take` fails once the repository's unlock TTL
178lapses, and a job can easily outlive an unlock — re-reading would mean a long
179run silently skips log masking, which is exactly the run whose log is most
180likely to contain something. The values are already in this process's vault,
181so this is not new exposure.
182
183The honest gap: an expired lease can double-run a job whose runner is alive but
184unreachable. The container keeps going and the requeued run may be claimed
185elsewhere. That is what a lease without fencing buys; CI steps are assumed
186idempotent.
187
188In-memory rather than columns on `CiRun` because Toasty migrations do not exist
189yet — DEPLOY.md §Operations and TODO.md both flag that schema changes don't
190auto-apply to the live database. Adding `claimed_by`/`lease_expires_at` to
191`CiRun` (`models.rs:65`) would need a manual migration on hagrid.
192
193It also costs nothing: anvild is the only dispatcher, so a lease has no reason
194to outlive it, and the anvild-crash case is already handled —
195`requeue_interrupted` (`ci.rs:245`) re-queues everything left `running` at
196startup. The sweep covers the new case, a runner that dies mid-job.
197
198## Architecture
199
200The build host is arm64; hagrid is x86_64. Two separate concerns:
201
202**What the shipped image is.** Already solved: `deploy/build.sh` cross-compiles
203with zigbuild and `compose.yaml` pins `platforms: [linux/amd64]`. Nothing to do,
204though note the `Dockerfile`'s `RUN apt-get …` does execute amd64 binaries
205under emulation. Since `anvild` is a static musl binary,
206`gcr.io/distroless/static:nonroot` would make that build pure `COPY` and
207genuinely emulation-free. Nice-to-have.
208
209**What jobs run on.** On an M2, a multi-arch image resolves to arm64, so
210`cargo test` tests an architecture you don't ship. That is what M2 fixes.
211
212A job's platform comes from `platform` in `.anvil/ci.toml`, falling back to
213`[ci] platform`, falling back to the claiming runner's native architecture —
214so an instance that sets neither behaves exactly as it did before. It reaches
215`CreateContainerOptions.platform` and `CreateImageOptions.platform`, and the
216value is validated as `os/arch[/variant]` at parse time (a bare `amd64` would
217otherwise reach Docker as an *operating system* named amd64).
218
219Per-pipeline you then choose: native arm64 for lint and unit tests, amd64 under
220Rosetta for anything arch-sensitive. Turn Rosetta on in Docker Desktop; it is
221far faster than QEMU for amd64 Linux binaries.
222
223### Routing
224
225`platform` is also the scheduling dimension. Every claim and heartbeat records
226the runner's advertised architecture in `Dispatch` (`jobs.rs`), which expires
227after `RUNNER_TTL` — 5 minutes, longer than both the 55s claim poll and the 30s
228heartbeat, so an idle runner and a runner mid-build both stay visible. A
229claiming runner is then offered, in queue order:
230
2311. runs that name its platform, or name none at all;
2322. then runs whose platform *no currently connected runner* is native to.
233
234Tier 2 is what keeps one arm64 Mac usable as the only runner for pipelines that
235declare `linux/amd64`: nobody can run them natively, so it emulates them rather
236than leaving them queued forever. Add an amd64 runner and the Mac stops taking
237those jobs the moment the new runner's first claim registers it — no
238configuration, which is the property the dial-out model was chosen for.
239
240The corollary worth stating plainly: `platform` is not a promise of native
241execution. It is a promise about *what the job runs*, which the runner enforces
242by inspecting the image it ended up with (`anvil-docker::check_platform`) and
243failing the job if the architecture is not the one asked for. Where it runs is
244a scheduling preference. A run's log header says which it got:
245
246```
247platform: linux/amd64 (emulated on linux/arm64)
248```
249
250### Seeing who is connected
251
252`/-/admin/runners` (admin-only, 404 for everyone else) lists every runner
253`Dispatch` still counts as present: name, advertised platform, worker version,
254how long since it last spoke, how long it has been connected, and the runs it
255holds right now, each linked to its CI page. Underneath it is the same map
256routing reads, so the page and the dispatcher can never disagree about who is
257out there. It also prints the queue depth, because the two together are the
258whole diagnosis: runners and no queue is a healthy idle instance, a queue and
259no runners is a stuck one, and a queue with every runner busy is neither — it
260is capacity.
261
262**There is no separate heartbeat for presence, on purpose.** Liveness is the
263traffic a working runner already generates: an idle one re-registers itself
264every time its parked claim expires and it dials back in (`CLAIM_POLL`, 55s),
265and a busy one every `HEARTBEAT_INTERVAL` (30s) for as long as its job runs.
266Between them there is no state a runner can be in where it is useful and
267silent, so a dedicated ping would only add a way for a runner to *look* alive
268while claiming nothing.
269
270What that costs is resolution, and the page is explicit about it rather than
271hiding it. A runner is shown "late" once it has been quiet for a whole claim
272poll plus a heartbeat (85s) — long enough that a runner merely parked in a
273long poll never reads as late. It stays listed, and keeps being routed to,
274until `RUNNER_TTL` (5 min), because that is exactly what the dispatcher still
275believes; the page's job is to show that belief, not to invent a second one.
276So a machine that loses power disappears from routing in up to five minutes and
277reads as late within ninety seconds. The page reloads itself every 15s.
278
279## Isolation on macOS
280
281Docker on macOS is a Linux VM (LinuxKit under Docker Desktop, Lima under
282Colima/OrbStack) and every container runs inside it, so the sandbox is enforced
283by the same kernel primitives as on Linux: `--cap-drop=ALL` and
284`no-new-privileges` are capabilities and prctl, the pids/memory/cpu caps are
285cgroups v2.
286
287It is a boundary *better* than hagrid's. `untrusted-mode.md` §1 notes that a
288kernel or runc escape defeats the sandbox; on the Mac that escape reaches a
289disposable Linux VM, not the host.
290
291Two things to get right:
292
293- **`Docker::connect_with_socket_defaults()` (`docker.rs:15`) will not find the
294 socket.** It hardcodes `/var/run/docker.sock`; Docker Desktop only creates
295 that symlink when "Allow the default Docker socket to be used" is ticked, and
296 the real path is `~/.docker/run/docker.sock` (Colima differs again). The
297 runner uses `Docker::connect_with_defaults()`, which honours `DOCKER_HOST`.
298- **Run the runner natively under launchd, not in a container.** Containerizing
299 it means mounting the socket into it, rebuilding the root-equivalent hole
300 this change removes from hagrid.
301
302Size the VM's RAM deliberately: per-job `memory_mb` is carved out of a fixed
303allocation. The dispatcher runs one job at a time, so this is slack rather than
304a constraint.
305
306### Running one on a Mac mini
307
308```sh
309# On the Mac, in a checkout of anvil:
310cargo build --release -p anvil-worker
311sudo cp target/release/anvil-worker /usr/local/bin/
312
313# Docker Desktop → Settings → General:
314# ✓ Use Rosetta for x86_64/amd64 emulation on Apple Silicon
315# Without it, a linux/amd64 job runs under QEMU — correct, and much slower.
316
317cp deploy/worker/com.anvil.worker.plist ~/Library/LaunchAgents/
318# Fill in --url, --name, ANVIL_RUNNER_TOKEN and DOCKER_HOST, then:
319chmod 600 ~/Library/LaunchAgents/com.anvil.worker.plist
320launchctl load -w ~/Library/LaunchAgents/com.anvil.worker.plist
321tail -f /tmp/anvil-worker.log # "anvil-worker macmini (linux/arm64) → …"
322```
323
324The forge side needs `[ci] runner_token` set to the same secret; until it is,
325every claim gets a 503 saying so and queued runs sit.
326
327The plist is a **LaunchAgent**, not a LaunchDaemon, because Docker Desktop's
328socket only exists inside the logged-in user's session — a root daemon starts
329before Docker and never finds it. The consequence to know about: the runner is
330only up while that user is logged in, and a Mac that sleeps stops claiming.
331That is the failure mode the dial-out design chose (a sleeping Mac reads as "no
332runner available"), and after `RUNNER_TTL` its architecture stops counting as
333present, so anything routed to it falls back to another runner.
334
335`--name` matters: the default reads `$HOSTNAME`, which launchd does not set, so
336an unnamed runner is called `runner`.
337
338## Two runners locally
339
340`compose.override.yaml` brings up `runner-1` and `runner-2` next to the local
341forge, so the parts of this design that only appear with more than one runner —
342concurrent pipelines, and the "no connected runner is native to this platform"
343half of routing — are testable without a second machine:
344
345```sh
346./deploy/build.sh --debug --worker # stages the anvild and anvil-worker binaries
347docker compose up -d --build
348docker compose logs -f runner-1 runner-2
349# anvil-worker dev-1 (linux/amd64) → http://anvil:3000
350```
351
352They reach the forge as `http://anvil:3000` over the compose network (the
353browser-facing `base_url` does not resolve inside a container) and authenticate
354with the `runner_token` committed in `deploy/anvil.dev.toml`. Both are amd64
355here, so to watch the fallback tier work, give a pipeline `platform:
356linux/arm64` and confirm it still gets claimed. Adding a third is a copy of the
357four-line service block with a new name.
358
359These runners **are** containerized, and that is the one thing production must
360never copy: the socket mount is the root-equivalent hold that moving CI off the
361forge removed. It is acceptable locally only because the same file already
362mounts that socket into anvil for agent sessions, so the machine's trust
363boundary is unchanged. The deployed `compose.yaml` grants neither.
364
365## Two processes on the build host
366
367Image building **cannot be a CI job**. Job containers get no Docker socket by
368design (`execute`'s doc comment at `:369`, and `untrusted-mode.md` §1), and
369that invariant is the whole broker model. So the build host runs two things at
370two trust levels:
371
3721. **`anvil-worker`** — claims jobs, runs them sandboxed, no Docker access
373 *inside* the job container.
3742. **a deploy agent** — fired on green CI, runs *outside* any sandbox with full
375 Docker access, does `docker build --platform linux/amd64` and `docker push`
376 to `registry.vibe.richardscollin.com`, then triggers hagrid to pull.
377
378Keeping these separate is what preserves the broker model. Folding the second
379into the first would give pipeline authors a path to the daemon.
380
381The deploy agent is out of scope for M1 — the existing `deploy_webhook` still
382works, with the receiver moved to the build host and reached over Tailscale.
383M2 folds it into the same dial-out channel as a privileged "publish" job kind,
384authorized server-side by the `is_deploy_target` check that already scopes CD
385to exactly one repository (`config.rs:357`). That removes the last inbound
386requirement.
387
388## Scope
389
390**Agent sessions stay local-socket and stay out.** `anvil-agent` calls
391`anvil_ci::docker::connect()` in five places (`lib.rs:141`,
392`supervisor.rs:55,278,414,439`) and `docker.rs` is explicitly shared plumbing,
393so the crate split touches them — but a session is interactive (tmux attach,
394exec streaming, resize, `pump_transcript`), which is a far harder protocol than
395fire-and-forget CI. They are `enabled = false` and absent from
396`deploy/anvil.toml` entirely, so nothing in production regresses.
397
398`ensure_image` and `connect` therefore need a home both crates can reach. They
399move to a thin `anvil-docker` crate rather than being duplicated —
400`ensure_image`'s local-fallback pull logic is subtle enough to be worth having
401once.
402
403The corollary: turning agent sessions on for hagrid later means either putting
404the socket back, or moving sessions onto the runner protocol too.
405
406## Registry cleanup
407
408`ensure_image`'s tolerance of a failed pull exists because `anvil-runner:latest`
409"exists in no registry" (`docker.rs:22-24`) and is built straight into the local
410daemon store. With `registry.vibe.richardscollin.com` up, push the image there
411and the fallback stops being load-bearing.
412
413It should stay in the code regardless — a runner on a fresh machine wants a
414clear failure when the pull fails and nothing is cached — but the comment
415explaining *why* it exists needs rewriting.
416
417## Milestones
418
419**M1 — the split. Done.** `anvil-job`, `anvil-docker` and `anvil-worker`
420crates, the five endpoints, in-memory leases, `[ci] runner_token`,
421`connect_with_defaults`, `run_worker` → `run_dispatcher`, and no socket mount in
422the deployed container. Deploys keep using the existing webhook.
423
424**M2 — platform. Done.** `platform` in `.anvil/ci.toml`, a `[ci] platform`
425default, `CiConfig::resolve_platform`, the two-tier routing above (backed by a
426runner registry in `Dispatch`), the emulation note in the run header, and an
427architecture check on the image the runner actually got.
428
429**M3 — publish jobs.** The deploy agent folds into the dial-out channel,
430authorized by `is_deploy_target`. No inbound path to the build host remains.
431
432## Not yet done
433
434- **Live logs.** Newly worth doing, still not done. The result POST is a single
435 write; streaming needs chunked append with offsets and a UI that tolerates
436 gaps.
437- **Per-runner credentials.** One shared secret means one revocation for all
438 runners, and no `last_used_at`. Wants the API-token write scope first.
439- **Runner labels.** `platform` is the only scheduling dimension in M2. Tags
440 ("has-postgres", "big-memory") are the obvious next axis and are not designed.
441- **A runners page.** `Dispatch::runners()` now knows every runner connected in
442 the last five minutes and what it is. Nothing renders it, so "is my Mac
443 actually claiming?" is still answered by reading logs.
444- **Routing is per-claim, not per-queue.** A runner that can take nothing sleeps
445 until the next wake; it does not reserve the job it declined. With two runners
446 and a job only one can run natively, the other simply keeps polling — correct,
447 but it means a queue can look busy while a runner looks idle.
448- **Concurrency.** Nothing bounds how many jobs are in flight beyond how many
449 runners exist, and nothing stops one runner claiming repeatedly. The
450 in-process runner's "one job at a time" was a property of the loop, and it is
451 gone; a `max_concurrent` equivalent for CI does not exist.
452- **Concurrency, again.** Platform routing makes a second runner useful, which
453 makes the missing `max_concurrent` more pressing rather than less.