collin/anvil
RenderedSource
| 1 | # Remote runners |
| 2 | |
| 3 | Status: **M1 + M2 implemented** (2026-08-24). Supersedes the in-process CI |
| 4 | executor that used to live in `crates/anvil-ci`. |
| 5 | |
| 6 | anvil used to *be* the runner: `run_worker` drained the queue in-process and |
| 7 | `execute` created the job container on the local Docker socket. That works, and |
| 8 | it pins CI to whichever machine anvil runs on — which is hagrid, a droplet |
| 9 | small enough that a release build OOMs it (see [DEPLOY.md](../DEPLOY.md) §3). |
| 10 | So CI had to run somewhere else. |
| 11 | |
| 12 | A **runner** is a separate binary that dials out to anvil, claims a job, runs it |
| 13 | in a sandboxed container on its own Docker daemon, and posts the result back. |
| 14 | anvild stops executing anything. |
| 15 | |
| 16 | ``` |
| 17 | hagrid (droplet, 1 GB) build host (Mac mini M2) |
| 18 | +------------------------+ +---------------------------+ |
| 19 | | anvild | | anvil-worker (launchd) | |
| 20 | | queue, tree, vault | <--------- | long-poll claim | |
| 21 | | artifacts, webhook | https | POST result | |
| 22 | | | ---------> | | | |
| 23 | | NO docker socket | | v | |
| 24 | +------------------------+ | [ job container ] | |
| 25 | | cap-drop, no socket | |
| 26 | +---------------------------+ |
| 27 | ``` |
| 28 | |
| 29 | ## Why dial-out |
| 30 | |
| 31 | The alternative was forwarding the build host's Docker socket to hagrid and |
| 32 | pointing anvil at it. Dial-out wins on four counts, in order of importance: |
| 33 | |
| 34 | 1. **No inbound path from the public VPS into the home LAN.** A forwarded |
| 35 | Docker socket is unauthenticated root on the machine that owns it. Punching |
| 36 | that from an internet-facing droplet into a home network makes hagrid's |
| 37 | compromise the build host's compromise. |
| 38 | 2. **NAT and sleep/wake are free.** The build host is a desktop machine behind |
| 39 | a residential connection. A dialing client reconnects; a dialed-into server |
| 40 | needs a tunnel supervised on both ends. |
| 41 | 3. **The failure mode is legible.** A sleeping Mac becomes "no runner |
| 42 | available," not a Docker API error surfacing as `[runner error]` in a job |
| 43 | log. |
| 44 | 4. **Runners become plural.** Once anvil addresses runners rather than a |
| 45 | socket, a second one is configuration, not architecture — which is how the |
| 46 | arm64/amd64 split below gets solved properly. |
| 47 | |
| 48 | ## What this buys hagrid |
| 49 | |
| 50 | anvild stops needing Docker at all. `compose.yaml` carries no |
| 51 | `/var/run/docker.sock` mount and no `group_add`, which deletes |
| 52 | the warning in [DEPLOY.md](../DEPLOY.md) §4 about the container holding |
| 53 | root-equivalent control of the host. The internet-facing process stops being a |
| 54 | host-escape vector. |
| 55 | |
| 56 | This is contingent on agent sessions staying off — see [Scope](#scope). |
| 57 | |
| 58 | ## What moves |
| 59 | |
| 60 | | Stays in `anvild` | Moved to `anvil-worker` | |
| 61 | | ------------------------------------------ | ---------------------------- | |
| 62 | | queue, `enqueue`/`requeue_interrupted` | `execute` | |
| 63 | | repo + tree resolution, `build_tar` | `collect_artifacts` | |
| 64 | | `parse_pipeline`, image allowlist check | `download_tar` | |
| 65 | | the secret vault (`app.vault.take`) | `parse_meta_tar` | |
| 66 | | step-script assembly, `single_quote` | the `ArtifactSink` trait | |
| 67 | | `store_artifact`, the swap, `gc_artifacts` | | |
| 68 | | `mask_secrets` (applied on receipt) | | |
| 69 | | the deploy webhook | | |
| 70 | |
| 71 | `store_artifact` deliberately stayed behind. The runner uploads the raw tar it |
| 72 | pulled out of the container and anvil decides what to do with it, so on-disk |
| 73 | layout — and `browse`, which turns a tarball into a servable directory tree — |
| 74 | never becomes a runner's call. The upload response reports the stored size, so |
| 75 | the per-run artifact budget is still charged what actually landed. |
| 76 | |
| 77 | `run_worker` becomes `run_dispatcher`: same queue drain, but instead of calling |
| 78 | `execute` it parks the job until a runner claims it. |
| 79 | |
| 80 | Script assembly stays server-side deliberately. The runner then never parses |
| 81 | `.anvil/ci.yml` and holds no pipeline model — it receives an image, a script, |
| 82 | sandbox caps, and a tar. That keeps the wire format stable as the pipeline |
| 83 | schema grows. |
| 84 | |
| 85 | ### The crates |
| 86 | |
| 87 | | Crate | Holds | |
| 88 | | --------------- | ------------------------------------------------------ | |
| 89 | | `anvil-job` | the wire format, and nothing else — serde only | |
| 90 | | `anvil-docker` | `connect`/`ensure_image`, shared with `anvil-agent` | |
| 91 | | `anvil-worker` | the runner binary: claim loop, client, executor | |
| 92 | |
| 93 | `anvil-job` exists so the runner does not link `anvil-core` — and therefore |
| 94 | toasty, SQLite, gix and the rest of the forge — just to learn the shape of a |
| 95 | job. `anvil-docker` exists so `anvil-agent` need not depend on the runner. |
| 96 | |
| 97 | The binary is **`anvil-worker`**, not `anvil-runner`, because |
| 98 | `anvil-runner:latest` is already the *image* CI jobs and agent sessions run in |
| 99 | (see [agent-sessions.md](agent-sessions.md)). Prose says "runner" for the |
| 100 | concept; the crate avoids the collision. |
| 101 | |
| 102 | ## The protocol |
| 103 | |
| 104 | Plain HTTP against the existing axum server, under the `/-/` system namespace, |
| 105 | so it inherits Caddy's TLS and needs no new listener. |
| 106 | |
| 107 | | Endpoint | Method | Purpose | |
| 108 | | -------------------------------------------- | ------ | -------------------------------- | |
| 109 | | `/-/runner/claim` | POST | long-poll; 204 on timeout | |
| 110 | | `/-/runner/jobs/{run_id}/checkout.tar` | GET | the uploaded checkout | |
| 111 | | `/-/runner/jobs/{run_id}/heartbeat` | POST | extend the lease | |
| 112 | | `/-/runner/jobs/{run_id}/artifacts/{name}` | POST | one artifact tar | |
| 113 | | `/-/runner/jobs/{run_id}/result` | POST | exit code + log + artifact meta | |
| 114 | |
| 115 | The claim response carries everything `execute` takes as arguments today: |
| 116 | |
| 117 | ```json |
| 118 | { |
| 119 | "run_id": 42, |
| 120 | "image": "rust:1.95-bookworm", |
| 121 | "platform": "linux/amd64", |
| 122 | "script": "set -e\n...", |
| 123 | "secrets": [{"name": "CARGO_TOKEN", "value": "..."}], |
| 124 | "artifacts": [{"name": "docs", "path": "target/doc", "browse": true}], |
| 125 | "sandbox": { |
| 126 | "memory_mb": 2048, "cpus": 2.0, "pids_limit": 512, |
| 127 | "timeout_secs": 1800, "network": true, "run_as": "" |
| 128 | } |
| 129 | } |
| 130 | ``` |
| 131 | |
| 132 | The checkout is a separate `GET` rather than a base64 field, so a large tree |
| 133 | doesn't inflate a JSON body by a third. Secrets ride in the claim body over |
| 134 | TLS, never on a separately-fetchable URL. |
| 135 | |
| 136 | Artifacts upload individually before `result`, for the same reason. |
| 137 | |
| 138 | ### Logs |
| 139 | |
| 140 | The runner posts the whole log with `result`, which is exactly what the |
| 141 | in-process runner did: `append_log` was only ever called when a run ended (and |
| 142 | on the secrets-failure exit). The log accumulated in memory and landed in one |
| 143 | write. The run page was not live before and is no less live now — a header |
| 144 | naming the runner is now written at claim time, so a `running` run at least |
| 145 | shows something. |
| 146 | |
| 147 | Live logs are a genuine follow-up, and a remote runner makes them *easier* to |
| 148 | justify (there is now a producer that could stream). Out of scope here. |
| 149 | |
| 150 | ### Auth |
| 151 | |
| 152 | A shared secret in `[ci] runner_token`, sent as `X-Anvil-Runner-Token`, |
| 153 | constant-time compared. The runner's self-asserted name rides alongside in |
| 154 | `X-Anvil-Runner-Name` — it labels runs and keys leases, and is explicitly not |
| 155 | a credential: everyone holding the token is one principal. |
| 156 | |
| 157 | The four per-job endpoints additionally require the caller to hold that run's |
| 158 | lease, so a valid token gets you *a* job rather than everyone else's. A lease |
| 159 | mismatch answers 409, not 403: the caller is a legitimate runner whose claim |
| 160 | simply expired. |
| 161 | |
| 162 | This matches the existing `deploy_secret` pattern rather than inventing a |
| 163 | credential type. It is deliberate: API tokens are read-only and Bearer-only on |
| 164 | GET/HEAD (see [untrusted-mode.md](untrusted-mode.md)), a runner must POST, and |
| 165 | the write scope is still on [TODO.md](../TODO.md). Per-runner DB-backed tokens |
| 166 | with `last_used_at` are the right end state; a single-tenant forge with one |
| 167 | runner does not need them to start. |
| 168 | |
| 169 | ### Leases |
| 170 | |
| 171 | Held **in memory** on the server (`anvil_core::jobs::Dispatch`): |
| 172 | `run_id -> (runner_name, expires_at, secrets)`, bumped by `heartbeat` every 30s |
| 173 | against a 120s TTL, swept every 30s. An expired lease returns the run to |
| 174 | `queued`. |
| 175 | |
| 176 | A job's secrets are stashed on its lease rather than re-read from the vault |
| 177 | when the result lands. `Vault::take` fails once the repository's unlock TTL |
| 178 | lapses, and a job can easily outlive an unlock — re-reading would mean a long |
| 179 | run silently skips log masking, which is exactly the run whose log is most |
| 180 | likely to contain something. The values are already in this process's vault, |
| 181 | so this is not new exposure. |
| 182 | |
| 183 | The honest gap: an expired lease can double-run a job whose runner is alive but |
| 184 | unreachable. The container keeps going and the requeued run may be claimed |
| 185 | elsewhere. That is what a lease without fencing buys; CI steps are assumed |
| 186 | idempotent. |
| 187 | |
| 188 | In-memory rather than columns on `CiRun` because Toasty migrations do not exist |
| 189 | yet — DEPLOY.md §Operations and TODO.md both flag that schema changes don't |
| 190 | auto-apply to the live database. Adding `claimed_by`/`lease_expires_at` to |
| 191 | `CiRun` (`models.rs:65`) would need a manual migration on hagrid. |
| 192 | |
| 193 | It also costs nothing: anvild is the only dispatcher, so a lease has no reason |
| 194 | to outlive it, and the anvild-crash case is already handled — |
| 195 | `requeue_interrupted` (`ci.rs:245`) re-queues everything left `running` at |
| 196 | startup. The sweep covers the new case, a runner that dies mid-job. |
| 197 | |
| 198 | ## Architecture |
| 199 | |
| 200 | The build host is arm64; hagrid is x86_64. Two separate concerns: |
| 201 | |
| 202 | **What the shipped image is.** Already solved: `deploy/build.sh` cross-compiles |
| 203 | with zigbuild and `compose.yaml` pins `platforms: [linux/amd64]`. Nothing to do, |
| 204 | though note the `Dockerfile`'s `RUN apt-get …` does execute amd64 binaries |
| 205 | under emulation. Since `anvild` is a static musl binary, |
| 206 | `gcr.io/distroless/static:nonroot` would make that build pure `COPY` and |
| 207 | genuinely emulation-free. Nice-to-have. |
| 208 | |
| 209 | **What jobs run on.** On an M2, `rust:1.95-bookworm` resolves to arm64, so |
| 210 | `cargo test` tests an architecture you don't ship. That is what M2 fixes. |
| 211 | |
| 212 | A job's platform comes from `platform:` in `.anvil/ci.yml`, falling back to |
| 213 | `[ci] platform`, falling back to the claiming runner's native architecture — |
| 214 | so an instance that sets neither behaves exactly as it did before. It reaches |
| 215 | `CreateContainerOptions.platform` and `CreateImageOptions.platform`, and the |
| 216 | value is validated as `os/arch[/variant]` at parse time (a bare `amd64` would |
| 217 | otherwise reach Docker as an *operating system* named amd64). |
| 218 | |
| 219 | Per-pipeline you then choose: native arm64 for lint and unit tests, amd64 under |
| 220 | Rosetta for anything arch-sensitive. Turn Rosetta on in Docker Desktop; it is |
| 221 | far faster than QEMU for amd64 Linux binaries. |
| 222 | |
| 223 | ### Routing |
| 224 | |
| 225 | `platform` is also the scheduling dimension. Every claim and heartbeat records |
| 226 | the runner's advertised architecture in `Dispatch` (`jobs.rs`), which expires |
| 227 | after `RUNNER_TTL` — 5 minutes, longer than both the 55s claim poll and the 30s |
| 228 | heartbeat, so an idle runner and a runner mid-build both stay visible. A |
| 229 | claiming runner is then offered, in queue order: |
| 230 | |
| 231 | 1. runs that name its platform, or name none at all; |
| 232 | 2. then runs whose platform *no currently connected runner* is native to. |
| 233 | |
| 234 | Tier 2 is what keeps one arm64 Mac usable as the only runner for pipelines that |
| 235 | declare `linux/amd64`: nobody can run them natively, so it emulates them rather |
| 236 | than leaving them queued forever. Add an amd64 runner and the Mac stops taking |
| 237 | those jobs the moment the new runner's first claim registers it — no |
| 238 | configuration, which is the property the dial-out model was chosen for. |
| 239 | |
| 240 | The corollary worth stating plainly: `platform:` is not a promise of native |
| 241 | execution. It is a promise about *what the job runs*, which the runner enforces |
| 242 | by inspecting the image it ended up with (`anvil-docker::check_platform`) and |
| 243 | failing the job if the architecture is not the one asked for. Where it runs is |
| 244 | a scheduling preference. A run's log header says which it got: |
| 245 | |
| 246 | ``` |
| 247 | platform: linux/amd64 (emulated on linux/arm64) |
| 248 | ``` |
| 249 | |
| 250 | ## Isolation on macOS |
| 251 | |
| 252 | Docker on macOS is a Linux VM (LinuxKit under Docker Desktop, Lima under |
| 253 | Colima/OrbStack) and every container runs inside it, so the sandbox is enforced |
| 254 | by the same kernel primitives as on Linux: `--cap-drop=ALL` and |
| 255 | `no-new-privileges` are capabilities and prctl, the pids/memory/cpu caps are |
| 256 | cgroups v2. |
| 257 | |
| 258 | It is a boundary *better* than hagrid's. `untrusted-mode.md` §1 notes that a |
| 259 | kernel or runc escape defeats the sandbox; on the Mac that escape reaches a |
| 260 | disposable Linux VM, not the host. |
| 261 | |
| 262 | Two things to get right: |
| 263 | |
| 264 | - **`Docker::connect_with_socket_defaults()` (`docker.rs:15`) will not find the |
| 265 | socket.** It hardcodes `/var/run/docker.sock`; Docker Desktop only creates |
| 266 | that symlink when "Allow the default Docker socket to be used" is ticked, and |
| 267 | the real path is `~/.docker/run/docker.sock` (Colima differs again). The |
| 268 | runner uses `Docker::connect_with_defaults()`, which honours `DOCKER_HOST`. |
| 269 | - **Run the runner natively under launchd, not in a container.** Containerizing |
| 270 | it means mounting the socket into it, rebuilding the root-equivalent hole |
| 271 | this change removes from hagrid. |
| 272 | |
| 273 | Size the VM's RAM deliberately: per-job `memory_mb` is carved out of a fixed |
| 274 | allocation. The dispatcher runs one job at a time, so this is slack rather than |
| 275 | a constraint. |
| 276 | |
| 277 | ### Running one on a Mac mini |
| 278 | |
| 279 | ```sh |
| 280 | # On the Mac, in a checkout of anvil: |
| 281 | cargo build --release -p anvil-worker |
| 282 | sudo cp target/release/anvil-worker /usr/local/bin/ |
| 283 | |
| 284 | # Docker Desktop → Settings → General: |
| 285 | # ✓ Use Rosetta for x86_64/amd64 emulation on Apple Silicon |
| 286 | # Without it, a linux/amd64 job runs under QEMU — correct, and much slower. |
| 287 | |
| 288 | cp deploy/worker/com.anvil.worker.plist ~/Library/LaunchAgents/ |
| 289 | # Fill in --url, --name, ANVIL_RUNNER_TOKEN and DOCKER_HOST, then: |
| 290 | chmod 600 ~/Library/LaunchAgents/com.anvil.worker.plist |
| 291 | launchctl load -w ~/Library/LaunchAgents/com.anvil.worker.plist |
| 292 | tail -f /tmp/anvil-worker.log # "anvil-worker macmini (linux/arm64) → …" |
| 293 | ``` |
| 294 | |
| 295 | The forge side needs `[ci] runner_token` set to the same secret; until it is, |
| 296 | every claim gets a 503 saying so and queued runs sit. |
| 297 | |
| 298 | The plist is a **LaunchAgent**, not a LaunchDaemon, because Docker Desktop's |
| 299 | socket only exists inside the logged-in user's session — a root daemon starts |
| 300 | before Docker and never finds it. The consequence to know about: the runner is |
| 301 | only up while that user is logged in, and a Mac that sleeps stops claiming. |
| 302 | That is the failure mode the dial-out design chose (a sleeping Mac reads as "no |
| 303 | runner available"), and after `RUNNER_TTL` its architecture stops counting as |
| 304 | present, so anything routed to it falls back to another runner. |
| 305 | |
| 306 | `--name` matters: the default reads `$HOSTNAME`, which launchd does not set, so |
| 307 | an unnamed runner is called `runner`. |
| 308 | |
| 309 | ## Two processes on the build host |
| 310 | |
| 311 | Image building **cannot be a CI job**. Job containers get no Docker socket by |
| 312 | design (`execute`'s doc comment at `:369`, and `untrusted-mode.md` §1), and |
| 313 | that invariant is the whole broker model. So the build host runs two things at |
| 314 | two trust levels: |
| 315 | |
| 316 | 1. **`anvil-worker`** — claims jobs, runs them sandboxed, no Docker access |
| 317 | *inside* the job container. |
| 318 | 2. **a deploy agent** — fired on green CI, runs *outside* any sandbox with full |
| 319 | Docker access, does `docker build --platform linux/amd64` and `docker push` |
| 320 | to `registry.vibe.richardscollin.com`, then triggers hagrid to pull. |
| 321 | |
| 322 | Keeping these separate is what preserves the broker model. Folding the second |
| 323 | into the first would give pipeline authors a path to the daemon. |
| 324 | |
| 325 | The deploy agent is out of scope for M1 — the existing `deploy_webhook` still |
| 326 | works, with the receiver moved to the build host and reached over Tailscale. |
| 327 | M2 folds it into the same dial-out channel as a privileged "publish" job kind, |
| 328 | authorized server-side by the `is_deploy_target` check that already scopes CD |
| 329 | to exactly one repository (`config.rs:357`). That removes the last inbound |
| 330 | requirement. |
| 331 | |
| 332 | ## Scope |
| 333 | |
| 334 | **Agent sessions stay local-socket and stay out.** `anvil-agent` calls |
| 335 | `anvil_ci::docker::connect()` in five places (`lib.rs:141`, |
| 336 | `supervisor.rs:55,278,414,439`) and `docker.rs` is explicitly shared plumbing, |
| 337 | so the crate split touches them — but a session is interactive (tmux attach, |
| 338 | exec streaming, resize, `pump_transcript`), which is a far harder protocol than |
| 339 | fire-and-forget CI. They are `enabled = false` and absent from |
| 340 | `deploy/anvil.toml` entirely, so nothing in production regresses. |
| 341 | |
| 342 | `ensure_image` and `connect` therefore need a home both crates can reach. They |
| 343 | move to a thin `anvil-docker` crate rather than being duplicated — |
| 344 | `ensure_image`'s local-fallback pull logic is subtle enough to be worth having |
| 345 | once. |
| 346 | |
| 347 | The corollary: turning agent sessions on for hagrid later means either putting |
| 348 | the socket back, or moving sessions onto the runner protocol too. |
| 349 | |
| 350 | ## Registry cleanup |
| 351 | |
| 352 | `ensure_image`'s tolerance of a failed pull exists because `anvil-runner:latest` |
| 353 | "exists in no registry" (`docker.rs:22-24`) and is built straight into the local |
| 354 | daemon store. With `registry.vibe.richardscollin.com` up, push the image there |
| 355 | and the fallback stops being load-bearing. |
| 356 | |
| 357 | It should stay in the code regardless — a runner on a fresh machine wants a |
| 358 | clear failure when the pull fails and nothing is cached — but the comment |
| 359 | explaining *why* it exists needs rewriting. |
| 360 | |
| 361 | ## Milestones |
| 362 | |
| 363 | **M1 — the split. Done.** `anvil-job`, `anvil-docker` and `anvil-worker` |
| 364 | crates, the five endpoints, in-memory leases, `[ci] runner_token`, |
| 365 | `connect_with_defaults`, `run_worker` → `run_dispatcher`, and no socket mount in |
| 366 | the deployed container. Deploys keep using the existing webhook. |
| 367 | |
| 368 | **M2 — platform. Done.** `platform:` in `.anvil/ci.yml`, a `[ci] platform` |
| 369 | default, `CiConfig::resolve_platform`, the two-tier routing above (backed by a |
| 370 | runner registry in `Dispatch`), the emulation note in the run header, and an |
| 371 | architecture check on the image the runner actually got. |
| 372 | |
| 373 | **M3 — publish jobs.** The deploy agent folds into the dial-out channel, |
| 374 | authorized by `is_deploy_target`. No inbound path to the build host remains. |
| 375 | |
| 376 | ## Not yet done |
| 377 | |
| 378 | - **Live logs.** Newly worth doing, still not done. The result POST is a single |
| 379 | write; streaming needs chunked append with offsets and a UI that tolerates |
| 380 | gaps. |
| 381 | - **Per-runner credentials.** One shared secret means one revocation for all |
| 382 | runners, and no `last_used_at`. Wants the API-token write scope first. |
| 383 | - **Runner labels.** `platform` is the only scheduling dimension in M2. Tags |
| 384 | ("has-postgres", "big-memory") are the obvious next axis and are not designed. |
| 385 | - **A runners page.** `Dispatch::runners()` now knows every runner connected in |
| 386 | the last five minutes and what it is. Nothing renders it, so "is my Mac |
| 387 | actually claiming?" is still answered by reading logs. |
| 388 | - **Routing is per-claim, not per-queue.** A runner that can take nothing sleeps |
| 389 | until the next wake; it does not reserve the job it declined. With two runners |
| 390 | and a job only one can run natively, the other simply keeps polling — correct, |
| 391 | but it means a queue can look busy while a runner looks idle. |
| 392 | - **Concurrency.** Nothing bounds how many jobs are in flight beyond how many |
| 393 | runners exist, and nothing stops one runner claiming repeatedly. The |
| 394 | in-process runner's "one job at a time" was a property of the loop, and it is |
| 395 | gone; a `max_concurrent` equivalent for CI does not exist. |
| 396 | - **Concurrency, again.** Platform routing makes a second runner useful, which |
| 397 | makes the missing `max_concurrent` more pressing rather than less. |