collin/anvil
RenderedSource
| 1 | # Remote runners |
| 2 | |
| 3 | Status: **M1 implemented** (2026-08-24). Supersedes the in-process CI |
| 4 | executor that used to live in `crates/anvil-ci`. |
| 5 | |
| 6 | anvil used to *be* the runner: `run_worker` drained the queue in-process and |
| 7 | `execute` created the job container on the local Docker socket. That works, and |
| 8 | it pins CI to whichever machine anvil runs on — which is hagrid, a droplet |
| 9 | small enough that a release build OOMs it (see [DEPLOY.md](../DEPLOY.md) §3). |
| 10 | So CI had to run somewhere else. |
| 11 | |
| 12 | A **runner** is a separate binary that dials out to anvil, claims a job, runs it |
| 13 | in a sandboxed container on its own Docker daemon, and posts the result back. |
| 14 | anvild stops executing anything. |
| 15 | |
| 16 | ``` |
| 17 | hagrid (droplet, 1 GB) build host (Mac mini M2) |
| 18 | +------------------------+ +---------------------------+ |
| 19 | | anvild | | anvil-worker (launchd) | |
| 20 | | queue, tree, vault | <--------- | long-poll claim | |
| 21 | | artifacts, webhook | https | POST result | |
| 22 | | | ---------> | | | |
| 23 | | NO docker socket | | v | |
| 24 | +------------------------+ | [ job container ] | |
| 25 | | cap-drop, no socket | |
| 26 | +---------------------------+ |
| 27 | ``` |
| 28 | |
| 29 | ## Why dial-out |
| 30 | |
| 31 | The alternative was forwarding the build host's Docker socket to hagrid and |
| 32 | pointing anvil at it. Dial-out wins on four counts, in order of importance: |
| 33 | |
| 34 | 1. **No inbound path from the public VPS into the home LAN.** A forwarded |
| 35 | Docker socket is unauthenticated root on the machine that owns it. Punching |
| 36 | that from an internet-facing droplet into a home network makes hagrid's |
| 37 | compromise the build host's compromise. |
| 38 | 2. **NAT and sleep/wake are free.** The build host is a desktop machine behind |
| 39 | a residential connection. A dialing client reconnects; a dialed-into server |
| 40 | needs a tunnel supervised on both ends. |
| 41 | 3. **The failure mode is legible.** A sleeping Mac becomes "no runner |
| 42 | available," not a Docker API error surfacing as `[runner error]` in a job |
| 43 | log. |
| 44 | 4. **Runners become plural.** Once anvil addresses runners rather than a |
| 45 | socket, a second one is configuration, not architecture — which is how the |
| 46 | arm64/amd64 split below gets solved properly. |
| 47 | |
| 48 | ## What this buys hagrid |
| 49 | |
| 50 | anvild stops needing Docker at all. `deploy/run.sh` drops |
| 51 | `-v /var/run/docker.sock:/var/run/docker.sock` and `--group-add`, which deletes |
| 52 | the warning in [DEPLOY.md](../DEPLOY.md) §4 about the container holding |
| 53 | root-equivalent control of the host. The internet-facing process stops being a |
| 54 | host-escape vector. |
| 55 | |
| 56 | This is contingent on agent sessions staying off — see [Scope](#scope). |
| 57 | |
| 58 | ## What moves |
| 59 | |
| 60 | | Stays in `anvild` | Moved to `anvil-worker` | |
| 61 | | ------------------------------------------ | ---------------------------- | |
| 62 | | queue, `enqueue`/`requeue_interrupted` | `execute` | |
| 63 | | repo + tree resolution, `build_tar` | `collect_artifacts` | |
| 64 | | `parse_pipeline`, image allowlist check | `download_tar` | |
| 65 | | the secret vault (`app.vault.take`) | `parse_meta_tar` | |
| 66 | | step-script assembly, `single_quote` | the `ArtifactSink` trait | |
| 67 | | `store_artifact`, the swap, `gc_artifacts` | | |
| 68 | | `mask_secrets` (applied on receipt) | | |
| 69 | | the deploy webhook | | |
| 70 | |
| 71 | `store_artifact` deliberately stayed behind. The runner uploads the raw tar it |
| 72 | pulled out of the container and anvil decides what to do with it, so on-disk |
| 73 | layout — and `browse`, which turns a tarball into a servable directory tree — |
| 74 | never becomes a runner's call. The upload response reports the stored size, so |
| 75 | the per-run artifact budget is still charged what actually landed. |
| 76 | |
| 77 | `run_worker` becomes `run_dispatcher`: same queue drain, but instead of calling |
| 78 | `execute` it parks the job until a runner claims it. |
| 79 | |
| 80 | Script assembly stays server-side deliberately. The runner then never parses |
| 81 | `.anvil/ci.yml` and holds no pipeline model — it receives an image, a script, |
| 82 | sandbox caps, and a tar. That keeps the wire format stable as the pipeline |
| 83 | schema grows. |
| 84 | |
| 85 | ### The crates |
| 86 | |
| 87 | | Crate | Holds | |
| 88 | | --------------- | ------------------------------------------------------ | |
| 89 | | `anvil-job` | the wire format, and nothing else — serde only | |
| 90 | | `anvil-docker` | `connect`/`ensure_image`, shared with `anvil-agent` | |
| 91 | | `anvil-worker` | the runner binary: claim loop, client, executor | |
| 92 | |
| 93 | `anvil-job` exists so the runner does not link `anvil-core` — and therefore |
| 94 | toasty, SQLite, gix and the rest of the forge — just to learn the shape of a |
| 95 | job. `anvil-docker` exists so `anvil-agent` need not depend on the runner. |
| 96 | |
| 97 | The binary is **`anvil-worker`**, not `anvil-runner`, because |
| 98 | `anvil-runner:latest` is already the *image* CI jobs and agent sessions run in |
| 99 | (see [agent-sessions.md](agent-sessions.md)). Prose says "runner" for the |
| 100 | concept; the crate avoids the collision. |
| 101 | |
| 102 | ## The protocol |
| 103 | |
| 104 | Plain HTTP against the existing axum server, under the `/-/` system namespace, |
| 105 | so it inherits Caddy's TLS and needs no new listener. |
| 106 | |
| 107 | | Endpoint | Method | Purpose | |
| 108 | | -------------------------------------------- | ------ | -------------------------------- | |
| 109 | | `/-/runner/claim` | POST | long-poll; 204 on timeout | |
| 110 | | `/-/runner/jobs/{run_id}/checkout.tar` | GET | the uploaded checkout | |
| 111 | | `/-/runner/jobs/{run_id}/heartbeat` | POST | extend the lease | |
| 112 | | `/-/runner/jobs/{run_id}/artifacts/{name}` | POST | one artifact tar | |
| 113 | | `/-/runner/jobs/{run_id}/result` | POST | exit code + log + artifact meta | |
| 114 | |
| 115 | The claim response carries everything `execute` takes as arguments today: |
| 116 | |
| 117 | ```json |
| 118 | { |
| 119 | "run_id": 42, |
| 120 | "image": "rust:1.95-bookworm", |
| 121 | "platform": "linux/amd64", |
| 122 | "script": "set -e\n...", |
| 123 | "secrets": [{"name": "CARGO_TOKEN", "value": "..."}], |
| 124 | "artifacts": [{"name": "docs", "path": "target/doc", "browse": true}], |
| 125 | "sandbox": { |
| 126 | "memory_mb": 2048, "cpus": 2.0, "pids_limit": 512, |
| 127 | "timeout_secs": 1800, "network": true, "run_as": "" |
| 128 | } |
| 129 | } |
| 130 | ``` |
| 131 | |
| 132 | The checkout is a separate `GET` rather than a base64 field, so a large tree |
| 133 | doesn't inflate a JSON body by a third. Secrets ride in the claim body over |
| 134 | TLS, never on a separately-fetchable URL. |
| 135 | |
| 136 | Artifacts upload individually before `result`, for the same reason. |
| 137 | |
| 138 | ### Logs |
| 139 | |
| 140 | The runner posts the whole log with `result`, which is exactly what the |
| 141 | in-process runner did: `append_log` was only ever called when a run ended (and |
| 142 | on the secrets-failure exit). The log accumulated in memory and landed in one |
| 143 | write. The run page was not live before and is no less live now — a header |
| 144 | naming the runner is now written at claim time, so a `running` run at least |
| 145 | shows something. |
| 146 | |
| 147 | Live logs are a genuine follow-up, and a remote runner makes them *easier* to |
| 148 | justify (there is now a producer that could stream). Out of scope here. |
| 149 | |
| 150 | ### Auth |
| 151 | |
| 152 | A shared secret in `[ci] runner_token`, sent as `X-Anvil-Runner-Token`, |
| 153 | constant-time compared. The runner's self-asserted name rides alongside in |
| 154 | `X-Anvil-Runner-Name` — it labels runs and keys leases, and is explicitly not |
| 155 | a credential: everyone holding the token is one principal. |
| 156 | |
| 157 | The four per-job endpoints additionally require the caller to hold that run's |
| 158 | lease, so a valid token gets you *a* job rather than everyone else's. A lease |
| 159 | mismatch answers 409, not 403: the caller is a legitimate runner whose claim |
| 160 | simply expired. |
| 161 | |
| 162 | This matches the existing `deploy_secret` pattern rather than inventing a |
| 163 | credential type. It is deliberate: API tokens are read-only and Bearer-only on |
| 164 | GET/HEAD (see [untrusted-mode.md](untrusted-mode.md)), a runner must POST, and |
| 165 | the write scope is still on [TODO.md](../TODO.md). Per-runner DB-backed tokens |
| 166 | with `last_used_at` are the right end state; a single-tenant forge with one |
| 167 | runner does not need them to start. |
| 168 | |
| 169 | ### Leases |
| 170 | |
| 171 | Held **in memory** on the server (`anvil_core::jobs::Dispatch`): |
| 172 | `run_id -> (runner_name, expires_at, secrets)`, bumped by `heartbeat` every 30s |
| 173 | against a 120s TTL, swept every 30s. An expired lease returns the run to |
| 174 | `queued`. |
| 175 | |
| 176 | A job's secrets are stashed on its lease rather than re-read from the vault |
| 177 | when the result lands. `Vault::take` fails once the repository's unlock TTL |
| 178 | lapses, and a job can easily outlive an unlock — re-reading would mean a long |
| 179 | run silently skips log masking, which is exactly the run whose log is most |
| 180 | likely to contain something. The values are already in this process's vault, |
| 181 | so this is not new exposure. |
| 182 | |
| 183 | The honest gap: an expired lease can double-run a job whose runner is alive but |
| 184 | unreachable. The container keeps going and the requeued run may be claimed |
| 185 | elsewhere. That is what a lease without fencing buys; CI steps are assumed |
| 186 | idempotent. |
| 187 | |
| 188 | In-memory rather than columns on `CiRun` because Toasty migrations do not exist |
| 189 | yet — DEPLOY.md §Operations and TODO.md both flag that schema changes don't |
| 190 | auto-apply to the live database. Adding `claimed_by`/`lease_expires_at` to |
| 191 | `CiRun` (`models.rs:65`) would need a manual migration on hagrid. |
| 192 | |
| 193 | It also costs nothing: anvild is the only dispatcher, so a lease has no reason |
| 194 | to outlive it, and the anvild-crash case is already handled — |
| 195 | `requeue_interrupted` (`ci.rs:245`) re-queues everything left `running` at |
| 196 | startup. The sweep covers the new case, a runner that dies mid-job. |
| 197 | |
| 198 | ## Architecture |
| 199 | |
| 200 | The build host is arm64; hagrid is x86_64. Two separate concerns: |
| 201 | |
| 202 | **What the shipped image is.** Already solved: `deploy/build.sh:33` cross- |
| 203 | compiles with zigbuild and builds `--platform linux/amd64`. Nothing to do, |
| 204 | though note the `Dockerfile`'s `RUN apt-get …` does execute amd64 binaries |
| 205 | under emulation. Since `anvild` is a static musl binary, |
| 206 | `gcr.io/distroless/static:nonroot` would make that build pure `COPY` and |
| 207 | genuinely emulation-free. Nice-to-have. |
| 208 | |
| 209 | **What jobs run on.** On an M2, `rust:1.95-bookworm` resolves to arm64, so |
| 210 | `cargo test` tests an architecture you don't ship. Fix: a `platform:` key in |
| 211 | `.anvil/ci.yml`, plumbed to `CreateContainerOptions.platform` |
| 212 | (bollard 0.18 `container.rs:107`) and `CreateImageOptions.platform` |
| 213 | (`image.rs:72`) — `ensure_image` currently leaves both at `Default::default()`. |
| 214 | Enable Docker Desktop's Rosetta option; it is far faster than QEMU for amd64 |
| 215 | Linux binaries. |
| 216 | |
| 217 | Per-pipeline you then choose: native arm64 for lint and unit tests, amd64 under |
| 218 | Rosetta for anything arch-sensitive. |
| 219 | |
| 220 | Once runners are plural, `platform` also becomes a scheduling hint — a job |
| 221 | declaring `linux/amd64` prefers a runner that is natively amd64. Not needed |
| 222 | with one runner; the field is forward-compatible with it. |
| 223 | |
| 224 | ## Isolation on macOS |
| 225 | |
| 226 | Docker on macOS is a Linux VM (LinuxKit under Docker Desktop, Lima under |
| 227 | Colima/OrbStack) and every container runs inside it, so the sandbox is enforced |
| 228 | by the same kernel primitives as on Linux: `--cap-drop=ALL` and |
| 229 | `no-new-privileges` are capabilities and prctl, the pids/memory/cpu caps are |
| 230 | cgroups v2. |
| 231 | |
| 232 | It is a boundary *better* than hagrid's. `untrusted-mode.md` §1 notes that a |
| 233 | kernel or runc escape defeats the sandbox; on the Mac that escape reaches a |
| 234 | disposable Linux VM, not the host. |
| 235 | |
| 236 | Two things to get right: |
| 237 | |
| 238 | - **`Docker::connect_with_socket_defaults()` (`docker.rs:15`) will not find the |
| 239 | socket.** It hardcodes `/var/run/docker.sock`; Docker Desktop only creates |
| 240 | that symlink when "Allow the default Docker socket to be used" is ticked, and |
| 241 | the real path is `~/.docker/run/docker.sock` (Colima differs again). The |
| 242 | runner uses `Docker::connect_with_defaults()`, which honours `DOCKER_HOST`. |
| 243 | - **Run the runner natively under launchd, not in a container.** Containerizing |
| 244 | it means mounting the socket into it, rebuilding the root-equivalent hole |
| 245 | this change removes from hagrid. |
| 246 | |
| 247 | Size the VM's RAM deliberately: per-job `memory_mb` is carved out of a fixed |
| 248 | allocation. The dispatcher runs one job at a time, so this is slack rather than |
| 249 | a constraint. |
| 250 | |
| 251 | ## Two processes on the build host |
| 252 | |
| 253 | Image building **cannot be a CI job**. Job containers get no Docker socket by |
| 254 | design (`execute`'s doc comment at `:369`, and `untrusted-mode.md` §1), and |
| 255 | that invariant is the whole broker model. So the build host runs two things at |
| 256 | two trust levels: |
| 257 | |
| 258 | 1. **`anvil-worker`** — claims jobs, runs them sandboxed, no Docker access |
| 259 | *inside* the job container. |
| 260 | 2. **a deploy agent** — fired on green CI, runs *outside* any sandbox with full |
| 261 | Docker access, does `docker build --platform linux/amd64` and `docker push` |
| 262 | to `registry.vibe.richardscollin.com`, then triggers hagrid to pull. |
| 263 | |
| 264 | Keeping these separate is what preserves the broker model. Folding the second |
| 265 | into the first would give pipeline authors a path to the daemon. |
| 266 | |
| 267 | The deploy agent is out of scope for M1 — the existing `deploy_webhook` still |
| 268 | works, with the receiver moved to the build host and reached over Tailscale. |
| 269 | M2 folds it into the same dial-out channel as a privileged "publish" job kind, |
| 270 | authorized server-side by the `is_deploy_target` check that already scopes CD |
| 271 | to exactly one repository (`config.rs:357`). That removes the last inbound |
| 272 | requirement. |
| 273 | |
| 274 | ## Scope |
| 275 | |
| 276 | **Agent sessions stay local-socket and stay out.** `anvil-agent` calls |
| 277 | `anvil_ci::docker::connect()` in five places (`lib.rs:141`, |
| 278 | `supervisor.rs:55,278,414,439`) and `docker.rs` is explicitly shared plumbing, |
| 279 | so the crate split touches them — but a session is interactive (tmux attach, |
| 280 | exec streaming, resize, `pump_transcript`), which is a far harder protocol than |
| 281 | fire-and-forget CI. They are `enabled = false` and absent from |
| 282 | `deploy/anvil.toml` entirely, so nothing in production regresses. |
| 283 | |
| 284 | `ensure_image` and `connect` therefore need a home both crates can reach. They |
| 285 | move to a thin `anvil-docker` crate rather than being duplicated — |
| 286 | `ensure_image`'s local-fallback pull logic is subtle enough to be worth having |
| 287 | once. |
| 288 | |
| 289 | The corollary: turning agent sessions on for hagrid later means either putting |
| 290 | the socket back, or moving sessions onto the runner protocol too. |
| 291 | |
| 292 | ## Registry cleanup |
| 293 | |
| 294 | `ensure_image`'s tolerance of a failed pull exists because `anvil-runner:latest` |
| 295 | "exists in no registry" (`docker.rs:22-24`) and is built straight into the local |
| 296 | daemon store. With `registry.vibe.richardscollin.com` up, push the image there |
| 297 | and the fallback stops being load-bearing. |
| 298 | |
| 299 | It should stay in the code regardless — a runner on a fresh machine wants a |
| 300 | clear failure when the pull fails and nothing is cached — but the comment |
| 301 | explaining *why* it exists needs rewriting. |
| 302 | |
| 303 | ## Milestones |
| 304 | |
| 305 | **M1 — the split. Done.** `anvil-job`, `anvil-docker` and `anvil-worker` |
| 306 | crates, the five endpoints, in-memory leases, `[ci] runner_token`, |
| 307 | `connect_with_defaults`, `run_worker` → `run_dispatcher`, socket mount dropped |
| 308 | from `deploy/run.sh`. Deploys keep using the existing webhook. |
| 309 | |
| 310 | **M2 — platform.** The plumbing is already live: `JobSpec.platform` reaches |
| 311 | `CreateContainerOptions` and `CreateImageOptions`, and a runner advertises its |
| 312 | native platform when it claims. What is missing is a *source* — a `platform:` |
| 313 | key in `.anvil/ci.yml` (and probably a `[ci] platform` default), plus routing a |
| 314 | job to a runner that has that architecture natively. |
| 315 | |
| 316 | **M3 — publish jobs.** The deploy agent folds into the dial-out channel, |
| 317 | authorized by `is_deploy_target`. No inbound path to the build host remains. |
| 318 | |
| 319 | ## Not yet done |
| 320 | |
| 321 | - **Live logs.** Newly worth doing, still not done. The result POST is a single |
| 322 | write; streaming needs chunked append with offsets and a UI that tolerates |
| 323 | gaps. |
| 324 | - **Per-runner credentials.** One shared secret means one revocation for all |
| 325 | runners, and no `last_used_at`. Wants the API-token write scope first. |
| 326 | - **Runner labels.** `platform` is the only scheduling dimension in M2. Tags |
| 327 | ("has-postgres", "big-memory") are the obvious next axis and are not designed. |
| 328 | - **Concurrency.** Nothing bounds how many jobs are in flight beyond how many |
| 329 | runners exist, and nothing stops one runner claiming repeatedly. The |
| 330 | in-process runner's "one job at a time" was a property of the loop, and it is |
| 331 | gone; a `max_concurrent` equivalent for CI does not exist. |
| 332 | - **`ensure_image`'s local fallback vs `platform`.** A cached image of the |
| 333 | wrong architecture satisfies the fallback, since it inspects presence and not |
| 334 | arch. Only bites an offline runner asked to cross-build. |