collin/anvil
RenderedSource
| 1 | # Remote runners |
| 2 | |
| 3 | Status: **M1 + M2 implemented** (2026-08-24). Supersedes the in-process CI |
| 4 | executor that used to live in `crates/anvil-ci`. |
| 5 | |
| 6 | anvil used to *be* the runner: `run_worker` drained the queue in-process and |
| 7 | `execute` created the job container on the local Docker socket. That works, and |
| 8 | it pins CI to whichever machine anvil runs on — which is hagrid, a droplet |
| 9 | small enough that a release build OOMs it (see [DEPLOY.md](../DEPLOY.md) §3). |
| 10 | So CI had to run somewhere else. |
| 11 | |
| 12 | A **runner** is a separate binary that dials out to anvil, claims a job, runs it |
| 13 | in a sandboxed container on its own Docker daemon, and posts the result back. |
| 14 | anvild stops executing anything. |
| 15 | |
| 16 | ``` |
| 17 | hagrid (droplet, 1 GB) build host (Mac mini M2) |
| 18 | +------------------------+ +---------------------------+ |
| 19 | | anvild | | anvil-worker (launchd) | |
| 20 | | queue, tree, vault | <--------- | long-poll claim | |
| 21 | | artifacts, webhook | https | POST result | |
| 22 | | | ---------> | | | |
| 23 | | NO docker socket | | v | |
| 24 | +------------------------+ | [ job container ] | |
| 25 | | cap-drop, no socket | |
| 26 | +---------------------------+ |
| 27 | ``` |
| 28 | |
| 29 | ## Why dial-out |
| 30 | |
| 31 | The alternative was forwarding the build host's Docker socket to hagrid and |
| 32 | pointing anvil at it. Dial-out wins on four counts, in order of importance: |
| 33 | |
| 34 | 1. **No inbound path from the public VPS into the home LAN.** A forwarded |
| 35 | Docker socket is unauthenticated root on the machine that owns it. Punching |
| 36 | that from an internet-facing droplet into a home network makes hagrid's |
| 37 | compromise the build host's compromise. |
| 38 | 2. **NAT and sleep/wake are free.** The build host is a desktop machine behind |
| 39 | a residential connection. A dialing client reconnects; a dialed-into server |
| 40 | needs a tunnel supervised on both ends. |
| 41 | 3. **The failure mode is legible.** A sleeping Mac becomes "no runner |
| 42 | available," not a Docker API error surfacing as `[runner error]` in a job |
| 43 | log. |
| 44 | 4. **Runners become plural.** Once anvil addresses runners rather than a |
| 45 | socket, a second one is configuration, not architecture — which is how the |
| 46 | arm64/amd64 split below gets solved properly. |
| 47 | |
| 48 | ## What this buys hagrid |
| 49 | |
| 50 | anvild stops needing Docker at all. `compose.yaml` carries no |
| 51 | `/var/run/docker.sock` mount and no `group_add`, which deletes |
| 52 | the warning in [DEPLOY.md](../DEPLOY.md) §4 about the container holding |
| 53 | root-equivalent control of the host. The internet-facing process stops being a |
| 54 | host-escape vector. |
| 55 | |
| 56 | This is contingent on agent sessions staying off — see [Scope](#scope). |
| 57 | |
| 58 | ## What moves |
| 59 | |
| 60 | | Stays in `anvild` | Moved to `anvil-worker` | |
| 61 | | ------------------------------------------ | ---------------------------- | |
| 62 | | queue, `enqueue`/`requeue_interrupted` | `execute` | |
| 63 | | repo + tree resolution, `build_tar` | `collect_artifacts` | |
| 64 | | `parse_pipeline`, image allowlist check | `download_tar` | |
| 65 | | the secret vault (`app.vault.take`) | `parse_meta_tar` | |
| 66 | | step-script assembly, `single_quote` | the `ArtifactSink` trait | |
| 67 | | `store_artifact`, the swap, `gc_artifacts` | | |
| 68 | | `mask_secrets` (applied on receipt) | | |
| 69 | | the deploy webhook | | |
| 70 | |
| 71 | `store_artifact` deliberately stayed behind. The runner uploads the raw tar it |
| 72 | pulled out of the container and anvil decides what to do with it, so on-disk |
| 73 | layout — and `browse`, which turns a tarball into a servable directory tree — |
| 74 | never becomes a runner's call. The upload response reports the stored size, so |
| 75 | the per-run artifact budget is still charged what actually landed. |
| 76 | |
| 77 | `run_worker` becomes `run_dispatcher`: same queue drain, but instead of calling |
| 78 | `execute` it parks the job until a runner claims it. |
| 79 | |
| 80 | Script assembly stays server-side deliberately. The runner then never parses |
| 81 | `.anvil/ci.yml` and holds no pipeline model — it receives an image, a script, |
| 82 | sandbox caps, and a tar. That keeps the wire format stable as the pipeline |
| 83 | schema grows. |
| 84 | |
| 85 | ### The crates |
| 86 | |
| 87 | | Crate | Holds | |
| 88 | | --------------- | ------------------------------------------------------ | |
| 89 | | `anvil-job` | the wire format, and nothing else — serde only | |
| 90 | | `anvil-docker` | `connect`/`ensure_image`, shared with `anvil-agent` | |
| 91 | | `anvil-worker` | the runner binary: claim loop, client, executor | |
| 92 | |
| 93 | `anvil-job` exists so the runner does not link `anvil-core` — and therefore |
| 94 | toasty, SQLite, gix and the rest of the forge — just to learn the shape of a |
| 95 | job. `anvil-docker` exists so `anvil-agent` need not depend on the runner. |
| 96 | |
| 97 | The binary is **`anvil-worker`**, not `anvil-runner`, because |
| 98 | `anvil-runner:latest` is already the *image* CI jobs and agent sessions run in |
| 99 | (see [agent-sessions.md](agent-sessions.md)). Prose says "runner" for the |
| 100 | concept; the crate avoids the collision. |
| 101 | |
| 102 | ## The protocol |
| 103 | |
| 104 | Plain HTTP against the existing axum server, under the `/-/` system namespace, |
| 105 | so it inherits Caddy's TLS and needs no new listener. |
| 106 | |
| 107 | | Endpoint | Method | Purpose | |
| 108 | | -------------------------------------------- | ------ | -------------------------------- | |
| 109 | | `/-/runner/claim` | POST | long-poll; 204 on timeout | |
| 110 | | `/-/runner/jobs/{run_id}/checkout.tar` | GET | the uploaded checkout | |
| 111 | | `/-/runner/jobs/{run_id}/heartbeat` | POST | extend the lease | |
| 112 | | `/-/runner/jobs/{run_id}/artifacts/{name}` | POST | one artifact tar | |
| 113 | | `/-/runner/jobs/{run_id}/result` | POST | exit code + log + artifact meta | |
| 114 | |
| 115 | The claim response carries everything `execute` takes as arguments today: |
| 116 | |
| 117 | ```json |
| 118 | { |
| 119 | "run_id": 42, |
| 120 | "image": "anvil-runner:rust", |
| 121 | "platform": "linux/amd64", |
| 122 | "script": "set -e\n...", |
| 123 | "secrets": [{"name": "CARGO_TOKEN", "value": "..."}], |
| 124 | "artifacts": [{"name": "docs", "path": "target/doc", "browse": true}], |
| 125 | "sandbox": { |
| 126 | "memory_mb": 2048, "cpus": 2.0, "pids_limit": 512, |
| 127 | "timeout_secs": 1800, "network": true, "run_as": "" |
| 128 | } |
| 129 | } |
| 130 | ``` |
| 131 | |
| 132 | The checkout is a separate `GET` rather than a base64 field, so a large tree |
| 133 | doesn't inflate a JSON body by a third. Secrets ride in the claim body over |
| 134 | TLS, never on a separately-fetchable URL. |
| 135 | |
| 136 | Artifacts upload individually before `result`, for the same reason. |
| 137 | |
| 138 | ### Logs |
| 139 | |
| 140 | The runner posts the whole log with `result`, which is exactly what the |
| 141 | in-process runner did: `append_log` was only ever called when a run ended (and |
| 142 | on the secrets-failure exit). The log accumulated in memory and landed in one |
| 143 | write. The run page was not live before and is no less live now — a header |
| 144 | naming the runner is now written at claim time, so a `running` run at least |
| 145 | shows something. |
| 146 | |
| 147 | Live logs are a genuine follow-up, and a remote runner makes them *easier* to |
| 148 | justify (there is now a producer that could stream). Out of scope here. |
| 149 | |
| 150 | ### Auth |
| 151 | |
| 152 | A shared secret in `[ci] runner_token`, sent as `X-Anvil-Runner-Token`, |
| 153 | constant-time compared. The runner's self-asserted name rides alongside in |
| 154 | `X-Anvil-Runner-Name` — it labels runs and keys leases, and is explicitly not |
| 155 | a credential: everyone holding the token is one principal. |
| 156 | |
| 157 | The four per-job endpoints additionally require the caller to hold that run's |
| 158 | lease, so a valid token gets you *a* job rather than everyone else's. A lease |
| 159 | mismatch answers 409, not 403: the caller is a legitimate runner whose claim |
| 160 | simply expired. |
| 161 | |
| 162 | This matches the existing `deploy_secret` pattern rather than inventing a |
| 163 | credential type. It is deliberate: API tokens are read-only and Bearer-only on |
| 164 | GET/HEAD (see [untrusted-mode.md](untrusted-mode.md)), a runner must POST, and |
| 165 | the write scope is still on [TODO.md](../TODO.md). Per-runner DB-backed tokens |
| 166 | with `last_used_at` are the right end state; a single-tenant forge with one |
| 167 | runner does not need them to start. |
| 168 | |
| 169 | ### Leases |
| 170 | |
| 171 | Held **in memory** on the server (`anvil_core::jobs::Dispatch`): |
| 172 | `run_id -> (runner_name, expires_at, secrets)`, bumped by `heartbeat` every 30s |
| 173 | against a 120s TTL, swept every 30s. An expired lease returns the run to |
| 174 | `queued`. |
| 175 | |
| 176 | A job's secrets are stashed on its lease rather than re-read from the vault |
| 177 | when the result lands. `Vault::take` fails once the repository's unlock TTL |
| 178 | lapses, and a job can easily outlive an unlock — re-reading would mean a long |
| 179 | run silently skips log masking, which is exactly the run whose log is most |
| 180 | likely to contain something. The values are already in this process's vault, |
| 181 | so this is not new exposure. |
| 182 | |
| 183 | The honest gap: an expired lease can double-run a job whose runner is alive but |
| 184 | unreachable. The container keeps going and the requeued run may be claimed |
| 185 | elsewhere. That is what a lease without fencing buys; CI steps are assumed |
| 186 | idempotent. |
| 187 | |
| 188 | In-memory rather than columns on `CiRun` because Toasty migrations do not exist |
| 189 | yet — DEPLOY.md §Operations and TODO.md both flag that schema changes don't |
| 190 | auto-apply to the live database. Adding `claimed_by`/`lease_expires_at` to |
| 191 | `CiRun` (`models.rs:65`) would need a manual migration on hagrid. |
| 192 | |
| 193 | It also costs nothing: anvild is the only dispatcher, so a lease has no reason |
| 194 | to outlive it, and the anvild-crash case is already handled — |
| 195 | `requeue_interrupted` (`ci.rs:245`) re-queues everything left `running` at |
| 196 | startup. The sweep covers the new case, a runner that dies mid-job. |
| 197 | |
| 198 | ## Architecture |
| 199 | |
| 200 | The build host is arm64; hagrid is x86_64. Two separate concerns: |
| 201 | |
| 202 | **What the shipped image is.** Already solved: `deploy/build.sh` cross-compiles |
| 203 | with zigbuild and `compose.yaml` pins `platforms: [linux/amd64]`. Nothing to do, |
| 204 | though note the `Dockerfile`'s `RUN apt-get …` does execute amd64 binaries |
| 205 | under emulation. Since `anvild` is a static musl binary, |
| 206 | `gcr.io/distroless/static:nonroot` would make that build pure `COPY` and |
| 207 | genuinely emulation-free. Nice-to-have. |
| 208 | |
| 209 | **What jobs run on.** On an M2, a multi-arch image resolves to arm64, so |
| 210 | `cargo test` tests an architecture you don't ship. That is what M2 fixes. |
| 211 | |
| 212 | A job's platform comes from `platform:` in `.anvil/ci.yml`, falling back to |
| 213 | `[ci] platform`, falling back to the claiming runner's native architecture — |
| 214 | so an instance that sets neither behaves exactly as it did before. It reaches |
| 215 | `CreateContainerOptions.platform` and `CreateImageOptions.platform`, and the |
| 216 | value is validated as `os/arch[/variant]` at parse time (a bare `amd64` would |
| 217 | otherwise reach Docker as an *operating system* named amd64). |
| 218 | |
| 219 | Per-pipeline you then choose: native arm64 for lint and unit tests, amd64 under |
| 220 | Rosetta for anything arch-sensitive. Turn Rosetta on in Docker Desktop; it is |
| 221 | far faster than QEMU for amd64 Linux binaries. |
| 222 | |
| 223 | ### Routing |
| 224 | |
| 225 | `platform` is also the scheduling dimension. Every claim and heartbeat records |
| 226 | the runner's advertised architecture in `Dispatch` (`jobs.rs`), which expires |
| 227 | after `RUNNER_TTL` — 5 minutes, longer than both the 55s claim poll and the 30s |
| 228 | heartbeat, so an idle runner and a runner mid-build both stay visible. A |
| 229 | claiming runner is then offered, in queue order: |
| 230 | |
| 231 | 1. runs that name its platform, or name none at all; |
| 232 | 2. then runs whose platform *no currently connected runner* is native to. |
| 233 | |
| 234 | Tier 2 is what keeps one arm64 Mac usable as the only runner for pipelines that |
| 235 | declare `linux/amd64`: nobody can run them natively, so it emulates them rather |
| 236 | than leaving them queued forever. Add an amd64 runner and the Mac stops taking |
| 237 | those jobs the moment the new runner's first claim registers it — no |
| 238 | configuration, which is the property the dial-out model was chosen for. |
| 239 | |
| 240 | The corollary worth stating plainly: `platform:` is not a promise of native |
| 241 | execution. It is a promise about *what the job runs*, which the runner enforces |
| 242 | by inspecting the image it ended up with (`anvil-docker::check_platform`) and |
| 243 | failing the job if the architecture is not the one asked for. Where it runs is |
| 244 | a scheduling preference. A run's log header says which it got: |
| 245 | |
| 246 | ``` |
| 247 | platform: linux/amd64 (emulated on linux/arm64) |
| 248 | ``` |
| 249 | |
| 250 | ### Seeing who is connected |
| 251 | |
| 252 | `/-/admin/runners` (admin-only, 404 for everyone else) lists every runner |
| 253 | `Dispatch` still counts as present: name, advertised platform, worker version, |
| 254 | how long since it last spoke, how long it has been connected, and the runs it |
| 255 | holds right now, each linked to its CI page. Underneath it is the same map |
| 256 | routing reads, so the page and the dispatcher can never disagree about who is |
| 257 | out there. It also prints the queue depth, because the two together are the |
| 258 | whole diagnosis: runners and no queue is a healthy idle instance, a queue and |
| 259 | no runners is a stuck one, and a queue with every runner busy is neither — it |
| 260 | is capacity. |
| 261 | |
| 262 | **There is no separate heartbeat for presence, on purpose.** Liveness is the |
| 263 | traffic a working runner already generates: an idle one re-registers itself |
| 264 | every time its parked claim expires and it dials back in (`CLAIM_POLL`, 55s), |
| 265 | and a busy one every `HEARTBEAT_INTERVAL` (30s) for as long as its job runs. |
| 266 | Between them there is no state a runner can be in where it is useful and |
| 267 | silent, so a dedicated ping would only add a way for a runner to *look* alive |
| 268 | while claiming nothing. |
| 269 | |
| 270 | What that costs is resolution, and the page is explicit about it rather than |
| 271 | hiding it. A runner is shown "late" once it has been quiet for a whole claim |
| 272 | poll plus a heartbeat (85s) — long enough that a runner merely parked in a |
| 273 | long poll never reads as late. It stays listed, and keeps being routed to, |
| 274 | until `RUNNER_TTL` (5 min), because that is exactly what the dispatcher still |
| 275 | believes; the page's job is to show that belief, not to invent a second one. |
| 276 | So a machine that loses power disappears from routing in up to five minutes and |
| 277 | reads as late within ninety seconds. The page reloads itself every 15s. |
| 278 | |
| 279 | ## Isolation on macOS |
| 280 | |
| 281 | Docker on macOS is a Linux VM (LinuxKit under Docker Desktop, Lima under |
| 282 | Colima/OrbStack) and every container runs inside it, so the sandbox is enforced |
| 283 | by the same kernel primitives as on Linux: `--cap-drop=ALL` and |
| 284 | `no-new-privileges` are capabilities and prctl, the pids/memory/cpu caps are |
| 285 | cgroups v2. |
| 286 | |
| 287 | It is a boundary *better* than hagrid's. `untrusted-mode.md` §1 notes that a |
| 288 | kernel or runc escape defeats the sandbox; on the Mac that escape reaches a |
| 289 | disposable Linux VM, not the host. |
| 290 | |
| 291 | Two things to get right: |
| 292 | |
| 293 | - **`Docker::connect_with_socket_defaults()` (`docker.rs:15`) will not find the |
| 294 | socket.** It hardcodes `/var/run/docker.sock`; Docker Desktop only creates |
| 295 | that symlink when "Allow the default Docker socket to be used" is ticked, and |
| 296 | the real path is `~/.docker/run/docker.sock` (Colima differs again). The |
| 297 | runner uses `Docker::connect_with_defaults()`, which honours `DOCKER_HOST`. |
| 298 | - **Run the runner natively under launchd, not in a container.** Containerizing |
| 299 | it means mounting the socket into it, rebuilding the root-equivalent hole |
| 300 | this change removes from hagrid. |
| 301 | |
| 302 | Size the VM's RAM deliberately: per-job `memory_mb` is carved out of a fixed |
| 303 | allocation. The dispatcher runs one job at a time, so this is slack rather than |
| 304 | a constraint. |
| 305 | |
| 306 | ### Running one on a Mac mini |
| 307 | |
| 308 | ```sh |
| 309 | # On the Mac, in a checkout of anvil: |
| 310 | cargo build --release -p anvil-worker |
| 311 | sudo cp target/release/anvil-worker /usr/local/bin/ |
| 312 | |
| 313 | # Docker Desktop → Settings → General: |
| 314 | # ✓ Use Rosetta for x86_64/amd64 emulation on Apple Silicon |
| 315 | # Without it, a linux/amd64 job runs under QEMU — correct, and much slower. |
| 316 | |
| 317 | cp deploy/worker/com.anvil.worker.plist ~/Library/LaunchAgents/ |
| 318 | # Fill in --url, --name, ANVIL_RUNNER_TOKEN and DOCKER_HOST, then: |
| 319 | chmod 600 ~/Library/LaunchAgents/com.anvil.worker.plist |
| 320 | launchctl load -w ~/Library/LaunchAgents/com.anvil.worker.plist |
| 321 | tail -f /tmp/anvil-worker.log # "anvil-worker macmini (linux/arm64) → …" |
| 322 | ``` |
| 323 | |
| 324 | The forge side needs `[ci] runner_token` set to the same secret; until it is, |
| 325 | every claim gets a 503 saying so and queued runs sit. |
| 326 | |
| 327 | The plist is a **LaunchAgent**, not a LaunchDaemon, because Docker Desktop's |
| 328 | socket only exists inside the logged-in user's session — a root daemon starts |
| 329 | before Docker and never finds it. The consequence to know about: the runner is |
| 330 | only up while that user is logged in, and a Mac that sleeps stops claiming. |
| 331 | That is the failure mode the dial-out design chose (a sleeping Mac reads as "no |
| 332 | runner available"), and after `RUNNER_TTL` its architecture stops counting as |
| 333 | present, so anything routed to it falls back to another runner. |
| 334 | |
| 335 | `--name` matters: the default reads `$HOSTNAME`, which launchd does not set, so |
| 336 | an unnamed runner is called `runner`. |
| 337 | |
| 338 | ## Two runners locally |
| 339 | |
| 340 | `compose.override.yaml` brings up `runner-1` and `runner-2` next to the local |
| 341 | forge, so the parts of this design that only appear with more than one runner — |
| 342 | concurrent pipelines, and the "no connected runner is native to this platform" |
| 343 | half of routing — are testable without a second machine: |
| 344 | |
| 345 | ```sh |
| 346 | ./deploy/build.sh --debug --worker # stages deploy/anvild and deploy/anvil-worker |
| 347 | docker compose up -d --build |
| 348 | docker compose logs -f runner-1 runner-2 |
| 349 | # anvil-worker dev-1 (linux/amd64) → http://anvil:3000 |
| 350 | ``` |
| 351 | |
| 352 | They reach the forge as `http://anvil:3000` over the compose network (the |
| 353 | browser-facing `base_url` does not resolve inside a container) and authenticate |
| 354 | with the `runner_token` committed in `deploy/anvil.dev.toml`. Both are amd64 |
| 355 | here, so to watch the fallback tier work, give a pipeline `platform: |
| 356 | linux/arm64` and confirm it still gets claimed. Adding a third is a copy of the |
| 357 | four-line service block with a new name. |
| 358 | |
| 359 | These runners **are** containerized, and that is the one thing production must |
| 360 | never copy: the socket mount is the root-equivalent hold that moving CI off the |
| 361 | forge removed. It is acceptable locally only because the same file already |
| 362 | mounts that socket into anvil for agent sessions, so the machine's trust |
| 363 | boundary is unchanged. The deployed `compose.yaml` grants neither. |
| 364 | |
| 365 | ## Two processes on the build host |
| 366 | |
| 367 | Image building **cannot be a CI job**. Job containers get no Docker socket by |
| 368 | design (`execute`'s doc comment at `:369`, and `untrusted-mode.md` §1), and |
| 369 | that invariant is the whole broker model. So the build host runs two things at |
| 370 | two trust levels: |
| 371 | |
| 372 | 1. **`anvil-worker`** — claims jobs, runs them sandboxed, no Docker access |
| 373 | *inside* the job container. |
| 374 | 2. **a deploy agent** — fired on green CI, runs *outside* any sandbox with full |
| 375 | Docker access, does `docker build --platform linux/amd64` and `docker push` |
| 376 | to `registry.vibe.richardscollin.com`, then triggers hagrid to pull. |
| 377 | |
| 378 | Keeping these separate is what preserves the broker model. Folding the second |
| 379 | into the first would give pipeline authors a path to the daemon. |
| 380 | |
| 381 | The deploy agent is out of scope for M1 — the existing `deploy_webhook` still |
| 382 | works, with the receiver moved to the build host and reached over Tailscale. |
| 383 | M2 folds it into the same dial-out channel as a privileged "publish" job kind, |
| 384 | authorized server-side by the `is_deploy_target` check that already scopes CD |
| 385 | to exactly one repository (`config.rs:357`). That removes the last inbound |
| 386 | requirement. |
| 387 | |
| 388 | ## Scope |
| 389 | |
| 390 | **Agent sessions stay local-socket and stay out.** `anvil-agent` calls |
| 391 | `anvil_ci::docker::connect()` in five places (`lib.rs:141`, |
| 392 | `supervisor.rs:55,278,414,439`) and `docker.rs` is explicitly shared plumbing, |
| 393 | so the crate split touches them — but a session is interactive (tmux attach, |
| 394 | exec streaming, resize, `pump_transcript`), which is a far harder protocol than |
| 395 | fire-and-forget CI. They are `enabled = false` and absent from |
| 396 | `deploy/anvil.toml` entirely, so nothing in production regresses. |
| 397 | |
| 398 | `ensure_image` and `connect` therefore need a home both crates can reach. They |
| 399 | move to a thin `anvil-docker` crate rather than being duplicated — |
| 400 | `ensure_image`'s local-fallback pull logic is subtle enough to be worth having |
| 401 | once. |
| 402 | |
| 403 | The corollary: turning agent sessions on for hagrid later means either putting |
| 404 | the socket back, or moving sessions onto the runner protocol too. |
| 405 | |
| 406 | ## Registry cleanup |
| 407 | |
| 408 | `ensure_image`'s tolerance of a failed pull exists because `anvil-runner:latest` |
| 409 | "exists in no registry" (`docker.rs:22-24`) and is built straight into the local |
| 410 | daemon store. With `registry.vibe.richardscollin.com` up, push the image there |
| 411 | and the fallback stops being load-bearing. |
| 412 | |
| 413 | It should stay in the code regardless — a runner on a fresh machine wants a |
| 414 | clear failure when the pull fails and nothing is cached — but the comment |
| 415 | explaining *why* it exists needs rewriting. |
| 416 | |
| 417 | ## Milestones |
| 418 | |
| 419 | **M1 — the split. Done.** `anvil-job`, `anvil-docker` and `anvil-worker` |
| 420 | crates, the five endpoints, in-memory leases, `[ci] runner_token`, |
| 421 | `connect_with_defaults`, `run_worker` → `run_dispatcher`, and no socket mount in |
| 422 | the deployed container. Deploys keep using the existing webhook. |
| 423 | |
| 424 | **M2 — platform. Done.** `platform:` in `.anvil/ci.yml`, a `[ci] platform` |
| 425 | default, `CiConfig::resolve_platform`, the two-tier routing above (backed by a |
| 426 | runner registry in `Dispatch`), the emulation note in the run header, and an |
| 427 | architecture check on the image the runner actually got. |
| 428 | |
| 429 | **M3 — publish jobs.** The deploy agent folds into the dial-out channel, |
| 430 | authorized by `is_deploy_target`. No inbound path to the build host remains. |
| 431 | |
| 432 | ## Not yet done |
| 433 | |
| 434 | - **Live logs.** Newly worth doing, still not done. The result POST is a single |
| 435 | write; streaming needs chunked append with offsets and a UI that tolerates |
| 436 | gaps. |
| 437 | - **Per-runner credentials.** One shared secret means one revocation for all |
| 438 | runners, and no `last_used_at`. Wants the API-token write scope first. |
| 439 | - **Runner labels.** `platform` is the only scheduling dimension in M2. Tags |
| 440 | ("has-postgres", "big-memory") are the obvious next axis and are not designed. |
| 441 | - **A runners page.** `Dispatch::runners()` now knows every runner connected in |
| 442 | the last five minutes and what it is. Nothing renders it, so "is my Mac |
| 443 | actually claiming?" is still answered by reading logs. |
| 444 | - **Routing is per-claim, not per-queue.** A runner that can take nothing sleeps |
| 445 | until the next wake; it does not reserve the job it declined. With two runners |
| 446 | and a job only one can run natively, the other simply keeps polling — correct, |
| 447 | but it means a queue can look busy while a runner looks idle. |
| 448 | - **Concurrency.** Nothing bounds how many jobs are in flight beyond how many |
| 449 | runners exist, and nothing stops one runner claiming repeatedly. The |
| 450 | in-process runner's "one job at a time" was a property of the loop, and it is |
| 451 | gone; a `max_concurrent` equivalent for CI does not exist. |
| 452 | - **Concurrency, again.** Platform routing makes a second runner useful, which |
| 453 | makes the missing `max_concurrent` more pressing rather than less. |