anvilsign in

collin/anvil

RenderedSource

Remote runners

Status: M1 implemented (2026-08-24). Supersedes the in-process CI executor that used to live in crates/anvil-ci.

anvil used to be the runner: run_worker drained the queue in-process and execute created the job container on the local Docker socket. That works, and it pins CI to whichever machine anvil runs on — which is hagrid, a droplet small enough that a release build OOMs it (see DEPLOY.md §3). So CI had to run somewhere else.

A runner is a separate binary that dials out to anvil, claims a job, runs it in a sandboxed container on its own Docker daemon, and posts the result back. anvild stops executing anything.

   hagrid (droplet, 1 GB)                 build host (Mac mini M2)
  +------------------------+            +---------------------------+
  |  anvild                |            |  anvil-worker (launchd)   |
  |   queue, tree, vault   | <--------- |   long-poll claim         |
  |   artifacts, webhook   |  https     |   POST result             |
  |                        | ---------> |        |                  |
  |  NO docker socket      |            |        v                  |
  +------------------------+            |   [ job container ]       |
                                        |   cap-drop, no socket     |
                                        +---------------------------+

Why dial-out

The alternative was forwarding the build host's Docker socket to hagrid and pointing anvil at it. Dial-out wins on four counts, in order of importance:

  1. No inbound path from the public VPS into the home LAN. A forwarded Docker socket is unauthenticated root on the machine that owns it. Punching that from an internet-facing droplet into a home network makes hagrid's compromise the build host's compromise.
  2. NAT and sleep/wake are free. The build host is a desktop machine behind a residential connection. A dialing client reconnects; a dialed-into server needs a tunnel supervised on both ends.
  3. The failure mode is legible. A sleeping Mac becomes "no runner available," not a Docker API error surfacing as [runner error] in a job log.
  4. Runners become plural. Once anvil addresses runners rather than a socket, a second one is configuration, not architecture — which is how the arm64/amd64 split below gets solved properly.

What this buys hagrid

anvild stops needing Docker at all. deploy/run.sh drops -v /var/run/docker.sock:/var/run/docker.sock and --group-add, which deletes the warning in DEPLOY.md §4 about the container holding root-equivalent control of the host. The internet-facing process stops being a host-escape vector.

This is contingent on agent sessions staying off — see Scope.

What moves

Stays in anvildMoved to anvil-worker
queue, enqueue/requeue_interruptedexecute
repo + tree resolution, build_tarcollect_artifacts
parse_pipeline, image allowlist checkdownload_tar
the secret vault (app.vault.take)parse_meta_tar
step-script assembly, single_quotethe ArtifactSink trait
store_artifact, the swap, gc_artifacts
mask_secrets (applied on receipt)
the deploy webhook

store_artifact deliberately stayed behind. The runner uploads the raw tar it pulled out of the container and anvil decides what to do with it, so on-disk layout — and browse, which turns a tarball into a servable directory tree — never becomes a runner's call. The upload response reports the stored size, so the per-run artifact budget is still charged what actually landed.

run_worker becomes run_dispatcher: same queue drain, but instead of calling execute it parks the job until a runner claims it.

Script assembly stays server-side deliberately. The runner then never parses .anvil/ci.yml and holds no pipeline model — it receives an image, a script, sandbox caps, and a tar. That keeps the wire format stable as the pipeline schema grows.

The crates

CrateHolds
anvil-jobthe wire format, and nothing else — serde only
anvil-dockerconnect/ensure_image, shared with anvil-agent
anvil-workerthe runner binary: claim loop, client, executor

anvil-job exists so the runner does not link anvil-core — and therefore toasty, SQLite, gix and the rest of the forge — just to learn the shape of a job. anvil-docker exists so anvil-agent need not depend on the runner.

The binary is anvil-worker, not anvil-runner, because anvil-runner:latest is already the image CI jobs and agent sessions run in (see agent-sessions.md). Prose says "runner" for the concept; the crate avoids the collision.

The protocol

Plain HTTP against the existing axum server, under the /-/ system namespace, so it inherits Caddy's TLS and needs no new listener.

EndpointMethodPurpose
/-/runner/claimPOSTlong-poll; 204 on timeout
/-/runner/jobs/{run_id}/checkout.tarGETthe uploaded checkout
/-/runner/jobs/{run_id}/heartbeatPOSTextend the lease
/-/runner/jobs/{run_id}/artifacts/{name}POSTone artifact tar
/-/runner/jobs/{run_id}/resultPOSTexit code + log + artifact meta

The claim response carries everything execute takes as arguments today:

{
  "run_id": 42,
  "image": "rust:1.95-bookworm",
  "platform": "linux/amd64",
  "script": "set -e\n...",
  "secrets": [{"name": "CARGO_TOKEN", "value": "..."}],
  "artifacts": [{"name": "docs", "path": "target/doc", "browse": true}],
  "sandbox": {
    "memory_mb": 2048, "cpus": 2.0, "pids_limit": 512,
    "timeout_secs": 1800, "network": true, "run_as": ""
  }
}

The checkout is a separate GET rather than a base64 field, so a large tree doesn't inflate a JSON body by a third. Secrets ride in the claim body over TLS, never on a separately-fetchable URL.

Artifacts upload individually before result, for the same reason.

Logs

The runner posts the whole log with result, which is exactly what the in-process runner did: append_log was only ever called when a run ended (and on the secrets-failure exit). The log accumulated in memory and landed in one write. The run page was not live before and is no less live now — a header naming the runner is now written at claim time, so a running run at least shows something.

Live logs are a genuine follow-up, and a remote runner makes them easier to justify (there is now a producer that could stream). Out of scope here.

Auth

A shared secret in [ci] runner_token, sent as X-Anvil-Runner-Token, constant-time compared. The runner's self-asserted name rides alongside in X-Anvil-Runner-Name — it labels runs and keys leases, and is explicitly not a credential: everyone holding the token is one principal.

The four per-job endpoints additionally require the caller to hold that run's lease, so a valid token gets you a job rather than everyone else's. A lease mismatch answers 409, not 403: the caller is a legitimate runner whose claim simply expired.

This matches the existing deploy_secret pattern rather than inventing a credential type. It is deliberate: API tokens are read-only and Bearer-only on GET/HEAD (see untrusted-mode.md), a runner must POST, and the write scope is still on TODO.md. Per-runner DB-backed tokens with last_used_at are the right end state; a single-tenant forge with one runner does not need them to start.

Leases

Held in memory on the server (anvil_core::jobs::Dispatch): run_id -> (runner_name, expires_at, secrets), bumped by heartbeat every 30s against a 120s TTL, swept every 30s. An expired lease returns the run to queued.

A job's secrets are stashed on its lease rather than re-read from the vault when the result lands. Vault::take fails once the repository's unlock TTL lapses, and a job can easily outlive an unlock — re-reading would mean a long run silently skips log masking, which is exactly the run whose log is most likely to contain something. The values are already in this process's vault, so this is not new exposure.

The honest gap: an expired lease can double-run a job whose runner is alive but unreachable. The container keeps going and the requeued run may be claimed elsewhere. That is what a lease without fencing buys; CI steps are assumed idempotent.

In-memory rather than columns on CiRun because Toasty migrations do not exist yet — DEPLOY.md §Operations and TODO.md both flag that schema changes don't auto-apply to the live database. Adding claimed_by/lease_expires_at to CiRun (models.rs:65) would need a manual migration on hagrid.

It also costs nothing: anvild is the only dispatcher, so a lease has no reason to outlive it, and the anvild-crash case is already handled — requeue_interrupted (ci.rs:245) re-queues everything left running at startup. The sweep covers the new case, a runner that dies mid-job.

Architecture

The build host is arm64; hagrid is x86_64. Two separate concerns:

What the shipped image is. Already solved: deploy/build.sh:33 cross- compiles with zigbuild and builds --platform linux/amd64. Nothing to do, though note the Dockerfile's RUN apt-get … does execute amd64 binaries under emulation. Since anvild is a static musl binary, gcr.io/distroless/static:nonroot would make that build pure COPY and genuinely emulation-free. Nice-to-have.

What jobs run on. On an M2, rust:1.95-bookworm resolves to arm64, so cargo test tests an architecture you don't ship. Fix: a platform: key in .anvil/ci.yml, plumbed to CreateContainerOptions.platform (bollard 0.18 container.rs:107) and CreateImageOptions.platform (image.rs:72) — ensure_image currently leaves both at Default::default(). Enable Docker Desktop's Rosetta option; it is far faster than QEMU for amd64 Linux binaries.

Per-pipeline you then choose: native arm64 for lint and unit tests, amd64 under Rosetta for anything arch-sensitive.

Once runners are plural, platform also becomes a scheduling hint — a job declaring linux/amd64 prefers a runner that is natively amd64. Not needed with one runner; the field is forward-compatible with it.

Isolation on macOS

Docker on macOS is a Linux VM (LinuxKit under Docker Desktop, Lima under Colima/OrbStack) and every container runs inside it, so the sandbox is enforced by the same kernel primitives as on Linux: --cap-drop=ALL and no-new-privileges are capabilities and prctl, the pids/memory/cpu caps are cgroups v2.

It is a boundary better than hagrid's. untrusted-mode.md §1 notes that a kernel or runc escape defeats the sandbox; on the Mac that escape reaches a disposable Linux VM, not the host.

Two things to get right:

  • Docker::connect_with_socket_defaults() (docker.rs:15) will not find the socket. It hardcodes /var/run/docker.sock; Docker Desktop only creates that symlink when "Allow the default Docker socket to be used" is ticked, and the real path is ~/.docker/run/docker.sock (Colima differs again). The runner uses Docker::connect_with_defaults(), which honours DOCKER_HOST.
  • Run the runner natively under launchd, not in a container. Containerizing it means mounting the socket into it, rebuilding the root-equivalent hole this change removes from hagrid.

Size the VM's RAM deliberately: per-job memory_mb is carved out of a fixed allocation. The dispatcher runs one job at a time, so this is slack rather than a constraint.

Two processes on the build host

Image building cannot be a CI job. Job containers get no Docker socket by design (execute's doc comment at :369, and untrusted-mode.md §1), and that invariant is the whole broker model. So the build host runs two things at two trust levels:

  1. anvil-worker — claims jobs, runs them sandboxed, no Docker access inside the job container.
  2. a deploy agent — fired on green CI, runs outside any sandbox with full Docker access, does docker build --platform linux/amd64 and docker push to registry.vibe.richardscollin.com, then triggers hagrid to pull.

Keeping these separate is what preserves the broker model. Folding the second into the first would give pipeline authors a path to the daemon.

The deploy agent is out of scope for M1 — the existing deploy_webhook still works, with the receiver moved to the build host and reached over Tailscale. M2 folds it into the same dial-out channel as a privileged "publish" job kind, authorized server-side by the is_deploy_target check that already scopes CD to exactly one repository (config.rs:357). That removes the last inbound requirement.

Scope

Agent sessions stay local-socket and stay out. anvil-agent calls anvil_ci::docker::connect() in five places (lib.rs:141, supervisor.rs:55,278,414,439) and docker.rs is explicitly shared plumbing, so the crate split touches them — but a session is interactive (tmux attach, exec streaming, resize, pump_transcript), which is a far harder protocol than fire-and-forget CI. They are enabled = false and absent from deploy/anvil.toml entirely, so nothing in production regresses.

ensure_image and connect therefore need a home both crates can reach. They move to a thin anvil-docker crate rather than being duplicated — ensure_image's local-fallback pull logic is subtle enough to be worth having once.

The corollary: turning agent sessions on for hagrid later means either putting the socket back, or moving sessions onto the runner protocol too.

Registry cleanup

ensure_image's tolerance of a failed pull exists because anvil-runner:latest "exists in no registry" (docker.rs:22-24) and is built straight into the local daemon store. With registry.vibe.richardscollin.com up, push the image there and the fallback stops being load-bearing.

It should stay in the code regardless — a runner on a fresh machine wants a clear failure when the pull fails and nothing is cached — but the comment explaining why it exists needs rewriting.

Milestones

M1 — the split. Done. anvil-job, anvil-docker and anvil-worker crates, the five endpoints, in-memory leases, [ci] runner_token, connect_with_defaults, run_worker → run_dispatcher, socket mount dropped from deploy/run.sh. Deploys keep using the existing webhook.

M2 — platform. The plumbing is already live: JobSpec.platform reaches CreateContainerOptions and CreateImageOptions, and a runner advertises its native platform when it claims. What is missing is a source — a platform: key in .anvil/ci.yml (and probably a [ci] platform default), plus routing a job to a runner that has that architecture natively.

M3 — publish jobs. The deploy agent folds into the dial-out channel, authorized by is_deploy_target. No inbound path to the build host remains.

Not yet done

  • Live logs. Newly worth doing, still not done. The result POST is a single write; streaming needs chunked append with offsets and a UI that tolerates gaps.
  • Per-runner credentials. One shared secret means one revocation for all runners, and no last_used_at. Wants the API-token write scope first.
  • Runner labels. platform is the only scheduling dimension in M2. Tags ("has-postgres", "big-memory") are the obvious next axis and are not designed.
  • Secrets on the runner host. They now cross the network and sit in plaintext in a container on a machine anvil does not own. This needs a paragraph in untrusted-mode.md §1 — the exposure is no longer bounded by hagrid.
  • Concurrency. Nothing bounds how many jobs are in flight beyond how many runners exist, and nothing stops one runner claiming repeatedly. The in-process runner's "one job at a time" was a property of the loop, and it is gone; a max_concurrent equivalent for CI does not exist.
  • ensure_image's local fallback vs platform. A cached image of the wrong architecture satisfies the fallback, since it inspects presence and not arch. Only bites an offline runner asked to cross-build.