anvilsign in

collin/anvil

RenderedSource

Remote runners

Status: M1 + M2 implemented (2026-08-24). Supersedes the in-process CI executor that used to live in crates/anvil-ci.

anvil used to be the runner: run_worker drained the queue in-process and execute created the job container on the local Docker socket. That works, and it pins CI to whichever machine anvil runs on — which is hagrid, a droplet small enough that a release build OOMs it (see DEPLOY.md §3). So CI had to run somewhere else.

A runner is a separate binary that dials out to anvil, claims a job, runs it in a sandboxed container on its own Docker daemon, and posts the result back. anvild stops executing anything.

   hagrid (droplet, 1 GB)                 build host (Mac mini M2)
  +------------------------+            +---------------------------+
  |  anvild                |            |  anvil-worker (launchd)   |
  |   queue, tree, vault   | <--------- |   long-poll claim         |
  |   artifacts, webhook   |  https     |   POST result             |
  |                        | ---------> |        |                  |
  |  NO docker socket      |            |        v                  |
  +------------------------+            |   [ job container ]       |
                                        |   cap-drop, no socket     |
                                        +---------------------------+

Why dial-out

The alternative was forwarding the build host's Docker socket to hagrid and pointing anvil at it. Dial-out wins on four counts, in order of importance:

  1. No inbound path from the public VPS into the home LAN. A forwarded Docker socket is unauthenticated root on the machine that owns it. Punching that from an internet-facing droplet into a home network makes hagrid's compromise the build host's compromise.
  2. NAT and sleep/wake are free. The build host is a desktop machine behind a residential connection. A dialing client reconnects; a dialed-into server needs a tunnel supervised on both ends.
  3. The failure mode is legible. A sleeping Mac becomes "no runner available," not a Docker API error surfacing as [runner error] in a job log.
  4. Runners become plural. Once anvil addresses runners rather than a socket, a second one is configuration, not architecture — which is how the arm64/amd64 split below gets solved properly.

What this buys hagrid

anvild stops needing Docker at all. compose.yaml carries no /var/run/docker.sock mount and no group_add, which deletes the warning in DEPLOY.md §4 about the container holding root-equivalent control of the host. The internet-facing process stops being a host-escape vector.

This is contingent on agent sessions staying off — see Scope.

What moves

Stays in anvildMoved to anvil-worker
queue, enqueue/requeue_interruptedexecute
repo + tree resolution, build_tarcollect_artifacts
parse_pipeline, image allowlist checkdownload_tar
the secret vault (app.vault.take)parse_meta_tar
step-script assembly, single_quotethe ArtifactSink trait
store_artifact, the swap, gc_artifacts
mask_secrets (applied on receipt)
the deploy webhook

store_artifact deliberately stayed behind. The runner uploads the raw tar it pulled out of the container and anvil decides what to do with it, so on-disk layout — and browse, which turns a tarball into a servable directory tree — never becomes a runner's call. The upload response reports the stored size, so the per-run artifact budget is still charged what actually landed.

run_worker becomes run_dispatcher: same queue drain, but instead of calling execute it parks the job until a runner claims it.

Script assembly stays server-side deliberately. The runner then never parses .anvil/ci.toml and holds no pipeline model — it receives an image, a script, sandbox caps, and a tar. That keeps the wire format stable as the pipeline schema grows.

The crates

CrateHolds
anvil-jobthe wire format, and nothing else — serde only
anvil-dockerconnect/ensure_image, shared with anvil-agent
anvil-workerthe runner binary: claim loop, client, executor

anvil-job exists so the runner does not link anvil-core — and therefore toasty, SQLite, gix and the rest of the forge — just to learn the shape of a job. anvil-docker exists so anvil-agent need not depend on the runner.

The binary is anvil-worker, not anvil-runner, because anvil-runner:latest is already the image CI jobs and agent sessions run in (see agent-sessions.md). Prose says "runner" for the concept; the crate avoids the collision.

The protocol

Plain HTTP against the existing axum server, under the /-/ system namespace, so it inherits Caddy's TLS and needs no new listener.

EndpointMethodPurpose
/-/runner/claimPOSTlong-poll; 204 on timeout
/-/runner/jobs/{run_id}/checkout.tarGETthe uploaded checkout
/-/runner/jobs/{run_id}/heartbeatPOSTextend the lease
/-/runner/jobs/{run_id}/artifacts/{name}POSTone artifact tar
/-/runner/jobs/{run_id}/resultPOSTexit code + log + artifact meta

The claim response carries everything execute takes as arguments today:

{
  "run_id": 42,
  "image": "anvil-runner:rust",
  "platform": "linux/amd64",
  "script": "set -e\n...",
  "secrets": [{"name": "CARGO_TOKEN", "value": "..."}],
  "artifacts": [{"name": "docs", "path": "target/doc", "browse": true}],
  "sandbox": {
    "memory_mb": 2048, "cpus": 2.0, "pids_limit": 512,
    "timeout_secs": 1800, "network": true, "run_as": ""
  }
}

The checkout is a separate GET rather than a base64 field, so a large tree doesn't inflate a JSON body by a third. Secrets ride in the claim body over TLS, never on a separately-fetchable URL.

Artifacts upload individually before result, for the same reason.

Logs

The runner posts the whole log with result, which is exactly what the in-process runner did: append_log was only ever called when a run ended (and on the secrets-failure exit). The log accumulated in memory and landed in one write. The run page was not live before and is no less live now — a header naming the runner is now written at claim time, so a running run at least shows something.

Live logs are a genuine follow-up, and a remote runner makes them easier to justify (there is now a producer that could stream). Out of scope here.

Auth

A shared secret in [ci] runner_token, sent as X-Anvil-Runner-Token, constant-time compared. The runner's self-asserted name rides alongside in X-Anvil-Runner-Name — it labels runs and keys leases, and is explicitly not a credential: everyone holding the token is one principal.

The four per-job endpoints additionally require the caller to hold that run's lease, so a valid token gets you a job rather than everyone else's. A lease mismatch answers 409, not 403: the caller is a legitimate runner whose claim simply expired.

This matches the existing deploy_secret pattern rather than inventing a credential type. It is deliberate: API tokens are read-only and Bearer-only on GET/HEAD (see untrusted-mode.md), a runner must POST, and the write scope is still on TODO.md. Per-runner DB-backed tokens with last_used_at are the right end state; a single-tenant forge with one runner does not need them to start.

Leases

Held in memory on the server (anvil_core::jobs::Dispatch): run_id -> (runner_name, expires_at, secrets), bumped by heartbeat every 30s against a 120s TTL, swept every 30s. An expired lease returns the run to queued.

A job's secrets are stashed on its lease rather than re-read from the vault when the result lands. Vault::take fails once the repository's unlock TTL lapses, and a job can easily outlive an unlock — re-reading would mean a long run silently skips log masking, which is exactly the run whose log is most likely to contain something. The values are already in this process's vault, so this is not new exposure.

The honest gap: an expired lease can double-run a job whose runner is alive but unreachable. The container keeps going and the requeued run may be claimed elsewhere. That is what a lease without fencing buys; CI steps are assumed idempotent.

In-memory rather than columns on CiRun because Toasty migrations do not exist yet — DEPLOY.md §Operations and TODO.md both flag that schema changes don't auto-apply to the live database. Adding claimed_by/lease_expires_at to CiRun (models.rs:65) would need a manual migration on hagrid.

It also costs nothing: anvild is the only dispatcher, so a lease has no reason to outlive it, and the anvild-crash case is already handled — requeue_interrupted (ci.rs:245) re-queues everything left running at startup. The sweep covers the new case, a runner that dies mid-job.

Architecture

The build host is arm64; hagrid is x86_64. Two separate concerns:

What the shipped image is. Already solved: the builder stage in docker/Dockerfile cross-compiles with zigbuild on the build host's own architecture, and both it and compose.yaml pin the runtime to amd64. Nothing to do, though note the runtime stages' RUN apt-get … does execute amd64 binaries under emulation. Since anvild is a static musl binary, gcr.io/distroless/static:nonroot would make that build pure COPY and genuinely emulation-free. Nice-to-have.

What jobs run on. On an M2, a multi-arch image resolves to arm64, so cargo test tests an architecture you don't ship. That is what M2 fixes.

A job's platform comes from platform in .anvil/ci.toml, falling back to [ci] platform, falling back to the claiming runner's native architecture — so an instance that sets neither behaves exactly as it did before. It reaches CreateContainerOptions.platform and CreateImageOptions.platform, and the value is validated as os/arch[/variant] at parse time (a bare amd64 would otherwise reach Docker as an operating system named amd64).

Per-pipeline you then choose: native arm64 for lint and unit tests, amd64 under Rosetta for anything arch-sensitive. Turn Rosetta on in Docker Desktop; it is far faster than QEMU for amd64 Linux binaries.

Routing

platform is also the scheduling dimension. Every claim and heartbeat records the runner's advertised architecture in Dispatch (jobs.rs), which expires after RUNNER_TTL — 5 minutes, longer than both the 55s claim poll and the 30s heartbeat, so an idle runner and a runner mid-build both stay visible. A claiming runner is then offered, in queue order:

  1. runs that name its platform, or name none at all;
  2. then runs whose platform no currently connected runner is native to.

Tier 2 is what keeps one arm64 Mac usable as the only runner for pipelines that declare linux/amd64: nobody can run them natively, so it emulates them rather than leaving them queued forever. Add an amd64 runner and the Mac stops taking those jobs the moment the new runner's first claim registers it — no configuration, which is the property the dial-out model was chosen for.

The corollary worth stating plainly: platform is not a promise of native execution. It is a promise about what the job runs, which the runner enforces by inspecting the image it ended up with (anvil-docker::check_platform) and failing the job if the architecture is not the one asked for. Where it runs is a scheduling preference. A run's log header says which it got:

platform: linux/amd64 (emulated on linux/arm64)

Seeing who is connected

/-/admin/runners (admin-only, 404 for everyone else) lists every runner Dispatch still counts as present: name, advertised platform, worker version, how long since it last spoke, how long it has been connected, and the runs it holds right now, each linked to its CI page. Underneath it is the same map routing reads, so the page and the dispatcher can never disagree about who is out there. It also prints the queue depth, because the two together are the whole diagnosis: runners and no queue is a healthy idle instance, a queue and no runners is a stuck one, and a queue with every runner busy is neither — it is capacity.

There is no separate heartbeat for presence, on purpose. Liveness is the traffic a working runner already generates: an idle one re-registers itself every time its parked claim expires and it dials back in (CLAIM_POLL, 55s), and a busy one every HEARTBEAT_INTERVAL (30s) for as long as its job runs. Between them there is no state a runner can be in where it is useful and silent, so a dedicated ping would only add a way for a runner to look alive while claiming nothing.

What that costs is resolution, and the page is explicit about it rather than hiding it. A runner is shown "late" once it has been quiet for a whole claim poll plus a heartbeat (85s) — long enough that a runner merely parked in a long poll never reads as late. It stays listed, and keeps being routed to, until RUNNER_TTL (5 min), because that is exactly what the dispatcher still believes; the page's job is to show that belief, not to invent a second one. So a machine that loses power disappears from routing in up to five minutes and reads as late within ninety seconds. The page reloads itself every 15s.

Isolation on macOS

Docker on macOS is a Linux VM (LinuxKit under Docker Desktop, Lima under Colima/OrbStack) and every container runs inside it, so the sandbox is enforced by the same kernel primitives as on Linux: --cap-drop=ALL and no-new-privileges are capabilities and prctl, the pids/memory/cpu caps are cgroups v2.

It is a boundary better than hagrid's. untrusted-mode.md §1 notes that a kernel or runc escape defeats the sandbox; on the Mac that escape reaches a disposable Linux VM, not the host.

Two things to get right:

  • Docker::connect_with_socket_defaults() (docker.rs:15) will not find the socket. It hardcodes /var/run/docker.sock; Docker Desktop only creates that symlink when "Allow the default Docker socket to be used" is ticked, and the real path is ~/.docker/run/docker.sock (Colima differs again). The runner uses Docker::connect_with_defaults(), which honours DOCKER_HOST.
  • Run the runner natively under launchd, not in a container. Containerizing it means mounting the socket into it, rebuilding the root-equivalent hole this change removes from hagrid.

Size the VM's RAM deliberately: per-job memory_mb is carved out of a fixed allocation. The dispatcher runs one job at a time, so this is slack rather than a constraint.

Running one on a Mac mini

# On the Mac, in a checkout of anvil:
cargo build --release -p anvil-worker
sudo cp target/release/anvil-worker /usr/local/bin/

# Docker Desktop → Settings → General:
#   ✓ Use Rosetta for x86_64/amd64 emulation on Apple Silicon
# Without it, a linux/amd64 job runs under QEMU — correct, and much slower.

cp deploy/worker/com.anvil.worker.plist ~/Library/LaunchAgents/
# Fill in --url, --name, ANVIL_RUNNER_TOKEN and DOCKER_HOST, then:
chmod 600 ~/Library/LaunchAgents/com.anvil.worker.plist
launchctl load -w ~/Library/LaunchAgents/com.anvil.worker.plist
tail -f /tmp/anvil-worker.log     # "anvil-worker macmini (linux/arm64) → …"

The forge side needs [ci] runner_token set to the same secret; until it is, every claim gets a 503 saying so and queued runs sit.

The plist is a LaunchAgent, not a LaunchDaemon, because Docker Desktop's socket only exists inside the logged-in user's session — a root daemon starts before Docker and never finds it. The consequence to know about: the runner is only up while that user is logged in, and a Mac that sleeps stops claiming. That is the failure mode the dial-out design chose (a sleeping Mac reads as "no runner available"), and after RUNNER_TTL its architecture stops counting as present, so anything routed to it falls back to another runner.

--name matters: the default reads $HOSTNAME, which launchd does not set, so an unnamed runner is called runner.

Two runners locally

compose.override.yaml brings up runner-1 and runner-2 next to the local forge, so the parts of this design that only appear with more than one runner — concurrent pipelines, and the "no connected runner is native to this platform" half of routing — are testable without a second machine:

docker compose up -d --build
docker compose logs -f runner-1 runner-2
# anvil-worker dev-1 (linux/amd64) → http://anvil:3000

They reach the forge as http://anvil:3000 over the compose network (the browser-facing base_url does not resolve inside a container) and authenticate with the runner_token committed in deploy/anvil.dev.toml. Both are amd64 here, so to watch the fallback tier work, give a pipeline platform: linux/arm64 and confirm it still gets claimed. Adding a third is a copy of the four-line service block with a new name.

These runners are containerized, and that is the one thing production must never copy: the socket mount is the root-equivalent hold that moving CI off the forge removed. It is acceptable locally only because the same file already mounts that socket into anvil for agent sessions, so the machine's trust boundary is unchanged. The deployed compose.yaml grants neither.

Two processes on the build host

Image building cannot be a CI job. Job containers get no Docker socket by design (execute's doc comment at :369, and untrusted-mode.md §1), and that invariant is the whole broker model. So the build host runs two things at two trust levels:

  1. anvil-worker — claims jobs, runs them sandboxed, no Docker access inside the job container.
  2. a deploy agent — fired on green CI, runs outside any sandbox with full Docker access, does docker build --platform linux/amd64 and docker push to registry.vibe.richardscollin.com, then triggers hagrid to pull.

Keeping these separate is what preserves the broker model. Folding the second into the first would give pipeline authors a path to the daemon.

The deploy agent is out of scope for M1 — the existing deploy_webhook still works, with the receiver moved to the build host and reached over Tailscale. M2 folds it into the same dial-out channel as a privileged "publish" job kind, authorized server-side by the is_deploy_target check that already scopes CD to exactly one repository (config.rs:357). That removes the last inbound requirement.

Scope

Agent sessions stay local-socket and stay out. anvil-agent calls anvil_ci::docker::connect() in five places (lib.rs:141, supervisor.rs:55,278,414,439) and docker.rs is explicitly shared plumbing, so the crate split touches them — but a session is interactive (tmux attach, exec streaming, resize, pump_transcript), which is a far harder protocol than fire-and-forget CI. They are enabled = false and absent from deploy/anvil.toml entirely, so nothing in production regresses.

ensure_image and connect therefore need a home both crates can reach. They move to a thin anvil-docker crate rather than being duplicated — ensure_image's local-fallback pull logic is subtle enough to be worth having once.

The corollary: turning agent sessions on for hagrid later means either putting the socket back, or moving sessions onto the runner protocol too.

Registry cleanup

ensure_image's tolerance of a failed pull exists because anvil-runner:latest "exists in no registry" (docker.rs:22-24) and is built straight into the local daemon store. With registry.vibe.richardscollin.com up, push the image there and the fallback stops being load-bearing.

It should stay in the code regardless — a runner on a fresh machine wants a clear failure when the pull fails and nothing is cached — but the comment explaining why it exists needs rewriting.

Milestones

M1 — the split. Done. anvil-job, anvil-docker and anvil-worker crates, the five endpoints, in-memory leases, [ci] runner_token, connect_with_defaults, run_worker → run_dispatcher, and no socket mount in the deployed container. Deploys keep using the existing webhook.

M2 — platform. Done. platform in .anvil/ci.toml, a [ci] platform default, CiConfig::resolve_platform, the two-tier routing above (backed by a runner registry in Dispatch), the emulation note in the run header, and an architecture check on the image the runner actually got.

M3 — publish jobs. The deploy agent folds into the dial-out channel, authorized by is_deploy_target. No inbound path to the build host remains.

Not yet done

  • Live logs. Newly worth doing, still not done. The result POST is a single write; streaming needs chunked append with offsets and a UI that tolerates gaps.
  • Per-runner credentials. One shared secret means one revocation for all runners, and no last_used_at. Wants the API-token write scope first.
  • Runner labels. platform is the only scheduling dimension in M2. Tags ("has-postgres", "big-memory") are the obvious next axis and are not designed.
  • A runners page. Dispatch::runners() now knows every runner connected in the last five minutes and what it is. Nothing renders it, so "is my Mac actually claiming?" is still answered by reading logs.
  • Routing is per-claim, not per-queue. A runner that can take nothing sleeps until the next wake; it does not reserve the job it declined. With two runners and a job only one can run natively, the other simply keeps polling — correct, but it means a queue can look busy while a runner looks idle.
  • Concurrency. Nothing bounds how many jobs are in flight beyond how many runners exist, and nothing stops one runner claiming repeatedly. The in-process runner's "one job at a time" was a property of the loop, and it is gone; a max_concurrent equivalent for CI does not exist.
  • Concurrency, again. Platform routing makes a second runner useful, which makes the missing max_concurrent more pressing rather than less.