collin/anvil
RenderedSource
Remote runners
Status: design (2026-08-24). Supersedes the in-process CI executor in
crates/anvil-ci.
Today anvil is the runner: run_worker (anvil-ci/src/lib.rs:73) drains the
queue in-process and execute (:387) creates the job container directly on
the local Docker socket. The crate header states the model outright — "anvil is
the broker: it is the only Docker client."
That works, and it pins CI to whichever machine anvil runs on. anvil runs on hagrid, a droplet small enough that a release build OOMs it (see DEPLOY.md §3). So CI has to run somewhere else.
A runner is a separate binary that dials out to anvil, claims a job, runs it in a sandboxed container on its own Docker daemon, and posts the result back. anvild stops executing anything.
hagrid (droplet, 1 GB) build host (Mac mini M2)
+------------------------+ +---------------------------+
| anvild | | anvil-worker (launchd) |
| queue, tree, vault | <--------- | long-poll claim |
| artifacts, webhook | https | POST result |
| | ---------> | | |
| NO docker socket | | v |
+------------------------+ | [ job container ] |
| cap-drop, no socket |
+---------------------------+
Why dial-out
The alternative was forwarding the build host's Docker socket to hagrid and pointing anvil at it. Dial-out wins on four counts, in order of importance:
- No inbound path from the public VPS into the home LAN. A forwarded Docker socket is unauthenticated root on the machine that owns it. Punching that from an internet-facing droplet into a home network makes hagrid's compromise the build host's compromise.
- NAT and sleep/wake are free. The build host is a desktop machine behind a residential connection. A dialing client reconnects; a dialed-into server needs a tunnel supervised on both ends.
- The failure mode is legible. A sleeping Mac becomes "no runner
available," not a Docker API error surfacing as
[runner error]in a job log. - Runners become plural. Once anvil addresses runners rather than a socket, a second one is configuration, not architecture — which is how the arm64/amd64 split below gets solved properly.
What this buys hagrid
anvild stops needing Docker at all. deploy/run.sh drops
-v /var/run/docker.sock:/var/run/docker.sock and --group-add, which deletes
the warning in DEPLOY.md §4 about the container holding
root-equivalent control of the host. The internet-facing process stops being a
host-escape vector.
This is contingent on agent sessions staying off — see Scope.
What moves
Stays in anvild | Moves to anvil-worker |
|---|---|
queue, enqueue/requeue_interrupted | execute |
repo + tree resolution, build_tar | collect_artifacts |
parse_pipeline, image allowlist check | store_artifact |
the secret vault (app.vault.take) | download_tar |
step-script assembly, single_quote | parse_meta_tar |
artifact storage, the swap, gc_artifacts | docker.rs |
mask_secrets (applied on receipt) | |
| the deploy webhook |
run_worker becomes run_dispatcher: same queue drain, but instead of calling
execute it parks the job until a runner claims it.
Script assembly stays server-side deliberately. The runner then never parses
.anvil/ci.yml and holds no pipeline model — it receives an image, a script,
sandbox caps, and a tar. That keeps the wire format stable as the pipeline
schema grows.
The crate name
The binary is anvil-worker, not anvil-runner, because
anvil-runner:latest is already the image CI jobs and agent sessions run in
(see agent-sessions.md). Prose says "runner" for the
concept; the crate avoids the collision.
The protocol
Plain HTTP against the existing axum server, under the /-/ system namespace,
so it inherits Caddy's TLS and needs no new listener.
| Endpoint | Method | Purpose |
|---|---|---|
/-/runner/claim | POST | long-poll; 204 on timeout |
/-/runner/jobs/{run_id}/checkout.tar | GET | the uploaded checkout |
/-/runner/jobs/{run_id}/heartbeat | POST | extend the lease |
/-/runner/jobs/{run_id}/artifacts/{name} | POST | one artifact tar |
/-/runner/jobs/{run_id}/result | POST | exit code + log + artifact meta |
The claim response carries everything execute takes as arguments today:
{
"run_id": 42,
"image": "rust:1.95-bookworm",
"platform": "linux/amd64",
"script": "set -e\n...",
"secrets": [{"name": "CARGO_TOKEN", "value": "..."}],
"artifacts": [{"name": "docs", "path": "target/doc", "browse": true}],
"sandbox": {
"memory_mb": 2048, "cpus": 2.0, "pids_limit": 512,
"timeout_secs": 1800, "network": true, "run_as": ""
}
}
The checkout is a separate GET rather than a base64 field, so a large tree
doesn't inflate a JSON body by a third. Secrets ride in the claim body over
TLS, never on a separately-fetchable URL.
Artifacts upload individually before result, for the same reason.
Logs
The runner posts the whole log with result, which is exactly current
behaviour: append_log is called only twice in process — once on the
secrets-failure exit (:161) and once when the run ends (:233). The log
accumulates in memory and lands in one write. The run page is not live today
and does not become less live.
Live logs are a genuine follow-up, and a remote runner makes them easier to justify (there is now a producer that could stream). Out of scope here.
Auth
A shared secret in [ci] runner_token, sent as X-Anvil-Runner-Token,
constant-time compared.
This matches the existing deploy_secret pattern rather than inventing a
credential type. It is deliberate: API tokens are read-only and Bearer-only on
GET/HEAD (see untrusted-mode.md), a runner must POST, and
the write scope is still on TODO.md. Per-runner DB-backed tokens
with last_used_at are the right end state; a single-tenant forge with one
runner does not need them to start.
Leases
Held in memory on the server: run_id -> (runner_name, expires_at), bumped
by heartbeat, swept periodically. An expired lease returns the run to
queued.
In-memory rather than columns on CiRun because Toasty migrations do not exist
yet — DEPLOY.md §Operations and TODO.md both flag that schema changes don't
auto-apply to the live database. Adding claimed_by/lease_expires_at to
CiRun (models.rs:65) would need a manual migration on hagrid.
It also costs nothing: anvild is the only dispatcher, so a lease has no reason
to outlive it, and the anvild-crash case is already handled —
requeue_interrupted (ci.rs:245) re-queues everything left running at
startup. The sweep covers the new case, a runner that dies mid-job.
Architecture
The build host is arm64; hagrid is x86_64. Two separate concerns:
What the shipped image is. Already solved: deploy/build.sh:33 cross-
compiles with zigbuild and builds --platform linux/amd64. Nothing to do,
though note the Dockerfile's RUN apt-get … does execute amd64 binaries
under emulation. Since anvild is a static musl binary,
gcr.io/distroless/static:nonroot would make that build pure COPY and
genuinely emulation-free. Nice-to-have.
What jobs run on. On an M2, rust:1.95-bookworm resolves to arm64, so
cargo test tests an architecture you don't ship. Fix: a platform: key in
.anvil/ci.yml, plumbed to CreateContainerOptions.platform
(bollard 0.18 container.rs:107) and CreateImageOptions.platform
(image.rs:72) — ensure_image currently leaves both at Default::default().
Enable Docker Desktop's Rosetta option; it is far faster than QEMU for amd64
Linux binaries.
Per-pipeline you then choose: native arm64 for lint and unit tests, amd64 under Rosetta for anything arch-sensitive.
Once runners are plural, platform also becomes a scheduling hint — a job
declaring linux/amd64 prefers a runner that is natively amd64. Not needed
with one runner; the field is forward-compatible with it.
Isolation on macOS
Docker on macOS is a Linux VM (LinuxKit under Docker Desktop, Lima under
Colima/OrbStack) and every container runs inside it, so the sandbox is enforced
by the same kernel primitives as on Linux: --cap-drop=ALL and
no-new-privileges are capabilities and prctl, the pids/memory/cpu caps are
cgroups v2.
It is a boundary better than hagrid's. untrusted-mode.md §1 notes that a
kernel or runc escape defeats the sandbox; on the Mac that escape reaches a
disposable Linux VM, not the host.
Two things to get right:
Docker::connect_with_socket_defaults()(docker.rs:15) will not find the socket. It hardcodes/var/run/docker.sock; Docker Desktop only creates that symlink when "Allow the default Docker socket to be used" is ticked, and the real path is~/.docker/run/docker.sock(Colima differs again). The runner usesDocker::connect_with_defaults(), which honoursDOCKER_HOST.- Run the runner natively under launchd, not in a container. Containerizing it means mounting the socket into it, rebuilding the root-equivalent hole this change removes from hagrid.
Size the VM's RAM deliberately: per-job memory_mb is carved out of a fixed
allocation. The dispatcher runs one job at a time, so this is slack rather than
a constraint.
Two processes on the build host
Image building cannot be a CI job. Job containers get no Docker socket by
design (execute's doc comment at :369, and untrusted-mode.md §1), and
that invariant is the whole broker model. So the build host runs two things at
two trust levels:
anvil-worker— claims jobs, runs them sandboxed, no Docker access inside the job container.- a deploy agent — fired on green CI, runs outside any sandbox with full
Docker access, does
docker build --platform linux/amd64anddocker pushtoregistry.vibe.richardscollin.com, then triggers hagrid to pull.
Keeping these separate is what preserves the broker model. Folding the second into the first would give pipeline authors a path to the daemon.
The deploy agent is out of scope for M1 — the existing deploy_webhook still
works, with the receiver moved to the build host and reached over Tailscale.
M2 folds it into the same dial-out channel as a privileged "publish" job kind,
authorized server-side by the is_deploy_target check that already scopes CD
to exactly one repository (config.rs:357). That removes the last inbound
requirement.
Scope
Agent sessions stay local-socket and stay out. anvil-agent calls
anvil_ci::docker::connect() in five places (lib.rs:141,
supervisor.rs:55,278,414,439) and docker.rs is explicitly shared plumbing,
so the crate split touches them — but a session is interactive (tmux attach,
exec streaming, resize, pump_transcript), which is a far harder protocol than
fire-and-forget CI. They are enabled = false and absent from
deploy/anvil.toml entirely, so nothing in production regresses.
ensure_image and connect therefore need a home both crates can reach. They
move to a thin anvil-docker crate rather than being duplicated —
ensure_image's local-fallback pull logic is subtle enough to be worth having
once.
The corollary: turning agent sessions on for hagrid later means either putting the socket back, or moving sessions onto the runner protocol too.
Registry cleanup
ensure_image's tolerance of a failed pull exists because anvil-runner:latest
"exists in no registry" (docker.rs:22-24) and is built straight into the local
daemon store. With registry.vibe.richardscollin.com up, push the image there
and the fallback stops being load-bearing.
It should stay in the code regardless — a runner on a fresh machine wants a clear failure when the pull fails and nothing is cached — but the comment explaining why it exists needs rewriting.
Milestones
M1 — the split. anvil-docker and anvil-worker crates, the five
endpoints, in-memory leases, [ci] runner_token, connect_with_defaults,
run_worker → run_dispatcher, socket mount dropped from deploy/run.sh.
Deploys keep using the existing webhook.
M2 — platform. platform: in .anvil/ci.yml, plumbed through both bollard
options. Runner advertises its native platform at claim time.
M3 — publish jobs. The deploy agent folds into the dial-out channel,
authorized by is_deploy_target. No inbound path to the build host remains.
Not yet done
- Live logs. Newly worth doing, still not done. The result POST is a single write; streaming needs chunked append with offsets and a UI that tolerates gaps.
- Per-runner credentials. One shared secret means one revocation for all
runners, and no
last_used_at. Wants the API-token write scope first. - Runner labels.
platformis the only scheduling dimension in M2. Tags ("has-postgres", "big-memory") are the obvious next axis and are not designed. - Secrets on the runner host. They now cross the network and sit in
plaintext in a container on a machine anvil does not own. This needs a
paragraph in
untrusted-mode.md§1 — the exposure is no longer bounded by hagrid. - Concurrency. The dispatcher still hands out one job at a time, inherited
from
run_worker. Multiple runners make that the binding constraint rather than a sensible default.