collin/anvil
RenderedSource
Remote runners
Status: M1 implemented (2026-08-24). Supersedes the in-process CI
executor that used to live in crates/anvil-ci.
anvil used to be the runner: run_worker drained the queue in-process and
execute created the job container on the local Docker socket. That works, and
it pins CI to whichever machine anvil runs on — which is hagrid, a droplet
small enough that a release build OOMs it (see DEPLOY.md §3).
So CI had to run somewhere else.
A runner is a separate binary that dials out to anvil, claims a job, runs it in a sandboxed container on its own Docker daemon, and posts the result back. anvild stops executing anything.
hagrid (droplet, 1 GB) build host (Mac mini M2)
+------------------------+ +---------------------------+
| anvild | | anvil-worker (launchd) |
| queue, tree, vault | <--------- | long-poll claim |
| artifacts, webhook | https | POST result |
| | ---------> | | |
| NO docker socket | | v |
+------------------------+ | [ job container ] |
| cap-drop, no socket |
+---------------------------+
Why dial-out
The alternative was forwarding the build host's Docker socket to hagrid and pointing anvil at it. Dial-out wins on four counts, in order of importance:
- No inbound path from the public VPS into the home LAN. A forwarded Docker socket is unauthenticated root on the machine that owns it. Punching that from an internet-facing droplet into a home network makes hagrid's compromise the build host's compromise.
- NAT and sleep/wake are free. The build host is a desktop machine behind a residential connection. A dialing client reconnects; a dialed-into server needs a tunnel supervised on both ends.
- The failure mode is legible. A sleeping Mac becomes "no runner
available," not a Docker API error surfacing as
[runner error]in a job log. - Runners become plural. Once anvil addresses runners rather than a socket, a second one is configuration, not architecture — which is how the arm64/amd64 split below gets solved properly.
What this buys hagrid
anvild stops needing Docker at all. deploy/run.sh drops
-v /var/run/docker.sock:/var/run/docker.sock and --group-add, which deletes
the warning in DEPLOY.md §4 about the container holding
root-equivalent control of the host. The internet-facing process stops being a
host-escape vector.
This is contingent on agent sessions staying off — see Scope.
What moves
Stays in anvild | Moved to anvil-worker |
|---|---|
queue, enqueue/requeue_interrupted | execute |
repo + tree resolution, build_tar | collect_artifacts |
parse_pipeline, image allowlist check | download_tar |
the secret vault (app.vault.take) | parse_meta_tar |
step-script assembly, single_quote | the ArtifactSink trait |
store_artifact, the swap, gc_artifacts | |
mask_secrets (applied on receipt) | |
| the deploy webhook |
store_artifact deliberately stayed behind. The runner uploads the raw tar it
pulled out of the container and anvil decides what to do with it, so on-disk
layout — and browse, which turns a tarball into a servable directory tree —
never becomes a runner's call. The upload response reports the stored size, so
the per-run artifact budget is still charged what actually landed.
run_worker becomes run_dispatcher: same queue drain, but instead of calling
execute it parks the job until a runner claims it.
Script assembly stays server-side deliberately. The runner then never parses
.anvil/ci.yml and holds no pipeline model — it receives an image, a script,
sandbox caps, and a tar. That keeps the wire format stable as the pipeline
schema grows.
The crates
| Crate | Holds |
|---|---|
anvil-job | the wire format, and nothing else — serde only |
anvil-docker | connect/ensure_image, shared with anvil-agent |
anvil-worker | the runner binary: claim loop, client, executor |
anvil-job exists so the runner does not link anvil-core — and therefore
toasty, SQLite, gix and the rest of the forge — just to learn the shape of a
job. anvil-docker exists so anvil-agent need not depend on the runner.
The binary is anvil-worker, not anvil-runner, because
anvil-runner:latest is already the image CI jobs and agent sessions run in
(see agent-sessions.md). Prose says "runner" for the
concept; the crate avoids the collision.
The protocol
Plain HTTP against the existing axum server, under the /-/ system namespace,
so it inherits Caddy's TLS and needs no new listener.
| Endpoint | Method | Purpose |
|---|---|---|
/-/runner/claim | POST | long-poll; 204 on timeout |
/-/runner/jobs/{run_id}/checkout.tar | GET | the uploaded checkout |
/-/runner/jobs/{run_id}/heartbeat | POST | extend the lease |
/-/runner/jobs/{run_id}/artifacts/{name} | POST | one artifact tar |
/-/runner/jobs/{run_id}/result | POST | exit code + log + artifact meta |
The claim response carries everything execute takes as arguments today:
{
"run_id": 42,
"image": "rust:1.95-bookworm",
"platform": "linux/amd64",
"script": "set -e\n...",
"secrets": [{"name": "CARGO_TOKEN", "value": "..."}],
"artifacts": [{"name": "docs", "path": "target/doc", "browse": true}],
"sandbox": {
"memory_mb": 2048, "cpus": 2.0, "pids_limit": 512,
"timeout_secs": 1800, "network": true, "run_as": ""
}
}
The checkout is a separate GET rather than a base64 field, so a large tree
doesn't inflate a JSON body by a third. Secrets ride in the claim body over
TLS, never on a separately-fetchable URL.
Artifacts upload individually before result, for the same reason.
Logs
The runner posts the whole log with result, which is exactly what the
in-process runner did: append_log was only ever called when a run ended (and
on the secrets-failure exit). The log accumulated in memory and landed in one
write. The run page was not live before and is no less live now — a header
naming the runner is now written at claim time, so a running run at least
shows something.
Live logs are a genuine follow-up, and a remote runner makes them easier to justify (there is now a producer that could stream). Out of scope here.
Auth
A shared secret in [ci] runner_token, sent as X-Anvil-Runner-Token,
constant-time compared. The runner's self-asserted name rides alongside in
X-Anvil-Runner-Name — it labels runs and keys leases, and is explicitly not
a credential: everyone holding the token is one principal.
The four per-job endpoints additionally require the caller to hold that run's lease, so a valid token gets you a job rather than everyone else's. A lease mismatch answers 409, not 403: the caller is a legitimate runner whose claim simply expired.
This matches the existing deploy_secret pattern rather than inventing a
credential type. It is deliberate: API tokens are read-only and Bearer-only on
GET/HEAD (see untrusted-mode.md), a runner must POST, and
the write scope is still on TODO.md. Per-runner DB-backed tokens
with last_used_at are the right end state; a single-tenant forge with one
runner does not need them to start.
Leases
Held in memory on the server (anvil_core::jobs::Dispatch):
run_id -> (runner_name, expires_at, secrets), bumped by heartbeat every 30s
against a 120s TTL, swept every 30s. An expired lease returns the run to
queued.
A job's secrets are stashed on its lease rather than re-read from the vault
when the result lands. Vault::take fails once the repository's unlock TTL
lapses, and a job can easily outlive an unlock — re-reading would mean a long
run silently skips log masking, which is exactly the run whose log is most
likely to contain something. The values are already in this process's vault,
so this is not new exposure.
The honest gap: an expired lease can double-run a job whose runner is alive but unreachable. The container keeps going and the requeued run may be claimed elsewhere. That is what a lease without fencing buys; CI steps are assumed idempotent.
In-memory rather than columns on CiRun because Toasty migrations do not exist
yet — DEPLOY.md §Operations and TODO.md both flag that schema changes don't
auto-apply to the live database. Adding claimed_by/lease_expires_at to
CiRun (models.rs:65) would need a manual migration on hagrid.
It also costs nothing: anvild is the only dispatcher, so a lease has no reason
to outlive it, and the anvild-crash case is already handled —
requeue_interrupted (ci.rs:245) re-queues everything left running at
startup. The sweep covers the new case, a runner that dies mid-job.
Architecture
The build host is arm64; hagrid is x86_64. Two separate concerns:
What the shipped image is. Already solved: deploy/build.sh:33 cross-
compiles with zigbuild and builds --platform linux/amd64. Nothing to do,
though note the Dockerfile's RUN apt-get … does execute amd64 binaries
under emulation. Since anvild is a static musl binary,
gcr.io/distroless/static:nonroot would make that build pure COPY and
genuinely emulation-free. Nice-to-have.
What jobs run on. On an M2, rust:1.95-bookworm resolves to arm64, so
cargo test tests an architecture you don't ship. Fix: a platform: key in
.anvil/ci.yml, plumbed to CreateContainerOptions.platform
(bollard 0.18 container.rs:107) and CreateImageOptions.platform
(image.rs:72) — ensure_image currently leaves both at Default::default().
Enable Docker Desktop's Rosetta option; it is far faster than QEMU for amd64
Linux binaries.
Per-pipeline you then choose: native arm64 for lint and unit tests, amd64 under Rosetta for anything arch-sensitive.
Once runners are plural, platform also becomes a scheduling hint — a job
declaring linux/amd64 prefers a runner that is natively amd64. Not needed
with one runner; the field is forward-compatible with it.
Isolation on macOS
Docker on macOS is a Linux VM (LinuxKit under Docker Desktop, Lima under
Colima/OrbStack) and every container runs inside it, so the sandbox is enforced
by the same kernel primitives as on Linux: --cap-drop=ALL and
no-new-privileges are capabilities and prctl, the pids/memory/cpu caps are
cgroups v2.
It is a boundary better than hagrid's. untrusted-mode.md §1 notes that a
kernel or runc escape defeats the sandbox; on the Mac that escape reaches a
disposable Linux VM, not the host.
Two things to get right:
Docker::connect_with_socket_defaults()(docker.rs:15) will not find the socket. It hardcodes/var/run/docker.sock; Docker Desktop only creates that symlink when "Allow the default Docker socket to be used" is ticked, and the real path is~/.docker/run/docker.sock(Colima differs again). The runner usesDocker::connect_with_defaults(), which honoursDOCKER_HOST.- Run the runner natively under launchd, not in a container. Containerizing it means mounting the socket into it, rebuilding the root-equivalent hole this change removes from hagrid.
Size the VM's RAM deliberately: per-job memory_mb is carved out of a fixed
allocation. The dispatcher runs one job at a time, so this is slack rather than
a constraint.
Two processes on the build host
Image building cannot be a CI job. Job containers get no Docker socket by
design (execute's doc comment at :369, and untrusted-mode.md §1), and
that invariant is the whole broker model. So the build host runs two things at
two trust levels:
anvil-worker— claims jobs, runs them sandboxed, no Docker access inside the job container.- a deploy agent — fired on green CI, runs outside any sandbox with full
Docker access, does
docker build --platform linux/amd64anddocker pushtoregistry.vibe.richardscollin.com, then triggers hagrid to pull.
Keeping these separate is what preserves the broker model. Folding the second into the first would give pipeline authors a path to the daemon.
The deploy agent is out of scope for M1 — the existing deploy_webhook still
works, with the receiver moved to the build host and reached over Tailscale.
M2 folds it into the same dial-out channel as a privileged "publish" job kind,
authorized server-side by the is_deploy_target check that already scopes CD
to exactly one repository (config.rs:357). That removes the last inbound
requirement.
Scope
Agent sessions stay local-socket and stay out. anvil-agent calls
anvil_ci::docker::connect() in five places (lib.rs:141,
supervisor.rs:55,278,414,439) and docker.rs is explicitly shared plumbing,
so the crate split touches them — but a session is interactive (tmux attach,
exec streaming, resize, pump_transcript), which is a far harder protocol than
fire-and-forget CI. They are enabled = false and absent from
deploy/anvil.toml entirely, so nothing in production regresses.
ensure_image and connect therefore need a home both crates can reach. They
move to a thin anvil-docker crate rather than being duplicated —
ensure_image's local-fallback pull logic is subtle enough to be worth having
once.
The corollary: turning agent sessions on for hagrid later means either putting the socket back, or moving sessions onto the runner protocol too.
Registry cleanup
ensure_image's tolerance of a failed pull exists because anvil-runner:latest
"exists in no registry" (docker.rs:22-24) and is built straight into the local
daemon store. With registry.vibe.richardscollin.com up, push the image there
and the fallback stops being load-bearing.
It should stay in the code regardless — a runner on a fresh machine wants a clear failure when the pull fails and nothing is cached — but the comment explaining why it exists needs rewriting.
Milestones
M1 — the split. Done. anvil-job, anvil-docker and anvil-worker
crates, the five endpoints, in-memory leases, [ci] runner_token,
connect_with_defaults, run_worker → run_dispatcher, socket mount dropped
from deploy/run.sh. Deploys keep using the existing webhook.
M2 — platform. The plumbing is already live: JobSpec.platform reaches
CreateContainerOptions and CreateImageOptions, and a runner advertises its
native platform when it claims. What is missing is a source — a platform:
key in .anvil/ci.yml (and probably a [ci] platform default), plus routing a
job to a runner that has that architecture natively.
M3 — publish jobs. The deploy agent folds into the dial-out channel,
authorized by is_deploy_target. No inbound path to the build host remains.
Not yet done
- Live logs. Newly worth doing, still not done. The result POST is a single write; streaming needs chunked append with offsets and a UI that tolerates gaps.
- Per-runner credentials. One shared secret means one revocation for all
runners, and no
last_used_at. Wants the API-token write scope first. - Runner labels.
platformis the only scheduling dimension in M2. Tags ("has-postgres", "big-memory") are the obvious next axis and are not designed. - Concurrency. Nothing bounds how many jobs are in flight beyond how many
runners exist, and nothing stops one runner claiming repeatedly. The
in-process runner's "one job at a time" was a property of the loop, and it is
gone; a
max_concurrentequivalent for CI does not exist. ensure_image's local fallback vsplatform. A cached image of the wrong architecture satisfies the fallback, since it inspects presence and not arch. Only bites an offline runner asked to cross-build.