collin/anvil
RenderedSource
Remote runners
Status: M1 + M2 implemented (2026-08-24). Supersedes the in-process CI
executor that used to live in crates/anvil-ci.
anvil used to be the runner: run_worker drained the queue in-process and
execute created the job container on the local Docker socket. That works, and
it pins CI to whichever machine anvil runs on — which is hagrid, a droplet
small enough that a release build OOMs it (see DEPLOY.md §3).
So CI had to run somewhere else.
A runner is a separate binary that dials out to anvil, claims a job, runs it in a sandboxed container on its own Docker daemon, and posts the result back. anvild stops executing anything.
hagrid (droplet, 1 GB) build host (Mac mini M2)
+------------------------+ +---------------------------+
| anvild | | anvil-worker (launchd) |
| queue, tree, vault | <--------- | long-poll claim |
| artifacts, webhook | https | POST result |
| | ---------> | | |
| NO docker socket | | v |
+------------------------+ | [ job container ] |
| cap-drop, no socket |
+---------------------------+
Why dial-out
The alternative was forwarding the build host's Docker socket to hagrid and pointing anvil at it. Dial-out wins on four counts, in order of importance:
- No inbound path from the public VPS into the home LAN. A forwarded Docker socket is unauthenticated root on the machine that owns it. Punching that from an internet-facing droplet into a home network makes hagrid's compromise the build host's compromise.
- NAT and sleep/wake are free. The build host is a desktop machine behind a residential connection. A dialing client reconnects; a dialed-into server needs a tunnel supervised on both ends.
- The failure mode is legible. A sleeping Mac becomes "no runner
available," not a Docker API error surfacing as
[runner error]in a job log. - Runners become plural. Once anvil addresses runners rather than a socket, a second one is configuration, not architecture — which is how the arm64/amd64 split below gets solved properly.
What this buys hagrid
anvild stops needing Docker at all. compose.yaml carries no
/var/run/docker.sock mount and no group_add, which deletes
the warning in DEPLOY.md §4 about the container holding
root-equivalent control of the host. The internet-facing process stops being a
host-escape vector.
This is contingent on agent sessions staying off — see Scope.
What moves
Stays in anvild | Moved to anvil-worker |
|---|---|
queue, enqueue/requeue_interrupted | execute |
repo + tree resolution, build_tar | collect_artifacts |
parse_pipeline, image allowlist check | download_tar |
the secret vault (app.vault.take) | parse_meta_tar |
step-script assembly, single_quote | the ArtifactSink trait |
store_artifact, the swap, gc_artifacts | |
mask_secrets (applied on receipt) | |
| the deploy webhook |
store_artifact deliberately stayed behind. The runner uploads the raw tar it
pulled out of the container and anvil decides what to do with it, so on-disk
layout — and browse, which turns a tarball into a servable directory tree —
never becomes a runner's call. The upload response reports the stored size, so
the per-run artifact budget is still charged what actually landed.
run_worker becomes run_dispatcher: same queue drain, but instead of calling
execute it parks the job until a runner claims it.
Script assembly stays server-side deliberately. The runner then never parses
.anvil/ci.toml and holds no pipeline model — it receives an image, a script,
sandbox caps, and a tar. That keeps the wire format stable as the pipeline
schema grows.
The crates
| Crate | Holds |
|---|---|
anvil-job | the wire format, and nothing else — serde only |
anvil-docker | connect/ensure_image, shared with anvil-agent |
anvil-worker | the runner binary: claim loop, client, executor |
anvil-job exists so the runner does not link anvil-core — and therefore
toasty, SQLite, gix and the rest of the forge — just to learn the shape of a
job. anvil-docker exists so anvil-agent need not depend on the runner.
The binary is anvil-worker, not anvil-runner, because
anvil-runner:latest is already the image CI jobs and agent sessions run in
(see agent-sessions.md). Prose says "runner" for the
concept; the crate avoids the collision.
The protocol
Plain HTTP against the existing axum server, under the /-/ system namespace,
so it inherits Caddy's TLS and needs no new listener.
| Endpoint | Method | Purpose |
|---|---|---|
/-/runner/claim | POST | long-poll; 204 on timeout |
/-/runner/jobs/{run_id}/checkout.tar | GET | the uploaded checkout |
/-/runner/jobs/{run_id}/heartbeat | POST | extend the lease |
/-/runner/jobs/{run_id}/artifacts/{name} | POST | one artifact tar |
/-/runner/jobs/{run_id}/result | POST | exit code + log + artifact meta |
The claim response carries everything execute takes as arguments today:
{
"run_id": 42,
"image": "anvil-runner:rust",
"platform": "linux/amd64",
"script": "set -e\n...",
"secrets": [{"name": "CARGO_TOKEN", "value": "..."}],
"artifacts": [{"name": "docs", "path": "target/doc", "browse": true}],
"sandbox": {
"memory_mb": 2048, "cpus": 2.0, "pids_limit": 512,
"timeout_secs": 1800, "network": true, "run_as": ""
}
}
The checkout is a separate GET rather than a base64 field, so a large tree
doesn't inflate a JSON body by a third. Secrets ride in the claim body over
TLS, never on a separately-fetchable URL.
Artifacts upload individually before result, for the same reason.
Logs
The runner posts the whole log with result, which is exactly what the
in-process runner did: append_log was only ever called when a run ended (and
on the secrets-failure exit). The log accumulated in memory and landed in one
write. The run page was not live before and is no less live now — a header
naming the runner is now written at claim time, so a running run at least
shows something.
Live logs are a genuine follow-up, and a remote runner makes them easier to justify (there is now a producer that could stream). Out of scope here.
Auth
A shared secret in [ci] runner_token, sent as X-Anvil-Runner-Token,
constant-time compared. The runner's self-asserted name rides alongside in
X-Anvil-Runner-Name — it labels runs and keys leases, and is explicitly not
a credential: everyone holding the token is one principal.
The four per-job endpoints additionally require the caller to hold that run's lease, so a valid token gets you a job rather than everyone else's. A lease mismatch answers 409, not 403: the caller is a legitimate runner whose claim simply expired.
This matches the existing deploy_secret pattern rather than inventing a
credential type. It is deliberate: API tokens are read-only and Bearer-only on
GET/HEAD (see untrusted-mode.md), a runner must POST, and
the write scope is still on TODO.md. Per-runner DB-backed tokens
with last_used_at are the right end state; a single-tenant forge with one
runner does not need them to start.
Leases
Held in memory on the server (anvil_core::jobs::Dispatch):
run_id -> (runner_name, expires_at, secrets), bumped by heartbeat every 30s
against a 120s TTL, swept every 30s. An expired lease returns the run to
queued.
A job's secrets are stashed on its lease rather than re-read from the vault
when the result lands. Vault::take fails once the repository's unlock TTL
lapses, and a job can easily outlive an unlock — re-reading would mean a long
run silently skips log masking, which is exactly the run whose log is most
likely to contain something. The values are already in this process's vault,
so this is not new exposure.
The honest gap: an expired lease can double-run a job whose runner is alive but unreachable. The container keeps going and the requeued run may be claimed elsewhere. That is what a lease without fencing buys; CI steps are assumed idempotent.
In-memory rather than columns on CiRun because Toasty migrations do not exist
yet — DEPLOY.md §Operations and TODO.md both flag that schema changes don't
auto-apply to the live database. Adding claimed_by/lease_expires_at to
CiRun (models.rs:65) would need a manual migration on hagrid.
It also costs nothing: anvild is the only dispatcher, so a lease has no reason
to outlive it, and the anvild-crash case is already handled —
requeue_interrupted (ci.rs:245) re-queues everything left running at
startup. The sweep covers the new case, a runner that dies mid-job.
Architecture
The build host is arm64; hagrid is x86_64. Two separate concerns:
What the shipped image is. Already solved: deploy/build.sh cross-compiles
with zigbuild and compose.yaml pins platforms: [linux/amd64]. Nothing to do,
though note the Dockerfile's RUN apt-get … does execute amd64 binaries
under emulation. Since anvild is a static musl binary,
gcr.io/distroless/static:nonroot would make that build pure COPY and
genuinely emulation-free. Nice-to-have.
What jobs run on. On an M2, a multi-arch image resolves to arm64, so
cargo test tests an architecture you don't ship. That is what M2 fixes.
A job's platform comes from platform in .anvil/ci.toml, falling back to
[ci] platform, falling back to the claiming runner's native architecture —
so an instance that sets neither behaves exactly as it did before. It reaches
CreateContainerOptions.platform and CreateImageOptions.platform, and the
value is validated as os/arch[/variant] at parse time (a bare amd64 would
otherwise reach Docker as an operating system named amd64).
Per-pipeline you then choose: native arm64 for lint and unit tests, amd64 under Rosetta for anything arch-sensitive. Turn Rosetta on in Docker Desktop; it is far faster than QEMU for amd64 Linux binaries.
Routing
platform is also the scheduling dimension. Every claim and heartbeat records
the runner's advertised architecture in Dispatch (jobs.rs), which expires
after RUNNER_TTL — 5 minutes, longer than both the 55s claim poll and the 30s
heartbeat, so an idle runner and a runner mid-build both stay visible. A
claiming runner is then offered, in queue order:
- runs that name its platform, or name none at all;
- then runs whose platform no currently connected runner is native to.
Tier 2 is what keeps one arm64 Mac usable as the only runner for pipelines that
declare linux/amd64: nobody can run them natively, so it emulates them rather
than leaving them queued forever. Add an amd64 runner and the Mac stops taking
those jobs the moment the new runner's first claim registers it — no
configuration, which is the property the dial-out model was chosen for.
The corollary worth stating plainly: platform is not a promise of native
execution. It is a promise about what the job runs, which the runner enforces
by inspecting the image it ended up with (anvil-docker::check_platform) and
failing the job if the architecture is not the one asked for. Where it runs is
a scheduling preference. A run's log header says which it got:
platform: linux/amd64 (emulated on linux/arm64)
Seeing who is connected
/-/admin/runners (admin-only, 404 for everyone else) lists every runner
Dispatch still counts as present: name, advertised platform, worker version,
how long since it last spoke, how long it has been connected, and the runs it
holds right now, each linked to its CI page. Underneath it is the same map
routing reads, so the page and the dispatcher can never disagree about who is
out there. It also prints the queue depth, because the two together are the
whole diagnosis: runners and no queue is a healthy idle instance, a queue and
no runners is a stuck one, and a queue with every runner busy is neither — it
is capacity.
There is no separate heartbeat for presence, on purpose. Liveness is the
traffic a working runner already generates: an idle one re-registers itself
every time its parked claim expires and it dials back in (CLAIM_POLL, 55s),
and a busy one every HEARTBEAT_INTERVAL (30s) for as long as its job runs.
Between them there is no state a runner can be in where it is useful and
silent, so a dedicated ping would only add a way for a runner to look alive
while claiming nothing.
What that costs is resolution, and the page is explicit about it rather than
hiding it. A runner is shown "late" once it has been quiet for a whole claim
poll plus a heartbeat (85s) — long enough that a runner merely parked in a
long poll never reads as late. It stays listed, and keeps being routed to,
until RUNNER_TTL (5 min), because that is exactly what the dispatcher still
believes; the page's job is to show that belief, not to invent a second one.
So a machine that loses power disappears from routing in up to five minutes and
reads as late within ninety seconds. The page reloads itself every 15s.
Isolation on macOS
Docker on macOS is a Linux VM (LinuxKit under Docker Desktop, Lima under
Colima/OrbStack) and every container runs inside it, so the sandbox is enforced
by the same kernel primitives as on Linux: --cap-drop=ALL and
no-new-privileges are capabilities and prctl, the pids/memory/cpu caps are
cgroups v2.
It is a boundary better than hagrid's. untrusted-mode.md §1 notes that a
kernel or runc escape defeats the sandbox; on the Mac that escape reaches a
disposable Linux VM, not the host.
Two things to get right:
Docker::connect_with_socket_defaults()(docker.rs:15) will not find the socket. It hardcodes/var/run/docker.sock; Docker Desktop only creates that symlink when "Allow the default Docker socket to be used" is ticked, and the real path is~/.docker/run/docker.sock(Colima differs again). The runner usesDocker::connect_with_defaults(), which honoursDOCKER_HOST.- Run the runner natively under launchd, not in a container. Containerizing it means mounting the socket into it, rebuilding the root-equivalent hole this change removes from hagrid.
Size the VM's RAM deliberately: per-job memory_mb is carved out of a fixed
allocation. The dispatcher runs one job at a time, so this is slack rather than
a constraint.
Running one on a Mac mini
# On the Mac, in a checkout of anvil:
cargo build --release -p anvil-worker
sudo cp target/release/anvil-worker /usr/local/bin/
# Docker Desktop → Settings → General:
# ✓ Use Rosetta for x86_64/amd64 emulation on Apple Silicon
# Without it, a linux/amd64 job runs under QEMU — correct, and much slower.
cp deploy/worker/com.anvil.worker.plist ~/Library/LaunchAgents/
# Fill in --url, --name, ANVIL_RUNNER_TOKEN and DOCKER_HOST, then:
chmod 600 ~/Library/LaunchAgents/com.anvil.worker.plist
launchctl load -w ~/Library/LaunchAgents/com.anvil.worker.plist
tail -f /tmp/anvil-worker.log # "anvil-worker macmini (linux/arm64) → …"
The forge side needs [ci] runner_token set to the same secret; until it is,
every claim gets a 503 saying so and queued runs sit.
The plist is a LaunchAgent, not a LaunchDaemon, because Docker Desktop's
socket only exists inside the logged-in user's session — a root daemon starts
before Docker and never finds it. The consequence to know about: the runner is
only up while that user is logged in, and a Mac that sleeps stops claiming.
That is the failure mode the dial-out design chose (a sleeping Mac reads as "no
runner available"), and after RUNNER_TTL its architecture stops counting as
present, so anything routed to it falls back to another runner.
--name matters: the default reads $HOSTNAME, which launchd does not set, so
an unnamed runner is called runner.
Two runners locally
compose.override.yaml brings up runner-1 and runner-2 next to the local
forge, so the parts of this design that only appear with more than one runner —
concurrent pipelines, and the "no connected runner is native to this platform"
half of routing — are testable without a second machine:
./deploy/build.sh --debug --worker # stages the anvild and anvil-worker binaries
docker compose up -d --build
docker compose logs -f runner-1 runner-2
# anvil-worker dev-1 (linux/amd64) → http://anvil:3000
They reach the forge as http://anvil:3000 over the compose network (the
browser-facing base_url does not resolve inside a container) and authenticate
with the runner_token committed in deploy/anvil.dev.toml. Both are amd64
here, so to watch the fallback tier work, give a pipeline platform: linux/arm64 and confirm it still gets claimed. Adding a third is a copy of the
four-line service block with a new name.
These runners are containerized, and that is the one thing production must
never copy: the socket mount is the root-equivalent hold that moving CI off the
forge removed. It is acceptable locally only because the same file already
mounts that socket into anvil for agent sessions, so the machine's trust
boundary is unchanged. The deployed compose.yaml grants neither.
Two processes on the build host
Image building cannot be a CI job. Job containers get no Docker socket by
design (execute's doc comment at :369, and untrusted-mode.md §1), and
that invariant is the whole broker model. So the build host runs two things at
two trust levels:
anvil-worker— claims jobs, runs them sandboxed, no Docker access inside the job container.- a deploy agent — fired on green CI, runs outside any sandbox with full
Docker access, does
docker build --platform linux/amd64anddocker pushtoregistry.vibe.richardscollin.com, then triggers hagrid to pull.
Keeping these separate is what preserves the broker model. Folding the second into the first would give pipeline authors a path to the daemon.
The deploy agent is out of scope for M1 — the existing deploy_webhook still
works, with the receiver moved to the build host and reached over Tailscale.
M2 folds it into the same dial-out channel as a privileged "publish" job kind,
authorized server-side by the is_deploy_target check that already scopes CD
to exactly one repository (config.rs:357). That removes the last inbound
requirement.
Scope
Agent sessions stay local-socket and stay out. anvil-agent calls
anvil_ci::docker::connect() in five places (lib.rs:141,
supervisor.rs:55,278,414,439) and docker.rs is explicitly shared plumbing,
so the crate split touches them — but a session is interactive (tmux attach,
exec streaming, resize, pump_transcript), which is a far harder protocol than
fire-and-forget CI. They are enabled = false and absent from
deploy/anvil.toml entirely, so nothing in production regresses.
ensure_image and connect therefore need a home both crates can reach. They
move to a thin anvil-docker crate rather than being duplicated —
ensure_image's local-fallback pull logic is subtle enough to be worth having
once.
The corollary: turning agent sessions on for hagrid later means either putting the socket back, or moving sessions onto the runner protocol too.
Registry cleanup
ensure_image's tolerance of a failed pull exists because anvil-runner:latest
"exists in no registry" (docker.rs:22-24) and is built straight into the local
daemon store. With registry.vibe.richardscollin.com up, push the image there
and the fallback stops being load-bearing.
It should stay in the code regardless — a runner on a fresh machine wants a clear failure when the pull fails and nothing is cached — but the comment explaining why it exists needs rewriting.
Milestones
M1 — the split. Done. anvil-job, anvil-docker and anvil-worker
crates, the five endpoints, in-memory leases, [ci] runner_token,
connect_with_defaults, run_worker → run_dispatcher, and no socket mount in
the deployed container. Deploys keep using the existing webhook.
M2 — platform. Done. platform in .anvil/ci.toml, a [ci] platform
default, CiConfig::resolve_platform, the two-tier routing above (backed by a
runner registry in Dispatch), the emulation note in the run header, and an
architecture check on the image the runner actually got.
M3 — publish jobs. The deploy agent folds into the dial-out channel,
authorized by is_deploy_target. No inbound path to the build host remains.
Not yet done
- Live logs. Newly worth doing, still not done. The result POST is a single write; streaming needs chunked append with offsets and a UI that tolerates gaps.
- Per-runner credentials. One shared secret means one revocation for all
runners, and no
last_used_at. Wants the API-token write scope first. - Runner labels.
platformis the only scheduling dimension in M2. Tags ("has-postgres", "big-memory") are the obvious next axis and are not designed. - A runners page.
Dispatch::runners()now knows every runner connected in the last five minutes and what it is. Nothing renders it, so "is my Mac actually claiming?" is still answered by reading logs. - Routing is per-claim, not per-queue. A runner that can take nothing sleeps until the next wake; it does not reserve the job it declined. With two runners and a job only one can run natively, the other simply keeps polling — correct, but it means a queue can look busy while a runner looks idle.
- Concurrency. Nothing bounds how many jobs are in flight beyond how many
runners exist, and nothing stops one runner claiming repeatedly. The
in-process runner's "one job at a time" was a property of the loop, and it is
gone; a
max_concurrentequivalent for CI does not exist. - Concurrency, again. Platform routing makes a second runner useful, which
makes the missing
max_concurrentmore pressing rather than less.