collin/anvil
RenderedSource
Threat model: running anvil with untrusted users
anvil is built as a single-tenant, owner-operated forge: the operator and the people with accounts are assumed to trust each other (a person, a family, a small team). This document records what would have to be true before opening registration (or repo write access) to people you don't trust, ranked by severity. It is the output of the security-audit pass; keep it updated as items land.
Supported stance: single-tenant / owner-operated. Untrusted multi-tenancy is not supported until at least items 1–4 below are closed.
1. CI: arbitrary code execution by design
Anyone who can push to a repo with a .anvil/ci.yml runs arbitrary code on
your hardware. That is the point of CI, so the question is only how well the
blast radius is contained.
Broker model (implemented). anvil itself is the only Docker client. The job container gets:
- no Docker socket, no bind mounts, no volumes — the checkout is uploaded into the container as a tar via the Docker API, so the job can never reach anvil's data directory or the host filesystem;
--cap-drop=ALLandno-new-privilegesunconditionally;- pids / memory(+swap) / cpu caps (
ci.pids_limit,ci.memory_mb,ci.cpus; defaults 512 / 2048 MiB / 2); - a wall-clock timeout (
ci.timeout_secs, default 30 minutes) after which the container is force-removed; - optionally no network (
ci.network = false) and a non-root user (ci.run_as = "1000:1000") — most real builds need network and many base images assume root, so these default to permissive; - an image allowlist (
ci.allowed_images) — empty allows any image, which is fine single-tenant; set it before letting strangers push.
Deliberately not done: read-only rootfs (the workspace lives in the
container filesystem precisely so no volume is ever attached; builds also
write $HOME caches), and egress filtering (network is all-or-nothing).
Residual risk / stronger tier. Containers share the host kernel; a kernel or runc escape defeats all of the above. For genuinely hostile tenants run the jobs under gVisor/Kata/Firecracker (a runtime flag on the broker — the "isolated workers" idea), and add per-user CI-minute and disk quotas (image pulls consume host disk). Until then, CI for untrusted users should stay off.
2. Stored XSS via served content
Repo browsing renders escaped text (Maud auto-escapes; highlighting emits sanitized HTML), so hostile file content does not execute in the forge origin today.
Pages hosting is the exception by design: it serves attacker-authored
HTML/JS. On a single-origin deployment, a pages site runs in the same origin
as the forge UI — its JS could read forge pages and drive authenticated
requests in a visitor's session. Mitigations in place: session cookies are
HttpOnly (no token theft) and all mutating routes require the CSRF token.
But same-origin JS can still read the token off a fetched page, so for
untrusted users pages must move to a separate origin (e.g.
*.pages.example.com), as GitHub does. The same applies to any future "raw
blob" endpoint: serve text/plain + nosniff + a restrictive CSP, or a
separate origin.
3. Git resource exhaustion
A hostile pusher can send decompression bombs (tiny pack, enormous objects),
deep delta chains, or millions of refs; a hostile cloner can request expensive
packs repeatedly. Needed before untrusted use: per-repo and per-user storage
quotas, an upload size cap on receive-pack, timeouts/memory bounds on pack
ingestion and pack generation, and a cap on advertised refs. (The CI tar
materializer also loads a full checkout into memory — bounded today only by
push quotas not existing.)
4. Open registration anti-abuse
Registration is currently operator-controlled (CLI), which is the real
mitigation. Opening it requires: email verification, rate limiting on signup /
login / repo creation, a CAPTCHA or proof-of-work, per-user quotas (repos,
storage, CI minutes), and an admin ban/cleanup path. Reserved usernames and
the /-/ route namespace already prevent route-shadowing squats.
5. Authorization granularity
Access today is owner-or-admin, repo public-or-private. Fine single-tenant; multi-user collaboration needs collaborator roles (read/write/admin per repo), per-repo deploy keys, and scoped tokens instead of full-account SSH keys. The deploy webhook is already scoped to exactly one configured repo.
6. Webhook SSRF (future)
User-configurable webhooks don't exist yet (the deploy webhook URL is operator-set in the config file, not user data). When they land: resolve and block private/link-local/metadata ranges (and re-check on redirect), pin DNS, cap response sizes and time, and never reflect response bodies to users.
Already right (keep it that way)
- Private repos 404 for non-readers — no existence leak (
resolve_repo). - Reserved usernames +
/-/namespace for app routes. - CI checkout is tar-uploaded, never bind-mounted; job containers get no socket; sandbox defaults are on (see §1).
- CD webhook gated to a single configured repo + shared-secret header.
- Passwords: argon2. Sessions:
HttpOnly+SameSite=Lax+Secure(auto whenbase_urlis https). CSRF: HMAC synchronizer token, constant-time compare, on all mutating forms; htmx requests carry it viahx-headers. - SSH auth by exact public-key match; unknown keys rejected.