anvilsign in

collin/anvil

RenderedSource

1# Threat model: running anvil with untrusted users
2
3anvil is built as a **single-tenant, owner-operated forge**: the operator and
4the people with accounts are assumed to trust each other (a person, a family, a
5small team). This document records what would have to be true before opening
6registration (or repo write access) to people you *don't* trust, ranked by
7severity. It is the output of the security-audit pass; keep it updated as
8items land.
9
10**Supported stance:** single-tenant / owner-operated. Untrusted multi-tenancy
11is *not* supported until at least items 1–4 below are closed. Agent sessions
12(§7) raise the bar further and are off by default.
13
14---
15
16## 1. CI: arbitrary code execution by design
17
18Anyone who can push to a repo with a `.anvil/ci.yml` runs arbitrary code on
19your hardware. That is the *point* of CI, so the question is only how well the
20blast radius is contained.
21
22**Broker model (implemented).** anvil itself is the only Docker client. The
23job container gets:
24
25- **no Docker socket, no bind mounts, no volumes** — the checkout is uploaded
26 into the container as a tar via the Docker API, so the job can never reach
27 anvil's data directory or the host filesystem;
28- **`--cap-drop=ALL` and `no-new-privileges`** unconditionally;
29- **pids / memory(+swap) / cpu caps** (`ci.pids_limit`, `ci.memory_mb`,
30 `ci.cpus`; defaults 512 / 2048 MiB / 2);
31- **a wall-clock timeout** (`ci.timeout_secs`, default 30 minutes) after which
32 the container is force-removed;
33- optionally **no network** (`ci.network = false`) and a **non-root user**
34 (`ci.run_as = "1000:1000"`) — most real builds need network and many base
35 images assume root, so these default to permissive;
36- an **image allowlist** (`ci.allowed_images`) — empty allows any image, which
37 is fine single-tenant; set it before letting strangers push.
38
39**Secrets in a pipeline.** A run can request repository secrets (see
40[secrets.md](secrets.md)), which arrive as environment variables inside that
41same container. Anyone who can push to the repo can therefore read every secret
42it declares, by editing the pipeline; log masking stops accidents, not intent.
43The compensating control is that anvil cannot decrypt them at all unless the
44owner has unlocked the repo, so the exposure window is bounded by the unlock
45TTL rather than being permanent.
46
47**Secrets leave the forge.** Since execution moved to remote runners (see
48[remote-runners.md](remote-runners.md)), a job's secrets are sent to the runner
49with its job and sit in plaintext in a container on a machine anvil does not
50own. A runner host is therefore as trusted as the forge itself, and the shared
51`[ci] runner_token` means any holder of that one secret can claim any job — and
52so receive whatever secrets it declares. Two consequences worth stating plainly:
53runners belong on hardware you control, and per-runner credentials (with the
54ability to revoke one without revoking all) are a prerequisite for anything
55looser.
56
57**Deliberately not done:** read-only rootfs (the workspace lives in the
58container filesystem precisely so no volume is ever attached; builds also
59write `$HOME` caches), and egress *filtering* (network is all-or-nothing).
60
61**Residual risk / stronger tier.** Containers share the host kernel; a kernel
62or runc escape defeats all of the above. For genuinely hostile tenants run the
63jobs under gVisor/Kata/Firecracker (a runtime flag on the broker — the
64"isolated workers" idea), and add per-user CI-minute and disk quotas (image
65pulls consume host disk). Until then, CI for untrusted users should stay off.
66
67## 2. Stored XSS via served content
68
69Repo browsing renders escaped text (Maud auto-escapes; highlighting emits
70sanitized HTML), so hostile file *content* does not execute in the forge
71origin today.
72
73**Pages hosting is the exception by design**: it serves attacker-authored
74HTML/JS. On a single-origin deployment, a pages site runs in the same origin
75as the forge UI — its JS could read forge pages and drive authenticated
76requests in a visitor's session. Mitigations in place: session cookies are
77`HttpOnly` (no token theft) and all mutating routes require the CSRF token.
78But same-origin JS can still *read* the token off a fetched page, so for
79untrusted users pages must move to a **separate origin** (e.g.
80`*.pages.example.com`), as GitHub does. The same applies to any future "raw
81blob" endpoint: serve `text/plain` + `nosniff` + a restrictive CSP, or a
82separate origin.
83
84## 3. Git resource exhaustion
85
86A hostile pusher can send decompression bombs (tiny pack, enormous objects),
87deep delta chains, or millions of refs; a hostile cloner can request expensive
88packs repeatedly. Needed before untrusted use: per-repo and per-user storage
89quotas, an upload size cap on `receive-pack`, timeouts/memory bounds on pack
90ingestion and pack generation, and a cap on advertised refs. (The CI tar
91materializer also loads a full checkout into memory — bounded today only by
92push quotas not existing.)
93
94## 4. Open registration anti-abuse
95
96Registration is currently operator-controlled (CLI), which is the real
97mitigation. Opening it requires: email verification, rate limiting on signup /
98login / repo creation, a CAPTCHA or proof-of-work, per-user quotas (repos,
99storage, CI minutes), and an admin ban/cleanup path. Reserved usernames and
100the `/-/` route namespace already prevent route-shadowing squats.
101
102## 5. Authorization granularity
103
104Access today is owner-or-admin, repo public-or-private. Fine single-tenant;
105multi-user collaboration needs collaborator roles (read/write/admin per repo),
106per-repo deploy keys, and scoped tokens instead of full-account SSH keys. The
107deploy webhook is already scoped to exactly one configured repo.
108
109## 6. Webhook SSRF (future)
110
111User-configurable webhooks don't exist yet (the deploy webhook URL is
112operator-set in the config file, not user data). When they land: resolve and
113block private/link-local/metadata ranges (and re-check on redirect), pin DNS,
114cap response sizes and time, and never reflect response bodies to users.
115
116## 7. Agent sessions: a model reading repo content as instructions
117
118Off by default (`agent.enabled`), and it should stay off unless you trust
119everyone who can push. See [agent-sessions.md](agent-sessions.md) for the
120design.
121
122This is **not** a variant of §1. CI runs code the pusher *wrote*: hostile, but
123authored. An agent session runs a model that reads repository content, file
124names, issue text and TODO items and may treat any of it as instruction. A
125prompt injected into a README is therefore a path to arbitrary tool use inside
126the session — and the session has network access it cannot do without.
127
128**What carries over from §1.** The container is created through the same broker:
129no Docker socket, no bind mounts, no volumes, `--cap-drop=ALL`,
130`no-new-privileges`, pids/memory/cpu caps. Sessions additionally run as an
131unprivileged user (`agent`, uid 1000) rather than root, which CI does not. The
132workspace and the agent's credentials both arrive as tar uploads through the
133Docker API, so nothing on anvil's filesystem is reachable.
134
135**What is new.**
136
137- **Network is mandatory.** `ci.network = false` has no counterpart here: the
138 session must reach the model API. Egress filtering does not exist, so an
139 injected instruction can exfiltrate anything in the workspace.
140- **`--dangerously-skip-permissions` is passed deliberately.** The container is
141 the security boundary; a permission prompt inside a container the agent
142 already owns end-to-end buys nothing. The consequence is that containment is
143 entirely the container's job — there is no second line of defence inside it.
144- **Credentials sit in the container.** `agent.credentials_dir` is copied into
145 every session, so any code the agent runs — including code the repository
146 supplied — can read the model credentials. Scope that account accordingly.
147- **A push credential is coming and is not yet scoped.** M1 seeds files only and
148 hands out no credential at all. When the container gets a real clone it will
149 need one, and restricting it to `refs/heads/agent/*` requires a ref filter in
150 receive-pack that does not exist. Until it does, a session credential could
151 write any branch, `main` included.
152- **No rate limiting.** `agent.max_concurrent` bounds how many run at once, not
153 how many are started. Once push/issue triggers land, automated pushes could
154 queue sessions indefinitely; restricting who can push is the only control
155 today.
156
157**Before untrusted use:** all of §§1–4, plus per-user session quotas, egress
158filtering or an allowlisted proxy, the ref-scoped push credential, and a
159credential per repository rather than one instance-wide.
160
161---
162
163## Already right (keep it that way)
164
165- Private repos 404 for non-readers — no existence leak (`resolve_repo`).
166- Reserved usernames + `/-/` namespace for app routes.
167- CI checkout is tar-uploaded, never bind-mounted; job containers get no
168 socket; sandbox defaults are on (see §1).
169- CD webhook gated to a single configured repo + shared-secret header.
170- Passwords: argon2. Sessions: `HttpOnly` + `SameSite=Lax` + `Secure` (auto
171 when `base_url` is https). CSRF: HMAC synchronizer token, constant-time
172 compare, on all mutating forms; htmx requests carry it via `hx-headers`.
173- SSH auth by exact public-key match; unknown keys rejected.