anvilsign in

collin/anvil · 35ea3553

Serialize job claims; note secrets leaving the forge in the threat model

Collin Richards · 2026-08-24 08:08 UTC · 35ea3553765a36712c2c8c15719662969a443c4b · parent 12a87324 · browse files

modifiedcrates/anvil-ci/src/lib.rs+5 −0
⋯ 198 unchanged lines
199199 /// handed out, and the next queued run is tried. That keeps a single bad
200200 /// pipeline from wedging the queue.
201201 pub async fn claim_next(app: &App, runner: &str) -> Option<JobSpec> {
202+ // Held across the whole attempt: listing the queue and marking a run
203+ // running are separate awaits, so without it two runners polling at the
204+ // same moment could both be handed the same job.
205+ let _claiming = app.jobs.claim_guard().await;
206+
202207 let queued = match ci::queued_ids(&app.db).await {
203208 Ok(ids) => ids,
204209 Err(e) => {
⋯ 577 unchanged lines
modifiedcrates/anvil-core/src/jobs.rs+11 −0
⋯ 58 unchanged lines
5959 struct Inner {
6060 leases: Mutex<HashMap<i64, Lease>>,
6161 wake: tokio::sync::Notify,
62+ /// Serializes claim attempts. Reading the queue and marking a run
63+ /// `running` are separate awaits, so two runners polling at once could
64+ /// otherwise both be handed the same job.
65+ claim: tokio::sync::Mutex<()>,
6266 }
6367
6468 impl Dispatch {
⋯ 15 unchanged lines
8084 let _ = tokio::time::timeout(timeout, self.inner.wake.notified()).await;
8185 }
8286
87+ /// Hold for the whole of a claim attempt, so the queue read and the
88+ /// `mark_running` that follows it are atomic with respect to other
89+ /// runners.
90+ pub async fn claim_guard(&self) -> tokio::sync::MutexGuard<'_, ()> {
91+ self.inner.claim.lock().await
92+ }
93+
8394 /// Record that `runner` holds `run_id`, stashing the job's secrets for
8495 /// masking when the result arrives.
8596 pub fn claim(&self, run_id: i64, runner: &str, secrets: Vec<(String, String)>) {
⋯ 83 unchanged lines
modifieddocs/remote-runners.md+0 −4
⋯ 324 unchanged lines
325325 runners, and no `last_used_at`. Wants the API-token write scope first.
326326 - **Runner labels.** `platform` is the only scheduling dimension in M2. Tags
327327 ("has-postgres", "big-memory") are the obvious next axis and are not designed.
328-- **Secrets on the runner host.** They now cross the network and sit in
329- plaintext in a container on a machine anvil does not own. This needs a
330- paragraph in `untrusted-mode.md` §1 — the exposure is no longer bounded by
331- hagrid.
332328 - **Concurrency.** Nothing bounds how many jobs are in flight beyond how many
333329 runners exist, and nothing stops one runner claiming repeatedly. The
334330 in-process runner's "one job at a time" was a property of the loop, and it is
⋯ 4 unchanged lines
modifieddocs/untrusted-mode.md+10 −0
⋯ 43 unchanged lines
4444 owner has unlocked the repo, so the exposure window is bounded by the unlock
4545 TTL rather than being permanent.
4646
47+**Secrets leave the forge.** Since execution moved to remote runners (see
48+[remote-runners.md](remote-runners.md)), a job's secrets are sent to the runner
49+with its job and sit in plaintext in a container on a machine anvil does not
50+own. A runner host is therefore as trusted as the forge itself, and the shared
51+`[ci] runner_token` means any holder of that one secret can claim any job — and
52+so receive whatever secrets it declares. Two consequences worth stating plainly:
53+runners belong on hardware you control, and per-runner credentials (with the
54+ability to revoke one without revoking all) are a prerequisite for anything
55+looser.
56+
4757 **Deliberately not done:** read-only rootfs (the workspace lives in the
4858 container filesystem precisely so no volume is ever attached; builds also
4959 write `$HOME` caches), and egress *filtering* (network is all-or-nothing).
⋯ 114 unchanged lines