anvilsign in

collin/anvil

RenderedSource

CI artifacts — design

Status: implemented (2026-06-10). Companion to the CI runner in crates/anvil-ci and the threat model in docs/untrusted-mode.md. Deviations from the original sketch are noted inline.

Goals

  • A CI run can produce artifacts (binaries, reports, generated docs).
  • Artifacts are keyed by commit and served from anvil: downloadable from the run page and the commit page, with a "latest on branch" alias.
  • A directory artifact can opt into being browsable — served as a static site rather than downloaded. The canonical test case is rustdoc: a big generated HTML subtree, viewable at a stable latest-on-branch URL.
  • Jobs can declare how to extract metadata from artifacts (sizes, version strings, test/coverage numbers) so the UI can surface it next to the run/commit without downloading anything.
  • The broker security model is unchanged: the job container still gets no socket, no mounts, no volumes.

How the broker gets files out of the container

The checkout already goes in via the Docker API (upload_to_container, a tar). Artifacts come out the same way: after wait_container returns and before the container is removed, the broker calls download_from_container (GET /containers/{id}/archive?path=...) for each declared path. That works on a stopped container, needs no shared filesystem, and keeps anvil the only Docker client.

Rules:

  • Declared paths are resolved under /workspace; absolute paths and .. are rejected at parse time.
  • A path that is a directory arrives as a tar and is stored as <name>.tar.gz; a single file is stored as-is.
  • Collection happens on success and failure (test reports matter most on red runs) but not after a timeout kill (the container is already gone). Each artifact records which it was.
  • Per-artifact and per-run size caps come from config (below). The download stream is aborted, and the artifact skipped with a logged note, when a cap is exceeded — never the run failed retroactively.

Declaring artifacts in .anvil/ci.yml

image: anvil-runner:rust
steps:
  - name: build
    run: cargo build --release
  - name: test
    run: cargo test --workspace

artifacts:
  - name: anvild                     # unique per pipeline; [a-zA-Z0-9._-]+
    path: target/release/anvild     # file → download; dir → tar.gz download
    meta:
      version: ./target/release/anvild --version
      size: stat -c %s target/release/anvild
  - name: coverage
    path: coverage/
    meta:
      line_pct: jq -r .line_pct coverage/summary.json
  - name: doc
    path: target/doc/               # rustdoc HTML subtree
    browse: true                    # serve as a static site, don't download

meta is a map of key → shell command. The commands run inside the job container (appended to the script after the steps, still set -e-free — each is best-effort), because artifact content is untrusted and must never be executed or parsed on the host. Each command's stdout is trimmed and capped (1 KiB); the resulting key→value map is stored as JSON on the artifact row. A failed extractor stores nothing for that key and appends a note to the log.

Implementation note: each extractor's stdout lands in its own file under /tmp/anvil-meta/<artifact>/<key> (written by a generated script trailer that runs even when a step fails — the steps execute in a subshell whose exit code is preserved). The broker downloads that directory via the same archive mechanism as artifacts; file-per-value avoids shell JSON-escaping entirely.

Storage

data_dir/artifacts/{repo_id}/{commit}/{name}            # file artifact
data_dir/artifacts/{repo_id}/{commit}/{name}.tar.gz     # dir, download-only
data_dir/artifacts/{repo_id}/{commit}/{name}/...        # dir, browse: true
  • Keyed by repo_id (stable across renames) and full commit sha.
  • Browsable directory artifacts are stored extracted so requests are plain file reads (no per-request untar); download-only directories stay tar.gz. Extraction rejects entries that escape the artifact root (.., absolute, symlinks) — the tar comes from an untrusted container.
  • A re-run of the same commit overwrites that commit's directory.
  • New Toasty model CiArtifact: id, run_id, repo_id, commit, name, size, is_dir, browse, meta (JSON string), created_at. (Toasty migrations still don't exist — this lands as a new table, which db::connect only creates on a fresh DB; same caveat as every schema change so far.)

Serving

  • Run page: artifact list (name, size, metadata chips) with download links — or a "browse" link for browse: true artifacts.
  • GET /{owner}/{repo}/ci/{run_id}/artifacts/{name} — direct download.
  • GET /{owner}/{repo}/artifacts/{rev}/{name} — alias: resolve rev (branch or commit) to the latest run with that artifact on that commit, redirect to the run-scoped URL. This gives "latest on branch" for free since CI runs on every push tip.
  • GET /{owner}/{repo}/artifacts/{rev}/{name}/{*path} — browsable artifacts only: serve files from the extracted subtree exactly like pages.rs serves a pages branch (extension→content-type map, nosniff, index.html resolution with trailing-slash redirect so rustdoc's relative links work). E.g. /{owner}/{repo}/artifacts/main/doc/anvil_core/ is always the default branch's latest rustdoc.
  • Commit page and per-commit CI badges link through to the run's artifacts.
  • Visibility follows the repo, like pages and CI logs.
  • Download artifacts get Content-Disposition: attachment + application/octet-stream + nosniff. Browsable artifacts serve inline by design; that is the same stored-XSS-on-forge-origin exposure as pages (threat model §2) — acceptable single-tenant, and both move to a separate origin together before untrusted users.

Retention / GC

Config, all under [ci]:

artifact_max_mb = 256        # per artifact
artifact_run_max_mb = 512    # per run, summed
artifact_quota_mb = 4096     # per repo, summed; 0 = unlimited

GC is deterministic and runs after each run's artifacts are stored: while the repo is over artifact_quota_mb, delete the oldest commit-directory (and its rows) — except directories that are the newest artifact-bearing commit of any branch head, which are pinned. No background sweeper, no clocks to test; the invariant holds whenever an artifact lands.

Orphan cleanup (repo deleted → remove artifacts/{repo_id}) will hook into repo deletion when that feature exists — anvil currently has no way to delete a repository at all, so there is nothing to hook yet.

Out of scope (deliberately)

  • Cross-run caching (e.g. cargo registry/target caching) — different problem, different lifetime, mounts would pierce the sandbox.
  • Artifact upload from outside CI (release uploads) — maybe later, different authz.
  • Dedup/content-addressing — at our scale, per-commit copies are fine.

Implementation order

  1. Schema + config: CiArtifact model, [ci] caps, parse artifacts: in ci.rs (with path/name validation + tests).
  2. Broker: collect declared paths via download_from_container, store to disk, write rows; meta-extractor trailer + .anvil-meta.json pickup.
  3. Web: run-page list + download route, then the {rev} alias route, then the browse route (reusing the content-type/index helpers from pages.rs). Test with rustdoc on this repo (cargo doc → target/doc, browse: true).
  4. GC + repo-deletion hook.