anvilsign in

collin/anvil · 1f463de5

docs: CI artifact system design; check off browsing + rename TODOs

Collin Richards · 2026-06-10 08:28 UTC · 1f463de586fbfeb46c9a7f1d5a63230df5b3f3a7 · parent f32fb115 · browse files

modifiedTODO.md+15 −2
11 # Misc TODO
22
3-- [ ] do we have a way to view the source in a repo by branch or at a given commit?
4-- [ ] design a CI artifact system, keyed by commit
3+- [x] do we have a way to view the source in a repo by branch or at a given commit?
4+ - yes: `/{owner}/{repo}/tree/{rev}` always accepted any branch/tag/commit;
5+ what was missing was UI. Added: branch/tag switcher dropdown on the repo
6+ and tree pages, "browse files" link on the commit page, percent-encoded
7+ ref names so branches with `/` work, and an unborn-HEAD fallback so a
8+ repo whose HEAD names a missing branch no longer renders as empty.
9+- [ ] implement the CI artifact system _(designed — see `docs/ci-artifacts.md`;
10+ implementation order is at the bottom of that doc)_
511 - every push triggers CI on the tip commit (already true); a run should be
612 able to produce artifacts
713 - artifacts are stored per-commit and served from anvil (download from the
⋯ 4 unchanged lines
1218 - sketch needed: where artifacts live on disk, retention/GC, how
1319 `.anvil/ci.yml` declares artifact paths + metadata extractors, and how
1420 the broker gets files out of the sandboxed container
21+ - use rustdoc as an example when testing the feature: it generates a big
22+ HTML subtree, so support rendering/serving a whole HTML artifact subtree
23+ the way the pages feature does (rustdoc output as a served site)
1524 - [ ] repo mirroring to/from GitHub
1625 - push mirror: pushes to an anvil repo get forwarded to a configured
1726 GitHub remote
1827 - pull mirror (maybe): a repo that virtually mirrors a GitHub repo and
1928 just displays it here — periodically fetched, read-only on the anvil
2029 side
30+- [ ] per-repository issue tracker
31+- [x] rename the `anvil` crate to `anvil_cli` _(named it `anvil-cli` to match
32+ the workspace's hyphenated crate names; the binary is still `anvild`, so
33+ deploy scripts and docs needed no changes)_
2134
2235 ---
2336
⋯ 137 unchanged lines
addeddocs/ci-artifacts.md+158 −0
1+# CI artifacts — design
2+
3+Status: **design sketch, not implemented.** Companion to the CI runner in
4+`crates/anvil-ci` and the threat model in `docs/untrusted-mode.md`.
5+
6+## Goals
7+
8+- A CI run can produce artifacts (binaries, reports, generated docs).
9+- Artifacts are keyed by commit and served from anvil: downloadable from the
10+ run page and the commit page, with a "latest on branch" alias.
11+- A directory artifact can opt into being *browsable* — served as a static
12+ site rather than downloaded. The canonical test case is rustdoc: a big
13+ generated HTML subtree, viewable at a stable latest-on-branch URL.
14+- Jobs can declare how to extract metadata from artifacts (sizes, version
15+ strings, test/coverage numbers) so the UI can surface it next to the
16+ run/commit without downloading anything.
17+- The broker security model is unchanged: the job container still gets no
18+ socket, no mounts, no volumes.
19+
20+## How the broker gets files out of the container
21+
22+The checkout already goes *in* via the Docker API (`upload_to_container`, a
23+tar). Artifacts come *out* the same way: after `wait_container` returns and
24+before the container is removed, the broker calls `download_from_container`
25+(`GET /containers/{id}/archive?path=...`) for each declared path. That works
26+on a stopped container, needs no shared filesystem, and keeps anvil the only
27+Docker client.
28+
29+Rules:
30+
31+- Declared paths are resolved under `/workspace`; absolute paths and `..` are
32+ rejected at parse time.
33+- A path that is a directory arrives as a tar and is stored as
34+ `<name>.tar.gz`; a single file is stored as-is.
35+- Collection happens on success **and** failure (test reports matter most on
36+ red runs) but not after a timeout kill (the container is already gone).
37+ Each artifact records which it was.
38+- Per-artifact and per-run size caps come from config (below). The download
39+ stream is aborted, and the artifact skipped with a logged note, when a cap
40+ is exceeded — never the run failed retroactively.
41+
42+## Declaring artifacts in `.anvil/ci.yml`
43+
44+```yaml
45+image: rust:1.95-bookworm
46+steps:
47+ - name: build
48+ run: cargo build --release
49+ - name: test
50+ run: cargo test --workspace
51+
52+artifacts:
53+ - name: anvild # unique per pipeline; [a-zA-Z0-9._-]+
54+ path: target/release/anvild # file → download; dir → tar.gz download
55+ meta:
56+ version: ./target/release/anvild --version
57+ size: stat -c %s target/release/anvild
58+ - name: coverage
59+ path: coverage/
60+ meta:
61+ line_pct: jq -r .line_pct coverage/summary.json
62+ - name: doc
63+ path: target/doc/ # rustdoc HTML subtree
64+ browse: true # serve as a static site, don't download
65+```
66+
67+`meta` is a map of key → shell command. The commands run **inside the job
68+container** (appended to the script after the steps, still `set -e`-free —
69+each is best-effort), because artifact content is untrusted and must never be
70+executed or parsed on the host. Each command's stdout is trimmed and capped
71+(1 KiB); the resulting key→value map is stored as JSON on the artifact row.
72+A failed extractor stores nothing for that key and appends a note to the log.
73+
74+Implementation note: the extractor output travels in a well-known file the
75+broker downloads (e.g. `/workspace/.anvil-meta.json`, written by a generated
76+trailer in the script), so it rides the same archive mechanism as artifacts
77+and needs no log parsing.
78+
79+## Storage
80+
81+```
82+data_dir/artifacts/{repo_id}/{commit}/{name} # file artifact
83+data_dir/artifacts/{repo_id}/{commit}/{name}.tar.gz # dir, download-only
84+data_dir/artifacts/{repo_id}/{commit}/{name}/... # dir, browse: true
85+```
86+
87+- Keyed by `repo_id` (stable across renames) and full commit sha.
88+- Browsable directory artifacts are stored *extracted* so requests are plain
89+ file reads (no per-request untar); download-only directories stay tar.gz.
90+ Extraction rejects entries that escape the artifact root (`..`, absolute,
91+ symlinks) — the tar comes from an untrusted container.
92+- A re-run of the same commit overwrites that commit's directory.
93+- New Toasty model `CiArtifact`: `id`, `run_id`, `repo_id`, `commit`, `name`,
94+ `size`, `is_dir`, `browse`, `meta` (JSON string), `created_at`. (Toasty
95+ migrations still don't exist — this lands as a new table, which
96+ `db::connect` only creates on a fresh DB; same caveat as every schema
97+ change so far.)
98+
99+## Serving
100+
101+- Run page: artifact list (name, size, metadata chips) with download links —
102+ or a "browse" link for `browse: true` artifacts.
103+- `GET /{owner}/{repo}/ci/{run_id}/artifacts/{name}` — direct download.
104+- `GET /{owner}/{repo}/artifacts/{rev}/{name}` — alias: resolve `rev` (branch
105+ or commit) to the latest run with that artifact on that commit, redirect to
106+ the run-scoped URL. This gives "latest on branch" for free since CI runs on
107+ every push tip.
108+- `GET /{owner}/{repo}/artifacts/{rev}/{name}/{*path}` — browsable artifacts
109+ only: serve files from the extracted subtree exactly like `pages.rs` serves
110+ a `pages` branch (extension→content-type map, `nosniff`, `index.html`
111+ resolution with trailing-slash redirect so rustdoc's relative links work).
112+ E.g. `/{owner}/{repo}/artifacts/main/doc/anvil_core/` is always the default
113+ branch's latest rustdoc.
114+- Commit page and per-commit CI badges link through to the run's artifacts.
115+- Visibility follows the repo, like pages and CI logs.
116+- Download artifacts get `Content-Disposition: attachment` +
117+ `application/octet-stream` + `nosniff`. Browsable artifacts serve inline by
118+ design; that is the same stored-XSS-on-forge-origin exposure as pages
119+ (threat model §2) — acceptable single-tenant, and both move to a separate
120+ origin together before untrusted users.
121+
122+## Retention / GC
123+
124+Config, all under `[ci]`:
125+
126+```toml
127+artifact_max_mb = 256 # per artifact
128+artifact_run_max_mb = 512 # per run, summed
129+artifact_quota_mb = 4096 # per repo, summed; 0 = unlimited
130+```
131+
132+GC is deterministic and runs after each run's artifacts are stored: while the
133+repo is over `artifact_quota_mb`, delete the oldest commit-directory (and its
134+rows) — except directories that are the newest artifact-bearing commit of any
135+branch head, which are pinned. No background sweeper, no clocks to test; the
136+invariant holds whenever an artifact lands.
137+
138+Orphan cleanup (repo deleted → remove `artifacts/{repo_id}`) hooks into repo
139+deletion alongside the existing git-dir removal.
140+
141+## Out of scope (deliberately)
142+
143+- Cross-run caching (e.g. cargo registry/target caching) — different problem,
144+ different lifetime, mounts would pierce the sandbox.
145+- Artifact upload from outside CI (release uploads) — maybe later, different
146+ authz.
147+- Dedup/content-addressing — at our scale, per-commit copies are fine.
148+
149+## Implementation order
150+
151+1. Schema + config: `CiArtifact` model, `[ci]` caps, parse `artifacts:` in
152+ `ci.rs` (with path/name validation + tests).
153+2. Broker: collect declared paths via `download_from_container`, store to
154+ disk, write rows; meta-extractor trailer + `.anvil-meta.json` pickup.
155+3. Web: run-page list + download route, then the `{rev}` alias route, then
156+ the browse route (reusing the content-type/index helpers from `pages.rs`).
157+ Test with rustdoc on this repo (`cargo doc` → `target/doc`, `browse: true`).
158+4. GC + repo-deletion hook.