The gate harness (xtask harness)
The harness runs the repo's CI gates and decides, for each one, whether it passed, failed, did not run, or is unproven — and refuses to report success in any case where it cannot tell.
Design authority: docs/ci-test-architecture-v2.md §7.
This page is the operator's view.
cargo run --release -p xtask -- harness --list # the registry
cargo run --release -p xtask -- harness # run the T1 gates
cargo run --release -p xtask -- harness --only reject # run one
cargo run --release -p xtask -- harness --verify-falsifiers # incremental
cargo run --release -p xtask -- harness --verify-falsifiers --all # re-prove every gate
cargo run --release -p xtask -- harness --explain-inputs conformance
scripts/gates-for-change.sh --dry-run # which gates a change needs
Why it exists
The v0.19.13 release burned nine consecutive preflight attempts, none for a product reason, and the audit that followed found 23 gate defects. The shape of almost all of them was the same: a gate that could not fail.
xtaskexited 0 on an unknown subcommand — a typo'd gate name in CI was a permanently green no-op.verify-all-web.shusedif node … | tail -8; then, testingtail's exit status. The console e2e gate could not fail.grep -qE "0 fail"matched inside"10 fail", so a run ending0 pass / 12 failpassed.SKIPcounted aspass; nightly reported29 passed, 0 failedwith three examples never built.doc-examples.shwithtotal=0printed0/0 … GATE: PASS.ty/tests/reject.rsasserts>= 13against an actual 63 — deleting 50 corpus files keeps it green.
Every rule below is a direct answer to one of those.
States
| State | Meaning | Suite effect |
|---|---|---|
PASS | ran, every assertion held, assertions > 0, and the count matched exactly | PASS |
FAIL | an assertion broke, the budget was exceeded, the gate was vacuous, or the body could not be spawned | FAIL (exit 1) |
NOT RUN | registered and selected, but no usable verdict | UNKNOWN (exit 3) |
UNPROVEN | passed, but its falsifying mutation is unproven or stale | UNKNOWN (exit 3) |
NOT APPLICABLE | outside the selected tier/platform, or deselected by --only | none |
BLOCKED | declared unrunnable, with an issue and an expiry | none — until the expiry, then FAIL |
Three properties follow, and they are the point:
- A suite containing
NOT RUNorUNPROVENcan never renderPASS. A run that cannot say whether a gate passed has not passed. --onlyproducesNOT APPLICABLE, neverNOT RUN. Deliberate selection is not an unknown. Conflating them makes local runs emitUNKNOWNconstantly, which trains people to ignore the one state that means "we do not know".- Rows come from the registry, not from the run. A gate cannot disappear by not executing. This kills "SKIP counted as pass" at the root.
Ordering: prove falsifiers BEFORE regenerating the coverage ledger
docs/coverage/falsifier-proofs.json is an input to xtask coverage-ledger
— a surface covered by a gate whose mutation is recorded PROVEN scores
Falsified (4) rather than Asserted (3). So proving a falsifier legitimately
changes the ledger, and running the two in the wrong order leaves the checked-in
ledger stale.
cargo run --release -q -p xtask -- harness --verify-falsifiers # 1. prove
cargo run --release -q -p xtask -- coverage-ledger # 2. regenerate
cargo run --release -q -p xtask -- coverage-ledger --check # 3. confirm
This is not a wrinkle to work around — it is the ledger noticing that the
coverage claim actually improved. --check reporting STALE after a falsifier
sweep is the mechanism working.
BLOCKED — a deadline, not a parking space
A soft skip is a gate that quietly stops asserting and keeps reporting
non-failure for ever. That is the class this harness exists to kill, which is
why BLOCKED did not exist here at first even though the design declared it.
It is admitted only with four teeth, all enforced in code:
pub static BLOCKED: &[Blocked] = &[Blocked::new(
"gate-name",
"https://github.com/anzellai/sky/issues/NNN", // required, non-empty
"2026-12-31", // required, YYYY-MM-DD
"the structural obstacle — not \"flaky\", not \"todo\"",
)];
- Declared at compile time.
Blocked::newisconst fn; an empty issue, a missing reason or a non-YYYY-MM-DDexpiry fails the build, exactly asMutations::new(&[])does. - It expires by itself. From the declared date the gate renders
FAIL, with nobody in the loop to forget. A malformed date reads as expired, not far-future — a typo must not buy an unbounded block. - It never renders
PASS. - Its surfaces count as UNCOVERED in the coverage ledger
(
GateState::counts_as_cover()is false). This is the property that removes the incentive: blocking never preserves a coverage number, it lowers one.
The suite verdict is neutral before expiry, deliberately. A state that turns CI permanently red is a state people delete rather than fix, and deleting the row restores exactly the invisible absence that rendering rows from the registry was meant to end.
A registry test forbids blocking any product-tier gate. Blocking a T0-T4 gate silently removes coverage, so it must land together with the ledger row that shows the surface going uncovered, and the test relaxed deliberately in the same commit.
selftest-blocked is a permanent witness whose body would pass — so what it
proves is that a gate which would pass still does not render PASS while it is
blocked. With an empty BLOCKED list nothing would exercise the mechanism and
it would rot like the gates it polices.
Every gate declares a falsifier, and the compiler enforces it
mutations: Mutations::new(&[Mutation {
id: "reject.neutralise-axis",
description: "neutralise the axis under test so the file type-checks",
kind: MutationKind::ReplaceOnce { path: "…", from: "add 1 2 3", to: "add 1 2" },
}]),
Mutations::new is a const fn whose assert! is const-evaluated, so an empty
set fails the build:
error[E0080]: evaluation panicked: every gate must declare at least one
falsifying mutation (a gate that cannot fail is worse than no gate)
--verify-falsifiers then proves the mutation actually bites:
- run the gate — the baseline must be green, or the result is
INCONCLUSIVE(a red baseline says nothing about what the mutation did); - apply the mutation — exact-once replacement; a pattern that is missing or
ambiguous is refused, because a mutation that silently did nothing would
report
VACUOUSand be misread as a gate defect; - run again under the same
killpg-backed budget — red ⇒PROVEN, green ⇒VACUOUS; - revert, guaranteed — the patch reverts in
Drop, including on panic.
Proofs are recorded in docs/coverage/falsifier-proofs.json, one record per
gate: as declared only when every declared mutation behaved as declared (a
gate with two mutations used to be recorded by whichever ran last). A gate whose
proof is missing, older than the window, or whose inputs changed since it was
taken renders UNPROVEN under --require-proofs.
Incremental proofs — a proof is re-taken only when its inputs change
A full sweep re-proves every gate: each baseline and each mutated run, with an
xtask rebuild per Rust-source mutation. Measured locally it is over an hour,
and almost all of it re-establishes proofs nothing touched. So a local
--verify-falsifiers is incremental. Each record carries an inputs_hash,
the digest of everything that decides whether the proof still holds:
- the gate's registration — name, tier, platforms, budget, exact assertion
count, expectation, and every mutation's id, description, target,
fromandtotext; - the mutation targets, byte for byte;
- the gate's body source — its function and the transitive closure of the
helper items it calls in
harness/bodies.rs,harness/layer2.rsandmain.rs(comments stripped, so an edited comment does not force a re-proof), plus everyxtaskmodule it calls into (crate::corpus::…is all ofsrc/corpus/); - its fixtures and scripts — every tracked path named by a string literal in
that closure or in those modules, and every tracked path a named script names
in turn (one level:
conformance.shpulls in thelib/*.shit sources and the suites it runs). A script's references into the toolchain (rust/,runtime-go/,sky-stdlib/,docs/, …) are not followed — the compiler as a whole is what--allis for; - the falsifier runner (
harness/falsify.rs,harness/child.rs).
Only tracked files count, by their working-tree content, so a digest is a
function of what a commit contains and CI computes the same one. The proof
ledger itself is never an input (the coverage-ledger gate reads it).
A run re-proves a gate when its digest changed, when it has no record, when the record is not as declared or names a retired mutation, when it is out of the window, or when either side has no digest (a legacy record, or inputs that could not be resolved — an error always means "do the work"). Every other gate is carried, and the run says so:
RE-RUN 2 gate(s), CARRIED 39 gate(s) (inputs unchanged since their recorded proof)
re-run canary: the canary is re-run by every falsifier run
re-run apps-ledger: its inputs changed since the proof
carried roundtrip: proof taken 0d ago, inputs digest unchanged
The canary runs in every falsifier run, whatever the selection. --all carries
nothing; the nightly and the release pass it, so a whole-system interaction the
digest does not model is still re-proven before a tag.
--explain-inputs <gate> prints a gate's input paths, its body closure and its
digest — use it to see why a gate re-proved, or to check a new gate's inputs.
Because --require-proofs (the release tier lines) rejects a proof whose digest
no longer matches the tree, a change to a gate's inputs must land with the
re-proved ledger: run the incremental --verify-falsifiers (it is the last
step scripts/gates-for-change.sh plans) and commit
docs/coverage/falsifier-proofs.json.
Where the proofs are checked
A recorded proof is evidence about a (gate, mutation) pair, and two things can rot it. Both are now caught, at two depths:
-
From-pattern drift, statically, per-PR. If a mutation's
ReplaceOnce.fromstring drifts out of its target file (a literal reworded, a!= nilflipped tonil !=), the recordedPROVENcredits a falsification that can no longer be reproduced.coverage-ledger --checkasserts every proven proof'sfromstill occurs exactly once in its target — a grep over the cited files, no build — so a drift reddens the per-PRconfig-gatesjob, not just a release. -
Does-it-still-bite, deeply, nightly.
--verify-falsifiersactually applies each mutation and confirms the gate reddens. It rebuildsxtaskper Rust-source mutation and runs every selected gate twice (baseline + mutated), so it is the deepest and slowest check here. It runs with--allin the nightlyfalsifier-verificationjob (nightly-sweep.yml,--tier t1) and in the release workflow'sgate-falsifiers-1…-6jobs, whose--onlylists together cover every registered gate of every tier. It accepts--tier/--onlyto verify one slice at a time; with neither it sweeps the whole registry. Locally it is incremental (above).
Which gates to run: the narrow set locally, the full suite at merge and release
- Per change, locally:
scripts/gates-for-change.sh. It diffs against a base (defaultorigin/main; committed, staged, unstaged and untracked paths all count), maps each path to the narrowest gates that exercise it, prints the plan and runs it (--dry-runprints only). For example:runtime-go/rt/**→go test ./rt/...+ the browser tier + the Sky.Live client e2e;rust/crates/project/src/spa_*→cargo test -p project+ the Sky.Spa e2e +spa-diff-fuzz;rust/crates/{syntax,hir,ty,lower,codegen}/**→ the crate's tests + the corpus gates + coerce-floor + T2 + build-run + the example sweep + conformance;docs/**→doc-examples.sh;examples/<x>/**→ roundtrip +build-run --only=<x>+ the sweep. A path no rule claims is printed asUNMAPPED. The last step is always the incremental falsifier run. - At a merge to
mainor a release tag: the release workflow's full suite (.github/workflows/release.yml), which must be green before the tag. It runs every tier — workspace tests, T1-T4, every falsifier proof with--all, the census--checks, the full clean-slate example sweep (build and run), build-run, conformance, verify-cli, the browser tier, every e2e script, the doc examples andgo test -race— in concurrent jobs thatrelease:waits on. Nothing is deferred to the nightly.tests/workflows_parse.rsfails the build if agate-*job is disabled or drops out ofrelease: needs:.
The gate build cache
The sweep, the browser gate, the ui-showcase gate, the e2e scripts and
build-run's sky build paths share scripts/lib/gate-build-cache.sh, a
content-addressed cache of built projects. The key is the compiler binary's
hash (which carries the baked embed fingerprint), the project's absolute path
and full content (everything but build output and run-time state), the sky
arguments and artefact paths, go version plus the Go env that changes codegen,
and the SKY_* / CGO_* environment. A hit restores what a clean build of
identical inputs stored (an APFS clone on macOS); anything else builds.
It never serves what a fresh build would not produce: a changed source file, a
rebuilt compiler, other arguments or another toolchain miss; only a --clean
build stores; a project with floating [go.dependencies] (network-resolved
"latest") is never cached; the callers still refuse a stale compiler
(require_fresh_compiler). SKY_GATE_CACHE=off disables it (the default on a
GitHub Actions runner, where each job builds a project once),
SKY_GATE_CACHE_DIR moves it, SKY_GATE_CACHE_MAX_MB (default 8192) bounds
it, oldest-used first, and nothing is stored while the disk has under
SKY_GATE_CACHE_MIN_FREE_MB (default 20480) free. e2e fixtures build in stable per-worktree directories
under the cache directory, because the key includes the project path.
scripts/build.sh no longer wipes the Go build cache on every run past 5 GB;
it wipes it only under disk pressure (over 5 GB with under 30 GB free), or past
20 GB — the policy example-sweep.sh already used.
The canary
One gate — canary — is deliberately vacuous and paired with a no-op patch.
A correct runner must report it VACUOUS. Reporting PROVEN means the harness
applied its patch somewhere the gate never read, or is not reading the verdict
from the run it just performed — in which case every other PROVEN is worthless.
It is the only place a passing gate is the success signal, and the only construction that catches a verifier whose every answer is "green".
Budgets are enforced by killpg, not by hope
Gate bodies run in a child process placed in its own process group
(process_group(0)). On budget expiry the harness sends SIGTERM then
SIGKILL to the group.
This is not stylistic. Gates spawn go builds, servers, PTYs and browsers:
- A thread's children are not reachable as a group, so "kill the process group" is unimplementable from a thread — a timeout would leak a process holding a port into every later gate.
- An orphaned worker can write a result after its gate was recorded FAIL, corrupting a later gate's verdict — a wrong green, attributed to the wrong gate. Results are therefore generation-stamped, and a result whose generation does not match the gate being awaited is discarded.
Measured negative control: replacing the killpg with a plain kill of the
direct child leaves the body's sleep 600 grandchild alive.
Timeouts live in the harness and never in GNU timeout, which is absent on
every macOS runner — the exact hole that left conformance.sh unbounded there.
Wrapped verifiers emit JSON; the gate reads the file
No verifier is rewritten. Each gains a --json <path> mode, and the gate asserts
on the file — never on scraped stdout, which is where the unanchored-grep
class comes from.
scripts/conformance.sh --json <path>writes a manifest of(suite, exit_code, per-suite Sky.Test report)and deliberately does not aggregate. It emits only values it controls and never parses JSON; a shell that parses its own output is howgrep -qE "0 fail"came to match"10 fail". Theconformancegate aggregates in a real JSON parser.scripts/verify-cli.sh --json <path> --rebuildwrites one record per entry.--rebuildis mandatory for gate use: without it the script only builds an example whose binary is missing, so it certifies whatever artefact an earlier run left behind, and no source mutation can falsify it.
Per-case data comes from the Sky.Test JSON reporter — see
testing.md.
Assertion counts are exact
Every gate pins an exact expected count, never a >=. A corpus that shrinks
is a failure with an actionable message, not a quieter green.
Pinning exact counts found a real discrepancy on its first run: v2 §5.4 records
conformance at 772 cases, from a static count of Test.test leaves. The
harness measured 770. Both are correct about different things —
StoreConformanceTest.sky:75 and StoreCrudConformanceTest.sky:68 each declare
a Test.test "setup" leaf inside the Err arm of case setup () of, which
materialises only when the DB setup fails. The gate pins the number that runs.
Concurrency
Gates run sequentially. Deliberate: the failure that motivated this whole
mandate is a parallel sweep spawning thousands of xcrun processes and
exhausting the per-uid process table (measured 2,167 of 2,472), which kills
mem-guard's ability to fork. Parallelism and the persistent semaphore belong
with the CI-topology phase, when there are measured runner numbers to budget
against.
Exit codes
| Code | Meaning |
|---|---|
| 0 | PASS |
| 1 | FAIL |
| 2 | usage error (including an unknown gate name — never an empty selection that passes) |
| 3 | UNKNOWN — a NOT RUN or UNPROVEN gate |