diff --git a/docs/INFRA-REGISTER.md b/docs/INFRA-REGISTER.md index 137c9af3..6762b55e 100644 --- a/docs/INFRA-REGISTER.md +++ b/docs/INFRA-REGISTER.md @@ -4783,6 +4783,33 @@ the wrapper's harness children orphaned + running") — so IR-80 ships callers, `e2e-leaked-daemons-shared-box` (memory). - **Ripe when:** arm 1 now; arms 2-4 on the operator's answer; arm 5 at the next `golden.yml` touch after any of them. · **Size:** arm 1 a measurement, arm 4 one census line, arm 5 two literals. +- **AMENDMENT 2026-09-09 (two terms this entry did not name when it landed at `d32d5c4c`).** + 1. **`ci.yml`'s `changes` job runs `unit` on BOTH runners for every push to `main`.** The classify + step at `.github/workflows/ci.yml:49` emits `code=true` for every non-`pull_request` event, so + the docs-only skip that PR runs #206/#207 demonstrated is **`pull_request`-only** — read from + `ci.yml` at `main` by doyle, who names it his own error after twice ruling the opposite from the + PR runs alone. Consequence for this entry's wall clock: run **34337797758** (the ff of + `b66a9612`, a docs-only delta over golden-green `a2f335f8`) ran `unit (Windows)` 09:59:57Z → + 10:40:48Z and **completed FAILURE at the 40-minute job wall** (step 6 `Unit tests` 10:08:23 → + 10:40:03) — a red on `main` at a sha whose content cannot fail a unit test. It is an IR-85 + face, not a flake row. doyle's cancel of it returned `Cannot cancel a workflow run that is + completed`: the run ended on its own and was ruled VOID after the fact, recorded that way + deliberately, because a register that credits a controlled action nobody performed teaches the + next reader that the box was under control. + 2. **A near-full volume is an environment term for the slowdown**, alongside the Defender + first-touch tax and the operator-desktop load already recorded. **Deliberately unquantified:** + the discriminator window that would have apportioned it was VOID, because the one leg that + completed ran under both a CI job and a falling disk. See **IR-90**. +- **WHAT THE VOIDED DISCRIMINATOR ESTABLISHED, negatively (hertz, 2026-09-09).** The fs-heavy + slowdown REPRODUCES at `04e32c8c`, which predates the +11,388-line head growth: spt-daemon + `concurrent_writes` 105.420 s, `two_tier_sync` 58.362 s, spt-store `clone_copies` 65.657 s, + `different_monics` 73.967 s, `syncmerge reconciled_write` 86.765 s, against the 09-06 baseline of + 22.4 / 18.7 / 17.7 / 18.2 / 27.3 s — **3.2x to 4.7x on all five, at the OLD sha**. So head growth + is **not NECESSARY** for the slowdown. That is the whole of it: it does NOT measure how much the + environment explains, and it is not evidence about the head arm at all — that leg ran inside run + 34337797758's window and on a volume that reached 0.018 GiB free, and the head arm never produced + a comparable pair. **Arm 1 stays OPEN**, its re-run deferred until **IR-90**'s guard exists and a + window with no `main` push and no CI job on the box can be scheduled. ### IR-86 — golden's 32 GiB floor is BELOW the measured 67.4 GiB Windows suite footprint, so a start floor passes a box that cannot finish @@ -4971,3 +4998,75 @@ the wrapper's harness children orphaned + running") — so IR-80 ships callers, the product"; the layer is named only by the one measurement that decides it (same listener, LAN vs Tailscale), which the failure text prints as a recipe. A guard that names a cause it did not measure is the defect class this whole entry documents. + +### IR-90 — a full disk on the self-hosted box reds a rig as a PRODUCT refusal, not a build error, and no rig or gate records the free space that would falsify it + +- **Status:** OPEN, filed by hertz 2026-09-09 from the disk-full incident on hfenduleam. · + **Origin:** the IR-85 discriminator lane's head arm, killed by the volume rather than by the code + under test. +- **Symptom:** `spt-daemon::sync concurrent_writes_reconcile_on_elected_node_and_converge` FAILED at + 67.588 s with a panic in our own test at `crates\spt-daemon\tests\sync.rs:198`: + `pull: Custom { kind: Other, error: "sync refused: bundle failed: git -C \tracked-b\.seed.git + bundle create \scratch\serve\serve-pull-6.bundle ^d12a7134... a-doyle failed (exit Some(1)): + fatal: sha1 file '' write error. Out of diskspace\nerror: pack-objects died" }`, then a + second panic at `:213` (`pull thread: Any { .. }`) as the harness thread unwound. +- **Cause:** `C:` was at **0.018 GiB free of 1862.02 GiB (0.00%)** at that instant + (`Win32_LogicalDisk`, 10:57Z). `git bundle create` could not write; the daemon's serve path turned + that into its real product refusal string `sync refused: bundle failed`; the test asserted on the + refusal. **Every layer behaved correctly, and the report reads as a sync regression at the sha + under test.** +- **Why it is worse than the disk faces already known:** the two banked faces + (`disk-full-reds-as-lnk1318-pdb-error`) are TOOLCHAIN costumes — `LNK1318` at link, and rustc + I/O before any link — and both route a reader to "the box is sick". This one routes to a CODE + OWNER: it names our file, our line and our refusal string, and the disk word sits in the FOURTH + nested clause behind a git exit code. It arrived mid-discriminator with an old sha and a new sha + side by side, where the cheapest reading — "the head arm failed, the old arm passed" — is a + head regression that does not exist. +- **Blast radius, same incident:** the volume also carried a live `unit (self-hosted, Windows, + hfenduleam)` job (run 34341010297) compiling into it, which completed FAILURE on its own and was + ruled VOID after the fact. `spt daemon status` went to peer pump last tick 185 s with 24 brain + subscribers stall-evicted (last evict 10:51:22Z, inside the disk window), with an operator restart + under consideration — a restart that risks the releases#287 shell stranding and that, had it + appeared to help, would have taught everyone the wrong cause. +- **⚠ Two things that looked like disk fallout and are NOT**, recorded so this entry does not + overclaim. (a) `serve list` returning `SERVE_UNCONFIRMED`: doyle read it from source — + `servehost.rs` and `KIND_SERVE_REQUEST` are ABSENT at 0.67.0/0.67.1 and were minted by 0.68.0, and + the resident broker is still the 0.67.0 image after the brain-only flip, so the CLI's 10 s bound + times out **by design** against an older daemon. (b) `peer reachability: DEGRADED, 4 of 7 peers + unreachable for 882775 s` — 10.2 days, predating the incident entirely. Even the pump term is + claimed as CONCURRENT, not caused: hertz called the pump recovered off ONE post-reclaim sample + (54 s) and doyle falsified it two minutes later (133 s), because a single reading of a monotonic + "last tick N s ago" counter cannot separate STILL TICKING from TICKED ONCE, and a fresh project + index proves the coordinator loop, not the pump. +- **Remedy — one guard, two placements, and it REFUSES rather than runs:** + 1. **Rig start.** For the two-host rig and any test rig that shells out to `git bundle` or writes a + store: read free space on the volume holding the rig's temp root and `target`, and under a floor + fail immediately with a message that says DISK and prints the number — never enter the + ceremony. Floor: start at **10 GiB**, above the largest single artifact these rigs write and far + below any healthy state of this box; tune only with a measurement. + 2. **Failure text**, for the case where space runs out MID-run and no start check can catch it: + when a shelled-out git/store operation fails, append the current free space to the error before + it becomes a product refusal string, so the panic that reaches a human already carries the + falsifier. +- **WHY BOTH, measured rather than argued (hertz, 2026-09-09, the #209 lane):** a floor read ONCE at + rig start would have PASSED that run and the run still ended nearly empty — `cargo check` started + at **136.32 GiB** free, the test-profile build of 221 binaries took it to **40.18 GiB in ten + minutes**, the legs finished at **16.24 GiB**, and it read **12.27 GiB** ninety seconds later. A + start-only guard is a guard against yesterday's disk. That last ~4 GiB fell with those legs already + terminal; the candidates are the runner's CI job and its `_work` tree, unsampled and therefore + UNATTRIBUTED. +- **Explicitly NOT the remedy:** a bigger disk, or a reap schedule. ~195 GiB of headroom was consumed + by ordinary work, ACCRUED across a ~45-minute build phase; **no reading bounds a rate** — the + free-space observations either side are endpoints of an accrual, and a rate derived from them + retargets a hunt (deployah's went to runaway logs, VSS and torrent preallocation on the strength of + a `>150 MB/s` derived that way, while the measured live box-wide write rate was ~1.5 MB/s). What + consumed it: two `cargo nextest run -E <5-test filter>` lanes at **82.88** and **64.44 GiB** — a + filter narrows the RUN, never the BUILD, and both built 221 test binaries under the `test` profile + — plus 40.93 GiB of `C:\actions-runner\_work` and 9.8 GiB of `%TEMP%`. Any headroom this box + has is two cold pools away from gone, so the guard must be a REFUSAL, not a budget. +- **Kin:** **IR-85** (the same box's fs-heavy slowdown; a near-full volume is now a named environment + term there), **IR-86** (a start floor that passes a box which cannot finish — the same defect one + layer up, and this entry is its mid-run half), **IR-76** (the golden runner is the operator's + desktop), `disk-full-reds-as-lnk1318-pdb-error` (memory; this is its third face and its standing + gap), `test-profile-pool-outgrows-the-disk-floor`, `free-space-floor-blocks-golden`. +- **Ripe when:** now. · **Size:** one assertion plus one error-context append, both test-side.