# Infra register — CI / build-pipeline debt

Operator-ruled 2026-08-02: infrastructure and CI-pipeline work items live HERE, not on the
`spt-bs-releases` board. The board carries product surface the operator triages; this register
carries what the gater triages. **Mandate: doyle sweeps this file at every milestone intake and
every release close, and composes ripe entries into waves/milestone riders.** An entry leaves
this file only by being built (link the lane) or being retired with a stated reason.

Entry format: status · origin · what/why · trigger condition (what makes it ripe) · size guess.

Last sweep: 2026-08-18, KEYSTONE (#182) INTAKE — the prior sweep's four composition rulings
EXECUTED (wave map: `KEYSTONE-182-JIT.md`, base main @`8248bc3`): (1) IR-9 pin lane DISPATCHED to
hertz (W0 item 1; interim kitsubito-clippy-authoritative dies when it lands). (2) IR-21 remedy
(2) + IR-39 vs `f24e732` DISPOSED: the branch is a BEACHHEAD (golden.yml prebuild step = the
prebuild-made-a-rule arm, ONE fail-fast site, one ledger row) — hertz rebases it off its
abandoned parent `b483699` (the id-collision draft; content landed renumbered as IR-46 @`8248bc3`)
and lands it in W0; the class remainder (IR-39's SHARED precondition helper + 24-site sibling_bin
sweep, the 33 cross-package build-edge expressions) folds into the test-hygiene lane. (3)
Test-hygiene family DECIDED-ACTIVATED as a dedicated hertz lane inside the KEYSTONE window,
sequenced after W0 (members: IR-13/23/36/37/38 + IR-21 clarity half + IR-39 helper +
twohost.rs:394 doc comment). (4) First-execution-cells discipline restated in the JIT's
golden-head step. Board members #84/#85 ride hertz's W0; #166/#57/#185 are todlando's W1. IR-46
id-collision RULED this intake (renumber-and-land, parallel issuance not authored disagreement;
deployah landed it @`8248bc3`). No new entries filed; no entries retired.

Prior sweep: 2026-08-19, NAMEPLATE (#181) / v0.56.0 RELEASE CLOSE — shipped c91 @`60d74ea`
(tag == golden-tested sha, run 32209922535; one respin, both first-golden reds ruled rig defects).
Filed IR-42 (pool-claim writes / build enforces — BUILT in this same commit, AGENTS.md line),
IR-43 (knock NoReply past its 30s carrier bound), IR-44 (perch-sentinel comment overclaims
preservation), IR-45 (twohost rig home premise; SETTLED + RETIRED
2026-08-19 — pump paths resolve under the per-run temp root, see the entry; #189 corrected on the
board, comment 5337519936). Two cycle findings ruled RUNBOOK-homed and
landed in this commit rather than as entries: the first-execution-cells intake question
(RELEASE-RUNBOOK golden-head intake — name the never-executed cells before the run) and
deployah's sweep-vs-cascade mechanism (RELEASE-RUNBOOK board step — under golden CI the cascade
is DRIVEN via `state <mref> acceptance`, never swept). ⚠ LABELLED HOLE — CLOSED UNDERIVABLE
(2026-08-19): the close commune's batch list named "alchemy create-races"; its content did not
survive the author's context reset and was not recoverable from #181, the JIT records, or memory.
Deployah answered the query: he ran ZERO create ops at the cut (could not have witnessed a create
race), a fresh probe over the milestone window shows no duplicate mints (the 4-issues-in-2s batch
mint is batching, not duplication), and the only surviving trace is doyle's own pre-reset message
naming the item as already-known — a pointer, not a sighting. Item DROPPED; the hole stands as
the record. Deployah's sweep-vs-cascade ship-path trap (#181 comment 5337402639, runbook-homed
above) is explicitly NOT this item's content — do not fold it in. Composition: IR-9's pin lane
(`rust-toolchain.toml` @ 1.96.0) goes to hertz AT KEYSTONE #182 INTAKE per the 2026-08-05
ruling; IR-21 remedy (2) + IR-39's precondition helper compose with hertz's standing
fixture-prebuild-hardening branch `f24e732` — disposition at the same intake; the test-hygiene
family (IR-13/23/36/37/38 + IR-21's clarity half) stays the dedicated post-batch lane candidate,
decision at intake; IR-29's proving run + IR-30's instrument lanes ride the next golden batch.
IR-2's trigger explicitly NOT met (queued unlanded lanes exist: IR-29/IR-30 instruments,
four-arm refusal eprintln, f24e732). IR-14 hygiene movement: `.worktrees/nameplate-asm-2bd36f1`
reaped this sweep (+66.98 GB by FS delta 120.53→187.51; claim `gate-w7-courtesy` base `fd3dc5a`
in main = finished lane; zero inbound reparse points) — stale-lane audit itself still open.

Prior sweep: 2026-08-05, LOCKSMITH tranche-2 (#141) CLOSE — golden run 30971976024 green on all 9
jobs, main ff'd to `0a25b77`, v0.55.0. IR-40 filed (stale-resume-brief + early-informant class).
IR-9's decision trigger FIRED 2026-08-05: doyle read the runner-account versions off this run's
`test` legs and RULED — pin in-repo via `rust-toolchain.toml` @ 1.96.0, kitsubito's clippy leg
authoritative in the interim; pin lane to hertz at next intake (see the entry). IR-41 filed the
same night (queued main run superseded without a record; runbook step 3 corrected in the same
commit). IR-1/IR-4's golden-only steps (link probe both boxes, toolchain
print both legs) had their FIRST EXERCISE here, discharging the "unexercised until a golden run"
caveat at the CI-RIDER LANE STATE foot section. IR-35's re-measure rode the batch (`c65b838`).
Owlery-noun thin lane SCOPED and dispatched to hertz for the next batch (class A only, two sites;
class B on-disk rename is an explicit non-goal — see IR-40's kin discipline for why the boundary is
written into the brief rather than left to judgement).

---

## OPEN

### IR-1 — Quiet predicate needs a network axis (tailscale RTT probe)
- **Status:** open, LANE EXISTS UNLANDED — see [[CI-RIDER LANE STATE]] at the foot of this file;
  carried by `bf8c4a2` → `REQ-CI-LINK-HEALTH-PROBE`, mapping CONFIRMED by builder hertz 2026-08-03
  (by content: tailscale ping ×5, med/max RTT rows into the bench ledger, three arms
  success/NO-REPLY/UNAVAILABLE, always exit 0 — instrument, not gate). NOTE `bf8c4a2` is NOT
  single-purpose: it also corrects the free-space preflight floor read
  (`REQ-CI-FREE-SPACE-PREFLIGHT`, the ci-runner-has-no-warm-target shape) — no 1:1 commit→IR map
  for this commit · **Origin:** golden/bench-wiring red triage 2026-08-02 (ex releases#126)
- **What/why:** the shared-runner quiet predicate (zero non-terminal runs + no local
  cargo/rustc/nextest by parent chain) is process-shaped; both axes passed on a box whose only
  link was degrading (321s for a 1s checkout, bidirectional 10s QUIC dial timeouts). A tailscale
  RTT probe to the peer box before two-host rendezvous, carried in the bench ledger, would have
  called run 30771155390's red in seconds. Evidence: the arm-1 count table (PUMP_PEER_FAIL
  a 0→3→0, b 8→22→8 across green/red/rerun).
- **Permanent, not stopgap:** operator-confirmed 2026-08-02 that kitsubito cannot be provided
  ethernet — wifi-only indefinitely, so the link cannot be hardened and the predicate must see
  link health.
- **Ripe when:** next CI-touching wave, or the next network-shaped golden red — whichever first.
- **Size:** small (one probe step + ledger row + predicate doc).
- **Composed:** LOCKSMITH (#132) CI-rider cluster, hertz thin lane — 2026-08-03. GREENLIT with
  #132 and **DISPATCHED to hertz 2026-08-03** (direct brief; the register is the spec, there is no
  board issue). Leaves the register only when the lane lands or the entry is retired.
- **FIRST FIELD USE, and it DISCRIMINATED (2026-08-04, golden 30873007187 attempt 1):** the probe
  (landed via the rider lane, riding `4b37512`) read 4–7ms RTT healthy on both twohost legs
  minutes before both legs redded — REFUTING the degraded-link read for that red and steering
  triage to the real mechanism (the [[IR-29]] serve-window race) instead of a link chase. The
  instrument's first catch was a correct NEGATIVE — exactly the call run 30771155390 needed and
  could not make.

### IR-2 — Settle the warm-runner CARGO_INCREMENTAL delta
- **Status:** open · **Origin:** #103/#108 bench-wiring lane 2026-08-02 (ex releases#127)
- **What/why:** the #103 measurement (−29.8% wall, −5.65 GB/target, n=3) is COLD-build only.
  Golden's runner `_work` target persists warm, where incremental is exactly what keeps it cheap;
  CARGO_INCREMENTAL=0 was applied only to the genuine cold build (n1-gate pinned old-broker
  cache) + local rig recipes (docs/GOLDEN-CI.md). Open question: does incremental still pay on
  the warm runner, weighed against 5.65 GB/target on a box with LNK1318 free-space history?
- **Method (hertz):** one full golden each way on a quiet box, outside a milestone, compared
  per-step from the bench ledger.
- **Ripe when:** a quiet between-milestone window with no queued lanes (the measurement burns
  two golden windows).
- **Size:** medium (two proving runs + verdict + possible leg flips).

### IR-3 — Daemon-level guard: broker-net wakeup rate bounded across endpoint churn
- **Status:** open · **Origin:** releases#125 remediation, todlando REQ call 1 (ex releases#128)
- **What/why:** the swarm-discovery GC-spin burned two cores for two weeks visible only in a
  process table — no suite assertion sees the class. Wanted: a daemon-level assertion that
  broker-net workers stay quiescent across repeated endpoint create/destroy churn.
- **Design constraint (pre-ruled):** assert on WAKEUP RATE / voluntary ctxt-switch delta over
  the churn window, NOT %CPU — CPU thresholds flake under CI load; the defect signature
  (~100 Hz per orphan loop) is load-independent. Mint the REQ at activation.
- **Near-product:** this is runtime-defect visibility, the most product-adjacent entry here —
  a candidate rider on any daemon-lifecycle milestone.
- **Ripe when:** the next milestone touching spt-net endpoint lifecycle or daemon supervision.
- **Size:** medium (churn harness + counter plumbing + flake-safe assertion).

### IR-4 — Lock-pin guard + lock-procedure rule + toolchain print (three riders, one lane)
- **Status:** open, LANE EXISTS UNLANDED — see [[CI-RIDER LANE STATE]] at the foot of this file;
  candidate commits `1275e47` + `49d4805` + `8ed006b`, mapping NOT yet confirmed by its builder ·
  **Origin:** releases#125 fix-lane intake hold (ex releases#129 + riders)
- **What/why, three parts that land together:**
  1. **xtask check leg:** assert Cargo.lock resolves swarm-discovery to git rev
     `89a2200d54a4e3cab2f46cc75ebff49a1fb07614` while the patch is load-bearing; the check's
     message states its own drop condition (upstream ships a post-PR#27 release AND iroh's pin
     reaches it). Without it, stanza removal or a routine iroh bump silently returns the lock to
     the spinning crate and nothing reds.
  2. **Procedure rule for lock-touching lanes** (docs): targeted `cargo update -p <crate>` only,
     never full re-resolve; count changed `[[package]]` blocks AND diff per-block edges (set-identical
     hid 8 windows-sys edge movers at acaaa4f); two resolutions disagreeing = toolchain drift —
     stop and compare against CI before shipping either lock; hand-edited lock acceptable iff
     `cargo check --workspace --locked` passes.
  3. **Toolchain-version print step in golden** (cargo/rustc versions, both OS legs): the acaaa4f
     comparison against CI was impossible because no run log prints a version. One-grep audit.
- **Ripe when:** next CI-touching wave; part 1 sooner if any iroh bump is proposed.
- **Size:** small-medium (one xtask leg, one docs section, one workflow step).
- **Composed:** LOCKSMITH (#132) CI-rider cluster, hertz thin lane — 2026-08-03. GREENLIT with
  #132 and **DISPATCHED to hertz 2026-08-03** (all three parts land together).

### IR-5 — Shared nextest summary parser
- **Status:** open · **Origin:** BAROMETER triage (standing recommendation, pre-register)
- **What/why:** two agents in one day wrote `[0-9]+ tests run` parsers that read "1 test run"
  (singular) as zero — a guard fed by a broken parser condemns valid rounds. One shared,
  singular-aware parser (single-source discriminant) for every consumer of nextest summaries.
- **Ripe-when REWORDED 2026-08-04 (todlando audit): the old trigger was unfireable as worded.** At
  `b7b00c3` the in-tree population of nextest-SUMMARY parsers is ZERO — the two scripts that read
  nextest output (`g6-curve.ps1:83-90` per-test lines, `g6-postbounce.ps1:44` display-only grep)
  neither parse counts nor carry the defect, so "next wave touching any gate script that reads
  nextest output" could fire on a non-defective script while the real population (agent-authored
  throwaway parsers, which never enter the tree) stays out of reach. New trigger: **the next time
  anyone — agent or lane — needs a nextest summary COUNT**, the shared parser is built FIRST and the
  need consumes it; rig briefs should name it so throwaways stop being authored.
- **Size:** small.
- **Composed:** LOCKSMITH (#132) CI-rider cluster, CONDITIONAL — lands iff the hertz thin lane
  touches gate scripts; otherwise stays open here. 2026-08-03. Carried in hertz's 2026-08-03
  dispatch brief as a **conditional** rider; hertz reports whether the condition fired. **Not yet
  known to be building** — an unreported condition leaves this entry open, not landed.

### IR-6 — Membership logging on subnet gates
- **Status:** open · **Origin:** BAROMETER triage (standing recommendation, pre-register)
- **What/why:** counts beside results, membership beside counts — gate logs that state a count
  without naming the population keep producing unreadable reds. Standardize membership
  enumeration in gate output.
- **Ripe when:** next wave touching gate scripts / CI legs that report counts.
- **Size:** small.
- **Composed:** LOCKSMITH (#132) CI-rider cluster, CONDITIONAL — lands iff the hertz thin lane
  touches gate scripts; otherwise stays open here. 2026-08-03. Carried in hertz's 2026-08-03
  dispatch brief as a **conditional** rider; hertz reports whether the condition fired. **Not yet
  known to be building** — an unreported condition leaves this entry open, not landed.

### IR-7 — Phase A rigs leak a daemon+brain pair on Windows (exe-lock kills notify relink)
- **Status:** open · **Origin:** BAROMETER post-publish triage (ex releases#124 — full mechanism on the closed issue)
- **What/why:** two Phase A rigs launch daemons that escape the job object via WMI-rung autostart
  (double-space unquoted cmdline fingerprint; SPT_HOME in the wrapper cmdline is the attribution
  key); the leaked pair holds target/debug/spt.exe and kills every golden job reaching the notify
  relink. CI reaps as tourniquet (fa6e597); in-job test launch is the fix.
- **Ripe when:** next wave touching the Phase A rigs or daemon autostart path.
- **Size:** medium.

### IR-8 — reap-census scoped_survivors=0 is blind to unreadable-path holders
- **Status:** BUILT 2026-08-18 on KEYSTONE #182 hygiene lane. · **Origin:** BAROMETER triage
  (ex releases#122)
- **What/why:** a zero that cannot see is not a zero — census scoping skips procs whose exe path is
  unreadable, so the survivors count can report clean while a holder lives. Needs a positive control
  / explicit unreadable bucket in the verdict line (unreadable_path count exists; the ZERO must
  refuse when it is nonzero).
- **Defining specimen (golden 30782259675, hfenduleam test job, 2026-08-03):**
  `CI-REAP summary: killed=5 kill_failed=1 scoped_survivors=0` — an admitted kill failure printed
  beside a zero-survivors claim on the same verdict line. The held image surfaced one step later:
  run-scoped tmp cleanup denied 5/5 attempts on `...\relshell\svcmock.exe` (2nd appearance of the
  svcmock hold; 1st @7a3c08c, pre-kill-auth). hertz's addendum: an image-held survivor also blocks
  WRITES to the exe path — the same class manufactures build/relink access-denied reds that mask as
  build problems, not just cleanup warnings. Not per-run: the same-sha green rerun's leg read
  `kill_failed=0 scoped_survivors=0` throughout (hertz, 30784469908) — intermittent sighting,
  second of its class, not a deterministic fixture property.
- **The Linux twin is strictly worse (todlando audit 2026-08-04, vs `b7b00c3`):** `reap-census.sh`
  has NO `kill_failed` anywhere (0 occurrences vs 2 in the `.ps1`) — its kill loop increments
  `killed` only in the success branch with no else, so a failed kill increments nothing and prints
  nothing. The specimen that made this class VISIBLE on Windows would be INVISIBLE on Linux.
  Population precision so this is not overclaimed: ESRCH is benign (already gone; the kill-time
  re-resolve makes it the common case); the vanishing case is EPERM against another account's
  process. The `.sh` `scoped_survivors` DOES come from a post-reap census re-measure, so survivors
  are measured — the hole is failed kills and the unreadable bucket, not the survivor count.
  Windows precision from the same audit: `unreadable_path` IS on the CI-CENSUS line (:157) but the
  CI-REAP verdict line (:243, :246) still carries only killed/kill_failed/scoped_survivors — the
  zero still does not refuse, exactly this entry's ask.
- **Built evidence:** both verdict lines now carry `unreadable_path`; any nonzero unreadable
  family population renders `scoped_survivors=UNPROVEN` rather than a false zero. Linux also
  counts and reports failed kills. The shared predicate is mutation-pinned by
  `reap-census-selftest.sh`; the PowerShell implementation parses cleanly.
- **Ripe when:** next census/reap script wave (natural pair with IR-7's lane) — now BOTH platforms.
- **Size:** small.

### IR-9 — Golden boxes run different clippy versions
- **Status:** OPEN — RULED, awaiting its build lane. Measurement half LANDED AND EXERCISED:
  the toolchain print (`8ed006b` → `REQ-CI-TOOLCHAIN-VERSION-PRINT`) ran on both legs of golden
  run 30971976024 (`test` jobs 92198170694/92198170700 — the step is scoped to `test`, not
  n1-gate; grep token `TOOLCHAIN `). Runner-account facts, read by doyle 2026-08-05:
  hfenduleam `cargo/rustc 1.93.0` + `clippy 0.1.93`, kitsubito `cargo/rustc 1.96.0` +
  `clippy 0.1.96` — the interactive prior confirmed, and the load-bearing NEW fact is
  `stable (default)` on BOTH: neither box is pinned, so any box-side update re-opens the skew.
- **RULING (doyle, 2026-08-05, on the runner-account facts as the 2026-08-03 hold required):**
  (1) align-by-event REJECTED — with both boxes on unpinned stable, alignment decays silently;
  the skew is a mechanism and the fix must be one too. (2) declare-one-leg REJECTED as the
  terminal state — it repairs lint authority but leaves the legs resolving with different cargo
  versions, and the acaaa4f lock-attribution question this family started from is a resolver
  question. (3) **PIN IN-REPO: `rust-toolchain.toml`, `channel = "1.96.0"`** — both runner
  accounts already resolve their toolchain through rustup (witnessed by the `active=` line), so
  the pin self-applies with zero per-box maintenance, and the judge's version becomes a property
  of the TESTED SHA — the same object golden CI already guarantees. Bumps become reviewed lane
  commits, tested by the golden run they ride; the hfenduleam leg proves 1.96.0 on the pin
  lane's own run. (4) INTERIM until the pin lands: kitsubito's clippy leg is AUTHORITATIVE for
  lint disputes — not because newer is stricter (skew direction stays unasserted) but because
  1.96 is the version the pin names, so interim and terminal rulings agree.
  · **Origin:** BAROMETER (ex releases#121); re-confirmed on the #125 fix lane
  (builder's Windows clippy vs kitsubito's rust-1.96.0 lints); see [[CI-RIDER LANE STATE]]
- **What/why:** a Windows-clean lane can land a lint that only reds on the Linux leg — toolchain
  skew makes local clippy evidence non-transferable. Align versions or declare the authoritative
  leg. Natural companion to IR-4's toolchain-version print step.
- **Ripe when:** RIPE NOW — the pin lane (add `rust-toolchain.toml` @ 1.96.0 + confirm nothing in
  CI or the pool machinery keys on the box-default toolchain path) goes to hertz AT NEXT INTAKE,
  not before: he holds the owlery-noun lane and three unlanded pool claims tonight. Entry closes
  BUILT when the pin lane lands.
- **Size:** small.
- **Composed:** LOCKSMITH (#132) CI-rider cluster (rides IR-4's toolchain-version print step) —
  2026-08-03. GREENLIT with #132 and **DISPATCHED to hertz 2026-08-03**; that dispatch delivered
  the measurement half (landed via [[CI-RIDER LANE STATE]], first exercised on run 30971976024).

### IR-10 — Wave gate runs the CONSUMERS of any predicate it changes
- **Status:** open · **Origin:** BAROMETER gate craft (ex releases#119)
- **What/why:** legs chosen from changed crates miss the predicate's callers; a composed red is
  triaged at the wave's own tip first. This is the gate-population rule made binding in the gate
  runbook + scripts rather than living in memory.
- **Ripe when:** next gate-runbook/docs wave.
- **Size:** small (docs + gate-script checklist).

### IR-11 — find-cwd-holders.ps1: --headless discriminator as a column
- **Status:** open · **Origin:** worktree-pin triage tooling (ex releases#118)
- **What/why:** the holder-triage script buries the headless-vs-interactive discriminator in prose;
  as a column it makes the orphan-vs-own-shell call one glance.
- **Ripe when:** any rig-tooling wave; trivial rider.
- **Size:** tiny.

### IR-12 — xtask contract-drift gate misses a stale manifest.schema.json
- **Status:** BUILT 2026-08-18 on KEYSTONE #182 hygiene lane (the narrowed blind arm only). · **Origin:** ex
  releases#116
- **What/why (corrected — the consequence sentence was false at `b7b00c3`):** the literal claim
  holds — `xtask check` does not regenerate-and-compare the schema — but staleness does NOT "ship a
  wrong public contract silently": `checked_in_schema_is_current`
  (crates/spt-runtime/src/manifest.rs:2363, `int->REQ-DOCS-5`, landed `be4e46c` 2026-06-05, BEFORE
  this entry's last sweep — docs lagging code) asserts full content equality of the checked-in
  `manifest.schema.json` against what the derives generate, CRLF-normalised, `SPT_BLESS=1` as the
  regenerate path. It is REACHED (lib unit test; ci.yml:114 runs `-E kind(lib)+kind(bin)` on push;
  golden Phase A re-runs the workspace). Exactly ONE schema file exists in the tree; `docs_bundle`
  (xtask main.rs:609-618) COPIES it at build time, so no second stored copy can drift. Chain closes:
  derives → checked-in (unit-gated) → bundle (copied, not stored).
- **What remains open, and the entry narrows to it:** `check_llms_links` (xtask main.rs:521)
  hardcodes `manifest.schema.json` and `llms-full.txt` as always-existing, so the link check can
  never see them MISSING — a blind arm, not a drift hole.
- **The discrimination is now PROVEN, not presumed (todlando 2026-08-04, burn-the-build arms in
  isolated worktree ir12-mutation @`11169c1`, env verified clear of `SPT_BLESS` FIRST — set, the
  test short-circuits into a WRITE and a mutation arm silently self-heals green; that env check is
  now part of the rig recipe):** arm 0 baseline PASS by NAME (1 test run); arm 1 semantic single
  byte (title `…manifest`→`…manifesX`, length unchanged, scripted edit with match-count refusal)
  REDS at manifest.rs:2374 exit 100 — "manifest.schema.json drifted from the derives — regenerate
  with SPT_BLESS=1", with the mutated token appearing exactly once in 96,248 B of assertion output
  so the byte is provably the only delta; arm 2 whitespace-only (CRLF→LF, 1014 endings) PASSES —
  and arm 0 is itself the stronger normalisation proof, since the on-disk file is CRLF while the
  generated string is LF, so an unmutated PASS is only possible because the test normalises; arm 3
  restore re-measured PASS at the pristine sha256. `checked_in_schema_is_current` is a REAL gate.
- **Built evidence:** `check_llms_links` resolves stored root assets to their
  real source paths and runs the `llms-full.txt` generator instead of returning
  hardcoded `true`. A unit cell proves a missing schema and install script are
  refused, then accepted only after the real source file exists.
- **Ripe when:** next xtask/docs-gate wave (now sized to the blind arm only).
- **Size:** tiny.

### IR-13 — Test-soundness follow-ups from the uniform-table sweep
- **Status:** BUILT 2026-08-18 on KEYSTONE #182 hygiene lane. · **Origin:** ex releases#58
- **What/why:** the remaining candidates from the closed issue were mutation-proved before changing
  tests. `hold_outranks_everything` now walks all four `Opportunity` arms (including
  `HoldRelease`); changing held+release to start reds on that arm. Pure walk tests cover all four
  `QuiesceOutcome` arms and all five `EffectKind` arms; making `KillUnconfirmed` clear or
  `Registry` ephemeral reds independently. The join test now binds all six `JoinFail` variants to
  both hold retention and the emitted failure class/retry payload; misclassifying `WrongCode`
  reds. The access candidates had already landed in the W4 follow-up and its dead-seam removal,
  so the hygiene lane does not manufacture a second test vocabulary beside them.

### IR-14 — .worktrees audit: "how many worktrees are there" has four defensible answers, and the difference is not junk
- **Status:** open · **Origin:** ex releases#56 (was flag: NEEDS-OPERATOR in eval — operator input
  now sought directly when the entry ripens, not via board flag)
- **What/why:** the project `.worktrees/` dir accumulates content beyond what git tracks; audit +
  reap recipe + a hygiene rule for lane close-out. Teardown discipline per docs and memory (classify
  before delete, outbound links first).
- ⚠ **The header of this entry previously read "14 untracked orphan dirs vs 25 git-tracked". Both
  numbers were stale AND UNDATED**, so nobody could tell drift from error. Every count below is dated
  and carries its command.

#### THE COUNT DISAGREEMENT IS THE FINDING (measured 2026-08-03, todlando + doyle, main @`3efd7e6`)

Three people measuring "the worktrees" got three answers. None was wrong; they answered three
different questions, and nothing in the tree states which one is meant:

| answer | question it actually answers | command |
|---|---|---|
| **73** | all ENTRIES under `.worktrees/` | `ls -A .worktrees \| wc -l` |
| **52** | DIRECTORIES under `.worktrees/` | `ls -dA .worktrees/*/` |
| **38** | registered worktrees INCLUDING the root checkout | `git worktree list` |
| **37** | registered worktrees under `.worktrees/` | above, minus the root |

**52 − 37 = 15 unregistered directories.** The 73 − 52 = **21 loose FILES** are covered below.

**A count is only as good as its question.** Treat "how many worktrees" as under-specified until the
answer names its population — the same defect that made a naive `grep -rn` from the project root
inflate a code count by **34.9×** (426 tracked `.rs` vs 14865 on disk excluding all `target/`), and
made that grep run past 120s while `git ls-files | xargs grep` returned instantly. **Scan roots and
population definitions are the same class of error.** Use `git grep` / `git ls-files` for tracked
content and `git worktree list --porcelain` for worktrees.

#### THE 15 UNREGISTERED DIRECTORIES ARE FOUR DIFFERENT KINDS OF THING

⛔ **CLASS A — LIVE BUILD POOLS, NOT ORPHANS. DO NOT DELETE.** 3 dirs, 13.2 GB. Each is the TARGET of
a junction that a REGISTERED worktree uses as its `target/`:

| dir | inbound junction from | size | claim |
|---|---|---|---|
| `gate-target` | assembly-doorbell, doorbell-w1, doorbell-w2, doorbell-w3 | EMPTY | none |
| `gate-target-render` | golden-render | 5.84 GB | `POOL-OWNER.json` → golden-render |
| `gate-target-w4doc` | w4-cli-doc | 7.36 GB | none |

**These are exactly the directories that read as obviously junk** — no `.git`, no source, leftover
names — and deleting one destroys a registered worktree's build pool and leaves a dangling junction.
**Polarity was checked BEFORE classification, which is what caught it:** all 15 top-level dirs are
REAL directories, none is itself a reparse point; the six junctions are **INBOUND**, at
`<registered-worktree>/target`. See [[worktree-target-junction]], [[gate-worktree-target-disk]].

Two findings inside class A, neither of them a deletion question:
- **`gate-target` has FOUR registered trees junctioned into ONE pool, with no `POOL-OWNER.json` at
  all** — so nothing would refuse a second LIVE lane there. That is releases#103's exact hazard
  sitting armed. The pool is also empty: someone reclaimed it and left four junctions aimed at a hole.
- **`gate-target-render`'s `POOL-OWNER.json` carries only `owner_tree` and `written_by` — no pid and
  no birth stamp.** A claim that cannot distinguish a live lane from a finished one is missing the one
  property the claim mechanism exists to provide.

**CLASS B — EMPTIED SKELETONS, ONE UNIFORM SHAPE.** 8 dirs, **0 files**: `acl-core`, `engine-room`,
`ff-fastfollow`, `gate-23d5ceb`, `gate-ff`, `golden-a`, `golden-b`, `w2b-sender-stamp`. Every one is
exactly `crates/spt-daemon/` and nothing else. **Eight independent removals stopping at the SAME
relative path is a mechanism, not litter.**

**MECHANISM CORROBORATED LIVE, 2026-08-03:** a read-only PEB sweep of process CWDs during a running
`spt-daemon` suite found ~20 processes — `spt_daemon-<hash>` test harnesses, `PING`, `cmd` — whose
CWD was **exactly `<worktree>/crates/spt-daemon/`**, the precise path all eight skeletons froze at. A
`git worktree remove` racing a straggler there deletes everything else and leaves that chain pinned.
The processes churn fast (all 17 sampled pids were gone within 30s, one already showing **pid reuse**
— see [[pid-reuse-across-reboot]] for why a pid alone is never an identity), so the pin is a race, not
a steady state, and it recurs on every daemon-suite worktree.

**CLASS C — FULLY EMPTY, no inbound junction.** 3 dirs, 0 bytes: `gate-41`, `gate-559632e`,
`release-runbook-main-advance`. Class B with even the chain gone.

**CLASS D — DELIBERATE SCRATCH, NOT RESIDUE.** 1 dir, 12 files, 50 KB: `_patches` — three `.patch`
files with their `.untracked` manifests, plus `ir18-gate-logs/`. Modified the day before the audit,
i.e. someone's working state.

3 + 8 + 3 + 1 = 15.

#### PIN STATE: NO SKELETON IS HELD TODAY (with a stated blind spot)

Read-only PEB CWD sweep, 2026-08-03: **zero processes hold a CWD in any class-B or class-C
directory** — every live pin was in `gate-ec5f38a` and `gate-main`, both REGISTERED and both running
rigs at the time. So removal of B and C would succeed today; nothing is retrying it.

⚠ **The sweep read 412 of 593 processes; 181 were unreadable** (elevated/system, no
`PROCESS_VM_READ`). The no-pin result therefore holds over the READABLE population only. Cheap to
re-run elevated before acting — [[absence-needs-sibling-probe]].

#### THE LOOSE FILES: A SHARED LOCATION WITH NO STATED CONTRACT

**21 loose files sit in `.worktrees/`** — gate logs (`gate-*.log`, `gate-559632e-log.txt`), rig
scripts (`g6-curve.ps1`, `g6-postbounce.ps1`, `gate-w5-*.ps1`), and `gate-w5-notes.md`. **Nobody
declared `.worktrees/` a log drop; it became one.** Same class as the memory index: a shared location
with no stated contract accumulates whatever anyone puts there, and **the first person to tidy it
cannot tell residue from someone's working state** — class D is that risk already realised. The
hygiene rule this entry owes should name where gate logs and rig scripts belong, not only how to reap
worktrees.

#### RECLAIM ARITHMETIC — AND WHY THIS ENTRY IS NOT THE DISK FIX

Classes B, C and D together are **under 51 KB**. All 13.2 GB of the unregistered population is class
A, behind live junctions. **An orphan sweep is not a disk-space remedy**, and reading it as one sends
you at the wrong target: on the audit date, free space was **12.5 GB against the 32 GB golden floor**
([[free-space-floor-blocks-golden]]) while the two largest pools on the box were `gate-ec5f38a/target`
(14.33 GB) and `gate-main/target` (28.12 GB) — **both REGISTERED, so both outside the orphan
population entirely.** The disk question and the orphan question have different populations; answering
one correctly says nothing about the other.

- **Ripe when:** between-milestone idle window (it is a dev-box chore, zero product risk) — but the
  class-A pool findings are armed hazards and do not wait for it.
- **Size:** small for the sweep; the hygiene rule and the pool-claim gaps are separate small items.
- **Field instance (hertz sweep, 2026-08-04 — an unclaimed pool found rather than reasoned about):**
  `.worktrees\gate-target-w4doc` held **7,907,558,098 B with NO POOL-OWNER.json at all**. Not a stale
  claim, not a foreign claim — no claim. Exactly this entry's shape, first measured instance. Reaped
  under doyle's ruling as part of [[IR-27]]'s two-step teardown; the measurement stands on its own.

### IR-17 — Deadline-burn bring-up family: a test burns its full window while the daemon/brain never comes up
- **Status:** open · **Origin:** golden 30782259675 red triage 2026-08-03
- **What/why:** two different tests on two runner hosts now share one signature — bring-up misses
  its ONLINE window under leg load and the test burns its ENTIRE deadline before the PRECONDITION
  panic: `spt::resident_service_e2e::a_declared_service_rises_with_the_daemon_and_reaches_the_cli`
  (hfenduleam, FAIL 123.99s, "PRECONDITION: the daemon never came up", run 30782259675, green on
  same-sha rerun 30784469908) and `spt::activity_link_push_e2e` (kitsubito, full 30s, brain stderr
  EMPTY — broker spawned, brain never emitted, run 30607903133 specimen, green on same-sha rerun;
  **the empty-stderr leg of THIS sighting is UNVERIFIED-BY-NEGATIVE-CONTROL** — nobody has checked
  whether that test's brain writes stderr on a passing run, and the resident_service twin of this
  observation was retracted for exactly that; cheap to settle next time someone has kitsubito).
  Family reading, not one flake: same missed-window shape, host-independent, both mid-leg under
  load. Cross-ref KNOWN-HAZARDS 5.13 — bring-up has a hard ONLINE budget with known sensitivity to
  anything that stalls it (the blanket-fsync canary; `attach_wedge_e2e` guards that budget).
- **State line (the discriminating datum, resident_service red):** `daemon_up=false boot_alive=true
  boot_pid=Some(8920) rel_started=false broker_survived=false` — the boot process was ALIVE at
  panic time and the daemon never reached up. Any characterization run records this line per run,
  not the verdict (hertz protocol 2026-08-03).
- **Read of that line, CORRECTED 2026-08-03 (hertz):** it is NOT "started-but-never-bound" (the
  earlier reading here) and not "never-started" — it is **started, did work, then killed** —
  CANDIDATE via [[IR-18]]. The daemon spawned its boot service and that service reached the CLI
  (the spool holds its message), and `broker_survived=false`. The DISCRIMINATING field is
  `broker_survived`: false in every killed round of the positive control, true in every healthy
  round. **An earlier evidence leg here is RETRACTED (hertz, same night, off the negative arm he
  ran before reporting; deployah, who had carried the leg into four register sites, swept every
  placement): "empty daemon.stderr.log = TerminateProcess signature" was false — the stderr
  block is empty on PASSING runs too, so it carries zero information in either direction; it
  failed the discriminator question and does not support the kill read.** The kill read now rests
  on the control's field-for-field reproduction and the IR-18 mechanism, not on the log.
  Consequences: (a) the bring-up window is **exonerated for this specimen** — measuring it would
  have measured nothing; (b) the candidate mechanism and its evidence now live in [[IR-18]]
  (a sibling test's bare breadcrumb tree-kill), CANDIDATE — not reproduced, not proven;
  (c) **the characterization population changed, and the rate run is CANCELLED** (doyle ruling
  2026-08-03) — `resident_service_e2e` run ALONE has no sibling to collide with, so the mechanism
  predicts 0/N and that number answers no live question; the only sound population was the test
  INSIDE the Phase A parallel pool, which costs a full Phase A leg per sample, and the fix is
  warranted by the source-verified hazard class regardless of the specimen's rate. No rate figure
  will exist for this specimen — do not later read its absence as a low rate; (d) **positive
  control still runs, falling out of the
  hypothesis:** kill the spawned daemon by pid mid-bring-up — after boot-service spawn, before
  ready — and it must reproduce the observed line field for field (`daemon_up=false
  boot_alive=true broker_survived=false`). If the rig cannot make that shape on
  demand, a 0/N from the untouched arm is worth nothing and must be reported as worth nothing.
  The host-INDEPENDENT family claim (kitsubito's `activity_link_push_e2e`) is untouched by this:
  it has no such kill site named, so the family survives even if this specimen leaves it.
- **POSITIVE CONTROL RAN 2026-08-03 (hertz, prebuilt a5042ec binary, box confirmed clear —
  0 open runs, 145.95 GB free): shape reproduced 3/3, deterministic.** Negative arm n=2: 2/2 PASS
  (5.59s, 5.46s), `daemon_up=true boot_alive=true rel_started=true broker_survived=true
  survived_teardown=true`. Positive arm (daemon killed by parent-scoped descent after
  boot-service spawn, before ready) n=3: 3/3 FAIL (93.95/93.25/93.50s), state line identical to
  the CI red in EVERY field except boot_pid (a pid, must differ), same panic, same site
  (resident_service_e2e.rs:382). Structural fact the arm settled, and the refutation risk that
  made it worth running: the boot service rises BEFORE brain.ready becomes readable — had the
  order been reversed the arm would have produced daemon_up=true and refuted the mechanism.
  CEILING MET, NOT EXCEEDED: the control used Stop-Process -Force (TerminateProcess — the same
  primitive as taskkill /F), so it proves a forced daemon kill after breakaway-service spawn
  produces this exact line ON DEMAND; it says nothing about WHO issued one in golden 30782259675.
  Candidate mechanism, shape reproduced on demand — "root caused" written by nobody. Duration
  note, recorded not explained: 93.5s local (idle box) vs 123.99s CI (Phase A parallelism), both
  burning the same three 45s windows — consistent, not verified.
- **Same-host sub-observation (hertz, host-CONSTANT — narrower claim, kept separate):** both
  reds of 2026-08-03 sat on hfenduleam Windows Phase A and burned their full windows:
  `a_tree_teardown_reaches_a_grandchild_the_service_spawned` (run 30776330383, FAIL 10.176s = the
  full poll deadline, vs a 0.19–0.66s pass band — 15–50x out) + the resident_service row above.
  MEMBERSHIP PROVISIONAL for the teardown row: it carries a candidate mechanism the bring-up burn
  does not share — the pid-only oracle IR-15 just replaced. If the oracle caused it, that row
  leaves the family and host-constant collapses to one sighting. The host-INDEPENDENT pair above
  is the claim that survives someone fixing this box; the sub-observation is what is actionable
  about hfenduleam (a live-daemon host) meanwhile. **UPDATE 2026-08-04 (golden 30873007187
  attempt 1 — SECOND sighting, and the provisional question above is now ANSWERED): the
  pid-only-oracle candidate is REFUTED for this row** — IR-15's authenticated repin rode the
  failing tree (a5042ec ancestor of 4b37512) and the row failed anyway (10.225s, "grandchild
  51848 outlived the tree teardown"), with is_stamped() asserted before the kill so the verdict
  passed the authenticated gate. The row therefore STAYS a family member with its mechanism OPEN,
  narrowed by hertz's log forensics (competence-controlled: the capture channel demonstrably
  prints): no DETACH_BREAKAWAY_DENIED (breakaway rung taken), no SERVICE_JOB_UNAVAILABLE (job
  created + child enrolled), no SERVICE_TREE_KILL_INCOMPLETE (TerminateJobObject returned
  success) — everything enrolled died; THE SURVIVOR WAS NEVER ENROLLED. Two hypotheses,
  inseparable in existing evidence: H1 test defect — the grandchild SELECTION is a bare
  th32ParentProcessID match over a recycling pid space (IR-15's mechanism one level up: the
  verdict was authenticated, the selection was not); H2 real enrollment gap. hertz's selection
  probe (IsProcessInJob at selection time — decisive; birth stamp one-directional; image; same-
  snapshot match count) is authored+compiled and makes the NEXT occurrence self-deciding. Row
  state: INSTRUMENTED-AWAITING-FIRE; a green retires nothing (doyle-ruled). **UPDATE 2026-08-04
  (USHER att3):** row green at 17f95a7 (retires nothing, per the ruling). New shape evidence while
  waiting for the probe: BOTH archived reds sit AT the 10s poll bound (daemon.rs:2757) — 10.176s and
  10.225s — while todlando's 200 off-CI passes under live-fleet load ran 1-4s, never near it. The
  post-fix red is therefore BOUND-EXPIRY-shaped (kill fired; the pinned identity stayed
  not-provably-gone past 10s — TerminateJobObject is async), not kill-wrong-shaped. The bound
  question (is 10s right for a loaded runner) is parked with hertz's USHER package item 4 and must
  not touch the kill path or spawn flags before the probe fires. See also the FLAKE-LEDGER row for
  the false-repeat correction (the 10.176s red was the PRE-a5042ec bare-pid mechanism).
- **Declared read (2026-08-03, rules the next red):** a same-sha green rerun is protocol-conclusive
  for the GATE, not the class. No third silent sample — the next occurrence of this signature gets
  an instrumented resident-service bring-up investigation (hertz), not a rerun.
- **Ripe when:** next occurrence of the signature (immediate instrumented dispatch), or a hertz
  test-hardening wave (bring-up phase-timing instrumentation, per-phase deadline attribution).
- **Size:** medium (bring-up instrumentation + attribution; the fix depends on what it shows).

### IR-16 — kill_tree discards TerminateJobObject/TerminateProcess returns (silent partial kill)
- **Status:** BUILT 2026-08-03 (landed @62c5623 on the LOCKSMITH t1 lane; rides golden sha
  b7b00c3, run 30860770146 green) · **Origin:** ex releases#130 (golden 30776330383 red triage;
  deployah's log read
  + todlando's flagged-not-asserted arm). Operator-classified infra 2026-08-02: daemon-kill
  internals are agent-facing, not operator-facing surface.
- **What/why:** `DetachedChild::kill_tree` (daemon.rs:1433 vicinity) calls
  `unsafe { TerminateJobObject(self.job, 1) }` and discards the return; TerminateProcess likewise
  unaudited. A job that was created and assigned at spawn but whose TERMINATION fails at kill time
  yields the grandchild-survives symptom with zero log signal — the spawn-side
  SERVICE_JOB_UNAVAILABLE announcement is correctly absent (that arm is instrumented and was
  refuted for the specimen by a positively-controlled zero). Wanted: (1) check + log both
  termination returns loudly on failure; (2) the degraded-arm decision — fallback process-table
  tree-walk kill or explicit refusal, never a silent partial kill wearing REQ-RESIDENT-SERVICE's
  unconditional promise. Do NOT add self.job==0 instrumentation (already loud; a second weaker
  rule beside a working one). Discrimination pairing with [[IR-15]]: a red whose captured pid is
  still ping.exe with no SERVICE_JOB_UNAVAILABLE in scope = this entry's arm.
- **Ripe when:** RIPE NOW — IR-15 landed BUILT 2026-08-03 (its rig-side half), so this kill-side
  half is the outstanding instrument; land with the next daemon-teardown wave or sooner.
- **Size:** small.
- **Composed:** LOCKSMITH (#132) — todlando product lane, with the `detached_no_inherit_env`
  rename as cosmetic rider (IR-20's load-bearing-pid caution applies to any rig touch) —
  2026-08-03. GREENLIT with #132 and **DISPATCHED to todlando 2026-08-03**, carrying IR-20's
  load-bearing-pid precondition as a stated check (todlando reports whether the rig is touched
  at all rather than silently skipping it).
- **BUILT — fulfilment checked against the wanted list, not merely mapped to a commit** (doyle,
  2026-08-04): `62c5623` satisfies (1) — both termination returns are now checked and the loss is
  NAMED on failure, as the kill-time twin of the spawn-time `SERVICE_JOB_UNAVAILABLE` — and (2) by
  the ruled LOUD REFUSAL arm rather than a tree-walk fallback, which [[IR-18]] measured blind in
  exactly this failure state. The `self.job == 0` prohibition was honored: the new check is
  `self.job != 0 && TerminateJobObject(...) == 0`, not a second weaker rule beside the working one.
  The unix `ESRCH` quiet arm — which no Windows gate can see — is covered by `ec5f38a`. The
  `detached_no_inherit_env` cosmetic rider did NOT land and was dropped at `de6a01d` on a MEASURED
  false premise, not deferred. A Windows-specific finding rides the fix as a comment and is worth
  keeping: `TerminateProcess` against a handle to an already-exited process returns 0 with
  `GetLastError` 5 (`ERROR_ACCESS_DENIED`), so a naive check there fires on the ORDINARY path.

### IR-18 — resident_service_e2e kills breadcrumb pids BARE (stale-breadcrumb tree-kill; kill-side twin of IR-15)
- **Status:** BUILT 2026-08-04 — the named acceptance carrier RAN and the condition is DISCHARGED,
  stated as the measurement rather than as the green board it sat on: LOCKSMITH's golden run
  30860770146 (head `b7b00c3`, conclusion success) executed the changed test IN-POOL on BOTH
  platforms — `spt::resident_service_e2e a_declared_service_rises_with_the_daemon_and_reaches_the_cli`
  PASS 17.773s (100/2637) on hfenduleam/Windows and PASS 17.895s (1364/2616) on kitsubito/Linux.
  The pool was checked for the test rather than inferred from the job colour, because a green leg
  that never ran the test is a competence-controlled zero, not an acquittal; `resident_service_e2e`
  contains exactly ONE `#[test]`, so the test that ran IS the test the lane rewrote (229 lines of
  that file changed in `6e1a962`). No red returned to doyle. `.worktrees/ir18` is released for
  teardown by this verdict · **Origin:** hertz
  characterization of the [[IR-17]] resident_service red, 2026-08-03 — CANDIDATE mechanism, NOT a
  reproduction
- **Lane:** hertz `fix/ir18-authenticated-teardown` — MERGED to main @6e1a962 (PR #142; gated at
  b48e35b, rebased 6e1a962 on 1e520b4 with zero code delta — doyle re-derived `git diff -- crates/`
  empty across the rebase, so the gate verdict and local behavioral evidence transfer). Gated by
  doyle: full diff review (one blocking finding — the scope_lost assert condition contradicted its
  own guard ruling — fixed and delta-verified at one line), mutation-proven locally both
  directions (B: forced boot_pid=None → population sweep names the leak no per-pid check sees,
  FAIL 101; A: wrong expected_exe → REFUSED-foreign-image reds loudly on the derivation check).
  EVIDENCE SPLIT, stated so the check marks cannot carry it (hertz's absent-leg callout): lane CI
  30792210008 was 5/5 green per job but ci.yml's test leg is kind(lib)+kind(bin) — the changed
  integration test NEVER RAN there (verified by log grep, 0 hits / 3023 lines); CI proved
  compile-everywhere (clippy all-targets = the pool-population check for a common/ module),
  traceability, lib/bin clean. The BEHAVIORAL evidence is local: three clean runs (all verdicts
  Killed, population empty) + the two mutation reds. reap.rs diff additive-only (zero deleted
  lines, verified) — cross-test risk bounded to compilation, which CI covered. Evidence custody:
  gate + mutation logs at `.worktrees/_patches/ir18-gate-logs/` (5 files); `.worktrees/ir18` is
  KEPT deliberately until the golden verdict (NOT an orphan — do not reap; a Phase A red wants the
  tree and logs in place, not rebuilt); its pool claim is released.
- **What/why:** `crates/spt/tests/resident_service_e2e.rs:46` defines its own reaper —
  `taskkill /PID <pid> /F /T`: bare pid, force, whole TREE, no identity check — and feeds it three
  BREADCRUMB-derived pids in cleanup (`boot_pid`, `rel_service_pid` at :360, and the pid read out
  of `brain.ready` at :363). The suite already ships the authenticated tool other tests use and
  this one does not: `common::reap::authenticated_kill(label, pid, expected_exe, observed)` at
  `crates/spt/tests/common/reap.rs:116`. Inside the Phase A parallel pool on a pid-churning box
  this is the KNOWN-HAZARDS stale-breadcrumb tree-kill class, live and **symmetric**: this test can
  take a sibling's daemon and a sibling can take this test's. Same defect [[IR-15]] just fixed, on
  the other side — IR-15 authenticated the READ (is the thing I pinned gone?), the KILL is still
  bare pid. Kill what you pinned, not the number it happens to hold.
- **Verified at source by deployah 2026-08-03** (relayed claims re-derived, all held): the bare
  kill, its three breadcrumb feeds, the shipped-but-unused authenticated helper, and the
  `CREATE_BREAKAWAY_FROM_JOB` spawn (`daemon.rs:1300`).
- **Evidence rating — why it is a candidate and not a cause:** it predicts every field of the one
  observed state line. The daemon is killed after spawning its boot service and before binding:
  `broker_survived=false` (the discriminating field — false in every killed control round, true in
  every healthy one) + `boot_alive=true boot_pid=Some(8920)` with the service's message in the
  spool. (The "empty stderr = silent death" leg that originally sat here is RETRACTED — see
  [[IR-17]]'s control record; stderr is empty on healthy runs too and discriminates nothing.)
  The service outlives its daemon because `detached_no_inherit_env` spawns it
  `CREATE_BREAKAWAY_FROM_JOB`, so it is NOT in the daemon's job and a daemon kill ORPHANS it
  rather than reaping it — plausibly also the five denied `svcmock.exe` attempts in that job's
  reap summary. **The [[IR-17]] positive control (2026-08-03) reproduced the shape 3/3
  deterministically with a forced kill at that window — establishing the SHAPE on demand, not the
  AGENT: nothing identifies who issued a kill in golden 30782259675.** The bare-pid tree-kill from
  a concurrent test remains the candidate agent, on the source-verified hazard class and reap.rs's
  documented prior casualty. **Not reproduced in the wild. Not proven. "Root caused" is not
  written here by anyone.** The [[IR-17]] rate run is cancelled, so no rate will ever back this —
  the fix stands on the source-verified hazard class alone. THE DEFECT IN MINIATURE, observed as a
  side effect of the control (hertz 2026-08-03): the three killed rounds leaked 12 processes
  (6 svcmock + 6 spt.exe, two services per round) and the test's own teardown reaped NONE — when
  the daemon dies, `rel_pid` is `None` so that service is never even a kill target, and the
  breakaway child outlives everything. A teardown that leaks two processes per round currently
  PASSES: the concrete case for `target_gone` as the gate's positive half. (Leak cleaned scoped,
  each pid re-authenticated by image-path prefix + creation window at kill time, not from the
  minute-old enumeration.) ONE LAYER DEEPER (deployah, same night, at source): the reap loop
  (:359) is `[boot_pid, rel_service_pid].into_iter().flatten()` — `.flatten()` DROPS None, and
  `mock_pid` (:100-106) collapses absent/permission-denied/IO-error/garbage into that None via
  `.ok()?`/`.ok()` — the reap.rs:59-63 collapse in a THIRD location, inside this very test. So the
  leak is both "service never started" AND "pid unreadable, therefore never a kill target":
  absence of knowledge read as absence of target, in a teardown, again.
- **GATE SPEC AMENDED PRE-BUILD (deployah found the hole, doyle-ruled 2026-08-03 — read this
  BEFORE building the lane):** the two halves as first specified (trap-class-absent +
  `target_gone` per pid) both operate on pids we HAVE — a process whose pid is None is invisible
  to BOTH, so the specified gate PASSES hertz's own control-run leak (12 processes, 6+6, 2/round,
  teardown reaped none). A gate its own reproduction case passes is not yet an instrument. THIRD
  HALF, REQUIRED: a POPULATION assertion at teardown — zero surviving staged-service or
  test-owned spt processes, by selectors external to the pid bookkeeping. SELECTOR RESPELLED
  (hertz 2026-08-03, doyle-approved — the first spelling, "parent-pid descent from the test
  process", would have caught ZERO of the 12: teardown kills the daemon before any check runs, the
  BFS over the current (pid,ppid) table breaks at the dead middle hop, and descendants(test_pid)
  returns empty exactly in the leak case — an absence read as a clean, the shape the lane exists
  to kill; the control's own cleanup is the evidence, it had to use image+creation-window because
  descent was already broken). THE TWO SELECTORS AS BUILT: (1) svcmock half by IMAGE PATH —
  stage_adapter copies the service binary under this run's unique tempdir home, so "any pid whose
  exe_path canonicalizes under home" names every staged service this run started and nothing else
  on the box; no bookkeeping, immune to reuse and broken chains. (2) spt.exe half by SEEDED
  descent + image — image alone is forbidden (target/debug/spt.exe is shared with concurrent
  tests and live perches: the machine-wide selector class); seed with the ancestor set captured
  WHILE ALIVE ({test, broker from the Child handle, brain from brain.ready — none via mock_pid}),
  union descendants, keep only exe_path==spt_bin. Reuse exposure on seeds is assertion-only (a
  red, never a kill) and the image filter closes it. Both halves catch the 12-leak conformance
  case; the first spelling caught none of it. PLUS THE [[IR-8]] BUCKET, REQUIRED (deployah,
  same night — the third instance of that sentence tonight, this time inside our own instrument):
  both selectors match on IMAGE PATH, and an exe path that cannot be READ yields no match — a
  leaked svcmock holding an unreadable path is invisible to both halves and the count comes back a
  zero that cannot see. Not hypothetical: IR-8's defining specimen is 5/5-denied on this exact
  image, on this box. So the population assertion counts unreadable-path processes as their OWN
  bucket and the zero-survivors claim REFUSES when that bucket is nonzero — zero seen AND zero
  unseeable, or the assertion states it could not see. Built in from the start, not retrofitted
  (may land in the lane's second commit at hertz's discretion; the requirement is that it lands in
  the lane). This instantiates IR-8's remedy test-side; the CI reap-census script half of IR-8
  stays open.
- **Fix (thin lane, hertz — doyle-ruled 2026-08-03):** adopt `authenticated_kill` at all three
  sites. (Board #131's `CREATE_NO_WINDOW` rider is DROPPED from this lane — hertz falsified the
  filed remedy at source 2026-08-03: daemon.rs has exactly ONE CreateProcessW (:1187) and
  BASE_FLAGS (:1105) already carries 0x0800_0000 = CREATE_NO_WINDOW on both rungs including the
  ACCESS_DENIED fallback (which drops only BREAKAWAY), so the specified edit ORs an already-set
  bit — a bit-for-bit identical flag word that would close a board item while its symptom, if
  real, continues. #131 handed back to board triage with the finding and a candidate direction:
  DETACHED_PROCESS gives the service no console, so a console child the SERVICE spawns flagless
  allocates a NEW visible window — the grandchild, not the daemon's spawn, is the candidate
  surface; wants its own diagnosis from an observed window. ~~The `detached_no_inherit_env` rename
  ("detached" actually means breakaway-from-job) is a candidate cosmetic rider on [[IR-16]]'s
  product lane~~ — **RENAME DROPPED 2026-08-03, ITS PREMISE MEASURED FALSE** (todlando, on the
  IR-16 lane; doyle re-verified at source before ruling). "detached" is ACCURATE: `BASE_FLAGS`
  (daemon.rs:1103-1105) is `DETACHED_PROCESS | CREATE_NEW_PROCESS_GROUP | CREATE_NO_WINDOW`, and
  `0x0000_0008` IS DETACHED_PROCESS — the function sets BOTH postures and the name states one of
  them. The name is not wrong, it is INCOMPLETE (silent about `CREATE_BREAKAWAY_FROM_JOB`, which
  arrives as `extra_flags` at daemon.rs:1506), so renaming on the stated premise would have traded
  an accurate word for one that drops a flag the function really sets. Independently fatal: the
  premise, if true, covers the sibling `detached_no_inherit` equally — 17 references across
  daemon.rs/deelevate.rs/shellhost.rs — and renaming one of two siblings of ONE posture leaves the
  tree MORE inconsistent, not less. **What lands instead:** one doc-comment line on EACH function
  naming the full effective flag word, which fixes the real complaint (neither name says breakaway)
  without touching a call site. Kept here rather than deleted because this entry was the claim's
  only home, and a reader who found the rename gone with no reason would re-derive it.) PER-PID GROUND TRUTH for expected_exe (hertz — SELF-CORRECTED
  at source before it ever landed; his first derivation, "Copy runs from adapters_dir(), Pointer
  runs from srcs/", was WRONG and registration mode is NOT the discriminator): servicehost.rs:1328
  takes `install_dir` from `record.source_dir` on the Started arm for EVERY service, Copy and
  Pointer alike, and Copy copies only the manifest + strings/ (registry.rs:395-404), never a
  binary — `adapters_dir()` holds no executable in either mode. BOTH services run from their OWN
  staged src dir: `staged_bin(<adapter's staged src dir>, "svcmock")`. The trap survives with a
  different mechanism: two adapters staged into two DIFFERENT source dirs are two different files
  on disk, so one expected_exe still cannot serve both. A lane that trusted the mode split would
  derive a path with no binary at it. How it was caught, kept because it is the reusable part: not
  by rereading the mode table but by finding where install_dir is COMPUTED instead of trusting an
  already-published inference — the mode split was true of MANIFESTS and had been generalized to
  BINARIES without checking the consumer; and the derivation-asserted-at-observe() instrument
  would have caught it at runtime regardless — the instrument did its job before it ever ran,
  argued against its own author. SCOPE_LOST RULED A GUARD, NOT A FINDING (hertz design change,
  doyle-approved): a seed pid recycled before the sweep makes the sweep DECLINE that subtree
  loudly, not fail — failing would turn ordinary pid churn on a busy runner into a red against a
  clean teardown (the IR-15 false-red class re-minted inside its own fix); coverage holds without
  it because the staged services are caught by image-path-under-home and the daemon + brain by NEW
  direct per-pid target_gone checks (their pids this test never loses) — the ancestry half is
  reach, not load-bearing. `unreadable` stays FATAL as ruled.
  Sequencing: positive control
  first (it is the instrument that could still refute the mechanism), then the lane. Does NOT ride
  [[IR-16]] (doyle-ruled same day): that is the product-side kill_tree audit and stays a separate
  lane under the dispatch split; the two cross-ref, they do not merge.
- **Lane trap, flagged before build (hertz 2026-08-03 — the false-clean shape):** `expected_exe`
  differs PER PID. `boot_pid` and `rel_service_pid` run the staged svcmock image, NOT spt.exe —
  passing spt_bin for those refuses every kill as foreign-image and LEAKS both services while the
  reap reads hardened. Only `brain.ready`'s pid takes spt_bin. Also: `observe()` boot_pid at :210
  while it is provably ours, so the reuse check has creation-time teeth instead of degrading to
  image+ancestry.
- **Gate the lane on killing, not on refusing (deployah 2026-08-03; REFINED by hertz same night
  from reap.rs's own contract):** the hardened version fails safe in the WRONG direction —
  refuse-everything leaks both services while printing exactly the reap line a reviewer wants to
  see, strictly worse than the bare kill it replaces and invisible in the same log. But a BLANKET
  zero-REFUSED gate is wrong too: reap.rs's module doc (:18–20) states refusal is the CORRECT
  happy-path outcome when the test already stopped its daemon (`Refused("gone")` — leak insurance,
  not primary teardown), so the blanket gate reds healthy runs, someone loosens it to go green, and
  the loosened version is exactly the one blind to foreign-image — the detector dies by
  maintenance, wearing a green. The refusal reasons split, and only one class is a defect:
  SOUND (expected): `gone`, `reused`, `breadcrumb-moved`, `self`, `ancestor` — target provably not
  there. TRAP (the false-clean state): `foreign-image`, `unreadable-image`, `unproven-identity` —
  something WAS there and authentication declined; `foreign-image` is precisely what a wrong
  per-pid expected_exe produces on every pid, every run. THE GATE: in the arm that must kill,
  assert no verdict in the trap class — structurally on the returned `Verdict` (Killed |
  Refused(&'static str), stable tokens by design, reap.rs:48), not by scraping output; the
  `REAP[label]: … verdict=` line stays as CI-log forensics. The gate's comment MUST say
  `Refused("gone")` is expected, or the next reader "fixes" it. Positive counterpart that actually
  proves the reap: assert `target_gone(pid, expected_exe, observed)` (reap.rs:215 — no-knowledge
  never read as death) per service pid; bare-liveness "no survivors" is the false-clean named above.
  TOKEN SET VERIFIED COMPLETE-PLUS-ONE (deployah, refuse() sites enumerated at source): the eight
  tokens above sit at :128 self, :132 ancestor, :136 breadcrumb-moved, :146 gone, :149
  unproven-identity, :155 reused, :166 unreadable-image, :172 foreign-image — each in the right
  class — and there is a NINTH neither list held: `Refused("no-breadcrumb")` at :243, returned
  before authenticated_kill is reached. Doyle-ruled 2026-08-03, both halves: (1) `no-breadcrumb`
  classifies SOUND — **SCOPED (deployah source-read, same night): it is two states wearing one
  name.** `breadcrumb_daemon_pid` (reap.rs:59–63) collapses file-absent, permission-denied, any IO
  error, AND a husked/unparseable partial write into one `None` via two `.ok()`s — so the token is
  ABSENCE OF KNOWLEDGE, not proof of absence, and it is SOUND only because the calling test has
  already asserted its daemon stopped. Every other SOUND token carries positive proof; a LIVE
  daemon whose pid file is unreadable or husked emits the identical token and leaks wearing
  "nothing to reap" — the helper's one fail-open path inside a fail-closed design (its own module
  doc: no-knowledge is never read as death), and the same sentence as [[IR-8]]: a zero that cannot
  see is not a zero. CORRECT REMEDY (doyle-ruled, deferred — not this lane, unreachable from this
  test's pids): split at source into `no-breadcrumb` (absent) vs `unreadable-breadcrumb` (TRAP,
  beside unreadable-image) — IR-8's unreadable-bucket fix in a second location; land the pair in
  the census/reap wave IR-8 already names, one remedy for two entries. Precondition on the split
  (hertz): common/reap.rs compiles into EVERY spt integration test binary, so enumerate the callers
  that actually REACH `reap_breadcrumb_daemon` before the token split lands — those are the tests
  whose teardown semantics change; own lane, not a rider. ENUMERATION DONE (deployah, same night,
  read-only): 21 caller tests, every one passing &spt_bin as expected_exe (the per-pid trap is
  specific to resident_service_e2e's svcmock pids — it does not generalize to this population);
  20 callers fire-and-forget the Verdict; the one value-assertion, reaper_guard.rs:192-196, pins
  `Refused("no-breadcrumb")` for an ABSENT breadcrumb only — absent KEEPS the token under the
  split, so that assertion survives unchanged (:183's stale-pid check is the loose
  `matches!(Refused(_))`). The split is ADDITIVE: one source split, one new token, zero caller
  rewrites. THE GAP THAT LET THE COLLAPSE HIDE: reaper_guard covers eight refusal shapes but has
  NO test for an unreadable or malformed daemon.pid — the one state where absence-of-knowledge
  reads as absence-of-target is the one state without a conformance test; (b) ships with exactly
  that evidence (a garbage daemon.pid and an unreadable one, each asserting the new token).
  ASSIGNED: deployah owns the (b) lane (doyle-ruled 2026-08-03), SEQUENCED AFTER hertz's IR-18
  lane — the exhaustive panic-catch-all lands first, so the new token arrives forced-classified
  per this entry's own design, and the (b) lane classifies it TRAP in the same change. SCOPE
  RE-RULED (2026-08-03, on deployah's third-site find): the collapse now has THREE known sites
  (reap.rs:59-63, resident_service_e2e's mock_pid :100-106, and the shape wherever a test reads a
  self-written pid file) — so (b) lands ONE shared read-a-pid-breadcrumb helper that distinguishes
  absent from unreadable ONCE, adopted at the enumerated sites, not a per-site token split; the
  single-source-discriminant rule, applied before a fourth site mints itself.
  **THIRD SITE CONFIRMED AT SOURCE (deployah, read-only 2026-08-04) — it was a predicted SHAPE and
  is now an address:** `crates/spt/tests/resident_service_e2e.rs:136-137`, `brain_ready`, which
  `.ok()?`s BOTH the read and the JSON parse — two collapses in two lines, the self-written-pid-file
  shape this entry predicted. What makes it the worst of the three rather than the third of three:
  its output feeds `brain_seed`, a population-sweep SEED, so an unreadable `brain.ready` silently
  drops a whole subtree from the sweep. The other two sites lose one pid; this one degrades the very
  instrument that is supposed to cover them, and it does so while the sweep still reports clean —
  [[IR-8]]'s sentence again, now inside the instrument rather than beside it. The (b) lane's helper
  therefore has to be adopted at this site to be worth landing at all. Still deployah's,
  still after hertz's lane; hertz's lane does NOT wait for it (the population assertion covers the
  leak class independent of pid bookkeeping — that is its virtue). (b) INTAKE ITEM (doyle gate on
  the IR-18 lane, 2026-08-03, deferred there deliberately): the lane's verdict-gate SOUND-arm
  comment reads "provably not there", which overclaims for `no-breadcrumb` (absence-of-knowledge,
  not proof — the scoped reading belongs at that arm); comment-only, and (b) rewrites that arm
  when `unreadable-breadcrumb` splits out, so it lands there rather than costing a solo respin.
  SECOND (b) INTAKE ITEM (hertz self-noted at re-gate, doyle-ruled deferred): cleanliness is
  currently stated at THREE sites (two asserts + `is_clean`) — the drift that produced the
  scope_lost condition bug (a failed edit followed by a narrower successful one replaced the
  message but not the condition; the artifact of that failure mode IS message/code disagreement).
  (b) consolidates to conditions derived from one source, weighing the two-message diagnostic
  split it would cost. The
  exhaustive panic-catch-all match is what makes the split safe to do later: the new token lands
  loud on first fire in every consuming gate, forced to be classified, never silently passed. Until then the scoped reading above is the
  gate's contract — and the LANE BUILDER writes that scoped reading into the match's SOUND-arm
  comment (the sketch's "nothing to read is nothing to kill" is the exact inference the source
  forbids; the comment must not teach the false step); (2) the gate closes over the
  token set — SPELLED BUILDABLY (hertz correction 2026-08-03; the earlier "no wildcard arm" wording
  is an instruction rustc rejects, since the refusal reason is a `&'static str` (:55) and a string
  match always requires a catch-all): the match names the six SOUND tokens as a pass arm, the three
  TRAP tokens as a panic arm, and the REQUIRED catch-all arm is not a wildcard PASS — it panics
  naming the unknown token ("unclassified reap refusal token — classify it in this match"), so
  upstream drift announces itself the first run it fires instead of sliding into a silent pass.
  Runtime fail-closed, not compile-time: compile-time closure needs the reason to become an enum in
  common/reap.rs, an upstream change touching every consumer — ruled NOT smuggled into this lane
  (candidate for its own thin lane if ever wanted; doyle concurs runtime fail-closed is
  proportionate). Direction of the default, kept in the record: a missed trap state is an invisible
  leak wearing the reap line a reviewer wants to see; an unclassified harmless token is a loud red
  costing one edit. Cheap and loud beats silent and wrong. (Count hygiene, hertz self-flagged: a
  single-line grep for the refuse() sites returns five of nine — three wrap across lines; the nine
  stands on reading, not the grep. The sweep-count trap in miniature.)
- **Prior observed instance of the class, in-tree (hertz 2026-08-03 — stated to understate):**
  reap.rs:3–11 documents the measurement and a casualty: pid reuse on the Windows gate box at
  2.2s minimum / p50 11.7s under a Phase-A battery, and a concurrent test process killed by a bare
  kill on a stale breadcrumb pid, dying with bare exit 1 and no panic — the observed
  `worker_lifecycle_e2e` red in golden job 90742055246 (that test's intermittent class is board
  releases#32, still in-flight). Different test, different job — it does NOT reproduce this entry's
  specimen and promotes nothing; what it settles is that the class is real IN THIS SUITE with a
  named victim, and the authenticated reaper is the remedy someone already built for it —
  resident_service_e2e simply never adopted it. Two shape-predictions it supplies, made before this
  specimen existed: the p50 11.7s reuse window is short against this test's 45s waits and ~124s
  runtime (window wide open). (A second claimed match — "bare-exit-1-no-panic matches the empty
  daemon.stderr.log" — is RETRACTED as a category error, hertz+deployah 2026-08-03: reap.rs's
  phrase describes the VICTIM TEST PROCESS dying, not a daemon's log; two processes, two
  artifacts.) Predictions, not proof; the specimen stays CANDIDATE — shape since reproduced on
  demand, agent unidentified (see the control record above).
- **Rig near-miss, caught before first execution (hertz 2026-08-03) — kept as design justification,
  not an anecdote:** the positive-control arm's first draft selected its victim machine-wide
  (`Win32_Process Name='spt.exe'` + cmdline match `daemon run`, then force-kill). hfenduleam is the
  Windows CI runner as well as a dev box, so that filter matches CI's own test daemons: run while a
  leg was live, it would have force-killed a running job's daemon and reported the result as a
  measurement — fabricating a red in someone else's job, the exact class this entry documents.
  Corrected to a parent-scoped selector (daemon = child of the test pid; service = child of that
  daemon); `CREATE_BREAKAWAY_FROM_JOB` does not weaken parent-id attribution, since breakaway leaves
  the JOB, not the parent record — which is also why the orphaned svcmock stays attributable to the
  daemon that spawned it. **The load bought the catch** (blocked on the box ⇒ re-read instead of
  ran). Bearing on the fix: the INSTRUMENT built to study a bare-selector reaper was itself written
  with a bare selector. The remedy therefore cannot be "be careful" — the authenticated path must be
  the only one within reach in this suite.
- **Ripe when:** RIPE NOW, and not gated on the rate run that was cancelled.
- **Size:** small.

### IR-19 — Docs-only pushes to main run full unit legs (classifier is PR-only) — intent unverified
- **Status:** RETIRED 2026-08-18 — intentional for main pushes, not a classifier defect. · **Origin:** hertz observation 2026-08-03 (register-push
  cadence gated his rig behind repeated hfenduleam unit legs); mechanism verified by doyle at
  source same night.
- **What/why:** `ci.yml`'s `changes` job classifies docs-only diffs ONLY for `pull_request` events
  — push events hardcode `code=true` (ci.yml:45–48, an explicit branch, not a fallthrough), so
  every docs-only push to main spins both full unit legs. Tonight: five register docs pushes each
  queued a Windows unit leg on hfenduleam; concurrency kept one running + one pending (main is
  never cancelled in-progress, pending runs supersede — the recorded behavior matches ci.yml:9–13).
  FINDING SETTLED (deployah, same night, from the workflow): this is not intent — it is a scope
  the classifier never had; the `*.md` / `traceable-reqs.toml` case arms are only ever REACHED on
  pull_request (the early-return precedes them), and golden.yml's own hardcoded `code=true` is
  unrelated (golden triggers only on golden/**). REMEDY STILL OPEN on a named tension: extending
  the classifier to pushes leaves a docs-only main TIP with no run of its own — fine if
  tested-sha==merged-sha means "the code at this sha was tested" (it was, at the last code sha),
  not fine if any gate or reader takes "main tip has a green run" as the check. One person reads
  the consuming gates before any remedy (doyle, at the next CI-touching wave) — three assuming is
  how this class ships.
- **Resolution:** the consuming contract already answers the held question:
  `REQ-CI-DOCS-ONLY-THIN` is PR-only and explicitly says ADR-0050 supersedes
  it for main pushes. Main-tip evidence remains full by design; no classifier
  branch was widened.
- **Interim rule (doyle, same night):** batch register edits into one push instead of landing them
  as they occur — the cadence is a real gate on whoever is queued behind the runner.
- **Ripe when:** next CI-touching wave, after the intent question is answered on the REQ record.
- **Size:** small (one conditional, if the answer is "extend").
- **Composed:** LOCKSMITH (#132) — doyle reads the consuming gates during the batch's CI-touching
  wave (the intent question), BEFORE any classifier remedy. 2026-08-03.

### IR-20 — resident_service_e2e's `spt daemon stop --force` does not stop the daemon; teardown is complete only because the kills are
- **Status:** open · **Origin:** hertz 2026-08-03, measured on the IR-18 lane's own clean run
  (hfenduleam, log kept); routed text, doyle-landed.
- **MEASURED:** teardown runs `spt(&["daemon","stop","--force"])` then reaps. At reaper fire, every
  target was still ALIVE and every verdict was `Killed`, not `gone` — verdicts=[boot Killed,
  rel Killed, brain Killed]; the brain's kill line names it `(child process of PID 46896)` — the
  daemon — so the daemon was still resident too and was taken down by the `Child` handle, not the
  stop. A stop that worked would have left the reaper `gone`. This rules OUT the innocent reading
  (stop doesn't manage services): the stop did not reach the DAEMON either.
- **NOT ESTABLISHED — two rungs, neither discriminated:** (1) the rig seeds a `doyle` perch with
  `std::process::id()`, satisfying `ceremony_agent_ground`'s pid-ancestry rung for every child of
  the harness; (2) the rig never scrubs `OWL_SESSION_ID`/`SPT_AGENT_ID`/`SPT_ENDPOINT_ID` from its
  teardown commands, and the measuring run launched from a live agent session, so the env rung was
  live too. `--force` overrides neither. The stop's output is discarded at the call site
  (`let _ = spt(...)`) — no diagnostic names a denial, and no claim is made that one fired.
  Settling it is one run with the stop's stderr captured — on the CI runner AS WELL AS the dev box,
  since the rungs fire on different machines.
- **Why it matters beyond this rig:** REQ-TEST-RIG-DAEMON-TEARDOWN-PROVEN's own failure shape
  inside a rig that passes — teardown completeness rested entirely on breadcrumb kills that were,
  until IR-18, bare `taskkill /F /T` on unauthenticated pids; the most trustworthy-LOOKING
  component (an explicit `--force` stop) was doing nothing while the least trustworthy one did all
  the work. Exactly why the population gate exists rather than per-pid checks alone.
- **Remedy caution (the sentence that must not drop):** the prescribed fix (seed pid 0, scrub the
  three markers — REQ-TEST-RIG-DAEMON-TEARDOWN-PROVEN's own gate) is NOT free here: this rig's
  `doyle` perch pid is load-bearing for the shell bind-by-token legs, so the two remedies are not
  interchangeable — a lane taking this must check which legs depend on the pid before changing it,
  or it fixes the stop and breaks the REQ-INSTALL-11 legs the rig exists for. Kin:
  REQ-BROKER-STOP-ENDPOINT-DENY (the refusal being worked around).
- **Ripe when:** a hertz test-hardening wave, or alongside the (b) helper lane's touch of this rig.
- **Size:** small (stderr-captured discrimination run × two machines, then the scoped remedy).

## IN-FLIGHT ON THE BOARD (not re-homed — live WIP lanes)

- releases#93 (servicehost swap test pid-inequality), releases#47 (thin-lane unit(Windows)
  isolation class), releases#32 (worker_lifecycle_e2e Phase A intermittent) — infra-shaped but WIP;
  a live lane is never re-homed mid-flight. When each closes, any residue lands here as a new entry.

## BUILT / RETIRED

### IR-15 — provably_gone is pid-only on Windows: teardown tests can fabricate reds
- **Status:** BUILT 2026-08-03 · **Origin:** golden 30776330383 red on the #125 fix lane (2026-08-02)
- **Lane:** hertz `golden/ir15` @a5042ec — REQ-TEST-LIVENESS-ORACLE-AUTHENTICATED (impl+unit): one
  shared authenticated death-oracle test helper backed by `spt_procident::process_identity`
  (identity pinned at find time, re-verified at assert), both pid-only polling sites repinned
  (daemon.rs teardown test; endpoint_lifecycle.rs relay_pid). Ruled polarity held: Absent⇒gone,
  Present(same start)⇒not gone, Present(different)⇒gone, Unproven⇒NOT gone, missing-stamp
  degrades to pid-only loudly — every unknown errs toward the recoverable red, zero new
  false-green paths. Linux caveat documented in the helper (10ms jiffies: Present(different) is
  narrowing, not decisive; same-tick reuse errs red, safe).
- **Golden:** run 30782259675 red on `resident_service_e2e` bring-up — delta exonerated (unrelated
  family, now [[IR-17]]; same job minted [[IR-8]]'s defining specimen). Same-sha rerun 30784469908
  GREEN; main ff'd to a5042ec.
- **A/B discriminator (hertz):** 80/80 green both arms, quiet + churn phases; churn arm
  positive-controlled (pid-allocator wrap observed <400ms, so reuse pressure was REAL in the churn
  arm); rule-of-three bounds the original flake at ~7.5%/run/sha. Specimen 30776330383 remains
  not-reproduced, cause unidentified — the repin removes the false-red MECHANISM; it does not
  adjudicate the specimen. [[IR-16]]'s kill-side arm (TerminateJobObject return discarded) stays
  a live discriminating instrument for any recurrence — and since 2026-08-03 it is no longer the
  only one: [[IR-18]] (the test's OWN bare breadcrumb tree-kill) is the second kill-side candidate,
  test-side rather than product-side. A recurrence must discriminate between them, not assume
  either: IR-16's arm is a job whose termination silently half-fails (survivor still holds the
  captured image); IR-18's is a victim killed by a bare pid it no longer owns —
  `broker_survived=false` with the boot service still ALIVE and its message spooled. (An earlier
  spelling of IR-18's signature here said "empty stderr"; retracted — empty stderr is present on
  healthy runs and discriminates nothing.)

### IR-21 — CLASS: a helper binary's build is never requested, only its LOCATION is, so any narrow invocation manufactures a red that belongs to the rig
- **Status:** BUILT 2026-08-18 on KEYSTONE #182 hygiene lane. · **Origin:** todlando 2026-08-03. Filed as one instance, rewritten as a CLASS the
  same day when a second member appeared, rewritten AGAIN when the mechanism was measured — the
  first two versions described symptoms and got the remedy wrong. **Corrected a fourth time the same
  day**: remedy (1) claimed the 11 same-package sites were missing a build edge, and todlando
  falsified that by measurement while executing it (replicated independently by doyle). The error was
  doyle's to carry — it was ruled on, not just written. The general rule that refutes it was already
  in this entry's own CI-member section; diagnosing a class does not inoculate you against drawing
  its opposite consequence one section later.
- **Numbers below re-derived on `main` @`3efd7e6` over `git ls-files` (tracked files only). Each
  carries its command; re-run rather than cite.**
- ⚠ **Scan-root hazard, measured 2026-08-03 and worse than first reported:** a naive recursive
  `grep -r` from the project root inflates by **~35×**, not the ~5× an earlier draft of this entry
  claimed — 426 tracked `.rs` against 14,865 on disk with `target/` excluded, because `.worktrees/`
  holds full copies of the tree. **And the unscoped grep does not COMPLETE** (still running at a
  120s timeout while `git ls-files` returns instantly), so the failure mode is not only a wrong
  number but a command that reads as hung on a tree where dozens of worktrees are normal. Use
  `git grep` or `git ls-files | xargs grep`; both are scoped to tracked files by construction.

#### THE MECHANISM (measured, both ends)

**Build end — `cargo test` with a narrow target selector compiles a `[[bin]]` AS A TEST HARNESS and
never emits the plain executable.** Measured: `cargo test -p spt --bins` on a clean pool produced
five executables, every one hash-suffixed under `target/<profile>/deps/`, and ZERO plain exes in
`target/<profile>/`. The bin was built. The file the test looks for was not written.

**Consumer end — the resolver borrows a guaranteed binary's path to locate an unguaranteed one.**
All 29 copies of `sibling_bin` in `crates/spt/tests/` reduce to two textual variants (22 + 7) of one
body:

```rust
fn sibling_bin(name: &str) -> PathBuf {
    PathBuf::from(env!("CARGO_BIN_EXE_spt"))
        .with_file_name(format!("{name}{}", std::env::consts::EXE_SUFFIX))
}
```

`CARGO_BIN_EXE_spt` is used **only as a directory anchor**. Cargo therefore sees a dependency on
`spt` and on nothing else; the `{name}` half is a string join it cannot observe. The dependency edge
does not exist at any level a build system could act on — which is why "use the wider command" is
not the fix and why the failure is invisible on a warm pool where some earlier run happened to leave
the file behind.

The unit-test resolver at `crates/spt/src/cli.rs:25601` reaches the SAME directory from a different
anchor (`current_exe()` → `deps/` → parent) for the same reason: its own doc comment records that
unit tests get no `CARGO_BIN_EXE_*` at all.

**That mechanism explains both original members at once, including the polarity inversion that made
them look like different bugs:** whether `cargo test -p spt` or `cargo test -p spt --bins` happens
to leave a plain exe behind is incidental to a dependency neither command was told about.

#### SCOPE — WHERE IT BITES, AND WHERE IT DOES NOT

**It bites NARROW invocations — `-p <pkg>`, `--bins`, `--lib` — i.e. gate rigs and local runs.**
Both members were hit on the LOCKSMITH t1 lane against fresh throwaway pools; each was discharged
in-lane as a rig artifact. That is the whole field population to date.

**GOLDEN IS ACQUITTED, and the acquittal is measured rather than assumed.** `cargo test --workspace
--no-run` emits all 13 plain binaries — mock-session, mock-shell, capture-player,
console-mode-probe, service_fixture and the rest — so a workspace build produces the cross-package
helpers by construction. Confirmed under nextest, the tool CI actually runs, with a real compile.

⚠ **The measurement that establishes this is not the obvious one.** Deleting the plain exe and
watching it reappear is NOT evidence of a rebuild: cargo's uplift restores a hardlink from an intact
`deps/` artifact, with link count 3 and an UNCHANGED mtime, which looks exactly like a build and
says nothing about a pool that never had the file — and a pool that never had it is the only kind a
rig runs on. The sound probe removes the plain exe AND both `deps/` artifacts, confirms all three
absent, then watches the target actually recompile (mock-adapter, 5.77s, fresh mtime). The first two
attempts were setup, not measurement.

⚠ **The suspicion this acquits is structurally well-founded, so record the answer, not just the
verdict.** `twohost-a`/`twohost-b` are `needs: test`, so the only mock-adapter prebuild runs AFTER
the phases that use it; `target/` is gitignored, so checkout never cleans it on a persistent
self-hosted workdir. Every precondition for artifact leakage is present — the build simply does not
need it. The next person who reads that job graph will form the same hypothesis; this is the answer
waiting for them.

#### THE ONE CI-SIDE MEMBER THAT IS REAL

`crates/spt/src/cli.rs::adapter_translate_proof_gates_on_commit` is a **unit** test (`kind(bin)`),
and `CARGO_BIN_EXE_*` is set only for integration tests and benches of the declaring package. A unit
consumer therefore CANNOT express the dependency through the env var at all — it has no choice but
to resolve by path. That is why `.github/workflows/ci.yml:105-106` carries a hand-written
`cargo build -p spt --bin translate_proof_fixture` before the unit lane, with a comment
(`ci.yml:102-104`) stating this exact mechanism.

**Deleting that step would red the unit lane on a clean pool, and it would read as a code red.**
Somebody already hit this class, fixed their own leg correctly, and never turned it into a rule —
the comment at `ci.yml:102` is IR-21 written a milestone early, in the one place only its author
would find it.

#### POPULATION

**13 bin targets in the workspace** (`cargo metadata --no-deps`, not a grep — an auto-target under
`src/bin/` and `xtask` are both invisible to a `[[bin]]` grep). Two are not helpers (`spt`, `xtask`);
the other **11 are test helpers**, across three packages:

| package | helpers |
|---|---|
| `adapters/mock` | mock-session, mock-shell, capture-player, console-mode-probe |
| `crates/spt-daemon` | dispatch_fixture, service_fixture, xlate_choreo_fixture |
| `crates/spt` | translate_proof_fixture, post_step_fixture, gh_fixture, git_fixture |

**44 literal `sibling_bin("…")` call sites**, all in `crates/spt/tests/`, served by **29 copied
resolvers**. Split by whether the BUILD IS ALREADY GUARANTEED — which is not the same question as
whether the env var is available, and an earlier version of this entry conflated the two:

| class | sites | detail | build guaranteed? |
|---|---|---|---|
| **cross-package** | **33** | mock-session 26, mock-shell 6, service_fixture 1 | **NO — the hazard members** |
| **same-package** | **11** | translate_proof_fixture 7, git_fixture 2, post_step_fixture 1, gh_fixture 1 | **YES — already, by construction** |

⚠ **MEASURED TWICE, and it falsifies what this entry said on 2026-08-03 before this revision:** cargo
builds **every bin target of a package whenever it builds ANY integration test of that package** —
that same act is what sets `CARGO_BIN_EXE_*` in the first place. So the 11 same-package sites were
never unexpressed in a way that could bite, and the hazard population is **33 cross-package sites
plus the one unit-test member = 34**, not 44.

- **Probe 1 (todlando, root pool):** deleted all four `translate_proof_fixture` artifacts (both
  hash-suffixed harness exes, the `deps/` plain exe, the uplifted plain exe), confirmed absent, then
  built ONE UNRELATED and UNMODIFIED integration test of the same package —
  `cargo test -p spt --test attach_wedge_e2e --no-run`, exit 0, 14.45s. Both plain exes returned with
  fresh mtimes; the hash-suffixed harness exes stayed absent, which is what distinguishes a bin
  DEPENDENCY build from a `--bins` harness build.
- **Probe 2 (doyle, `.worktrees/gate-ec5f38a` pool, independent replication with a different fixture
  and a different probe test):** deleted `gh_fixture`'s plain exe, its `deps/` plain exe AND its
  hash-suffixed harness exe, confirmed all three absent, then built `--test json_emit --no-run`
  (exit 0) — a test that never names `gh_fixture`. The plain pair returned at a FRESH mtime (09:13
  against the 09:00 it carried before), so this is a real build and not the hardlink uplift this
  entry warns about elsewhere; the harness exe stayed absent.
- **Neither probe needs a baseline arm:** the probe test is unmodified and references nothing under
  edit, so what it measures is cargo's behaviour, not anyone's change.
- **Corroborated by this entry's own field data:** neither original member was a same-package
  integration site — member 1 is a UNIT test, member 2 is CROSS-package. The class never had a
  same-package integration member, and the CI-member section below already stated the governing rule
  (`CARGO_BIN_EXE_*` is set only for integration tests and benches of the declaring package) one
  section before the remedy drew the opposite consequence from it.

Command: `git ls-files '*.rs' | xargs grep -hon 'sibling_bin("[a-z_-]*"' | sed 's/.*sibling_bin("//;
s/"//' | sort | uniq -c`.

⚠ An earlier report gave 48 sites and a 6/11/37 split. Take the table above: it counts only literal
call sites in tracked files and it ships its command. The 11 is the same 11 in both counts.

**The LOCATION half of this class is already closed** by `crates/spt-term/tests/support/fixture_bin.rs`
— a shared resolver rather than 29 copies. It does not close the BUILD half, and should not be
mistaken for having done so.

#### REMEDY SHAPE (not ruled)

The distinction that matters is location vs. build:

1. **The 11 same-package sites are NOT hazard members and need no build fix.** Their build edge
   already exists (see the two probes in POPULATION above); `env!("CARGO_BIN_EXE_<name>")` for the
   fixture itself would add nothing to it. Converting them is a **CLARITY** change, worth doing on
   its own smaller merits — the path becomes the one cargo actually emitted rather than a string-join
   of a directory anchor and `EXE_SUFFIX`, five copies of `sibling_bin` stop existing (29 → 24), and
   it completes a migration already paid for: `crates/spt/tests/fixtures/translate_proof_fixture.rs:7-12`
   records that the fixture was re-homed into `spt` precisely to obtain
   `CARGO_BIN_EXE_translate_proof_fixture`, and then all 7 call sites resolved by path anyway. **It is
   not a build fix and must not be filed as one.**
2. **The 33 cross-package sites cannot**, by cargo's design. Their options are an asserted build in
   the test's own setup, a `dev-dependencies` artifact dependency, or an explicit documented
   prebuild — the ci.yml:106 shape, made a rule instead of a local fix.
3. **Standardising on one wider invocation is NOT a remedy.** It changes which pools happen to work;
   it does not create the dependency edge, and the two original members had opposite polarity under
   exactly that theory.
4. Collapsing the 29 resolver copies is worth doing on the `fixture_bin.rs` model, but on its own it
   makes the class HARDER to see — one shared resolver still anchored on `CARGO_BIN_EXE_spt` hides
   the 33 genuinely unexpressed dependencies among its 44 call sites behind one function.

#### SUPERSEDED FRAMING, KEPT SO IT IS NOT RE-DERIVED
- **Status of the original filing:** open · **Origin:** todlando 2026-08-03, both members hit on the
  LOCKSMITH t1 lane against fresh throwaway pools; each discharged in-lane as a rig artifact, filed
  here so the next clean rig does not re-diagnose them as code reds.
- The first two versions of this entry framed the class as "the rig's command does not build what
  the test needs" and proposed standardising on a wider invocation. **Both are superseded by the
  measured mechanism above** — the dependency is not under-expressed, it is INEXPRESSIBLE in the
  form these call sites use, so no choice of invocation creates it. Kept only as the two FIELD
  MEASUREMENTS that produced the class, which remain true:
- **MEMBER 1 — MEASURED:** `cli::tests::adapter_translate_proof_gates_on_commit` failed on the
  first `cargo test -p spt --bins` run. Its fixture binary `translate_proof_fixture` (a
  `tests/`-homed `[[bin]]`) was ABSENT from the pool — `ls` on the path returned No such file.
  Building it explicitly and re-running the IDENTICAL command PASSED, after which the full `--bins`
  suite passed 584/585 with only the releases#117 probe red (that one RED by design). So the
  discriminator is the fixture's presence, not the tree: same command, same sha, red then green
  across one `cargo build` of the fixture.
- **MEMBER 2 — MEASURED:** `cargo test -p spt` does not build `mock-adapter --bin mock-session`, so
  `attach_wedge_e2e` panics `"the dummy-harness program must be built"`. Note the polarity is
  INVERTED against member 1 — there the narrower `--bins` was the defective invocation and
  `cargo test -p spt` the correct one; here `cargo test -p spt` is itself insufficient. So the
  class is NOT "use the wider command"; it is that the dependency is not expressed to the build at
  all, and which invocation happens to work is incidental.
- **What/why (still true):** on a WARM pool the helper is already there from some earlier run and
  the test passes, so the defect is invisible exactly where most people work and fires only on a
  clean pool — i.e. on a GATE RIG, which is the one place a false red costs the most.
- **The sweep this entry once called its first step HAS BEEN RUN** (todlando 2026-08-03) and its
  result is the POPULATION section above. It is no longer outstanding.

- **Why it is register debt and not a lane bug:** nothing in the product is wrong. The gap is
  between what a test needs built and what the rig's command builds, and the fix belongs to the
  tests' declarations, not to whoever is running a gate that day. Sibling rule:
  [[gate-clean-target-not-incremental]].
- **Built evidence:** golden now names every cross-package fixture prebuild
  (`mock-session`, `mock-shell`, `capture-player`, `console-mode-probe`,
  `service_fixture`) explicitly; `xtask binedge-check` reports zero missing
  edges. All 29 local `sibling_bin` resolvers collapsed into
  `tests/common/mod.rs`; the shared precondition names the missing fixture and
  its package-correct build command before any product timeout.
- **Ripe when:** the next gate-rig or CI-touching wave. Cheap, and it pays for itself the first time
  it stops someone chasing a phantom red on a clean pool.
- **Size:** remedy (1) is small AND optional — it buys clarity and subtraction, never a build edge;
  medium for (2), which is a design call before it is an edit and is the only remedy that closes the
  class.

### IR-22 — An inherited identity env var fails a test, and the diagnostic names them ONE AT A TIME so a correct fix reads as no fix
- **Status:** open · **Origin:** todlando 2026-08-03, chasing what looked like an `attach_wedge_e2e`
  code red on the LOCKSMITH t1 lane; root-caused to the runner's own process environment.
- **MEASURED:** `attach_wedge_e2e` failed for an INHERITED PROCESS-GLOBAL and nothing in the tree:
  the daemon-stop refusal fired on the running session's own identity env. It named
  `$OWL_SESSION_ID`; clearing that made it name `$SPT_ENDPOINT_ID`. With `OWL_SESSION_ID`,
  `SPT_ENDPOINT_ID`, `SPT_AGENT_ID` and `SPT_SESSION_ID` all cleared: exit 0, 1 passed.
- **The finding is the DIAGNOSTIC SHAPE, not the env hygiene.** Naming one variable at a time means
  a correct partial fix produces an identical-looking failure, so the natural reading of "I cleared
  it and it still fails" is that the clearing did not work — when in fact each step was right and
  the message had simply moved on to the next name. A refusal that can only ever name one member of
  a set it is checking teaches the person debugging it the wrong lesson. Compare the same class in
  [[IR-15]]/[[IR-18]] terms: the instrument is competent and the report is not.
- **Why it matters beyond one test:** any agent running suites from a live spt session carries these
  vars, so this fires for every builder on a perched session and for nobody running from a bare
  shell — which is precisely the split between how builders work and how CI runs.
- **Candidate remedies (not ruled):** have the refusal name EVERY identity var it found set, in one
  line, rather than the first; and/or have the affected tests clear the identity set in their own
  setup so a perched session is not a special environment. The first is the one that pays off
  outside this test.
- **Ripe when:** next CI/test-hygiene wave. **Size:** small.

### IR-23 — `endpoint_teardown_authority_e2e`'s two tests collide with EACH OTHER through the machine-global spt home
- **Status:** BUILT 2026-08-18 on KEYSTONE #182 hygiene lane. · **Origin:** todlando 2026-08-03, LOCKSMITH t1 lane.
- **MEASURED:** both tests copy a psyche binary fixture into `perch::spt_home()/srcs/dummyharness/`
  (`crates/spt/tests/endpoint_teardown_authority_e2e.rs:399`), which is process-global, so run in
  parallel inside one binary each holds the file the other wants: `os error 32`, "used by another
  process". **Signature: WHICH of the two fails alternates between runs.** Under
  `--test-threads=1`: 2 passed, exit 0 — and the wall clock drops from 182s to 14s.
- **What/why:** two defects in one, and they should not be conflated. (1) The tests are not
  isolated from each other. (2) The staging path is the MACHINE-GLOBAL spt home rather than the
  test's own temp home — on a box that is also a CI runner, so the blast radius is not confined to
  the suite. (2) is the one worth fixing; (1) is a symptom of it.
- **Note on why this is not already caught:** the standing rule is that integration tests run
  through nextest, which gives each test its own process and hides the collision entirely. So this
  is latent on the sanctioned path and only bites the bare `cargo test` path — the rule is working
  and masking a real defect at the same time, which is why the entry exists rather than a shrug.
- **Candidate remedy (not ruled):** stage the fixture into the test's own temp home. Sweep for
  siblings first — any other test writing under `perch::spt_home()` rather than a temp home shares
  the shape, and nobody has counted them.
- **Built evidence:** adapter source, manifest registration, and psyche fixture
  staging now use the rig TempDir. The two process-global resolver users are
  explicitly serialized, and the bare `cargo test` path passes both cells with
  `--test-threads=2` in 13.82s after the declared fixture prebuild.
- **Ripe when:** next test-hygiene wave; the 182s -> 14s figure makes it pay for itself on the
  bare path. **Size:** small per test, unknown until the sweep.

### IR-24 — `reap::terminate_job` discards `TerminateJobObject`'s return, alone among its own module's siblings
- **Status:** BUILT — landed `3d12cf3` on main (2026-08-03), fulfilment verified against this
  entry's own wanted list by doyle 2026-08-04, not merely mapped to the commit: the return is
  audited and the loss NAMED (`REAP_JOB_TERMINATE_FAIL: brain subtree NOT reaped (<os error>); …
  the daemon holds no other handle to reach them` — which also ANSWERS the cost-of-loss question
  this entry said nobody had written down); the "first thing to check" was checked and the ruling
  came out OPPOSITE to IR-16's site in one respect, documented as measured asymmetry in the doc
  comment: `TerminateJobObject` has NO ordinary failing path (zero only for a broken handle) so a
  zero is a real loss, while `TerminateProcess` fails `ERROR_ACCESS_DENIED` on the ordinary
  already-exited path and is deliberately NOT checked — the premise is unit-PINNED both directions
  (`terminate_job_return_discriminates_broken_handle_from_ordinary_states`, cfg(windows), with a
  memberless-job false-positive arm). No `job != 0` re-check (second-weaker-rule prohibition
  honored, same as IR-16). Golden evidence: rides `4b37512` (run 30873007187 attempt 5, 9/9
  green; the premise unit observed PASS in-pool on the Windows leg, Linux leg green — the
  cross-platform quiet-`ESRCH` concern stays with `ec5f38a`'s coverage as before). Path
  correction to this entry's MEASURED line: the site is `crates/spt-daemon/src/reap.rs` (brain
  subtree reaper), not `crates/spt/src/reap.rs` as first filed ·
  **Origin:** todlando 2026-08-03, swept while building [[IR-16]]; deliberately
  NOT folded into that lane (doyle ruling, same day) and filed instead.
- **MEASURED:** `crates/spt/src/reap.rs:207-211` calls `TerminateJobObject` and discards the result,
  the identical shape [[IR-16]] closed at `spt-daemon`'s `kill_tree`. Its own module's
  create/assign siblings at `:178` and `:202` ARE instrumented, so this call is the odd one out
  where it lives — the module already decided that these returns are worth reading.
- **Why it was NOT folded into IR-16:** different subject. IR-16's site is a SUPERVISED SERVICE's
  teardown, where the loss is "a supervised service's descendants survive" and the promise it
  breaks is `REQ-RESIDENT-SERVICE`'s tree claim. This site is the BRAIN SUBTREE's kill-on-close
  job, which has its own lifetime, its own caller and its own answer to "what does a failed
  termination cost here" — and that answer has not been written down by anyone. Folding it in would
  have meant settling that question in passing, inside a commit about the supervisor, which is how a
  second ruling gets smuggled into a lane scoped to one.
- **What it needs that IR-16's fix does not supply:** IR-16 ruled a LOUD REFUSAL (option B) on the
  ground that the alternative — a process-table tree-walk fallback — is measured blind in exactly
  that failure state ([[IR-18]]: descent breaks at the dead middle hop) and would trade a silent
  loss for a false clean. Whether the same reasoning holds here depends on whether this site kills
  its direct process in the same breath, which is what makes the middle hop dead by construction
  there. **That is the first thing to check, and it is not assumed.**
- **Ripe when:** any wave touching reap/teardown; it inherits IR-16's vocabulary and its
  discrimination note, so the second one is cheaper than the first. **Size:** small, once the
  cost-of-loss question is answered for this subject.

### IR-25 — `spt-daemon --lib` reds 2-in-3 under concurrent load with bare `cargo test`, and holds green under nextest
- **Status:** open · **Origin:** doyle 2026-08-03, found while gating LOCKSMITH tranche 1 in an
  isolated worktree; chased to a mechanism and scoped OUT of the release path before the head was
  assembled.
- **MEASURED, four arms, same box, same hour, matched load** (a second `cargo test -p spt-store
  --lib` loop running throughout; `spt-store` itself stayed green in every round, so the box was not
  generically failing):

  | tree | instrument | result |
  |---|---|---|
  | lane `ec5f38a` | `cargo test --lib`, at rest | 821/821 ×4 |
  | lane `ec5f38a` | `cargo test --lib`, under load | **2 of 3 RED** (142s, 117s) |
  | lane `ec5f38a` minus its 5 new tests | `cargo test --lib`, under load | 3 of 3 green (117s, 84s, 105s) |
  | main `932e14b` | `cargo test --lib`, under load | 3 of 3 green (97s, 119s, 104s) |
  | lane `ec5f38a` | **`cargo nextest run`, under load** | **3 of 3 green, 821/821** (88s, 85s, 72s) |

- **Three DIFFERENT victims across two loaded runs**, all process/timing-shaped, all PRESENT AT MAIN
  and green there: `servicehost::…a_service_that_ignores_the_stop_marker_is_force_killed_and_confirmed_dead`
  (panics `ForceKilled must MEAN the owned child is gone` right after `SERVICE_KILL_UNCONFIRMED`),
  `applyhost::…broker_reports_its_compiled_image_version_over_ipc`, and
  `daemon::…a_tree_teardown_reaches_a_grandchild_the_service_spawned`.
- **MECHANISM: in-process thread interference, not product code.** Bare `cargo test` runs a crate's
  tests as threads in ONE process. The lane's five new tests are process-spawning kill-tree tests;
  under load they starve confirm windows of NEIGHBOURING tests that were always marginal. Removing
  exactly those five flips 2-of-3-red to 0-of-3-green at the same sha, in the same binary, at the
  same durations — one of the green runs sits at 117s, precisely a duration that had produced two
  failures with them present.
- **IT DOES NOT REACH GOLDEN.** nextest gives every test its own process, and golden runs nextest.
  The population is GATE RIGS AND LOCAL RUNS — the same population as [[IR-21]], reached by a
  different mechanism, which is why these are two entries and not one.
- ⚠ **The instrument switch is the whole finding.** Measured with bare `cargo test`, this reads as a
  lane blocker; measured with the tool CI actually runs, it is a rig-scoped nuisance. Anyone
  re-opening this must state WHICH instrument produced their observation before quoting a verdict.
  The same tool gap acquitted [[IR-21]]'s golden question the same day.
- ⚠ **The pre-existing fragility is NOT thereby closed.** Those three tests were marginal before this
  lane and the lane only exposed them; "make the new tests quieter" would re-hide a real weakness.
  Hardening them is test-side work (hertz's lane by the dispatch split), not the builder's.
- **Two rig defects of the gater's, recorded because they nearly cost the verdict:** (1) a first
  re-run at rest came back 4/4 green and proves NOTHING — a sequential idle probe cannot express a
  load-sensitive failure, so that is a competence-controlled zero, not an acquittal; (2) the first
  baseline attempt ran its two arms UNSYNCHRONISED (the load generator finished while the measured
  arm was still on run 1), which would have compared a loaded lane against an idle main and read as
  "the lane broke it". A discriminator whose arms ran under different conditions cannot discriminate.
- **Ripe when:** alongside [[IR-21]] on the next gate-rig or CI-touching wave, or immediately if
  anyone starts gating on bare `cargo test` on a loaded box. **Size:** small to document the rig
  rule (gate with nextest); medium to harden the three marginal tests.

### IR-26 — POOL-OWNER claims authorize takeover from a DEAD holder, so releases#103's hazard reaches through the guard
- **Status:** open, LANE BUILT BUT UNCOMMITTED (see the warning below — this is not a figure of
  speech) · **Origin:** hertz 2026-08-03/04, found by its own test rather than by reading, and
  reported unfiled for composition. Composition ruling is doyle's, 2026-08-04.
- **What/why:** a POOL-OWNER claim named a holder pid; lane liveness was read from whether that pid
  was alive. hertz MEASURED the failure on this box: **3 of 3 claims had dead holders while one of
  those lanes was live** — so the guard authorized exactly the interleaving takeover releases#103
  exists to prevent. The remedy records the lane's GIT IDENTITY (branch + the base sha it carried at
  claim time) and reads lane state from ancestry: in flight while the tip is not contained in
  `origin/main`; settled once merged, once the branch is gone, or once the branch no longer carries
  the claimed base (the branch-level twin of a recycled pid). Holder pid demotes to advisory — a
  LIVE holder still refuses, a dead one no longer authorizes.
- **The strongest finding in it, and it is a filter-competence one:** without the base anchor, a
  missing branch read as "settled, branch gone" **out of a directory that was not a repository at
  all**. An unanchored Settled verdict is a clean zero produced by a filter that cannot express the
  question. A Settled verdict now requires the base object to be present in the reading repo;
  unanchored is UNKNOWN. Sibling of [[zero-match-filter-reads-as-absent]] in the pool domain.
- ⚠ **POLARITY, stated rather than buried:** the new arm (dead holder + unlanded lane => REFUSE)
  converts a false-TAKEOVER into a false-REFUSAL. That is the better failure but not a free one, and
  its triggering case is not exotic — it is an agent going down mid-lane, which happened to hertz
  itself this same batch and cost [[IR-1]]/[[IR-4]]/[[IR-9]] their landing. **The open gating
  question, put to hertz 2026-08-04 and NOT yet answered:** from the refusal state, does
  `pool-release` still clear a claim whose holder is dead and whose lane is unlanded, or does the new
  arm refuse that too? A guard whose only escape is the blanket `SPT_POOL_UNCHECKED=1` trains people
  onto the blanket override. The refusal must NAME the remedy, not only the witness, and the remedy
  must be exercised FROM the refusal state rather than reasoned about ([[remedy-must-run-from-refusal-state]]).
  If the answer is "it refuses", that is a design change and returns to doyle before the lane lands.
- **COMMITTED 2026-08-04 (was uncommitted when first measured):** `ci/poolowner-lane-claim` =
  `8ce40d9` (the change) + `98cfafe` (the remedy fix below), both off `b7b00c3`. The original
  measurement found ZERO commits and six modified files in `.worktrees/hertz-poolowner`, i.e. every
  reported number had been measured against a tree no git object captured; the builder committed on
  being told. Recorded because the near-loss is the lesson, not the correction.
- **THE REMEDY DID NOT RUN — measured from the refusal state, which is the only place the question
  can be asked.** The refusal named its witness AND a command, and executing that command verbatim
  failed with the IDENTICAL `SPT_POOL_FOREIGN` block that printed it: **xtask depends on spt-store,
  so every command that builds the tool goes through the pool being refused.** This is the same dead
  end the unclaimed-foreign arm already hit in the field and fixed by leading its line with the
  hatch — adding a second refusal arm re-broke it immediately. FIXED at `98cfafe`: the in-flight
  remedy now leads with the hatch and names `pool-claim`, the command that actually resolves
  ownership. Full chain re-measured on hfenduleam: refuse → verbatim remedy fails → hatch-prefixed
  form succeeds → build proceeds.
- **Answer to "does `pool-release` clear a dead-holder + unlanded claim", two-part:** it DOES — it
  never consults the verdict, it rewrites `lane=None` — **but it is NOT the exit**, because after
  release the same build is refused again on the pre-existing unclaimed-foreign arm ("no lane has
  claimed this pool"). Release moves you from one refusal to another. `pool-claim` is the exit.
- **The guard is a POPULATION, not a row** (builder's generalization, adopted by doyle as the
  standard for this class): one test poses all three refusal arms and asserts that any remedy naming
  a command leads with the hatch, PLUS membership — three posed, three refused — so no arm passes by
  never being reached. The rule had been learned once already and the next arm added broke it, which
  is the definition of something that needs a population test rather than a better comment. 26/26.
- ⚠⚠ **FLEET PROPERTY — the rig refuted itself, and this is the entry's most important line.** The
  first arm did NOT refuse, it TOOK OVER: the arriving worktree was at `b7b00c3`, so the build script
  that ran was ITS OWN, pre-identity. **Enforcement code comes from the ARRIVING tree, never from
  the claim.** So until this lands everywhere, an old tree still takes a live lane's pool, and the
  guard is only ever as new as the tree that arrives. Found by measurement; reasoning would have
  reported a refusal the rig never produced.
- **RULING on the fleet property (doyle 2026-08-04) — a stamp VERSION LINE is REFUSED, and not on
  taste: it cannot work.** The tree that must refuse is the OLD one, and an old build script does not
  read a new stamp field because it does not know the field exists. You cannot make a stale enforcer
  honor a new rule by writing something new into the record it enforces against — **the reader is
  the problem, not the record.** A version line would help only a CURRENT tree explain itself while
  doing nothing to the population it is aimed at, i.e. it would retire the worry without retiring the
  risk. Ruled instead, in priority order: **(1)** land it and state the migration window honestly —
  the arm protects a lane only when the ARRIVING tree carries the change, coverage grows as trees
  turn over, never retroactive; **(2)** NAME THE LOSS YOU CANNOT PREVENT ([[IR-16]]'s pattern) — the
  claiming side may be able to notice afterwards that its pool was taken and say so loudly, turning
  silent artifact corruption into a named event; scope separately and MEASURE whether the claiming
  side can observe it at all before committing; **(3)** SHRINK THE POPULATION — old enforcers are old
  trees, this box carries ~45 worktrees and most are stale, so reaping them is real mitigation and
  cheaper than code. (2) and (3) do NOT ride this lane.
- **The unanchored row, as asked:** `a_claim_from_another_repository_reads_unknown_not_settled` was
  RED before the base-object requirement existed. Verbatim: `left: Settled { reason: "lane branch
  build/lane-x no longer exists" }  right: Unknown`, out of a tempdir that was not a repository at
  all. It fails without the requirement and passes with it — a real negative control, not a row that
  only ever passed.
- **Builder's evidence (as reported, against that uncommitted tree):** spt-poolguard 25/25 with 7 new
  rows, one per precedence cell, two carrying their own negative control; pool_guard_canary 3/3;
  `cargo check -p xtask -p spt-store --all-targets` clean; clippy clean on the three touched crates;
  `traceable-reqs check` exit 0 at 745/745 with new `REQ-POOL-LANE-IDENTITY` complete at doc+impl+unit.
- **Composed:** its OWN thin lane, riding the same next golden batch as the CI-rider cluster but NOT
  merged into it (doyle ruling 2026-08-04). A red must stay attributable: the riders change what the
  pipeline measures, this changes a build-time REFUSAL, and a guard whose failure mode is refusing
  builds is the worst thing to make inseparable from a pipeline change. IR-14's four-trees-into-one-
  unclaimed-pool finding stays OUT of this lane and stays IR-14's, on the same ground [[IR-24]] was
  kept out of [[IR-16]] — folding it in settles its question in passing.
- **Ripe when:** next CI/infra batch, gated on the polarity question above being answered first.
  **Size:** small once the remedy question is settled.
- **FALSE-LIVE, first specimen of the blocking polarity (hertz sweep, measured 2026-08-04):**
  `.worktrees\hertz-ci-riders\target\POOL-OWNER.json` named `holder_pid 31380` with
  `holder_started_at 134302739211259713` = 2026-08-03 16:38:41. Pid 31380 on that box at read time was
  `claude.exe -n "doyle @ HFENDULEAM (spt-core/)"`, created 17:31:11, FILETIME 134302770715019370.
  Same boot, ~52 minutes apart, stamps MISMATCH. The claiming process died and the OS reissued its pid
  to the gater's own session. Every previously measured recycle produced a false-DEAD (reads free
  while the lane is live — the corrupting direction); this one produces a **false-LIVE**: a bare-pid
  guard reads the claim as held by a running holder and refuses, naming doyle as the holder of hertz's
  pool — the blocking direction, first time caught. The `holder_started_at` arm shipped in `8ce40d9`
  discriminates it correctly — that field caught defending, not asserted to defend
  ([[stable-anchor-is-not-a-recycling-defense]]).
- **The mirror image, same box, same hour:** `.worktrees\ir24-reap\target` named `holder_pid 33612`,
  DEAD — while that lane was genuinely LIVE, with `cargo clippy -q -p spt-daemon --lib --tests`
  (pid 22452) running under todlando's session as it was read. Two claims, opposite errors, inside one
  hour: liveness misreads in BOTH directions on the same box in the same hour. Holder liveness has no
  sound polarity as a lane-state oracle. Ancestry decides ([[pool-claim-holder-death-is-not-lane-state]]).
- **THE BARE-CLAIM GAP — a claim the stamp arm cannot adjudicate at all.** The main checkout's own pool
  (`<root>\target`, 72,001,016,671 B = 67.06 GiB) carries
  `{"owner_tree": "...\\spt-core", "written_by": "spt-poolguard"}` — no `lane_label`, no `holder_pid`,
  no `holder_started_at`, no `lane_branch`, no `lane_base`. `read_owner` maps that to `lane: None`, so
  every arm the lane-identity work added is unreachable for it: it cannot be read as in-flight,
  settled, or recycled. **FIVE of the eight** claims found on this box are this shape — `<root>\target`,
  `<root>\target-seam` (below), `gate-target-render`, `ir21-xtask\target`, `locksmith-t1\target` —
  against three carrying a lane identity (`hertz-ci-riders`, `hertz-poolowner`, `ir24-reap`). A claim
  written by a plain build rather than by `pool-claim` is structurally un-adjudicable, and it is the
  MAJORITY shape in the field. Whatever IR-26 does next has to say what an identity-less claim means —
  "the stamp arm handles it" is false for most claims that actually exist.
- **The fifth bare claim was invisible to the first inventory, and the miss is its own lesson:**
  `<root>\target-seam`, **52.06 GiB**, real dir, gitignored on its own line (`/target-seam`, added
  `5ae68f8`, v0.12.1 era), created 2026-06-18, last written 2026-08-02 — nothing in the tracked tree
  references it (grep over rs/toml/md/yml/ps1/sh: zero hits outside this file). The inventory predicate
  was "a directory named `target` inside a worktree"; a second root-level pool under a different name
  is not expressible in that filter, so it returned a clean zero that was reported as a population
  ([[verdict-from-probe-competence]] in the field). Corrects the earlier "72 GB root pool is the
  biggest object by 9x": the two root pools are 67.06 + 52.06 = 119.1 GiB of a 133.43 GiB box pool
  total, so the 44-tree sweep covered ~61% of pool bytes, not the population. **Ruling (doyle,
  2026-08-04): target-seam stands for now** — free space is ~115 GB against the 32 GB golden floor, so
  there is no pressure; it was written to yesterday, so "cold" is one day deep and its writer is
  unidentified. It becomes a reap candidate at the next free-space squeeze, after a fresh last-write
  check and identification of what wrote it on 08-02 — reaping a pool whose writer you cannot name is
  how the v0.51.0 fabricated-red class starts.
- **CLAIM-OBSERVABILITY, MEASURED (hertz, 2026-08-04, three-arm rig with competence control):** a
  displaced claimant CANNOT distinguish takeover from never-claimed — the two states are byte-identical
  on disk (arm 1 vs arm 2 identical; arm 3 proves the rig separates states that differ, so the null is
  real). The takeover is silent on BOTH sides: the taker's `pool-claim` prints ordinary success naming
  no prior owner. Mechanism at `b75258d`: `pool_claim` never calls `read_owner` — it constructs its own
  `PoolOwner` and `fs::write`s unconditionally. The displaced lane's only channel is its next build's
  refusal, which reports a STATE ("this pool belongs to B"), never a transition, and arrives pull-only
  and arbitrarily late while interleaved-artifact harm is already underway. Labelled holes: measured
  against the origin/main xtask binary (lane-tip carry is a source read, not a measurement);
  pool-sweep-as-channel not exercised; single-process holder pid in all arms.
- **Bootstrap property, not a defect (hertz, measured 2026-08-03 while claiming pools for the lanes
  3-4 rebases):** a lane DELIVERING claim-identity cannot claim its own pool WITH that identity until
  it lands — the prebuilt origin/main `xtask.exe` writes old-shaped stamps (owner_tree, lane_label,
  holder pid+stamp; no lane_branch, no lane_base). Enforcer turnover surfacing at CLAIM time rather
  than build time: coverage grows only as trees turn over. Recorded because it reads as a bug the
  first time someone meets it in the field. Related lane-craft fact from the same queue: a rebase
  changes SHAS necessarily but hunk OFFSETS only incidentally — todlando's main.rs hunks stayed at
  @20/@53/@70 across the f2e4516→2f427b2 rebase; treat pre-rebase overlap measurements as stale by
  construction and re-measure, but expect the geometry to usually hold.
- **Bootstrap property FIELD INSTANCE, main checkout (todlando 2026-08-04, IR-12 rig):** the
  prebuilt `target/debug/xtask.exe` in the MAIN checkout pool was stale against main's own source
  at `11169c1` — it printed `(holder pid …, born …)` and wrote a FLAT identity-less claim
  (owner_tree/lane_label/holder_pid/holder_started_at/written_by) where main's source
  (main.rs:2282) prints `(branch …, base …; advisory holder pid …)` and writes the nested lane.
  A lane claimed with a stale tool records NO git identity and demotes its own adjudication back
  to the refused bare-pid predicate. DETECTION that worked and is the reusable part: read the
  claim line the tool PRINTS against the source just read — a mismatch names the stale enforcer
  before it writes. REMEDY at claim time: rebuild xtask FROM THE LANE'S WORKTREE before
  `pool-claim` (which the protocol's claim-from-your-own-worktree rule already implies; this
  instance is why it is load-bearing and not ceremony).
- **RULING (doyle, 2026-08-04) — the claim path must run the same guard the build path runs, and a
  takeover becomes an event.** The documented design already says a finished lane is "taken over
  loudly, not refused"; a silent unconditional overwrite violates the spelling the protocol shipped
  under. Remedy, both halves, one small lane: (ii) `pool-claim` reads the prior owner before writing —
  prior lane LIVE (ancestry + stamp arm, never bare pid) ⇒ refuse exactly as build.rs would; prior
  lane finished ⇒ proceed AND announce the displacement, naming the displaced lane in the success
  line; plus (i) a `displaced` field (prior owner + takeover stamp) rides the new claim, so the fact
  survives in the artifact the late-reading claimant already reads. The identity-less bare claim
  (majority field shape, above) must be handled explicitly by the same change: a bare claim carries no
  lane to adjudicate, so takeover of a bare claim always proceeds-and-announces, and the announcement
  names the claim as identity-less — never silently, never refused. Ripe: next CI/infra batch, same
  lane as or beside the IR-26 remedy. Size: small.

### IR-27 — A junction-pooled rig is TWO objects and teardown only ever removes one
- **Status:** open, found in the wild 2026-08-04 · **Origin:** hertz's 44-tree worktree sweep
  (IR-26 follow-on 3) · **Cross-ref:** [[IR-14]] (unclaimed pools), [[worktree-target-junction]]
- **What/why:** a rig built as `worktree\target -> junction -> .worktrees\<pool>` has its build cache
  in a directory that is NOT the worktree and NOT in `git worktree list`. `git worktree remove` deletes
  the junction and reports success; the pool survives with zero inbound links and nothing naming it.
  Measured on this box: reaping `golden-render` and `w4-cli-doc` the obvious way would have reclaimed
  ~26 MB of worktree and stranded **14.17 GB** of pool (`gate-target-render` 6,265,448,382 B +
  `gate-target-w4doc` 7,907,558,098 B). Neither pool appears in any worktree listing.
- **The second shape, same mechanism, already realised:** `.worktrees\gate-target` was found at
  **0 files / 0 bytes with FOUR live inbound junctions** (assembly-doorbell, doorbell-w1/w2/w3). Someone
  had already reaped that pool and left four dangling-but-valid links behind. So the defect runs both
  ways: remove the worktree and the pool is orphaned; remove the pool and the links are orphaned.
  Either half alone leaves a lie on disk.
- **Why it is not just "be careful":** the classification step people already know
  (`Get-Item -Force` OUTBOUND, INBOUND reparse sweep) tells you HOW to delete an object you have
  already decided to delete. It does not tell you that a second object exists. The missing thing is an
  enumeration that reaches pools no worktree names.
- **Remedy shape (CORRECTED by builder retraction, 2026-08-04 — the enumeration already exists and
  shipped):** `xtask pool-sweep` (`c6515f1`, `REQ-POOL-GC-ORPHAN-RECLAIM`, ancestor of `b7b00c3` by
  merge-base) walks the root for dirs AND links, counts inbound links over the whole sweep before
  forming a verdict, classifies InUse/Owned/Orphaned, and reaps orphans on `--reap`. The gap is not
  detection, it is TIMING: a junction-pooled tree's pool is `Owned` and not reclaimable right up to
  the moment its worktree is removed, and `Orphaned` only after — so the reclaim always lands in a
  LATER sweep, and nothing triggers one. `gate-target` (0 bytes, four dangling inbound links) is a
  completed instance of that deferral. Remedy is a post-teardown pass or a teardown that runs the
  sweep itself, not new enumeration. (An earlier draft of this entry proposed building the
  enumeration; retracted by its author before ripening — the tool refuting it was one run away.)
- **Ripe when:** next CI/infra batch. **Size:** small (one enumeration + report, no new policy).
- **Discipline this cost, stated so it is not re-learned:** teardown must be TWO EXPLICIT STEPS per
  junction-pooled tree — junction deleted AS A LINK (`Directory.Delete(path, recursive: false)`), then
  the pool AS A TREE after an inbound sweep confirms it is linkless. Executed that way on 2026-08-04:
  28 worktrees, 4 pools, 0 skipped, 0 pinned, 17.04 GiB measured on a box with concurrent writers.

### IR-28 — A wave-gap rider is visible to the request board and invisible to a milestone-narrative notes pass
- **Status:** open · **Origin:** deployah, v0.53.0 release close 2026-08-04.
- **What/why:** the v0.53.0 changelog omitted `184f2ac` (the releases#125 two-core CPU-burn fix, a
  wave-gap rider that rode the commit range but was never LOCKSMITH scope) while the alchemy release
  verb promoted #125 to DONE and named it publicly as shipped — two published surfaces contradicting
  each other until the post-publish docs-only repair (`9841592`). The rule was already right (the
  runbook audits the changelog against the COMMIT RANGE); the EXECUTION missed it, because a notes
  pass written from the milestone narrative never enumerates the range. The mechanism, not the
  incident: any rider is board-visible and narrative-invisible.
- **Remedy shape:** a mechanical range-vs-notes diff at release time — list `vPREV..HEAD` commits of
  user-facing type (fix/feat), match each against a changelog mention, print the unmatched. An xtask
  step or runbook checklist line; the repair path (docs-only amend + `gh release edit`, no retag) is
  now precedent twice over.
- **Ripe when:** next release-tooling or runbook wave. **Size:** small.

### IR-29 — twohost ladder: role A fires the done-barrier BEFORE its last cross-node rung, so the re-pull's serve window is won by timing, not guaranteed
- **Status:** open, FIX AUTHORED-AND-COMPILED unlanded — hertz `test/twohost-serve-window`
  @`17bbbd8` off `11169c1` (+33/-18 twohost.rs, relocation proven byte-identical by
  comment-stripped set-diff, cargo check + clippy -D warnings green; proving two-host run still
  owed and rides the same window as #145's closed-subnet knock row) · **Origin:** golden
  30873007187 attempt 1 red, doyle timeline triage + hertz source diagnosis 2026-08-04.
- **Mechanism (hertz falsified the gater's first spelling, kept because the correction is the
  entry):** NOT "B exits early" and NOT a missing barrier. B's contract (twohost.rs:1116-1121,
  "hold until A finishes its side — done-file as the ladder's completion barrier") is right and B
  obeys it. A BREAKS it: A pushes done.txt (:1819-1832, payload literally "ladder complete on A")
  then runs the ENTIRE leg-D digest rung including the re-pull assert (:1959) before its own
  `stop.store` (:1966). B tears down exactly when A said it could; the re-pull then races B's
  post-signal teardown. The panic's hint text ("may be running a version without cross-node
  digest") is pre-authored and was NOT evidence of version skew.
- **Margin series, six draws of one race (deployah, markers cited by TEXT and positive-controlled
  against a known draw before producing an unknown one):** golden 30860770146 +60s · a1 −0.6s
  (the RED) · a2 +59.4s · a3 +60s · a4 +59.59s · a5 +63.89s. BIMODAL, not drifting: greens
  cluster within seconds, one collapse through zero. A green does not mean the window got safer;
  sample count characterizes a bimodal race, not margin size — the slack is incidental, not
  designed. Local load on hfenduleam during A's leg D eats the margin directly (role A runs
  there; the done-file is already pushed), which is why builds hold through twohost legs.
- **⚠ THE SERIES IS CLOSED AT SIX DRAWS — ITS DERIVATION IS UNRECOVERABLE, AND NO SEVENTH NUMBER
  MAY BE ADDED (ruled by doyle 2026-08-05, at the tranche-2 golden).** Asked to extend it with a
  draw from run 30971976024, deployah positive-controlled the markers FIRST and the control
  REFUSED: grepping `margin|deadline|serve window|slack|headroom|budget` against a control run
  known to carry good twohost evidence (30940180764) returned ZERO. Source confirms it —
  **`margin` is not an emitted token anywhere**: absent from `.github/workflows/golden.yml`, and
  every `crates/` hit is unrelated (terminal right-margins in broker/resize tests,
  `ROUND_DRAIN_MARGIN` in `pump/mod.rs`). So the six figures were DERIVED from log timestamps
  against a bound that was never written down, by a measurer whose context has since been cleared;
  the gater who recorded them never held the formula either. **Consequence, stated so nobody
  re-derives it:** any new figure computed from these logs would be a DIFFERENT quantity wearing
  the same name, and would enter a series whose value is its bimodality — the one structure a
  mismatched sample corrupts invisibly. This run therefore gets a LABELLED HOLE, not a number.
  What IS reportable for it, with endpoints named rather than a name reused: **ladder span**, node
  start → `TWOHOST role B: ladder complete` = **183.50s** on the batch vs **259.21s** on the
  control. Marker parity established under an IDENTICAL expression on both sides (25 = 25, after a
  first pass that compared two different greps and would have read as parity by coincidence).
- **The fix this makes obvious:** the twohost margin must become an EMITTED token with its two
  events named at the emit site, so the next draw is read rather than reconstructed. A derived
  quantity whose formula lives only in a measurer's context is one context reset away from being
  unfalsifiable — which is exactly what happened here.
- **Fix shape (doyle-approved):** move the done-push block to AFTER leg D, immediately ahead of
  A's stop — one honest barrier made true, no second completion signal (a redundant pair is how
  the next ordering bug lands after the SECOND one). `[int->REQ-REACH-1]` travels with the block.
- **Cascade note, so a future triage does not count it:** in a1, twohost-b's gated-CLI red
  (twohost_cli.rs:189, 910s) was the CASCADE — role-A's gated driver step SKIPPED after the
  ladder death, so B waited for a row never produced. A red that postdates its carrier leg's
  failure carries zero information about its own subject.
- **Ripe when:** proving run at the next two-host window. **Size:** landed-sized already.

### IR-30 — GOLDEN 4b37512 (#145): five attempts, four distinct Windows victims, one sha — the random-victim family measured end to end, and the instruments it left behind
- **Status:** open (carries the authored-unlanded instrument ledger + the family sightings; the
  #145 gate itself CONCLUDED GREEN attempt 5, main ff'd, tested==merged) · **Origin:** doyle,
  night of 2026-08-04, run 30873007187 attempts 1-5.
- **The record:** four reds, four DIFFERENT daemon-spawning tests, Windows leg only, Linux green
  at the same sha every attempt: a1 `daemon::tests::a_tree_teardown_reaches_a_grandchild…`
  (10.225s full deadline, 2nd golden sighting — see [[IR-17]]'s updated sub-observation); a2
  `registry_lifecycle::multichunk_feed_applies_with_exactly_one_snapshot_write` (0-vs-1 at
  mono_ms 77, THIRD sighting of the 2026-07-22 seed; hertz source-named the window: converge
  polls rows landed at merge, assert reads `snapshot_writes` incremented in `write_snapshots`,
  unjoined-thread gap between — predicts under-count only, matching all sightings; quiet-arm
  control 40/40 PASS 0 leaky on a still box); a3 `brain_split::broker_survives_brain_kill…`
  ("supervisor did not respawn the brain", 32.8s window burn; characterization queued); a4
  `resume_no_control_steal_e2e::brain_respawn_keeps_every_session_controller…` — UNDER the
  ruled quiesce-partial window with a measured-clean box, and its panic NAMED A MECHANISM ITS
  OWN DATA CONTRADICTS (pre-authored Failure-A text; gained=[0,15,17] through ONE
  resume_sessions while a steal displaces the SET; session 0 already 4x behind BEFORE the
  window; the rig's child-liveness immunity was a COMMENT nothing measured). Every red
  lane-independent by per-row delta test run fresh each time, never transferred.
- **Instruments authored off it (all thin test/obs lanes off `11169c1`, compiled where stated,
  landing rides the next golden batch):** hertz `test/twohost-serve-window` @`17bbbd8`
  ([[IR-29]]); hertz grandchild SELECTION probe (+166/-8 daemon.rs test-mod: IsProcessInJob at
  selection time — decisive H1/H2 splitter, birth-stamp one-directional, image, same-snapshot
  match COUNT; compiled + deliberate-break positive control); hertz `test/resume-steal-taxonomy`
  (+109/-6: producer counter `Broker::session_output_seq` bracketing the consumer window,
  authenticated child pin at t1 BEFORE teardown — read at assert time would convict the rig of
  its own cleanup — three-way verdict PRECONDITION / measured-Failure-A-with-producer-story /
  absent-producer-is-its-own-story); todlando `obs/resume-attach-intent` @`9c9e6c5`
  — **a DEAD-PATH emit that never fired in production; removed 2026-08-04.** The
  `RESUME_ATTACH_INTENT` breadcrumb sat in `resume_sessions`, which has had NO production
  caller since `03c7109` (2026-07-09): the daemon's respawn path is `run_brain` ->
  `resume_session_cursors`, cursor-only, no attach. Its absence from a field log therefore read
  as "no resume happened" when it meant "uncalled function" — a clean-zero manufactory aimed at
  the very #123 hunt it was built to serve. **Do not hunt for this token.** The instrument now
  lives at the choosers the daemon actually executes, as `ATTACH_INTENT_CHOSEN` with a distinct
  `site=` per chooser (`gap_resume` | `serve_request` | `shell_channel`), carrying the emit
  discipline this one always had — intent bound once, logged and passed from that same binding,
  wildcard-free label. No REQ, per the OBS-rider precedent.
  - _Corrected 2026-08-04 (todlando, in the re-aim rider lane): this row previously described
    the breadcrumb as a live instrument — "intent bound once, logged and passed from the same
    binding, per-call epoch makes set-vs-one readable off the log". That sentence was true
    about the emit's CONSTRUCTION and false about its REACH, and a reader consulting this
    register during the #123 hunt would have been sent after a token that cannot fire.
    Replaced rather than annotated, per the register's correction convention; the superseded
    wording lives in git._
- **Rulings that outlive the night:** a green retires NO intermittent row — grandchild and the
  a4 row are INSTRUMENTED-AWAITING-FIRE, the next occurrence carries its own verdict; the
  environmental family's leg 3 (runner off the fleet host) stays operator-ruled NON-OPTION
  ("stop proposing it") — sightings accumulate in the family record, evidence files itself,
  re-asking is direction; the quiesce-partial window is the sanctioned tool and it WORKED as a
  discriminator (a4's red under quiesce is what forced the rig-soundness finding).
- **Instrument notes minted on the way, cheap to lose:** (a) `gh run view --log` refuses for a
  COMPLETED attempt while any later attempt is live (gates on the RUN) — recover per-job with
  `gh api repos/<o>/<r>/actions/jobs/<id>/logs`; (b) a process-census RESIDENT COUNT is not
  trendable (23→17→15 in 90s of inbox-poll CLI churn) — the CLASSIFICATION is the measurement,
  the count is noise; (c) nextest's per-run LEAKY count is a free process-leakage series already
  in every log (a1=12/a3=10/a4=13, and the multichunk row alone leaks 0/40 — the leaks are the
  row's NEIGHBOURS); (d) diagnostic-token greps derive from the FAILING FRAME's own emit sites,
  never from the register entry discussing the symptom — the gater's own
  `REAP_JOB_TERMINATE_FAIL` zero over a path that emits `SERVICE_TREE_KILL_INCOMPLETE` was VOID,
  and the emit sites document their own vocabulary (tree_kill_incomplete_line's doc names its
  grep-partner); (e) **"golden green ⇒ box free" is FALSE — the ff-push that concludes a golden
  STARTS the automatic post-merge ci run on the same box** (hertz, measured to the second:
  Runner.Worker up 00:13:02 local = run 30886843672 job start 07:13:02Z, ~2min after the ff), so
  the box empties after the POST-MERGE run, not after the golden; the box has a fourth actor no
  agent steps off, woken by every main push (todlando's spelling), and post-golden sequencing
  must gate on `gh run view --json jobs` of the post-merge run before any timing-sensitive work.
- **Ripe when:** instrument lanes land with the next golden batch; family sightings stay until
  the family record closes them. **Size:** the lanes are authored; landing is gate-time only.

### IR-31 — Per-lane worktree targets multiply disk cost by lane count, and pool-claim is structurally blind to the aggregate
- **Status:** open · **Origin:** measured box event, HFENDULEAM 2026-08-04 ~05:00 local, during
  USHER lane concurrency (hertz mechanism statement + both agents' reclaim arithmetic).
- **What/why:** the box hit **0.00 bytes free** mid-build. Failure signatures it manufactured
  look like toolchain or code defects, not disk: `rustc-LLVM ERROR: IO failure on output stream:
  no space on device`, `LNK1318: Unexpected PDB error; LIMIT (12)`, `LNK1108: cannot write file`
  — a red wearing a linker's face (doyle's U2 gate rig ate exactly this; verdict legs already
  green survived, the suite leg had to be re-run). Contributions measured, not inferred: hertz
  ~70 GB across four lane targets (er-seams **42.66 GB** cold `--workspace --tests`,
  selection-probe 25.80, poolowner 1.03, ci-riders 0.91), doyle gate rig 15 GB cold full build.
  Reclaim arithmetic reconciles both and neither alone: 0.00 → 25.47 GB (hertz reaps
  selection-probe) → ~40 GB (doyle reaps rig). **The mechanism (hertz, ruled register-worthy):
  every isolated lane worktree carries its OWN `target/`, so disk cost is multiplied by lane
  count, while pool-claim — built to arbitrate a SHARED pool — sees none of it; a claim on
  main's pool is ceremonial for a lane that never writes there. The tool we use to reason about
  build-cache contention cannot see the resource that actually ran out.** Eight-plus worktrees
  at 25–43 GB per cold full build is the shape of the next occurrence, and it will not announce
  itself through pool-claim. Adjacent discipline failure the same night, named so the entry
  carries it: a gater firing a cold full-workspace rig into a window with two live builder lanes
  is the pre-flight question-1 failure (right-size the run) — targeted legs and warm pools
  first.
- **Remedy shape (sketch, not ruled):** (a) a box-level free-space floor as a RIG STEP at lane
  start — the claim verb is the natural seat (print box free + du of known lane targets at
  `pool-claim`, warn under a floor); kin to the CI-side `REQ-CI-FREE-SPACE-PREFLIGHT` (IR-1's
  companion fix), which covers runners but not agent lane rigs; (b) teardown-on-gate-close is
  already a rig step (the gate-target-disposal rule) — the gap is the AGGREGATE view across
  lanes nobody owns; (c) possibly a `pool-census` xtask verb listing every `.worktrees/*/target`
  with sizes, so the sweep is one command instead of a du walk each agent re-derives.
- **Ripe when:** next CI/rig-touching wave, or the next ENOSPC-signature red — whichever first.
- **Size:** small-medium (claim-verb print + floor; census verb optional).
- **Field addendum (2026-08-04 second event, same day filed — ripeness condition met):** the box
  fell under the CI runner's 32 GiB free-space floor twice more (12:31 main @9d65e65 docs-only,
  12:50 PR core#144 @b110bc8) while todlando's F-lane built. Both Windows unit legs REFUSED with
  `RESOURCE=disk drive=C:\ free_bytes=…` — the `REQ-CI-FREE-SPACE-PREFLIGHT` floor did exactly its
  job: a docs-only main red named the resource instead of wearing a linker's face, and triage was
  one log read instead of an RCA. That is the instrument's first field catch; the CI side of this
  entry is PROVEN. The agent-lane side stays open: local trough measured 8.8 GB free mid-F-lane.
  Reclaim, measured before/after per the teardown rule: doyle reaped gated u1+u2 lane targets
  (POOL-OWNER claims verified own+dead-holder, inbound reparse sweep clean, worktrees kept)
  8.8 → 59.5 GB (+50.7); hertz reaped er-seams target (pool-release first, same discipline)
  59.46 → 101.16 GB (+42.66). Refined mechanism statement (hertz, this event): the disk floor is a
  per-BOX resource our per-POOL instrument is structurally blind to — pool-claim answers a
  question about contention that is not the question the box ran out of.
- **AXIS ADDENDUM (deployah, golden 30928816784 / USHER `fc7fad1`, 2026-08-04) — one floor reading
  is a SNAPSHOT, and the swing is bigger than the margin you fire on.** Two facts this entry's
  remedy (a) has to be built against, both measured rather than reasoned:
  (1) **A pre-fire PROCESS census cannot see the floor at all.** Mine came back clean minutes before
  the run — zero test-path `spt.exe` residents, zero `cargo`/`rustc`/`link`/`cl` — and both Windows
  legs then died in the disk preflight (`drive=C:\ free_bytes=26099576832 floor_bytes=34359738368`,
  `RESOURCE=disk`, exit 1) with a 128-line log carrying ZERO `Compiling`/`PASS`/`FAIL`/`Summary`
  lines: no test signal, and the merge chain never indicted. Both readings were TRUE at the same
  instant — the gater's finished assembly gates were sitting on disk as a DIRECTORY, not running as
  a process. A process-axis probe is blind to a finished build's cost by construction; "re-check
  closer to the fire" would have changed nothing, because the axis was never read.
  (2) **The floor number itself moves by tens of GB inside one run window.** Same box, same drive,
  same run, pulled from its own logs: `16:22:55` test-Windows free=26099576832 REFUSED · `16:23:07`
  n1-gate-Windows free=26096607232 REFUSED · `16:46:54` twohost-a free=71033049088 PASSED and then
  ran 11 minutes to success. ~45 GB returned with NO deliberate teardown, and my own reading at
  17:13:45Z was 61352714240 — ~10 GB BELOW what twohost-a saw 27 minutes earlier. The observed swing
  amplitude EXCEEDED the margin I was about to fire on (25.14 GiB).
  So the rule remedy (a) must encode: measure BOTH axes (process residents AND free space against
  the workflow's own 32 GiB floor), and reclaim until the margin exceeds the observed SWING, not one
  sample. Triage corollary, equally load-bearing: twohost-a runs the same preflight on the same C:
  and PASSED, so a preflight refusal is a threshold event on a moving number — never evidence the
  box cannot run the work, and never by itself a reason to indict a merge chain.
- **THIRD ADDENDUM (deployah, v0.54.0 cut, 2026-08-04) — the sweep answers "0 B reclaimable"
  TRUTHFULLY on a box that cannot run its own work, because the reclaimable population and the
  owned-lane population are DISJOINT.** This hazard blocked a release twice before it was seen.
  Both Windows legs at the shipped sha `86f0d84` refused at the guard before compiling anything —
  `ci.yml` at `free_bytes=28167569408` (26.2 GiB) and `release.yml` at `27690885120` (25.8 GiB),
  both against `floor_bytes=34359738368`. `assemble` skipped in consequence, so no draft release
  existed: the release was hard-blocked on disk, not on code. Neither leg produced a test verdict,
  so both are LABELLED HOLES, not reds (ruled recorded-and-proceed by doyle, the shipped-sha delta
  being 5 files and zero `.rs`).
  **The finding is not the disk.** `xtask pool-sweep --root .worktrees` reported, correctly,
  `total 66.43 GB across 3 pool(s); 0 B reclaimable in this sweep's scope` while the box sat ~6 GB
  under its own floor. The sweep is behaving as designed — it refuses to reap owned, live lanes.
  But the floor reasons over an AGGREGATE the sweep is forbidden to touch, so the designed
  instrument, run at the exact moment of refusal, tells an operator there is nothing to reclaim.
  That is true and useless. The gap between "the sweep is correct" and "the box can run work" is
  the defect; it is not fixed by making the sweep more aggressive.
  **Aggregate measured under `.worktrees`** (real target bytes, junction-classified first): 87.79 GB
  across 9 lanes — 2.7x the entire 32 GiB floor. Registered pools: `w1t2-relink-force` 54.67 GB,
  `obs-resume-attach-intent` 10.18, `w1t2-subnet-status` 1.58. Carrying targets the sweep does not
  count as claimed pools: `hertz-teardown-bound` 11.25, `hertz-liveresolve-diag` 4.83,
  `hertz-psyche-bound` 4.19, `hertz-poolowner` 1.03, `hertz-ci-riders` 0.91.
  **The 54.67 GB single lane is NOT waste** (doyle's classification, ratified at this filing): a
  verb lane at that size is workspace-all-targets e2e cost. No lane is misbehaving. That is what
  makes this structural rather than a cleanup task — every lane is individually justified and the
  sum still exceeds the floor.
  **Reclaim taken, and its boundary.** 82.12 GB, from `spt-core/target` — the release driver's OWN
  main-checkout pool, NOT a lane; free 23.23 → 105.34 GB. Classified before removal per the
  teardown rule: OUTBOUND `Get-Item -Force` reported a real directory rather than a reparse point,
  the INBOUND sweep found zero reparse points aimed at it, `CARGO_TARGET_DIR` was unset (no env
  aliasing), the target SUBTREE only was reaped, both sides measured.
  **Snapshot behavior reproduced inside this measurement window**, corroborating the axis addendum
  above at a smaller amplitude: free read 25.49 GB when first measured and 23.23 GB at the reap
  minutes later — it drifted 2.26 GB DOWNWARD while the operator was deciding what to do about it.
  **Priced side effect, for the next reader:** reaping the main pool took the prebuilt `xtask` with
  it, so the next lane claim rebuilds it first — the lane-check shortcut is cold until then. It cost
  a 2m11s cold rebuild inside this release's own publish step.
  **What this does NOT license:** reaping another agent's owned lane to clear a floor. The sweep's
  refusal is correct and stays. What is missing is an AGGREGATE-AWARE signal — the box knowing its
  lanes sum past the floor *before* a run is dispatched into a guard that will decline it.
- **FOURTH ADDENDUM (doyle, 2026-08-04 late evening — fourth event in one day, and the first
  RULING on this entry).** C: hit **0.44 GB free** during doyle's lane-2 stack gate and todlando's
  re-aim leg set. Signatures manufactured this time, all initially read as something else:
  `os error 112` mid-rlib-archive surfacing as xtask "building spt failed" (exit 101); a
  servicehost force-kill unit red at a sha gated GREEN on the same rig an hour earlier; todlando's
  6-of-7 leg table red with `LNK1318` while only traceable-reqs survived. Both agents' verdicts
  from the window were VOIDED and re-run, not re-read.
  **Discriminator correction (hertz, ratified at this filing):** "treqs alone survives" is a
  POSITIVE TELL, never a clearing test — a full disk reds BEFORE any link step (rustc writing
  rlibs), and treqs can red for its own reasons; signature absence says NOTHING. The only
  falsifier for "this red was the disk" is free space AT RUN TIME, and no log on the box recorded
  it, which made every red from the window unanswerable after the fact.
  **RULING (doyle):** every rig/gate script records `FREE-AT-START` / `FREE-AT-END` in its own
  SUMMARY — one line each, the datum that makes a disk confound answerable post-hoc. Effective
  immediately for hand-authored rigs; hertz builds it into the shared rig harness at next touch
  (assigned, not unasked). This is remedy (a) narrowed to its cheapest load-bearing slice.
  **Reclaim, all FS-delta-measured:** todlando +59.4 GB (obs-resume-attach-intent 8.56 +
  w1t2-relink-force 50.83; 0.44 → 59.83), hertz +1.79 GB (ci-riders + poolowner; sum-of-lengths
  said 1.94, the FS delta governs; 59.09 → 60.87), doyle +38.8 GB (w1t2-subnet-status +
  w1t2-perch-gc + w1t3-node-verb targets; 60.75 → 99.54). Full discipline each: outbound
  classification, inbound reparse sweep (zero), CARGO_TARGET_DIR confirmed unset, subtree-only.
  **Mechanism sharpened by this event:** the standing swing IS the finished-lane population —
  ~110 GB of pools belonging to lanes already GATED AND PUSHED, held for hours because disposal
  fires at GATE close (a rig step) while nothing fires at LANE finish; a pushed lane awaiting
  assembly holds its pool invisibly. Corollary corrections carried: todlando withdrew his
  name-based pool census (10x off; ancestry, not directory names, classifies a pool as closed),
  and hertz-selection-probe's worktree removal REFUSED Permission denied — IR-38's second holder
  class recurring, left for retry rather than forced.

### IR-32 — Docs-drift gate blesses its own gen: an empty-emitting producer agrees with itself perfectly
- **Status:** open · **Origin:** todlando mid-lane stop-and-report, U1 (#144) 2026-08-04. Filed by
  the builder to the gater deliberately — register shape is doyle's to own.
- **What/why:** with a defective binary in the tree (stack-overflowed before reaching its own
  code, empty stdout, non-fatal to the generator), `cargo run -p xtask -- gen` wrote
  `docs-site/src/cli/reference.md` reduced to EIGHT lines — title, do-not-edit banner, empty code
  fence — deleting 3193 lines of published CLI reference. `xtask check` then exited 0, because
  the drift gate compares what the binary emits NOW against what the file holds, and gen had just
  overwritten the file with exactly that emit. **A producer that emits nothing agrees with itself
  perfectly.** Caught only because the builder grepped the regenerated page for his new flag
  names and treated the clean zero as suspicious ([[zero-match-filter-reads-as-absent]] shape);
  the gate itself would never have said a word, and the gutted public reference ships green.
  The mechanism generalizes past the stack overflow that exposed it: ANY failure mode that makes
  `spt --help` produce empty stdout while exiting non-fatally to the generator gets the same
  green. Kin to [[cli-command-docs-drift]] (the gate this defeats) — the gate detects DRIFT
  between binary and page, and is structurally blind to a page just regenerated from a broken
  binary.
- **Remedy shape (builder's sketch, sound; not yet ruled on seat):** a floor in `xtask gen` —
  refuse to write a page whose help block is empty, or whose output is implausibly shorter than
  the file it replaces — so the generator fails loudly instead of producing a stub the gate
  blesses. Gen is the right seat: check runs in CI on a fresh emit too, so a check-side floor
  alone still lets a local gen gut the working tree.
- **Ripe when:** next docs-tooling or xtask-touching lane; SOONER if any lane regenerates the CLI
  reference before the floor exists (a gate on that lane must assert page LENGTH, not gen's exit
  code, until this is built).
- **Size:** small (one refusal + a length-plausibility floor in gen).

### IR-33 — Debug-build clap Command tree runs near the main-thread stack ceiling; the margin is THREE net new args, bisected at `8f291e1` (the original ~6 was measured on U1's tree and is superseded)
- **Status:** open · **Origin:** todlando U1 (#144) 2026-08-04, delta-tested mid-lane.
- **What/why:** eight extra hidden bool args across the five knock seats grew the derive-built
  Command tree past what the debug binary's main thread stack can construct: EVERY invocation
  (`spt --version` included) died with `thread 'main' has overflowed its stack`, exit
  -1073741571, before reaching any of its own code. Delta-tested, not assumed: the same command
  on the main-pool binary from main printed normally. U1 repaired its own trigger (raw-argv
  pre-scan for retired flags before clap parses; lane ends net +2 args) — but the CEILING
  remains: the tree is now close enough that a handful of net new arguments in a debug build
  reproduce a total binary outage — **and the handful is THREE, not six; see the correction
  below before planning against this entry.** **#5's verb-surface rework is the next lane that adds
  arguments and it is much bigger than U1** — this entry exists so #5 is planned knowing the
  margin, not discovering it. Note the failure's face: it presents as a broken binary, and via
  [[IR-32]] it presents as a silently gutted docs page — neither names the stack.
- **CORRECTION — the margin is THREE, bisected at `8f291e1` (todlando's U3 lane
  `build/usher-u3-verb-surface`, own pool, debug profile; filed by deployah 2026-08-04).** The
  original ~6 was measured on U1's tree and was right when written; it is wrong now, and the
  difference is the difference between "plan carefully" and "one flag anywhere takes the binary
  down". N hidden bool args added to the ROOT `Cli` derive, tree restored from git between every
  row, `--version` as the invocation (it carries no work of its own, so whatever it costs IS tree
  construction):

      N=0  exit 0                OK  (control — unmodified lane binary answers `spt 0.53.0`)
      N=1  exit 0                OK
      N=2  exit 0                OK
      N=3  exit -1073741571      STACK-OVERFLOW
      N=4 / N=6 / N=8            STACK-OVERFLOW

  PLACEMENT DIFFERENTIAL, run because #5 adds args at `EndpointCmd` seats rather than at the root
  and a root-only number could have measured the wrong thing: same N injected inside a NESTED
  endpoint seat gives `SEAT N=2` OK, `SEAT N=3` STACK-OVERFLOW. **Identical ceiling — placement does
  not move it**, so three is the number for the shape #5 actually builds. (Rig hazard worth keeping:
  the first seat attempt produced a COMPILE error, E0027, because the dispatch destructures that
  variant exhaustively — taken at face value it would have read as "the seat is fine". Both rows
  above are from the fixed rig.) CONSEQUENCE FOR #5: the ratified MIN spelling ends net +2, i.e. it
  would have shipped with ONE argument of headroom, and any later lane adding a single flag anywhere
  in the tree would then take the whole binary down, `--version` included. Remedy (b) below stops
  being "buys headroom" and becomes the precondition for #5 shipping safely.
- **Remedy shape (sketch, not ruled):** (a) cheapest tooth — a debug-build smoke that runs
  `spt --help` and asserts non-empty stdout + exit 0, which also backstops IR-32's trigger;
  (b) raise the main thread stack for the binary (build config), buying headroom without
  restructuring; (c) structural — box/flatten the derive tree so construction cost stops scaling
  with arg count. (a) is a rider candidate for any lane; (b)/(c) want a real measurement of
  where the ceiling sits before choosing.
- **Ripe when:** #5 (U3 verb surface) PLANNING — this is a planning input, not just a build item;
  the smoke tooth (a) is ripe for the next CI-touching lane regardless.
- **Size:** small for (a); medium for (b)/(c) with the measurement.

### IR-34 — nextest LEAK flag on the zombie-fixture row is dominant-but-intermittent (19/20 measured); any gate reading it as a signal flaps
- **Status:** open · **Origin:** doyle #142 F-lane gate 2026-08-04 (first sighting in a gate run);
  characterized same day by hertz from retained logs, ownership falsified by todlando.
- **What/why:** `broker::tests::windows_session_is_zombie_sees_a_handle_held_corpse_as_dead` (the
  ADR-0041 zombie-detection fixture, whose subject is a corpse process kept alive-looking by a HELD
  HANDLE) flags nextest LEAK on most draws but not all — measured 19 LEAK of 20 on a tree carrying
  no #142 content (`test/grandchild-selection-probe` @`695398b`, based on `4b37512`: lib arm 10/10
  LEAK across 822-test runs; full-suite arm 9/10 across 1044-test runs). The single clean PASS
  kills "intrinsic therefore always": leakiness is DOMINANT BUT INTERMITTENT. Consequences:
  (1) a gate or reader that treats this row's LEAK as a defect signal flaps at roughly 1 in 20;
  (2) one non-leaky run is NEVER evidence that something fixed it (kin:
  [[intermittent-green-is-zero-information]]); (3) the verdict method that closed it is the
  template — two independent legs, neither load-bearing alone: the suspect lane's diff grepped
  zero hits on the row's subject (zombie / handle_held / OpenProcess / corpse), AND a 20-draw
  baseline on a tree predating the suspect change. This entry exists so the next gater who sees
  the flag reads one register line instead of running that RCA cold.
- **Remedy shape (sketch, not ruled):** annotate the row as expected-leak in nextest config
  (per-test `leak-timeout` override or documented allowlist) so the flag stops presenting as
  signal; alternatively a fixture-side close of the held handle on the clean path if ADR-0041's
  arrangement permits — hertz's call, test/CI lane.
- **Ripe when:** hertz's next nextest-config-touching lane; blocks nothing today.
- **Size:** small.

### IR-35 — hfenduleam Windows test-leg victim rate: one teardown/liveness family, measured at ~half of all executions
- **Status:** open · **Origin:** USHER #150 golden triage 2026-08-04 (doyle; all rates deployah-measured).
- **The numbers (15 golden runs 2026-08-02..08-04 + the USHER cycle, deduped on `(run, box,
  started_at)` — GitHub partial reruns COPY untouched jobs into the new attempt with their ORIGINAL
  `started_at` and carried conclusion, so a carried failure is the SAME observation; dedupe before
  any rate math):** Windows (hfenduleam) 13 red / 22 executions = 59% leg-red, splitting 10/22 = 45%
  test-victim + 3/22 infra (IR-31 class). Linux (kitsubito) 2/20 = 10%, 1 victim. Every victim-red
  leg had EXACTLY ONE victim (8/8 legs). Clean-cycle arithmetic ≈ 0.55 × 0.95 ≈ half — a golden
  cycle at these rates is a coin flip. knock145 (the previous green golden) took FIVE Windows
  executions to its green; USHER took four. **A both-green draw is a draw, not evidence the family
  closed.**
- **The family:** 11 victim legs, 8 distinct tests, ONE cluster — process teardown / kill-reach /
  respawn-controller survival, concentrated on hfenduleam. Per-victim dispositions live in the USHER
  #150 triage ledger (releases#150 record). The only surviving cross-sha EXACT repeat after source
  verification is `resume_no_control_steal_e2e` (:488, byte-identical @2138b16 + @4b37512) — and
  both its reds PREDATE the post-#145 producer-side instrumentation, so what repeats is an AMBIGUOUS
  observation (stolen controller vs starved child) twice; only a post-#145 red discriminates. The
  `a_tree_teardown` pair was a FALSE repeat — two mechanisms wearing one assert string (see
  FLAKE-LEDGER row; the counting lesson: assert-body identity must include the polled predicate's
  SUBJECT).
- **Off-CI discriminator (todlando 2026-08-04):** 0/400 (A 0/200 + B 0/200, interleaved,
  `--test-threads 1`, filters positive-controlled, by-path sweeps every iteration) on the warm
  d449ce5 pool UNDER live-fleet load (CPU 46-59%, 13-18 spt processes). Points the family at the
  RUNNER ENVIRONMENT rather than a product race any Windows box expresses — not exoneration (absence
  of reproduction is not proof of absence), but where the next hour goes. A-alone N=500 on the
  instrumented tree ran post-close (result on releases#150).
- **Remedy lanes already moving:** hertz 4-item test-rework package dispatched 2026-08-04 (psyche
  bound split, live_resolve rig instrumentation, green-capture probe profile, resident_service/
  teardown hardening); prod `wait_bounded` tree-kill-on-timeout filed to board as EVAL. This entry
  is the RATE's home — it leaves the register when the operator has ruled on the rate (accept vs
  hold-for-hertz) AND the hertz package has landed with a re-measured victim rate.
- **Ripe when:** operator brief at USHER close (immediate); re-measure after the hertz package lands.
- **Size:** visibility entry + the re-measure; fix cost carried by the hertz package.

### IR-36 — `output_bounded` is a 51-copy clone estate with a named-but-empty shared home; hoist-and-delete wants its own lane
- **Status:** BUILT 2026-08-18 on KEYSTONE #182 hygiene lane. · **Origin:** hertz measurement
  2026-08-04 (at `7465ac1`, counted before listing, listed untruncated), during the USHER package
  item-2 sweep ruling; doyle ruled file-don't-sweep.
- **What/why:** 51 `fn output_bounded` definitions under `crates/spt/tests` — 50 per-file copies
  plus `tests/common/mod.rs:53`, which is NOT pub, has ZERO callers outside its own module, and
  whose doc comment states outright that the estate's copies predate the shared module. The lossy
  timeout arm (panic loses everything captured, names nothing) is in 38 sites across 36 files.
  The copies have already drifted: two variants of the deadline expect-string exist in the estate
  (`"captured spt call must complete within its deadline"` in common/mod.rs:60 vs the
  no-deadline-clause spelling in per-file copies), so a grep keyed to either string undercounts.
  The right fix is a hoist-and-delete refactor (make the common copy pub, adopt it estate-wide,
  carry the item-2 diagnostic arm once) — a 50-file test-estate touch that must NOT ride golden
  cadence as a rider or hide inside a scoped diagnostics item (ADR-0050 thin-lane rationale;
  hertz's USHER item 2 stays scoped to live_resolve, which keeps ONE improved copy as the model).
- **Built evidence:** one public owning helper replaced all 50 file-local
  clones. It owns and tree-kills a timed-out child, waits for reap, drains both
  pipes, and includes the deadline plus partial stdout/stderr in the panic.
  A dedicated regression observes both diagnostics and proves the timed-out PID
  is gone.
- **Ripe when:** a dedicated post-batch test-hygiene lane (natural companion to IR-13's
  mutation-proof loop and IR-23's home-collision fix — same hertz wave shape).
- **Size:** medium by file count, small by risk (mechanical hoist; no assertion changes).
- **Hygiene-lane scope addition (todlando, at #6, 2026-08-04):** a CI prebuild of
  `cargo build -p mock-adapter --bin mock-shell` (derived from tracked `.github/` content) turns
  SIX binedge baseline rows green at once instead of growing the burn-down list one consumer at a
  time — belongs to this lane, not to any thin verb lane. [[IR-38]] (live-binary e2e reap sweep)
  rides the same lane.
- **Built evidence:** all 51 file-local definitions are gone; 158 call sites use the one public
  `common::output_bounded`. The shared helper owns the child, kills and reaps its process tree on
  timeout, drains both pipes, and reports partial stdout/stderr. `bounded_output_e2e` mutation-pins
  both timeout ownership and retained diagnostics.
- **Declared validation extra:** the lane's docs-link gate exposed that `llms_target_exists` treated
  generated targets as absent before generation. Its small refactor + unit keeps source and
  generated link targets distinct; it is deliberate validation fallout, not part of IR-36.

### IR-37 — `traceable-reqs check` is coverage-only: tag PLACEMENT is unenforced on any multi-tagged REQ, and the more tags a REQ accumulates the less any one is checked
- **Status:** UPSTREAM CLOSED THROUGH RELEASE 2026-08-19 — traceable-reqs PR #19 squash-merged
  @`82d8b14` after doyle's targeted re-review PASS (defect A item-tables + 9 red-first regression
  cells + sighted control; defect B enclosing-item fallback DECLARED in SPEC with cost; 267/267,
  blind repros exit 0); v0.4.0 cut (version bump @`c9a8eb5`, tag pushed, 3 platform assets,
  release notes authored; refs sent to hertz). REMAINING and dispatched to hertz post-#182-landing:
  spt-core consumes v0.4.0 — CI version bump + placement config (enforce-on, module_banner=accept);
  banner cleanup rides a later lane. Entry closes when the consume lane lands. · **Origin:** hertz
  measurement 2026-08-04 (filed by doyle; the measurement is
  hertz's), found when his own pre/post `check` comparison around a tag-adjacency fix came back
  EXIT=0 on BOTH sides — the green he had cited beside the fix validated nothing.
- **What/why:** measured on `test/teardown-bound-shape`: a `[unit->REQ-RESIDENT-SERVICE]` tag
  SEPARATED from its `#[test]` fn (a const landed between tag block and fn — an AGENTS.md rule-1
  violation) scores EXIT=0 identically to the fixed adjacency. Positive control explains the
  mechanism: that REQ carries 57 unit tags across five files, so 56 other sites decide the
  coverage verdict and no single tag's placement can turn it red. Consequence: AGENTS.md rule 1
  ("tag on or immediately above the real evidence") is agent discipline only — drifted tags decay
  silently while coverage stays green, and heavily-tagged REQs decay fastest. This is the
  [[IR-32]] family shape (a gate that cannot see its own blind spot) applied to the traceability
  gate itself.
- **LIMIT, with falsifier (hertz's, verbatim in kind):** behaviour on a SINGLE-tagged REQ was NOT
  established. Separating the only tag of a single-tagged REQ at the stage under test would settle
  whether placement is unenforced outright (still exit 0 ⇒ the checker never reads adjacency) or
  merely un-enforceable at scale (exit 1 ⇒ the checker sees absence, and only multi-tag redundancy
  masks drift). Run that discriminator BEFORE designing any remedy — the two outcomes want
  different fixes (a placement rule in the checker vs a per-tag-nearest-item heuristic).
- **Transfer:** doyle accepted this lane's independent EXIT-0 remeasurement and confirmed it
  matches todlando's settled 2026-08-04 discriminator. Upstream remedy: enforce adjacency for
  each individual `doc`/`impl`/`unit`/`int` tag, with separated-tag negatives covering
  single-tag, interposed-const, file-top-banner, and one-of-many displacement. This repo consumes
  the released checker only.
- **Close condition:** doyle reviews and merges the upstream lane, cuts the `traceable-reqs`
  release, and spt-core consumes that release. Issue/branch/PR/commit refs append here when the
  upstream lane lands.
- **Size:** measurement complete; remedy scope is upstream-owned and dispatched.
- **DISCRIMINATOR RUN 2026-08-04 (todlando, authorized specimen on `build/w1t2-shell-relink-force`;
  filed by doyle — the measurement is todlando's): the answer is the EXIT-0 ARM.** Single-tagged
  int stage (`REQ-SHELL-RELINK-FORCE`), tag SEPARATED from its fn (parked at file top):
  `traceable-reqs check` EXIT 0 and the REQ still reports `+int`. Tag reverted to adjacency:
  EXIT 0, `+int`. Both exits redirect-read, no pipes. **Placement is UNENFORCED OUTRIGHT — the
  checker never reads adjacency; multi-tag redundancy was never the mechanism, it only widened the
  blind spot.** Remedy option (a) is the live one: a real placement rule in the checker (upstream
  experimplate), not a per-tag-nearest-item heuristic. Measurement caveat, carried at the
  measurer's own insistence: an intermediate probe printed `TAG_ADJACENT_AFTER_REVERT=0` and that
  was the PROBE wrong (`grep -B 1` above the fn lands on `#[test]`, not the tag above it), not a
  failed revert — the revert was verified directly (exactly one tag, above `#[test]`). The
  discriminator result is the record; the intermediate zero is not.
- **REMEDY DISPATCHED 2026-08-18 (doyle):** upstream lane opened in BigscreenVR/traceable-reqs
  (the actual upstream remote; the checkout at `~/Documents/projects/traceable-reqs` tracks it —
  "experimplate" above named the tool's origin project, not the repo) — checker-side placement
  rule, default-on in `check`, positive controls all four stages + separated-tag negatives
  including the single-tag displaced discriminator shape and the interposed-const shape.
  Requested by hertz as the IR-37 prerequisite of his test-hygiene family lane (KEYSTONE #182).
  spt-core consumes the upstream RELEASE only — no local shadow checker. Issue/branch/PR refs
  land here when the lane reports.

### IR-38 — CLASS: an e2e that ends with a live daemon-spawning binary wedges the NEXT build in its pool, and the diagnostic names the wrong lane
- **Status:** BUILT 2026-08-18 on KEYSTONE #182 hygiene lane. · **Origin:** todlando 2026-08-04,
  `shell_relink_force_e2e` on the #6 lane (filed by doyle; the finding and the fix are todlando's).
- **What/why:** the e2e ends with a LIVE binary by design; on Windows a live process holds its
  image open, so the NEXT cargo invocation in that target dir dies with
  `failed to remove file target\debug\spt.exe: Access is denied (os error 5)` — surfacing as
  exit 101 with NO test summary, because the failure names the BUILD and the FILE, never the test
  that leaked the holder. Actual holders measured: two lane-local `spt daemon run --detached` +
  `daemon brain` processes spawned under the test's temp SPT_HOME by a child CLI's
  `ensure_running`, still alive from an earlier run. **Reap BY IMAGE PATH** — the fleet's other
  eleven spt processes run from the installed binary; a name-based sweep takes your own daemon.
  The live shell-spawn population is eleven e2e files. Nine already end through product teardown
  or their authenticated home-scoped reap. The two intentionally live-at-success siblings now
  kill and prove the final resident gone AFTER all identity/wake assertions:
  `shell_relink_force_e2e` and `shell_stale_online_e2e`.
- **Kin:** [[IR-7]] (phase-A daemon+brain pair leak), the exe-lock reap-by-path rule, and the
  wrong-lane-diagnostic family ([[IR-21]]'s location-not-build class).
- **SECOND HOLDER CLASS, measured 2026-08-04 (doyle, own gate rig):** an **ORPHANED CONHOST with an
  inherited CWD** inside the test tree blocks `git worktree remove` with Permission-denied on a
  DIRECTORY handle — invisible to any exe-path scan (conhost runs from System32; the parent that
  spawned it was already dead). Found via `.github/ci/find-cwd-holders.ps1` (IR-11's instrument);
  killed by pid after parent-dead classification; 34.78 GB reclaimed. Mechanism chain: the
  windowless spawn path masks DETACHED_PROCESS off, the child owns a console, conhost inherits the
  child's cwd and can OUTLIVE it. Boundary (hertz's, correct): this find and his field-4 companion
  are the SAME MECHANISM FAMILY measured on DIFFERENT populations by different instruments — his
  capture measured a ppid-match COUNT (floor 2, companion identity UNMEASURED, his stated limit),
  and this conhost is the first FIELD identity evidence in the family, on THIS population. Do not
  cite his row as identity data it never carried. (His box-wide accretion count 77→80 is a third
  population again.) Teardown rule
  addendum: a disposal leg must GATE on the removal's exit code (a rig that logs "reclaimed 0.01
  GB" from a failed remove fabricates its own success) and sweep CWD HOLDERS, not only image
  locks — two populations, either alone is a clean zero on the other.
- **Ripe when:** the test-hygiene lane (IR-13/IR-23/IR-36 family) — sweep every e2e that ends
  with a live spawning binary, apply the same reap-after-assertions shape.
- **Size:** small per test; population unknown until swept.
- **Built evidence:** focused runs of both live-at-success tests pass back-to-back in one target
  pool, followed by a build invocation from that same pool; the second test's final kill is polled
  to proven process death rather than treated as fire-and-forget.

### IR-39 — CLASS: the missing-fixture-bin defect has TWO failure faces, and one of them impersonates a lifecycle defect of the subject under test
- **Status:** BUILT 2026-08-18 on KEYSTONE #182 hygiene lane. · **Origin:** hertz + todlando
  2026-08-04, same evening from opposite sides (filed by doyle; hertz relayed todlando's
  suggestion without endorsing scope).
- **What/why:** a rig whose narrowed build (`-p <pkg> --test <one>` / `-p spt --bin spt`) skips a
  cross-package fixture bin fails in one of two ways depending on the consuming site.
  Face 1 (cheap): `translate_proof_fixture` panics "must be built: <path>" — names the artifact
  and the recipe, costs one build. Face 2 (expensive): `sibling_bin("mock-session")` resolves to a
  nonexistent exe, nothing spawns, and the test reds at a PRECONDITION assert whose message
  ("warm bring-up spawned=false", left None / right Some("offline")) reads as a lifecycle defect
  of the subject — measured cost 186s + a triage inside a lent box window (hertz's F1 red arm,
  attempt 1). The build gap itself is the known cross-package row of the fixture-bin population;
  what this entry carries is the DIAGNOSTIC asymmetry.
- **Ask:** a fixture-bin existence precondition that NAMES the missing exe and its build recipe
  wherever a rig consumes one — `sibling_bin` (and the shared
  `crates/spt-term/tests/support/fixture_bin.rs` resolver already carries the recipe string as an
  argument, the shape to copy) refuses with the artifact + recipe instead of letting the consuming
  test fail downstream in its own vocabulary.
- **Kin:** [[IR-21]] (wrong-lane-diagnostic family: the failure names the wrong actor), the
  cargo-builds-package-bins-for-integration-tests rule (cross-package row is the only unguaranteed
  build), golden.yml's dedicated fixture prebuild step.
- **Ripe when:** the test-hygiene lane (IR-13/IR-23/IR-36/IR-38 family) — same sweep population;
  or standalone small if that lane keeps slipping.
- **Size:** small — one shared helper + a sweep of local `sibling_bin` resolver copies (24 known).
- **Addendum (doyle, 2026-08-04 late): third and fourth sightings inside 24 h.** Face 1 twice
  more the same evening — todlando's teach-lane leg set (his own translate_proof_fixture, caught
  by his prebuild), and doyle's cold gate rig for the same lane (first-pass spt-bins red at
  136/611, correctly withheld from the lane verdict as rig evidence, fixed by `cargo build -p spt
  --bins` + rerun 611/611). Four sightings in one day across three agents and both faces: the
  ripeness condition ("the test-hygiene lane keeps slipping") is under measurable pressure — the
  prebuild remedy is being re-derived per-rig per-agent, which is the recurring cost the shared
  precondition helper exists to delete.
- **Built evidence:** all 29 file-local `sibling_bin` resolvers now route through one shared
  precondition. A missing fixture names its path and the package-correct `cargo build -p <owner>
  --bin <name>` recovery command before product assertions or timeout paths run; a focused
  negative test pins that diagnostic.

### IR-40 — CLASS: NO envelope-side ordering signal on a resume brief tracks content recency, and a CI informant's per-job line is a claim about that STEP, not the run
- **Status:** open · **Origin:** three sightings in one night, 2026-08-04/05 — hertz, doyle,
  todlando (filed by doyle; the artifacts are hertz's and todlando's, the CI arm is doyle's with
  deployah's framing). Two arms of ONE family: a stale record that carries a confident ordering
  signal, and a status claim that arrives before the thing it describes.

- **ARM A — resume briefs. What/why:** a session can receive TWO start-of-session briefs, and the
  one routed LATER can carry the STALER project-context. Measured on hertz's pair (same node id):
  routed_at_ms `1785897407373` > `1785896642480`, spill-filename epoch `1785897764588` >
  `1785896658159`, and vector sequence `1242529/1242531` > `1241542/1241544` — **all three
  monotonic orderings rank the staler-content brief as newer.** todlando's pair is the NEGATIVE
  arm: his later brief (`…1785897765196`, routed `…577476`, vector `1242756/7`) genuinely WAS the
  fresher one, so the same heuristic happened to pick right.
- **The finding is UNCORRELATED, not inverted** — and that is the whole entry. An inverted signal
  is usable once you know to flip it; an uncorrelated one reads as reliable until the pair where it
  costs you. Neither agent could have known which pair he held without checking content against
  measurement. **Do not name routed_at_ms specifically** — naming the stamp invites the repair that
  fails ("the timestamp is unreliable, use the sequence number"). The vector is the most dangerous
  of the three: a stamp looks like a clock and invites suspicion, a sequence number looks like an
  ordering and does not.
- **Visible failure mode if a brief is obeyed as instructed** (hertz's, measured): re-open
  `releases#123` — work already ruled closed — and under-report his own claimed-pool count by one,
  the row sitting under a live gate. doyle's own sighting the same night: a brief describing main
  @`92b2482` with a teach lane "mid-flight" hours after every lane had gated.
- **PRE-REFUSED DISCRIMINATOR, recorded so it is not rediscovered as a finding:** file SIZE ranks
  the fresher brief correctly on BOTH pairs (hertz 16308 > 12649; todlando 14138 > 10418) — 2 of 2,
  the only signal surviving both samples. **Do not use it.** Size measures RICHNESS, not recency;
  the two correlated only because todlando's thin brief was a correction fragment, and hertz's
  STALER brief was a full 12649-byte dump. Two full dumps hours apart go flat or inverted. Handed
  over pre-refused by its own finder.
- **ARM B — CI informants (doyle, 2026-08-05, run 30971976024).** `CI-KITSUBITO` sent
  "CI SUCCESS … sha 0a25b77" while the run was measurably `status=in_progress` with
  `conclusion=""` — the EMPTY STRING, not `success`. The informant fires from a per-job `notify`
  step, and at send time `notify` itself had no conclusion. All eight substantive jobs were green,
  so the line was **right in the end but EARLY**. deployah's framing, kept verbatim in kind: *an
  informant line that is right in the end but early is worse than one that is simply wrong, because
  a wrong signal gets distrusted while an early-but-correct one gets promoted to the conclusion.*
  todlando's sentence is the one to keep at the point of use: **a per-job notify step's message is
  a claim about that STEP, not about the RUN.**
- **ARM B MECHANISM — hertz, source read 2026-08-05, zero box cost. It is NOT a race; it fires
  early on EVERY run, and the entry above understated it.** `notify` is a job INSIDE the run
  (`golden.yml:1138`), and a job cannot observe its own run's conclusion — so `conclusion: ""` is
  not a timing artifact, it is **the only value that can exist when the informant speaks. The
  informant has never once made a statement about a run conclusion.** What `ci-notify.sh` actually
  asserts is narrower and worth naming exactly: a verdict computed PURELY from six `needs` results
  (changes, traceability, test, n1-gate, twohost-a, twohost-b — the two matrices collapse to one
  result each) being neither failure nor cancelled. Nothing more.
  **The blind spot is exactly ONE job: `notify` itself, the 9th — it cannot appear in its own
  `needs` list. And that one job has a WITNESSED red path**, documented in the script's own header
  (2026-07-27, twice on kitsubito): verdict computed, then `spt send` HUNG ~5 minutes until the
  job's own 5-minute budget cancelled it, and that cancellation reddened the run. So *"informant
  says SUCCESS, run concludes FAILURE"* is a real observed sequence, not a hypothetical. The 30s
  `SEND_BOUND_SECS` added since bounds the hang but does not close the class — it makes the
  informant's own failure FAST rather than impossible.
  ⇒ The informant's SUCCESS covers 8 of 9 jobs and is blind to the 9th, and the 9th is the only one
  with an observed failure mode that reds the run. This is why the ff predicate is the run's
  conclusion field and never the informant line.
- **RECOVERY, and it is the same for both arms: never rank the claims — re-derive from the source
  object.** For briefs, hertz's four falsifiers are the method (main sha, claimed-pool count, lane
  tip, issue state) and they work because each is a content claim checkable in ONE command against
  a source NEITHER brief controls. For CI, read the RUN's own `conclusion` field, and prefer
  independent reads of the run object (`gh run watch --exit-status` polls the run, so its exit is a
  second measurement rather than an echo of `notify`). On the batch this was caught in, three
  sources agreed with the informant excluded from all three before main was fast-forwarded.
- **Standing rule this entry exists to enforce:** a resume brief's RULINGS are durable; its STATE
  is a HYPOTHESIS to re-derive before acting.
- **Kin:** [[IR-21]] and [[IR-39]] (the diagnostic names the wrong actor / the wrong subject),
  [[IR-32]] (a gate that cannot see its own blind spot), [[IR-34]] (an intermittent marker read as
  a signal).
- **Ripe when:** the next commune/psyche-tier touch for arm A (the routing layer is spt-core's own
  `spool`/psyche ingest, so a remedy is in-repo, not upstream); for arm B, the next `golden.yml`
  informant touch — the cheap fix is that the notify step names its own scope in the message body
  (STEP vs RUN) rather than emitting a bare "CI SUCCESS".
- **Size:** arm B small (message text + scope word in one workflow step). Arm A unsized —
  establishing whether ANY durable content-recency signal can ride the envelope is the measurement,
  and until it runs the answer is "re-derive, do not rank."

### IR-41 — A QUEUED main run is superseded without a record; `cancel-in-progress: false` protects only STARTED runs
- **Status:** open — mechanism CONFIRMED (doyle ruling 2026-08-05); the runbook sentence it
  falsified is already corrected (`docs/RELEASE-RUNBOOK.md` step 3, same commit as this entry)
  · **Origin:** deployah, v0.55.0 publish night, filed by him explicitly as UNPROVEN with the
  manual-cancel alternative not ruled out; confirmed by doyle on the run objects.
- **What/why:** `ci.yml` sets `cancel-in-progress: false` on main and the runbook read that as
  "main's concurrency policy never cancels this run; that preserves the record." FALSE for the
  queued phase: a concurrency group holds at most one running + one pending run, and a newer push
  REPLACES the pending run regardless of the flag — the flag governs only whether a STARTED run
  is cancelled. A superseded queued run leaves NO record: zero jobs, conclusion `cancelled`,
  nothing measured about its sha.
- **Evidence (all read off the run objects, not the informant):** run 30973909279 at `327f1f8`
  (push, main) — created 04:01:08Z, **zero jobs ever started**, cancelled 04:03:49Z, ONE SECOND
  after `e8805f7`'s run 30974041521 entered the group (created 04:03:48Z), while `0a25b77`'s run
  30973770039 was still in flight (done 04:07:59Z). One-running-one-pending, newer push landed,
  pending run died. The 1s coupling to an unrelated push is the discriminator against deployah's
  own manual-cancel alternative — a human cancel co-timed to the second with a push it had no
  view of is not a credible mechanism; supersession fires on exactly that trigger by design.
- **The hole this leaves:** under a rapid push sequence, an intermediate sha on MAIN can have no
  thin-CI verdict AT ALL, and the absence reads as nothing rather than as supersession. Anyone
  auditing "did sha X pass main CI" gets a hole where the runbook promised a record. Absence of
  a run is now a labelled state, not evidence about the sha.
- **Ripe when:** next ci.yml/runbook touch. Candidate remedies to weigh THEN, not now: accept and
  document (done — the runbook now carries the mechanism), or give main's group per-sha keys
  (`group: ci-${{ github.sha }}`) so runs never share a group — at the cost of concurrent main
  runs competing for the boxes, which is exactly what the group exists to prevent. The trade is
  real; do not fold it into a drive-by.
- **Size:** docs half DONE in this commit; ci.yml half small but load-bearing — needs its own
  lane if taken.
- **Kin:** [[IR-40]] (a record that is right-in-the-end-but-early vs a record that never comes to
  exist — both read as clean states unless labelled), [[cancelled-measurement-leaves-labelled-hole]].

### IR-42 — `pool-claim` writes a record; only the BUILD enforces — and AGENTS.md invited the misread
- **Status:** BUILT 2026-08-19 — the AGENTS.md corrective line lands in the SAME COMMIT as this
  entry (the register batch commit), which is the entry's stated exit condition. The optional
  code-side nicety (claim verb prints the incumbent record it overwrites, information only)
  remains unclaimed — a candidate rider on the IR-26 remedy lane, which already touches
  `pool_claim`'s read-before-write path, NOT a reason to keep this entry open ·
  **Origin:** doyle + todlando independently at source, 2026-08-19, during the #193 gate.
- **Mechanism (source-verified):** `xtask pool-claim` (crates/xtask/src/main.rs:2232-2294 at the
  read sha) parses its args, reads the lane's own git identity, builds a `PoolOwner` and calls
  `spt_poolguard::write_owner` UNCONDITIONALLY (:2280) — no `read_owner`, no verdict, no
  comparison against the incumbent claim anywhere in the function. Claiming is last-writer-wins
  by construction: two lanes can each "hold" a pool in sequence with only the last write
  surviving, and the displaced lane learns nothing at displacement time. All four enforcement
  arms — `Refuse` (SPT_POOL_FOREIGN), `Takeover` (loud), `Unproven`, `HatchOpen` — live in
  crates/spt-store/build.rs:31-113 and speak only at the next BUILD. (Same object [[IR-26]]'s
  CLAIM-OBSERVABILITY measurement saw from the displaced side; this entry is the CLAIMING side
  plus the doc that invited the misread.)
- **Measured consequence:** 2026-08-19, doyle's gate claim over todlando's unlanded
  redeem-echo-diag claim — no refusal, no notice (three sequential overwrites that session, one
  over a lane with an unlanded commit). todlando then predicted from AGENTS.md's refusal
  sentence that the gate's claim "will be REFUSED" — a false warning to a gater mid-run. The
  refusal sentence sat directly under the "Claim a pool at lane start" instruction, inviting the
  build-time semantics to be read onto the claim verb.
- **Corrective landed:** one AGENTS.md sentence — the claim WRITES a record and never refuses;
  predict refusals from builds, never from `pool-claim`.

### IR-43 — knock answer carrier: `NoReply` landed at 60.08s against a stated 30s carrier deadline
- **Status:** open, OBSERVATION — mechanism unmeasured, filed exactly as wide as the datum ·
  **Origin:** KNOCK-169 RCA layer 1 (todlando, read-only), golden 32191618042 @`93c1130`,
  twohost-b, 2026-08-18.
- **What/why:** B's `request_answer` yielded `NoReply` at 22:50:08.3605Z — 60.08s after A's
  death — while the carrier deadline is stated at 30s. Recorded in the RCA as a side observation
  with no claim; the cell's red itself was a CASCADE of A's earlier death (answerop.rs:73-81
  produces the courtesy BEFORE returning, so the reply was never sent — that half is settled and
  is NOT this entry). What this entry carries: a transport/liveness bound that reported at 2x its
  stated value. Candidate shapes, neither asserted: two sequential 30s legs each honoring its own
  bound (a composed wait wearing one bound's name), or a bound armed after a wait it does not
  cover (the [[IR-30]] nethost.rs permit-wait shape: `sem.acquire_owned()` outside the timeout).
- **Ask:** derive where the 30s is armed and what path yields a 60s `NoReply` — one code read
  from the emit site, before any instrument.
- **Ripe when:** next knock-carrier touch, or the next `NoReply`-shaped red — whichever first.
- **Size:** small (code read; a token only if the read forks).

### IR-44 — perch-sentinel comment overclaims: serialization is not preservation, and the comment teaches the false step
- **Status:** open · **Origin:** #188 RCA lock read (todlando flagged, doyle confirmed the lock
  read), 2026-08-18; the collapse-of-branches argument is on #188 and in the RCA record.
- **What/why:** `lock_perch_sentinel` (info.rs:778 at `b20e770`, the diag sha — re-derive the
  line on main before editing) takes a cross-process fs2 lock on a stable per-perch `.info.lock`
  sentinel, and its comment claims it serializes "ALL info.json writers so a whole-record write
  and a locked RMW can never lose each other's update (`REQ-HAZARD-INFO-RMW-LOST-UPDATE`)" —
  TRUE for interleaving, FALSE for preservation. `write_info_unlocked` is module-private with
  exactly three callers, all taking the sentinel first (bypass refuted by visibility), but every
  caller of the public `write_info` composes its record BEFORE the lock is taken (:842) — **the
  lock makes that write atomic; it cannot make it preserving.** Only `mutate_info` /
  `establish_locked` read under the hold and can preserve. Measured field case: census write #1
  (`twohost.rs:736`) clobbered `controlled` while correctly locked. A reader trusting the
  comment infers lost-update immunity the funnel does not provide — the #188 hunt burned a
  branch on exactly that inference before the collapse argument killed it.
- **Ask:** respell the comment — serializes writers; a composed `write_info` is atomic, never
  preserving; preservation requires `mutate_info`/`establish_locked` — and weigh whether
  `REQ-HAZARD-INFO-RMW-LOST-UPDATE`'s doc wording carries the same overclaim.
- **Ripe when:** next spt-store/info.rs touch. **Size:** tiny (comment + possibly one REQ doc
  line).

### IR-45 — twohost rig home premise, SETTLED: pump paths resolve under the per-run TEMP root (rig is disk-hermetic; the "live fleet roster" reading is retired)
- **Status:** RETIRED 2026-08-19 — ask answered by the one code read at main @`78a9a16`; residue
  re-homed (see Answer). · **Origin:** #188/#189 RCA readouts (todlando, doyle's 2a-2d), runs
  32097943571 + 32108557362 + 32111921251, 2026-08-18.
- **What/why, two measured halves in tension:** (1) B's pump dialed a THIRD node 155-159 times
  per run at kitsubito's OWN tailscale IP (`100.98.197.12`), different hex + different ephemeral
  port each run — a short-lived local endpoint re-minting identity between runs, present in both
  the golden and the diag run. Read at the time as: the rig uses the CANONICAL home
  (`canonical_pump_paths` → `perch::spt_home()`), so B's roster is kitsubito's live fleet
  roster — environmental coupling of a CI rig to host fleet state. (2) The ER perch measured in
  the chain runs resolves under a per-run TEMP root
  (`…/_temp/spt-test-tmp-<run>…/owlery/engine-room`), NOT the canonical home. Both measurements
  stand; they answer different artifacts (pump roster vs perches), and nothing on record settles
  which home the PUMP paths actually resolve. #189's third-node evidence and its
  cache-leg-never-heals reading rest on that premise (the tension is filed on the #188 record as
  an open question against #189).
- **Answer (2026-08-19, code read at `78a9a16`, unambiguous — no run token needed):** both role
  tests set `SPT_HOME` to a fresh per-run `TempDir` as their FIRST act (`twohost.rs:813-814` role
  B, `:1501-1502` role A), before any store touch. `canonical_pump_paths` resolves every path via
  `perch::spt_home()` (`twohost.rs:396-414`), which honors the `SPT_HOME` override first
  (`perch.rs:34-41`); the pump's roster source `presence::registry_snapshot_dir()` →
  `perch::identity_dir()` sits under the same root (`presence.rs:83-85`). The rig is
  DISK-HERMETIC per run; both prior measurements reconcile (ER perch observed under the temp
  root, pump paths resolve there too). The misleader was the helper's name + doc comment ("real
  homes, not test roots"), which describe production path LAYOUT, not the resolved ROOT.
- **Residue disposition:** (1) #189's third-node evidence was corrected on the board (comment
  5337519936) — the roster row arrived at RUNTIME into a temp-rooted store, ingress mechanism
  OPEN on #189, no longer explained-environmental. (2) KEYSTONE #182's hygiene lane repaired the
  `twohost.rs:394` comment to say exactly what the helper does: production path layout rooted in
  the process's fresh per-run `SPT_HOME`.

### IR-46 — the Windows disk-floor preflight asserts an INSTANT; workspace free space moves tens of GB inside the hour, so a green floor is not a claim about run headroom
- **Status:** open — mechanism CONFIRMED by a measurement series on hfenduleam; doyle ruled it
  register-shaped 2026-08-18 · **Origin:** deployah, NAMEPLATE #181 golden-head intake, on the
  run that the floor red-floored · **Filed as IR-46 after an id collision:** this row was authored
  as IR-42 and held unpushed under the v0.56.0 tag freeze, during which the release-close sweep
  (`78a9a16`) issued IR-42..45 to other findings. Parallel issuance, not authored disagreement —
  and noted here so a later reader chasing "IR-42" in this row's history does not hunt for a lost
  version of it.
- **What/why:** `golden.yml`'s Windows preflight computes `$freeBytes` once at job start and
  hard-fails under a `32GB` floor. As a fast-fail that is correct and it did its job — it refused
  before compiling rather than dying deep in a link step. The defect is what a PASS is then read
  to mean. The reading is a point sample of a quantity that moves by tens of GB unattended, so a
  green preflight licenses a run whose headroom was never measured. Nothing downstream re-checks.
- **Evidence — five readings of `C:` free on one box inside roughly one hour, 2026-08-18:**
  - `03:01:54Z` — `free_bytes=2666205184` (2.48 GiB) against `floor_bytes=34359738368`, the
    preflight's OWN reading; run `32093894524`, both `n1-gate` and `test` on Windows/hfenduleam
    failed at this same step, before any compilation.
  - `~03:2xZ` — ~2.50 GB, independent `Get-PSDrive` read, corroborating the preflight.
  - pre-reap — **33.70 GB**, immediately before deployah deleted anything.
  - post-reap-1 — 76.42 GB (42.72 GB reclaimed: `claude_skill_owl/target`, `golden-w1/target`).
  - pre-reap-2 — **73.81 GB**, i.e. 2.61 GB consumed unattended between the two reaps; then
    post-reap-2 129.70 GB (55.89 GB reclaimed: `.worktrees/nameplate-w1/target`).
- **LABELLED HOLE — do not let this harden:** roughly **31 GB was released between the preflight
  failure and the pre-reap reading by something OTHER than the reclaim**. Released-by-unknown.
  The runner cleaning its workspace after the job died is *plausible* and is NOT the recorded
  cause — it was never measured, and no one looked while it was happening. Recorded as a hole on
  purpose (doyle's explicit ask at filing): a plausible cause written down as the cause would make
  this row read as explained when the actual mechanism of the 31 GB is unknown.
- **The hazard this leaves:** a run can clear the floor at preflight and starve mid-build, and
  mid-build disk starvation does not present as disk — it surfaces as a link failure, a truncated
  artifact, or a rustc ICE, i.e. as a defect in the tree under test. The failure mode inverts the
  gate's purpose: the preflight's whole point is to keep box conditions from being read as lane
  reds, and a passing preflight actively argues the opposite. Note the polarity — this row is
  about the PASS, not the FAIL. The observed red was honest.
- **Ripe when:** next `golden.yml` touch. Remedies to weigh THEN, not as a drive-by: re-assert the
  floor at job END (turns a starvation into a labelled disk verdict instead of a fake lane red);
  or raise the floor to cover a full build's peak rather than its entry condition; or sample free
  space across the run and emit the minimum as a token. All three cost run time; none is obvious.
- **Size:** small in `golden.yml`, load-bearing in what a green run is taken to prove.
- **Kin:** [[IR-41]] and [[IR-40]] (a state that reads as a clean verdict unless it is labelled —
  here a PASS that is read as headroom it never measured), [[measure-the-box-before-the-instrument]],
  [[cancelled-measurement-leaves-labelled-hole]].

### IR-47 — ci-notify treats a missing co-author trailer as a quiet info line; silent attribution loss reads as "there was none"
- **Status:** open — patch AUTHORED and stashed unlanded since 2026-07-29
  (`.worktrees/_patches/ci-notify-missing-trailer.patch` + its `test-ci-notify.sh` harness in the
  sibling `.untracked` dir); surfaced by doyle's 2026-08-19 idle-queue verification of that stash ·
  **Origin:** the AGENTS.md trailer mandate's own hazard family (the space-spelling trailer is
  structurally invisible to git's tokenizer, so confident zeros already read as attribution loss
  once); this row is the NOTIFY-side twin.
- **What/why:** `.github/ci/ci-notify.sh` parses the head commit for the line-anchored
  `Co-authored by: <agent>` trailer to add a second notification recipient. When the trailer is
  absent or unparseable, the current script emits one stdout info line ("doyle only") and moves
  on — correct as a non-fatal outcome, wrong as a SILENT one: a lane that lost its attribution
  (typo'd trailer, squash that dropped the body, hyphenated spelling) notifies doyle alone on
  every run and nothing anywhere says an agent stopped being told about their own lane's verdicts.
  The stashed patch keeps absence non-fatal but promotes it to a `::warning` workflow annotation
  (visible on the run summary), and splits the legitimate quiet case (co-author IS doyle) from the
  loss case so the two stop sharing one message.
- **Evidence:** patch verified 2026-08-19 to still NOT be on main — the warning string is absent
  from `.github/ci/ci-notify.sh` at the current tree. Patch content re-read at verification; it
  applies against the script's post-`SEND_BOUND_SECS` shape.
- **Ripe when:** next `golden.yml`/ci-notify touch — natural co-rider with [[IR-46]]'s remedy
  window (same file family, same "weigh then, not drive-by" rule). The stashed test harness rides
  with it.
- **Size:** ~10 lines in one shell script plus its test.
- **Kin:** [[IR-40]] (an unlabelled state read as a clean verdict — here a quiet info line read as
  "no co-author existed"), the AGENTS.md `%(trailers:)` tokenizer mandate (the sibling silent-zero
  on the AUDIT side).

### IR-48 — the nextest.toml parity cell tolerates CRLF only by parser accident; a newline-sensitive arm added later reds fresh Windows checkouts alone
- **Status:** open — filed 2026-08-19 (doyle route, hertz verdict: REGISTER as preventive
  hardening, not current defect) · **Origin:** sibling sweep after the W3 gate CRLF finding —
  parent mechanism: `include_str!` embeds working-tree bytes verbatim, `core.autocrlf=true`
  smudges fresh checkouts CRLF, and the author's never-re-smudged tree stays LF — builder green,
  every fresh rig/golden checkout red.
- **What/why:** `crates/xtask/src/main.rs` `the_checked_in_config_is_in_parity` include_str!s
  `.config/nextest.toml` and feeds `phase_a_overrides_missing_from_ci_windows`. TODAY this is
  causally CRLF-tolerant — the parser is `str::lines()`+trim (`override_filters`, main.rs:434-440),
  and the 52/52 fresh-rig green at the hygiene gate (2026-08-19, this box) is therefore causal, not
  luck. The hazard is the MISSING CONTROL: nothing pins the tolerance, so a future
  newline-sensitive arm in the parity predicate reds only on fresh Windows checkouts — the
  builder-green/rig-red inversion, deterministic but reading as flake.
- **Remedy (hertz's, platform-independent):** feed the parity predicate an LF fixture AND the same
  fixture converted to CRLF, assert identical missing-sets (including one unmirrored negative), so
  a newline-sensitive arm reds on EVERY host rather than only fresh Windows checkouts.
- **Ripe when:** next xtask parity-cell touch; explicitly OUTSIDE the family-B lane (hertz
  scoping). **Size:** one fixture-pair cell.
- **Kin:** the W3 gate finding this swept out of (skeleton cells, fixed at the fixture edge in
  lane commit e99633f); [[IR-40]] (an untested tolerance read as a guarantee).

### CI-RIDER LANE STATE — LANDED 2026-08-04 (post-v0.53.0 merge queue); history below kept for its mechanisms
- **RESOLVED:** the lane rebased clean onto the post-tag queue and landed ff-only as
  `2b33a47`/`2a7b016`/`d63f4ce`/`19d7f79` (content byte-identical to `1275e47..bf8c4a2` by lane-diff
  blob hash). IR-1 and IR-4 are BUILT AND LANDED; IR-9's measurement half is landed with its trigger
  now armed — the toolchain print reports runner-account versions on the NEXT GOLDEN RUN, which is
  when doyle's 2026-08-03 ruling (decide from runner-account versions, never the interactive prior)
  becomes executable. The two golden-only steps (link probe both boxes, toolchain print both legs)
  remain unexercised until that run — the landing does not change that caveat.
- **As recorded pre-landing (2026-08-04, doyle):** `ci/locksmith-riders` was at `bf8c4a2` with FOUR
  commits not contained in `origin/main` and not present in LOCKSMITH's golden head `b7b00c3`. The
  worktree `.worktrees/hertz-ci-riders` was clean, so the work existed and was simply unlanded:
  `1275e47` (xtask: assert a load-bearing patch pin is still in force, both ways it lapses) ·
  `49d4805` (docs/ci: a lock-touching lane reads edges, not just the package set) ·
  `8ed006b` (ci/golden: print the toolchain that judged the run, both legs) ·
  `bf8c4a2` (ci/golden: measure the link before rendezvous, read the floor after the checkout that
  clears it — the floor read AFTER its own reclaim is [[ci-runner-has-no-warm-target]]'s shape).
- **Why it is recorded rather than quietly re-dispatched:** [[IR-1]]/[[IR-4]]/[[IR-9]] were carried
  as DISPATCHED, which is now false in BOTH directions — the work is further along than dispatched,
  and it also did not ship. The cause was the same agent outage that cost `#123` its build, and the
  operator's standing rule from that drop applies here too: an outage must not be able to remove
  work from a batch silently. Board requests got a drop comment; register entries get this.
- **MAPPING CONFIRMED BY ITS BUILDER 2026-08-04, by REQ id rather than by recollection** — hertz
  noted first that its own context had been cleared between building the lane and answering, and
  declined to testify from memory about the original instrument. Everything here is either measured
  that day or read off the commits:
  `1275e47` → `REQ-CI-LOAD-BEARING-PATCH-PIN` ([[IR-4]] part 1, the mechanical pin guard in xtask) ·
  `49d4805` → `REQ-LOCK-TOUCHING-LANE-PROCEDURE` ([[IR-4]] part 2, procedure in docs/GOLDEN-CI.md) ·
  `8ed006b` → `REQ-CI-TOOLCHAIN-VERSION-PRINT` ([[IR-4]] part 3, carrying [[IR-9]]) ·
  `bf8c4a2` → `REQ-CI-LINK-HEALTH-PROBE` ([[IR-1]], the network axis).
- **CORRECTION TO THE GATER'S EARLIER WORDING, from the builder:** `8ed006b` does NOT discharge
  [[IR-9]] — it discharges IR-9's MEASUREMENT half only. IR-9's other half (align the two boxes, or
  declare one authoritative clippy leg) is held by doyle's own 2026-08-03 ruling: decide once the
  step reports RUNNER-ACCOUNT versions, never on the interactive-account prior. IR-9 therefore stays
  OPEN with its trigger now REACHABLE, which is a different state from "waiting".
- **Gate evidence, instrument NAMED:** `cargo nextest run -p xtask` in `.worktrees/hertz-ci-riders`
  on hfenduleam, box quiet with no runs in flight — 34 run, 34 passed, 0 skipped. Nextest rather
  than bare `cargo test`, per [[IR-25]]. ⚠ **What that does NOT cover, stated by the builder rather
  than discovered later:** the two golden steps the lane ADDS (link probe on both boxes, toolchain
  print on both legs) cannot be exercised outside a golden run. The probe's three arms
  (five-sample success, no-reply, absent-CLI) were exercised on the real boxes at commit time, and
  that is reported as a CLAIM RECORDED AT COMMIT TIME, not as a re-verification — the kitsubito arm
  could not be re-checked from here. This is why the lane wants a golden rather than a quiet ff.
- **[[IR-5]]/[[IR-6]] CONDITIONAL RIDERS: CONDITION NOT MET — the answer is NO and it closes the
  question for this lane.** Measured file set `6d291e0..bf8c4a2`: `.gitattributes`,
  `.github/bench/link-probe.ps1`, `.github/bench/link-probe.sh`, `.github/workflows/golden.yml`,
  `crates/xtask/src/main.rs`, `crates/xtask/tests/bench_row_parity.rs`, `docs/GOLDEN-CI.md`,
  `traceable-reqs.toml`. Nothing under `.github/ci/`, where the nextest-summary readers and
  count-reporting scripts live. IR-5: no consumer of nextest output touched. IR-6: no EXISTING gate
  output changed — though the NEW output was voluntarily written to IR-6's rule (the LINK line
  prints `samples=N/M` and the individual rtts, membership beside the count, in both shells). That
  SHRINKS IR-6's future scope by two lines; it does not discharge it.
- **Three mechanisms lifted from this lane's REQ titles, worth more than the lane:**
  (a) **An instrument must not be able to RED the run, and each shell breaks that differently** —
  bash: `| head` under `pipefail` SIGPIPEs the producer to 141, so a print step reds (trim with
  parameter expansion instead); pwsh under GitHub's wrapper: a MISSING COMMAND is a terminating
  error, so rustup's absence must be TESTED with `Get-Command`, never caught. One requirement, two
  constructions, and neither is portable reasoning about the other.
  (b) **A negative-control mutation must be COUNTED before it is trusted** — the absent-CLI arm was
  first "measured" against a tree where the mutation had not taken, silently re-measuring the
  unmutated arm and reading as a pass.
  (c) **Jitter, not the median, is the signal on a link that mostly works** — the motivating red's
  link ran 10..83ms and 3..75ms while IDLE, so a median-only probe reads healthy straight through
  the failure. Both med and max rows ride the ledger.
- **Disposition:** composes onto the next golden batch as a thin lane, NOT slipped onto main. It
  changes the golden pipeline itself, so it wants a golden rather than a quiet ff — and while
  v0.53.0 is untagged, any push to `main` moves the runbook's bare `git tag` off the tested sha.
