# Infra register — CI / build-pipeline debt

Operator-ruled 2026-08-02: infrastructure and CI-pipeline work items live HERE, not on the
`spt-bs-releases` board. The board carries product surface the operator triages; this register
carries what the gater triages. **Mandate: doyle sweeps this file at every milestone intake and
every release close, and composes ripe entries into waves/milestone riders.** An entry leaves
this file only by being built (link the lane) or being retired with a stated reason.

Entry format: status · origin · what/why · trigger condition (what makes it ripe) · size guess.

Last sweep: 2026-09-06, v0.67.1 RELEASE CLOSE + WEBSERVE (#272) INTAKE (one delta pass; the 08-30
full pass stands) — shipped counter 103, tag == main == tested `04e32c8c`, golden 34017906638 att4
9/9 on a QUIET box after three reds (:473 attach_link structural test defect; wtlock :147 under
builder load; resident_service_e2e :664 ledgered leak row, 3rd occurrence) — the quiet-box arm is
IR-76's origin. Enumeration by a read-only subagent into `REGISTER-SWEEP-0671-DRAFT.md` (75
entries, 10 trigger-text candidates; IR-31 has NO `###` heading — its body sits under IR-30's,
flagged, not normalized). Rulings: **IR-46 trigger FIRED and was MISSED** — `golden.yml` was
touched by `04e32c8c` itself (2-line HEAVY reclass on the respin) and no remedy was weighed; a
respin edit under a pre-registered hard stop is not a remedy window, so the entry gains that datum
and is COMPOSED forward. **Composed: the floor trio IR-46 + IR-59 (log-the-floor half) + IR-73
(8 literal-first sites) as ONE hertz workflow rider on WEBSERVE** (`ci.yml`/`golden.yml`/
`release.yml`, thin PR off the golden path, before the WEBSERVE golden head so that head runs
under it); **IR-59's log-the-floor line ALSO rides IR-76's three-arm pre-flight on W0's driver
template** (todlando authors; free space recorded beside the floor check). **IR-66 residual (two
first-chunk-needle members) = hertz WEBSERVE test rider** (two one-line edits). **IR-12 residual =
W3 drift-gate rider** (hertz already holds W3's drift-gate riders). IR-67 satisfied at this intake
by construction (`--no-fail-fast` in the W0 driver template; todlando's 2026-09-06 correction).
IR-34 unchanged, lane-linked to hertz's `test/attach-relink-barrier` @`b976cc09` (thin PR queued
behind the release runs). NOT fired: IR-3 (a new stream family is neither endpoint lifecycle nor
daemon supervision), IR-48, IR-52 (DOCS-NITS item 2 is a parent-option placement filing, not a
deep-help divergence — adjacent, noted), IR-64 (att4 did not die at preflight; no operator ruling).
Board mints this arc: #277 (io-events seq reset), #278 (adapter update never prunes strings/),
#279 (ER_HOSTED_PROBE on every bind), #280 (MSG_IN never published on the relay edge). Design
rulings on the grill branch: ADR-0056 Am.1 (router precedence), ADR-0057 Am.1 (served root =
`adapters/<adapter>/web/`), ADR-0058 Am.1 (#17 auto-serve, reference-served). No entries retired.

Prior sweep: 2026-08-30, NOW-SIGNAL (#23) / v0.67.0 RELEASE CLOSE — shipped counter 102, tag ==
main-at-tag == tested `da71b785`, 9/9 golden att2 (att1 sole red = FLAKE-LEDGER :664
teardown-leak row, ruled ledgered-class at triage — signature predates the diff — rerun
non-vacuity proven by cell re-execution; ledger row +1 with new data, stays OPEN; hertz
IR-34/35 cluster strengthened). Same-day second sweep, so most composition state is the
morning's; deltas only: IR-59 gained the SEVENTH face (parallel-lane fill rate: 236.6 GB /
3.5 h across four wave pools drove the box to 0.01 GB mid-close; per-wave pool reaping rule).
IR-12 gained the stale-binary kin note (xtask reference-drift arm reads the ON-DISK spt binary;
structural remedy = the now-mandatory workspace-bins prebuild leg, adopted this arc after it
also caught W2's spacerun red). IR-63's conhost litter REPRODUCED on cue: 14 fresh CWD pins on
the four party trees at the same `crates/spt-daemon` relative path, killed by pid via the PEB
probe, all four worktrees then removed FIRST TRY — the pin-check-first party rule held
(+58.2 GB; earlier emergency reap of the same wave's finished-lane targets +186.6 GB).
Test-craft banked to memory, not entries: `cargo test -- <bare> --exact` runs NOTHING and exits
0 (todlando's structural fix: names carry the filter token); Win32_Process CommandLine filters
CANNOT answer a CWD-pin question (todlando's false-positive self-catch). Board mints this arc:
#244/#246/#247/#254 (BACKLOG). #16/#17 discharged drops (detach + EVAL). IR-73's ci.yml/
release.yml floor-site reorders did NOT ride this arc's workflow commits — still owed, next
touch. No entries retired.

Prior sweep: 2026-08-30, SEMAPHORE (#242) / v0.66.0 RELEASE CLOSE — shipped counter 101, tag ==
main == tested `d931dd63`, 9/9 at one sha across two attempts; full-register pass executed from
`REGISTER-SWEEP-242-DRAFT.md` (72 entries enumerated pre-golden, load-bearing claims re-measured
in-tree). Rulings: IR-69 close-out re-census RAN and FOUND the entry's own predicted truncation
blind spot LIVE (338 cli.rs + 2 main.rs shipping sites never censused; addendum on the entry,
stays OPEN; residual conversion + generator fix + seam enforcement = releases#243, scoping
reconciled on the issue; common-word matcher RULED a separate hertz test-side rider). IR-17
three corrections by replacement (08-03 kitsubito "empty stderr" = METER ARTIFACT — panel
capture postdates the specimen by 16 days; kitsubito family member CLOSED under the RCA-242
exe-hash-on-ready-path mechanism, behaviour-change request minted as releases#244; Windows
specimen + hfenduleam sub-observation stay the live population). Daemon-leak cluster
IR-7/17/20/34/35/63 LANE-LINKED to hertz's queued fixup lane (briefs in his spool cite the entry
ids). IR-57 first-real-assembly use appended (ASM-241: 6 picks, 5 MATCH / 1 LOSS reconciled
byte-for-byte to the deliberate resolution; its earlier "figure correction rides next batch"
note was STALE — the correction landed inline 2026-08-29, retired here). IR-46 gained the #242
golden additions (per-job floor instants; the `/runs/<id>/jobs` run_attempt trap with
`started_at` discriminator + `/attempts/N/jobs` authority; raw-bytes rerun preflight;
non-vacuous rerun check). IR-59 gained the SIXTH face (r4 LNK1318 at the job's internal
low-water; within-job floor re-read rides the next workflow commit; box ~51 GB non-recovering
consumption → audit slot). FILED: IR-73 (ci.yml/release.yml literal-first floor sites off the
golden path, 8 sites), IR-74 (kitsubito 21,643 /tmp endpoint homes, hertz's measurement), IR-75
(kitsubito treqs 127 vacuous leg — install-or-drop; IR-37 deliberately NOT reopened, different
surface). Composed: IR-14/26/27/49 as ONE between-milestone pool/worktree audit slot (window
open now); IR-11 + IR-33(a) as next-intake small riders; IR-59's log-the-floor half + IR-57's
scripted audit stand as the next intake's tooling pair. IR-54 evidence noted: the #242 cut's
unshaped-head intake miss (self-caught by deployah, shape-on-top `e4f52047`, 8th consecutive
shape-inside-candidate) is another construction-not-discipline datum. No entries retired.

Prior sweep: 2026-08-27, IO-PARSER (#22) INTAKE (wave map: `IO-PARSER-22-JIT.md` + #22 comment).
Rulings: IR-66 COMPOSED as the parallel hertz rider (the two attach-cell later-needle treatments,
attach.rs:561/:672 — the "rides next intake" note comes due here). IR-52 conditional rider
CARRIED FORWARD on this milestone's docs lane, same terms as WAX-SEAL (depth fix lands iff the
lane touches the docs-site CLI reference generator, else stays open). IR-67 COMPOSED as JIT gate
discipline (no wave battery leans on a `-p spt --bins` leg as integration coverage); the entry
stays open until its named construction fix. IR-55 stays armed with hertz; IR-6 LOCKSMITH
composition still unreported, stays; IR-2 trigger unevaluated this sweep (teardown just cleared
the unlanded-lane field — re-check next sweep). No entries retired; none filed.

Prior sweep: 2026-08-23, WAX-SEAL (#21) INTAKE (wave map: `WAX-SEAL-21-JIT.md` + #21 comment
5390092115). Rulings: IR-52 COMPOSED as a CONDITIONAL W4 rider on the wax-seal docs lane —
the milestone mints new nested `spt seal` CLI verbs, exactly the shallow-render class the entry
names; the depth fix lands iff that lane touches the docs-site CLI reference generator, else the
entry stays open here. IR-2 trigger NOT met (todlando's PR set parked unlanded for the next
golden chain). IR-55 stays ARMED-FOR-CAPTURE with hertz. IR-6 conditional composition from
LOCKSMITH still unreported — stays open. IR-56/57/58/59/60/61 are discipline/craft entries or
await their named triggers; IR-62 is hertz-class test/rig work, unscheduled. No entries retired;
none filed.

Prior sweep: 2026-08-19, CONCIERGE (#183) INTAKE (wave map: `CONCIERGE-183-JIT.md` + #183
comment 5347768797) — rows read from the LOCAL register stack (`4799031→ecd640c→8d4c224`;
origin/main lacks IR-49/50/51 until the chain lands — absence ≠ not-filed). Rulings: IR-50
COMPOSED — dispatched to hertz same day (sink-path-helper-FIRST sequencing per the entry;
`engine_room_bringup_e2e.rs` excluded, its panel sites ride the gated #199 lane), with the
engineroom.rs:145 misnomer as a rider on the same lane. IR-47 = candidate co-rider IF this
chain touches golden.yml/ci-notify, else holds. IR-48 holds (next xtask parity-cell touch).
IR-49 holds (next poolguard touch or first landed-lane takeover request). IR-51 holds —
gated on #199 attribution; a net-off fix must NOT precede attribution. IR-2 trigger NOT met
(queued unlanded lanes exist). No entries retired; none filed — the intake window's product
finding (rc terminal-Exit omission, hertz RCA: proven invariant violation, load excluded
92/92, H2 candidate unconfirmed) routed to the BOARD as releases#201, the correct venue for
product surface.

Prior sweep: 2026-08-18, KEYSTONE (#182) INTAKE — the prior sweep's four composition rulings
EXECUTED (wave map: `KEYSTONE-182-JIT.md`, base main @`8248bc3`): (1) IR-9 pin lane DISPATCHED to
hertz (W0 item 1; interim kitsubito-clippy-authoritative dies when it lands). (2) IR-21 remedy
(2) + IR-39 vs `f24e732` DISPOSED: the branch is a BEACHHEAD (golden.yml prebuild step = the
prebuild-made-a-rule arm, ONE fail-fast site, one ledger row) — hertz rebases it off its
abandoned parent `b483699` (the id-collision draft; content landed renumbered as IR-46 @`8248bc3`)
and lands it in W0; the class remainder (IR-39's SHARED precondition helper + 24-site sibling_bin
sweep, the 33 cross-package build-edge expressions) folds into the test-hygiene lane. (3)
Test-hygiene family DECIDED-ACTIVATED as a dedicated hertz lane inside the KEYSTONE window,
sequenced after W0 (members: IR-13/23/36/37/38 + IR-21 clarity half + IR-39 helper +
twohost.rs:394 doc comment). (4) First-execution-cells discipline restated in the JIT's
golden-head step. Board members #84/#85 ride hertz's W0; #166/#57/#185 are todlando's W1. IR-46
id-collision RULED this intake (renumber-and-land, parallel issuance not authored disagreement;
deployah landed it @`8248bc3`). No new entries filed; no entries retired.

Prior sweep: 2026-08-19, NAMEPLATE (#181) / v0.56.0 RELEASE CLOSE — shipped c91 @`60d74ea`
(tag == golden-tested sha, run 32209922535; one respin, both first-golden reds ruled rig defects).
Filed IR-42 (pool-claim writes / build enforces — BUILT in this same commit, AGENTS.md line),
IR-43 (knock NoReply past its 30s carrier bound), IR-44 (perch-sentinel comment overclaims
preservation), IR-45 (twohost rig home premise; SETTLED + RETIRED
2026-08-19 — pump paths resolve under the per-run temp root, see the entry; #189 corrected on the
board, comment 5337519936). Two cycle findings ruled RUNBOOK-homed and
landed in this commit rather than as entries: the first-execution-cells intake question
(RELEASE-RUNBOOK golden-head intake — name the never-executed cells before the run) and
deployah's sweep-vs-cascade mechanism (RELEASE-RUNBOOK board step — under golden CI the cascade
is DRIVEN via `state <mref> acceptance`, never swept). ⚠ LABELLED HOLE — CLOSED UNDERIVABLE
(2026-08-19): the close commune's batch list named "alchemy create-races"; its content did not
survive the author's context reset and was not recoverable from #181, the JIT records, or memory.
Deployah answered the query: he ran ZERO create ops at the cut (could not have witnessed a create
race), a fresh probe over the milestone window shows no duplicate mints (the 4-issues-in-2s batch
mint is batching, not duplication), and the only surviving trace is doyle's own pre-reset message
naming the item as already-known — a pointer, not a sighting. Item DROPPED; the hole stands as
the record. Deployah's sweep-vs-cascade ship-path trap (#181 comment 5337402639, runbook-homed
above) is explicitly NOT this item's content — do not fold it in. Composition: IR-9's pin lane
(`rust-toolchain.toml` @ 1.96.0) goes to hertz AT KEYSTONE #182 INTAKE per the 2026-08-05
ruling; IR-21 remedy (2) + IR-39's precondition helper compose with hertz's standing
fixture-prebuild-hardening branch `f24e732` — disposition at the same intake; the test-hygiene
family (IR-13/23/36/37/38 + IR-21's clarity half) stays the dedicated post-batch lane candidate,
decision at intake; IR-29's proving run + IR-30's instrument lanes ride the next golden batch.
IR-2's trigger explicitly NOT met (queued unlanded lanes exist: IR-29/IR-30 instruments,
four-arm refusal eprintln, f24e732). IR-14 hygiene movement: `.worktrees/nameplate-asm-2bd36f1`
reaped this sweep (+66.98 GB by FS delta 120.53→187.51; claim `gate-w7-courtesy` base `fd3dc5a`
in main = finished lane; zero inbound reparse points) — stale-lane audit itself still open.

Prior sweep: 2026-08-05, LOCKSMITH tranche-2 (#141) CLOSE — golden run 30971976024 green on all 9
jobs, main ff'd to `0a25b77`, v0.55.0. IR-40 filed (stale-resume-brief + early-informant class).
IR-9's decision trigger FIRED 2026-08-05: doyle read the runner-account versions off this run's
`test` legs and RULED — pin in-repo via `rust-toolchain.toml` @ 1.96.0, kitsubito's clippy leg
authoritative in the interim; pin lane to hertz at next intake (see the entry). IR-41 filed the
same night (queued main run superseded without a record; runbook step 3 corrected in the same
commit). IR-1/IR-4's golden-only steps (link probe both boxes, toolchain
print both legs) had their FIRST EXERCISE here, discharging the "unexercised until a golden run"
caveat at the CI-RIDER LANE STATE foot section. IR-35's re-measure rode the batch (`c65b838`).
Owlery-noun thin lane SCOPED and dispatched to hertz for the next batch (class A only, two sites;
class B on-disk rename is an explicit non-goal — see IR-40's kin discipline for why the boundary is
written into the brief rather than left to judgement).

---

## OPEN

### IR-1 — Quiet predicate needs a network axis (tailscale RTT probe)
- **Status:** open, LANE EXISTS UNLANDED — see [[CI-RIDER LANE STATE]] at the foot of this file;
  carried by `bf8c4a2` → `REQ-CI-LINK-HEALTH-PROBE`, mapping CONFIRMED by builder hertz 2026-08-03
  (by content: tailscale ping ×5, med/max RTT rows into the bench ledger, three arms
  success/NO-REPLY/UNAVAILABLE, always exit 0 — instrument, not gate). NOTE `bf8c4a2` is NOT
  single-purpose: it also corrects the free-space preflight floor read
  (`REQ-CI-FREE-SPACE-PREFLIGHT`, the ci-runner-has-no-warm-target shape) — no 1:1 commit→IR map
  for this commit · **Origin:** golden/bench-wiring red triage 2026-08-02 (ex releases#126)
- **What/why:** the shared-runner quiet predicate (zero non-terminal runs + no local
  cargo/rustc/nextest by parent chain) is process-shaped; both axes passed on a box whose only
  link was degrading (321s for a 1s checkout, bidirectional 10s QUIC dial timeouts). A tailscale
  RTT probe to the peer box before two-host rendezvous, carried in the bench ledger, would have
  called run 30771155390's red in seconds. Evidence: the arm-1 count table (PUMP_PEER_FAIL
  a 0→3→0, b 8→22→8 across green/red/rerun).
- **Permanent, not stopgap:** operator-confirmed 2026-08-02 that kitsubito cannot be provided
  ethernet — wifi-only indefinitely, so the link cannot be hardened and the predicate must see
  link health.
- **Ripe when:** next CI-touching wave, or the next network-shaped golden red — whichever first.
- **Size:** small (one probe step + ledger row + predicate doc).
- **Composed:** LOCKSMITH (#132) CI-rider cluster, hertz thin lane — 2026-08-03. GREENLIT with
  #132 and **DISPATCHED to hertz 2026-08-03** (direct brief; the register is the spec, there is no
  board issue). Leaves the register only when the lane lands or the entry is retired.
- **FIRST FIELD USE, and it DISCRIMINATED (2026-08-04, golden 30873007187 attempt 1):** the probe
  (landed via the rider lane, riding `4b37512`) read 4–7ms RTT healthy on both twohost legs
  minutes before both legs redded — REFUTING the degraded-link read for that red and steering
  triage to the real mechanism (the [[IR-29]] serve-window race) instead of a link chase. The
  instrument's first catch was a correct NEGATIVE — exactly the call run 30771155390 needed and
  could not make.

### IR-2 — Settle the warm-runner CARGO_INCREMENTAL delta
- **Status:** RETIRED 2026-08-28, measurement run, verdict NO FLIP — incremental stays ON.
  hertz paired goldens at main fed965f8: ON run 33167693879 GREEN; OFF run 33172022108 RED
  (Windows input_ack_deadlock SetupFailed at a 3.0s IPC read deadline — the box-load family,
  discarded with the leg). Disk: OFF target 7.49 GB vs ON 11.40 GB = 3.91 GB / 34.3% saved
  (smaller than the entry's 5.65 GB cold figure). Timing NOT causal-quality: sequential runs in
  one persistent workspace let OFF inherit ON's cache (order contamination dominates the
  apparent OFF speedups); Windows build-spt-bin +3.8%, Linux -5.1%. Retire reasoning: the
  DECISION is answerable now — OFF produced no green golden, the timing question needs 4+
  counterbalanced fresh-cache windows to answer cleanly, and the disk motivation has weakened
  (box at ~148 GB free under the teardown discipline; the LNK1318 pressure era predates it).
  Runs/artifacts preserved on the record. Reopen only if runner disk pressure returns as a
  recurring floor-class red.
- **Was:** open · **Origin:** #103/#108 bench-wiring lane 2026-08-02 (ex releases#127)
- **What/why:** the #103 measurement (−29.8% wall, −5.65 GB/target, n=3) is COLD-build only.
  Golden's runner `_work` target persists warm, where incremental is exactly what keeps it cheap;
  CARGO_INCREMENTAL=0 was applied only to the genuine cold build (n1-gate pinned old-broker
  cache) + local rig recipes (docs/GOLDEN-CI.md). Open question: does incremental still pay on
  the warm runner, weighed against 5.65 GB/target on a box with LNK1318 free-space history?
- **Method (hertz):** one full golden each way on a quiet box, outside a milestone, compared
  per-step from the bench ledger.
- **Ripe when:** a quiet between-milestone window with no queued lanes (the measurement burns
  two golden windows).
- **Size:** medium (two proving runs + verdict + possible leg flips).

### IR-3 — Daemon-level guard: broker-net wakeup rate bounded across endpoint churn
- **Status:** open · **Origin:** releases#125 remediation, todlando REQ call 1 (ex releases#128)
- **What/why:** the swarm-discovery GC-spin burned two cores for two weeks visible only in a
  process table — no suite assertion sees the class. Wanted: a daemon-level assertion that
  broker-net workers stay quiescent across repeated endpoint create/destroy churn.
- **Design constraint (pre-ruled):** assert on WAKEUP RATE / voluntary ctxt-switch delta over
  the churn window, NOT %CPU — CPU thresholds flake under CI load; the defect signature
  (~100 Hz per orphan loop) is load-independent. Mint the REQ at activation.
- **Near-product:** this is runtime-defect visibility, the most product-adjacent entry here —
  a candidate rider on any daemon-lifecycle milestone.
- **Ripe when:** the next milestone touching spt-net endpoint lifecycle or daemon supervision.
- **Size:** medium (churn harness + counter plumbing + flake-safe assertion).

### IR-4 — Lock-pin guard + lock-procedure rule + toolchain print (three riders, one lane)
- **Status:** open, LANE EXISTS UNLANDED — see [[CI-RIDER LANE STATE]] at the foot of this file;
  candidate commits `1275e47` + `49d4805` + `8ed006b`, mapping NOT yet confirmed by its builder ·
  **Origin:** releases#125 fix-lane intake hold (ex releases#129 + riders)
- **What/why, three parts that land together:**
  1. **xtask check leg:** assert Cargo.lock resolves swarm-discovery to git rev
     `89a2200d54a4e3cab2f46cc75ebff49a1fb07614` while the patch is load-bearing; the check's
     message states its own drop condition (upstream ships a post-PR#27 release AND iroh's pin
     reaches it). Without it, stanza removal or a routine iroh bump silently returns the lock to
     the spinning crate and nothing reds.
  2. **Procedure rule for lock-touching lanes** (docs): targeted `cargo update -p <crate>` only,
     never full re-resolve; count changed `[[package]]` blocks AND diff per-block edges (set-identical
     hid 8 windows-sys edge movers at acaaa4f); two resolutions disagreeing = toolchain drift —
     stop and compare against CI before shipping either lock; hand-edited lock acceptable iff
     `cargo check --workspace --locked` passes.
  3. **Toolchain-version print step in golden** (cargo/rustc versions, both OS legs): the acaaa4f
     comparison against CI was impossible because no run log prints a version. One-grep audit.
- **Ripe when:** next CI-touching wave; part 1 sooner if any iroh bump is proposed.
- **Size:** small-medium (one xtask leg, one docs section, one workflow step).
- **Composed:** LOCKSMITH (#132) CI-rider cluster, hertz thin lane — 2026-08-03. GREENLIT with
  #132 and **DISPATCHED to hertz 2026-08-03** (all three parts land together).

### IR-5 — Shared nextest summary parser
- **Status:** open · **Origin:** BAROMETER triage (standing recommendation, pre-register)
- **What/why:** two agents in one day wrote `[0-9]+ tests run` parsers that read "1 test run"
  (singular) as zero — a guard fed by a broken parser condemns valid rounds. One shared,
  singular-aware parser (single-source discriminant) for every consumer of nextest summaries.
- **Ripe-when REWORDED 2026-08-04 (todlando audit): the old trigger was unfireable as worded.** At
  `b7b00c3` the in-tree population of nextest-SUMMARY parsers is ZERO — the two scripts that read
  nextest output (`g6-curve.ps1:83-90` per-test lines, `g6-postbounce.ps1:44` display-only grep)
  neither parse counts nor carry the defect, so "next wave touching any gate script that reads
  nextest output" could fire on a non-defective script while the real population (agent-authored
  throwaway parsers, which never enter the tree) stays out of reach. New trigger: **the next time
  anyone — agent or lane — needs a nextest summary COUNT**, the shared parser is built FIRST and the
  need consumes it; rig briefs should name it so throwaways stop being authored.
- **Size:** small.
- **Composed:** LOCKSMITH (#132) CI-rider cluster, CONDITIONAL — lands iff the hertz thin lane
  touches gate scripts; otherwise stays open here. 2026-08-03. Carried in hertz's 2026-08-03
  dispatch brief as a **conditional** rider; hertz reports whether the condition fired. **Not yet
  known to be building** — an unreported condition leaves this entry open, not landed.

### IR-6 — Membership logging on subnet gates
- **Status:** open · **Origin:** BAROMETER triage (standing recommendation, pre-register)
- **What/why:** counts beside results, membership beside counts — gate logs that state a count
  without naming the population keep producing unreadable reds. Standardize membership
  enumeration in gate output.
- **Ripe when:** next wave touching gate scripts / CI legs that report counts.
- **Size:** small.
- **Composed:** LOCKSMITH (#132) CI-rider cluster, CONDITIONAL — lands iff the hertz thin lane
  touches gate scripts; otherwise stays open here. 2026-08-03. Carried in hertz's 2026-08-03
  dispatch brief as a **conditional** rider; hertz reports whether the condition fired. **Not yet
  known to be building** — an unreported condition leaves this entry open, not landed.

### IR-7 — Phase A rigs leak a daemon+brain pair on Windows (exe-lock kills notify relink)
- **Status:** open · **Origin:** BAROMETER post-publish triage (ex releases#124 — full mechanism on the closed issue)
- **What/why:** two Phase A rigs launch daemons that escape the job object via WMI-rung autostart
  (double-space unquoted cmdline fingerprint; SPT_HOME in the wrapper cmdline is the attribution
  key); the leaked pair holds target/debug/spt.exe and kills every golden job reaching the notify
  relink. CI reaps as tourniquet (fa6e597); in-job test launch is the fix.
- **LANE-LINKED 2026-08-30 (#242 close sweep):** hertz's queued daemon-leak fixup lane carries
  this entry (brief cites IR-7/17/20/34/35/63); leaves the register when that lane lands.
- **Ripe when:** next wave touching the Phase A rigs or daemon autostart path.
- **Size:** medium.

### IR-8 — reap-census scoped_survivors=0 is blind to unreadable-path holders
- **Status:** BUILT 2026-08-18 on KEYSTONE #182 hygiene lane. · **Origin:** BAROMETER triage
  (ex releases#122)
- **What/why:** a zero that cannot see is not a zero — census scoping skips procs whose exe path is
  unreadable, so the survivors count can report clean while a holder lives. Needs a positive control
  / explicit unreadable bucket in the verdict line (unreadable_path count exists; the ZERO must
  refuse when it is nonzero).
- **Defining specimen (golden 30782259675, hfenduleam test job, 2026-08-03):**
  `CI-REAP summary: killed=5 kill_failed=1 scoped_survivors=0` — an admitted kill failure printed
  beside a zero-survivors claim on the same verdict line. The held image surfaced one step later:
  run-scoped tmp cleanup denied 5/5 attempts on `...\relshell\svcmock.exe` (2nd appearance of the
  svcmock hold; 1st @7a3c08c, pre-kill-auth). hertz's addendum: an image-held survivor also blocks
  WRITES to the exe path — the same class manufactures build/relink access-denied reds that mask as
  build problems, not just cleanup warnings. Not per-run: the same-sha green rerun's leg read
  `kill_failed=0 scoped_survivors=0` throughout (hertz, 30784469908) — intermittent sighting,
  second of its class, not a deterministic fixture property.
- **The Linux twin is strictly worse (todlando audit 2026-08-04, vs `b7b00c3`):** `reap-census.sh`
  has NO `kill_failed` anywhere (0 occurrences vs 2 in the `.ps1`) — its kill loop increments
  `killed` only in the success branch with no else, so a failed kill increments nothing and prints
  nothing. The specimen that made this class VISIBLE on Windows would be INVISIBLE on Linux.
  Population precision so this is not overclaimed: ESRCH is benign (already gone; the kill-time
  re-resolve makes it the common case); the vanishing case is EPERM against another account's
  process. The `.sh` `scoped_survivors` DOES come from a post-reap census re-measure, so survivors
  are measured — the hole is failed kills and the unreadable bucket, not the survivor count.
  Windows precision from the same audit: `unreadable_path` IS on the CI-CENSUS line (:157) but the
  CI-REAP verdict line (:243, :246) still carries only killed/kill_failed/scoped_survivors — the
  zero still does not refuse, exactly this entry's ask.
- **Built evidence:** both verdict lines now carry `unreadable_path`; any nonzero unreadable
  family population renders `scoped_survivors=UNPROVEN` rather than a false zero. Linux also
  counts and reports failed kills. The shared predicate is mutation-pinned by
  `reap-census-selftest.sh`; the PowerShell implementation parses cleanly.
- **Ripe when:** next census/reap script wave (natural pair with IR-7's lane) — now BOTH platforms.
- **Size:** small.

### IR-9 — Golden boxes run different clippy versions
- **Status:** pin LANDED 2026-08-18 at `35968156` (`build: pin Rust 1.96.0`), then
  audited **BUILT** at the 2026-08-29 v0.65.0 close sweep; the pin is an ancestor of
  golden-green cut SHA `4d6007ac`. The earlier measurement half had already landed and run:
  the toolchain print (`8ed006b` → `REQ-CI-TOOLCHAIN-VERSION-PRINT`) exercised both legs of
  golden run 30971976024 (`test` jobs 92198170694/92198170700 — scoped to `test`, not n1-gate;
  grep token `TOOLCHAIN `). Runner-account facts, read by doyle 2026-08-05: hfenduleam
  `cargo/rustc 1.93.0` + `clippy 0.1.93`, kitsubito `cargo/rustc 1.96.0` +
  `clippy 0.1.96`; `stable (default)` on both proved the skew would otherwise decay silently.
- **RULING (doyle, 2026-08-05, on the runner-account facts as the 2026-08-03 hold required):**
  (1) align-by-event REJECTED — with both boxes on unpinned stable, alignment decays silently;
  the skew is a mechanism and the fix must be one too. (2) declare-one-leg REJECTED as the
  terminal state — it repairs lint authority but leaves the legs resolving with different cargo
  versions, and the acaaa4f lock-attribution question this family started from is a resolver
  question. (3) **PIN IN-REPO: `rust-toolchain.toml`, `channel = "1.96.0"`** — both runner
  accounts already resolve their toolchain through rustup (witnessed by the `active=` line), so
  the pin self-applies with zero per-box maintenance, and the judge's version becomes a property
  of the TESTED SHA — the same object golden CI already guarantees. Bumps become reviewed lane
  commits, tested by the golden run they ride; the hfenduleam leg proves 1.96.0 on the pin
  lane's own run. (4) INTERIM until the pin lands: kitsubito's clippy leg is AUTHORITATIVE for
  lint disputes — not because newer is stricter (skew direction stays unasserted) but because
  1.96 is the version the pin names, so interim and terminal rulings agree.
  · **Origin:** BAROMETER (ex releases#121); re-confirmed on the #125 fix lane
  (builder's Windows clippy vs kitsubito's rust-1.96.0 lints); see [[CI-RIDER LANE STATE]]
- **What/why:** a Windows-clean lane can land a lint that only reds on the Linux leg — toolchain
  skew makes local clippy evidence non-transferable. Align versions or declare the authoritative
  leg. Natural companion to IR-4's toolchain-version print step.
- **Built evidence / fulfilment check:** `rust-toolchain.toml` contains
  `channel = "1.96.0"` and rode the tested cut. A source sweep at close found no CI or pool
  machinery keyed to a box-default toolchain **path**: golden invokes `cargo`, `rustc`, `clippy`,
  and `rustup show active-toolchain` through PATH; pool ownership keys on source-tree/lane identity,
  not rustup installation paths. `release.yml`'s `$HOME/.cargo/bin/mdbook` is an installed tool
  location, not a Rust toolchain selector; its builds still invoke PATH-resolved `cargo`.
- **Size:** built.
- **Composed:** LOCKSMITH (#132) CI-rider cluster (rides IR-4's toolchain-version print step) —
  2026-08-03. GREENLIT with #132 and **DISPATCHED to hertz 2026-08-03**; that dispatch delivered
  the measurement half (landed via [[CI-RIDER LANE STATE]], first exercised on run 30971976024).

### IR-10 — Wave gate runs the CONSUMERS of any predicate it changes
- **Status:** open · **Origin:** BAROMETER gate craft (ex releases#119)
- **What/why:** legs chosen from changed crates miss the predicate's callers; a composed red is
  triaged at the wave's own tip first. This is the gate-population rule made binding in the gate
  runbook + scripts rather than living in memory.
- **Ripe when:** next gate-runbook/docs wave.
- **Size:** small (docs + gate-script checklist).

### IR-11 — find-cwd-holders.ps1: --headless discriminator as a column
- **Status:** open · **Origin:** worktree-pin triage tooling (ex releases#118)
- **What/why:** the holder-triage script buries the headless-vs-interactive discriminator in prose;
  as a column it makes the orphan-vs-own-shell call one glance.
- **Ripe when:** any rig-tooling wave; trivial rider.
- **Size:** tiny.

### IR-12 — xtask contract-drift gate misses a stale manifest.schema.json
- **Status:** BUILT 2026-08-18 on KEYSTONE #182 hygiene lane (the narrowed blind arm only). · **Origin:** ex
  releases#116
- **KIN NOTE RETIRED (2026-09-06, hertz; doyle measure-first ruling):** the stale-binary
  premise was a rig artifact, not a reproduced product hole. At `9f809f8d`, in isolated Linux
  worktree `hertz-ir12-mutation`, build `spt` once, mutate only its root CLI help string, do not
  rebuild, then run `cargo run -p xtask -- check` alone: **exit 1**, combined output **3435 bytes**,
  ending `xtask check: docs-site/src/cli/reference.md drifted from the binary's --help` (followed
  by the regeneration instruction). The binary hash changed during check; the stale input was
  rebuilt by `gen(true)` → `spt_bin` → `cargo build -q -p spt --bin spt` before comparison.
  `CARGO_TARGET_DIR=UNSET`, `SPT_BLESS=UNSET`; Cargo's metadata target and xtask's read path both
  resolve to this worktree's `target/debug/spt`. Source mutation was restored byte-for-byte.
  Acceptance files: `~/spt-evidence/ir12-mutation-20260906/{check.exit,check.raw,env.txt}` on kitsubito.
  The original rig mechanism remains unproven (old-mtime restoration or two target directories
  are candidates, not findings). No product fix or permanent test added. IR-12 leaves the
  register at the WEBSERVE close sweep.
  The mandatory workspace-bins prebuild leg stays in every gate driver on its own merit (caught W2's spacerun red the night it was adopted); it is no longer justified by this note.
- **What/why (corrected — the consequence sentence was false at `b7b00c3`):** the literal claim
  holds — `xtask check` does not regenerate-and-compare the schema — but staleness does NOT "ship a
  wrong public contract silently": `checked_in_schema_is_current`
  (crates/spt-runtime/src/manifest.rs:2363, `int->REQ-DOCS-5`, landed `be4e46c` 2026-06-05, BEFORE
  this entry's last sweep — docs lagging code) asserts full content equality of the checked-in
  `manifest.schema.json` against what the derives generate, CRLF-normalised, `SPT_BLESS=1` as the
  regenerate path. It is REACHED (lib unit test; ci.yml:114 runs `-E kind(lib)+kind(bin)` on push;
  golden Phase A re-runs the workspace). Exactly ONE schema file exists in the tree; `docs_bundle`
  (xtask main.rs:609-618) COPIES it at build time, so no second stored copy can drift. Chain closes:
  derives → checked-in (unit-gated) → bundle (copied, not stored).
- **What remains open, and the entry narrows to it:** `check_llms_links` (xtask main.rs:521)
  hardcodes `manifest.schema.json` and `llms-full.txt` as always-existing, so the link check can
  never see them MISSING — a blind arm, not a drift hole.
- **The discrimination is now PROVEN, not presumed (todlando 2026-08-04, burn-the-build arms in
  isolated worktree ir12-mutation @`11169c1`, env verified clear of `SPT_BLESS` FIRST — set, the
  test short-circuits into a WRITE and a mutation arm silently self-heals green; that env check is
  now part of the rig recipe):** arm 0 baseline PASS by NAME (1 test run); arm 1 semantic single
  byte (title `…manifest`→`…manifesX`, length unchanged, scripted edit with match-count refusal)
  REDS at manifest.rs:2374 exit 100 — "manifest.schema.json drifted from the derives — regenerate
  with SPT_BLESS=1", with the mutated token appearing exactly once in 96,248 B of assertion output
  so the byte is provably the only delta; arm 2 whitespace-only (CRLF→LF, 1014 endings) PASSES —
  and arm 0 is itself the stronger normalisation proof, since the on-disk file is CRLF while the
  generated string is LF, so an unmutated PASS is only possible because the test normalises; arm 3
  restore re-measured PASS at the pristine sha256. `checked_in_schema_is_current` is a REAL gate.
- **Built evidence:** `check_llms_links` resolves stored root assets to their
  real source paths and runs the `llms-full.txt` generator instead of returning
  hardcoded `true`. A unit cell proves a missing schema and install script are
  refused, then accepted only after the real source file exists.
- **Ripe when:** next xtask/docs-gate wave (now sized to the blind arm only).
- **Size:** tiny.

### IR-13 — Test-soundness follow-ups from the uniform-table sweep
- **Status:** BUILT 2026-08-18 on KEYSTONE #182 hygiene lane. · **Origin:** ex releases#58
- **What/why:** the remaining candidates from the closed issue were mutation-proved before changing
  tests. `hold_outranks_everything` now walks all four `Opportunity` arms (including
  `HoldRelease`); changing held+release to start reds on that arm. Pure walk tests cover all four
  `QuiesceOutcome` arms and all five `EffectKind` arms; making `KillUnconfirmed` clear or
  `Registry` ephemeral reds independently. The join test now binds all six `JoinFail` variants to
  both hold retention and the emitted failure class/retry payload; misclassifying `WrongCode`
  reds. The access candidates had already landed in the W4 follow-up and its dead-seam removal,
  so the hygiene lane does not manufacture a second test vocabulary beside them.

### IR-14 — .worktrees audit: "how many worktrees are there" has four defensible answers, and the difference is not junk
- **Status:** open · **Origin:** ex releases#56 (was flag: NEEDS-OPERATOR in eval — operator input
  now sought directly when the entry ripens, not via board flag)
- **What/why:** the project `.worktrees/` dir accumulates content beyond what git tracks; audit +
  reap recipe + a hygiene rule for lane close-out. Teardown discipline per docs and memory (classify
  before delete, outbound links first).
- ⚠ **The header of this entry previously read "14 untracked orphan dirs vs 25 git-tracked". Both
  numbers were stale AND UNDATED**, so nobody could tell drift from error. Every count below is dated
  and carries its command.

#### THE COUNT DISAGREEMENT IS THE FINDING (measured 2026-08-03, todlando + doyle, main @`3efd7e6`)

Three people measuring "the worktrees" got three answers. None was wrong; they answered three
different questions, and nothing in the tree states which one is meant:

| answer | question it actually answers | command |
|---|---|---|
| **73** | all ENTRIES under `.worktrees/` | `ls -A .worktrees \| wc -l` |
| **52** | DIRECTORIES under `.worktrees/` | `ls -dA .worktrees/*/` |
| **38** | registered worktrees INCLUDING the root checkout | `git worktree list` |
| **37** | registered worktrees under `.worktrees/` | above, minus the root |

**52 − 37 = 15 unregistered directories.** The 73 − 52 = **21 loose FILES** are covered below.

**A count is only as good as its question.** Treat "how many worktrees" as under-specified until the
answer names its population — the same defect that made a naive `grep -rn` from the project root
inflate a code count by **34.9×** (426 tracked `.rs` vs 14865 on disk excluding all `target/`), and
made that grep run past 120s while `git ls-files | xargs grep` returned instantly. **Scan roots and
population definitions are the same class of error.** Use `git grep` / `git ls-files` for tracked
content and `git worktree list --porcelain` for worktrees.

#### THE 15 UNREGISTERED DIRECTORIES ARE FOUR DIFFERENT KINDS OF THING

⛔ **CLASS A — LIVE BUILD POOLS, NOT ORPHANS. DO NOT DELETE.** 3 dirs, 13.2 GB. Each is the TARGET of
a junction that a REGISTERED worktree uses as its `target/`:

| dir | inbound junction from | size | claim |
|---|---|---|---|
| `gate-target` | assembly-doorbell, doorbell-w1, doorbell-w2, doorbell-w3 | EMPTY | none |
| `gate-target-render` | golden-render | 5.84 GB | `POOL-OWNER.json` → golden-render |
| `gate-target-w4doc` | w4-cli-doc | 7.36 GB | none |

**These are exactly the directories that read as obviously junk** — no `.git`, no source, leftover
names — and deleting one destroys a registered worktree's build pool and leaves a dangling junction.
**Polarity was checked BEFORE classification, which is what caught it:** all 15 top-level dirs are
REAL directories, none is itself a reparse point; the six junctions are **INBOUND**, at
`<registered-worktree>/target`. See [[worktree-target-junction]], [[gate-worktree-target-disk]].

Two findings inside class A, neither of them a deletion question:
- **`gate-target` has FOUR registered trees junctioned into ONE pool, with no `POOL-OWNER.json` at
  all** — so nothing would refuse a second LIVE lane there. That is releases#103's exact hazard
  sitting armed. The pool is also empty: someone reclaimed it and left four junctions aimed at a hole.
- **`gate-target-render`'s `POOL-OWNER.json` carries only `owner_tree` and `written_by` — no pid and
  no birth stamp.** A claim that cannot distinguish a live lane from a finished one is missing the one
  property the claim mechanism exists to provide.

**CLASS B — EMPTIED SKELETONS, ONE UNIFORM SHAPE.** 8 dirs, **0 files**: `acl-core`, `engine-room`,
`ff-fastfollow`, `gate-23d5ceb`, `gate-ff`, `golden-a`, `golden-b`, `w2b-sender-stamp`. Every one is
exactly `crates/spt-daemon/` and nothing else. **Eight independent removals stopping at the SAME
relative path is a mechanism, not litter.**

**MECHANISM CORROBORATED LIVE, 2026-08-03:** a read-only PEB sweep of process CWDs during a running
`spt-daemon` suite found ~20 processes — `spt_daemon-<hash>` test harnesses, `PING`, `cmd` — whose
CWD was **exactly `<worktree>/crates/spt-daemon/`**, the precise path all eight skeletons froze at. A
`git worktree remove` racing a straggler there deletes everything else and leaves that chain pinned.
The processes churn fast (all 17 sampled pids were gone within 30s, one already showing **pid reuse**
— see [[pid-reuse-across-reboot]] for why a pid alone is never an identity), so the pin is a race, not
a steady state, and it recurs on every daemon-suite worktree.

**CLASS C — FULLY EMPTY, no inbound junction.** 3 dirs, 0 bytes: `gate-41`, `gate-559632e`,
`release-runbook-main-advance`. Class B with even the chain gone.

**CLASS D — DELIBERATE SCRATCH, NOT RESIDUE.** 1 dir, 12 files, 50 KB: `_patches` — three `.patch`
files with their `.untracked` manifests, plus `ir18-gate-logs/`. Modified the day before the audit,
i.e. someone's working state.

3 + 8 + 3 + 1 = 15.

#### PIN STATE: NO SKELETON IS HELD TODAY (with a stated blind spot)

Read-only PEB CWD sweep, 2026-08-03: **zero processes hold a CWD in any class-B or class-C
directory** — every live pin was in `gate-ec5f38a` and `gate-main`, both REGISTERED and both running
rigs at the time. So removal of B and C would succeed today; nothing is retrying it.

⚠ **The sweep read 412 of 593 processes; 181 were unreadable** (elevated/system, no
`PROCESS_VM_READ`). The no-pin result therefore holds over the READABLE population only. Cheap to
re-run elevated before acting — [[absence-needs-sibling-probe]].

#### THE LOOSE FILES: A SHARED LOCATION WITH NO STATED CONTRACT

**21 loose files sit in `.worktrees/`** — gate logs (`gate-*.log`, `gate-559632e-log.txt`), rig
scripts (`g6-curve.ps1`, `g6-postbounce.ps1`, `gate-w5-*.ps1`), and `gate-w5-notes.md`. **Nobody
declared `.worktrees/` a log drop; it became one.** Same class as the memory index: a shared location
with no stated contract accumulates whatever anyone puts there, and **the first person to tidy it
cannot tell residue from someone's working state** — class D is that risk already realised. The
hygiene rule this entry owes should name where gate logs and rig scripts belong, not only how to reap
worktrees.

#### RECLAIM ARITHMETIC — AND WHY THIS ENTRY IS NOT THE DISK FIX

Classes B, C and D together are **under 51 KB**. All 13.2 GB of the unregistered population is class
A, behind live junctions. **An orphan sweep is not a disk-space remedy**, and reading it as one sends
you at the wrong target: on the audit date, free space was **12.5 GB against the 32 GB golden floor**
([[free-space-floor-blocks-golden]]) while the two largest pools on the box were `gate-ec5f38a/target`
(14.33 GB) and `gate-main/target` (28.12 GB) — **both REGISTERED, so both outside the orphan
population entirely.** The disk question and the orphan question have different populations; answering
one correctly says nothing about the other.

- **AUDIT EXECUTED 2026-08-30 (doyle, the v0.66.0→NOW-SIGNAL between-milestone window; full
  output `wt-census-2026-08-30.txt` in the session scratchpad, method = this entry's own):**
  four answers BEFORE the sweep: 85 entries / 85 dirs / 76 registered incl root / 75 under
  `.worktrees` → 10 unregistered, 0 loose files (the 21 loose files of 08-03 are gone).
  **The 08-03 headline populations have INVERTED:** zero pools and zero junctions exist under
  `.worktrees` at all (every registered tree reads `target=none`; the reparse sweep found no
  inbound junction) — class A is EMPTY, so the armed pool hazards this entry named are
  currently unpopulated. Unregistered classified: 4 fully-empty (C), 3 `crates/spt-daemon`
  skeletons (B — the mechanism's fingerprint again), `_patches` (D, kept — working scratch),
  and TWO content-bearing checkout remnants (gate-signet-w1 10.8MB, gate-w1-a8f04aff 27MB) =
  registry-pruned dirs held by the conhost CWD pins IR-63's third face names. **Sweep
  executed on the gater's own rigs:** evidence preserved FIRST
  (`spt-preserve\audit-2026-08-30\`: gate-w1-a8f04aff's 229 non-HEAD files as a verified
  tarball; er-instrument's 115-line UNCOMMITTED instrument diff as a patch — that worktree
  belongs to hertz's lane and was NOT touched beyond the read-only copy), 16 pinning conhosts
  killed by pid after a fresh probe, 9 dead rigs removed. AFTER: **76 dirs = 75 registered +
  `_patches`, fully reconciled; zero unexplained entries, zero pins.** Reclaim was ~38MB —
  reconfirming this entry's arithmetic that the orphan sweep is hygiene, not a disk remedy.
  Landed-classification caveat: `git cherry` patch-id reads conflict-resolved picks
  as UNLANDED (emit-single-write, fix-206/208/209 all show UNLANDED with shipped content), so
  the census's landed column is a datum, never a reap authorization.
- **HYGIENE RULE (the text this entry owed, written 2026-08-30 while the loose-file population
  is zero — the rule exists BEFORE it refills):** `.worktrees/` holds exactly two kinds of
  entry: **registered worktrees** and **`_patches/`** (the one declared scratch location).
  Nothing else. Gate/rig artifacts follow their lifetime: (1) driver scripts, exit/raw files,
  and logs live INSIDE their rig's worktree for the rig's life — they leave WITH it, via the
  preservation step (copy what the verdict cites into `spt-preserve/<occasion>/`, verified by
  cmp or a listed tarball, restore cost stated) and never accumulate beside the rigs; (2)
  anything meant to outlive its rig goes to the repo ROOT beside the verdicts and RCAs it
  supports (the census/finding/JIT convention) or into `spt-preserve/` — never loose in
  `.worktrees/`; (3) a file found loose in `.worktrees/` is treated as UNCLASSIFIED WORKING
  STATE — moved to `_patches/` with a dated note, never deleted on sight (class D is the risk
  realised). Lane close-out = `git worktree remove` + prune + the conhost CWD-pin check
  ([[IR-63]] third face) when removal refuses; a registry-pruned dir left behind is a defect
  to sweep, not a norm.
- **Ripe when:** between-milestone idle window (it is a dev-box chore, zero product risk) — but the
  class-A pool findings are armed hazards and do not wait for it.
- **Size:** small for the sweep; the hygiene rule and the pool-claim gaps are separate small items.
- **Field instance (hertz sweep, 2026-08-04 — an unclaimed pool found rather than reasoned about):**
  `.worktrees\gate-target-w4doc` held **7,907,558,098 B with NO POOL-OWNER.json at all**. Not a stale
  claim, not a foreign claim — no claim. Exactly this entry's shape, first measured instance. Reaped
  under doyle's ruling as part of [[IR-27]]'s two-step teardown; the measurement stands on its own.

### IR-17 — Deadline-burn bring-up family: a test burns its full window while the daemon/brain never comes up
- **Status:** open · **Origin:** golden 30782259675 red triage 2026-08-03
- **What/why:** two different tests on two runner hosts now share one signature — bring-up misses
  its ONLINE window under leg load and the test burns its ENTIRE deadline before the PRECONDITION
  panic: `spt::resident_service_e2e::a_declared_service_rises_with_the_daemon_and_reaches_the_cli`
  (hfenduleam, FAIL 123.99s, "PRECONDITION: the daemon never came up", run 30782259675, green on
  same-sha rerun 30784469908) and `spt::activity_link_push_e2e` (kitsubito, full 30s, run
  30607903133 specimen, green on same-sha rerun; its "brain stderr EMPTY" leg and its family
  membership are both CORRECTED in the 2026-08-30 sweep update below — the kitsubito member is
  CLOSED under a named mechanism).
  Family reading, not one flake: same missed-window shape, host-independent, both mid-leg under
  load. Cross-ref KNOWN-HAZARDS 5.13 — bring-up has a hard ONLINE budget with known sensitivity to
  anything that stalls it (the blanket-fsync canary; `attach_wedge_e2e` guards that budget).
- **State line (the discriminating datum, resident_service red):** `daemon_up=false boot_alive=true
  boot_pid=Some(8920) rel_started=false broker_survived=false` — the boot process was ALIVE at
  panic time and the daemon never reached up. Any characterization run records this line per run,
  not the verdict (hertz protocol 2026-08-03).
- **Read of that line, CORRECTED 2026-08-03 (hertz):** it is NOT "started-but-never-bound" (the
  earlier reading here) and not "never-started" — it is **started, did work, then killed** —
  CANDIDATE via [[IR-18]]. The daemon spawned its boot service and that service reached the CLI
  (the spool holds its message), and `broker_survived=false`. The DISCRIMINATING field is
  `broker_survived`: false in every killed round of the positive control, true in every healthy
  round. **An earlier evidence leg here is RETRACTED (hertz, same night, off the negative arm he
  ran before reporting; deployah, who had carried the leg into four register sites, swept every
  placement): "empty daemon.stderr.log = TerminateProcess signature" was false — the stderr
  block is empty on PASSING runs too, so it carries zero information in either direction; it
  failed the discriminator question and does not support the kill read.** The kill read now rests
  on the control's field-for-field reproduction and the IR-18 mechanism, not on the log.
  Consequences: (a) the bring-up window is **exonerated for this specimen** — measuring it would
  have measured nothing; (b) the candidate mechanism and its evidence now live in [[IR-18]]
  (a sibling test's bare breadcrumb tree-kill), CANDIDATE — not reproduced, not proven;
  (c) **the characterization population changed, and the rate run is CANCELLED** (doyle ruling
  2026-08-03) — `resident_service_e2e` run ALONE has no sibling to collide with, so the mechanism
  predicts 0/N and that number answers no live question; the only sound population was the test
  INSIDE the Phase A parallel pool, which costs a full Phase A leg per sample, and the fix is
  warranted by the source-verified hazard class regardless of the specimen's rate. No rate figure
  will exist for this specimen — do not later read its absence as a low rate; (d) **positive
  control still runs, falling out of the
  hypothesis:** kill the spawned daemon by pid mid-bring-up — after boot-service spawn, before
  ready — and it must reproduce the observed line field for field (`daemon_up=false
  boot_alive=true broker_survived=false`). If the rig cannot make that shape on
  demand, a 0/N from the untouched arm is worth nothing and must be reported as worth nothing.
  The host-INDEPENDENT family claim (kitsubito's `activity_link_push_e2e`) is untouched by this:
  it has no such kill site named, so the family survives even if this specimen leaves it.
- **POSITIVE CONTROL RAN 2026-08-03 (hertz, prebuilt a5042ec binary, box confirmed clear —
  0 open runs, 145.95 GB free): shape reproduced 3/3, deterministic.** Negative arm n=2: 2/2 PASS
  (5.59s, 5.46s), `daemon_up=true boot_alive=true rel_started=true broker_survived=true
  survived_teardown=true`. Positive arm (daemon killed by parent-scoped descent after
  boot-service spawn, before ready) n=3: 3/3 FAIL (93.95/93.25/93.50s), state line identical to
  the CI red in EVERY field except boot_pid (a pid, must differ), same panic, same site
  (resident_service_e2e.rs:382). Structural fact the arm settled, and the refutation risk that
  made it worth running: the boot service rises BEFORE brain.ready becomes readable — had the
  order been reversed the arm would have produced daemon_up=true and refuted the mechanism.
  CEILING MET, NOT EXCEEDED: the control used Stop-Process -Force (TerminateProcess — the same
  primitive as taskkill /F), so it proves a forced daemon kill after breakaway-service spawn
  produces this exact line ON DEMAND; it says nothing about WHO issued one in golden 30782259675.
  Candidate mechanism, shape reproduced on demand — "root caused" written by nobody. Duration
  note, recorded not explained: 93.5s local (idle box) vs 123.99s CI (Phase A parallelism), both
  burning the same three 45s windows — consistent, not verified.
- **Same-host sub-observation (hertz, host-CONSTANT — narrower claim, kept separate):** both
  reds of 2026-08-03 sat on hfenduleam Windows Phase A and burned their full windows:
  `a_tree_teardown_reaches_a_grandchild_the_service_spawned` (run 30776330383, FAIL 10.176s = the
  full poll deadline, vs a 0.19–0.66s pass band — 15–50x out) + the resident_service row above.
  MEMBERSHIP PROVISIONAL for the teardown row: it carries a candidate mechanism the bring-up burn
  does not share — the pid-only oracle IR-15 just replaced. If the oracle caused it, that row
  leaves the family and host-constant collapses to one sighting. The host-INDEPENDENT pair above
  is the claim that survives someone fixing this box; the sub-observation is what is actionable
  about hfenduleam (a live-daemon host) meanwhile. **UPDATE 2026-08-04 (golden 30873007187
  attempt 1 — SECOND sighting, and the provisional question above is now ANSWERED): the
  pid-only-oracle candidate is REFUTED for this row** — IR-15's authenticated repin rode the
  failing tree (a5042ec ancestor of 4b37512) and the row failed anyway (10.225s, "grandchild
  51848 outlived the tree teardown"), with is_stamped() asserted before the kill so the verdict
  passed the authenticated gate. The row therefore STAYS a family member with its mechanism OPEN,
  narrowed by hertz's log forensics (competence-controlled: the capture channel demonstrably
  prints): no DETACH_BREAKAWAY_DENIED (breakaway rung taken), no SERVICE_JOB_UNAVAILABLE (job
  created + child enrolled), no SERVICE_TREE_KILL_INCOMPLETE (TerminateJobObject returned
  success) — everything enrolled died; THE SURVIVOR WAS NEVER ENROLLED. Two hypotheses,
  inseparable in existing evidence: H1 test defect — the grandchild SELECTION is a bare
  th32ParentProcessID match over a recycling pid space (IR-15's mechanism one level up: the
  verdict was authenticated, the selection was not); H2 real enrollment gap. hertz's selection
  probe (IsProcessInJob at selection time — decisive; birth stamp one-directional; image; same-
  snapshot match count) is authored+compiled and makes the NEXT occurrence self-deciding. Row
  state: INSTRUMENTED-AWAITING-FIRE; a green retires nothing (doyle-ruled). **UPDATE 2026-08-04
  (USHER att3):** row green at 17f95a7 (retires nothing, per the ruling). New shape evidence while
  waiting for the probe: BOTH archived reds sit AT the 10s poll bound (daemon.rs:2757) — 10.176s and
  10.225s — while todlando's 200 off-CI passes under live-fleet load ran 1-4s, never near it. The
  post-fix red is therefore BOUND-EXPIRY-shaped (kill fired; the pinned identity stayed
  not-provably-gone past 10s — TerminateJobObject is async), not kill-wrong-shaped. The bound
  question (is 10s right for a loaded runner) is parked with hertz's USHER package item 4 and must
  not touch the kill path or spawn flags before the probe fires. See also the FLAKE-LEDGER row for
  the false-repeat correction (the 10.176s red was the PRE-a5042ec bare-pid mechanism).
- **Declared read (2026-08-03, rules the next red):** a same-sha green rerun is protocol-conclusive
  for the GATE, not the class. No third silent sample — the next occurrence of this signature gets
  an instrumented resident-service bring-up investigation (hertz), not a rerun.
- **SWEEP UPDATE 2026-08-30 (#242/v0.66.0 close — three corrections, correct-by-replacement):**
  (1) **The 2026-08-03 "brain stderr EMPTY" reading on the kitsubito specimen is a METER
  ARTIFACT, not a silent brain** (deployah measured 2026-08-29): the brain-panel stderr capture
  only LANDED 2026-08-19 (`ab657626` + `e6e188e1`), 16 days after the specimen — at the specimen's
  sha there was nothing to capture, so "empty" carried zero information; current bits label
  absence explicitly, so empty-vs-present cannot recur as a distinction. This replaces the
  "cheap to settle next time someone has kitsubito" item — there is nothing left to settle.
  (2) **The kitsubito family member is CLOSED under a NAMED mechanism:** RCA-242-R2-LINUX.md
  (v0.66.0 golden r2) measured `current_exe_hash()` on the READY PATH at ~10.1s per boot on
  kitsubito under load (~1.5s Windows) via BRAIN_PHASE breadcrumbs + harvest-loop repro — a
  bring-up that misses its ONLINE window under leg load is exactly what a ~10s synchronous hash
  inside the ready path manufactures. The behaviour-change request (hash off the ready path,
  ADR-0018 Q7) is minted on the board at this sweep as **releases#244** (BACKLOG, operator
  triages). The family's host-INDEPENDENT claim
  therefore collapses: the Windows `resident_service_e2e` specimen (IR-18 kill-shape candidate)
  and the hfenduleam sub-observation STAY OPEN as this entry's remaining live population.
  (3) **Lane-linked:** hertz's queued fixup lane brief names this entry directly
  (activity_link_push_e2e 4-cell fixup: false precondition message + progress-bounded readiness
  wait, #235 shape). Overlap with correction (2) checked at link time: his cells are TEST-side
  (message truth + wait shape) and stand regardless of the product-side hash fix.
- **Ripe when:** next occurrence of the signature (immediate instrumented dispatch), or a hertz
  test-hardening wave (bring-up phase-timing instrumentation, per-phase deadline attribution).
- **Size:** medium (bring-up instrumentation + attribution; the fix depends on what it shows).

### IR-16 — kill_tree discards TerminateJobObject/TerminateProcess returns (silent partial kill)
- **Status:** BUILT 2026-08-03 (landed @62c5623 on the LOCKSMITH t1 lane; rides golden sha
  b7b00c3, run 30860770146 green) · **Origin:** ex releases#130 (golden 30776330383 red triage;
  deployah's log read
  + todlando's flagged-not-asserted arm). Operator-classified infra 2026-08-02: daemon-kill
  internals are agent-facing, not operator-facing surface.
- **What/why:** `DetachedChild::kill_tree` (daemon.rs:1433 vicinity) calls
  `unsafe { TerminateJobObject(self.job, 1) }` and discards the return; TerminateProcess likewise
  unaudited. A job that was created and assigned at spawn but whose TERMINATION fails at kill time
  yields the grandchild-survives symptom with zero log signal — the spawn-side
  SERVICE_JOB_UNAVAILABLE announcement is correctly absent (that arm is instrumented and was
  refuted for the specimen by a positively-controlled zero). Wanted: (1) check + log both
  termination returns loudly on failure; (2) the degraded-arm decision — fallback process-table
  tree-walk kill or explicit refusal, never a silent partial kill wearing REQ-RESIDENT-SERVICE's
  unconditional promise. Do NOT add self.job==0 instrumentation (already loud; a second weaker
  rule beside a working one). Discrimination pairing with [[IR-15]]: a red whose captured pid is
  still ping.exe with no SERVICE_JOB_UNAVAILABLE in scope = this entry's arm.
- **Ripe when:** RIPE NOW — IR-15 landed BUILT 2026-08-03 (its rig-side half), so this kill-side
  half is the outstanding instrument; land with the next daemon-teardown wave or sooner.
- **Size:** small.
- **Composed:** LOCKSMITH (#132) — todlando product lane, with the `detached_no_inherit_env`
  rename as cosmetic rider (IR-20's load-bearing-pid caution applies to any rig touch) —
  2026-08-03. GREENLIT with #132 and **DISPATCHED to todlando 2026-08-03**, carrying IR-20's
  load-bearing-pid precondition as a stated check (todlando reports whether the rig is touched
  at all rather than silently skipping it).
- **BUILT — fulfilment checked against the wanted list, not merely mapped to a commit** (doyle,
  2026-08-04): `62c5623` satisfies (1) — both termination returns are now checked and the loss is
  NAMED on failure, as the kill-time twin of the spawn-time `SERVICE_JOB_UNAVAILABLE` — and (2) by
  the ruled LOUD REFUSAL arm rather than a tree-walk fallback, which [[IR-18]] measured blind in
  exactly this failure state. The `self.job == 0` prohibition was honored: the new check is
  `self.job != 0 && TerminateJobObject(...) == 0`, not a second weaker rule beside the working one.
  The unix `ESRCH` quiet arm — which no Windows gate can see — is covered by `ec5f38a`. The
  `detached_no_inherit_env` cosmetic rider did NOT land and was dropped at `de6a01d` on a MEASURED
  false premise, not deferred. A Windows-specific finding rides the fix as a comment and is worth
  keeping: `TerminateProcess` against a handle to an already-exited process returns 0 with
  `GetLastError` 5 (`ERROR_ACCESS_DENIED`), so a naive check there fires on the ORDINARY path.

### IR-18 — resident_service_e2e kills breadcrumb pids BARE (stale-breadcrumb tree-kill; kill-side twin of IR-15)
- **Status:** BUILT 2026-08-04 — the named acceptance carrier RAN and the condition is DISCHARGED,
  stated as the measurement rather than as the green board it sat on: LOCKSMITH's golden run
  30860770146 (head `b7b00c3`, conclusion success) executed the changed test IN-POOL on BOTH
  platforms — `spt::resident_service_e2e a_declared_service_rises_with_the_daemon_and_reaches_the_cli`
  PASS 17.773s (100/2637) on hfenduleam/Windows and PASS 17.895s (1364/2616) on kitsubito/Linux.
  The pool was checked for the test rather than inferred from the job colour, because a green leg
  that never ran the test is a competence-controlled zero, not an acquittal; `resident_service_e2e`
  contains exactly ONE `#[test]`, so the test that ran IS the test the lane rewrote (229 lines of
  that file changed in `6e1a962`). No red returned to doyle. `.worktrees/ir18` is released for
  teardown by this verdict · **Origin:** hertz
  characterization of the [[IR-17]] resident_service red, 2026-08-03 — CANDIDATE mechanism, NOT a
  reproduction
- **Lane:** hertz `fix/ir18-authenticated-teardown` — MERGED to main @6e1a962 (PR #142; gated at
  b48e35b, rebased 6e1a962 on 1e520b4 with zero code delta — doyle re-derived `git diff -- crates/`
  empty across the rebase, so the gate verdict and local behavioral evidence transfer). Gated by
  doyle: full diff review (one blocking finding — the scope_lost assert condition contradicted its
  own guard ruling — fixed and delta-verified at one line), mutation-proven locally both
  directions (B: forced boot_pid=None → population sweep names the leak no per-pid check sees,
  FAIL 101; A: wrong expected_exe → REFUSED-foreign-image reds loudly on the derivation check).
  EVIDENCE SPLIT, stated so the check marks cannot carry it (hertz's absent-leg callout): lane CI
  30792210008 was 5/5 green per job but ci.yml's test leg is kind(lib)+kind(bin) — the changed
  integration test NEVER RAN there (verified by log grep, 0 hits / 3023 lines); CI proved
  compile-everywhere (clippy all-targets = the pool-population check for a common/ module),
  traceability, lib/bin clean. The BEHAVIORAL evidence is local: three clean runs (all verdicts
  Killed, population empty) + the two mutation reds. reap.rs diff additive-only (zero deleted
  lines, verified) — cross-test risk bounded to compilation, which CI covered. Evidence custody:
  gate + mutation logs at `.worktrees/_patches/ir18-gate-logs/` (5 files); `.worktrees/ir18` is
  KEPT deliberately until the golden verdict (NOT an orphan — do not reap; a Phase A red wants the
  tree and logs in place, not rebuilt); its pool claim is released.
- **What/why:** `crates/spt/tests/resident_service_e2e.rs:46` defines its own reaper —
  `taskkill /PID <pid> /F /T`: bare pid, force, whole TREE, no identity check — and feeds it three
  BREADCRUMB-derived pids in cleanup (`boot_pid`, `rel_service_pid` at :360, and the pid read out
  of `brain.ready` at :363). The suite already ships the authenticated tool other tests use and
  this one does not: `common::reap::authenticated_kill(label, pid, expected_exe, observed)` at
  `crates/spt/tests/common/reap.rs:116`. Inside the Phase A parallel pool on a pid-churning box
  this is the KNOWN-HAZARDS stale-breadcrumb tree-kill class, live and **symmetric**: this test can
  take a sibling's daemon and a sibling can take this test's. Same defect [[IR-15]] just fixed, on
  the other side — IR-15 authenticated the READ (is the thing I pinned gone?), the KILL is still
  bare pid. Kill what you pinned, not the number it happens to hold.
- **Verified at source by deployah 2026-08-03** (relayed claims re-derived, all held): the bare
  kill, its three breadcrumb feeds, the shipped-but-unused authenticated helper, and the
  `CREATE_BREAKAWAY_FROM_JOB` spawn (`daemon.rs:1300`).
- **Evidence rating — why it is a candidate and not a cause:** it predicts every field of the one
  observed state line. The daemon is killed after spawning its boot service and before binding:
  `broker_survived=false` (the discriminating field — false in every killed control round, true in
  every healthy one) + `boot_alive=true boot_pid=Some(8920)` with the service's message in the
  spool. (The "empty stderr = silent death" leg that originally sat here is RETRACTED — see
  [[IR-17]]'s control record; stderr is empty on healthy runs too and discriminates nothing.)
  The service outlives its daemon because `detached_no_inherit_env` spawns it
  `CREATE_BREAKAWAY_FROM_JOB`, so it is NOT in the daemon's job and a daemon kill ORPHANS it
  rather than reaping it — plausibly also the five denied `svcmock.exe` attempts in that job's
  reap summary. **The [[IR-17]] positive control (2026-08-03) reproduced the shape 3/3
  deterministically with a forced kill at that window — establishing the SHAPE on demand, not the
  AGENT: nothing identifies who issued a kill in golden 30782259675.** The bare-pid tree-kill from
  a concurrent test remains the candidate agent, on the source-verified hazard class and reap.rs's
  documented prior casualty. **Not reproduced in the wild. Not proven. "Root caused" is not
  written here by anyone.** The [[IR-17]] rate run is cancelled, so no rate will ever back this —
  the fix stands on the source-verified hazard class alone. THE DEFECT IN MINIATURE, observed as a
  side effect of the control (hertz 2026-08-03): the three killed rounds leaked 12 processes
  (6 svcmock + 6 spt.exe, two services per round) and the test's own teardown reaped NONE — when
  the daemon dies, `rel_pid` is `None` so that service is never even a kill target, and the
  breakaway child outlives everything. A teardown that leaks two processes per round currently
  PASSES: the concrete case for `target_gone` as the gate's positive half. (Leak cleaned scoped,
  each pid re-authenticated by image-path prefix + creation window at kill time, not from the
  minute-old enumeration.) ONE LAYER DEEPER (deployah, same night, at source): the reap loop
  (:359) is `[boot_pid, rel_service_pid].into_iter().flatten()` — `.flatten()` DROPS None, and
  `mock_pid` (:100-106) collapses absent/permission-denied/IO-error/garbage into that None via
  `.ok()?`/`.ok()` — the reap.rs:59-63 collapse in a THIRD location, inside this very test. So the
  leak is both "service never started" AND "pid unreadable, therefore never a kill target":
  absence of knowledge read as absence of target, in a teardown, again.
- **GATE SPEC AMENDED PRE-BUILD (deployah found the hole, doyle-ruled 2026-08-03 — read this
  BEFORE building the lane):** the two halves as first specified (trap-class-absent +
  `target_gone` per pid) both operate on pids we HAVE — a process whose pid is None is invisible
  to BOTH, so the specified gate PASSES hertz's own control-run leak (12 processes, 6+6, 2/round,
  teardown reaped none). A gate its own reproduction case passes is not yet an instrument. THIRD
  HALF, REQUIRED: a POPULATION assertion at teardown — zero surviving staged-service or
  test-owned spt processes, by selectors external to the pid bookkeeping. SELECTOR RESPELLED
  (hertz 2026-08-03, doyle-approved — the first spelling, "parent-pid descent from the test
  process", would have caught ZERO of the 12: teardown kills the daemon before any check runs, the
  BFS over the current (pid,ppid) table breaks at the dead middle hop, and descendants(test_pid)
  returns empty exactly in the leak case — an absence read as a clean, the shape the lane exists
  to kill; the control's own cleanup is the evidence, it had to use image+creation-window because
  descent was already broken). THE TWO SELECTORS AS BUILT: (1) svcmock half by IMAGE PATH —
  stage_adapter copies the service binary under this run's unique tempdir home, so "any pid whose
  exe_path canonicalizes under home" names every staged service this run started and nothing else
  on the box; no bookkeeping, immune to reuse and broken chains. (2) spt.exe half by SEEDED
  descent + image — image alone is forbidden (target/debug/spt.exe is shared with concurrent
  tests and live perches: the machine-wide selector class); seed with the ancestor set captured
  WHILE ALIVE ({test, broker from the Child handle, brain from brain.ready — none via mock_pid}),
  union descendants, keep only exe_path==spt_bin. Reuse exposure on seeds is assertion-only (a
  red, never a kill) and the image filter closes it. Both halves catch the 12-leak conformance
  case; the first spelling caught none of it. PLUS THE [[IR-8]] BUCKET, REQUIRED (deployah,
  same night — the third instance of that sentence tonight, this time inside our own instrument):
  both selectors match on IMAGE PATH, and an exe path that cannot be READ yields no match — a
  leaked svcmock holding an unreadable path is invisible to both halves and the count comes back a
  zero that cannot see. Not hypothetical: IR-8's defining specimen is 5/5-denied on this exact
  image, on this box. So the population assertion counts unreadable-path processes as their OWN
  bucket and the zero-survivors claim REFUSES when that bucket is nonzero — zero seen AND zero
  unseeable, or the assertion states it could not see. Built in from the start, not retrofitted
  (may land in the lane's second commit at hertz's discretion; the requirement is that it lands in
  the lane). This instantiates IR-8's remedy test-side; the CI reap-census script half of IR-8
  stays open.
- **Fix (thin lane, hertz — doyle-ruled 2026-08-03):** adopt `authenticated_kill` at all three
  sites. (Board #131's `CREATE_NO_WINDOW` rider is DROPPED from this lane — hertz falsified the
  filed remedy at source 2026-08-03: daemon.rs has exactly ONE CreateProcessW (:1187) and
  BASE_FLAGS (:1105) already carries 0x0800_0000 = CREATE_NO_WINDOW on both rungs including the
  ACCESS_DENIED fallback (which drops only BREAKAWAY), so the specified edit ORs an already-set
  bit — a bit-for-bit identical flag word that would close a board item while its symptom, if
  real, continues. #131 handed back to board triage with the finding and a candidate direction:
  DETACHED_PROCESS gives the service no console, so a console child the SERVICE spawns flagless
  allocates a NEW visible window — the grandchild, not the daemon's spawn, is the candidate
  surface; wants its own diagnosis from an observed window. ~~The `detached_no_inherit_env` rename
  ("detached" actually means breakaway-from-job) is a candidate cosmetic rider on [[IR-16]]'s
  product lane~~ — **RENAME DROPPED 2026-08-03, ITS PREMISE MEASURED FALSE** (todlando, on the
  IR-16 lane; doyle re-verified at source before ruling). "detached" is ACCURATE: `BASE_FLAGS`
  (daemon.rs:1103-1105) is `DETACHED_PROCESS | CREATE_NEW_PROCESS_GROUP | CREATE_NO_WINDOW`, and
  `0x0000_0008` IS DETACHED_PROCESS — the function sets BOTH postures and the name states one of
  them. The name is not wrong, it is INCOMPLETE (silent about `CREATE_BREAKAWAY_FROM_JOB`, which
  arrives as `extra_flags` at daemon.rs:1506), so renaming on the stated premise would have traded
  an accurate word for one that drops a flag the function really sets. Independently fatal: the
  premise, if true, covers the sibling `detached_no_inherit` equally — 17 references across
  daemon.rs/deelevate.rs/shellhost.rs — and renaming one of two siblings of ONE posture leaves the
  tree MORE inconsistent, not less. **What lands instead:** one doc-comment line on EACH function
  naming the full effective flag word, which fixes the real complaint (neither name says breakaway)
  without touching a call site. Kept here rather than deleted because this entry was the claim's
  only home, and a reader who found the rename gone with no reason would re-derive it.) PER-PID GROUND TRUTH for expected_exe (hertz — SELF-CORRECTED
  at source before it ever landed; his first derivation, "Copy runs from adapters_dir(), Pointer
  runs from srcs/", was WRONG and registration mode is NOT the discriminator): servicehost.rs:1328
  takes `install_dir` from `record.source_dir` on the Started arm for EVERY service, Copy and
  Pointer alike, and Copy copies only the manifest + strings/ (registry.rs:395-404), never a
  binary — `adapters_dir()` holds no executable in either mode. BOTH services run from their OWN
  staged src dir: `staged_bin(<adapter's staged src dir>, "svcmock")`. The trap survives with a
  different mechanism: two adapters staged into two DIFFERENT source dirs are two different files
  on disk, so one expected_exe still cannot serve both. A lane that trusted the mode split would
  derive a path with no binary at it. How it was caught, kept because it is the reusable part: not
  by rereading the mode table but by finding where install_dir is COMPUTED instead of trusting an
  already-published inference — the mode split was true of MANIFESTS and had been generalized to
  BINARIES without checking the consumer; and the derivation-asserted-at-observe() instrument
  would have caught it at runtime regardless — the instrument did its job before it ever ran,
  argued against its own author. SCOPE_LOST RULED A GUARD, NOT A FINDING (hertz design change,
  doyle-approved): a seed pid recycled before the sweep makes the sweep DECLINE that subtree
  loudly, not fail — failing would turn ordinary pid churn on a busy runner into a red against a
  clean teardown (the IR-15 false-red class re-minted inside its own fix); coverage holds without
  it because the staged services are caught by image-path-under-home and the daemon + brain by NEW
  direct per-pid target_gone checks (their pids this test never loses) — the ancestry half is
  reach, not load-bearing. `unreadable` stays FATAL as ruled.
  Sequencing: positive control
  first (it is the instrument that could still refute the mechanism), then the lane. Does NOT ride
  [[IR-16]] (doyle-ruled same day): that is the product-side kill_tree audit and stays a separate
  lane under the dispatch split; the two cross-ref, they do not merge.
- **Lane trap, flagged before build (hertz 2026-08-03 — the false-clean shape):** `expected_exe`
  differs PER PID. `boot_pid` and `rel_service_pid` run the staged svcmock image, NOT spt.exe —
  passing spt_bin for those refuses every kill as foreign-image and LEAKS both services while the
  reap reads hardened. Only `brain.ready`'s pid takes spt_bin. Also: `observe()` boot_pid at :210
  while it is provably ours, so the reuse check has creation-time teeth instead of degrading to
  image+ancestry.
- **Gate the lane on killing, not on refusing (deployah 2026-08-03; REFINED by hertz same night
  from reap.rs's own contract):** the hardened version fails safe in the WRONG direction —
  refuse-everything leaks both services while printing exactly the reap line a reviewer wants to
  see, strictly worse than the bare kill it replaces and invisible in the same log. But a BLANKET
  zero-REFUSED gate is wrong too: reap.rs's module doc (:18–20) states refusal is the CORRECT
  happy-path outcome when the test already stopped its daemon (`Refused("gone")` — leak insurance,
  not primary teardown), so the blanket gate reds healthy runs, someone loosens it to go green, and
  the loosened version is exactly the one blind to foreign-image — the detector dies by
  maintenance, wearing a green. The refusal reasons split, and only one class is a defect:
  SOUND (expected): `gone`, `reused`, `breadcrumb-moved`, `self`, `ancestor` — target provably not
  there. TRAP (the false-clean state): `foreign-image`, `unreadable-image`, `unproven-identity` —
  something WAS there and authentication declined; `foreign-image` is precisely what a wrong
  per-pid expected_exe produces on every pid, every run. THE GATE: in the arm that must kill,
  assert no verdict in the trap class — structurally on the returned `Verdict` (Killed |
  Refused(&'static str), stable tokens by design, reap.rs:48), not by scraping output; the
  `REAP[label]: … verdict=` line stays as CI-log forensics. The gate's comment MUST say
  `Refused("gone")` is expected, or the next reader "fixes" it. Positive counterpart that actually
  proves the reap: assert `target_gone(pid, expected_exe, observed)` (reap.rs:215 — no-knowledge
  never read as death) per service pid; bare-liveness "no survivors" is the false-clean named above.
  TOKEN SET VERIFIED COMPLETE-PLUS-ONE (deployah, refuse() sites enumerated at source): the eight
  tokens above sit at :128 self, :132 ancestor, :136 breadcrumb-moved, :146 gone, :149
  unproven-identity, :155 reused, :166 unreadable-image, :172 foreign-image — each in the right
  class — and there is a NINTH neither list held: `Refused("no-breadcrumb")` at :243, returned
  before authenticated_kill is reached. Doyle-ruled 2026-08-03, both halves: (1) `no-breadcrumb`
  classifies SOUND — **SCOPED (deployah source-read, same night): it is two states wearing one
  name.** `breadcrumb_daemon_pid` (reap.rs:59–63) collapses file-absent, permission-denied, any IO
  error, AND a husked/unparseable partial write into one `None` via two `.ok()`s — so the token is
  ABSENCE OF KNOWLEDGE, not proof of absence, and it is SOUND only because the calling test has
  already asserted its daemon stopped. Every other SOUND token carries positive proof; a LIVE
  daemon whose pid file is unreadable or husked emits the identical token and leaks wearing
  "nothing to reap" — the helper's one fail-open path inside a fail-closed design (its own module
  doc: no-knowledge is never read as death), and the same sentence as [[IR-8]]: a zero that cannot
  see is not a zero. CORRECT REMEDY (doyle-ruled, deferred — not this lane, unreachable from this
  test's pids): split at source into `no-breadcrumb` (absent) vs `unreadable-breadcrumb` (TRAP,
  beside unreadable-image) — IR-8's unreadable-bucket fix in a second location; land the pair in
  the census/reap wave IR-8 already names, one remedy for two entries. Precondition on the split
  (hertz): common/reap.rs compiles into EVERY spt integration test binary, so enumerate the callers
  that actually REACH `reap_breadcrumb_daemon` before the token split lands — those are the tests
  whose teardown semantics change; own lane, not a rider. ENUMERATION DONE (deployah, same night,
  read-only): 21 caller tests, every one passing &spt_bin as expected_exe (the per-pid trap is
  specific to resident_service_e2e's svcmock pids — it does not generalize to this population);
  20 callers fire-and-forget the Verdict; the one value-assertion, reaper_guard.rs:192-196, pins
  `Refused("no-breadcrumb")` for an ABSENT breadcrumb only — absent KEEPS the token under the
  split, so that assertion survives unchanged (:183's stale-pid check is the loose
  `matches!(Refused(_))`). The split is ADDITIVE: one source split, one new token, zero caller
  rewrites. THE GAP THAT LET THE COLLAPSE HIDE: reaper_guard covers eight refusal shapes but has
  NO test for an unreadable or malformed daemon.pid — the one state where absence-of-knowledge
  reads as absence-of-target is the one state without a conformance test; (b) ships with exactly
  that evidence (a garbage daemon.pid and an unreadable one, each asserting the new token).
  ASSIGNED: deployah owns the (b) lane (doyle-ruled 2026-08-03), SEQUENCED AFTER hertz's IR-18
  lane — the exhaustive panic-catch-all lands first, so the new token arrives forced-classified
  per this entry's own design, and the (b) lane classifies it TRAP in the same change. SCOPE
  RE-RULED (2026-08-03, on deployah's third-site find): the collapse now has THREE known sites
  (reap.rs:59-63, resident_service_e2e's mock_pid :100-106, and the shape wherever a test reads a
  self-written pid file) — so (b) lands ONE shared read-a-pid-breadcrumb helper that distinguishes
  absent from unreadable ONCE, adopted at the enumerated sites, not a per-site token split; the
  single-source-discriminant rule, applied before a fourth site mints itself.
  **THIRD SITE CONFIRMED AT SOURCE (deployah, read-only 2026-08-04) — it was a predicted SHAPE and
  is now an address:** `crates/spt/tests/resident_service_e2e.rs:136-137`, `brain_ready`, which
  `.ok()?`s BOTH the read and the JSON parse — two collapses in two lines, the self-written-pid-file
  shape this entry predicted. What makes it the worst of the three rather than the third of three:
  its output feeds `brain_seed`, a population-sweep SEED, so an unreadable `brain.ready` silently
  drops a whole subtree from the sweep. The other two sites lose one pid; this one degrades the very
  instrument that is supposed to cover them, and it does so while the sweep still reports clean —
  [[IR-8]]'s sentence again, now inside the instrument rather than beside it. The (b) lane's helper
  therefore has to be adopted at this site to be worth landing at all. Still deployah's,
  still after hertz's lane; hertz's lane does NOT wait for it (the population assertion covers the
  leak class independent of pid bookkeeping — that is its virtue). (b) INTAKE ITEM (doyle gate on
  the IR-18 lane, 2026-08-03, deferred there deliberately): the lane's verdict-gate SOUND-arm
  comment reads "provably not there", which overclaims for `no-breadcrumb` (absence-of-knowledge,
  not proof — the scoped reading belongs at that arm); comment-only, and (b) rewrites that arm
  when `unreadable-breadcrumb` splits out, so it lands there rather than costing a solo respin.
  SECOND (b) INTAKE ITEM (hertz self-noted at re-gate, doyle-ruled deferred): cleanliness is
  currently stated at THREE sites (two asserts + `is_clean`) — the drift that produced the
  scope_lost condition bug (a failed edit followed by a narrower successful one replaced the
  message but not the condition; the artifact of that failure mode IS message/code disagreement).
  (b) consolidates to conditions derived from one source, weighing the two-message diagnostic
  split it would cost. The
  exhaustive panic-catch-all match is what makes the split safe to do later: the new token lands
  loud on first fire in every consuming gate, forced to be classified, never silently passed. Until then the scoped reading above is the
  gate's contract — and the LANE BUILDER writes that scoped reading into the match's SOUND-arm
  comment (the sketch's "nothing to read is nothing to kill" is the exact inference the source
  forbids; the comment must not teach the false step); (2) the gate closes over the
  token set — SPELLED BUILDABLY (hertz correction 2026-08-03; the earlier "no wildcard arm" wording
  is an instruction rustc rejects, since the refusal reason is a `&'static str` (:55) and a string
  match always requires a catch-all): the match names the six SOUND tokens as a pass arm, the three
  TRAP tokens as a panic arm, and the REQUIRED catch-all arm is not a wildcard PASS — it panics
  naming the unknown token ("unclassified reap refusal token — classify it in this match"), so
  upstream drift announces itself the first run it fires instead of sliding into a silent pass.
  Runtime fail-closed, not compile-time: compile-time closure needs the reason to become an enum in
  common/reap.rs, an upstream change touching every consumer — ruled NOT smuggled into this lane
  (candidate for its own thin lane if ever wanted; doyle concurs runtime fail-closed is
  proportionate). Direction of the default, kept in the record: a missed trap state is an invisible
  leak wearing the reap line a reviewer wants to see; an unclassified harmless token is a loud red
  costing one edit. Cheap and loud beats silent and wrong. (Count hygiene, hertz self-flagged: a
  single-line grep for the refuse() sites returns five of nine — three wrap across lines; the nine
  stands on reading, not the grep. The sweep-count trap in miniature.)
- **Prior observed instance of the class, in-tree (hertz 2026-08-03 — stated to understate):**
  reap.rs:3–11 documents the measurement and a casualty: pid reuse on the Windows gate box at
  2.2s minimum / p50 11.7s under a Phase-A battery, and a concurrent test process killed by a bare
  kill on a stale breadcrumb pid, dying with bare exit 1 and no panic — the observed
  `worker_lifecycle_e2e` red in golden job 90742055246 (that test's intermittent class is board
  releases#32, still in-flight). Different test, different job — it does NOT reproduce this entry's
  specimen and promotes nothing; what it settles is that the class is real IN THIS SUITE with a
  named victim, and the authenticated reaper is the remedy someone already built for it —
  resident_service_e2e simply never adopted it. Two shape-predictions it supplies, made before this
  specimen existed: the p50 11.7s reuse window is short against this test's 45s waits and ~124s
  runtime (window wide open). (A second claimed match — "bare-exit-1-no-panic matches the empty
  daemon.stderr.log" — is RETRACTED as a category error, hertz+deployah 2026-08-03: reap.rs's
  phrase describes the VICTIM TEST PROCESS dying, not a daemon's log; two processes, two
  artifacts.) Predictions, not proof; the specimen stays CANDIDATE — shape since reproduced on
  demand, agent unidentified (see the control record above).
- **Rig near-miss, caught before first execution (hertz 2026-08-03) — kept as design justification,
  not an anecdote:** the positive-control arm's first draft selected its victim machine-wide
  (`Win32_Process Name='spt.exe'` + cmdline match `daemon run`, then force-kill). hfenduleam is the
  Windows CI runner as well as a dev box, so that filter matches CI's own test daemons: run while a
  leg was live, it would have force-killed a running job's daemon and reported the result as a
  measurement — fabricating a red in someone else's job, the exact class this entry documents.
  Corrected to a parent-scoped selector (daemon = child of the test pid; service = child of that
  daemon); `CREATE_BREAKAWAY_FROM_JOB` does not weaken parent-id attribution, since breakaway leaves
  the JOB, not the parent record — which is also why the orphaned svcmock stays attributable to the
  daemon that spawned it. **The load bought the catch** (blocked on the box ⇒ re-read instead of
  ran). Bearing on the fix: the INSTRUMENT built to study a bare-selector reaper was itself written
  with a bare selector. The remedy therefore cannot be "be careful" — the authenticated path must be
  the only one within reach in this suite.
- **Ripe when:** RIPE NOW, and not gated on the rate run that was cancelled.
- **Size:** small.

### IR-19 — Docs-only pushes to main run full unit legs (classifier is PR-only) — intent unverified
- **Status:** RETIRED 2026-08-18 — intentional for main pushes, not a classifier defect. · **Origin:** hertz observation 2026-08-03 (register-push
  cadence gated his rig behind repeated hfenduleam unit legs); mechanism verified by doyle at
  source same night.
- **What/why:** `ci.yml`'s `changes` job classifies docs-only diffs ONLY for `pull_request` events
  — push events hardcode `code=true` (ci.yml:45–48, an explicit branch, not a fallthrough), so
  every docs-only push to main spins both full unit legs. Tonight: five register docs pushes each
  queued a Windows unit leg on hfenduleam; concurrency kept one running + one pending (main is
  never cancelled in-progress, pending runs supersede — the recorded behavior matches ci.yml:9–13).
  FINDING SETTLED (deployah, same night, from the workflow): this is not intent — it is a scope
  the classifier never had; the `*.md` / `traceable-reqs.toml` case arms are only ever REACHED on
  pull_request (the early-return precedes them), and golden.yml's own hardcoded `code=true` is
  unrelated (golden triggers only on golden/**). REMEDY STILL OPEN on a named tension: extending
  the classifier to pushes leaves a docs-only main TIP with no run of its own — fine if
  tested-sha==merged-sha means "the code at this sha was tested" (it was, at the last code sha),
  not fine if any gate or reader takes "main tip has a green run" as the check. One person reads
  the consuming gates before any remedy (doyle, at the next CI-touching wave) — three assuming is
  how this class ships.
- **Resolution:** the consuming contract already answers the held question:
  `REQ-CI-DOCS-ONLY-THIN` is PR-only and explicitly says ADR-0050 supersedes
  it for main pushes. Main-tip evidence remains full by design; no classifier
  branch was widened.
- **Interim rule (doyle, same night):** batch register edits into one push instead of landing them
  as they occur — the cadence is a real gate on whoever is queued behind the runner.
- **Ripe when:** next CI-touching wave, after the intent question is answered on the REQ record.
- **Size:** small (one conditional, if the answer is "extend").
- **Composed:** LOCKSMITH (#132) — doyle reads the consuming gates during the batch's CI-touching
  wave (the intent question), BEFORE any classifier remedy. 2026-08-03.

### IR-20 — resident_service_e2e's `spt daemon stop --force` does not stop the daemon; teardown is complete only because the kills are
- **Status:** open · **Origin:** hertz 2026-08-03, measured on the IR-18 lane's own clean run
  (hfenduleam, log kept); routed text, doyle-landed.
- **MEASURED:** teardown runs `spt(&["daemon","stop","--force"])` then reaps. At reaper fire, every
  target was still ALIVE and every verdict was `Killed`, not `gone` — verdicts=[boot Killed,
  rel Killed, brain Killed]; the brain's kill line names it `(child process of PID 46896)` — the
  daemon — so the daemon was still resident too and was taken down by the `Child` handle, not the
  stop. A stop that worked would have left the reaper `gone`. This rules OUT the innocent reading
  (stop doesn't manage services): the stop did not reach the DAEMON either.
- **NOT ESTABLISHED — two rungs, neither discriminated:** (1) the rig seeds a `doyle` perch with
  `std::process::id()`, satisfying `ceremony_agent_ground`'s pid-ancestry rung for every child of
  the harness; (2) the rig never scrubs `OWL_SESSION_ID`/`SPT_AGENT_ID`/`SPT_ENDPOINT_ID` from its
  teardown commands, and the measuring run launched from a live agent session, so the env rung was
  live too. `--force` overrides neither. The stop's output is discarded at the call site
  (`let _ = spt(...)`) — no diagnostic names a denial, and no claim is made that one fired.
  Settling it is one run with the stop's stderr captured — on the CI runner AS WELL AS the dev box,
  since the rungs fire on different machines.
- **Why it matters beyond this rig:** REQ-TEST-RIG-DAEMON-TEARDOWN-PROVEN's own failure shape
  inside a rig that passes — teardown completeness rested entirely on breadcrumb kills that were,
  until IR-18, bare `taskkill /F /T` on unauthenticated pids; the most trustworthy-LOOKING
  component (an explicit `--force` stop) was doing nothing while the least trustworthy one did all
  the work. Exactly why the population gate exists rather than per-pid checks alone.
- **Remedy caution (the sentence that must not drop):** the prescribed fix (seed pid 0, scrub the
  three markers — REQ-TEST-RIG-DAEMON-TEARDOWN-PROVEN's own gate) is NOT free here: this rig's
  `doyle` perch pid is load-bearing for the shell bind-by-token legs, so the two remedies are not
  interchangeable — a lane taking this must check which legs depend on the pid before changing it,
  or it fixes the stop and breaks the REQ-INSTALL-11 legs the rig exists for. Kin:
  REQ-BROKER-STOP-ENDPOINT-DENY (the refusal being worked around).
- **LANE-LINKED 2026-08-30 (#242 close sweep):** hertz's queued daemon-leak fixup lane carries
  this entry (brief cites IR-7/17/20/34/35/63); leaves the register when that lane lands.
- **Ripe when:** a hertz test-hardening wave, or alongside the (b) helper lane's touch of this rig.
- **Size:** small (stderr-captured discrimination run × two machines, then the scoped remedy).

### IR-54 — golden-head-intake runbook: three properties this cut proved should be construction, not discipline
- **Status:** BUILT 2026-08-29 — PR #176 (`880b3b9c`, landed on main via `37368263`).
  · **Origin:** PORTER (#205) v0.59.0 cut, 2026-08-21 — doyle (gater) + deployah (driver),
  both mechanisms measured at the cut, neither cost this run anything because the driver caught
  them by hand. Filed at the release-close register sweep.
- **What/why (three runbook amendments, one lane):**
  1. **Release-shaping at ASSEMBLY intake, not post-golden.** The head arrived code-complete
     but not release-shaped (no version bump, stale CHANGELOG top section) — 5 of the last 6
     cuts. v0.59.0 fixed it by authoring version material BEFORE the golden run, which is what
     produced the four-name shape (ruled sha == main tip == golden ref == tag == `c62904e7`,
     first time in six cuts; the previous five each argued a post-golden delta inert). Amend
     `docs/golden-head-intake` so the gater's assembly checklist asks the release-shaping
     question at intake and the driver never authors on top.
  2. **Never-executed-cells list is a hand-off artifact.** The runbook makes it a
     gater-compiles / driver-re-checks item; PORTER's handoff omitted it and the driver had to
     ask. Name it in the assembly checklist beside the greenlit-form delta: which new cells'
     first CI execution is this golden, and where each HAS executed (lane gate / assembled
     head) — it converts a golden red from an RCA into a lookup.
  3. **Lockfile-refresh rule states the PROPERTY, not the command.** Step 1 prescribes
     `cargo metadata --offline`; the driver used `cargo update --workspace --offline` and
     proved the actual rule by diff (workspace members only, third-party pins untouched —
     14 first-party line pairs, 90 third-party unchanged). The property is the rule and the
     diff is its evidence (the standing by-diff rule exists because counting misread v0.39.4
     and v0.41.0 in opposite directions — and the `windows-sys` 0.59.0 collision was LIVE at
     this cut, the exact condition where counting fabricates a 15th "first-party" bump);
     demote the command to an example vehicle.
- **Built evidence:** `docs/RELEASE-RUNBOOK.md` now asks the release-shaping question at assembly
  intake, makes the never-executed-cells list part of the hand-off itself, and states lock refresh
  as a diff-proven property with commands demoted to example vehicles. PR #176 passed
  `traceable-reqs check`; the v0.65.0 cut had already exercised the release-shaped-head rule.
- **Size:** small (one runbook doc lane; no code).

## IN-FLIGHT ON THE BOARD (not re-homed — live WIP lanes)

- releases#93 (servicehost swap test pid-inequality), releases#47 (thin-lane unit(Windows)
  isolation class), releases#32 (worker_lifecycle_e2e Phase A intermittent) — infra-shaped but WIP;
  a live lane is never re-homed mid-flight. When each closes, any residue lands here as a new entry.

## BUILT / RETIRED

### IR-15 — provably_gone is pid-only on Windows: teardown tests can fabricate reds
- **Status:** BUILT 2026-08-03 · **Origin:** golden 30776330383 red on the #125 fix lane (2026-08-02)
- **Lane:** hertz `golden/ir15` @a5042ec — REQ-TEST-LIVENESS-ORACLE-AUTHENTICATED (impl+unit): one
  shared authenticated death-oracle test helper backed by `spt_procident::process_identity`
  (identity pinned at find time, re-verified at assert), both pid-only polling sites repinned
  (daemon.rs teardown test; endpoint_lifecycle.rs relay_pid). Ruled polarity held: Absent⇒gone,
  Present(same start)⇒not gone, Present(different)⇒gone, Unproven⇒NOT gone, missing-stamp
  degrades to pid-only loudly — every unknown errs toward the recoverable red, zero new
  false-green paths. Linux caveat documented in the helper (10ms jiffies: Present(different) is
  narrowing, not decisive; same-tick reuse errs red, safe).
- **Golden:** run 30782259675 red on `resident_service_e2e` bring-up — delta exonerated (unrelated
  family, now [[IR-17]]; same job minted [[IR-8]]'s defining specimen). Same-sha rerun 30784469908
  GREEN; main ff'd to a5042ec.
- **A/B discriminator (hertz):** 80/80 green both arms, quiet + churn phases; churn arm
  positive-controlled (pid-allocator wrap observed <400ms, so reuse pressure was REAL in the churn
  arm); rule-of-three bounds the original flake at ~7.5%/run/sha. Specimen 30776330383 remains
  not-reproduced, cause unidentified — the repin removes the false-red MECHANISM; it does not
  adjudicate the specimen. [[IR-16]]'s kill-side arm (TerminateJobObject return discarded) stays
  a live discriminating instrument for any recurrence — and since 2026-08-03 it is no longer the
  only one: [[IR-18]] (the test's OWN bare breadcrumb tree-kill) is the second kill-side candidate,
  test-side rather than product-side. A recurrence must discriminate between them, not assume
  either: IR-16's arm is a job whose termination silently half-fails (survivor still holds the
  captured image); IR-18's is a victim killed by a bare pid it no longer owns —
  `broker_survived=false` with the boot service still ALIVE and its message spooled. (An earlier
  spelling of IR-18's signature here said "empty stderr"; retracted — empty stderr is present on
  healthy runs and discriminates nothing.)

### IR-21 — CLASS: a helper binary's build is never requested, only its LOCATION is, so any narrow invocation manufactures a red that belongs to the rig
- **Status:** BUILT 2026-08-18 on KEYSTONE #182 hygiene lane. · **Origin:** todlando 2026-08-03. Filed as one instance, rewritten as a CLASS the
  same day when a second member appeared, rewritten AGAIN when the mechanism was measured — the
  first two versions described symptoms and got the remedy wrong. **Corrected a fourth time the same
  day**: remedy (1) claimed the 11 same-package sites were missing a build edge, and todlando
  falsified that by measurement while executing it (replicated independently by doyle). The error was
  doyle's to carry — it was ruled on, not just written. The general rule that refutes it was already
  in this entry's own CI-member section; diagnosing a class does not inoculate you against drawing
  its opposite consequence one section later.
- **Numbers below re-derived on `main` @`3efd7e6` over `git ls-files` (tracked files only). Each
  carries its command; re-run rather than cite.**
- ⚠ **Scan-root hazard, measured 2026-08-03 and worse than first reported:** a naive recursive
  `grep -r` from the project root inflates by **~35×**, not the ~5× an earlier draft of this entry
  claimed — 426 tracked `.rs` against 14,865 on disk with `target/` excluded, because `.worktrees/`
  holds full copies of the tree. **And the unscoped grep does not COMPLETE** (still running at a
  120s timeout while `git ls-files` returns instantly), so the failure mode is not only a wrong
  number but a command that reads as hung on a tree where dozens of worktrees are normal. Use
  `git grep` or `git ls-files | xargs grep`; both are scoped to tracked files by construction.

#### THE MECHANISM (measured, both ends)

**Build end — `cargo test` with a narrow target selector compiles a `[[bin]]` AS A TEST HARNESS and
never emits the plain executable.** Measured: `cargo test -p spt --bins` on a clean pool produced
five executables, every one hash-suffixed under `target/<profile>/deps/`, and ZERO plain exes in
`target/<profile>/`. The bin was built. The file the test looks for was not written.

**Consumer end — the resolver borrows a guaranteed binary's path to locate an unguaranteed one.**
All 29 copies of `sibling_bin` in `crates/spt/tests/` reduce to two textual variants (22 + 7) of one
body:

```rust
fn sibling_bin(name: &str) -> PathBuf {
    PathBuf::from(env!("CARGO_BIN_EXE_spt"))
        .with_file_name(format!("{name}{}", std::env::consts::EXE_SUFFIX))
}
```

`CARGO_BIN_EXE_spt` is used **only as a directory anchor**. Cargo therefore sees a dependency on
`spt` and on nothing else; the `{name}` half is a string join it cannot observe. The dependency edge
does not exist at any level a build system could act on — which is why "use the wider command" is
not the fix and why the failure is invisible on a warm pool where some earlier run happened to leave
the file behind.

The unit-test resolver at `crates/spt/src/cli.rs:25601` reaches the SAME directory from a different
anchor (`current_exe()` → `deps/` → parent) for the same reason: its own doc comment records that
unit tests get no `CARGO_BIN_EXE_*` at all.

**That mechanism explains both original members at once, including the polarity inversion that made
them look like different bugs:** whether `cargo test -p spt` or `cargo test -p spt --bins` happens
to leave a plain exe behind is incidental to a dependency neither command was told about.

#### SCOPE — WHERE IT BITES, AND WHERE IT DOES NOT

**It bites NARROW invocations — `-p <pkg>`, `--bins`, `--lib` — i.e. gate rigs and local runs.**
Both members were hit on the LOCKSMITH t1 lane against fresh throwaway pools; each was discharged
in-lane as a rig artifact. That is the whole field population to date.

**GOLDEN IS ACQUITTED, and the acquittal is measured rather than assumed.** `cargo test --workspace
--no-run` emits all 13 plain binaries — mock-session, mock-shell, capture-player,
console-mode-probe, service_fixture and the rest — so a workspace build produces the cross-package
helpers by construction. Confirmed under nextest, the tool CI actually runs, with a real compile.

⚠ **The measurement that establishes this is not the obvious one.** Deleting the plain exe and
watching it reappear is NOT evidence of a rebuild: cargo's uplift restores a hardlink from an intact
`deps/` artifact, with link count 3 and an UNCHANGED mtime, which looks exactly like a build and
says nothing about a pool that never had the file — and a pool that never had it is the only kind a
rig runs on. The sound probe removes the plain exe AND both `deps/` artifacts, confirms all three
absent, then watches the target actually recompile (mock-adapter, 5.77s, fresh mtime). The first two
attempts were setup, not measurement.

⚠ **The suspicion this acquits is structurally well-founded, so record the answer, not just the
verdict.** `twohost-a`/`twohost-b` are `needs: test`, so the only mock-adapter prebuild runs AFTER
the phases that use it; `target/` is gitignored, so checkout never cleans it on a persistent
self-hosted workdir. Every precondition for artifact leakage is present — the build simply does not
need it. The next person who reads that job graph will form the same hypothesis; this is the answer
waiting for them.

#### THE ONE CI-SIDE MEMBER THAT IS REAL

`crates/spt/src/cli.rs::adapter_translate_proof_gates_on_commit` is a **unit** test (`kind(bin)`),
and `CARGO_BIN_EXE_*` is set only for integration tests and benches of the declaring package. A unit
consumer therefore CANNOT express the dependency through the env var at all — it has no choice but
to resolve by path. That is why `.github/workflows/ci.yml:105-106` carries a hand-written
`cargo build -p spt --bin translate_proof_fixture` before the unit lane, with a comment
(`ci.yml:102-104`) stating this exact mechanism.

**Deleting that step would red the unit lane on a clean pool, and it would read as a code red.**
Somebody already hit this class, fixed their own leg correctly, and never turned it into a rule —
the comment at `ci.yml:102` is IR-21 written a milestone early, in the one place only its author
would find it.

#### POPULATION

**13 bin targets in the workspace** (`cargo metadata --no-deps`, not a grep — an auto-target under
`src/bin/` and `xtask` are both invisible to a `[[bin]]` grep). Two are not helpers (`spt`, `xtask`);
the other **11 are test helpers**, across three packages:

| package | helpers |
|---|---|
| `adapters/mock` | mock-session, mock-shell, capture-player, console-mode-probe |
| `crates/spt-daemon` | dispatch_fixture, service_fixture, summarizer_fixture, xlate_choreo_fixture (⚠ this row was mis-"corrected" 2026-08-22 to claim `xlate_choreo_fixture` no longer existed — IT DOES, as an AUTODISCOVERED `src/bin/` target with no `[[bin]]` stanza; the original row was right and merely predated `summarizer_fixture`. See IR-58) |
| `crates/spt` | translate_proof_fixture, post_step_fixture, gh_fixture, git_fixture |

**44 literal `sibling_bin("…")` call sites**, all in `crates/spt/tests/`, served by **29 copied
resolvers**. Split by whether the BUILD IS ALREADY GUARANTEED — which is not the same question as
whether the env var is available, and an earlier version of this entry conflated the two:

| class | sites | detail | build guaranteed? |
|---|---|---|---|
| **cross-package** | **33** | mock-session 26, mock-shell 6, service_fixture 1 | **NO — the hazard members** |
| **same-package** | **11** | translate_proof_fixture 7, git_fixture 2, post_step_fixture 1, gh_fixture 1 | **YES — already, by construction** |

⚠ **MEASURED TWICE, and it falsifies what this entry said on 2026-08-03 before this revision:** cargo
builds **every bin target of a package whenever it builds ANY integration test of that package** —
that same act is what sets `CARGO_BIN_EXE_*` in the first place. So the 11 same-package sites were
never unexpressed in a way that could bite, and the hazard population is **33 cross-package sites
plus the one unit-test member = 34**, not 44.

- **Probe 1 (todlando, root pool):** deleted all four `translate_proof_fixture` artifacts (both
  hash-suffixed harness exes, the `deps/` plain exe, the uplifted plain exe), confirmed absent, then
  built ONE UNRELATED and UNMODIFIED integration test of the same package —
  `cargo test -p spt --test attach_wedge_e2e --no-run`, exit 0, 14.45s. Both plain exes returned with
  fresh mtimes; the hash-suffixed harness exes stayed absent, which is what distinguishes a bin
  DEPENDENCY build from a `--bins` harness build.
- **Probe 2 (doyle, `.worktrees/gate-ec5f38a` pool, independent replication with a different fixture
  and a different probe test):** deleted `gh_fixture`'s plain exe, its `deps/` plain exe AND its
  hash-suffixed harness exe, confirmed all three absent, then built `--test json_emit --no-run`
  (exit 0) — a test that never names `gh_fixture`. The plain pair returned at a FRESH mtime (09:13
  against the 09:00 it carried before), so this is a real build and not the hardlink uplift this
  entry warns about elsewhere; the harness exe stayed absent.
- **Neither probe needs a baseline arm:** the probe test is unmodified and references nothing under
  edit, so what it measures is cargo's behaviour, not anyone's change.
- **Corroborated by this entry's own field data:** neither original member was a same-package
  integration site — member 1 is a UNIT test, member 2 is CROSS-package. The class never had a
  same-package integration member, and the CI-member section below already stated the governing rule
  (`CARGO_BIN_EXE_*` is set only for integration tests and benches of the declaring package) one
  section before the remedy drew the opposite consequence from it.

Command: `git ls-files '*.rs' | xargs grep -hon 'sibling_bin("[a-z_-]*"' | sed 's/.*sibling_bin("//;
s/"//' | sort | uniq -c`.

⚠ An earlier report gave 48 sites and a 6/11/37 split. Take the table above: it counts only literal
call sites in tracked files and it ships its command. The 11 is the same 11 in both counts.

**The LOCATION half of this class is already closed** by `crates/spt-term/tests/support/fixture_bin.rs`
— a shared resolver rather than 29 copies. It does not close the BUILD half, and should not be
mistaken for having done so.

#### REMEDY SHAPE (not ruled)

The distinction that matters is location vs. build:

1. **The 11 same-package sites are NOT hazard members and need no build fix.** Their build edge
   already exists (see the two probes in POPULATION above); `env!("CARGO_BIN_EXE_<name>")` for the
   fixture itself would add nothing to it. Converting them is a **CLARITY** change, worth doing on
   its own smaller merits — the path becomes the one cargo actually emitted rather than a string-join
   of a directory anchor and `EXE_SUFFIX`, five copies of `sibling_bin` stop existing (29 → 24), and
   it completes a migration already paid for: `crates/spt/tests/fixtures/translate_proof_fixture.rs:7-12`
   records that the fixture was re-homed into `spt` precisely to obtain
   `CARGO_BIN_EXE_translate_proof_fixture`, and then all 7 call sites resolved by path anyway. **It is
   not a build fix and must not be filed as one.**
2. **The 33 cross-package sites cannot**, by cargo's design. Their options are an asserted build in
   the test's own setup, a `dev-dependencies` artifact dependency, or an explicit documented
   prebuild — the ci.yml:106 shape, made a rule instead of a local fix.
3. **Standardising on one wider invocation is NOT a remedy.** It changes which pools happen to work;
   it does not create the dependency edge, and the two original members had opposite polarity under
   exactly that theory.
4. Collapsing the 29 resolver copies is worth doing on the `fixture_bin.rs` model, but on its own it
   makes the class HARDER to see — one shared resolver still anchored on `CARGO_BIN_EXE_spt` hides
   the 33 genuinely unexpressed dependencies among its 44 call sites behind one function.

#### SUPERSEDED FRAMING, KEPT SO IT IS NOT RE-DERIVED
- **Status of the original filing:** open · **Origin:** todlando 2026-08-03, both members hit on the
  LOCKSMITH t1 lane against fresh throwaway pools; each discharged in-lane as a rig artifact, filed
  here so the next clean rig does not re-diagnose them as code reds.
- The first two versions of this entry framed the class as "the rig's command does not build what
  the test needs" and proposed standardising on a wider invocation. **Both are superseded by the
  measured mechanism above** — the dependency is not under-expressed, it is INEXPRESSIBLE in the
  form these call sites use, so no choice of invocation creates it. Kept only as the two FIELD
  MEASUREMENTS that produced the class, which remain true:
- **MEMBER 1 — MEASURED:** `cli::tests::adapter_translate_proof_gates_on_commit` failed on the
  first `cargo test -p spt --bins` run. Its fixture binary `translate_proof_fixture` (a
  `tests/`-homed `[[bin]]`) was ABSENT from the pool — `ls` on the path returned No such file.
  Building it explicitly and re-running the IDENTICAL command PASSED, after which the full `--bins`
  suite passed 584/585 with only the releases#117 probe red (that one RED by design). So the
  discriminator is the fixture's presence, not the tree: same command, same sha, red then green
  across one `cargo build` of the fixture.
- **MEMBER 2 — MEASURED:** `cargo test -p spt` does not build `mock-adapter --bin mock-session`, so
  `attach_wedge_e2e` panics `"the dummy-harness program must be built"`. Note the polarity is
  INVERTED against member 1 — there the narrower `--bins` was the defective invocation and
  `cargo test -p spt` the correct one; here `cargo test -p spt` is itself insufficient. So the
  class is NOT "use the wider command"; it is that the dependency is not expressed to the build at
  all, and which invocation happens to work is incidental.
- **What/why (still true):** on a WARM pool the helper is already there from some earlier run and
  the test passes, so the defect is invisible exactly where most people work and fires only on a
  clean pool — i.e. on a GATE RIG, which is the one place a false red costs the most.
- **The sweep this entry once called its first step HAS BEEN RUN** (todlando 2026-08-03) and its
  result is the POPULATION section above. It is no longer outstanding.

- **Why it is register debt and not a lane bug:** nothing in the product is wrong. The gap is
  between what a test needs built and what the rig's command builds, and the fix belongs to the
  tests' declarations, not to whoever is running a gate that day. Sibling rule:
  [[gate-clean-target-not-incremental]].
- **Built evidence:** golden now names every cross-package fixture prebuild
  (`mock-session`, `mock-shell`, `capture-player`, `console-mode-probe`,
  `service_fixture`) explicitly; `xtask binedge-check` reports zero missing
  edges. All 29 local `sibling_bin` resolvers collapsed into
  `tests/common/mod.rs`; the shared precondition names the missing fixture and
  its package-correct build command before any product timeout.
- **Ripe when:** the next gate-rig or CI-touching wave. Cheap, and it pays for itself the first time
  it stops someone chasing a phantom red on a clean pool.
- **Size:** remedy (1) is small AND optional — it buys clarity and subtraction, never a build edge;
  medium for (2), which is a design call before it is an edit and is the only remedy that closes the
  class.

### IR-22 — An inherited identity env var fails a test, and the diagnostic names them ONE AT A TIME so a correct fix reads as no fix
- **Status:** open · **Origin:** todlando 2026-08-03, chasing what looked like an `attach_wedge_e2e`
  code red on the LOCKSMITH t1 lane; root-caused to the runner's own process environment.
- **MEASURED:** `attach_wedge_e2e` failed for an INHERITED PROCESS-GLOBAL and nothing in the tree:
  the daemon-stop refusal fired on the running session's own identity env. It named
  `$OWL_SESSION_ID`; clearing that made it name `$SPT_ENDPOINT_ID`. With `OWL_SESSION_ID`,
  `SPT_ENDPOINT_ID`, `SPT_AGENT_ID` and `SPT_SESSION_ID` all cleared: exit 0, 1 passed.
- **The finding is the DIAGNOSTIC SHAPE, not the env hygiene.** Naming one variable at a time means
  a correct partial fix produces an identical-looking failure, so the natural reading of "I cleared
  it and it still fails" is that the clearing did not work — when in fact each step was right and
  the message had simply moved on to the next name. A refusal that can only ever name one member of
  a set it is checking teaches the person debugging it the wrong lesson. Compare the same class in
  [[IR-15]]/[[IR-18]] terms: the instrument is competent and the report is not.
- **Why it matters beyond one test:** any agent running suites from a live spt session carries these
  vars, so this fires for every builder on a perched session and for nobody running from a bare
  shell — which is precisely the split between how builders work and how CI runs.
- **Candidate remedies (not ruled):** have the refusal name EVERY identity var it found set, in one
  line, rather than the first; and/or have the affected tests clear the identity set in their own
  setup so a perched session is not a special environment. The first is the one that pays off
  outside this test.
- **Ripe when:** next CI/test-hygiene wave. **Size:** small.

### IR-23 — `endpoint_teardown_authority_e2e`'s two tests collide with EACH OTHER through the machine-global spt home
- **Status:** BUILT 2026-08-18 on KEYSTONE #182 hygiene lane. · **Origin:** todlando 2026-08-03, LOCKSMITH t1 lane.
- **MEASURED:** both tests copy a psyche binary fixture into `perch::spt_home()/srcs/dummyharness/`
  (`crates/spt/tests/endpoint_teardown_authority_e2e.rs:399`), which is process-global, so run in
  parallel inside one binary each holds the file the other wants: `os error 32`, "used by another
  process". **Signature: WHICH of the two fails alternates between runs.** Under
  `--test-threads=1`: 2 passed, exit 0 — and the wall clock drops from 182s to 14s.
- **What/why:** two defects in one, and they should not be conflated. (1) The tests are not
  isolated from each other. (2) The staging path is the MACHINE-GLOBAL spt home rather than the
  test's own temp home — on a box that is also a CI runner, so the blast radius is not confined to
  the suite. (2) is the one worth fixing; (1) is a symptom of it.
- **Note on why this is not already caught:** the standing rule is that integration tests run
  through nextest, which gives each test its own process and hides the collision entirely. So this
  is latent on the sanctioned path and only bites the bare `cargo test` path — the rule is working
  and masking a real defect at the same time, which is why the entry exists rather than a shrug.
- **Candidate remedy (not ruled):** stage the fixture into the test's own temp home. Sweep for
  siblings first — any other test writing under `perch::spt_home()` rather than a temp home shares
  the shape, and nobody has counted them.
- **Built evidence:** adapter source, manifest registration, and psyche fixture
  staging now use the rig TempDir. The two process-global resolver users are
  explicitly serialized, and the bare `cargo test` path passes both cells with
  `--test-threads=2` in 13.82s after the declared fixture prebuild.
- **Ripe when:** next test-hygiene wave; the 182s -> 14s figure makes it pay for itself on the
  bare path. **Size:** small per test, unknown until the sweep.

### IR-24 — `reap::terminate_job` discards `TerminateJobObject`'s return, alone among its own module's siblings
- **Status:** BUILT — landed `3d12cf3` on main (2026-08-03), fulfilment verified against this
  entry's own wanted list by doyle 2026-08-04, not merely mapped to the commit: the return is
  audited and the loss NAMED (`REAP_JOB_TERMINATE_FAIL: brain subtree NOT reaped (<os error>); …
  the daemon holds no other handle to reach them` — which also ANSWERS the cost-of-loss question
  this entry said nobody had written down); the "first thing to check" was checked and the ruling
  came out OPPOSITE to IR-16's site in one respect, documented as measured asymmetry in the doc
  comment: `TerminateJobObject` has NO ordinary failing path (zero only for a broken handle) so a
  zero is a real loss, while `TerminateProcess` fails `ERROR_ACCESS_DENIED` on the ordinary
  already-exited path and is deliberately NOT checked — the premise is unit-PINNED both directions
  (`terminate_job_return_discriminates_broken_handle_from_ordinary_states`, cfg(windows), with a
  memberless-job false-positive arm). No `job != 0` re-check (second-weaker-rule prohibition
  honored, same as IR-16). Golden evidence: rides `4b37512` (run 30873007187 attempt 5, 9/9
  green; the premise unit observed PASS in-pool on the Windows leg, Linux leg green — the
  cross-platform quiet-`ESRCH` concern stays with `ec5f38a`'s coverage as before). Path
  correction to this entry's MEASURED line: the site is `crates/spt-daemon/src/reap.rs` (brain
  subtree reaper), not `crates/spt/src/reap.rs` as first filed ·
  **Origin:** todlando 2026-08-03, swept while building [[IR-16]]; deliberately
  NOT folded into that lane (doyle ruling, same day) and filed instead.
- **MEASURED:** `crates/spt/src/reap.rs:207-211` calls `TerminateJobObject` and discards the result,
  the identical shape [[IR-16]] closed at `spt-daemon`'s `kill_tree`. Its own module's
  create/assign siblings at `:178` and `:202` ARE instrumented, so this call is the odd one out
  where it lives — the module already decided that these returns are worth reading.
- **Why it was NOT folded into IR-16:** different subject. IR-16's site is a SUPERVISED SERVICE's
  teardown, where the loss is "a supervised service's descendants survive" and the promise it
  breaks is `REQ-RESIDENT-SERVICE`'s tree claim. This site is the BRAIN SUBTREE's kill-on-close
  job, which has its own lifetime, its own caller and its own answer to "what does a failed
  termination cost here" — and that answer has not been written down by anyone. Folding it in would
  have meant settling that question in passing, inside a commit about the supervisor, which is how a
  second ruling gets smuggled into a lane scoped to one.
- **What it needs that IR-16's fix does not supply:** IR-16 ruled a LOUD REFUSAL (option B) on the
  ground that the alternative — a process-table tree-walk fallback — is measured blind in exactly
  that failure state ([[IR-18]]: descent breaks at the dead middle hop) and would trade a silent
  loss for a false clean. Whether the same reasoning holds here depends on whether this site kills
  its direct process in the same breath, which is what makes the middle hop dead by construction
  there. **That is the first thing to check, and it is not assumed.**
- **Ripe when:** any wave touching reap/teardown; it inherits IR-16's vocabulary and its
  discrimination note, so the second one is cheaper than the first. **Size:** small, once the
  cost-of-loss question is answered for this subject.

### IR-25 — `spt-daemon --lib` reds 2-in-3 under concurrent load with bare `cargo test`, and holds green under nextest
- **Status:** open · **Origin:** doyle 2026-08-03, found while gating LOCKSMITH tranche 1 in an
  isolated worktree; chased to a mechanism and scoped OUT of the release path before the head was
  assembled.
- **MEASURED, four arms, same box, same hour, matched load** (a second `cargo test -p spt-store
  --lib` loop running throughout; `spt-store` itself stayed green in every round, so the box was not
  generically failing):

  | tree | instrument | result |
  |---|---|---|
  | lane `ec5f38a` | `cargo test --lib`, at rest | 821/821 ×4 |
  | lane `ec5f38a` | `cargo test --lib`, under load | **2 of 3 RED** (142s, 117s) |
  | lane `ec5f38a` minus its 5 new tests | `cargo test --lib`, under load | 3 of 3 green (117s, 84s, 105s) |
  | main `932e14b` | `cargo test --lib`, under load | 3 of 3 green (97s, 119s, 104s) |
  | lane `ec5f38a` | **`cargo nextest run`, under load** | **3 of 3 green, 821/821** (88s, 85s, 72s) |

- **Three DIFFERENT victims across two loaded runs**, all process/timing-shaped, all PRESENT AT MAIN
  and green there: `servicehost::…a_service_that_ignores_the_stop_marker_is_force_killed_and_confirmed_dead`
  (panics `ForceKilled must MEAN the owned child is gone` right after `SERVICE_KILL_UNCONFIRMED`),
  `applyhost::…broker_reports_its_compiled_image_version_over_ipc`, and
  `daemon::…a_tree_teardown_reaches_a_grandchild_the_service_spawned`.
- **MECHANISM: in-process thread interference, not product code.** Bare `cargo test` runs a crate's
  tests as threads in ONE process. The lane's five new tests are process-spawning kill-tree tests;
  under load they starve confirm windows of NEIGHBOURING tests that were always marginal. Removing
  exactly those five flips 2-of-3-red to 0-of-3-green at the same sha, in the same binary, at the
  same durations — one of the green runs sits at 117s, precisely a duration that had produced two
  failures with them present.
- **IT DOES NOT REACH GOLDEN.** nextest gives every test its own process, and golden runs nextest.
  The population is GATE RIGS AND LOCAL RUNS — the same population as [[IR-21]], reached by a
  different mechanism, which is why these are two entries and not one.
- ⚠ **The instrument switch is the whole finding.** Measured with bare `cargo test`, this reads as a
  lane blocker; measured with the tool CI actually runs, it is a rig-scoped nuisance. Anyone
  re-opening this must state WHICH instrument produced their observation before quoting a verdict.
  The same tool gap acquitted [[IR-21]]'s golden question the same day.
- ⚠ **The pre-existing fragility is NOT thereby closed.** Those three tests were marginal before this
  lane and the lane only exposed them; "make the new tests quieter" would re-hide a real weakness.
  Hardening them is test-side work (hertz's lane by the dispatch split), not the builder's.
- **Two rig defects of the gater's, recorded because they nearly cost the verdict:** (1) a first
  re-run at rest came back 4/4 green and proves NOTHING — a sequential idle probe cannot express a
  load-sensitive failure, so that is a competence-controlled zero, not an acquittal; (2) the first
  baseline attempt ran its two arms UNSYNCHRONISED (the load generator finished while the measured
  arm was still on run 1), which would have compared a loaded lane against an idle main and read as
  "the lane broke it". A discriminator whose arms ran under different conditions cannot discriminate.
- **Ripe when:** alongside [[IR-21]] on the next gate-rig or CI-touching wave, or immediately if
  anyone starts gating on bare `cargo test` on a loaded box. **Size:** small to document the rig
  rule (gate with nextest); medium to harden the three marginal tests.

### IR-26 — POOL-OWNER claims authorize takeover from a DEAD holder, so releases#103's hazard reaches through the guard
- **Status:** open, LANE BUILT BUT UNCOMMITTED (see the warning below — this is not a figure of
  speech) · **Origin:** hertz 2026-08-03/04, found by its own test rather than by reading, and
  reported unfiled for composition. Composition ruling is doyle's, 2026-08-04.
- **What/why:** a POOL-OWNER claim named a holder pid; lane liveness was read from whether that pid
  was alive. hertz MEASURED the failure on this box: **3 of 3 claims had dead holders while one of
  those lanes was live** — so the guard authorized exactly the interleaving takeover releases#103
  exists to prevent. The remedy records the lane's GIT IDENTITY (branch + the base sha it carried at
  claim time) and reads lane state from ancestry: in flight while the tip is not contained in
  `origin/main`; settled once merged, once the branch is gone, or once the branch no longer carries
  the claimed base (the branch-level twin of a recycled pid). Holder pid demotes to advisory — a
  LIVE holder still refuses, a dead one no longer authorizes.
- **The strongest finding in it, and it is a filter-competence one:** without the base anchor, a
  missing branch read as "settled, branch gone" **out of a directory that was not a repository at
  all**. An unanchored Settled verdict is a clean zero produced by a filter that cannot express the
  question. A Settled verdict now requires the base object to be present in the reading repo;
  unanchored is UNKNOWN. Sibling of [[zero-match-filter-reads-as-absent]] in the pool domain.
- ⚠ **POLARITY, stated rather than buried:** the new arm (dead holder + unlanded lane => REFUSE)
  converts a false-TAKEOVER into a false-REFUSAL. That is the better failure but not a free one, and
  its triggering case is not exotic — it is an agent going down mid-lane, which happened to hertz
  itself this same batch and cost [[IR-1]]/[[IR-4]]/[[IR-9]] their landing. **The open gating
  question, put to hertz 2026-08-04 and NOT yet answered:** from the refusal state, does
  `pool-release` still clear a claim whose holder is dead and whose lane is unlanded, or does the new
  arm refuse that too? A guard whose only escape is the blanket `SPT_POOL_UNCHECKED=1` trains people
  onto the blanket override. The refusal must NAME the remedy, not only the witness, and the remedy
  must be exercised FROM the refusal state rather than reasoned about ([[remedy-must-run-from-refusal-state]]).
  If the answer is "it refuses", that is a design change and returns to doyle before the lane lands.
- **COMMITTED 2026-08-04 (was uncommitted when first measured):** `ci/poolowner-lane-claim` =
  `8ce40d9` (the change) + `98cfafe` (the remedy fix below), both off `b7b00c3`. The original
  measurement found ZERO commits and six modified files in `.worktrees/hertz-poolowner`, i.e. every
  reported number had been measured against a tree no git object captured; the builder committed on
  being told. Recorded because the near-loss is the lesson, not the correction.
- **THE REMEDY DID NOT RUN — measured from the refusal state, which is the only place the question
  can be asked.** The refusal named its witness AND a command, and executing that command verbatim
  failed with the IDENTICAL `SPT_POOL_FOREIGN` block that printed it: **xtask depends on spt-store,
  so every command that builds the tool goes through the pool being refused.** This is the same dead
  end the unclaimed-foreign arm already hit in the field and fixed by leading its line with the
  hatch — adding a second refusal arm re-broke it immediately. FIXED at `98cfafe`: the in-flight
  remedy now leads with the hatch and names `pool-claim`, the command that actually resolves
  ownership. Full chain re-measured on hfenduleam: refuse → verbatim remedy fails → hatch-prefixed
  form succeeds → build proceeds.
- **Answer to "does `pool-release` clear a dead-holder + unlanded claim", two-part:** it DOES — it
  never consults the verdict, it rewrites `lane=None` — **but it is NOT the exit**, because after
  release the same build is refused again on the pre-existing unclaimed-foreign arm ("no lane has
  claimed this pool"). Release moves you from one refusal to another. `pool-claim` is the exit.
- **The guard is a POPULATION, not a row** (builder's generalization, adopted by doyle as the
  standard for this class): one test poses all three refusal arms and asserts that any remedy naming
  a command leads with the hatch, PLUS membership — three posed, three refused — so no arm passes by
  never being reached. The rule had been learned once already and the next arm added broke it, which
  is the definition of something that needs a population test rather than a better comment. 26/26.
- ⚠⚠ **FLEET PROPERTY — the rig refuted itself, and this is the entry's most important line.** The
  first arm did NOT refuse, it TOOK OVER: the arriving worktree was at `b7b00c3`, so the build script
  that ran was ITS OWN, pre-identity. **Enforcement code comes from the ARRIVING tree, never from
  the claim.** So until this lands everywhere, an old tree still takes a live lane's pool, and the
  guard is only ever as new as the tree that arrives. Found by measurement; reasoning would have
  reported a refusal the rig never produced.
- **RULING on the fleet property (doyle 2026-08-04) — a stamp VERSION LINE is REFUSED, and not on
  taste: it cannot work.** The tree that must refuse is the OLD one, and an old build script does not
  read a new stamp field because it does not know the field exists. You cannot make a stale enforcer
  honor a new rule by writing something new into the record it enforces against — **the reader is
  the problem, not the record.** A version line would help only a CURRENT tree explain itself while
  doing nothing to the population it is aimed at, i.e. it would retire the worry without retiring the
  risk. Ruled instead, in priority order: **(1)** land it and state the migration window honestly —
  the arm protects a lane only when the ARRIVING tree carries the change, coverage grows as trees
  turn over, never retroactive; **(2)** NAME THE LOSS YOU CANNOT PREVENT ([[IR-16]]'s pattern) — the
  claiming side may be able to notice afterwards that its pool was taken and say so loudly, turning
  silent artifact corruption into a named event; scope separately and MEASURE whether the claiming
  side can observe it at all before committing; **(3)** SHRINK THE POPULATION — old enforcers are old
  trees, this box carries ~45 worktrees and most are stale, so reaping them is real mitigation and
  cheaper than code. (2) and (3) do NOT ride this lane.
- **The unanchored row, as asked:** `a_claim_from_another_repository_reads_unknown_not_settled` was
  RED before the base-object requirement existed. Verbatim: `left: Settled { reason: "lane branch
  build/lane-x no longer exists" }  right: Unknown`, out of a tempdir that was not a repository at
  all. It fails without the requirement and passes with it — a real negative control, not a row that
  only ever passed.
- **Builder's evidence (as reported, against that uncommitted tree):** spt-poolguard 25/25 with 7 new
  rows, one per precedence cell, two carrying their own negative control; pool_guard_canary 3/3;
  `cargo check -p xtask -p spt-store --all-targets` clean; clippy clean on the three touched crates;
  `traceable-reqs check` exit 0 at 745/745 with new `REQ-POOL-LANE-IDENTITY` complete at doc+impl+unit.
- **Composed:** its OWN thin lane, riding the same next golden batch as the CI-rider cluster but NOT
  merged into it (doyle ruling 2026-08-04). A red must stay attributable: the riders change what the
  pipeline measures, this changes a build-time REFUSAL, and a guard whose failure mode is refusing
  builds is the worst thing to make inseparable from a pipeline change. IR-14's four-trees-into-one-
  unclaimed-pool finding stays OUT of this lane and stays IR-14's, on the same ground [[IR-24]] was
  kept out of [[IR-16]] — folding it in settles its question in passing.
- **Ripe when:** next CI/infra batch, gated on the polarity question above being answered first.
  **Size:** small once the remedy question is settled.
- **FALSE-LIVE, first specimen of the blocking polarity (hertz sweep, measured 2026-08-04):**
  `.worktrees\hertz-ci-riders\target\POOL-OWNER.json` named `holder_pid 31380` with
  `holder_started_at 134302739211259713` = 2026-08-03 16:38:41. Pid 31380 on that box at read time was
  `claude.exe -n "doyle @ HFENDULEAM (spt-core/)"`, created 17:31:11, FILETIME 134302770715019370.
  Same boot, ~52 minutes apart, stamps MISMATCH. The claiming process died and the OS reissued its pid
  to the gater's own session. Every previously measured recycle produced a false-DEAD (reads free
  while the lane is live — the corrupting direction); this one produces a **false-LIVE**: a bare-pid
  guard reads the claim as held by a running holder and refuses, naming doyle as the holder of hertz's
  pool — the blocking direction, first time caught. The `holder_started_at` arm shipped in `8ce40d9`
  discriminates it correctly — that field caught defending, not asserted to defend
  ([[stable-anchor-is-not-a-recycling-defense]]).
- **The mirror image, same box, same hour:** `.worktrees\ir24-reap\target` named `holder_pid 33612`,
  DEAD — while that lane was genuinely LIVE, with `cargo clippy -q -p spt-daemon --lib --tests`
  (pid 22452) running under todlando's session as it was read. Two claims, opposite errors, inside one
  hour: liveness misreads in BOTH directions on the same box in the same hour. Holder liveness has no
  sound polarity as a lane-state oracle. Ancestry decides ([[pool-claim-holder-death-is-not-lane-state]]).
- **THE BARE-CLAIM GAP — a claim the stamp arm cannot adjudicate at all.** The main checkout's own pool
  (`<root>\target`, 72,001,016,671 B = 67.06 GiB) carries
  `{"owner_tree": "...\\spt-core", "written_by": "spt-poolguard"}` — no `lane_label`, no `holder_pid`,
  no `holder_started_at`, no `lane_branch`, no `lane_base`. `read_owner` maps that to `lane: None`, so
  every arm the lane-identity work added is unreachable for it: it cannot be read as in-flight,
  settled, or recycled. **FIVE of the eight** claims found on this box are this shape — `<root>\target`,
  `<root>\target-seam` (below), `gate-target-render`, `ir21-xtask\target`, `locksmith-t1\target` —
  against three carrying a lane identity (`hertz-ci-riders`, `hertz-poolowner`, `ir24-reap`). A claim
  written by a plain build rather than by `pool-claim` is structurally un-adjudicable, and it is the
  MAJORITY shape in the field. Whatever IR-26 does next has to say what an identity-less claim means —
  "the stamp arm handles it" is false for most claims that actually exist.
- **The fifth bare claim was invisible to the first inventory, and the miss is its own lesson:**
  `<root>\target-seam`, **52.06 GiB**, real dir, gitignored on its own line (`/target-seam`, added
  `5ae68f8`, v0.12.1 era), created 2026-06-18, last written 2026-08-02 — nothing in the tracked tree
  references it (grep over rs/toml/md/yml/ps1/sh: zero hits outside this file). The inventory predicate
  was "a directory named `target` inside a worktree"; a second root-level pool under a different name
  is not expressible in that filter, so it returned a clean zero that was reported as a population
  ([[verdict-from-probe-competence]] in the field). Corrects the earlier "72 GB root pool is the
  biggest object by 9x": the two root pools are 67.06 + 52.06 = 119.1 GiB of a 133.43 GiB box pool
  total, so the 44-tree sweep covered ~61% of pool bytes, not the population. **Ruling (doyle,
  2026-08-04): target-seam stands for now** — free space is ~115 GB against the 32 GB golden floor, so
  there is no pressure; it was written to yesterday, so "cold" is one day deep and its writer is
  unidentified. It becomes a reap candidate at the next free-space squeeze, after a fresh last-write
  check and identification of what wrote it on 08-02 — reaping a pool whose writer you cannot name is
  how the v0.51.0 fabricated-red class starts.
- **CLAIM-OBSERVABILITY, MEASURED (hertz, 2026-08-04, three-arm rig with competence control):** a
  displaced claimant CANNOT distinguish takeover from never-claimed — the two states are byte-identical
  on disk (arm 1 vs arm 2 identical; arm 3 proves the rig separates states that differ, so the null is
  real). The takeover is silent on BOTH sides: the taker's `pool-claim` prints ordinary success naming
  no prior owner. Mechanism at `b75258d`: `pool_claim` never calls `read_owner` — it constructs its own
  `PoolOwner` and `fs::write`s unconditionally. The displaced lane's only channel is its next build's
  refusal, which reports a STATE ("this pool belongs to B"), never a transition, and arrives pull-only
  and arbitrarily late while interleaved-artifact harm is already underway. Labelled holes: measured
  against the origin/main xtask binary (lane-tip carry is a source read, not a measurement);
  pool-sweep-as-channel not exercised; single-process holder pid in all arms.
- **Bootstrap property, not a defect (hertz, measured 2026-08-03 while claiming pools for the lanes
  3-4 rebases):** a lane DELIVERING claim-identity cannot claim its own pool WITH that identity until
  it lands — the prebuilt origin/main `xtask.exe` writes old-shaped stamps (owner_tree, lane_label,
  holder pid+stamp; no lane_branch, no lane_base). Enforcer turnover surfacing at CLAIM time rather
  than build time: coverage grows only as trees turn over. Recorded because it reads as a bug the
  first time someone meets it in the field. Related lane-craft fact from the same queue: a rebase
  changes SHAS necessarily but hunk OFFSETS only incidentally — todlando's main.rs hunks stayed at
  @20/@53/@70 across the f2e4516→2f427b2 rebase; treat pre-rebase overlap measurements as stale by
  construction and re-measure, but expect the geometry to usually hold.
- **Bootstrap property FIELD INSTANCE, main checkout (todlando 2026-08-04, IR-12 rig):** the
  prebuilt `target/debug/xtask.exe` in the MAIN checkout pool was stale against main's own source
  at `11169c1` — it printed `(holder pid …, born …)` and wrote a FLAT identity-less claim
  (owner_tree/lane_label/holder_pid/holder_started_at/written_by) where main's source
  (main.rs:2282) prints `(branch …, base …; advisory holder pid …)` and writes the nested lane.
  A lane claimed with a stale tool records NO git identity and demotes its own adjudication back
  to the refused bare-pid predicate. DETECTION that worked and is the reusable part: read the
  claim line the tool PRINTS against the source just read — a mismatch names the stale enforcer
  before it writes. REMEDY at claim time: rebuild xtask FROM THE LANE'S WORKTREE before
  `pool-claim` (which the protocol's claim-from-your-own-worktree rule already implies; this
  instance is why it is load-bearing and not ceremony).
- **RULING (doyle, 2026-08-04) — the claim path must run the same guard the build path runs, and a
  takeover becomes an event.** The documented design already says a finished lane is "taken over
  loudly, not refused"; a silent unconditional overwrite violates the spelling the protocol shipped
  under. Remedy, both halves, one small lane: (ii) `pool-claim` reads the prior owner before writing —
  prior lane LIVE (ancestry + stamp arm, never bare pid) ⇒ refuse exactly as build.rs would; prior
  lane finished ⇒ proceed AND announce the displacement, naming the displaced lane in the success
  line; plus (i) a `displaced` field (prior owner + takeover stamp) rides the new claim, so the fact
  survives in the artifact the late-reading claimant already reads. The identity-less bare claim
  (majority field shape, above) must be handled explicitly by the same change: a bare claim carries no
  lane to adjudicate, so takeover of a bare claim always proceeds-and-announces, and the announcement
  names the claim as identity-less — never silently, never refused. Ripe: next CI/infra batch, same
  lane as or beside the IR-26 remedy. Size: small.

### IR-27 — A junction-pooled rig is TWO objects and teardown only ever removes one
- **Status:** open, found in the wild 2026-08-04 · **Origin:** hertz's 44-tree worktree sweep
  (IR-26 follow-on 3) · **Cross-ref:** [[IR-14]] (unclaimed pools), [[worktree-target-junction]]
- **What/why:** a rig built as `worktree\target -> junction -> .worktrees\<pool>` has its build cache
  in a directory that is NOT the worktree and NOT in `git worktree list`. `git worktree remove` deletes
  the junction and reports success; the pool survives with zero inbound links and nothing naming it.
  Measured on this box: reaping `golden-render` and `w4-cli-doc` the obvious way would have reclaimed
  ~26 MB of worktree and stranded **14.17 GB** of pool (`gate-target-render` 6,265,448,382 B +
  `gate-target-w4doc` 7,907,558,098 B). Neither pool appears in any worktree listing.
- **The second shape, same mechanism, already realised:** `.worktrees\gate-target` was found at
  **0 files / 0 bytes with FOUR live inbound junctions** (assembly-doorbell, doorbell-w1/w2/w3). Someone
  had already reaped that pool and left four dangling-but-valid links behind. So the defect runs both
  ways: remove the worktree and the pool is orphaned; remove the pool and the links are orphaned.
  Either half alone leaves a lie on disk.
- **Why it is not just "be careful":** the classification step people already know
  (`Get-Item -Force` OUTBOUND, INBOUND reparse sweep) tells you HOW to delete an object you have
  already decided to delete. It does not tell you that a second object exists. The missing thing is an
  enumeration that reaches pools no worktree names.
- **Remedy shape (CORRECTED by builder retraction, 2026-08-04 — the enumeration already exists and
  shipped):** `xtask pool-sweep` (`c6515f1`, `REQ-POOL-GC-ORPHAN-RECLAIM`, ancestor of `b7b00c3` by
  merge-base) walks the root for dirs AND links, counts inbound links over the whole sweep before
  forming a verdict, classifies InUse/Owned/Orphaned, and reaps orphans on `--reap`. The gap is not
  detection, it is TIMING: a junction-pooled tree's pool is `Owned` and not reclaimable right up to
  the moment its worktree is removed, and `Orphaned` only after — so the reclaim always lands in a
  LATER sweep, and nothing triggers one. `gate-target` (0 bytes, four dangling inbound links) is a
  completed instance of that deferral. Remedy is a post-teardown pass or a teardown that runs the
  sweep itself, not new enumeration. (An earlier draft of this entry proposed building the
  enumeration; retracted by its author before ripening — the tool refuting it was one run away.)
- **Ripe when:** next CI/infra batch. **Size:** small (one enumeration + report, no new policy).
- **Discipline this cost, stated so it is not re-learned:** teardown must be TWO EXPLICIT STEPS per
  junction-pooled tree — junction deleted AS A LINK (`Directory.Delete(path, recursive: false)`), then
  the pool AS A TREE after an inbound sweep confirms it is linkless. Executed that way on 2026-08-04:
  28 worktrees, 4 pools, 0 skipped, 0 pinned, 17.04 GiB measured on a box with concurrent writers.

### IR-28 — A wave-gap rider is visible to the request board and invisible to a milestone-narrative notes pass
- **Status:** BUILT 2026-08-29 — PR #176 (`880b3b9c`, landed on main via `37368263`).
  · **Origin:** deployah, v0.53.0 release close 2026-08-04.
- **What/why:** the v0.53.0 changelog omitted `184f2ac` (the releases#125 two-core CPU-burn fix, a
  wave-gap rider that rode the commit range but was never LOCKSMITH scope) while the alchemy release
  verb promoted #125 to DONE and named it publicly as shipped — two published surfaces contradicting
  each other until the post-publish docs-only repair (`9841592`). The rule was already right (the
  runbook audits the changelog against the COMMIT RANGE); the EXECUTION missed it, because a notes
  pass written from the milestone narrative never enumerates the range. The mechanism, not the
  incident: any rider is board-visible and narrative-invisible.
- **Built evidence:** `docs/RELEASE-RUNBOOK.md` now requires a mechanical `vPREV..HEAD`
  user-facing commit list, a changelog match beside every row, and an explicit accounting of every
  unmatched row before the golden run. The repair precedent remains docs-only amend +
  `gh release edit`, never retag. PR #176 passed `traceable-reqs check`.

### IR-29 — twohost ladder: role A fires the done-barrier BEFORE its last cross-node rung, so the re-pull's serve window is won by timing, not guaranteed
- **Status:** open, FIX AUTHORED-AND-COMPILED unlanded — hertz `test/twohost-serve-window`
  @`17bbbd8` off `11169c1` (+33/-18 twohost.rs, relocation proven byte-identical by
  comment-stripped set-diff, cargo check + clippy -D warnings green; proving two-host run still
  owed and rides the same window as #145's closed-subnet knock row) · **Origin:** golden
  30873007187 attempt 1 red, doyle timeline triage + hertz source diagnosis 2026-08-04.
- **SECOND FACE, measured 2026-08-24 (WAX-SEAL #21 W3 gate climb @ 035c3fe6, doyle):** the class
  recurs on a NEW rung, and this face is UNWINNABLE rather than racy. The W3 seal rung S1 mints
  on A ~0.3s before A's ladder completes (mint mono 80843ms, ladder end ~81150ms, walked from
  the capture), while the seal pump worker's cadence is 500ms — the pusher exits before its
  first eligible post-mint tick, so ZERO feed pushes ever fire (both logs hold zero seal-feed
  lines) and role B burns its full 900s window on "seal: A's record replicated into B's store".
  Full-re-presentation heals only while a pump is RESIDENT; the rig kills it. Product not
  implicated (local legs + design review all green at the sha). Dispatched to hertz 2026-08-24
  with two fix shapes (mint early with ~120 ticks of resident-pump margin, or a
  reverse-confirmation barrier on A's done-push); the climb re-runs on his fix. Lesson for
  future rungs, stated as construction: a rung whose delivery is CADENCED must not sit inside
  the terminal margin of the process that pushes it.
- **Mechanism (hertz falsified the gater's first spelling, kept because the correction is the
  entry):** NOT "B exits early" and NOT a missing barrier. B's contract (twohost.rs:1116-1121,
  "hold until A finishes its side — done-file as the ladder's completion barrier") is right and B
  obeys it. A BREAKS it: A pushes done.txt (:1819-1832, payload literally "ladder complete on A")
  then runs the ENTIRE leg-D digest rung including the re-pull assert (:1959) before its own
  `stop.store` (:1966). B tears down exactly when A said it could; the re-pull then races B's
  post-signal teardown. The panic's hint text ("may be running a version without cross-node
  digest") is pre-authored and was NOT evidence of version skew.
- **Margin series, six draws of one race (deployah, markers cited by TEXT and positive-controlled
  against a known draw before producing an unknown one):** golden 30860770146 +60s · a1 −0.6s
  (the RED) · a2 +59.4s · a3 +60s · a4 +59.59s · a5 +63.89s. BIMODAL, not drifting: greens
  cluster within seconds, one collapse through zero. A green does not mean the window got safer;
  sample count characterizes a bimodal race, not margin size — the slack is incidental, not
  designed. Local load on hfenduleam during A's leg D eats the margin directly (role A runs
  there; the done-file is already pushed), which is why builds hold through twohost legs.
- **⚠ THE SERIES IS CLOSED AT SIX DRAWS — ITS DERIVATION IS UNRECOVERABLE, AND NO SEVENTH NUMBER
  MAY BE ADDED (ruled by doyle 2026-08-05, at the tranche-2 golden).** Asked to extend it with a
  draw from run 30971976024, deployah positive-controlled the markers FIRST and the control
  REFUSED: grepping `margin|deadline|serve window|slack|headroom|budget` against a control run
  known to carry good twohost evidence (30940180764) returned ZERO. Source confirms it —
  **`margin` is not an emitted token anywhere**: absent from `.github/workflows/golden.yml`, and
  every `crates/` hit is unrelated (terminal right-margins in broker/resize tests,
  `ROUND_DRAIN_MARGIN` in `pump/mod.rs`). So the six figures were DERIVED from log timestamps
  against a bound that was never written down, by a measurer whose context has since been cleared;
  the gater who recorded them never held the formula either. **Consequence, stated so nobody
  re-derives it:** any new figure computed from these logs would be a DIFFERENT quantity wearing
  the same name, and would enter a series whose value is its bimodality — the one structure a
  mismatched sample corrupts invisibly. This run therefore gets a LABELLED HOLE, not a number.
  What IS reportable for it, with endpoints named rather than a name reused: **ladder span**, node
  start → `TWOHOST role B: ladder complete` = **183.50s** on the batch vs **259.21s** on the
  control. Marker parity established under an IDENTICAL expression on both sides (25 = 25, after a
  first pass that compared two different greps and would have read as parity by coincidence).
- **The fix this makes obvious:** the twohost margin must become an EMITTED token with its two
  events named at the emit site, so the next draw is read rather than reconstructed. A derived
  quantity whose formula lives only in a measurer's context is one context reset away from being
  unfalsifiable — which is exactly what happened here.
- **Fix shape (doyle-approved):** move the done-push block to AFTER leg D, immediately ahead of
  A's stop — one honest barrier made true, no second completion signal (a redundant pair is how
  the next ordering bug lands after the SECOND one). `[int->REQ-REACH-1]` travels with the block.
- **Cascade note, so a future triage does not count it:** in a1, twohost-b's gated-CLI red
  (twohost_cli.rs:189, 910s) was the CASCADE — role-A's gated driver step SKIPPED after the
  ladder death, so B waited for a row never produced. A red that postdates its carrier leg's
  failure carries zero information about its own subject.
- **Ripe when:** proving run at the next two-host window. **Size:** landed-sized already.

### IR-30 — GOLDEN 4b37512 (#145): five attempts, four distinct Windows victims, one sha — the random-victim family measured end to end, and the instruments it left behind
- **Status:** open for the FAMILY SIGHTINGS ONLY — the instrument lanes LANDED as planned with
  the next golden batches (measured 2026-08-25, doyle, `git merge-base --is-ancestor` + first
  tag: serve-window `17bbbd8`, selection probe `695398ba`, resume taxonomy `58337879` all in
  **v0.54.0** (USHER, merge `390204a3` et al.); attach-intent re-aim `37afa571` in **v0.55.0**).
  Any later record citing these as "authored-unlanded riders" (the v0.63.0 FIELD-SEAL plan did,
  carried forward from this ledger) is stale — this row is the correction. The #145 gate itself
  CONCLUDED GREEN attempt 5, main ff'd, tested==merged · **Origin:** doyle, night of 2026-08-04,
  run 30873007187 attempts 1-5.
- **The record:** four reds, four DIFFERENT daemon-spawning tests, Windows leg only, Linux green
  at the same sha every attempt: a1 `daemon::tests::a_tree_teardown_reaches_a_grandchild…`
  (10.225s full deadline, 2nd golden sighting — see [[IR-17]]'s updated sub-observation); a2
  `registry_lifecycle::multichunk_feed_applies_with_exactly_one_snapshot_write` (0-vs-1 at
  mono_ms 77, THIRD sighting of the 2026-07-22 seed; hertz source-named the window: converge
  polls rows landed at merge, assert reads `snapshot_writes` incremented in `write_snapshots`,
  unjoined-thread gap between — predicts under-count only, matching all sightings; quiet-arm
  control 40/40 PASS 0 leaky on a still box); a3 `brain_split::broker_survives_brain_kill…`
  ("supervisor did not respawn the brain", 32.8s window burn; characterization queued); a4
  `resume_no_control_steal_e2e::brain_respawn_keeps_every_session_controller…` — UNDER the
  ruled quiesce-partial window with a measured-clean box, and its panic NAMED A MECHANISM ITS
  OWN DATA CONTRADICTS (pre-authored Failure-A text; gained=[0,15,17] through ONE
  resume_sessions while a steal displaces the SET; session 0 already 4x behind BEFORE the
  window; the rig's child-liveness immunity was a COMMENT nothing measured). Every red
  lane-independent by per-row delta test run fresh each time, never transferred.
- **Instruments authored off it (all thin test/obs lanes off `11169c1`, compiled where stated;
  ALL LANDED — v0.54.0/v0.55.0, see Status):** hertz `test/twohost-serve-window` @`17bbbd8`
  ([[IR-29]]); hertz grandchild SELECTION probe (+166/-8 daemon.rs test-mod: IsProcessInJob at
  selection time — decisive H1/H2 splitter, birth-stamp one-directional, image, same-snapshot
  match COUNT; compiled + deliberate-break positive control); hertz `test/resume-steal-taxonomy`
  (+109/-6: producer counter `Broker::session_output_seq` bracketing the consumer window,
  authenticated child pin at t1 BEFORE teardown — read at assert time would convict the rig of
  its own cleanup — three-way verdict PRECONDITION / measured-Failure-A-with-producer-story /
  absent-producer-is-its-own-story); todlando `obs/resume-attach-intent` @`9c9e6c5`
  — **a DEAD-PATH emit that never fired in production; removed 2026-08-04.** The
  `RESUME_ATTACH_INTENT` breadcrumb sat in `resume_sessions`, which has had NO production
  caller since `03c7109` (2026-07-09): the daemon's respawn path is `run_brain` ->
  `resume_session_cursors`, cursor-only, no attach. Its absence from a field log therefore read
  as "no resume happened" when it meant "uncalled function" — a clean-zero manufactory aimed at
  the very #123 hunt it was built to serve. **Do not hunt for this token.** The instrument now
  lives at the choosers the daemon actually executes, as `ATTACH_INTENT_CHOSEN` with a distinct
  `site=` per chooser (`gap_resume` | `serve_request` | `shell_channel`), carrying the emit
  discipline this one always had — intent bound once, logged and passed from that same binding,
  wildcard-free label. No REQ, per the OBS-rider precedent.
  - _Corrected 2026-08-04 (todlando, in the re-aim rider lane): this row previously described
    the breadcrumb as a live instrument — "intent bound once, logged and passed from the same
    binding, per-call epoch makes set-vs-one readable off the log". That sentence was true
    about the emit's CONSTRUCTION and false about its REACH, and a reader consulting this
    register during the #123 hunt would have been sent after a token that cannot fire.
    Replaced rather than annotated, per the register's correction convention; the superseded
    wording lives in git._
- **Rulings that outlive the night:** a green retires NO intermittent row — grandchild and the
  a4 row are INSTRUMENTED-AWAITING-FIRE, the next occurrence carries its own verdict; the
  environmental family's leg 3 (runner off the fleet host) stays operator-ruled NON-OPTION
  ("stop proposing it") — sightings accumulate in the family record, evidence files itself,
  re-asking is direction; the quiesce-partial window is the sanctioned tool and it WORKED as a
  discriminator (a4's red under quiesce is what forced the rig-soundness finding).
- **Instrument notes minted on the way, cheap to lose:** (a) `gh run view --log` refuses for a
  COMPLETED attempt while any later attempt is live (gates on the RUN) — recover per-job with
  `gh api repos/<o>/<r>/actions/jobs/<id>/logs`; (b) a process-census RESIDENT COUNT is not
  trendable (23→17→15 in 90s of inbox-poll CLI churn) — the CLASSIFICATION is the measurement,
  the count is noise; (c) nextest's per-run LEAKY count is a free process-leakage series already
  in every log (a1=12/a3=10/a4=13, and the multichunk row alone leaks 0/40 — the leaks are the
  row's NEIGHBOURS); (d) diagnostic-token greps derive from the FAILING FRAME's own emit sites,
  never from the register entry discussing the symptom — the gater's own
  `REAP_JOB_TERMINATE_FAIL` zero over a path that emits `SERVICE_TREE_KILL_INCOMPLETE` was VOID,
  and the emit sites document their own vocabulary (tree_kill_incomplete_line's doc names its
  grep-partner); (e) **"golden green ⇒ box free" is FALSE — the ff-push that concludes a golden
  STARTS the automatic post-merge ci run on the same box** (hertz, measured to the second:
  Runner.Worker up 00:13:02 local = run 30886843672 job start 07:13:02Z, ~2min after the ff), so
  the box empties after the POST-MERGE run, not after the golden; the box has a fourth actor no
  agent steps off, woken by every main push (todlando's spelling), and post-golden sequencing
  must gate on `gh run view --json jobs` of the post-merge run before any timing-sensitive work.
- **Ripe when:** ~~instrument lanes land~~ (DONE, v0.54.0/v0.55.0); family sightings stay until
  the family record closes them — the entry now waits on FIRES, not landings. **Size:** nothing
  left to build; the next family sighting carries its own verdict via the landed instruments.
- **LINUX-FACE SIGHTING + ITS RESOLUTION (doyle, v0.63.0 FIELD-SEAL goldens r1-r3 + proof runs,
  2026-08-25 — the family's first measured MECHANISM, and it is environmental-amplifier ×
  latent-test-defect, not a product fault):** three goldens, three DISTINCT kitsubito victims,
  8/9 green each, victims migrating across same-sha reruns (r2/r3), every red a LONG death
  (9.9s/59.4s/62.7s) of a fast cell — the family signature on the OTHER OS. Root-caused to ONE
  window whose WIDTH varies with load: the brain's `spawn_session_pid` Spawned-wait
  consumes-and-discards output racing ahead of the reply (KNOWN-HAZARDS 6.9, same-session face,
  amended at `dbe3daad`); the broker execs the PTY child ~2ms BEFORE writing `Spawned`, so a
  fast first chunk can be eaten. Under kitsubito's then-active kernel-audit backpressure
  (see the kitsubito-audit entry below) the broker DISPATCH thread could stall arbitrarily in
  that gap — the ~2ms window stretched to seconds, swallowing ANY marker prefix of ANY
  spawn-wait test: random victims, long deaths, spawn-intensity × serialization scaling
  (deployah's measured pair), migration across reruns. Box remediation narrowed the window back
  to ~2ms, leaving exactly one knife-edge victim (a first-chunk needle, near-deterministic red,
  62.7s natural-life death), fixed test-side in `dbe3daad` (v0.63.0's head). Full RCA + the
  observer-effect instrument record: releases#225 comments 5417343026/5418844404.
  **Standing lesson for the family:** an environmental slowdown is an accidental mitigation —
  REMEDIATING a box can EXPOSE latent knife-edge races (the storm had hidden this one for
  releases). And a "random victim family" should be tested against ONE window with
  load-dependent width before positing per-victim mechanisms. The Windows face above remains
  open on its own instruments; the two faces now have one shared candidate SHAPE (a race window
  amplified by box load) with different windows.
- **Status:** open · **Origin:** measured box event, HFENDULEAM 2026-08-04 ~05:00 local, during
  USHER lane concurrency (hertz mechanism statement + both agents' reclaim arithmetic).
- **What/why:** the box hit **0.00 bytes free** mid-build. Failure signatures it manufactured
  look like toolchain or code defects, not disk: `rustc-LLVM ERROR: IO failure on output stream:
  no space on device`, `LNK1318: Unexpected PDB error; LIMIT (12)`, `LNK1108: cannot write file`
  — a red wearing a linker's face (doyle's U2 gate rig ate exactly this; verdict legs already
  green survived, the suite leg had to be re-run). Contributions measured, not inferred: hertz
  ~70 GB across four lane targets (er-seams **42.66 GB** cold `--workspace --tests`,
  selection-probe 25.80, poolowner 1.03, ci-riders 0.91), doyle gate rig 15 GB cold full build.
  Reclaim arithmetic reconciles both and neither alone: 0.00 → 25.47 GB (hertz reaps
  selection-probe) → ~40 GB (doyle reaps rig). **The mechanism (hertz, ruled register-worthy):
  every isolated lane worktree carries its OWN `target/`, so disk cost is multiplied by lane
  count, while pool-claim — built to arbitrate a SHARED pool — sees none of it; a claim on
  main's pool is ceremonial for a lane that never writes there. The tool we use to reason about
  build-cache contention cannot see the resource that actually ran out.** Eight-plus worktrees
  at 25–43 GB per cold full build is the shape of the next occurrence, and it will not announce
  itself through pool-claim. Adjacent discipline failure the same night, named so the entry
  carries it: a gater firing a cold full-workspace rig into a window with two live builder lanes
  is the pre-flight question-1 failure (right-size the run) — targeted legs and warm pools
  first.
- **Remedy shape (sketch, not ruled):** (a) a box-level free-space floor as a RIG STEP at lane
  start — the claim verb is the natural seat (print box free + du of known lane targets at
  `pool-claim`, warn under a floor); kin to the CI-side `REQ-CI-FREE-SPACE-PREFLIGHT` (IR-1's
  companion fix), which covers runners but not agent lane rigs; (b) teardown-on-gate-close is
  already a rig step (the gate-target-disposal rule) — the gap is the AGGREGATE view across
  lanes nobody owns; (c) possibly a `pool-census` xtask verb listing every `.worktrees/*/target`
  with sizes, so the sweep is one command instead of a du walk each agent re-derives.
- **Ripe when:** next CI/rig-touching wave, or the next ENOSPC-signature red — whichever first.
- **Size:** small-medium (claim-verb print + floor; census verb optional).
- **Field addendum (2026-08-04 second event, same day filed — ripeness condition met):** the box
  fell under the CI runner's 32 GiB free-space floor twice more (12:31 main @9d65e65 docs-only,
  12:50 PR core#144 @b110bc8) while todlando's F-lane built. Both Windows unit legs REFUSED with
  `RESOURCE=disk drive=C:\ free_bytes=…` — the `REQ-CI-FREE-SPACE-PREFLIGHT` floor did exactly its
  job: a docs-only main red named the resource instead of wearing a linker's face, and triage was
  one log read instead of an RCA. That is the instrument's first field catch; the CI side of this
  entry is PROVEN. The agent-lane side stays open: local trough measured 8.8 GB free mid-F-lane.
  Reclaim, measured before/after per the teardown rule: doyle reaped gated u1+u2 lane targets
  (POOL-OWNER claims verified own+dead-holder, inbound reparse sweep clean, worktrees kept)
  8.8 → 59.5 GB (+50.7); hertz reaped er-seams target (pool-release first, same discipline)
  59.46 → 101.16 GB (+42.66). Refined mechanism statement (hertz, this event): the disk floor is a
  per-BOX resource our per-POOL instrument is structurally blind to — pool-claim answers a
  question about contention that is not the question the box ran out of.
- **AXIS ADDENDUM (deployah, golden 30928816784 / USHER `fc7fad1`, 2026-08-04) — one floor reading
  is a SNAPSHOT, and the swing is bigger than the margin you fire on.** Two facts this entry's
  remedy (a) has to be built against, both measured rather than reasoned:
  (1) **A pre-fire PROCESS census cannot see the floor at all.** Mine came back clean minutes before
  the run — zero test-path `spt.exe` residents, zero `cargo`/`rustc`/`link`/`cl` — and both Windows
  legs then died in the disk preflight (`drive=C:\ free_bytes=26099576832 floor_bytes=34359738368`,
  `RESOURCE=disk`, exit 1) with a 128-line log carrying ZERO `Compiling`/`PASS`/`FAIL`/`Summary`
  lines: no test signal, and the merge chain never indicted. Both readings were TRUE at the same
  instant — the gater's finished assembly gates were sitting on disk as a DIRECTORY, not running as
  a process. A process-axis probe is blind to a finished build's cost by construction; "re-check
  closer to the fire" would have changed nothing, because the axis was never read.
  (2) **The floor number itself moves by tens of GB inside one run window.** Same box, same drive,
  same run, pulled from its own logs: `16:22:55` test-Windows free=26099576832 REFUSED · `16:23:07`
  n1-gate-Windows free=26096607232 REFUSED · `16:46:54` twohost-a free=71033049088 PASSED and then
  ran 11 minutes to success. ~45 GB returned with NO deliberate teardown, and my own reading at
  17:13:45Z was 61352714240 — ~10 GB BELOW what twohost-a saw 27 minutes earlier. The observed swing
  amplitude EXCEEDED the margin I was about to fire on (25.14 GiB).
  So the rule remedy (a) must encode: measure BOTH axes (process residents AND free space against
  the workflow's own 32 GiB floor), and reclaim until the margin exceeds the observed SWING, not one
  sample. Triage corollary, equally load-bearing: twohost-a runs the same preflight on the same C:
  and PASSED, so a preflight refusal is a threshold event on a moving number — never evidence the
  box cannot run the work, and never by itself a reason to indict a merge chain.
- **THIRD ADDENDUM (deployah, v0.54.0 cut, 2026-08-04) — the sweep answers "0 B reclaimable"
  TRUTHFULLY on a box that cannot run its own work, because the reclaimable population and the
  owned-lane population are DISJOINT.** This hazard blocked a release twice before it was seen.
  Both Windows legs at the shipped sha `86f0d84` refused at the guard before compiling anything —
  `ci.yml` at `free_bytes=28167569408` (26.2 GiB) and `release.yml` at `27690885120` (25.8 GiB),
  both against `floor_bytes=34359738368`. `assemble` skipped in consequence, so no draft release
  existed: the release was hard-blocked on disk, not on code. Neither leg produced a test verdict,
  so both are LABELLED HOLES, not reds (ruled recorded-and-proceed by doyle, the shipped-sha delta
  being 5 files and zero `.rs`).
  **The finding is not the disk.** `xtask pool-sweep --root .worktrees` reported, correctly,
  `total 66.43 GB across 3 pool(s); 0 B reclaimable in this sweep's scope` while the box sat ~6 GB
  under its own floor. The sweep is behaving as designed — it refuses to reap owned, live lanes.
  But the floor reasons over an AGGREGATE the sweep is forbidden to touch, so the designed
  instrument, run at the exact moment of refusal, tells an operator there is nothing to reclaim.
  That is true and useless. The gap between "the sweep is correct" and "the box can run work" is
  the defect; it is not fixed by making the sweep more aggressive.
  **Aggregate measured under `.worktrees`** (real target bytes, junction-classified first): 87.79 GB
  across 9 lanes — 2.7x the entire 32 GiB floor. Registered pools: `w1t2-relink-force` 54.67 GB,
  `obs-resume-attach-intent` 10.18, `w1t2-subnet-status` 1.58. Carrying targets the sweep does not
  count as claimed pools: `hertz-teardown-bound` 11.25, `hertz-liveresolve-diag` 4.83,
  `hertz-psyche-bound` 4.19, `hertz-poolowner` 1.03, `hertz-ci-riders` 0.91.
  **The 54.67 GB single lane is NOT waste** (doyle's classification, ratified at this filing): a
  verb lane at that size is workspace-all-targets e2e cost. No lane is misbehaving. That is what
  makes this structural rather than a cleanup task — every lane is individually justified and the
  sum still exceeds the floor.
  **Reclaim taken, and its boundary.** 82.12 GB, from `spt-core/target` — the release driver's OWN
  main-checkout pool, NOT a lane; free 23.23 → 105.34 GB. Classified before removal per the
  teardown rule: OUTBOUND `Get-Item -Force` reported a real directory rather than a reparse point,
  the INBOUND sweep found zero reparse points aimed at it, `CARGO_TARGET_DIR` was unset (no env
  aliasing), the target SUBTREE only was reaped, both sides measured.
  **Snapshot behavior reproduced inside this measurement window**, corroborating the axis addendum
  above at a smaller amplitude: free read 25.49 GB when first measured and 23.23 GB at the reap
  minutes later — it drifted 2.26 GB DOWNWARD while the operator was deciding what to do about it.
  **Priced side effect, for the next reader:** reaping the main pool took the prebuilt `xtask` with
  it, so the next lane claim rebuilds it first — the lane-check shortcut is cold until then. It cost
  a 2m11s cold rebuild inside this release's own publish step.
  **What this does NOT license:** reaping another agent's owned lane to clear a floor. The sweep's
  refusal is correct and stays. What is missing is an AGGREGATE-AWARE signal — the box knowing its
  lanes sum past the floor *before* a run is dispatched into a guard that will decline it.
- **FOURTH ADDENDUM (doyle, 2026-08-04 late evening — fourth event in one day, and the first
  RULING on this entry).** C: hit **0.44 GB free** during doyle's lane-2 stack gate and todlando's
  re-aim leg set. Signatures manufactured this time, all initially read as something else:
  `os error 112` mid-rlib-archive surfacing as xtask "building spt failed" (exit 101); a
  servicehost force-kill unit red at a sha gated GREEN on the same rig an hour earlier; todlando's
  6-of-7 leg table red with `LNK1318` while only traceable-reqs survived. Both agents' verdicts
  from the window were VOIDED and re-run, not re-read.
  **Discriminator correction (hertz, ratified at this filing):** "treqs alone survives" is a
  POSITIVE TELL, never a clearing test — a full disk reds BEFORE any link step (rustc writing
  rlibs), and treqs can red for its own reasons; signature absence says NOTHING. The only
  falsifier for "this red was the disk" is free space AT RUN TIME, and no log on the box recorded
  it, which made every red from the window unanswerable after the fact.
  **RULING (doyle):** every rig/gate script records `FREE-AT-START` / `FREE-AT-END` in its own
  SUMMARY — one line each, the datum that makes a disk confound answerable post-hoc. Effective
  immediately for hand-authored rigs; hertz builds it into the shared rig harness at next touch
  (assigned, not unasked). This is remedy (a) narrowed to its cheapest load-bearing slice.
  **Reclaim, all FS-delta-measured:** todlando +59.4 GB (obs-resume-attach-intent 8.56 +
  w1t2-relink-force 50.83; 0.44 → 59.83), hertz +1.79 GB (ci-riders + poolowner; sum-of-lengths
  said 1.94, the FS delta governs; 59.09 → 60.87), doyle +38.8 GB (w1t2-subnet-status +
  w1t2-perch-gc + w1t3-node-verb targets; 60.75 → 99.54). Full discipline each: outbound
  classification, inbound reparse sweep (zero), CARGO_TARGET_DIR confirmed unset, subtree-only.
  **Mechanism sharpened by this event:** the standing swing IS the finished-lane population —
  ~110 GB of pools belonging to lanes already GATED AND PUSHED, held for hours because disposal
  fires at GATE close (a rig step) while nothing fires at LANE finish; a pushed lane awaiting
  assembly holds its pool invisibly. Corollary corrections carried: todlando withdrew his
  name-based pool census (10x off; ancestry, not directory names, classifies a pool as closed),
  and hertz-selection-probe's worktree removal REFUSED Permission denied — IR-38's second holder
  class recurring, left for retry rather than forced.

- **FIFTH ADDENDUM (deployah, v0.63.0 FIELD-SEAL golden, 2026-08-25) — SCOPED DELIBERATELY TO ONE
  CONSEQUENCE: reaping a pool UNCLAIMS its tree, and the next builder meets a refusal that is the
  rule working.** (The disk-floor mechanism from the same event is doyle's to file at his
  release-close sweep — one voice each, by division agreed in-window. This addendum stops at the
  pool-claim consequence and deliberately does not restate the floor finding.)
  Five FIELD-SEAL lane pools were reaped to clear the golden floor — gate rig `gate-222-ca30a79e`
  plus lanes `fix-222`, `w2-seal-ux`, `w3-ingest`, `w4-exit-subject`; **148.48 GB measured by FS
  delta**, free 1.37 → 149.84 GB, worktree checkouts kept, `main/target` untouched by ruling.
  **The consequence to record:** `POOL-OWNER.json` lives INSIDE `target/`, so reaping the subtree
  takes the claim with it. Those five trees are now **UNCLAIMED — which is fresh-not-foreign**, a
  different state from the `SPT_POOL_FOREIGN` refusal. Whoever resumes one of those lanes (a hertz
  fixup, a golden-red rebuild) must run `cargo run -p xtask -- pool-claim --pool <dir> --label
  <lane>` **from that worktree** before the first build, or `crates/spt-store/build.rs` refuses.
  **A refusal there is the rule working, not a defect** — filed precisely so a later reader does
  not open an issue against the build script for behaving correctly. Naming it costs one line;
  the alternative is an RCA against our own guard.
  **Prediction-vs-meter, corroborating hertz's fourth-addendum note at ~75x its scale:** pre-reap
  sum-of-lengths sizing predicted 163.9 GB; the FS delta paid 148.48 GB, a 15.4 GB shortfall left
  UNEXPLAINED here rather than rationalized (allocation granularity, sparse/compressed extents and
  concurrent writes all sit between the two meters). **The FS delta governs** — hertz ruled this at
  1.94-vs-actual and it holds at three digits. Report the meter, never the estimate that agreed
  with your plan.
  **Discipline executed, for the audit trail:** outbound classification re-read AT REAP TIME (all
  five real dirs, zero LinkType); inbound sweep over **2854** reparse points under `projects\`
  found zero aimed at any victim; `CARGO_TARGET_DIR` confirmed empty in process, **Machine and
  User** scope (the arm that leaves no directory entry to notice afterwards); process census
  RE-RUN at reap time rather than carried from a peer's snapshot — zero builders, zero processes
  imaged out of a victim pool, both live `spt` images resolving to the installed path. The peer's
  snapshot proved correct; it was re-run because a snapshot is not a standing property, which is
  the reason it was offered with that caveat.

### IR-32 — Docs-drift gate blesses its own gen: an empty-emitting producer agrees with itself perfectly
- **Status:** open · **Origin:** todlando mid-lane stop-and-report, U1 (#144) 2026-08-04. Filed by
  the builder to the gater deliberately — register shape is doyle's to own.
- **What/why:** with a defective binary in the tree (stack-overflowed before reaching its own
  code, empty stdout, non-fatal to the generator), `cargo run -p xtask -- gen` wrote
  `docs-site/src/cli/reference.md` reduced to EIGHT lines — title, do-not-edit banner, empty code
  fence — deleting 3193 lines of published CLI reference. `xtask check` then exited 0, because
  the drift gate compares what the binary emits NOW against what the file holds, and gen had just
  overwritten the file with exactly that emit. **A producer that emits nothing agrees with itself
  perfectly.** Caught only because the builder grepped the regenerated page for his new flag
  names and treated the clean zero as suspicious ([[zero-match-filter-reads-as-absent]] shape);
  the gate itself would never have said a word, and the gutted public reference ships green.
  The mechanism generalizes past the stack overflow that exposed it: ANY failure mode that makes
  `spt --help` produce empty stdout while exiting non-fatally to the generator gets the same
  green. Kin to [[cli-command-docs-drift]] (the gate this defeats) — the gate detects DRIFT
  between binary and page, and is structurally blind to a page just regenerated from a broken
  binary.
- **Remedy shape (builder's sketch, sound; not yet ruled on seat):** a floor in `xtask gen` —
  refuse to write a page whose help block is empty, or whose output is implausibly shorter than
  the file it replaces — so the generator fails loudly instead of producing a stub the gate
  blesses. Gen is the right seat: check runs in CI on a fresh emit too, so a check-side floor
  alone still lets a local gen gut the working tree.
- **Ripe when:** next docs-tooling or xtask-touching lane; SOONER if any lane regenerates the CLI
  reference before the floor exists (a gate on that lane must assert page LENGTH, not gen's exit
  code, until this is built).
- **Size:** small (one refusal + a length-plausibility floor in gen).

### IR-33 — Debug-build clap Command tree runs near the main-thread stack ceiling; the margin is THREE net new args, bisected at `8f291e1` (the original ~6 was measured on U1's tree and is superseded)
- **Status:** open · **Origin:** todlando U1 (#144) 2026-08-04, delta-tested mid-lane.
- **What/why:** eight extra hidden bool args across the five knock seats grew the derive-built
  Command tree past what the debug binary's main thread stack can construct: EVERY invocation
  (`spt --version` included) died with `thread 'main' has overflowed its stack`, exit
  -1073741571, before reaching any of its own code. Delta-tested, not assumed: the same command
  on the main-pool binary from main printed normally. U1 repaired its own trigger (raw-argv
  pre-scan for retired flags before clap parses; lane ends net +2 args) — but the CEILING
  remains: the tree is now close enough that a handful of net new arguments in a debug build
  reproduce a total binary outage — **and the handful is THREE, not six; see the correction
  below before planning against this entry.** **#5's verb-surface rework is the next lane that adds
  arguments and it is much bigger than U1** — this entry exists so #5 is planned knowing the
  margin, not discovering it. Note the failure's face: it presents as a broken binary, and via
  [[IR-32]] it presents as a silently gutted docs page — neither names the stack.
- **CORRECTION — the margin is THREE, bisected at `8f291e1` (todlando's U3 lane
  `build/usher-u3-verb-surface`, own pool, debug profile; filed by deployah 2026-08-04).** The
  original ~6 was measured on U1's tree and was right when written; it is wrong now, and the
  difference is the difference between "plan carefully" and "one flag anywhere takes the binary
  down". N hidden bool args added to the ROOT `Cli` derive, tree restored from git between every
  row, `--version` as the invocation (it carries no work of its own, so whatever it costs IS tree
  construction):

      N=0  exit 0                OK  (control — unmodified lane binary answers `spt 0.53.0`)
      N=1  exit 0                OK
      N=2  exit 0                OK
      N=3  exit -1073741571      STACK-OVERFLOW
      N=4 / N=6 / N=8            STACK-OVERFLOW

  PLACEMENT DIFFERENTIAL, run because #5 adds args at `EndpointCmd` seats rather than at the root
  and a root-only number could have measured the wrong thing: same N injected inside a NESTED
  endpoint seat gives `SEAT N=2` OK, `SEAT N=3` STACK-OVERFLOW. **Identical ceiling — placement does
  not move it**, so three is the number for the shape #5 actually builds. (Rig hazard worth keeping:
  the first seat attempt produced a COMPILE error, E0027, because the dispatch destructures that
  variant exhaustively — taken at face value it would have read as "the seat is fine". Both rows
  above are from the fixed rig.) CONSEQUENCE FOR #5: the ratified MIN spelling ends net +2, i.e. it
  would have shipped with ONE argument of headroom, and any later lane adding a single flag anywhere
  in the tree would then take the whole binary down, `--version` included. Remedy (b) below stops
  being "buys headroom" and becomes the precondition for #5 shipping safely.
- **Remedy shape (sketch, not ruled):** (a) cheapest tooth — a debug-build smoke that runs
  `spt --help` and asserts non-empty stdout + exit 0, which also backstops IR-32's trigger;
  (b) raise the main thread stack for the binary (build config), buying headroom without
  restructuring; (c) structural — box/flatten the derive tree so construction cost stops scaling
  with arg count. (a) is a rider candidate for any lane; (b)/(c) want a real measurement of
  where the ceiling sits before choosing.
- **Ripe when:** #5 (U3 verb surface) PLANNING — this is a planning input, not just a build item;
  the smoke tooth (a) is ripe for the next CI-touching lane regardless.
- **Size:** small for (a); medium for (b)/(c) with the measurement.

### IR-34 — nextest LEAK flag on the zombie-fixture row is dominant-but-intermittent (19/20 measured); any gate reading it as a signal flaps
- **Status:** open · **Origin:** doyle #142 F-lane gate 2026-08-04 (first sighting in a gate run);
  characterized same day by hertz from retained logs, ownership falsified by todlando.
- **What/why:** `broker::tests::windows_session_is_zombie_sees_a_handle_held_corpse_as_dead` (the
  ADR-0041 zombie-detection fixture, whose subject is a corpse process kept alive-looking by a HELD
  HANDLE) flags nextest LEAK on most draws but not all — measured 19 LEAK of 20 on a tree carrying
  no #142 content (`test/grandchild-selection-probe` @`695398b`, based on `4b37512`: lib arm 10/10
  LEAK across 822-test runs; full-suite arm 9/10 across 1044-test runs). The single clean PASS
  kills "intrinsic therefore always": leakiness is DOMINANT BUT INTERMITTENT. Consequences:
  (1) a gate or reader that treats this row's LEAK as a defect signal flaps at roughly 1 in 20;
  (2) one non-leaky run is NEVER evidence that something fixed it (kin:
  [[intermittent-green-is-zero-information]]); (3) the verdict method that closed it is the
  template — two independent legs, neither load-bearing alone: the suspect lane's diff grepped
  zero hits on the row's subject (zombie / handle_held / OpenProcess / corpse), AND a 20-draw
  baseline on a tree predating the suspect change. This entry exists so the next gater who sees
  the flag reads one register line instead of running that RCA cold.
- **Remedy shape (sketch, not ruled):** annotate the row as expected-leak in nextest config
  (per-test `leak-timeout` override or documented allowlist) so the flag stops presenting as
  signal; alternatively a fixture-side close of the held handle on the clean path if ADR-0041's
  arrangement permits — hertz's call, test/CI lane.
- **LANE-LINKED 2026-08-30 (#242 close sweep):** hertz's queued daemon-leak fixup lane carries
  this entry (brief cites IR-7/17/20/34/35/63); leaves the register when that lane lands.
- **Ripe when:** hertz's next nextest-config-touching lane; blocks nothing today.
- **Size:** small.

### IR-35 — hfenduleam Windows test-leg victim rate: one teardown/liveness family, measured at ~half of all executions
- **Status:** open · **Origin:** USHER #150 golden triage 2026-08-04 (doyle; all rates deployah-measured).
- **The numbers (15 golden runs 2026-08-02..08-04 + the USHER cycle, deduped on `(run, box,
  started_at)` — GitHub partial reruns COPY untouched jobs into the new attempt with their ORIGINAL
  `started_at` and carried conclusion, so a carried failure is the SAME observation; dedupe before
  any rate math):** Windows (hfenduleam) 13 red / 22 executions = 59% leg-red, splitting 10/22 = 45%
  test-victim + 3/22 infra (IR-31 class). Linux (kitsubito) 2/20 = 10%, 1 victim. Every victim-red
  leg had EXACTLY ONE victim (8/8 legs). Clean-cycle arithmetic ≈ 0.55 × 0.95 ≈ half — a golden
  cycle at these rates is a coin flip. knock145 (the previous green golden) took FIVE Windows
  executions to its green; USHER took four. **A both-green draw is a draw, not evidence the family
  closed.**
- **The family:** 11 victim legs, 8 distinct tests, ONE cluster — process teardown / kill-reach /
  respawn-controller survival, concentrated on hfenduleam. Per-victim dispositions live in the USHER
  #150 triage ledger (releases#150 record). The only surviving cross-sha EXACT repeat after source
  verification is `resume_no_control_steal_e2e` (:488, byte-identical @2138b16 + @4b37512) — and
  both its reds PREDATE the post-#145 producer-side instrumentation, so what repeats is an AMBIGUOUS
  observation (stolen controller vs starved child) twice; only a post-#145 red discriminates. The
  `a_tree_teardown` pair was a FALSE repeat — two mechanisms wearing one assert string (see
  FLAKE-LEDGER row; the counting lesson: assert-body identity must include the polled predicate's
  SUBJECT).
- **Off-CI discriminator (todlando 2026-08-04):** 0/400 (A 0/200 + B 0/200, interleaved,
  `--test-threads 1`, filters positive-controlled, by-path sweeps every iteration) on the warm
  d449ce5 pool UNDER live-fleet load (CPU 46-59%, 13-18 spt processes). Points the family at the
  RUNNER ENVIRONMENT rather than a product race any Windows box expresses — not exoneration (absence
  of reproduction is not proof of absence), but where the next hour goes. A-alone N=500 on the
  instrumented tree ran post-close (result on releases#150).
- **Remedy lanes already moving:** hertz 4-item test-rework package dispatched 2026-08-04 (psyche
  bound split, live_resolve rig instrumentation, green-capture probe profile, resident_service/
  teardown hardening); prod `wait_bounded` tree-kill-on-timeout filed to board as EVAL. This entry
  is the RATE's home — it leaves the register when the operator has ruled on the rate (accept vs
  hold-for-hertz) AND the hertz package has landed with a re-measured victim rate.
- **LANE-LINKED 2026-08-30 (#242 close sweep):** hertz's queued daemon-leak fixup lane carries
  this entry (brief cites IR-7/17/20/34/35/63); the rate re-measure rides that lane's landing.
- **Ripe when:** operator brief at USHER close (immediate); re-measure after the hertz package lands.
- **Size:** visibility entry + the re-measure; fix cost carried by the hertz package.

### IR-36 — `output_bounded` is a 51-copy clone estate with a named-but-empty shared home; hoist-and-delete wants its own lane
- **Status:** BUILT 2026-08-18 on KEYSTONE #182 hygiene lane. · **Origin:** hertz measurement
  2026-08-04 (at `7465ac1`, counted before listing, listed untruncated), during the USHER package
  item-2 sweep ruling; doyle ruled file-don't-sweep.
- **What/why:** 51 `fn output_bounded` definitions under `crates/spt/tests` — 50 per-file copies
  plus `tests/common/mod.rs:53`, which is NOT pub, has ZERO callers outside its own module, and
  whose doc comment states outright that the estate's copies predate the shared module. The lossy
  timeout arm (panic loses everything captured, names nothing) is in 38 sites across 36 files.
  The copies have already drifted: two variants of the deadline expect-string exist in the estate
  (`"captured spt call must complete within its deadline"` in common/mod.rs:60 vs the
  no-deadline-clause spelling in per-file copies), so a grep keyed to either string undercounts.
  The right fix is a hoist-and-delete refactor (make the common copy pub, adopt it estate-wide,
  carry the item-2 diagnostic arm once) — a 50-file test-estate touch that must NOT ride golden
  cadence as a rider or hide inside a scoped diagnostics item (ADR-0050 thin-lane rationale;
  hertz's USHER item 2 stays scoped to live_resolve, which keeps ONE improved copy as the model).
- **Built evidence:** one public owning helper replaced all 50 file-local
  clones. It owns and tree-kills a timed-out child, waits for reap, drains both
  pipes, and includes the deadline plus partial stdout/stderr in the panic.
  A dedicated regression observes both diagnostics and proves the timed-out PID
  is gone.
- **Ripe when:** a dedicated post-batch test-hygiene lane (natural companion to IR-13's
  mutation-proof loop and IR-23's home-collision fix — same hertz wave shape).
- **Size:** medium by file count, small by risk (mechanical hoist; no assertion changes).
- **Hygiene-lane scope addition (todlando, at #6, 2026-08-04):** a CI prebuild of
  `cargo build -p mock-adapter --bin mock-shell` (derived from tracked `.github/` content) turns
  SIX binedge baseline rows green at once instead of growing the burn-down list one consumer at a
  time — belongs to this lane, not to any thin verb lane. [[IR-38]] (live-binary e2e reap sweep)
  rides the same lane.
- **Built evidence:** all 51 file-local definitions are gone; 158 call sites use the one public
  `common::output_bounded`. The shared helper owns the child, kills and reaps its process tree on
  timeout, drains both pipes, and reports partial stdout/stderr. `bounded_output_e2e` mutation-pins
  both timeout ownership and retained diagnostics.
- **Declared validation extra:** the lane's docs-link gate exposed that `llms_target_exists` treated
  generated targets as absent before generation. Its small refactor + unit keeps source and
  generated link targets distinct; it is deliberate validation fallout, not part of IR-36.

### IR-37 — `traceable-reqs check` is coverage-only: tag PLACEMENT is unenforced on any multi-tagged REQ, and the more tags a REQ accumulates the less any one is checked
- **Status:** UPSTREAM CLOSED THROUGH RELEASE 2026-08-19 — traceable-reqs PR #19 squash-merged
  @`82d8b14` after doyle's targeted re-review PASS (defect A item-tables + 9 red-first regression
  cells + sighted control; defect B enclosing-item fallback DECLARED in SPEC with cost; 267/267,
  blind repros exit 0); v0.4.0 cut (version bump @`c9a8eb5`, tag pushed, 3 platform assets,
  release notes authored; refs sent to hertz). REMAINING and dispatched to hertz post-#182-landing:
  spt-core consumes v0.4.0 — CI version bump + placement config (enforce-on, module_banner=accept);
  banner cleanup rides a later lane. Entry closes when the consume lane lands. · **Origin:** hertz
  measurement 2026-08-04 (filed by doyle; the measurement is
  hertz's), found when his own pre/post `check` comparison around a tag-adjacency fix came back
  EXIT=0 on BOTH sides — the green he had cited beside the fix validated nothing.
- **What/why:** measured on `test/teardown-bound-shape`: a `[unit->REQ-RESIDENT-SERVICE]` tag
  SEPARATED from its `#[test]` fn (a const landed between tag block and fn — an AGENTS.md rule-1
  violation) scores EXIT=0 identically to the fixed adjacency. Positive control explains the
  mechanism: that REQ carries 57 unit tags across five files, so 56 other sites decide the
  coverage verdict and no single tag's placement can turn it red. Consequence: AGENTS.md rule 1
  ("tag on or immediately above the real evidence") is agent discipline only — drifted tags decay
  silently while coverage stays green, and heavily-tagged REQs decay fastest. This is the
  [[IR-32]] family shape (a gate that cannot see its own blind spot) applied to the traceability
  gate itself.
- **LIMIT, with falsifier (hertz's, verbatim in kind):** behaviour on a SINGLE-tagged REQ was NOT
  established. Separating the only tag of a single-tagged REQ at the stage under test would settle
  whether placement is unenforced outright (still exit 0 ⇒ the checker never reads adjacency) or
  merely un-enforceable at scale (exit 1 ⇒ the checker sees absence, and only multi-tag redundancy
  masks drift). Run that discriminator BEFORE designing any remedy — the two outcomes want
  different fixes (a placement rule in the checker vs a per-tag-nearest-item heuristic).
- **Transfer:** doyle accepted this lane's independent EXIT-0 remeasurement and confirmed it
  matches todlando's settled 2026-08-04 discriminator. Upstream remedy: enforce adjacency for
  each individual `doc`/`impl`/`unit`/`int` tag, with separated-tag negatives covering
  single-tag, interposed-const, file-top-banner, and one-of-many displacement. This repo consumes
  the released checker only.
- **Close condition:** doyle reviews and merges the upstream lane, cuts the `traceable-reqs`
  release, and spt-core consumes that release. Issue/branch/PR/commit refs append here when the
  upstream lane lands.
- **Size:** measurement complete; remedy scope is upstream-owned and dispatched.
- **DISCRIMINATOR RUN 2026-08-04 (todlando, authorized specimen on `build/w1t2-shell-relink-force`;
  filed by doyle — the measurement is todlando's): the answer is the EXIT-0 ARM.** Single-tagged
  int stage (`REQ-SHELL-RELINK-FORCE`), tag SEPARATED from its fn (parked at file top):
  `traceable-reqs check` EXIT 0 and the REQ still reports `+int`. Tag reverted to adjacency:
  EXIT 0, `+int`. Both exits redirect-read, no pipes. **Placement is UNENFORCED OUTRIGHT — the
  checker never reads adjacency; multi-tag redundancy was never the mechanism, it only widened the
  blind spot.** Remedy option (a) is the live one: a real placement rule in the checker (upstream
  experimplate), not a per-tag-nearest-item heuristic. Measurement caveat, carried at the
  measurer's own insistence: an intermediate probe printed `TAG_ADJACENT_AFTER_REVERT=0` and that
  was the PROBE wrong (`grep -B 1` above the fn lands on `#[test]`, not the tag above it), not a
  failed revert — the revert was verified directly (exactly one tag, above `#[test]`). The
  discriminator result is the record; the intermediate zero is not.
- **REMEDY DISPATCHED 2026-08-18 (doyle):** upstream lane opened in BigscreenVR/traceable-reqs
  (the actual upstream remote; the checkout at `~/Documents/projects/traceable-reqs` tracks it —
  "experimplate" above named the tool's origin project, not the repo) — checker-side placement
  rule, default-on in `check`, positive controls all four stages + separated-tag negatives
  including the single-tag displaced discriminator shape and the interposed-const shape.
  Requested by hertz as the IR-37 prerequisite of his test-hygiene family lane (KEYSTONE #182).
  spt-core consumes the upstream RELEASE only — no local shadow checker. Issue/branch/PR refs
  land here when the lane reports.

### IR-38 — CLASS: an e2e that ends with a live daemon-spawning binary wedges the NEXT build in its pool, and the diagnostic names the wrong lane
- **Status:** BUILT 2026-08-18 on KEYSTONE #182 hygiene lane. · **Origin:** todlando 2026-08-04,
  `shell_relink_force_e2e` on the #6 lane (filed by doyle; the finding and the fix are todlando's).
- **What/why:** the e2e ends with a LIVE binary by design; on Windows a live process holds its
  image open, so the NEXT cargo invocation in that target dir dies with
  `failed to remove file target\debug\spt.exe: Access is denied (os error 5)` — surfacing as
  exit 101 with NO test summary, because the failure names the BUILD and the FILE, never the test
  that leaked the holder. Actual holders measured: two lane-local `spt daemon run --detached` +
  `daemon brain` processes spawned under the test's temp SPT_HOME by a child CLI's
  `ensure_running`, still alive from an earlier run. **Reap BY IMAGE PATH** — the fleet's other
  eleven spt processes run from the installed binary; a name-based sweep takes your own daemon.
  The live shell-spawn population is eleven e2e files. Nine already end through product teardown
  or their authenticated home-scoped reap. The two intentionally live-at-success siblings now
  kill and prove the final resident gone AFTER all identity/wake assertions:
  `shell_relink_force_e2e` and `shell_stale_online_e2e`.
- **Kin:** [[IR-7]] (phase-A daemon+brain pair leak), the exe-lock reap-by-path rule, and the
  wrong-lane-diagnostic family ([[IR-21]]'s location-not-build class).
- **SECOND HOLDER CLASS, measured 2026-08-04 (doyle, own gate rig):** an **ORPHANED CONHOST with an
  inherited CWD** inside the test tree blocks `git worktree remove` with Permission-denied on a
  DIRECTORY handle — invisible to any exe-path scan (conhost runs from System32; the parent that
  spawned it was already dead). Found via `.github/ci/find-cwd-holders.ps1` (IR-11's instrument);
  killed by pid after parent-dead classification; 34.78 GB reclaimed. Mechanism chain: the
  windowless spawn path masks DETACHED_PROCESS off, the child owns a console, conhost inherits the
  child's cwd and can OUTLIVE it. Boundary (hertz's, correct): this find and his field-4 companion
  are the SAME MECHANISM FAMILY measured on DIFFERENT populations by different instruments — his
  capture measured a ppid-match COUNT (floor 2, companion identity UNMEASURED, his stated limit),
  and this conhost is the first FIELD identity evidence in the family, on THIS population. Do not
  cite his row as identity data it never carried. (His box-wide accretion count 77→80 is a third
  population again.) Teardown rule
  addendum: a disposal leg must GATE on the removal's exit code (a rig that logs "reclaimed 0.01
  GB" from a failed remove fabricates its own success) and sweep CWD HOLDERS, not only image
  locks — two populations, either alone is a clean zero on the other.
- **Ripe when:** the test-hygiene lane (IR-13/IR-23/IR-36 family) — sweep every e2e that ends
  with a live spawning binary, apply the same reap-after-assertions shape.
- **Size:** small per test; population unknown until swept.
- **Built evidence:** focused runs of both live-at-success tests pass back-to-back in one target
  pool, followed by a build invocation from that same pool; the second test's final kill is polled
  to proven process death rather than treated as fire-and-forget.

### IR-39 — CLASS: the missing-fixture-bin defect has TWO failure faces, and one of them impersonates a lifecycle defect of the subject under test
- **Status:** BUILT 2026-08-18 on KEYSTONE #182 hygiene lane. · **Origin:** hertz + todlando
  2026-08-04, same evening from opposite sides (filed by doyle; hertz relayed todlando's
  suggestion without endorsing scope).
- **What/why:** a rig whose narrowed build (`-p <pkg> --test <one>` / `-p spt --bin spt`) skips a
  cross-package fixture bin fails in one of two ways depending on the consuming site.
  Face 1 (cheap): `translate_proof_fixture` panics "must be built: <path>" — names the artifact
  and the recipe, costs one build. Face 2 (expensive): `sibling_bin("mock-session")` resolves to a
  nonexistent exe, nothing spawns, and the test reds at a PRECONDITION assert whose message
  ("warm bring-up spawned=false", left None / right Some("offline")) reads as a lifecycle defect
  of the subject — measured cost 186s + a triage inside a lent box window (hertz's F1 red arm,
  attempt 1). The build gap itself is the known cross-package row of the fixture-bin population;
  what this entry carries is the DIAGNOSTIC asymmetry.
- **Ask:** a fixture-bin existence precondition that NAMES the missing exe and its build recipe
  wherever a rig consumes one — `sibling_bin` (and the shared
  `crates/spt-term/tests/support/fixture_bin.rs` resolver already carries the recipe string as an
  argument, the shape to copy) refuses with the artifact + recipe instead of letting the consuming
  test fail downstream in its own vocabulary.
- **Kin:** [[IR-21]] (wrong-lane-diagnostic family: the failure names the wrong actor), the
  cargo-builds-package-bins-for-integration-tests rule (cross-package row is the only unguaranteed
  build), golden.yml's dedicated fixture prebuild step.
- **Ripe when:** the test-hygiene lane (IR-13/IR-23/IR-36/IR-38 family) — same sweep population;
  or standalone small if that lane keeps slipping.
- **Size:** small — one shared helper + a sweep of local `sibling_bin` resolver copies (24 known).
- **Addendum (doyle, 2026-08-04 late): third and fourth sightings inside 24 h.** Face 1 twice
  more the same evening — todlando's teach-lane leg set (his own translate_proof_fixture, caught
  by his prebuild), and doyle's cold gate rig for the same lane (first-pass spt-bins red at
  136/611, correctly withheld from the lane verdict as rig evidence, fixed by `cargo build -p spt
  --bins` + rerun 611/611). Four sightings in one day across three agents and both faces: the
  ripeness condition ("the test-hygiene lane keeps slipping") is under measurable pressure — the
  prebuild remedy is being re-derived per-rig per-agent, which is the recurring cost the shared
  precondition helper exists to delete.
- **Built evidence:** all 29 file-local `sibling_bin` resolvers now route through one shared
  precondition. A missing fixture names its path and the package-correct `cargo build -p <owner>
  --bin <name>` recovery command before product assertions or timeout paths run; a focused
  negative test pins that diagnostic.

### IR-40 — CLASS: NO envelope-side ordering signal on a resume brief tracks content recency, and a CI informant's per-job line is a claim about that STEP, not the run
- **Status:** open · **Origin:** three sightings in one night, 2026-08-04/05 — hertz, doyle,
  todlando (filed by doyle; the artifacts are hertz's and todlando's, the CI arm is doyle's with
  deployah's framing). Two arms of ONE family: a stale record that carries a confident ordering
  signal, and a status claim that arrives before the thing it describes.

- **ARM A — resume briefs. What/why:** a session can receive TWO start-of-session briefs, and the
  one routed LATER can carry the STALER project-context. Measured on hertz's pair (same node id):
  routed_at_ms `1785897407373` > `1785896642480`, spill-filename epoch `1785897764588` >
  `1785896658159`, and vector sequence `1242529/1242531` > `1241542/1241544` — **all three
  monotonic orderings rank the staler-content brief as newer.** todlando's pair is the NEGATIVE
  arm: his later brief (`…1785897765196`, routed `…577476`, vector `1242756/7`) genuinely WAS the
  fresher one, so the same heuristic happened to pick right.
- **The finding is UNCORRELATED, not inverted** — and that is the whole entry. An inverted signal
  is usable once you know to flip it; an uncorrelated one reads as reliable until the pair where it
  costs you. Neither agent could have known which pair he held without checking content against
  measurement. **Do not name routed_at_ms specifically** — naming the stamp invites the repair that
  fails ("the timestamp is unreliable, use the sequence number"). The vector is the most dangerous
  of the three: a stamp looks like a clock and invites suspicion, a sequence number looks like an
  ordering and does not.
- **Visible failure mode if a brief is obeyed as instructed** (hertz's, measured): re-open
  `releases#123` — work already ruled closed — and under-report his own claimed-pool count by one,
  the row sitting under a live gate. doyle's own sighting the same night: a brief describing main
  @`92b2482` with a teach lane "mid-flight" hours after every lane had gated.
- **PRE-REFUSED DISCRIMINATOR, recorded so it is not rediscovered as a finding:** file SIZE ranks
  the fresher brief correctly on BOTH pairs (hertz 16308 > 12649; todlando 14138 > 10418) — 2 of 2,
  the only signal surviving both samples. **Do not use it.** Size measures RICHNESS, not recency;
  the two correlated only because todlando's thin brief was a correction fragment, and hertz's
  STALER brief was a full 12649-byte dump. Two full dumps hours apart go flat or inverted. Handed
  over pre-refused by its own finder.
- **ARM B — CI informants (doyle, 2026-08-05, run 30971976024).** `CI-KITSUBITO` sent
  "CI SUCCESS … sha 0a25b77" while the run was measurably `status=in_progress` with
  `conclusion=""` — the EMPTY STRING, not `success`. The informant fires from a per-job `notify`
  step, and at send time `notify` itself had no conclusion. All eight substantive jobs were green,
  so the line was **right in the end but EARLY**. deployah's framing, kept verbatim in kind: *an
  informant line that is right in the end but early is worse than one that is simply wrong, because
  a wrong signal gets distrusted while an early-but-correct one gets promoted to the conclusion.*
  todlando's sentence is the one to keep at the point of use: **a per-job notify step's message is
  a claim about that STEP, not about the RUN.**
- **ARM B MECHANISM — hertz, source read 2026-08-05, zero box cost. It is NOT a race; it fires
  early on EVERY run, and the entry above understated it.** `notify` is a job INSIDE the run
  (`golden.yml:1138`), and a job cannot observe its own run's conclusion — so `conclusion: ""` is
  not a timing artifact, it is **the only value that can exist when the informant speaks. The
  informant has never once made a statement about a run conclusion.** What `ci-notify.sh` actually
  asserts is narrower and worth naming exactly: a verdict computed PURELY from six `needs` results
  (changes, traceability, test, n1-gate, twohost-a, twohost-b — the two matrices collapse to one
  result each) being neither failure nor cancelled. Nothing more.
  **The blind spot is exactly ONE job: `notify` itself, the 9th — it cannot appear in its own
  `needs` list. And that one job has a WITNESSED red path**, documented in the script's own header
  (2026-07-27, twice on kitsubito): verdict computed, then `spt send` HUNG ~5 minutes until the
  job's own 5-minute budget cancelled it, and that cancellation reddened the run. So *"informant
  says SUCCESS, run concludes FAILURE"* is a real observed sequence, not a hypothetical. The 30s
  `SEND_BOUND_SECS` added since bounds the hang but does not close the class — it makes the
  informant's own failure FAST rather than impossible.
  ⇒ The informant's SUCCESS covers 8 of 9 jobs and is blind to the 9th, and the 9th is the only one
  with an observed failure mode that reds the run. This is why the ff predicate is the run's
  conclusion field and never the informant line.
- **RECOVERY, and it is the same for both arms: never rank the claims — re-derive from the source
  object.** For briefs, hertz's four falsifiers are the method (main sha, claimed-pool count, lane
  tip, issue state) and they work because each is a content claim checkable in ONE command against
  a source NEITHER brief controls. For CI, read the RUN's own `conclusion` field, and prefer
  independent reads of the run object (`gh run watch --exit-status` polls the run, so its exit is a
  second measurement rather than an echo of `notify`). On the batch this was caught in, three
  sources agreed with the informant excluded from all three before main was fast-forwarded.
- **Standing rule this entry exists to enforce:** a resume brief's RULINGS are durable; its STATE
  is a HYPOTHESIS to re-derive before acting.
- **Kin:** [[IR-21]] and [[IR-39]] (the diagnostic names the wrong actor / the wrong subject),
  [[IR-32]] (a gate that cannot see its own blind spot), [[IR-34]] (an intermittent marker read as
  a signal).
- **Ripe when:** the next commune/psyche-tier touch for arm A (the routing layer is spt-core's own
  `spool`/psyche ingest, so a remedy is in-repo, not upstream); for arm B, the next `golden.yml`
  informant touch — the cheap fix is that the notify step names its own scope in the message body
  (STEP vs RUN) rather than emitting a bare "CI SUCCESS".
- **Size:** arm B small (message text + scope word in one workflow step). Arm A unsized —
  establishing whether ANY durable content-recency signal can ride the envelope is the measurement,
  and until it runs the answer is "re-derive, do not rank."

### IR-41 — A QUEUED main run is superseded without a record; `cancel-in-progress: false` protects only STARTED runs
- **Status:** open — mechanism CONFIRMED (doyle ruling 2026-08-05); the runbook sentence it
  falsified is already corrected (`docs/RELEASE-RUNBOOK.md` step 3, same commit as this entry)
  · **Origin:** deployah, v0.55.0 publish night, filed by him explicitly as UNPROVEN with the
  manual-cancel alternative not ruled out; confirmed by doyle on the run objects.
- **What/why:** `ci.yml` sets `cancel-in-progress: false` on main and the runbook read that as
  "main's concurrency policy never cancels this run; that preserves the record." FALSE for the
  queued phase: a concurrency group holds at most one running + one pending run, and a newer push
  REPLACES the pending run regardless of the flag — the flag governs only whether a STARTED run
  is cancelled. A superseded queued run leaves NO record: zero jobs, conclusion `cancelled`,
  nothing measured about its sha.
- **Evidence (all read off the run objects, not the informant):** run 30973909279 at `327f1f8`
  (push, main) — created 04:01:08Z, **zero jobs ever started**, cancelled 04:03:49Z, ONE SECOND
  after `e8805f7`'s run 30974041521 entered the group (created 04:03:48Z), while `0a25b77`'s run
  30973770039 was still in flight (done 04:07:59Z). One-running-one-pending, newer push landed,
  pending run died. The 1s coupling to an unrelated push is the discriminator against deployah's
  own manual-cancel alternative — a human cancel co-timed to the second with a push it had no
  view of is not a credible mechanism; supersession fires on exactly that trigger by design.
- **The hole this leaves:** under a rapid push sequence, an intermediate sha on MAIN can have no
  thin-CI verdict AT ALL, and the absence reads as nothing rather than as supersession. Anyone
  auditing "did sha X pass main CI" gets a hole where the runbook promised a record. Absence of
  a run is now a labelled state, not evidence about the sha.
- **Ripe when:** next ci.yml/runbook touch. Candidate remedies to weigh THEN, not now: accept and
  document (done — the runbook now carries the mechanism), or give main's group per-sha keys
  (`group: ci-${{ github.sha }}`) so runs never share a group — at the cost of concurrent main
  runs competing for the boxes, which is exactly what the group exists to prevent. The trade is
  real; do not fold it into a drive-by.
- **Size:** docs half DONE in this commit; ci.yml half small but load-bearing — needs its own
  lane if taken.
- **Kin:** [[IR-40]] (a record that is right-in-the-end-but-early vs a record that never comes to
  exist — both read as clean states unless labelled), [[cancelled-measurement-leaves-labelled-hole]].

### IR-42 — `pool-claim` writes a record; only the BUILD enforces — and AGENTS.md invited the misread
- **Status:** BUILT 2026-08-19 — the AGENTS.md corrective line lands in the SAME COMMIT as this
  entry (the register batch commit), which is the entry's stated exit condition. The optional
  code-side nicety (claim verb prints the incumbent record it overwrites, information only)
  remains unclaimed — a candidate rider on the IR-26 remedy lane, which already touches
  `pool_claim`'s read-before-write path, NOT a reason to keep this entry open ·
  **Origin:** doyle + todlando independently at source, 2026-08-19, during the #193 gate.
- **Mechanism (source-verified):** `xtask pool-claim` (crates/xtask/src/main.rs:2232-2294 at the
  read sha) parses its args, reads the lane's own git identity, builds a `PoolOwner` and calls
  `spt_poolguard::write_owner` UNCONDITIONALLY (:2280) — no `read_owner`, no verdict, no
  comparison against the incumbent claim anywhere in the function. Claiming is last-writer-wins
  by construction: two lanes can each "hold" a pool in sequence with only the last write
  surviving, and the displaced lane learns nothing at displacement time. All four enforcement
  arms — `Refuse` (SPT_POOL_FOREIGN), `Takeover` (loud), `Unproven`, `HatchOpen` — live in
  crates/spt-store/build.rs:31-113 and speak only at the next BUILD. (Same object [[IR-26]]'s
  CLAIM-OBSERVABILITY measurement saw from the displaced side; this entry is the CLAIMING side
  plus the doc that invited the misread.)
- **Measured consequence:** 2026-08-19, doyle's gate claim over todlando's unlanded
  redeem-echo-diag claim — no refusal, no notice (three sequential overwrites that session, one
  over a lane with an unlanded commit). todlando then predicted from AGENTS.md's refusal
  sentence that the gate's claim "will be REFUSED" — a false warning to a gater mid-run. The
  refusal sentence sat directly under the "Claim a pool at lane start" instruction, inviting the
  build-time semantics to be read onto the claim verb.
- **Corrective landed:** one AGENTS.md sentence — the claim WRITES a record and never refuses;
  predict refusals from builds, never from `pool-claim`.

### IR-43 — knock answer carrier: `NoReply` landed at 60.08s against a stated 30s carrier deadline
- **Status:** open, OBSERVATION — mechanism unmeasured, filed exactly as wide as the datum ·
  **Origin:** KNOCK-169 RCA layer 1 (todlando, read-only), golden 32191618042 @`93c1130`,
  twohost-b, 2026-08-18.
- **What/why:** B's `request_answer` yielded `NoReply` at 22:50:08.3605Z — 60.08s after A's
  death — while the carrier deadline is stated at 30s. Recorded in the RCA as a side observation
  with no claim; the cell's red itself was a CASCADE of A's earlier death (answerop.rs:73-81
  produces the courtesy BEFORE returning, so the reply was never sent — that half is settled and
  is NOT this entry). What this entry carries: a transport/liveness bound that reported at 2x its
  stated value. Candidate shapes, neither asserted: two sequential 30s legs each honoring its own
  bound (a composed wait wearing one bound's name), or a bound armed after a wait it does not
  cover (the [[IR-30]] nethost.rs permit-wait shape: `sem.acquire_owned()` outside the timeout).
- **Ask:** derive where the 30s is armed and what path yields a 60s `NoReply` — one code read
  from the emit site, before any instrument.
- **Ripe when:** next knock-carrier touch, or the next `NoReply`-shaped red — whichever first.
- **Size:** small (code read; a token only if the read forks).

### IR-44 — perch-sentinel comment overclaims: serialization is not preservation, and the comment teaches the false step
- **Status:** open · **Origin:** #188 RCA lock read (todlando flagged, doyle confirmed the lock
  read), 2026-08-18; the collapse-of-branches argument is on #188 and in the RCA record.
- **What/why:** `lock_perch_sentinel` (info.rs:778 at `b20e770`, the diag sha — re-derive the
  line on main before editing) takes a cross-process fs2 lock on a stable per-perch `.info.lock`
  sentinel, and its comment claims it serializes "ALL info.json writers so a whole-record write
  and a locked RMW can never lose each other's update (`REQ-HAZARD-INFO-RMW-LOST-UPDATE`)" —
  TRUE for interleaving, FALSE for preservation. `write_info_unlocked` is module-private with
  exactly three callers, all taking the sentinel first (bypass refuted by visibility), but every
  caller of the public `write_info` composes its record BEFORE the lock is taken (:842) — **the
  lock makes that write atomic; it cannot make it preserving.** Only `mutate_info` /
  `establish_locked` read under the hold and can preserve. Measured field case: census write #1
  (`twohost.rs:736`) clobbered `controlled` while correctly locked. A reader trusting the
  comment infers lost-update immunity the funnel does not provide — the #188 hunt burned a
  branch on exactly that inference before the collapse argument killed it.
- **Ask:** respell the comment — serializes writers; a composed `write_info` is atomic, never
  preserving; preservation requires `mutate_info`/`establish_locked` — and weigh whether
  `REQ-HAZARD-INFO-RMW-LOST-UPDATE`'s doc wording carries the same overclaim.
- **Ripe when:** next spt-store/info.rs touch. **Size:** tiny (comment + possibly one REQ doc
  line).

### IR-45 — twohost rig home premise, SETTLED: pump paths resolve under the per-run TEMP root (rig is disk-hermetic; the "live fleet roster" reading is retired)
- **Status:** RETIRED 2026-08-19 — ask answered by the one code read at main @`78a9a16`; residue
  re-homed (see Answer). · **Origin:** #188/#189 RCA readouts (todlando, doyle's 2a-2d), runs
  32097943571 + 32108557362 + 32111921251, 2026-08-18.
- **What/why, two measured halves in tension:** (1) B's pump dialed a THIRD node 155-159 times
  per run at kitsubito's OWN tailscale IP (`100.98.197.12`), different hex + different ephemeral
  port each run — a short-lived local endpoint re-minting identity between runs, present in both
  the golden and the diag run. Read at the time as: the rig uses the CANONICAL home
  (`canonical_pump_paths` → `perch::spt_home()`), so B's roster is kitsubito's live fleet
  roster — environmental coupling of a CI rig to host fleet state. (2) The ER perch measured in
  the chain runs resolves under a per-run TEMP root
  (`…/_temp/spt-test-tmp-<run>…/owlery/engine-room`), NOT the canonical home. Both measurements
  stand; they answer different artifacts (pump roster vs perches), and nothing on record settles
  which home the PUMP paths actually resolve. #189's third-node evidence and its
  cache-leg-never-heals reading rest on that premise (the tension is filed on the #188 record as
  an open question against #189).
- **Answer (2026-08-19, code read at `78a9a16`, unambiguous — no run token needed):** both role
  tests set `SPT_HOME` to a fresh per-run `TempDir` as their FIRST act (`twohost.rs:813-814` role
  B, `:1501-1502` role A), before any store touch. `canonical_pump_paths` resolves every path via
  `perch::spt_home()` (`twohost.rs:396-414`), which honors the `SPT_HOME` override first
  (`perch.rs:34-41`); the pump's roster source `presence::registry_snapshot_dir()` →
  `perch::identity_dir()` sits under the same root (`presence.rs:83-85`). The rig is
  DISK-HERMETIC per run; both prior measurements reconcile (ER perch observed under the temp
  root, pump paths resolve there too). The misleader was the helper's name + doc comment ("real
  homes, not test roots"), which describe production path LAYOUT, not the resolved ROOT.
- **Residue disposition:** (1) #189's third-node evidence was corrected on the board (comment
  5337519936) — the roster row arrived at RUNTIME into a temp-rooted store, ingress mechanism
  OPEN on #189, no longer explained-environmental. (2) KEYSTONE #182's hygiene lane repaired the
  `twohost.rs:394` comment to say exactly what the helper does: production path layout rooted in
  the process's fresh per-run `SPT_HOME`.

### IR-46 — the Windows disk-floor preflight asserts an INSTANT; workspace free space moves tens of GB inside the hour, so a green floor is not a claim about run headroom
- **Status:** open — mechanism CONFIRMED by a measurement series on hfenduleam; doyle ruled it
  register-shaped 2026-08-18 · **Origin:** deployah, NAMEPLATE #181 golden-head intake, on the
  run that the floor red-floored · **Filed as IR-46 after an id collision:** this row was authored
  as IR-42 and held unpushed under the v0.56.0 tag freeze, during which the release-close sweep
  (`78a9a16`) issued IR-42..45 to other findings. Parallel issuance, not authored disagreement —
  and noted here so a later reader chasing "IR-42" in this row's history does not hunt for a lost
  version of it.
- **What/why:** `golden.yml`'s Windows preflight computes `$freeBytes` once at job start and
  hard-fails under a `32GB` floor. As a fast-fail that is correct and it did its job — it refused
  before compiling rather than dying deep in a link step. The defect is what a PASS is then read
  to mean. The reading is a point sample of a quantity that moves by tens of GB unattended, so a
  green preflight licenses a run whose headroom was never measured. Nothing downstream re-checks.
- **Evidence — five readings of `C:` free on one box inside roughly one hour, 2026-08-18:**
  - `03:01:54Z` — `free_bytes=2666205184` (2.48 GiB) against `floor_bytes=34359738368`, the
    preflight's OWN reading; run `32093894524`, both `n1-gate` and `test` on Windows/hfenduleam
    failed at this same step, before any compilation.
  - `~03:2xZ` — ~2.50 GB, independent `Get-PSDrive` read, corroborating the preflight.
  - pre-reap — **33.70 GB**, immediately before deployah deleted anything.
  - post-reap-1 — 76.42 GB (42.72 GB reclaimed: `claude_skill_owl/target`, `golden-w1/target`).
  - pre-reap-2 — **73.81 GB**, i.e. 2.61 GB consumed unattended between the two reaps; then
    post-reap-2 129.70 GB (55.89 GB reclaimed: `.worktrees/nameplate-w1/target`).
- **LABELLED HOLE — do not let this harden:** roughly **31 GB was released between the preflight
  failure and the pre-reap reading by something OTHER than the reclaim**. Released-by-unknown.
  The runner cleaning its workspace after the job died is *plausible* and is NOT the recorded
  cause — it was never measured, and no one looked while it was happening. Recorded as a hole on
  purpose (doyle's explicit ask at filing): a plausible cause written down as the cause would make
  this row read as explained when the actual mechanism of the 31 GB is unknown.
- **The hazard this leaves:** a run can clear the floor at preflight and starve mid-build, and
  mid-build disk starvation does not present as disk — it surfaces as a link failure, a truncated
  artifact, or a rustc ICE, i.e. as a defect in the tree under test. The failure mode inverts the
  gate's purpose: the preflight's whole point is to keep box conditions from being read as lane
  reds, and a passing preflight actively argues the opposite. Note the polarity — this row is
  about the PASS, not the FAIL. The observed red was honest.
- **#242/v0.66.0 GOLDEN ADDITIONS (deployah measured 2026-08-29, banked at the 08-30 close
  sweep):** (i) per-job floor figures are SEPARATE INSTANTS — quote per job, never one figure for
  both (attempt 1: `test` 26521477120 vs `n1-gate` 26535051264 free bytes, 8s apart, one run);
  (ii) **CORRECTED TRAP** (retracting deployah's earlier "attempt-2 hides jobs" version): the
  default `/runs/<id>/jobs` endpoint stamps `run_attempt=<current>` on EVERY job including ones
  that never re-ran (4 of 8 mis-assigned on this run) — it is the RUN's attempt, not the JOB's,
  so a stale green can read as re-earned; `started_at` is the discriminator and
  `/runs/<id>/attempts/N/jobs` the authority. Also `rerun --failed` re-runs DEPENDENT jobs
  (twohost re-ran despite a green attempt 1); (iii) a rerun preflight reads RAW free BYTES with
  the CI step's own predicate, not display GB; (iv) a rerun green counts as floor-RCA-closing
  evidence only after a step-level NON-VACUOUS check (checkout SUCCESS + workload steps ran) —
  the v0.64.0 checkout-SKIPPED shape means a green job can carry zero code signal.
- **Ripe when:** next `golden.yml` touch. Remedies to weigh THEN, not as a drive-by: re-assert the
  floor at job END (turns a starvation into a labelled disk verdict instead of a fake lane red);
  or raise the floor to cover a full build's peak rather than its entry condition; or sample free
  space across the run and emit the minimum as a token. All three cost run time; none is obvious.
- **Size:** small in `golden.yml`, load-bearing in what a green run is taken to prove.
- **Kin:** [[IR-41]] and [[IR-40]] (a state that reads as a clean verdict unless it is labelled —
  here a PASS that is read as headroom it never measured), [[measure-the-box-before-the-instrument]],
  [[cancelled-measurement-leaves-labelled-hole]].

- **2026-09-06 addendum (doyle, v0.67.1 close sweep): trigger FIRED at `04e32c8c` — `golden.yml` touched (2-line HEAVY reclass of `brain_resume_conn_deadlock` on the respin) with no remedy weighed. Not a drive-by by choice: the touch rode a respin under a pre-registered hard stop, which is no remedy window. COMPOSED with IR-59 (log-the-floor) + IR-73 (literal-first sites) as one hertz workflow rider on WEBSERVE #272; the remedy weighing happens on THAT lane, before the WEBSERVE golden head.**

<!-- [doc->REQ-CI-FREE-SPACE-PREFLIGHT] -->
- **2026-09-06 remedy ruling (doyle, WEBSERVE rider):** retain 32 GiB; record raw-byte
  `FLOOR_START` and `FLOOR_END` PASS/RED tokens per job/runner, each labelled `sample=INSTANT`.
  End DISK assertions run `if: always()` before teardown, including after failed workloads.
  A fresh `FLOOR_DOCS` assertion immediately before docs-drift composes with IR-73's second half.
  Raising the floor is refused without a measured peak; continuous sampling adds lifecycle
  machinery while remaining blind between samples. Start/end are a dataset, not a minimum or
  headroom guarantee. Golden execution evidence waits for the WEBSERVE golden head; the thin
  PR must first show actual START/END output on its own green CI run.

### IR-47 — ci-notify treats a missing co-author trailer as a quiet info line; silent attribution loss reads as "there was none"
- **Status:** open — patch AUTHORED and stashed unlanded since 2026-07-29
  (`.worktrees/_patches/ci-notify-missing-trailer.patch` + its `test-ci-notify.sh` harness in the
  sibling `.untracked` dir); surfaced by doyle's 2026-08-19 idle-queue verification of that stash ·
  **Origin:** the AGENTS.md trailer mandate's own hazard family (the space-spelling trailer is
  structurally invisible to git's tokenizer, so confident zeros already read as attribution loss
  once); this row is the NOTIFY-side twin.
- **What/why:** `.github/ci/ci-notify.sh` parses the head commit for the line-anchored
  `Co-authored by: <agent>` trailer to add a second notification recipient. When the trailer is
  absent or unparseable, the current script emits one stdout info line ("doyle only") and moves
  on — correct as a non-fatal outcome, wrong as a SILENT one: a lane that lost its attribution
  (typo'd trailer, squash that dropped the body, hyphenated spelling) notifies doyle alone on
  every run and nothing anywhere says an agent stopped being told about their own lane's verdicts.
  The stashed patch keeps absence non-fatal but promotes it to a `::warning` workflow annotation
  (visible on the run summary), and splits the legitimate quiet case (co-author IS doyle) from the
  loss case so the two stop sharing one message.
- **Evidence:** patch verified 2026-08-19 to still NOT be on main — the warning string is absent
  from `.github/ci/ci-notify.sh` at the current tree. Patch content re-read at verification; it
  applies against the script's post-`SEND_BOUND_SECS` shape.
- **Ripe when:** next `golden.yml`/ci-notify touch — natural co-rider with [[IR-46]]'s remedy
  window (same file family, same "weigh then, not drive-by" rule). The stashed test harness rides
  with it.
- **Size:** ~10 lines in one shell script plus its test.
- **Kin:** [[IR-40]] (an unlabelled state read as a clean verdict — here a quiet info line read as
  "no co-author existed"), the AGENTS.md `%(trailers:)` tokenizer mandate (the sibling silent-zero
  on the AUDIT side).

### IR-48 — the nextest.toml parity cell tolerates CRLF only by parser accident; a newline-sensitive arm added later reds fresh Windows checkouts alone
- **Status:** open — filed 2026-08-19 (doyle route, hertz verdict: REGISTER as preventive
  hardening, not current defect) · **Origin:** sibling sweep after the W3 gate CRLF finding —
  parent mechanism: `include_str!` embeds working-tree bytes verbatim, `core.autocrlf=true`
  smudges fresh checkouts CRLF, and the author's never-re-smudged tree stays LF — builder green,
  every fresh rig/golden checkout red.
- **What/why:** `crates/xtask/src/main.rs` `the_checked_in_config_is_in_parity` include_str!s
  `.config/nextest.toml` and feeds `phase_a_overrides_missing_from_ci_windows`. TODAY this is
  causally CRLF-tolerant — the parser is `str::lines()`+trim (`override_filters`, main.rs:434-440),
  and the 52/52 fresh-rig green at the hygiene gate (2026-08-19, this box) is therefore causal, not
  luck. The hazard is the MISSING CONTROL: nothing pins the tolerance, so a future
  newline-sensitive arm in the parity predicate reds only on fresh Windows checkouts — the
  builder-green/rig-red inversion, deterministic but reading as flake.
- **Remedy (hertz's, platform-independent):** feed the parity predicate an LF fixture AND the same
  fixture converted to CRLF, assert identical missing-sets (including one unmirrored negative), so
  a newline-sensitive arm reds on EVERY host rather than only fresh Windows checkouts.
- **Ripe when:** next xtask parity-cell touch; explicitly OUTSIDE the family-B lane (hertz
  scoping). **Size:** one fixture-pair cell.
- **Kin:** the W3 gate finding this swept out of (skeleton cells, fixed at the fixture edge in
  lane commit e99633f); [[IR-40]] (an untested tolerance read as a guarantee).

### IR-49 — poolguard's landed-lane predicate is ancestry-only, so a cherry-picked member lane can never read Settled and its pool can never be taken over
- **Status:** open — filed 2026-08-19 (todlando measurement at the keystone-w1/fix-197 reap,
  doyle ruling; doyle re-verified the predicate at source before filing) · **Origin:** todlando's
  post-landing landed-check used the guard's own predicate and it contradicted the (correct)
  landed fact — the checker was wrong with the guard, which is exactly how the guard will be
  wrong alone.
- **What/why:** `git_lane_state`'s final arm answers Settled/InFlight purely by
  `merge-base --is-ancestor <lane tip> <integration head>`
  (`crates/spt-poolguard/src/lib.rs:424-436`, probe at `:474-479`). Under ADR-0050 golden CI
  (thin member lanes cherry-picked onto a stage branch, ff-only main) a MEMBER lane's own branch
  tip never becomes an ancestor of main. Measured at origin/main `9ea595c`:
  `build/keystone-182-w2-sealed-multi` @`ba2d998` and `fix/197-absent-code-uncounted` @`5eb6a3b`
  both read NOT-ancestor while `git cherry -v` reports every lane commit `-` (already upstream).
  So `LaneState::Settled` is unreachable for any cherry-pick-assembled member lane whose branch
  still exists, the takeover arm can never fire, and the pool refuses forever — including with a
  dead holder. That contradicts the AGENTS.md mandate ("a merged branch is taken over") because
  the guard's notion of merged is ancestry-only and the golden model structurally never produces
  ancestry for member branches. The branch-vanished (`:391`) and re-pointed (`:399`) arms still
  fire, and the STAGE lane works (its tip IS an ancestor) — the hole is landed-member-branch-
  still-exists only. It did not bite at the 2026-08-19 reaps solely because deleting a pool
  deletes its claim; it bites the first time someone wants to TAKE OVER a landed lane's pool
  rather than reap it.
- **Remedy (todlando's candidate, doyle-endorsed direction):** when ancestry answers false, fall
  back to patch-id containment (`git cherry` / `rev-list --cherry-mark` against the claimed
  base) before concluding InFlight. **Known residual to state in the fix's docs/tests:** a
  conflict-adjusted pick changes the patch-id, so such a lane still reads InFlight; its release
  path is branch deletion (the vanished arm), which is at least reachable. Tests want a
  cherry-picked-landed fixture plus a conflict-adjusted negative.
- **Ripe when:** next poolguard touch, or the first landed-lane pool takeover request.
- **Size:** one fallback arm in `git_lane_state` + two fixtures.
- **Kin:** [[IR-26]] (dead-holder takeover AUTHORIZED when it shouldn't be — this is the mirror:
  takeover UNREACHABLE when it should fire), [[IR-42]] (claim writes, build enforces), the
  AGENTS.md pool mandate this contradicts.

### IR-50 — e2e failure panels read the PRE-REDIRECT stderr capture; every daemon-side diagnostic lands in a sink the rig deletes unread
- **Status:** CLOSED — BUILT 2026-08-19; adoption-completion lane
  `docs/ir50-register-close` @`e02f565f` (hertz, doyle-gated 2026-08-20) LANDED in
  v0.59.0 as a PORTER rider (pick `e702d4b7` tip, main @`c62904e7`; register sweep
  2026-08-21) — lane `fix/ir50-stderr-sink-census` @`1a2f6a2`+`2ac9543`
  (hertz; base `0c86b2d`; rides the CONCIERGE #183 chain; these shas are the 2026-08-19
  message-only reword of `cb4cd9d`+`7336e03` — trees proved identical by tree-id, so every
  measurement below carries): exported
  `spt_daemon::stderrlog::sink_path(home)` with `stderr_log_path` routed through it +
  equality unit; `daemon_stderr_panel` in tests/common renders BOTH channels and labels a
  read failure with channel + exact looked-at path + OS error (absent-sink regression pins
  it — the IR-40 unlabelled-absence class was caught at gate and fixed in `2ac9543`);
  censused every daemon-run captured-stderr panel outside `engine_room_bringup_e2e` (rides
  #199 lane) and `er_briefing_*` (rides #164 lane); engineroom.rs:145 misnomer rider landed
  in the same lane. Gate figure 781/781 at `2ac9543` (measured at the pre-reword tree,
  which is byte-identical; first-report 765 was transcription,
  corrected with raw line + JSON provenance). CLASS FOLD (doyle ruling, no fresh row):
  `er_briefing_presented_e2e` hit this exact class during its own authoring (blank panels
  on a real loud line) and was fixed locally in `72efb6a` reading both channels.
  **REGISTER CLOSED 2026-08-19:** branch `docs/ir50-register-close` links the
  golden-tested implementation at `4661bc9d` and completes helper adoption in the five measured
  residual files: `er_brief_once_per_session_e2e.rs`, `er_briefing_presented_e2e.rs`,
  `er_sequestered_cwd_e2e.rs`, `endpoint_autostart_e2e.rs`, and `n1_pairing.rs`.
  `engine_room_bringup_e2e` remains the explicit exclusion, recorded as todlando's direct
  sink-path read from the #199 lane. ·
  **Originally filed:** 2026-08-19 (todlando RCA inside the #199 investigation; the immediate
  four-panel fix in `engine_room_bringup_e2e.rs` is his and rides the #199 lane — THIS entry is
  the class census, hertz-class) · **Origin:** the #199 evidence void — 40 rca-v2 runs held zero
  daemon-side evidence for either erhost face; `ENGINE_ROOM_SPAWN_FAIL` (broker.rs:5658) was
  written to a log the rig destroyed at teardown. Absence discipline honoured: the SUCCESS-path
  line (`ENGINE_ROOM_BROUGHT_UP`) was also absent from all 40, proving the channel dead rather
  than the event absent.
- **What/why:** product `stderrlog::install` (broker at cli.rs:7633, brain at brainproc.rs:200)
  repoints the process STD_ERROR_HANDLE within its first statements (`redirect_stderr_to`,
  stderrlog.rs:101 — Windows `SetStdHandle` + `mem::forget`), so a rig's `.stderr(File)` capture
  holds only the pre-redirect window; from that line on every diagnostic lands in
  `SPT_HOME/logs/daemon.stderr.log` (stderrlog.rs:34/42) inside the rig's temp home, destroyed
  unread at teardown. Panels are also mislabelled ("brain stderr" holding the broker's first
  lines). REQ-DAEMON-STDERR-PERSIST fixed this blindness PRODUCT-side ("the incident-night
  RCA-blind gap", cli.rs:7626-7631); no rig was ever taught to read the sink, so the fix does
  not reach any harness. Class: every e2e rig that spawns `spt daemon run` and prints a
  captured-stderr panel has the same void.
- **Remedy:** census all such rigs; failure panels additionally dump
  `SPT_HOME/logs/daemon.stderr.log` with the path read FROM `spt_daemon::stderrlog` (never
  re-spelled), and the labels corrected ("broker stderr (pre-redirect)" / "daemon stderr sink").
  `engine_room_bringup_e2e.rs`'s panel sites land in the #199 lane (todlando, scoped there —
  measured shape: FIVE producers incl. the :167 precondition panic, SEVEN print sites, seven
  relabels; the early "four" was an estimate the file refuted). ⚠ Rigs can only import
  `STDERR_LOG_BASENAME`; the `"logs"` dir segment has no exported const, so every rig
  re-spells it — the census lane should add an exported path helper (e.g.
  `stderrlog::sink_path(home)`) to spt-daemon FIRST, then consume it everywhere. The census
  and remainder are hertz-class.
- **Ripe when:** hertz's queue after current items, or the next e2e red printing an empty
  stderr panel.
- **Size:** census + mechanical panel edits per rig.
- **Kin:** [[IR-40]] (an absent signal read as a clean verdict), REQ-DAEMON-STDERR-PERSIST (the
  product half of the same incident), [[IR-8]] (a censused zero that cannot see its blind spot).

### IR-51 — engine-room e2e daemon runs NET-ENABLED on real interfaces; rig hermeticity is disk-deep only
- **Status:** open — filed 2026-08-19 (todlando's x40 run-4 sink read, the first live read any
  rig has had of the daemon sink; doyle ruled own-id rather than riding #199, at the finder's
  own request) · **Origin:** unasked-for find inside the #199 instrument's first catch.
- **What/why:** the rig's temp `SPT_HOME` buys DISK hermeticity ([[IR-45]]'s settlement — which
  explicitly covered paths, not network) but nothing turns the net off: the run-4 sink shows
  the rig's own daemon with `BRAIN_NET_CONSUMERS_UP` (dispatcher + peer pump started), three
  `NET_FAMILY_GATE: binding IPv4-only`, and three `PAIR_MEET_UP:erhome` carrying the box's REAL
  Tailscale (100.68.35.65) and LAN (192.168.1.81) addresses — live peer work during an e2e —
  plus an `INBOUND_REACHABILITY` warning whose firewall rule admits the CI-runner's binary path
  (`C:\actions-runner\_work\spt-bs-core\spt-bs-core\target\debug\spt.exe`) while this daemon
  ran from a dev tree. Consequences: test behavior conditioned on shared-box network state;
  rate-shaped flakes; e2e runs radiating real traffic. SAME SINK, RECORDED NOT DIAGNOSED: conn
  churn at ~150 ms cadence (327 write-start / 324 transport-close, conn ids past 300 in 58 s) —
  unattributed. ⚠ CANDIDATE INTERACTION, explicitly NOT established (finder's own framing
  honoured): the run-4 face (30 s bound blown, zero session events in the window, launch Ok, no
  error anywhere) matches the SHAPE of the filed offline-but-resolvable-peer pump-stall
  mechanism (dial does not fast-fail). Discrimination open; nothing here is a #199 attribution.
- **Remedy:** a rig-honoured net-off switch (peer pump + discovery disabled under test unless
  the test is ABOUT them) or loopback-only binding; census which e2e rigs start net consumers.
  Sequencing rule: do NOT fix hermeticity into #199's face before the face is attributed — a
  green bought by turning the net off would bury the mechanism unread.
- **Ripe when:** #199 attribution answers whether net state is causal; else the next
  hermeticity pass.
- **Size:** product switch + rig adoption + census.
- **Kin:** [[IR-45]] (the disk half), the subnet-pump dial-does-not-fast-fail mechanism (filed,
  pump-stall RCA), [[IR-38]] (a live daemon's side effects outliving the test's intent),
  releases#125 gc-spin (a conn-churn shape candidate).

### IR-53 — failed-address expiry unit panics during the first 10m01s after Windows boot
- **Status:** BUILT 2026-08-21 on `fix/ir53-cold-boot-instant` — LANDED in v0.59.0
  (PORTER rider, pick `7b26f28d` on `assembly/porter-205`, main @`c62904e7`; register
  sweep 2026-08-21 at release close) · **Origin:** todlando,
  found and falsifiably measured mid-releases#204; this rider fixes the latent CI-pipeline
  flake before PORTER golden.
- **Measured:** `failedaddr::tests::an_expired_record_stops_matching_and_is_swept` ages an
  `Instant` with `Instant::now().checked_sub(FAILED_ADDR_TTL + 1s).expect(...)`, where
  `FAILED_ADDR_TTL = 600s`. Windows `Instant` counts from boot, so a host younger than 601s
  cannot represent the requested stamp and panics `monotonic clock older than the TTL`.
  Todlando predicted and observed the red before the threshold, then observed the identical
  tree green after it; two subsequent daemon sweeps held 860/860.
- **Remedy:** when the monotonic clock cannot represent the age, return through the loud
  `SKIP_FAILED_ADDR_EXPIRED_SWEEP` arm naming the exact unproved property. On every warm host
  the original direct aging and both expiry assertions still run unchanged; there is no
  ten-minute sleep and no silent pass on the cold arm.
- **Class census (untruncated):** four test-side
  `Instant::now().checked_sub(...TTL/deadline...).expect(...)` patterns total. This cell is
  the sole 601s member. Three siblings remain recorded, not changed in this rider:
  `broker.rs:8432-8434` and `broker.rs:9198-9200` age
  `CONTROLLER_WRITE_DEADLINE + 1s` (6s), while `broker.rs:10385-10388` ages
  `BRAIN_WRITE_DEADLINE + 1s` (16s). Their much narrower cold-boot windows are the same
  mechanism class and should use a controlled clock seam when next touched.
- **Kin:** releases#204 (finder lane), [[IR-25]] (daemon-lib flake population), and
  [[IR-40]] (a confident verdict without the evidence needed to make it).

### IR-52 — docs-site CLI reference renders SHALLOW; help text below the rendered depth is outside the drift gate, unit-held per-REQ or held by nothing
- **THIRD INSTANCE, observed 2026-08-24 (todlando, WAX-SEAL #21 W4; not acted on):** the new
  `spt api seal verify|describe` FLAGS sit below the generator's rendered depth — the entry's
  exact class. The W4 user page hand-documents them, which is per-page self-defense, not a
  gate. Conditional-rider condition did NOT fire at W4 (no generator touch); the entry's own
  trigger stands for whichever lane next touches the generator.
- **Status:** BUILT 2026-08-24 on SIGNET #218 W1 T6 (todlando; doyle's intake made it the wave's
  rider when the new `seal enroll-authenticator` verb forced the generator touch that armed the
  trigger): `xtask` reference-gen now DESCENDS THE FULL COMMAND TREE recursively (the durable
  remedy), regen published the 22 previously-unrendered leaves + the new section, and a depth
  regression is a drift RED by construction. Self-holding units censused per the remedy: the
  SURFACE-SECTION-SITED walk and both MONIC-TRIGGER-SECTION sited-help units are now DOUBLE-held
  (published surface under the drift gate + content pins kept) — kept, not retired. · **Origin:**
  todlando's #160 build report, fact flagged for the gate rather than folded silently.
  (Pre-build status: open, filed 2026-08-19 — doyle, W3 #160 gate prep; second measured instance
  of a class REQ-CLI-SURFACE-SECTION-SITED's own title already named for its sites.)
- **What/why:** the generated `docs-site/src/cli/reference.md` renders `spt endpoint monic`
  but does not descend to `monic add` / `monic update`, so the #160 strike of the maintained
  kind enumeration at those two flag docs never appeared in the published reference and the
  docs-drift gate CANNOT see the strike — the REQ-CLI-MONIC-TRIGGER-SECTION rendered-help
  unit is the only hold. General mechanism: any help text below the generator's descent
  depth is invisible to the drift gate; a deep help edit (or regression) ships unpublished
  and ungated unless some REQ's unit happens to pin it. First measured instance:
  REQ-CLI-SURFACE-SECTION-SITED names "the three-deep sites the docs-site drift gate cannot
  reach" and closes them with its own pinned walk — per-REQ self-defense, not a gate.
  Consequence beyond drift: adapter builders build BLIND from the public docs (DRI
  protocol), so a deep-help divergence hands every adapter a thinner contract than the
  operator's terminal shows.
- **Remedy:** either xtask reference-gen descends the full command tree (deep help becomes
  published surface the drift gate already covers), or a generic rendered-help walk over all
  sites at all depths joins the drift gate; census which REQ units currently self-hold deep
  sites (known: SURFACE-SECTION-SITED walk, MONIC-TRIGGER-SECTION sited-help unit).
- **Ripe when:** next xtask docs-gen touch, or the first adapter filing traceable to a
  deep-help divergence.
- **Size:** xtask render change + regen + census of self-holding units.
- **Kin:** the REQ-CLI-SURFACE-SECTION-SITED pinned walk (the per-REQ closure shape this
  entry generalizes), [[IR-47]] only if its golden.yml/ci-notify surface rides the same
  docs-gate window.

### CI-RIDER LANE STATE — LANDED 2026-08-04 (post-v0.53.0 merge queue); history below kept for its mechanisms
- **RESOLVED:** the lane rebased clean onto the post-tag queue and landed ff-only as
  `2b33a47`/`2a7b016`/`d63f4ce`/`19d7f79` (content byte-identical to `1275e47..bf8c4a2` by lane-diff
  blob hash). IR-1 and IR-4 are BUILT AND LANDED; IR-9's measurement half is landed with its trigger
  now armed — the toolchain print reports runner-account versions on the NEXT GOLDEN RUN, which is
  when doyle's 2026-08-03 ruling (decide from runner-account versions, never the interactive prior)
  becomes executable. The two golden-only steps (link probe both boxes, toolchain print both legs)
  remain unexercised until that run — the landing does not change that caveat.
- **As recorded pre-landing (2026-08-04, doyle):** `ci/locksmith-riders` was at `bf8c4a2` with FOUR
  commits not contained in `origin/main` and not present in LOCKSMITH's golden head `b7b00c3`. The
  worktree `.worktrees/hertz-ci-riders` was clean, so the work existed and was simply unlanded:
  `1275e47` (xtask: assert a load-bearing patch pin is still in force, both ways it lapses) ·
  `49d4805` (docs/ci: a lock-touching lane reads edges, not just the package set) ·
  `8ed006b` (ci/golden: print the toolchain that judged the run, both legs) ·
  `bf8c4a2` (ci/golden: measure the link before rendezvous, read the floor after the checkout that
  clears it — the floor read AFTER its own reclaim is [[ci-runner-has-no-warm-target]]'s shape).
- **Why it is recorded rather than quietly re-dispatched:** [[IR-1]]/[[IR-4]]/[[IR-9]] were carried
  as DISPATCHED, which is now false in BOTH directions — the work is further along than dispatched,
  and it also did not ship. The cause was the same agent outage that cost `#123` its build, and the
  operator's standing rule from that drop applies here too: an outage must not be able to remove
  work from a batch silently. Board requests got a drop comment; register entries get this.
- **MAPPING CONFIRMED BY ITS BUILDER 2026-08-04, by REQ id rather than by recollection** — hertz
  noted first that its own context had been cleared between building the lane and answering, and
  declined to testify from memory about the original instrument. Everything here is either measured
  that day or read off the commits:
  `1275e47` → `REQ-CI-LOAD-BEARING-PATCH-PIN` ([[IR-4]] part 1, the mechanical pin guard in xtask) ·
  `49d4805` → `REQ-LOCK-TOUCHING-LANE-PROCEDURE` ([[IR-4]] part 2, procedure in docs/GOLDEN-CI.md) ·
  `8ed006b` → `REQ-CI-TOOLCHAIN-VERSION-PRINT` ([[IR-4]] part 3, carrying [[IR-9]]) ·
  `bf8c4a2` → `REQ-CI-LINK-HEALTH-PROBE` ([[IR-1]], the network axis).
- **CORRECTION TO THE GATER'S EARLIER WORDING, from the builder:** `8ed006b` does NOT discharge
  [[IR-9]] — it discharges IR-9's MEASUREMENT half only. IR-9's other half (align the two boxes, or
  declare one authoritative clippy leg) is held by doyle's own 2026-08-03 ruling: decide once the
  step reports RUNNER-ACCOUNT versions, never on the interactive-account prior. IR-9 therefore stays
  OPEN with its trigger now REACHABLE, which is a different state from "waiting".
- **Gate evidence, instrument NAMED:** `cargo nextest run -p xtask` in `.worktrees/hertz-ci-riders`
  on hfenduleam, box quiet with no runs in flight — 34 run, 34 passed, 0 skipped. Nextest rather
  than bare `cargo test`, per [[IR-25]]. ⚠ **What that does NOT cover, stated by the builder rather
  than discovered later:** the two golden steps the lane ADDS (link probe on both boxes, toolchain
  print on both legs) cannot be exercised outside a golden run. The probe's three arms
  (five-sample success, no-reply, absent-CLI) were exercised on the real boxes at commit time, and
  that is reported as a CLAIM RECORDED AT COMMIT TIME, not as a re-verification — the kitsubito arm
  could not be re-checked from here. This is why the lane wants a golden rather than a quiet ff.
- **[[IR-5]]/[[IR-6]] CONDITIONAL RIDERS: CONDITION NOT MET — the answer is NO and it closes the
  question for this lane.** Measured file set `6d291e0..bf8c4a2`: `.gitattributes`,
  `.github/bench/link-probe.ps1`, `.github/bench/link-probe.sh`, `.github/workflows/golden.yml`,
  `crates/xtask/src/main.rs`, `crates/xtask/tests/bench_row_parity.rs`, `docs/GOLDEN-CI.md`,
  `traceable-reqs.toml`. Nothing under `.github/ci/`, where the nextest-summary readers and
  count-reporting scripts live. IR-5: no consumer of nextest output touched. IR-6: no EXISTING gate
  output changed — though the NEW output was voluntarily written to IR-6's rule (the LINK line
  prints `samples=N/M` and the individual rtts, membership beside the count, in both shells). That
  SHRINKS IR-6's future scope by two lines; it does not discharge it.
- **Three mechanisms lifted from this lane's REQ titles, worth more than the lane:**
  (a) **An instrument must not be able to RED the run, and each shell breaks that differently** —
  bash: `| head` under `pipefail` SIGPIPEs the producer to 141, so a print step reds (trim with
  parameter expansion instead); pwsh under GitHub's wrapper: a MISSING COMMAND is a terminating
  error, so rustup's absence must be TESTED with `Get-Command`, never caught. One requirement, two
  constructions, and neither is portable reasoning about the other.
  (b) **A negative-control mutation must be COUNTED before it is trusted** — the absent-CLI arm was
  first "measured" against a tree where the mutation had not taken, silently re-measuring the
  unmutated arm and reading as a pass.
  (c) **Jitter, not the median, is the signal on a link that mostly works** — the motivating red's
  link ran 10..83ms and 3..75ms while IDLE, so a median-only probe reads healthy straight through
  the failure. Both med and max rows ride the ledger.
- **Disposition:** composes onto the next golden batch as a thin lane, NOT slipped onto main. It
  changes the golden pipeline itself, so it wants a golden rather than a quiet ff — and while
  v0.53.0 is untagged, any push to `main` moves the runbook's bare `git tag` off the tested sha.

### IR-55 — the `spt --bins` suite wedges partway through on leaked `findstr .` PTY children; the tests at the frontier are casualties, not culprits

- **Status:** OPEN, **ARMED-FOR-CAPTURE**, owned by **hertz**. Test/rig defect by the dispatch
  split (test/CI rework is never the product builder's). Filed by doyle 2026-08-22 from the
  TURNKEY #212 W3 gate on releases#213.
- **CORRECTED BY REPLACEMENT the same day.** This entry first named
  `cli::tests::adapter_profile_verbs_local_only` as the wedging test, because that is where the
  streamed frontier stopped. **That was wrong, and the correction is the useful part of the
  entry:** run in isolation, on the same pool and binary and with the same env, that test passes
  in **0.02s**. So do the other two the suite later stalled on —
  `cli::tests::shell_channels_relay_sensory_and_text_file` (0.20s) and
  `cli::tests::resolve_proof_target_override_reads_on_disk` (0.00s), each 1 passed / 652 filtered.
  **A frontier names where progress STOPPED, not what CAUSED it.** Naming the last test printed
  is the same error shape as reading a stack frame as a root cause, and it would have sent hertz
  to rewrite three innocent tests.
- **What is actually measured:**
  - `cargo test -p spt --bins` unfiltered sat **24 minutes** at 13.5s CPU — blocked, not spinning.
  - Re-run with `--test-threads=1` and streaming output: stalled after **183** completed.
  - Re-run again with that frontier test skipped: progressed to **335** completed, then stalled
    with **two** tests simultaneously past libtest's 60-second notice. Skipping one casualty
    simply moves the frontier, which is itself evidence the fault is not in any one test.
  - Every wedged harness carried `findstr.exe .` children under `conhost.exe --headless`.
    **`findstr` with no file argument reads STDIN and blocks forever.**
  - The frontier tests are **byte-identical at `c62904e7` and at the #213 tip `3d69f77c`**
    (function bodies hashed at both blobs), so nothing in that lane introduced this.
- **Reading, stated as a reading rather than a conclusion:** an EARLIER test spawns the `findstr`
  stand-in as a PTY child and does not reap it; the children accumulate across the run and a later
  test blocks behind them. Which test leaks is NOT yet identified and this entry does not guess —
  the next step is to bisect for the leaker rather than to touch any test at the frontier. Five
  orphaned `findstr` processes with creation times spread across the day were observed on this box
  independently of any single run, which is consistent with accumulation across many suite runs.
- **Hypothesis RAISED AND NOT SUPPORTED, recorded so nobody re-runs it:** the gate scrubs
  `OWL_SESSION_ID` / `SPT_AGENT_ID` / `SPT_ENDPOINT_ID`, and a scrubbed identity was suspected of
  routing a normally-REFUSED test down a live PTY path. The isolation runs above were performed
  **with the scrub applied** and passed, so the scrub alone does not produce the hang. It remains
  unexplained why the same unfiltered command completed at 652 passed / 1 failed for another agent
  on the same tree; do not treat the scrub as the differentiator without a fresh one-variable run.
- **Why it is register-grade:** golden CI runs the `spt` package. A test that never returns cannot
  red — it consumes the single self-hosted runner slot until the job timeout, which `golden.yml`'s
  50-minute cap bounds rather than fixes. A wedge that fires only sometimes is worse than a red,
  because the suite's green record stays intact and nobody audits it — the same self-concealing
  shape as `exit-code-after-a-pipe-is-the-tails`.
- **Trigger condition (ripe):** any milestone whose gate wants an unfiltered `spt --bins` sweep,
  which is every one of them. Ripe now.
- **REMEDY, MEASURED 2026-08-22 and it changes this entry's urgency:** the same 653 tests under
  `cargo nextest run -p spt --bins --no-fail-fast` complete in **22.2 seconds, 653 passed, 0
  skipped, exit 0** — no slow markers, no timeouts — on the very tree and pool where
  `cargo test -p spt --bins` wedged indefinitely twice. **nextest runs each test in its OWN
  process**, so a leaked `findstr` child can block only itself. This is the discriminating
  measurement rather than a workaround reached for to get a green: **the suite is not broken; the
  SINGLE-PROCESS harness is what lets one test's leaked child wedge every test after it.** Golden
  CI already runs nextest, which is the likely reason this has never surfaced there — the exposure
  is to anyone running bare `cargo test` on the `spt` package, which is every local gate.
- **Count reconciliation, recorded so the figures are not read as a discrepancy:** runs of this
  suite report 652 passed / 1 failed where the one failure is
  `cli::tests::adapter_translate_proof_gates_on_commit` — the `--bins`-only artifact, which
  refuses because `translate_proof_fixture.exe` is not built under a bins-only invocation and says
  so in its own message. Prebuilding it (`cargo build -p spt --bin translate_proof_fixture`) makes
  it pass: 652 + 1 = 653. Prebuild the fixture rather than carrying a known-red through every gate.
- **⚠ AMENDED 2026-08-22 — DO NOT PREBUILD FROM THIS BULLET'S EXAMPLE.** The line above names ONE
  fixture because one fixture is what *this* wedge needed; read as the prebuild recipe it covers a
  fraction of the population, and it was read that way twice on one day. **The census already exists
  in IR-21 above** (`adapters/mock`: mock-session, mock-shell, capture-player, console-mode-probe ·
  `crates/spt-daemon`: dispatch_fixture, service_fixture, summarizer_fixture · `crates/spt`:
  translate_proof_fixture, post_step_fixture, gh_fixture, git_fixture), together with the split that
  matters: **cross-package sites are the hazard members; same-package builds are guaranteed by
  construction.** Field instances, both 2026-08-22 TURNKEY #212 W4: todlando's lane red on three
  `spt-daemon::attach_resize_capture` cells wanting `adapters/mock`'s **capture-player** (a
  cross-package member), and doyle's head gate prebuilding the remembered pair while never running a
  `spt-daemon` or `spt-term` suite at all. **Prebuild `--workspace --bins`, or read the
  `build_hint` the helper already composes** (`crates/spt-term/tests/support/fixture_bin.rs:42`) —
  five test files gate on that panic. Neither of us lacked the census; we prebuilt from memory while
  our own register held the list.
- **FIXTURE PREBUILD, RECURRENCE 2026-08-29 (todlando, emission lane):** the rule was
  already banked AND I am one of its authors, and I still ate three fixture reds in a row —
  mock-session, then capture-player, then mock-shell — each tripping in about 0.01s with a panic
  that reads exactly like a test failure. Having the rule is not the same as applying it completely:
  I prebuilt from RECALL (`-p spt --bins`) rather than from the rule's stronger banked form, and the
  tell is that my fix arrived in three instalments instead of once. **TWO FACES, both structural
  rather than careless.** (1) THE FIXTURE MAY LIVE IN ANOTHER PACKAGE — mock-session,
  capture-player, mock-shell and console-mode-probe are all `adapters/mock` (`-p mock-adapter`), so
  no amount of `-p spt --bins` ever produces them. (2) ENUMERATING FIXTURES BY GREPPING THEIR
  PRE-BUILD STRINGS IS STRUCTURALLY INCOMPLETE — it finds the call sites that embed a literal
  `pre-build: cargo build ...` string and CANNOT find mock-shell, whose helper composes the command
  at runtime (`sibling_bin(name)` → `fixture_package(name)`), so there is no literal to match. A
  predicate over source text cannot see a string the program builds at runtime; same blindness
  class as the emission census's TOKEN-colon-vs-TOKEN-space predicate, met twice in one session.
  **CLASS FIX:** never hand-list fixtures — build the whole fixture package
  (`cargo build -p mock-adapter --bins`) or `--workspace --bins` before any
  `--test`/`--bins`/filtered leg. Chasing named fixtures one red at a time only ever finds them one
  at a time, and each one costs a full battery run.
- **Named negative + settle-class result (hertz, 2026-08-23):** current main at
  `4dd699cb7867979ca0497e7c74d8d0aaf881e9e9` did not reproduce the wedge on HFENDULEAM from
  isolated checkout `.worktrees/ir55-findstr-bisect`. Test binaries were freshly built before the
  diagnostic population and reused for the settle sweep; the gate environment scrub was **not**
  applied; the installed daemon was live. The fleet snapshot taken immediately **after**, not
  during, the sweep showed BIGNET 7 nodes / 20 endpoints, SPT_MANTLE 3 / 19, SPT_DEV 2 / 19,
  3-of-7 peers connected, `degraded-partial`. Four startup cells that spawn `findstr` passed alone
  and reaped their observed ConPTY children; a full single-threaded run passed 663/663 with a
  before/after census delta of `IR55_NEW=0`. The ratified settle-class sweep then ran five
  consecutive bare `cargo test -p spt --bins` invocations: all 663/663 green in 420.35s total,
  with zero `findstr.exe` residents before and after. This is a named negative data point, not a
  discharge of the intermittent filing-day measurements.
- **Next-wedge capture protocol — do not improvise or clean first:** preserve the exact bare
  invocation and snapshot the environment (box, checkout SHA/worktree, fresh-versus-reused test
  binary provenance, gate-scrub state, live-daemon state, and fleet roster); stream test output to
  durable storage so the last started and completed test names identify the frontier; census
  `Win32_Process` immediately before the run and again while wedged, recording each `findstr.exe`
  PID, parent PID, creation time, executable path, and command line. **Do not kill any child before
  the second census is banked.** The 22 residents found during the 2026-08-23 investigation were
  path-verified as `C:\Windows\System32\findstr.exe`, shown to predate its live runs, banked, then
  reaped; the pre-sweep and post-sweep resident counts were both zero.
- **Size guess:** small if one test is missing a reap or an EOF for its stand-in child; medium if
  the shared PTY test scaffolding needs the reap, in which case every sibling spawning the same
  stand-in wants it. The RUN-LEVEL exposure has an answer today (nextest); the LEAK itself is
  still worth fixing, because a leaked child per run accumulates on any box that runs the suite.
- **Gater's craft, worth more than the entry itself:**
  (a) The first wedge was undiagnosable because the gate script wrote every leg as
  `cargo test … | tail -N > file` — nothing reaches disk until the command completes, so a wedged
  leg produces an **empty file** and the process table is the only witness. **Stream to a file and
  tail the FILE.**
  (b) To NAME a frontier, run `--test-threads=1`: libtest prints each test's name BEFORE running
  it, so the last incomplete line is where progress stopped.
  (c) **Then prove the frontier test is actually at fault by running it alone** — which is the
  step this entry originally skipped, and the one that turned a wrong entry into a right one.

### IR-56 — `pool-claim` harvests lane identity from CWD, not from `--pool`; a claim run from the wrong tree writes a WRONG RECORD with a success message

- **Status:** BUILT 2026-08-24 for the SIGNET #218 golden batch on
  `fix/ir56-pool-claim-cwd`, owned by **hertz**. Tooling/rig defect by the dispatch split. Filed by
  doyle 2026-08-22 from todlando's TURNKEY #212 W4 Lane 1 field instance — the first confirmation
  of the IR-42 mechanism from a LANE rather than from a source read.
- **What happens:** `cargo run -p xtask -- pool-claim --pool <dir> --label <lane>` records the lane
  identity it harvests from the **current working directory** — `owner_tree`, `lane_branch`, and the
  base sha — not from the tree `--pool` names. Run it from the project root with `--pool` aimed at a
  worktree and the record reads `owner_tree = <project root>, lane_branch = <root's branch>` while
  every artifact in that pool belongs to the worktree. The verb prints its ordinary success line
  naming the wrong owner, and returns 0.
- **Why it is silent, which is the whole entry:** the claim verb WRITES and never adjudicates
  (IR-42) — every enforcement arm lives in `crates/spt-store/build.rs` and speaks only at the next
  BUILD. So a wrong record has no symptom at the moment it is created; the first signal is a build
  refusal minutes later, on a lane whose `--pool` argument was correct all along. **`--pool` being
  correct is not evidence the record is.** Nothing on the command line shows you cwd, which is the
  one input that decided the outcome.
- **Recovery is the documented hatch, and the bootstrap is why:** the claim tool builds THROUGH the
  pool it is claiming, so a pool already refusing cannot be re-claimed by the ordinary route. The
  refusal text says so itself — meaning the bootstrap case is known, and a wrong-cwd claim is the
  ordinary way a lane lands in it.
- **The rule, as a must-do:** claim from the lane's own worktree so the identity the verb harvests
  is the lane's — `( cd <lane> && cargo run -p xtask -- pool-claim --pool "$PWD/target" … )`, or any
  form that makes cwd explicit. Then READ THE PRINTED RECORD BACK and confirm `owner_tree` and
  `branch` name the lane: the success line carries the record, so the check costs nothing.
- **Trigger condition:** ripe now — it is a doc/UX fix on a verb we run every lane. Ripest alongside
  any other `xtask` pool-verb touch.
- **Size guess:** small. Either make the verb REFUSE when cwd is not inside the tree containing
  `--pool` (loud, and it cannot be wrong), or print the harvested identity as a distinct confirmed
  line. The refusing form is preferred: it removes the reading step rather than adding one.
- **Built evidence, corrected at review:** `pool-claim` resolves the git worktree containing
  `--pool` from the pool's nearest existing ancestor and compares it with the cwd worktree before
  writing. An accidental cross-worktree call returns 2, names both worktrees and the remedy, and
  leaves the not-yet-existing pool absent. The sanctioned sequential takeover carries an explicit
  `--foreign-pool`, says that it is harvesting the arriving cwd identity, writes that identity, and
  returns 0. Both poolguard refusal remedies now print that flag. An executable xtask test drives
  both arms against two real temporary git repositories, pinning VERB-ADMITS-REMEDY rather than
  merely testing path equality. This remains identity-harvest validation only: the claim still
  writes without adjudicating admission, so IR-42's enforcement boundary remains in the build.

### IR-57 — an assembly pick's conflict resolution can silently DROP lines, and every gate run at assembly time is blind to it

- **Status:** OPEN, owned by **doyle** (gater craft, not a code defect — no lane to dispatch).
  Filed 2026-08-22 from doyle's own defect on the TURNKEY #212 assembly head.
- **What happened:** assembly head `5cce533d` **did not compile**, and was carried across a session
  boundary as a ready head with "treqs exit 0" attached to it. `crates/spt-store/src/access.rs` had
  an unclosed `mod tests`: the cherry-pick of the #206 lane commit `b2194aa9` conflicted against the
  #209 test block — both add functions to the same region, and the two sides share the trailing
  `);` / `}` / `}` — and the resolution deleted the markers without restoring the FIRST side's
  function closer. Exactly two lines lost: `        );` and `    }`.
- **Every instrument in the assembly path reported success:** `git cherry-pick` completed with no
  conflict remaining; `traceable-reqs check` returned 797/797, 0 findings, exit 0 — **it parses
  tags and never invokes the compiler**; and the lanes themselves were green and stayed provably
  clean (the #206 lane's brace balance is correct at all five of its commits). **Lane-green plus
  conflict-free is not a claim about the assembled head.** Same lesson as the #182 v2 assembly
  (`E0425`×4) by a different mechanism: that one was a semantic composition break between two
  correct hunks, this one is a fidelity LOSS during resolution.
- **Two must-dos, both cheap:**
  1. **Compile the assembled head before it is handed anywhere — including forward to your own next
     session.** A head is not ready on a treqs exit; it is ready on a build. Any sentence carrying a
     head sha to another agent must name the instrument that proved it.
  2. **Audit pick fidelity on CHANGED LINES for every pick on the chain**, not just the one that
     tripped you: hash `git show --format= <sha> | grep -E '^[+-]' | grep -vE '^(\+\+\+|---)'` for
     the lane source and for the pick, and compare. Whole-diff comparison is useless here — hunk
     headers and context legitimately drift once the head's copy of the file has moved.
- **Read COUNT and HASH together, because they fail in opposite directions.** Of 24 picks on this
  chain two did not match. `b2194aa9 → 4d6deed1` differed in COUNT (302 lane / 300 pick by the header-excluding pipeline above; **this entry originally cited 310/308, which is the RAW `^[+-]` count INCLUDING the 8 `+++`/`---` header lines of its 4 files** — the figures were taken without the second grep the entry itself prescribes. The DELTA is 2 under either meter, which is why the wrong figures never surfaced: they supported a correct conclusion, so nothing pressed them. Corrected todlando 2026-08-29, measured both ways in-tree) — the real
  defect. `f27f15c6 → 5865cc98` had the SAME count and a different HASH — benign: a `CONTEXT.md`
  paragraph the head had already amended for #209, whose merged result correctly carries both lanes'
  sentences. A count check MISSES the first class entirely; a hash check FLAGS the second as if it
  were a defect. Neither alone classifies a pick.
- **Repair shape used, recorded because it is the cheap one:** reset to the commit before the bad
  pick, re-pick the lane source, resolve the same conflict correctly, then replay the remaining picks
  in chain order with the fidelity check on each. Verify the repair changed nothing else with a
  WHOLE-TREE diff against the old head — it must show exactly the restored lines and no other file.
  Where a later pick re-conflicts, check its pre-pick blobs against the old chain
  (`git rev-parse <old-pick>~1:<file>`); if they match, the old chain's post-pick blob is a
  transferable, already-reviewed resolution.
- **AND GATE THE REPAIR'S OWN SUBJECT.** The rebuilt head passed six legs — clippy `-D warnings`,
  sweep 660/660, treqs 803/803, xtask OK — while the restored lines live in `spt-store`, whose LIB
  tests `-p spt --bins` never runs. Six greens, and not one executed the function the repair
  restored; it was proven to COMPILE and assumed to PASS. Proven afterwards by name (4/4) and by the
  full `spt-store --lib` suite (519/519). **A repair's gate must include the suite that owns the
  repaired file**, which is not necessarily the suite the head gate runs.
- **FIRST LIVE CATCH, and it was the author's own resolution (2026-08-29, one hour after the verb landed):** rebasing the IR-59 rider over the IR-57 rider produced four conflicts — three additive keep-both, one usage line both sides had edited. `xtask pick-audit` read the rebase as **LOSS, 282 lane / 280 pick**. The two missing lines were a `[[requirements]]` header and its blank line in `traceable-reqs.toml`, so `id = REQ-DISK-FLOOR-PREFLIGHT` landed INSIDE the previous table as a duplicate key: **`traceable-reqs check` exit 2, the registry unparseable, every reading after that edit checking nothing**. Restored, the same audit reads **DRIFT 282/282** whose one differing line is the usage string whose base copy legitimately gained the other verb — both verdict classes, on one resolution, in the order the entry predicts. The instrument's first real save was from the hazard it was filed for, against the agent who built it.
- **FIRST REAL-ASSEMBLY USE (2026-08-29, #241 emission head, ASM-241-GATE-VERDICT.md):** `xtask
  pick-audit --range 57a0d27c..HEAD --lane 4d6007ac..d0fdd58d` over 6 picks — **5 MATCH
  (digest-identical), 1 LOSS, 0 drift, 0 unclassified**, exit 1 forcing the accounting. The LOSS
  was exactly the one human-resolved conflict (`cfb9d9f8 <- 10f12e4c`, count 432/402), and the
  delta RECONCILED to the resolution rather than assumed: lane stat 15 files/235+/197- vs pick
  14 files/205+/197-, difference byte-for-byte the deliberately-dropped 30-line interim IR-69
  register block superseded by main's consolidated entry. No false positives, no silent pass.
  Entry stays OPEN as living procedure; both live uses (the rebase catch above, this assembly)
  produced the verdict classes in the order the entry predicts.
- **Trigger condition:** ripe at the next assembly — this is procedure, and its cost is one command
  per pick.
- **Size guess:** small as a scripted check in `xtask` (fidelity audit over a pick range); zero as
  discipline, which is how it is being applied now.

### IR-58 — a `[[bin]]` grep is not a census: cargo AUTODISCOVERS `src/bin/*.rs` targets, and a hand-built bin list produced a confident false "this bin does not exist"

- **Status:** OPEN, unowned. Filed by doyle 2026-08-22; **its original filing was WRONG and is
  retracted in full below.** The entry is kept because the way it was wrong is the finding.
- **RETRACTED CLAIM, stated plainly so no reader acts on it:** this entry first said that
  `crates/spt/Cargo.toml:18-21` and `crates/spt/src/cli.rs:30355` forbid a bin-name collision with
  `xlate_choreo_fixture`, "a bin that does not exist", and that spt-daemon's third helper had been
  renamed to `summarizer_fixture`. **`xlate_choreo_fixture` EXISTS** — at
  `crates/spt-daemon/src/bin/xlate_choreo_fixture.rs`, verified at the assembly head. The collision
  rule's counterpart is real, both prose sites are CORRECT, and IR-21's original table row was
  correct too; it merely predated `summarizer_fixture`. Nothing in that prose needs fixing.
- **How the false claim was produced, which is the entry:** the census was built by grepping
  `^\[\[bin\]\]` across `Cargo.toml`s. **Cargo AUTODISCOVERS `src/bin/*.rs` as bin targets with no
  stanza at all**, so a stanza grep cannot see them, and it fails SILENTLY — it returns a clean,
  well-formed list that is simply short. From that list it read as established fact that a named bin
  had been renamed away, and that "fact" was then written into a register table, an IR entry, and a
  message to the builder. **A pattern-built population encodes the pattern's assumption; the count
  looks like a census and is a property of the pattern.**
  (Found by todlando, resolving one of doyle's figures against his own tree under the pointer rule.)
- **A CENSUS IS ALSO A PROPERTY OF A TREE.** Stanza counts differed legitimately across trees the
  same hour — 13 at the assembly head, 11 at `dc1c7532` — because `summarizer_fixture` exists at one
  and not the other. Two correct censuses of the same workspace disagree unless each names its sha.
- **The must-do, and it removes the list rather than lengthening it:** never hand-maintain a bin
  roster. Run `cargo build --workspace --bins` and let cargo enumerate its own targets — the same
  enumeration the test harness resolves against, so it cannot drift from what a fixture lookup
  expects. Stanza-derived lists, remembered pairs, and register tables are all the roster problem at
  different depths; only cargo's own enumeration is the population.
- **Kin:** IR-55's amended prebuild bullet (prebuild the census, not a remembered pair) and IR-42
  (a verb that writes without adjudicating). The shared shape is an instrument returning an orderly
  answer about a set it never fully saw.
- **Trigger condition:** ripe on any `xtask` check work — a duplicate-bin-name check over cargo's
  target enumeration (NOT over manifest stanzas) is a handful of lines and cannot go stale.
- **Size guess:** small. The prose fix originally proposed here is WITHDRAWN — there was nothing
  wrong with the prose.

### IR-59 — pool arithmetic: this box holds TWO cold pools, not four; and a build that exhausts the volume reds as a LINKER defect that names no disk

- **Status:** OPEN, unowned — it is arithmetic and discipline, not a code lane. Filed by doyle
  2026-08-22 from a live disk-floor abort during the TURNKEY #212 W4 gate; footprint figures measured
  by todlando the same hour.
- **SECOND FACE, measured 2026-08-23/24 (WAX-SEAL #21 W2 gate, doyle — the violator this time):**
  both predicted mechanisms fired at once, WITHOUT the floor guard because the gate legs ran as a
  plain script. (1) doyle minted a THIRD cold pool (gate rig) while todlando's lane pool was LIVE —
  the two-pool arithmetic, violated by the gater. (2) The lane pool had silently accumulated
  **104.8 GB** over repeated full-workspace sweeps (nothing bounds or even measures a pool's
  accumulation until the volume dies). The volume hit **3 MB free**; the reds wore exactly the
  entry's predicted costume — `LNK1318: Unexpected PDB error; LIMIT (12)` naming no disk — plus a
  second face worth keeping: `traceable-reqs check` died as a PANIC printing to stdout
  (`os error 112`), so a REGISTRY gate red can also be the disk's. Both gate verdicts VOIDED
  (clippy's full compile had exited 0 before space died — the code was never the question).
  Recovery measured: gate pool reap 9.6 GB, lane pool clean 100.2 GB, drive back to 109.6 GB free.
  **Remedies adopted at the incident:** the gate never mints its own pool while a builder lane is
  live on the box — it takes the lane pool over SEQUENTIALLY at pool-release (warm, loud, the
  releases#103 pattern); and gate/lane scripts get the disk-floor preflight the golden runner
  already carries (this entry's part 2, still unbuilt).
- **THIRD FACE, measured 2026-08-24 (WAX-SEAL #21 golden window, doyle):** two new mechanisms, both
  about the REMEDY rather than the red. (1) The wax-seal-w1 lane pool **re-accumulated to 152–169 GB
  (two meters, IR-46 caveat) within ONE DAY of the 08-23 clean** and rode the volume back down to
  4.92 GB free — a reap is an instant, not a state; nothing yet bounds a pool between reaps.
  (2) **The remedy itself arrived garbled through relay:** the golden push was held "waiting on the
  IR-59 reboot" when this entry prescribes no reboot anywhere — tens of releases have shipped without
  one; the operator refused the premise and the register's own recorded remedy (reap finished-lane
  pools) cleared it in ten minutes. A remedy relayed as a DIFFERENT remedy is the relay trap wearing
  infra clothes: the entry is the author — read it before adopting a hold. Recovery measured: five
  finished-lane pools reaped after ancestry classification (`git cherry` catches rebased lands that
  a plain `merge-base --is-ancestor` calls unlanded — gate-w4l1's 90.9 GB pool was reapable only by
  that meter), 4.92 → 260.88 GB free (~256 GB free-delta vs ~270–287 GB sum-of-lengths). Two main
  runs red inside the low-disk window (09:17Z/09:41Z) were re-run, not re-read, per this entry's
  must-do.
- **FOURTH FACE, measured 2026-08-24 same day (SIGNET #218 W2 gate, doyle — the gater's own pool
  this time):** the MAIN pool went **13.7 GB → 158.3 GB in ONE DAY** absorbing four full gate
  batteries across two lane tips, and dragged the volume to **18 GB free mid-gate** — so the
  entry's mechanism is not a lane-pool problem, it is EVERY pool under repeated full sweeps, the
  gate rig included. Two riders worth the ink: (1) the golden runner's **free-space floor step
  produced its first live catch** — a thin CI leg refused at 14:14 local over exactly this window
  with zero test signal, and the register's re-run-not-re-read rule resolved it in one command;
  (2) the reap-and-rebuild trade ("cleaning a live pool buys one window at rebuild price") was
  taken deliberately mid-gate and cost ~35 minutes cold — cheaper than one uninterpretable red.
  A red measured at 18 GB free in this gate got NO read at all, and the next sweep's red was a
  DIFFERENT test: low disk manufactures random victims before it manufactures link errors.
- **The event:** a gate's preflight refused with `GATE_ABORTED_DISK_FLOOR` at **4.75 GB free on a
  1863 GB volume**. Four cold pools plus the milestone head's pool were standing at once. The floor
  guard did its job — without it a full workspace build would have started with under 5 GB of
  headroom and produced a red belonging to nothing.
- **THE ARITHMETIC, and it corrects the intuition that worktrees are the cost.** Measured across the
  whole `.worktrees` tree: **45 worktrees = 0.79 GB total** (~18 MB each). **3 pools = 130.45 GB.**
  Worktrees are **0.6%** of the footprint; pools are **99.4%**. **One cold pool is worth roughly 2,500
  worktrees.** A rule keyed on worktree count would cost a day of `git worktree remove` work for under
  a gigabyte, risk removing a lane someone still wants, and leave the lever untouched.
- **THE CEILING, as arithmetic rather than hygiene:** at **45–75 GB per cold pool** on this workspace
  against a **40 GB floor**, this box supports **TWO live pools comfortably, THREE only if one is
  small**. The two rules that already exist need no replacement, only this number attached:
  - **Reap a pool when its lane is finished** → 45–75 GB back immediately.
  - **Do not run two builds at once** → not only the contention argument, but **75–150 GB of
    simultaneous allocation** against that floor.
- **A FULL VOLUME REDS AS A LINKER DEFECT.** `LINK: fatal error LNK1318: Unexpected PDB error`
  (variously `OK (0)` or `LIMIT (12)`), often beside `LNK4209: debugging information corrupt`. It
  reads as a corrupt artifact or a broken toolchain and is neither — it is the linker unable to grow
  a large `.pdb`. **Nothing in the failure text says "disk".** Do not key recognition on the
  parenthesised code; it varies.
- **THE DANGER WINDOW IS THE TAIL OF A COLD BUILD, not a steady state you can check once.** The
  2026-08-22 instance died at the link of `spt` — the last and largest artifact of a **74 GB** cold
  pool — while that same build was draining the volume out from under itself. A free-space reading
  taken before the build would have looked fine.
- **Positive tell, with its own refutation attached:** `traceable-reqs check` alone surviving a leg
  table is suggestive, because it is the only leg that never links. It is a POSITIVE tell ONLY — a
  full disk can red before any link step, so the signature's ABSENCE clears nothing.
- **THE FALSIFIER IS FREE SPACE AT THE TIME OF THE RUN**, not the shape of the leg table.
- **Must-do, both cheap:**
  1. **A free-space reading goes in the FIRST LINE of any build-failure report** — not in the
     controls, not as a follow-up. Two agents spent hours on a mechanism for this failure and neither
     ran `df`; the falsifier was one command away throughout.
  2. **Rig and gate logs should record free space at run start.** Today they do not, which makes "was
     the disk full" **unanswerable after the fact for every red already in hand**. Any red produced
     under ~1 GB free is UNINTERPRETABLE and must be **re-run, not re-read**.
- **Measurement caveat, observed twice in one hour:** sum-of-file-lengths and free-space delta
  disagree by several percent when anything else on the box is writing (98.07 GB of subtrees → 92.25
  GB reclaimed; 47.28 GB → 40.81 GB). Report BOTH, derive neither from the other, and treat any single
  free-space number as an instant rather than headroom — see IR-46.
- **FIFTH FACE, measured 2026-08-29 at the CONDUIT #236/v0.65.0 cut (deployah):**
  the finished milestone's sequential lane-sharing accumulated **140.51 GB**, while a freshly
  rebuilt pool measured **7.78 GB** — about **18× one build's working set** retained across lane
  hand-offs with no reap between them. This is the budgetable rate behind that cut's disk-floor
  red: sequential execution prevents concurrent ownership corruption, but it does not bound
  historical artifacts inside the shared pool. Treat takeover and reclamation as separate
  operations; a pool may be safe to reuse and still be too large to keep.
- **SEVENTH FACE, measured 2026-08-30 (v0.67.0 milestone close — the PARALLEL-LANE fill
  rate):** the #23 wave arc drove C: from ~240 GB free to **0.01 GB in ~3.5 hours**: four
  concurrent/serial lane+gate pools totalled **236.6 GB** (ns23-w1 87.7 + ns23-113 72.3 +
  gate rig 57.3 + ns23-w3 19.1) — effectively the entire ir57 reclaim re-consumed inside one
  milestone's build window. The head gate and two thin-CI runs redded with the entry's exact
  costume ("could not compile" + treqs panic, zero disk words) at 0.01 GB and all re-earned
  green after reclaim (+186.6 GB, finished-lane targets only, single-actor rule held under
  racing consent). Budget rule this face adds: a MULTI-WAVE milestone on one box books
  ~60-90 GB PER LANE-PLUS-GATE cycle, so the pool budget is per-wave, not per-milestone —
  reap each wave's pool at its LAND, not at the milestone party (the party inherited only
  ~58 GB because the emergency had already forced the other ~179).
- **SIXTH FACE, measured 2026-08-30 (#242 golden r4, Windows test job, LNK1318 at the PDB
  write):** the floor is asserted ONCE at job START, but the job's internal LOW-WATER sits at its
  LAST heavy step — the docs-drift build, ~28 minutes in. r4's floor read 65.1 GiB honest at
  23:07; the box was at ~37–39 GB by the 23:35 link failure. LNK1318 on PDB write = this entry's
  documented low-disk costume; the incremental-corruption candidate was retired by preservation
  snapshot (no xtask state existed at failure). **Durable fix = a floor RE-READ (or reclaim)
  immediately before the docs-drift build** — the WITHIN-job sibling of the across-job
  read-after-checkout fix, rides the next workflow commit. Riders: incremental hygiene as
  disk-pressure maintenance (the dir measured 4.56 GB), and the box's **~51 GB non-recovering
  post-run consumption** is this entry's pool-weight face — classify at the IR-14/26/27/49 audit
  slot (composed this sweep).
- **Composed for next intake (doyle, v0.65.0 close):** the still-open log-the-floor build half
  rides with [[IR-57]]'s scripted pick-fidelity audit as the next milestone's two tooling riders.
- **Trigger condition:** ripe now for the log-the-floor half (it rides any gate-script or CI touch);
  the arithmetic is discipline and applies immediately.
- **Size guess:** small — one line in each rig/gate preflight to record free space alongside the
  existing floor check.

### IR-60 — a wrapper that CAPTURES a leg's exit status ends with the capture, so the wrapper's own status is the capture's, not the leg's

- **Status:** OPEN, unowned — script-shape discipline, not a code lane. Filed by doyle 2026-08-22
  during the TURNKEY #212 W5 lane; mechanism measured by todlando, who reported it first as a
  possible harness misreport and then RETRACTED that reading himself on a minimal probe.
- **The reading that started it:** a background leg was summarised as `exit code 0` while the leg's
  own exit file read `100`. Reported as-is, that is a task runner lying about a red — the worst
  possible direction for a purposeful-red claim, since it would turn two reds that DID fire into two
  reds reported as not having fired.
- **The measurement, minimal and on this box:**

      ( exit 100 ); echo $? > probe.exit; FINAL=$?
      leg status captured to file : 100
      wrapper's own final status  : 0

  The wrapper's shape is `( cd <lane> && cargo nextest ... > leg.raw 2>&1 ); echo $? > leg.exit`.
  **The last command is the `echo`, which succeeds**, so the wrapper genuinely exits 0 while the leg
  genuinely exited 100. THE RUNNER TOLD THE TRUTH ABOUT THE WRAPPER. There is nothing to file against
  the task runner, and it would have been filed.
- **The class, and this is why it is worth an entry rather than a fix:** it is the truncation-pipe
  family — the meter measured something real, it just was not the thing about to be quoted. **Three
  shapes now share ONE trigger, and the trigger is not "is this a pipe":**
  1. a truncation pipe (`cmd | tail`, `cmd | head`) — `$?` measures the truncator, and `head` closing
     the pipe can MANUFACTURE a status (SIGPIPE → 101), so `PIPESTATUS` is not protection either;
  2. a flattened `$?` read after any intervening command;
  3. **a wrapper whose last command is the capture itself.**
  **The trigger in all three is the REPORTING** — every instance happened while trimming or capturing
  output in order to QUOTE it as evidence.
- **The practice that held, and the one to keep:** the leg's own file is the leg's verdict. Both reds
  were real, and the only reason that was knowable is that every leg wrote its exit to its own file
  rather than to the wrapper's status.
- **Remedy, in preference order:** (1) never read a wrapper's status as a leg's — read the leg's exit
  FILE, which the per-leg rule already requires; (2) if a wrapper's own status must be meaningful,
  end it by re-raising the captured status (`exit "$(cat leg.exit)"`) or set the status before the
  capture is the last word; (3) for anything that becomes evidence, redirect to a file and read the
  file.
- **Trigger condition:** ripe now — it rides the next gate-script touch, and the gate scripts in this
  milestone already write per-leg exit files, which is what makes this discipline rather than debt.
- **Size guess:** small — a convention line in the gate-script preamble; no product code.

### IR-61 — spacerun's test-module skip carries a SECOND literal tracker; the two answers can drift, and today only a cell stands between them

- **Status:** OPEN, unowned, accepted-with-mitigation. Filed by doyle 2026-08-22 at the TURNKEY #212
  head gate, as the named residual of the latch repair (the repair itself rides the milestone).
- **What the debt is.** `crates/xtask/src/spacerun.rs` now skips a column-0 `#[cfg(test)]` module by
  BRACE DEPTH and resumes after its close, rather than latching the scan to end-of-file. Counting
  braces safely means knowing when a brace is inside a string, a raw string or a char literal — and
  the scanner ALREADY tracks literals for its rendering pass. It does not reuse that tracking. There
  are now TWO answers in one file to "am I inside a literal", and **two answers to one question is a
  disagreement waiting for a reader who fixes one of them.**
- **Why it was accepted rather than refactored on the spot, stated so the trade is auditable:** the
  two trackers answer genuinely different questions — the rendering walk is PER LINE and forward
  from an opening quote ("where does this literal start and end"), while the skip needs "am I inside
  a literal right now" carried CONTINUOUSLY across lines. Unifying them is a real refactor of a check
  that is already gated, already carries five fixed defects, and sits on the tree being handed to
  golden. The narrow change with an OBSERVABLE failure mode beat the correct-shaped change with a
  wide blast radius, on that tree, on that day.
- **The mitigation, and its exact limit.** Three brace cells fail the moment the two trackers
  disagree about a raw string, a normal string or a char literal, plus a cell that reds when the skip
  stops consulting a literal tracker at all (`match lit` -> `match Lit::None`). **That is a cell
  standing in for a refactor.** It catches drift in the three forms it names and nothing else — a
  fourth literal form, or a change to only one tracker in a form neither cell covers, passes.
- **The general shape, which is why this is a register entry rather than a comment:** a check whose
  own correctness depends on a second implementation of a thing it already implements is one
  refactor away from the defect class it exists to prevent. This same file has already produced
  three instrument defects (an indented-marker latch, a column-0 latch, and a suppression that
  decided what got PARSED rather than what got REPORTED). The pattern is not carelessness; it is a
  scanner accumulating special cases.
- **Trigger condition:** the next SUBSTANTIVE change to spacerun's scanning — any new skip, any new
  literal form, any change to either tracker. At that point unify, rather than adding a third answer.
- **Size guess:** small to medium — one continuous literal-state walk serving both the render pass
  and the skip, with the existing corpus and brace cells as the regression net.

### IR-62 — an e2e daemon binds well-known ports and collides with the resident fleet on shared runners

Filed 2026-08-23 (doyle), from the #212 r2 golden's Windows Phase A red — one witnessed
instance, mechanism verified from the run's own capture, filed on the mechanism per the
register's standard.

**Mechanism.** `endpoint_autostart_e2e::saved_endpoint_replays_on_daemon_restart` failed its
own PRECONDITION ("daemon B must come up with a fresh brain") because daemon B's broker came up
degraded: `NODE_KEY_FAIL: identity unavailable` (net-less broker, no retry) and
`DOCS_SERVER_BIND_FAIL` port 5474 `os error 10048` — the docs port was held by a CO-RESIDENT
daemon. HFENDULEAM is live infra: the resident fleet's daemon legitimately holds well-known
ports, and any e2e that brings up a daemon with default port bindings is in a race with it BY
CONSTRUCTION. An isolated SPT_HOME isolates the store and broker socket, NOT globally-numbered
TCP ports.

**Classification history.** r2: red (35.4s, precondition panic). Run 1 and r3 on the same box:
green. Folded into r3 under a pre-registered predicate (reds twice → dedicated triage); it
greened, so this stays an environment-shaped intermittent, NOT a product defect and NOT
closed — the collision window is real and will re-fire under the right co-residence timing.

**Remedy direction (not built).** Test daemons should bind ephemeral/rig-scoped ports for
every advisory surface (the docs server is advisory — a bind failure should degrade the rig
loudly, not poison an unrelated cell's precondition), or the precondition should name the
port-collision cause distinctly so the red self-classifies. Either lane is test/CI work
(hertz's), triggered the next time this class fires anywhere.

**Kin.** IR-51 (e2e daemon net-enabled on real interfaces — hermeticity is disk-deep only);
the known not-ours red classes list in the #212 hand-off.

**BUILT 2026-08-24 (hertz, SIGNET #218 batch; PR #158).** The class fired its second witnessed
instance the same day — doyle's W2 gate rig, `engine_room_bringup_e2e` precondition panic with
the IR-50 panel naming `DOCS_SERVER_BIND_FAIL` port 5474 `os error 10048` — which armed this
entry's own trigger. Hertz repro'd deterministically (held the port; brain still reached
BRAIN_UP), pre-ranked three mechanisms before probing, and built the ephemeral-advisory form:
rig-only `SPT_TEST_EPHEMERAL_ADVISORY_PORTS=1` makes the daemon's docs listener bind port 0
(production config/`SPT_DOCS_PORT`/docs-url untouched; the flag matches the literal "1" only).
The golden workflow's test job sets it globally on both self-hosted legs, the standalone ER rig
sets it too. Scope note kept honest per his own ranking: the docs collision is retired; if a
co-diagnostic bind elsewhere still poisons a precondition, that is a residual face of THIS
entry, not a closed question.

### IR-63 — suite-mix sweeps leak job-escaped autostart daemons that LOCK the pool's spt.exe; and the same box conditions manufacture one random victim per sweep

- **Status:** OPEN, mitigated-by-rig-step — the reap is adopted discipline, not yet construction.
  Filed by doyle 2026-08-24 at the SIGNET #218 gates; first measured by todlando the same day
  (his W1 build: four reap rounds of 2, 8, 1, 1 processes, count-by-path stated each round).
- **Mechanism, two faces of one condition:** (1) e2e sweeps that exercise autostart leave
  job-escaped daemons running `target/debug/spt.exe`; the NEXT build or `xtask check` then dies
  on `os error 5` removing the exe — a rig red wearing a build defect's clothes. nextest's own
  `leaky` flags corroborate (4–8 per sweep measured). (2) The same co-residence (leaked daemons +
  resident fleet + runner traffic) manufactures ONE red per full sweep with a DIFFERENT victim
  each time: doyle's W2 gate measured FOUR distinct one-off victims in four same-sha sweeps
  (ER bringup, composite+bootstrap pair under CI load, resident_service, brain_resume), every
  one green alone and/or in a sibling sweep. The e2e-leaked-daemons rule holds: a different test
  dying each run on one sha is ONE environment cause, and hardening members never closes it.
- **Mitigation adopted (rig step, both todlando's lanes and doyle's gate runner):** reap
  `spt.exe` BY PATH (`*spt-core\target*`) before EVERY cargo invocation and after every sweep;
  report the count each time so a zero is a claim, not silence.
- **Remedy direction (unbuilt):** the durable form is construction, not discipline — either
  nextest wrap/xtask verb that performs the by-path reap as a pre-step on this box's rigs, or
  autostart e2es gain teardown that provably outlives job escape (kin IR-38's built wedge-namer,
  IR-62's ephemeral ports which retire one collision axis, IR-35's victim-rate ledger).
- **THIRD FACE, measured 2026-08-30 (doyle, ir57/asm-241 reaps):** the leaked population is not
  only daemons — battery legs leave ORPHANED HEADLESS `conhost.exe` processes whose CWD sits in
  the worktree's `crates/spt-daemon` (where the leg ran), and a CWD pin blocks `Remove-Item`/
  `git worktree remove` on the whole tree with "being used by another process" naming the DIR.
  This is the mechanism behind every recent "handle-pinned, retry later" worktree removal: 31
  such conhosts found at once (Aug 26–29 vintages) pinning FIVE dead worktrees. cargo/rustc
  absence does NOT clear the suspect list — the conhost outlives its client. Holder census
  tool: a ~30-line NtQueryInformationProcess CWD probe (PEB→ProcessParameters→CurrentDirectory)
  over all pids, filtered on the path — names every holder in one pass where no handle.exe
  exists. 13 killed scoped to the two consented reaps; 18 remain pinning gate-w1-a8f04aff (12),
  io-parser-w1 (5), gate-w1-786d2381 (1 after kills) — sweep them at the IR-14/26/27/49 audit
  slot before those removals.
- **LANE-LINKED 2026-08-30 (#242 close sweep):** hertz's queued daemon-leak fixup lane carries
  this entry (brief cites IR-7/17/20/34/35/63); leaves the register when that lane lands.
- **Trigger condition:** next rig/gate-script construction touch, or the first golden red that
  classifies to face (1).
- **Size guess:** small-medium — one wrapper seam plus adopting it in the gate/lane scripts.

### IR-64 — HFENDULEAM's disk floor is set by non-CI bulk: ~845 GB of operator payload leaves the golden box ~2 GB of slack against its own 32 GB preflight

- **Status:** OPEN · **Origin:** doyle, v0.63.0 close sweep 2026-08-26; mechanism first measured
  at the FIELD-SEAL golden window (deployah's IR-31 fifth addendum carries the POOL side of the
  same event by agreed division — this entry is the OTHER reservoir, deliberately not his).
- **Mechanism:** the 1.86 TB C: carries ~614 GB Steam + ~230 GB Downloads (operator payload, not
  CI state). With that floor fixed, normal lane traffic alone walks the box under the 32 GB
  golden preflight — a full-sweep pool WEIGHS ~90–112 GB steady state (IR-59's measurement) and
  five FIELD-SEAL lanes cost 148.48 GB (IR-31 fifth addendum), so ONE milestone's pools exceed
  the entire free margin. Reaping buys windows, not headroom: deployah measured free fall
  167 → 108 GB within hours of his reap, ~59 GB re-consumed by lane builds. The recurring shape:
  every milestone pays a reap-and-measure tax to rent space the box does not structurally have.
- **Why register-worthy rather than "clean up more":** agent-side discipline (IR-31's budgeting,
  pool reaps, teardown steps) is already adopted and still only rents windows — the reservoir
  that would durably move the floor is operator-owned bulk no agent may touch. Naming it here is
  the boundary: agents keep budgeting IN POOLS (not GB); moving Steam/Downloads (or adding a
  disk, or pinning golden to a box without operator payload) is an OPERATOR decision this entry
  exists to put in front of them once, with numbers, instead of re-deriving the floor each cut.
- **Ripe when:** operator rules on the bulk (move/expand/accept-the-tax), OR the first golden
  that dies at the 32 GB preflight despite adopted pool discipline (IR-59's LNK1180-class red).
- **Size:** zero code; one operator decision + at most a runbook line naming the chosen floor.

### IR-65 — kitsubito kernel-audit backpressure: a tailscale-snap AppArmor denial storm + no auditd turned the default audit backlog into a CI-wide stall amplifier (REMEDIATED AT BOX; re-check triggers named)

- **Status:** REMEDIATED-AT-BOX 2026-08-25 (doyle, operator-authorized root — ruling pinned
  releases#225 comment 5418800776); entry stays OPEN as the re-check record because the storm
  SOURCE persists and two named events can silently revert the fix.
- **Mechanism:** `snap.tailscale.tailscaled` (1.92.5) polls /proc and takes an AppArmor
  ptrace-read DENIED per poll — a continuous kernel-audit record storm scaling with process
  count (spawn-heavy serialized CI legs amplify their own storm). With NO auditd installed the
  records rode printk (kauditd throttling) and the 8192 kernel backlog overran
  (lost=1,537,743); the default `backlog_wait_time` 60000ms turns a full backlog into A 60s
  SLEEP INSIDE ANY AUDITED SYSCALL — the broker dispatch stall that stretched IR-30's
  Linux-face race window (see IR-30's Linux-face addendum; deaths clustered 59–63s ↔ this
  knob's value).
- **Remediation (verified at the EFFECTIVE layer, not the fragment):** auditd installed +
  active (storm consumed, backlog drains to 0, lost flat), backlog_limit 32768,
  backlog_wait_time 0 (full backlog may DROP, never STALL). ⚠ THE TRAP PAID FOR ONCE: the first
  knob write (`50-backlog.rules`) silently LOST the augenrules merge — apt's own
  `/etc/audit/rules.d/audit.rules` sorts LAST (digits before letters) and its `-b 8192` /
  `--backlog_wait_time 60000` won; a whole proof run executed under the defaults while the
  fragment grepped perfect. Values now live in the merge-WINNING file; verified `auditctl -s`.
- **Re-check triggers (the reason this entry stays):** (1) tailscale snap update — the plug
  landscape may change, the storm may stop or grow; (2) auditd package update/reinstall — may
  rewrite `rules.d/audit.rules` and re-lose the merge; (3) box reimage. On any of these: one
  `auditctl -s` (expect 32768/0) + one `journalctl -k | grep 'backlog limit'` (expect silence).
- **Ripe when:** a re-check trigger fires. **Size:** two read-only commands per check.

### IR-66 — first-chunk-needle test class: two latent members remain after the v0.63.0 fix (attach.rs:561, :672)

- **Status:** BUILT — confirmed present at `17815c9c` (hertz 2026-09-06); leaves the register at the WEBSERVE close sweep. Hertz rider PR #161, landed ff at `d04b922d` (2026-08-27, IO-PARSER #22
  intake rider). Both members got the later-needle treatment (TICK39 delayed past burst on both
  OS arms; alt-screen entered before delayed ALT_VIEWPORT_MARKER). Gate: doyle — diff-scope
  review; Windows isolated worktree 3× TICK39 PASS + clippy + treqs; Linux CLEAN worktree at the
  PR sha on kitsubito 3× both cells PASS (independent of the builder's dirty-shared-tree proof,
  disclosure on record). · **Origin:** doyle tree-wide census at `dbe3daad`
  (releases#225 comment 5418844404 — the census that corrected my own resume.rs-scoped
  overclaim), after the class's first member was RCA'd and fixed in v0.63.0's head.
- **Mechanism (cite, not restate):** KNOWN-HAZARDS 6.9 same-session face, amended at
  `dbe3daad` — `spawn_session_pid`'s Spawned-wait consumes-and-discards a first output chunk
  that races the reply (~2ms window on a healthy Linux box; load-stretched under backpressure,
  see IR-30 Linux-face + IR-65). A test whose needle exists ONLY in the child's first chunk
  asserts winning a race the contract does not promise.
- **The two members:** `crates/spt-daemon/tests/attach.rs:561` (TICK0..39 instant burst then
  `cat`; needle TICK39 read directly after spawn — the whole burst can sit in the first
  chunks) and `:672` `cross_node_cold_attach_to_alt_screen_gets_clean_repaint` (unix-only;
  `printf '…ALT_VIEWPORT_MARKER'; cat` — needle in the first chunk). Failure shape differs from
  the fixed member: both children idle on `cat`, so a lost prefix is a HANG killed by nextest's
  backstop at exactly 240s (slow-timeout 60s × terminate-after 4) — not a natural-life death.
  Verified NOT in the class: daemon_e2e.rs:209 (needle = echo of post-spawn input), attach.rs:995
  + brain_swap.rs:170 (attach/replay reads, not spawn-waits).
- **Fix shape (ruled at the v0.63.0 close):** hertz lane, NEXT intake — the same one-line
  later-needle / robust-read treatment the resume seed got in `dbe3daad`; deliberately NOT
  landed mid-r4. Both cells passed r4 (0.010–0.058s) — the members are latent, not failing.
- **Ripe when:** next milestone intake (hertz test lane) or the first golden red at ~240s on an
  attach cell — either way the mechanism and fix are pre-derived, triage should cite this entry
  and skip the RCA. **Size:** two one-line test edits.

### IR-67 — wave-battery gap: a `-p spt --bins` leg compiles bin targets as harnesses and runs ZERO integration/e2e tests, so batteries that lean on it carry silent no-coverage legs

- **Status:** BUILT — remedy applied across the IO-PARSER #22 milestone (2026-08-27/28): every
  wave battery named explicit `-p spt --test <suite>` e2e legs beside the `--bins` unit leg
  (GW1..GW5 evidence), every builder evidence report carried a named real e2e per the dispatch
  template, and the one filter mistake (a `--test` name not matching its file) REFUSED loudly
  rather than silently skipping — the failure mode this entry exists to kill. Battery template =
  the dispatch text now carried forward in gate craft.
- **Was:** OPEN · **Origin:** doyle, banked at the SIGNET #218 gates 2026-08-25 ("gate
  battery gap learned"), filed at this sweep per the standing rule.
- **Mechanism:** `cargo test/nextest -p spt --bins` builds `[[bin]]` targets as TEST HARNESSES
  (unit tests inside bins only) — `tests/` integration and e2e suites of the crate NEVER run,
  and the leg's green output is indistinguishable from coverage. Kin of the two banked pool
  faces (`--bins` never emits fixture exes; fresh-pool fixture reds) but this face is about the
  BATTERY TEMPLATE: my SIGNET wave batteries carried a `-p spt --bins` leg believed to cover
  the crate.
- **Remedy:** at next intake, the wave-battery template gains an explicit `tests/` leg for the
  spt crate (nextest `-p spt --test <suites>` or unfiltered `-p spt` where pool history
  permits), and any battery doc that lists `--bins` as a coverage leg gets the one-line caveat.
- **Ripe when:** next milestone intake (the battery template is touched at every intake).
  **Size:** template lines only.

### IR-68 — PSYCHE_INGEST_FAIL cause class unproven: git index.lock did NOT reproduce the ingest failure it was blamed for

- **Status:** OPEN, investigation-shaped · **Origin:** todlando measurement at the W2 int-leg rig
  (IO-PARSER #22, PR #163 at `769fb0a9`, 2026-08-27), reported measure-first; filed by doyle.
- **The contradiction:** the six silent `PSYCHE_INGEST_FAIL:todlando` lines observed at the
  drop-dir probe (2026-08-25) were attributed to shared-checkout git `index.lock` contention. The
  W2 rig planted NON-EMPTY `index.lock` files in BOTH places git takes one (bare git-dir + every
  `<git_dir>/worktrees/<name>/`, mirroring `branchstore::sweep_stale_index_locks`, plant count
  asserted >= 2) — and the ingest COMMITTED ANYWAY (tier write `Written`; `route_slices`
  propagates `commit_live(...)?`, so a blocked checkpoint would have failed it). Non-empty was
  deliberate: KH 1.3 reaps only 0-byte locks, and `pulse_tick` runs no boot sweep. So on this
  store an `index.lock` does not block the checkpoint, and the field incident's mechanism is
  UNKNOWN, not merely unconfirmed. The rig's doc comment records the attempt; the shipped int leg
  fails the ingest at the drop READ instead (upstream of the store).
- **Strongest negative evidence, pinned to the executable rig:**
  `crates/spt-daemon/tests/commune_io_events_int.rs:95-107` documents that the lock mechanism was
  tried **first**, planted non-empty locks in both the bare git-dir and every linked-worktree
  location, asserted that the population was non-empty, and still observed `Written`. Its own
  sentence is the required scope boundary: “This leg does not claim to reproduce a mechanism it
  measured as inert.” The working injection is instead drop-file → directory replacement, which
  fails the upstream read deterministically and proves only the `COMMUNE_FAIL` event contract.
- **NOT claimed:** that the field incident was misattributed to ingest failure generally — the
  files did survive with failed-ingest log lines; what is unproven is the index.lock CAUSE.
- **Ripe when:** next PSYCHE_INGEST_FAIL sighting in the field (capture the failing store state
  before touching it), or a dedicated probe slot. W2's COMMUNE_FAIL event now pushes a NAMED
  reason on every failure, so the next occurrence self-reports its cause — read that first.
  **Size:** investigation; no product change until the mechanism is pinned.

### IR-69 — inherited stderr tokens can tear between format fragments; 26 test consumers parse that surface as structured truth

- **Status:** OPEN, product remedy complete and parked GREEN-AND-READY at `d0fdd58d` on
  `feat/emit-single-write`, owned by **todlando**, awaiting golden-batch assembly; board request
  releases#241 carries the product remedy and full lane evidence in comment 5461613480. This entry
  is the exposure map and census discipline, not the implementation. · **Origin:** CONDUIT #236 r3
  golden attempt 2,
  2026-08-29: `endpoint_autostart_e2e` missed its contiguous
  `ENDPOINT_AUTOSTART:gwauto` keystone although the matching fresh session id proved the replay
  happened. Attempt 3 was clean on the same discarded SHA after 139 GB was reaped; unequal load
  makes the pair corroborating, never a rate.
- **Mechanism, recovered from the literal torn bytes:** two daemon processes inherited one stderr
  pipe. `BRAIN_UP` and `ENDPOINT_AUTOSTART` are each one `eprintln!`, but `write_fmt` may reach the
  handle once per format fragment; the per-process stderr lock cannot serialize the other process.
  The observed interleave split the autostart token between `ENDPOINT_AUTOSTART:` and `gwauto`.
  This is fragment-level cross-process interleaving, not a failed replay and not a two-thread race.
- **Durable remedy, ruled product-side:** render a complete diagnostic line into one buffer and
  issue the complete rendered text, newline included, as **ONE handle call**. An OS short-write
  may retry the late tail so completeness wins over a silently truncated line; the deterministic
  property pins application-side fragmentation, not syscall count under that rare retry. It makes
  no claim that the OS write is atomic on Windows, and no CI re-fire against a 1-of-2 observation
  may stand in for it. Widening one test's grep is refused: it greens one consumer while every
  sibling token remains exposed.
- **Cut-SHA census** (todlando generator, independently hash-verified by hertz), measured from
  `4d6007ac4ee3d0991f4bc60b7e1a55825aad1aaf` via `git show`, not a dirty worktree:
  **486 real emitter sites / 383 distinct tokens / 26 consumer test files**. The generator emitted
  487 site rows, but one is a known phantom: `sealverb.rs:456` is a `format!` inside `map_err`
  that the four-line lookback attributed to a neighbouring real `eprintln!`. Corrected real emitters
  by crate: spt-daemon 270, spt 194, spt-runtime 8, spt-live 5, spt-net 5, spt-store 4; by macro:
  `eprintln!` 484, `println!` 1, `eprint!` 1. `DRIVEN_BY` is the sole stdout row and therefore
  outside the stderr remedy. A consumer means the test both mentions `TOKEN:` and reads a
  stderr/log surface; bare-token matching inflated the population to 72 files by admitting prose.
  This is an **exposure map**, not a rate and not a claim that all 26 have torn.
- **Artifacts (the preserved raw generator outputs contain 487 site rows):** cut census SHA-256:
  sites `f10d6bfe5e0399f98945cf63bb5719a01a62e1251be84165c73030fa9bb224f5`,
  tokens `80b0d1835d9d46376b2f200b1c50be4ca0e185a135168444414255f9308ed92b`,
  consumers `249a2e33df43a30570ca82b5bd200d8b6e2bf07aa8250fcc8bace47fe23e3cbc`.
  The sealed sites file's macro column is wrong on 7 of its 487 raw rows; the corrected raw split
  is `eprintln!` 485 / `println!` 1 / `eprint!` 1. The sealed tokens file has no final newline, so
  `wc -l` reports 382 although it contains 383 logical records. The sealed files remain unedited;
  these corrections are textual.
  Generator `81df82f01ef33a82307920d9804a2394b787ef2ffbe3b41d2d7dda560d9a62a9`
  was recorded as asserting newline-terminated output, but the sealed tokens artifact refutes that
  claim for at least one output. Its remaining guards cover logical record counts, a known-positive
  `SUBSCRIBE_DECISION` sentinel, four-line macro lookback, and truncation only at a
  `#[cfg(test)]` that opens a module.
- **Census hazards paid for while constructing and reconciling the population:** the provisional
  instrument moved `277/243 → 286 → 412 → 487 raw → 486 real` as four silent assumptions were
  found: same-line matching misses multi-line macros; truncating at the first `#[cfg(test)]` drops
  later shipping code; lookback must prefer the token line's own macro over a neighbouring arm;
  and even that lookback can promote a nested token literal beneath a real macro into a phantom
  second site. A file without a final newline also makes `wc -l` silently undercount logical
  records by one. Every future census must name its SHA, assert a known member, count records
  inside the generator, and reconcile emitted rows to real sites.
- **Separate anchored-matcher finding:** `IDLE`, `DISPATCH`, `BUSY`, and `BOUND` are common words
  matched as substrings in exposed consumers. One-write emission does not prevent an unrelated log
  line from satisfying them; those predicates need anchored token matching. Keep this separate from
  the tear remedy so neither defect is claimed to close the other.
- **Kin:** [[IR-50]] (the sink must be read before its absence means anything), [[IR-40]] (a
  confident signal that does not describe the source truth), [[IR-8]] (a plausible census whose
  blind spot returns a clean answer).
- **CLOSE-OUT ADDENDUM 2026-08-30 (#242/v0.66.0 sweep — the entry's own close-out re-census RAN
  and FOUND the predicted blind spot LIVE; entry stays OPEN):** re-census at the landed sha
  `ec6da9b0` (independent containment-classified meter, reconciled with todlando's independent
  meter — full write-up `EMISSION-RESIDUAL-CENSUS-FINDING.md`): the generator truncated
  `crates/spt/src/cli.rs` at its FIRST module-opening `#[cfg(test)]` (line 3007, module closes
  3164, file runs to 39118) and never scanned the rest — **338 TOKEN-shaped shipping sites remain
  bare (332 colon-form, i.e. misses under the sealed census's own predicate, + 6 space-form)**,
  plus `main.rs` 2 sites behind the same mechanism; exactly TWO files tree-wide carry the
  triggering property (first module-opening cfg(test) precedes shipping code) — the generator
  fix's owed blind-spot measurement, answered. The census-hazards bullet above already NAMES this
  third-face truncation hazard; the module-opening fix narrowed the trigger and kept the blind
  spot for files where a test module CLOSES and shipping resumes — IR-8 shape, predicted by this
  entry's own text. The shipped #241 conversions are UNAFFECTED (495 removals ≈ censused 486 +
  space family; the SEEN population is closed; golden green stands). Residual conversion +
  generator fix + enforcement-at-the-seam = **releases#243** (BACKLOG, operator triages);
  delivery.rs:149/313 are post-cut MINTS relative to the census sha (todlando's own, settled
  attribution), deliberately outside #243's census-mechanism scope — and ALREADY CONVERTED: the
  fix rode his breadcrumb branch (`60a12056`) into the r4 head and shipped IN-TAG (todlando
  flagged the stale "his thin lane converts post-cut" claim 2026-08-30; re-measured — zero bare
  TOKEN `eprintln!` in delivery.rs at both `d931dd63` and origin/main). No lane is owed for it. Re-runnable meter
  `EMISSION-RESIDUAL-CENSUS-METER.py` (repo root, untracked; version-control rides #243's lane).
  **Common-word matcher RULED at this sweep (discharges the "decide" clause below):** the four
  common-word consumers (`IDLE`/`DISPATCH`/`BUSY`/`BOUND` substring predicates) get a SEPARATE
  hertz test-side rider — anchored token matching is consumer-predicate work, NOT folded into
  #243 (keeps #243 product-scoped) and not claimed closed by one-write emission; lane-link when
  hertz's spool drains.
- **Closes when:** releases#243's residual conversion + generator fix land and a re-census at
  that sha reads zero shipping residue under both delimiter forms. **Size:** product lane
  complete elsewhere; register follow-through small.

### IR-70 — a hand-built `IoBus` silently omitted later sinks while every gate stayed green

- **Status:** OPEN, remedy unscheduled; the v0.65.0 respin fixed the witnessed call site on PR #173.
  This entry owns the recurrence class, not that shipped correction. · **Origin:** releases#234,
  commit `c5459795167`, CONDUIT #236 respin.
- **Mechanism:** `publish_commune_io` constructed `IoBus` itself and registered only the sink it
  knew. `default_bus` already existed as the assembly point and its module contract explicitly says
  emitters must know only `IoBus::publish`, so adding a later sink costs one registration rather
  than an emitter sweep. The hand-built publisher recreated exactly the drift that contract warned
  against: `COMMUNE` and `COMMUNE_FAIL` reached the old sink but never the adapter log, leaving
  `spt api io-events` permanently short on two documented kinds.
- **Why the complete battery was green:** the e2e was vacuous on the broken kinds; units synthesized
  rows without traversing `publish_commune_io`; and traceability checked attached tags, not whether
  the tagged path exercised the real publisher. Three green instruments shared one blind seam.
  The respin routed the site through `crate::iobus::default_bus`, and its source comment now records
  the failure shape.
- **Durable remedy direction (todlando mechanism, not scheduled):** enforce the assembly point so
  the next hand-built bus is impossible rather than found — make production publishers obtain the
  composed bus through one construction API, and gate the real publish path for every documented
  kind. A search or comment is not enforcement; another emitter can satisfy both while rebuilding
  a partial sink list.
- **Kin:** [[IR-39]] (a green fixture never reached the missing dependency), [[IR-37]]
  (traceability tags prove coverage bookkeeping, not behavioral truth), [[IR-58]] (a hand-built
  population returns an orderly but incomplete answer).
- **Ripe when:** the next `IoBus` construction/API touch, or a new sink registration. **Size:**
  small-to-medium assembly-point enforcement plus a real-path population test.

### IR-71 — controller-seat release has no persisted breadcrumb, so tests cannot distinguish propagation lag from a missing detach

- **Status:** OPEN, observability gap; no remedy scheduled. · **Origin:** CONDUIT #236 r3
  `er_brief_once_per_session_e2e` intermittent, classified test-side and fixed on PR #174.
- **What/why:** killing rc1 reaps the controller process, but the broker is another process and
  notices the socket close later. On that detach edge it clears `driven_by` and `controlled` in
  `info.json`; rc2's preflight reads that file locally and can race the write, refuse normally, and
  exit 0 without dialing the broker. The test originally spawned rc2 immediately and had no
  observable precondition separating “release is propagating” from “release never happened.”
- **The missing surface:** broker lifecycle state records `session-detach was_controller=true`, but
  there is no always-on persisted breadcrumb that states the controller stamp was cleared, names
  the resulting `driven_by`/`controlled` state, or measures detach-to-persist latency. The repaired
  test therefore polls `info.json` itself, prints elapsed time (0.075s on the first local real run),
  and uses a named 60s timeout whose failure promotes the finding from timing to missing release.
  Its deliberately latched-seat unit proves the barrier can red.
- **Remedy direction:** add one structured release breadcrumb after the persisted clear, carrying
  session/connection identity and the resulting control state; if latency is carried, measure it
  from the detach edge. It must ride the persisted daemon sink and remain distinct from
  `session-detach`, which proves the connection event but not the file write.
- **Kin:** [[IR-50]] (the sink a test must actually read), [[IR-40]] (an event timestamp is not a
  completion signal), REQ-HAZARD-CONTROL-STAMP-CONVERGENCE.
- **Ripe when:** the next controller lifecycle or structured-breadcrumb touch, or a field timeout
  from PR #174's barrier. **Size:** small emitter plus one contract-level assertion.

### IR-72 — no product read verb surfaces which process holds a perch or controller seat

- **Status:** OPEN, observability gap; no remedy scheduled. · **Origin:** CONDUIT #236 r3
  two-host RCA and PR #174 (`70c1a303`), corrected by doyle after IR-18 and daemon-status were
  ruled out as different classes.
- **Source-verified boundary:** the holder identity already exists in durable/internal records
  (`InfoJson.pid`, `parent_pid`, and the broker's session pid), and product internals read it for
  liveness, teardown, and self-detection. No product **read verb** returns the process that holds a
  named perch/seat. PID-bearing teardown messages are failure outcomes, not an inspection surface;
  daemon status is process-wide, not perch- or seat-scoped. This also does not collapse into
  [[IR-18]], where an existing internal `read_pid` collapses absent and unreadable records.
- **Paid-for consequence 1 — rigs manufacture custody:** PR #174's two-host repair had to add
  `PidHolder` in `crates/spt/tests/twohost_cli.rs:401-437`, spawn a long-lived sibling process,
  write that pid directly into each synthetic perch, retain the child handle, and kill+wait it on
  drop. The rig can make a known holder; it cannot ask the product which holder the product sees.
- **Paid-for consequence 2 — ambiguity stayed latent until behavior flipped:** the r3 fixture
  seeded three perches with one test-process pid. Self-detection then enumerated a directory whose
  first matching row was a filesystem-order coin, so a green could name the wrong perch. The
  repaired resolver refuses an ambiguous candidate set instead of guessing, and the fixture's
  `assert_only_ancestor_candidate` now re-reads every `info.json` rig-side
  (`twohost_cli.rs:439-457`). Refusal is enforcement, not observability: it proves a tie but still
  offers no supported verb that names each holding process.
- **Remedy direction:** add one machine-readable, named-perch inspection surface that reports the
  custody fields the product actually used — holder pid and role/source (bind relay, stable harness
  parent, brokered controller/session) — without asking callers to parse private `info.json` or
  scan the process table. PID alone is recyclable; where a birth/image/session stamp exists, carry
  it so the output does not become a new bare-pid oracle.
- **Kin:** [[IR-18]] (read primitive loses error class, explicitly distinct), [[IR-71]] (release
  completion lacks a persisted breadcrumb), REQ-HAZARD-SELF-DETECT-TIE (refuse ambiguous ancestry
  rather than select by enumeration order).
- **Ripe when:** the next roster/status read-model change or another rig needs to reap a named
  holder. **Size:** small read-model/API addition, medium if seat and perch custody need separate
  typed variants.

### IR-73 — ci.yml and release.yml still read the free-space floor PRE-checkout (literal-first sites off the golden path); and any job's floor is stale by its last heavy step

- **Status:** OPEN, filed by doyle 2026-08-30 at the #242/v0.66.0 close sweep (deployah flag
  2026-08-29, widen lane). · **Origin:** the #242 cut lost two golden attempts to exactly this
  REQ-CI-FREE-SPACE-PREFLIGHT trap on `golden.yml`'s n1-gate (floor read BEFORE checkout, so the
  checkout's own cost lands after the assertion); `golden.yml` is fully reshaped (`ac7d2609`,
  both jobs, checkout-protective reaps verified to stay pre-checkout) — but `ci.yml` (5 tagged
  sites) and `release.yml` (3 tagged sites) keep the literal-first shape, deliberately NOT
  widened off the back of a golden.yml-scoped ruling.
- **What/why:** review each of the 8 sites' comments for checkout-protective purpose (the
  `ac7d2609` method), then reorder the read after checkout or RECORD why not, per site. The
  "other jobs keep the literal-first shape" clause is the trap's carrier: every un-reviewed site
  is a queued golden-attempt loss on a loaded box.
- **Second half (from #242 r4, same family):** a single job-start floor read is stale by the
  job's LAST heavy step — the r4 LNK1318 low-water face recorded on [[IR-59]] (sixth face);
  the within-job re-read before the docs-drift build rides the next workflow commit alongside
  this entry's reorders.
- **Kin:** [[IR-46]] (a floor asserts an instant), [[IR-59]] (full volume reds as a linker
  defect), REQ-CI-FREE-SPACE-PREFLIGHT.
- **Ripe when:** next touch of `ci.yml`/`release.yml`, or the next intake's tooling-rider slot
  (compose with IR-59/IR-57's standing pair). · **Size:** small — reorder or comment per site,
  8 sites.

### IR-74 — kitsubito holds 21,643 /tmp endpoint homes: box-side accumulation nothing bounds or measures

- **Status:** OPEN, filed by doyle 2026-08-30 at the #242/v0.66.0 close sweep; measurement by
  hertz (2026-08-29). · **Origin:** rig work on kitsubito found `/tmp` carrying 21,643 endpoint
  homes — per-run spt homes that outlive their runs.
- **What/why:** sibling of the [[IR-31]]/[[IR-59]]/[[IR-64]] family (box accumulation the
  preflights never see): each is small, the population is not, and /tmp on a long-lived box is
  never reaped by CI. Un-reaped homes also widen every by-path census a rig runs. Needs a
  bounded lifecycle (per-run temp root reaped at run end, or a box-side sweep with a stated
  retention), and the reap must state its count so a zero is a claim.
- **Ripe when:** next kitsubito rig touch, or the between-milestone box-audit slot.
- **Size:** small (one sweep script + adoption in the box's rig steps); the measurement rerun is
  one `find | wc`.

### IR-75 — traceable-reqs is not installed on kitsubito, so every Linux gate driver's treqs leg exits 127 (vacuous, wrapper-green shape)

- **Status:** OPEN, filed by doyle 2026-08-30 at the #242/v0.66.0 close sweep (deployah
  co-flagged at intake). · **Origin:** ASM-241 Linux battery — `ASM_treqs.exit` = 127, raw
  `env: 'traceable-reqs': No such file or directory`; leg vacuous, covered that day by the
  Windows treqs = 0 at the identical tree (treqs is a platform-independent text scan).
- **What/why:** a 127 leg is exactly the wrapper-green shape [[IR-60]] documents — a driver that
  doesn't read the exit FILE would report the leg green. Fix is either: install traceable-reqs on
  kitsubito (then the leg is real), or DROP the leg from Linux gate drivers and state the
  Windows-covers-it rule in the driver comment. Do NOT reopen [[IR-37]] for this — IR-37's
  consume lane is spt-core CI (GitHub runners, green all week); this is BOX tooling, a different
  surface (ruled at the sweep against the enumerating agent's overcautious flag).
- **Ripe when:** next Linux gate-driver touch or kitsubito box-audit slot.
- **Size:** trivial (install or one driver edit + comment).

### IR-76 — hfenduleam is the golden Windows runner AND the builders' box: a local workspace nextest fully contained a golden Phase A and manufactured a load-timing red, and the job-start census cannot see it

- **Status:** OPEN, filed by doyle 2026-09-06 during v0.67.1 golden triage (run 34017906638
  att2). · **Origin:** the Windows Phase A red on
  `spt-store::wtlock_two_process_int two_processes_commit_into_one_worktree_without_failing`
  (B hit the 10 s worktree-lock bound while A's six serial commits took 13.5 s; 12 other
  cells ran >10 s in the same phase). Load audit by todlando from his own `.raw`
  CreationTimeUtc / `.exit` LastWriteTimeUtc: his #276 gate's 1240-cell workspace nextest
  (`--no-fail-fast`, test phase 08:50:06–08:58:00Z) FULLY CONTAINED the golden Phase A
  (08:52:38–08:56:49Z), every core saturated. He had been dispatched into that window and
  never told the box was under a golden hold; perri and hertz had been.
- **What/why:** two mechanisms, both unmeasured by anything in CI today. (1) The box is
  shared by design — every agent's gate rig, adapter build and workspace nextest runs on the
  same host as the self-hosted golden Windows runner, so a golden Windows leg is a QUIET
  WINDOW only if every agent knows it is open; nothing says so, and a dispatch brief that
  starts a gate is exactly how a builder lands inside it. (2) The job-start
  `reap-census.ps1` counts the SPT FAMILY only (`family_total=12 scoped=0 unscoped=12` at
  att2's 08:39:46Z start, with a builder's prebuild live on the box); cargo/nextest/rustc
  load is outside its predicate, so the one instrument that runs at the right moment is
  blind to the load that matters. Every load-timing red on Windows therefore triages as a
  random victim with the cause off the record — the [[IR-63]] "one random victim per sweep"
  face, now with a named source.
- **Fix shape:** (a) driver templates (every gate/battery driver an agent authors here)
  gain a PRE-FLIGHT leg with THREE outcomes, never two (todlando's amendment at filing: an
  empty `gh run list` is the same bytes for CLEAR and for a broken meter — wrong filter,
  expired token, rate limit, wrong repo default all return exit 0 and no rows — so a
  two-arm guard fails OPEN into the exact hazard it guards, the `--commit <short-sha>`
  confident-EMPTY class): run `gh run list --workflow golden --status in_progress`; rows
  present -> HOLD, print the run id; no rows AND a positive control (`gh run list -L 1`
  unfiltered, or `gh auth status`) returns a row -> CLEAR; control empty or erroring ->
  REFUSE and say the meter is broken. A refusal, not a warning, so a wrapper-green shape
  cannot swallow it ([[IR-60]]); the dispatch-brief line in (b) is likewise a MEASURED
  clear with a live meter, never the absence of a sentence; (b) dispatch briefs that start a gate NAME the hold state of the box; (c)
  the job-start census gains a second predicate — cargo/nextest/rustc/link processes by
  image name, count + total CPU — reported beside the family count so a red's triage can
  place foreign load from the log alone; (d) the golden hand-off contract
  (docs/golden-head-intake) states the quiet-window rule once.
- **Kin:** [[IR-63]] (random victim per sweep), [[IR-59]] (box conditions the preflights
  never see), [[IR-64]] (hfenduleam's non-CI payload), FLAKE-LEDGER rows for
  `wtlock_two_process_int` and `attach_link_push_e2e` (2026-09-06).
- **Ripe when:** next driver-template touch (a/b are text; a is one leg), next `ci.yml`
  touch for (c) — compose with [[IR-73]]'s reorders; (d) rides the next runbook edit.
- **Size:** small — one pre-flight leg + one census predicate + two doc sentences.

### IR-77 — nothing stamps the interval between a daemon child's spawn and the brain's first log line, so a readiness-deadline red cannot say WHERE the time went

- **Status:** OPEN, filed by hertz 2026-09-07 from the W1 #249 kitsubito battery at the gated
  sha `8d980fdf`; doyle-ruled at filing to be a register entry rather than a lane rider. ·
  **Origin:** FLAKE-LEDGER `resident_service_e2e` :453 PRECONDITION (53 s) and
  `resume_no_control_steal_e2e` (46 s), co-victims of ONE window — nextest exit 100, ONE
  Summary `2999 tests run: 2997 passed (8 slow, 1 leaky), 2 failed, 1 skipped`. Evidence
  preserved at `.spt/preserved/w1-kitsubito-8d980fdf/nextest.raw`, sha256
  `9c456e21a6a0f75d8b0375ac648f232dba07f3887ec5e8dc232daef36253b51e`, hash-verified against
  the kitsubito original.
- **What/why:** both reds are a 45 s wait on `brain.ready` that expired, and the panel each
  one prints shows the tree ALIVE — `BRAIN_UP`, `BRAIN_PHASE:announce done in 1ms`,
  `BRAIN_PHASE:resume done in 0ms`, `SERVICE_STARTED` for both services, both later reaped
  with an empty survivor set. So the daemon came up; it came up after the clock. WHERE those
  45 s went is not recoverable from any artifact we keep: the first timestamped line in the
  daemon's own sink is emitted by a process that has already started, and nothing stamps the
  interval from the test's `Command::spawn` to it. The consequence is that every red of this
  family is triaged by inference — v0.66.0 measured a ~10.1 s exe-hash on the ready path, and
  that number is a CANDIDATE for the gap and can never be more than a candidate while the
  interval is unmeasured. Remedy: one monotonic breadcrumb at daemon-child entry (before any
  work), a second at the point `brain.ready` is stamped, both into the stderr sink the panel
  already renders — then a red of this family reports its own split without a rerun.
- **Kin:** [[IR-71]] (a seat release with no breadcrumb, same shape one layer over),
  FLAKE-LEDGER `resident_service_e2e` :453 and `resume_no_control_steal_e2e` (2026-09-07),
  and the HEAVY census stanza in `.config/nextest.toml`, which hardens the RECIPE these reds
  came from and deliberately leaves the deadlines alone — a budget retune is the same race
  with a different number, and it would also destroy the only signal this entry wants.
- **Ripe when:** the next touch of the daemon boot path or of `brainproc`'s ready stamp; also
  ripe as a rider on any lane that re-opens this test family.
- **Size:** small — two `emit_line_err!` breadcrumbs and the sentence in the panel that says
  what they mean.

### IR-78 — `SPT_TEST_EPHEMERAL_ADVISORY_PORTS` silently overrides an explicitly set `SPT_DOCS_PORT`

- **Status:** OPEN, filed by hertz 2026-09-07, MEASURED by doyle the same night on his W2
  field pair (both boxes). · **Origin:** the 5474 rig-hygiene sweep in
  `test/rig-advisory-ports-and-heavy-class`, which sets the rig flag at 37 `spt daemon run`
  spawn sites.
- **What/why:** `resolve_daemon_docs_port` (crates/spt-daemon/src/docshost.rs) returns `0`
  whenever the rig flag is set, BEFORE the `SPT_DOCS_PORT` env override is consulted. With
  both set, doyle's rig daemons took ephemeral ports (`DOCS_SERVER_UP` on 55369 Windows,
  44015 Linux) while `SPT_DOCS_PORT=5480` was ignored without a word. The precedence itself is
  defensible — the rig flag exists precisely so a test tree cannot take a well-known port —
  but a SET override discarded in SILENCE is the shape IR-37 files under a different name: the
  operator is left believing the value they set is the value in force. Remedy is a line, not a
  redesign: emit one notice naming both env vars when the flag wins. No test cell is proposed
  with it — none of the 31 swept binaries asserts on a docs port, so a cell here would assert
  a behaviour nothing consumes.
- **Kin:** [[IR-37]] (a set input silently ignored), the 5474 hygiene sweep, and the
  `DOCS_SERVER_BIND_FAIL: port 5474: Address already in use (os error 98)` lines in
  `.spt/preserved/w1-kitsubito-8d980fdf/nextest.raw` that started the sweep.
- **NOT measured, stated so nobody reads it as covered:** doyle measured that the WMI
  auto-start rung DOES carry the caller's environment (rig daemon pid 50088 on hfenduleam:
  `DAEMON_LAUNCH_VIA_WMI`, its environ carrying `SPT_HOME`, `SPT_DOCS_PORT` and the rig flag
  as his shell set them, `brain.ready` in the rig home). The schtasks (at-logon) rung was NOT
  measured, and neither was the unix path.
- **Ripe when:** the next `docshost.rs` touch, or the first time someone sets `SPT_DOCS_PORT`
  on a rig and believes it.
- **Size:** tiny — one emitted notice.

### IR-79 — 22 of the 31 rig binaries leak their daemon tree on a failing assert: teardown is a statement, not a guard

- **Status:** OPEN, filed by hertz 2026-09-07 on doyle's dispatch during PR #198's CI window,
  from todlando's red 6 (a stale same-port daemon from his own previous run answering the next
  one). · **Origin:** PR #198 moved 31 rigs to `SPT_TEST_EPHEMERAL_ADVISORY_PORTS=1`. Under an
  ephemeral port a leaked daemon can no longer ANSWER the next run, so that PR is right as it
  stands and this entry is not a defect in it; what remains is the orphan process itself — it pins
  `spt.exe` (the Windows delete-and-rebuild hazard) and holds an `SPT_HOME` that nothing reaps. · **SCOPE BROADENED 07:20Z:** the title says "on a failing
  assert" because that is the face it was filed from; the SECOND FACE bullet below shows the same
  mechanism firing on a nextest TIMEOUT and leaking THREADS rather than a daemon. Read the title as
  the handle, not the boundary — the trigger is any exit that does not reach the teardown statement.
- **What/why:** measured over the 31, classifying each binary by whether its daemon teardown
  survives a panic (a failing `assert!` unwinds, so only a guard or an ordering discipline saves
  the child):
  - **panic-safe, 9:** `impl Drop` guard — `endpoint_autostart_e2e`,
    `knock_mutual_cross_node_e2e`, `twohost_cli`; `catch_unwind` teardown —
    `activity_link_push_e2e`, `attach_link_push_e2e`, `wake_resume_bind_e2e`;
    teardown-then-assert with zero exposed asserts — `endpoint_teardown_authority_e2e`,
    `er_briefing_session_scoped_e2e`, `er_sequestered_cwd_e2e`.
  - **leaks on a failing assert, 22** (exposed asserts / total asserts in the test body):
    `projindex_writer_e2e` 21/21, `projindex_reader_e2e` 18/18,
    `live_adapt_translation_swap_e2e` 15/38, `brain_split` 12/12,
    `er_briefing_presented_e2e` 10/28, `rc_attach_truth` 8/29, `brain_respawn_rename` 6/6,
    `dummy_harness_e2e` 4/11, `idle_edge_drain_e2e` 4/11, `multi_subnet_bringup_e2e` 4/20,
    `bind_honest_cross_perch_e2e` 3/6, `idle_edge_seal_e2e` 3/17, `resident_service_e2e` 3/24,
    `attach_wedge_e2e` 2/8, `bind_cwd_project_e2e` 2/8, `daemon_refresh_e2e` 2/11,
    `er_brief_once_per_session_e2e` 2/13, `n1_pairing` 2/5, `resume_template_e2e` 2/10,
    `run_no_dup_session_e2e` 2/14, `er_briefing_presentation_e2e` 1/4,
    `livehost_bootgate_e2e` 1/3.
  **FIRST WAVE of the Drop-guard generalization lane (doyle-ruled 2026-09-07, composed at the next
  register sweep, not now): `projindex_writer_e2e`, `projindex_reader_e2e`, `brain_split`,
  `brain_respawn_rename`.** These four expose EVERY assert they have — their teardown is the last
  statement in the body, so any red at all leaks, which makes them both the worst cases and the
  cleanest proofs that a guard works. The remedy already exists in this tree —
  todlando's `DaemonReaper`, a `Drop` guard armed BEFORE the first CLI call — and the follow-up
  lane is to generalise it into `crates/spt/tests/common` and adopt it at these 22 sites, which is
  also the only shape that covers a `SIGKILL`-free timeout kill by nextest.
- **METHOD, and its limits, so the count can be re-derived and challenged:** the classifier reads
  each `#[test]` body, resolves file-local helper fns whose own body tears down (so a teardown
  called through `sweep()` counts), and reports every `assert!`/`panic!` positioned before the LAST
  teardown call in that body. It therefore (a) misses asserts written inline inside a closure or a
  macro argument rather than at statement position, (b) treats the last teardown as THE teardown —
  a partial earlier teardown still leaves the tree, so the true exposure is >= this count, and
  (c) does not model `?` or early `return`. Spot-verified by hand on `resident_service_e2e`, whose
  3 exposed asserts are real: they fire in the `wait_until` legs well before the `sweep()` at :435
  that the "── ASSERTIONS ──" block follows.
- **TONIGHT'S LIVE SAMPLE, so this is not an academic count (doyle's hfenduleam census 06:43Z):**
  10 leaked `spt.exe` from todlando's pool, pairing by start time, from his pre-yield legs — a
  single evening's rig work on one box. Each pins the binary against a rebuild and holds an
  `SPT_HOME` nothing reaps.
- **SECOND FACE — the same hazard in THREAD clothes, and it blocks the RUNNER, not just the box
  (measured by todlando 07:17Z, relayed by doyle 07:20Z; I did not measure it myself and record it
  as his testimony):** a nextest TIMEOUT on a `twohost_web` cell left the cell's child process
  holding the broker + listener THREADS. Consequences, with his numbers:
  - nextest itself blocked **14 minutes** at **0.61 CPU-seconds** — a wall-clock hang with
    essentially no CPU, which is the signature of a parent waiting on a child's pipe rather than
    of work being done;
  - **`a.exit` was never written**, while **every verdict was already in `a.raw`** — so the leg
    read as INCOMPLETE at the exact moment its results were complete. A reader who trusts the
    exit file over the raw would call this a hung or failed leg and rerun it;
  - **killing the path-verified orphan let the parent finish at once** — which is the causal test,
    not a correlation: the kill is the intervention and the unblock is the response.
  **Why it belongs in THIS entry rather than a new one:** the leaked thing is a thread inside a
  child, not a daemon, but the mechanism is identical — a teardown that is a STATEMENT does not run
  when the body does not reach it, and a timeout kill reaches the body even less reliably than a
  panic does. The `Drop`-guard lane covers both faces with one remedy, which is doyle's ruling and
  the reason no separate entry is opened. It also raises the lane's value: the daemon face costs a
  pinned binary and an unreaped `SPT_HOME`, this face costs **14 minutes of a serialized golden
  leg** and manufactures a false INCOMPLETE.
  **Reading rule this hands the gate, worth stating because it is cheap:** when a leg's `.exit` is
  missing, read the `.raw` for a Summary BEFORE concluding the leg hung — kin to
  [[a-stopped-local-ssh-does-not-stop-its-remote-command]], where the raw was likewise the honest
  record and the wrapper was not.
- **Kin:** FLAKE-LEDGER `resident_service_e2e` :664 teardown LEAK row (3 occurrences, the Windows
  face of the same class), [[IR-34]], [[IR-74]] (kitsubito's 21,643 `/tmp` endpoint homes — this is
  one of the producers), and todlando's red 6.
- **Ripe when:** the next lane that touches `crates/spt/tests/common` — the guard belongs there,
  not copied 22 times.
- **Size:** medium — one guard in `common`, then 22 mechanical adoptions, each provable by
  panicking the body under a temporary test and watching the census go quiet.

### IR-80 — five leaky cells on ONE module (`brainproc` / `supervise_brain`), Windows only: a candidate PRODUCT leak on the promotion/rollback path, or a cluster of names — nobody has looked yet

- **Status:** OPEN, filed by hertz 2026-09-07 07:40Z on doyle's ruling (07:38Z) that this is its
  OWN entry and **not** a third face of [[IR-79]] — IR-79 is rig teardown, this is a candidate
  defect in the PRODUCT path, so its owner may be todlando rather than the rig lane. ·
  **Origin:** noticed while measuring the #198 HEAVY scan-root gap ([[IR-37]] rider 2); the leak
  roster was read out of the baseline golden log for an unrelated reason.
- **⚠ WHAT THIS ENTRY IS, STATED FIRST SO IT IS NOT OVERREAD: a CLUSTER OF NAMES.** Five cells
  that leaked share one module and one code path by their identifiers. Nothing has been measured
  about WHAT is left alive, and a shared module is not a shared cause. Do not cite this as a
  product defect until the census below has run.
- **The observation, exact.** Golden run **34017906638**, sha `04e32c8c95cf`, both boxes, Phase A.
  Windows (hfenduleam) reported **9 leaky** cells; Linux (kitsubito) **1** (`livehost::tests::
  legacy_psyche_sweep_guard_is_id_specific_and_fail_safe`, unrelated). **Five of the Windows nine
  are the same module:**

  | cell | kind | time |
  |---|---|---|
  | `spt-daemon brainproc::tests::clear_before_spawn_defeats_exact_generation_stale_file` (:2200) | lib | 0.457s |
  | `spt-daemon brainproc::tests::ready_but_old_gen_never_drains_does_not_promote_rolls_back` (:1834) | lib | 0.554s |
  | `spt-daemon brainproc::tests::stale_generation_minus_one_ready_never_promotes` (:2155) | lib | 0.731s |
  | `spt-daemon brainproc::tests::trial_kills_alive_never_ready_candidate_before_rollback` (:2109) | lib | 0.582s |
  | `spt-daemon::false_promote ready_candidate_does_not_promote_until_the_wedged_old_gen_conn_drains` | int | 1.744s |

  Denominator for the lib half: `brainproc`'s test module holds **26** `#[test]`/`#[tokio::test]`
  cells, so **4 of 26 leaked, and all four are the trial/promote/rollback cells by name.** The
  fifth is the integration rig that drives the same path.
- **Why it is worth an entry even unmeasured:** four of the five are `kind(lib)`, so they run in
  **`ci.yml` on every push**, not only in golden — if something really is left alive it is being
  left alive on the shared Windows box many times a day. That is the same standing cost as
  [[IR-79]]'s daemon face, arrived at from the opposite direction.
- **HYPOTHESIS, labelled as such and NOT measured.** `supervise_brain` (:935) takes an injected
  `spawn_child: impl FnMut(...) -> io::Result<Child>` (:941) and reaps candidates with bare
  `let _ = child.kill();` at :998, :1015, :1055, :1071, :1081 — **`kill()` with no `wait()`.** On
  Windows termination is asynchronous, and a killed-but-unreaped child can still hold the
  inherited stdout/stderr pipe past the test's exit, which is exactly what nextest reports as
  LEAK. The production spawner's own doc comment at :1190 states the intent — "dies with no
  orphaning. NOT `spawn_detached`" — so if this is the mechanism, the code's intent and its
  behaviour have diverged on one platform. **Every clause of this bullet is a reading of source,
  not an observation of a live process. It may be entirely wrong.**
- **FIRST MEASUREMENT — and doyle sharpened it 07:41Z from a census into a FALSIFIER, which is the
  form to run:** on a free Windows box after W2, run
  `trial_kills_alive_never_ready_candidate_before_rollback` **alone** and read one thing —
  **is the killed pid still present, with the pipe handle open?** That is a yes/no against the
  hypothesis above, not an open look at what happens to be running, and it is the difference
  between a reading that settles the entry and a reading that produces more names. Capture pid,
  image path and parent alongside the answer so a YES is immediately actionable, but the ANSWER is
  the deliverable. One cell, one read. **Owner stays UNASSIGNED until it is taken** (doyle) — a
  product path and a rig artifact are not distinguishable from here, and assigning before the read
  would pick one by guess.
- **Limits, so the entry cannot be overread:** nextest `LEAK` is **not a failure** — none of these
  cells reds anything, and the run was green. Windows-only in **one** run; no trend established.
  The clustering is by identifier and shared module, which is suggestive and is not causation.
  I have not run any of these cells. Everything above is source plus one golden log — no box, no
  build.
- **Kin:** [[IR-79]] (children outliving the cell — the rig-teardown face of the same symptom, and
  the reason doyle ruled these must stay SEPARATE: same symptom, different suspected owner),
  [[IR-37]] rider 2 (the measurement that surfaced it), FLAKE-LEDGER's `resident_service_e2e` :664
  teardown LEAK row.
- **Ripe when:** the first free Windows box after W2 lands — it is a single-cell census, so it
  fits any gap.
- **Size:** unknown until the census. Tiny if it is a missing `wait()`; not tiny if the promotion
  path leaves a real candidate brain alive on rollback.

---

#### THE LEAK SET IS NOISY — measured across three shas, 2026-09-07 (hertz)

**This retracts the "8 -> 5 leaky, direction favourable" reading I flagged at 09:31Z.** It is not a
fix and not an improvement; it is threshold noise, and the census this draft was built on inherits
that noise.

Windows `ci` unit leg (`--workspace -E 'kind(lib) + kind(bin)'`), three consecutive shas:

| cell | ff4b405d | e3bd53d4 | 401a19ad |
|---|---|---|---|
| brainproc::clear_before_spawn_defeats_exact_generation_stale_file | ● | ● | ● |
| brainproc::ready_but_old_gen_never_drains_does_not_promote_rolls_back | ● | ● | ● |
| brainproc::trial_kills_alive_never_ready_candidate_before_rollback | ● | ● | ● |
| spt-runtime runtime::bounded_run_kills_on_timeout | ● | ● | ● |
| brainproc::stale_generation_minus_one_ready_never_promotes | ● | ● | — |
| spt-live digest::extractor_timeout_errors | ● | ● | — |
| spt-live history::fetcher_timeout_errors | ● | ● | — |
| spt-daemon broker::windows_session_is_zombie_sees_a_handle_held_corpse_as_dead | ● | — | — |
| spt-daemon shellwake::kill_waker_at_still_kills_a_matching_pair | — | ● | ● |
| **total** | **8** | **8** | **5** |

**Causality check, which is what makes this a retraction rather than a hypothesis:** the three
commits between e3bd53d4 and 401a19ad touch **zero files** in brainproc, spt-live, spt-runtime,
shellwake or broker (`git diff --name-only` over those paths returns nothing). The count fell by 3
with no change to any leaking module. So the drop cannot be a fix; the membership simply moves.
Nine cells appear in the union, only **four** are present at all three shas, and one
(`shellwake::kill_waker_at`) *appeared* rather than vanished — a count that falls while a new member
joins is the signature of a threshold, not of a repair.

#### Three consequences for this draft

1. **A count is not a quality signal here, in either direction.** Any future "leaky went down"
   claim needs the NAME diff and a causality check against the changed files, or it is noise
   reported as progress. I made exactly that error and it survived about twenty minutes.
2. **The five-cell cluster is really "three stable + one intermittent".**
   `stale_generation_minus_one_ready_never_promotes` is absent at 401a19ad. The cluster framing
   still holds — three brainproc cells leak at every sha and they are all one module — but the
   membership count must not be quoted as fixed.
3. **doyle's falsifier beats my census, and this is the evidence.** The census (count the leaky
   cells, name the cluster) is exactly what the noise destroys. His falsifier — *after
   `trial_kills_alive_never_ready_candidate_before_rollback` on Windows, is the killed pid still
   present with the pipe handle open?* — names ONE cell and probes a MECHANISM, and that cell is
   **leaky at all three shas**, so it is the most stable target available. A mechanism probe is
   immune to the threshold that moves the population. Owner stays UNASSIGNED until that read.

#### One structural finding the noise does NOT touch
**All nine union members spawn a child and then kill it, let it time out, or inspect its corpse** —
`trial_kills_alive`, `kill_waker_at`, `bounded_run_kills_on_timeout`, `extractor_timeout_errors`,
`fetcher_timeout_errors`, `windows_session_is_zombie_sees_a_handle_held_corpse_as_dead`, and the
three brainproc spawn/promote/rollback cells. Not one unrelated cell is in the set. Membership
fluctuates; the *kind* of cell does not. That is a real signal about where child-process teardown
is unreliable on Windows, and it is the same territory as [[ir79]] (rigs leaking daemons on a
failing assert) and [[ir81]] (kill scoping) — three drafts converging on one seam.

---

#### THE 12-CELL UNION — 4 shas x 2 OSes, membership beside the invariant (hertz, 2026-09-07 10:15Z)

Machine-generated from the eight ci unit logs, not transcribed (a hand-typed 12x8 matrix is where
transcription errors live). `X` = reported LEAK at that sha on that box.

| cell | W ff4b405d | W e3bd53d4 | W 401a19ad | W f3c8495b | L ff4b405d | L e3bd53d4 | L 401a19ad | L f3c8495b |
|---|---|---|---|---|---|---|---|---|
| `spt-daemon brainproc::tests::clear_before_spawn_defeats_exact_generation_stale_file` | X | X | X | X | . | . | . | . |
| `spt-daemon brainproc::tests::ready_but_old_gen_never_drains_does_not_promote_rolls_back` | X | X | X | X | . | . | . | . |
| `spt-daemon brainproc::tests::stale_generation_minus_one_ready_never_promotes` | X | X | . | X | . | . | . | . |
| `spt-daemon brainproc::tests::trial_kills_alive_never_ready_candidate_before_rollback` | X | X | X | X | . | . | . | . |
| `spt-daemon broker::tests::windows_session_is_zombie_sees_a_handle_held_corpse_as_dead` | X | . | . | X | . | . | . | . |
| `spt-daemon livehost::tests::legacy_psyche_sweep_guard_is_id_specific_and_fail_safe` | . | . | . | . | X | X | X | . |
| `spt-daemon shellhost::tests::kill_shell_at_still_kills_a_matching_pair` | . | . | . | X | . | . | . | . |
| `spt-daemon shellwake::tests::kill_waker_at_still_kills_a_matching_pair` | . | X | X | X | . | . | . | . |
| `spt-live digest::tests::extractor_timeout_errors` | X | X | . | X | . | . | . | . |
| `spt-live history::tests::fetcher_timeout_errors` | X | X | . | . | . | . | . | . |
| `spt-runtime runtime::tests::bounded_run_kills_on_timeout` | X | X | X | X | . | . | . | . |
| `spt-store proc::tests::process_cmdline_reads_a_live_arg_marker` | . | . | . | . | X | X | X | X |
| **leaky count** | 8 | 8 | 5 | 9 | 2 | 2 | 2 | 1 |

Union = 12 cells across 4 shas x 2 OSes. Present at EVERY Windows sha: 4. Present at every Linux sha: 1.

#### Three readings, in increasing order of what they support

**1. The churn is total and it is the whole point.** Windows counts run 8, 8, 5, 9 — the HIGHEST
value is at `f3c8495b`, the tip carrying every fix F1-F16. Only **4 of 12** cells are present at
every Windows sha; only **1 of 2** at every Linux sha. Read as a quality signal the series says W2
made leaking worse, which is exactly as wrong as my retracted "8 -> 5, favourable" read and wrong
for the same reason. **A leaky count is not a signal in either direction.**

**2. NEW, and I had not said it before building this table: the two OS sets are DISJOINT.** Ten
cells leak only on Windows, two only on Linux, and **not one cell leaks on both**. Windows carries
8-9 leaky cells per run against Linux's 1-2. So this is not one flaky population sampled twice —
it is two populations with no overlap, and the Windows one is roughly five times larger. Any
explanation that treats "leaky tests" as a single phenomenon has to account for a clean partition
by OS.

**This is an OPEN QUESTION and the draft does not answer it. I do not have that explanation yet.**
Recorded deliberately without a candidate cause: the partition is a measurement, and the temptation
to pair a striking measurement with a plausible story is how the "8 -> 5 is favourable" reading got
into the record in the first place. Whoever takes this should arrive at a cause by reading, not by
inheriting one from me.

**SCOPED CAUSE, IN SOURCE, FOR ONE FAMILY ONLY (doyle ruled 10:22Z: source beats story; a source
cause covering one family is a measurement and is filed as one).** The brainproc family leaks on
Windows and not Linux because *its own fixture is asymmetric*: `long_child()` (brainproc.rs:1359)
spawns `Command::new("cmd").args(["/C","ping","-n","30","127.0.0.1"])` on Windows but
`Command::new("sleep").arg("30")` on Unix — a shell wrapper on one box, a bare process on the other.
The Windows candidate the supervisor kills is `cmd`; the process burning 29 seconds is its child.
That is two lines of source, not an inference from the symptom. **And it is NOT the whole partition:**
`process_cmdline_reads_a_live_arg_marker` uses a shell on BOTH arms (`sh -c "sleep 30; : marker"` /
`cmd /C "ping … & rem marker"`) yet leaks on **Linux only** — the shell mechanism predicts it should
leak on both boxes and it does not. So one family has a cause I can point at in the source, the
partition as a whole remains OPEN, and the counterexample stands in this paragraph rather than a
footnote so no one carries the cause further than it goes.

**THE PRODUCT PATH HAS NO INTERMEDIARY; THE FIXTURE DOES.** I chased the obvious escalation — if the
real supervisor also killed a shell wrapper, then "never two live brains"
(REQ-HAZARD-BROKER-PROCESS-ISOLATION, the invariant this very cell guards) could be violated on
Windows in production while the test stayed green. It cannot: `spawn_brain_child` (brainproc.rs:1224)
does `Command::new(exe)` — the brain is spawned DIRECTLY, no shell, no grandchild, so a single-pid
kill reaches it. **This is a test-fixture artifact, not a product defect**, and the fixture's shell
is a Windows-only convenience for "a process that blocks ~30s" that introduced a layer the real path
does not have. The negative is the useful half: it keeps the crashloop invariant out of IR-81 (a).

**Why the cell passed review for as long as it did.** It asserts `pid != 0 && !pid_alive(pid)` before
rollback, and that assertion is TRUE — `cmd` really is dead. The process still doing work is the one
it never looks at. The leak and the green are not in tension; they are the same fact from two ends.
**An assertion over a pid cannot see a tree.**

**3. The invariant holds at 12 of 12, and it survived the population being shuffled three times.**
Every union member spawns a child process and then kills it, lets it time out, or inspects its
corpse — `trial_kills_alive`, `kill_shell_at`, `kill_waker_at`, `bounded_run_kills_on_timeout`,
`extractor_timeout_errors`, `fetcher_timeout_errors`, `windows_session_is_zombie_..._corpse_as_dead`,
`legacy_psyche_sweep_guard_...`, `process_cmdline_reads_a_live_arg_marker`, and the three brainproc
spawn/promote/rollback cells. Membership churns freely across four shas and two OSes; the KIND of
cell has not varied once. **This is the IR-80 headline** and it is better supported now than when it
rested on a stable-looking five-cell cluster, precisely because the population underneath it has
been reshuffled and the invariant did not move.

#### A MECHANISM, read out of the source of the two Linux cells — calibrated, not generalized
`process_cmdline_reads_a_live_arg_marker` spawns **`sh -c "sleep 30; : marker"`**, and the test's
own comment says the trailing `; :` is deliberate — it keeps the shell RESIDENT so its
`/proc/<pid>/cmdline` still carries the marker instead of being tail-exec-replaced by `sleep`. The
teardown is `child.kill()` then `child.wait()`. **But `child` is the SHELL, and `sleep 30` is the
shell's own child.** Killing the shell orphans the sleeper, which outlives the test by up to 30
seconds — which is precisely what nextest reports as a LEAK. The Windows arm has the same shape:
`cmd /C "ping -n 30 127.0.0.1 >NUL & rem marker"`, where `cmd` forks `ping`.
So the very construct that makes the test work (a resident shell, so the cmdline stays readable) is
what guarantees a grandchild that a single-pid kill cannot reach.

**What is verified and what is not, stated separately because the difference decides who owns this.**
VERIFIED by reading source: this shape in BOTH Linux-only cells. CONSISTENT but not proven: the
Windows family kills the same way — `kill_waker_at_still_kills_a_matching_pair` tears down with
`kill_shell_pid(ours.id())`, a single-pid kill — so the mechanism is available to it, but I have not
read every spawn helper. UNVERIFIED: the brainproc, broker and spt-live cells; those spawn real
brains and time out extractors, and may leak for unrelated reasons.

**Why this matters beyond IR-80:** a single-pid kill that cannot reach a grandchild is exactly
[[ir81]] arm (a) — identity and tree-awareness pushed into `kill_pid_tree`/`kill_pid` so no caller
can express the unreachable kill. If the mechanism above generalizes, IR-79 (rigs leaking daemons
on a failing assert), IR-80 (this) and IR-81 (kill scoping) are not three findings that share a
seam, they are three FACES of one defect: **the codebase kills pids, and the things it needs dead
are trees.** That is a claim worth testing, not asserting — and doyle's one-cell mechanism probe
is still the right instrument, now with a sharper question: after the kill, is the surviving
process the CHILD or the GRANDCHILD?

---

#### AMENDMENT 1 — "a pid is not a tree" has THREE faces, and only one of them is a kill

*(hertz 2026-09-07, ruled by doyle 11:17Z: write this as an IR-80 amendment naming the three faces
rather than three separate entries, so a reader hunting any one face lands on all three.)*

The register sentence this lane has been running on is doyle's: **a pid is not an identity and a pid
is not a tree.** IR-80 and IR-81 were both filed against the KILL side, which made the sentence read
like a rule about killing. It is not. It is a rule about **which process you are talking about**, and
killing is only the face we happened to find first. Three faces are now measured, each found by a
different agent working a different problem:

| face | the operation | what going one level too high produces | measured |
|---|---|---|---|
| **KILL** | `kill_pid(pid)` on a wrapper | the wrapper dies, the grandchild keeps working — the leak | IR-80's 12-cell union; `proc.rs:430`'s own doc states it |
| **IDENTITY** | `kill_pid_tree(remembered_pid)` with no image read | you kill *a* process with that number, not the one you meant | IR-81 / #285, `broker.rs:8102` vs `servicehost.rs:651` |
| **ENV READ** | `psutil.Process(child.pid).environ()` | the wrapper's env certifies a launch the battery would have failed | hertz 2026-09-07, self-test of `.github/bench/launch-battery.py` |

**The third face, in full, because it is the newest and the least intuitive.** Reading back the
environment of the process you spawned is the standard remedy for "did my scrub/export actually
land" — it is what this project's own memory banked after the W0 perch-identity leak. It is
insufficient, and it fails in the flattering direction. A wrapper (`env VAR=x`, `cargo`, `nextest`,
`bash -lc`) modifies the environment it hands the process **below** it, while its own `environ()`
still shows exactly what you passed in. So the read-back agrees with your intent, prints PROVEN, and
says nothing whatsoever about the process under test. Measured: the first version of
`launch-battery.py` passed BOTH deliberately-negative tests. The corrected version walks descendants
and, in the refusal, prints the wrapper as CLEAN beside the grandchild as VIOLATOR — that contrast
is the face made visible in one line of output.

**Why this belongs in IR-80 rather than in a new entry.** All three are the same error committed at
the same seam: an operation is aimed at a pid when the thing it needs to reach is a tree (or a
specific member of one). The remedies rhyme — walk the tree, read the identity, report the scope you
actually covered — and a reader who finds any one face is one paragraph away from the other two.
Filing them separately would have hidden that, which is the concrete cost doyle's ruling avoids.

**What this amendment does NOT claim.** It does not widen IR-80's population: the 12-cell union
table and its Windows/Linux partition are unchanged, and the env face is not a test leak. It does
not add a remedy owner — the ENV face's fix is a tool (`.github/bench/launch-battery.py`), not product code, and the KILL and
IDENTITY faces keep the owners they already have (#285 / todlando post-W2). It is a naming, and its
whole value is that the next instance of this error gets recognized as an instance instead of being
filed fresh.

**Standing prediction, so this is falsifiable rather than tidy:** a fourth face exists wherever the
codebase asks a question *of* a pid that is really a question *about* a tree. The candidates I would
look at first are (a) liveness — `process_exists(pid)` answering "is the work still running" when
the work is a grandchild, and (b) resource attribution — reading a pid's handles/memory to decide
whether a lane is finished. Neither is measured; both are named so that finding one counts as
confirmation and finding none over the next few incidents counts against the generalization.

### IR-81 — a process kill can be written scoped or machine-wide, and nothing makes the scoped form the only reachable one: both codebases carry the right pattern beside the wrong one

- **Status:** OPEN, drafted by hertz 2026-09-07 08:38Z on doyle's dispatch (08:36Z), from the
  fleet-daemon death on hfenduleam at 08:03:15Z ([[RCA-FLEET-DAEMON-14444]], cause still UNNAMED).
  Census body: `docs/PID-KILL-CENSUS.md`. · **Origin:** doyle asked which sites kill a REMEMBERED
  pid without re-verifying identity AT KILL TIME. The census answered that, and then answered a
  better question nobody asked: **why the unguarded sites exist at all when the guarded ones sit
  feet away.**
- **⚠ THIS ENTRY IS NOT "FIX THESE TWO SITES."** Both individual fixes are small and are dispatched
  elsewhere (broker.rs:8102 → todlando's lane post-W2; live-relay-int.sh:78 → filed to perri). This
  entry is about the thing that PRODUCED them and will produce the next one: the scoped form and the
  machine-wide form are equally easy to write, equally plausible on review, and only one of them is
  correct on a shared box.
- **THE PAIRED EXHIBITS — same repo, same hand, feet apart.**

  | | correct, and it says why | incorrect |
  |---|---|---|
  | **spt-core** | `spt-daemon/src/servicehost.rs:651` — `provably_gone` first, then `exe_path(pid)` compared to the parked image via `same_image`, then a FRESH post-kill read. Its comment IS the rule: *"The pid is live. WHO is it? The image path decides — never the number."* | `spt-daemon/src/broker.rs:8102` — kills `spid`, a pid REMEMBERED off the session record, through `kill_pid_tree`. `zombie_verdict` consults liveness BY NUMBER (`process_exists`/`is_process_alive`), `adapter_labeled` (a property of the RECORD, not the live process), descendants and grace. **The live process's image is never read.** |
  | **spt-claude-code** | `ci/launcher/bind-int.sh:50` — `wmic process where "name='claude-spt.exe' and commandline like '%$ID%'"`, scoped to this run's unique id, comment *"never wall-a's"*. Also `multi-subnet-bringup-int.sh:117` ($C3_ID) and `wake-survival-int.sh:64` (`-match '$PROBE'`). | `ci/psyche/live-relay-int.sh:78` — `for p in $(tasklist \| grep -i claude-spt \| awk '{print $2}'); do taskkill //PID "$p" //T //F; done` — **every** `claude-spt.exe` on the box, tree-force, no run-scoping. Comment: *"by marker pid then name."* |

  **The knowledge is not missing in either codebase. It is present, written down, and adjacent.**
  Three correct sites to one wrong one in the adapter; a rule-stating comment in core. What is
  missing is any mechanism that makes the wrong spelling hard to reach.
- **Why "just fix the sites" is the wrong remedy:** it has already been tried implicitly — someone
  wrote the scoped form three times, which is what a team looks like when it knows the rule and
  still ships the exception. A fourth correct site does not prevent a fifth incorrect one. The
  remedy has to change what is REACHABLE, not what is written.
- **REMEDY — RULED 2026-09-07 08:40Z (doyle): (a) AND (b), not a choice between them. Two owners,
  two lanes, and they cover different halves of the hazard.**
  1. **(a) PRODUCT — push the identity requirement into the primitive.** `spt_store::proc::kill_pid_tree`
     / `kill_pid` take an expected identity (image path, or a `same_image` closure) and refuse
     without it, so a caller **cannot express** the unguarded kill; `servicehost.rs:651` stops being
     the exception among callers and becomes the shape of all of them. **Owner: todlando, post-W2.**
     ⇄ **This IS the fix shape of the `broker.rs:8102` board BUGFIX — that item and this entry
     cross-reference each other; neither is complete alone.** Reaches Rust callers only.
  2. **(b) GATE — `xtask check` refuses the unscoped spelling in CI scripts.** A `tasklist` /
     `Get-Process` enumeration piped into a kill, or a bare `taskkill` with no
     `commandline like` / `-match` scope in the same block, is a build failure. **Owner: hertz,
     rides the IR-37 thin PR after W2 lands.** This is the half remedy (a) can NEVER reach — the
     unguarded adapter site is a shell script, and no Rust signature constrains a shell script.
  3. **(c) Fix the two sites and stop.** Recorded as CONSIDERED AND LOSING, not omitted: it has
     already been tried implicitly — the adapter's author wrote the scoped form three times and
     shipped the exception anyway — so a fourth correct site does not prevent a fifth wrong one.
     It leaves the next instance free to appear.
  **The (a)/(b) split is the entry's real content:** the hazard lives in two languages, and a remedy
  in one of them is a half-measure that will read as a fix.
- **STANDING LIMIT, so this entry is not overread as an incident cause:** none of these sites
  explains the 08:03:15Z death. broker.rs:8102 would have been saved from firing by
  `has_live_descendants` (the daemon had a live brain) — accidental protection, not deliberate —
  and every claude-spt site enumerates `claude-spt.exe` while the dead process was `spt.exe`. The
  census found a real hazard while looking for a different one. **Do not let this entry close the
  incident.** The instrument at `C:\Users\decid\.spt-watch\daemon-watch.log` is what will answer
  that, or fail to.
- **Kin:** [[IR-79]] and [[IR-80]] (children outliving their cell; the leak family), the
  [[RCA-FLEET-DAEMON-14444]] timeline, `servicehost.rs`'s own "image path decides" comment as the
  in-tree statement of the rule, and the paired-exhibit method itself — a correct and an incorrect
  site in one repo is stronger evidence about PROCESS than either site is about code.
- **Ripe when:** the next touch of `spt-store/src/proc.rs` (for remedy 1) or of `xtask check`'s
  gate family (for remedy 2). The adapter half is perri's, filed separately, and does not wait
  on this.
- **Size:** small for remedy 2, medium for remedy 1 (signature + every caller). The census that
  justifies either is already written and does not need redoing.

---

#### ADAPTER-SIDE ARM: CLOSED (claude-spt, perri, 2026-09-07)

**Status: closed in the consumer repo, test-only, does not close IR-81 here.**

- **Commit `9c87372`** (spt-claude-code). Broad `tasklist | grep claude-spt | taskkill` replaced
  with an **id-scoped wmic** call, **name-pinned so wmic cannot self-match**; the three
  remembered-pid kills additionally got a **kill-time recheck**, covering the stale-identity face
  as well as the unscoped-pattern face.
- **Guard: `tests/ci-kill-scoping.sh`, traced as `REQ-HAZARD-CI-KILL-SCOPING`.** Gate green.
  This is the check-shape arm (b) proposes for spt-core, already landed once — precedent, not
  merely agreement.
- **Scope: test/CI infrastructure only.** No product code changed. (doyle, 08:50Z.)
- **`live-relay-int.sh:78` is retained as the historical RED exhibit** — the one genuinely
  machine-wide kill found in the whole fleet census. It is the reason this hazard is written as a
  hazard rather than a style note; do not quietly replace the exhibit with the fixed line.

**Population lesson carried into arm (b)'s census.** I filed ONE site; perri confirmed **four**, and
reports the idiom was already present in that repo **3x** before the filing. A filing that names one
line under-counts its class by 3-4x in a single repo. **Census the ACT — any taskkill / pkill /
wmic-delete reachable from a script root — never the literal string that led me to the first site.**
State the walked roots in the check's own message, and give the gate a negative control that goes
RED on a known-bad line before its green is trusted.

**Still open here, unchanged:** arm (a), Rust-side identity pushed into `kill_pid_tree` / `kill_pid`
so no Rust caller can express the unguarded kill (todlando, post-W2, cross-referenced with the
broker.rs:8102 board BUGFIX); and arm (b), the xtask script gate (mine, rides IR-37).

---

#### SCOPING ARM (a) — the product path has no intermediary; the fixture does
*(hertz 2026-09-07, ruled by doyle 10:22Z off the IR-80 leak work.)*

> **A pid is not an identity and a pid is not a tree. IR-80 needs CALLERS, not code.**
> *(The register's sentence — doyle 10:23Z. Arm (a) fixes the identity half in
> `kill_pid`/`kill_pid_tree`; the tree half needs no new machinery, because `kill_pid_tree` already
> exists at proc.rs:430 with the defect stated in its own doc. This is why IR-79, IR-80 and IR-81
> converge on one seam without collapsing into one finding.)*

Measured, not inferred: **`spawn_brain_child` (crates/spt-daemon/src/brainproc.rs:1224) does
`Command::new(exe)`** — the real brain is spawned DIRECTLY, with no shell and no launcher between
the daemon and the process it supervises. So on the crashloop/rollback path a single-pid kill
*does* reach the thing it aims at, and the "orphaned grandchild" mechanism found in the IR-80 leak
cells is **a property of the TEST FIXTURE**, whose `long_child()` (brainproc.rs:1359) goes through
`cmd /C` on Windows, not of the product.

**What that scopes.** Arm (a) is NOT about the brain-supervision path and must not be sold as
protecting `REQ-HAZARD-BROKER-PROCESS-ISOLATION` ("never two live brains") — that invariant is not
at risk here, and claiming it would be borrowing urgency this arm has not earned. Arm (a) stays
pointed at **the broker.rs:8102 reap, where the image is never read** — a kill aimed by pid alone
at a process whose identity was never verified. That is the real unguarded form, and it is a
different failure (kill the wrong process) from the one IR-80 surfaced (fail to kill the right
one).

**The two failures are different, which is why the drafts converge without collapsing.** Arm (a) is
*kill the wrong process* (identity never read); IR-80 is *fail to kill the right one* (tree never
walked). The identity half is todlando's, post-W2. The tree half is already solved in-tree —
`kill_pid_tree` at proc.rs:430, whose doc states the defect outright ("a single-pid `kill_pid` leaves
the wrapper's harness children orphaned + running") — so IR-80 ships callers, not machinery. See
[[RIDER 4]] for the two test-side callers.

### IR-82 — a daemon death leaves NO record of who died: autostart decides on a socket ping, says "no daemon", and the successor overwrites the only handle on the corpse

- **Status:** RULED 2026-09-07 11:07Z (doyle), drafted by hertz 11:04Z on doyle's dispatch (11:03Z: *"the autostart
  death-cause item is an IR, not a board item — mint it IR-82 in your thin PR"*). Register on main
  ends at IR-78; 79/80/81 are drafts in the same PR. · **Origin:** the fleet-daemon death on
  hfenduleam at 08:03:14.87Z ([[RCA-FLEET-DAEMON-14444]]) whose cause is **still UNNAMED after a day
  of measurement** — and the reason it is unnamed is this entry. Filed as a defect in
  OBSERVABILITY, not in the death.
- **⚠ THIS IS NOT "find the killer."** The killer hunt is the RCA and stays open there. This entry
  is about the fact that a killer hunt had to be run at all from a cold start, with a watcher script
  written by hand after the event, because the product recorded nothing about the transition at the
  moment it noticed it.

- **THE MEASUREMENT — all four facts read at source, no build (gate was live on both boxes).**

  1. **The daemon writes its pid.** `crates/spt-daemon/src/daemon.rs:471` —
     `let _ = std::fs::write(daemon_pid_path(), std::process::id().to_string());`, into
     `<spt_home>/daemon.pid` (`crates/spt-daemon/src/endpoint.rs:81`). Unconditional overwrite, no
     read-before-write, no rotation, no history.
  2. **Autostart never looks at it.** `ensure_running_outcome`
     (`crates/spt-daemon/src/daemon.rs:686-724`) decides on `is_running()`, which is
     `seedmap::ping(&seed_socket_name()).is_ok()` (`daemon.rs:598-600`) — **a socket ping and
     nothing else.** The decision consults no pid, no mtime, no prior state.
  3. **The breadcrumb it emits carries no identity.** `daemon.rs:721` —
     `DAEMON_AUTOSTART: no daemon and no standing operator stop — starting one`. Deliberately
     "always-on, ids-free" per its own comment, because its job is to be COUNTED (the convoy
     serialization observable that `daemon_stop_convoy_e2e` asserts at :178/:237). That design is
     correct for what it was built for and is exactly why it answers nothing here.
  4. **Then the successor destroys the evidence.** The new daemon reaches :471 and overwrites
     `daemon.pid` with its own pid. The dead pid, and the file mtime that bracketed its life, are
     gone — overwritten by the very event that should have recorded them.

- **THE SHARP FORM — the file is read ONLY in the state where it can tell you nothing.** The only
  two readers of `daemon_pid_path()` in the tree are `crates/spt/src/cli.rs:8581` and `:8754` (both inside `cmd_daemon_status`, :8573), and
  **both are gated on `running`** (`if running { let pid = read(...) }`). So `daemon.pid` is consulted
  exclusively when a daemon is alive — when its contents are guaranteed fresh and merely restate
  what the ping already proved — and is ignored in the one state where its staleness is the only
  evidence anyone has. Its doc comment (`endpoint.rs:82-85`) says *"a stale pid here is harmless
  because callers probe the socket, never this file."* **That is true for LIVENESS and false for
  FORENSICS**, and the entry is that the codebase has only ever considered the first reading.

- **WHAT IT COST, concretely, on 2026-09-07.** Daemon pid 14444 died at 08:03:14.87Z inside the ci
  Windows unit leg's nextest LIST phase; successor 48232 cold-started at 08:03:20Z and
  `DAEMON_RESTART_RESUME`d every session, orphaning live agents' monitors and readers and handing at
  least one agent a stale brief. Neither peer ran a kill. To get ANY signal at all I had to hand-write
  an external watcher (`C:\Users\decid\.spt-watch\watch-daemon.ps1`, 1s cadence, identity re-verified
  by image path AND creation time every poll) **after** the death, which by construction cannot
  observe the event it was written for. 6600+ polls later the successor is alive and the original
  death is still uncharacterized. A product-side record of the transition would have cost bytes.

- **WHY THIS IS NOT THE CONVOY BREADCRUMB'S JOB, and must not be bolted onto it.**
  `DAEMON_AUTOSTART` is load-bearing as a COUNTABLE line: `daemon_stop_convoy_e2e` asserts *exactly
  one* appears under a storm (:178) and counts N under a deliberate race (:237). Adding identity
  fields to that line risks the count and buries a forensic record inside a line whose contract is
  its cardinality. **The record belongs in its own line and its own artifact.**

- **REMEDY — RULED 2026-09-07 11:07Z (doyle), APPROVED as proposed with two refinements. Small,
  product-side, no wire change.** doyle re-derived the census independently at `f3c8495b` (one
  writer `daemon.rs:471`, two readers `cli.rs:8581` + `:8754`, both inside `if running`) before
  ruling — the falsifiable claim below has been checked, not taken.
  1. **Read before you clobber.** At `daemon.rs:471`, read the existing `daemon.pid` and its mtime
     BEFORE the write; if it holds a pid that is not ours, emit a distinct, greppable line naming
     the predecessor pid and how long ago the file was last touched — e.g.
     `DAEMON_SUCCEEDS_PID: prior=<pid> prior_pidfile_age_ms=<n> reason=<socket-ping-failed>`. The
     value is that the successor is the only process in the system that is guaranteed to run at
     exactly the moment the question becomes askable.
  2. **Keep one generation.** Rename the existing `daemon.pid` to `daemon.pid.prev` instead of
     overwriting it (best-effort, same directory, same failure posture as today's `let _ =`). One
     generation is enough to answer "which pid did I replace, and when did it last write?" and costs
     no rotation policy.
  3. **REFINED BY THE RULING — the discriminator is a FIELD, not a second line.** I proposed a
     separate "what the ping saw" record; doyle collapsed it: `prior=<pid|none>` already separates
     the two incidents, because `none` **is** the fresh-home case. One line, one field, no second
     breadcrumb to keep in sync.
  4. **ADDED BY THE RULING — `alive=<bool>`, from a by-number liveness READ at write time.** With
     the prior pid in hand, read whether that number is still live and put the answer on the line.
     It splits the two incidents that otherwise look identical: **`alive` prior = two daemons racing
     the seed bind** (the convoy case), **dead prior = the corpse** (the 14444 case). This is a
     READ, not a kill — #285's identity-before-kill rule binds kills and is not weakened here; a
     by-number read that informs a log line commits no action on the pid.
  5. **UNIT, specified by the ruling:** seed a pidfile with a fake pid and an aged mtime, run the
     write path, assert the emitted fields (`prior`, `age_ms`, `alive`) AND the `daemon.pid.prev`
     bytes. The `.prev` bytes are part of the assertion, not an afterthought — the generation kept
     is the half that survives the process.
  6. **NOT proposed, and the ruling kept it out:** a liveness change. `is_running()` must stay the socket ping — this entry does
     not argue the pidfile should become a liveness signal, which is the exact mistake
     `endpoint.rs:82-85` was written to prevent. The pidfile stays advisory; it just stops being
     silently destroyed.

- **LANE — SPLIT BY THE RULING, and the two halves ship separately.** The **IR-82 ENTRY** (this
  text) rides hertz's IR-37 thin PR alongside IR-79/80/81. The **PRODUCT REMEDY** rides **#285's
  lane** (todlando, post-W2) because it is the same pid-identity gate: #285 makes a kill read an
  identity, this makes a successor read the identity it is about to erase. One lane, one owner,
  one review of the pid-identity question. Do not implement the remedy in the thin PR.

- **Kin:** [[RCA-FLEET-DAEMON-14444]] (the open incident this serves — IR-82 does not close it and
  must not be read as closing it), [[IR-81]] (a pid is not an identity — the same lesson from the
  KILL side; here it is the RECORD side), [[IR-80]]/[[IR-79]] (processes outliving what should have
  accounted for them). The through-line doyle named at 10:22Z holds: *a pid is not an identity and a
  pid is not a tree* — and IR-82 adds that **a pid nobody wrote down is not evidence.**

- **Falsifiable claim, so it can be checked rather than believed:** at `f3c8495b` the tree contains
  exactly ONE writer of `daemon_pid_path()` (`daemon.rs:471`) and exactly TWO readers
  (`cli.rs:8581`, `cli.rs:8754`), both inside a `running` gate. If a future reader appears in an
  unguarded path, this entry's premise weakens and it should be re-measured, not re-asserted.

### IR-83 — a two-host rig addressed over the tailnet cannot receive INBOUND on Windows, and the whole rig population had only ever run the admitted direction

- **Status:** OPEN, minted by hertz 2026-09-07 11:41Z on doyle's dispatch (11:40Z). **Provenance: the
  mechanism below is DOYLE'S measurement, not mine** — eight probe arms, each with sender, listener
  and output file, recorded in `GATE-W2-272-CHECKLIST.md` rows 11:36Z–11:39Z. I am off cargo on both
  boxes and have re-measured none of it; this entry is the register form of his finding plus the
  structural reading it supports. Board face: **F18**. Register slot: IR-83 (79/80/81/82 are the
  same thin PR).
- **THE FINDING, in one line:** two-host rigs must address peers by **LAN**, never the tailnet.

- **HALF ONE — the tailnet ACL is per-node and ONE-WAY.** `hfenduleam` admits 9 sources and
  `kitsubito` is **not** among them; `kitsubito` admits `hfenduleam`. So a rig cell needing INBOUND
  to the Windows box over `100.x` fails **at the receiver**, and it fails there **whatever the host
  firewall says** — which is why a host-firewall investigation can be run to completion, come back
  clean, and explain nothing.

- **HALF TWO — Windows defaults `BlockInbound`, and a hash-named test exe has no rule.** Test
  binaries are built with hash-suffixed names, so a rule written against the *program* is keyed to a
  name that changes on the next build. The stale `spt_net-d037…` rule is the precedent already in
  the tree: a rule that was correct once, is present, greps fine, and admits nothing. **The durable
  form is PORT-scoped, not program-scoped:** ports `7480-7499`, `remoteip` = the peer's LAN address,
  added by the operator.

- **WHY IT SURFACED ONLY NOW, and this is the part that generalizes.** The W2 helper cell is the
  **first cell in the entire rig population that ever needed `kitsubito -> hfenduleam` inbound.**
  W1's four xbox cells and golden's old-twohost job are all `hfenduleam -> kitsubito` — the admitted
  direction. So the population had a **direction monoculture**, and a one-way ACL is invisible to a
  suite that only ever runs the admitted way. Every green in that population was consistent with the
  ACL being wide open and equally consistent with it being one-way; the suite could not tell the
  difference, and nobody had reason to ask. **A capability exercised in only one direction is not
  covered, it is merely unexamined** — the same shape as an assertion that only ever sees the
  passing arm.

- **THE SECOND STRUCTURAL POINT — a rule keyed to a MUTABLE identity is the lane's own through-line.**
  A per-hash program rule is an exception bound to an identity the build changes underneath it,
  exactly as [[IR-81]] is about a kill bound to a pid the OS may reassign and IR-80 Amendment 1's ENV
  READ face is about a read bound to a process identity `exec` can change. Port-scoping replaces a
  mutable identity with a stable property. That is the same remedy shape as reading an image path
  before a kill and reporting an environment in-band: **stop keying on the thing that moves.** I am
  flagging the rhyme, not claiming a fourth face — this is infrastructure config, not a code path,
  and the amendment's standing prediction is about pid-shaped questions.

- **FALSIFIABLE CLAIM (doyle's, recorded verbatim as the entry's test):** with the rig on **LAN
  addresses AND the port-scoped rule present**, the helper cell's owner raw carries `WEB_SERVE_FOR`;
  with **either** missing, it carries none. Both arms matter — this is a conjunction, so a single
  green with one of the two absent would refute it, and a red with both present would too.

- **WHAT THIS ENTRY DOES NOT CLAIM.** It does not say the host firewall was ever misconfigured; the
  ACL half fails at the receiver independently. It does not close F17 or explain the vacuous-green
  run (that was the mid-flight script rewrite — different cause, different fix, and the two were
  live in the same hour, which is precisely why each needed its own named mechanism rather than one
  story). It proposes no product change: both halves are rig addressing and operator-added firewall
  config.

- **REMEDY.** (1) Rig-side: address peers by LAN in two-host rigs, and make the tailnet address
  unreachable-by-construction rather than merely unused, so a future rig cannot quietly pick it up.
  (2) Operator-side: the port-scoped inbound rule (`7480-7499`, `remoteip` = peer LAN), replacing
  per-hash program rules; the stale `spt_net-d037…` rule should be removed in the same pass so it
  stops reading as coverage. (3) Population-side — **RULED 11:42Z (doyle), with a precision that
  changes the shape:** this is NOT a new cell. **The W2 helper cell IS the first reverse-direction
  cell**, and on LAN it exercises `kitsubito -> hfenduleam` inbound for real once the rule exists.
  The remedy is therefore to **keep at least one reverse cell in the xbox battery permanently and
  NAME it as the direction witness**, so its role is a stated property rather than an accident of
  what W2 happened to need. A suite that only runs one way cannot discover the next one-way ACL
  either, and the way that recurs is not "nobody wrote the cell" but "the cell that had it stopped
  being run."

  **The hazard that follows from naming it, stated here so the naming is not decorative:** a witness
  cell that is SKIPPED stops being a witness while still reporting green — the exact confusion this
  lane already paid for, where a must-skip `role_b` PASS was indistinguishable from four cells that
  no-opped. So the direction witness needs its skip to be loud: if it is filtered out, phase-reclassed,
  or short-circuits on a missing precondition, that must read as ABSENT COVERAGE and not as a pass.
  Whoever names it owns that half too. **Concrete enforcement, doyle-ruled 11:44Z and written up
  as the RIDER 5 ADDENDUM:** the xbox driver's HELPER WITNESS verdict line already refuses a
  `0.0x s` helper PASS and a `b.raw` without `WEB_SERVE_FOR outcome=registered`; it must ALSO
  print `NOT-A-WITNESS` when the cell is absent from the LIST leg. Those first two are
  *ran-but-vacuous*; absence is *never-ran*, and a verdict built only from the first two passes
  it by omission — every refusal condition evaluates against output that does not exist, finds
  nothing to object to, and falls through to green.

- **GATING CLAUSE — added 2026-09-07 12:56Z on doyle's dispatch (gate reshape, same minute).** An
  inbound-to-Windows cell in a two-host rig is **gated on the port-scoped rule being present**, and
  without it the cell must **SKIP LOUDLY**: it may not run, may not report PASS, and may not be
  silently filtered — its absence must render as ABSENT COVERAGE in the verdict, on the RIDER 5
  ADDENDUM's `NOT-A-WITNESS:NEVER-RAN` face. The precondition is machine-readable, so the driver
  reads `netsh` at start and the gate is a measured fact rather than a remembered one; a cell whose
  precondition is unmet and which therefore emits nothing is exactly the case a verdict assembled
  only from ran-but-vacuous refusals passes by omission. **This clause is what stops the rule's
  ABSENCE from reading as the rig's health.**

  **How the reshaped W2 F17 gate applies it (doyle 12:56Z, recorded, not re-measured by me — I am
  still off cargo on both boxes).** With the rule absent 75+ minutes and no operator activity, and
  golden carrying no cross-box `twohost_web`, the shortfall is a **rig blocker, not a golden one**.
  So the HELPER WITNESS was reshaped to the **ONE-BOX pair on BOTH OSes** — doyle's gate tree on
  Windows, a kitsubito clone at the tip on Linux with the rig shipped as the frozen copy — on the
  basis that the 03:31Z one-box run at `f3c8495b` reproduced F17 end to end (A `outcome=unanswered`,
  B `DISPATCH:4:Unknown`), i.e. the one-box pair drives the real dispatcher on both roles.
  Mutation D must turn the classify unit cell red **and** the one-box helper cell red
  (owner `family=Unknown`>0, `registered`=0) while deny/fetch stay green; revert dirty 0. The
  cross-box pair still runs as the A->B regression (deny/fetch must PASS), and while the rule is
  absent its helper/range TIMEOUTs are classified **F18-INFRA** — becoming a witness automatically
  if the rule lands before the run, because the driver reads `netsh` at start.
  **The cross-box helper witness therefore stays an OPEN rider on the operator's rule, not a land
  blocker.** Note the shape this preserves: the reshape moves the witness to a pair that CAN run,
  and leaves the direction it cannot exercise named and open — it does not let the one-box green
  stand in for the direction coverage that is still missing. A one-box pair proves the dispatcher;
  it cannot prove an inbound ACL, and nothing here claims it does.

  **PROBE RESOLUTION — doyle 12:59Z, ADOPTED INTO THIS GATE (not deferred to the rider).** The
  precondition read is now THREE-VALUE in `gate-w2-f17.sh`: `Rule Name` in stdout -> PRESENT,
  `No rules match` -> ABSENT, empty stdout or anything else -> PROBE-FAILED. **PROBE-FAILED
  classifies the cross-box helper/range as `UNCLASSIFIED: this run cannot tell INFRA from a real
  red` — never F18-INFRA** — and xbox-D's skip line names the state. The raw probe (exit code +
  first 3 lines) is printed into the driver log, so the classification is auditable after the fact
  rather than inferred from what got bucketed. Exercised standalone on all three arms: real
  `netsh` = exit 1 + `No rules match` -> ABSENT; forced bad args -> PROBE-FAILED; empty stdout ->
  PROBE-FAILED.

  **Recorded because it corrects the asker (me), not the implementer:** I proposed the three-value
  shape AND, in the same message, proposed discriminating on the exit code. doyle's standalone
  exercise measured `netsh` exiting **1 on a clean ABSENT**, so an exit-code discriminator would
  have classified every legitimately-absent rule as an instrument fault — turning the run's most
  common honest state into PROBE-FAILED and, with it, `UNCLASSIFIED`. The discriminator that landed
  is STDOUT CONTENT. The general form belongs in the entry because it is the reason the third arm
  needed measuring at all: **a CLI's exit code is a claim about the QUERY, not about the ANSWER**,
  and an empty-but-correct result is reported non-zero by a great many tools. A third value added
  to a probe needs its own evidence, obtained by making all three arms fire on purpose — which is
  this lane's standing rule that a tool whose success path has never been made to fail on purpose
  is untested, applied to a tool's FAILURE path instead.

- **Kin:** [[IR-81]] and IR-80 Amendment 1 (keying on a mutable identity), the F17 riders (a witness
  that never executed the direction it certifies).

### IR-84 — the one `PUMP_PEER_FAIL` arm the rigs actually fire was the only one carrying no clock, so the stall that produced it could be counted but never timed

- **Status:** OPEN, filed by hertz 2026-09-09 with the instrument that closes its first half. ·
  **Origin:** the 65 s two-host helper stall (design written before the run,
  `docs/design/RIDER-65S-DESIGN.md`). The number was branch-claimed 2026-09-08 by name only — IR-85
  records that it carried no entry text in any register file in any worktree — and this is that text.
- **Symptom:** `two_host_web_helper_role_a` takes **61.956–69.096 s** across five runs while its
  siblings in the same binary and the same run take **0.135–0.753 s**. Every ladder rig —
  win-onebox, linux-onebox, cross-box — logs **exactly ONE**
  `PUMP_PEER_FAIL:<node>:peer reply-read: no progress within budget — dropping peer (brain IPC read
  deadline elapsed)` on B before completing. **3/3, one each, never on A** (doyle's GATE-W3 readout).
- **Cause of the DIAGNOSTIC failure, which is what this entry is about:** that line is emitted from
  the `Err(e)` arm of `peer_leg_outcome` (`crates/spt-daemon/src/pump/mod.rs`), and at
  `b0b67aaa` it read `emit_line_err!("PUMP_PEER_FAIL:{peer_hex}:{e}")` — **no timestamp**. Measured
  at that sha rather than recalled: the file carries **four** `PUMP_PEER_FAIL` emit sites, **three**
  of them stamped `wall_ms=… mono_ms=…` (the submit arm, the no-route arm, and the
  `PRESENCE_DIAL_FAILED` arm), and this one was **the only unstamped site left**. It is also the
  arm the rigs actually fire — a `PEER_REPLY_READ_BUDGET` expiry reclassified out of `TimedOut` by
  `brain::reclassify_peer_reply_err` so it drops one peer rather than poisoning the round. **The
  single arm a reader most needs to time was the single arm that could not be timed.**
- **What that cost, stated as the open question it left:** the arithmetic does not close and this
  entry does not close it. `PUMP_PEER_IO_TIMEOUT` is **30 s** (`:118`) and `SUPERVISE_BACKOFF_BASE`
  is **5 s** (`:121`) — read as constants, not recalled — so one deadline plus one supervise floor
  accounts for **35 s** against **61.956–69.096 s** observed. **~27–34 s is unaccounted**, itself
  close to another 30, which is exactly the coincidence the design document warns against; and the
  count of **exactly one** fail forbids the easy reading that two deadlines elapsed. Three candidates
  stay open and are NOT ranked here: a first ~30 s burned before the deadline's clock starts; a
  second budget stacking on the first; or the restart's re-prime costing materially more than its
  5 s floor. **Separating them needs the deadline's own timestamps against the fail line** — which
  is precisely what the log did not carry.
- **⚠ WHAT THIS ENTRY DOES NOT CLAIM.** The instrument does not explain the stall, does not shorten
  it, and changes no product behaviour: the same peer is dropped at the same instant with the same
  backoff. It makes the stall **answerable from a log that already contains the failure**, instead of
  requiring a rebuild to learn when the budget started. `PUMP_PEER_IO_TIMEOUT` is a hard const with
  no env knob, so varying it costs a rebuild — deliberately deferred until the stamps say which half
  of the arithmetic is missing.
- **Kin and the recurrence that matters:** this is the **SECOND** time an unstamped
  `PUMP_PEER_FAIL` cost a hunt. The `PRESENCE_DIAL_FAILED` arm carries a comment naming the
  "2026-07-14 `PUMP_PEER_FAIL`-unstamped seed, folded here", and **REQ-PUMP-STAGE-TRUTH** already
  requires every peer failure "stamped (wall+mono) and peer-attributed", explicitly subsuming that
  seed. So the contract was written, three of four sites complied, and the fourth stayed silent for
  eight weeks — **a convention enforced by no mechanism degrades one arm at a time, and every reader
  of a complying arm reports the convention as held**, the same shape as the `validate_docs_dir`
  refusal-arm rider in the same design document. Also kin: **IR-83** (the rig population had only
  ever run the admitted direction), `REQ-HAZARD-PUMP-IPC-DEADLINE` (a `TimedOut` brain-IPC read is a
  supervised restart, never a per-peer retry), ADR-0039 Decision 5.
- **Ripe when:** now, and it lands with this entry — the stamp is one emit line, tagged
  `[impl->REQ-PUMP-STAGE-TRUTH]` at the site. · **Size:** one emit line plus its comment.
  **Follow-up, not this lane:** a cell that walks every `PUMP_PEER_FAIL` arm and asserts each one
  stamps, so the fifth site added cannot be born silent.
### IR-85 — the Windows self-hosted box runs its fs-heavy tests 3-4x slower than a week ago; two CI wall clocks were sized for the old box, and a red run kept hiding it

- **Status:** OPEN, filed by hertz 2026-09-09 at the #272/v0.68.0 golden r3/r4 arc, on doyle's
  dispatch. Number ruled by doyle from his own census of main@`a2f335f8` (register ended at IR-83;
  IR-84 was branch-claimed by name only, with no entry text in any register file in any worktree —
  that gap is now closed: **IR-84 was filed 2026-09-09 in PR #213**, so this sentence records the
  state at `a2f335f8` and is no longer a reason to go looking for a missing entry).
  **Absorbs two drafts:** doyle's `IR-DRAFT-windows-fs-heavy-slowdown-and-golden-wall` (his IR-A +
  IR-B) and his earlier `IR-NEXT` (operator-desktop load), both retired by reference — this is the
  single entry. Every job/step and per-test timing below is **doyle's**, read via the job/step API.
  · **Origin:** #272 golden r3 attempt 2 was CANCELLED by a wall clock with both test phases green.
- **⚠ THE CAP CHANGE IN `a2f335f8` IS NOT THE FIX AND MUST NOT BE READ AS ONE.** It bounds the run;
  it repairs nothing. A future reader who finds an 80-minute wall and no entry here would reasonably
  conclude the problem was solved. It was only made visible.

- **A GREEN RUN COSTS MORE THAN A RED ONE, WHICH IS WHY THIS WENT UNSEEN.** A failing run
  short-circuits past the wall a green run has to cross. r3 att1 finished in **48m39s only because
  it FAILED at Phase B**; att2, green, hit 49m59s and was cancelled. Every earlier docs-drift skip
  therefore presented as "upstream failure" — the wall was never the reported cause of anything, and
  the surviving evidence systematically flattered the budget.

- **THE WALL.** Windows golden `test` job (48 steps), fully GREEN: **30m53s** (2026-08-30, run
  33296634901) and **33m21s** (2026-09-06, run 34017906638). At the v0.68.0 head against the
  then-current 50-minute cap — **r3 att2**, job 102352551368 at `f6110c2a`: Phase A 17m14s, Phase B
  20m25s, doctests 1m04s, clippy 3m31s, **CANCELLED at step 30 at 49m59s with everything green**;
  steps 31-42 (installer, docs floor, both docs-drift gates, dormancy) never ran. r2 att1
  (`25e60015`, 09-08 18:47Z) reached step 34 at 44m30s. Steps 30..39 cost ~3m30s on the 09-06 green
  (notify 47s, installer 13s, docs-drift 1m54s). Predicted green need ~**56 min**; cap raised to
  **80** (need + ~40% for day-to-day variance) in `a2f335f8`, ci unit 25 -> 40 in the same commit.
- **r4, GREEN — the first complete measurement at the head: 54m35s** (07:12:26Z -> 08:07:01Z),
  **25m25s** headroom under 80; docs-drift (step 38) 2m23s; Phase A 15m51s, Phase B 22m56s. The ~56
  prediction held to within 1.5 min, so the cap is sized on evidence, not generosity. **What that
  does NOT establish:** it was a sizing forecast, and its holding says nothing about the cause
  diagnosis below. Nor are att2-vs-r4 per-phase deltas a trend — att2 was cancelled mid-run, so only
  its phase legs are comparable at all, and the spread they sit in is the same variance the 40% is
  there to absorb.

- **THE SLOWDOWN, per test, same box, same tests.** `spt-daemon::sync` concurrent_writes
  **22.4s -> 74.8s**, two_tier_sync 18.7 -> 67.4; `spt-store` monic clone_copies 17.7 -> 62.9,
  different_monics 18.2 -> 64.9; syncmerge reconciled_write 27.3 -> 51.2. **~25 spt-store/spt-daemon
  tests now exceed 30s on Windows Phase A at `f6110c2a`, against 1-2s each on kitsubito.**
  Phase-level: Phase A nextest 170s (09-06) -> 642s (att1) -> 1034s (att2); Phase B 710s -> 1164s.
  The suite did not get more expensive — Linux is unchanged.

- **DATING SAYS ENVIRONMENT DOMINATES, AND BY HOW MUCH.** main's own ci Windows `unit` job:
  11-12 min on 09-06 (runs 34040416870, 34041526195, 34043323577) -> 13-20 min on 09-07 -> 12-24 min
  on 09-08, hitting **22 min at `e4444413`** (run 34261096301) three minutes under its own 25-minute
  wall. **Gradual over days, under no single gate change, is the shape of an environment term, not a
  commit's.** Head growth is real but MINOR: +138 Phase A tests, +20 Phase B, HEAVY 34 -> 35.

- **THE BOX IS AN OPERATOR DESKTOP (folded in from the retired IR-NEXT), and that turns every fixed
  wall-clock budget in the Windows suite into a coin.** From the r2 arc, three attempts at ONE sha
  (run 34262154550 @ `25e60015`): Phase B red each time on a **different single cell or none** — a1
  234/234 (job red on disk floors only), a2 `spt::webserve_attachment_e2e` arm 12, a3
  `spt-daemon::mesh_recovery roster_route_survives_a_transient_dial_failure_with_discovery_disabled`
  (15.0s `converge()` budget, cell 15.715s; the same cell 9.8s / 7.2s on a1/a2). Phase A — 3,346
  spawn-dominated unit cells — slowed **monotonically 448.7 -> 495.1 -> 542.6s at that one sha**.
  Per-cell a3/a2 over 73 Phase B cells >= 1s: median 1.05x, mean 1.34x, 19 cells >= 1.5x, worst 5.2x
  (`endpoint_lifecycle poll_vs_reap` 1.1 -> 5.8s) — **BURSTY, not uniform**. A rotating single victim
  across attempts at one sha is ONE environment cause; hardening victims one at a time never closes
  it (paid before: `e2e-leaked-daemons-shared-box`). The budgets that lost were ~1.5x the fast
  observation — inside the box's measured variance, so they were coins that had been landing right.
- **Box census (doyle 01:13-01:19Z, no cargo/rustc/nextest running, CPU 31%, ~1.1 of 16 cores busy):**
  `qbittorrent.exe` seeding since 09-08 10:20Z (box tx 21.3 MB/s over 5s); fleet `spt daemon brain`
  with 516 GB read since 09-07 08:03Z (~3.5 MB/s steady); Defender real-time ON, `MsMpEng` at
  67.8/62.8/48.9/32.2/12.2% of a core over 5s **on the idle box**; a fresh 35 MB exe pays
  2092/994/1171/1043 ms on FIRST execution against 31/263/19/260 ms on the second — **and every CI
  attempt rebuilds every test binary fresh.** At 06:45Z on 09-09, with the twohost legs running:
  MsMpEng 89% CPU / 991 MB WS, qbittorrent pid 47056 holding 7817 CPU-seconds (2.2 h) since
  09-08 03:20, free 134 GiB (275 -> 197 -> 131 across the three r3 dispatches).
- **⚠ AN UNREADABLE ROW IS NOT AN ABSENT ONE.** The Defender exclusion list cannot be read
  unelevated on this box: `Get-MpPreference` returns the literal string
  `N/A: Must be an administrator to view exclusions` **as the ExclusionPath VALUE**, so a
  `-contains` test reads ABSENT and is meaningless; the HKLM `Windows Defender\Exclusions\Paths`
  read throws `SecurityException`. Nobody may report the runner directory as unexcluded from an
  unelevated shell.

- **HYPOTHESIS ALREADY KILLED, so nobody re-runs it:** the Windows-only `ADAPTER_WEB_PENDING`
  reconcile failure (servehost nudge) **cannot** explain this — spt-store monic and contextstore
  never touch that path.

- **WHAT IS STILL NOT ESTABLISHED (labelled, so it is not inherited as fact).** The dating argument
  establishes environment-DOMINANT and bounds head growth as the minor term. It does **not**
  apportion the environment term itself: Defender vs the third-party torrent load vs the fleet
  brain's steady read vs anything else is unsplit, and **"MsMpEng at 89%" remains a correlate
  measured beside the slowdown, not a proven cause.** The discriminator lane below is what settles
  head-vs-environment on evidence rather than on the shape of a drift curve.

- **REMEDY — none landed. This entry is the debt, and its middle arms need an operator.** Sequence
  matters; run them in this order:
  1. **DISCRIMINATOR LANE (hertz, one box, ~1.5-2 h).** Run the five named tests at `04e32c8c` and
     at `f6110c2a` on hfenduleam. **Same-slow at both = environment; slow only at the head = head
     growth.** Cheap and decidable, and it must precede any operator ask — do not spend an elevation
     request on a hypothesis a one-lane measurement can test. Design note, because it is the part
     that makes the number trustworthy: the arms run **INTERLEAVED** (A/B/A/B/A/B, 3 reps each), not
     all-A-then-all-B, so drift that hits the whole box cancels in the comparison instead of landing
     on one arm. **A measurement that only works if everyone behaves is not a measurement** — the
     torrent client and Defender are running throughout and are the thing under test, not noise
     anyone gets to remove. Confounder already excluded: the three test-bearing files
     (`spt-daemon/tests/sync.rs`, `spt-store/src/monic.rs`, `spt-store/src/syncmerge.rs`) are
     **byte-identical blobs at both shas**, so a difference cannot be the tests changing; the two
     crates around them are not (+11,388 lines over 48 files).
  2. **Remove the third-party load first.** No torrent client on the CI box during golden windows.
     It is the cheapest variable to remove, and removing it first makes arm 3's benefit measurable
     instead of confounded.
  3. **OPERATOR ASK — Defender path exclusions** for `C:\actions-runner\_work` and the gate pools,
     read from an ADMIN shell first (see the unreadable-row warning above), then re-measure the
     first-touch tax on a file **under that path** — the scratchpad measurement is outside the
     runner dir, so it proves the mechanism's size, not the runner's exposure. Unsettable
     unelevated; same shape as **IR-89**'s elevated firewall rule.
  4. **CI census, cheap and durable:** print `Get-MpComputerStatus` RealTimeProtectionEnabled plus
     the ExclusionPath read **verbatim, refusal text included**, in the Windows job's census step,
     so a run's own record says what Defender it ran under.
  5. **Re-measure both caps after any arm lands — the arm a future reader will skip, so it is loud
     here.** A cap sized against a degraded box is correct only while the box stays degraded.
     Leaving 80 in place after a repair silently restores the original hazard: a wall so generous it
     no longer catches a wedged test, which is the job the golden cap was added to do in the first
     place (the 2026-06-03 handoff.rs ConPTY stall, 22 unbounded hosted minutes).
- **Kin:** [[IR-76]] (the golden runner IS the builders' box — the structural reason a desktop's
  load reaches CI at all), [[IR-64]] (box-level facts only the operator can move), [[IR-89]] (the
  other elevated box-rule ask), `defender-first-touch-tax-on-fresh-test-binaries` and
  `e2e-leaked-daemons-shared-box` (memory).
- **Ripe when:** arm 1 now; arms 2-4 on the operator's answer; arm 5 at the next `golden.yml` touch
  after any of them. · **Size:** arm 1 a measurement, arm 4 one census line, arm 5 two literals.
- **AMENDMENT 2026-09-09 (two terms this entry did not name when it landed at `d32d5c4c`).**
  1. **`ci.yml`'s `changes` job runs `unit` on BOTH runners for every push to `main`.** The classify
     step at `.github/workflows/ci.yml:49` emits `code=true` for every non-`pull_request` event, so
     the docs-only skip that PR runs #206/#207 demonstrated is **`pull_request`-only** — read from
     `ci.yml` at `main` by doyle, who names it his own error after twice ruling the opposite from the
     PR runs alone. Consequence for this entry's wall clock: run **34337797758** (the ff of
     `b66a9612`, a docs-only delta over golden-green `a2f335f8`) ran `unit (Windows)` 09:59:57Z →
     10:40:48Z and was **CANCELLED at the 40-minute job wall** (job timeout; conclusion `cancelled`,
     step 6 `Unit tests` 10:08:23 → 10:40:03) — a red on `main` at a sha whose content cannot fail a
     unit test. It is an IR-85 face, not a flake row.
  2. **A near-full volume is an environment term for the slowdown**, alongside the Defender
     first-touch tax and the operator-desktop load already recorded. **Deliberately unquantified:**
     the discriminator window that would have apportioned it was VOID, because the one leg that
     completed ran under both a CI job and a falling disk. See **IR-90**.
- **WHAT THE VOIDED DISCRIMINATOR ESTABLISHED, negatively (hertz, 2026-09-09).** The fs-heavy
  slowdown REPRODUCES at `04e32c8c`, which predates the +11,388-line head growth: spt-daemon
  `concurrent_writes` 105.420 s, `two_tier_sync` 58.362 s, spt-store `clone_copies` 65.657 s,
  `different_monics` 73.967 s, `syncmerge reconciled_write` 86.765 s, against the 09-06 baseline of
  22.4 / 18.7 / 17.7 / 18.2 / 27.3 s — **3.2x to 4.7x on all five, at the OLD sha**. So head growth
  is **not NECESSARY** for the slowdown. That is the whole of it: it does NOT measure how much the
  environment explains, and it is not evidence about the head arm at all — that leg ran inside run
  34337797758's window and on a volume that reached 0.018 GiB free, and the head arm never produced
  a comparable pair. **Arm 1 stays OPEN**, its re-run deferred until **IR-90**'s guard exists and a
  window with no `main` push and no CI job on the box can be scheduled.

### IR-86 — golden's 32 GiB floor is BELOW the measured 67.4 GiB Windows suite footprint, so a start floor passes a box that cannot finish

- **Status:** OPEN, filed by doyle 2026-09-08 at the #272/v0.68.0 golden r2 triage. · **Origin:**
  run 34262154550 @ `25e60015` (shaped on `e4444413`), Windows test job 102182665033: every product
  step green (Phase A 3346/3346, Phase B 234/234, Summary 2, FAIL 0), the ONLY reds were the two
  in-job floor gates. `FLOOR_START` 97,649,786,880 B PASS at 18:47:23Z -> `FLOOR_DOCS`
  25,284,501,504 B RED at 19:31:55Z -> `FLOOR_END` RED. Consumed in-job: 67.4 GiB.
- **Mechanism:** `golden.yml` derives 32 GiB as "the observed tens-of-GB full-suite footprint
  rounded up to the next binary boundary, preserving headroom for one complete run" (comment
  above the Windows start floor). The measured footprint is 67.4 GiB, so a start reading anywhere
  in [32, ~100) GiB passes and the job then walks under the floor by construction: box baseline
  91 GiB minus 67.4 = 23.5 < 32. This is a shortfall of the DERIVATION, not a rate, not a product
  red, and not the instant-vs-sustained face IR-46 files (that face is fixed by the FLOOR_DOCS
  re-read, which is what caught it). The start floor must be footprint + end-floor (~100 GiB) or
  it asserts nothing about finishing.
- **Coverage consequence (the load-bearing part):** the Windows `Docs drift gate (CLI ref + llms
  links)` step is sequenced BEHIND `FLOOR_DOCS`, so it SKIPPED and r2 held NO Windows CLI-ref
  axis at that sha until the rerun. A floor red that skips a gate is a coverage hole wearing a
  disk red.
- **Kin:** [[IR-46]] (a floor asserts an instant), [[IR-59]] (pool arithmetic, LNK-class reds
  that name no disk), [[IR-64]] (the operator-payload reservoir that sets the box baseline; its
  ripe-when arm — a golden that dies at the floor despite adopted pool discipline — FIRED here
  in the in-job form), [[IR-73]] (second half BUILT: the within-job re-read exists and is what
  produced the FLOOR_DOCS reading), REQ-CI-FREE-SPACE-PREFLIGHT.
- **Coupling (deployah 23:54Z, from `golden.yml` at `25e60015`):** the Windows `DISK docs floor`
  (:616) and `Docs drift gate — windows` (:629) carry only `if: runner.os == 'Windows'`, no
  `always()`, while `DISK end floor` does (`always() && runner.os == 'Windows'`, ~:705). Actions'
  default is "previous steps succeeded", so ANY Phase B red skips the docs axis by construction —
  observed three times at this sha (floor red, ttl cell, ttl cell). The docs gate is therefore never
  independent evidence: it can only be banked when Phase B is green in the same job. Remedy line:
  gate the two docs steps on `!cancelled()` plus the success of the steps they actually depend on
  (checkout/build), not on Phase B — a rig change, new sha, hertz workflow rider slot.
- **Remedy shape:** re-derive the start floor from the measured footprint (67.4 GiB Windows,
  Linux to be measured) plus the 32 GiB end floor, print the footprint (start minus end free)
  on every run so the number re-measures itself, and keep the 32 GiB literal only for
  `FLOOR_END`.
- **Ripe when:** next `golden.yml` touch (hertz workflow rider slot). · **Size:** small — one
  literal per start-floor site + one printed subtraction.

### IR-87 — two-host ceremony halves are INDEPENDENT jobs on DIFFERENT runners (A Windows/hfenduleam, B Linux/kitsubito); B's 900 s budget runs on B's own clock, so any older queued Windows job starves A and reds B deterministically, and the panic text blames pairing

- **Status:** OPEN, filed by doyle 2026-09-08 at the #272 golden r2 triage. · **Origin:**
  run 34262154550, twohost-b job 102207296960 (log 81,692 B, sha256 67403f04…8030): role B built,
  rostered, polled PAIR_MEET_UP every 30 s from 19:34:30Z, and hit the rig_wait barrier
  "pairing: A rostered via the daemon-hosted responder" at its 900 s deadline 19:49:17Z
  (`crates/spt-daemon/tests/twohost.rs`, exit 101, 2 passed 1 failed). twohost-a (102207296985)
  was QUEUED from 19:32:13Z with no runner: the Windows runner went to thin ci run 34261096301
  (post-merge on main @ `e4444413`) the second the golden test job released it at 19:31:56Z.
- **Mechanism:** `golden.yml` sequences twohost-a and twohost-b after `test`, but they are two
  jobs on two runners (measured 2026-09-08 20:56Z via the jobs API `runner_name`: twohost-a on
  hfenduleam, twohost-b on kitsubito), and GitHub's queue does not honour a workflow's intra-run
  job ordering against a FOREIGN run on either runner. B starts, its whole budget ticks against a partner that has not been
  scheduled, and `rig_wait` panics `never converged on the rig: {what}` — a text that reads as a
  pairing defect. The lanes were not at fault, and twohost-a's outcome was predetermined once B
  exited (a later A pass would be vacuous).
- **Kin:** [[IR-76]] (hfenduleam is golden runner AND builders' box; the job-start census cannot
  see a competing GitHub run either), [[IR-62]] (rig collisions on shared runners), the W1 two-host
  rig, `a-barrier-cannot-ride-a-carrier-its-sender-outlives` (memory, same family: a barrier needs
  the receiver's own answer).
- **Remedy shape (pick at rider time):** (1) start barrier on B keyed to A's job having STARTED
  (a run-scoped artifact or a `needs`+matrix collapse into ONE job that spawns both roles), or
  (2) B's deadline clock starts at first A-rostered signal with a separate, longer "A never
  appeared" budget, or (3) the panic names the missing partner's job state so triage does not
  chase pairing. Any option: the two-host halves must never share a runner queue with a foreign
  run while one half waits.
- **Ripe when:** next `golden.yml`/twohost rig touch. · **Size:** small-medium (job shape or
  barrier rework + panic text).

### IR-88 — `pool-release` via `cargo run -p xtask` REBUILDS xtask INTO the pool it is releasing, so a reaped pool regrows ~2.8 GB in silence

- **Status:** OPEN, filed by doyle 2026-09-08 from hertz's reap report (20:00Z) during the #272
  golden r2 repair. · **Origin:** hertz reaped hertz-lane4 / hertz-repin / hertz-percell-id target
  subtrees (Length-sum upper bounds 68.16 / 47.11 / 8.79 GB) and then ran `cargo run -p xtask --
  pool-release …` from each worktree: cargo rebuilt xtask into the just-reaped `target/` (2.8 GB
  back in lane4, partial in repin), nothing printed said so, and the descendancy stopper killed
  the second before it finished. Removed by hand afterwards; the post-reap free reading on the
  record (275.95 GiB at 19:55:51Z) predates the regrowth.
- **Mechanism:** the documented verb (AGENTS.md: "drop it with `pool-release`") is a `cargo run`,
  and cargo's default target for that run IS the pool being released. The verb's own success path
  recreates the state the operator just measured away; a teardown step that runs it AFTER a reap
  undoes part of the reap and leaves the recorded free reading stale.
- **Remedy shape:** either (a) `pool-release` documented and scripted as a PREBUILT `xtask.exe`
  invoked against `--pool <dir>` from another pool, or (b) the verb refuses when its own
  `CARGO_TARGET_DIR`/default target equals `--pool` and prints the prebuilt form, or (c) the reap
  recipe orders release BEFORE the subtree delete. Memory banked 2026-09-08
  (`pool-release-rebuilds-xtask-into-the-pool-you-just-reaped`); this entry is the durable home.
- **POSITIVE CONTROL (hertz, 2026-09-09, ruling-3 gate-r3 teardown):** the same release run
  from a PREBUILT `xtask.exe` invoked from outside the pool regrew NOTHING — the pool's own
  `xtask.exe` still carried its original 21:00:42 mtime afterwards, and the reap that followed
  reclaimed 63.71 GiB (132.03 -> 195.74 GiB free) against a 64.88 GiB Length-sum. So the
  regrowth is a property of the `cargo run` VEHICLE, not of the verb, and remedy (a) is proven
  rather than merely proposed.
- **Kin:** [[IR-31]] (pool budgeting), [[IR-56]] (pool-claim identity from CWD — same verb
  family, same "the tool acts on where it stands" shape), releases#103.
- **Ripe when:** next xtask pool-verb touch or the teardown runbook edit. · **Size:** tiny —
  one guard or one doc line.

### IR-89 — hfenduleam's Windows Firewall drops cold inbound UDP to the runner-built test exes, so the two-host rig's first B→A claim (W2 helper) reds for its whole 900 s budget and the panic blames A

- **Status:** BOX HALF **APPLIED** 2026-09-08; WORKFLOW HALF **BUILT** 2026-09-09 — PR #211
  (`9b96d7e7` + this amendment), **LANDED-pending-golden**: the `golden.yml` steps are UNEXERCISED
  until the next golden run, because thin CI compiles the cells and skips them. Filed by doyle 2026-09-08
  at the #272 golden r2 terminal triage; the operator-blocked flag is retired here, because
  the operator acted.
  The operator applied BOTH layers that evening — the Windows inbound rule (~23:41Z) and the
  tailnet ACL grant (~23:56Z) — and the re-probe with the same rule-less pwsh listener read
  **3/3 over Tailscale on 7483 AND 3/3 on 7489 at 23:57:50Z** from `100.98.197.12`, with a
  post-terminal probe 3/3 on both ports at attempt 4. v0.68.0 then shipped at `a2f335f8` with
  both two-host legs GREEN on golden r4, so the box half is closed by measurement and not by
  assertion. **STILL OPEN, and it is the durable half:** the WORKFLOW guard (hertz) — the
  pre-ceremony B→A UDP probe at `d882297f` that reds `INBOUND_BLOCKED` in 10 s with its own
  name instead of burning 900 s blaming pairing, plus retiring the dead `_work\spt-core`
  rules note in the runner runbook. Until that lands, the next box (or the next ACL edit)
  re-opens the 900 s face with nothing to name it. · **Origin:** run 34262154550 @ `25e60015`, twohost-a 102229928746
  (`two_host_web_helper_role_a`, :968, 900.40 s) and twohost-b 102229928689
  (`two_host_web_role_b`, :559, 910.20 s): B logged 75 × "A not ready … broker QUIC op exceeded
  the 10s bound (peer unresponsive)" from 20:49:49Z to 21:04:37Z and was never ADMITTED. Ladder
  green both sides; both halves started 20:41:23Z on their own runners (not IR-87's starvation).
- **Mechanism (MEASURED 21:11Z):** the box's firewall policy is `BlockInbound,AllowOutbound` on
  all three profiles (Ethernet and Tailscale both Private). Inbound allow rules exist only for
  exes under `Documents\projects\spt-core\target` (interactive Allow clicks) and under the DEAD
  pre-rename runner path `_work\spt-core\spt-core`; exactly ONE under `_work\spt-bs-core` (its
  `target\debug\spt.exe`, the CLI, not the binder — deployah's count, mine had said zero), none for
  any `twohost_web-*.exe`. `Get-NetFirewallProfile` shows `DefaultInboundAction = NotConfigured`,
  whose effective default IS block: the block is Windows' default, not a configured policy, so there
  is no policy to "restore". The runner is a service (`.\decid`), so the Allow dialog that minted the
  old rules cannot appear. A's helper cell 3 bound `broker udp 7483` inside
  `_work\spt-bs-core\spt-bs-core\target\debug\deps\twohost_web-f73bc44be13353fb.exe`; B dialled
  `100.68.35.65:7483` cold. Probe: a rule-less pwsh UDP listener on hfenduleam:7483 received 0 of
  3 datagrams from kitsubito `100.98.197.12`; the reverse control (python listener on
  kitsubito:7483) received 3 of 3 from hfenduleam.
- **Why three faces from one cause:** the shared-key rig (pre-`e4444413`) reached A only through
  holes A's OUT-dialling cells (7480–7482) had opened for `id_a` — six 10 s bounds until B's
  magicsock fell onto one of those paths, then ADMITTED (the 63 s stall, helper-stall memory), or
  onto a live same-key sibling that never replied (the 21 min hang in 34239258523). Per-cell A
  identity (`id_a_for`, `e4444413`) removed the accidental route by construction, leaving only
  the cold path, which the firewall drops. hertz's one-box discriminator "vanishes" was true on
  one box because one box never crosses the firewall. Product and rig are both exonerated; the
  rig's `a_addr` comment already names W2 as the first A-ward claim.
- **Remedy (two halves):** (1) BOX, operator-owned, elevated on hfenduleam:
  `New-NetFirewallRule -DisplayName "spt-ci two-host rig UDP-In (kitsubito only)" -Direction Inbound -Action Allow -Protocol UDP -LocalPort 7460-7499 -RemoteAddress 100.98.197.12 -Profile Any`
  (rig ports at the sha: ladder 7460/7461; web `PORT_OFFSET` 20 → A cells 7480–7483, B 7481; the
  A-ward surface today is exactly UDP 7480–7483 — deployah's tighter range — and the whole range
  moves with any `SPT_TWO_HOST_PORT_A` override, so the rule must move with it);
  program-path rules are the wrong shape because the exe hash changes per build. Verify with the
  same probe (3/3) before any rerun. (2) WORKFLOW, hertz: a B→A UDP probe step in the twohost
  jobs before the ceremony, so this reds in 10 s with its own name instead of 900 s blaming
  pairing; and retire the dead `_work\spt-core` rules note in the runner runbook.
- **Kin:** [[IR-83]] (the tailnet-ACL face of a cross-box inbound block — KIN, **NOT the same
  finding**: that entry is the tailnet policy layer, this one is the Windows firewall layer, and
  the SECOND LAYER note below is exactly the measurement that separates them), [[IR-87]] (the
  other way a half waits 900 s on a partner), [[IR-76]] (golden runner
  is the operator's box), [[IR-64]] (box-level facts only the operator can move),
  `twohost-web-helper-stall-shared-a-identity-stale-path` (memory; its CONFIRMED arm is re-read
  by this entry).
- **AMENDMENT 2026-09-09 — the WORKFLOW half is BUILT and FIELD-MEASURED; it is LANDED-pending-golden
  (hertz).** The guard is two cells in `twohost_web.rs` plus a step per half in `golden.yml`, and it
  grew four design corrections that are worth more than the guard itself, because each one is a way a
  probe can certify the fault it exists to catch.
  1. **A FIXED SYMMETRIC WINDOW ADDS A THIRD CAUSE THAT LOOKS LIKE THE OTHER TWO.** The step boundary
     orders the probe on ONE host; it does not align two, and after the ladder each half pays its own
     wrapper and `cargo test` (on Windows a fingerprint scan plus Defender's first touch of a fresh
     exe). MEASURED: the first hand-run in the A=kitsubito direction (13:14:13 → 13:14:51Z) redded
     `INBOUND_BLOCKED` at 10.24 s with text asserting "a BOX rule, not a product or rig fault" — and
     the cause was `cargo`'s start-up on the SENDER putting its first datagram after the receiver's
     10 s window closed. An innocent box, the guard's own words against it. What stopped it being
     filed as a kitsubito block was that the box facts contradicted the text (`tailscale debug
     netmap` on kitsubito lists `100.68.35.65/32` among its permitted inbound sources, and `ufw` is
     inactive) — a failure text that contradicts a measured box is the text's defect. So the
     rendezvous is DATA: A listens up to the rig budget and beacons once a second; B binds first,
     waits for a beacon, and only then starts a ten-second clock, so B's red has already excluded
     "A is not up". That is a FOURTH outcome with its own name, `PROBE_NO_BEACON`, never
     `INBOUND_BLOCKED`.
  2. **THE LISTENING SOCKET MUST NEVER SEND.** Any outbound datagram from the measured port opens
     stateful return state for it, so B's probe then arrives SOLICITED and crosses under both layers
     — this entry's own "solicited return traffic works under either fault" applied to the guard
     rather than to the ceremony. MEASURED: with the beacon sharing A's listening socket, the
     out-of-grant control went **GREEN at 13:28Z** on a port a listen-only run had measured BLOCKED
     at 13:20Z; a further green at 13:31Z with the beacon moved but the ack still leaving that socket
     was residual state from 13:28 — true BY ELIMINATION rather than by assertion: nothing in the
     13:31 run sent from 7509 before the first arrival, so the only outbound that could have opened
     that port's return state belongs to the previous run. Beacon and ack both leave an ephemeral socket now, and B sends
     from an ephemeral port too, so no run repeats a 4-tuple a previous run opened.
  3. **THE ACK IS ADDRESSED TO B'S RIG PORT, not to the probe's source port.** MEASURED: acking to
     `from` produced a run that CONTRADICTED ITSELF — A said `INBOUND OK` and B said
     `INBOUND_BLOCKED`, same run, 13:41:33 → 13:41:50Z — because A's ack leaves its ephemeral
     socket, is therefore not return traffic of the flow B opened, and a receiver whose inbound rule
     is a port RANGE drops it. The rig port is the only address either host is reachable on cold.
  4. **THE FALSIFIER, and it is the operator's own rule used as an instrument.** Forcing the probe
     port to **7509** (`SPT_TWO_HOST_PORT_A=7480`), one port outside the granted udp 7460-7499, reds
     the guard on demand with no elevation and no policy edit. Run TWICE BACK TO BACK at the final
     code, 13:43:41 → 13:44:23Z and 13:44:25 → 13:45:07Z: **RED, RED** — A after 40 s having sent
     39 beacons, B in its ten seconds having sent 48 and 47 datagrams with no ack, both exit 101.
     One red would only prove a state timeout expired; two consecutive reds prove the design writes
     none. The grant boundary is therefore measured from inside the test binary: 7489 crosses, 7509
     does not, same hosts, same binary, one env var apart.
  5. **ONE ACK IS NOT ENOUGH — the same split verdict arrives by packet loss** (doyle, reviewing the
     built lane). A acking ONCE and breaking leaves B's red resting on a single unrepeated UDP
     datagram: lose it and A prints `INBOUND OK` while B reds `INBOUND_BLOCKED`, blaming a firewall
     for ordinary loss — correction 3's defect reached by a different road. A now keeps draining for
     a short window after its own verdict is already settled and acks EVERY probe in it, while B
     stops on the first ack. The A-side assertion is unchanged: it was decided by the first datagram.
  6. **THE PROBE STEPS CARRY THEIR OWN BUDGET, 300 s, set EXPLICITLY** beside ceremony steps that say
     900. It bounds RENDEZVOUS skew, not pairing, so a blocked link costs 300 s ONCE instead of 900
     per cell per half. Left unset it would have been `from_env`'s invisible default sitting next to
     a visible 900 on the step below — a number the next reader would have inferred wrongly, which is
     this entry's whole failure mode in miniature.
- **FIELD PROOF, 2026-09-09, and it is the ONLY proof this guard has:** `dir 1` (A=hfenduleam,
  B=kitsubito) 13:45:14 → 13:45:18Z and `dir 2` (A=kitsubito, B=hfenduleam) 13:45:41 → 13:45:58Z,
  both GREEN with the ack after ONE datagram, both exits 0, both hosts running byte-identical source
  (md5 `f31779339266f8034dcd72917c56a37e`, checked on both sides). Free space 174.38 GiB, flat across
  all four arms.
  ⚠ **THE `golden.yml` STEPS THEMSELVES ARE UNEXERCISED UNTIL THE NEXT GOLDEN RUN.** Thin CI
  compiles the cells and SKIPS them — `Rig::from_env` returns `None` without `SPT_TWO_HOST`, so a
  green thin run says the code builds and nothing about the steps. The wall times above are the field
  proof; treat the workflow half as landed-pending-golden and not as verified by its own PR.
- **Ripe when:** the next GOLDEN run, which is the first thing that exercises the steps — the
  code half is built and field-measured (PR #211). (The box half was ripe "now" against
  `25e60015` and is done; #272 golden acceptance is no longer blocked by it.) · **Size:** one
  workflow step remaining; the elevated command is spent.
- **SECOND LAYER (measured 23:41–23:50Z, after the operator applied the Windows rule exactly as
  asked and the re-probe STILL read 0/3):** the tailnet ACL. The same `python.exe` listener (own
  program rule) received 3/3 from kitsubito over the LAN (`192.168.1.168 → 192.168.1.81:7483`) and
  0/3 over Tailscale (`100.98.197.12 → 100.68.35.65`), TCP connect over Tailscale times out too,
  solicited return over Tailscale works, ShieldsUp false both ends. `tailscale debug netmap` on
  hfenduleam: one PacketFilter rule, 18 permitted inbound sources, kitsubito ABSENT; kitsubito is
  `tag:eye-tracking-resource` owned by a different tailnet user, and its own filter DOES permit
  hfenduleam. Asymmetric policy: member device → tagged resource allowed, reverse denied. The
  Windows rule was NECESSARY (rule-less exe, LAN probe 0/2 earlier) and NOT SUFFICIENT. Remedy
  half (1) gains an ACL grant, operator-owned in the admin console:
  `{"action":"accept","src":["tag:eye-tracking-resource"],"dst":["hfenduleam:7460-7499"],"proto":"udp"}`.
  Rider 4's failure text names both layers and the LAN-vs-Tailscale probe as the discriminator.
  Lesson for the entry: a cross-box "unresponsive" has at least THREE layers (Windows rule, tailnet
  policy, the exe's own bind); probe through the SAME path the rig uses, and re-probe after every
  single change — the first fix reading as "applied" is not the probe reading 3/3.
- **Corroboration + reading traps (deployah 23:50Z, from the box):** Windows rule found present by
  FILTER search (`netsh advfirewall firewall show rule name=all dir=in verbose`), not by name — the
  operator's name is hyphenated "two-host" and a `twohost` grep reads it ABSENT. Netmap: the single
  filter permits ALL ports and ALL protocols from its 18 sources, so the grant is SOURCE-scoped, not
  port-scoped — "this peer is not a permitted source at all", which is why TCP timed out beside UDP.
  Two permitted sources (100.98.213.33, 100.98.214.87) share kitsubito's 100.98/16 and read as hits
  to an eyeball scan for "100.98."; test membership of the full /32, never a prefix. His first netmap
  read extracted a non-existent field (`SrcIPs`; the real key is `Srcs`), printed an empty list and
  minted a confident "ABSENT" — an empty extraction cannot witness absence; the 18-count is from the
  corrected read. Prefs on the box: ShieldsUp false, NoStatefulFiltering true, RouteAll true,
  NetfilterMode 2 — nothing there explains the drop.
- **Probe-rider design rule (hertz, d882297f):** the pre-ceremony probe reds `INBOUND_BLOCKED` for
  EITHER layer and does not pretend to know which — its job is "the box cannot receive, stop triaging
  the product"; the layer is named only by the one measurement that decides it (same listener, LAN
  vs Tailscale), which the failure text prints as a recipe. A guard that names a cause it did not
  measure is the defect class this whole entry documents.

### IR-90 — a full disk on the self-hosted box reds a rig as a PRODUCT refusal, not a build error, and no rig or gate records the free space that would falsify it

- **Status:** OPEN, filed by hertz 2026-09-09 from the disk-full incident on hfenduleam. ·
  **Origin:** the IR-85 discriminator lane's head arm, killed by the volume rather than by the code
  under test.
- **Symptom:** `spt-daemon::sync concurrent_writes_reconcile_on_elected_node_and_converge` FAILED at
  67.588 s with a panic in our own test at `crates\spt-daemon\tests\sync.rs:198`:
  `pull: Custom { kind: Other, error: "sync refused: bundle failed: git -C <tmp>\tracked-b\.seed.git
  bundle create <tmp>\scratch\serve\serve-pull-6.bundle ^d12a7134... a-doyle failed (exit Some(1)):
  fatal: sha1 file '<stdout>' write error. Out of diskspace\nerror: pack-objects died" }`, then a
  second panic at `:213` (`pull thread: Any { .. }`) as the harness thread unwound.
- **Cause:** `C:` was at **0.018 GiB free of 1862.02 GiB (0.00%)** at that instant
  (`Win32_LogicalDisk`, 10:57Z). `git bundle create` could not write; the daemon's serve path turned
  that into its real product refusal string `sync refused: bundle failed`; the test asserted on the
  refusal. **Every layer behaved correctly, and the report reads as a sync regression at the sha
  under test.**
- **Why it is worse than the disk faces already known:** the two banked faces
  (`disk-full-reds-as-lnk1318-pdb-error`) are TOOLCHAIN costumes — `LNK1318` at link, and rustc
  I/O before any link — and both route a reader to "the box is sick". This one routes to a CODE
  OWNER: it names our file, our line and our refusal string, and the disk word sits in the FOURTH
  nested clause behind a git exit code. It arrived mid-discriminator with an old sha and a new sha
  side by side, where the cheapest reading — "the head arm failed, the old arm passed" — is a
  head regression that does not exist.
- **Blast radius, same incident:** the volume also carried a live `unit (self-hosted, Windows,
  hfenduleam)` job (run 34341010297) compiling into it, which completed FAILURE on its own and was
  ruled VOID after the fact. doyle's cancel of THAT run returned `Cannot cancel a workflow run that
  is completed` — recorded because a register that credits a controlled action nobody performed
  teaches the next reader that the box was under control. (The earlier run 34337797758, in
  **IR-85**'s amendment, is a different run with a different ending: CANCELLED at the job wall. One
  register putting one cancel string on the wrong run is the same defect, one run over.) `spt daemon status` went to peer pump last tick 185 s with 24 brain
  subscribers stall-evicted (last evict 10:51:22Z, inside the disk window), with an operator restart
  under consideration — a restart that risks the releases#287 shell stranding and that, had it
  appeared to help, would have taught everyone the wrong cause.
- **⚠ Two things that looked like disk fallout and are NOT**, recorded so this entry does not
  overclaim. (a) `serve list` returning `SERVE_UNCONFIRMED`: doyle read it from source —
  `servehost.rs` and `KIND_SERVE_REQUEST` are ABSENT at 0.67.0/0.67.1 and were minted by 0.68.0, and
  the resident broker is still the 0.67.0 image after the brain-only flip, so the CLI's 10 s bound
  times out **by design** against an older daemon. (b) `peer reachability: DEGRADED, 4 of 7 peers
  unreachable for 882775 s` — 10.2 days, predating the incident entirely. Even the pump term is
  claimed as CONCURRENT, not caused: hertz called the pump recovered off ONE post-reclaim sample
  (54 s) and doyle falsified it two minutes later (133 s), because a single reading of a monotonic
  "last tick N s ago" counter cannot separate STILL TICKING from TICKED ONCE, and a fresh project
  index proves the coordinator loop, not the pump.
- **Remedy — one guard, two placements, and it REFUSES rather than runs:**
  1. **Rig start.** For the two-host rig and any test rig that shells out to `git bundle` or writes a
     store: read free space on the volume holding the rig's temp root and `target`, and under a floor
     fail immediately with a message that says DISK and prints the number — never enter the
     ceremony. Floor: start at **10 GiB**, above the largest single artifact these rigs write and far
     below any healthy state of this box; tune only with a measurement.
  2. **Failure text**, for the case where space runs out MID-run and no start check can catch it:
     when a shelled-out git/store operation fails, append the current free space to the error before
     it becomes a product refusal string, so the panic that reaches a human already carries the
     falsifier.
- **WHY BOTH, measured rather than argued (hertz, 2026-09-09, the #209 lane):** a floor read ONCE at
  rig start would have PASSED that run and the run still ended nearly empty — `cargo check` started
  at **136.32 GiB** free, the test-profile build of 221 binaries took it to **40.18 GiB in ten
  minutes**, the legs finished at **16.24 GiB**, and it read **12.27 GiB** ninety seconds later. A
  start-only guard is a guard against yesterday's disk. That last ~4 GiB fell with those legs already
  terminal; the candidates are the runner's CI job and its `_work` tree, unsampled and therefore
  UNATTRIBUTED.
- **The record itself, and the only one that caught the dip (hertz, 2026-09-09):**
  `.spt/preserved/hertz-leak-670-legs/disk-full-trace.log` — **139 samples, 11:52:46Z–13:57:07Z, one
  per minute**, max **238.41 GiB** @12:45:20Z, min **12.53 GiB** @12:12:45Z. Two samples fall under
  15 GiB — **14.85** @12:12:11Z and **12.53** @12:12:45Z — both inside the #209 leak lane's
  build-and-teardown window, i.e. AFTER a start floor would have passed the run at **136.15 GiB**
  (the window's first sample). Two caveats, both binding on any reading of it: **the true minimum is
  unbounded below 12.53 — the sampler cannot see between its samples**, and **it attributes nothing:
  hertz's own legs and the runner's CI job are indistinguishable in it.** It is the SECOND instrument
  on this incident and the only one covering the dip: doyle's 30 s trace, preserved at
  `.spt/preserved/ir90-free-space-trace-2026-09-09/free-sampler-12-13Z-to-12-28Z.txt`, runs
  12:15:37Z–12:31:10Z and so begins after hertz's reap — it records the RECOVERY, this one records
  the FALL. sha256 `d58ce838b677e62d4a017e77524f899fda81e7bacf73a0e2a05a1192e60abaf2`
  (doyle's, `8bdae01bc151597070d7011443a2cdc74614232dbc962faec1b1b1833c71f22a`).
- **Explicitly NOT the remedy:** a bigger disk, or a reap schedule. ~195 GiB of headroom was consumed
  by ordinary work, ACCRUED across a ~45-minute build phase; **no reading bounds a rate** — the
  free-space observations either side are endpoints of an accrual, and a rate derived from them
  retargets a hunt (deployah's went to runaway logs, VSS and torrent preallocation on the strength of
  a `>150 MB/s` derived that way, while the measured live box-wide write rate was ~1.5 MB/s). What
  consumed it: two `cargo nextest run -E <5-test filter>` lanes at **82.88** and **64.44 GiB** — a
  filter narrows the RUN, never the BUILD, and both built 221 test binaries under the `test` profile
  — plus 40.93 GiB of `C:\actions-runner\_work` and 9.8 GiB of `%TEMP%`. Any headroom this box
  has is two cold pools away from gone, so the guard must be a REFUSAL, not a budget.
- **Kin:** **IR-85** (the same box's fs-heavy slowdown; a near-full volume is now a named environment
  term there), **IR-86** (a start floor that passes a box which cannot finish — the same defect one
  layer up, and this entry is its mid-run half), **IR-76** (the golden runner is the operator's
  desktop), `disk-full-reds-as-lnk1318-pdb-error` (memory; this is its third face and its standing
  gap), `test-profile-pool-outgrows-the-disk-floor`, `free-space-floor-blocks-golden`.
- **Ripe when:** now. · **Size:** one assertion plus one error-context append, both test-side.

### IR-91 — 561 untracked files sat at the repo root, 18 of them cited by tracked files, and the census that classified them used a corpus that made the discriminator invisible

- **Status:** CLOSED by the lane that filed it (the root-scratch classify PR), hertz 2026-09-09 on
  doyle's census and rulings. · **Origin:** the `.spt/` ignore lane (PR #212) named 561 untracked
  root files as explicitly out of scope; this is that half.
- **The population, measured at `2037bcb8`:** `git status --porcelain --untracked-files=all` = **561
  `??` rows, all at the repo root, zero in subdirectories**, 50,034,994 bytes (47.7 MiB). By
  extension: 268 `.raw`, 184 `.exit`, 87 `.md`, 7 `.out`, 4 `.log`, 4 `.done`, 2 `.txt`,
  2 `.stackdump`, 2 `.sh`, 1 `.py`. **A `git add -A` at the root staged every one of them**, which
  is the same landmine `.spt/` carried before #212 and the reason this is one arc in two lanes.
- **⚠ THE FINDING, and it is about the CENSUS, not the files:** the first classification put two
  cited files in the "bury" class. `W3-196-CENSUS.md` and `WEBSERVE-272-JIT.md` were recorded as
  "cited only from other root scratch" — citing files `W3-196-JIT.md` and `W3-272-MEASUREMENTS.md`,
  which are **TRACKED**. The census corpus was tracked-only, so **every hit it produced was a
  tracked citation by construction**; the two were then relabelled by their PATH, because a `.md`
  sitting at the repo root reads as scratch. **The discriminator is TRACKED-OR-NOT, never
  ROOT-OR-NOT** — and a corpus that can only return tracked hits cannot tell you which of its own
  hits are tracked, because it has already thrown away the distinction. Burying those two would have
  left two tracked files citing paths that no longer existed: the `launch-battery.py` mechanism from
  **IR-90**'s sibling lane with the sides swapped, caught here only because the burial list was
  re-verified against the corpus a second time by a different reader.
- **The second finding, which withdrew the lane's own destinations:** the repo root **already holds
  140 tracked files, 134 of them `.md`, 25 with `JIT` in the name** — `IDLE-EDGE-JIT.md`,
  `W3-196-JIT.md`, `WEBSERVE-272-W2-JIT.md`, every `M*-PLAN.md`, `MESH-D*-PLAN.md`,
  `RESTORATION-D*-PLAN.md`. **The house convention is measured, not assumed: a plan or record lives
  at the repo root and is TRACKED.** The first ruling would have relocated the cited files into
  `docs/intake/`, `docs/design/` and `docs/gate-records/` — splitting siblings across two homes
  (`WEBSERVE-272-W2-JIT.md` at root, `WEBSERVE-272-W3-DRIFT-RIDERS.md` in `docs/intake/`) and
  repointing 18 citation lines to buy that split. **Destinations withdrawn; the 18 are tracked where
  they are.** Zero moves, zero repoints, zero sha-transfer risk. A lane that tidies against a
  measured convention is not tidying.
- **THE RULE, which is what this entry is for:** *a plan or record cited from a tracked file is
  itself tracked, at the path the citation already names.* The root is a legal home. What decides a
  file's fate is whether the repo points at it — not where it sits, not what it is called, and not
  whether it looks like scratch.
- **What the lane did:** **18 Class A** files `git add`ed in place. **473 Class B** gate legs
  (`.raw`/`.exit`/`.done`/`.log`/`.out`/`.stackdump`/`.sh`/`.txt`, zero tracked citations, each
  word-bounded against the whole 911-file tracked corpus) **MOVED — not deleted** — to
  `.spt/preserved/root-gate-legs-2026-09-09/` with a sha256 manifest, every file hashed before and
  after the move. **70 Class C** uncited `.md` likewise to `.spt/preserved/root-md-2026-09-09/`.
  Then five **root-anchored** ignore lines: `/*.raw` `/*.exit` `/*.done` `/*.out` `/*.stackdump`.
- **Why those five and not more.** Legs are the one class whose FUTURE instances are scratch **by
  construction**, so ignoring them institutionalizes no loss. `/*.md`, `/*.log`, `/*.sh` and
  `/*.txt` are deliberately NOT ignored: a plan or script dropped at the root must keep showing as
  `??` so the next census classifies it — **that is exactly the pattern that produced this repo's
  140-file tracked root corpus**, and an ignore line there would hide the next one. The anchoring is
  load-bearing and was proved with a negative control: `crates/probe.raw` is **not** ignored, which
  an unanchored `*.raw` would have swallowed.
- **Kin:** the `.spt/` ignore lane (PR #212 — same arc, and its `launch-battery.py` case is why
  nothing here was deleted), **IR-90** (a citation into `.spt/preserved/` is machine-bound evidence,
  never a repo path), `ignoring-a-directory-buries-what-the-repo-cites-in-it` (memory).
- **Ripe when:** landed with this entry. · **Size:** one `.gitignore` stanza, 18 adds, 543
  preserved files, no deletions.
