
### IR-85 — the Windows self-hosted box runs its fs-heavy tests 3-4x slower than a week ago; two CI wall clocks were sized for the old box, and a red run kept hiding it

- **Status:** OPEN, filed by hertz 2026-09-09 at the #272/v0.68.0 golden r3/r4 arc, on doyle's
  dispatch. Number ruled by doyle from his own census of main@`a2f335f8` (register ended at IR-83;
  IR-84 was branch-claimed by name only, with no entry text in any register file in any worktree).
  **Absorbs two drafts:** doyle's `IR-DRAFT-windows-fs-heavy-slowdown-and-golden-wall` (his IR-A +
  IR-B) and his earlier `IR-NEXT` (operator-desktop load), both retired by reference — this is the
  single entry. Every job/step and per-test timing below is **doyle's**, read via the job/step API.
  · **Origin:** #272 golden r3 attempt 2 was CANCELLED by a wall clock with both test phases green.
- **⚠ THE CAP CHANGE IN `a2f335f8` IS NOT THE FIX AND MUST NOT BE READ AS ONE.** It bounds the run;
  it repairs nothing. A future reader who finds an 80-minute wall and no entry here would reasonably
  conclude the problem was solved. It was only made visible.

- **A GREEN RUN COSTS MORE THAN A RED ONE, WHICH IS WHY THIS WENT UNSEEN.** A failing run
  short-circuits past the wall a green run has to cross. r3 att1 finished in **48m39s only because
  it FAILED at Phase B**; att2, green, hit 49m59s and was cancelled. Every earlier docs-drift skip
  therefore presented as "upstream failure" — the wall was never the reported cause of anything, and
  the surviving evidence systematically flattered the budget.

- **THE WALL.** Windows golden `test` job (48 steps), fully GREEN: **30m53s** (2026-08-30, run
  33296634901) and **33m21s** (2026-09-06, run 34017906638). At the v0.68.0 head against the
  then-current 50-minute cap — **r3 att2**, job 102352551368 at `f6110c2a`: Phase A 17m14s, Phase B
  20m25s, doctests 1m04s, clippy 3m31s, **CANCELLED at step 30 at 49m59s with everything green**;
  steps 31-42 (installer, docs floor, both docs-drift gates, dormancy) never ran. r2 att1
  (`25e60015`, 09-08 18:47Z) reached step 34 at 44m30s. Steps 30..39 cost ~3m30s on the 09-06 green
  (notify 47s, installer 13s, docs-drift 1m54s). Predicted green need ~**56 min**; cap raised to
  **80** (need + ~40% for day-to-day variance) in `a2f335f8`, ci unit 25 -> 40 in the same commit.
- **r4, GREEN — the first complete measurement at the head: 54m35s** (07:12:26Z -> 08:07:01Z),
  **25m25s** headroom under 80; docs-drift (step 38) 2m23s; Phase A 15m51s, Phase B 22m56s. The ~56
  prediction held to within 1.5 min, so the cap is sized on evidence, not generosity. **What that
  does NOT establish:** it was a sizing forecast, and its holding says nothing about the cause
  diagnosis below. Nor are att2-vs-r4 per-phase deltas a trend — att2 was cancelled mid-run, so only
  its phase legs are comparable at all, and the spread they sit in is the same variance the 40% is
  there to absorb.

- **THE SLOWDOWN, per test, same box, same tests.** `spt-daemon::sync` concurrent_writes
  **22.4s -> 74.8s**, two_tier_sync 18.7 -> 67.4; `spt-store` monic clone_copies 17.7 -> 62.9,
  different_monics 18.2 -> 64.9; syncmerge reconciled_write 27.3 -> 51.2. **~25 spt-store/spt-daemon
  tests now exceed 30s on Windows Phase A at `f6110c2a`, against 1-2s each on kitsubito.**
  Phase-level: Phase A nextest 170s (09-06) -> 642s (att1) -> 1034s (att2); Phase B 710s -> 1164s.
  The suite did not get more expensive — Linux is unchanged.

- **DATING SAYS ENVIRONMENT DOMINATES, AND BY HOW MUCH.** main's own ci Windows `unit` job:
  11-12 min on 09-06 (runs 34040416870, 34041526195, 34043323577) -> 13-20 min on 09-07 -> 12-24 min
  on 09-08, hitting **22 min at `e4444413`** (run 34261096301) three minutes under its own 25-minute
  wall. **Gradual over days, under no single gate change, is the shape of an environment term, not a
  commit's.** Head growth is real but MINOR: +138 Phase A tests, +20 Phase B, HEAVY 34 -> 35.

- **THE BOX IS AN OPERATOR DESKTOP (folded in from the retired IR-NEXT), and that turns every fixed
  wall-clock budget in the Windows suite into a coin.** From the r2 arc, three attempts at ONE sha
  (run 34262154550 @ `25e60015`): Phase B red each time on a **different single cell or none** — a1
  234/234 (job red on disk floors only), a2 `spt::webserve_attachment_e2e` arm 12, a3
  `spt-daemon::mesh_recovery roster_route_survives_a_transient_dial_failure_with_discovery_disabled`
  (15.0s `converge()` budget, cell 15.715s; the same cell 9.8s / 7.2s on a1/a2). Phase A — 3,346
  spawn-dominated unit cells — slowed **monotonically 448.7 -> 495.1 -> 542.6s at that one sha**.
  Per-cell a3/a2 over 73 Phase B cells >= 1s: median 1.05x, mean 1.34x, 19 cells >= 1.5x, worst 5.2x
  (`endpoint_lifecycle poll_vs_reap` 1.1 -> 5.8s) — **BURSTY, not uniform**. A rotating single victim
  across attempts at one sha is ONE environment cause; hardening victims one at a time never closes
  it (paid before: `e2e-leaked-daemons-shared-box`). The budgets that lost were ~1.5x the fast
  observation — inside the box's measured variance, so they were coins that had been landing right.
- **Box census (doyle 01:13-01:19Z, no cargo/rustc/nextest running, CPU 31%, ~1.1 of 16 cores busy):**
  `qbittorrent.exe` seeding since 09-08 10:20Z (box tx 21.3 MB/s over 5s); fleet `spt daemon brain`
  with 516 GB read since 09-07 08:03Z (~3.5 MB/s steady); Defender real-time ON, `MsMpEng` at
  67.8/62.8/48.9/32.2/12.2% of a core over 5s **on the idle box**; a fresh 35 MB exe pays
  2092/994/1171/1043 ms on FIRST execution against 31/263/19/260 ms on the second — **and every CI
  attempt rebuilds every test binary fresh.** At 06:45Z on 09-09, with the twohost legs running:
  MsMpEng 89% CPU / 991 MB WS, qbittorrent pid 47056 holding 7817 CPU-seconds (2.2 h) since
  09-08 03:20, free 134 GiB (275 -> 197 -> 131 across the three r3 dispatches).
- **⚠ AN UNREADABLE ROW IS NOT AN ABSENT ONE.** The Defender exclusion list cannot be read
  unelevated on this box: `Get-MpPreference` returns the literal string
  `N/A: Must be an administrator to view exclusions` **as the ExclusionPath VALUE**, so a
  `-contains` test reads ABSENT and is meaningless; the HKLM `Windows Defender\Exclusions\Paths`
  read throws `SecurityException`. Nobody may report the runner directory as unexcluded from an
  unelevated shell.

- **HYPOTHESIS ALREADY KILLED, so nobody re-runs it:** the Windows-only `ADAPTER_WEB_PENDING`
  reconcile failure (servehost nudge) **cannot** explain this — spt-store monic and contextstore
  never touch that path.

- **WHAT IS STILL NOT ESTABLISHED (labelled, so it is not inherited as fact).** The dating argument
  establishes environment-DOMINANT and bounds head growth as the minor term. It does **not**
  apportion the environment term itself: Defender vs the third-party torrent load vs the fleet
  brain's steady read vs anything else is unsplit, and **"MsMpEng at 89%" remains a correlate
  measured beside the slowdown, not a proven cause.** The discriminator lane below is what settles
  head-vs-environment on evidence rather than on the shape of a drift curve.

- **REMEDY — none landed. This entry is the debt, and its middle arms need an operator.** Sequence
  matters; run them in this order:
  1. **DISCRIMINATOR LANE (hertz, one box, ~1.5-2 h).** Run the five named tests at `04e32c8c` and
     at `f6110c2a` on hfenduleam. **Same-slow at both = environment; slow only at the head = head
     growth.** Cheap and decidable, and it must precede any operator ask — do not spend an elevation
     request on a hypothesis a one-lane measurement can test. Design note, because it is the part
     that makes the number trustworthy: the arms run **INTERLEAVED** (A/B/A/B/A/B, 3 reps each), not
     all-A-then-all-B, so drift that hits the whole box cancels in the comparison instead of landing
     on one arm. **A measurement that only works if everyone behaves is not a measurement** — the
     torrent client and Defender are running throughout and are the thing under test, not noise
     anyone gets to remove. Confounder already excluded: the three test-bearing files
     (`spt-daemon/tests/sync.rs`, `spt-store/src/monic.rs`, `spt-store/src/syncmerge.rs`) are
     **byte-identical blobs at both shas**, so a difference cannot be the tests changing; the two
     crates around them are not (+11,388 lines over 48 files).
  2. **Remove the third-party load first.** No torrent client on the CI box during golden windows.
     It is the cheapest variable to remove, and removing it first makes arm 3's benefit measurable
     instead of confounded.
  3. **OPERATOR ASK — Defender path exclusions** for `C:\actions-runner\_work` and the gate pools,
     read from an ADMIN shell first (see the unreadable-row warning above), then re-measure the
     first-touch tax on a file **under that path** — the scratchpad measurement is outside the
     runner dir, so it proves the mechanism's size, not the runner's exposure. Unsettable
     unelevated; same shape as **IR-82**'s elevated firewall rule.
  4. **CI census, cheap and durable:** print `Get-MpComputerStatus` RealTimeProtectionEnabled plus
     the ExclusionPath read **verbatim, refusal text included**, in the Windows job's census step,
     so a run's own record says what Defender it ran under.
  5. **Re-measure both caps after any arm lands — the arm a future reader will skip, so it is loud
     here.** A cap sized against a degraded box is correct only while the box stays degraded.
     Leaving 80 in place after a repair silently restores the original hazard: a wall so generous it
     no longer catches a wedged test, which is the job the golden cap was added to do in the first
     place (the 2026-06-03 handoff.rs ConPTY stall, 22 unbounded hosted minutes).
- **Kin:** [[IR-76]] (the golden runner IS the builders' box — the structural reason a desktop's
  load reaches CI at all), [[IR-64]] (box-level facts only the operator can move), [[IR-82]] (the
  other elevated box-rule ask), `defender-first-touch-tax-on-fresh-test-binaries` and
  `e2e-leaked-daemons-shared-box` (memory).
- **Ripe when:** arm 1 now; arms 2-4 on the operator's answer; arm 5 at the next `golden.yml` touch
  after any of them. · **Size:** arm 1 a measurement, arm 4 one census line, arm 5 two literals.
