### IR-85 — the Windows self-hosted box runs its fs-heavy tests 3-4x slower than a week ago; two CI wall clocks were sized for the old box, and a red run kept hiding it

- **Status:** FILED 2026-09-09 (hertz), on doyle's dispatch during the #272 golden r3/r4 arc. Number
  ruled by doyle from his own census of main@`a2f335f8`: the register ends at **IR-83**, and
  **IR-84 is claimed by the branch `fix/ir84-pump-peer-budget-instrument` by NAME ONLY — no entry
  text in any register file in any worktree** — so 85 is free and promised nowhere.
  **Absorbs doyle's parallel draft** `IR-DRAFT-windows-fs-heavy-slowdown-and-golden-wall.md` (his
  IR-A + IR-B), retired by reference; this is the single entry. Every job/step and per-test timing
  below is **doyle's**, read via the job/step API. · **Origin:** #272 golden r3 attempt 2 was
  CANCELLED by a wall clock with both test phases green.
- **⚠ THE CAP CHANGE IN `a2f335f8` IS NOT THE FIX AND MUST NOT BE READ AS ONE.** It bounds the run;
  it repairs nothing. A future reader who finds an 80-minute wall and no entry here would reasonably
  conclude the problem was solved. It was only made visible.

- **A GREEN RUN COSTS MORE THAN A RED ONE, WHICH IS WHY THIS WENT UNSEEN.** A failing run
  short-circuits past the wall a green run has to cross. r3 att1 finished in **48m39s only because
  it FAILED at Phase B**; att2, green, hit 49m59s and was cancelled. So every earlier docs-drift
  skip presented as "upstream failure" — the wall was never the reported cause of anything, and the
  surviving evidence systematically flattered the budget.

- **THE WALL.** Windows golden `test` job (48 steps), fully GREEN:

  | when | run | wall |
  |---|---|---|
  | 2026-08-30 | 33296634901 | 30m53s |
  | 2026-09-06 | 34017906638 | 33m21s |

  At the v0.68.0 head against the then-current **50**-minute cap — **r3 att2**, job 102352551368 at
  `f6110c2a`: Phase A 17m14s, Phase B 20m25s, doctests 1m04s, clippy 3m31s, **CANCELLED at step 30 at
  49m59s with everything green**; steps 31-42 (installer, docs floor, both docs-drift gates,
  dormancy) never ran. r2 att1 (`25e60015`, 09-08 18:47Z) reached step 34 at 44m30s. Steps 30..39
  cost ~3m30s on the 09-06 green (notify 47s, installer 13s, docs-drift 1m54s).

  Predicted green need ~**56 min**; cap raised to **80** (need + ~40% for day-to-day variance) in
  `a2f335f8`, ci unit 25 -> 40 in the same commit.

  **r4, GREEN — the first complete measurement at the head: 54m35s** (07:12:26Z -> 08:07:01Z),
  **25m25s** headroom under 80; docs-drift (step 38) 2m23s; Phase A 15m51s, Phase B 22m56s. The ~56
  prediction held to within 1.5 min, so the cap is sized on evidence, not generosity. **What that
  does NOT establish:** it was a sizing forecast, and its holding says nothing about the cause
  diagnosis below. Nor are att2-vs-r4 per-phase deltas a trend — att2 was cancelled mid-run, so only
  its phase legs are comparable at all, and the spread they sit in is the same variance the 40% is
  there to absorb.

- **THE SLOWDOWN, per test, same box, same tests.** `spt-daemon::sync` concurrent_writes
  **22.4s -> 74.8s**, two_tier_sync 18.7 -> 67.4; `spt-store` monic clone_copies 17.7 -> 62.9,
  different_monics 18.2 -> 64.9; syncmerge reconciled_write 27.3 -> 51.2. **~25 spt-store/spt-daemon
  tests now exceed 30s on Windows Phase A at `f6110c2a`, against 1-2s each on kitsubito.** Phase-level:
  Phase A nextest 170s (09-06) -> 642s (att1) -> 1034s (att2); Phase B 710s -> 1164s. The suite did
  not get more expensive — Linux is unchanged.

- **DATING SAYS ENVIRONMENT DOMINATES, AND BY HOW MUCH.** main's own ci Windows `unit` job:
  11-12 min on 09-06 (runs 34040416870, 34041526195, 34043323577) -> 13-20 min on 09-07 -> 12-24 min
  on 09-08, hitting **22 min at `e4444413`** (run 34261096301) three minutes under its own 25-minute
  wall. **Gradual over days, under no single gate change, is the shape of an environment term, not a
  commit's.** Head growth is real but MINOR: +138 Phase A tests, +20 Phase B, HEAVY 34 -> 35.
  Box at 06:45Z 09-09 with the twohost legs running: **MsMpEng 89% CPU / 991 MB WS**; **qbittorrent
  pid 47056 holding 7817 CPU-seconds (2.2 h) since 09-08 03:20**; free 134 GiB (275 -> 197 -> 131
  across the three r3 dispatches).

- **HYPOTHESIS ALREADY KILLED, so nobody re-runs it:** the Windows-only `ADAPTER_WEB_PENDING`
  reconcile failure (servehost nudge) **cannot** explain this — spt-store monic and contextstore
  never touch that path.

- **WHAT IS STILL NOT ESTABLISHED (labelled, so it is not inherited as fact).** The dating argument
  above establishes environment-DOMINANT and bounds head growth as the minor term. It does **not**
  apportion the environment term itself: Defender vs the third-party torrent load vs anything else
  on the box is unsplit, and **"MsMpEng at 89%" remains a correlate measured beside the slowdown,
  not a proven cause.** The discriminator lane below is what settles head-vs-environment on
  evidence rather than on the shape of a drift curve.

- **REMEDY — none landed. This entry is the debt, and its middle arms need an operator.** Sequence
  matters; run them in this order:
  1. **DISCRIMINATOR LANE (hertz, post-publish, one box, one hour).** Run the five named tests at
     `04e32c8c` and at `f6110c2a` from warm pools on hfenduleam. **Same-slow at both = environment;
     slow only at the head = head growth.** This is cheap, it is decidable, and it must precede any
     operator ask — do not spend an elevation request on a hypothesis a one-hour lane can test.
  2. **Remove the third-party load first.** No torrent client on the CI box during golden windows
     (qbittorrent pid 47056 above; kin to a runner-desktop-contention entry that still needs
     filing). It is the cheapest variable to remove, and removing it first makes arm 3's benefit
     measurable instead of confounded.
  3. **OPERATOR ASK — Defender path exclusions** for `C:\actions-runner\_work` and the gate pools.
     Unreadable and unsettable unelevated on this box; same shape as **IR-82**'s elevated firewall
     rule. Already measured beside it: fresh test binaries pay a 1.0-2.1s first-touch tax against
     20-260ms warm, with MsMpEng at 68% on an *idle* box.
  4. **Re-measure both caps after any arm lands — the arm a future reader will skip, so it is loud
     here.** A cap sized against a degraded box is correct only while the box stays degraded.
     Leaving 80 in place after a repair silently restores the original hazard: a wall so generous it
     no longer catches a wedged test, which is the job the golden cap was added to do in the first
     place (the 2026-06-03 handoff.rs ConPTY stall, 22 unbounded hosted minutes).
