STEP-LEVEL CONTROL, measured by me at 14:11-14:14Z on the jobs/steps endpoint, three arms. This kills the duration question permanently and gives your harvest a falsifiable prediction BEFORE the log lands. PHASE A (step 19, "Test - Phase A (light pool; bounded parallelism) - windows"): r2 gate 34481993681 job 102886952516 13:36:29->13:57:35 21m06s FAILURE (job total 30m12s) arm 2 34474627303 job 102862459074 12:12:50->12:34:37 21m47s success (job total 66m19s) record 34445961595 job 102770789323 06:52:07->07:18:15 26m08s FAILURE (job total 34m32s) So r2's FAILING Phase A ran 41s SHORTER than arm 2's PASSING Phase A, and 5m02s shorter than the record red's. PHASE A DURATION DOES NOT SEPARATE RED FROM GREEN. Your withdrawal of "did not run to completion" was right, and this is the replacement: the 36-minute job gap is entirely the ladder BELOW Phase A, not the testing. WHAT ARM 2 RAN THAT r2 DID NOT (why 66m19s vs 30m12s, structurally): Phase B windows 12:34:37->12:59:50 25m13s success <- SKIPPED on r2 and on the record Doctests windows 12:59:50->13:00:51 1m01s success <- SKIPPED on both reds Clippy windows 13:00:51->13:05:09 4m18s success <- SKIPPED on both reds notify E2E 13:05:56->13:06:56 1m00s success <- SKIPPED on both reds installer E2E 13:07:00->13:07:18 18s success <- SKIPPED on both reds Fixture build is flat across all three: r2 7m30s, arm 2 7m13s, record 6m52s. IDENTICAL STEP SHAPE ON BOTH REDS: step 19 fails, steps 20-42 cascade to skipped at the same instant, cleanup 43-49 runs (DISK end floor, reap+census, sandbox removal, bench ledger all success). The r2 red and the record red match at STEP level. That is a SHAPE match and NOT a cause match -- the failing CELL inside Phase A is what decides Finding 1, and only the log answers that. I am not calling it a reproduction and neither should the panel. PREDICTION FOR YOUR HARVEST, stated so it can be WRONG (CONTRACT FINAL, boundary map authoritative): the expected Summary count on the r2 Windows leg is ONE, not two. Phase B SKIPPED, so the job's supposed-to-have execution count is 1. Its population should sit near arm 2's Phase A figure of 3381; arm 2's 238 was its PHASE B population, and a 238-shaped Summary on r2 would be anomalous because Phase B never ran. If you find TWO Summaries on this leg, that is SURPLUS over the boundary map = VOID, and it wants investigating before any FAIL line is read. This is the exact case where holding "expected = 2" would have made a clean single-Summary leg look like missing evidence -- which is why the boundary map is authoritative and the count never was. Your instrument report is accepted and it is the better half of that message: a detector that notifies ONLY ON EXIT with an all-jobs-terminal predicate is silent through precisely the event it exists to catch, and twohost-a/b materializing 3s after the Windows leg died pushed its exit an hour out. Worse is concl=[] empty at every tick INCLUDING after a failure -- on exit it would have handed you an empty list reading as CLEAN. Banking that: AN EMPTY COLLECTION FROM A BROKEN COLLECTOR IS NOT A CLEAN RESULT, and silence is not evidence of nothing happening. todlando's repair spec is right and I endorse it as the gater: notify on INDIVIDUAL JOB TRANSITIONS, not whole-run completion, and VALIDATE the repaired collector against THIS already-failed Windows job -- a collector that cannot see a known red is not repaired. Until it is repaired, fresh jobs-endpoint reads with an observation timestamp on every report, which is what you did at 14:10:24Z and what I did at 14:11:43Z (my read confirms your refreshed numbers exactly: 8 rows, Windows test FAILURE 13:27:58->13:58:10Z, Linux test SUCCESS, n1-gate Linux SUCCESS, twohost-a/b in_progress). UNCHANGED: classification mine and still UNCLASSIFIED, 34445961595 the verdict of record, release HELD, v0.69.0 uncleared, core#218 unmerged, no tag, no publish, no board work. Next from you is the panel; take the time to map it properly, the twohost legs have an hour to run and nothing is waiting on speed.