---
name: req-confusion-audit-run
description: "Full-manifest requirement confusion audit (operator-directed 2026-07-28) — lia owns the run; status, design rulings, and stall/recovery state"
metadata: 
  node_type: memory
  type: project
  originSessionId: b64c7b8a-be77-4b68-9195-cb9bdae97e8e
  modified: 2026-07-29T13:28:08.331Z
---

# REQ confusion audit — full spt-core manifest (2026-07-28)

**Operator directive:** install latest BigscreenVR/traceable-reqs release, run the confusion audit workflow — FULL roster, entire run handed to **lia**. Doyle = installer + gater only.

## Installed (done, verified)
- traceable-reqs 0.1.4 → **0.2.0** @ `%LOCALAPPDATA%\Programs\traceable-reqs\traceable-reqs.exe`.
- Bundled `req-confusion-audit` skill → `~/.claude/skills/req-confusion-audit/` and `req-confusion-score.js` → `~/.claude/workflows/` — from repo **main** (v0.2.0 tag predates them); install skill refreshed too.

## Run design (doyle-gated)
- 597 reqs (548 active) → C(597,2)=177,906 pairs. Titles are mini-specs (mean 960 chars!) so lia block-partitioned: seed-fixed random 12-blocks → 24-req sub-rosters per ~150-pair chunk, 1250 chunks, 9 batches, ~3 workflows concurrent. APPROVED (full coverage, not prefiltering).
- REDIRECT enforced: scouts stay on schema returns, NO write tools; the real wall was the workflow's ~20k-pair aggregate return → lia harvests dense rows from journal.jsonl coordinator-side. Modified copy of req-confusion-score.js (per-chunk roster) invoked by scriptPath.
- Batch-local consolidation = non-authoritative (compacted roster); FINAL cross-batch consolidation over union of candidates ≥30 (~12k pairs, ~50 full-evidence slices) before any ranked list. 6 anchors in every chunk; observed scorer drift ±20 pts → per-chunk anchor rescale. NS never 0.
- Pilot PASS: 522/522 dense, 197 tok/pair → ~36M tokens projected.

## Findings so far
- **357/597 reqs have ZERO doc-stage evidence** (389 don't require doc stage) — id+title-only scoring, report must flag.
- Pilot top pair: REQ-WAN-SPT-HOSTED-DELIVERY / REQ-IDLE-PARKED-DELIVERY @ 60 ("same root guarantee").
- `traceable-reqs check` on spt-core: 0 findings, ~30s.

## ⚠ STALL (2026-07-28 ~01:03→) — lia session limit
- Fleet (agent-*.jsonl) stopped 01:03; her main session hit **"session limit · resets 4:20am PT"** at 02:48 (banner again 03:17). Wake attempts 02:47/03:17/04:25 all produced NO turn — perch row ONLINE, session dead (row≠life, again). Post-reset 04:25 wake also failed.
- Batches 0-2 died mid-flight; completed chunks are journal-cached — **resume via scriptPath + resumeFromRunId, never restart** (recovery brief queued to lia, drains on her next real wake).
- **CONFIRMED DEAD-IN-SESSION (05:44 verdict):** post-reset wakes 04:25:49 and 05:25:49 (exact-hour timer) both phantom — file touched, tail still the 03:17 system entry, zero turns, zero fleet. Self-recovery did NOT happen; operator bounced her hosting session during the day.
- **✅ stuck-mechanism RCA'd + FIXED adapter-side (perri, claude-spt 0.25.25):** CC's synthetic limit banner rendered as an Agent digest entry ⇒ interrupt_watch read the limited session as "resumed" and stood down once CC went hook-silent. Fix: banner = second stuck-marker vocabulary (limit turn never a resume, window-scan + heal-once shared with Esc path). No core action.
- **✅ RESUMED (14:58 PDT 2026-07-28):** stop was pure session-limit, no workflow error. Harvest complete: 226 chunk files from 4 run-dir journals, verified 25,410/177,906 pairs (14.28%), NS 152,496. Gates: 178 OK / 10 SHORT / 38 DEGENERATE (17% — cause under investigation, re-dispatched not trusted). **Operator cadence ruling: one batch per ~5h session window** → ~8 windows remain (~30M tokens), multi-day run. Lia drives a missing-chunk dispatcher off the VERIFIED matrix (not resumeFromRunId) — partial batches merge cleanly. Doyle riders: degenerate-cause line in next ping; persist matrix+gate ledger to disk each slot; rescore attribution in final report.

## Slot progress
- Slot 1 (15:4x 07-28): 150/150 dense, 161 agents, 0 errors, 5.98M tok, 40min — single-workflow pacing fits a 5h window with headroom. Coverage 52,230/177,906 = 29.36%. ~6 slots remain.
- ⚖ Gate-3 REFINEMENT RATIFIED (doyle): original degenerate test conflated mid-pile (lazy scout) with FLOOR-pile (honest shape — most random pairs in a 597-req manifest truly unrelated). Evidence: anchor concordance 0.93 median in both groups + identical tail richness ⇒ the 38 "degenerates" were judging correctly. New Gate 3: fail mid-pile (≥20) or no-tail pile; pass floor-piles with disclosed 'flat-floor' flag (58 carry it). No rescore on this basis; attribution unchanged.
- Lia self-caught: ns-vector truncation left 10 chunks' in-loop density check blind (disk verification backstop caught it; emitter now asserts len match). Per-slot persistence in place (scores/*.jsonl + manifest/todo/next_args from VERIFIED state).
- ⏳ operator lever pending: drop 49 placeholders = −28,028 pairs (~1 slot). Full roster stands until explicit GO.
- Slot 2 checkpoint (15:56 07-28, limit cutoff at 93%): coverage 60,690/177,906 = **34.11%**, gate ledger 424 OK / 10 SHORT / 826 chunks remain; 69 flat-floor flags. Mid-run harvest proven safe (64 chunks recovered live; post-cutoff completions stay in journals for next sweep; run ids in runs.txt). **RESUME.md in her run dir = cold-pickup authority** (4-step resume + full decision record).
- Slot 2 final (19:58 07-28 boundary ping): 90/150 chunks before wall, post-cutoff harvest got ALL 90 (25 on top of the 64 mid-flight) — **session death = pause, never loss, proven end-to-end**. Coverage 64,110/177,906 = **36.04%**, 448 OK / 10 SHORT / 802 chunks remain, 78 flat-floor. Slot 3 dispatched. Burn: window ≈ 90-150 chunks / 3.3-6M tok ⇒ **6-7 slots remain**.
- Slot 4 final (01:46 PDT 07-29 boundary ping): 99/150 chunks before session-limit wall (no workflow fault), post-reset harvest recovered ALL 99. Coverage 97,410/177,906 = **54.75%** (past halfway), NS 80,496, ledger 681 OK / 10 SHORT / 569 chunks remain. Slot 5 dispatched (150 chunks / 22,320 pairs, run wf_e20a0ac3-c3f). SHORT handling ratified: shorts ride the coverage-driven selector automatically (not a cleanup step); 6 chunks short ×2 attempts — any surviving to end-of-coverage get **coordinator-scored per the skill, cells marked coordinator-scored in matrix** (attribution rider applies). ~4 slots left, then Gate 4 cross-batch consolidation.
- Slot 3 final (boundary ping, ~03:0x 07-29): CLEAN 150/150 chunks, 22,320/22,320 pairs dense, 153 agents, 0 errors, 5.63M tok, 33min — full window, no cap hit. Coverage 84,666/177,906 = **47.59%**, 592 OK / 6 SHORT / 658 chunks remain. Slot 4 dispatched (150 chunks / 22,356 pairs, run wf_cc1c2e87-398). Two lia self-caught durability fixes (no scored data affected): (1) hand-maintained RESUME.md state went stale → state now generated by snapshot.py into STATE.md from verified on-disk rows; (2) ⭐ mid-slot snapshot re-queued the SAME in-flight 150 chunks (in-flight reads MISSING on disk; dupes merge silently under max-keeping) → inflight.json written at dispatch, excluded from selection, cleared at harvest.

- Slot 7 final (05:51 07-29): 59/60 chunks, chunk 593 short carried. Coverage 125,814 = **70.72%**. Coordinator-scoring stragglers WORKED: SHORT 10→4→1, chronic ids gone. Slot 8 CLEAN (06:2x ping): 150/150, 22,284 pairs, 158 agents, 0 errors, 5.73M tok, 34min. Coverage 147,198 = **82.74%**, ledger 1,030 OK / ZERO short / ZERO degenerate. FINAL slot dispatched: all 220 remaining chunks / 32,028 pairs, run wf_a0cd57a6-de3 (100% in one go if window holds).
- ⚖ **GATE 4 RULED (doyle, 06:0x 07-29): S1 entry ≥30** (+31% rescaled-up pairs = mechanism working; ±20 drift blinds ≥40 entry; ~2.3M extra tokens accepted vs stamping the 3,795-pair 30-39 band single-scout-provisional). 3 stages approved: S1 all ~8,500 cands ≥30, random 100/slice, ~86 slices; S2 ≥40 (~800 expected); S3 single-authority pass ≥50 → sole ranked-list source. RIDERS: (b) slice slope outside 0.75–1.75 envelope = FLAG+HOLD, never stretch (mechanized in slice_frame.py); (c) S2 output >~2,000 = STOP, distribution to doyle before S3 (written into RESUME.md). Pipeline BUILT idle: build_slices.py, consolidate.js, harvest.py routes to consolidated/<stage>/. Reference deviation doyle-approved: 12 refs = 6 original + 6 new top-weighted (62,55,40,35,15,5), all coordinator judgments from id+title+doc, disclosed as such in report. S1 fires when coverage closes.

## At completion
- Deliverables: coverage line first, 597×597 matrix (NS-distinct) + heatmap, ranked ≥40 list (consolidated attribution), Phase D diagnoses, Phase E proposals. STOP at E — apply/commit need separate operator authorization.
- deployah only enters if a manifest change wants a version.
