# Phase B — confusability scoring of the spt-core requirement registry

**Status: SHIPPED ON RAW SCORES.** Ruled by doyle, 2026-07-31 (option B).
**Deliverable:** `phaseB/scores.jsonl` — 14,397 scored pairs.
**This document is binding on every use of those scores.** The constraints below are not
commentary; they are the conditions under which the numbers mean anything. Do not quote a
figure from this run without the caveat attached to it here.

---

## 1. What was measured, and the two coverage numbers

Every requirement pair **inside the ratified `[[groups]]` partition** was scored for
confusability — how plausibly a tagger could attach evidence to the wrong one of the two.

| | |
|---|---|
| Scored / in-scope | **14,397 / 14,397 = 100.00%** |
| In-scope / total possible | **14,397 / 204,480 = 7.04%** |
| NP — cross-group, never scored **by design** | 190,083 |
| NS — in scope but unscored | **0** |

**Both coverage numbers travel together, always.** Quoting the 100.00% alone would imply the
registry was swept; it was not. 7.04% of the pair space was examined, because the partition
deliberately excludes cross-group pairs. **NP is not zero and is not NS** — a cross-group pair
was never looked at, which is a different claim from "looked at and found unconfusable."

Registry: 640 requirements, 231 groups (34 subject + 197 spanning), manifest `traceable-reqs.toml`
at commit b4af804. Partition provenance and its own recall caveats: `PARTITION-PROPOSAL.md`.

**Standing registry finding, unresolved and not mine to fix:** 361 of 640 requirements (56.4%)
have no doc-stage evidence, so they were scored from **id + title alone**. That is the weakest
input exactly where cross-scope confusion is worst. Owner: the registry owner.

---

## 2. Results — bands, not a ranking

| Band | Pairs |
|---|---|
| ≥70 | **85** |
| 50–69 | **434** |
| 40–49 | **611** |
| 20–39 | 4,071 |
| <20 | 9,196 |

**These are SOFT boundaries, not thresholds** (§3). A pair at 68 and a pair at 71 are not
reliably distinguishable by this instrument. **Nothing in this run is citable as a ranking.**
Band membership is a triage signal for human review, nothing more.

The ≥70 band is listed in Appendix A **alphabetically and unranked**, deliberately, so that the
listing order cannot be mistaken for a ranking. 64 of the 85 carry a scout rationale in
`scores.jsonl`.

---

## 3. Reliability — both numbers, and why they disagree

Two independent measurements of what happens when the same pair is scored twice:

| Source | Sample | Result |
|---|---|---|
| Raw pass duplicate observations | 448 pairs re-scored (3.1% of the set) | 442 of 448 disagree; **mean \|delta\| 14.3**, median 14, max 50 |
| Anchored run, 24 pairs × 23 chunks | largest sample in the audit | **median re-score spread ~40 points**, per-pair stdev median 11.3, max spread 67 |

**They disagree, and the disagreement is not resolved.** The ~40-point figure comes from scouts
working candidate-heavy chunks, so it may measure *those scouts* rather than the raw pass's;
the 448-duplicate **\|delta\| 14.3 remains the closer estimate for the raw scores shipped here.**
Both numbers are stated because picking the flattering one would be dishonest.

Consequences, binding:

- **Band boundaries are SOFT.** ≥40 / ≥50 / ≥70 are triage cuts, not thresholds.
- **The 448-duplicate figure is a FLOOR quoted with its 3.1% denominator**, never a bare rate.
- Of those 448 re-scores, 22 cross ≥40, 4 cross ≥50, 4 cross ≥70. **The 4s sit on a denominator
  too small to support a ratio** — quote them as counts with the denominator, never as a
  percentage.

---

## 4. Scores are per-chunk RAW and NOT cross-chunk comparable

The scoring workflow's consolidation leg — the step that puts candidates on one calibration —
**never successfully ran.** Two attempts, both quarantined (§6). So a 65 assigned by one chunk's
scout and a 65 assigned by another's are **not established to be the same 65.**

This is the single largest limitation on the shipped numbers. Comparisons *within* a chunk are
better supported than comparisons *across* chunks, and the shipped file does not record which
chunk each pair came from as a score qualifier — treat all cross-pair comparisons as soft.

---

## 5. The ≥30 blind spot

Consolidation, had it worked, could only ever have **confirmed or demoted** a pair the raw pass
already placed at ≥30. **A pair scored 29 could never be rescued by it.** Given re-score deltas
of 14.3 (or worse, §3), pairs in the **20–39 band could plausibly belong ≥40** and were never
re-examined. This is a property of the workflow's `consolidateThreshold`, not a choice made in
this audit. **No ranking or coverage claim may be quoted as if the ≥30 cut were free.**

---

## 6. Quarantine register — what exists but must never be used

| Run | Observations | Status |
|---|---|---|
| Unanchored consolidation `wf_06c8fb62-fc4` | 3,894 | **QUARANTINED WHOLESALE** |
| Anchored consolidation `wf_5c00d1d6-703` | 2,969 | **QUARANTINED WHOLESALE** |

**Neither ships. Neither overrides raw. Neither may be partially salvaged.**

Both runs were *mechanically* clean — full coverage, zero invented pairs, perfect alignment. Both
failed on **calibration**:

- **Unanchored:** 17 chunks interleaved from one candidate pool are *exchangeable*, so their true
  medians must be near-identical. Raw per-chunk medians spanned 35–40 (stdev **1.3**). The
  consolidation pass on **the same pairs** spanned 3.0–40.0 (stdev **9.0**) — a 7× inflation of
  the very variance consolidation exists to remove. It produced 24 mutually incompatible scales.
- **Anchored:** failed its **pre-declared** acceptance gate — stdev of per-chunk anchor medians
  **7.14 against a bound of 3.0**, declared and sent to the ruling authority *before launch* and
  computed *before* any score movement was examined.

**No partial salvage, on principle.** In the unanchored run the agents whose medians sat closest
to raw included the worst template-degraded scorer. Selecting agents by whether their answers
agreed with raw is fitting, not measuring — agreement never promotes a result to proven.

---

## 7. Open question — do NOT cite this run as settling it

**Hypothesis:** the consolidation leg fails because candidate-only chunks are homogeneous — the
scout sees no low-score tail, and each scout re-centres its own scale.

**Status: OPEN. Neither confirmed nor refuted.** The anchored run was intended as the test and
**was underpowered by its own construction**: 24 anchors in a 149-pair chunk gave only **6.0%
sub-30 content against the raw pass's 82.6%** — it restored **7.3%** of the low tail the
hypothesis is about. The gate failure therefore condemns *that run*, and does **not** refute the
hypothesis.

**No downstream document may cite the anchored run as a refutation of the homogeneity
hypothesis.**

---

## 8. Possible future work — NOT authorized

A composition-matched re-test of the homogeneity hypothesis. **Not authorized; parked.** Running
it is a fresh, operator-visible spend decision.

Required composition, spelled out so the design is not re-derived under pressure later:

- chunks must reproduce the raw pass's distribution — **~83% sub-30 content**, not ~6%
- anchors identical across every chunk, spanning the full range, interleaved rather than blocked
- an acceptance gate **declared numerically before launch**, computed before any score movement
  is examined, with **one** binding statistic so the verdict cannot be shopped across statistics
- expected size: substantially larger than either attempt here, since each chunk must carry ~5×
  more non-candidate pairs than candidates

---

## 9. Method findings worth keeping (they cost real money to learn)

1. **Key scout returns on ALIGNMENT against the dispatched input, never on count.** 3 of 96
   chunks returned exactly the 150 lines requested while silently dropping pairs and inventing
   others — count-invariant substitution, invisible to any count-based check. 6 dropped, 5
   invented. One agent reported `evaluated: 2582` while returning 20 pairs: **self-reported
   counts are decoration.**
2. **An invented score must never override the dispatched scout's.** Mandate = union of chunks an
   agent demonstrably worked (≥2 pairs); threshold declared, not tuned, and verified insensitive
   (observed counts 1, 1, 10, 10, 141+ — nothing in 2..9, so any threshold in 2..10 is identical).
3. **Above ~4,096 items, harvest from the run journal — never return across the workflow VM
   boundary.** The main scoring run completed all 96 chunks and then threw
   `array length 14397 exceeds the maximum of 4096`; every score was recovered from the journal
   and nothing was re-scored.
4. **A retry ladder cannot tell an account-limit error from a short chunk.** One dispatch burned
   384 agents and 442k tokens re-firing into a hard session limit for zero information. Limit
   errors should be fatal-stop.
5. **Exchangeable chunks make between-chunk variance a calibration instrument.** Interleaving
   globally is what made the consolidation failures measurable at all.

---

## Appendix A — the ≥70 band (85 pairs, alphabetical, UNRANKED)

Band membership is soft (§3) and this listing is deliberately not ordered by score.

| Requirement A | Requirement B | Score |
|---|---|---|
| REQ-ACL-SUBJECT-CHAIN | REQ-ACL-SUBNET-MODE-CAPTURE | 72 |
| REQ-ACL-VIEW-DRILLDOWN | REQ-ACL-VIEW-ROSTER | 75 |
| REQ-ADAPTER-UPDATE-INPLACE | REQ-DAEMON-SERVICE-INSTALL | 75 |
| REQ-ADAPTER-UPDATE-INPLACE | REQ-INSTALL-12 | 77 |
| REQ-ADAPTER-UPDATE-INPLACE | REQ-UPD-2 | 75 |
| REQ-API-1 | REQ-UPDATE-FINISH-COMMUNE-FLUSH | 75 |
| REQ-API-4 | REQ-INSTALL-11 | 91 |
| REQ-API-4 | REQ-READY-AGENT-RESUME | 70 |
| REQ-BIND-PSYCHE-CUSTODY-SQUAT-GUARD | REQ-HAZARD-TEMPLATE-ARGV-FILL | 75 |
| REQ-BRAIN-RESUME-NO-CONN-DEADLOCK | REQ-BRAIN-UPDATE-RESTART-CLEAN-CLOSE | 95 |
| REQ-BROKER-ATTACH-JOURNAL-RESILIENT | REQ-UPD-7 | 75 |
| REQ-CI-FREE-SPACE-PREFLIGHT | REQ-INFRA-1 | 70 |
| REQ-CLI-4 | REQ-EP-2 | 75 |
| REQ-CLI-4 | REQ-HAZARD-DEFERRED-MANIFEST | 75 |
| REQ-CLI-JSON | REQ-DAEMON-5 | 75 |
| REQ-CLI-WIN-VT-ENABLE | REQ-WHOAMI-EXPLICIT-SID-REFUSAL | 75 |
| REQ-CONN-BLACKHOLE-LIFECYCLE-HARNESS | REQ-INJECT-MULTILINE-INTEGRITY | 75 |
| REQ-CREATE-BIND-REST-ACTIVE | REQ-SELF-ID-TRUST-INJECTED-ENV | 70 |
| REQ-DAEMON-6 | REQ-ENDPOINT-MESSAGE-ONLY-DISPLAY | 75 |
| REQ-DAEMON-7 | REQ-HAZARD-DAEMON-STOP-REAP | 75 |
| REQ-DAEMON-BITS-AMBIGUITY | REQ-HAZARD-TRANSLATE-FAULT-PERMANENT-DEATH | 83 |
| REQ-DAEMON-SERVICE-INSTALL | REQ-INST-13 | 75 |
| REQ-DOCS-5 | REQ-EP-2 | 75 |
| REQ-ECHO-DROP-DIR-RESOLVE | REQ-HAZARD-WORKER-PATH | 75 |
| REQ-EFFECTIVE-INSTANCE-STATE | REQ-ENDPOINT-TEARDOWN-AUTHORITY | 75 |
| REQ-ENDPOINT-LIST-MERGE-LOCAL | REQ-ENDPOINT-LIST-NODE-GROUPED | 91 |
| REQ-ENDPOINT-LIST-PALETTE | REQ-EP-4 | 75 |
| REQ-ENDPOINT-LIST-PROJECT-COL | REQ-PICKER-5 | 75 |
| REQ-ENDPOINT-LIST-PROJECT-COL | REQ-WORKER-PICKER-EXCLUDED | 70 |
| REQ-ENDPOINT-MESSAGE-ONLY-DISPLAY | REQ-SUBNET-COUNT-ROUTABLE | 75 |
| REQ-ENDPOINT-STOP-OFFLINE | REQ-SELF-DETECT-PARENT-PID | 75 |
| REQ-ER-BRINGUP-ATTEMPT-BOUND | REQ-ER-RC-INTENT-LOCKS | 75 |
| REQ-FRONT-1 | REQ-HAZARD-DAEMON-HOSTED-LIVENESS | 75 |
| REQ-GOSSIP-PROJECT-DERIVE-ONCE | REQ-HAZARD-MESH-BOOTSTRAP-TRAP | 70 |
| REQ-HAZARD-ATTACH-WEDGE | REQ-SHELL-LIST-DERIVED-PROVENANCE | 75 |
| REQ-HAZARD-BOUNDARY-READY-STRAND | REQ-UPDATE-PROMOTE-DRAINED | 75 |
| REQ-HAZARD-BROKER-FLOOR-LOCK-POISON | REQ-HAZARD-INFO-RMW-LOST-UPDATE | 82 |
| REQ-HAZARD-CEREMONY-CLOCK-STEP | REQ-HAZARD-PAIR-TRANSCRIPT-BIND | 75 |
| REQ-HAZARD-DAEMON-STOP-REAP | REQ-HAZARD-RC-EOF | 80 |
| REQ-HAZARD-DRIVEN-BY-IDLE-REMOTE-EVICT | REQ-HAZARD-STALE-SIGNOFF-SENTINEL | 75 |
| REQ-HAZARD-DRIVEN-BY-SELFHEAL | REQ-RELAY-DEATH-CONVERGENCE | 75 |
| REQ-HAZARD-EBUSY-RENAME | REQ-HAZARD-INFO-JSON-TORN-READ | 75 |
| REQ-HAZARD-HANDOFF-ARGV-COMPAT | REQ-SESSIONS-LOG-ENDPOINT-ATTRIBUTION | 75 |
| REQ-HAZARD-LIVEHOST-NONRESIDENT | REQ-HAZARD-WORKER-PATH | 75 |
| REQ-HAZARD-LIVEHOST-NONRESIDENT | REQ-PUBLIC-ERROR-SURFACES | 70 |
| REQ-HAZARD-PAIR-SEED-ROTATION | REQ-REL-3 | 75 |
| REQ-HAZARD-RC-ATTACH-TRUTH | REQ-HAZARD-RC-EOF | 82 |
| REQ-HAZARD-SINGLE-PATH-SOURCE | REQ-SEAM-HISTORY | 75 |
| REQ-HAZARD-SOFT-CLEANUP | REQ-RC-DISPLAY-SOLE-WRITER | 75 |
| REQ-HAZARD-WMI-DAEMON-WINDOW | REQ-INST-13 | 75 |
| REQ-HAZARD-WMI-DAEMON-WINDOW | REQ-PICKER-ADAPTER-DESCRIPTION | 75 |
| REQ-HOST-RUN-1 | REQ-PSYCHE-CRASHLOOP-BACKOFF-SHUTDOWN | 75 |
| REQ-HOST-RUN-2 | REQ-LIVENESS-ORACLE-SOUND | 70 |
| REQ-INST-10 | REQ-INSTALL-6 | 75 |
| REQ-INST-13 | REQ-INSTALL-1 | 75 |
| REQ-INST-2 | REQ-PUBLIC-ERROR-SURFACES | 75 |
| REQ-INST-5 | REQ-INSTALL-6 | 75 |
| REQ-INSTALL-13 | REQ-INSTALL-9 | 75 |
| REQ-INSTALL-7 | REQ-UPDATE-FETCH-CURRENT-UX | 75 |
| REQ-MANIFEST-6 | REQ-RUN-ID-REUSES-ADAPTER | 86 |
| REQ-MESH-1 | REQ-WAN-SEND-DELIVERY | 75 |
| REQ-MESH-6 | REQ-PICKER-NODE-GROUPING | 75 |
| REQ-MSG-2 | REQ-SHELL-3 | 75 |
| REQ-ONEWAY-STREAM-TERMINAL | REQ-REGISTRY-APPLY-TRANSACTIONAL | 70 |
| REQ-OPID-MINTER-NAMESPACE | REQ-PSYCHE-NESTED-RESOLUTION | 70 |
| REQ-PAIR-3 | REQ-SUBNET-7 | 75 |
| REQ-PEERADDR-INVARIANT | REQ-PUMP-STAGE-TRUTH | 70 |
| REQ-PICKER-1 | REQ-PICKER-NODE-GROUPING | 75 |
| REQ-PICKER-5 | REQ-PICKER-CHOOSE-DEDUP-ALL | 75 |
| REQ-PICKER-CHOOSE-DEDUP-ALL | REQ-PICKER-PURGE-SHORTCUT | 75 |
| REQ-PICKER-OFFLINE-NO-VIEW | REQ-PICKER-ONLINE-ACTION | 75 |
| REQ-PICKER-RESUME-CONTEXT-PANEL | REQ-RESUME-ROW-PER-PROJECT | 75 |
| REQ-PSYCHE-EPHEMERAL-DRIVER | REQ-PSYCHE-LEGACY-RESIDENT-SWEEP | 75 |
| REQ-PSYCHE-LEGACY-RESIDENT-SWEEP | REQ-SPAWN-COLLISION-GUARD-LIVE-DUP | 75 |
| REQ-RC-DISPLAY-SOLE-WRITER | REQ-RUN-EMPTY-CREATE | 75 |
| REQ-RC-HARNESS-ONLY-REFUSAL | REQ-RESUME-UNBOUND-STAMP | 70 |
| REQ-RC-HARNESS-ONLY-REFUSAL | REQ-WHOAMI-EXPLICIT-SID-REFUSAL | 74 |
| REQ-RC-SINGLE-PUMP-BRAIN | REQ-TERM-3 | 75 |
| REQ-RESIZE-INPUT-MODE-INTEGRITY | REQ-SCREENGRID-REPAINT-MODE-REPLAY | 70 |
| REQ-RESUME-HARNESS-SESSION-ID | REQ-SELF-ID-TRUST-INJECTED-ENV | 75 |
| REQ-SEAM-UPDATE | REQ-UPD-2 | 75 |
| REQ-SELF-ID-TRUST-INJECTED-ENV | REQ-START-3 | 75 |
| REQ-SHELL-3 | REQ-WHOAMI-EXPLICIT-SID-REFUSAL | 75 |
| REQ-UPD-5 | REQ-UPD-8 | 75 |
| REQ-UPDATE-ADAPTERS-VERB | REQ-UPDATE-RESTART-SAFE-SWAP | 70 |

---

## Reproduction

All paths relative to the audit run directory; nothing in this audit was written into the
spt-core repository.

```
python harvest_phaseb.py            # rebuild scores.jsonl from the run journals (md5-stable)
python align_phaseb.py              # per-chunk alignment audit -> alignment.json
python gate_anchored.py             # re-evaluate the pre-declared anchored gate
python verify_phaseb_inputs.py      # input integrity (NOT `md5sum -c` — see that file's docstring)
python check_baseline.py            # frozen-tree census (covers consolidated/ and slices/ only)
```

Artifacts: `phaseB/scores.jsonl`, `coverage.json`, `gaps.json`, `alignment.json`,
`anchor-set.json`, `anchor-gate.json`, `candidates-raw.jsonl`, `PHASEB-INPUT-MD5.txt`.
