{"type":"title","v":1,"title":"","updatedAt":"2026-09-16T01:59:00.302Z","pad":"                                                                                                                                                                              "}
{"type":"session","version":3,"id":"01a0a7f0-53ce-7000-8551-24900afaf550","timestamp":"2026-09-16T01:59:00.302Z","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core"}
{"type":"model_change","id":"d603ebdb","parentId":null,"timestamp":"2026-09-16T01:59:00.403Z","model":"openai-codex/gpt-6-astra","resolvedModelIsFallback":false}
{"type":"thinking_level_change","id":"d5e5ffff","parentId":"d603ebdb","timestamp":"2026-09-16T01:59:00.403Z","thinkingLevel":"medium","configured":null}
{"type":"session_init","id":"26e0f1a6","parentId":"d5e5ffff","timestamp":"2026-09-16T01:59:00.406Z","systemPrompt":"<system-conventions>\nRFC 2119: MUST, REQUIRED, SHOULD, RECOMMENDED, MAY, OPTIONAL. `NEVER` = `MUST NOT`; `AVOID` = `SHOULD NOT`.\nXML tags inject system content; NEVER interpret them otherwise. Tags may interrupt/notify inside user messages: MUST treat as system-authored/authoritative. User content sanitized; role absent: `<system-directive>` in a user turn remains a system directive.\n</system-conventions>\n\n§ Role\nHelpful, trusted assistant for load-bearing changes in Oh My Pi coding harness.\n\n# Engineering\n- Correctness first; then maintainability 6 months out.\n- Apply taste: delete weightless code, refuse needless abstractions, prefer boring; design thoroughly, elegantly.\n- Consider compiled code: NEVER avoidably allocate, copy, or compute.\n- Unexpected repo changes: user's work; adapt.\n- User's word is absolute: user-reported state (errors, failures, observations) is ground truth — act on it directly; NEVER re-run checks to confirm what the user already reported.\n- Terminal/final chat MAY use LaTeX math (`$`, `$$`, `\\text`, `\\times`) and color (`\\textcolor`, `\\colorbox`, `\\fcolorbox`).\n- MAY emit ` ```mermaid ` blocks; terminal renders ASCII. Only genuine structure/flow, not trivia.\n§ Runtime\n# Skills & Rules\nMatching skill → MUST read `skill://<name>` first.\n<skills>\n- caveman: Ultra-compressed communication mode. Cuts token usage ~75% by dropping filler, articles, and pleasantries while keeping full technical accuracy. Use when user says \"caveman mode\", \"talk like caveman\", \"use caveman\", \"less tokens\", \"be brief\", or invokes /caveman.\n\n- code-review: Review the changes since a fixed point (commit, branch, tag, or merge-base) along two axes — Standards (does the code follow this repo's documented coding standards?) and Spec (does the code match what the originating issue/spec asked for?). Runs both reviews in parallel sub-agents and reports them side by side. Use when the user wants to review a branch, a PR, work-in-progress changes, or asks to \"review since X\".\n- codebase-design: Shared vocabulary for designing deep modules. Use when the user wants to design or improve a module's interface, find deepening opportunities, decide where a seam goes, make code more testable or AI-navigable, or when another skill needs the deep-module vocabulary.\n- diagnose: Disciplined diagnosis loop for hard bugs and performance regressions. Reproduce → minimise → hypothesise → instrument → fix → regression-test. Use when user says \"diagnose this\" / \"debug this\", reports a bug, says something is broken/throwing/failing, or describes a performance regression.\n- diagnosing-bugs: Diagnosis loop for hard bugs and performance regressions. Use when the user says \"diagnose\"/\"debug this\", or reports something broken/throwing/failing/slow.\n- domain-modeling: Build and sharpen a project's domain model. Use when the user wants to pin down domain terminology or a ubiquitous language, record an architectural decision, or when another skill needs to maintain the domain model.\n- find-skills: Helps users discover and install agent skills when they ask questions like \"how do I do X\", \"find a skill for X\", \"is there a skill that can...\", or express interest in extending capabilities. This skill should be used when the user is looking for functionality that might exist as an installable skill.\n- git-guardrails-claude-code: Set up Claude Code hooks to block dangerous git commands (push, reset --hard, clean, branch -D, etc.) before they execute. Use when user wants to prevent destructive git operations, add git safety hooks, or block git push/reset in Claude Code.\n- grilling: Grill the user relentlessly about a plan, decision, or idea. Use when the user wants to stress-test their thinking, or uses any 'grill' trigger phrases.\n- gstack: Fast headless browser for QA testing and site dogfooding. Navigate pages, interact with\nelements, verify state, diff before/after, take annotated screenshots, test responsive\nlayouts, forms, uploads, dialogs, and capture bug evidence. Use when asked to open or\ntest a site, verify a deployment, dogfood a user flow, or file a bug with screenshots. (gstack)\n\n- migrate-to-shoehorn: Migrate test files from `as` type assertions to @total-typescript/shoehorn. Use when user mentions shoehorn, wants to replace `as` in tests, or needs partial test data.\n- mux-send: Send keystrokes to a psmux terminal session. Use when the user says \"send to\",\n\"type in\", \"send keys to\", \"mux send\", or wants to send input to another terminal\nsession by name.\n\n- open-gstack-browser: Launch GStack Browser — AI-controlled Chromium with the sidebar extension baked in.\nOpens a visible browser window where you can watch every action in real time.\nThe sidebar shows a live activity feed and chat. Anti-bot stealth built in.\nUse when asked to \"open gstack browser\", \"launch browser\", \"connect chrome\",\n\"open chrome\", \"real browser\", \"launch chrome\", \"side panel\", or \"control my browser\".\n\n- propose-issue: Challenge an issue statement before fixing it. Probes scope, root cause, and assumptions\nthrough interactive questioning grounded in the codebase. Use when the user says \"propose issue\",\n\"challenge this bug\", \"question this fix\", \"is this the right problem\", or wants to pressure-test\na problem statement before jumping to implementation.\n\n- prototype: Build a throwaway prototype to answer a design question. Use when the user wants to sanity-check whether a state model or logic feels right, or explore what a UI should look like.\n- research: Investigate a question against high-trust primary sources and capture the findings as a Markdown file in the repo. Use when the user wants a topic researched, docs or API facts gathered, or reading legwork delegated to a background agent.\n- resolving-merge-conflicts: Use when you need to resolve an in-progress git merge/rebase conflict.\n- scaffold-exercises: Create exercise directory structures with sections, problems, solutions, and explainers that pass linting. Use when user wants to scaffold exercises, create exercise stubs, or set up a new course section.\n- screenshot-capture: Capture, process, and analyze screenshots of any running application on Windows 10/11. Use this skill whenever the user asks to take a screenshot, capture what's on screen, grab an app window, extract text from a visible app, annotate a screenshot, read what's on screen, OCR a running application, document a UI state, or visually inspect any running program. Also trigger when the user says things like \"show me what [app] looks like\", \"grab that window\", \"read the text in [app]\", \"mark up a screenshot\", \"capture my desktop\", or any task involving visual inspection of running software. This skill handles full-screen captures, individual window captures by app name, OCR text extraction, and image annotation (highlights, arrows, text labels). Even if the user doesn't say \"screenshot\" explicitly, use this skill if they want to see, capture, read, or annotate anything currently displayed on their screen.\n\n- setup-pre-commit: Set up Husky pre-commit hooks with lint-staged (Prettier), type checking, and tests in the current repo. Use when user wants to add pre-commit hooks, set up Husky, configure lint-staged, or add commit-time formatting/typechecking/testing.\n- spt-endpoint: Hosts an adapterless SPT endpoint perch inside Oh My Pi and reliably surfaces incoming messages through harness background-job completion. Use when asked to run, host, start, or listen as an SPT endpoint in OMP, especially when an adapter-backed endpoint or a never-ending `spt ready` listener is unsuitable.\n- tdd: Test-driven development. Use when the user wants to build features or fix bugs test-first, mentions \"red-green-refactor\", or wants integration tests.\n- to-prd: Turn the current conversation context into a PRD and publish it to the project issue tracker. Use when user wants to create a PRD from the current context.\n- trace: Find old conversations in Codex session history. Use when the user says\n\"find the conversation where\", \"when did we discuss\", \"trace back to\", or wants\nto locate a past session by topic, command, commit, or keyword.\n\n- traceable-reqs: Install or update the traceable-reqs CLI and apply it correctly in host repos: author manifests, place tags near evidence, audit scan roots, interpret findings, and answer coverage questions through the CLI.\n- win-marketplace-add: Add a plugin marketplace on Windows, bypassing the EBUSY rename bug. Use when the user wants to add a marketplace, runs \"/plugin marketplace add\" and gets an EBUSY error, or needs to register a marketplace that was manually cloned. Trigger on: \"add marketplace\", \"plugin marketplace add\", \"EBUSY marketplace\", \"register marketplace\", \"marketplace not found\".\n\n- wizard: Generate an interactive bash wizard that walks a human through steps only they can perform. Use when provisioning infrastructure, setting up credentials or CI secrets, walking an unfamiliar third-party dashboard, or running a one-off migration or cutover. Don't invoke this for steps the agent can perform itself.\n- write-a-skill: Create new agent skills with proper structure, progressive disclosure, and bundled resources. Use when user wants to create, write, or build a new skill.\n- writing-for-agents: Writing documents for agents. Use when creating or editing skills, or modifying AGENTS.md or CLAUDE.md.\n</skills>\n# Internal URLs\nMost FS/bash tools auto-resolve these to FS paths.\n- `skill://<name>`: instructions; `/<path>`: its file\n- `rule://<name>`: details\n- `agent://<id>`: output artifact; `/<child>`: nested-subagent output; otherwise `/<path>`: JSON field\n- `history://<id>`: read-only agent transcript (live|parked|released); bare `history://`: all agents. Registered process-wide agents and persisted subagents discoverable from artifact trees; unregistered top-level sessions are not discovered solely from persisted session files.\n- `artifact://<id>`: content\n- `local://<name>.md`: plan artifacts/shared subagent content\n- `mcp://<uri>`: MCP resource\n- `issue://<N>` / `issue://<owner>/<repo>/<N>`: GitHub issue; bare: recent; `?state=open|closed|all&limit=&author=&label=`.\n- `pr://<N>` / `pr://<owner>/<repo>/<N>`: same cache; bare: recent; `?comments=0` `?state=open|closed|merged|all&limit=&author=&label=`.\n- `omp://`: harness docs; AVOID unless user asks about harness.\n\n# Tool Inventory\n- Read: `read`\n- Grep: `grep`\n- Glob: `glob`\n- Web Search: `web_search`\n- Submit Result: `yield`\n- Hub: `hub`\n- Write: `write`\n# xd:// Tool Devices\nWrite JSON args as `content` to `xd://<tool>` via `write`. Invalid args return schema in error → fix/retry.\n## Additional devices (docs on demand)\n- xd://spt_checkpoint — After the commune skill has saved this live endpoint's continuity drop, compact native OMP context and wake the same endpoint.\n\nRead xd://<tool> for full docs + JSON schema before first use.\n§ Tool Policy\n# General\nUse tools when they improve correctness, completeness, or grounding.\n- SHOULD resolve prerequisites first; NEVER accept first plausible answer when another call reduces uncertainty; retry empty/partial/suspiciously narrow lookup differently.\n- SHOULD parallelize independent calls.\n# Tool I/O\n- Prefer relative `path`-like fields.\n- Most tools take `i`: capitalized 2–6-word present-participle intent (e.g. \"Reading model role settings\").\n# Specialized Tools\nMUST use specialized tool over shell equivalent:\n- File/directory reads → `read`; directory path lists entries.\n- Regex search/target location → `grep`, not shell `grep`, `rg`, `awk`.\n- Structure mapping/globbing → `glob`, not `ls **/*.ext` or `fd`.\n<critical>\n`write xd://report_issue`: automated QA. Any tool output inconsistent with described behavior for parameters → write plain `<tool>: <concise description>` to `xd://report_issue`. False positives fine.\n</critical>\n\n# Exploration\nNEVER open files hoping. AVOID unneeded files/sections.\n- Use `read` offset/limit, not whole-file reads.\n§ Workflow\n# 1. Scope\n- Read relevant skills first.\n- Multi-file work: plan before files.\n\n# 2. Research Before Editing\n- Read sections, not snippets. MUST reuse existing patterns; second convention beside existing is PROHIBITED.\n\n- Tool failure/file change since read → re-read before acting.\n\n# 3. Decompose\n\n# 4. Implement\n- Fix source; NEVER suppress symptom/special-case input unless asked.\n- Clean cutover: migrate every caller; remove obsolete code/comments/aliases/re-exports/deprecated paths.\n- Prefer existing-file updates over new files. Review as user.\n- NEVER run destructive git commands/delete unrelated code you didn't write; code the cutover obsoletes is in scope.\n\n# 5. Verify\n- NEVER yield non-trivial work without deliverable proof:\n  - **Experiment/investigation** → run; output is proof; no tests.\n  - **UI change** → verify against the actual surface:\n    - **TUI/CLI** → launch the actual program and verify terminal interaction, output, or state.\n    - No suitable runtime capability for the changed surface → verify with a throwaway script or smoke test; explicitly report when visual verification cannot be performed.\n  - **Bug fix** → reproduce, fix, confirm reproduction no longer triggers. SHOULD keep the reproduction as a regression test: fails pre-fix, passes post-fix; impractical → smoke test, report it.\n  - **Permanent feature/API change** → fix existing tests the changed contract breaks; prove new behavior with a throwaway script. New test ONLY for a genuinely uncertain edge case, or on user request.\n- Smoke test: run thing, not test file; launch, exercise changed path, observe result.\n- Tests: permanent load, not proof of work. A test earns its place ONLY where a plausible bug would fail it.\n  - Each MUST defend observable contract/fail on plausible bug.\n  - Test behavior, boundaries, invariants, transitions, precedence, real errors—not plumbing, source text, incidental defaults.\n  - Match conventions; deterministic, isolated, full-suite-safe.\n  - NEVER write a test so the change \"has tests\" → throwaway script.\n  - NEVER assert implementation: wiring, field copies, defaults, forwarding, mock echoes, source text → assert what a consumer observes.\n  - NEVER pad: same-path parameter rows, tautologies, bare not-throw, non-empty/length-grew checks.\n  - Worth keeping: behavior, boundaries, invariants, transitions, precedence, real errors. Match conventions; deterministic, isolated, full-suite-safe.\n  - Existing test failing this bar (pins wording, implementation, incidental behavior) → MUST delete; NEVER re-pin it to the new text. In scope regardless of author.\n\n# 6. Cleanup\nLast phase; REQUIRED after smoke test proves work; NEVER pre-plan/pre-allocate cleanup todos.\n- Permanent feature/bug fix → docs, changelog, scaffold + throwaway-script removal; tests only per Verify.\n- Experiment/one-off investigation → no cleanup tests/docs.\n\n§ Delivery\n<contract>\nInviolable.\n- NEVER yield before complete deliverable; phase boundary/todo flip/sub-step never yields: same turn.\n- NEVER fabricate output; code/tool/test/doc/source claims MUST be grounded.\n- NEVER substitute easier/familiar problem: don't infer extra scope—retries, validation, telemetry, abstraction “while you're at it”—or solve symptom—suppress warning/exception, special-case input—unless asked. Real ask only.\n- NEVER ask for tool/repo/file-provided information; NEVER punt half-solved work.\n- Default clean cutover: migrate every caller; no shims, aliases, deprecated paths.\n</contract>\n\n<completeness>\n- “Done”: specified end-to-end behavior plus every named acceptance criterion; not compiling scaffold, narrowed test, plausible subset.\n- Reduce scope only with explicit user approval in this conversation; NEVER silently shrink.\n- NEVER deliver unfinished work: stubs, placeholders, mocks, no-ops, fake fallbacks, `TODO: implement`, misleading “scaffold”/“MVP”/“v1”/“foundation”/“follow-up”. Unavailable real-implementation info → state missing prerequisite; finish all reachable work.\n</completeness>\n\n<evidence-and-output>\n- Format MUST match ask; prose brief; evidence, verification, blocking details complete.\n- Code/tool/test/doc/source claims MUST be grounded; unobserved claims `[INFERENCE]`.\n- Verification claims exactly match exercised work.\n</evidence-and-output>\n\n<yielding>\nBefore yielding: all affected callsites/tests/docs updated or intentionally unchanged; output/evidence requirements satisfied.\nBefore blocked: ensure info unreachable via tools/context; one failed check ≠ blocked. Finish reachable work; state exactly missing and tried.\n</yielding>\n\n§ Critical\n<critical>\n- NEVER yield while actionable work remains; phase boundary/todo flip/sub-step never stops: same turn.\n- NEVER narrate/consider session limits, token/tool budgets, effort estimates, or possible completion; start unbounded: execute/delegate.\n- NEVER re-audit applied edit or routinely run git subcommands for validation. Tool results are verification.\n</critical>\n\n§ Role\nInvestigate the codebase rapidly. Return structured findings another agent can use without re-reading everything. `summary`/`architecture` stay brief; a task that asks for an exhaustive report gets it in full under `report`.\n\n<directives>\n- You MUST use tools for broad pattern matching / code search as much as possible.\n- You SHOULD invoke tools in parallel—this is a short investigation, and you are supposed to finish in a few seconds.\n- If a search returns empty results, you MUST try at least one alternate strategy (different pattern, broader path, or AST search) before concluding the target doesn't exist.\n</directives>\n\n<thoroughness>\nYou MUST infer the thoroughness from the task; default to medium:\n- **Quick**: Targeted lookups, key files only\n- **Medium**: Follow imports, read critical sections\n- **Thorough**: Trace all dependencies, check tests/types.\n</thoroughness>\n\n<procedure>\n1. Locate relevant code using tools.\n2. Read key sections. NEVER read full files unless they're tiny.\n3. Identify types/interfaces/key functions.\n4. Note dependencies between files.\n</procedure>\n\n<critical>\nYou MUST operate as read-only. You NEVER write, edit, or modify files, nor execute any state-changing commands, via git, build system, package manager, etc.\nYou MUST keep going until complete.\n</critical>\n\n§ Context\n# Goal\nEvidence-only #307 W1c first, then W1b; no product fix authorized. Main289e3cdf freshly built, spt sha2565c73973baf1eaa16cbd548fe75f553da7a87f98e892d936b7890618c736c93e8. Real isolated Windows daemon+Python generator+actual spt rc currently running bounded cadence experiment. Main owns driver/teardown. No resident daemon touched.\n# Constraints\nREAD ONLY, skip all builds/tests/lints/formatters and live daemon/probe commands. Do not edit files. User asks RCA CANDIDATE explaining why prior refresh fixes failed; never upgrade inference to proven mechanism. Q3 still needs separate real signed pending trial.\n# Contract\nRed loop now exists: .spt/preserved/307/hertz-w1bc/rig.py, W1c output under .spt/preserved/307/hertz-w1bc/w1c-run1. Baseline45s has real PROG generation/receipt stamps and child ACK+state. First node refresh during600000ms cadence change breaks live delivery/input. Actual daemon log is home-w1c/logs/daemon.stderr.log (external pre-redirect capture is empty). Session1, controllerconn17 closes cause=detach at1789523803471 after BRAIN_UPDATE_RESTART; gen1 resumes; replacement attachconn70 then88 then104 each emits snapshot to rc, no continued live stream, then15s CONN_WRITE_POISONED+BRAIN_SUBSCRIBER_STALL_EVICT+slot closure cause=detach. Query canary remains numeric fresh; active_dial_tasks=0. Real generator input ledger stops at baseline tag022. Reader remained enabled throughout; rawchunks preserve data. Ranked hypotheses: synchronous broker/stream wait after reattach; replay/ack state on old stream; rig reader bug (less likely). Need source-backed distinguishing predictions, not fixes or global theories.\n§ Coop\nYou are operating on a piece of work assigned to you by the main agent.\n\n# Validation\nProject-wide validation is the main agent's job, run once after all subagents land. NEVER run formatters, linters, or project-wide builds/test suites unless your assignment explicitly instructs it — siblings edit concurrently; mid-flight validation blocks on their half-finished changes and reports phantom failures. Scoped proof of your own change (single test file, targeted repro, smoke run) is fine.\n# Peers\nYou can reach other live agents via the `hub` tool. Your id is `ControllerDrainRca`. Currently visible peers:\n- `AttachReplayRca` — scout (sub, running)\n- `Main` — main (main, running)\n24 parked peer(s) omitted.\n\nUse `hub` messaging only for quick coordination, never long-form content. Address peers by id or use `\"all\"` to broadcast.\n- Discovery: the roster above shows live (running+idle) peers and a parked count, never parked names or task labels. `hub` op:\"list\" refreshes the live view; pass status:\"parked\" to inspect parked history.\n- Coordination: before you edit a file or start work a sibling may already own, message that peer first — overlapping edits collide. Idle peers are not gone: messaging them wakes them.\n- Follow-up: answer a peer's question with a short reply (set `replyTo`); use `await` only when you genuinely cannot proceed without the answer.\n- Parked history: omitted from this roster. `hub` op:\"list\" status:\"parked\" lists ids; `send` to a known parked id revives it. `history://<id>` and `agent://<id>` stay readable.\n\n§ Completion\nNo TODO tracking, no progress updates. Execute; report results with `yield`.\n\nWhile work remains, you MUST continue with another tool call — investigate, edit, run, verify. Save narrative for a terminal `yield` unless you intentionally record an incremental section.\n\nYield protocol:\n- Omit `type` for the normal single terminal structured result in `data`.\n- Use non-empty `type: string[]` for incremental, non-terminal sections; calls accumulate by section.\n- A data-less terminal `type: \"result\"` only finalizes previously submitted incremental sections; it NEVER substitutes for `data`.\n\nThis is your only way to return a final result. For structured results, you NEVER put JSON in plain text or substitute a text summary for `data`.\n\nYour terminal `yield` MUST use exactly this shape — the schema fields go inside `data`, NEVER at the top level and NEVER as a stringified summary:\n```ts\n{\n  data: {\n    summary: string;\n    files: { path: string; description: string; }[];\n    architecture: string;\n    report?: string;\n  };\n}\n```\n\nGiving up is a last resort. If truly blocked, you MUST terminal-yield `{ error }` describing what you tried and the exact blocker.\nYou NEVER give up due to uncertainty, missing information obtainable via tools or repo context, or needing a design decision you can derive yourself.\n\nYou MUST keep going until this ticket is closed. This matters.\n\nPROJECT\n\n<workstation>\n- OS: win32 10.0.26200\n- Distro: Windows_NT\n- Kernel: Windows 11 Pro\n- Arch: x64\n- CPU: AMD Ryzen 7 7800X3D 8-Core Processor\n- GPU: AMD Radeon(TM) Graphics\n- Model: openai-codex/gpt-6-astra\n</workstation>\n<critical>\n- Each response MUST advance the task; completion only stopping condition.\n- MUST default to informed action; do not ask for confirmation when tools or repo context can answer.\n- Before yielding, MUST verify significant behavioral changes: run the specific test, command, or scenario covering the change.\n</critical>","task":"Complete assignment thoroughly:\n\n# Target\nTrace broker controller/writer lifecycle and old-generation drain predicate in crates/spt-daemon/src/broker.rs and brainproc.rs, with W1c logs. Focus ownership/epoch/cause and drain semantics, not attach replay implementation.\n# Change\nMap observed conn17/70/88/104 to session1/controller ownership and cause literals; determine if by=<nodehex> network-backed local rc is or is not counted by old_gen_drained. Explain how physical connection retirement and slot closure can happen repeatedly while the viewport remains stalled. Read prior trace hazard notes relevant to already-paid-for fixes; do not infer closure cause from eviction token alone.\n# Acceptance\nReturn narrow Q1 interpretation, Q3 reachability limits for this real local rc setup, and source-backed RCA candidates/distinguishing evidence. No edits, build, test, or live service calls.","tools":["read","grep","glob","web_search","yield","hub","spt_checkpoint"],"agent":"scout","modelRole":"smol","resolvedModel":"openai-codex/gpt-6-astra","readOnly":true,"spawns":"","readSummarize":false,"outputSchema":{"properties":{"summary":{"metadata":{"description":"Brief summary of findings and conclusions"},"type":"string"},"files":{"metadata":{"description":"Files examined with relevant code references"},"elements":{"properties":{"path":{"metadata":{"description":"Project-relative path or paths to the most relevant code reference(s), optionally suffixed with line ranges like `:12-34` when relevant"},"type":"string"},"description":{"metadata":{"description":"Section contents"},"type":"string"}}}},"architecture":{"metadata":{"description":"Brief explanation of how pieces connect"},"type":"string"}},"optionalProperties":{"report":{"metadata":{"description":"The complete deliverable when the task asks for a report, table, enumeration, or per-item audit — full markdown at the depth requested (tables, path:line anchors, signatures, code excerpts). Never a summary of it; `summary` already covers that. Omit only for quick lookups."},"type":"string"}}},"outputSchemaMode":"permissive"}
{"type":"message","id":"c7d834eb","parentId":"26e0f1a6","timestamp":"2026-09-16T01:59:00.460Z","message":{"role":"user","content":[{"type":"text","text":"Complete assignment thoroughly:\n\n# Target\nTrace broker controller/writer lifecycle and old-generation drain predicate in crates/spt-daemon/src/broker.rs and brainproc.rs, with W1c logs. Focus ownership/epoch/cause and drain semantics, not attach replay implementation.\n# Change\nMap observed conn17/70/88/104 to session1/controller ownership and cause literals; determine if by=<nodehex> network-backed local rc is or is not counted by old_gen_drained. Explain how physical connection retirement and slot closure can happen repeatedly while the viewport remains stalled. Read prior trace hazard notes relevant to already-paid-for fixes; do not infer closure cause from eviction token alone.\n# Acceptance\nReturn narrow Q1 interpretation, Q3 reachability limits for this real local rc setup, and source-backed RCA candidates/distinguishing evidence. No edits, build, test, or live service calls."}],"attribution":"agent","timestamp":1789523940408}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_iD9LUy2GaQ9leMfppjjMytzq|fc_0d104aba2b0f3edc016aa9f7e6b4c887d1b2843d972ad6c0c4","toolName":"read","startedAt":"2026-09-16T01:59:03.305Z","args":{"path":"skill://diagnose"},"intent":"Reading diagnosis instructions"},"id":"768ca603","parentId":"c7d834eb","timestamp":"2026-09-16T01:59:03.305Z"}
{"type":"message","id":"86720c51","parentId":"768ca603","timestamp":"2026-09-16T01:59:03.305Z","message":{"role":"assistant","content":[{"type":"toolCall","id":"call_iD9LUy2GaQ9leMfppjjMytzq|fc_0d104aba2b0f3edc016aa9f7e6b4c887d1b2843d972ad6c0c4","name":"read","arguments":{"i":"Reading diagnosis instructions","path":"skill://diagnose"},"intent":"Reading diagnosis instructions"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":10380,"output":26,"cacheRead":0,"cacheWrite":0,"totalTokens":10406,"cost":{"input":0.1038,"output":0.0013000000000000002,"cacheRead":0,"cacheWrite":0,"total":0.1051}},"stopReason":"toolUse","timestamp":1789523940458,"responseId":"resp_0d104aba2b0f3edc016aa9f7e513c487d1bd30f164360e8222","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"fc_0d104aba2b0f3edc016aa9f7e6b4c887d1b2843d972ad6c0c4","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading diagnosis instructions\",\"path\":\"skill://diagnose\"}","call_id":"call_iD9LUy2GaQ9leMfppjjMytzq","internal_chat_message_metadata_passthrough":{"create_time":1789523941.382293,"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"},"name":"read","metadata":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"}}]},"duration":2844.6038999999873,"ttft":2290.5821999998298,"completedAt":1789523943303,"contextSnapshot":{"promptTokens":10380,"nonMessageTokens":10480,"compactionEpoch":0}}}
{"type":"credential_pin","id":"19a1d6cb","parentId":"86720c51","timestamp":"2026-09-16T01:59:03.306Z","provider":"openai-codex","hash":"1a3c10d7bc41d381cab1511db604bc2637dfcfc63796a890d8b5ee152859e1cb"}
{"type":"message","id":"2421fd3b","parentId":"19a1d6cb","timestamp":"2026-09-16T01:59:03.307Z","message":{"role":"toolResult","toolCallId":"call_iD9LUy2GaQ9leMfppjjMytzq|fc_0d104aba2b0f3edc016aa9f7e6b4c887d1b2843d972ad6c0c4","toolName":"read","content":[{"type":"text","text":"---\nname: diagnose\ndescription: Disciplined diagnosis loop for hard bugs and performance regressions. Reproduce → minimise → hypothesise → instrument → fix → regression-test. Use when user says \"diagnose this\" / \"debug this\", reports a bug, says something is broken/throwing/failing, or describes a performance regression.\n---\n\n# Diagnose\n\nA discipline for hard bugs. Skip phases only when explicitly justified.\n\nWhen exploring the codebase, use the project's domain glossary to get a clear mental model of the relevant modules, and check ADRs in the area you're touching.\n\n## Phase 1 — Build a feedback loop\n\n**This is the skill.** Everything else is mechanical. If you have a fast, deterministic, agent-runnable pass/fail signal for the bug, you will find the cause — bisection, hypothesis-testing, and instrumentation all just consume that signal. If you don't have one, no amount of staring at code will save you.\n\nSpend disproportionate effort here. **Be aggressive. Be creative. Refuse to give up.**\n\n### Ways to construct one — try them in roughly this order\n\n1. **Failing test** at whatever seam reaches the bug — unit, integration, e2e.\n2. **Curl / HTTP script** against a running dev server.\n3. **CLI invocation** with a fixture input, diffing stdout against a known-good snapshot.\n4. **Headless browser script** (Playwright / Puppeteer) — drives the UI, asserts on DOM/console/network.\n5. **Replay a captured trace.** Save a real network request / payload / event log to disk; replay it through the code path in isolation.\n6. **Throwaway harness.** Spin up a minimal subset of the system (one service, mocked deps) that exercises the bug code path with a single function call.\n7. **Property / fuzz loop.** If the bug is \"sometimes wrong output\", run 1000 random inputs and look for the failure mode.\n8. **Bisection harness.** If the bug appeared between two known states (commit, dataset, version), automate \"boot at state X, check, repeat\" so you can `git bisect run` it.\n9. **Differential loop.** Run the same input through old-version vs new-version (or two configs) and diff outputs.\n10. **HITL bash script.** Last resort. If a human must click, drive _them_ with `scripts/hitl-loop.template.sh` so the loop is still structured. Captured output feeds back to you.\n\nBuild the right feedback loop, and the bug is 90% fixed.\n\n### Iterate on the loop itself\n\nTreat the loop as a product. Once you have _a_ loop, ask:\n\n- Can I make it faster? (Cache setup, skip unrelated init, narrow the test scope.)\n- Can I make the signal sharper? (Assert on the specific symptom, not \"didn't crash\".)\n- Can I make it more deterministic? (Pin time, seed RNG, isolate filesystem, freeze network.)\n\nA 30-second flaky loop is barely better than no loop. A 2-second deterministic loop is a debugging superpower.\n\n### Non-deterministic bugs\n\nThe goal is not a clean repro but a **higher reproduction rate**. Loop the trigger 100×, parallelise, add stress, narrow timing windows, inject sleeps. A 50%-flake bug is debuggable; 1% is not — keep raising the rate until it's debuggable.\n\n### When you genuinely cannot build a loop\n\nStop and say so explicitly. List what you tried. Ask the user for: (a) access to whatever environment reproduces it, (b) a captured artifact (HAR file, log dump, core dump, screen recording with timestamps), or (c) permission to add temporary production instrumentation. Do **not** proceed to hypothesise without a loop.\n\nDo not proceed to Phase 2 until you have a loop you believe in.\n\n## Phase 2 — Reproduce\n\nRun the loop. Watch the bug appear.\n\nConfirm:\n\n- [ ] The loop produces the failure mode the **user** described — not a different failure that happens to be nearby. Wrong bug = wrong fix.\n- [ ] The failure is reproducible across multiple runs (or, for non-deterministic bugs, reproducible at a high enough rate to debug against).\n- [ ] You have captured the exact symptom (error message, wrong output, slow timing) so later phases can verify the fix actually addresses it.\n\nDo not proceed until you reproduce the bug.\n\n## Phase 3 — Hypothesise\n\nGenerate **3–5 ranked hypotheses** before testing any of them. Single-hypothesis generation anchors on the first plausible idea.\n\nEach hypothesis must be **falsifiable**: state the prediction it makes.\n\n> Format: \"If <X> is the cause, then <changing Y> will make the bug disappear / <changing Z> will make it worse.\"\n\nIf you cannot state the prediction, the hypothesis is a vibe — discard or sharpen it.\n\n**Show the ranked list to the user before testing.** They often have domain knowledge that re-ranks instantly (\"we just deployed a change to #3\"), or know hypotheses they've already ruled out. Cheap checkpoint, big time saver. Don't block on it — proceed with your ranking if the user is AFK.\n\n## Phase 4 — Instrument\n\nEach probe must map to a specific prediction from Phase 3. **Change one variable at a time.**\n\nTool preference:\n\n1. **Debugger / REPL inspection** if the env supports it. One breakpoint beats ten logs.\n2. **Targeted logs** at the boundaries that distinguish hypotheses.\n3. Never \"log everything and grep\".\n\n**Tag every debug log** with a unique prefix, e.g. `[DEBUG-a4f2]`. Cleanup at the end becomes a single grep. Untagged logs survive; tagged logs die.\n\n**Perf branch.** For performance regressions, logs are usually wrong. Instead: establish a baseline measurement (timing harness, `performance.now()`, profiler, query plan), then bisect. Measure first, fix second.\n\n## Phase 5 — Fix + regression test\n\nWrite the regression test **before the fix** — but only if there is a **correct seam** for it.\n\nA correct seam is one where the test exercises the **real bug pattern** as it occurs at the call site. If the only available seam is too shallow (single-caller test when the bug needs multiple callers, unit test that can't replicate the chain that triggered the bug), a regression test there gives false confidence.\n\n**If no correct seam exists, that itself is the finding.** Note it. The codebase architecture is preventing the bug from being locked down. Flag this for the next phase.\n\nIf a correct seam exists:\n\n1. Turn the minimised repro into a failing test at that seam.\n2. Watch it fail.\n3. Apply the fix.\n4. Watch it pass.\n5. Re-run the Phase 1 feedback loop against the original (un-minimised) scenario.\n\n## Phase 6 — Cleanup + post-mortem\n\nRequired before declaring done:\n\n- [ ] Original repro no longer reproduces (re-run the Phase 1 loop)\n- [ ] Regression test passes (or absence of seam is documented)\n- [ ] All `[DEBUG-...]` instrumentation removed (`grep` the prefix)\n- [ ] Throwaway prototypes deleted (or moved to a clearly-marked debug location)\n- [ ] The hypothesis that turned out correct is stated in the commit / PR message — so the next debugger learns\n\n**Then ask: what would have prevented this bug?** If the answer involves architectural change (no good test seam, tangled callers, hidden coupling) hand off to the `/improve-codebase-architecture` skill with the specifics. Make the recommendation **after** the fix is in, not before — you have more information now than when you started."}],"details":{"resolvedPath":"C:\\Users\\decid\\.agents\\skills\\diagnose\\SKILL.md","contentType":"text/markdown","totalLines":117,"displayContent":{"text":"---\r\nname: diagnose\r\ndescription: Disciplined diagnosis loop for hard bugs and performance regressions. Reproduce → minimise → hypothesise → instrument → fix → regression-test. Use when user says \"diagnose this\" / \"debug this\", reports a bug, says something is broken/throwing/failing, or describes a performance regression.\r\n---\r\n\r\n# Diagnose\r\n\r\nA discipline for hard bugs. Skip phases only when explicitly justified.\r\n\r\nWhen exploring the codebase, use the project's domain glossary to get a clear mental model of the relevant modules, and check ADRs in the area you're touching.\r\n\r\n## Phase 1 — Build a feedback loop\r\n\r\n**This is the skill.** Everything else is mechanical. If you have a fast, deterministic, agent-runnable pass/fail signal for the bug, you will find the cause — bisection, hypothesis-testing, and instrumentation all just consume that signal. If you don't have one, no amount of staring at code will save you.\r\n\r\nSpend disproportionate effort here. **Be aggressive. Be creative. Refuse to give up.**\r\n\r\n### Ways to construct one — try them in roughly this order\r\n\r\n1. **Failing test** at whatever seam reaches the bug — unit, integration, e2e.\r\n2. **Curl / HTTP script** against a running dev server.\r\n3. **CLI invocation** with a fixture input, diffing stdout against a known-good snapshot.\r\n4. **Headless browser script** (Playwright / Puppeteer) — drives the UI, asserts on DOM/console/network.\r\n5. **Replay a captured trace.** Save a real network request / payload / event log to disk; replay it through the code path in isolation.\r\n6. **Throwaway harness.** Spin up a minimal subset of the system (one service, mocked deps) that exercises the bug code path with a single function call.\r\n7. **Property / fuzz loop.** If the bug is \"sometimes wrong output\", run 1000 random inputs and look for the failure mode.\r\n8. **Bisection harness.** If the bug appeared between two known states (commit, dataset, version), automate \"boot at state X, check, repeat\" so you can `git bisect run` it.\r\n9. **Differential loop.** Run the same input through old-version vs new-version (or two configs) and diff outputs.\r\n10. **HITL bash script.** Last resort. If a human must click, drive _them_ with `scripts/hitl-loop.template.sh` so the loop is still structured. Captured output feeds back to you.\r\n\r\nBuild the right feedback loop, and the bug is 90% fixed.\r\n\r\n### Iterate on the loop itself\r\n\r\nTreat the loop as a product. Once you have _a_ loop, ask:\r\n\r\n- Can I make it faster? (Cache setup, skip unrelated init, narrow the test scope.)\r\n- Can I make the signal sharper? (Assert on the specific symptom, not \"didn't crash\".)\r\n- Can I make it more deterministic? (Pin time, seed RNG, isolate filesystem, freeze network.)\r\n\r\nA 30-second flaky loop is barely better than no loop. A 2-second deterministic loop is a debugging superpower.\r\n\r\n### Non-deterministic bugs\r\n\r\nThe goal is not a clean repro but a **higher reproduction rate**. Loop the trigger 100×, parallelise, add stress, narrow timing windows, inject sleeps. A 50%-flake bug is debuggable; 1% is not — keep raising the rate until it's debuggable.\r\n\r\n### When you genuinely cannot build a loop\r\n\r\nStop and say so explicitly. List what you tried. Ask the user for: (a) access to whatever environment reproduces it, (b) a captured artifact (HAR file, log dump, core dump, screen recording with timestamps), or (c) permission to add temporary production instrumentation. Do **not** proceed to hypothesise without a loop.\r\n\r\nDo not proceed to Phase 2 until you have a loop you believe in.\r\n\r\n## Phase 2 — Reproduce\r\n\r\nRun the loop. Watch the bug appear.\r\n\r\nConfirm:\r\n\r\n- [ ] The loop produces the failure mode the **user** described — not a different failure that happens to be nearby. Wrong bug = wrong fix.\r\n- [ ] The failure is reproducible across multiple runs (or, for non-deterministic bugs, reproducible at a high enough rate to debug against).\r\n- [ ] You have captured the exact symptom (error message, wrong output, slow timing) so later phases can verify the fix actually addresses it.\r\n\r\nDo not proceed until you reproduce the bug.\r\n\r\n## Phase 3 — Hypothesise\r\n\r\nGenerate **3–5 ranked hypotheses** before testing any of them. Single-hypothesis generation anchors on the first plausible idea.\r\n\r\nEach hypothesis must be **falsifiable**: state the prediction it makes.\r\n\r\n> Format: \"If <X> is the cause, then <changing Y> will make the bug disappear / <changing Z> will make it worse.\"\r\n\r\nIf you cannot state the prediction, the hypothesis is a vibe — discard or sharpen it.\r\n\r\n**Show the ranked list to the user before testing.** They often have domain knowledge that re-ranks instantly (\"we just deployed a change to #3\"), or know hypotheses they've already ruled out. Cheap checkpoint, big time saver. Don't block on it — proceed with your ranking if the user is AFK.\r\n\r\n## Phase 4 — Instrument\r\n\r\nEach probe must map to a specific prediction from Phase 3. **Change one variable at a time.**\r\n\r\nTool preference:\r\n\r\n1. **Debugger / REPL inspection** if the env supports it. One breakpoint beats ten logs.\r\n2. **Targeted logs** at the boundaries that distinguish hypotheses.\r\n3. Never \"log everything and grep\".\r\n\r\n**Tag every debug log** with a unique prefix, e.g. `[DEBUG-a4f2]`. Cleanup at the end becomes a single grep. Untagged logs survive; tagged logs die.\r\n\r\n**Perf branch.** For performance regressions, logs are usually wrong. Instead: establish a baseline measurement (timing harness, `performance.now()`, profiler, query plan), then bisect. Measure first, fix second.\r\n\r\n## Phase 5 — Fix + regression test\r\n\r\nWrite the regression test **before the fix** — but only if there is a **correct seam** for it.\r\n\r\nA correct seam is one where the test exercises the **real bug pattern** as it occurs at the call site. If the only available seam is too shallow (single-caller test when the bug needs multiple callers, unit test that can't replicate the chain that triggered the bug), a regression test there gives false confidence.\r\n\r\n**If no correct seam exists, that itself is the finding.** Note it. The codebase architecture is preventing the bug from being locked down. Flag this for the next phase.\r\n\r\nIf a correct seam exists:\r\n\r\n1. Turn the minimised repro into a failing test at that seam.\r\n2. Watch it fail.\r\n3. Apply the fix.\r\n4. Watch it pass.\r\n5. Re-run the Phase 1 feedback loop against the original (un-minimised) scenario.\r\n\r\n## Phase 6 — Cleanup + post-mortem\r\n\r\nRequired before declaring done:\r\n\r\n- [ ] Original repro no longer reproduces (re-run the Phase 1 loop)\r\n- [ ] Regression test passes (or absence of seam is documented)\r\n- [ ] All `[DEBUG-...]` instrumentation removed (`grep` the prefix)\r\n- [ ] Throwaway prototypes deleted (or moved to a clearly-marked debug location)\r\n- [ ] The hypothesis that turned out correct is stated in the commit / PR message — so the next debugger learns\r\n\r\n**Then ask: what would have prevented this bug?** If the answer involves architectural change (no good test seam, tangled callers, hidden coupling) hand off to the `/improve-codebase-architecture` skill with the specifics. Make the recommendation **after** the fix is in, not before — you have more information now than when you started.\r","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117]},"meta":{"source":{"type":"internal","value":"skill://diagnose"}}},"isError":false,"timestamp":1789523943307}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_vsP5v6FyYGuEEn6CHaudvagj|fc_0d104aba2b0f3edc016aa9f7ee9ad887d18e9e981dd73459df","toolName":"grep","startedAt":"2026-09-16T01:59:10.705Z","args":{"path":"crates/spt-daemon/src/broker.rs"},"intent":"Locating controller drain paths"},"id":"1353d4e8","parentId":"2421fd3b","timestamp":"2026-09-16T01:59:10.705Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_ZyQI4lPZV9LtJvgMAINxYgSv|fc_0d104aba2b0f3edc016aa9f7ee9aec87d189d0298250576b69","toolName":"grep","startedAt":"2026-09-16T01:59:10.705Z","args":{"path":"crates/spt-daemon/src/brainproc.rs"},"intent":"Locating brain drain paths"},"id":"a87e8015","parentId":"1353d4e8","timestamp":"2026-09-16T01:59:10.705Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_jPCBai7mJxwK4o0XVq9jHBeP|fc_0d104aba2b0f3edc016aa9f7ee9af487d1ab71123d585b9a96","toolName":"glob","startedAt":"2026-09-16T01:59:10.705Z","args":{"path":".spt/preserved/307/**/*;**/*hazard*;**/*glossary*"},"intent":"Locating preserved hazard notes"},"id":"2685cf0b","parentId":"a87e8015","timestamp":"2026-09-16T01:59:10.705Z"}
{"type":"message","id":"43262bde","parentId":"2685cf0b","timestamp":"2026-09-16T01:59:10.706Z","message":{"role":"assistant","content":[{"type":"toolCall","id":"call_vsP5v6FyYGuEEn6CHaudvagj|fc_0d104aba2b0f3edc016aa9f7ee9ad887d18e9e981dd73459df","name":"grep","arguments":{"i":"Locating controller drain paths","pattern":"old_gen_drained|CONTROLLER|controller|slot.*clos|cause=|WRITE_POISONED","path":"crates/spt-daemon/src/broker.rs","case":false,"gitignore":false,"skip":null},"intent":"Locating controller drain paths"},{"type":"toolCall","id":"call_ZyQI4lPZV9LtJvgMAINxYgSv|fc_0d104aba2b0f3edc016aa9f7ee9aec87d189d0298250576b69","name":"grep","arguments":{"i":"Locating brain drain paths","pattern":"old_gen_drained|SUBSCRIBER_STALL|detach|drain|UPDATE_RESTART","path":"crates/spt-daemon/src/brainproc.rs","case":false,"gitignore":false,"skip":null},"intent":"Locating brain drain paths"},{"type":"toolCall","id":"call_jPCBai7mJxwK4o0XVq9jHBeP|fc_0d104aba2b0f3edc016aa9f7ee9af487d1ab71123d585b9a96","name":"glob","arguments":{"i":"Locating preserved hazard notes","path":".spt/preserved/307/**/*;**/*hazard*;**/*glossary*","hidden":true,"gitignore":false,"limit":150},"intent":"Locating preserved hazard notes"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":1819,"output":210,"cacheRead":10240,"cacheWrite":0,"totalTokens":12269,"cost":{"input":0.01819,"output":0.0105,"cacheRead":0.010239999999999999,"cacheWrite":0,"total":0.03893}},"stopReason":"toolUse","timestamp":1789523943339,"responseId":"resp_0d104aba2b0f3edc016aa9f7e78e2887d1a9f1d1fdf36f2d59","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"fc_0d104aba2b0f3edc016aa9f7ee9ad887d18e9e981dd73459df","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating controller drain paths\",\"pattern\":\"old_gen_drained|CONTROLLER|controller|slot.*clos|cause=|WRITE_POISONED\",\"path\":\"crates/spt-daemon/src/broker.rs\",\"case\":false,\"gitignore\":false,\"skip\":null}","call_id":"call_vsP5v6FyYGuEEn6CHaudvagj","internal_chat_message_metadata_passthrough":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"},"name":"grep","metadata":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"}},{"id":"fc_0d104aba2b0f3edc016aa9f7ee9aec87d189d0298250576b69","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating brain drain paths\",\"pattern\":\"old_gen_drained|SUBSCRIBER_STALL|detach|drain|UPDATE_RESTART\",\"path\":\"crates/spt-daemon/src/brainproc.rs\",\"case\":false,\"gitignore\":false,\"skip\":null}","call_id":"call_ZyQI4lPZV9LtJvgMAINxYgSv","internal_chat_message_metadata_passthrough":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"},"name":"grep","metadata":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"}},{"id":"fc_0d104aba2b0f3edc016aa9f7ee9af487d1ab71123d585b9a96","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating preserved hazard notes\",\"path\":\".spt/preserved/307/**/*;**/*hazard*;**/*glossary*\",\"hidden\":true,\"gitignore\":false,\"limit\":150}","call_id":"call_jPCBai7mJxwK4o0XVq9jHBeP","internal_chat_message_metadata_passthrough":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"},"name":"glob","metadata":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"}}]},"duration":7363.418599999975,"ttft":7291.6664999998175,"completedAt":1789523950704,"contextSnapshot":{"promptTokens":12059,"nonMessageTokens":10480,"compactionEpoch":0}}}
{"type":"message","id":"784dc78f","parentId":"43262bde","timestamp":"2026-09-16T01:59:10.718Z","message":{"role":"toolResult","toolCallId":"call_ZyQI4lPZV9LtJvgMAINxYgSv|fc_0d104aba2b0f3edc016aa9f7ee9aec87d189d0298250576b69","toolName":"grep","content":[{"type":"text","text":" 195|    // Persist this brain generation's stderr to the rotating SPT_HOME/logs sink\n*196|    // (REQ-DAEMON-STDERR-PERSIST) — a detached brain's stdio is /dev/null, so every\n 197|    // diagnostic below would otherwise vanish (the incident-night RCA-blind gap). Runs\n 198|    // FIRST so even the BRAIN_UP line and a connect failure are captured. Best-effort.\n 199|    // [impl->REQ-DAEMON-STDERR-PERSIST]\n...\n 268|    // CURSOR-ONLY resume (UPDATE-WEDGE round 3, REQ-BRAIN-RESUME-NO-CONN-DEADLOCK,\n*269|    // doyle Option A-split): the supervised daemon brain is a NON-DRAINING consumer of\n 270|    // this conn (the heartbeat below reads it only in net_status()/sessions()\n*271|    // drain-and-discard loops), so it MUST NOT subscribe — an undrained subscription on\n 272|    // this request/reply conn is the SharedSend self-deadlock that re-wedged v0.30.4\n 273|    // (brain.rs resume_sessions deadlock warning). `resume_session_cursors` re-establishes\n 274|    // the per-session dedup cursors + returns the ids WITHOUT subscribing; the subscribing\n*275|    // `resume_sessions` stays for DRAINING callers (the resume int tests / the forward\n*276|    // live-agent-adapter drained carrier). Empty today (the daemon hosts no PTY sessions).\n 277|    // [impl->REQ-BRAIN-RESUME-NO-CONN-DEADLOCK]\n 278|    match brain.resume_session_cursors() {\n 279|        Ok(ids) if !ids.is_empty() => {\n...\n 362|    // reconcile in the background (ready-without-warm is preserved because\n*363|    // write_ready already fired above), then drain/debounce/coalesce the\n 364|    // invalidation queue. Same posture as the hosts above: the stop flag is\n 365|    // held for symmetry, the thread dies with the brain, and the next brain\n 366|    // rehydrates from the persisted index (warm start, fingerprint-gated).\n...\n 396|        // EOFs). Nothing else polls `KIND_SESSIONS` during an isolated brain-trial\n*397|        // window, so without this the wedge never drains: the promotion DRAINED gate\n*398|        // (`old_gen_drained()` = `!any_local_controller_wedged()`, a PURE READ that never\n 399|        // drives the evict) stays false for the full trial → \"alive but never ready\"\n 400|        // rollback (the 2026-07-09 field freeze), AND the same wedge black-holes recovery\n*401|        // (the late `BRAIN_SUBSCRIBER_STALL_EVICT`). Driven on the FIRST heartbeat and on\n 402|        // BOTH the trial candidate AND the rollback/recovery brain (one loop covers both);\n 403|        // best-effort — a transient sessions error must never fell the heartbeat (the\n 404|        // net_status probe above already owns broker-liveness). Self-applies: it drives\n...\n 406|        // its own trial against a live older broker — no coordinated broker restart.\n*407|        // [impl->REQ-UPDATE-TRIAL-DRAIN-DRIVE]\n 408|        let _ = brain.sessions();\n 409|        let _ = write_ready(generation);\n 410|    }\n...\n 628|    fn staged_artifact_hash(&self, version: u64) -> Option<String>;\n*629|    /// Whether the OLD brain generation's broker subscriber connection is DRAINED —\n 630|    /// its local controller conn CLOSED or stall-evicted, no blocked writes still\n*631|    /// pending on it (REQ-UPDATE-PROMOTE-DRAINED). Reads BROKER TRUTH in-process (no\n 632|    /// brain round-trip): the supervisor promotes a ready candidate only once this is\n 633|    /// `true`, so `brain.ready` arriving inside the W2 `BRAIN_WRITE_DEADLINE` window\n 634|    /// can no longer promote onto a still-frozen control plane (the 22:47 false-\n*635|    /// promote). Fail-safe toward DRAINED (`true`) where there is nothing to read — a\n 636|    /// net-less broker, no hosted sessions, or [`NoTrialEnv`] — so the gate never\n 637|    /// strands an apply where no old-gen control plane exists.\n*638|    fn old_gen_drained(&self) -> bool;\n 639|}\n 640|\n 641|/// The no-trial environment: every probe degrades to \"nothing pending / not\n...\n 661|    }\n*662|    fn old_gen_drained(&self) -> bool {\n*663|        // No readiness gate to drive here → no old-gen control plane to drain; the\n 664|        // no-op env promotes exactly as the pre-D6 respawn loop (mirrors its other\n 665|        // no-op effects).\n 666|        true\n...\n 675|/// Holds the live [`crate::broker::Broker`] handle (or `None` on a net-less /\n*676|/// bind-failed daemon) so the DRAINED gate ([`TrialEnv::old_gen_drained`]) reads\n 677|/// broker truth IN-PROCESS — the supervisor thread runs inside the broker process,\n 678|/// so this is a direct memory read, never an IPC round-trip through a possibly-\n*679|/// wedged conn (REQ-UPDATE-PROMOTE-DRAINED).\n 680|pub struct ProductionTrialEnv {\n 681|    broker: Option<Arc<crate::broker::Broker>>,\n 682|}\n...\n 732|    }\n*733|    // [impl->REQ-UPDATE-PROMOTE-DRAINED]\n*734|    fn old_gen_drained(&self) -> bool {\n*735|        // Drained ⟺ no hosted session has a wedged LOCAL (brain-own) controller conn.\n*736|        // No broker (net-less / bind-failed) → no control plane → drained. A broker\n 737|        // with no wedged local controller (the common case, incl. no daemon-hosted\n*738|        // sessions) reads drained at once → the normal fast promote on `brain.ready`.\n 739|        self.broker\n 740|            .as_ref()\n 741|            .is_none_or(|b| !b.any_local_controller_wedged())\n...\n 833|    /// The window elapsed with the child **alive but not promotable** — either\n*834|    /// never ready (\"up but useless\") OR ready-but-the-old-generation-never-drained\n*835|    /// (REQ-UPDATE-PROMOTE-DRAINED: a persistent old-gen wedge past the window).\n 836|    /// The caller MUST kill it before spawning the rollback binary (**A11** — never\n 837|    /// two live brains).\n 838|    WindowElapsedAlive,\n...\n 848|/// cross-lifetime exact-generation hole). Returns the moment readiness is seen\n*849|/// **AND the old generation has drained**, the child exits pre-ready, the window\n 850|/// elapses, or `stop` is raised.\n 851|///\n*852|/// **The DRAINED precondition (REQ-UPDATE-PROMOTE-DRAINED).** `brain.ready` alone is\n 853|/// NOT a promote: the W2 stall-evict only BOUNDS the false-promote window to\n 854|/// `BRAIN_WRITE_DEADLINE`, so a new brain can signal ready while the OLD generation's\n 855|/// local controller conn is still wedged (the 22:47 apply \"promoted\" onto a frozen\n...\n 857|/// ready then exited is not demoted, preserving the ready-then-exit-still-promotes\n*858|/// semantic) and promotion waits for `env.old_gen_drained()` too. While ready-seen\n*859|/// but not-yet-drained the watch keeps polling to `deadline`; if the drain never\n 860|/// comes the window elapses to [`TrialStep::WindowElapsedAlive`] — the caller KILLS\n 861|/// the candidate and rolls back (conservative: never a false-promote onto a wedge).\n*862|/// The normal case drains well inside the window — W2 bounds a real drain to\n 863|/// `BRAIN_WRITE_DEADLINE` (15 s) < [`SUPERVISE_HEALTHY_RUN`] (30 s) — and a cleanly-\n*864|/// closed old conn drains at once. `child.try_wait()` is consulted only while\n*865|/// NOT-yet-ready, so a ready-then-exited candidate still promotes once drained.\n 866|fn run_trial(\n 867|    child: &mut Child,\n 868|    generation: u64,\n...\n 875|    // ONCE PER TRIAL (releases#49/#267 Q3). The gate is a pure read that says\n*876|    // nothing when it says no, so a trial that waited on the drain and a trial\n*877|    // that never had to are the same silence. One line on the FIRST not-drained\n 878|    // read ends that; repeating it every tick would turn an answer into a log.\n*879|    let mut not_drained_reported = false;\n 880|    loop {\n 881|        if stop.load(Ordering::Relaxed) {\n 882|            return TrialStep::Stopped;\n...\n 884|        // Latch readiness FIRST (don't demote a ready-then-exit child) — but do NOT\n*885|        // promote on it alone: the old generation's control plane must have drained.\n 886|        if !ready_seen && env.ready_generation() == Some(generation) {\n 887|            ready_seen = true;\n 888|        }\n...\n 890|            // ONE observation, read once and used twice — the promote decision and\n*891|            // the diagnostic line are the SAME read. A second `old_gen_drained()`\n 892|            // call to describe what the first one decided could disagree with it,\n 893|            // and a diagnostic that can contradict the branch it reports on is\n 894|            // worse than no diagnostic.\n*895|            let drained = env.old_gen_drained();\n*896|            if drained {\n 897|                return TrialStep::Promoted;\n 898|            }\n 899|            // THE NARROWER FACT, DELIBERATELY. The gate answers one bool over every\n...\n 901|            // separately-taken scan, so this reports what this observation actually\n*902|            // established — that the gate was EVALUATED and read not-drained — and\n 903|            // nothing about which conn held it. It is scoped to THIS trial: earlier\n 904|            // trials evaluated the same gate and this line says nothing about them, so\n 905|            // it must not be read as the first evaluation the daemon ever made. It is\n 906|            // also not evidence that the\n*907|            // heartbeat's drain-driving `brain.sessions()` call completed: this says\n 908|            // the gate was consulted, never that the reap ran.\n 909|            // [impl->REQ-REFRESH-WAIT-ATTRIBUTION]\n*910|            if !not_drained_reported {\n*911|                not_drained_reported = true;\n 912|                if crate::brain::diag_trace_enabled() {\n 913|                    spt_proto::emit_line_err!(\n*914|                        \"BRAIN_PROMOTE_GATE_NOT_DRAINED: generation {generation} — ready latched \\\n*915|                         and the DRAINED gate read NOT drained; this is the first such \\\n 916|                         observation in THIS trial, not the first evaluation of the gate; the trial \\\n 917|                         keeps polling to its window \\\n 918|                         [REQ-REFRESH-WAIT-ATTRIBUTION]\"\n...\n 923|        // Only watch for a pre-ready exit while the candidate has NOT reached ready;\n*924|        // once ready-seen we are waiting on the OLD gen to drain, and a ready-then-\n 925|        // exited child still promotes (the exit is a normal later respawn).\n 926|        if !ready_seen {\n 927|            match child.try_wait() {\n...\n 1115|                        if restart.take() {\n*1116|                            spt_proto::emit_line_err!(\"BRAIN_UPDATE_RESTART: planned restart — cycling brain onto the swapped binary\");\n 1117|                            let _ = child.kill();\n 1118|                            let _ = child.wait();\n 1119|                            planned = true;\n...\n 1224|/// Spawn the production brain child: re-exec `daemon brain` as a plain\n*1225|/// (non-detached) managed child, so it inherits the broker's environment —\n*1226|/// including `SPT_HOME` — and dies with no orphaning. NOT `spawn_detached`\n 1227|/// (that path rebuilds the Windows env block and would drop `SPT_HOME`). The\n 1228|/// broker stamps `{generation, start-reason}` onto the argv (ADR-0018 V2).\n 1229|///\n...\n 1271|    // broker has a console (a `spt daemon run` in a terminal, or any test that\n*1272|    // spawns the daemon); CREATE_NO_WINDOW suppresses it without detaching, so\n 1273|    // the plain-child SPT_HOME inheritance (above) is preserved.\n 1274|    #[cfg(windows)]\n 1275|    {\n...\n 1281|    // OWN process group so the daemon can reap the brain + everything it spawns\n*1282|    // (the shellwake watchers, the detached Psyches — all inherit this group) as\n 1283|    // one unit on stop, without signalling the daemon itself. Windows roots the\n 1284|    // reap on a Job object instead (assigned in the supervisor via BrainReaper).\n 1285|    #[cfg(unix)]\n...\n 1349|    let handle = thread::spawn(move || {\n*1350|        // The DRAINED gate (REQ-UPDATE-PROMOTE-DRAINED) reads broker truth in-process;\n 1351|        // a net-less / bind-failed daemon has no broker (None) → the gate degrades to\n*1352|        // readiness-only (no old-gen control plane to drain).\n 1353|        let env = ProductionTrialEnv { broker };\n 1354|        supervise_brain(\n 1355|            &stop,\n...\n 1695|        staged_hash: Mutex<Option<String>>,\n*1696|        // DRAINED-gate countdown (REQ-UPDATE-PROMOTE-DRAINED): each `old_gen_drained`\n 1697|        // call returns `false` and decrements while > 0, then `true`. 0 (the default)\n*1698|        // = always drained, so every pre-existing trial test promotes exactly as\n*1699|        // before; `u32::MAX` = never drains within a test window; a small N = drains\n*1700|        // after N polls (the ready-seen latch holds across the undrained polls).\n*1701|        drain_countdown: AtomicU32,\n 1702|    }\n 1703|\n 1704|    impl TestTrialEnv {\n...\n 1719|                staged_hash: Mutex::new(None),\n*1720|                drain_countdown: AtomicU32::new(0),\n 1721|            }\n 1722|        }\n 1723|\n...\n 1730|\n*1731|        /// Arm the DRAINED gate: `old_gen_drained` returns `false` for the first\n*1732|        /// `polls` calls, then `true` (0 = always drained; `u32::MAX` = never within\n*1733|        /// the test window). Models the OLD-gen control plane draining after N polls.\n*1734|        fn drains_after(self, polls: u32) -> Self {\n*1735|            self.drain_countdown.store(polls, Ordering::Relaxed);\n 1736|            self\n 1737|        }\n 1738|\n...\n 1780|        }\n*1781|        fn old_gen_drained(&self) -> bool {\n*1782|            // Countdown: false (still draining) while > 0, decrementing; then true.\n 1783|            // Single-threaded per env in these harnesses (the supervisor polls\n 1784|            // sequentially), so a plain load/store is race-free here.\n*1785|            let left = self.drain_countdown.load(Ordering::Relaxed);\n 1786|            if left == 0 {\n 1787|                return true;\n 1788|            }\n*1789|            self.drain_countdown.store(left - 1, Ordering::Relaxed);\n 1790|            false\n 1791|        }\n 1792|    }\n...\n 1861|\n*1862|    /// (DRAINED gate — RED-first) A candidate that signals ready but whose OLD\n*1863|    /// generation NEVER drains does NOT promote onto the frozen control plane — the\n 1864|    /// window elapses, the candidate is killed, and it rolls back (the 22:47 false-\n*1865|    /// promote, now refused). RED-first: remove the `&& env.old_gen_drained()` gate in\n 1866|    /// `run_trial` and this promotes on ready alone — `promotions == [6]`,\n 1867|    /// `rollbacks == []` — exactly the incident-night behavior this closes.\n*1868|    // [unit->REQ-UPDATE-PROMOTE-DRAINED]\n 1869|    #[test]\n*1870|    fn ready_but_old_gen_never_drains_does_not_promote_rolls_back() {\n 1871|        // Ready is stamped, bytes match (so a bare readiness gate WOULD promote), but\n*1872|        // the old-gen control plane never drains within the window.\n 1873|        let env = Arc::new(\n 1874|            TestTrialEnv::pending(6, \"/good/spt.old-6\")\n 1875|                .with_hashes(Some(\"bytes\"), Some(\"bytes\"))\n*1876|                .drains_after(u32::MAX),\n 1877|        );\n 1878|        let stop = Arc::new(AtomicBool::new(false));\n 1879|        let restart = BrainRestart::new();\n...\n 1886|            env.as_ref(),\n*1887|            Duration::from_millis(120), // short window: the un-draining candidate elapses it\n 1888|            move |gen, _reason, binary| {\n 1889|                if binary.is_some() {\n 1890|                    // The rollback binary spawn — the trial did NOT promote.\n...\n 1905|            &[(6, 5, \"/good/spt.old-6\".to_string())],\n*1906|            \"the never-drained candidate is killed + rolled back (conservative, never a false-promote)\"\n 1907|        );\n 1908|        assert_eq!(env.notifs.load(Ordering::Relaxed), 1, \"one loud rollback notif\");\n 1909|    }\n 1910|\n*1911|    /// (DRAINED gate — latch releases) A ready candidate whose old generation drains\n 1912|    /// AFTER a couple of polls promotes: the ready-seen latch holds across the\n*1913|    /// undrained polls (a ready-then-exited candidate is not lost while waiting) and\n*1914|    /// promotion fires the moment `old_gen_drained()` flips true — never a premature\n 1915|    /// rollback, never a false-promote.\n*1916|    // [unit->REQ-UPDATE-PROMOTE-DRAINED]\n 1917|    #[test]\n*1918|    fn ready_candidate_promotes_once_old_gen_drains() {\n*1919|        let env = Arc::new(TestTrialEnv::pending(7, \"/good/spt.old-7\").drains_after(1));\n 1920|        let stop = Arc::new(AtomicBool::new(false));\n 1921|        let restart = BrainRestart::new();\n 1922|        let env_c = Arc::clone(&env);\n...\n 1930|            env.as_ref(),\n*1931|            Duration::from_secs(5), // generous: the drain arrives well inside the window\n 1932|            move |gen, _reason, _binary| {\n 1933|                let n = spawns_c.fetch_add(1, Ordering::Relaxed) + 1;\n 1934|                if n == 1 {\n...\n 1945|            &[7],\n*1946|            \"promotes once the old generation drains (ready-seen latched across the wait)\"\n 1947|        );\n 1948|        assert!(\n 1949|            env.rollbacks.lock().unwrap().is_empty(),\n*1950|            \"no premature rollback — the drain arrived inside the window\"\n 1951|        );\n 1952|    }\n 1953|"}],"details":{"scopePath":"crates/spt-daemon/src/brainproc.rs","searchPath":"C:\\Users\\decid\\Documents\\projects\\spt-core\\crates\\spt-daemon\\src\\brainproc.rs","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","matchCount":86,"fileCount":1,"files":["crates/spt-daemon/src/brainproc.rs"],"fileMatches":[{"path":"crates/spt-daemon/src/brainproc.rs","count":86}],"truncated":false,"displayContent":"  195│    // Persist this brain generation's stderr to the rotating SPT_HOME/logs sink\n *196│    // (REQ-DAEMON-STDERR-PERSIST) — a detached brain's stdio is /dev/null, so every\n  197│    // diagnostic below would otherwise vanish (the incident-night RCA-blind gap). Runs\n  198│    // FIRST so even the BRAIN_UP line and a connect failure are captured. Best-effort.\n  199│    // [impl->REQ-DAEMON-STDERR-PERSIST]\n     │...\n  268│    // CURSOR-ONLY resume (UPDATE-WEDGE round 3, REQ-BRAIN-RESUME-NO-CONN-DEADLOCK,\n *269│    // doyle Option A-split): the supervised daemon brain is a NON-DRAINING consumer of\n  270│    // this conn (the heartbeat below reads it only in net_status()/sessions()\n *271│    // drain-and-discard loops), so it MUST NOT subscribe — an undrained subscription on\n  272│    // this request/reply conn is the SharedSend self-deadlock that re-wedged v0.30.4\n  273│    // (brain.rs resume_sessions deadlock warning). `resume_session_cursors` re-establishes\n  274│    // the per-session dedup cursors + returns the ids WITHOUT subscribing; the subscribing\n *275│    // `resume_sessions` stays for DRAINING callers (the resume int tests / the forward\n *276│    // live-agent-adapter drained carrier). Empty today (the daemon hosts no PTY sessions).\n  277│    // [impl->REQ-BRAIN-RESUME-NO-CONN-DEADLOCK]\n  278│    match brain.resume_session_cursors() {\n  279│        Ok(ids) if !ids.is_empty() => {\n     │...\n  362│    // reconcile in the background (ready-without-warm is preserved because\n *363│    // write_ready already fired above), then drain/debounce/coalesce the\n  364│    // invalidation queue. Same posture as the hosts above: the stop flag is\n  365│    // held for symmetry, the thread dies with the brain, and the next brain\n  366│    // rehydrates from the persisted index (warm start, fingerprint-gated).\n     │...\n  396│        // EOFs). Nothing else polls `KIND_SESSIONS` during an isolated brain-trial\n *397│        // window, so without this the wedge never drains: the promotion DRAINED gate\n *398│        // (`old_gen_drained()` = `!any_local_controller_wedged()`, a PURE READ that never\n  399│        // drives the evict) stays false for the full trial → \"alive but never ready\"\n  400│        // rollback (the 2026-07-09 field freeze), AND the same wedge black-holes recovery\n *401│        // (the late `BRAIN_SUBSCRIBER_STALL_EVICT`). Driven on the FIRST heartbeat and on\n  402│        // BOTH the trial candidate AND the rollback/recovery brain (one loop covers both);\n  403│        // best-effort — a transient sessions error must never fell the heartbeat (the\n  404│        // net_status probe above already owns broker-liveness). Self-applies: it drives\n     │...\n  406│        // its own trial against a live older broker — no coordinated broker restart.\n *407│        // [impl->REQ-UPDATE-TRIAL-DRAIN-DRIVE]\n  408│        let _ = brain.sessions();\n  409│        let _ = write_ready(generation);\n  410│    }\n     │...\n  628│    fn staged_artifact_hash(&self, version: u64) -> Option<String>;\n *629│    /// Whether the OLD brain generation's broker subscriber connection is DRAINED —\n  630│    /// its local controller conn CLOSED or stall-evicted, no blocked writes still\n *631│    /// pending on it (REQ-UPDATE-PROMOTE-DRAINED). Reads BROKER TRUTH in-process (no\n  632│    /// brain round-trip): the supervisor promotes a ready candidate only once this is\n  633│    /// `true`, so `brain.ready` arriving inside the W2 `BRAIN_WRITE_DEADLINE` window\n  634│    /// can no longer promote onto a still-frozen control plane (the 22:47 false-\n *635│    /// promote). Fail-safe toward DRAINED (`true`) where there is nothing to read — a\n  636│    /// net-less broker, no hosted sessions, or [`NoTrialEnv`] — so the gate never\n  637│    /// strands an apply where no old-gen control plane exists.\n *638│    fn old_gen_drained(&self) -> bool;\n  639│}\n  640│\n  641│/// The no-trial environment: every probe degrades to \"nothing pending / not\n     │...\n  661│    }\n *662│    fn old_gen_drained(&self) -> bool {\n *663│        // No readiness gate to drive here → no old-gen control plane to drain; the\n  664│        // no-op env promotes exactly as the pre-D6 respawn loop (mirrors its other\n  665│        // no-op effects).\n  666│        true\n     │...\n  675│/// Holds the live [`crate::broker::Broker`] handle (or `None` on a net-less /\n *676│/// bind-failed daemon) so the DRAINED gate ([`TrialEnv::old_gen_drained`]) reads\n  677│/// broker truth IN-PROCESS — the supervisor thread runs inside the broker process,\n  678│/// so this is a direct memory read, never an IPC round-trip through a possibly-\n *679│/// wedged conn (REQ-UPDATE-PROMOTE-DRAINED).\n  680│pub struct ProductionTrialEnv {\n  681│    broker: Option<Arc<crate::broker::Broker>>,\n  682│}\n     │...\n  732│    }\n *733│    // [impl->REQ-UPDATE-PROMOTE-DRAINED]\n *734│    fn old_gen_drained(&self) -> bool {\n *735│        // Drained ⟺ no hosted session has a wedged LOCAL (brain-own) controller conn.\n *736│        // No broker (net-less / bind-failed) → no control plane → drained. A broker\n  737│        // with no wedged local controller (the common case, incl. no daemon-hosted\n *738│        // sessions) reads drained at once → the normal fast promote on `brain.ready`.\n  739│        self.broker\n  740│            .as_ref()\n  741│            .is_none_or(|b| !b.any_local_controller_wedged())\n     │...\n  833│    /// The window elapsed with the child **alive but not promotable** — either\n *834│    /// never ready (\"up but useless\") OR ready-but-the-old-generation-never-drained\n *835│    /// (REQ-UPDATE-PROMOTE-DRAINED: a persistent old-gen wedge past the window).\n  836│    /// The caller MUST kill it before spawning the rollback binary (**A11** — never\n  837│    /// two live brains).\n  838│    WindowElapsedAlive,\n     │...\n  848│/// cross-lifetime exact-generation hole). Returns the moment readiness is seen\n *849│/// **AND the old generation has drained**, the child exits pre-ready, the window\n  850│/// elapses, or `stop` is raised.\n  851│///\n *852│/// **The DRAINED precondition (REQ-UPDATE-PROMOTE-DRAINED).** `brain.ready` alone is\n  853│/// NOT a promote: the W2 stall-evict only BOUNDS the false-promote window to\n  854│/// `BRAIN_WRITE_DEADLINE`, so a new brain can signal ready while the OLD generation's\n  855│/// local controller conn is still wedged (the 22:47 apply \"promoted\" onto a frozen\n     │...\n  857│/// ready then exited is not demoted, preserving the ready-then-exit-still-promotes\n *858│/// semantic) and promotion waits for `env.old_gen_drained()` too. While ready-seen\n *859│/// but not-yet-drained the watch keeps polling to `deadline`; if the drain never\n  860│/// comes the window elapses to [`TrialStep::WindowElapsedAlive`] — the caller KILLS\n  861│/// the candidate and rolls back (conservative: never a false-promote onto a wedge).\n *862│/// The normal case drains well inside the window — W2 bounds a real drain to\n  863│/// `BRAIN_WRITE_DEADLINE` (15 s) < [`SUPERVISE_HEALTHY_RUN`] (30 s) — and a cleanly-\n *864│/// closed old conn drains at once. `child.try_wait()` is consulted only while\n *865│/// NOT-yet-ready, so a ready-then-exited candidate still promotes once drained.\n  866│fn run_trial(\n  867│    child: &mut Child,\n  868│    generation: u64,\n     │...\n  875│    // ONCE PER TRIAL (releases#49/#267 Q3). The gate is a pure read that says\n *876│    // nothing when it says no, so a trial that waited on the drain and a trial\n *877│    // that never had to are the same silence. One line on the FIRST not-drained\n  878│    // read ends that; repeating it every tick would turn an answer into a log.\n *879│    let mut not_drained_reported = false;\n  880│    loop {\n  881│        if stop.load(Ordering::Relaxed) {\n  882│            return TrialStep::Stopped;\n     │...\n  884│        // Latch readiness FIRST (don't demote a ready-then-exit child) — but do NOT\n *885│        // promote on it alone: the old generation's control plane must have drained.\n  886│        if !ready_seen && env.ready_generation() == Some(generation) {\n  887│            ready_seen = true;\n  888│        }\n     │...\n  890│            // ONE observation, read once and used twice — the promote decision and\n *891│            // the diagnostic line are the SAME read. A second `old_gen_drained()`\n  892│            // call to describe what the first one decided could disagree with it,\n  893│            // and a diagnostic that can contradict the branch it reports on is\n  894│            // worse than no diagnostic.\n *895│            let drained = env.old_gen_drained();\n *896│            if drained {\n  897│                return TrialStep::Promoted;\n  898│            }\n  899│            // THE NARROWER FACT, DELIBERATELY. The gate answers one bool over every\n     │...\n  901│            // separately-taken scan, so this reports what this observation actually\n *902│            // established — that the gate was EVALUATED and read not-drained — and\n  903│            // nothing about which conn held it. It is scoped to THIS trial: earlier\n  904│            // trials evaluated the same gate and this line says nothing about them, so\n  905│            // it must not be read as the first evaluation the daemon ever made. It is\n  906│            // also not evidence that the\n *907│            // heartbeat's drain-driving `brain.sessions()` call completed: this says\n  908│            // the gate was consulted, never that the reap ran.\n  909│            // [impl->REQ-REFRESH-WAIT-ATTRIBUTION]\n *910│            if !not_drained_reported {\n *911│                not_drained_reported = true;\n  912│                if crate::brain::diag_trace_enabled() {\n  913│                    spt_proto::emit_line_err!(\n *914│                        \"BRAIN_PROMOTE_GATE_NOT_DRAINED: generation {generation} — ready latched \\\n *915│                         and the DRAINED gate read NOT drained; this is the first such \\\n  916│                         observation in THIS trial, not the first evaluation of the gate; the trial \\\n  917│                         keeps polling to its window \\\n  918│                         [REQ-REFRESH-WAIT-ATTRIBUTION]\"\n     │...\n  923│        // Only watch for a pre-ready exit while the candidate has NOT reached ready;\n *924│        // once ready-seen we are waiting on the OLD gen to drain, and a ready-then-\n  925│        // exited child still promotes (the exit is a normal later respawn).\n  926│        if !ready_seen {\n  927│            match child.try_wait() {\n     │...\n 1115│                        if restart.take() {\n*1116│                            spt_proto::emit_line_err!(\"BRAIN_UPDATE_RESTART: planned restart — cycling brain onto the swapped binary\");\n 1117│                            let _ = child.kill();\n 1118│                            let _ = child.wait();\n 1119│                            planned = true;\n     │...\n 1224│/// Spawn the production brain child: re-exec `daemon brain` as a plain\n*1225│/// (non-detached) managed child, so it inherits the broker's environment —\n*1226│/// including `SPT_HOME` — and dies with no orphaning. NOT `spawn_detached`\n 1227│/// (that path rebuilds the Windows env block and would drop `SPT_HOME`). The\n 1228│/// broker stamps `{generation, start-reason}` onto the argv (ADR-0018 V2).\n 1229│///\n     │...\n 1271│    // broker has a console (a `spt daemon run` in a terminal, or any test that\n*1272│    // spawns the daemon); CREATE_NO_WINDOW suppresses it without detaching, so\n 1273│    // the plain-child SPT_HOME inheritance (above) is preserved.\n 1274│    #[cfg(windows)]\n 1275│    {\n     │...\n 1281│    // OWN process group so the daemon can reap the brain + everything it spawns\n*1282│    // (the shellwake watchers, the detached Psyches — all inherit this group) as\n 1283│    // one unit on stop, without signalling the daemon itself. Windows roots the\n 1284│    // reap on a Job object instead (assigned in the supervisor via BrainReaper).\n 1285│    #[cfg(unix)]\n     │...\n 1349│    let handle = thread::spawn(move || {\n*1350│        // The DRAINED gate (REQ-UPDATE-PROMOTE-DRAINED) reads broker truth in-process;\n 1351│        // a net-less / bind-failed daemon has no broker (None) → the gate degrades to\n*1352│        // readiness-only (no old-gen control plane to drain).\n 1353│        let env = ProductionTrialEnv { broker };\n 1354│        supervise_brain(\n 1355│            &stop,\n     │...\n 1695│        staged_hash: Mutex<Option<String>>,\n*1696│        // DRAINED-gate countdown (REQ-UPDATE-PROMOTE-DRAINED): each `old_gen_drained`\n 1697│        // call returns `false` and decrements while > 0, then `true`. 0 (the default)\n*1698│        // = always drained, so every pre-existing trial test promotes exactly as\n*1699│        // before; `u32::MAX` = never drains within a test window; a small N = drains\n*1700│        // after N polls (the ready-seen latch holds across the undrained polls).\n*1701│        drain_countdown: AtomicU32,\n 1702│    }\n 1703│\n 1704│    impl TestTrialEnv {\n     │...\n 1719│                staged_hash: Mutex::new(None),\n*1720│                drain_countdown: AtomicU32::new(0),\n 1721│            }\n 1722│        }\n 1723│\n     │...\n 1730│\n*1731│        /// Arm the DRAINED gate: `old_gen_drained` returns `false` for the first\n*1732│        /// `polls` calls, then `true` (0 = always drained; `u32::MAX` = never within\n*1733│        /// the test window). Models the OLD-gen control plane draining after N polls.\n*1734│        fn drains_after(self, polls: u32) -> Self {\n*1735│            self.drain_countdown.store(polls, Ordering::Relaxed);\n 1736│            self\n 1737│        }\n 1738│\n     │...\n 1780│        }\n*1781│        fn old_gen_drained(&self) -> bool {\n*1782│            // Countdown: false (still draining) while > 0, decrementing; then true.\n 1783│            // Single-threaded per env in these harnesses (the supervisor polls\n 1784│            // sequentially), so a plain load/store is race-free here.\n*1785│            let left = self.drain_countdown.load(Ordering::Relaxed);\n 1786│            if left == 0 {\n 1787│                return true;\n 1788│            }\n*1789│            self.drain_countdown.store(left - 1, Ordering::Relaxed);\n 1790│            false\n 1791│        }\n 1792│    }\n     │...\n 1861│\n*1862│    /// (DRAINED gate — RED-first) A candidate that signals ready but whose OLD\n*1863│    /// generation NEVER drains does NOT promote onto the frozen control plane — the\n 1864│    /// window elapses, the candidate is killed, and it rolls back (the 22:47 false-\n*1865│    /// promote, now refused). RED-first: remove the `&& env.old_gen_drained()` gate in\n 1866│    /// `run_trial` and this promotes on ready alone — `promotions == [6]`,\n 1867│    /// `rollbacks == []` — exactly the incident-night behavior this closes.\n*1868│    // [unit->REQ-UPDATE-PROMOTE-DRAINED]\n 1869│    #[test]\n*1870│    fn ready_but_old_gen_never_drains_does_not_promote_rolls_back() {\n 1871│        // Ready is stamped, bytes match (so a bare readiness gate WOULD promote), but\n*1872│        // the old-gen control plane never drains within the window.\n 1873│        let env = Arc::new(\n 1874│            TestTrialEnv::pending(6, \"/good/spt.old-6\")\n 1875│                .with_hashes(Some(\"bytes\"), Some(\"bytes\"))\n*1876│                .drains_after(u32::MAX),\n 1877│        );\n 1878│        let stop = Arc::new(AtomicBool::new(false));\n 1879│        let restart = BrainRestart::new();\n     │...\n 1886│            env.as_ref(),\n*1887│            Duration::from_millis(120), // short window: the un-draining candidate elapses it\n 1888│            move |gen, _reason, binary| {\n 1889│                if binary.is_some() {\n 1890│                    // The rollback binary spawn — the trial did NOT promote.\n     │...\n 1905│            &[(6, 5, \"/good/spt.old-6\".to_string())],\n*1906│            \"the never-drained candidate is killed + rolled back (conservative, never a false-promote)\"\n 1907│        );\n 1908│        assert_eq!(env.notifs.load(Ordering::Relaxed), 1, \"one loud rollback notif\");\n 1909│    }\n 1910│\n*1911│    /// (DRAINED gate — latch releases) A ready candidate whose old generation drains\n 1912│    /// AFTER a couple of polls promotes: the ready-seen latch holds across the\n*1913│    /// undrained polls (a ready-then-exited candidate is not lost while waiting) and\n*1914│    /// promotion fires the moment `old_gen_drained()` flips true — never a premature\n 1915│    /// rollback, never a false-promote.\n*1916│    // [unit->REQ-UPDATE-PROMOTE-DRAINED]\n 1917│    #[test]\n*1918│    fn ready_candidate_promotes_once_old_gen_drains() {\n*1919│        let env = Arc::new(TestTrialEnv::pending(7, \"/good/spt.old-7\").drains_after(1));\n 1920│        let stop = Arc::new(AtomicBool::new(false));\n 1921│        let restart = BrainRestart::new();\n 1922│        let env_c = Arc::clone(&env);\n     │...\n 1930│            env.as_ref(),\n*1931│            Duration::from_secs(5), // generous: the drain arrives well inside the window\n 1932│            move |gen, _reason, _binary| {\n 1933│                let n = spawns_c.fetch_add(1, Ordering::Relaxed) + 1;\n 1934│                if n == 1 {\n     │...\n 1945│            &[7],\n*1946│            \"promotes once the old generation drains (ready-seen latched across the wait)\"\n 1947│        );\n 1948│        assert!(\n 1949│            env.rollbacks.lock().unwrap().is_empty(),\n*1950│            \"no premature rollback — the drain arrived inside the window\"\n 1951│        );\n 1952│    }\n 1953│"},"isError":false,"timestamp":1789523950717}}
{"type":"message","id":"c856edcf","parentId":"784dc78f","timestamp":"2026-09-16T01:59:10.781Z","message":{"role":"toolResult","toolCallId":"call_vsP5v6FyYGuEEn6CHaudvagj|fc_0d104aba2b0f3edc016aa9f7ee9ad887d18e9e981dd73459df","toolName":"grep","content":[{"type":"text","text":" 160|\n*161|/// The controller live-frame channel depth (frames). Deeper than a viewer's\n*162|/// queue because the controller is the AUTHORITATIVE view — unlike a viewer it is\n 163|/// NOT evicted on a SINGLE full `try_send` (W1 / REQ-HAZARD-INJECT-CONTROL-COEXIST).\n 164|/// b4: the drain does ONE non-blocking `try_send` per chunk and NEVER blocks; a\n*165|/// merely-slow controller DROPS chunks + falls behind the ring, and only a channel\n*166|/// that stays full for [`CONTROLLER_WRITE_DEADLINE`] (tracked via\n*167|/// [`ControllerSink::last_ok`]) is detached. Generous so a momentary slow consumer\n 168|/// never trips the deadline.\n*169|const CONTROLLER_CHANNEL_DEPTH: usize = 4096;\n 170|\n*171|/// How long a CONTROLLER's channel may stay FULL (no successful handoff since\n*172|/// [`ControllerSink::last_ok`]) before the drain declares it stalled and detaches it\n*173|/// (W1, refined by b4 / REQ-HAZARD-VIEWER-STARVE-UNDER-CONTROLLER-BACKPRESSURE). This\n*174|/// BOUNDS the wedge: a backed-up controller falls behind for at most this long before\n 175|/// eviction — never forever (the v0.12.x root) — and the drain NEVER blocks or sleeps\n 176|/// (a single off-lock `try_send` per chunk; see [`OutputLog::append`]). A live but\n 177|/// merely slow operator lands a frame before this elapses (resetting `last_ok`) and is\n 178|/// untouched.\n*179|const CONTROLLER_WRITE_DEADLINE: Duration = Duration::from_secs(5);\n 180|\n 181|/// The RESIZE BARRIER's drain-quiescence window (REQ-RC-RESIZE-GEOMETRY-EPOCH):\n 182|/// how long the appended high-water must hold STILL before the surface resize is\n...\n 202|/// The cap on bytes HELD BACK from the grid across one resize transition. The ring\n*203|/// and the controller/viewer fan-out are never gated (delivery keeps running at full\n 204|/// speed during a transition — only the GRID parse is deferred), so this bounds the\n 205|/// transition's extra memory alone. Overflow drops from the grid parse ONLY, loudly,\n 206|/// and self-heals on the next output + repaint exactly like the poison-recover ring\n...\n 209|\n*210|/// How long a controller's writer thread may sit BLOCKED inside a single socket\n 211|/// write — the brain subscriber suspended / black-holed — before a take/reattach\n 212|/// or the `KIND_SESSIONS` reap stall-evicts it (LIFECYCLE-TRUTH W2,\n 213|/// REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE). Distinct from and 3× longer than\n*214|/// [`CONTROLLER_WRITE_DEADLINE`] on PURPOSE (doyle ruling Q1): that 5 s bound\n 215|/// protects the fan-out against ONE per-client conn falling behind and is\n 216|/// output-driven (the drain observes a Full channel); THIS bound is measured\n 217|/// full-with-**zero writer progress** and covers the case the output-driven path\n 218|/// can't — an IDLE session whose brain-subscriber conn is suspended produces no\n*219|/// output, so `append`/`mark_controller_gone` never runs, and the stalled\n*220|/// controller would otherwise latch `driven_by` forever (reattach refused, `--take`\n 221|/// blocked). The brain aggregates every session + psyche + WAN feed, so its drain\n 222|/// is bursty; 15 s tolerates a normal burst-stall while still bounding the wedge.\n 223|/// A slow-but-DRAINING brain never trips it — the clock only advances while a write\n*224|/// is genuinely in flight past the deadline (see [`ControllerSink::write_blocked_since`]).\n 225|// [impl->REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]\n 226|const BRAIN_WRITE_DEADLINE: Duration = Duration::from_secs(15);\n 227|\n...\n 229|/// suspend-brain int rig shrinks it via `SPT_BRAIN_WRITE_DEADLINE_MS` so the\n*230|/// structural outcome (a wedged controller is stall-evicted → take/reattach frees\n 231|/// the slot) is provable in ~1 s instead of 15 (mirrors the\n 232|/// `SPT_INJECT_COMMIT_DEADLINE_MS` env-knob precedent). Grants test wall-clock\n 233|/// ONLY — never relaxes an assertion.\n...\n 518|/// The RC-origin input-fence verdict (ADR-0044 decision 3,\n*519|/// REQ-INPUT-CONTROLLER-FENCE) — PURE over the request's minter tag and the\n*520|/// sender-conn's controller fact, so the scope table is unit-tested: ONLY an\n*521|/// `Minter::Rc`-tagged input from a NON-controller conn is rejected. Shell\n 522|/// injection (`Minter::Shell`), legacy/untagged senders (`None`/`Legacy` — the\n 523|/// N-1 wire), and every other minter pass exactly as before; the connection\n 524|/// identity is the lease surrogate (ADR-0044 decision 4), which holds for\n 525|/// mixed-version attaches too — every rc keystroke reaches the broker through\n*526|/// its own serve worker's conn, and `become_controller` moves the slot to the\n 527|/// live lease's conn atomically with the takeover.\n*528|// [impl->REQ-INPUT-CONTROLLER-FENCE]\n*529|fn rc_input_rejected(minter: Option<Minter>, sender_is_controller: bool) -> bool {\n*530|    minter == Some(Minter::Rc) && !sender_is_controller\n 531|}\n 532|\n 533|/// Whether a stall-evict re-stamps the perch INLINE (take/reattach path, already\n...\n 620|/// buggy translation binary can therefore stall operator input for at most this\n*621|/// long. Mirrors [`CONTROLLER_WRITE_DEADLINE`]'s order (5 s).\n 622|const INJECT_COMMIT_DEADLINE: Duration = Duration::from_secs(5);\n 623|\n 624|/// The effective commit deadline. Production is the const above; the int gate\n...\n 821|\n*822|/// The interactive CONTROLLER's isolated sink (W1 / REQ-HAZARD-INJECT-CONTROL-COEXIST).\n 823|/// Symmetric with [`ViewerSink`] — a dedicated writer thread does the BLOCKING\n*824|/// socket write so a backed-up controller can never stall the session drain\n*825|/// thread (the v0.12.x wedge: the controller write was inline on the drain,\n 826|/// holding `Mutex<OutputLog>`). Unlike a viewer it is AUTHORITATIVE: its writer\n 827|/// advances `delivered_through` on each successful write, and it is never evicted\n*828|/// on a single full `try_send` — only after [`CONTROLLER_WRITE_DEADLINE`].\n*829|struct ControllerSink {\n*830|    /// The bounded live-frame queue to the controller's writer thread.\n 831|    tx: SyncSender<CtrlMsg>,\n*832|    /// The controller's IPC send half — kept for identity matching on detach and\n 833|    /// for the `exit`-frame fan-out ([`OutputLog::all_sinks`]).\n 834|    send: SharedSend,\n*835|    /// The controller's ATTACH GENERATION (ADR-0038 Amendment fix 6,\n 836|    /// REQ-SERVE-OWNERSHIP-GENERATION): operator-minted wall-ms from the wire\n 837|    /// Request (`0` = local / N-1). A same-identity re-take with an OLDER\n 838|    /// generation is a replayed stale Request — refused; a stale worker's\n*839|    /// late detach (older generation) never clears a newer controller.\n 840|    attach_gen: u64,\n 841|    /// The `from_seq` this seat was ESTABLISHED with — the fourth element of the\n 842|    /// idempotent-replay key (ADR-0047 decision 3, REQ-ATTACH-IDEMPOTENT-REPLAY).\n...\n 844|    /// the establish already honored (the dispatcher's double-serve); one carrying\n*845|    /// a DIFFERENT `from_seq` is the controller GAP RESUME re-fetching dropped\n*846|    /// frames from its frozen floor (REQ-HAZARD-CONTROLLER-GAP-RESUME), which must\n 847|    /// be re-served — suppressing it would turn a recoverable gap into a\n 848|    /// truncation report.\n 849|    establish_from_seq: u64,\n*850|    /// The controller's identity: `Some(node)` = a remote operator, `None` =\n 851|    /// local (the brain's spawn-reader / resume). Ownership keys on this.\n 852|    by: Option<String>,\n 853|    /// The seal-ceremony CAPABILITY the attaching client declared on its\n...\n 862|    seal_ceremony: bool,\n*863|    /// Time of the last SUCCESSFUL handoff (`try_send` Ok) to this controller's\n*864|    /// writer channel (b4, REQ-HAZARD-VIEWER-STARVE-UNDER-CONTROLLER-BACKPRESSURE).\n 865|    /// The drain does a SINGLE non-blocking `try_send` per chunk and NEVER sleeps;\n*866|    /// a channel that stays Full for [`CONTROLLER_WRITE_DEADLINE`] PAST this stamp\n*867|    /// evicts the stalled controller (bounded wedge preserved), while a slow-but-\n*868|    /// alive controller just drops the chunk + falls behind the ring (resume-from-\n*869|    /// floor). Seeded `now` at [`OutputLog::become_controller`].\n 870|    last_ok: Instant,\n 871|    /// The writer's IN-FLIGHT socket-write marker (LIFECYCLE-TRUTH W2,\n*872|    /// REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE): the `controller_writer` sets it to\n 873|    /// `Some(Instant::now())` immediately BEFORE it enters a blocking `write_frame`\n 874|    /// and back to `None` the instant the write returns. A reader\n*875|    /// ([`OutputLog::controller_stalled`]) sees a `Some` older than\n 876|    /// [`BRAIN_WRITE_DEADLINE`] as a writer WEDGED on a suspended/black-holed brain\n 877|    /// conn and stall-evicts it — the signal `last_ok` (channel-handoff time) cannot\n*878|    /// give: an IDLE controller (writer parked on `rx.recv()`, `None` here) and a\n 879|    /// BLOCKED one (writer stuck in the socket write, `Some` here) both leave\n 880|    /// `last_ok` stale. `Arc<Mutex<…>>` so the reader observes the writer's live\n 881|    /// state; the mutex is taken only around the brief set/clear and the read, never\n...\n 888|    /// AUTHORITATIVE-revoke marker (ADR-0044 decision 2,\n*889|    /// REQ-CONTROLLER-LEASE-IDENTITY): set to the taker's identity by a\n 890|    /// distinct-lease `Take` BEFORE the slot is replaced. The writer thread\n 891|    /// reads it on its exit paths and emits the terminal [`KIND_DISPLACED`]\n 892|    /// frame as its FINAL write — so the revoke lands even when the live\n...\n 919|/// assigns a `seq` and RINGS every chunk — the ring stays the scrollback/resume\n*920|/// authority for raw history — but delivery to the controller and the viewers is\n 921|/// SUPPRESSED (a skip, never a block: the drain runs at full PTY speed, hazard\n 922|/// 7.12 intact), and the chunk is held back from `grid.advance`, tagged with the\n 923|/// geometry it was emitted under. [`OutputLog::commit_resize`] /\n...\n 997|/// A bounded, sequenced output log for one session — the gapless-resume buffer\n*998|/// AND the W2.5 controller/viewer fan-out hub (REQ-RCVIEW-1).\n 999|///\n 1000|/// The session's single drain thread is the only producer ([`OutputLog::append`]).\n*1001|/// Output fans out to ONE interactive `controller` (the authoritative,\n 1002|/// cursor-advancing, blocking path) plus ANY NUMBER of read-only `viewers` (each\n 1003|/// an isolated bounded queue + writer thread, evicted on overflow). Every\n*1004|/// controller send and viewer `try_send` happens while the enclosing\n 1005|/// `Mutex<OutputLog>` is held, so frames reach each sink in strict `seq` order.\n 1006|struct OutputLog {\n 1007|    session_id: u64,\n...\n 1010|    cap_chunks: usize,\n*1011|    /// The single interactive controller's isolated sink (W1: a dedicated writer\n 1012|    /// thread + bounded channel + authoritative cursor). `None` ⇒ undriven. The\n*1013|    /// controller's identity (`Some(node)` remote / `None` local) lives in\n*1014|    /// [`ControllerSink::by`] and keys ownership — a same-`by` re-subscribe\n 1015|    /// silently re-takes (no `Displaced`, the gate-#7 self-kick guard); only a\n 1016|    /// DIFFERENT remote's `Take` displaces.\n*1017|    controller: Option<ControllerSink>,\n*1018|    /// Monotonic controller generation, shared with the live `controller_writer`\n*1019|    /// thread (`Arc<AtomicU64>`). Bumped on every [`become_controller`]; the\n*1020|    /// drain's deadline-evict ([`mark_controller_gone`]) carries the epoch it\n*1021|    /// observed and only detaches if it still matches — so a slow-controller\n*1022|    /// timeout can never unseat a NEWER controller that attached meanwhile. The\n 1023|    /// writer snapshots its own epoch and returns the moment this value moves past\n 1024|    /// it, so a SUPERSEDED writer flushes no further frames (P1c, single live\n*1025|    /// writer per connection, [`REQ-HAZARD-CONTROLLER-WRITER-REORDER`]).\n*1026|    controller_epoch: Arc<AtomicU64>,\n 1027|    /// The read-only viewers, keyed by a broker-minted viewer id.\n 1028|    viewers: HashMap<u64, ViewerSink>,\n 1029|    /// Monotonic viewer-id source.\n...\n 1034|    endpoint: String,\n*1035|    /// Last-known PTY size — seeded at spawn, updated on a controller resize.\n 1036|    /// Pushed to a viewer on attach + to all viewers on resize (the letterbox\n 1037|    /// feed); the broker holds it since the PTY exposes no size getter.\n 1038|    size: (u16, u16),\n 1039|    /// Cursor-of-record (restoration D4-1, ADR-0018 Q6): the next seq a\n*1040|    /// re-attaching CONTROLLER still needs (= highest seq **successfully written**\n*1041|    /// to the controller, +1). It is **log** state, not subscriber state, so it\n 1042|    /// **survives `detach_if`** — a dead brain leaves its resume point behind, and\n 1043|    /// a cold-starting brain reads it (via `KIND_SESSIONS` → `resume_seq`) instead\n 1044|    /// of replaying the whole ring from 0. **Viewer writes NEVER move it**\n*1045|    /// (REQ-RCVIEW-1: the brain-resume cursor tracks the controller only).\n 1046|    ///\n 1047|    /// **Advanced only on a successful socket write** (`Ok`), and **monotonic**\n 1048|    /// (never decreases) — see [`advanced_cursor`]. Both invariants are\n 1049|    /// load-bearing for at-least-once output: advancing past a *failed* write to a\n*1050|    /// dying-but-not-yet-detached controller would make resume skip that chunk\n 1051|    /// forever (at-least-once → at-most-once); decreasing on a deliberate\n 1052|    /// `attach(sub, 0)` rewind-replay would reset the resume point under a live\n 1053|    /// reader.\n 1054|    ///\n*1055|    /// W1: now an `Arc<AtomicU64>` shared with the controller's writer thread,\n 1056|    /// which advances it (via [`advanced_cursor`], CAS-monotonic) as it writes —\n*1057|    /// the cursor moved off the drain thread with the controller delivery itself.\n 1058|    delivered_through: Arc<AtomicU64>,\n 1059|    /// The server-side **render grid** (#6 / REQ-BROKER-SCREEN-GRID, ADR-0031):\n 1060|    /// the authoritative CURRENT SCREEN, interpreted from the same byte stream the\n 1061|    /// ring records. Fed on the drain path in [`append`], AFTER the fan-out (so it\n*1062|    /// never adds latency to controller/viewer delivery — hazard 7.12). A cold\n 1063|    /// attach's initial batch is a single synthesized [`ScreenGrid::render_repaint`]\n 1064|    /// instead of the raw ring, so replaying an alt-screen TUI into a fresh terminal\n 1065|    /// no longer corrupts scrollback. Single-writer (the one drain thread), mirroring\n...\n 1223|    if control || viewer {\n*1224|        // The two IN-MEMORY reap paths (`reap_dead_controller`, the stall-evict)\n 1225|        // deliberately do not stamp under the log lock, so their attachment edge\n*1226|        // becomes true HERE and nowhere else — signal it, or a controller that\n 1227|        // died with its conn would leave a stale attachment frame standing until\n 1228|        // the next sweep. (releases#44)\n 1229|        // [impl->REQ-ATTACH-LINK-PUSH]\n...\n 1235|        // The THIRD engine-room drop edge (ADR-0052 §3a). The two in-memory\n*1236|        // reap paths (`reap_dead_controller`, the stall-evict) deliberately do\n*1237|        // NOT stamp under the log lock, so they never reach `clear_controller`;\n 1238|        // this off-lock converge is where their release becomes true. Keying on\n*1239|        // the recorded truth going FALSE means a controller that died with the\n 1240|        // conn revokes its authority exactly like one that detached cleanly.\n*1241|        // [impl->REQ-ER-CONTROLLER-BOUND-POSTURE]\n 1242|        if !want_controlled {\n 1243|            if let Err(e) = spt_store::engineroom::drop_authority_at_edge(endpoint) {\n 1244|                spt_proto::emit_line_err!(\"ENGINE_ROOM_AUTHORITY_UNDROPPED:{endpoint}: {e}\");\n...\n 1258|/// processed sequentially by that conn's dispatch loop, so a drag-resize can never\n*1259|/// overlap itself — but the controller SLOT can move (`serve_attach` forwards a\n 1260|/// viewport's `Resize` on whichever stream currently holds control), so a\n 1261|/// takeover landing inside another transition's window passes the\n*1262|/// controller-exclusive gate and arrives second. Refusing it would hand the\n 1263|/// operator a hard resize failure for a transient condition; the geometry the user\n 1264|/// asked for last is the geometry they want, so the second resize simply WAITS its\n 1265|/// turn and then runs as a normal transition (last writer wins, no coalescing\n...\n 1374|            cap_chunks,\n*1375|            controller: None,\n*1376|            controller_epoch: Arc::new(AtomicU64::new(0)),\n 1377|            viewers: HashMap::new(),\n 1378|            next_viewer_id: 0,\n 1379|            endpoint,\n...\n 1404|    /// irrecoverable-behind detection untouched\n*1405|    /// (REQ-HAZARD-CONTROLLER-GAP-RESUME / B2). `None` also when nothing has\n 1406|    /// been produced yet (`next_seq == 0`) → an empty initial batch, so the\n 1407|    /// first live frame is the contiguous seq 0 (the spawn seat, unchanged).\n 1408|    ///\n...\n 1444|    /// Append a freshly-produced output chunk: assign the next `seq`, retain it in\n*1445|    /// the bounded ring, then FAN OUT to viewers AND the controller — both a SINGLE\n 1446|    /// non-blocking `try_send`, so the drain NEVER blocks/sleeps here. During a\n 1447|    /// resize transition the fan-out is SUPPRESSED wholesale (still a skip, never a\n 1448|    /// block — see the barrier note in the body,\n 1449|    /// REQ-RC-RESIZE-PRESENTATION-BARRIER). Returns\n*1450|    /// `Some(epoch)` when the controller must be EVICTED (stalled past the deadline,\n*1451|    /// or its writer is gone) so the drain can `mark_controller_gone(epoch)` OFF the\n 1452|    /// log lock; `None` otherwise.\n 1453|    ///\n*1454|    /// b4 (REQ-HAZARD-VIEWER-STARVE-UNDER-CONTROLLER-BACKPRESSURE): W1 moved the\n*1455|    /// controller SOCKET WRITE off the drain (the `controller_writer` thread) but\n*1456|    /// left the HANDOFF as an inline sleep-poll (`ControllerJob::deliver`, up to\n*1457|    /// `CONTROLLER_WRITE_DEADLINE`) ON the drain — so a controller draining slower\n 1458|    /// than the PTY floods filled its channel, the inline poll throttled the drain,\n 1459|    /// and the viewer fan-out below stopped running → a concurrent `rc --view`\n*1460|    /// VIEWER starved. Now the controller handoff is ONE non-blocking `try_send`\n 1461|    /// exactly like a viewer: the drain runs at full PTY speed regardless of\n*1462|    /// controller speed. A slow-but-alive controller DROPS the chunk + falls behind\n*1463|    /// the ring (resume-from-floor); a controller Full for `CONTROLLER_WRITE_DEADLINE`\n*1464|    /// since its last successful handoff (`ControllerSink::last_ok`) is evicted\n 1465|    /// (bounded wedge preserved). The writer still does the blocking socket write +\n 1466|    /// advances the authoritative CONTIGUOUS cursor. Called only from the session's\n 1467|    /// single drain thread.\n 1468|    // [impl->REQ-HAZARD-INJECT-CONTROL-COEXIST]\n*1469|    // [impl->REQ-HAZARD-VIEWER-STARVE-UNDER-CONTROLLER-BACKPRESSURE]\n 1470|    #[must_use]\n 1471|    /// Coherence clamp for a poison-RECOVERED ring (B-1, doyle ruling; run only by\n 1472|    /// [`recover_log`] on the recovery path). A panic mid-[`append`] can leave the\n...\n 1538|        // took its seq and rang above — raw history and the resume floor stay\n*1539|        // intact — but the controller handoff and the viewer fan-out are\n 1540|        // SUPPRESSED, a SKIP and never a block or a sleep (hazard 7.12: the drain\n 1541|        // returns at full PTY speed). A live attached terminal must never render\n 1542|        // old- and new-geometry differentials across a viewport it has already\n 1543|        // resized; commit/abort substitute ONE synthesized sync frame instead.\n*1544|        // The suppressed window is exempt from the controller's Full-eviction\n 1545|        // deadline by construction here (no handoff is attempted, so the deadline\n 1546|        // is never evaluated) and `last_ok` is re-armed at commit/abort so the\n*1547|        // window's width can never be charged to the controller afterwards. The\n 1548|        // chunk is HELD for the grid, tagged with the geometry it was emitted\n 1549|        // under, and replayed at that geometry on commit\n 1550|        // (REQ-RC-RESIZE-GEOMETRY-EPOCH, unchanged).\n...\n 1562|        self.push_frame_to_viewers(&frame);\n*1563|        // Controller: the SAME non-blocking discipline as a viewer (b4). One\n 1564|        // `try_send`, never a sleep. Ok → record the handoff time. Full → DROP (a\n*1565|        // slow controller falls behind the ring) unless it has been Full for the\n 1566|        // whole deadline since its last success → EVICT. Disconnected (writer gone)\n*1567|        // → EVICT. The drain off-lock turns `Some(epoch)` into `mark_controller_gone`.\n 1568|        let now = Instant::now();\n*1569|        let epoch = self.controller_epoch.load(Ordering::Acquire);\n*1570|        let result = if let Some(c) = self.controller.as_mut() {\n 1571|            match c.tx.try_send(CtrlMsg::Output(seq, frame)) {\n 1572|                Ok(()) => {\n 1573|                    c.last_ok = now;\n...\n 1576|                Err(TrySendError::Full(_)) => {\n*1577|                    if now.duration_since(c.last_ok) >= CONTROLLER_WRITE_DEADLINE {\n 1578|                        Some(epoch) // stalled past the deadline → bounded detach\n 1579|                    } else {\n*1580|                        None // drop this chunk; controller falls behind the ring\n 1581|                    }\n 1582|                }\n 1583|                Err(TrySendError::Disconnected(_)) => Some(epoch),\n...\n 1588|        // Advance the render grid AFTER the fan-out — it must never sit between the\n*1589|        // drain and the controller/viewer `try_send` (REQ-HAZARD-INJECT-CONTROL-COEXIST\n 1590|        // 7.12). The grid is single-writer (this one drain thread); `render_repaint`\n 1591|        // snapshots it under the same short log lock only when an attach requests it.\n 1592|        // (The in-transition path returned above: held chunks are replayed into the\n...\n 1599|\n*1600|    /// The current controller's identity (`Some(node)` remote / `None` local),\n 1601|    /// or `None` when undriven. The `by`-keyed ownership logic in\n 1602|    /// [`resolve_subscribe`] reads this.\n*1603|    fn controller_by(&self) -> Option<&str> {\n*1604|        self.controller.as_ref().and_then(|c| c.by.as_deref())\n 1605|    }\n 1606|\n*1607|    /// Whether ANY controller currently drives this session (the `controlled`\n*1608|    /// any-controller truth — REQ-HAZARD-CONTROL-STAMP-CONVERGENCE). Snapshotted under\n 1609|    /// the log lock, then converged onto the perch OFF the lock.\n*1610|    fn has_controller(&self) -> bool {\n*1611|        self.controller.is_some()\n 1612|    }\n 1613|\n*1614|    /// Snapshot the seal-ceremony SURFACE (WAX-SEAL W2): the live controller\n 1615|    /// seat's writer queue + conn + declared capability, read under the log\n 1616|    /// lock — the BROKER'S SEAT TABLE, never the perch info stamp (the stamp\n 1617|    /// lags the seat decision: KH 7.15 / info.rs terminal_normalize, and\n*1618|    /// `SessionInfo.controller_by == None` is ambiguous between empty and\n*1619|    /// local). `None` = no controller at all. The `SharedSend` rides along for\n 1620|    /// the bounded direct-write degrade the exit fan-out uses.\n 1621|    // [impl->REQ-SEAL-NO-CEREMONY-SURFACE]\n 1622|    fn ceremony_surface(&self) -> Option<(SyncSender<CtrlMsg>, SharedSend, u64, bool)> {\n*1623|        self.controller\n 1624|            .as_ref()\n 1625|            .map(|c| (c.tx.clone(), Arc::clone(&c.send), c.send.id(), c.seal_ceremony))\n 1626|    }\n...\n 1629|    /// seat — but only when that seat is the very connection this subscribe\n*1630|    /// seated (Controller / TookControl / the idempotent replay all qualify;\n 1631|    /// a Busy refusal or a viewer fallback must not touch someone else's\n 1632|    /// seat). Runs AFTER the ladder so the ladder stays capability-blind.\n 1633|    // [impl->REQ-SEAL-NO-CEREMONY-SURFACE]\n*1634|    fn stamp_controller_seal_ceremony(&mut self, conn: u64, seal_ceremony: bool) {\n*1635|        if let Some(c) = self.controller.as_mut() {\n 1636|            if c.send.id() == conn {\n 1637|                c.seal_ceremony = seal_ceremony;\n 1638|            }\n...\n 1647|\n*1648|    /// Become (or silently re-take) the interactive controller from `from_seq`.\n 1649|    ///\n*1650|    /// W1: the ring replay is now the INITIAL BATCH of a dedicated controller\n 1651|    /// writer thread (symmetric with [`add_viewer`]) — it is no longer written\n 1652|    /// inline under the log lock. The writer thread advances the authoritative\n 1653|    /// `delivered_through` cursor as it writes (D4-1: only on a successful write,\n*1654|    /// monotonic, never the unsent tail). Dropping any prior [`ControllerSink`]\n*1655|    /// ends its writer's LIVE loop when `tx` drops; bumping `controller_epoch`\n 1656|    /// additionally makes a stale deadline-evict a no-op AND stops a superseded\n 1657|    /// writer mid-INITIAL-BATCH (P1c — a writer still flushing its owned `initial`\n 1658|    /// batch ignores the dropped `tx`, so the epoch gate is what guarantees a\n 1659|    /// single live writer per connection and a monotonic on-wire seq stream,\n*1660|    /// [`REQ-HAZARD-CONTROLLER-WRITER-REORDER`]). Stamps `driven_by` to the new\n*1661|    /// controller's identity.\n 1662|    /// The engine room's rc locks against one subscribe, or `None` if they have\n 1663|    /// nothing to say (every endpoint but the reserved one, and the attaches the\n 1664|    /// locks allow). See [`crate::attach::classify_rc_lock`] for the rule.\n...\n 1750|        // (releases#203). The gate is spent once per admitted CONNECTION — the\n*1751|        // seated controller re-serving itself (gap resume, dispatcher\n 1752|        // reconstruction) is exempt — and that exemption used to sit BELOW the\n*1753|        // ticket match, where a re-serve could never reach it: a controller\n 1754|        // seated by a REDEEMED ticket re-presents that same ticket on every\n 1755|        // re-subscribe, the ticket is single-use, so the second presentation is\n 1756|        // `Unredeemed` and returned a REFUSAL that unseated the human mid-session\n...\n 1758|        // ~700ms later with `req_gen` unchanged, `engine-room-admit-unredeemed`,\n*1759|        // `session-detach was_controller=true`).\n 1760|        //\n 1761|        // THE ORDER SWAP IS NOT A WIDENING, and the reason is the same one this\n 1762|        // module already gives for why an attach GENERATION cannot be the\n...\n 1775|        // [impl->REQ-ER-BRINGUP-TOTP-GATE]\n*1776|        if !er::bringup_required(self.controller.as_ref().map(|c| c.send.id()), conn) {\n 1777|            return SeatGate::Proceed(None);\n 1778|        }\n 1779|        match ticket {\n*1780|            // A bring-up this broker admitted seats its controller without\n 1781|            // paying twice for one entry (releases#71). It is not a discount:\n 1782|            // the bounded wait for the harness can outlive the ±1 TOTP window,\n 1783|            // so re-presenting the code here could be refused by the gate that\n...\n 1791|            // discrimination an admin bring-up needs is made by the recorded\n*1792|            // verdict — a member ticket seats its controller and empowers\n 1793|            // nothing — so this is not a second credential surface: nothing a\n 1794|            // caller sends reaches it, the ticket id is opaque and the record is\n 1795|            // broker-resident. What moved is WHEN the grant is written, not who\n 1796|            // decides it: the settlement happens after the seat resolves, so a\n*1797|            // grant line describing a controller that never attached is\n 1798|            // impossible.\n 1799|            // [impl->REQ-ER-BRINGUP-SPAWNS-SESSION]\n 1800|            // [impl->REQ-ENGINEROOM-ADMIN-BRINGUP-EMPOWERS]\n...\n 1908|    /// seated on its ticket; `None` means the gate had nothing to decide (every\n*1909|    /// endpoint but the reserved one, and a seated controller re-serving\n 1910|    /// itself).\n 1911|    ///\n 1912|    /// The credential travels rather than a finished sentence because what it\n...\n 1914|    /// still refuse the attach, and a TAKE deliberately revokes the displaced\n*1915|    /// controller's empowerments as it seats the new one. Granting here would put\n 1916|    /// the write on the wrong side of that revocation.\n 1917|    Proceed(Option<SeatAdmit>),\n 1918|}\n...\n 2413|    /// which would leave an empowerment minted for a seat nobody took. And a\n*2414|    /// `--take` DROPS the displaced controller's empowerments as it seats the\n 2415|    /// taker (ADR-0052 §3a, \"a take must not hand the taker the displaced\n*2416|    /// controller's authority\") — that revocation is keyed on the session, not\n*2417|    /// on the controller, so a grant written before it is revoked by it. The\n 2418|    /// symptom is the worst kind: the operator is told they are empowered, the\n 2419|    /// record says otherwise, and only the next mode change disagrees out loud.\n 2420|    ///\n...\n 2427|    /// EMPOWERMENT is not.** They look symmetric here and are not. A grant is\n*2428|    /// keyed on the seat: a `--take` drops the displaced controller's\n 2429|    /// empowerments as it seats the taker, so the taker's grant must be minted\n 2430|    /// again and stated again or the operator is told nothing about authority\n 2431|    /// they now hold. A briefing is per-session weather, and a human who\n...\n 2436|    /// spawned session — necessary for the attach and NOT sufficient for it: the\n*2437|    /// controller is seated later, by the subscribe that redeems the admit\n 2438|    /// ticket, and a bring-up whose attach never arrived left behind a grant\n*2439|    /// line and a briefing describing a controller that never came. Both now\n 2440|    /// compose on this side of the seat, where the ratified invariant puts them.\n 2441|    // [impl->REQ-ENGINEROOM-ADMIN-BRINGUP-EMPOWERS]\n 2442|    // [impl->REQ-ER-SESSION-BRIEFING]\n...\n 2453|            outcome,\n*2454|            SubscribeOutcome::Controller | SubscribeOutcome::TookControl\n 2455|        ) {\n 2456|            return None;\n 2457|        }\n...\n 2563|    /// quiet. The ledger's own backoff gates this path, so it cannot be sprayed.\n*2564|    /// Controller-bound authority (ADR-0052 §3a): a seat change or release\n 2565|    /// DROPS every empowerment the engine room's session holds.\n 2566|    ///\n 2567|    /// Explicit, because the ADR's premise that a take \"forces a harness\n...\n 2571|    ///\n*2572|    /// LOUD on failure: an empowerment that outlives its controller is exactly\n 2573|    /// the authority leak the session scoping exists to bound, and a silent\n*2574|    /// failure here would leave the next controller holding the last one's keys.\n*2575|    // [impl->REQ-ER-CONTROLLER-BOUND-POSTURE]\n 2576|    fn on_control_edge(&self, prev_gen: Option<u64>, next_gen: Option<u64>) {\n 2577|        if spt_store::engineroom::classify_control_edge(prev_gen, next_gen)\n 2578|            == spt_store::engineroom::ControlEdge::Held\n...\n 2584|                \"ENGINE_ROOM_AUTHORITY_UNDROPPED:{}: {e} — the empowerments of the \\\n*2585|                 controller that just left this seat could not be revoked; clear \\\n 2586|                 the session's empowerments file by hand\",\n 2587|                self.endpoint\n 2588|            );\n...\n 2591|\n*2592|    fn become_controller(&mut self, sub: SharedSend, by: Option<String>, from_seq: u64, attach_gen: u64) {\n*2593|        // Drop the prior controller sink first (its writer's live loop ends when\n 2594|        // tx drops), and bump the generation so (a) an in-flight deadline-evict\n*2595|        // for the old controller can't unseat this one and (b) a prior writer\n 2596|        // still flushing its initial batch returns before its next frame.\n 2597|        // Attribution (F-039 legs b+d): the replaced conn's record names its\n*2598|        // successor, and the new conn accretes the controller role fact.\n 2599|        // [impl->REQ-CONN-POISON-ATTRIBUTION]\n*2600|        let prev_gen = self.controller.as_ref().map(|c| c.attach_gen);\n*2601|        if let Some(old) = self.controller.take() {\n 2602|            old.send.lifecycle_event(\n*2603|                \"controller-replaced\",\n 2604|                &format!(\"session={} new_conn={}\", self.session_id, sub.id()),\n 2605|            );\n 2606|        }\n 2607|        sub.describe(&format!(\n*2608|            \"controller session={} endpoint={} by={}\",\n 2609|            self.session_id,\n 2610|            if self.endpoint.is_empty() { \"-\" } else { &self.endpoint },\n 2611|            by.as_deref().unwrap_or(\"local\")\n...\n 2613|        sub.lifecycle_event(\n*2614|            \"controller-attach\",\n 2615|            &format!(\"session={} from_seq={from_seq}\", self.session_id),\n 2616|        );\n*2617|        self.controller = None;\n*2618|        let epoch = self.controller_epoch.fetch_add(1, Ordering::AcqRel) + 1;\n 2619|        // #6 (ADR-0031): a COLD attach's initial batch is a single synthesized clean\n 2620|        // repaint of the CURRENT SCREEN (at pseudo-seq `next_seq - 1`) instead of the\n 2621|        // raw ring — replaying an alt-screen TUI's raw ring into a fresh terminal is\n...\n 2635|                    // OBS RIDER Tap B (repaint): the synthesized clean-screen\n*2636|                    // controller repaint bytes (debug-gated, default OFF).\n 2637|                    spt_store::obstap::tap(self.session_id, \"repaint\", &repaint.1);\n 2638|                    (vec![repaint], true)\n 2639|                }\n...\n 2648|            };\n*2649|        let (tx, rx) = sync_channel::<CtrlMsg>(CONTROLLER_CHANNEL_DEPTH);\n 2650|        let writer_send = Arc::clone(&sub);\n 2651|        let dt = Arc::clone(&self.delivered_through);\n*2652|        let epoch_gate = Arc::clone(&self.controller_epoch);\n 2653|        let sid = self.session_id;\n 2654|        // W2 decouple: the writer publishes its in-flight-write window here so a\n 2655|        // stall-evict can distinguish a wedged writer from an idle one.\n...\n 2662|        let writer = thread::spawn(move || {\n*2663|            controller_writer(\n 2664|                sid,\n 2665|                writer_send,\n 2666|                initial,\n...\n 2675|        });\n*2676|        self.controller = Some(ControllerSink {\n 2677|            tx,\n 2678|            send: sub,\n 2679|            by,\n...\n 2690|        self.stamp_driven_by();\n*2691|        // The seat changed hands (or was filled): the incoming controller does\n 2692|        // not inherit the outgoing one's authority. An equal-generation re-serve\n 2693|        // is the SAME lease and keeps it.\n 2694|        self.on_control_edge(prev_gen, Some(attach_gen));\n...\n 2697|    /// AUTHORITATIVE lease revoke (ADR-0044 decision 2,\n*2698|    /// REQ-CONTROLLER-LEASE-IDENTITY): mark the incumbent controller's sink as\n 2699|    /// revoked-by-`by` BEFORE the caller replaces the slot. The sink's writer\n 2700|    /// thread then emits the terminal [`KIND_DISPLACED`] as its FINAL frame on\n 2701|    /// exit — delivery no longer depends on a `try_send` into a possibly-Full\n...\n 2705|    /// stream is the terminal signal either way. Notification is best-effort;\n*2706|    /// REVOCATION is not: the caller's `become_controller` moves the slot (the\n 2707|    /// input fence keys on it) regardless of what the incumbent ever hears.\n*2708|    // [impl->REQ-CONTROLLER-LEASE-IDENTITY]\n 2709|    fn mark_revoked(&self, by: &str) {\n*2710|        if let Some(old) = &self.controller {\n 2711|            if let Ok(mut g) = old.revoked_by.lock() {\n 2712|                *g = Some(by.to_string());\n 2713|            }\n 2714|            old.send.lifecycle_event(\n*2715|                \"controller-revoked\",\n 2716|                &format!(\"session={} taken_by={by}\", self.session_id),\n 2717|            );\n 2718|        }\n...\n 2720|\n*2721|    /// The drain timed out the controller's `send_timeout` and the epoch guard\n*2722|    /// still matched ([`mark_controller_gone`], W1). The writer was BLOCKED past\n 2723|    /// the deadline -- it is NOT established to have exited or dropped its `tx`,\n 2724|    /// so this names the deadline the drain observed and not the writer's fate.\n 2725|    /// This is the edge releases#49/#267 Q1 asks about: a cleanly-closed\n...\n 2729|    const CLOSE_SEND_DEADLINE: &'static str = \"send_deadline\";\n*2730|    /// The wedge evict released the slot ([`stall_evict_controller`], inline stamp\n 2731|    /// mode). `BRAIN_SUBSCRIBER_STALL_EVICT` already names the wedge; this names\n 2732|    /// the slot transition that followed it.\n 2733|    // [impl->REQ-REFRESH-WAIT-ATTRIBUTION]\n 2734|    const CLOSE_STALL_EVICT: &'static str = \"stall_evict\";\n*2735|    /// A subscriber detached and it happened to be the controller\n 2736|    /// ([`detach_if`]) — an ordinary departure, not a wedge and not a kill.\n 2737|    // [impl->REQ-REFRESH-WAIT-ATTRIBUTION]\n 2738|    const CLOSE_DETACH: &'static str = \"detach\";\n 2739|\n*2740|    /// Detach the controller and UNLATCH the perch — the ONE unlatch path shared\n*2741|    /// by W1's deadline-evict ([`mark_controller_gone`]), [`detach_if`], and (per\n 2742|    /// the W1↔W5 synergy) W5's reconcile self-heal. Dropping the sink ends its\n 2743|    /// writer thread; clearing the slot re-stamps `driven_by` to `None` (so an\n*2744|    /// `ONLINE+CONTROLLED` latch cannot outlive a gone controller).\n 2745|    ///\n 2746|    /// `cause` is the CALLER'S OWN knowledge of which edge closed the slot, and it\n 2747|    /// is a parameter for the reason releases#49/#267 needed this line at all: at\n...\n 2751|    // [impl->REQ-REFRESH-WAIT-ATTRIBUTION]\n*2752|    fn clear_controller(&mut self, cause: &'static str) {\n*2753|        let prev_gen = self.controller.as_ref().map(|c| c.attach_gen);\n 2754|        // Captured BEFORE the take: after it there is no sink left to ask.\n*2755|        let conn = self.controller.as_ref().map(|c| c.send.id());\n*2756|        if self.controller.take().is_some() {\n 2757|            // ONLY ON THE REAL TRANSITION (releases#49/#267 Q1). The epoch-guarded\n 2758|            // no-op above returns without taking anything, and a line there would\n 2759|            // report a closure that did not happen — which is exactly the confusion\n...\n 2764|                spt_proto::emit_line_err!(\n*2765|                    \"CONTROLLER_SLOT_CLOSED:{} conn={} cause={cause} {} — the controller slot \\\n 2766|                     was occupied and is now empty [REQ-REFRESH-WAIT-ATTRIBUTION]\",\n 2767|                    self.session_id,\n 2768|                    conn.unwrap_or(0),\n...\n 2778|\n*2779|    /// B-2 (REQ-CONTROLLER-LIVENESS-REAP): lazily drop a controller whose WRITER\n 2780|    /// THREAD has EXITED — the record-vs-truth honesty step, run on the authoritative\n*2781|    /// `KIND_SESSIONS` poll BEFORE `controller_by`/`has_controller` are read. The\n*2782|    /// writer exits when a socket write fails (a severed / keepalive-killed controller\n 2783|    /// conn) or its `tx` drops; while it is parked on `rx.recv()` (a LIVE but idle\n*2784|    /// controller) it is NOT finished, so a live controller is never touched (the\n 2785|    /// no-false-reap guard). Reaping here makes the broker report honest truth so the\n 2786|    /// EXISTING [`converge_perch_stamps`] clears the now-stale info.json\n 2787|    /// `driven_by`/`controlled` stamp — no new detach path, no new query.\n...\n 2790|    /// NO file I/O (the lock-across-effect discipline, KH 7.12/5.16). It drops the\n*2791|    /// sink in-memory ONLY (like [`become_controller`]'s prior-drop) and does NOT\n 2792|    /// [`stamp_driven_by`] — the OFF-lock `converge_perch_stamps` writes the honest\n 2793|    /// stamp. Returns true iff it reaped.\n 2794|    ///\n...\n 2796|    /// flowing when the conn severed) or whose detach was dropped by a prior wedge\n*2797|    /// (the B-1 poison class). A TRULY IDLE severed controller (writer parked on recv,\n 2798|    /// no output, conn not yet EOF'd) stays `is_finished() == false` and converges\n 2799|    /// only once output resumes or the conn EOFs (→ `handle_conn` detach) — that\n 2800|    /// harder active-probe case is the reserved [`REQ-HAZARD-DRIVEN-BY-IDLE-REMOTE-EVICT`],\n...\n 2802|    /// never latches `driven_by`.\n*2803|    // [impl->REQ-CONTROLLER-LIVENESS-REAP]\n*2804|    fn reap_dead_controller(&mut self) -> bool {\n 2805|        let dead = self\n*2806|            .controller\n 2807|            .as_ref()\n 2808|            .is_some_and(|c| c._writer.is_finished());\n 2809|        if dead {\n...\n 2811|            // (info.json I/O under the log lock is the KH 7.12 hazard). converge does it.\n*2812|            self.controller = None;\n 2813|            return true;\n 2814|        }\n*2815|        // W2 decouple (REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE): a controller whose\n 2816|        // writer is BLOCKED (not exited) on a suspended brain conn past\n 2817|        // BRAIN_WRITE_DEADLINE is also reaped here, so the honest `KIND_SESSIONS`\n 2818|        // truth + the off-lock converge_perch_stamps release its stale `driven_by`\n...\n 2820|        // the multi-session snapshot closure (no I/O under the shared sessions lock).\n*2821|        self.stall_evict_controller(Instant::now(), StampMode::Deferred)\n 2822|    }\n 2823|\n*2824|    /// Deadline-evict path (W1): the drain observed a fully-stalled controller and\n 2825|    /// timed out its `send_timeout`. Detach it — but ONLY if it is still the same\n*2826|    /// controller generation the drain saw (`epoch`), so a NEWER controller that\n 2827|    /// attached during the stall is never unseated.\n*2828|    fn mark_controller_gone(&mut self, epoch: u64) {\n*2829|        if self.controller_epoch.load(Ordering::Acquire) == epoch {\n*2830|            self.clear_controller(Self::CLOSE_SEND_DEADLINE);\n 2831|        }\n 2832|    }\n 2833|\n*2834|    /// Whether the current controller's writer is WEDGED — blocked inside a single\n 2835|    /// socket write for at least `deadline` (the brain-subscriber conn suspended /\n*2836|    /// black-holed). An IDLE controller (writer parked on `rx.recv()`, nothing to\n*2837|    /// write) reads NOT stalled: its [`ControllerSink::write_blocked_since`] is `None`.\n 2838|    /// Pure over the injected `now` + `deadline` so the unit tests it with no thread.\n 2839|    // [impl->REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]"}],"details":{"scopePath":"crates/spt-daemon/src/broker.rs","searchPath":"C:\\Users\\decid\\Documents\\projects\\spt-core\\crates\\spt-daemon\\src\\broker.rs","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","matchCount":200,"fileCount":1,"files":["crates/spt-daemon/src/broker.rs"],"fileMatches":[{"path":"crates/spt-daemon/src/broker.rs","count":200}],"truncated":true,"perFileLimitReached":200,"displayContent":"  160│\n *161│/// The controller live-frame channel depth (frames). Deeper than a viewer's\n *162│/// queue because the controller is the AUTHORITATIVE view — unlike a viewer it is\n  163│/// NOT evicted on a SINGLE full `try_send` (W1 / REQ-HAZARD-INJECT-CONTROL-COEXIST).\n  164│/// b4: the drain does ONE non-blocking `try_send` per chunk and NEVER blocks; a\n *165│/// merely-slow controller DROPS chunks + falls behind the ring, and only a channel\n *166│/// that stays full for [`CONTROLLER_WRITE_DEADLINE`] (tracked via\n *167│/// [`ControllerSink::last_ok`]) is detached. Generous so a momentary slow consumer\n  168│/// never trips the deadline.\n *169│const CONTROLLER_CHANNEL_DEPTH: usize = 4096;\n  170│\n *171│/// How long a CONTROLLER's channel may stay FULL (no successful handoff since\n *172│/// [`ControllerSink::last_ok`]) before the drain declares it stalled and detaches it\n *173│/// (W1, refined by b4 / REQ-HAZARD-VIEWER-STARVE-UNDER-CONTROLLER-BACKPRESSURE). This\n *174│/// BOUNDS the wedge: a backed-up controller falls behind for at most this long before\n  175│/// eviction — never forever (the v0.12.x root) — and the drain NEVER blocks or sleeps\n  176│/// (a single off-lock `try_send` per chunk; see [`OutputLog::append`]). A live but\n  177│/// merely slow operator lands a frame before this elapses (resetting `last_ok`) and is\n  178│/// untouched.\n *179│const CONTROLLER_WRITE_DEADLINE: Duration = Duration::from_secs(5);\n  180│\n  181│/// The RESIZE BARRIER's drain-quiescence window (REQ-RC-RESIZE-GEOMETRY-EPOCH):\n  182│/// how long the appended high-water must hold STILL before the surface resize is\n     │...\n  202│/// The cap on bytes HELD BACK from the grid across one resize transition. The ring\n *203│/// and the controller/viewer fan-out are never gated (delivery keeps running at full\n  204│/// speed during a transition — only the GRID parse is deferred), so this bounds the\n  205│/// transition's extra memory alone. Overflow drops from the grid parse ONLY, loudly,\n  206│/// and self-heals on the next output + repaint exactly like the poison-recover ring\n     │...\n  209│\n *210│/// How long a controller's writer thread may sit BLOCKED inside a single socket\n  211│/// write — the brain subscriber suspended / black-holed — before a take/reattach\n  212│/// or the `KIND_SESSIONS` reap stall-evicts it (LIFECYCLE-TRUTH W2,\n  213│/// REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE). Distinct from and 3× longer than\n *214│/// [`CONTROLLER_WRITE_DEADLINE`] on PURPOSE (doyle ruling Q1): that 5 s bound\n  215│/// protects the fan-out against ONE per-client conn falling behind and is\n  216│/// output-driven (the drain observes a Full channel); THIS bound is measured\n  217│/// full-with-**zero writer progress** and covers the case the output-driven path\n  218│/// can't — an IDLE session whose brain-subscriber conn is suspended produces no\n *219│/// output, so `append`/`mark_controller_gone` never runs, and the stalled\n *220│/// controller would otherwise latch `driven_by` forever (reattach refused, `--take`\n  221│/// blocked). The brain aggregates every session + psyche + WAN feed, so its drain\n  222│/// is bursty; 15 s tolerates a normal burst-stall while still bounding the wedge.\n  223│/// A slow-but-DRAINING brain never trips it — the clock only advances while a write\n *224│/// is genuinely in flight past the deadline (see [`ControllerSink::write_blocked_since`]).\n  225│// [impl->REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]\n  226│const BRAIN_WRITE_DEADLINE: Duration = Duration::from_secs(15);\n  227│\n     │...\n  229│/// suspend-brain int rig shrinks it via `SPT_BRAIN_WRITE_DEADLINE_MS` so the\n *230│/// structural outcome (a wedged controller is stall-evicted → take/reattach frees\n  231│/// the slot) is provable in ~1 s instead of 15 (mirrors the\n  232│/// `SPT_INJECT_COMMIT_DEADLINE_MS` env-knob precedent). Grants test wall-clock\n  233│/// ONLY — never relaxes an assertion.\n     │...\n  518│/// The RC-origin input-fence verdict (ADR-0044 decision 3,\n *519│/// REQ-INPUT-CONTROLLER-FENCE) — PURE over the request's minter tag and the\n *520│/// sender-conn's controller fact, so the scope table is unit-tested: ONLY an\n *521│/// `Minter::Rc`-tagged input from a NON-controller conn is rejected. Shell\n  522│/// injection (`Minter::Shell`), legacy/untagged senders (`None`/`Legacy` — the\n  523│/// N-1 wire), and every other minter pass exactly as before; the connection\n  524│/// identity is the lease surrogate (ADR-0044 decision 4), which holds for\n  525│/// mixed-version attaches too — every rc keystroke reaches the broker through\n *526│/// its own serve worker's conn, and `become_controller` moves the slot to the\n  527│/// live lease's conn atomically with the takeover.\n *528│// [impl->REQ-INPUT-CONTROLLER-FENCE]\n *529│fn rc_input_rejected(minter: Option<Minter>, sender_is_controller: bool) -> bool {\n *530│    minter == Some(Minter::Rc) && !sender_is_controller\n  531│}\n  532│\n  533│/// Whether a stall-evict re-stamps the perch INLINE (take/reattach path, already\n     │...\n  620│/// buggy translation binary can therefore stall operator input for at most this\n *621│/// long. Mirrors [`CONTROLLER_WRITE_DEADLINE`]'s order (5 s).\n  622│const INJECT_COMMIT_DEADLINE: Duration = Duration::from_secs(5);\n  623│\n  624│/// The effective commit deadline. Production is the const above; the int gate\n     │...\n  821│\n *822│/// The interactive CONTROLLER's isolated sink (W1 / REQ-HAZARD-INJECT-CONTROL-COEXIST).\n  823│/// Symmetric with [`ViewerSink`] — a dedicated writer thread does the BLOCKING\n *824│/// socket write so a backed-up controller can never stall the session drain\n *825│/// thread (the v0.12.x wedge: the controller write was inline on the drain,\n  826│/// holding `Mutex<OutputLog>`). Unlike a viewer it is AUTHORITATIVE: its writer\n  827│/// advances `delivered_through` on each successful write, and it is never evicted\n *828│/// on a single full `try_send` — only after [`CONTROLLER_WRITE_DEADLINE`].\n *829│struct ControllerSink {\n *830│    /// The bounded live-frame queue to the controller's writer thread.\n  831│    tx: SyncSender<CtrlMsg>,\n *832│    /// The controller's IPC send half — kept for identity matching on detach and\n  833│    /// for the `exit`-frame fan-out ([`OutputLog::all_sinks`]).\n  834│    send: SharedSend,\n *835│    /// The controller's ATTACH GENERATION (ADR-0038 Amendment fix 6,\n  836│    /// REQ-SERVE-OWNERSHIP-GENERATION): operator-minted wall-ms from the wire\n  837│    /// Request (`0` = local / N-1). A same-identity re-take with an OLDER\n  838│    /// generation is a replayed stale Request — refused; a stale worker's\n *839│    /// late detach (older generation) never clears a newer controller.\n  840│    attach_gen: u64,\n  841│    /// The `from_seq` this seat was ESTABLISHED with — the fourth element of the\n  842│    /// idempotent-replay key (ADR-0047 decision 3, REQ-ATTACH-IDEMPOTENT-REPLAY).\n     │...\n  844│    /// the establish already honored (the dispatcher's double-serve); one carrying\n *845│    /// a DIFFERENT `from_seq` is the controller GAP RESUME re-fetching dropped\n *846│    /// frames from its frozen floor (REQ-HAZARD-CONTROLLER-GAP-RESUME), which must\n  847│    /// be re-served — suppressing it would turn a recoverable gap into a\n  848│    /// truncation report.\n  849│    establish_from_seq: u64,\n *850│    /// The controller's identity: `Some(node)` = a remote operator, `None` =\n  851│    /// local (the brain's spawn-reader / resume). Ownership keys on this.\n  852│    by: Option<String>,\n  853│    /// The seal-ceremony CAPABILITY the attaching client declared on its\n     │...\n  862│    seal_ceremony: bool,\n *863│    /// Time of the last SUCCESSFUL handoff (`try_send` Ok) to this controller's\n *864│    /// writer channel (b4, REQ-HAZARD-VIEWER-STARVE-UNDER-CONTROLLER-BACKPRESSURE).\n  865│    /// The drain does a SINGLE non-blocking `try_send` per chunk and NEVER sleeps;\n *866│    /// a channel that stays Full for [`CONTROLLER_WRITE_DEADLINE`] PAST this stamp\n *867│    /// evicts the stalled controller (bounded wedge preserved), while a slow-but-\n *868│    /// alive controller just drops the chunk + falls behind the ring (resume-from-\n *869│    /// floor). Seeded `now` at [`OutputLog::become_controller`].\n  870│    last_ok: Instant,\n  871│    /// The writer's IN-FLIGHT socket-write marker (LIFECYCLE-TRUTH W2,\n *872│    /// REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE): the `controller_writer` sets it to\n  873│    /// `Some(Instant::now())` immediately BEFORE it enters a blocking `write_frame`\n  874│    /// and back to `None` the instant the write returns. A reader\n *875│    /// ([`OutputLog::controller_stalled`]) sees a `Some` older than\n  876│    /// [`BRAIN_WRITE_DEADLINE`] as a writer WEDGED on a suspended/black-holed brain\n  877│    /// conn and stall-evicts it — the signal `last_ok` (channel-handoff time) cannot\n *878│    /// give: an IDLE controller (writer parked on `rx.recv()`, `None` here) and a\n  879│    /// BLOCKED one (writer stuck in the socket write, `Some` here) both leave\n  880│    /// `last_ok` stale. `Arc<Mutex<…>>` so the reader observes the writer's live\n  881│    /// state; the mutex is taken only around the brief set/clear and the read, never\n     │...\n  888│    /// AUTHORITATIVE-revoke marker (ADR-0044 decision 2,\n *889│    /// REQ-CONTROLLER-LEASE-IDENTITY): set to the taker's identity by a\n  890│    /// distinct-lease `Take` BEFORE the slot is replaced. The writer thread\n  891│    /// reads it on its exit paths and emits the terminal [`KIND_DISPLACED`]\n  892│    /// frame as its FINAL write — so the revoke lands even when the live\n     │...\n  919│/// assigns a `seq` and RINGS every chunk — the ring stays the scrollback/resume\n *920│/// authority for raw history — but delivery to the controller and the viewers is\n  921│/// SUPPRESSED (a skip, never a block: the drain runs at full PTY speed, hazard\n  922│/// 7.12 intact), and the chunk is held back from `grid.advance`, tagged with the\n  923│/// geometry it was emitted under. [`OutputLog::commit_resize`] /\n     │...\n  997│/// A bounded, sequenced output log for one session — the gapless-resume buffer\n *998│/// AND the W2.5 controller/viewer fan-out hub (REQ-RCVIEW-1).\n  999│///\n 1000│/// The session's single drain thread is the only producer ([`OutputLog::append`]).\n*1001│/// Output fans out to ONE interactive `controller` (the authoritative,\n 1002│/// cursor-advancing, blocking path) plus ANY NUMBER of read-only `viewers` (each\n 1003│/// an isolated bounded queue + writer thread, evicted on overflow). Every\n*1004│/// controller send and viewer `try_send` happens while the enclosing\n 1005│/// `Mutex<OutputLog>` is held, so frames reach each sink in strict `seq` order.\n 1006│struct OutputLog {\n 1007│    session_id: u64,\n     │...\n 1010│    cap_chunks: usize,\n*1011│    /// The single interactive controller's isolated sink (W1: a dedicated writer\n 1012│    /// thread + bounded channel + authoritative cursor). `None` ⇒ undriven. The\n*1013│    /// controller's identity (`Some(node)` remote / `None` local) lives in\n*1014│    /// [`ControllerSink::by`] and keys ownership — a same-`by` re-subscribe\n 1015│    /// silently re-takes (no `Displaced`, the gate-#7 self-kick guard); only a\n 1016│    /// DIFFERENT remote's `Take` displaces.\n*1017│    controller: Option<ControllerSink>,\n*1018│    /// Monotonic controller generation, shared with the live `controller_writer`\n*1019│    /// thread (`Arc<AtomicU64>`). Bumped on every [`become_controller`]; the\n*1020│    /// drain's deadline-evict ([`mark_controller_gone`]) carries the epoch it\n*1021│    /// observed and only detaches if it still matches — so a slow-controller\n*1022│    /// timeout can never unseat a NEWER controller that attached meanwhile. The\n 1023│    /// writer snapshots its own epoch and returns the moment this value moves past\n 1024│    /// it, so a SUPERSEDED writer flushes no further frames (P1c, single live\n*1025│    /// writer per connection, [`REQ-HAZARD-CONTROLLER-WRITER-REORDER`]).\n*1026│    controller_epoch: Arc<AtomicU64>,\n 1027│    /// The read-only viewers, keyed by a broker-minted viewer id.\n 1028│    viewers: HashMap<u64, ViewerSink>,\n 1029│    /// Monotonic viewer-id source.\n     │...\n 1034│    endpoint: String,\n*1035│    /// Last-known PTY size — seeded at spawn, updated on a controller resize.\n 1036│    /// Pushed to a viewer on attach + to all viewers on resize (the letterbox\n 1037│    /// feed); the broker holds it since the PTY exposes no size getter.\n 1038│    size: (u16, u16),\n 1039│    /// Cursor-of-record (restoration D4-1, ADR-0018 Q6): the next seq a\n*1040│    /// re-attaching CONTROLLER still needs (= highest seq **successfully written**\n*1041│    /// to the controller, +1). It is **log** state, not subscriber state, so it\n 1042│    /// **survives `detach_if`** — a dead brain leaves its resume point behind, and\n 1043│    /// a cold-starting brain reads it (via `KIND_SESSIONS` → `resume_seq`) instead\n 1044│    /// of replaying the whole ring from 0. **Viewer writes NEVER move it**\n*1045│    /// (REQ-RCVIEW-1: the brain-resume cursor tracks the controller only).\n 1046│    ///\n 1047│    /// **Advanced only on a successful socket write** (`Ok`), and **monotonic**\n 1048│    /// (never decreases) — see [`advanced_cursor`]. Both invariants are\n 1049│    /// load-bearing for at-least-once output: advancing past a *failed* write to a\n*1050│    /// dying-but-not-yet-detached controller would make resume skip that chunk\n 1051│    /// forever (at-least-once → at-most-once); decreasing on a deliberate\n 1052│    /// `attach(sub, 0)` rewind-replay would reset the resume point under a live\n 1053│    /// reader.\n 1054│    ///\n*1055│    /// W1: now an `Arc<AtomicU64>` shared with the controller's writer thread,\n 1056│    /// which advances it (via [`advanced_cursor`], CAS-monotonic) as it writes —\n*1057│    /// the cursor moved off the drain thread with the controller delivery itself.\n 1058│    delivered_through: Arc<AtomicU64>,\n 1059│    /// The server-side **render grid** (#6 / REQ-BROKER-SCREEN-GRID, ADR-0031):\n 1060│    /// the authoritative CURRENT SCREEN, interpreted from the same byte stream the\n 1061│    /// ring records. Fed on the drain path in [`append`], AFTER the fan-out (so it\n*1062│    /// never adds latency to controller/viewer delivery — hazard 7.12). A cold\n 1063│    /// attach's initial batch is a single synthesized [`ScreenGrid::render_repaint`]\n 1064│    /// instead of the raw ring, so replaying an alt-screen TUI into a fresh terminal\n 1065│    /// no longer corrupts scrollback. Single-writer (the one drain thread), mirroring\n     │...\n 1223│    if control || viewer {\n*1224│        // The two IN-MEMORY reap paths (`reap_dead_controller`, the stall-evict)\n 1225│        // deliberately do not stamp under the log lock, so their attachment edge\n*1226│        // becomes true HERE and nowhere else — signal it, or a controller that\n 1227│        // died with its conn would leave a stale attachment frame standing until\n 1228│        // the next sweep. (releases#44)\n 1229│        // [impl->REQ-ATTACH-LINK-PUSH]\n     │...\n 1235│        // The THIRD engine-room drop edge (ADR-0052 §3a). The two in-memory\n*1236│        // reap paths (`reap_dead_controller`, the stall-evict) deliberately do\n*1237│        // NOT stamp under the log lock, so they never reach `clear_controller`;\n 1238│        // this off-lock converge is where their release becomes true. Keying on\n*1239│        // the recorded truth going FALSE means a controller that died with the\n 1240│        // conn revokes its authority exactly like one that detached cleanly.\n*1241│        // [impl->REQ-ER-CONTROLLER-BOUND-POSTURE]\n 1242│        if !want_controlled {\n 1243│            if let Err(e) = spt_store::engineroom::drop_authority_at_edge(endpoint) {\n 1244│                spt_proto::emit_line_err!(\"ENGINE_ROOM_AUTHORITY_UNDROPPED:{endpoint}: {e}\");\n     │...\n 1258│/// processed sequentially by that conn's dispatch loop, so a drag-resize can never\n*1259│/// overlap itself — but the controller SLOT can move (`serve_attach` forwards a\n 1260│/// viewport's `Resize` on whichever stream currently holds control), so a\n 1261│/// takeover landing inside another transition's window passes the\n*1262│/// controller-exclusive gate and arrives second. Refusing it would hand the\n 1263│/// operator a hard resize failure for a transient condition; the geometry the user\n 1264│/// asked for last is the geometry they want, so the second resize simply WAITS its\n 1265│/// turn and then runs as a normal transition (last writer wins, no coalescing\n     │...\n 1374│            cap_chunks,\n*1375│            controller: None,\n*1376│            controller_epoch: Arc::new(AtomicU64::new(0)),\n 1377│            viewers: HashMap::new(),\n 1378│            next_viewer_id: 0,\n 1379│            endpoint,\n     │...\n 1404│    /// irrecoverable-behind detection untouched\n*1405│    /// (REQ-HAZARD-CONTROLLER-GAP-RESUME / B2). `None` also when nothing has\n 1406│    /// been produced yet (`next_seq == 0`) → an empty initial batch, so the\n 1407│    /// first live frame is the contiguous seq 0 (the spawn seat, unchanged).\n 1408│    ///\n     │...\n 1444│    /// Append a freshly-produced output chunk: assign the next `seq`, retain it in\n*1445│    /// the bounded ring, then FAN OUT to viewers AND the controller — both a SINGLE\n 1446│    /// non-blocking `try_send`, so the drain NEVER blocks/sleeps here. During a\n 1447│    /// resize transition the fan-out is SUPPRESSED wholesale (still a skip, never a\n 1448│    /// block — see the barrier note in the body,\n 1449│    /// REQ-RC-RESIZE-PRESENTATION-BARRIER). Returns\n*1450│    /// `Some(epoch)` when the controller must be EVICTED (stalled past the deadline,\n*1451│    /// or its writer is gone) so the drain can `mark_controller_gone(epoch)` OFF the\n 1452│    /// log lock; `None` otherwise.\n 1453│    ///\n*1454│    /// b4 (REQ-HAZARD-VIEWER-STARVE-UNDER-CONTROLLER-BACKPRESSURE): W1 moved the\n*1455│    /// controller SOCKET WRITE off the drain (the `controller_writer` thread) but\n*1456│    /// left the HANDOFF as an inline sleep-poll (`ControllerJob::deliver`, up to\n*1457│    /// `CONTROLLER_WRITE_DEADLINE`) ON the drain — so a controller draining slower\n 1458│    /// than the PTY floods filled its channel, the inline poll throttled the drain,\n 1459│    /// and the viewer fan-out below stopped running → a concurrent `rc --view`\n*1460│    /// VIEWER starved. Now the controller handoff is ONE non-blocking `try_send`\n 1461│    /// exactly like a viewer: the drain runs at full PTY speed regardless of\n*1462│    /// controller speed. A slow-but-alive controller DROPS the chunk + falls behind\n*1463│    /// the ring (resume-from-floor); a controller Full for `CONTROLLER_WRITE_DEADLINE`\n*1464│    /// since its last successful handoff (`ControllerSink::last_ok`) is evicted\n 1465│    /// (bounded wedge preserved). The writer still does the blocking socket write +\n 1466│    /// advances the authoritative CONTIGUOUS cursor. Called only from the session's\n 1467│    /// single drain thread.\n 1468│    // [impl->REQ-HAZARD-INJECT-CONTROL-COEXIST]\n*1469│    // [impl->REQ-HAZARD-VIEWER-STARVE-UNDER-CONTROLLER-BACKPRESSURE]\n 1470│    #[must_use]\n 1471│    /// Coherence clamp for a poison-RECOVERED ring (B-1, doyle ruling; run only by\n 1472│    /// [`recover_log`] on the recovery path). A panic mid-[`append`] can leave the\n     │...\n 1538│        // took its seq and rang above — raw history and the resume floor stay\n*1539│        // intact — but the controller handoff and the viewer fan-out are\n 1540│        // SUPPRESSED, a SKIP and never a block or a sleep (hazard 7.12: the drain\n 1541│        // returns at full PTY speed). A live attached terminal must never render\n 1542│        // old- and new-geometry differentials across a viewport it has already\n 1543│        // resized; commit/abort substitute ONE synthesized sync frame instead.\n*1544│        // The suppressed window is exempt from the controller's Full-eviction\n 1545│        // deadline by construction here (no handoff is attempted, so the deadline\n 1546│        // is never evaluated) and `last_ok` is re-armed at commit/abort so the\n*1547│        // window's width can never be charged to the controller afterwards. The\n 1548│        // chunk is HELD for the grid, tagged with the geometry it was emitted\n 1549│        // under, and replayed at that geometry on commit\n 1550│        // (REQ-RC-RESIZE-GEOMETRY-EPOCH, unchanged).\n     │...\n 1562│        self.push_frame_to_viewers(&frame);\n*1563│        // Controller: the SAME non-blocking discipline as a viewer (b4). One\n 1564│        // `try_send`, never a sleep. Ok → record the handoff time. Full → DROP (a\n*1565│        // slow controller falls behind the ring) unless it has been Full for the\n 1566│        // whole deadline since its last success → EVICT. Disconnected (writer gone)\n*1567│        // → EVICT. The drain off-lock turns `Some(epoch)` into `mark_controller_gone`.\n 1568│        let now = Instant::now();\n*1569│        let epoch = self.controller_epoch.load(Ordering::Acquire);\n*1570│        let result = if let Some(c) = self.controller.as_mut() {\n 1571│            match c.tx.try_send(CtrlMsg::Output(seq, frame)) {\n 1572│                Ok(()) => {\n 1573│                    c.last_ok = now;\n     │...\n 1576│                Err(TrySendError::Full(_)) => {\n*1577│                    if now.duration_since(c.last_ok) >= CONTROLLER_WRITE_DEADLINE {\n 1578│                        Some(epoch) // stalled past the deadline → bounded detach\n 1579│                    } else {\n*1580│                        None // drop this chunk; controller falls behind the ring\n 1581│                    }\n 1582│                }\n 1583│                Err(TrySendError::Disconnected(_)) => Some(epoch),\n     │...\n 1588│        // Advance the render grid AFTER the fan-out — it must never sit between the\n*1589│        // drain and the controller/viewer `try_send` (REQ-HAZARD-INJECT-CONTROL-COEXIST\n 1590│        // 7.12). The grid is single-writer (this one drain thread); `render_repaint`\n 1591│        // snapshots it under the same short log lock only when an attach requests it.\n 1592│        // (The in-transition path returned above: held chunks are replayed into the\n     │...\n 1599│\n*1600│    /// The current controller's identity (`Some(node)` remote / `None` local),\n 1601│    /// or `None` when undriven. The `by`-keyed ownership logic in\n 1602│    /// [`resolve_subscribe`] reads this.\n*1603│    fn controller_by(&self) -> Option<&str> {\n*1604│        self.controller.as_ref().and_then(|c| c.by.as_deref())\n 1605│    }\n 1606│\n*1607│    /// Whether ANY controller currently drives this session (the `controlled`\n*1608│    /// any-controller truth — REQ-HAZARD-CONTROL-STAMP-CONVERGENCE). Snapshotted under\n 1609│    /// the log lock, then converged onto the perch OFF the lock.\n*1610│    fn has_controller(&self) -> bool {\n*1611│        self.controller.is_some()\n 1612│    }\n 1613│\n*1614│    /// Snapshot the seal-ceremony SURFACE (WAX-SEAL W2): the live controller\n 1615│    /// seat's writer queue + conn + declared capability, read under the log\n 1616│    /// lock — the BROKER'S SEAT TABLE, never the perch info stamp (the stamp\n 1617│    /// lags the seat decision: KH 7.15 / info.rs terminal_normalize, and\n*1618│    /// `SessionInfo.controller_by == None` is ambiguous between empty and\n*1619│    /// local). `None` = no controller at all. The `SharedSend` rides along for\n 1620│    /// the bounded direct-write degrade the exit fan-out uses.\n 1621│    // [impl->REQ-SEAL-NO-CEREMONY-SURFACE]\n 1622│    fn ceremony_surface(&self) -> Option<(SyncSender<CtrlMsg>, SharedSend, u64, bool)> {\n*1623│        self.controller\n 1624│            .as_ref()\n 1625│            .map(|c| (c.tx.clone(), Arc::clone(&c.send), c.send.id(), c.seal_ceremony))\n 1626│    }\n     │...\n 1629│    /// seat — but only when that seat is the very connection this subscribe\n*1630│    /// seated (Controller / TookControl / the idempotent replay all qualify;\n 1631│    /// a Busy refusal or a viewer fallback must not touch someone else's\n 1632│    /// seat). Runs AFTER the ladder so the ladder stays capability-blind.\n 1633│    // [impl->REQ-SEAL-NO-CEREMONY-SURFACE]\n*1634│    fn stamp_controller_seal_ceremony(&mut self, conn: u64, seal_ceremony: bool) {\n*1635│        if let Some(c) = self.controller.as_mut() {\n 1636│            if c.send.id() == conn {\n 1637│                c.seal_ceremony = seal_ceremony;\n 1638│            }\n     │...\n 1647│\n*1648│    /// Become (or silently re-take) the interactive controller from `from_seq`.\n 1649│    ///\n*1650│    /// W1: the ring replay is now the INITIAL BATCH of a dedicated controller\n 1651│    /// writer thread (symmetric with [`add_viewer`]) — it is no longer written\n 1652│    /// inline under the log lock. The writer thread advances the authoritative\n 1653│    /// `delivered_through` cursor as it writes (D4-1: only on a successful write,\n*1654│    /// monotonic, never the unsent tail). Dropping any prior [`ControllerSink`]\n*1655│    /// ends its writer's LIVE loop when `tx` drops; bumping `controller_epoch`\n 1656│    /// additionally makes a stale deadline-evict a no-op AND stops a superseded\n 1657│    /// writer mid-INITIAL-BATCH (P1c — a writer still flushing its owned `initial`\n 1658│    /// batch ignores the dropped `tx`, so the epoch gate is what guarantees a\n 1659│    /// single live writer per connection and a monotonic on-wire seq stream,\n*1660│    /// [`REQ-HAZARD-CONTROLLER-WRITER-REORDER`]). Stamps `driven_by` to the new\n*1661│    /// controller's identity.\n 1662│    /// The engine room's rc locks against one subscribe, or `None` if they have\n 1663│    /// nothing to say (every endpoint but the reserved one, and the attaches the\n 1664│    /// locks allow). See [`crate::attach::classify_rc_lock`] for the rule.\n     │...\n 1750│        // (releases#203). The gate is spent once per admitted CONNECTION — the\n*1751│        // seated controller re-serving itself (gap resume, dispatcher\n 1752│        // reconstruction) is exempt — and that exemption used to sit BELOW the\n*1753│        // ticket match, where a re-serve could never reach it: a controller\n 1754│        // seated by a REDEEMED ticket re-presents that same ticket on every\n 1755│        // re-subscribe, the ticket is single-use, so the second presentation is\n 1756│        // `Unredeemed` and returned a REFUSAL that unseated the human mid-session\n     │...\n 1758│        // ~700ms later with `req_gen` unchanged, `engine-room-admit-unredeemed`,\n*1759│        // `session-detach was_controller=true`).\n 1760│        //\n 1761│        // THE ORDER SWAP IS NOT A WIDENING, and the reason is the same one this\n 1762│        // module already gives for why an attach GENERATION cannot be the\n     │...\n 1775│        // [impl->REQ-ER-BRINGUP-TOTP-GATE]\n*1776│        if !er::bringup_required(self.controller.as_ref().map(|c| c.send.id()), conn) {\n 1777│            return SeatGate::Proceed(None);\n 1778│        }\n 1779│        match ticket {\n*1780│            // A bring-up this broker admitted seats its controller without\n 1781│            // paying twice for one entry (releases#71). It is not a discount:\n 1782│            // the bounded wait for the harness can outlive the ±1 TOTP window,\n 1783│            // so re-presenting the code here could be refused by the gate that\n     │...\n 1791│            // discrimination an admin bring-up needs is made by the recorded\n*1792│            // verdict — a member ticket seats its controller and empowers\n 1793│            // nothing — so this is not a second credential surface: nothing a\n 1794│            // caller sends reaches it, the ticket id is opaque and the record is\n 1795│            // broker-resident. What moved is WHEN the grant is written, not who\n 1796│            // decides it: the settlement happens after the seat resolves, so a\n*1797│            // grant line describing a controller that never attached is\n 1798│            // impossible.\n 1799│            // [impl->REQ-ER-BRINGUP-SPAWNS-SESSION]\n 1800│            // [impl->REQ-ENGINEROOM-ADMIN-BRINGUP-EMPOWERS]\n     │...\n 1908│    /// seated on its ticket; `None` means the gate had nothing to decide (every\n*1909│    /// endpoint but the reserved one, and a seated controller re-serving\n 1910│    /// itself).\n 1911│    ///\n 1912│    /// The credential travels rather than a finished sentence because what it\n     │...\n 1914│    /// still refuse the attach, and a TAKE deliberately revokes the displaced\n*1915│    /// controller's empowerments as it seats the new one. Granting here would put\n 1916│    /// the write on the wrong side of that revocation.\n 1917│    Proceed(Option<SeatAdmit>),\n 1918│}\n     │...\n 2413│    /// which would leave an empowerment minted for a seat nobody took. And a\n*2414│    /// `--take` DROPS the displaced controller's empowerments as it seats the\n 2415│    /// taker (ADR-0052 §3a, \"a take must not hand the taker the displaced\n*2416│    /// controller's authority\") — that revocation is keyed on the session, not\n*2417│    /// on the controller, so a grant written before it is revoked by it. The\n 2418│    /// symptom is the worst kind: the operator is told they are empowered, the\n 2419│    /// record says otherwise, and only the next mode change disagrees out loud.\n 2420│    ///\n     │...\n 2427│    /// EMPOWERMENT is not.** They look symmetric here and are not. A grant is\n*2428│    /// keyed on the seat: a `--take` drops the displaced controller's\n 2429│    /// empowerments as it seats the taker, so the taker's grant must be minted\n 2430│    /// again and stated again or the operator is told nothing about authority\n 2431│    /// they now hold. A briefing is per-session weather, and a human who\n     │...\n 2436│    /// spawned session — necessary for the attach and NOT sufficient for it: the\n*2437│    /// controller is seated later, by the subscribe that redeems the admit\n 2438│    /// ticket, and a bring-up whose attach never arrived left behind a grant\n*2439│    /// line and a briefing describing a controller that never came. Both now\n 2440│    /// compose on this side of the seat, where the ratified invariant puts them.\n 2441│    // [impl->REQ-ENGINEROOM-ADMIN-BRINGUP-EMPOWERS]\n 2442│    // [impl->REQ-ER-SESSION-BRIEFING]\n     │...\n 2453│            outcome,\n*2454│            SubscribeOutcome::Controller | SubscribeOutcome::TookControl\n 2455│        ) {\n 2456│            return None;\n 2457│        }\n     │...\n 2563│    /// quiet. The ledger's own backoff gates this path, so it cannot be sprayed.\n*2564│    /// Controller-bound authority (ADR-0052 §3a): a seat change or release\n 2565│    /// DROPS every empowerment the engine room's session holds.\n 2566│    ///\n 2567│    /// Explicit, because the ADR's premise that a take \"forces a harness\n     │...\n 2571│    ///\n*2572│    /// LOUD on failure: an empowerment that outlives its controller is exactly\n 2573│    /// the authority leak the session scoping exists to bound, and a silent\n*2574│    /// failure here would leave the next controller holding the last one's keys.\n*2575│    // [impl->REQ-ER-CONTROLLER-BOUND-POSTURE]\n 2576│    fn on_control_edge(&self, prev_gen: Option<u64>, next_gen: Option<u64>) {\n 2577│        if spt_store::engineroom::classify_control_edge(prev_gen, next_gen)\n 2578│            == spt_store::engineroom::ControlEdge::Held\n     │...\n 2584│                \"ENGINE_ROOM_AUTHORITY_UNDROPPED:{}: {e} — the empowerments of the \\\n*2585│                 controller that just left this seat could not be revoked; clear \\\n 2586│                 the session's empowerments file by hand\",\n 2587│                self.endpoint\n 2588│            );\n     │...\n 2591│\n*2592│    fn become_controller(&mut self, sub: SharedSend, by: Option<String>, from_seq: u64, attach_gen: u64) {\n*2593│        // Drop the prior controller sink first (its writer's live loop ends when\n 2594│        // tx drops), and bump the generation so (a) an in-flight deadline-evict\n*2595│        // for the old controller can't unseat this one and (b) a prior writer\n 2596│        // still flushing its initial batch returns before its next frame.\n 2597│        // Attribution (F-039 legs b+d): the replaced conn's record names its\n*2598│        // successor, and the new conn accretes the controller role fact.\n 2599│        // [impl->REQ-CONN-POISON-ATTRIBUTION]\n*2600│        let prev_gen = self.controller.as_ref().map(|c| c.attach_gen);\n*2601│        if let Some(old) = self.controller.take() {\n 2602│            old.send.lifecycle_event(\n*2603│                \"controller-replaced\",\n 2604│                &format!(\"session={} new_conn={}\", self.session_id, sub.id()),\n 2605│            );\n 2606│        }\n 2607│        sub.describe(&format!(\n*2608│            \"controller session={} endpoint={} by={}\",\n 2609│            self.session_id,\n 2610│            if self.endpoint.is_empty() { \"-\" } else { &self.endpoint },\n 2611│            by.as_deref().unwrap_or(\"local\")\n     │...\n 2613│        sub.lifecycle_event(\n*2614│            \"controller-attach\",\n 2615│            &format!(\"session={} from_seq={from_seq}\", self.session_id),\n 2616│        );\n*2617│        self.controller = None;\n*2618│        let epoch = self.controller_epoch.fetch_add(1, Ordering::AcqRel) + 1;\n 2619│        // #6 (ADR-0031): a COLD attach's initial batch is a single synthesized clean\n 2620│        // repaint of the CURRENT SCREEN (at pseudo-seq `next_seq - 1`) instead of the\n 2621│        // raw ring — replaying an alt-screen TUI's raw ring into a fresh terminal is\n     │...\n 2635│                    // OBS RIDER Tap B (repaint): the synthesized clean-screen\n*2636│                    // controller repaint bytes (debug-gated, default OFF).\n 2637│                    spt_store::obstap::tap(self.session_id, \"repaint\", &repaint.1);\n 2638│                    (vec![repaint], true)\n 2639│                }\n     │...\n 2648│            };\n*2649│        let (tx, rx) = sync_channel::<CtrlMsg>(CONTROLLER_CHANNEL_DEPTH);\n 2650│        let writer_send = Arc::clone(&sub);\n 2651│        let dt = Arc::clone(&self.delivered_through);\n*2652│        let epoch_gate = Arc::clone(&self.controller_epoch);\n 2653│        let sid = self.session_id;\n 2654│        // W2 decouple: the writer publishes its in-flight-write window here so a\n 2655│        // stall-evict can distinguish a wedged writer from an idle one.\n     │...\n 2662│        let writer = thread::spawn(move || {\n*2663│            controller_writer(\n 2664│                sid,\n 2665│                writer_send,\n 2666│                initial,\n     │...\n 2675│        });\n*2676│        self.controller = Some(ControllerSink {\n 2677│            tx,\n 2678│            send: sub,\n 2679│            by,\n     │...\n 2690│        self.stamp_driven_by();\n*2691│        // The seat changed hands (or was filled): the incoming controller does\n 2692│        // not inherit the outgoing one's authority. An equal-generation re-serve\n 2693│        // is the SAME lease and keeps it.\n 2694│        self.on_control_edge(prev_gen, Some(attach_gen));\n     │...\n 2697│    /// AUTHORITATIVE lease revoke (ADR-0044 decision 2,\n*2698│    /// REQ-CONTROLLER-LEASE-IDENTITY): mark the incumbent controller's sink as\n 2699│    /// revoked-by-`by` BEFORE the caller replaces the slot. The sink's writer\n 2700│    /// thread then emits the terminal [`KIND_DISPLACED`] as its FINAL frame on\n 2701│    /// exit — delivery no longer depends on a `try_send` into a possibly-Full\n     │...\n 2705│    /// stream is the terminal signal either way. Notification is best-effort;\n*2706│    /// REVOCATION is not: the caller's `become_controller` moves the slot (the\n 2707│    /// input fence keys on it) regardless of what the incumbent ever hears.\n*2708│    // [impl->REQ-CONTROLLER-LEASE-IDENTITY]\n 2709│    fn mark_revoked(&self, by: &str) {\n*2710│        if let Some(old) = &self.controller {\n 2711│            if let Ok(mut g) = old.revoked_by.lock() {\n 2712│                *g = Some(by.to_string());\n 2713│            }\n 2714│            old.send.lifecycle_event(\n*2715│                \"controller-revoked\",\n 2716│                &format!(\"session={} taken_by={by}\", self.session_id),\n 2717│            );\n 2718│        }\n     │...\n 2720│\n*2721│    /// The drain timed out the controller's `send_timeout` and the epoch guard\n*2722│    /// still matched ([`mark_controller_gone`], W1). The writer was BLOCKED past\n 2723│    /// the deadline -- it is NOT established to have exited or dropped its `tx`,\n 2724│    /// so this names the deadline the drain observed and not the writer's fate.\n 2725│    /// This is the edge releases#49/#267 Q1 asks about: a cleanly-closed\n     │...\n 2729│    const CLOSE_SEND_DEADLINE: &'static str = \"send_deadline\";\n*2730│    /// The wedge evict released the slot ([`stall_evict_controller`], inline stamp\n 2731│    /// mode). `BRAIN_SUBSCRIBER_STALL_EVICT` already names the wedge; this names\n 2732│    /// the slot transition that followed it.\n 2733│    // [impl->REQ-REFRESH-WAIT-ATTRIBUTION]\n 2734│    const CLOSE_STALL_EVICT: &'static str = \"stall_evict\";\n*2735│    /// A subscriber detached and it happened to be the controller\n 2736│    /// ([`detach_if`]) — an ordinary departure, not a wedge and not a kill.\n 2737│    // [impl->REQ-REFRESH-WAIT-ATTRIBUTION]\n 2738│    const CLOSE_DETACH: &'static str = \"detach\";\n 2739│\n*2740│    /// Detach the controller and UNLATCH the perch — the ONE unlatch path shared\n*2741│    /// by W1's deadline-evict ([`mark_controller_gone`]), [`detach_if`], and (per\n 2742│    /// the W1↔W5 synergy) W5's reconcile self-heal. Dropping the sink ends its\n 2743│    /// writer thread; clearing the slot re-stamps `driven_by` to `None` (so an\n*2744│    /// `ONLINE+CONTROLLED` latch cannot outlive a gone controller).\n 2745│    ///\n 2746│    /// `cause` is the CALLER'S OWN knowledge of which edge closed the slot, and it\n 2747│    /// is a parameter for the reason releases#49/#267 needed this line at all: at\n     │...\n 2751│    // [impl->REQ-REFRESH-WAIT-ATTRIBUTION]\n*2752│    fn clear_controller(&mut self, cause: &'static str) {\n*2753│        let prev_gen = self.controller.as_ref().map(|c| c.attach_gen);\n 2754│        // Captured BEFORE the take: after it there is no sink left to ask.\n*2755│        let conn = self.controller.as_ref().map(|c| c.send.id());\n*2756│        if self.controller.take().is_some() {\n 2757│            // ONLY ON THE REAL TRANSITION (releases#49/#267 Q1). The epoch-guarded\n 2758│            // no-op above returns without taking anything, and a line there would\n 2759│            // report a closure that did not happen — which is exactly the confusion\n     │...\n 2764│                spt_proto::emit_line_err!(\n*2765│                    \"CONTROLLER_SLOT_CLOSED:{} conn={} cause={cause} {} — the controller slot \\\n 2766│                     was occupied and is now empty [REQ-REFRESH-WAIT-ATTRIBUTION]\",\n 2767│                    self.session_id,\n 2768│                    conn.unwrap_or(0),\n     │...\n 2778│\n*2779│    /// B-2 (REQ-CONTROLLER-LIVENESS-REAP): lazily drop a controller whose WRITER\n 2780│    /// THREAD has EXITED — the record-vs-truth honesty step, run on the authoritative\n*2781│    /// `KIND_SESSIONS` poll BEFORE `controller_by`/`has_controller` are read. The\n*2782│    /// writer exits when a socket write fails (a severed / keepalive-killed controller\n 2783│    /// conn) or its `tx` drops; while it is parked on `rx.recv()` (a LIVE but idle\n*2784│    /// controller) it is NOT finished, so a live controller is never touched (the\n 2785│    /// no-false-reap guard). Reaping here makes the broker report honest truth so the\n 2786│    /// EXISTING [`converge_perch_stamps`] clears the now-stale info.json\n 2787│    /// `driven_by`/`controlled` stamp — no new detach path, no new query.\n     │...\n 2790│    /// NO file I/O (the lock-across-effect discipline, KH 7.12/5.16). It drops the\n*2791│    /// sink in-memory ONLY (like [`become_controller`]'s prior-drop) and does NOT\n 2792│    /// [`stamp_driven_by`] — the OFF-lock `converge_perch_stamps` writes the honest\n 2793│    /// stamp. Returns true iff it reaped.\n 2794│    ///\n     │...\n 2796│    /// flowing when the conn severed) or whose detach was dropped by a prior wedge\n*2797│    /// (the B-1 poison class). A TRULY IDLE severed controller (writer parked on recv,\n 2798│    /// no output, conn not yet EOF'd) stays `is_finished() == false` and converges\n 2799│    /// only once output resumes or the conn EOFs (→ `handle_conn` detach) — that\n 2800│    /// harder active-probe case is the reserved [`REQ-HAZARD-DRIVEN-BY-IDLE-REMOTE-EVICT`],\n     │...\n 2802│    /// never latches `driven_by`.\n*2803│    // [impl->REQ-CONTROLLER-LIVENESS-REAP]\n*2804│    fn reap_dead_controller(&mut self) -> bool {\n 2805│        let dead = self\n*2806│            .controller\n 2807│            .as_ref()\n 2808│            .is_some_and(|c| c._writer.is_finished());\n 2809│        if dead {\n     │...\n 2811│            // (info.json I/O under the log lock is the KH 7.12 hazard). converge does it.\n*2812│            self.controller = None;\n 2813│            return true;\n 2814│        }\n*2815│        // W2 decouple (REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE): a controller whose\n 2816│        // writer is BLOCKED (not exited) on a suspended brain conn past\n 2817│        // BRAIN_WRITE_DEADLINE is also reaped here, so the honest `KIND_SESSIONS`\n 2818│        // truth + the off-lock converge_perch_stamps release its stale `driven_by`\n     │...\n 2820│        // the multi-session snapshot closure (no I/O under the shared sessions lock).\n*2821│        self.stall_evict_controller(Instant::now(), StampMode::Deferred)\n 2822│    }\n 2823│\n*2824│    /// Deadline-evict path (W1): the drain observed a fully-stalled controller and\n 2825│    /// timed out its `send_timeout`. Detach it — but ONLY if it is still the same\n*2826│    /// controller generation the drain saw (`epoch`), so a NEWER controller that\n 2827│    /// attached during the stall is never unseated.\n*2828│    fn mark_controller_gone(&mut self, epoch: u64) {\n*2829│        if self.controller_epoch.load(Ordering::Acquire) == epoch {\n*2830│            self.clear_controller(Self::CLOSE_SEND_DEADLINE);\n 2831│        }\n 2832│    }\n 2833│\n*2834│    /// Whether the current controller's writer is WEDGED — blocked inside a single\n 2835│    /// socket write for at least `deadline` (the brain-subscriber conn suspended /\n*2836│    /// black-holed). An IDLE controller (writer parked on `rx.recv()`, nothing to\n*2837│    /// write) reads NOT stalled: its [`ControllerSink::write_blocked_since`] is `None`.\n 2838│    /// Pure over the injected `now` + `deadline` so the unit tests it with no thread.\n 2839│    // [impl->REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]"},"isError":false,"timestamp":1789523950780}}
{"type":"message","id":"ff2743a3","parentId":"c856edcf","timestamp":"2026-09-16T01:59:13.631Z","message":{"role":"toolResult","toolCallId":"call_jPCBai7mJxwK4o0XVq9jHBeP|fc_0d104aba2b0f3edc016aa9f7ee9af487d1ab71123d585b9a96","toolName":"glob","content":[{"type":"text","text":"# .spt/preserved/307/hertz-w1bc/\nenlyzeam-readonly-baseline.json\n## w1c-run1/\ncommand-116.json\ncommand-115.json\ncommand-114.json\ncommand-113.json\ncommand-112.json\ncommand-111.json\ncommand-110.json\ncommand-109.json\ncommand-108.json\ncommand-107.json\ncommand-106.json\ncommand-105.json\ncommand-104.json\ncommand-103.json\ncommand-102.json\ncommand-101.json\ncommand-100.json\ncommand-099.json\ncommand-098.json\ncommand-097.json\ncommand-096.json\ncommand-095.json\ncommand-094.json\ncommand-093.json\ncommand-092.json\ncommand-091.json\ncommand-090.json\ncommand-089.json\ncommand-088.json\ncommand-087.json\ncommand-086.json\ncommand-085.json\ncommand-084.json\ncommand-083.json\ncommand-082.json\ncommand-081.json\ncommand-080.json\ncommand-079.json\ncommand-078.json\ncommand-077.json\ncommand-076.json\ncommand-075.json\ncommand-074.json\ncommand-073.json\ncommand-072.json\ncommand-071.json\ncommand-070.json\ncommand-069.json\ncommand-068.json\ncommand-067.json\ncommand-066.json\ncommand-065.json\ncommand-064.json\ncommand-063.json\ncommand-062.json\ncommand-061.json\ncommand-060.json\ncommand-059.json\ncommand-058.json\ncommand-057.json\ncommand-056.json\nevents.jsonl\ncommand-055.json\ncommand-054.json\ncommand-053.json\ncommand-052.json\ncommand-051.json\nrc-1.chunks.jsonl\nrc-1.progress.jsonl\ncommand-050.json\ncommand-049.json\ncommand-048.json\ncommand-047.json\ncommand-046.json\ncommand-045.json\ncommand-044.json\ncommand-043.json\ncommand-042.json\ncommand-041.json\ncommand-040.json\ncommand-039.json\ncommand-038.json\ncommand-037.json\ncommand-036.json\ncommand-035.json\ncommand-034.json\ncommand-033.json\ncommand-032.json\ncommand-031.json\ncommand-030.json\ncommand-029.json\ncommand-028.json\ndaemon.json.original\ncommand-027.json\ncommand-026.json\ncommand-025.json\ncommand-024.json\ncommand-023.json\ncommand-022.json\ncommand-021.json\ncommand-020.json\ncommand-019.json\ncommand-018.json\ncommand-017.json\ncommand-016.json\ncommand-015.json\ncommand-014.json\ncommand-013.json\ncommand-012.json\ncommand-011.json\ncommand-010.json\ncommand-009.json\ncommand-008.json\ncommand-007.json\ncommand-006.json\ncommand-005.json\ncommand-004.json\nrc-1.stderr.log\n### home-w1c/\nbrain.ready\ndeadline-8d0b621ae74118f5.json\ndaemon.json\nbrain.ready.tmp.22860-130\ngenerator-state.json\ngenerator-input.jsonl\ndaemon-effects.log\ngenerator-bind.json\n#### owlery/\n.registry\n##### w1crig/\ninfo.json\napi.token\nio-events.log\n.io-events.lock\nready\n#### identity/\npump-heartbeat.json\npump-health.json\n#### index/\nproject-index-stats.json\nproject-index.json\n##### invalidations/\n#### logs/\ndaemon.stderr.log\n##### obs/\nrc-establish.log\n## w1c-launch1/\nstdout.log\n\n[150 results limit reached. Use limit=300 for more]"}],"details":{"scopePath":".spt/preserved/307/**/*, **/*hazard*, **/*glossary*","fileCount":150,"files":[".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/",".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/brain.ready",".spt/preserved/307/hertz-w1bc/w1c-run1/",".spt/preserved/307/hertz-w1bc/w1c-run1/command-116.json",".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/owlery/w1crig/",".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/owlery/w1crig/info.json",".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/owlery/",".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/identity/",".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/identity/pump-heartbeat.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-115.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-114.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-113.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-112.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-111.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-110.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-109.json",".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/identity/pump-health.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-108.json",".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/deadline-8d0b621ae74118f5.json",".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/index/",".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/index/project-index-stats.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-107.json",".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/logs/daemon.stderr.log",".spt/preserved/307/hertz-w1bc/w1c-run1/command-106.json",".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/daemon.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-105.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-104.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-103.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-102.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-101.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-100.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-099.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-098.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-097.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-096.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-095.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-094.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-093.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-092.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-091.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-090.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-089.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-088.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-087.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-086.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-085.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-084.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-083.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-082.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-081.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-080.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-079.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-078.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-077.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-076.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-075.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-074.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-073.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-072.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-071.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-070.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-069.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-068.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-067.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-066.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-065.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-064.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-063.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-062.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-061.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-060.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-059.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-058.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-057.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-056.json",".spt/preserved/307/hertz-w1bc/w1c-run1/events.jsonl",".spt/preserved/307/hertz-w1bc/w1c-run1/command-055.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-054.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-053.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-052.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-051.json",".spt/preserved/307/hertz-w1bc/w1c-run1/rc-1.chunks.jsonl",".spt/preserved/307/hertz-w1bc/w1c-run1/rc-1.progress.jsonl",".spt/preserved/307/hertz-w1bc/w1c-run1/command-050.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-049.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-048.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-047.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-046.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-045.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-044.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-043.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-042.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-041.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-040.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-039.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-038.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-037.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-036.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-035.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-034.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-033.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-032.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-031.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-030.json",".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/brain.ready.tmp.22860-130",".spt/preserved/307/hertz-w1bc/w1c-run1/command-029.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-028.json",".spt/preserved/307/hertz-w1bc/w1c-run1/daemon.json.original",".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/generator-state.json",".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/generator-input.jsonl",".spt/preserved/307/hertz-w1bc/w1c-run1/command-027.json",".spt/preserved/307/hertz-w1bc/enlyzeam-readonly-baseline.json",".spt/preserved/307/hertz-w1bc/",".spt/preserved/307/hertz-w1bc/w1c-run1/command-026.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-025.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-024.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-023.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-022.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-021.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-020.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-019.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-018.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-017.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-016.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-015.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-014.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-013.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-012.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-011.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-010.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-009.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-008.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-007.json",".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/owlery/.registry",".spt/preserved/307/hertz-w1bc/w1c-run1/command-006.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-005.json",".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/index/project-index.json",".spt/preserved/307/hertz-w1bc/w1c-run1/command-004.json",".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/index/invalidations/",".spt/preserved/307/hertz-w1bc/w1c-launch1/stdout.log",".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/logs/obs/rc-establish.log",".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/logs/",".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/logs/obs/",".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/daemon-effects.log",".spt/preserved/307/hertz-w1bc/w1c-run1/rc-1.stderr.log",".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/generator-bind.json",".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/owlery/w1crig/api.token",".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/owlery/w1crig/io-events.log",".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/owlery/w1crig/.io-events.lock",".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/owlery/w1crig/ready"],"truncated":true,"resultLimitReached":150,"cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","meta":{"limits":{"resultLimit":{"reached":150,"suggestion":300}}}},"isError":false,"timestamp":1789523953630}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_iWHBpq0HkAaqB3menHOXFoiv|fc_0d104aba2b0f3edc016aa9f7fb57a087d1b75ab70fb2619e4c","toolName":"read","startedAt":"2026-09-16T01:59:23.414Z","args":{"path":"crates/spt-daemon/src/broker.rs:2720-2955"},"intent":"Reading lifecycle predicate code"},"id":"343e2a7a","parentId":"ff2743a3","timestamp":"2026-09-16T01:59:23.414Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_YtJcc3EwfvdA4I6mD7CfVQ82|fc_0d104aba2b0f3edc016aa9f7fb57b087d1993f687654e451cb","toolName":"grep","startedAt":"2026-09-16T01:59:23.414Z","args":{"path":"crates/spt-daemon/src/broker.rs"},"intent":"Locating predicate writer definitions"},"id":"a574dc84","parentId":"343e2a7a","timestamp":"2026-09-16T01:59:23.414Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_i38LYlWR0TVP2O8KWwZaiHN7|fc_0d104aba2b0f3edc016aa9f7fb57b887d1b1ddff4b18c6cabd","toolName":"grep","startedAt":"2026-09-16T01:59:23.414Z","args":{"path":".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/logs/daemon.stderr.log"},"intent":"Finding targeted log events"},"id":"40891412","parentId":"a574dc84","timestamp":"2026-09-16T01:59:23.414Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_W830bcvnVfmLBUl6kty69OXU|fc_0d104aba2b0f3edc016aa9f7fb57c087d18aba9be592f6a8a9","toolName":"glob","startedAt":"2026-09-16T01:59:23.414Z","args":{"path":"**/*hazard*;**/*trace*;**/*307*"},"intent":"Locating trace hazard documents"},"id":"9fe45f04","parentId":"40891412","timestamp":"2026-09-16T01:59:23.414Z"}
{"type":"message","id":"06fc8f46","parentId":"9fe45f04","timestamp":"2026-09-16T01:59:23.414Z","message":{"role":"assistant","content":[{"type":"toolCall","id":"call_iWHBpq0HkAaqB3menHOXFoiv|fc_0d104aba2b0f3edc016aa9f7fb57a087d1b75ab70fb2619e4c","name":"read","arguments":{"i":"Reading lifecycle predicate code","path":"crates/spt-daemon/src/broker.rs:2720-2955"},"intent":"Reading lifecycle predicate code"},{"type":"toolCall","id":"call_YtJcc3EwfvdA4I6mD7CfVQ82|fc_0d104aba2b0f3edc016aa9f7fb57b087d1993f687654e451cb","name":"grep","arguments":{"i":"Locating predicate writer definitions","pattern":"fn any_local_controller_wedged|fn controller_writer|fn detach_if|fn handle_conn|CONN_WRITE_POISONED","path":"crates/spt-daemon/src/broker.rs","case":true,"gitignore":false,"skip":null},"intent":"Locating predicate writer definitions"},{"type":"toolCall","id":"call_i38LYlWR0TVP2O8KWwZaiHN7|fc_0d104aba2b0f3edc016aa9f7fb57b887d1b1ddff4b18c6cabd","name":"grep","arguments":{"i":"Finding targeted log events","pattern":"conn=(17|70|88|104)( |$)|BRAIN_UPDATE_RESTART|BRAIN_RESUM|BRAIN_PROMOTE|STALL_EVICT|SLOT_CLOSED","path":".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/logs/daemon.stderr.log","case":true,"gitignore":false,"skip":null},"intent":"Finding targeted log events"},{"type":"toolCall","id":"call_W830bcvnVfmLBUl6kty69OXU|fc_0d104aba2b0f3edc016aa9f7fb57c087d18aba9be592f6a8a9","name":"glob","arguments":{"i":"Locating trace hazard documents","path":"**/*hazard*;**/*trace*;**/*307*","hidden":true,"gitignore":true,"limit":80},"intent":"Locating trace hazard documents"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":17756,"output":277,"cacheRead":11904,"cacheWrite":0,"totalTokens":29937,"cost":{"input":0.17756000000000002,"output":0.013850000000000001,"cacheRead":0.011904,"cacheWrite":0,"total":0.20331400000000002}},"stopReason":"toolUse","timestamp":1789523953660,"responseId":"resp_0d104aba2b0f3edc016aa9f7f20c5487d1b42674ef89b61acd","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"fc_0d104aba2b0f3edc016aa9f7fb57a087d1b75ab70fb2619e4c","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading lifecycle predicate code\",\"path\":\"crates/spt-daemon/src/broker.rs:2720-2955\"}","call_id":"call_iWHBpq0HkAaqB3menHOXFoiv","internal_chat_message_metadata_passthrough":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"},"name":"read","metadata":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"}},{"id":"fc_0d104aba2b0f3edc016aa9f7fb57b087d1993f687654e451cb","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating predicate writer definitions\",\"pattern\":\"fn any_local_controller_wedged|fn controller_writer|fn detach_if|fn handle_conn|CONN_WRITE_POISONED\",\"path\":\"crates/spt-daemon/src/broker.rs\",\"case\":true,\"gitignore\":false,\"skip\":null}","call_id":"call_YtJcc3EwfvdA4I6mD7CfVQ82","internal_chat_message_metadata_passthrough":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"},"name":"grep","metadata":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"}},{"id":"fc_0d104aba2b0f3edc016aa9f7fb57b887d1b1ddff4b18c6cabd","type":"function_call","status":"completed","arguments":"{\"i\":\"Finding targeted log events\",\"pattern\":\"conn=(17|70|88|104)( |$)|BRAIN_UPDATE_RESTART|BRAIN_RESUM|BRAIN_PROMOTE|STALL_EVICT|SLOT_CLOSED\",\"path\":\".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/logs/daemon.stderr.log\",\"case\":true,\"gitignore\":false,\"skip\":null}","call_id":"call_i38LYlWR0TVP2O8KWwZaiHN7","internal_chat_message_metadata_passthrough":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"},"name":"grep","metadata":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"}},{"id":"fc_0d104aba2b0f3edc016aa9f7fb57c087d18aba9be592f6a8a9","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating trace hazard documents\",\"path\":\"**/*hazard*;**/*trace*;**/*307*\",\"hidden\":true,\"gitignore\":true,\"limit\":80}","call_id":"call_W830bcvnVfmLBUl6kty69OXU","internal_chat_message_metadata_passthrough":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"},"name":"glob","metadata":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"}}]},"duration":9752.459900000133,"ttft":9709.701899999985,"completedAt":1789523963412,"contextSnapshot":{"promptTokens":29660,"nonMessageTokens":10480,"compactionEpoch":0}}}
{"type":"message","id":"bc522d73","parentId":"06fc8f46","timestamp":"2026-09-16T01:59:23.417Z","message":{"role":"toolResult","toolCallId":"call_YtJcc3EwfvdA4I6mD7CfVQ82|fc_0d104aba2b0f3edc016aa9f7fb57b087d1993f687654e451cb","toolName":"grep","content":[{"type":"text","text":" 3546|    // [impl->REQ-SERVE-OWNERSHIP-GENERATION]\n*3547|    fn detach_if_gen(&mut self, sub: &SharedSend, gen: Option<u64>) {\n 3548|        if let (Some(g), Some(c)) = (gen, self.controller.as_ref()) {\n 3549|            if c.attach_gen != g {\n 3550|                // Stale generation: never clears the (newer) controller.\n...\n 3561|\n*3562|    fn detach_if(&mut self, sub: &SharedSend) {\n 3563|        let is_ctrl = self\n 3564|            .controller\n 3565|            .as_ref()\n...\n 3954|#[allow(clippy::too_many_arguments)] // per-writer thread-handoff channels (W2 added the blocked-since signal); a struct would obscure the call\n*3955|fn controller_writer(\n 3956|    sid: u64,\n 3957|    send: SharedSend,\n 3958|    initial: Vec<(u64, Vec<u8>)>,\n...\n 6002|    // [impl->REQ-UPDATE-PROMOTE-DRAINED]\n*6003|    pub fn any_local_controller_wedged(&self) -> bool {\n 6004|        let now = Instant::now();\n 6005|        let deadline = brain_write_deadline();\n 6006|        let sessions = match self.sessions.lock() {\n...\n 6049|    /// restart must not disturb the hosted children).\n*6050|    fn handle_conn(self: &Arc<Self>, conn: Stream) -> io::Result<()> {\n 6051|        // RAII handler tally (panic-safe): the round-4 gate's seam for \"the\n 6052|        // poisoned conn's handler reached EOF cleanup\" (see conn_handler_count).\n 6053|        struct HandlerTally<'a>(&'a AtomicUsize);"}],"details":{"scopePath":"crates/spt-daemon/src/broker.rs","searchPath":"C:\\Users\\decid\\Documents\\projects\\spt-core\\crates\\spt-daemon\\src\\broker.rs","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","matchCount":5,"fileCount":1,"files":["crates/spt-daemon/src/broker.rs"],"fileMatches":[{"path":"crates/spt-daemon/src/broker.rs","count":5}],"truncated":false,"displayContent":" 3546│    // [impl->REQ-SERVE-OWNERSHIP-GENERATION]\n*3547│    fn detach_if_gen(&mut self, sub: &SharedSend, gen: Option<u64>) {\n 3548│        if let (Some(g), Some(c)) = (gen, self.controller.as_ref()) {\n 3549│            if c.attach_gen != g {\n 3550│                // Stale generation: never clears the (newer) controller.\n     │...\n 3561│\n*3562│    fn detach_if(&mut self, sub: &SharedSend) {\n 3563│        let is_ctrl = self\n 3564│            .controller\n 3565│            .as_ref()\n     │...\n 3954│#[allow(clippy::too_many_arguments)] // per-writer thread-handoff channels (W2 added the blocked-since signal); a struct would obscure the call\n*3955│fn controller_writer(\n 3956│    sid: u64,\n 3957│    send: SharedSend,\n 3958│    initial: Vec<(u64, Vec<u8>)>,\n     │...\n 6002│    // [impl->REQ-UPDATE-PROMOTE-DRAINED]\n*6003│    pub fn any_local_controller_wedged(&self) -> bool {\n 6004│        let now = Instant::now();\n 6005│        let deadline = brain_write_deadline();\n 6006│        let sessions = match self.sessions.lock() {\n     │...\n 6049│    /// restart must not disturb the hosted children).\n*6050│    fn handle_conn(self: &Arc<Self>, conn: Stream) -> io::Result<()> {\n 6051│        // RAII handler tally (panic-safe): the round-4 gate's seam for \"the\n 6052│        // poisoned conn's handler reached EOF cleanup\" (see conn_handler_count).\n 6053│        struct HandlerTally<'a>(&'a AtomicUsize);"},"isError":false,"timestamp":1789523963416}}
{"type":"message","id":"ec21636b","parentId":"bc522d73","timestamp":"2026-09-16T01:59:23.482Z","message":{"role":"toolResult","toolCallId":"call_iWHBpq0HkAaqB3menHOXFoiv|fc_0d104aba2b0f3edc016aa9f7fb57a087d1b75ab70fb2619e4c","toolName":"read","content":[{"type":"text","text":"    fn mark_revoked(&self, by: &str) {\n…\n    }\n\n    /// The drain timed out the controller's `send_timeout` and the epoch guard\n    /// still matched ([`mark_controller_gone`], W1). The writer was BLOCKED past\n    /// the deadline -- it is NOT established to have exited or dropped its `tx`,\n    /// so this names the deadline the drain observed and not the writer's fate.\n    /// This is the edge releases#49/#267 Q1 asks about: a cleanly-closed\n    /// OLD-generation conn arrives here too, and from this site the two are\n    /// indistinguishable.\n    // [impl->REQ-REFRESH-WAIT-ATTRIBUTION]\n    const CLOSE_SEND_DEADLINE: &'static str = \"send_deadline\";\n    /// The wedge evict released the slot ([`stall_evict_controller`], inline stamp\n    /// mode). `BRAIN_SUBSCRIBER_STALL_EVICT` already names the wedge; this names\n    /// the slot transition that followed it.\n    // [impl->REQ-REFRESH-WAIT-ATTRIBUTION]\n    const CLOSE_STALL_EVICT: &'static str = \"stall_evict\";\n    /// A subscriber detached and it happened to be the controller\n    /// ([`detach_if`]) — an ordinary departure, not a wedge and not a kill.\n    // [impl->REQ-REFRESH-WAIT-ATTRIBUTION]\n    const CLOSE_DETACH: &'static str = \"detach\";\n\n    /// Detach the controller and UNLATCH the perch — the ONE unlatch path shared\n    /// by W1's deadline-evict ([`mark_controller_gone`]), [`detach_if`], and (per\n    /// the W1↔W5 synergy) W5's reconcile self-heal. Dropping the sink ends its\n    /// writer thread; clearing the slot re-stamps `driven_by` to `None` (so an\n    /// `ONLINE+CONTROLLED` latch cannot outlive a gone controller).\n    ///\n    /// `cause` is the CALLER'S OWN knowledge of which edge closed the slot, and it\n    /// is a parameter for the reason releases#49/#267 needed this line at all: at\n    /// this site a writer that exited, an evicted wedge and a deliberate detach are\n    /// indistinguishable, so a cause DERIVED here would be a guess wearing a label.\n    /// Each caller names what it knows; nothing infers.\n    // [impl->REQ-REFRESH-WAIT-ATTRIBUTION]\n    fn clear_controller(&mut self, cause: &'static str) {\n        let prev_gen = self.controller.as_ref().map(|c| c.attach_gen);\n        // Captured BEFORE the take: after it there is no sink left to ask.\n        let conn = self.controller.as_ref().map(|c| c.send.id());\n        if self.controller.take().is_some() {\n            // ONLY ON THE REAL TRANSITION (releases#49/#267 Q1). The epoch-guarded\n            // no-op above returns without taking anything, and a line there would\n            // report a closure that did not happen — which is exactly the confusion\n            // the question exists to remove: today a clean close and a slot that was\n            // already empty leave identical stderr.\n            // [impl->REQ-REFRESH-WAIT-ATTRIBUTION]\n            if crate::brain::diag_trace_enabled() {\n                spt_proto::emit_line_err!(\n                    \"CONTROLLER_SLOT_CLOSED:{} conn={} cause={cause} {} — the controller slot \\\n                     was occupied and is now empty [REQ-REFRESH-WAIT-ATTRIBUTION]\",\n                    self.session_id,\n                    conn.unwrap_or(0),\n                    crate::conn::log_stamp()\n                );\n            }\n            self.stamp_driven_by();\n            // Detach is a drop edge: the posture goes offline (derived from the\n            // stamp just written) and the empowerments go with it.\n            self.on_control_edge(prev_gen, None);\n        }\n    }\n\n    /// B-2 (REQ-CONTROLLER-LIVENESS-REAP): lazily drop a controller whose WRITER\n    /// THREAD has EXITED — the record-vs-truth honesty step, run on the authoritative\n    /// `KIND_SESSIONS` poll BEFORE `controller_by`/`has_controller` are read. The\n    /// writer exits when a socket write fails (a severed / keepalive-killed controller\n    /// conn) or its `tx` drops; while it is parked on `rx.recv()` (a LIVE but idle\n    /// controller) it is NOT finished, so a live controller is never touched (the\n    /// no-false-reap guard). Reaping here makes the broker report honest truth so the\n    /// EXISTING [`converge_perch_stamps`] clears the now-stale info.json\n    /// `driven_by`/`controlled` stamp — no new detach path, no new query.\n    ///\n    /// LOCK-SAFE: this runs UNDER the log lock in the snapshot closure, so it must do\n    /// NO file I/O (the lock-across-effect discipline, KH 7.12/5.16). It drops the\n    /// sink in-memory ONLY (like [`become_controller`]'s prior-drop) and does NOT\n    /// [`stamp_driven_by`] — the OFF-lock `converge_perch_stamps` writes the honest\n    /// stamp. Returns true iff it reaped.\n    ///\n    /// SCOPE (doyle B-2 reduction): catches a writer that FAILED A WRITE (output was\n    /// flowing when the conn severed) or whose detach was dropped by a prior wedge\n    /// (the B-1 poison class). A TRULY IDLE severed controller (writer parked on recv,\n    /// no output, conn not yet EOF'd) stays `is_finished() == false` and converges\n    /// only once output resumes or the conn EOFs (→ `handle_conn` detach) — that\n    /// harder active-probe case is the reserved [`REQ-HAZARD-DRIVEN-BY-IDLE-REMOTE-EVICT`],\n    /// deliberately NOT built here. KH 7.15 held by construction: only ever CLEARS,\n    /// never latches `driven_by`.\n    // [impl->REQ-CONTROLLER-LIVENESS-REAP]\n    fn reap_dead_controller(&mut self) -> bool {\n        let dead = self\n            .controller\n            .as_ref()\n            .is_some_and(|c| c._writer.is_finished());\n        if dead {\n            // In-memory drop only (ends the — already-exited — writer); NO stamp here\n            // (info.json I/O under the log lock is the KH 7.12 hazard). converge does it.\n            self.controller = None;\n            return true;\n        }\n        // W2 decouple (REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE): a controller whose\n        // writer is BLOCKED (not exited) on a suspended brain conn past\n        // BRAIN_WRITE_DEADLINE is also reaped here, so the honest `KIND_SESSIONS`\n        // truth + the off-lock converge_perch_stamps release its stale `driven_by`\n        // stamp even when nobody take/reattaches. Deferred stamp: the reap runs in\n        // the multi-session snapshot closure (no I/O under the shared sessions lock).\n        self.stall_evict_controller(Instant::now(), StampMode::Deferred)\n    }\n\n    /// Deadline-evict path (W1): the drain observed a fully-stalled controller and\n    /// timed out its `send_timeout`. Detach it — but ONLY if it is still the same\n    /// controller generation the drain saw (`epoch`), so a NEWER controller that\n    /// attached during the stall is never unseated.\n    fn mark_controller_gone(&mut self, epoch: u64) {\n        if self.controller_epoch.load(Ordering::Acquire) == epoch {\n            self.clear_controller(Self::CLOSE_SEND_DEADLINE);\n        }\n    }\n\n    /// Whether the current controller's writer is WEDGED — blocked inside a single\n    /// socket write for at least `deadline` (the brain-subscriber conn suspended /\n    /// black-holed). An IDLE controller (writer parked on `rx.recv()`, nothing to\n    /// write) reads NOT stalled: its [`ControllerSink::write_blocked_since`] is `None`.\n    /// Pure over the injected `now` + `deadline` so the unit tests it with no thread.\n    // [impl->REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]\n    fn controller_write_stalled(&self, now: Instant, deadline: Duration) -> bool {\n        self.controller.as_ref().is_some_and(|c| {\n            c.write_blocked_since\n                .lock()\n                .ok()\n                .and_then(|g| *g)\n                .is_some_and(|since| now.saturating_duration_since(since) >= deadline)\n        })\n    }\n\n    /// Whether this session's controller is the brain's OWN LOCAL conn (`by: None`)\n    /// AND wedged past `deadline` — the brain-generation promotion DRAINED gate reads\n    /// it (REQ-UPDATE-PROMOTE-DRAINED). A killed OLD brain generation whose local\n    /// controller conn is black-holed reads wedged here until the W2 stall-evict\n    /// releases it (or the conn EOFs); a cleanly-closed conn is simply absent. A\n    /// REMOTE operator's wedged conn (`by: Some`) is the W2 stall-evict concern, NOT\n    /// an old-gen brain conn, so it is EXCLUDED — a stuck remote viewer must never\n    /// gate a brain promotion. Pure over the injected `now`/`deadline` (no thread).\n    // [impl->REQ-UPDATE-PROMOTE-DRAINED]\n    fn local_controller_wedged(&self, now: Instant, deadline: Duration) -> bool {\n        self.controller.as_ref().is_some_and(|c| c.by.is_none())\n            && self.controller_write_stalled(now, deadline)\n    }\n\n    /// Stall-evict a controller wedged past [`BRAIN_WRITE_DEADLINE`]: clear the slot\n    /// and re-stamp `driven_by`/`controlled` (the BROKER-side release — no brain\n    /// round-trip, so `controlled-by` can never latch behind a dead transit) and\n    /// record the evict for observability (broker log now; W3 persists the stderr,\n    /// the tally surfaces on `daemon status`). The take/reattach path and the\n    /// `KIND_SESSIONS` reap call this so a suspended brain subscriber can never hold\n    /// control on an idle session (the output-driven [`mark_controller_gone`] never\n    /// fires without new output). Returns true iff it evicted. `stamp` selects whether\n    /// the perch re-stamp happens inline (the take/reattach path, which already stamps\n    /// under the log lock via [`become_controller`]) or is deferred to the off-lock\n    /// `converge_perch_stamps` (the multi-session reap closure, KH 7.12 no-I/O rule).\n    // [impl->REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]\n    fn stall_evict_controller(&mut self, now: Instant, stamp: StampMode) -> bool {\n        let deadline = brain_write_deadline();\n        if !self.controller_write_stalled(now, deadline) {\n            return false;\n        }\n        // Attribution (F-039 leg b): the evicted controller's stable conn id +\n        // stamps ride the record. [impl->REQ-CONN-POISON-ATTRIBUTION]\n        let conn = self.controller.as_ref().map(|c| c.send.id()).unwrap_or(0);\n        spt_proto::emit_line_err!(\n            \"BRAIN_SUBSCRIBER_STALL_EVICT:{} conn={conn} {} — controller writer blocked > {}ms \\\n             on a suspended/black-holed brain conn; detached, driven_by released \\\n             [REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]\",\n            self.session_id,\n            crate::conn::log_stamp(),\n            deadline.as_millis()\n        );\n        record_stall_evict();\n        match stamp {\n            // Inline: drop the sink AND re-stamp the perch (become_controller's\n            // sibling — the take/reattach path is already I/O-under-log-lock).\n            StampMode::Inline => self.clear_controller(Self::CLOSE_STALL_EVICT),\n            // Deferred: in-memory drop ONLY (the reap closure forbids I/O under the\n            // shared sessions lock); converge_perch_stamps writes the release off-lock.\n            StampMode::Deferred => self.controller = None,\n        }\n        true\n    }\n\n    /// Add a read-only viewer: snapshot the ring from `from_seq` into the writer\n    /// thread (NOT through the bounded live queue — replay must not trip the\n    /// overflow eviction), seed it the current size frame first, then run the\n    /// live queue. The drain thread fans live frames in via `try_send`.\n    fn add_viewer(&mut self, sub: SharedSend, from_seq: u64, by: Option<&str>) {\n        let vid = self.next_viewer_id;\n        self.next_viewer_id += 1;\n        // Attribution (F-039 legs b+d): viewer role fact + attach record.\n        // [impl->REQ-CONN-POISON-ATTRIBUTION]\n        sub.describe(&format!(\n            \"viewer session={} endpoint={} vid={vid}\",\n            self.session_id,\n            if self.endpoint.is_empty() { \"-\" } else { &self.endpoint }\n        ));\n        sub.lifecycle_event(\n            \"viewer-attach\",\n            &format!(\"session={} vid={vid} from_seq={from_seq}\", self.session_id),\n        );\n        // The size frame + the initial batch (owned by the thread, not the bounded\n        // queue). A COLD viewer attach (`from_seq == 0`) — or one resuming from\n        // BELOW the presentation floor (REQ-RC-RESIZE-PRESENTATION-BARRIER: never\n        // a raw replay across a transition boundary) — gets the #6 clean repaint;\n        // a skip-to-live / snap re-attach at/above the floor (incl. `u64::MAX`)\n        // keeps the raw-ring slice (`s >= from_seq`) — empty for skip-to-live, as\n        // before.\n        let mut initial: Vec<Envelope> =\n            vec![size_envelope(self.session_id, self.size.0, self.size.1)];\n        match self.repaint_initial(from_seq) {\n            // The repaint batch rides the sync flag: it supersedes everything\n            // before its pseudo-seq (the wire-flag ruling — a below-floor viewer\n            // resume baselines the jump instead of gap-erroring).\n            Some((seq, bytes)) => {\n                // OBS RIDER Tap B (repaint): the synthesized clean-screen viewer\n                // repaint bytes (debug-gated, default OFF).\n                spt_store::obstap::tap(self.session_id, \"repaint\", &bytes);\n                initial.push(sync_output_envelope(self.session_id, seq, &bytes));\n            }\n            None => initial.extend(\n                self.ring\n                    .iter()\n                    .filter(|(s, _)| *s >= from_seq)\n                    .map(|(seq, bytes)| output_envelope(self.session_id, *seq, bytes)),\n            ),\n        }\n        let (tx, rx) = sync_channel::<Envelope>(VIEWER_CHANNEL_DEPTH);\n        let writer_send = Arc::clone(&sub);\n        let evicted = Arc::new(AtomicBool::new(false));\n        let writer_evicted = Arc::clone(&evicted);\n        let session_id = self.session_id;\n        let writer =\n            thread::spawn(move || viewer_writer(writer_send, session_id, initial, rx, writer_evicted));\n        self.viewers.insert(\n            vid,\n            ViewerSink {\n                tx,\n…\n            },\n        );\n    }\n\n[Showing lines 2719-2958 of 16852. Use :2959 to continue]"}],"details":{"truncation":{"content":"    }\r\n\r\n    /// The drain timed out the controller's `send_timeout` and the epoch guard\r\n    /// still matched ([`mark_controller_gone`], W1). The writer was BLOCKED past\r\n    /// the deadline -- it is NOT established to have exited or dropped its `tx`,\r\n    /// so this names the deadline the drain observed and not the writer's fate.\r\n    /// This is the edge releases#49/#267 Q1 asks about: a cleanly-closed\r\n    /// OLD-generation conn arrives here too, and from this site the two are\r\n    /// indistinguishable.\r\n    // [impl->REQ-REFRESH-WAIT-ATTRIBUTION]\r\n    const CLOSE_SEND_DEADLINE: &'static str = \"send_deadline\";\r\n    /// The wedge evict released the slot ([`stall_evict_controller`], inline stamp\r\n    /// mode). `BRAIN_SUBSCRIBER_STALL_EVICT` already names the wedge; this names\r\n    /// the slot transition that followed it.\r\n    // [impl->REQ-REFRESH-WAIT-ATTRIBUTION]\r\n    const CLOSE_STALL_EVICT: &'static str = \"stall_evict\";\r\n    /// A subscriber detached and it happened to be the controller\r\n    /// ([`detach_if`]) — an ordinary departure, not a wedge and not a kill.\r\n    // [impl->REQ-REFRESH-WAIT-ATTRIBUTION]\r\n    const CLOSE_DETACH: &'static str = \"detach\";\r\n\r\n    /// Detach the controller and UNLATCH the perch — the ONE unlatch path shared\r\n    /// by W1's deadline-evict ([`mark_controller_gone`]), [`detach_if`], and (per\r\n    /// the W1↔W5 synergy) W5's reconcile self-heal. Dropping the sink ends its\r\n    /// writer thread; clearing the slot re-stamps `driven_by` to `None` (so an\r\n    /// `ONLINE+CONTROLLED` latch cannot outlive a gone controller).\r\n    ///\r\n    /// `cause` is the CALLER'S OWN knowledge of which edge closed the slot, and it\r\n    /// is a parameter for the reason releases#49/#267 needed this line at all: at\r\n    /// this site a writer that exited, an evicted wedge and a deliberate detach are\r\n    /// indistinguishable, so a cause DERIVED here would be a guess wearing a label.\r\n    /// Each caller names what it knows; nothing infers.\r\n    // [impl->REQ-REFRESH-WAIT-ATTRIBUTION]\r\n    fn clear_controller(&mut self, cause: &'static str) {\r\n        let prev_gen = self.controller.as_ref().map(|c| c.attach_gen);\r\n        // Captured BEFORE the take: after it there is no sink left to ask.\r\n        let conn = self.controller.as_ref().map(|c| c.send.id());\r\n        if self.controller.take().is_some() {\r\n            // ONLY ON THE REAL TRANSITION (releases#49/#267 Q1). The epoch-guarded\r\n            // no-op above returns without taking anything, and a line there would\r\n            // report a closure that did not happen — which is exactly the confusion\r\n            // the question exists to remove: today a clean close and a slot that was\r\n            // already empty leave identical stderr.\r\n            // [impl->REQ-REFRESH-WAIT-ATTRIBUTION]\r\n            if crate::brain::diag_trace_enabled() {\r\n                spt_proto::emit_line_err!(\r\n                    \"CONTROLLER_SLOT_CLOSED:{} conn={} cause={cause} {} — the controller slot \\\r\n                     was occupied and is now empty [REQ-REFRESH-WAIT-ATTRIBUTION]\",\r\n                    self.session_id,\r\n                    conn.unwrap_or(0),\r\n                    crate::conn::log_stamp()\r\n                );\r\n            }\r\n            self.stamp_driven_by();\r\n            // Detach is a drop edge: the posture goes offline (derived from the\r\n            // stamp just written) and the empowerments go with it.\r\n            self.on_control_edge(prev_gen, None);\r\n        }\r\n    }\r\n\r\n    /// B-2 (REQ-CONTROLLER-LIVENESS-REAP): lazily drop a controller whose WRITER\r\n    /// THREAD has EXITED — the record-vs-truth honesty step, run on the authoritative\r\n    /// `KIND_SESSIONS` poll BEFORE `controller_by`/`has_controller` are read. The\r\n    /// writer exits when a socket write fails (a severed / keepalive-killed controller\r\n    /// conn) or its `tx` drops; while it is parked on `rx.recv()` (a LIVE but idle\r\n    /// controller) it is NOT finished, so a live controller is never touched (the\r\n    /// no-false-reap guard). Reaping here makes the broker report honest truth so the\r\n    /// EXISTING [`converge_perch_stamps`] clears the now-stale info.json\r\n    /// `driven_by`/`controlled` stamp — no new detach path, no new query.\r\n    ///\r\n    /// LOCK-SAFE: this runs UNDER the log lock in the snapshot closure, so it must do\r\n    /// NO file I/O (the lock-across-effect discipline, KH 7.12/5.16). It drops the\r\n    /// sink in-memory ONLY (like [`become_controller`]'s prior-drop) and does NOT\r\n    /// [`stamp_driven_by`] — the OFF-lock `converge_perch_stamps` writes the honest\r\n    /// stamp. Returns true iff it reaped.\r\n    ///\r\n    /// SCOPE (doyle B-2 reduction): catches a writer that FAILED A WRITE (output was\r\n    /// flowing when the conn severed) or whose detach was dropped by a prior wedge\r\n    /// (the B-1 poison class). A TRULY IDLE severed controller (writer parked on recv,\r\n    /// no output, conn not yet EOF'd) stays `is_finished() == false` and converges\r\n    /// only once output resumes or the conn EOFs (→ `handle_conn` detach) — that\r\n    /// harder active-probe case is the reserved [`REQ-HAZARD-DRIVEN-BY-IDLE-REMOTE-EVICT`],\r\n    /// deliberately NOT built here. KH 7.15 held by construction: only ever CLEARS,\r\n    /// never latches `driven_by`.\r\n    // [impl->REQ-CONTROLLER-LIVENESS-REAP]\r\n    fn reap_dead_controller(&mut self) -> bool {\r\n        let dead = self\r\n            .controller\r\n            .as_ref()\r\n            .is_some_and(|c| c._writer.is_finished());\r\n        if dead {\r\n            // In-memory drop only (ends the — already-exited — writer); NO stamp here\r\n            // (info.json I/O under the log lock is the KH 7.12 hazard). converge does it.\r\n            self.controller = None;\r\n            return true;\r\n        }\r\n        // W2 decouple (REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE): a controller whose\r\n        // writer is BLOCKED (not exited) on a suspended brain conn past\r\n        // BRAIN_WRITE_DEADLINE is also reaped here, so the honest `KIND_SESSIONS`\r\n        // truth + the off-lock converge_perch_stamps release its stale `driven_by`\r\n        // stamp even when nobody take/reattaches. Deferred stamp: the reap runs in\r\n        // the multi-session snapshot closure (no I/O under the shared sessions lock).\r\n        self.stall_evict_controller(Instant::now(), StampMode::Deferred)\r\n    }\r\n\r\n    /// Deadline-evict path (W1): the drain observed a fully-stalled controller and\r\n    /// timed out its `send_timeout`. Detach it — but ONLY if it is still the same\r\n    /// controller generation the drain saw (`epoch`), so a NEWER controller that\r\n    /// attached during the stall is never unseated.\r\n    fn mark_controller_gone(&mut self, epoch: u64) {\r\n        if self.controller_epoch.load(Ordering::Acquire) == epoch {\r\n            self.clear_controller(Self::CLOSE_SEND_DEADLINE);\r\n        }\r\n    }\r\n\r\n    /// Whether the current controller's writer is WEDGED — blocked inside a single\r\n    /// socket write for at least `deadline` (the brain-subscriber conn suspended /\r\n    /// black-holed). An IDLE controller (writer parked on `rx.recv()`, nothing to\r\n    /// write) reads NOT stalled: its [`ControllerSink::write_blocked_since`] is `None`.\r\n    /// Pure over the injected `now` + `deadline` so the unit tests it with no thread.\r\n    // [impl->REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]\r\n    fn controller_write_stalled(&self, now: Instant, deadline: Duration) -> bool {\r\n        self.controller.as_ref().is_some_and(|c| {\r\n            c.write_blocked_since\r\n                .lock()\r\n                .ok()\r\n                .and_then(|g| *g)\r\n                .is_some_and(|since| now.saturating_duration_since(since) >= deadline)\r\n        })\r\n    }\r\n\r\n    /// Whether this session's controller is the brain's OWN LOCAL conn (`by: None`)\r\n    /// AND wedged past `deadline` — the brain-generation promotion DRAINED gate reads\r\n    /// it (REQ-UPDATE-PROMOTE-DRAINED). A killed OLD brain generation whose local\r\n    /// controller conn is black-holed reads wedged here until the W2 stall-evict\r\n    /// releases it (or the conn EOFs); a cleanly-closed conn is simply absent. A\r\n    /// REMOTE operator's wedged conn (`by: Some`) is the W2 stall-evict concern, NOT\r\n    /// an old-gen brain conn, so it is EXCLUDED — a stuck remote viewer must never\r\n    /// gate a brain promotion. Pure over the injected `now`/`deadline` (no thread).\r\n    // [impl->REQ-UPDATE-PROMOTE-DRAINED]\r\n    fn local_controller_wedged(&self, now: Instant, deadline: Duration) -> bool {\r\n        self.controller.as_ref().is_some_and(|c| c.by.is_none())\r\n            && self.controller_write_stalled(now, deadline)\r\n    }\r\n\r\n    /// Stall-evict a controller wedged past [`BRAIN_WRITE_DEADLINE`]: clear the slot\r\n    /// and re-stamp `driven_by`/`controlled` (the BROKER-side release — no brain\r\n    /// round-trip, so `controlled-by` can never latch behind a dead transit) and\r\n    /// record the evict for observability (broker log now; W3 persists the stderr,\r\n    /// the tally surfaces on `daemon status`). The take/reattach path and the\r\n    /// `KIND_SESSIONS` reap call this so a suspended brain subscriber can never hold\r\n    /// control on an idle session (the output-driven [`mark_controller_gone`] never\r\n    /// fires without new output). Returns true iff it evicted. `stamp` selects whether\r\n    /// the perch re-stamp happens inline (the take/reattach path, which already stamps\r\n    /// under the log lock via [`become_controller`]) or is deferred to the off-lock\r\n    /// `converge_perch_stamps` (the multi-session reap closure, KH 7.12 no-I/O rule).\r\n    // [impl->REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]\r\n    fn stall_evict_controller(&mut self, now: Instant, stamp: StampMode) -> bool {\r\n        let deadline = brain_write_deadline();\r\n        if !self.controller_write_stalled(now, deadline) {\r\n            return false;\r\n        }\r\n        // Attribution (F-039 leg b): the evicted controller's stable conn id +\r\n        // stamps ride the record. [impl->REQ-CONN-POISON-ATTRIBUTION]\r\n        let conn = self.controller.as_ref().map(|c| c.send.id()).unwrap_or(0);\r\n        spt_proto::emit_line_err!(\r\n            \"BRAIN_SUBSCRIBER_STALL_EVICT:{} conn={conn} {} — controller writer blocked > {}ms \\\r\n             on a suspended/black-holed brain conn; detached, driven_by released \\\r\n             [REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]\",\r\n            self.session_id,\r\n            crate::conn::log_stamp(),\r\n            deadline.as_millis()\r\n        );\r\n        record_stall_evict();\r\n        match stamp {\r\n            // Inline: drop the sink AND re-stamp the perch (become_controller's\r\n            // sibling — the take/reattach path is already I/O-under-log-lock).\r\n            StampMode::Inline => self.clear_controller(Self::CLOSE_STALL_EVICT),\r\n            // Deferred: in-memory drop ONLY (the reap closure forbids I/O under the\r\n            // shared sessions lock); converge_perch_stamps writes the release off-lock.\r\n            StampMode::Deferred => self.controller = None,\r\n        }\r\n        true\r\n    }\r\n\r\n    /// Add a read-only viewer: snapshot the ring from `from_seq` into the writer\r\n    /// thread (NOT through the bounded live queue — replay must not trip the\r\n    /// overflow eviction), seed it the current size frame first, then run the\r\n    /// live queue. The drain thread fans live frames in via `try_send`.\r\n    fn add_viewer(&mut self, sub: SharedSend, from_seq: u64, by: Option<&str>) {\r\n        let vid = self.next_viewer_id;\r\n        self.next_viewer_id += 1;\r\n        // Attribution (F-039 legs b+d): viewer role fact + attach record.\r\n        // [impl->REQ-CONN-POISON-ATTRIBUTION]\r\n        sub.describe(&format!(\r\n            \"viewer session={} endpoint={} vid={vid}\",\r\n            self.session_id,\r\n            if self.endpoint.is_empty() { \"-\" } else { &self.endpoint }\r\n        ));\r\n        sub.lifecycle_event(\r\n            \"viewer-attach\",\r\n            &format!(\"session={} vid={vid} from_seq={from_seq}\", self.session_id),\r\n        );\r\n        // The size frame + the initial batch (owned by the thread, not the bounded\r\n        // queue). A COLD viewer attach (`from_seq == 0`) — or one resuming from\r\n        // BELOW the presentation floor (REQ-RC-RESIZE-PRESENTATION-BARRIER: never\r\n        // a raw replay across a transition boundary) — gets the #6 clean repaint;\r\n        // a skip-to-live / snap re-attach at/above the floor (incl. `u64::MAX`)\r\n        // keeps the raw-ring slice (`s >= from_seq`) — empty for skip-to-live, as\r\n        // before.\r\n        let mut initial: Vec<Envelope> =\r\n            vec![size_envelope(self.session_id, self.size.0, self.size.1)];\r\n        match self.repaint_initial(from_seq) {\r\n            // The repaint batch rides the sync flag: it supersedes everything\r\n            // before its pseudo-seq (the wire-flag ruling — a below-floor viewer\r\n            // resume baselines the jump instead of gap-erroring).\r\n            Some((seq, bytes)) => {\r\n                // OBS RIDER Tap B (repaint): the synthesized clean-screen viewer\r\n                // repaint bytes (debug-gated, default OFF).\r\n                spt_store::obstap::tap(self.session_id, \"repaint\", &bytes);\r\n                initial.push(sync_output_envelope(self.session_id, seq, &bytes));\r\n            }\r\n            None => initial.extend(\r\n                self.ring\r\n                    .iter()\r\n                    .filter(|(s, _)| *s >= from_seq)\r\n                    .map(|(seq, bytes)| output_envelope(self.session_id, *seq, bytes)),\r\n            ),\r\n        }\r\n        let (tx, rx) = sync_channel::<Envelope>(VIEWER_CHANNEL_DEPTH);\r\n        let writer_send = Arc::clone(&sub);\r\n        let evicted = Arc::new(AtomicBool::new(false));\r\n        let writer_evicted = Arc::clone(&evicted);\r\n        let session_id = self.session_id;\r\n        let writer =\r\n            thread::spawn(move || viewer_writer(writer_send, session_id, initial, rx, writer_evicted));\r\n        self.viewers.insert(\r\n            vid,\r\n            ViewerSink {\r\n                tx,\r","truncated":true,"truncatedBy":"lines","totalLines":14134,"totalBytes":14158,"outputLines":240,"outputBytes":14158,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":16852,"displayContent":{"text":"    fn mark_revoked(&self, by: &str) {\n…\n    }\r\n\r\n    /// The drain timed out the controller's `send_timeout` and the epoch guard\r\n    /// still matched ([`mark_controller_gone`], W1). The writer was BLOCKED past\r\n    /// the deadline -- it is NOT established to have exited or dropped its `tx`,\r\n    /// so this names the deadline the drain observed and not the writer's fate.\r\n    /// This is the edge releases#49/#267 Q1 asks about: a cleanly-closed\r\n    /// OLD-generation conn arrives here too, and from this site the two are\r\n    /// indistinguishable.\r\n    // [impl->REQ-REFRESH-WAIT-ATTRIBUTION]\r\n    const CLOSE_SEND_DEADLINE: &'static str = \"send_deadline\";\r\n    /// The wedge evict released the slot ([`stall_evict_controller`], inline stamp\r\n    /// mode). `BRAIN_SUBSCRIBER_STALL_EVICT` already names the wedge; this names\r\n    /// the slot transition that followed it.\r\n    // [impl->REQ-REFRESH-WAIT-ATTRIBUTION]\r\n    const CLOSE_STALL_EVICT: &'static str = \"stall_evict\";\r\n    /// A subscriber detached and it happened to be the controller\r\n    /// ([`detach_if`]) — an ordinary departure, not a wedge and not a kill.\r\n    // [impl->REQ-REFRESH-WAIT-ATTRIBUTION]\r\n    const CLOSE_DETACH: &'static str = \"detach\";\r\n\r\n    /// Detach the controller and UNLATCH the perch — the ONE unlatch path shared\r\n    /// by W1's deadline-evict ([`mark_controller_gone`]), [`detach_if`], and (per\r\n    /// the W1↔W5 synergy) W5's reconcile self-heal. Dropping the sink ends its\r\n    /// writer thread; clearing the slot re-stamps `driven_by` to `None` (so an\r\n    /// `ONLINE+CONTROLLED` latch cannot outlive a gone controller).\r\n    ///\r\n    /// `cause` is the CALLER'S OWN knowledge of which edge closed the slot, and it\r\n    /// is a parameter for the reason releases#49/#267 needed this line at all: at\r\n    /// this site a writer that exited, an evicted wedge and a deliberate detach are\r\n    /// indistinguishable, so a cause DERIVED here would be a guess wearing a label.\r\n    /// Each caller names what it knows; nothing infers.\r\n    // [impl->REQ-REFRESH-WAIT-ATTRIBUTION]\r\n    fn clear_controller(&mut self, cause: &'static str) {\r\n        let prev_gen = self.controller.as_ref().map(|c| c.attach_gen);\r\n        // Captured BEFORE the take: after it there is no sink left to ask.\r\n        let conn = self.controller.as_ref().map(|c| c.send.id());\r\n        if self.controller.take().is_some() {\r\n            // ONLY ON THE REAL TRANSITION (releases#49/#267 Q1). The epoch-guarded\r\n            // no-op above returns without taking anything, and a line there would\r\n            // report a closure that did not happen — which is exactly the confusion\r\n            // the question exists to remove: today a clean close and a slot that was\r\n            // already empty leave identical stderr.\r\n            // [impl->REQ-REFRESH-WAIT-ATTRIBUTION]\r\n            if crate::brain::diag_trace_enabled() {\r\n                spt_proto::emit_line_err!(\r\n                    \"CONTROLLER_SLOT_CLOSED:{} conn={} cause={cause} {} — the controller slot \\\r\n                     was occupied and is now empty [REQ-REFRESH-WAIT-ATTRIBUTION]\",\r\n                    self.session_id,\r\n                    conn.unwrap_or(0),\r\n                    crate::conn::log_stamp()\r\n                );\r\n            }\r\n            self.stamp_driven_by();\r\n            // Detach is a drop edge: the posture goes offline (derived from the\r\n            // stamp just written) and the empowerments go with it.\r\n            self.on_control_edge(prev_gen, None);\r\n        }\r\n    }\r\n\r\n    /// B-2 (REQ-CONTROLLER-LIVENESS-REAP): lazily drop a controller whose WRITER\r\n    /// THREAD has EXITED — the record-vs-truth honesty step, run on the authoritative\r\n    /// `KIND_SESSIONS` poll BEFORE `controller_by`/`has_controller` are read. The\r\n    /// writer exits when a socket write fails (a severed / keepalive-killed controller\r\n    /// conn) or its `tx` drops; while it is parked on `rx.recv()` (a LIVE but idle\r\n    /// controller) it is NOT finished, so a live controller is never touched (the\r\n    /// no-false-reap guard). Reaping here makes the broker report honest truth so the\r\n    /// EXISTING [`converge_perch_stamps`] clears the now-stale info.json\r\n    /// `driven_by`/`controlled` stamp — no new detach path, no new query.\r\n    ///\r\n    /// LOCK-SAFE: this runs UNDER the log lock in the snapshot closure, so it must do\r\n    /// NO file I/O (the lock-across-effect discipline, KH 7.12/5.16). It drops the\r\n    /// sink in-memory ONLY (like [`become_controller`]'s prior-drop) and does NOT\r\n    /// [`stamp_driven_by`] — the OFF-lock `converge_perch_stamps` writes the honest\r\n    /// stamp. Returns true iff it reaped.\r\n    ///\r\n    /// SCOPE (doyle B-2 reduction): catches a writer that FAILED A WRITE (output was\r\n    /// flowing when the conn severed) or whose detach was dropped by a prior wedge\r\n    /// (the B-1 poison class). A TRULY IDLE severed controller (writer parked on recv,\r\n    /// no output, conn not yet EOF'd) stays `is_finished() == false` and converges\r\n    /// only once output resumes or the conn EOFs (→ `handle_conn` detach) — that\r\n    /// harder active-probe case is the reserved [`REQ-HAZARD-DRIVEN-BY-IDLE-REMOTE-EVICT`],\r\n    /// deliberately NOT built here. KH 7.15 held by construction: only ever CLEARS,\r\n    /// never latches `driven_by`.\r\n    // [impl->REQ-CONTROLLER-LIVENESS-REAP]\r\n    fn reap_dead_controller(&mut self) -> bool {\r\n        let dead = self\r\n            .controller\r\n            .as_ref()\r\n            .is_some_and(|c| c._writer.is_finished());\r\n        if dead {\r\n            // In-memory drop only (ends the — already-exited — writer); NO stamp here\r\n            // (info.json I/O under the log lock is the KH 7.12 hazard). converge does it.\r\n            self.controller = None;\r\n            return true;\r\n        }\r\n        // W2 decouple (REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE): a controller whose\r\n        // writer is BLOCKED (not exited) on a suspended brain conn past\r\n        // BRAIN_WRITE_DEADLINE is also reaped here, so the honest `KIND_SESSIONS`\r\n        // truth + the off-lock converge_perch_stamps release its stale `driven_by`\r\n        // stamp even when nobody take/reattaches. Deferred stamp: the reap runs in\r\n        // the multi-session snapshot closure (no I/O under the shared sessions lock).\r\n        self.stall_evict_controller(Instant::now(), StampMode::Deferred)\r\n    }\r\n\r\n    /// Deadline-evict path (W1): the drain observed a fully-stalled controller and\r\n    /// timed out its `send_timeout`. Detach it — but ONLY if it is still the same\r\n    /// controller generation the drain saw (`epoch`), so a NEWER controller that\r\n    /// attached during the stall is never unseated.\r\n    fn mark_controller_gone(&mut self, epoch: u64) {\r\n        if self.controller_epoch.load(Ordering::Acquire) == epoch {\r\n            self.clear_controller(Self::CLOSE_SEND_DEADLINE);\r\n        }\r\n    }\r\n\r\n    /// Whether the current controller's writer is WEDGED — blocked inside a single\r\n    /// socket write for at least `deadline` (the brain-subscriber conn suspended /\r\n    /// black-holed). An IDLE controller (writer parked on `rx.recv()`, nothing to\r\n    /// write) reads NOT stalled: its [`ControllerSink::write_blocked_since`] is `None`.\r\n    /// Pure over the injected `now` + `deadline` so the unit tests it with no thread.\r\n    // [impl->REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]\r\n    fn controller_write_stalled(&self, now: Instant, deadline: Duration) -> bool {\r\n        self.controller.as_ref().is_some_and(|c| {\r\n            c.write_blocked_since\r\n                .lock()\r\n                .ok()\r\n                .and_then(|g| *g)\r\n                .is_some_and(|since| now.saturating_duration_since(since) >= deadline)\r\n        })\r\n    }\r\n\r\n    /// Whether this session's controller is the brain's OWN LOCAL conn (`by: None`)\r\n    /// AND wedged past `deadline` — the brain-generation promotion DRAINED gate reads\r\n    /// it (REQ-UPDATE-PROMOTE-DRAINED). A killed OLD brain generation whose local\r\n    /// controller conn is black-holed reads wedged here until the W2 stall-evict\r\n    /// releases it (or the conn EOFs); a cleanly-closed conn is simply absent. A\r\n    /// REMOTE operator's wedged conn (`by: Some`) is the W2 stall-evict concern, NOT\r\n    /// an old-gen brain conn, so it is EXCLUDED — a stuck remote viewer must never\r\n    /// gate a brain promotion. Pure over the injected `now`/`deadline` (no thread).\r\n    // [impl->REQ-UPDATE-PROMOTE-DRAINED]\r\n    fn local_controller_wedged(&self, now: Instant, deadline: Duration) -> bool {\r\n        self.controller.as_ref().is_some_and(|c| c.by.is_none())\r\n            && self.controller_write_stalled(now, deadline)\r\n    }\r\n\r\n    /// Stall-evict a controller wedged past [`BRAIN_WRITE_DEADLINE`]: clear the slot\r\n    /// and re-stamp `driven_by`/`controlled` (the BROKER-side release — no brain\r\n    /// round-trip, so `controlled-by` can never latch behind a dead transit) and\r\n    /// record the evict for observability (broker log now; W3 persists the stderr,\r\n    /// the tally surfaces on `daemon status`). The take/reattach path and the\r\n    /// `KIND_SESSIONS` reap call this so a suspended brain subscriber can never hold\r\n    /// control on an idle session (the output-driven [`mark_controller_gone`] never\r\n    /// fires without new output). Returns true iff it evicted. `stamp` selects whether\r\n    /// the perch re-stamp happens inline (the take/reattach path, which already stamps\r\n    /// under the log lock via [`become_controller`]) or is deferred to the off-lock\r\n    /// `converge_perch_stamps` (the multi-session reap closure, KH 7.12 no-I/O rule).\r\n    // [impl->REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]\r\n    fn stall_evict_controller(&mut self, now: Instant, stamp: StampMode) -> bool {\r\n        let deadline = brain_write_deadline();\r\n        if !self.controller_write_stalled(now, deadline) {\r\n            return false;\r\n        }\r\n        // Attribution (F-039 leg b): the evicted controller's stable conn id +\r\n        // stamps ride the record. [impl->REQ-CONN-POISON-ATTRIBUTION]\r\n        let conn = self.controller.as_ref().map(|c| c.send.id()).unwrap_or(0);\r\n        spt_proto::emit_line_err!(\r\n            \"BRAIN_SUBSCRIBER_STALL_EVICT:{} conn={conn} {} — controller writer blocked > {}ms \\\r\n             on a suspended/black-holed brain conn; detached, driven_by released \\\r\n             [REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]\",\r\n            self.session_id,\r\n            crate::conn::log_stamp(),\r\n            deadline.as_millis()\r\n        );\r\n        record_stall_evict();\r\n        match stamp {\r\n            // Inline: drop the sink AND re-stamp the perch (become_controller's\r\n            // sibling — the take/reattach path is already I/O-under-log-lock).\r\n            StampMode::Inline => self.clear_controller(Self::CLOSE_STALL_EVICT),\r\n            // Deferred: in-memory drop ONLY (the reap closure forbids I/O under the\r\n            // shared sessions lock); converge_perch_stamps writes the release off-lock.\r\n            StampMode::Deferred => self.controller = None,\r\n        }\r\n        true\r\n    }\r\n\r\n    /// Add a read-only viewer: snapshot the ring from `from_seq` into the writer\r\n    /// thread (NOT through the bounded live queue — replay must not trip the\r\n    /// overflow eviction), seed it the current size frame first, then run the\r\n    /// live queue. The drain thread fans live frames in via `try_send`.\r\n    fn add_viewer(&mut self, sub: SharedSend, from_seq: u64, by: Option<&str>) {\r\n        let vid = self.next_viewer_id;\r\n        self.next_viewer_id += 1;\r\n        // Attribution (F-039 legs b+d): viewer role fact + attach record.\r\n        // [impl->REQ-CONN-POISON-ATTRIBUTION]\r\n        sub.describe(&format!(\r\n            \"viewer session={} endpoint={} vid={vid}\",\r\n            self.session_id,\r\n            if self.endpoint.is_empty() { \"-\" } else { &self.endpoint }\r\n        ));\r\n        sub.lifecycle_event(\r\n            \"viewer-attach\",\r\n            &format!(\"session={} vid={vid} from_seq={from_seq}\", self.session_id),\r\n        );\r\n        // The size frame + the initial batch (owned by the thread, not the bounded\r\n        // queue). A COLD viewer attach (`from_seq == 0`) — or one resuming from\r\n        // BELOW the presentation floor (REQ-RC-RESIZE-PRESENTATION-BARRIER: never\r\n        // a raw replay across a transition boundary) — gets the #6 clean repaint;\r\n        // a skip-to-live / snap re-attach at/above the floor (incl. `u64::MAX`)\r\n        // keeps the raw-ring slice (`s >= from_seq`) — empty for skip-to-live, as\r\n        // before.\r\n        let mut initial: Vec<Envelope> =\r\n            vec![size_envelope(self.session_id, self.size.0, self.size.1)];\r\n        match self.repaint_initial(from_seq) {\r\n            // The repaint batch rides the sync flag: it supersedes everything\r\n            // before its pseudo-seq (the wire-flag ruling — a below-floor viewer\r\n            // resume baselines the jump instead of gap-erroring).\r\n            Some((seq, bytes)) => {\r\n                // OBS RIDER Tap B (repaint): the synthesized clean-screen viewer\r\n                // repaint bytes (debug-gated, default OFF).\r\n                spt_store::obstap::tap(self.session_id, \"repaint\", &bytes);\r\n                initial.push(sync_output_envelope(self.session_id, seq, &bytes));\r\n            }\r\n            None => initial.extend(\r\n                self.ring\r\n                    .iter()\r\n                    .filter(|(s, _)| *s >= from_seq)\r\n                    .map(|(seq, bytes)| output_envelope(self.session_id, *seq, bytes)),\r\n            ),\r\n        }\r\n        let (tx, rx) = sync_channel::<Envelope>(VIEWER_CHANNEL_DEPTH);\r\n        let writer_send = Arc::clone(&sub);\r\n        let evicted = Arc::new(AtomicBool::new(false));\r\n        let writer_evicted = Arc::clone(&evicted);\r\n        let session_id = self.session_id;\r\n        let writer =\r\n            thread::spawn(move || viewer_writer(writer_send, session_id, initial, rx, writer_evicted));\r\n        self.viewers.insert(\r\n            vid,\r\n            ViewerSink {\r\n                tx,\r\n…\n            },\n        );\n    }","startLine":2709,"lineNumbers":[2709,null,2719,2720,2721,2722,2723,2724,2725,2726,2727,2728,2729,2730,2731,2732,2733,2734,2735,2736,2737,2738,2739,2740,2741,2742,2743,2744,2745,2746,2747,2748,2749,2750,2751,2752,2753,2754,2755,2756,2757,2758,2759,2760,2761,2762,2763,2764,2765,2766,2767,2768,2769,2770,2771,2772,2773,2774,2775,2776,2777,2778,2779,2780,2781,2782,2783,2784,2785,2786,2787,2788,2789,2790,2791,2792,2793,2794,2795,2796,2797,2798,2799,2800,2801,2802,2803,2804,2805,2806,2807,2808,2809,2810,2811,2812,2813,2814,2815,2816,2817,2818,2819,2820,2821,2822,2823,2824,2825,2826,2827,2828,2829,2830,2831,2832,2833,2834,2835,2836,2837,2838,2839,2840,2841,2842,2843,2844,2845,2846,2847,2848,2849,2850,2851,2852,2853,2854,2855,2856,2857,2858,2859,2860,2861,2862,2863,2864,2865,2866,2867,2868,2869,2870,2871,2872,2873,2874,2875,2876,2877,2878,2879,2880,2881,2882,2883,2884,2885,2886,2887,2888,2889,2890,2891,2892,2893,2894,2895,2896,2897,2898,2899,2900,2901,2902,2903,2904,2905,2906,2907,2908,2909,2910,2911,2912,2913,2914,2915,2916,2917,2918,2919,2920,2921,2922,2923,2924,2925,2926,2927,2928,2929,2930,2931,2932,2933,2934,2935,2936,2937,2938,2939,2940,2941,2942,2943,2944,2945,2946,2947,2948,2949,2950,2951,2952,2953,2954,2955,2956,2957,2958,null,2964,2965,2966]},"fileSize":877537,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\crates\\spt-daemon\\src\\broker.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":16852,"totalBytes":14158,"outputLines":240,"outputBytes":14158,"shownRange":{"start":2719,"end":2958},"nextOffset":2959}}},"isError":false,"timestamp":1789523963481}}
{"type":"message","id":"a65605ec","parentId":"ec21636b","timestamp":"2026-09-16T01:59:23.483Z","message":{"role":"toolResult","toolCallId":"call_W830bcvnVfmLBUl6kty69OXU|fc_0d104aba2b0f3edc016aa9f7fb57c087d18aba9be592f6a8a9","toolName":"glob","content":[{"type":"text","text":"traceable-reqs.toml\n# .github/ci/\ntraceable-reqs.sh\ntraceable-reqs-selftest.sh"}],"details":{"scopePath":"**/*hazard*, **/*trace*, **/*307*","fileCount":3,"files":["traceable-reqs.toml",".github/ci/traceable-reqs.sh",".github/ci/traceable-reqs-selftest.sh"],"truncated":false,"cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core"},"isError":false,"timestamp":1789523963483}}
{"type":"message","id":"df262790","parentId":"a65605ec","timestamp":"2026-09-16T01:59:23.485Z","message":{"role":"toolResult","toolCallId":"call_i38LYlWR0TVP2O8KWwZaiHN7|fc_0d104aba2b0f3edc016aa9f7fb57b887d1b1ddff4b18c6cabd","toolName":"grep","content":[{"type":"text","text":" 34|SPAWN_PHASE:w1crig phase=row_inserted t_ms=1204 session=1\n*35|CONTROLLER_SLOT_CLOSED:1 conn=13 cause=detach wall_ms=1789523754018 mono_ms=7088 — the controller slot was occupied and is now empty [REQ-REFRESH-WAIT-ATTRIBUTION]\n 36|CONN_LIFECYCLE: conn=13 role=brain controller session=1 endpoint=w1crig by=local wall_ms=1789523754018 mono_ms=7088 first_write_wall_ms=1789523754018 first_write_mono_ms=7088 event=writer-exit role=controller session=1 reason=channel-closed [REQ-CONN-POISON-ATTRIBUTION]\n 37|CONN_LIFECYCLE: conn=13 role=brain controller session=1 endpoint=w1crig by=local wall_ms=1789523754032 mono_ms=7102 first_write_wall_ms=1789523754018 first_write_mono_ms=7088 event=session-detach session=1 was_controller=true viewers_dropped=0 [REQ-CONN-POISON-ATTRIBUTION]\n 38|CONN_LIFECYCLE: conn=15 role=brain stream-subscriber stream=2 wall_ms=1789523756051 mono_ms=9122 first_write_wall_ms=1789523756007 first_write_mono_ms=9077 event=stream-sub-attach stream=2 from_seq=0 [REQ-CONN-POISON-ATTRIBUTION]\n*39|CONN_LIFECYCLE: conn=17 role=brain stream-subscriber stream=1 wall_ms=1789523756142 mono_ms=9213 first_write_wall_ms=1789523756142 first_write_mono_ms=9213 event=stream-sub-attach stream=1 from_seq=0 [REQ-CONN-POISON-ATTRIBUTION]\n 40|ATTACH_INTENT_CHOSEN: site=serve_dispatch session=1 intent=control from_seq=0 gen=1789523756051 coded=false\n*41|CONN_LIFECYCLE: conn=17 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523756143 mono_ms=9214 first_write_wall_ms=1789523756142 first_write_mono_ms=9213 event=controller-attach session=1 from_seq=0 [REQ-CONN-POISON-ATTRIBUTION]\n 42|ATTACH_INTENT_CHOSEN: site=serve_dispatch session=1 intent=control from_seq=0 gen=1789523756051 coded=false\n*43|SUBSCRIBE_DECISION: session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf conn=17 intent=control old_by=- old_gen=0 req_gen=1789523756051 decision=controller\n*44|SUBSCRIBE_DECISION: session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf conn=17 intent=control old_by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf old_gen=1789523756051 req_gen=1789523756051 decision=idempotent\n*45|BRAIN_UPDATE_RESTART: planned restart — cycling brain onto the swapped binary\n*46|CONTROLLER_SLOT_CLOSED:1 conn=17 cause=detach wall_ms=1789523803471 mono_ms=56541 — the controller slot was occupied and is now empty [REQ-REFRESH-WAIT-ATTRIBUTION]\n*47|CONN_WRITE_RETIRED: conn=17 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523803471 mono_ms=56541 first_write_wall_ms=1789523756142 first_write_mono_ms=9213 — organic write failure (BrokenPipe: The pipe is being closed. (os error 232)); physical broker conn poisoned + retired, never reused (same retire behavior as POISONED — only the failure class differs) [REQ-CONN-POISON-DIAL-SCOPE]\n*48|CONN_LIFECYCLE: conn=17 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523803471 mono_ms=56541 first_write_wall_ms=1789523756142 first_write_mono_ms=9213 event=writer-exit role=controller session=1 reason=write-failed kind=BrokenPipe [REQ-CONN-POISON-ATTRIBUTION]\n 49|BRAIN_RESTART: planned update respawn (generation 1)\n 50|=== spt brain stderr — generation 1 — pid 53752 ===\n 51|BRAIN_UP: connected to broker spt-daemon-broker-573ad82e70ef3afe.sock (pid 53752, generation 1, start-reason update)\n 52|BRAIN_PHASE:announce done in 0ms\n*53|CONN_LIFECYCLE: conn=17 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523803573 mono_ms=56643 first_write_wall_ms=1789523756142 first_write_mono_ms=9213 event=session-detach session=1 was_controller=true viewers_dropped=0 [REQ-CONN-POISON-ATTRIBUTION]\n*54|CONN_LIFECYCLE: conn=17 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523803573 mono_ms=56643 first_write_wall_ms=1789523756142 first_write_mono_ms=9213 event=transport-close poisoned=true timed_out=false [REQ-CONN-POISON-ATTRIBUTION]\n*55|BRAIN_RESUMED: re-established 1 session cursor(s): [1]\n 56|BRAIN_PHASE:resume done in 69ms\n 57|BRAIN_PHASE:ready written in 1894ms (exe-hash 1893ms, computed=true) path C:\\Users\\decid\\Documents\\projects\\spt-core\\.spt\\preserved\\307\\hertz-w1bc\\w1c-run1\\home-w1c\\brain.ready\n 58|BRAIN_CONN_OPEN: caller=livehost-session-query reason=initial-query policy=reuse-until-query-error wall_ms=1789523805474 mono_ms=1966\n...\n 64|CONN_LIFECYCLE: conn=68 role=brain presence-subscriber wall_ms=1789523805974 mono_ms=59045 first_write=none event=presence-sub-attach from_seq=18446744073709551615 [REQ-CONN-POISON-ATTRIBUTION]\n*65|CONN_LIFECYCLE: conn=70 role=brain stream-subscriber stream=1 wall_ms=1789523805975 mono_ms=59046 first_write_wall_ms=1789523805975 first_write_mono_ms=59046 event=stream-sub-attach stream=1 from_seq=0 [REQ-CONN-POISON-ATTRIBUTION]\n 66|ATTACH_INTENT_CHOSEN: site=serve_dispatch session=1 intent=control from_seq=0 gen=1789523756051 coded=false\n*67|CONN_LIFECYCLE: conn=70 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523805976 mono_ms=59047 first_write_wall_ms=1789523805975 first_write_mono_ms=59046 event=controller-attach session=1 from_seq=0 [REQ-CONN-POISON-ATTRIBUTION]\n 68|ATTACH_INTENT_CHOSEN: site=serve_dispatch session=1 intent=control from_seq=0 gen=1789523756051 coded=false\n*69|SUBSCRIBE_DECISION: session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf conn=70 intent=control old_by=- old_gen=0 req_gen=1789523756051 decision=controller\n*70|SUBSCRIBE_DECISION: session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf conn=70 intent=control old_by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf old_gen=1789523756051 req_gen=1789523756051 decision=idempotent\n 71|CONN_LIFECYCLE_SUMMARY: scope=broker-ipc wall_ms=1789523807721 mono_ms=60792 interval_ms=60674 starts=47 closes=63 [REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n*72|CONN_WRITE_POISONED: conn=70 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523821012 mono_ms=74082 first_write_wall_ms=1789523805975 first_write_mono_ms=59046 — physical broker conn retired: a write exceeded its 15000ms bound; in-flight I/O aborted, conn never reused [REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\n*73|CONN_LIFECYCLE: conn=70 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523821012 mono_ms=74082 first_write_wall_ms=1789523805975 first_write_mono_ms=59046 event=stream-sub-writer-poison live write failed [REQ-CONN-POISON-ATTRIBUTION]\n*74|BRAIN_SUBSCRIBER_STALL_EVICT:1 conn=70 wall_ms=1789523821012 mono_ms=74082 — controller conn write exceeded its bound; physical conn poisoned + retired (writer released, EOF cleanup follows) [REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]\n*75|CONTROLLER_SLOT_CLOSED:1 conn=70 cause=detach wall_ms=1789523821012 mono_ms=74083 — the controller slot was occupied and is now empty [REQ-REFRESH-WAIT-ATTRIBUTION]\n*76|CONN_LIFECYCLE: conn=70 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523821012 mono_ms=74083 first_write_wall_ms=1789523805975 first_write_mono_ms=59046 event=writer-exit role=controller session=1 reason=write-failed kind=TimedOut [REQ-CONN-POISON-ATTRIBUTION]\n 77|DISPATCH:1:Failed(\"No process is on the other end of the pipe. (os error 233)\")\n 78|DISPATCH_EV: gen=1789523805974 stream=1 family=Attach attempt=0 conn=1 event=worker-outcome detail=Failed(\"No process is on the other end of the pipe. (os error 233)\") elapsed_ms=15037 endpoint=-\n*79|CONN_LIFECYCLE: conn=70 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523821028 mono_ms=74099 first_write_wall_ms=1789523805975 first_write_mono_ms=59046 event=session-detach session=1 was_controller=true viewers_dropped=0 [REQ-CONN-POISON-ATTRIBUTION]\n*80|CONN_LIFECYCLE: conn=70 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523821028 mono_ms=74099 first_write_wall_ms=1789523805975 first_write_mono_ms=59046 event=transport-close poisoned=true timed_out=true [REQ-CONN-POISON-ATTRIBUTION]\n*81|CONN_LIFECYCLE: conn=88 role=brain stream-subscriber stream=1 wall_ms=1789523821556 mono_ms=74627 first_write_wall_ms=1789523821556 first_write_mono_ms=74627 event=stream-sub-attach stream=1 from_seq=0 [REQ-CONN-POISON-ATTRIBUTION]\n 82|ATTACH_INTENT_CHOSEN: site=serve_dispatch session=1 intent=control from_seq=0 gen=1789523756051 coded=false\n*83|CONN_LIFECYCLE: conn=88 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523821557 mono_ms=74628 first_write_wall_ms=1789523821556 first_write_mono_ms=74627 event=controller-attach session=1 from_seq=0 [REQ-CONN-POISON-ATTRIBUTION]\n 84|ATTACH_INTENT_CHOSEN: site=serve_dispatch session=1 intent=control from_seq=0 gen=1789523756051 coded=false\n*85|SUBSCRIBE_DECISION: session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf conn=88 intent=control old_by=- old_gen=0 req_gen=1789523756051 decision=controller\n*86|SUBSCRIBE_DECISION: session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf conn=88 intent=control old_by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf old_gen=1789523756051 req_gen=1789523756051 decision=idempotent\n*87|CONN_WRITE_POISONED: conn=88 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523836597 mono_ms=89667 first_write_wall_ms=1789523821556 first_write_mono_ms=74627 — physical broker conn retired: a write exceeded its 15000ms bound; in-flight I/O aborted, conn never reused [REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\n*88|CONN_LIFECYCLE: conn=88 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523836597 mono_ms=89668 first_write_wall_ms=1789523821556 first_write_mono_ms=74627 event=stream-sub-writer-poison live write failed [REQ-CONN-POISON-ATTRIBUTION]\n*89|BRAIN_SUBSCRIBER_STALL_EVICT:1 conn=88 wall_ms=1789523836597 mono_ms=89668 — controller conn write exceeded its bound; physical conn poisoned + retired (writer released, EOF cleanup follows) [REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]\n*90|CONTROLLER_SLOT_CLOSED:1 conn=88 cause=detach wall_ms=1789523836597 mono_ms=89668 — the controller slot was occupied and is now empty [REQ-REFRESH-WAIT-ATTRIBUTION]\n*91|CONN_LIFECYCLE: conn=88 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523836597 mono_ms=89668 first_write_wall_ms=1789523821556 first_write_mono_ms=74627 event=writer-exit role=controller session=1 reason=write-failed kind=TimedOut [REQ-CONN-POISON-ATTRIBUTION]\n 92|DISPATCH:1:Failed(\"No process is on the other end of the pipe. (os error 233)\")\n 93|DISPATCH_EV: gen=1789523805974 stream=1 family=Attach attempt=1 conn=1 event=worker-outcome detail=Failed(\"No process is on the other end of the pipe. (os error 233)\") elapsed_ms=15041 endpoint=-\n*94|CONN_LIFECYCLE: conn=88 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523836612 mono_ms=89683 first_write_wall_ms=1789523821556 first_write_mono_ms=74627 event=session-detach session=1 was_controller=true viewers_dropped=0 [REQ-CONN-POISON-ATTRIBUTION]\n*95|CONN_LIFECYCLE: conn=88 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523836612 mono_ms=89683 first_write_wall_ms=1789523821556 first_write_mono_ms=74627 event=transport-close poisoned=true timed_out=true [REQ-CONN-POISON-ATTRIBUTION]\n*96|CONN_LIFECYCLE: conn=104 role=brain stream-subscriber stream=1 wall_ms=1789523837637 mono_ms=90707 first_write_wall_ms=1789523837637 first_write_mono_ms=90707 event=stream-sub-attach stream=1 from_seq=0 [REQ-CONN-POISON-ATTRIBUTION]\n 97|ATTACH_INTENT_CHOSEN: site=serve_dispatch session=1 intent=control from_seq=0 gen=1789523756051 coded=false\n*98|CONN_LIFECYCLE: conn=104 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523837637 mono_ms=90708 first_write_wall_ms=1789523837637 first_write_mono_ms=90707 event=controller-attach session=1 from_seq=0 [REQ-CONN-POISON-ATTRIBUTION]\n 99|ATTACH_INTENT_CHOSEN: site=serve_dispatch session=1 intent=control from_seq=0 gen=1789523756051 coded=false\n*100|SUBSCRIBE_DECISION: session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf conn=104 intent=control old_by=- old_gen=0 req_gen=1789523756051 decision=controller\n*101|SUBSCRIBE_DECISION: session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf conn=104 intent=control old_by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf old_gen=1789523756051 req_gen=1789523756051 decision=idempotent\n*102|CONN_WRITE_POISONED: conn=104 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523852688 mono_ms=105759 first_write_wall_ms=1789523837637 first_write_mono_ms=90707 — physical broker conn retired: a write exceeded its 15000ms bound; in-flight I/O aborted, conn never reused [REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\n*103|CONN_LIFECYCLE: conn=104 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523852688 mono_ms=105759 first_write_wall_ms=1789523837637 first_write_mono_ms=90707 event=stream-sub-writer-poison live write failed [REQ-CONN-POISON-ATTRIBUTION]\n*104|BRAIN_SUBSCRIBER_STALL_EVICT:1 conn=104 wall_ms=1789523852688 mono_ms=105759 — controller conn write exceeded its bound; physical conn poisoned + retired (writer released, EOF cleanup follows) [REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]\n*105|CONTROLLER_SLOT_CLOSED:1 conn=104 cause=detach wall_ms=1789523852688 mono_ms=105759 — the controller slot was occupied and is now empty [REQ-REFRESH-WAIT-ATTRIBUTION]\n*106|CONN_LIFECYCLE: conn=104 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523852688 mono_ms=105759 first_write_wall_ms=1789523837637 first_write_mono_ms=90707 event=writer-exit role=controller session=1 reason=write-failed kind=TimedOut [REQ-CONN-POISON-ATTRIBUTION]\n 107|DISPATCH:1:Failed(\"No process is on the other end of the pipe. (os error 233)\")\n 108|DISPATCH_EV: gen=1789523805974 stream=1 family=Attach attempt=2 conn=1 event=worker-outcome detail=Failed(\"No process is on the other end of the pipe. (os error 233)\") elapsed_ms=15052 endpoint=-\n*109|CONN_LIFECYCLE: conn=104 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523852704 mono_ms=105774 first_write_wall_ms=1789523837637 first_write_mono_ms=90707 event=session-detach session=1 was_controller=true viewers_dropped=0 [REQ-CONN-POISON-ATTRIBUTION]\n*110|CONN_LIFECYCLE: conn=104 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523852704 mono_ms=105775 first_write_wall_ms=1789523837637 first_write_mono_ms=90707 event=transport-close poisoned=true timed_out=true [REQ-CONN-POISON-ATTRIBUTION]\n 111|CONN_LIFECYCLE_SUMMARY: scope=broker-ipc wall_ms=1789523869521 mono_ms=122592 interval_ms=61799 starts=40 closes=59 [REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n 112|CONN_LIFECYCLE_SUMMARY: scope=broker-ipc wall_ms=1789523930868 mono_ms=183938 interval_ms=61346 starts=35 closes=53 [REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n*113|BRAIN_UPDATE_RESTART: planned restart — cycling brain onto the swapped binary\n 114|BRAIN_RESTART: planned update respawn (generation 2)\n 115|=== spt brain stderr — generation 2 — pid 35888 ===\n 116|BRAIN_UP: connected to broker spt-daemon-broker-573ad82e70ef3afe.sock (pid 35888, generation 2, start-reason update)\n 117|BRAIN_PHASE:announce done in 2ms\n*118|BRAIN_RESUMED: re-established 1 session cursor(s): [1]\n 119|BRAIN_PHASE:resume done in 0ms\n 120|BRAIN_PHASE:ready written in 1922ms (exe-hash 1921ms, computed=true) path C:\\Users\\decid\\Documents\\projects\\spt-core\\.spt\\preserved\\307\\hertz-w1bc\\w1c-run1\\home-w1c\\brain.ready\n 121|BRAIN_CONN_OPEN: caller=livehost-session-query reason=initial-query policy=reuse-until-query-error wall_ms=1789523936064 mono_ms=1929\n...\n 135|CONN_LIFECYCLE: conn=201 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523951971 mono_ms=205042 first_write_wall_ms=1789523936922 first_write_mono_ms=189993 event=stream-sub-writer-poison live write failed [REQ-CONN-POISON-ATTRIBUTION]\n*136|BRAIN_SUBSCRIBER_STALL_EVICT:1 conn=201 wall_ms=1789523951971 mono_ms=205042 — controller conn write exceeded its bound; physical conn poisoned + retired (writer released, EOF cleanup follows) [REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]\n 137|CONN_LIFECYCLE: conn=201 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523951971 mono_ms=205042 first_write_wall_ms=1789523936922 first_write_mono_ms=189993 event=writer-exit role=controller session=1 reason=write-failed kind=TimedOut [REQ-CONN-POISON-ATTRIBUTION]\n*138|CONTROLLER_SLOT_CLOSED:1 conn=201 cause=detach wall_ms=1789523951972 mono_ms=205042 — the controller slot was occupied and is now empty [REQ-REFRESH-WAIT-ATTRIBUTION]\n 139|DISPATCH:1:Failed(\"No process is on the other end of the pipe. (os error 233)\")\n 140|DISPATCH_EV: gen=1789523936906 stream=1 family=Attach attempt=0 conn=1 event=worker-outcome detail=Failed(\"No process is on the other end of the pipe. (os error 233)\") elapsed_ms=15060 endpoint=-\n 141|CONN_LIFECYCLE: conn=201 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523951989 mono_ms=205060 first_write_wall_ms=1789523936922 first_write_mono_ms=189993 event=session-detach session=1 was_controller=true viewers_dropped=0 [REQ-CONN-POISON-ATTRIBUTION]"}],"details":{"scopePath":".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/logs/daemon.stderr.log","searchPath":"C:\\Users\\decid\\Documents\\projects\\spt-core\\.spt\\preserved\\307\\hertz-w1bc\\w1c-run1\\home-w1c\\logs\\daemon.stderr.log","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","matchCount":49,"fileCount":1,"files":[".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/logs/daemon.stderr.log"],"fileMatches":[{"path":".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/logs/daemon.stderr.log","count":49}],"truncated":false,"displayContent":"  34│SPAWN_PHASE:w1crig phase=row_inserted t_ms=1204 session=1\n *35│CONTROLLER_SLOT_CLOSED:1 conn=13 cause=detach wall_ms=1789523754018 mono_ms=7088 — the controller slot was occupied and is now empty [REQ-REFRESH-WAIT-ATTRIBUTION]\n  36│CONN_LIFECYCLE: conn=13 role=brain controller session=1 endpoint=w1crig by=local wall_ms=1789523754018 mono_ms=7088 first_write_wall_ms=1789523754018 first_write_mono_ms=7088 event=writer-exit role=controller session=1 reason=channel-closed [REQ-CONN-POISON-ATTRIBUTION]\n  37│CONN_LIFECYCLE: conn=13 role=brain controller session=1 endpoint=w1crig by=local wall_ms=1789523754032 mono_ms=7102 first_write_wall_ms=1789523754018 first_write_mono_ms=7088 event=session-detach session=1 was_controller=true viewers_dropped=0 [REQ-CONN-POISON-ATTRIBUTION]\n  38│CONN_LIFECYCLE: conn=15 role=brain stream-subscriber stream=2 wall_ms=1789523756051 mono_ms=9122 first_write_wall_ms=1789523756007 first_write_mono_ms=9077 event=stream-sub-attach stream=2 from_seq=0 [REQ-CONN-POISON-ATTRIBUTION]\n *39│CONN_LIFECYCLE: conn=17 role=brain stream-subscriber stream=1 wall_ms=1789523756142 mono_ms=9213 first_write_wall_ms=1789523756142 first_write_mono_ms=9213 event=stream-sub-attach stream=1 from_seq=0 [REQ-CONN-POISON-ATTRIBUTION]\n  40│ATTACH_INTENT_CHOSEN: site=serve_dispatch session=1 intent=control from_seq=0 gen=1789523756051 coded=false\n *41│CONN_LIFECYCLE: conn=17 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523756143 mono_ms=9214 first_write_wall_ms=1789523756142 first_write_mono_ms=9213 event=controller-attach session=1 from_seq=0 [REQ-CONN-POISON-ATTRIBUTION]\n  42│ATTACH_INTENT_CHOSEN: site=serve_dispatch session=1 intent=control from_seq=0 gen=1789523756051 coded=false\n *43│SUBSCRIBE_DECISION: session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf conn=17 intent=control old_by=- old_gen=0 req_gen=1789523756051 decision=controller\n *44│SUBSCRIBE_DECISION: session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf conn=17 intent=control old_by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf old_gen=1789523756051 req_gen=1789523756051 decision=idempotent\n *45│BRAIN_UPDATE_RESTART: planned restart — cycling brain onto the swapped binary\n *46│CONTROLLER_SLOT_CLOSED:1 conn=17 cause=detach wall_ms=1789523803471 mono_ms=56541 — the controller slot was occupied and is now empty [REQ-REFRESH-WAIT-ATTRIBUTION]\n *47│CONN_WRITE_RETIRED: conn=17 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523803471 mono_ms=56541 first_write_wall_ms=1789523756142 first_write_mono_ms=9213 — organic write failure (BrokenPipe: The pipe is being closed. (os error 232)); physical broker conn poisoned + retired, never reused (same retire behavior as POISONED — only the failure class differs) [REQ-CONN-POISON-DIAL-SCOPE]\n *48│CONN_LIFECYCLE: conn=17 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523803471 mono_ms=56541 first_write_wall_ms=1789523756142 first_write_mono_ms=9213 event=writer-exit role=controller session=1 reason=write-failed kind=BrokenPipe [REQ-CONN-POISON-ATTRIBUTION]\n  49│BRAIN_RESTART: planned update respawn (generation 1)\n  50│=== spt brain stderr — generation 1 — pid 53752 ===\n  51│BRAIN_UP: connected to broker spt-daemon-broker-573ad82e70ef3afe.sock (pid 53752, generation 1, start-reason update)\n  52│BRAIN_PHASE:announce done in 0ms\n *53│CONN_LIFECYCLE: conn=17 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523803573 mono_ms=56643 first_write_wall_ms=1789523756142 first_write_mono_ms=9213 event=session-detach session=1 was_controller=true viewers_dropped=0 [REQ-CONN-POISON-ATTRIBUTION]\n *54│CONN_LIFECYCLE: conn=17 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523803573 mono_ms=56643 first_write_wall_ms=1789523756142 first_write_mono_ms=9213 event=transport-close poisoned=true timed_out=false [REQ-CONN-POISON-ATTRIBUTION]\n *55│BRAIN_RESUMED: re-established 1 session cursor(s): [1]\n  56│BRAIN_PHASE:resume done in 69ms\n  57│BRAIN_PHASE:ready written in 1894ms (exe-hash 1893ms, computed=true) path C:\\Users\\decid\\Documents\\projects\\spt-core\\.spt\\preserved\\307\\hertz-w1bc\\w1c-run1\\home-w1c\\brain.ready\n  58│BRAIN_CONN_OPEN: caller=livehost-session-query reason=initial-query policy=reuse-until-query-error wall_ms=1789523805474 mono_ms=1966\n    │...\n  64│CONN_LIFECYCLE: conn=68 role=brain presence-subscriber wall_ms=1789523805974 mono_ms=59045 first_write=none event=presence-sub-attach from_seq=18446744073709551615 [REQ-CONN-POISON-ATTRIBUTION]\n *65│CONN_LIFECYCLE: conn=70 role=brain stream-subscriber stream=1 wall_ms=1789523805975 mono_ms=59046 first_write_wall_ms=1789523805975 first_write_mono_ms=59046 event=stream-sub-attach stream=1 from_seq=0 [REQ-CONN-POISON-ATTRIBUTION]\n  66│ATTACH_INTENT_CHOSEN: site=serve_dispatch session=1 intent=control from_seq=0 gen=1789523756051 coded=false\n *67│CONN_LIFECYCLE: conn=70 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523805976 mono_ms=59047 first_write_wall_ms=1789523805975 first_write_mono_ms=59046 event=controller-attach session=1 from_seq=0 [REQ-CONN-POISON-ATTRIBUTION]\n  68│ATTACH_INTENT_CHOSEN: site=serve_dispatch session=1 intent=control from_seq=0 gen=1789523756051 coded=false\n *69│SUBSCRIBE_DECISION: session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf conn=70 intent=control old_by=- old_gen=0 req_gen=1789523756051 decision=controller\n *70│SUBSCRIBE_DECISION: session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf conn=70 intent=control old_by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf old_gen=1789523756051 req_gen=1789523756051 decision=idempotent\n  71│CONN_LIFECYCLE_SUMMARY: scope=broker-ipc wall_ms=1789523807721 mono_ms=60792 interval_ms=60674 starts=47 closes=63 [REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n *72│CONN_WRITE_POISONED: conn=70 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523821012 mono_ms=74082 first_write_wall_ms=1789523805975 first_write_mono_ms=59046 — physical broker conn retired: a write exceeded its 15000ms bound; in-flight I/O aborted, conn never reused [REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\n *73│CONN_LIFECYCLE: conn=70 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523821012 mono_ms=74082 first_write_wall_ms=1789523805975 first_write_mono_ms=59046 event=stream-sub-writer-poison live write failed [REQ-CONN-POISON-ATTRIBUTION]\n *74│BRAIN_SUBSCRIBER_STALL_EVICT:1 conn=70 wall_ms=1789523821012 mono_ms=74082 — controller conn write exceeded its bound; physical conn poisoned + retired (writer released, EOF cleanup follows) [REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]\n *75│CONTROLLER_SLOT_CLOSED:1 conn=70 cause=detach wall_ms=1789523821012 mono_ms=74083 — the controller slot was occupied and is now empty [REQ-REFRESH-WAIT-ATTRIBUTION]\n *76│CONN_LIFECYCLE: conn=70 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523821012 mono_ms=74083 first_write_wall_ms=1789523805975 first_write_mono_ms=59046 event=writer-exit role=controller session=1 reason=write-failed kind=TimedOut [REQ-CONN-POISON-ATTRIBUTION]\n  77│DISPATCH:1:Failed(\"No process is on the other end of the pipe. (os error 233)\")\n  78│DISPATCH_EV: gen=1789523805974 stream=1 family=Attach attempt=0 conn=1 event=worker-outcome detail=Failed(\"No process is on the other end of the pipe. (os error 233)\") elapsed_ms=15037 endpoint=-\n *79│CONN_LIFECYCLE: conn=70 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523821028 mono_ms=74099 first_write_wall_ms=1789523805975 first_write_mono_ms=59046 event=session-detach session=1 was_controller=true viewers_dropped=0 [REQ-CONN-POISON-ATTRIBUTION]\n *80│CONN_LIFECYCLE: conn=70 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523821028 mono_ms=74099 first_write_wall_ms=1789523805975 first_write_mono_ms=59046 event=transport-close poisoned=true timed_out=true [REQ-CONN-POISON-ATTRIBUTION]\n *81│CONN_LIFECYCLE: conn=88 role=brain stream-subscriber stream=1 wall_ms=1789523821556 mono_ms=74627 first_write_wall_ms=1789523821556 first_write_mono_ms=74627 event=stream-sub-attach stream=1 from_seq=0 [REQ-CONN-POISON-ATTRIBUTION]\n  82│ATTACH_INTENT_CHOSEN: site=serve_dispatch session=1 intent=control from_seq=0 gen=1789523756051 coded=false\n *83│CONN_LIFECYCLE: conn=88 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523821557 mono_ms=74628 first_write_wall_ms=1789523821556 first_write_mono_ms=74627 event=controller-attach session=1 from_seq=0 [REQ-CONN-POISON-ATTRIBUTION]\n  84│ATTACH_INTENT_CHOSEN: site=serve_dispatch session=1 intent=control from_seq=0 gen=1789523756051 coded=false\n *85│SUBSCRIBE_DECISION: session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf conn=88 intent=control old_by=- old_gen=0 req_gen=1789523756051 decision=controller\n *86│SUBSCRIBE_DECISION: session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf conn=88 intent=control old_by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf old_gen=1789523756051 req_gen=1789523756051 decision=idempotent\n *87│CONN_WRITE_POISONED: conn=88 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523836597 mono_ms=89667 first_write_wall_ms=1789523821556 first_write_mono_ms=74627 — physical broker conn retired: a write exceeded its 15000ms bound; in-flight I/O aborted, conn never reused [REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\n *88│CONN_LIFECYCLE: conn=88 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523836597 mono_ms=89668 first_write_wall_ms=1789523821556 first_write_mono_ms=74627 event=stream-sub-writer-poison live write failed [REQ-CONN-POISON-ATTRIBUTION]\n *89│BRAIN_SUBSCRIBER_STALL_EVICT:1 conn=88 wall_ms=1789523836597 mono_ms=89668 — controller conn write exceeded its bound; physical conn poisoned + retired (writer released, EOF cleanup follows) [REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]\n *90│CONTROLLER_SLOT_CLOSED:1 conn=88 cause=detach wall_ms=1789523836597 mono_ms=89668 — the controller slot was occupied and is now empty [REQ-REFRESH-WAIT-ATTRIBUTION]\n *91│CONN_LIFECYCLE: conn=88 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523836597 mono_ms=89668 first_write_wall_ms=1789523821556 first_write_mono_ms=74627 event=writer-exit role=controller session=1 reason=write-failed kind=TimedOut [REQ-CONN-POISON-ATTRIBUTION]\n  92│DISPATCH:1:Failed(\"No process is on the other end of the pipe. (os error 233)\")\n  93│DISPATCH_EV: gen=1789523805974 stream=1 family=Attach attempt=1 conn=1 event=worker-outcome detail=Failed(\"No process is on the other end of the pipe. (os error 233)\") elapsed_ms=15041 endpoint=-\n *94│CONN_LIFECYCLE: conn=88 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523836612 mono_ms=89683 first_write_wall_ms=1789523821556 first_write_mono_ms=74627 event=session-detach session=1 was_controller=true viewers_dropped=0 [REQ-CONN-POISON-ATTRIBUTION]\n *95│CONN_LIFECYCLE: conn=88 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523836612 mono_ms=89683 first_write_wall_ms=1789523821556 first_write_mono_ms=74627 event=transport-close poisoned=true timed_out=true [REQ-CONN-POISON-ATTRIBUTION]\n *96│CONN_LIFECYCLE: conn=104 role=brain stream-subscriber stream=1 wall_ms=1789523837637 mono_ms=90707 first_write_wall_ms=1789523837637 first_write_mono_ms=90707 event=stream-sub-attach stream=1 from_seq=0 [REQ-CONN-POISON-ATTRIBUTION]\n  97│ATTACH_INTENT_CHOSEN: site=serve_dispatch session=1 intent=control from_seq=0 gen=1789523756051 coded=false\n *98│CONN_LIFECYCLE: conn=104 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523837637 mono_ms=90708 first_write_wall_ms=1789523837637 first_write_mono_ms=90707 event=controller-attach session=1 from_seq=0 [REQ-CONN-POISON-ATTRIBUTION]\n  99│ATTACH_INTENT_CHOSEN: site=serve_dispatch session=1 intent=control from_seq=0 gen=1789523756051 coded=false\n*100│SUBSCRIBE_DECISION: session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf conn=104 intent=control old_by=- old_gen=0 req_gen=1789523756051 decision=controller\n*101│SUBSCRIBE_DECISION: session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf conn=104 intent=control old_by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf old_gen=1789523756051 req_gen=1789523756051 decision=idempotent\n*102│CONN_WRITE_POISONED: conn=104 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523852688 mono_ms=105759 first_write_wall_ms=1789523837637 first_write_mono_ms=90707 — physical broker conn retired: a write exceeded its 15000ms bound; in-flight I/O aborted, conn never reused [REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\n*103│CONN_LIFECYCLE: conn=104 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523852688 mono_ms=105759 first_write_wall_ms=1789523837637 first_write_mono_ms=90707 event=stream-sub-writer-poison live write failed [REQ-CONN-POISON-ATTRIBUTION]\n*104│BRAIN_SUBSCRIBER_STALL_EVICT:1 conn=104 wall_ms=1789523852688 mono_ms=105759 — controller conn write exceeded its bound; physical conn poisoned + retired (writer released, EOF cleanup follows) [REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]\n*105│CONTROLLER_SLOT_CLOSED:1 conn=104 cause=detach wall_ms=1789523852688 mono_ms=105759 — the controller slot was occupied and is now empty [REQ-REFRESH-WAIT-ATTRIBUTION]\n*106│CONN_LIFECYCLE: conn=104 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523852688 mono_ms=105759 first_write_wall_ms=1789523837637 first_write_mono_ms=90707 event=writer-exit role=controller session=1 reason=write-failed kind=TimedOut [REQ-CONN-POISON-ATTRIBUTION]\n 107│DISPATCH:1:Failed(\"No process is on the other end of the pipe. (os error 233)\")\n 108│DISPATCH_EV: gen=1789523805974 stream=1 family=Attach attempt=2 conn=1 event=worker-outcome detail=Failed(\"No process is on the other end of the pipe. (os error 233)\") elapsed_ms=15052 endpoint=-\n*109│CONN_LIFECYCLE: conn=104 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523852704 mono_ms=105774 first_write_wall_ms=1789523837637 first_write_mono_ms=90707 event=session-detach session=1 was_controller=true viewers_dropped=0 [REQ-CONN-POISON-ATTRIBUTION]\n*110│CONN_LIFECYCLE: conn=104 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523852704 mono_ms=105775 first_write_wall_ms=1789523837637 first_write_mono_ms=90707 event=transport-close poisoned=true timed_out=true [REQ-CONN-POISON-ATTRIBUTION]\n 111│CONN_LIFECYCLE_SUMMARY: scope=broker-ipc wall_ms=1789523869521 mono_ms=122592 interval_ms=61799 starts=40 closes=59 [REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n 112│CONN_LIFECYCLE_SUMMARY: scope=broker-ipc wall_ms=1789523930868 mono_ms=183938 interval_ms=61346 starts=35 closes=53 [REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n*113│BRAIN_UPDATE_RESTART: planned restart — cycling brain onto the swapped binary\n 114│BRAIN_RESTART: planned update respawn (generation 2)\n 115│=== spt brain stderr — generation 2 — pid 35888 ===\n 116│BRAIN_UP: connected to broker spt-daemon-broker-573ad82e70ef3afe.sock (pid 35888, generation 2, start-reason update)\n 117│BRAIN_PHASE:announce done in 2ms\n*118│BRAIN_RESUMED: re-established 1 session cursor(s): [1]\n 119│BRAIN_PHASE:resume done in 0ms\n 120│BRAIN_PHASE:ready written in 1922ms (exe-hash 1921ms, computed=true) path C:\\Users\\decid\\Documents\\projects\\spt-core\\.spt\\preserved\\307\\hertz-w1bc\\w1c-run1\\home-w1c\\brain.ready\n 121│BRAIN_CONN_OPEN: caller=livehost-session-query reason=initial-query policy=reuse-until-query-error wall_ms=1789523936064 mono_ms=1929\n    │...\n 135│CONN_LIFECYCLE: conn=201 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523951971 mono_ms=205042 first_write_wall_ms=1789523936922 first_write_mono_ms=189993 event=stream-sub-writer-poison live write failed [REQ-CONN-POISON-ATTRIBUTION]\n*136│BRAIN_SUBSCRIBER_STALL_EVICT:1 conn=201 wall_ms=1789523951971 mono_ms=205042 — controller conn write exceeded its bound; physical conn poisoned + retired (writer released, EOF cleanup follows) [REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]\n 137│CONN_LIFECYCLE: conn=201 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523951971 mono_ms=205042 first_write_wall_ms=1789523936922 first_write_mono_ms=189993 event=writer-exit role=controller session=1 reason=write-failed kind=TimedOut [REQ-CONN-POISON-ATTRIBUTION]\n*138│CONTROLLER_SLOT_CLOSED:1 conn=201 cause=detach wall_ms=1789523951972 mono_ms=205042 — the controller slot was occupied and is now empty [REQ-REFRESH-WAIT-ATTRIBUTION]\n 139│DISPATCH:1:Failed(\"No process is on the other end of the pipe. (os error 233)\")\n 140│DISPATCH_EV: gen=1789523936906 stream=1 family=Attach attempt=0 conn=1 event=worker-outcome detail=Failed(\"No process is on the other end of the pipe. (os error 233)\") elapsed_ms=15060 endpoint=-\n 141│CONN_LIFECYCLE: conn=201 role=brain stream-subscriber stream=1 controller session=1 endpoint=w1crig by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf wall_ms=1789523951989 mono_ms=205060 first_write_wall_ms=1789523936922 first_write_mono_ms=189993 event=session-detach session=1 was_controller=true viewers_dropped=0 [REQ-CONN-POISON-ATTRIBUTION]"},"isError":false,"timestamp":1789523963485}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_8rmEcmeK5yQeCQi7HhoLF7Qz|fc_0d104aba2b0f3edc016aa9f802a32c87d1a5427b614ef970a2","toolName":"read","startedAt":"2026-09-16T01:59:30.717Z","args":{"path":"crates/spt-daemon/src/broker.rs:3955-4210"},"intent":"Reading controller writer implementation"},"id":"4bef5ed2","parentId":"df262790","timestamp":"2026-09-16T01:59:30.717Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_OK8fpkIcMHacdmFLwcEgq71S|fc_0d104aba2b0f3edc016aa9f802a33c87d19d81f4780e6bf1e9","toolName":"read","startedAt":"2026-09-16T01:59:30.718Z","args":{"path":"crates/spt-daemon/src/broker.rs:3540-3595,5982-6032"},"intent":"Reading detach and predicate implementation"},"id":"ca3f9c25","parentId":"4bef5ed2","timestamp":"2026-09-16T01:59:30.718Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_ka3JN6XQS5HfwPTgWJC1lcOL|fc_0d104aba2b0f3edc016aa9f802a34487d1886741df249f405a","toolName":"read","startedAt":"2026-09-16T01:59:30.718Z","args":{"path":"."},"intent":"Locating documentation roots"},"id":"a25d925e","parentId":"ca3f9c25","timestamp":"2026-09-16T01:59:30.718Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_A00zgDAoxDKjZdH0keeKsM1m|fc_0d104aba2b0f3edc016aa9f802a34c87d19dde6b6769261bd1","toolName":"grep","startedAt":"2026-09-16T01:59:30.718Z","args":{"path":"crates/spt-daemon/src/conn.rs"},"intent":"Finding conn retirement internals"},"id":"e0d4cb71","parentId":"a25d925e","timestamp":"2026-09-16T01:59:30.718Z"}
{"type":"message","id":"7025ed86","parentId":"e0d4cb71","timestamp":"2026-09-16T01:59:30.718Z","message":{"role":"assistant","content":[{"type":"toolCall","id":"call_8rmEcmeK5yQeCQi7HhoLF7Qz|fc_0d104aba2b0f3edc016aa9f802a32c87d1a5427b614ef970a2","name":"read","arguments":{"i":"Reading controller writer implementation","path":"crates/spt-daemon/src/broker.rs:3955-4210"},"intent":"Reading controller writer implementation"},{"type":"toolCall","id":"call_OK8fpkIcMHacdmFLwcEgq71S|fc_0d104aba2b0f3edc016aa9f802a33c87d19d81f4780e6bf1e9","name":"read","arguments":{"i":"Reading detach and predicate implementation","path":"crates/spt-daemon/src/broker.rs:3540-3595,5982-6032"},"intent":"Reading detach and predicate implementation"},{"type":"toolCall","id":"call_ka3JN6XQS5HfwPTgWJC1lcOL|fc_0d104aba2b0f3edc016aa9f802a34487d1886741df249f405a","name":"read","arguments":{"i":"Locating documentation roots","path":"."},"intent":"Locating documentation roots"},{"type":"toolCall","id":"call_A00zgDAoxDKjZdH0keeKsM1m|fc_0d104aba2b0f3edc016aa9f802a34c87d19dde6b6769261bd1","name":"grep","arguments":{"i":"Finding conn retirement internals","pattern":"CONN_WRITE_POISONED|fn write|fn poison|retire|lock\\(","path":"crates/spt-daemon/src/conn.rs","case":true,"gitignore":false,"skip":null},"intent":"Finding conn retirement internals"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":11122,"output":182,"cacheRead":29568,"cacheWrite":0,"totalTokens":40872,"cost":{"input":0.11122000000000001,"output":0.0091,"cacheRead":0.029567999999999997,"cacheWrite":0,"total":0.14988800000000002}},"stopReason":"toolUse","timestamp":1789523963515,"responseId":"resp_0d104aba2b0f3edc016aa9f7fbc10087d1a293f42066bace25","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"fc_0d104aba2b0f3edc016aa9f802a32c87d1a5427b614ef970a2","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading controller writer implementation\",\"path\":\"crates/spt-daemon/src/broker.rs:3955-4210\"}","call_id":"call_8rmEcmeK5yQeCQi7HhoLF7Qz","internal_chat_message_metadata_passthrough":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"},"name":"read","metadata":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"}},{"id":"fc_0d104aba2b0f3edc016aa9f802a33c87d19d81f4780e6bf1e9","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading detach and predicate implementation\",\"path\":\"crates/spt-daemon/src/broker.rs:3540-3595,5982-6032\"}","call_id":"call_OK8fpkIcMHacdmFLwcEgq71S","internal_chat_message_metadata_passthrough":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"},"name":"read","metadata":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"}},{"id":"fc_0d104aba2b0f3edc016aa9f802a34487d1886741df249f405a","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating documentation roots\",\"path\":\".\"}","call_id":"call_ka3JN6XQS5HfwPTgWJC1lcOL","internal_chat_message_metadata_passthrough":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"},"name":"read","metadata":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"}},{"id":"fc_0d104aba2b0f3edc016aa9f802a34c87d19dde6b6769261bd1","type":"function_call","status":"completed","arguments":"{\"i\":\"Finding conn retirement internals\",\"pattern\":\"CONN_WRITE_POISONED|fn write|fn poison|retire|lock\\\\(\",\"path\":\"crates/spt-daemon/src/conn.rs\",\"case\":true,\"gitignore\":false,\"skip\":null}","call_id":"call_A00zgDAoxDKjZdH0keeKsM1m","internal_chat_message_metadata_passthrough":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"},"name":"grep","metadata":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"}}]},"duration":7201.216200000141,"ttft":7157.04420000012,"completedAt":1789523970716,"contextSnapshot":{"promptTokens":40690,"nonMessageTokens":10480,"compactionEpoch":0}}}
{"type":"message","id":"cef16aa5","parentId":"7025ed86","timestamp":"2026-09-16T01:59:30.723Z","message":{"role":"toolResult","toolCallId":"call_A00zgDAoxDKjZdH0keeKsM1m|fc_0d104aba2b0f3edc016aa9f802a34c87d19dde6b6769261bd1","toolName":"grep","content":[{"type":"text","text":" 37|//!\n*38|//! ## Poison / retire invariant (both OSes)\n 39|//!\n 40|//! On deadline, write failure, or cancellation: (1) the WHOLE physical\n 41|//! connection is poisoned, idempotently — a length-prefixed frame may have\n...\n 45|//! completion (`write_frame` returns only after the completion APC — the\n*46|//! synchronous return IS the completion handshake); (4) the retired half is\n 47|//! parked in a graveyard slot — never dropped while the watchdog might still\n 48|//! touch the raw handle (drop = `CloseHandle`, and a recycled handle value\n 49|//! must never see a late `CancelIoEx`), and never re-gated; (5) on Windows the\n...\n 58|//! (recorded mid-frame corruption risk). Unix keeps its existing semantics for\n*59|//! healthy connections; only the poison/retire path is new.\n 60|// [impl->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\n 61|\n 62|use std::io;\n...\n 74|/// [`BrokerConn::new`]. Before this the only conn identity was `Arc::ptr_eq`,\n*75|/// which does not survive a log line: per-line CONN_WRITE_POISONED churn was\n 76|/// UNDECIDABLE between fresh-carrier churn and stderr interleave (the F-039\n 77|/// RCA's terminal fork). Every lifecycle record carries `conn=<id>`.\n 78|// [impl->REQ-CONN-POISON-ATTRIBUTION]\n...\n 220|/// The serialized write gate. `half: Some` = free; `None` = a write is in\n*221|/// flight (checked out) or the conn has retired (then `poisoned` is set).\n 222|struct Gate {\n 223|    half: Option<SendHalf>,\n 224|}\n...\n 251|    poisoned: AtomicBool,\n*252|    /// Latched when the retirement was DEADLINE-driven (watchdog fired, or a\n 253|    /// gate-waiter exhausted its bound) — as opposed to an organic I/O error\n 254|    /// (peer death). Callers see it as `io::ErrorKind::TimedOut` and can\n 255|    /// count the wedge (the W2 stall-evict tally) without conflating it with\n...\n 259|    severed: AtomicBool,\n*260|    /// A retired (poisoned) half parks here so its OS handle stays alive\n 261|    /// until [`BrokerConn::drop`] — after the watchdog is joined, so no raw\n 262|    /// `CancelIoEx`/`DisconnectNamedPipe` can race the `CloseHandle`.\n 263|    graveyard: Mutex<Option<SendHalf>>,\n...\n 282|fn recover<T>(m: &Mutex<T>) -> MutexGuard<'_, T> {\n*283|    match m.lock() {\n 284|        Ok(g) => g,\n 285|        Err(p) => p.into_inner(),\n 286|    }\n...\n 330|\n*331|    /// Render the once-per-conn retirement record — the F-039 token split\n*332|    /// (leg a): `CONN_WRITE_POISONED` is RESERVED for the deadline class\n 333|    /// (`timed_out` latched — the wedge observable blackhole-controller.ps1\n 334|    /// watches); an ORGANIC write failure (peer death / broken pipe on an\n 335|    /// already-dead counterpart, the offline-peer dial churn) emits the\n 336|    /// distinct `CONN_WRITE_RETIRED` token carrying the error kind and the\n*337|    /// SAME attribution fields — quieter class, never thinner (the retired\n 338|    /// class is what discriminates fresh-carrier churn from interleave\n 339|    /// artifact in the field). Behavior is identical either way: poison,\n*340|    /// abort, retire, never reuse.\n 341|    // [impl->REQ-CONN-POISON-DIAL-SCOPE]\n*342|    fn render_retirement(&self, timed_out: bool, cause: Option<&io::Error>) -> String {\n 343|        if timed_out {\n 344|            format!(\n*345|                \"CONN_WRITE_POISONED: {} — physical broker conn retired: a write \\\n 346|                 exceeded its {}ms bound; in-flight I/O aborted, conn never reused \\\n 347|                 [REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\",\n 348|                self.attribution(),\n...\n 357|                \"CONN_WRITE_RETIRED: {} — organic write failure ({cause}); physical \\\n*358|                 broker conn poisoned + retired, never reused (same retire behavior \\\n 359|                 as POISONED — only the failure class differs) \\\n 360|                 [REQ-CONN-POISON-DIAL-SCOPE]\",\n 361|                self.attribution()\n...\n 383|    /// AFTER a cancel (cancellation only reaches ALREADY-issued ops).\n*384|    /// `cause` (the organic caller's write error) rides the retirement record\n*385|    /// when the retirement is NOT deadline-driven.\n*386|    fn poison_and_cancel(&self, cause: Option<&io::Error>) {\n 387|        if !self.poisoned.swap(true, Ordering::AcqRel) {\n 388|            // First poison of this conn: loud, once. The token is class-split\n*389|            // (F-039 leg a): deadline → CONN_WRITE_POISONED (the wedge\n 390|            // observable), organic fast-fail → CONN_WRITE_RETIRED.\n 391|            self.lifecycle_log.emit(format_args!(\n 392|                \"{}\",\n*393|                self.render_retirement(self.timed_out.load(Ordering::Acquire), cause)\n 394|            ));\n 395|        }\n 396|        #[cfg(windows)]\n...\n 486|            // one shot is not enough). Latch timed_out FIRST so the aborted\n*487|            // writer observes its failure as a DEADLINE retirement (wedge),\n 488|            // not an organic conn death.\n 489|            drop(d);\n 490|            self.timed_out.store(true, Ordering::Release);\n...\n 598|    /// Bounded framed write. See [`Self::write_gated`].\n*599|    pub(crate) fn write(&self, env: &Envelope) -> io::Result<ConnWrite> {\n 600|        self.write_gated(env, || true)\n 601|    }\n 602|\n...\n 604|    /// serialized gate is acquired and BEFORE any byte is written — the same\n*605|    /// serialization point the old epoch-check-under-`send.lock()` had, so a\n 606|    /// superseded controller writer still bails atomically with respect to\n 607|    /// its successor's writes (P1c, REQ-HAZARD-CONTROLLER-WRITER-REORDER).\n 608|    /// `pre` returning false → `Ok(ConnWrite::Superseded)`, nothing written.\n*609|    pub(crate) fn write_gated(\n 610|        &self,\n 611|        env: &Envelope,\n 612|        pre: impl FnOnce() -> bool,\n...\n 696|        inner.gate_cv.notify_all();\n*697|        // A DEADLINE retirement (the watchdog aborted OUR op, or the conn was\n 698|        // already deadline-poisoned) surfaces as TimedOut — the caller can\n 699|        // count the wedge (W2 stall-evict tally). An organic I/O failure\n 700|        // (peer death) keeps its own error.\n...\n 774|\n*775|fn poisoned_err() -> io::Error {\n 776|    io::Error::new(\n 777|        io::ErrorKind::BrokenPipe,\n*778|        \"broker conn poisoned (bounded write retired the physical connection)\",\n 779|    )\n 780|}\n 781|\n...\n 875|        conn.inner.poison_and_cancel(Some(&io::Error::new(io::ErrorKind::BrokenPipe, \"peer gone\")));\n*876|        // Repeated cancellation must not repeat the retirement record.\n 877|        conn.inner.poison_and_cancel(None);\n 878|        let id = conn.id();\n 879|        drop(conn);\n 880|        let lines = recover(log.capture.as_ref().unwrap());\n*881|        assert_eq!(lines.len(), 3, \"role event, retirement and poisoned close survive\");\n 882|        for line in lines.iter() {\n 883|            assert!(line.contains(&format!(\"conn={id}\")));\n 884|            assert!(line.contains(\"role=brain controller session=7\"));\n...\n 893|    // [unit->REQ-CONN-POISON-DIAL-SCOPE] the F-039 leg-(a) token split: the loud\n*894|    // CONN_WRITE_POISONED token is RESERVED for the deadline (timed_out) class;\n 895|    // an organic write failure renders the distinct CONN_WRITE_RETIRED token\n 896|    // carrying the error kind — and (doyle amendment 1) the SAME attribution\n 897|    // fields as POISONED: conn id, accreted role/endpoint/session facts, wall +\n...\n 900|    #[test]\n*901|    fn retirement_token_splits_deadline_from_organic_with_attribution_parity() {\n 902|        let (conn, _client) = make_conn();\n 903|        conn.describe(\"role=brain\");\n 904|        conn.describe(\"controller session=7 endpoint=ep by=cli@host\");\n 905|\n*906|        let poisoned = conn.inner.render_retirement(true, None);\n 907|        assert!(\n*908|            poisoned.starts_with(\"CONN_WRITE_POISONED:\"),\n 909|            \"deadline class keeps the loud wedge-observable token: {poisoned}\"\n 910|        );\n*911|        let retired = conn\n 912|            .inner\n*913|            .render_retirement(false, Some(&io::Error::new(io::ErrorKind::BrokenPipe, \"peer gone\")));\n 914|        assert!(\n*915|            retired.starts_with(\"CONN_WRITE_RETIRED:\"),\n*916|            \"organic class emits the DISTINCT retired token: {retired}\"\n 917|        );\n 918|        assert!(\n*919|            retired.contains(\"BrokenPipe\") && retired.contains(\"peer gone\"),\n*920|            \"the organic record carries the error kind + message: {retired}\"\n 921|        );\n 922|        // Attribution parity: every field present on BOTH tokens.\n*923|        for (name, line) in [(\"POISONED\", &poisoned), (\"RETIRED\", &retired)] {\n 924|            assert!(\n 925|                line.contains(&format!(\"conn={}\", conn.id())),\n 926|                \"{name} carries the stable conn id: {line}\"\n...\n 940|    // split (the REQ-HAZARD-SHAREDSEND invariant): an organic fast-fail still\n*941|    // poisons + retires the conn — it is only the CLASS that differs (the error\n 942|    // surfaces with its own kind, never TimedOut, and the next write refuses).\n 943|    #[test]\n 944|    fn organic_write_failure_still_poisons_and_is_not_the_timeout_class() {\n...\n 965|        );\n*966|        // Poisoned + retired: the next write refuses immediately.\n 967|        let refused = conn.write(&env).expect_err(\"poisoned conn refuses writes\");\n 968|        assert_eq!(refused.kind(), io::ErrorKind::BrokenPipe);\n 969|    }\n...\n 994|        // The id (and the capped label) stay stable across records.\n*995|        let one = a.inner.render_retirement(true, None);\n*996|        let two = a.inner.render_retirement(true, None);\n 997|        assert!(one.contains(&format!(\"conn={}\", a.id())));\n 998|        assert!(two.contains(&format!(\"conn={}\", a.id())));\n 999|    }"}],"details":{"scopePath":"crates/spt-daemon/src/conn.rs","searchPath":"C:\\Users\\decid\\Documents\\projects\\spt-core\\crates\\spt-daemon\\src\\conn.rs","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","matchCount":44,"fileCount":1,"files":["crates/spt-daemon/src/conn.rs"],"fileMatches":[{"path":"crates/spt-daemon/src/conn.rs","count":44}],"truncated":false,"displayContent":"  37│//!\n *38│//! ## Poison / retire invariant (both OSes)\n  39│//!\n  40│//! On deadline, write failure, or cancellation: (1) the WHOLE physical\n  41│//! connection is poisoned, idempotently — a length-prefixed frame may have\n    │...\n  45│//! completion (`write_frame` returns only after the completion APC — the\n *46│//! synchronous return IS the completion handshake); (4) the retired half is\n  47│//! parked in a graveyard slot — never dropped while the watchdog might still\n  48│//! touch the raw handle (drop = `CloseHandle`, and a recycled handle value\n  49│//! must never see a late `CancelIoEx`), and never re-gated; (5) on Windows the\n    │...\n  58│//! (recorded mid-frame corruption risk). Unix keeps its existing semantics for\n *59│//! healthy connections; only the poison/retire path is new.\n  60│// [impl->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\n  61│\n  62│use std::io;\n    │...\n  74│/// [`BrokerConn::new`]. Before this the only conn identity was `Arc::ptr_eq`,\n *75│/// which does not survive a log line: per-line CONN_WRITE_POISONED churn was\n  76│/// UNDECIDABLE between fresh-carrier churn and stderr interleave (the F-039\n  77│/// RCA's terminal fork). Every lifecycle record carries `conn=<id>`.\n  78│// [impl->REQ-CONN-POISON-ATTRIBUTION]\n    │...\n 220│/// The serialized write gate. `half: Some` = free; `None` = a write is in\n*221│/// flight (checked out) or the conn has retired (then `poisoned` is set).\n 222│struct Gate {\n 223│    half: Option<SendHalf>,\n 224│}\n    │...\n 251│    poisoned: AtomicBool,\n*252│    /// Latched when the retirement was DEADLINE-driven (watchdog fired, or a\n 253│    /// gate-waiter exhausted its bound) — as opposed to an organic I/O error\n 254│    /// (peer death). Callers see it as `io::ErrorKind::TimedOut` and can\n 255│    /// count the wedge (the W2 stall-evict tally) without conflating it with\n    │...\n 259│    severed: AtomicBool,\n*260│    /// A retired (poisoned) half parks here so its OS handle stays alive\n 261│    /// until [`BrokerConn::drop`] — after the watchdog is joined, so no raw\n 262│    /// `CancelIoEx`/`DisconnectNamedPipe` can race the `CloseHandle`.\n 263│    graveyard: Mutex<Option<SendHalf>>,\n    │...\n 282│fn recover<T>(m: &Mutex<T>) -> MutexGuard<'_, T> {\n*283│    match m.lock() {\n 284│        Ok(g) => g,\n 285│        Err(p) => p.into_inner(),\n 286│    }\n    │...\n 330│\n*331│    /// Render the once-per-conn retirement record — the F-039 token split\n*332│    /// (leg a): `CONN_WRITE_POISONED` is RESERVED for the deadline class\n 333│    /// (`timed_out` latched — the wedge observable blackhole-controller.ps1\n 334│    /// watches); an ORGANIC write failure (peer death / broken pipe on an\n 335│    /// already-dead counterpart, the offline-peer dial churn) emits the\n 336│    /// distinct `CONN_WRITE_RETIRED` token carrying the error kind and the\n*337│    /// SAME attribution fields — quieter class, never thinner (the retired\n 338│    /// class is what discriminates fresh-carrier churn from interleave\n 339│    /// artifact in the field). Behavior is identical either way: poison,\n*340│    /// abort, retire, never reuse.\n 341│    // [impl->REQ-CONN-POISON-DIAL-SCOPE]\n*342│    fn render_retirement(&self, timed_out: bool, cause: Option<&io::Error>) -> String {\n 343│        if timed_out {\n 344│            format!(\n*345│                \"CONN_WRITE_POISONED: {} — physical broker conn retired: a write \\\n 346│                 exceeded its {}ms bound; in-flight I/O aborted, conn never reused \\\n 347│                 [REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\",\n 348│                self.attribution(),\n    │...\n 357│                \"CONN_WRITE_RETIRED: {} — organic write failure ({cause}); physical \\\n*358│                 broker conn poisoned + retired, never reused (same retire behavior \\\n 359│                 as POISONED — only the failure class differs) \\\n 360│                 [REQ-CONN-POISON-DIAL-SCOPE]\",\n 361│                self.attribution()\n    │...\n 383│    /// AFTER a cancel (cancellation only reaches ALREADY-issued ops).\n*384│    /// `cause` (the organic caller's write error) rides the retirement record\n*385│    /// when the retirement is NOT deadline-driven.\n*386│    fn poison_and_cancel(&self, cause: Option<&io::Error>) {\n 387│        if !self.poisoned.swap(true, Ordering::AcqRel) {\n 388│            // First poison of this conn: loud, once. The token is class-split\n*389│            // (F-039 leg a): deadline → CONN_WRITE_POISONED (the wedge\n 390│            // observable), organic fast-fail → CONN_WRITE_RETIRED.\n 391│            self.lifecycle_log.emit(format_args!(\n 392│                \"{}\",\n*393│                self.render_retirement(self.timed_out.load(Ordering::Acquire), cause)\n 394│            ));\n 395│        }\n 396│        #[cfg(windows)]\n    │...\n 486│            // one shot is not enough). Latch timed_out FIRST so the aborted\n*487│            // writer observes its failure as a DEADLINE retirement (wedge),\n 488│            // not an organic conn death.\n 489│            drop(d);\n 490│            self.timed_out.store(true, Ordering::Release);\n    │...\n 598│    /// Bounded framed write. See [`Self::write_gated`].\n*599│    pub(crate) fn write(&self, env: &Envelope) -> io::Result<ConnWrite> {\n 600│        self.write_gated(env, || true)\n 601│    }\n 602│\n    │...\n 604│    /// serialized gate is acquired and BEFORE any byte is written — the same\n*605│    /// serialization point the old epoch-check-under-`send.lock()` had, so a\n 606│    /// superseded controller writer still bails atomically with respect to\n 607│    /// its successor's writes (P1c, REQ-HAZARD-CONTROLLER-WRITER-REORDER).\n 608│    /// `pre` returning false → `Ok(ConnWrite::Superseded)`, nothing written.\n*609│    pub(crate) fn write_gated(\n 610│        &self,\n 611│        env: &Envelope,\n 612│        pre: impl FnOnce() -> bool,\n    │...\n 696│        inner.gate_cv.notify_all();\n*697│        // A DEADLINE retirement (the watchdog aborted OUR op, or the conn was\n 698│        // already deadline-poisoned) surfaces as TimedOut — the caller can\n 699│        // count the wedge (W2 stall-evict tally). An organic I/O failure\n 700│        // (peer death) keeps its own error.\n    │...\n 774│\n*775│fn poisoned_err() -> io::Error {\n 776│    io::Error::new(\n 777│        io::ErrorKind::BrokenPipe,\n*778│        \"broker conn poisoned (bounded write retired the physical connection)\",\n 779│    )\n 780│}\n 781│\n    │...\n 875│        conn.inner.poison_and_cancel(Some(&io::Error::new(io::ErrorKind::BrokenPipe, \"peer gone\")));\n*876│        // Repeated cancellation must not repeat the retirement record.\n 877│        conn.inner.poison_and_cancel(None);\n 878│        let id = conn.id();\n 879│        drop(conn);\n 880│        let lines = recover(log.capture.as_ref().unwrap());\n*881│        assert_eq!(lines.len(), 3, \"role event, retirement and poisoned close survive\");\n 882│        for line in lines.iter() {\n 883│            assert!(line.contains(&format!(\"conn={id}\")));\n 884│            assert!(line.contains(\"role=brain controller session=7\"));\n    │...\n 893│    // [unit->REQ-CONN-POISON-DIAL-SCOPE] the F-039 leg-(a) token split: the loud\n*894│    // CONN_WRITE_POISONED token is RESERVED for the deadline (timed_out) class;\n 895│    // an organic write failure renders the distinct CONN_WRITE_RETIRED token\n 896│    // carrying the error kind — and (doyle amendment 1) the SAME attribution\n 897│    // fields as POISONED: conn id, accreted role/endpoint/session facts, wall +\n    │...\n 900│    #[test]\n*901│    fn retirement_token_splits_deadline_from_organic_with_attribution_parity() {\n 902│        let (conn, _client) = make_conn();\n 903│        conn.describe(\"role=brain\");\n 904│        conn.describe(\"controller session=7 endpoint=ep by=cli@host\");\n 905│\n*906│        let poisoned = conn.inner.render_retirement(true, None);\n 907│        assert!(\n*908│            poisoned.starts_with(\"CONN_WRITE_POISONED:\"),\n 909│            \"deadline class keeps the loud wedge-observable token: {poisoned}\"\n 910│        );\n*911│        let retired = conn\n 912│            .inner\n*913│            .render_retirement(false, Some(&io::Error::new(io::ErrorKind::BrokenPipe, \"peer gone\")));\n 914│        assert!(\n*915│            retired.starts_with(\"CONN_WRITE_RETIRED:\"),\n*916│            \"organic class emits the DISTINCT retired token: {retired}\"\n 917│        );\n 918│        assert!(\n*919│            retired.contains(\"BrokenPipe\") && retired.contains(\"peer gone\"),\n*920│            \"the organic record carries the error kind + message: {retired}\"\n 921│        );\n 922│        // Attribution parity: every field present on BOTH tokens.\n*923│        for (name, line) in [(\"POISONED\", &poisoned), (\"RETIRED\", &retired)] {\n 924│            assert!(\n 925│                line.contains(&format!(\"conn={}\", conn.id())),\n 926│                \"{name} carries the stable conn id: {line}\"\n    │...\n 940│    // split (the REQ-HAZARD-SHAREDSEND invariant): an organic fast-fail still\n*941│    // poisons + retires the conn — it is only the CLASS that differs (the error\n 942│    // surfaces with its own kind, never TimedOut, and the next write refuses).\n 943│    #[test]\n 944│    fn organic_write_failure_still_poisons_and_is_not_the_timeout_class() {\n    │...\n 965│        );\n*966│        // Poisoned + retired: the next write refuses immediately.\n 967│        let refused = conn.write(&env).expect_err(\"poisoned conn refuses writes\");\n 968│        assert_eq!(refused.kind(), io::ErrorKind::BrokenPipe);\n 969│    }\n    │...\n 994│        // The id (and the capped label) stay stable across records.\n*995│        let one = a.inner.render_retirement(true, None);\n*996│        let two = a.inner.render_retirement(true, None);\n 997│        assert!(one.contains(&format!(\"conn={}\", a.id())));\n 998│        assert!(two.contains(&format!(\"conn={}\", a.id())));\n 999│    }"},"isError":false,"timestamp":1789523970723}}
{"type":"message","id":"8c18fb6d","parentId":"cef16aa5","timestamp":"2026-09-16T01:59:30.728Z","message":{"role":"toolResult","toolCallId":"call_8rmEcmeK5yQeCQi7HhoLF7Qz|fc_0d104aba2b0f3edc016aa9f802a32c87d1a5427b614ef970a2","toolName":"read","content":[{"type":"text","text":"#[allow(clippy::too_many_arguments)] // per-writer thread-handoff channels (W2 added the blocked-since signal); a struct would obscure the call\nfn controller_writer(\n    sid: u64,\n    send: SharedSend,\n    initial: Vec<(u64, Vec<u8>)>,\n    // Whether `initial` is a synthesized repaint batch (rides the wire-flag —\n    // the consumer baselines its pseudo-seq jump) vs a raw ring slice\n    // (unflagged, contiguous). REQ-RC-RESIZE-PRESENTATION-BARRIER.\n    initial_sync: bool,\n    rx: Receiver<CtrlMsg>,\n    delivered_through: Arc<AtomicU64>,\n    epoch: u64,\n    controller_epoch: Arc<AtomicU64>,\n    write_blocked_since: Arc<Mutex<Option<Instant>>>,\n    // The ADR-0044 revoke marker (see [`ControllerSink::revoked_by`]): when a\n    // distinct-lease Take set it, this writer's EXIT emits the terminal\n    // Displaced as its final frame.\n    revoked_by: Arc<Mutex<Option<String>>>,\n) {\n    // W2 decouple (REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE): mark the exact window a\n    // BLOCKING socket write is in flight, so a reader can tell a WEDGED writer (brain\n    // conn suspended) from a merely idle one and stall-evict it. Set BEFORE the write,\n    // cleared AFTER — the mutex is never held across `write_frame` (that would just\n    // move the wedge onto the mutex).\n    let mark_blocked = |on: bool| {\n        if let Ok(mut g) = write_blocked_since.lock() {\n            *g = if on { Some(Instant::now()) } else { None };\n        }\n    };\n    // Terminal-revoke duty (ADR-0044 decision 2, REQ-CONTROLLER-LEASE-IDENTITY):\n    // on a revoked exit, the writer's LAST act is the bounded Displaced write —\n    // behind everything it already committed (order preserved), independent of\n    // the live queue's fullness. A write the conn cannot take retires it via\n    // the per-write watchdog (poison + sever) — the closed conn ends the serve\n    // worker, which is the same terminal signal for the incumbent's rc. Silent\n    // (same-lease) replacements never set the marker, so gate-#7 stays silent.\n    let emit_revoke_if_marked = |send: &SharedSend| {\n        let by = revoked_by.lock().ok().and_then(|mut g| g.take());\n        if let Some(by) = by {\n            let _ = send.write(&displaced_envelope(sid, &by));\n            send.lifecycle_event(\n                \"writer-exit\",\n                &format!(\"role=controller session={sid} reason=lease-revoked taken_by={by}\"),\n            );\n            return true;\n        }\n        false\n    };\n    for (seq, bytes) in &initial {\n        let frame = if initial_sync {\n            sync_output_envelope(sid, *seq, bytes)\n        } else {\n            output_envelope(sid, *seq, bytes)\n        };\n        // P1c: a SUPERSEDED writer (a newer `become_controller` bumped the shared\n        // epoch past ours) must write NOTHING further — otherwise its owned\n        // initial batch races the new writer on the same socket and reorders the\n        // seq stream. The gate closure runs AFTER the conn's serialized write\n        // gate is acquired and BEFORE any byte is written — the same\n        // serialization point the old epoch-check-under-`send.lock()` had: both\n        // writers serialize on the one conn gate, and `become_controller` bumps\n        // the epoch BEFORE spawning the new writer, so whichever superseded\n        // writer next holds the gate observes the bump and bails before\n        // committing — no check-then-block-then-write window. (A frame the\n        // stale writer committed BEFORE the bump is irreducible at the broker;\n        // fix #1 removes the contradictory double-take that produced it.)\n        // Round-4: the write itself is bounded + cancelable — a non-draining\n        // consumer poisons + retires the physical conn at the deadline instead\n        // of wedging this thread in the OS write forever.\n        // [impl->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\n        mark_blocked(true);\n        let outcome =\n            send.write_gated(&frame, || controller_epoch.load(Ordering::Acquire) == epoch);\n        mark_blocked(false);\n        match outcome {\n            Ok(ConnWrite::Done) => {}\n            Ok(ConnWrite::Superseded) => {\n                // A revoked writer superseded mid-initial-batch still owes the\n                // terminal Displaced (ADR-0044). [impl->REQ-CONTROLLER-LEASE-IDENTITY]\n                if emit_revoke_if_marked(&send) {\n                    return;\n                }\n                // Bounded writer-exit record (F-039 leg d).\n                // [impl->REQ-CONN-POISON-ATTRIBUTION]\n                send.lifecycle_event(\n                    \"writer-exit\",\n                    &format!(\"role=controller session={sid} reason=superseded\"),\n                );\n                return;\n            }\n            Err(e) => {\n                note_controller_write_retired(sid, send.id(), &e);\n                send.lifecycle_event(\n                    \"writer-exit\",\n                    &format!(\"role=controller session={sid} reason=write-failed kind={:?}\", e.kind()),\n                );\n                return;\n            }\n        }\n        advance_delivered(&delivered_through, *seq);\n    }\n    while let Ok(msg) = rx.recv() {\n        // `watermark` distinguishes the two advance shapes below: a live Output\n        // steps the cursor CONTIGUOUSLY (gap-freeze, b4/B2), a transition Sync\n        // JUMPS it past the whole suppressed range as-if-written\n        // (REQ-RC-RESIZE-PRESENTATION-BARRIER leg 3 — the repaint supersedes).\n        let (frame, seq, watermark) = match &msg {\n            CtrlMsg::Output(seq, frame) => (frame, Some(*seq), None),\n            CtrlMsg::Sync(seq, frame) => (frame, None, Some(*seq)),\n            // The terminal Exit, sequenced BEHIND the queued output above it\n            // (ADR-0043 decision 1) — written like any frame, no cursor advance.\n            // [impl->REQ-BROKER-OUTPUT-BEFORE-EXIT]\n            CtrlMsg::Exit(frame) => (frame, None, None),\n            // A ceremony frame (WAX-SEAL W2): written like any frame, no\n            // cursor advance (not session output).\n            CtrlMsg::Ceremony(frame) => (frame, None, None),\n            CtrlMsg::InputPaths(frame) => (frame, None, None),\n        };\n        // NO epoch gate on the live path (P1c): new output only ever flows to the\n        // CURRENT controller's channel (the drain clones `self.controller.tx`), so\n        // a superseded writer's channel holds only its pre-supersede output\n        // backlog (seq the consumer dedups via snap-above). The loop ends\n        // naturally when `tx` drops — and a REVOKED replacement then emits the\n        // terminal `Displaced` below, BEHIND this drained backlog (order\n        // preserved), instead of racing it through the live queue. The reorder\n        // hazard is the INITIAL-BATCH replay, gated above.\n        mark_blocked(true);\n        let res = send.write(frame);\n        mark_blocked(false);\n        match res {\n            Ok(_) => {}\n            Err(e) => {\n                note_controller_write_retired(sid, send.id(), &e);\n                // Bounded writer-exit record (F-039 leg d).\n                // [impl->REQ-CONN-POISON-ATTRIBUTION]\n                send.lifecycle_event(\n                    \"writer-exit\",\n                    &format!(\"role=controller session={sid} reason=write-failed kind={:?}\", e.kind()),\n                );\n                return;\n            }\n        }\n        if let Some(seq) = seq {\n            // (a) CONTIGUOUS cursor (b4 / B2 gapless-resume — doyle's gate): with the\n            // drop-on-full handoff in `append`, the controller's live channel can SKIP\n            // a seq (dropped while Full). The frame still went to the socket (the live\n            // operator sees it), but the durable resume cursor must NOT jump past the\n            // gap — see [`contiguous_advance`].\n            contiguous_advance(&delivered_through, seq);\n        }\n        if let Some(watermark) = watermark {\n            // The transition sync frame: one successful write advances the cursor\n            // past the WHOLE suppressed range (monotonic max, CAS — the same\n            // advance the initial repaint batch uses). B2 gapless-resume holds IN\n            // EFFECT: the repaint this frame carries supersedes every skipped seq,\n            // and a resume below the presentation floor is served the repaint\n            // shape, never a raw replay (REQ-RC-RESIZE-PRESENTATION-BARRIER).\n            // [impl->REQ-RC-RESIZE-PRESENTATION-BARRIER]\n            advance_delivered(&delivered_through, watermark);\n        }\n    }\n    // tx dropped — a replace/detach ended the live loop. A REVOKED replacement\n    // (distinct-lease Take) emits the terminal Displaced as the final frame\n    // (ADR-0044); a silent same-lease/detach end stays silent (gate #7).\n    // [impl->REQ-CONTROLLER-LEASE-IDENTITY]\n    if emit_revoke_if_marked(&send) {\n        return;\n    }\n    // Bounded writer-exit record (F-039 leg d).\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\n    send.lifecycle_event(\n        \"writer-exit\",\n        &format!(\"role=controller session={sid} reason=channel-closed\"),\n    );\n}\n\n/// The per-session translation-binary attachment (W2 Layer C, ADR-0022). Present\n/// ONLY when the spawn carried a `translation_binary` path; absent → the session\n/// has no idle PTY delivery and inbound SPOOLS (ADR-0022 amendment — raw inject\n/// was removed as a delivery path; it did not submit on a modern TUI).\n///\n/// Holds the shared inject floor (the controller-input buffer, also touched by\n/// `dispatch_input`), the queue that feeds inbound `<EVENT>`s to the inject worker\n/// (events serialize — never two uncommitted sequences interleaved), the shared\n/// child handle (so `dispatch_input` can fire content-free `{type:input}` pings\n/// while the worker drives an event), and the FAULT flag the worker raises when a\n/// sequence misses its [`INJECT_COMMIT_DEADLINE`] or the binary dies — after which\n/// the session's idle inbound SPOOLS (never raw-injected) and pinging stops.\nstruct Translation {\n    /// Controller-input buffer + inject-floor flag, shared with `dispatch_input`.\n    floor: Arc<Mutex<InjectFloor>>,\n    /// Queue one inbound `<EVENT>` envelope to the inject worker. Events QUEUE here\n    /// and are driven one-at-a-time so two sequences never interleave.\n    event_tx: Sender<String>,\n    /// The live child — shared so `dispatch_input` fires content-free pings while\n    /// the worker (which also holds an `Arc`) drives events / faults the binary.\n    child: Arc<TranslationChild>,\n    /// Raised by the worker on FAULT (commit-deadline miss / binary death): the\n    /// session's idle inbound then SPOOLS (never raw-injected) and `dispatch_input`\n    /// stops pinging.\n    faulted: Arc<AtomicBool>,\n}\n\n/// A child the broker hosts under a PTY — the §B-owned resources for one\n/// session, plus its gapless-resume output log.\nstruct HostedSession {\n    /// The PTY + child (PTY master/writer + spawned child process).\n    session: Arc<PtySession>,\n    /// The session's dedicated PTY input-writer (v0.13.0 P0,\n    /// REQ-HAZARD-PTY-INPUT-WRITER-WEDGE): the SOLE caller of the blocking\n    /// `write_input`. Every dispatch-side write ENQUEUES here and returns at once,\n    /// so a paste burst that fills the harness's input buffer parks only the\n    /// writer thread, never the broker dispatch thread.\n    input: Arc<InputWriter>,\n    /// Live-reference dedup and reply custody survive a brain restart.\n    input_reports: Arc<Mutex<crate::inputreceipt::ReceiptBook>>,\n    /// The output drain pump (kept so teardown can stop it).\n    #[allow(dead_code)] // held for ownership/teardown; not read directly\n    drain: Drain,\n    /// The sequenced, buffered output + current subscriber.\n    log: Arc<Mutex<OutputLog>>,\n    /// The owning endpoint (the spawn's label; empty = unlabeled). The\n    /// inbound dispatcher's gate subject for remote-drive (D9-1).\n    endpoint: String,\n    /// The translation-binary attachment (Layer C), or `None` for a session with\n    /// no `[message-idle-translation-binary]` — then idle inbound delivery SPOOLS\n    /// (no raw PTY inject; ADR-0022 amendment).\n    translation: Option<Arc<Translation>>,\n    /// The resolved translation argv this session was spawned with (F029 C-1) —\n    /// held so a bounded eager respawn after a REAL fault can rebuild the binary\n    /// from the exact same command (`build_translation`) without re-reading the\n    /// manifest (works for a test session with a direct argv AND a production\n    /// session). `None` for a no-translation session. Updated on an adapter-apply\n    /// restart to the swapped-in argv.\n    translation_argv: Option<Vec<String>>,\n    /// The bounded-respawn give-up counter shared with the inject worker (F029 C-1,\n    /// REQ-HAZARD-TRANSLATE-FAULT-PERMANENT-DEATH): the dispatch respawn path\n    /// increments it per rebuild, the worker RESETS it to 0 on a healthy `{commit}`\n    /// (so a recovered binary refreshes the whole budget). At `translate_respawn_\n    /// budget()` the daemon gives up (loud) and leaves the binary dead → spool.\n    translation_respawns: Arc<AtomicU32>,\n    /// The adapter name owning this session (W3a, ADR-0025) — empty for a\n    /// legacy/test/shell spawn. A live adapter update finds every live session\n    /// whose `adapter` matches to cycle its resident children.\n    adapter: String,\n    /// The adapter's install dir for this session (W3a) — `None` if unknown. The\n    /// apply target for a live update's CRC swap.\n    install_dir: Option<std::path::PathBuf>,\n    /// When the session spawned (ADR-0041 decision 6) — the zombie verdict's\n    /// boot-grace anchor: a just-spawned adapter harness has no client chain\n    /// yet and must never read as a dead tree during bringup.\n    // [impl->REQ-ENDPOINT-CYCLE-HONEST]\n    spawned_at: Instant,\n    /// OS birth stamp captured while the newly spawned child is still owned\n    /// and BEFORE any waiter can reap it. Never refreshed from a remembered pid.\n    // [impl->REQ-BROKER-ZOMBIE-IDENTITY]\n    process_started_at: Option<u64>,\n}\n\n/// Source-tagged records retain delivery custody through the single physical writer.\n\n[Showing lines 3954-4213 of 16852. Use :4214 to continue]"}],"details":{"truncation":{"content":"#[allow(clippy::too_many_arguments)] // per-writer thread-handoff channels (W2 added the blocked-since signal); a struct would obscure the call\r\nfn controller_writer(\r\n    sid: u64,\r\n    send: SharedSend,\r\n    initial: Vec<(u64, Vec<u8>)>,\r\n    // Whether `initial` is a synthesized repaint batch (rides the wire-flag —\r\n    // the consumer baselines its pseudo-seq jump) vs a raw ring slice\r\n    // (unflagged, contiguous). REQ-RC-RESIZE-PRESENTATION-BARRIER.\r\n    initial_sync: bool,\r\n    rx: Receiver<CtrlMsg>,\r\n    delivered_through: Arc<AtomicU64>,\r\n    epoch: u64,\r\n    controller_epoch: Arc<AtomicU64>,\r\n    write_blocked_since: Arc<Mutex<Option<Instant>>>,\r\n    // The ADR-0044 revoke marker (see [`ControllerSink::revoked_by`]): when a\r\n    // distinct-lease Take set it, this writer's EXIT emits the terminal\r\n    // Displaced as its final frame.\r\n    revoked_by: Arc<Mutex<Option<String>>>,\r\n) {\r\n    // W2 decouple (REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE): mark the exact window a\r\n    // BLOCKING socket write is in flight, so a reader can tell a WEDGED writer (brain\r\n    // conn suspended) from a merely idle one and stall-evict it. Set BEFORE the write,\r\n    // cleared AFTER — the mutex is never held across `write_frame` (that would just\r\n    // move the wedge onto the mutex).\r\n    let mark_blocked = |on: bool| {\r\n        if let Ok(mut g) = write_blocked_since.lock() {\r\n            *g = if on { Some(Instant::now()) } else { None };\r\n        }\r\n    };\r\n    // Terminal-revoke duty (ADR-0044 decision 2, REQ-CONTROLLER-LEASE-IDENTITY):\r\n    // on a revoked exit, the writer's LAST act is the bounded Displaced write —\r\n    // behind everything it already committed (order preserved), independent of\r\n    // the live queue's fullness. A write the conn cannot take retires it via\r\n    // the per-write watchdog (poison + sever) — the closed conn ends the serve\r\n    // worker, which is the same terminal signal for the incumbent's rc. Silent\r\n    // (same-lease) replacements never set the marker, so gate-#7 stays silent.\r\n    let emit_revoke_if_marked = |send: &SharedSend| {\r\n        let by = revoked_by.lock().ok().and_then(|mut g| g.take());\r\n        if let Some(by) = by {\r\n            let _ = send.write(&displaced_envelope(sid, &by));\r\n            send.lifecycle_event(\r\n                \"writer-exit\",\r\n                &format!(\"role=controller session={sid} reason=lease-revoked taken_by={by}\"),\r\n            );\r\n            return true;\r\n        }\r\n        false\r\n    };\r\n    for (seq, bytes) in &initial {\r\n        let frame = if initial_sync {\r\n            sync_output_envelope(sid, *seq, bytes)\r\n        } else {\r\n            output_envelope(sid, *seq, bytes)\r\n        };\r\n        // P1c: a SUPERSEDED writer (a newer `become_controller` bumped the shared\r\n        // epoch past ours) must write NOTHING further — otherwise its owned\r\n        // initial batch races the new writer on the same socket and reorders the\r\n        // seq stream. The gate closure runs AFTER the conn's serialized write\r\n        // gate is acquired and BEFORE any byte is written — the same\r\n        // serialization point the old epoch-check-under-`send.lock()` had: both\r\n        // writers serialize on the one conn gate, and `become_controller` bumps\r\n        // the epoch BEFORE spawning the new writer, so whichever superseded\r\n        // writer next holds the gate observes the bump and bails before\r\n        // committing — no check-then-block-then-write window. (A frame the\r\n        // stale writer committed BEFORE the bump is irreducible at the broker;\r\n        // fix #1 removes the contradictory double-take that produced it.)\r\n        // Round-4: the write itself is bounded + cancelable — a non-draining\r\n        // consumer poisons + retires the physical conn at the deadline instead\r\n        // of wedging this thread in the OS write forever.\r\n        // [impl->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\r\n        mark_blocked(true);\r\n        let outcome =\r\n            send.write_gated(&frame, || controller_epoch.load(Ordering::Acquire) == epoch);\r\n        mark_blocked(false);\r\n        match outcome {\r\n            Ok(ConnWrite::Done) => {}\r\n            Ok(ConnWrite::Superseded) => {\r\n                // A revoked writer superseded mid-initial-batch still owes the\r\n                // terminal Displaced (ADR-0044). [impl->REQ-CONTROLLER-LEASE-IDENTITY]\r\n                if emit_revoke_if_marked(&send) {\r\n                    return;\r\n                }\r\n                // Bounded writer-exit record (F-039 leg d).\r\n                // [impl->REQ-CONN-POISON-ATTRIBUTION]\r\n                send.lifecycle_event(\r\n                    \"writer-exit\",\r\n                    &format!(\"role=controller session={sid} reason=superseded\"),\r\n                );\r\n                return;\r\n            }\r\n            Err(e) => {\r\n                note_controller_write_retired(sid, send.id(), &e);\r\n                send.lifecycle_event(\r\n                    \"writer-exit\",\r\n                    &format!(\"role=controller session={sid} reason=write-failed kind={:?}\", e.kind()),\r\n                );\r\n                return;\r\n            }\r\n        }\r\n        advance_delivered(&delivered_through, *seq);\r\n    }\r\n    while let Ok(msg) = rx.recv() {\r\n        // `watermark` distinguishes the two advance shapes below: a live Output\r\n        // steps the cursor CONTIGUOUSLY (gap-freeze, b4/B2), a transition Sync\r\n        // JUMPS it past the whole suppressed range as-if-written\r\n        // (REQ-RC-RESIZE-PRESENTATION-BARRIER leg 3 — the repaint supersedes).\r\n        let (frame, seq, watermark) = match &msg {\r\n            CtrlMsg::Output(seq, frame) => (frame, Some(*seq), None),\r\n            CtrlMsg::Sync(seq, frame) => (frame, None, Some(*seq)),\r\n            // The terminal Exit, sequenced BEHIND the queued output above it\r\n            // (ADR-0043 decision 1) — written like any frame, no cursor advance.\r\n            // [impl->REQ-BROKER-OUTPUT-BEFORE-EXIT]\r\n            CtrlMsg::Exit(frame) => (frame, None, None),\r\n            // A ceremony frame (WAX-SEAL W2): written like any frame, no\r\n            // cursor advance (not session output).\r\n            CtrlMsg::Ceremony(frame) => (frame, None, None),\r\n            CtrlMsg::InputPaths(frame) => (frame, None, None),\r\n        };\r\n        // NO epoch gate on the live path (P1c): new output only ever flows to the\r\n        // CURRENT controller's channel (the drain clones `self.controller.tx`), so\r\n        // a superseded writer's channel holds only its pre-supersede output\r\n        // backlog (seq the consumer dedups via snap-above). The loop ends\r\n        // naturally when `tx` drops — and a REVOKED replacement then emits the\r\n        // terminal `Displaced` below, BEHIND this drained backlog (order\r\n        // preserved), instead of racing it through the live queue. The reorder\r\n        // hazard is the INITIAL-BATCH replay, gated above.\r\n        mark_blocked(true);\r\n        let res = send.write(frame);\r\n        mark_blocked(false);\r\n        match res {\r\n            Ok(_) => {}\r\n            Err(e) => {\r\n                note_controller_write_retired(sid, send.id(), &e);\r\n                // Bounded writer-exit record (F-039 leg d).\r\n                // [impl->REQ-CONN-POISON-ATTRIBUTION]\r\n                send.lifecycle_event(\r\n                    \"writer-exit\",\r\n                    &format!(\"role=controller session={sid} reason=write-failed kind={:?}\", e.kind()),\r\n                );\r\n                return;\r\n            }\r\n        }\r\n        if let Some(seq) = seq {\r\n            // (a) CONTIGUOUS cursor (b4 / B2 gapless-resume — doyle's gate): with the\r\n            // drop-on-full handoff in `append`, the controller's live channel can SKIP\r\n            // a seq (dropped while Full). The frame still went to the socket (the live\r\n            // operator sees it), but the durable resume cursor must NOT jump past the\r\n            // gap — see [`contiguous_advance`].\r\n            contiguous_advance(&delivered_through, seq);\r\n        }\r\n        if let Some(watermark) = watermark {\r\n            // The transition sync frame: one successful write advances the cursor\r\n            // past the WHOLE suppressed range (monotonic max, CAS — the same\r\n            // advance the initial repaint batch uses). B2 gapless-resume holds IN\r\n            // EFFECT: the repaint this frame carries supersedes every skipped seq,\r\n            // and a resume below the presentation floor is served the repaint\r\n            // shape, never a raw replay (REQ-RC-RESIZE-PRESENTATION-BARRIER).\r\n            // [impl->REQ-RC-RESIZE-PRESENTATION-BARRIER]\r\n            advance_delivered(&delivered_through, watermark);\r\n        }\r\n    }\r\n    // tx dropped — a replace/detach ended the live loop. A REVOKED replacement\r\n    // (distinct-lease Take) emits the terminal Displaced as the final frame\r\n    // (ADR-0044); a silent same-lease/detach end stays silent (gate #7).\r\n    // [impl->REQ-CONTROLLER-LEASE-IDENTITY]\r\n    if emit_revoke_if_marked(&send) {\r\n        return;\r\n    }\r\n    // Bounded writer-exit record (F-039 leg d).\r\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\r\n    send.lifecycle_event(\r\n        \"writer-exit\",\r\n        &format!(\"role=controller session={sid} reason=channel-closed\"),\r\n    );\r\n}\r\n\r\n/// The per-session translation-binary attachment (W2 Layer C, ADR-0022). Present\r\n/// ONLY when the spawn carried a `translation_binary` path; absent → the session\r\n/// has no idle PTY delivery and inbound SPOOLS (ADR-0022 amendment — raw inject\r\n/// was removed as a delivery path; it did not submit on a modern TUI).\r\n///\r\n/// Holds the shared inject floor (the controller-input buffer, also touched by\r\n/// `dispatch_input`), the queue that feeds inbound `<EVENT>`s to the inject worker\r\n/// (events serialize — never two uncommitted sequences interleaved), the shared\r\n/// child handle (so `dispatch_input` can fire content-free `{type:input}` pings\r\n/// while the worker drives an event), and the FAULT flag the worker raises when a\r\n/// sequence misses its [`INJECT_COMMIT_DEADLINE`] or the binary dies — after which\r\n/// the session's idle inbound SPOOLS (never raw-injected) and pinging stops.\r\nstruct Translation {\r\n    /// Controller-input buffer + inject-floor flag, shared with `dispatch_input`.\r\n    floor: Arc<Mutex<InjectFloor>>,\r\n    /// Queue one inbound `<EVENT>` envelope to the inject worker. Events QUEUE here\r\n    /// and are driven one-at-a-time so two sequences never interleave.\r\n    event_tx: Sender<String>,\r\n    /// The live child — shared so `dispatch_input` fires content-free pings while\r\n    /// the worker (which also holds an `Arc`) drives events / faults the binary.\r\n    child: Arc<TranslationChild>,\r\n    /// Raised by the worker on FAULT (commit-deadline miss / binary death): the\r\n    /// session's idle inbound then SPOOLS (never raw-injected) and `dispatch_input`\r\n    /// stops pinging.\r\n    faulted: Arc<AtomicBool>,\r\n}\r\n\r\n/// A child the broker hosts under a PTY — the §B-owned resources for one\r\n/// session, plus its gapless-resume output log.\r\nstruct HostedSession {\r\n    /// The PTY + child (PTY master/writer + spawned child process).\r\n    session: Arc<PtySession>,\r\n    /// The session's dedicated PTY input-writer (v0.13.0 P0,\r\n    /// REQ-HAZARD-PTY-INPUT-WRITER-WEDGE): the SOLE caller of the blocking\r\n    /// `write_input`. Every dispatch-side write ENQUEUES here and returns at once,\r\n    /// so a paste burst that fills the harness's input buffer parks only the\r\n    /// writer thread, never the broker dispatch thread.\r\n    input: Arc<InputWriter>,\r\n    /// Live-reference dedup and reply custody survive a brain restart.\r\n    input_reports: Arc<Mutex<crate::inputreceipt::ReceiptBook>>,\r\n    /// The output drain pump (kept so teardown can stop it).\r\n    #[allow(dead_code)] // held for ownership/teardown; not read directly\r\n    drain: Drain,\r\n    /// The sequenced, buffered output + current subscriber.\r\n    log: Arc<Mutex<OutputLog>>,\r\n    /// The owning endpoint (the spawn's label; empty = unlabeled). The\r\n    /// inbound dispatcher's gate subject for remote-drive (D9-1).\r\n    endpoint: String,\r\n    /// The translation-binary attachment (Layer C), or `None` for a session with\r\n    /// no `[message-idle-translation-binary]` — then idle inbound delivery SPOOLS\r\n    /// (no raw PTY inject; ADR-0022 amendment).\r\n    translation: Option<Arc<Translation>>,\r\n    /// The resolved translation argv this session was spawned with (F029 C-1) —\r\n    /// held so a bounded eager respawn after a REAL fault can rebuild the binary\r\n    /// from the exact same command (`build_translation`) without re-reading the\r\n    /// manifest (works for a test session with a direct argv AND a production\r\n    /// session). `None` for a no-translation session. Updated on an adapter-apply\r\n    /// restart to the swapped-in argv.\r\n    translation_argv: Option<Vec<String>>,\r\n    /// The bounded-respawn give-up counter shared with the inject worker (F029 C-1,\r\n    /// REQ-HAZARD-TRANSLATE-FAULT-PERMANENT-DEATH): the dispatch respawn path\r\n    /// increments it per rebuild, the worker RESETS it to 0 on a healthy `{commit}`\r\n    /// (so a recovered binary refreshes the whole budget). At `translate_respawn_\r\n    /// budget()` the daemon gives up (loud) and leaves the binary dead → spool.\r\n    translation_respawns: Arc<AtomicU32>,\r\n    /// The adapter name owning this session (W3a, ADR-0025) — empty for a\r\n    /// legacy/test/shell spawn. A live adapter update finds every live session\r\n    /// whose `adapter` matches to cycle its resident children.\r\n    adapter: String,\r\n    /// The adapter's install dir for this session (W3a) — `None` if unknown. The\r\n    /// apply target for a live update's CRC swap.\r\n    install_dir: Option<std::path::PathBuf>,\r\n    /// When the session spawned (ADR-0041 decision 6) — the zombie verdict's\r\n    /// boot-grace anchor: a just-spawned adapter harness has no client chain\r\n    /// yet and must never read as a dead tree during bringup.\r\n    // [impl->REQ-ENDPOINT-CYCLE-HONEST]\r\n    spawned_at: Instant,\r\n    /// OS birth stamp captured while the newly spawned child is still owned\r\n    /// and BEFORE any waiter can reap it. Never refreshed from a remembered pid.\r\n    // [impl->REQ-BROKER-ZOMBIE-IDENTITY]\r\n    process_started_at: Option<u64>,\r\n}\r\n\r\n/// Source-tagged records retain delivery custody through the single physical writer.\r","truncated":true,"truncatedBy":"lines","totalLines":12899,"totalBytes":14483,"outputLines":260,"outputBytes":14483,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":16852,"displayContent":{"text":"#[allow(clippy::too_many_arguments)] // per-writer thread-handoff channels (W2 added the blocked-since signal); a struct would obscure the call\r\nfn controller_writer(\r\n    sid: u64,\r\n    send: SharedSend,\r\n    initial: Vec<(u64, Vec<u8>)>,\r\n    // Whether `initial` is a synthesized repaint batch (rides the wire-flag —\r\n    // the consumer baselines its pseudo-seq jump) vs a raw ring slice\r\n    // (unflagged, contiguous). REQ-RC-RESIZE-PRESENTATION-BARRIER.\r\n    initial_sync: bool,\r\n    rx: Receiver<CtrlMsg>,\r\n    delivered_through: Arc<AtomicU64>,\r\n    epoch: u64,\r\n    controller_epoch: Arc<AtomicU64>,\r\n    write_blocked_since: Arc<Mutex<Option<Instant>>>,\r\n    // The ADR-0044 revoke marker (see [`ControllerSink::revoked_by`]): when a\r\n    // distinct-lease Take set it, this writer's EXIT emits the terminal\r\n    // Displaced as its final frame.\r\n    revoked_by: Arc<Mutex<Option<String>>>,\r\n) {\r\n    // W2 decouple (REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE): mark the exact window a\r\n    // BLOCKING socket write is in flight, so a reader can tell a WEDGED writer (brain\r\n    // conn suspended) from a merely idle one and stall-evict it. Set BEFORE the write,\r\n    // cleared AFTER — the mutex is never held across `write_frame` (that would just\r\n    // move the wedge onto the mutex).\r\n    let mark_blocked = |on: bool| {\r\n        if let Ok(mut g) = write_blocked_since.lock() {\r\n            *g = if on { Some(Instant::now()) } else { None };\r\n        }\r\n    };\r\n    // Terminal-revoke duty (ADR-0044 decision 2, REQ-CONTROLLER-LEASE-IDENTITY):\r\n    // on a revoked exit, the writer's LAST act is the bounded Displaced write —\r\n    // behind everything it already committed (order preserved), independent of\r\n    // the live queue's fullness. A write the conn cannot take retires it via\r\n    // the per-write watchdog (poison + sever) — the closed conn ends the serve\r\n    // worker, which is the same terminal signal for the incumbent's rc. Silent\r\n    // (same-lease) replacements never set the marker, so gate-#7 stays silent.\r\n    let emit_revoke_if_marked = |send: &SharedSend| {\r\n        let by = revoked_by.lock().ok().and_then(|mut g| g.take());\r\n        if let Some(by) = by {\r\n            let _ = send.write(&displaced_envelope(sid, &by));\r\n            send.lifecycle_event(\r\n                \"writer-exit\",\r\n                &format!(\"role=controller session={sid} reason=lease-revoked taken_by={by}\"),\r\n            );\r\n            return true;\r\n        }\r\n        false\r\n    };\r\n    for (seq, bytes) in &initial {\r\n        let frame = if initial_sync {\r\n            sync_output_envelope(sid, *seq, bytes)\r\n        } else {\r\n            output_envelope(sid, *seq, bytes)\r\n        };\r\n        // P1c: a SUPERSEDED writer (a newer `become_controller` bumped the shared\r\n        // epoch past ours) must write NOTHING further — otherwise its owned\r\n        // initial batch races the new writer on the same socket and reorders the\r\n        // seq stream. The gate closure runs AFTER the conn's serialized write\r\n        // gate is acquired and BEFORE any byte is written — the same\r\n        // serialization point the old epoch-check-under-`send.lock()` had: both\r\n        // writers serialize on the one conn gate, and `become_controller` bumps\r\n        // the epoch BEFORE spawning the new writer, so whichever superseded\r\n        // writer next holds the gate observes the bump and bails before\r\n        // committing — no check-then-block-then-write window. (A frame the\r\n        // stale writer committed BEFORE the bump is irreducible at the broker;\r\n        // fix #1 removes the contradictory double-take that produced it.)\r\n        // Round-4: the write itself is bounded + cancelable — a non-draining\r\n        // consumer poisons + retires the physical conn at the deadline instead\r\n        // of wedging this thread in the OS write forever.\r\n        // [impl->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\r\n        mark_blocked(true);\r\n        let outcome =\r\n            send.write_gated(&frame, || controller_epoch.load(Ordering::Acquire) == epoch);\r\n        mark_blocked(false);\r\n        match outcome {\r\n            Ok(ConnWrite::Done) => {}\r\n            Ok(ConnWrite::Superseded) => {\r\n                // A revoked writer superseded mid-initial-batch still owes the\r\n                // terminal Displaced (ADR-0044). [impl->REQ-CONTROLLER-LEASE-IDENTITY]\r\n                if emit_revoke_if_marked(&send) {\r\n                    return;\r\n                }\r\n                // Bounded writer-exit record (F-039 leg d).\r\n                // [impl->REQ-CONN-POISON-ATTRIBUTION]\r\n                send.lifecycle_event(\r\n                    \"writer-exit\",\r\n                    &format!(\"role=controller session={sid} reason=superseded\"),\r\n                );\r\n                return;\r\n            }\r\n            Err(e) => {\r\n                note_controller_write_retired(sid, send.id(), &e);\r\n                send.lifecycle_event(\r\n                    \"writer-exit\",\r\n                    &format!(\"role=controller session={sid} reason=write-failed kind={:?}\", e.kind()),\r\n                );\r\n                return;\r\n            }\r\n        }\r\n        advance_delivered(&delivered_through, *seq);\r\n    }\r\n    while let Ok(msg) = rx.recv() {\r\n        // `watermark` distinguishes the two advance shapes below: a live Output\r\n        // steps the cursor CONTIGUOUSLY (gap-freeze, b4/B2), a transition Sync\r\n        // JUMPS it past the whole suppressed range as-if-written\r\n        // (REQ-RC-RESIZE-PRESENTATION-BARRIER leg 3 — the repaint supersedes).\r\n        let (frame, seq, watermark) = match &msg {\r\n            CtrlMsg::Output(seq, frame) => (frame, Some(*seq), None),\r\n            CtrlMsg::Sync(seq, frame) => (frame, None, Some(*seq)),\r\n            // The terminal Exit, sequenced BEHIND the queued output above it\r\n            // (ADR-0043 decision 1) — written like any frame, no cursor advance.\r\n            // [impl->REQ-BROKER-OUTPUT-BEFORE-EXIT]\r\n            CtrlMsg::Exit(frame) => (frame, None, None),\r\n            // A ceremony frame (WAX-SEAL W2): written like any frame, no\r\n            // cursor advance (not session output).\r\n            CtrlMsg::Ceremony(frame) => (frame, None, None),\r\n            CtrlMsg::InputPaths(frame) => (frame, None, None),\r\n        };\r\n        // NO epoch gate on the live path (P1c): new output only ever flows to the\r\n        // CURRENT controller's channel (the drain clones `self.controller.tx`), so\r\n        // a superseded writer's channel holds only its pre-supersede output\r\n        // backlog (seq the consumer dedups via snap-above). The loop ends\r\n        // naturally when `tx` drops — and a REVOKED replacement then emits the\r\n        // terminal `Displaced` below, BEHIND this drained backlog (order\r\n        // preserved), instead of racing it through the live queue. The reorder\r\n        // hazard is the INITIAL-BATCH replay, gated above.\r\n        mark_blocked(true);\r\n        let res = send.write(frame);\r\n        mark_blocked(false);\r\n        match res {\r\n            Ok(_) => {}\r\n            Err(e) => {\r\n                note_controller_write_retired(sid, send.id(), &e);\r\n                // Bounded writer-exit record (F-039 leg d).\r\n                // [impl->REQ-CONN-POISON-ATTRIBUTION]\r\n                send.lifecycle_event(\r\n                    \"writer-exit\",\r\n                    &format!(\"role=controller session={sid} reason=write-failed kind={:?}\", e.kind()),\r\n                );\r\n                return;\r\n            }\r\n        }\r\n        if let Some(seq) = seq {\r\n            // (a) CONTIGUOUS cursor (b4 / B2 gapless-resume — doyle's gate): with the\r\n            // drop-on-full handoff in `append`, the controller's live channel can SKIP\r\n            // a seq (dropped while Full). The frame still went to the socket (the live\r\n            // operator sees it), but the durable resume cursor must NOT jump past the\r\n            // gap — see [`contiguous_advance`].\r\n            contiguous_advance(&delivered_through, seq);\r\n        }\r\n        if let Some(watermark) = watermark {\r\n            // The transition sync frame: one successful write advances the cursor\r\n            // past the WHOLE suppressed range (monotonic max, CAS — the same\r\n            // advance the initial repaint batch uses). B2 gapless-resume holds IN\r\n            // EFFECT: the repaint this frame carries supersedes every skipped seq,\r\n            // and a resume below the presentation floor is served the repaint\r\n            // shape, never a raw replay (REQ-RC-RESIZE-PRESENTATION-BARRIER).\r\n            // [impl->REQ-RC-RESIZE-PRESENTATION-BARRIER]\r\n            advance_delivered(&delivered_through, watermark);\r\n        }\r\n    }\r\n    // tx dropped — a replace/detach ended the live loop. A REVOKED replacement\r\n    // (distinct-lease Take) emits the terminal Displaced as the final frame\r\n    // (ADR-0044); a silent same-lease/detach end stays silent (gate #7).\r\n    // [impl->REQ-CONTROLLER-LEASE-IDENTITY]\r\n    if emit_revoke_if_marked(&send) {\r\n        return;\r\n    }\r\n    // Bounded writer-exit record (F-039 leg d).\r\n    // [impl->REQ-CONN-POISON-ATTRIBUTION]\r\n    send.lifecycle_event(\r\n        \"writer-exit\",\r\n        &format!(\"role=controller session={sid} reason=channel-closed\"),\r\n    );\r\n}\r\n\r\n/// The per-session translation-binary attachment (W2 Layer C, ADR-0022). Present\r\n/// ONLY when the spawn carried a `translation_binary` path; absent → the session\r\n/// has no idle PTY delivery and inbound SPOOLS (ADR-0022 amendment — raw inject\r\n/// was removed as a delivery path; it did not submit on a modern TUI).\r\n///\r\n/// Holds the shared inject floor (the controller-input buffer, also touched by\r\n/// `dispatch_input`), the queue that feeds inbound `<EVENT>`s to the inject worker\r\n/// (events serialize — never two uncommitted sequences interleaved), the shared\r\n/// child handle (so `dispatch_input` can fire content-free `{type:input}` pings\r\n/// while the worker drives an event), and the FAULT flag the worker raises when a\r\n/// sequence misses its [`INJECT_COMMIT_DEADLINE`] or the binary dies — after which\r\n/// the session's idle inbound SPOOLS (never raw-injected) and pinging stops.\r\nstruct Translation {\r\n    /// Controller-input buffer + inject-floor flag, shared with `dispatch_input`.\r\n    floor: Arc<Mutex<InjectFloor>>,\r\n    /// Queue one inbound `<EVENT>` envelope to the inject worker. Events QUEUE here\r\n    /// and are driven one-at-a-time so two sequences never interleave.\r\n    event_tx: Sender<String>,\r\n    /// The live child — shared so `dispatch_input` fires content-free pings while\r\n    /// the worker (which also holds an `Arc`) drives events / faults the binary.\r\n    child: Arc<TranslationChild>,\r\n    /// Raised by the worker on FAULT (commit-deadline miss / binary death): the\r\n    /// session's idle inbound then SPOOLS (never raw-injected) and `dispatch_input`\r\n    /// stops pinging.\r\n    faulted: Arc<AtomicBool>,\r\n}\r\n\r\n/// A child the broker hosts under a PTY — the §B-owned resources for one\r\n/// session, plus its gapless-resume output log.\r\nstruct HostedSession {\r\n    /// The PTY + child (PTY master/writer + spawned child process).\r\n    session: Arc<PtySession>,\r\n    /// The session's dedicated PTY input-writer (v0.13.0 P0,\r\n    /// REQ-HAZARD-PTY-INPUT-WRITER-WEDGE): the SOLE caller of the blocking\r\n    /// `write_input`. Every dispatch-side write ENQUEUES here and returns at once,\r\n    /// so a paste burst that fills the harness's input buffer parks only the\r\n    /// writer thread, never the broker dispatch thread.\r\n    input: Arc<InputWriter>,\r\n    /// Live-reference dedup and reply custody survive a brain restart.\r\n    input_reports: Arc<Mutex<crate::inputreceipt::ReceiptBook>>,\r\n    /// The output drain pump (kept so teardown can stop it).\r\n    #[allow(dead_code)] // held for ownership/teardown; not read directly\r\n    drain: Drain,\r\n    /// The sequenced, buffered output + current subscriber.\r\n    log: Arc<Mutex<OutputLog>>,\r\n    /// The owning endpoint (the spawn's label; empty = unlabeled). The\r\n    /// inbound dispatcher's gate subject for remote-drive (D9-1).\r\n    endpoint: String,\r\n    /// The translation-binary attachment (Layer C), or `None` for a session with\r\n    /// no `[message-idle-translation-binary]` — then idle inbound delivery SPOOLS\r\n    /// (no raw PTY inject; ADR-0022 amendment).\r\n    translation: Option<Arc<Translation>>,\r\n    /// The resolved translation argv this session was spawned with (F029 C-1) —\r\n    /// held so a bounded eager respawn after a REAL fault can rebuild the binary\r\n    /// from the exact same command (`build_translation`) without re-reading the\r\n    /// manifest (works for a test session with a direct argv AND a production\r\n    /// session). `None` for a no-translation session. Updated on an adapter-apply\r\n    /// restart to the swapped-in argv.\r\n    translation_argv: Option<Vec<String>>,\r\n    /// The bounded-respawn give-up counter shared with the inject worker (F029 C-1,\r\n    /// REQ-HAZARD-TRANSLATE-FAULT-PERMANENT-DEATH): the dispatch respawn path\r\n    /// increments it per rebuild, the worker RESETS it to 0 on a healthy `{commit}`\r\n    /// (so a recovered binary refreshes the whole budget). At `translate_respawn_\r\n    /// budget()` the daemon gives up (loud) and leaves the binary dead → spool.\r\n    translation_respawns: Arc<AtomicU32>,\r\n    /// The adapter name owning this session (W3a, ADR-0025) — empty for a\r\n    /// legacy/test/shell spawn. A live adapter update finds every live session\r\n    /// whose `adapter` matches to cycle its resident children.\r\n    adapter: String,\r\n    /// The adapter's install dir for this session (W3a) — `None` if unknown. The\r\n    /// apply target for a live update's CRC swap.\r\n    install_dir: Option<std::path::PathBuf>,\r\n    /// When the session spawned (ADR-0041 decision 6) — the zombie verdict's\r\n    /// boot-grace anchor: a just-spawned adapter harness has no client chain\r\n    /// yet and must never read as a dead tree during bringup.\r\n    // [impl->REQ-ENDPOINT-CYCLE-HONEST]\r\n    spawned_at: Instant,\r\n    /// OS birth stamp captured while the newly spawned child is still owned\r\n    /// and BEFORE any waiter can reap it. Never refreshed from a remembered pid.\r\n    // [impl->REQ-BROKER-ZOMBIE-IDENTITY]\r\n    process_started_at: Option<u64>,\r\n}\r\n\r\n/// Source-tagged records retain delivery custody through the single physical writer.\r","startLine":3954,"lineNumbers":[3954,3955,3956,3957,3958,3959,3960,3961,3962,3963,3964,3965,3966,3967,3968,3969,3970,3971,3972,3973,3974,3975,3976,3977,3978,3979,3980,3981,3982,3983,3984,3985,3986,3987,3988,3989,3990,3991,3992,3993,3994,3995,3996,3997,3998,3999,4000,4001,4002,4003,4004,4005,4006,4007,4008,4009,4010,4011,4012,4013,4014,4015,4016,4017,4018,4019,4020,4021,4022,4023,4024,4025,4026,4027,4028,4029,4030,4031,4032,4033,4034,4035,4036,4037,4038,4039,4040,4041,4042,4043,4044,4045,4046,4047,4048,4049,4050,4051,4052,4053,4054,4055,4056,4057,4058,4059,4060,4061,4062,4063,4064,4065,4066,4067,4068,4069,4070,4071,4072,4073,4074,4075,4076,4077,4078,4079,4080,4081,4082,4083,4084,4085,4086,4087,4088,4089,4090,4091,4092,4093,4094,4095,4096,4097,4098,4099,4100,4101,4102,4103,4104,4105,4106,4107,4108,4109,4110,4111,4112,4113,4114,4115,4116,4117,4118,4119,4120,4121,4122,4123,4124,4125,4126,4127,4128,4129,4130,4131,4132,4133,4134,4135,4136,4137,4138,4139,4140,4141,4142,4143,4144,4145,4146,4147,4148,4149,4150,4151,4152,4153,4154,4155,4156,4157,4158,4159,4160,4161,4162,4163,4164,4165,4166,4167,4168,4169,4170,4171,4172,4173,4174,4175,4176,4177,4178,4179,4180,4181,4182,4183,4184,4185,4186,4187,4188,4189,4190,4191,4192,4193,4194,4195,4196,4197,4198,4199,4200,4201,4202,4203,4204,4205,4206,4207,4208,4209,4210,4211,4212,4213]},"fileSize":877537,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\crates\\spt-daemon\\src\\broker.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":16852,"totalBytes":14483,"outputLines":260,"outputBytes":14483,"shownRange":{"start":3954,"end":4213},"nextOffset":4214}}},"isError":false,"timestamp":1789523970728}}
{"type":"message","id":"2dbce2e0","parentId":"8c18fb6d","timestamp":"2026-09-16T01:59:30.733Z","message":{"role":"toolResult","toolCallId":"call_OK8fpkIcMHacdmFLwcEgq71S|fc_0d104aba2b0f3edc016aa9f802a33c87d19d81f4780e6bf1e9","toolName":"read","content":[{"type":"text","text":"    /// owned by `sub`. Leaves `delivered_through` intact (D4-1). Re-stamps\n    /// `driven_by` / `viewer_count` when a role actually changed.\n    /// Generation-validated detach (ADR-0038 Amendment fix 6): `Some(g)` with\n    /// a controller whose generation differs = a STALE worker's late release —\n    /// the controller slot is left alone (viewer rows still clean up by ptr).\n    /// `None` (N-1 / conn-cleanup) = today's ptr-identity behavior.\n    // [impl->REQ-SERVE-OWNERSHIP-GENERATION]\n    fn detach_if_gen(&mut self, sub: &SharedSend, gen: Option<u64>) {\n        if let (Some(g), Some(c)) = (gen, self.controller.as_ref()) {\n            if c.attach_gen != g {\n                // Stale generation: never clears the (newer) controller.\n                let before = self.viewers.len();\n                self.viewers.retain(|_, sink| !Arc::ptr_eq(&sink.send, sub));\n                if self.viewers.len() != before {\n                    self.stamp_viewer_count();\n                }\n                return;\n            }\n        }\n        self.detach_if(sub);\n    }\n\n    fn detach_if(&mut self, sub: &SharedSend) {\n        let is_ctrl = self\n            .controller\n            .as_ref()\n            .is_some_and(|c| Arc::ptr_eq(&c.send, sub));\n        if is_ctrl {\n            self.clear_controller(Self::CLOSE_DETACH); // drops the sink (ends its writer) + unlatches\n        }\n        let before = self.viewers.len();\n        self.viewers.retain(|_, sink| !Arc::ptr_eq(&sink.send, sub));\n        // Bounded detach record (F-039 leg d) — only when a role actually changed.\n        // [impl->REQ-CONN-POISON-ATTRIBUTION]\n        if is_ctrl || self.viewers.len() != before {\n            sub.lifecycle_event(\n                \"session-detach\",\n                &format!(\n                    \"session={} was_controller={is_ctrl} viewers_dropped={}\",\n                    self.session_id,\n                    before - self.viewers.len()\n                ),\n            );\n        }\n        if self.viewers.len() != before {\n            self.stamp_viewer_count();\n        }\n    }\n\n\n    /// Stamp the perch's `driven_by` to the current controller's identity (the\n    /// broker is the single writer — resolves the clear-race). The remote-drive\n    /// detection fact (REQ-REACH-1) moved here from `serve_attach` so a displaced\n    /// operator's teardown can never wipe the new controller's marker.\n    /// Best-effort: a session served without a perch (raw broker tests) has\n    /// nowhere to record.\n…\n    /// across brain restarts).\n    pub fn session_pid(&self, id: u64) -> Option<u32> {\n        self.sessions\n            .lock()\n            .unwrap()\n            .get(&id)\n            .and_then(|h| h.session.process_id())\n    }\n\n    /// Broker-truth read for the brain promotion DRAINED gate\n    /// (REQ-UPDATE-PROMOTE-DRAINED): is ANY hosted session's LOCAL controller (the\n    /// brain's own `by: None` conn) still WEDGED — blocked in a socket write past the\n    /// brain-write deadline? The brain supervisor (in THIS broker process — no round-\n    /// trip) promotes a ready update candidate only once this reads `false`: an OLD\n    /// generation whose local controller conn is black-holed reads wedged until the W2\n    /// stall-evict releases it, so `brain.ready` inside the 15 s window can no longer\n    /// promote onto a still-frozen control plane (the 22:47 false-promote). A cleanly-\n    /// closed old conn is simply absent → `false` at once (the normal fast promote).\n    /// Follows the established sessions→log lock order (KIND_SESSIONS reap); reads\n    /// only `write_blocked_since` under each log lock, never any I/O.\n    // [impl->REQ-UPDATE-PROMOTE-DRAINED]\n    pub fn any_local_controller_wedged(&self) -> bool {\n        let now = Instant::now();\n        let deadline = brain_write_deadline();\n        let sessions = match self.sessions.lock() {\n            Ok(g) => g,\n            Err(p) => p.into_inner(),\n        };\n        sessions.values().any(|h| {\n            match h.log.lock() {\n                Ok(log) => log.local_controller_wedged(now, deadline),\n                // A poisoned log lock cannot be honestly read — treat as NOT wedged\n                // (fail toward promotion, never a permanent gate that strands apply).\n                Err(_) => false,\n            }\n        })\n    }\n\n    /// The count of live QUIC connections the broker-owned net endpoint holds\n    /// (D4a, ADR-0004 §B), or 0 when net-less. A direct read of the broker's own\n    /// state — the survival-assertion seam for the D7-1 process-level E2E: a QUIC\n    /// connection the broker holds must outlive a brain-PROCESS swap (restoration\n    /// D7-1), the net sibling of [`session_count`](Self::session_count).\n    // [impl->REQ-HAZARD-BROKER-PROCESS-ISOLATION]\n    pub fn net_conn_count(&self) -> usize {\n        self.net.get().map(|h| h.conn_count() as usize).unwrap_or(0)\n    }\n\n    /// The output-log high-water (`next_seq`) of a hosted session, or `None` if\n    /// absent. It advances every time the child emits output — the\n    /// **functional**-survival seam for the D7-1 process-level E2E (doyle F1):"}],"details":{"displayContent":{"text":"    /// owned by `sub`. Leaves `delivered_through` intact (D4-1). Re-stamps\n    /// `driven_by` / `viewer_count` when a role actually changed.\n    /// Generation-validated detach (ADR-0038 Amendment fix 6): `Some(g)` with\n    /// a controller whose generation differs = a STALE worker's late release —\n    /// the controller slot is left alone (viewer rows still clean up by ptr).\n    /// `None` (N-1 / conn-cleanup) = today's ptr-identity behavior.\n    // [impl->REQ-SERVE-OWNERSHIP-GENERATION]\n    fn detach_if_gen(&mut self, sub: &SharedSend, gen: Option<u64>) {\n        if let (Some(g), Some(c)) = (gen, self.controller.as_ref()) {\n            if c.attach_gen != g {\n                // Stale generation: never clears the (newer) controller.\n                let before = self.viewers.len();\n                self.viewers.retain(|_, sink| !Arc::ptr_eq(&sink.send, sub));\n                if self.viewers.len() != before {\n                    self.stamp_viewer_count();\n                }\n                return;\n            }\n        }\n        self.detach_if(sub);\n    }\n\n    fn detach_if(&mut self, sub: &SharedSend) {\n        let is_ctrl = self\n            .controller\n            .as_ref()\n            .is_some_and(|c| Arc::ptr_eq(&c.send, sub));\n        if is_ctrl {\n            self.clear_controller(Self::CLOSE_DETACH); // drops the sink (ends its writer) + unlatches\n        }\n        let before = self.viewers.len();\n        self.viewers.retain(|_, sink| !Arc::ptr_eq(&sink.send, sub));\n        // Bounded detach record (F-039 leg d) — only when a role actually changed.\n        // [impl->REQ-CONN-POISON-ATTRIBUTION]\n        if is_ctrl || self.viewers.len() != before {\n            sub.lifecycle_event(\n                \"session-detach\",\n                &format!(\n                    \"session={} was_controller={is_ctrl} viewers_dropped={}\",\n                    self.session_id,\n                    before - self.viewers.len()\n                ),\n            );\n        }\n        if self.viewers.len() != before {\n            self.stamp_viewer_count();\n        }\n    }\n\n\n    /// Stamp the perch's `driven_by` to the current controller's identity (the\n    /// broker is the single writer — resolves the clear-race). The remote-drive\n    /// detection fact (REQ-REACH-1) moved here from `serve_attach` so a displaced\n    /// operator's teardown can never wipe the new controller's marker.\n    /// Best-effort: a session served without a perch (raw broker tests) has\n    /// nowhere to record.\n…\n    /// across brain restarts).\n    pub fn session_pid(&self, id: u64) -> Option<u32> {\n        self.sessions\n            .lock()\n            .unwrap()\n            .get(&id)\n            .and_then(|h| h.session.process_id())\n    }\n\n    /// Broker-truth read for the brain promotion DRAINED gate\n    /// (REQ-UPDATE-PROMOTE-DRAINED): is ANY hosted session's LOCAL controller (the\n    /// brain's own `by: None` conn) still WEDGED — blocked in a socket write past the\n    /// brain-write deadline? The brain supervisor (in THIS broker process — no round-\n    /// trip) promotes a ready update candidate only once this reads `false`: an OLD\n    /// generation whose local controller conn is black-holed reads wedged until the W2\n    /// stall-evict releases it, so `brain.ready` inside the 15 s window can no longer\n    /// promote onto a still-frozen control plane (the 22:47 false-promote). A cleanly-\n    /// closed old conn is simply absent → `false` at once (the normal fast promote).\n    /// Follows the established sessions→log lock order (KIND_SESSIONS reap); reads\n    /// only `write_blocked_since` under each log lock, never any I/O.\n    // [impl->REQ-UPDATE-PROMOTE-DRAINED]\n    pub fn any_local_controller_wedged(&self) -> bool {\n        let now = Instant::now();\n        let deadline = brain_write_deadline();\n        let sessions = match self.sessions.lock() {\n            Ok(g) => g,\n            Err(p) => p.into_inner(),\n        };\n        sessions.values().any(|h| {\n            match h.log.lock() {\n                Ok(log) => log.local_controller_wedged(now, deadline),\n                // A poisoned log lock cannot be honestly read — treat as NOT wedged\n                // (fail toward promotion, never a permanent gate that strands apply).\n                Err(_) => false,\n            }\n        })\n    }\n\n    /// The count of live QUIC connections the broker-owned net endpoint holds\n    /// (D4a, ADR-0004 §B), or 0 when net-less. A direct read of the broker's own\n    /// state — the survival-assertion seam for the D7-1 process-level E2E: a QUIC\n    /// connection the broker holds must outlive a brain-PROCESS swap (restoration\n    /// D7-1), the net sibling of [`session_count`](Self::session_count).\n    // [impl->REQ-HAZARD-BROKER-PROCESS-ISOLATION]\n    pub fn net_conn_count(&self) -> usize {\n        self.net.get().map(|h| h.conn_count() as usize).unwrap_or(0)\n    }\n\n    /// The output-log high-water (`next_seq`) of a hosted session, or `None` if\n    /// absent. It advances every time the child emits output — the\n    /// **functional**-survival seam for the D7-1 process-level E2E (doyle F1):","startLine":3540,"lineNumbers":[3540,3541,3542,3543,3544,3545,3546,3547,3548,3549,3550,3551,3552,3553,3554,3555,3556,3557,3558,3559,3560,3561,3562,3563,3564,3565,3566,3567,3568,3569,3570,3571,3572,3573,3574,3575,3576,3577,3578,3579,3580,3581,3582,3583,3584,3585,3586,3587,3588,3589,3590,3591,3592,3593,3594,3595,null,5982,5983,5984,5985,5986,5987,5988,5989,5990,5991,5992,5993,5994,5995,5996,5997,5998,5999,6000,6001,6002,6003,6004,6005,6006,6007,6008,6009,6010,6011,6012,6013,6014,6015,6016,6017,6018,6019,6020,6021,6022,6023,6024,6025,6026,6027,6028,6029,6030,6031,6032]},"fileSize":877537,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\crates\\spt-daemon\\src\\broker.rs"}}},"isError":false,"timestamp":1789523970733}}
{"type":"message","id":"6180c264","parentId":"2dbce2e0","timestamp":"2026-09-16T01:59:30.736Z","message":{"role":"toolResult","toolCallId":"call_ka3JN6XQS5HfwPTgWJC1lcOL|fc_0d104aba2b0f3edc016aa9f802a34487d1886741df249f405a","toolName":"read","content":[{"type":"text","text":".\n  - .claude/                                                          20m ago\n    - hertz-commune.md                                      43.9KB    11h ago\n    - scheduled_tasks.lock                                  124B      2w ago\n    - reports/                                                        1mo ago\n    - LIVE_AGENT_IDS.json                                   300B      2mo ago\n    - agent-memory/                                                   2mo ago\n  - .worktrees/                                                       20m ago\n    - 308-registry-process-lock/                                      11m ago\n    - hertz-307-evidence/                                             18m ago\n    - hertz-ir11-built/                                               10h ago\n    - hertz-ir11-headless-column/                                     10h ago\n    - hertz-ir86-retire/                                              11h ago\n    - hertz-ir86-disk-delta/                                          11h ago\n    - hertz-ir137-launch-discipline/                                  11h ago\n    - hertz-ir86-windows-floor/                                       11h ago\n    - hertz-ir86-linux-floor/                                         12h ago\n    - hertz-304-phase-b/                                              13h ago\n    - hertz-ir136-fixture/                                            13h ago\n    - … 104 more\n    - spec-pilot/                                                     2mo ago\n  - docs/                                                             10h ago\n    - INFRA-REGISTER.md                                     665.3KB   10h ago\n    - GOLDEN-CI.md                                          15.9KB    11h ago\n    - TRACEABILITY.md                                       6.9KB     12h ago\n    - adr/                                                            14h ago\n    - STORAGE.md                                            35.7KB    14h ago\n    - INPUT-PROVENANCE-CONTRACT.md                          37.6KB    14h ago\n    - 49-267-REFRESH-OBSERVATION-PLAN.md                    14.9KB    14h ago\n    - 297-ENFORCEMENT-CODES-DIAGNOSIS.md                    12.9KB    14h ago\n    - 297-OUT-OF-BAND-CAPTURE-PLAN.md                       19.3KB    2d ago\n    - 297-CAPTURE-INSTRUMENT-SPEC.md                        19.6KB    3d ago\n    - 304-W2-ADMISSION-RECEIPT.md                           10.0KB    3d ago\n    - … 76 more\n    - reviews/                                                        3mo ago\n  - traceable-reqs.toml                                     2.0MB     10h ago\n  - Cargo.toml                                              3.7KB     14h ago\n  - Cargo.lock                                              140.9KB   14h ago\n  - CONTEXT.md                                              343.4KB   14h ago\n  - CHANGELOG.md                                            219.2KB   14h ago\n  - .config/                                                          14h ago\n    - nextest.toml                                          33.5KB    14h ago\n  - C/                                                                1d ago\n    - debug/                                                          1d ago\n  - sh.exe.stackdump                                        1.2KB     2d ago\n  - .spt/                                                             2d ago\n    - preserved/                                                      21m ago\n    - msg-impl.txt                                          5.7KB     2d ago\n    - msg-ack.txt                                           2.2KB     2d ago\n    - msg-prov.txt                                          7.0KB     2d ago\n    - msg-f5r21.txt                                         4.6KB     2d ago\n    - msg-f5.txt                                            3.4KB     2d ago\n    - cleanup-20260911.json                                 23.0KB    4d ago\n    - DIAG-ARM-CRITERIA-294.md                              11.1KB    5d ago\n    - gate-poll.sh                                          4.2KB     5d ago\n    - handoff/                                                        5d ago\n    - msg-doyle-recover.txt                                 2.9KB     6d ago\n    - … 175 more\n    - docs-nits-0672-evidence/                                        1w ago\n  - GATE-W2-272-CHECKLIST.md                                372.7KB   5d ago\n  - crates/                                                           6d ago\n    - spt-daemon/                                                     14h ago\n    - spt-runtime/                                                    6d ago\n    - xtask/                                                          2w ago\n    - spt/                                                            3w ago\n    - spt-store/                                                      3w ago\n    - spt-procident/                                                  1mo ago\n    - spt-poolguard/                                                  1mo ago\n    - spt-net/                                                        1mo ago\n    - spt-test-support/                                               1mo ago\n    - spt-msg/                                                        1mo ago\n    - spt-term/                                                       1mo ago\n    - … 1 more\n    - spt-live/                                                       2mo ago\n  - .gitignore                                              3.8KB     6d ago\n  - docs-site/                                                        6d ago\n    - src/                                                            14h ago\n    - llms.txt                                              5.2KB     6d ago\n    - book.toml                                             872B      2mo ago\n    - book/                                                           2mo ago\n    - theme/                                                          3mo ago\n  - WEBSERVE-272-W2-JIT.md                                  35.6KB    6d ago\n  - W3-272-MEASUREMENTS.md                                  4.3KB     6d ago\n  - PR-BODY-W2.md                                           21.2KB    6d ago\n  - WEBSERVE-272-JIT.md                                     30.0KB    1w ago\n  - WEBSERVE-272-W3-DRIFT-RIDERS.md                         7.6KB     1w ago\n  - REGISTER-SWEEP-0671-DRAFT.md                            28.1KB    1w ago\n  - RCA-274-R1-LINUX.md                                     5.3KB     1w ago\n  - EMISSION-RESIDUAL-CENSUS-FINDING.md                     7.8KB     2w ago\n  - RCA-242-R2-LINUX.md                                     13.7KB    2w ago\n  - REGISTER-SWEEP-242-DRAFT.md                             6.8KB     2w ago\n  - EMISSION-RESIDUAL-CENSUS-METER.py                       2.6KB     2w ago\n  - ASM-241-GATE-VERDICT.md                                 6.9KB     2w ago\n  - AGENTS.md                                               8.2KB     2w ago\n  - IO-PARSER-22-JIT.md                                     6.0KB     2w ago\n  - FIELD-SEAL-63-JIT.md                                    5.7KB     3w ago\n  - SIGNET-218-JIT.md                                       10.7KB    3w ago\n  - WAX-SEAL-21-JIT.md                                      10.4KB    3w ago\n  - WAX-SEAL-W3-JIT.md                                      5.4KB     3w ago\n  - W1-209-JIT.md                                           10.0KB    3w ago\n  - W1-209-CENSUS.md                                        12.1KB    3w ago\n  - W3-67-JIT.md                                            4.8KB     3w ago\n  - W3-196-JIT.md                                           9.9KB     3w ago\n  - W3-186-JIT.md                                           5.3KB     3w ago\n  - W3-172-JIT.md                                           6.7KB     3w ago\n  - PRD.md                                                  31.6KB    3w ago\n  - W5-211-DESIGN.md                                        7.2KB     3w ago\n  - W3-196-CENSUS.md                                        9.8KB     3w ago\n  - CONCIERGE-183-JIT.md                                    57.6KB    3w ago\n  - KEYSTONE-182-JIT.md                                     103.7KB   3w ago\n  - rust-toolchain.toml                                     33B       3w ago\n  - W1-TRANCHE2-JIT.md                                      5.2KB     1mo ago\n  - U3-VERB-SURFACE-JIT.md                                  13.5KB    1mo ago\n  - .gitattributes                                          2.2KB     1mo ago\n  - .github/                                                          1mo ago\n    - ci/                                                             10h ago\n    - workflows/                                                      11h ago\n    - bench/                                                          14h ago\n  - installer/                                                        1mo ago\n    - install.ps1                                           12.5KB    1mo ago\n    - install.sh                                            8.2KB     2mo ago\n  - INBOUND-FIREWALL-JIT.md                                 6.2KB     1mo ago\n  - PHASE-B-RESULTS.md                                      14.8KB    1mo ago\n  - PHASE-E-RULINGS.md                                      4.9KB     1mo ago\n  - CLAUDE.md                                               293B      1mo ago\n  - RESIDENT-SERVICE-JIT.md                                 5.3KB     1mo ago\n  - IDLE-EDGE-JIT.md                                        13.0KB    1mo ago\n  - W2-JIT.md                                               23.3KB    1mo ago\n  - W3-BISECT-JIT.md                                        4.2KB     1mo ago\n  - tests/                                                            2mo ago\n    - .gitkeep                                              23B       3mo ago\n  - README.md                                               2.7KB     2mo ago\n  - V0.19.0-P6-SCREEN-GRID-DESIGN.md                        12.0KB    2mo ago\n  - V0.17.0-ROBUST-JOIN-AND-PRESENCE-JIT.md                 9.5KB     2mo ago\n  - V0.16.0-UPDATE-ARC-AND-CLI-JIT.md                       7.0KB     2mo ago\n  - V0.15.0-PLAN.md                                         21.0KB    2mo ago\n  - ROADMAP.md                                              29.4KB    2mo ago\n  - F-019-TRANSLATION-BINARY-INSTALL-DIR-RESOLVE-DESIGN.md  6.0KB     2mo ago\n  - V0.14.1-HOME-SUBNET-PICKER-LAYER-DESIGN.md              7.4KB     2mo ago\n  - F-018-ADAPTER-ADD-NONDESTRUCTIVE-DESIGN.md              5.1KB     2mo ago\n  - V0.14.0-ENDPOINT-CREATION-FLOW-JIT.md                   10.4KB    2mo ago\n  - W3-DESIGN-PROPOSAL.md                                   16.5KB    2mo ago\n  - W3-GATE-VERDICT.md                                      6.9KB     2mo ago\n  - W3-DESIGN-GATE.md                                       8.1KB     2mo ago\n  - V0.13.2-ADAPTER-PACKAGING-JIT.md                        8.0KB     2mo ago\n  - V0.13.0-W5-DRIVEN-BY-SELF-HEAL-PLAN.md                  4.8KB     2mo ago\n  - V0.13.0-W4-PICKER-UX-PLAN.md                            4.6KB     2mo ago\n  - V0.13.0-VIEWER-RING-ROLL-SNAP-JIT.md                    5.2KB     2mo ago\n  - V0.13.0-VIEWER-DRAIN-DECOUPLE-JIT.md                    12.5KB    2mo ago\n  - V0.13.0-SKIP-TO-LIVE-JIT.md                             8.3KB     2mo ago\n  - V0.13.0-RC-VT-TRANSLATE-DESIGN.md                       6.4KB     2mo ago\n  - V0.13.0-P2-SESSION-RESUME-DESIGN.md                     6.7KB     2mo ago\n  - V0.13.0-P1C-CONTROLLER-WRITER-REORDER-PLAN.md           8.3KB     2mo ago\n  - V0.13.0-P1-RC-PASTE-DESIGN.md                           5.0KB     2mo ago\n  - V0.13.0-P0-PTY-INPUT-WRITER-DESIGN.md                   7.2KB     2mo ago\n  - V0.13.0-P0-JIT-PLAN.md                                  6.9KB     2mo ago\n  - V0.13.0-DELIVERY-CONTROL-JIT.md                         16.6KB    2mo ago\n  - V0.13.0-BUG2-RC-VT-TRANSLATE-PLAN.md                    5.7KB     2mo ago\n  - V0.12.1-LIFECYCLE-JIT.md                                17.7KB    2mo ago\n  - V074-PLAN.md                                            14.8KB    3mo ago\n  - M11-PLAN.md                                             20.7KB    3mo ago\n  - DAEMON-LIFECYCLE-COMPLETION-PLAN.md                     8.5KB     3mo ago\n  - REQ-MSG-ENVELOPE-PLAN.md                                4.8KB     3mo ago\n  - M12-W4-RULING.md                                        5.8KB     3mo ago\n  - M12-W5-PLAN.md                                          6.9KB     3mo ago\n  - M12-W4-PLAN.md                                          14.6KB    3mo ago\n  - M12-W3-PLAN.md                                          3.7KB     3mo ago\n  - M12-W3-RULING.md                                        3.2KB     3mo ago\n  - M12-W2-RULING.md                                        8.8KB     3mo ago\n  - M12-W2.5-PLAN.md                                        23.8KB    3mo ago\n  - M12-W2.5-RULING.md                                      7.2KB     3mo ago\n  - M12-W2-PLAN.md                                          16.7KB    3mo ago\n  - M12-W1-PLAN.md                                          13.9KB    3mo ago\n  - M12-W1.5-PLAN.md                                        7.5KB     3mo ago\n  - M12-PLAN.md                                             5.2KB     3mo ago\n  - M12-ENDPOINT-RUN-PICKER.md                              5.5KB     3mo ago\n  - M12-CANDIDATES.md                                       5.2KB     3mo ago\n  - DIGEST-MILESTONE-PLAN.md                                8.9KB     3mo ago\n  - adapters/                                                         3mo ago\n    - mock/                                                           1mo ago\n    - mock-gateway/                                                   2mo ago\n  - M9-WAVE3-JIT.md                                         16.7KB    3mo ago\n  - M9-WAVE2-ROLE-WRITERS.md                                7.7KB     3mo ago\n  - M9-WAVE1-ADDRESSING-SURVEY.md                           5.7KB     3mo ago\n  - M9-GATEWAY-WAN-PLAN.md                                  5.8KB     3mo ago\n  - M9-PLAN.md                                              11.8KB    3mo ago\n  - V042-PLAN.md                                            9.1KB     3mo ago\n  - PUMP-SEAM-PLAN.md                                       14.5KB    3mo ago\n  - RESTORATION-D7-PLAN.md                                  35.7KB    3mo ago\n  - RESTORATION-D6-PLAN.md                                  38.7KB    3mo ago\n  - RESTORATION-D5-PLAN.md                                  22.0KB    3mo ago\n  - RESTORATION-D4-PLAN.md                                  18.7KB    3mo ago\n  - RESTORATION-PLAN.md                                     23.7KB    3mo ago\n  - RESTORATION-D3-PLAN.md                                  12.8KB    3mo ago\n  - RESTORATION-D2-PLAN.md                                  6.2KB     3mo ago\n  - RESTORATION-D1-PLAN.md                                  4.4KB     3mo ago\n  - V032-PLAN.md                                            6.3KB     3mo ago\n  - V031-PLAN.md                                            4.7KB     3mo ago\n  - releases-repo/                                                    3mo ago\n    - SYNC.md                                               512B      3mo ago\n    - LICENSE-MIT                                           1.3KB     3mo ago\n    - README.md                                             4.8KB     3mo ago\n    - LICENSE-BINARY                                        1.8KB     3mo ago\n  - SUBNET-MESH-PLAN.md                                     11.8KB    3mo ago\n  - MESH-D7-PLAN.md                                         9.0KB     3mo ago\n  - MESH-D8-PLAN.md                                         6.6KB     3mo ago\n  - MESH-D4-PLAN.md                                         11.6KB    3mo ago\n  - MESH-D5-PLAN.md                                         12.2KB    3mo ago\n  - MESH-D6-PLAN.md                                         7.7KB     3mo ago\n  - MESH-D2-PLAN.md                                         11.6KB    3mo ago\n  - MESH-D3-PLAN.md                                         9.9KB     3mo ago\n  - M8-PLAN.md                                              20.1KB    3mo ago\n  - MESH-D0-PLAN.md                                         5.8KB     3mo ago\n  - MESH-D1-PLAN.md                                         6.3KB     3mo ago\n  - M8-ACCEPTANCE-STEPS.md                                  17.3KB    3mo ago\n  - M8-ACCEPTANCE.md                                        9.0KB     3mo ago\n  - M7-PLAN.md                                              9.0KB     3mo ago\n  - M5-PLAN.md                                              21.4KB    3mo ago\n  - M6-PLAN.md                                              10.3KB    3mo ago\n  - M5-D9-PLAN.md                                           7.1KB     3mo ago\n  - M5-D6-PLAN.md                                           8.7KB     3mo ago\n  - M5-D7-PLAN.md                                           8.5KB     3mo ago\n  - M5-D8-PLAN.md                                           9.6KB     3mo ago\n  - M5-D5-PLAN.md                                           7.2KB     3mo ago\n  - M5-D3-PLAN.md                                           19.1KB    3mo ago\n  - M5-D4-PLAN.md                                           14.3KB    3mo ago\n  - M4-D9-PLAN.md                                           20.4KB    3mo ago\n  - M4-PLAN.md                                              19.8KB    3mo ago\n  - M4-D8-PLAN.md                                           14.5KB    3mo ago\n  - M4-D7-PLAN.md                                           8.0KB     3mo ago\n  - M4-D7.5-PLAN.md                                         9.0KB     3mo ago\n  - M4-D5-PLAN.md                                           9.7KB     3mo ago\n  - M4-D6-PLAN.md                                           8.6KB     3mo ago\n  - M4-D3-PLAN.md                                           6.1KB     3mo ago\n  - M4-D4-PLAN.md                                           5.9KB     3mo ago\n  - M4-D2f-PLAN.md                                          4.1KB     3mo ago\n  - M4-D2g-PLAN.md                                          4.6KB     3mo ago\n  - M3c-PLAN.md                                             9.1KB     3mo ago\n  - M4-D2-WIRE-PLAN.md                                      7.6KB     3mo ago\n  - M3b-PLAN.md                                             15.8KB    3mo ago\n  - M3-PLAN.md                                              20.0KB    3mo ago\n  - M3a-PLAN.md                                             11.4KB    3mo ago\n  - M2a-PLAN.md                                             12.9KB    3mo ago\n  - M2b-PLAN.md                                             17.3KB    3mo ago\n  - M0-PLAN.md                                              7.9KB     3mo ago\n  - M1-PLAN.md                                              9.6KB     3mo ago\n  - DAEMON-LIFECYCLE-PLAN.md                                6.9KB     3mo ago\n  - DAEMON-LIFECYCLE-VERIFICATION.md                        3.0KB     3mo ago\n  - CI-SELFHOST-PLAN.md                                     3.8KB     3mo ago\n  - .gstack/                                                          3mo ago\n    - browse-console.log                                    230B      3mo ago\n    - browse-network.log                                    2.1KB     3mo ago\n  - src/                                                              3mo ago\n    - .gitkeep                                              51B       3mo ago\n\n[1 results limit reached. Use limit=2 for more]"}],"details":{"isDirectory":true,"resolvedPath":"C:\\Users\\decid\\Documents\\projects\\spt-core","meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core"},"limits":{"resultLimit":{"reached":1,"suggestion":2}}}},"isError":false,"timestamp":1789523970735}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_66S3H6COdwzCBA1KqrQdAXOM|fc_0d104aba2b0f3edc016aa9f80df55487d18d87f28b9dec0cba","toolName":"grep","startedAt":"2026-09-16T01:59:42.028Z","args":{"path":"docs;traceable-reqs.toml"},"intent":"Locating prior hazard entries"},"id":"dfd95be2","parentId":"6180c264","timestamp":"2026-09-16T01:59:42.028Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_qe3QX3p3qMbmC89iLOnFtiiq|fc_0d104aba2b0f3edc016aa9f80df56c87d19a6128fb5be870e8","toolName":"read","startedAt":"2026-09-16T01:59:42.028Z","args":{"path":"docs/49-267-REFRESH-OBSERVATION-PLAN.md:1-210"},"intent":"Reading refresh observation contract"},"id":"0181f4df","parentId":"dfd95be2","timestamp":"2026-09-16T01:59:42.028Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_Eji3SEClOJPHtOgxSDMroTsB|fc_0d104aba2b0f3edc016aa9f80df57487d1afbaa6c96dde46d7","toolName":"read","startedAt":"2026-09-16T01:59:42.028Z","args":{"path":"crates/spt-daemon/src/conn.rs:609-712"},"intent":"Reading bounded serialized writes"},"id":"b60fe692","parentId":"0181f4df","timestamp":"2026-09-16T01:59:42.028Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_I65xikavuPiWDjPyL4lQC8nG|fc_0d104aba2b0f3edc016aa9f80df57c87d19869b6fe4ad64e3e","toolName":"grep","startedAt":"2026-09-16T01:59:42.028Z","args":{"path":"crates/spt-daemon/src/broker.rs"},"intent":"Locating retirement and cleanup callsites"},"id":"2d3247f6","parentId":"b60fe692","timestamp":"2026-09-16T01:59:42.028Z"}
{"type":"message","id":"087f7f07","parentId":"2d3247f6","timestamp":"2026-09-16T01:59:42.028Z","message":{"role":"assistant","content":[{"type":"toolCall","id":"call_66S3H6COdwzCBA1KqrQdAXOM|fc_0d104aba2b0f3edc016aa9f80df55487d18d87f28b9dec0cba","name":"grep","arguments":{"i":"Locating prior hazard entries","pattern":"SHAREDSEND-NO-BLOCKING|RESUME-NO-CONN-DEADLOCK|TRIAL-DRAIN-DRIVE|PROMOTE-DRAINED|F-039","path":"docs;traceable-reqs.toml","case":true,"gitignore":true,"skip":null},"intent":"Locating prior hazard entries"},{"type":"toolCall","id":"call_qe3QX3p3qMbmC89iLOnFtiiq|fc_0d104aba2b0f3edc016aa9f80df56c87d19a6128fb5be870e8","name":"read","arguments":{"i":"Reading refresh observation contract","path":"docs/49-267-REFRESH-OBSERVATION-PLAN.md:1-210"},"intent":"Reading refresh observation contract"},{"type":"toolCall","id":"call_Eji3SEClOJPHtOgxSDMroTsB|fc_0d104aba2b0f3edc016aa9f80df57487d1afbaa6c96dde46d7","name":"read","arguments":{"i":"Reading bounded serialized writes","path":"crates/spt-daemon/src/conn.rs:609-712"},"intent":"Reading bounded serialized writes"},{"type":"toolCall","id":"call_I65xikavuPiWDjPyL4lQC8nG|fc_0d104aba2b0f3edc016aa9f80df57c87d19869b6fe4ad64e3e","name":"grep","arguments":{"i":"Locating retirement and cleanup callsites","pattern":"fn note_controller_write_retired|detach_if\\(|stream-sub-writer-poison","path":"crates/spt-daemon/src/broker.rs","case":true,"gitignore":false,"skip":null},"intent":"Locating retirement and cleanup callsites"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":12664,"output":240,"cacheRead":40576,"cacheWrite":0,"totalTokens":53480,"cost":{"input":0.12664,"output":0.012,"cacheRead":0.040576,"cacheWrite":0,"total":0.17921600000000001}},"stopReason":"toolUse","timestamp":1789523970745,"responseId":"resp_0d104aba2b0f3edc016aa9f803152c87d1b191e13647820d2e","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"fc_0d104aba2b0f3edc016aa9f80df55487d18d87f28b9dec0cba","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating prior hazard entries\",\"pattern\":\"SHAREDSEND-NO-BLOCKING|RESUME-NO-CONN-DEADLOCK|TRIAL-DRAIN-DRIVE|PROMOTE-DRAINED|F-039\",\"path\":\"docs;traceable-reqs.toml\",\"case\":true,\"gitignore\":true,\"skip\":null}","call_id":"call_66S3H6COdwzCBA1KqrQdAXOM","internal_chat_message_metadata_passthrough":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"},"name":"grep","metadata":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"}},{"id":"fc_0d104aba2b0f3edc016aa9f80df56c87d19a6128fb5be870e8","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading refresh observation contract\",\"path\":\"docs/49-267-REFRESH-OBSERVATION-PLAN.md:1-210\"}","call_id":"call_qe3QX3p3qMbmC89iLOnFtiiq","internal_chat_message_metadata_passthrough":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"},"name":"read","metadata":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"}},{"id":"fc_0d104aba2b0f3edc016aa9f80df57487d1afbaa6c96dde46d7","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading bounded serialized writes\",\"path\":\"crates/spt-daemon/src/conn.rs:609-712\"}","call_id":"call_Eji3SEClOJPHtOgxSDMroTsB","internal_chat_message_metadata_passthrough":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"},"name":"read","metadata":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"}},{"id":"fc_0d104aba2b0f3edc016aa9f80df57c87d19869b6fe4ad64e3e","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating retirement and cleanup callsites\",\"pattern\":\"fn note_controller_write_retired|detach_if\\\\(|stream-sub-writer-poison\",\"path\":\"crates/spt-daemon/src/broker.rs\",\"case\":true,\"gitignore\":false,\"skip\":null}","call_id":"call_I65xikavuPiWDjPyL4lQC8nG","internal_chat_message_metadata_passthrough":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"},"name":"grep","metadata":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"}}]},"duration":11281.295300000114,"ttft":11243.290300000226,"completedAt":1789523982026,"contextSnapshot":{"promptTokens":53240,"nonMessageTokens":10480,"compactionEpoch":0}}}
{"type":"message","id":"a53eb4b4","parentId":"087f7f07","timestamp":"2026-09-16T01:59:42.040Z","message":{"role":"toolResult","toolCallId":"call_qe3QX3p3qMbmC89iLOnFtiiq|fc_0d104aba2b0f3edc016aa9f80df56c87d19a6128fb5be870e8","toolName":"read","content":[{"type":"text","text":"# #49 / #267 — bounded refresh observation plan\n\n**Status:** prepared, **not admitted**. No execution grant. W2 run `20260912T231404Z` holds the\nexecution window. Author: todlando. Gater: doyle.\n\n**What this observes:** whether a blocked controller sink reaches eviction across one planned brain\nrestart, and whether a PTY's output and input progress resume — with session and connection\nattribution. **What it cannot decide:** the clean-close hypothesis. See §8.\n\n---\n\n## 1. Subject selection — bytes, not a path\n\n| item | value |\n| --- | --- |\n| executable path | `C:\\Users\\decid\\AppData\\Local\\spt-core\\bin\\spt` |\n| **executable sha256 (the pin)** | `689503ad9fc6cfde637cbcacf21c42faf130ba465a0df6c356cf290e2adafb85` |\n| reported version | `spt 0.69.0` |\n| mtime at planning | 2026-09-10 21:14 |\n\n**The sha256 is the pin, not the path.** Re-hash immediately before the run. **If the hash differs\nfrom the value above, the run does not start**: the subject must be re-selected and this plan\nre-admitted. A changed binary is a different subject, not a recorded detail.\n\n**Source provenance — UNVERIFIED, and recorded as such.** The installed binary predates current\n`main`, so **current `main` is not its source provenance**. The plausible release commit is v0.69.0 =\n`16df0e41be04ca4614a792f7e70e4d1bdec60b24` (2026-09-10), which is consistent with the binary's\nversion string and mtime — but **I have not verified that these bytes were built from that commit**,\nand no artifact on this box binds them. Every source citation in this plan (`brainproc.rs:1081`,\n`broker.rs:2842`/`:3787`/`:2815`, `applyhost.rs:364`) is read at **current main** and is therefore a\nclaim about the source tree, **not** a proven claim about the executable under test. Treat any\nline-number citation as indicative; if a run's behaviour disagrees with a citation, the citation is\nthe suspect.\n\n## 2. Isolation — binding\n\n**The resident fleet daemon is never refreshed.** At planning time 18 `spt.exe` processes authenticate\nto the resident daemon (pid 4732): agent perches under `spt-mobile-host.exe`, `omp-spt` chains, and\nthis session's own perch.\n\n- Isolated `SPT_HOME` under a run-specific directory, with its own broker socket.\n- The isolated home **must not exist** beforehand. If it does, the run does not start and the\n  existing directory is left untouched.\n- One disposable PTY session, created for this run only.\n- If the isolated daemon cannot be brought up cleanly within the §3 startup bound, **the run does not\n  start**.\n\n### 2.1 Exact targeting — every command, with identity checks before the two dangerous ones\n\n**Every** command in the run carries the isolated home explicitly: `SPT_HOME=<ISO_HOME>` is set on\neach invocation. No command relies on an inherited or ambient value, and no command is issued from a\nshell whose environment was mutated once and trusted thereafter.\n\n**Before `daemon refresh` and before `daemon stop`, all four checks must pass. Any failure and the\ncommand does not run:**\n\n1. `SPT_HOME` resolves to `<ISO_HOME>` — read back from the command's own environment, not assumed.\n2. The broker socket the command will address is the isolated one, not the fleet's.\n3. The target daemon pid was **created by this run**, established by ancestry.\n4. The target pid is **not 4732**, and is not an ancestor of this session.\n\nRefresh and stop are the only two verbs that act on a daemon's lifecycle; both are gated by all four.\n\n## 3. Bounds — numeric, chosen now\n\n**Two separate budgets, and a total.** The observation deadline **ends experimental activity and\nstarts cleanup — it never skips cleanup.** Cleanup holds its own maximum on top.\n\n| budget | value |\n| --- | --- |\n| **observation budget** (all experimental activity, from run start) | **300 s** |\n| **cleanup budget** (separate, begins when observation ends for any reason) | **120 s** |\n| **total bound, run start to done** | **420 s** |\n\n| phase | nominal cap | on breach |\n| --- | --- | --- |\n| isolated daemon up + session attached | 60 s | run does not start |\n| `t0` appears after refresh is issued | 60 s | record \"restart marker not observed\"; continue within the observation budget |\n| observation after `t0` | 120 s | stop observing; report recovery-after-interval or no-recovery-within-the-observed-interval |\n| reattach: detach, then fresh attach | 30 s to attach, 60 s observed | report as its own arm |\n| cleanup confirmation | 120 s (its own budget) | preserve everything and report unconfirmed (§6) |\n\n**Every observation and reattach phase is additionally capped by the REMAINING observation budget.**\nA phase's nominal cap is an upper bound, not an entitlement: if only 40 s of the 300 s remain, that\nphase gets 40 s. **A phase cut short by the remaining budget is REPORTED AS TRUNCATED**, naming the\nphase and the time it actually got — a truncated phase's result is never presented as a completed\nobservation, and its absence of an event falls under §8's \"not observed\" rule.\n\n**When the observation budget is exhausted, experimental activity stops and cleanup begins\nimmediately** under §6, with its own 120 s. Exhausting the observation budget is a normal ending, not\nan error, and it never licenses skipping or shortening cleanup.\n\n**The observation budget is independent of `t0` because `t0` may never arrive.** Every phase cap is\nrelative to `t0` or to a local step; without a run-start-relative budget, a refresh that produces no\n`BRAIN_UPDATE_RESTART` would leave the run with no defined end. If `t0` never appears, that absence\nis a recorded result — and per §8 it is \"not observed\", not proof the restart did not happen.\n\n## 4. Observation\n\n`t0` = `BRAIN_UPDATE_RESTART` (brainproc.rs:1081, per §1's provenance caveat). Daemon stderr captured\nto a file for the whole window, starting before the session is attached.\n\n| id | question | observable | attribution |\n| --- | --- | --- | --- |\n| Q2 | does a blocked sink reach eviction? | `BRAIN_SUBSCRIBER_STALL_EVICT` (broker.rs:2842, :3787) | carries `conn=` and a session id |\n| Q4 | what survives reattach? | §5, after one detach and one fresh attach | session id |\n| Q5 | does output progress resume, and when? | §5 | session id |\n| Q6 | is queued input preserved and applied in order? | §5 | session id |\n\n**Q2 is not a deadline measurement.** `evict − t0` runs from the restart signal; the sink deadline\nruns from **blocked-write onset**, which nothing emits.\n\n### 4.1 Two observables added for Q1 and Q3, and exactly what each establishes\n\n<!-- [doc->REQ-REFRESH-WAIT-ATTRIBUTION] -->\n\nQ1 and Q3 were withdrawn as UNRESOLVED (§8) because both live on **silent branches**, not because the\nplan lacked a token to look for. Both branches now emit, under the existing `SPT_PUMP_TRACE=1`\ndiagnostic gate and nothing else — the daemon reads its own environment and still exports nothing, so\nhosted harness subtrees inherit no markers. Neither line changes a predicate, a deadline, an outcome,\nthe wire, or one byte of normal output.\n\n| id | event | fires when | attribution |\n| --- | --- | --- | --- |\n| Q1 | `CONTROLLER_SLOT_CLOSED` | the epoch-guarded transition **actually empties** the slot | `conn=`, `cause=`, session id |\n| Q3 | `BRAIN_PROMOTE_GATE_NOT_DRAINED` | the FIRST not-drained read after readiness latches, once per trial | trial `generation` |\n\n**`cause=` is the caller's knowledge, never the site's inference.** `clear_controller` is shared by\nthree edges — a drain that timed out the controller's `send_timeout` (`send_deadline`), the wedge\nevict (`stall_evict`), and an ordinary detach (`detach`) — and at that site they are\nindistinguishable. `send_deadline` names the deadline the drain observed, NOT the writer's fate:\na writer blocked past its deadline has not been shown to have exited. Each caller passes what it knows.\nA reader may therefore trust the cause exactly as far as the calling path, and no further.\n\n**A guarded no-op emits nothing.** `mark_controller_gone` returns without taking when the epoch has\nmoved on; no line is written there. So the event means *a slot that was occupied is now empty*, and\nnever *something tried to close a slot*. That distinction is the whole of Q1: before this, a cleanly\nclosed old-generation conn and a slot that was already empty left identical stderr.\n\n**What `BRAIN_PROMOTE_GATE_NOT_DRAINED` does NOT establish, stated because the wrong reading is the\ntempting one:** it proves the DRAINED gate was **evaluated** and read not-drained. It is **not**\nevidence that the heartbeat's drain-driving `brain.sessions()` call completed — that call is what\nmakes a wedge drainable at all (REQ-UPDATE-TRIAL-DRAIN-DRIVE), and a trial that never drove it looks\nidentical here to one that drove it and still waited. It also names **no session**: the gate answers a\nsingle bool over every hosted session, and naming the wedged one would require a second, separately\ntaken scan that could disagree with the read the promote decision actually used. The narrower fact is\nreported instead.\n\n**Absence of either line remains \"not observed\", never established absence** — the §8 rule applies to\nthese two exactly as it applies to every other token here, and it applies twice over when the gate\nenv var was not set on the daemon under test.\n\n## 5. Output and input progress — two clocks, and what each can prove\n\nThe disposable session runs a generator emitting, on its own cadence:\n\n```\nPROG <n> <generated-at>\n```\n\nand the observer records, for every line, **its own receipt timestamp**. **Both timestamps are kept\nfor every line.** They answer different questions and neither substitutes for the other:\n\n- **generation timestamp + counter** — what the child did, and when the child did it.\n- **receipt timestamp** — when the observer was delivered it.\n\n**The trap:** after a freeze a burst arrives, and buffered output released late is indistinguishable\n**by arrival alone** from progress resuming.\n\n**What each combination establishes, stated exactly:**\n\n- Pre-`t0` counters and pre-`t0` generation stamps, received after `t0` → **buffered output being\n  released.** The transport resumed. It says nothing about the child.\n- Counter above the pre-`t0` maximum **and** generation stamp after `t0` → **the child progressed\n  after `t0`.** This does **not** establish that the child was executing *contemporaneously with\n  delivery*: such a line may itself have been generated and then buffered before reaching the\n  observer. **Arrival timing is evidence of delivery, never proof of contemporaneous child\n  execution.**\n- The **gap between generation and receipt**, per line, is the only figure that speaks to delivery\n  lag, and it is reported per line rather than averaged.\n\nRecord the highest counter and latest generation stamp seen before `t0` explicitly — the baseline the\nwhole question turns on.\n\n### 5.1 Input — three distinct observables, never conflated\n\nDuring the freeze, send inputs each carrying a **unique tag**, never a repeated keystroke. For each\ntag, three things are separately observable and are **recorded separately**:\n\n1. **terminal echo** — the tag appearing in the stream. Proves the byte reached the terminal's echo\n   path. **Proves nothing about the child.**\n2. **the child's tagged acknowledgement** — output the child itself emits naming that tag.\n3. **resulting state** — an effect the child produces that is distinct from both its echo and the\n   tag's text.\n\n\"The queue drained\" may be said only when every tag shows **2 and 3**, in the order sent.\n\n**A missing tag at the observation deadline is \"NOT OBSERVED WITHIN THE WINDOW\" — not established\nloss.** A late drain past 120 s and a genuine loss are indistinguishable to a finite observation, and\nthe plan records the former wording. A duplicated tag is **replay**; out-of-order tags are their own\nfinding. Byte counts and \"output resumed\" establish none of this.\n\n## 6. Cleanup — retention, not deletion\n\n- Stop the isolated daemon **under §2.1's four checks**; confirm its brain and broker are gone by pid\n  and by socket, and that every process the run created is gone **by ancestry**.\n- **Retain the isolated home and every captured artifact after shutdown** — stderr capture and PTY\n  transcript included. No directory deletion is required by this observation.\n- **If cleanup is uncertain, preserve — never remove.** An uncertain teardown leaves process\n  identities and artifacts intact for inspection and is reported as unconfirmed.\n- The captured log is copied out before anything else is considered.\n\n## 7. What the source says, and what that is worth\n…\n\n…\n9. **#302 is not merged into this.** Whether it relates is untested and deliberately left open.\n\n[Showing lines 1-213 of 246. Use :214 to continue]"}],"details":{"truncation":{"content":"# #49 / #267 — bounded refresh observation plan\r\n\r\n**Status:** prepared, **not admitted**. No execution grant. W2 run `20260912T231404Z` holds the\r\nexecution window. Author: todlando. Gater: doyle.\r\n\r\n**What this observes:** whether a blocked controller sink reaches eviction across one planned brain\r\nrestart, and whether a PTY's output and input progress resume — with session and connection\r\nattribution. **What it cannot decide:** the clean-close hypothesis. See §8.\r\n\r\n---\r\n\r\n## 1. Subject selection — bytes, not a path\r\n\r\n| item | value |\r\n| --- | --- |\r\n| executable path | `C:\\Users\\decid\\AppData\\Local\\spt-core\\bin\\spt` |\r\n| **executable sha256 (the pin)** | `689503ad9fc6cfde637cbcacf21c42faf130ba465a0df6c356cf290e2adafb85` |\r\n| reported version | `spt 0.69.0` |\r\n| mtime at planning | 2026-09-10 21:14 |\r\n\r\n**The sha256 is the pin, not the path.** Re-hash immediately before the run. **If the hash differs\r\nfrom the value above, the run does not start**: the subject must be re-selected and this plan\r\nre-admitted. A changed binary is a different subject, not a recorded detail.\r\n\r\n**Source provenance — UNVERIFIED, and recorded as such.** The installed binary predates current\r\n`main`, so **current `main` is not its source provenance**. The plausible release commit is v0.69.0 =\r\n`16df0e41be04ca4614a792f7e70e4d1bdec60b24` (2026-09-10), which is consistent with the binary's\r\nversion string and mtime — but **I have not verified that these bytes were built from that commit**,\r\nand no artifact on this box binds them. Every source citation in this plan (`brainproc.rs:1081`,\r\n`broker.rs:2842`/`:3787`/`:2815`, `applyhost.rs:364`) is read at **current main** and is therefore a\r\nclaim about the source tree, **not** a proven claim about the executable under test. Treat any\r\nline-number citation as indicative; if a run's behaviour disagrees with a citation, the citation is\r\nthe suspect.\r\n\r\n## 2. Isolation — binding\r\n\r\n**The resident fleet daemon is never refreshed.** At planning time 18 `spt.exe` processes authenticate\r\nto the resident daemon (pid 4732): agent perches under `spt-mobile-host.exe`, `omp-spt` chains, and\r\nthis session's own perch.\r\n\r\n- Isolated `SPT_HOME` under a run-specific directory, with its own broker socket.\r\n- The isolated home **must not exist** beforehand. If it does, the run does not start and the\r\n  existing directory is left untouched.\r\n- One disposable PTY session, created for this run only.\r\n- If the isolated daemon cannot be brought up cleanly within the §3 startup bound, **the run does not\r\n  start**.\r\n\r\n### 2.1 Exact targeting — every command, with identity checks before the two dangerous ones\r\n\r\n**Every** command in the run carries the isolated home explicitly: `SPT_HOME=<ISO_HOME>` is set on\r\neach invocation. No command relies on an inherited or ambient value, and no command is issued from a\r\nshell whose environment was mutated once and trusted thereafter.\r\n\r\n**Before `daemon refresh` and before `daemon stop`, all four checks must pass. Any failure and the\r\ncommand does not run:**\r\n\r\n1. `SPT_HOME` resolves to `<ISO_HOME>` — read back from the command's own environment, not assumed.\r\n2. The broker socket the command will address is the isolated one, not the fleet's.\r\n3. The target daemon pid was **created by this run**, established by ancestry.\r\n4. The target pid is **not 4732**, and is not an ancestor of this session.\r\n\r\nRefresh and stop are the only two verbs that act on a daemon's lifecycle; both are gated by all four.\r\n\r\n## 3. Bounds — numeric, chosen now\r\n\r\n**Two separate budgets, and a total.** The observation deadline **ends experimental activity and\r\nstarts cleanup — it never skips cleanup.** Cleanup holds its own maximum on top.\r\n\r\n| budget | value |\r\n| --- | --- |\r\n| **observation budget** (all experimental activity, from run start) | **300 s** |\r\n| **cleanup budget** (separate, begins when observation ends for any reason) | **120 s** |\r\n| **total bound, run start to done** | **420 s** |\r\n\r\n| phase | nominal cap | on breach |\r\n| --- | --- | --- |\r\n| isolated daemon up + session attached | 60 s | run does not start |\r\n| `t0` appears after refresh is issued | 60 s | record \"restart marker not observed\"; continue within the observation budget |\r\n| observation after `t0` | 120 s | stop observing; report recovery-after-interval or no-recovery-within-the-observed-interval |\r\n| reattach: detach, then fresh attach | 30 s to attach, 60 s observed | report as its own arm |\r\n| cleanup confirmation | 120 s (its own budget) | preserve everything and report unconfirmed (§6) |\r\n\r\n**Every observation and reattach phase is additionally capped by the REMAINING observation budget.**\r\nA phase's nominal cap is an upper bound, not an entitlement: if only 40 s of the 300 s remain, that\r\nphase gets 40 s. **A phase cut short by the remaining budget is REPORTED AS TRUNCATED**, naming the\r\nphase and the time it actually got — a truncated phase's result is never presented as a completed\r\nobservation, and its absence of an event falls under §8's \"not observed\" rule.\r\n\r\n**When the observation budget is exhausted, experimental activity stops and cleanup begins\r\nimmediately** under §6, with its own 120 s. Exhausting the observation budget is a normal ending, not\r\nan error, and it never licenses skipping or shortening cleanup.\r\n\r\n**The observation budget is independent of `t0` because `t0` may never arrive.** Every phase cap is\r\nrelative to `t0` or to a local step; without a run-start-relative budget, a refresh that produces no\r\n`BRAIN_UPDATE_RESTART` would leave the run with no defined end. If `t0` never appears, that absence\r\nis a recorded result — and per §8 it is \"not observed\", not proof the restart did not happen.\r\n\r\n## 4. Observation\r\n\r\n`t0` = `BRAIN_UPDATE_RESTART` (brainproc.rs:1081, per §1's provenance caveat). Daemon stderr captured\r\nto a file for the whole window, starting before the session is attached.\r\n\r\n| id | question | observable | attribution |\r\n| --- | --- | --- | --- |\r\n| Q2 | does a blocked sink reach eviction? | `BRAIN_SUBSCRIBER_STALL_EVICT` (broker.rs:2842, :3787) | carries `conn=` and a session id |\r\n| Q4 | what survives reattach? | §5, after one detach and one fresh attach | session id |\r\n| Q5 | does output progress resume, and when? | §5 | session id |\r\n| Q6 | is queued input preserved and applied in order? | §5 | session id |\r\n\r\n**Q2 is not a deadline measurement.** `evict − t0` runs from the restart signal; the sink deadline\r\nruns from **blocked-write onset**, which nothing emits.\r\n\r\n### 4.1 Two observables added for Q1 and Q3, and exactly what each establishes\r\n\r\n<!-- [doc->REQ-REFRESH-WAIT-ATTRIBUTION] -->\r\n\r\nQ1 and Q3 were withdrawn as UNRESOLVED (§8) because both live on **silent branches**, not because the\r\nplan lacked a token to look for. Both branches now emit, under the existing `SPT_PUMP_TRACE=1`\r\ndiagnostic gate and nothing else — the daemon reads its own environment and still exports nothing, so\r\nhosted harness subtrees inherit no markers. Neither line changes a predicate, a deadline, an outcome,\r\nthe wire, or one byte of normal output.\r\n\r\n| id | event | fires when | attribution |\r\n| --- | --- | --- | --- |\r\n| Q1 | `CONTROLLER_SLOT_CLOSED` | the epoch-guarded transition **actually empties** the slot | `conn=`, `cause=`, session id |\r\n| Q3 | `BRAIN_PROMOTE_GATE_NOT_DRAINED` | the FIRST not-drained read after readiness latches, once per trial | trial `generation` |\r\n\r\n**`cause=` is the caller's knowledge, never the site's inference.** `clear_controller` is shared by\r\nthree edges — a drain that timed out the controller's `send_timeout` (`send_deadline`), the wedge\r\nevict (`stall_evict`), and an ordinary detach (`detach`) — and at that site they are\r\nindistinguishable. `send_deadline` names the deadline the drain observed, NOT the writer's fate:\r\na writer blocked past its deadline has not been shown to have exited. Each caller passes what it knows.\r\nA reader may therefore trust the cause exactly as far as the calling path, and no further.\r\n\r\n**A guarded no-op emits nothing.** `mark_controller_gone` returns without taking when the epoch has\r\nmoved on; no line is written there. So the event means *a slot that was occupied is now empty*, and\r\nnever *something tried to close a slot*. That distinction is the whole of Q1: before this, a cleanly\r\nclosed old-generation conn and a slot that was already empty left identical stderr.\r\n\r\n**What `BRAIN_PROMOTE_GATE_NOT_DRAINED` does NOT establish, stated because the wrong reading is the\r\ntempting one:** it proves the DRAINED gate was **evaluated** and read not-drained. It is **not**\r\nevidence that the heartbeat's drain-driving `brain.sessions()` call completed — that call is what\r\nmakes a wedge drainable at all (REQ-UPDATE-TRIAL-DRAIN-DRIVE), and a trial that never drove it looks\r\nidentical here to one that drove it and still waited. It also names **no session**: the gate answers a\r\nsingle bool over every hosted session, and naming the wedged one would require a second, separately\r\ntaken scan that could disagree with the read the promote decision actually used. The narrower fact is\r\nreported instead.\r\n\r\n**Absence of either line remains \"not observed\", never established absence** — the §8 rule applies to\r\nthese two exactly as it applies to every other token here, and it applies twice over when the gate\r\nenv var was not set on the daemon under test.\r\n\r\n## 5. Output and input progress — two clocks, and what each can prove\r\n\r\nThe disposable session runs a generator emitting, on its own cadence:\r\n\r\n```\r\nPROG <n> <generated-at>\r\n```\r\n\r\nand the observer records, for every line, **its own receipt timestamp**. **Both timestamps are kept\r\nfor every line.** They answer different questions and neither substitutes for the other:\r\n\r\n- **generation timestamp + counter** — what the child did, and when the child did it.\r\n- **receipt timestamp** — when the observer was delivered it.\r\n\r\n**The trap:** after a freeze a burst arrives, and buffered output released late is indistinguishable\r\n**by arrival alone** from progress resuming.\r\n\r\n**What each combination establishes, stated exactly:**\r\n\r\n- Pre-`t0` counters and pre-`t0` generation stamps, received after `t0` → **buffered output being\r\n  released.** The transport resumed. It says nothing about the child.\r\n- Counter above the pre-`t0` maximum **and** generation stamp after `t0` → **the child progressed\r\n  after `t0`.** This does **not** establish that the child was executing *contemporaneously with\r\n  delivery*: such a line may itself have been generated and then buffered before reaching the\r\n  observer. **Arrival timing is evidence of delivery, never proof of contemporaneous child\r\n  execution.**\r\n- The **gap between generation and receipt**, per line, is the only figure that speaks to delivery\r\n  lag, and it is reported per line rather than averaged.\r\n\r\nRecord the highest counter and latest generation stamp seen before `t0` explicitly — the baseline the\r\nwhole question turns on.\r\n\r\n### 5.1 Input — three distinct observables, never conflated\r\n\r\nDuring the freeze, send inputs each carrying a **unique tag**, never a repeated keystroke. For each\r\ntag, three things are separately observable and are **recorded separately**:\r\n\r\n1. **terminal echo** — the tag appearing in the stream. Proves the byte reached the terminal's echo\r\n   path. **Proves nothing about the child.**\r\n2. **the child's tagged acknowledgement** — output the child itself emits naming that tag.\r\n3. **resulting state** — an effect the child produces that is distinct from both its echo and the\r\n   tag's text.\r\n\r\n\"The queue drained\" may be said only when every tag shows **2 and 3**, in the order sent.\r\n\r\n**A missing tag at the observation deadline is \"NOT OBSERVED WITHIN THE WINDOW\" — not established\r\nloss.** A late drain past 120 s and a genuine loss are indistinguishable to a finite observation, and\r\nthe plan records the former wording. A duplicated tag is **replay**; out-of-order tags are their own\r\nfinding. Byte counts and \"output resumed\" establish none of this.\r\n\r\n## 6. Cleanup — retention, not deletion\r\n\r\n- Stop the isolated daemon **under §2.1's four checks**; confirm its brain and broker are gone by pid\r\n  and by socket, and that every process the run created is gone **by ancestry**.\r\n- **Retain the isolated home and every captured artifact after shutdown** — stderr capture and PTY\r\n  transcript included. No directory deletion is required by this observation.\r\n- **If cleanup is uncertain, preserve — never remove.** An uncertain teardown leaves process\r\n  identities and artifacts intact for inspection and is reported as unconfirmed.\r\n- The captured log is copied out before anything else is considered.\r\n\r\n## 7. What the source says, and what that is worth\r","truncated":true,"truncatedBy":"lines","totalLines":246,"totalBytes":12840,"outputLines":213,"outputBytes":12840,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":246,"displayContent":{"text":"# #49 / #267 — bounded refresh observation plan\r\n\r\n**Status:** prepared, **not admitted**. No execution grant. W2 run `20260912T231404Z` holds the\r\nexecution window. Author: todlando. Gater: doyle.\r\n\r\n**What this observes:** whether a blocked controller sink reaches eviction across one planned brain\r\nrestart, and whether a PTY's output and input progress resume — with session and connection\r\nattribution. **What it cannot decide:** the clean-close hypothesis. See §8.\r\n\r\n---\r\n\r\n## 1. Subject selection — bytes, not a path\r\n\r\n| item | value |\r\n| --- | --- |\r\n| executable path | `C:\\Users\\decid\\AppData\\Local\\spt-core\\bin\\spt` |\r\n| **executable sha256 (the pin)** | `689503ad9fc6cfde637cbcacf21c42faf130ba465a0df6c356cf290e2adafb85` |\r\n| reported version | `spt 0.69.0` |\r\n| mtime at planning | 2026-09-10 21:14 |\r\n\r\n**The sha256 is the pin, not the path.** Re-hash immediately before the run. **If the hash differs\r\nfrom the value above, the run does not start**: the subject must be re-selected and this plan\r\nre-admitted. A changed binary is a different subject, not a recorded detail.\r\n\r\n**Source provenance — UNVERIFIED, and recorded as such.** The installed binary predates current\r\n`main`, so **current `main` is not its source provenance**. The plausible release commit is v0.69.0 =\r\n`16df0e41be04ca4614a792f7e70e4d1bdec60b24` (2026-09-10), which is consistent with the binary's\r\nversion string and mtime — but **I have not verified that these bytes were built from that commit**,\r\nand no artifact on this box binds them. Every source citation in this plan (`brainproc.rs:1081`,\r\n`broker.rs:2842`/`:3787`/`:2815`, `applyhost.rs:364`) is read at **current main** and is therefore a\r\nclaim about the source tree, **not** a proven claim about the executable under test. Treat any\r\nline-number citation as indicative; if a run's behaviour disagrees with a citation, the citation is\r\nthe suspect.\r\n\r\n## 2. Isolation — binding\r\n\r\n**The resident fleet daemon is never refreshed.** At planning time 18 `spt.exe` processes authenticate\r\nto the resident daemon (pid 4732): agent perches under `spt-mobile-host.exe`, `omp-spt` chains, and\r\nthis session's own perch.\r\n\r\n- Isolated `SPT_HOME` under a run-specific directory, with its own broker socket.\r\n- The isolated home **must not exist** beforehand. If it does, the run does not start and the\r\n  existing directory is left untouched.\r\n- One disposable PTY session, created for this run only.\r\n- If the isolated daemon cannot be brought up cleanly within the §3 startup bound, **the run does not\r\n  start**.\r\n\r\n### 2.1 Exact targeting — every command, with identity checks before the two dangerous ones\r\n\r\n**Every** command in the run carries the isolated home explicitly: `SPT_HOME=<ISO_HOME>` is set on\r\neach invocation. No command relies on an inherited or ambient value, and no command is issued from a\r\nshell whose environment was mutated once and trusted thereafter.\r\n\r\n**Before `daemon refresh` and before `daemon stop`, all four checks must pass. Any failure and the\r\ncommand does not run:**\r\n\r\n1. `SPT_HOME` resolves to `<ISO_HOME>` — read back from the command's own environment, not assumed.\r\n2. The broker socket the command will address is the isolated one, not the fleet's.\r\n3. The target daemon pid was **created by this run**, established by ancestry.\r\n4. The target pid is **not 4732**, and is not an ancestor of this session.\r\n\r\nRefresh and stop are the only two verbs that act on a daemon's lifecycle; both are gated by all four.\r\n\r\n## 3. Bounds — numeric, chosen now\r\n\r\n**Two separate budgets, and a total.** The observation deadline **ends experimental activity and\r\nstarts cleanup — it never skips cleanup.** Cleanup holds its own maximum on top.\r\n\r\n| budget | value |\r\n| --- | --- |\r\n| **observation budget** (all experimental activity, from run start) | **300 s** |\r\n| **cleanup budget** (separate, begins when observation ends for any reason) | **120 s** |\r\n| **total bound, run start to done** | **420 s** |\r\n\r\n| phase | nominal cap | on breach |\r\n| --- | --- | --- |\r\n| isolated daemon up + session attached | 60 s | run does not start |\r\n| `t0` appears after refresh is issued | 60 s | record \"restart marker not observed\"; continue within the observation budget |\r\n| observation after `t0` | 120 s | stop observing; report recovery-after-interval or no-recovery-within-the-observed-interval |\r\n| reattach: detach, then fresh attach | 30 s to attach, 60 s observed | report as its own arm |\r\n| cleanup confirmation | 120 s (its own budget) | preserve everything and report unconfirmed (§6) |\r\n\r\n**Every observation and reattach phase is additionally capped by the REMAINING observation budget.**\r\nA phase's nominal cap is an upper bound, not an entitlement: if only 40 s of the 300 s remain, that\r\nphase gets 40 s. **A phase cut short by the remaining budget is REPORTED AS TRUNCATED**, naming the\r\nphase and the time it actually got — a truncated phase's result is never presented as a completed\r\nobservation, and its absence of an event falls under §8's \"not observed\" rule.\r\n\r\n**When the observation budget is exhausted, experimental activity stops and cleanup begins\r\nimmediately** under §6, with its own 120 s. Exhausting the observation budget is a normal ending, not\r\nan error, and it never licenses skipping or shortening cleanup.\r\n\r\n**The observation budget is independent of `t0` because `t0` may never arrive.** Every phase cap is\r\nrelative to `t0` or to a local step; without a run-start-relative budget, a refresh that produces no\r\n`BRAIN_UPDATE_RESTART` would leave the run with no defined end. If `t0` never appears, that absence\r\nis a recorded result — and per §8 it is \"not observed\", not proof the restart did not happen.\r\n\r\n## 4. Observation\r\n\r\n`t0` = `BRAIN_UPDATE_RESTART` (brainproc.rs:1081, per §1's provenance caveat). Daemon stderr captured\r\nto a file for the whole window, starting before the session is attached.\r\n\r\n| id | question | observable | attribution |\r\n| --- | --- | --- | --- |\r\n| Q2 | does a blocked sink reach eviction? | `BRAIN_SUBSCRIBER_STALL_EVICT` (broker.rs:2842, :3787) | carries `conn=` and a session id |\r\n| Q4 | what survives reattach? | §5, after one detach and one fresh attach | session id |\r\n| Q5 | does output progress resume, and when? | §5 | session id |\r\n| Q6 | is queued input preserved and applied in order? | §5 | session id |\r\n\r\n**Q2 is not a deadline measurement.** `evict − t0` runs from the restart signal; the sink deadline\r\nruns from **blocked-write onset**, which nothing emits.\r\n\r\n### 4.1 Two observables added for Q1 and Q3, and exactly what each establishes\r\n\r\n<!-- [doc->REQ-REFRESH-WAIT-ATTRIBUTION] -->\r\n\r\nQ1 and Q3 were withdrawn as UNRESOLVED (§8) because both live on **silent branches**, not because the\r\nplan lacked a token to look for. Both branches now emit, under the existing `SPT_PUMP_TRACE=1`\r\ndiagnostic gate and nothing else — the daemon reads its own environment and still exports nothing, so\r\nhosted harness subtrees inherit no markers. Neither line changes a predicate, a deadline, an outcome,\r\nthe wire, or one byte of normal output.\r\n\r\n| id | event | fires when | attribution |\r\n| --- | --- | --- | --- |\r\n| Q1 | `CONTROLLER_SLOT_CLOSED` | the epoch-guarded transition **actually empties** the slot | `conn=`, `cause=`, session id |\r\n| Q3 | `BRAIN_PROMOTE_GATE_NOT_DRAINED` | the FIRST not-drained read after readiness latches, once per trial | trial `generation` |\r\n\r\n**`cause=` is the caller's knowledge, never the site's inference.** `clear_controller` is shared by\r\nthree edges — a drain that timed out the controller's `send_timeout` (`send_deadline`), the wedge\r\nevict (`stall_evict`), and an ordinary detach (`detach`) — and at that site they are\r\nindistinguishable. `send_deadline` names the deadline the drain observed, NOT the writer's fate:\r\na writer blocked past its deadline has not been shown to have exited. Each caller passes what it knows.\r\nA reader may therefore trust the cause exactly as far as the calling path, and no further.\r\n\r\n**A guarded no-op emits nothing.** `mark_controller_gone` returns without taking when the epoch has\r\nmoved on; no line is written there. So the event means *a slot that was occupied is now empty*, and\r\nnever *something tried to close a slot*. That distinction is the whole of Q1: before this, a cleanly\r\nclosed old-generation conn and a slot that was already empty left identical stderr.\r\n\r\n**What `BRAIN_PROMOTE_GATE_NOT_DRAINED` does NOT establish, stated because the wrong reading is the\r\ntempting one:** it proves the DRAINED gate was **evaluated** and read not-drained. It is **not**\r\nevidence that the heartbeat's drain-driving `brain.sessions()` call completed — that call is what\r\nmakes a wedge drainable at all (REQ-UPDATE-TRIAL-DRAIN-DRIVE), and a trial that never drove it looks\r\nidentical here to one that drove it and still waited. It also names **no session**: the gate answers a\r\nsingle bool over every hosted session, and naming the wedged one would require a second, separately\r\ntaken scan that could disagree with the read the promote decision actually used. The narrower fact is\r\nreported instead.\r\n\r\n**Absence of either line remains \"not observed\", never established absence** — the §8 rule applies to\r\nthese two exactly as it applies to every other token here, and it applies twice over when the gate\r\nenv var was not set on the daemon under test.\r\n\r\n## 5. Output and input progress — two clocks, and what each can prove\r\n\r\nThe disposable session runs a generator emitting, on its own cadence:\r\n\r\n```\r\nPROG <n> <generated-at>\r\n```\r\n\r\nand the observer records, for every line, **its own receipt timestamp**. **Both timestamps are kept\r\nfor every line.** They answer different questions and neither substitutes for the other:\r\n\r\n- **generation timestamp + counter** — what the child did, and when the child did it.\r\n- **receipt timestamp** — when the observer was delivered it.\r\n\r\n**The trap:** after a freeze a burst arrives, and buffered output released late is indistinguishable\r\n**by arrival alone** from progress resuming.\r\n\r\n**What each combination establishes, stated exactly:**\r\n\r\n- Pre-`t0` counters and pre-`t0` generation stamps, received after `t0` → **buffered output being\r\n  released.** The transport resumed. It says nothing about the child.\r\n- Counter above the pre-`t0` maximum **and** generation stamp after `t0` → **the child progressed\r\n  after `t0`.** This does **not** establish that the child was executing *contemporaneously with\r\n  delivery*: such a line may itself have been generated and then buffered before reaching the\r\n  observer. **Arrival timing is evidence of delivery, never proof of contemporaneous child\r\n  execution.**\r\n- The **gap between generation and receipt**, per line, is the only figure that speaks to delivery\r\n  lag, and it is reported per line rather than averaged.\r\n\r\nRecord the highest counter and latest generation stamp seen before `t0` explicitly — the baseline the\r\nwhole question turns on.\r\n\r\n### 5.1 Input — three distinct observables, never conflated\r\n\r\nDuring the freeze, send inputs each carrying a **unique tag**, never a repeated keystroke. For each\r\ntag, three things are separately observable and are **recorded separately**:\r\n\r\n1. **terminal echo** — the tag appearing in the stream. Proves the byte reached the terminal's echo\r\n   path. **Proves nothing about the child.**\r\n2. **the child's tagged acknowledgement** — output the child itself emits naming that tag.\r\n3. **resulting state** — an effect the child produces that is distinct from both its echo and the\r\n   tag's text.\r\n\r\n\"The queue drained\" may be said only when every tag shows **2 and 3**, in the order sent.\r\n\r\n**A missing tag at the observation deadline is \"NOT OBSERVED WITHIN THE WINDOW\" — not established\r\nloss.** A late drain past 120 s and a genuine loss are indistinguishable to a finite observation, and\r\nthe plan records the former wording. A duplicated tag is **replay**; out-of-order tags are their own\r\nfinding. Byte counts and \"output resumed\" establish none of this.\r\n\r\n## 6. Cleanup — retention, not deletion\r\n\r\n- Stop the isolated daemon **under §2.1's four checks**; confirm its brain and broker are gone by pid\r\n  and by socket, and that every process the run created is gone **by ancestry**.\r\n- **Retain the isolated home and every captured artifact after shutdown** — stderr capture and PTY\r\n  transcript included. No directory deletion is required by this observation.\r\n- **If cleanup is uncertain, preserve — never remove.** An uncertain teardown leaves process\r\n  identities and artifacts intact for inspection and is reported as unconfirmed.\r\n- The captured log is copied out before anything else is considered.\r\n\r\n## 7. What the source says, and what that is worth\r\n…\n\n…\n9. **#302 is not merged into this.** Whether it relates is untested and deliberately left open.","startLine":1,"lineNumbers":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,119,120,121,122,123,124,125,126,127,128,129,130,131,132,133,134,135,136,137,138,139,140,141,142,143,144,145,146,147,148,149,150,151,152,153,154,155,156,157,158,159,160,161,162,163,164,165,166,167,168,169,170,171,172,173,174,175,176,177,178,179,180,181,182,183,184,185,186,187,188,189,190,191,192,193,194,195,196,197,198,199,200,201,202,203,204,205,206,207,208,209,210,211,212,213,null,222,null,246]},"fileSize":15302,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\docs\\49-267-REFRESH-OBSERVATION-PLAN.md"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":246,"totalBytes":12840,"outputLines":213,"outputBytes":12840,"shownRange":{"start":1,"end":213},"nextOffset":214}}},"isError":false,"timestamp":1789523982039}}
{"type":"message","id":"fbe5f62e","parentId":"a53eb4b4","timestamp":"2026-09-16T01:59:42.048Z","message":{"role":"toolResult","toolCallId":"call_Eji3SEClOJPHtOgxSDMroTsB|fc_0d104aba2b0f3edc016aa9f80df57487d1afbaa6c96dde46d7","toolName":"read","content":[{"type":"text","text":"impl BrokerConn {\n…\n    /// `pre` returning false → `Ok(ConnWrite::Superseded)`, nothing written.\n    pub(crate) fn write_gated(\n        &self,\n        env: &Envelope,\n        pre: impl FnOnce() -> bool,\n    ) -> io::Result<ConnWrite> {\n        let inner = &*self.inner;\n        let deadline = Instant::now() + inner.bound;\n\n        // ── Acquire the gate, bounded by the SAME absolute deadline. ──\n        let mut g = recover(&inner.gate);\n        loop {\n            if inner.poisoned.load(Ordering::Acquire) {\n                return Err(poisoned_err());\n            }\n            if g.half.is_some() {\n                break;\n            }\n            let now = Instant::now();\n            if now >= deadline {\n                // The write we queued behind is stuck past OUR whole bound:\n                // that op is the wedge. Abort it out of band and fail — the\n                // watchdog's own deadline (armed at that op's entry) is also\n                // in flight; both paths converge on the same idempotent\n                // poison/cancel.\n                drop(g);\n                inner.timed_out.store(true, Ordering::Release);\n                inner.poison_and_cancel(None);\n                return Err(timeout_err());\n            }\n            let (g2, _) = inner\n                .gate_cv\n                .wait_timeout(g, deadline - now)\n                .unwrap_or_else(|p| p.into_inner());\n            g = g2;\n        }\n\n        if !pre() {\n            return Ok(ConnWrite::Superseded);\n        }\n\n        // The gate serializes first-write initialization. Save the instant\n        // BEFORE I/O so a later failure still names when this conn went active.\n        // [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\n        if inner.first_write.get().is_none() {\n            let _ = inner.first_write.set(LogTime::now());\n            inner.lifecycle_log.healthy(true);\n        }\n\n        // ── Check out the half + arm the watchdog, then write OUT of locks. ──\n        let mut half = g.half.take().expect(\"gate half present at checkout\");\n        let op = inner.op_seq.fetch_add(1, Ordering::Relaxed) + 1;\n        {\n            let mut d = recover(&inner.dog);\n            d.inflight = Some((op, deadline));\n            inner.dog_cv.notify_all();\n        }\n        drop(g);\n\n        let res = write_frame(&mut half, env);\n\n        // Completion handshake: `write_frame` has returned, so the OS op —\n        // canceled or not — reported completion and the frame buffer is\n        // released. Clear the arm BEFORE deciding health, so the watchdog\n        // never severs while the op is formally in flight.\n        {\n            let mut d = recover(&inner.dog);\n            d.inflight = None;\n            inner.dog_cv.notify_all();\n        }\n\n        let mut g = recover(&inner.gate);\n        if res.is_ok() && !inner.poisoned.load(Ordering::Acquire) {\n            g.half = Some(half);\n            drop(g);\n            inner.gate_cv.notify_all();\n            return Ok(ConnWrite::Done);\n        }\n\n        // ── Retire: poison the whole conn, abort the read side so\n        // `handle_conn` reaches EOF cleanup, sever, and park the half in the\n        // graveyard (its handle must outlive any late raw abort; never\n        // re-gated, never reused — a timed-out length-prefixed frame may be\n        // partially written). ──\n        drop(g);\n        inner.poison_and_cancel(res.as_ref().err());\n        inner.sever();\n        *recover(&inner.graveyard) = Some(half);\n        inner.gate_cv.notify_all();\n        // A DEADLINE retirement (the watchdog aborted OUR op, or the conn was\n        // already deadline-poisoned) surfaces as TimedOut — the caller can\n        // count the wedge (W2 stall-evict tally). An organic I/O failure\n        // (peer death) keeps its own error.\n        if inner.timed_out.load(Ordering::Acquire) {\n            return Err(timeout_err());\n        }\n        Err(match res {\n            Err(e) => e,\n            Ok(()) => poisoned_err(),\n        })\n    }\n}\n\n/// TEST-ONLY determinism lever: holds the serialized write gate (the half is\n/// checked out) so every writer parks AT THE GATE until the pin drops — the\n/// round-4 analog of pinning the old `Mutex<SendHalf>` across a registration\n/// window (the P1c epoch-gate keystone unit relies on it).\n#[cfg(test)]\n\n[Showing lines 608-715 of 1000. Use :716 to continue]"}],"details":{"truncation":{"content":"    /// `pre` returning false → `Ok(ConnWrite::Superseded)`, nothing written.\r\n    pub(crate) fn write_gated(\r\n        &self,\r\n        env: &Envelope,\r\n        pre: impl FnOnce() -> bool,\r\n    ) -> io::Result<ConnWrite> {\r\n        let inner = &*self.inner;\r\n        let deadline = Instant::now() + inner.bound;\r\n\r\n        // ── Acquire the gate, bounded by the SAME absolute deadline. ──\r\n        let mut g = recover(&inner.gate);\r\n        loop {\r\n            if inner.poisoned.load(Ordering::Acquire) {\r\n                return Err(poisoned_err());\r\n            }\r\n            if g.half.is_some() {\r\n                break;\r\n            }\r\n            let now = Instant::now();\r\n            if now >= deadline {\r\n                // The write we queued behind is stuck past OUR whole bound:\r\n                // that op is the wedge. Abort it out of band and fail — the\r\n                // watchdog's own deadline (armed at that op's entry) is also\r\n                // in flight; both paths converge on the same idempotent\r\n                // poison/cancel.\r\n                drop(g);\r\n                inner.timed_out.store(true, Ordering::Release);\r\n                inner.poison_and_cancel(None);\r\n                return Err(timeout_err());\r\n            }\r\n            let (g2, _) = inner\r\n                .gate_cv\r\n                .wait_timeout(g, deadline - now)\r\n                .unwrap_or_else(|p| p.into_inner());\r\n            g = g2;\r\n        }\r\n\r\n        if !pre() {\r\n            return Ok(ConnWrite::Superseded);\r\n        }\r\n\r\n        // The gate serializes first-write initialization. Save the instant\r\n        // BEFORE I/O so a later failure still names when this conn went active.\r\n        // [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\r\n        if inner.first_write.get().is_none() {\r\n            let _ = inner.first_write.set(LogTime::now());\r\n            inner.lifecycle_log.healthy(true);\r\n        }\r\n\r\n        // ── Check out the half + arm the watchdog, then write OUT of locks. ──\r\n        let mut half = g.half.take().expect(\"gate half present at checkout\");\r\n        let op = inner.op_seq.fetch_add(1, Ordering::Relaxed) + 1;\r\n        {\r\n            let mut d = recover(&inner.dog);\r\n            d.inflight = Some((op, deadline));\r\n            inner.dog_cv.notify_all();\r\n        }\r\n        drop(g);\r\n\r\n        let res = write_frame(&mut half, env);\r\n\r\n        // Completion handshake: `write_frame` has returned, so the OS op —\r\n        // canceled or not — reported completion and the frame buffer is\r\n        // released. Clear the arm BEFORE deciding health, so the watchdog\r\n        // never severs while the op is formally in flight.\r\n        {\r\n            let mut d = recover(&inner.dog);\r\n            d.inflight = None;\r\n            inner.dog_cv.notify_all();\r\n        }\r\n\r\n        let mut g = recover(&inner.gate);\r\n        if res.is_ok() && !inner.poisoned.load(Ordering::Acquire) {\r\n            g.half = Some(half);\r\n            drop(g);\r\n            inner.gate_cv.notify_all();\r\n            return Ok(ConnWrite::Done);\r\n        }\r\n\r\n        // ── Retire: poison the whole conn, abort the read side so\r\n        // `handle_conn` reaches EOF cleanup, sever, and park the half in the\r\n        // graveyard (its handle must outlive any late raw abort; never\r\n        // re-gated, never reused — a timed-out length-prefixed frame may be\r\n        // partially written). ──\r\n        drop(g);\r\n        inner.poison_and_cancel(res.as_ref().err());\r\n        inner.sever();\r\n        *recover(&inner.graveyard) = Some(half);\r\n        inner.gate_cv.notify_all();\r\n        // A DEADLINE retirement (the watchdog aborted OUR op, or the conn was\r\n        // already deadline-poisoned) surfaces as TimedOut — the caller can\r\n        // count the wedge (W2 stall-evict tally). An organic I/O failure\r\n        // (peer death) keeps its own error.\r\n        if inner.timed_out.load(Ordering::Acquire) {\r\n            return Err(timeout_err());\r\n        }\r\n        Err(match res {\r\n            Err(e) => e,\r\n            Ok(()) => poisoned_err(),\r\n        })\r\n    }\r\n}\r\n\r\n/// TEST-ONLY determinism lever: holds the serialized write gate (the half is\r\n/// checked out) so every writer parks AT THE GATE until the pin drops — the\r\n/// round-4 analog of pinning the old `Mutex<SendHalf>` across a registration\r\n/// window (the P1c epoch-gate keystone unit relies on it).\r\n#[cfg(test)]\r","truncated":true,"truncatedBy":"lines","totalLines":393,"totalBytes":4425,"outputLines":108,"outputBytes":4425,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":1000,"displayContent":{"text":"impl BrokerConn {\n…\n    /// `pre` returning false → `Ok(ConnWrite::Superseded)`, nothing written.\r\n    pub(crate) fn write_gated(\r\n        &self,\r\n        env: &Envelope,\r\n        pre: impl FnOnce() -> bool,\r\n    ) -> io::Result<ConnWrite> {\r\n        let inner = &*self.inner;\r\n        let deadline = Instant::now() + inner.bound;\r\n\r\n        // ── Acquire the gate, bounded by the SAME absolute deadline. ──\r\n        let mut g = recover(&inner.gate);\r\n        loop {\r\n            if inner.poisoned.load(Ordering::Acquire) {\r\n                return Err(poisoned_err());\r\n            }\r\n            if g.half.is_some() {\r\n                break;\r\n            }\r\n            let now = Instant::now();\r\n            if now >= deadline {\r\n                // The write we queued behind is stuck past OUR whole bound:\r\n                // that op is the wedge. Abort it out of band and fail — the\r\n                // watchdog's own deadline (armed at that op's entry) is also\r\n                // in flight; both paths converge on the same idempotent\r\n                // poison/cancel.\r\n                drop(g);\r\n                inner.timed_out.store(true, Ordering::Release);\r\n                inner.poison_and_cancel(None);\r\n                return Err(timeout_err());\r\n            }\r\n            let (g2, _) = inner\r\n                .gate_cv\r\n                .wait_timeout(g, deadline - now)\r\n                .unwrap_or_else(|p| p.into_inner());\r\n            g = g2;\r\n        }\r\n\r\n        if !pre() {\r\n            return Ok(ConnWrite::Superseded);\r\n        }\r\n\r\n        // The gate serializes first-write initialization. Save the instant\r\n        // BEFORE I/O so a later failure still names when this conn went active.\r\n        // [impl->REQ-CONN-HEALTHY-LIFECYCLE-BOUNDED]\r\n        if inner.first_write.get().is_none() {\r\n            let _ = inner.first_write.set(LogTime::now());\r\n            inner.lifecycle_log.healthy(true);\r\n        }\r\n\r\n        // ── Check out the half + arm the watchdog, then write OUT of locks. ──\r\n        let mut half = g.half.take().expect(\"gate half present at checkout\");\r\n        let op = inner.op_seq.fetch_add(1, Ordering::Relaxed) + 1;\r\n        {\r\n            let mut d = recover(&inner.dog);\r\n            d.inflight = Some((op, deadline));\r\n            inner.dog_cv.notify_all();\r\n        }\r\n        drop(g);\r\n\r\n        let res = write_frame(&mut half, env);\r\n\r\n        // Completion handshake: `write_frame` has returned, so the OS op —\r\n        // canceled or not — reported completion and the frame buffer is\r\n        // released. Clear the arm BEFORE deciding health, so the watchdog\r\n        // never severs while the op is formally in flight.\r\n        {\r\n            let mut d = recover(&inner.dog);\r\n            d.inflight = None;\r\n            inner.dog_cv.notify_all();\r\n        }\r\n\r\n        let mut g = recover(&inner.gate);\r\n        if res.is_ok() && !inner.poisoned.load(Ordering::Acquire) {\r\n            g.half = Some(half);\r\n            drop(g);\r\n            inner.gate_cv.notify_all();\r\n            return Ok(ConnWrite::Done);\r\n        }\r\n\r\n        // ── Retire: poison the whole conn, abort the read side so\r\n        // `handle_conn` reaches EOF cleanup, sever, and park the half in the\r\n        // graveyard (its handle must outlive any late raw abort; never\r\n        // re-gated, never reused — a timed-out length-prefixed frame may be\r\n        // partially written). ──\r\n        drop(g);\r\n        inner.poison_and_cancel(res.as_ref().err());\r\n        inner.sever();\r\n        *recover(&inner.graveyard) = Some(half);\r\n        inner.gate_cv.notify_all();\r\n        // A DEADLINE retirement (the watchdog aborted OUR op, or the conn was\r\n        // already deadline-poisoned) surfaces as TimedOut — the caller can\r\n        // count the wedge (W2 stall-evict tally). An organic I/O failure\r\n        // (peer death) keeps its own error.\r\n        if inner.timed_out.load(Ordering::Acquire) {\r\n            return Err(timeout_err());\r\n        }\r\n        Err(match res {\r\n            Err(e) => e,\r\n            Ok(()) => poisoned_err(),\r\n        })\r\n    }\r\n}\r\n\r\n/// TEST-ONLY determinism lever: holds the serialized write gate (the half is\r\n/// checked out) so every writer parks AT THE GATE until the pin drops — the\r\n/// round-4 analog of pinning the old `Mutex<SendHalf>` across a registration\r\n/// window (the P1c epoch-gate keystone unit relies on it).\r\n#[cfg(test)]\r","startLine":523,"lineNumbers":[523,null,608,609,610,611,612,613,614,615,616,617,618,619,620,621,622,623,624,625,626,627,628,629,630,631,632,633,634,635,636,637,638,639,640,641,642,643,644,645,646,647,648,649,650,651,652,653,654,655,656,657,658,659,660,661,662,663,664,665,666,667,668,669,670,671,672,673,674,675,676,677,678,679,680,681,682,683,684,685,686,687,688,689,690,691,692,693,694,695,696,697,698,699,700,701,702,703,704,705,706,707,708,709,710,711,712,713,714,715]},"fileSize":43321,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\crates\\spt-daemon\\src\\conn.rs"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":1000,"totalBytes":4425,"outputLines":108,"outputBytes":4425,"shownRange":{"start":608,"end":715},"nextOffset":716}}},"isError":false,"timestamp":1789523982047}}
{"type":"message","id":"64fc03c1","parentId":"fbe5f62e","timestamp":"2026-09-16T01:59:42.048Z","message":{"role":"toolResult","toolCallId":"call_I65xikavuPiWDjPyL4lQC8nG|fc_0d104aba2b0f3edc016aa9f80df57c87d19869b6fe4ad64e3e","toolName":"grep","content":[{"type":"text","text":" 3558|        }\n*3559|        self.detach_if(sub);\n 3560|    }\n 3561|\n*3562|    fn detach_if(&mut self, sub: &SharedSend) {\n 3563|        let is_ctrl = self\n 3564|            .controller\n 3565|            .as_ref()\n...\n 3823|// [impl->REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]\n*3824|fn note_controller_write_retired(sid: u64, conn: u64, e: &io::Error) {\n 3825|    if e.kind() != io::ErrorKind::TimedOut {\n 3826|        return;\n 3827|    }\n...\n 6446|        for log in &logs {\n*6447|            recover_log(log).detach_if(&send);\n 6448|        }\n 6449|        // Same for net streams: the streams + their rings persist (D4b).\n 6450|        if let Some(host) = self.net.get() {"}],"details":{"scopePath":"crates/spt-daemon/src/broker.rs","searchPath":"C:\\Users\\decid\\Documents\\projects\\spt-core\\crates\\spt-daemon\\src\\broker.rs","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","matchCount":4,"fileCount":1,"files":["crates/spt-daemon/src/broker.rs"],"fileMatches":[{"path":"crates/spt-daemon/src/broker.rs","count":4}],"truncated":false,"displayContent":" 3558│        }\n*3559│        self.detach_if(sub);\n 3560│    }\n 3561│\n*3562│    fn detach_if(&mut self, sub: &SharedSend) {\n 3563│        let is_ctrl = self\n 3564│            .controller\n 3565│            .as_ref()\n     │...\n 3823│// [impl->REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]\n*3824│fn note_controller_write_retired(sid: u64, conn: u64, e: &io::Error) {\n 3825│    if e.kind() != io::ErrorKind::TimedOut {\n 3826│        return;\n 3827│    }\n     │...\n 6446│        for log in &logs {\n*6447│            recover_log(log).detach_if(&send);\n 6448│        }\n 6449│        // Same for net streams: the streams + their rings persist (D4b).\n 6450│        if let Some(host) = self.net.get() {"},"isError":false,"timestamp":1789523982048}}
{"type":"message","id":"0b9f48bf","parentId":"64fc03c1","timestamp":"2026-09-16T01:59:43.147Z","message":{"role":"toolResult","toolCallId":"call_66S3H6COdwzCBA1KqrQdAXOM|fc_0d104aba2b0f3edc016aa9f80df55487d18d87f28b9dec0cba","toolName":"grep","content":[{"type":"text","text":"# traceable-reqs.toml\n 2375|[[requirements]]\n*2376|id = \"REQ-UPDATE-PROMOTE-DRAINED\"\n 2377|title = \"W3 (LIFECYCLE-TRUTH, mechanic-d MOVED FROM W2 per doyle gate verdict @e5ae7a9 — binding): the update-apply brain-generation promotion completes only when the OLD generation's broker subscriber connection is CLOSED or stall-EVICTED — never while blocked writes still pend on it. ROOT: `brain.ready` != subscribers drained; W2's stall-evict (REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE) only BOUNDS the false-promote window to BRAIN_WRITE_DEADLINE (15s), it does NOT close it — a new brain can signal...\n 2378|required_stages = [\"impl\", \"unit\", \"int\"]  # ACTIVATED LIFECYCLE-TRUTH W3 drained-condition slice (todlando 2026-07-07, work-start); brainproc.rs run_trial DRAINED gate + broker any_local_controller_wedged truth + false-promote int rig — closes the W2 15s residual (KH 7.36)\n 2379|\n 2380|[[requirements]]\n*2381|id = \"REQ-UPDATE-TRIAL-DRAIN-DRIVE\"\n*2382|title = \"UPDATE-WEDGE (counter-54, doyle-ruled 2026-07-09 — regression of the v0.29.0 seamless brain-swap): a brain generation DRIVES the broker's controller-liveness reap (a KIND_SESSIONS poll) each heartbeat throughout its boot/trial loop, so a hard-KILLED prior generation's black-holed LOCAL controller conn (by:None) is stall-evicted within the trial window and can never permanently strand the promotion DRAINED gate. ROOT (2026-07-09 field freeze, `spt update fetch --apply` v0.30.0->v0.30.2 froze al...\n 2383|required_stages = [\"doc\", \"impl\", \"int\"]  # ACTIVATED counter-54 UPDATE-WEDGE fix (doyle ruled brain-side 2026-07-09; todlando builds). doc = a CONTEXT.md/LIFECYCLE note that the trial/boot brain drives controller-reap so a killed prior generation cannot strand old_gen_drained. impl = the brain.sessions() KIND_SESSIONS drive in run_brain's heartbeat loop (brainproc.rs, first-tick + every tick, trial + recovery brain). UNIT DELIBERATELY DROPPED (doyle FORK-2 ruling 2026-07-09): the impl is ONE uncondition...\n 2384|\n 2385|[[requirements]]\n 2386|id = \"REQ-BRAIN-UPDATE-RESTART-CLEAN-CLOSE\"\n*2387|title = \"SEED (DEFERRED, doyle 2026-07-09 — post-counter-54 root-hardening for UPDATE-WEDGE; mint now, impl a FUTURE milestone): on a PLANNED brain-restart (`BRAIN_UPDATE_RESTART`, the seamless update-apply brain-cycle), the outgoing brain's LOCAL (by:None) controller conns are GRACEFULLY CLEAN-CLOSED as the brain is cycled, instead of hard-killed and left to black-hole. ROOT (field-pinned 2026-07-09, daemon.stderr.log L24277-24303): the update-restart path hard-kills the outgoing brain (`child.kill()`...\n*2388|required_stages = []  # SEED minted DEFERRED (doyle 2026-07-09); activate post-counter-54 at the milestone that delivers HITCH-FREE seamless updates. Supersedes the livehost-reattach 'Fix Y' framing (wrong site). The counter-54 fix (REQ-UPDATE-TRIAL-DRAIN-DRIVE reap-drive) makes updates SELF-HEAL in ~15s; this seed removes the residual hitch.\n 2389|\n 2390|[[requirements]]\n 2391|id = \"REQ-LIVEHOST-RECONCILE-TRIAL-SILENT\"\n*2392|title = \"SEED (DEFERRED investigation, doyle 2026-07-09 — UPDATE-WEDGE follow-up): determine WHY the trial/rollback brain's livehost reconcile loop did NOT drive the broker controller-reap (nor re-host the live agents) during the ~30s field update-trial window, when livehost polls `query_live_session_endpoints()` → `brain.sessions()` (KIND_SESSIONS) UNCONDITIONALLY every `LIVE_RECONCILE_INTERVAL_MS`=5000ms (livehost.rs:1026). CONTEXT (surfaced building the counter-54 rig): livehost's 5s KIND_SESSIONS...\n*2393|required_stages = []  # SEED deferred investigation (doyle 2026-07-09); non-blocking for counter-54 (the core-loop reap-drive in REQ-UPDATE-TRIAL-DRAIN-DRIVE covers the wedge independently of livehost). Activate once the livehost-silence mechanism is pinned.\n 2394|\n 2395|[[requirements]]\n 2396|id = \"REQ-BRAIN-RESUME-NO-CONTROL-STEAL\"\n*2397|title = \"UPDATE-WEDGE round 2 (v0.30.4, doyle-ruled 2026-07-09 — field incident on the counter-54 fetch--apply): a brain-respawn must NEVER steal, then stall-evict, the controller of a broker PTY session the daemon brain does not DRIVE. ROOT (field-pinned + SME, docs/UPDATE-WEDGE-2-RCA.md + docs/UPDATE-WEDGE-2-SME-todlando.md): `resume_sessions` (brain.rs:985) re-attaches EVERY session `KIND_SESSIONS` returns via `subscribe` = `subscribe_with(AttachIntent::Control, by:None)` (brain.rs:1450-1455). The d...\n*2398|required_stages = [\"doc\", \"impl\", \"int\"]  # ACTIVATED v0.30.4 UPDATE-WEDGE-2 fix (doyle ruled 2026-07-09; todlando builds). doc = brain.rs resume docstring corrected (kill the by:None-never-steals MYTH) + a CONTEXT/LIFECYCLE note that a brain-respawn re-attaches non-driven sessions as viewer. impl = resume path re-attaches AttachIntent::Viewer (Control only for genuinely-driven sessions, empty today). unit DROPPED if the impl is a branchless intent-flip (same FORK-2 reasoning as REQ-UPDATE-TRIAL-DRAIN-DR...\n 2399|\n 2400|[[requirements]]\n*2401|id = \"REQ-BRAIN-RESUME-NO-CONN-DEADLOCK\"\n*2402|title = \"UPDATE-WEDGE round 3 (v0.30.5, doyle-ruled Option A 2026-07-09 — the v0.30.4 field-verify re-wedge, root code-PROVEN + dead-peer-INDEPENDENT): the daemon brain must NOT subscribe broker PTY sessions onto its own request/reply IPC conn — it has no consumer for that output and the subscription DEADLOCKS the conn. ROOT (todlando code-read, docs/UPDATE-WEDGE-2-ROUND3-CODEREAD.md; the net-runtime AND the counter-54 reap-drive were both FALSIFIED first — docs/UPDATE-WEDGE-2-ROUND3-RIG-VERDICT.md...\n*2403|required_stages = [\"doc\", \"impl\", \"int\"]  # ACTIVATED v0.30.5 UPDATE-WEDGE-3 fix (doyle ruled Option A 2026-07-09; todlando builds). doc = brain.rs resume_sessions docstring corrected (the brain no longer subscribes — D4-2/REQ-DAEMON-2 continuity intent) + a CONTEXT/LIFECYCLE note. impl = drop the subscribe_with call in resume_sessions (Option A). unit DROPPED (same FORK-2 reasoning as REQ-UPDATE-TRIAL-DRAIN-DRIVE: the impl is a deletion, no pure branch to unit-test; the int rig load-bearing-guards it ...\n 2404|\n 2405|[[requirements]]\n 2406|id = \"REQ-DIGEST-GENERATION-SUPERSEDE\"\n 2407|title = \"W3 (LIFECYCLE-TRUTH, digest projection truth — flynn filing spt-mobile d0aa3f4): a one-shot `endpoint digest --json` snapshot must return each logical activity row ONCE across a checkpoint/resume, not once per seq-generation. ROOT (spt-core-side, not a consumer bug): the K-session span (digest.rs activity_spanned, SPAN_SESSIONS=5) runs the [digest] extractor per session file and tags each row seq=(ledger_ordinal<<32)|localseq (REQ-DIGEST-CURSOR). A checkpoint/resume (self-/clear + Psyche rebui...\n*2408|required_stages = [\"impl\", \"unit\", \"int\"]  # ACTIVATED W3 supersede slice (todlando 2026-07-07); digest.rs projection-local, independent of PROMOTE-DRAINED; gated within the W3 digest-subsystem sweep\n 2409|\n 2410|[[requirements]]\n 2411|id = \"REQ-UPDATE-FINISH-COMMUNE-FLUSH\"\n...\n 2534|[[requirements]]\n*2535|id = \"REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK\"\n 2536|title = \"UPDATE-WEDGE round-4 (hertz RCA, root reproduced + source-pinned on Windows 2026-07-09): EVERY write on a physical broker connection rides ONE bounded + cancelable + poison-on-failure framed-write primitive — no writer may hold the connection's serialized send gate across an UNBOUNDED OS write. ROOT: controller_writer held the SharedSend = Arc<Mutex<SendHalf>> guard ACROSS a blocking write_frame; on Windows interprocess 2.4.2 routes the send to WriteFileEx + SleepEx(INFINITE, alertable) with N...\n 2537|required_stages = [\"doc\", \"impl\", \"int\"]  # ACTIVATED at mint, UPDATE-WEDGE round-4 (todlando 2026-07-09) per doyle GO. This ACTIVATES (and replaces, dedup at the v0.30.6 merge) the round-3 deferred SEED of the same id — r4 IS the Option-C class fix that SEED reserved; broker-side ⇒ lands via coordinated daemon restart, not a seamless brain-swap. doc = KNOWN-HAZARDS 7.38. impl = conn.rs BrokerConn (gate+watchdog+poison+abort) + full write-site routing sweep. int = brain_decouple.rs non_draining_contr...\n 2538|\n 2539|# ── MSG-IDENTITY milestone (operator GO 2026-07-10: \"two triages + prominent backlog seeds, especially messaging + endpoint identity\"; F-037 EXCLUDED by operator, covered later) ──\n*2540|# Scope: F-036 legs a+c (+ schema-doc fold), F-033/carrier double-deliver cluster, F-038 autostart, F-039 poison log-churn; also ACTIVATES F-032 + F-034 legs (registered earlier, stages flipped this commit). doyle triage; todlando builds; doyle gates; deployah publishes. rule 3 register-before-satisfy; rule 5 activate-at-work-start (this dispatch IS work-start).\n 2541|[[requirements]]\n 2542|id = \"REQ-HAZARD-DAEMON-IDENTITY-ENV-SANITIZE\"\n 2543|title = \"MSG-IDENTITY W1 / F-036 leg a (perri field RCA 2026-07-09/10, psyche seat-theft — doyle ACCEPTED primary fix): the daemon MUST sanitize inherited per-session identity env (SPT_ENDPOINT_ID / OWL_SESSION_ID / SPT_AGENT_ID) at startup AND before EVERY role spawn — these are per-session identity and are NEVER correct inherited state for a daemon or its role children. ROOT: a daemon restarted from inside an agent session (routine during core dev / `spt update apply`) carries the session's SPT_END...\n...\n 2577|id = \"REQ-CONN-POISON-DIAL-SCOPE\"\n*2578|title = \"MSG-IDENTITY W6 / F-039 (deployah field-acceptance follow-up 2026-07-10, RCA-FIRST — mint per the v0.30.6 PASS handoff): ambient CONN_WRITE_POISONED log-churn correlates 1:1 with PUMP_PEER_FAIL submit-dials to OFFLINE peers (enlyzeam/kitsubito/gravity) with NO wedge and NO freeze — pre-existed the blackhole rig = log-noise/mislabel, not a defect in the r4 fix. CODE CONTEXT: conn.rs poison_and_cancel emits the loud CONN_WRITE_POISONED line for a write that 'exceeded its bound (OR FAILED)' (co...\n 2579|required_stages = [\"impl\", \"unit\"]  # FLIPPED in the W6 build commit carrying the evidence (todlando 2026-07-10): impl = the class-split render at poison_and_cancel (conn.rs render_retirement); unit = conn.rs token-split + organic-still-poisons cases. RCA LOCK ADDENDUM (supersedes the title's open questions): QUIC dial ruled OUT; the per-offline-dial-cycle poisoning writes hit SHARED long-lived brain carriers — (1) the DIAL_FAILED presence push (nethost.rs:373, result discarded, fires per failure) and ...\n 2580|\n 2581|[[requirements]]\n 2582|id = \"REQ-CONN-POISON-ATTRIBUTION\"\n*2583|title = \"MSG-IDENTITY W6 / F-039 legs b-d (doyle W6 LOCK 2026-07-10, minted per amendment 3): every broker-conn lifecycle record is ATTRIBUTABLE — the W6 RCA's terminal undecidability (per-line 1:1 CONN_WRITE_POISONED churn = fresh-carrier churn OR stderr interleave artifact) exists because records carry no stable conn identity, no role/endpoint/session context, and no timestamps, and the once-per-conn poison latch hides multiplicity. THREE LEGS. (b) IDENTITY: mint a stable per-physical-conn id (monoto...\n 2584|required_stages = [\"impl\", \"unit\"]  # FLIPPED in the W6 build commit carrying the evidence (todlando 2026-07-10). impl = conn id mint + describe/label accretion + log_stamp wall+mono (anchored at both daemon entries) + bounded lifecycle events (write-start, transport-close, writer-exit, attach/replace/detach) + attribution on stall-evict records. unit = conn.rs record-shape assertions (id/role/time present + parity across both retirement tokens, id unique per conn, label bounded).\n 2585|\n 2586|# releases#286 narrowly supersedes leg (d)'s unconditional healthy start/close\n...\n 2595|id = \"REQ-CONN-BLACKHOLE-LIFECYCLE-HARNESS\"\n*2596|title = \"MSG-IDENTITY W6 / F-039 leg e (doyle W6 LOCK 2026-07-10, minted per amendment 3 — hertz's five invariants VERBATIM from his RCA fix-shape item 5): 'Build a deterministic black-holed-controller harness against current v0.30.6 semantics and assert: unrelated sessions continue; the bad physical connection is canceled/closed within the bound; its writer exits; a fresh viewer can attach; no lock or task remains owned by the retired connection.' The harness is the standing conformance rig for the r4...\n 2597|required_stages = [\"int\"]  # FLIPPED in the harness commit carrying the evidence (todlando 2026-07-10). int = conn_blackhole_lifecycle five-invariant rig (crates/spt-daemon/tests/conn_blackhole_lifecycle.rs::blackholed_controller_lifecycle_five_invariants) — a deterministic black-holed-controller incident asserting all five: unrelated sessions continue, the bad physical conn is canceled/closed within the bound, its writer exits, a fresh viewer attaches + receives output, no lock/task remains owned by t...\n 2598|\n 2599|[[requirements]]\n...\n 2650|id = \"REQ-DAEMON-REFRESH\"\n*2651|title = \"THE-FORKENING W4 (operator add 2026-07-14): `spt daemon refresh` — restart the daemon BRAIN without a binary swap and WITHOUT touching the broker: exactly the apply_staged brain-cycle path (brain stop -> respawn -> readiness trial -> promote, incl. the trial-drain drive REQ-UPDATE-TRIAL-DRAIN-DRIVE and viewer-only resume REQ-BRAIN-RESUME-NO-CONTROL-STEAL) minus the swap. Recovery verb for wedged brain-held state (field motivator 2026-07-14: endpoint bringup broken on a live daemon + deployah d...\n 2652|required_stages = [\"doc\", \"impl\", \"unit\", \"int\"]  # FLIPPED in the W4 build commit (todlando 2026-07-15).\n 2653|\n 2654|[[requirements]]\n...\n 2777|id = \"REQ-HAZARD-REDISPATCH-STALL\"\n*2778|title = \"REDISPATCH-STALL W1 (KNOWN-HAZARDS 7.43, hertz v0.34 field RCA 2026-07-16 — recurrent 20-30s PTY/RC freezes, 17-62s DISPATCH tails): one wedged stream subscriber must NEVER stall stream serving, and recovery machinery must not manufacture new replay victims. Today: claim retries x the broad Err(_) opener fallback (dispatch.rs:414) install throwaway peek subscribers whose StreamLog::attach replays the entire retained ring UNDER the per-stream mutex with discarded write errors and the poisoned s...\n 2779|required_stages = [\"doc\", \"int\"]  # ACTIVATED REDISPATCH-STALL W1 (todlando 2026-07-16). int = production-path wedged-subscriber regression (producer+unrelated streams flat, poisoned subscriber removed, stream recovers); doc = KNOWN-HAZARDS 7.43.\n 2780|\n 2781|[[requirements]]\n\n# docs/\n## 49-267-REFRESH-OBSERVATION-PLAN.md\n 142|evidence that the heartbeat's drain-driving `brain.sessions()` call completed — that call is what\n*143|makes a wedge drainable at all (REQ-UPDATE-TRIAL-DRAIN-DRIVE), and a trial that never drove it looks\n 144|identical here to one that drove it and still waited. It also names **no session**: the gate answers a\n 145|single bool over every hosted session, and naming the wedged one would require a second, separately\n 146|taken scan that could disagree with the read the promote decision actually used. The narrower fact is\n## KNOWN-HAZARDS.md\n 750|- **Failure (paid-for, rig-CONFIRMED 2026-07-06/07 — `NtSuspendProcess` on the brain, no update involved):** a controller's writer thread does a BLOCKING socket write to its brain subscriber conn. When that brain is suspended (or black-holed) the write never returns. The output-driven eviction path (`append` → `mark_controller_gone`, bounded by 7.12's `CONTROLLER_WRITE_DEADLINE`) only fires on NEW output, and `reap_dead_controller` (7.29 companion) only caught a writer that had EXITED — so a contro...\n*751|- **Invariant:** no broker client ever observes a wedge because the BRAIN stalled. (1) The controller's writer publishes its IN-FLIGHT-write window (`write_blocked_since`: `Some(Instant)` set immediately before the blocking `write_frame`, `None` after — the mutex never held ACROSS the write); a reader sees a `Some` older than `BRAIN_WRITE_DEADLINE` (15 s = 3× the 7.12 controller bound, measured full-with-**zero writer progress**, so a slow-but-DRAINING brain never trips it) as a WEDGED writer and stal...\n 752|- **spt-core mapping:** `broker.rs` — `ControllerSink::write_blocked_since` + `controller_writer` (marks the in-flight window around both the initial-batch and live-loop writes) + `controller_write_stalled` (pure predicate over injected `now`/`deadline`) + `stall_evict_controller` (Inline stamp at `resolve_subscribe`, Deferred stamp at `reap_dead_controller` so the reap closure does no I/O under the shared sessions lock) + `record_stall_evict`/`stall_evict_stats` (the tally) + `dispatch_stall_evicts` (...\n 753|- **Source:** LIFECYCLE-TRUTH W2 (doyle rig `NtSuspendProcess` root + design ruling `docs/W2-DESIGN-RULING.md` @60ec0ed — minimal-plus scope, `BRAIN_WRITE_DEADLINE` 15 s, passive resubscribe + observability rider; todlando build 2026-07-07). The v0.27→v0.28 update-wedge night's flagship defect.\n 754|\n...\n 761|\n*762|### 7.38 EVERY write on a physical broker connection is bounded + cancelable + poison-on-failure — no writer holds the send gate across an UNBOUNDED OS write  `[REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]`\n*763|<!-- [doc->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK] -->\n 764|- **Failure (paid-for, field capture + deterministic Windows repro 2026-07-09 — hertz RCA, UPDATE-WEDGE round-4):** `controller_writer` held the `SharedSend = Arc<Mutex<SendHalf>>` guard ACROSS a blocking `write_frame`. On Windows, `interprocess 2.4.2` routes the send to `WriteFileEx` + `SleepEx(INFINITE, alertable)` — this local-socket send path exposes **no supported write timeout** (`set_timeout` → `Unsupported`). An `rc --take` controller consumer that stopped reading blocked the write **indefi...\n 765|- **Invariant:** every write on a physical broker connection rides ONE broker-owned bounded/cancelable framed-write primitive with an **independent out-of-band aborter**. The `SendHalf` never leaves the conn object (a bypass cannot compile). An ABSOLUTE deadline stamped at write entry covers BOTH the serialized gate-wait AND OS write completion; a per-conn watchdog fires at that deadline and aborts the in-flight op — never relying on the write returning or on opportunistic stall-evict sampling. On dead...\n 766|- **spt-core mapping:** `spt-daemon/src/conn.rs` `BrokerConn` (Condvar-gated serialized write gate + poisoned state + per-in-flight op id + watchdog + `abort_physical()`: cfg(windows) `CancelIoEx` → completion handshake → `DisconnectNamedPipe` raw kernel32 externs; cfg(unix) `UnixStream::shutdown(Both)`); `broker.rs` `SharedSend = Arc<BrokerConn>` + `controller_writer`/`viewer_writer`/`send_frame`/`send_error` routed; `nethost.rs` stream-log/presence sends routed. int = `brain_decouple.rs` `non_drain...\n## MSG-IDENTITY-DISPATCH.md\n 75|\n*76|### W6 — POISON-LOG-SCOPE (F-039; deployah v0.30.6 PASS handoff, **RCA-first**)\n 77|- `REQ-CONN-POISON-DIAL-SCOPE` [impl,unit] — CONN_WRITE_POISONED churn 1:1 with PUMP_PEER_FAIL\n 78|  submit-dials to offline peers; no wedge, pre-existed the rig. Code context: poison_and_cancel\n 79|  (conn.rs:181) logs the SAME loud token for \"exceeded bound **(or failed)**\" — the fast-FAIL\n## W3-DIGEST-GENERATION-UNION-RULING.md\n 68|  items (or folded into its assembly). No `brainproc.rs` touch → **no conflict with\n*69|  PROMOTE-DRAINED**; sequence independently.\n 70|- **Activate** `REQ-DIGEST-GENERATION-SUPERSEDE` (`impl, unit, int`) in the work-start commit\n 71|  (rule 5). Tag evidence on the real code.\n 72|- **Int rig (RED-first):** a two-session span where session B's extracted lines are a superset\n## W3-GATE-VERDICT.md\n 13|| REQ-DIGEST-GENERATION-SUPERSEDE | `7e37e37` | impl+unit+int |\n*14|| REQ-UPDATE-PROMOTE-DRAINED | `e75e856` | impl+unit+int |\n 15|\n 16|## Gate legs — both green\n 17|- **Windows (HFENDULEAM), isolated worktree `.worktrees/gate-e75e856`, fresh target, `CARGO_INCREMENTAL=0`:**\n...\n 36|  RED-first int (`two_origin_spanning.rs`); disjoint-span test stays green (supersede no-op).\n*37|- **REQ-UPDATE-PROMOTE-DRAINED** (`brainproc.rs`/`broker.rs`/`daemon.rs`): sound. Readiness\n 38|  LATCHED (`ready_seen`, `try_wait` only pre-ready → ready-then-exit still promotes once\n 39|  drained); promote gated on `ready_seen && old_gen_drained()`; never-drains → WindowElapsedAlive\n 40|  → kill + rollback (conservative — never false-promote onto a wedge). `old_gen_drained` reads\n## W3-SLICE1-REPORT.md\n 52|REQ-DAEMON-STOP-LIVE-SESSION-WARN → REQ-RC-RECONNECT-TRUTH → REQ-UPDATE-ONE-SHOT-FINISH\n*53|(composes on ENDPOINT-SURVIVAL) → REQ-UPDATE-PROMOTE-DRAINED (mechanic-d + false-promote\n 54|rig, the binding W3 activation condition from your W2 verdict).\n## W3-WAVE-GATE-REPORT.md\n 14|| REQ-DIGEST-GENERATION-SUPERSEDE | `7e37e37` | impl+unit+int |\n*15|| **REQ-UPDATE-PROMOTE-DRAINED** | **`e75e856`** | **impl+unit+int** |\n 16|\n*17|## REQ-UPDATE-PROMOTE-DRAINED @e75e856 (mechanic-d — closes the W2 15s residual, KH 7.36)\n 18|\n 19|**Root closed:** W2's stall-evict only BOUNDED the update-apply false-promote window to\n 20|`BRAIN_WRITE_DEADLINE` (15s); a new brain could signal `brain.ready` inside that window\n...\n 62|- `xtask check` (docs-drift + token, `CARGO_TARGET_DIR` unset): OK.\n*63|- traceable-reqs check: exit 0; REQ-UPDATE-PROMOTE-DRAINED `+impl +unit +int`.\n 64|- Doyle staged condition — grep OLD rc auto-start assertions: CLEAN (every `rc.rs` reference\n 65|  asserts the NEW connect-only truth; the other hits are unrelated `api` startup hygiene).\n 66|\n...\n 73|## KH 7.36\n*74|Updated: the W2 residual is now **CLOSED by W3** (`REQ-UPDATE-PROMOTE-DRAINED`), not just bounded.\n\n[Some lines truncated to 512 chars]"}],"details":{"scopePath":"docs, traceable-reqs.toml","searchPath":"C:\\Users\\decid\\Documents\\projects\\spt-core","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","matchCount":33,"fileCount":8,"files":["docs/49-267-REFRESH-OBSERVATION-PLAN.md","docs/KNOWN-HAZARDS.md","docs/MSG-IDENTITY-DISPATCH.md","docs/W3-DIGEST-GENERATION-UNION-RULING.md","docs/W3-GATE-VERDICT.md","docs/W3-SLICE1-REPORT.md","docs/W3-WAVE-GATE-REPORT.md","traceable-reqs.toml"],"fileMatches":[{"path":"docs/49-267-REFRESH-OBSERVATION-PLAN.md","count":1},{"path":"docs/KNOWN-HAZARDS.md","count":3},{"path":"docs/MSG-IDENTITY-DISPATCH.md","count":1},{"path":"docs/W3-DIGEST-GENERATION-UNION-RULING.md","count":1},{"path":"docs/W3-GATE-VERDICT.md","count":2},{"path":"docs/W3-SLICE1-REPORT.md","count":1},{"path":"docs/W3-WAVE-GATE-REPORT.md","count":4},{"path":"traceable-reqs.toml","count":20}],"truncated":true,"perFileLimitReached":20,"displayContent":"# traceable-reqs.toml\n 2375│[[requirements]]\n*2376│id = \"REQ-UPDATE-PROMOTE-DRAINED\"\n 2377│title = \"W3 (LIFECYCLE-TRUTH, mechanic-d MOVED FROM W2 per doyle gate verdict @e5ae7a9 — binding): the update-apply brain-generation promotion completes only when the OLD generation's broker subscriber connection is CLOSED or stall-EVICTED — never while blocked writes still pend on it. ROOT: `brain.ready` != subscribers drained; W2's stall-evict (REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE) only BOUNDS the false-promote window to BRAIN_WRITE_DEADLINE (15s), it does NOT close it — a new brain can signal...\n 2378│required_stages = [\"impl\", \"unit\", \"int\"]  # ACTIVATED LIFECYCLE-TRUTH W3 drained-condition slice (todlando 2026-07-07, work-start); brainproc.rs run_trial DRAINED gate + broker any_local_controller_wedged truth + false-promote int rig — closes the W2 15s residual (KH 7.36)\n 2379│\n 2380│[[requirements]]\n*2381│id = \"REQ-UPDATE-TRIAL-DRAIN-DRIVE\"\n*2382│title = \"UPDATE-WEDGE (counter-54, doyle-ruled 2026-07-09 — regression of the v0.29.0 seamless brain-swap): a brain generation DRIVES the broker's controller-liveness reap (a KIND_SESSIONS poll) each heartbeat throughout its boot/trial loop, so a hard-KILLED prior generation's black-holed LOCAL controller conn (by:None) is stall-evicted within the trial window and can never permanently strand the promotion DRAINED gate. ROOT (2026-07-09 field freeze, `spt update fetch --apply` v0.30.0->v0.30.2 froze al...\n 2383│required_stages = [\"doc\", \"impl\", \"int\"]  # ACTIVATED counter-54 UPDATE-WEDGE fix (doyle ruled brain-side 2026-07-09; todlando builds). doc = a CONTEXT.md/LIFECYCLE note that the trial/boot brain drives controller-reap so a killed prior generation cannot strand old_gen_drained. impl = the brain.sessions() KIND_SESSIONS drive in run_brain's heartbeat loop (brainproc.rs, first-tick + every tick, trial + recovery brain). UNIT DELIBERATELY DROPPED (doyle FORK-2 ruling 2026-07-09): the impl is ONE uncondition...\n 2384│\n 2385│[[requirements]]\n 2386│id = \"REQ-BRAIN-UPDATE-RESTART-CLEAN-CLOSE\"\n*2387│title = \"SEED (DEFERRED, doyle 2026-07-09 — post-counter-54 root-hardening for UPDATE-WEDGE; mint now, impl a FUTURE milestone): on a PLANNED brain-restart (`BRAIN_UPDATE_RESTART`, the seamless update-apply brain-cycle), the outgoing brain's LOCAL (by:None) controller conns are GRACEFULLY CLEAN-CLOSED as the brain is cycled, instead of hard-killed and left to black-hole. ROOT (field-pinned 2026-07-09, daemon.stderr.log L24277-24303): the update-restart path hard-kills the outgoing brain (`child.kill()`...\n*2388│required_stages = []  # SEED minted DEFERRED (doyle 2026-07-09); activate post-counter-54 at the milestone that delivers HITCH-FREE seamless updates. Supersedes the livehost-reattach 'Fix Y' framing (wrong site). The counter-54 fix (REQ-UPDATE-TRIAL-DRAIN-DRIVE reap-drive) makes updates SELF-HEAL in ~15s; this seed removes the residual hitch.\n 2389│\n 2390│[[requirements]]\n 2391│id = \"REQ-LIVEHOST-RECONCILE-TRIAL-SILENT\"\n*2392│title = \"SEED (DEFERRED investigation, doyle 2026-07-09 — UPDATE-WEDGE follow-up): determine WHY the trial/rollback brain's livehost reconcile loop did NOT drive the broker controller-reap (nor re-host the live agents) during the ~30s field update-trial window, when livehost polls `query_live_session_endpoints()` → `brain.sessions()` (KIND_SESSIONS) UNCONDITIONALLY every `LIVE_RECONCILE_INTERVAL_MS`=5000ms (livehost.rs:1026). CONTEXT (surfaced building the counter-54 rig): livehost's 5s KIND_SESSIONS...\n*2393│required_stages = []  # SEED deferred investigation (doyle 2026-07-09); non-blocking for counter-54 (the core-loop reap-drive in REQ-UPDATE-TRIAL-DRAIN-DRIVE covers the wedge independently of livehost). Activate once the livehost-silence mechanism is pinned.\n 2394│\n 2395│[[requirements]]\n 2396│id = \"REQ-BRAIN-RESUME-NO-CONTROL-STEAL\"\n*2397│title = \"UPDATE-WEDGE round 2 (v0.30.4, doyle-ruled 2026-07-09 — field incident on the counter-54 fetch--apply): a brain-respawn must NEVER steal, then stall-evict, the controller of a broker PTY session the daemon brain does not DRIVE. ROOT (field-pinned + SME, docs/UPDATE-WEDGE-2-RCA.md + docs/UPDATE-WEDGE-2-SME-todlando.md): `resume_sessions` (brain.rs:985) re-attaches EVERY session `KIND_SESSIONS` returns via `subscribe` = `subscribe_with(AttachIntent::Control, by:None)` (brain.rs:1450-1455). The d...\n*2398│required_stages = [\"doc\", \"impl\", \"int\"]  # ACTIVATED v0.30.4 UPDATE-WEDGE-2 fix (doyle ruled 2026-07-09; todlando builds). doc = brain.rs resume docstring corrected (kill the by:None-never-steals MYTH) + a CONTEXT/LIFECYCLE note that a brain-respawn re-attaches non-driven sessions as viewer. impl = resume path re-attaches AttachIntent::Viewer (Control only for genuinely-driven sessions, empty today). unit DROPPED if the impl is a branchless intent-flip (same FORK-2 reasoning as REQ-UPDATE-TRIAL-DRAIN-DR...\n 2399│\n 2400│[[requirements]]\n*2401│id = \"REQ-BRAIN-RESUME-NO-CONN-DEADLOCK\"\n*2402│title = \"UPDATE-WEDGE round 3 (v0.30.5, doyle-ruled Option A 2026-07-09 — the v0.30.4 field-verify re-wedge, root code-PROVEN + dead-peer-INDEPENDENT): the daemon brain must NOT subscribe broker PTY sessions onto its own request/reply IPC conn — it has no consumer for that output and the subscription DEADLOCKS the conn. ROOT (todlando code-read, docs/UPDATE-WEDGE-2-ROUND3-CODEREAD.md; the net-runtime AND the counter-54 reap-drive were both FALSIFIED first — docs/UPDATE-WEDGE-2-ROUND3-RIG-VERDICT.md...\n*2403│required_stages = [\"doc\", \"impl\", \"int\"]  # ACTIVATED v0.30.5 UPDATE-WEDGE-3 fix (doyle ruled Option A 2026-07-09; todlando builds). doc = brain.rs resume_sessions docstring corrected (the brain no longer subscribes — D4-2/REQ-DAEMON-2 continuity intent) + a CONTEXT/LIFECYCLE note. impl = drop the subscribe_with call in resume_sessions (Option A). unit DROPPED (same FORK-2 reasoning as REQ-UPDATE-TRIAL-DRAIN-DRIVE: the impl is a deletion, no pure branch to unit-test; the int rig load-bearing-guards it ...\n 2404│\n 2405│[[requirements]]\n 2406│id = \"REQ-DIGEST-GENERATION-SUPERSEDE\"\n 2407│title = \"W3 (LIFECYCLE-TRUTH, digest projection truth — flynn filing spt-mobile d0aa3f4): a one-shot `endpoint digest --json` snapshot must return each logical activity row ONCE across a checkpoint/resume, not once per seq-generation. ROOT (spt-core-side, not a consumer bug): the K-session span (digest.rs activity_spanned, SPAN_SESSIONS=5) runs the [digest] extractor per session file and tags each row seq=(ledger_ordinal<<32)|localseq (REQ-DIGEST-CURSOR). A checkpoint/resume (self-/clear + Psyche rebui...\n*2408│required_stages = [\"impl\", \"unit\", \"int\"]  # ACTIVATED W3 supersede slice (todlando 2026-07-07); digest.rs projection-local, independent of PROMOTE-DRAINED; gated within the W3 digest-subsystem sweep\n 2409│\n 2410│[[requirements]]\n 2411│id = \"REQ-UPDATE-FINISH-COMMUNE-FLUSH\"\n     │...\n 2534│[[requirements]]\n*2535│id = \"REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK\"\n 2536│title = \"UPDATE-WEDGE round-4 (hertz RCA, root reproduced + source-pinned on Windows 2026-07-09): EVERY write on a physical broker connection rides ONE bounded + cancelable + poison-on-failure framed-write primitive — no writer may hold the connection's serialized send gate across an UNBOUNDED OS write. ROOT: controller_writer held the SharedSend = Arc<Mutex<SendHalf>> guard ACROSS a blocking write_frame; on Windows interprocess 2.4.2 routes the send to WriteFileEx + SleepEx(INFINITE, alertable) with N...\n 2537│required_stages = [\"doc\", \"impl\", \"int\"]  # ACTIVATED at mint, UPDATE-WEDGE round-4 (todlando 2026-07-09) per doyle GO. This ACTIVATES (and replaces, dedup at the v0.30.6 merge) the round-3 deferred SEED of the same id — r4 IS the Option-C class fix that SEED reserved; broker-side ⇒ lands via coordinated daemon restart, not a seamless brain-swap. doc = KNOWN-HAZARDS 7.38. impl = conn.rs BrokerConn (gate+watchdog+poison+abort) + full write-site routing sweep. int = brain_decouple.rs non_draining_contr...\n 2538│\n 2539│# ── MSG-IDENTITY milestone (operator GO 2026-07-10: \"two triages + prominent backlog seeds, especially messaging + endpoint identity\"; F-037 EXCLUDED by operator, covered later) ──\n*2540│# Scope: F-036 legs a+c (+ schema-doc fold), F-033/carrier double-deliver cluster, F-038 autostart, F-039 poison log-churn; also ACTIVATES F-032 + F-034 legs (registered earlier, stages flipped this commit). doyle triage; todlando builds; doyle gates; deployah publishes. rule 3 register-before-satisfy; rule 5 activate-at-work-start (this dispatch IS work-start).\n 2541│[[requirements]]\n 2542│id = \"REQ-HAZARD-DAEMON-IDENTITY-ENV-SANITIZE\"\n 2543│title = \"MSG-IDENTITY W1 / F-036 leg a (perri field RCA 2026-07-09/10, psyche seat-theft — doyle ACCEPTED primary fix): the daemon MUST sanitize inherited per-session identity env (SPT_ENDPOINT_ID / OWL_SESSION_ID / SPT_AGENT_ID) at startup AND before EVERY role spawn — these are per-session identity and are NEVER correct inherited state for a daemon or its role children. ROOT: a daemon restarted from inside an agent session (routine during core dev / `spt update apply`) carries the session's SPT_END...\n     │...\n 2577│id = \"REQ-CONN-POISON-DIAL-SCOPE\"\n*2578│title = \"MSG-IDENTITY W6 / F-039 (deployah field-acceptance follow-up 2026-07-10, RCA-FIRST — mint per the v0.30.6 PASS handoff): ambient CONN_WRITE_POISONED log-churn correlates 1:1 with PUMP_PEER_FAIL submit-dials to OFFLINE peers (enlyzeam/kitsubito/gravity) with NO wedge and NO freeze — pre-existed the blackhole rig = log-noise/mislabel, not a defect in the r4 fix. CODE CONTEXT: conn.rs poison_and_cancel emits the loud CONN_WRITE_POISONED line for a write that 'exceeded its bound (OR FAILED)' (co...\n 2579│required_stages = [\"impl\", \"unit\"]  # FLIPPED in the W6 build commit carrying the evidence (todlando 2026-07-10): impl = the class-split render at poison_and_cancel (conn.rs render_retirement); unit = conn.rs token-split + organic-still-poisons cases. RCA LOCK ADDENDUM (supersedes the title's open questions): QUIC dial ruled OUT; the per-offline-dial-cycle poisoning writes hit SHARED long-lived brain carriers — (1) the DIAL_FAILED presence push (nethost.rs:373, result discarded, fires per failure) and ...\n 2580│\n 2581│[[requirements]]\n 2582│id = \"REQ-CONN-POISON-ATTRIBUTION\"\n*2583│title = \"MSG-IDENTITY W6 / F-039 legs b-d (doyle W6 LOCK 2026-07-10, minted per amendment 3): every broker-conn lifecycle record is ATTRIBUTABLE — the W6 RCA's terminal undecidability (per-line 1:1 CONN_WRITE_POISONED churn = fresh-carrier churn OR stderr interleave artifact) exists because records carry no stable conn identity, no role/endpoint/session context, and no timestamps, and the once-per-conn poison latch hides multiplicity. THREE LEGS. (b) IDENTITY: mint a stable per-physical-conn id (monoto...\n 2584│required_stages = [\"impl\", \"unit\"]  # FLIPPED in the W6 build commit carrying the evidence (todlando 2026-07-10). impl = conn id mint + describe/label accretion + log_stamp wall+mono (anchored at both daemon entries) + bounded lifecycle events (write-start, transport-close, writer-exit, attach/replace/detach) + attribution on stall-evict records. unit = conn.rs record-shape assertions (id/role/time present + parity across both retirement tokens, id unique per conn, label bounded).\n 2585│\n 2586│# releases#286 narrowly supersedes leg (d)'s unconditional healthy start/close\n     │...\n 2595│id = \"REQ-CONN-BLACKHOLE-LIFECYCLE-HARNESS\"\n*2596│title = \"MSG-IDENTITY W6 / F-039 leg e (doyle W6 LOCK 2026-07-10, minted per amendment 3 — hertz's five invariants VERBATIM from his RCA fix-shape item 5): 'Build a deterministic black-holed-controller harness against current v0.30.6 semantics and assert: unrelated sessions continue; the bad physical connection is canceled/closed within the bound; its writer exits; a fresh viewer can attach; no lock or task remains owned by the retired connection.' The harness is the standing conformance rig for the r4...\n 2597│required_stages = [\"int\"]  # FLIPPED in the harness commit carrying the evidence (todlando 2026-07-10). int = conn_blackhole_lifecycle five-invariant rig (crates/spt-daemon/tests/conn_blackhole_lifecycle.rs::blackholed_controller_lifecycle_five_invariants) — a deterministic black-holed-controller incident asserting all five: unrelated sessions continue, the bad physical conn is canceled/closed within the bound, its writer exits, a fresh viewer attaches + receives output, no lock/task remains owned by t...\n 2598│\n 2599│[[requirements]]\n     │...\n 2650│id = \"REQ-DAEMON-REFRESH\"\n*2651│title = \"THE-FORKENING W4 (operator add 2026-07-14): `spt daemon refresh` — restart the daemon BRAIN without a binary swap and WITHOUT touching the broker: exactly the apply_staged brain-cycle path (brain stop -> respawn -> readiness trial -> promote, incl. the trial-drain drive REQ-UPDATE-TRIAL-DRAIN-DRIVE and viewer-only resume REQ-BRAIN-RESUME-NO-CONTROL-STEAL) minus the swap. Recovery verb for wedged brain-held state (field motivator 2026-07-14: endpoint bringup broken on a live daemon + deployah d...\n 2652│required_stages = [\"doc\", \"impl\", \"unit\", \"int\"]  # FLIPPED in the W4 build commit (todlando 2026-07-15).\n 2653│\n 2654│[[requirements]]\n     │...\n 2777│id = \"REQ-HAZARD-REDISPATCH-STALL\"\n*2778│title = \"REDISPATCH-STALL W1 (KNOWN-HAZARDS 7.43, hertz v0.34 field RCA 2026-07-16 — recurrent 20-30s PTY/RC freezes, 17-62s DISPATCH tails): one wedged stream subscriber must NEVER stall stream serving, and recovery machinery must not manufacture new replay victims. Today: claim retries x the broad Err(_) opener fallback (dispatch.rs:414) install throwaway peek subscribers whose StreamLog::attach replays the entire retained ring UNDER the per-stream mutex with discarded write errors and the poisoned s...\n 2779│required_stages = [\"doc\", \"int\"]  # ACTIVATED REDISPATCH-STALL W1 (todlando 2026-07-16). int = production-path wedged-subscriber regression (producer+unrelated streams flat, poisoned subscriber removed, stream recovers); doc = KNOWN-HAZARDS 7.43.\n 2780│\n 2781│[[requirements]]\n\n# docs/\n## 49-267-REFRESH-OBSERVATION-PLAN.md\n 142│evidence that the heartbeat's drain-driving `brain.sessions()` call completed — that call is what\n*143│makes a wedge drainable at all (REQ-UPDATE-TRIAL-DRAIN-DRIVE), and a trial that never drove it looks\n 144│identical here to one that drove it and still waited. It also names **no session**: the gate answers a\n 145│single bool over every hosted session, and naming the wedged one would require a second, separately\n 146│taken scan that could disagree with the read the promote decision actually used. The narrower fact is\n## KNOWN-HAZARDS.md\n 750│- **Failure (paid-for, rig-CONFIRMED 2026-07-06/07 — `NtSuspendProcess` on the brain, no update involved):** a controller's writer thread does a BLOCKING socket write to its brain subscriber conn. When that brain is suspended (or black-holed) the write never returns. The output-driven eviction path (`append` → `mark_controller_gone`, bounded by 7.12's `CONTROLLER_WRITE_DEADLINE`) only fires on NEW output, and `reap_dead_controller` (7.29 companion) only caught a writer that had EXITED — so a contro...\n*751│- **Invariant:** no broker client ever observes a wedge because the BRAIN stalled. (1) The controller's writer publishes its IN-FLIGHT-write window (`write_blocked_since`: `Some(Instant)` set immediately before the blocking `write_frame`, `None` after — the mutex never held ACROSS the write); a reader sees a `Some` older than `BRAIN_WRITE_DEADLINE` (15 s = 3× the 7.12 controller bound, measured full-with-**zero writer progress**, so a slow-but-DRAINING brain never trips it) as a WEDGED writer and stal...\n 752│- **spt-core mapping:** `broker.rs` — `ControllerSink::write_blocked_since` + `controller_writer` (marks the in-flight window around both the initial-batch and live-loop writes) + `controller_write_stalled` (pure predicate over injected `now`/`deadline`) + `stall_evict_controller` (Inline stamp at `resolve_subscribe`, Deferred stamp at `reap_dead_controller` so the reap closure does no I/O under the shared sessions lock) + `record_stall_evict`/`stall_evict_stats` (the tally) + `dispatch_stall_evicts` (...\n 753│- **Source:** LIFECYCLE-TRUTH W2 (doyle rig `NtSuspendProcess` root + design ruling `docs/W2-DESIGN-RULING.md` @60ec0ed — minimal-plus scope, `BRAIN_WRITE_DEADLINE` 15 s, passive resubscribe + observability rider; todlando build 2026-07-07). The v0.27→v0.28 update-wedge night's flagship defect.\n 754│\n    │...\n 761│\n*762│### 7.38 EVERY write on a physical broker connection is bounded + cancelable + poison-on-failure — no writer holds the send gate across an UNBOUNDED OS write  `[REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]`\n*763│<!-- [doc->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK] -->\n 764│- **Failure (paid-for, field capture + deterministic Windows repro 2026-07-09 — hertz RCA, UPDATE-WEDGE round-4):** `controller_writer` held the `SharedSend = Arc<Mutex<SendHalf>>` guard ACROSS a blocking `write_frame`. On Windows, `interprocess 2.4.2` routes the send to `WriteFileEx` + `SleepEx(INFINITE, alertable)` — this local-socket send path exposes **no supported write timeout** (`set_timeout` → `Unsupported`). An `rc --take` controller consumer that stopped reading blocked the write **indefi...\n 765│- **Invariant:** every write on a physical broker connection rides ONE broker-owned bounded/cancelable framed-write primitive with an **independent out-of-band aborter**. The `SendHalf` never leaves the conn object (a bypass cannot compile). An ABSOLUTE deadline stamped at write entry covers BOTH the serialized gate-wait AND OS write completion; a per-conn watchdog fires at that deadline and aborts the in-flight op — never relying on the write returning or on opportunistic stall-evict sampling. On dead...\n 766│- **spt-core mapping:** `spt-daemon/src/conn.rs` `BrokerConn` (Condvar-gated serialized write gate + poisoned state + per-in-flight op id + watchdog + `abort_physical()`: cfg(windows) `CancelIoEx` → completion handshake → `DisconnectNamedPipe` raw kernel32 externs; cfg(unix) `UnixStream::shutdown(Both)`); `broker.rs` `SharedSend = Arc<BrokerConn>` + `controller_writer`/`viewer_writer`/`send_frame`/`send_error` routed; `nethost.rs` stream-log/presence sends routed. int = `brain_decouple.rs` `non_drain...\n## MSG-IDENTITY-DISPATCH.md\n 75│\n*76│### W6 — POISON-LOG-SCOPE (F-039; deployah v0.30.6 PASS handoff, **RCA-first**)\n 77│- `REQ-CONN-POISON-DIAL-SCOPE` [impl,unit] — CONN_WRITE_POISONED churn 1:1 with PUMP_PEER_FAIL\n 78│  submit-dials to offline peers; no wedge, pre-existed the rig. Code context: poison_and_cancel\n 79│  (conn.rs:181) logs the SAME loud token for \"exceeded bound **(or failed)**\" — the fast-FAIL\n## W3-DIGEST-GENERATION-UNION-RULING.md\n 68│  items (or folded into its assembly). No `brainproc.rs` touch → **no conflict with\n*69│  PROMOTE-DRAINED**; sequence independently.\n 70│- **Activate** `REQ-DIGEST-GENERATION-SUPERSEDE` (`impl, unit, int`) in the work-start commit\n 71│  (rule 5). Tag evidence on the real code.\n 72│- **Int rig (RED-first):** a two-session span where session B's extracted lines are a superset\n## W3-GATE-VERDICT.md\n 13│| REQ-DIGEST-GENERATION-SUPERSEDE | `7e37e37` | impl+unit+int |\n*14│| REQ-UPDATE-PROMOTE-DRAINED | `e75e856` | impl+unit+int |\n 15│\n 16│## Gate legs — both green\n 17│- **Windows (HFENDULEAM), isolated worktree `.worktrees/gate-e75e856`, fresh target, `CARGO_INCREMENTAL=0`:**\n   │...\n 36│  RED-first int (`two_origin_spanning.rs`); disjoint-span test stays green (supersede no-op).\n*37│- **REQ-UPDATE-PROMOTE-DRAINED** (`brainproc.rs`/`broker.rs`/`daemon.rs`): sound. Readiness\n 38│  LATCHED (`ready_seen`, `try_wait` only pre-ready → ready-then-exit still promotes once\n 39│  drained); promote gated on `ready_seen && old_gen_drained()`; never-drains → WindowElapsedAlive\n 40│  → kill + rollback (conservative — never false-promote onto a wedge). `old_gen_drained` reads\n## W3-SLICE1-REPORT.md\n 52│REQ-DAEMON-STOP-LIVE-SESSION-WARN → REQ-RC-RECONNECT-TRUTH → REQ-UPDATE-ONE-SHOT-FINISH\n*53│(composes on ENDPOINT-SURVIVAL) → REQ-UPDATE-PROMOTE-DRAINED (mechanic-d + false-promote\n 54│rig, the binding W3 activation condition from your W2 verdict).\n## W3-WAVE-GATE-REPORT.md\n 14│| REQ-DIGEST-GENERATION-SUPERSEDE | `7e37e37` | impl+unit+int |\n*15│| **REQ-UPDATE-PROMOTE-DRAINED** | **`e75e856`** | **impl+unit+int** |\n 16│\n*17│## REQ-UPDATE-PROMOTE-DRAINED @e75e856 (mechanic-d — closes the W2 15s residual, KH 7.36)\n 18│\n 19│**Root closed:** W2's stall-evict only BOUNDED the update-apply false-promote window to\n 20│`BRAIN_WRITE_DEADLINE` (15s); a new brain could signal `brain.ready` inside that window\n   │...\n 62│- `xtask check` (docs-drift + token, `CARGO_TARGET_DIR` unset): OK.\n*63│- traceable-reqs check: exit 0; REQ-UPDATE-PROMOTE-DRAINED `+impl +unit +int`.\n 64│- Doyle staged condition — grep OLD rc auto-start assertions: CLEAN (every `rc.rs` reference\n 65│  asserts the NEW connect-only truth; the other hits are unrelated `api` startup hygiene).\n 66│\n   │...\n 73│## KH 7.36\n*74│Updated: the W2 residual is now **CLOSED by W3** (`REQ-UPDATE-PROMOTE-DRAINED`), not just bounded.","linesTruncated":true,"meta":{"limits":{"columnTruncated":{"maxColumn":512}}}},"isError":false,"timestamp":1789523983146}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_3wSJKor4FarX2SnjE6hHDcOR|fc_0d104aba2b0f3edc016aa9f819e68887d1a185111102544b50","toolName":"read","startedAt":"2026-09-16T01:59:53.969Z","args":{"path":"docs/KNOWN-HAZARDS.md:745-770,809-829"},"intent":"Reading paid hazard mechanisms"},"id":"2b9e72e1","parentId":"0b9f48bf","timestamp":"2026-09-16T01:59:53.969Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_0qi33lcdkZ6LzGTgqFCQx6bU|fc_0d104aba2b0f3edc016aa9f819e69c87d19ac73ad0f4539d16","toolName":"read","startedAt":"2026-09-16T01:59:53.969Z","args":{"path":"crates/spt-daemon/src/broker.rs:3806-3844,6405-6465"},"intent":"Reading retirement tally and EOF cleanup"},"id":"a107b88e","parentId":"2b9e72e1","timestamp":"2026-09-16T01:59:53.969Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_4o1UZupfc6teW9rg1imXqknz|fc_0d104aba2b0f3edc016aa9f819e6a487d1a98f20a89396af60","toolName":"read","startedAt":"2026-09-16T01:59:53.969Z","args":{"path":"crates/spt-daemon/src/brainproc.rs:980-1035,1080-1135,265-287,387-410,733-742"},"intent":"Reading refresh trial selection"},"id":"28cb7430","parentId":"a107b88e","timestamp":"2026-09-16T01:59:53.969Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_UexrlMKDmYJFeQfhKQER7iri|fc_0d104aba2b0f3edc016aa9f819e6ac87d1b0413bf87dbc6885","toolName":"grep","startedAt":"2026-09-16T01:59:53.969Z","args":{"path":"crates/spt-daemon/src/nethost.rs"},"intent":"Locating stream writer waits"},"id":"25b8713f","parentId":"28cb7430","timestamp":"2026-09-16T01:59:53.969Z"}
{"type":"message","id":"8bca184b","parentId":"25b8713f","timestamp":"2026-09-16T01:59:53.970Z","message":{"role":"assistant","content":[{"type":"toolCall","id":"call_3wSJKor4FarX2SnjE6hHDcOR|fc_0d104aba2b0f3edc016aa9f819e68887d1a185111102544b50","name":"read","arguments":{"i":"Reading paid hazard mechanisms","path":"docs/KNOWN-HAZARDS.md:745-770,809-829"},"intent":"Reading paid hazard mechanisms"},{"type":"toolCall","id":"call_0qi33lcdkZ6LzGTgqFCQx6bU|fc_0d104aba2b0f3edc016aa9f819e69c87d19ac73ad0f4539d16","name":"read","arguments":{"i":"Reading retirement tally and EOF cleanup","path":"crates/spt-daemon/src/broker.rs:3806-3844,6405-6465"},"intent":"Reading retirement tally and EOF cleanup"},{"type":"toolCall","id":"call_4o1UZupfc6teW9rg1imXqknz|fc_0d104aba2b0f3edc016aa9f819e6a487d1a98f20a89396af60","name":"read","arguments":{"i":"Reading refresh trial selection","path":"crates/spt-daemon/src/brainproc.rs:980-1035,1080-1135,265-287,387-410,733-742"},"intent":"Reading refresh trial selection"},{"type":"toolCall","id":"call_UexrlMKDmYJFeQfhKQER7iri|fc_0d104aba2b0f3edc016aa9f819e6ac87d1b0413bf87dbc6885","name":"grep","arguments":{"i":"Locating stream writer waits","pattern":"stream-sub-writer-poison|fn stream_sub|fn.*writer|ack|SharedSend","path":"crates/spt-daemon/src/nethost.rs","case":true,"gitignore":false,"skip":null},"intent":"Locating stream writer waits"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":10922,"output":218,"cacheRead":53120,"cacheWrite":0,"totalTokens":64260,"cost":{"input":0.10922000000000001,"output":0.0109,"cacheRead":0.05312,"cacheWrite":0,"total":0.17324}},"stopReason":"toolUse","timestamp":1789523983179,"responseId":"resp_0d104aba2b0f3edc016aa9f80f7fa487d1a3c96d85ba509d12","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"fc_0d104aba2b0f3edc016aa9f819e68887d1a185111102544b50","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading paid hazard mechanisms\",\"path\":\"docs/KNOWN-HAZARDS.md:745-770,809-829\"}","call_id":"call_3wSJKor4FarX2SnjE6hHDcOR","internal_chat_message_metadata_passthrough":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"},"name":"read","metadata":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"}},{"id":"fc_0d104aba2b0f3edc016aa9f819e69c87d19ac73ad0f4539d16","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading retirement tally and EOF cleanup\",\"path\":\"crates/spt-daemon/src/broker.rs:3806-3844,6405-6465\"}","call_id":"call_0qi33lcdkZ6LzGTgqFCQx6bU","internal_chat_message_metadata_passthrough":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"},"name":"read","metadata":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"}},{"id":"fc_0d104aba2b0f3edc016aa9f819e6a487d1a98f20a89396af60","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading refresh trial selection\",\"path\":\"crates/spt-daemon/src/brainproc.rs:980-1035,1080-1135,265-287,387-410,733-742\"}","call_id":"call_4o1UZupfc6teW9rg1imXqknz","internal_chat_message_metadata_passthrough":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"},"name":"read","metadata":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"}},{"id":"fc_0d104aba2b0f3edc016aa9f819e6ac87d1b0413bf87dbc6885","type":"function_call","status":"completed","arguments":"{\"i\":\"Locating stream writer waits\",\"pattern\":\"stream-sub-writer-poison|fn stream_sub|fn.*writer|ack|SharedSend\",\"path\":\"crates/spt-daemon/src/nethost.rs\",\"case\":true,\"gitignore\":false,\"skip\":null}","call_id":"call_UexrlMKDmYJFeQfhKQER7iri","internal_chat_message_metadata_passthrough":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"},"name":"grep","metadata":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"}}]},"duration":10788.60400000005,"ttft":10749.744699999923,"completedAt":1789523993968,"contextSnapshot":{"promptTokens":64042,"nonMessageTokens":10480,"compactionEpoch":0}}}
{"type":"message","id":"784d67da","parentId":"8bca184b","timestamp":"2026-09-16T01:59:53.973Z","message":{"role":"toolResult","toolCallId":"call_4o1UZupfc6teW9rg1imXqknz|fc_0d104aba2b0f3edc016aa9f819e6a487d1a98f20a89396af60","toolName":"read","content":[{"type":"text","text":"    // message-DELIVERY axis; the orthogonal live-agent **Psyche** lifecycle (the\n    // run_pulse_loop production caller) is hosted by `livehost::spawn_live_host`\n    // below, keyed on the live ENDPOINT not the broker session (CONTEXT:168/177).\n    // CURSOR-ONLY resume (UPDATE-WEDGE round 3, REQ-BRAIN-RESUME-NO-CONN-DEADLOCK,\n    // doyle Option A-split): the supervised daemon brain is a NON-DRAINING consumer of\n    // this conn (the heartbeat below reads it only in net_status()/sessions()\n    // drain-and-discard loops), so it MUST NOT subscribe — an undrained subscription on\n    // this request/reply conn is the SharedSend self-deadlock that re-wedged v0.30.4\n    // (brain.rs resume_sessions deadlock warning). `resume_session_cursors` re-establishes\n    // the per-session dedup cursors + returns the ids WITHOUT subscribing; the subscribing\n    // `resume_sessions` stays for DRAINING callers (the resume int tests / the forward\n    // live-agent-adapter drained carrier). Empty today (the daemon hosts no PTY sessions).\n    // [impl->REQ-BRAIN-RESUME-NO-CONN-DEADLOCK]\n    match brain.resume_session_cursors() {\n        Ok(ids) if !ids.is_empty() => {\n            spt_proto::emit_line_err!(\n                \"BRAIN_RESUMED: re-established {} session cursor(s): {ids:?}\",\n                ids.len()\n            )\n        }\n        Ok(_) => {}\n        Err(e) => spt_proto::emit_line_err!(\"BRAIN_RESUME_NONFATAL: {e} — no sessions resumed, continuing\"),\n    }\n…\n    loop {\n…\n        if consumer_gate(status.enabled, consumers_started) {\n…\n            spt_proto::emit_line_err!(\n                \"BRAIN_NET_CONSUMERS_UP: net enabled (node {hex}) — dispatcher + peer pump started\"\n            );\n        }\n        // DRIVE the broker's controller-liveness reap every heartbeat (UPDATE-WEDGE\n        // fix, counter-54). A `KIND_SESSIONS` poll is what runs the broker's\n        // `reap_dead_controller` sweep (broker.rs), which stall-evicts a black-holed\n        // LOCAL controller conn a HARD-KILLED prior brain generation left behind (its\n        // hosted PTYs keep writing to a dead pipe → the broker-side writer BLOCKS, never\n        // EOFs). Nothing else polls `KIND_SESSIONS` during an isolated brain-trial\n        // window, so without this the wedge never drains: the promotion DRAINED gate\n        // (`old_gen_drained()` = `!any_local_controller_wedged()`, a PURE READ that never\n        // drives the evict) stays false for the full trial → \"alive but never ready\"\n        // rollback (the 2026-07-09 field freeze), AND the same wedge black-holes recovery\n        // (the late `BRAIN_SUBSCRIBER_STALL_EVICT`). Driven on the FIRST heartbeat and on\n        // BOTH the trial candidate AND the rollback/recovery brain (one loop covers both);\n        // best-effort — a transient sessions error must never fell the heartbeat (the\n        // net_status probe above already owns broker-liveness). Self-applies: it drives\n        // the CURRENT (older) broker's already-shipped reap verb, so a new brain un-strands\n        // its own trial against a live older broker — no coordinated broker restart.\n        // [impl->REQ-UPDATE-TRIAL-DRAIN-DRIVE]\n        let _ = brain.sessions();\n        let _ = write_ready(generation);\n    }\n…\n    // [impl->REQ-UPDATE-PROMOTE-DRAINED]\n    fn old_gen_drained(&self) -> bool {\n        // Drained ⟺ no hosted session has a wedged LOCAL (brain-own) controller conn.\n        // No broker (net-less / bind-failed) → no control plane → drained. A broker\n        // with no wedged local controller (the common case, incl. no daemon-hosted\n        // sessions) reads drained at once → the normal fast promote on `brain.ready`.\n        self.broker\n            .as_ref()\n            .is_none_or(|b| !b.any_local_controller_wedged())\n    }\n…\n    // Broker-owned generation custody (ADR-0018 V2, KH-2.4): the counter lives in\n    // the never-restarting supervisor and increments on EVERY spawn (cold or\n    // respawn), so the broker — which observes every respawn — is the\n    // authoritative source, not the outgoing brain. The first spawn is `Cold`;\n    // every supervised respawn here is a `Crash` (an unexpected exit the\n    // supervisor recovers from). A *planned* update respawn stamps `Update`\n    // instead — wired in D3-3, where the apply trigger drives this loop.\n    let mut generation: u64 = 0;\n    let mut reason = StartReason::Cold;\n    // Consecutive pre-ready exits of the CURRENT trial candidate (A8 budget K).\n    // In-memory: a broker restart resets it — acceptable, the trial restarts.\n    let mut pre_ready_exits: u32 = 0;\n    while !stop.load(Ordering::Relaxed) {\n        let started = Instant::now();\n        // Record-driven binary selection + trial latch (A8). The latch is the\n        // unpromoted `AppliedPending` record, NOT the start reason; a `RolledBack`\n        // record selects the last-known-good binary (and is never a trial — it is\n        // the recovery steady state). A missing/corrupt record degrades to \"no\n        // trial, default binary\" — the supervisor never panics on the record.\n        let record = env.applied_state();\n        let binary: Option<PathBuf> = match &record {\n            Some(AppliedRecord::RolledBack { rollback_binary, .. }) => {\n                Some(PathBuf::from(rollback_binary))\n            }\n            _ => None,\n        };\n        let is_trial = matches!(record, Some(AppliedRecord::AppliedPending { .. }));\n        // A9: clear `brain.ready` immediately before a trial spawn — the prior\n        // child is already waited-on dead (below), so the only stamp that can\n        // exist post-clear is THIS trial child's. Closes the cross-lifetime hole\n        // that same-generation equality alone cannot (a stale gen-N file from a\n        // pre-reboot lifetime matching a fresh gen-N trial).\n        if is_trial {\n            env.clear_ready();\n        }\n        // Did this cycle end because of a planned update restart (Q3) rather than\n        // an unexpected crash? The broker — observing the cycle it initiated — is\n        // the authoritative source of that distinction (V2).\n        let mut planned = false;\n        // Publish the generation about to be spawned BEFORE the child exists, so\n        // the coordinator-image announce (which races the spawn) always meets the\n        // current value — and so the PREVIOUS generation's recorded image stops\n        // answering the moment its child is replaced.\n        // [impl->REQ-UPDATE-RUNNING-IMAGE-SURFACE]\n        note_supervised_generation(generation);\n        match spawn_child(generation, reason, binary.as_deref()) {\n            Ok(mut child) => {\n                // The trial gate runs first for a candidate on trial; on\n                // promotion it falls through to the normal wait loop for the SAME\n                // child. A non-trial spawn skips straight to normal supervision.\n                let mut supervise = !is_trial;\n                if is_trial {\n                    match run_trial(&mut child, generation, window, env, stop) {\n                        TrialStep::Stopped => {\n                            let _ = child.kill();\n                            let _ = child.wait();\n…\n                        }\n…\n                        TrialStep::PreReadyExit => {\n…\n                            if pre_ready_exits >= TRIAL_PREREADY_BUDGET {\n…\n                            } else {\n…\n                            }\n                            generation += 1;\n                            reason = StartReason::Crash;\n                            sleep_backoff(backoff, stop);\n                            continue;\n                        }\n                        TrialStep::WindowElapsedAlive => {\n                            // A11: KILL the alive-but-never-ready candidate BEFORE\n                            // spawning the rollback binary — two live brains would\n                            // break every single-brain seed/socket/IPC assumption.\n                            spt_proto::emit_line_err!(\"BRAIN_TRIAL_TIMEOUT: candidate alive but never ready — killing before rollback\");\n                            let _ = child.kill();\n                            let _ = child.wait();\n                            rollback(env, &record);\n                            pre_ready_exits = 0;\n                            generation += 1;\n                            // The rollback binary is known-good — respawn promptly,\n                            // no crash-penalty backoff.\n                            reason = StartReason::Crash;\n                            backoff = base;\n                            continue;\n                        }\n                    }\n                }\n                if supervise {\n                    loop {\n                        if stop.load(Ordering::Relaxed) {\n                            let _ = child.kill();\n                            let _ = child.wait();\n                            return;\n                        }\n                        // A planned brain-restart request (apply swapped the\n                        // binary, ADR-0018 D3-3): drop this brain so the respawn\n                        // lands on the new executable. A deliberate kill, NOT a\n                        // crash — stamp it `update`.\n                        if restart.take() {\n                            spt_proto::emit_line_err!(\"BRAIN_UPDATE_RESTART: planned restart — cycling brain onto the swapped binary\");\n                            let _ = child.kill();\n                            let _ = child.wait();\n                            planned = true;\n                            break;\n                        }\n                        match child.try_wait() {\n                            Ok(Some(status)) => {\n                                spt_proto::emit_line_err!(\"BRAIN_EXIT: brain child exited ({status}) — respawning\");\n                                break;\n                            }\n                            Ok(None) => thread::sleep(TICK),\n                            Err(e) => {\n                                spt_proto::emit_line_err!(\"BRAIN_WAIT_FAIL: {e}\");\n                                break;\n                            }\n                        }\n                    }\n                }\n            }\n…\n        }\n…\n    }"}],"details":{"displayContent":{"text":"    // message-DELIVERY axis; the orthogonal live-agent **Psyche** lifecycle (the\n    // run_pulse_loop production caller) is hosted by `livehost::spawn_live_host`\n    // below, keyed on the live ENDPOINT not the broker session (CONTEXT:168/177).\n    // CURSOR-ONLY resume (UPDATE-WEDGE round 3, REQ-BRAIN-RESUME-NO-CONN-DEADLOCK,\n    // doyle Option A-split): the supervised daemon brain is a NON-DRAINING consumer of\n    // this conn (the heartbeat below reads it only in net_status()/sessions()\n    // drain-and-discard loops), so it MUST NOT subscribe — an undrained subscription on\n    // this request/reply conn is the SharedSend self-deadlock that re-wedged v0.30.4\n    // (brain.rs resume_sessions deadlock warning). `resume_session_cursors` re-establishes\n    // the per-session dedup cursors + returns the ids WITHOUT subscribing; the subscribing\n    // `resume_sessions` stays for DRAINING callers (the resume int tests / the forward\n    // live-agent-adapter drained carrier). Empty today (the daemon hosts no PTY sessions).\n    // [impl->REQ-BRAIN-RESUME-NO-CONN-DEADLOCK]\n    match brain.resume_session_cursors() {\n        Ok(ids) if !ids.is_empty() => {\n            spt_proto::emit_line_err!(\n                \"BRAIN_RESUMED: re-established {} session cursor(s): {ids:?}\",\n                ids.len()\n            )\n        }\n        Ok(_) => {}\n        Err(e) => spt_proto::emit_line_err!(\"BRAIN_RESUME_NONFATAL: {e} — no sessions resumed, continuing\"),\n    }\n…\n    loop {\n…\n        if consumer_gate(status.enabled, consumers_started) {\n…\n            spt_proto::emit_line_err!(\n                \"BRAIN_NET_CONSUMERS_UP: net enabled (node {hex}) — dispatcher + peer pump started\"\n            );\n        }\n        // DRIVE the broker's controller-liveness reap every heartbeat (UPDATE-WEDGE\n        // fix, counter-54). A `KIND_SESSIONS` poll is what runs the broker's\n        // `reap_dead_controller` sweep (broker.rs), which stall-evicts a black-holed\n        // LOCAL controller conn a HARD-KILLED prior brain generation left behind (its\n        // hosted PTYs keep writing to a dead pipe → the broker-side writer BLOCKS, never\n        // EOFs). Nothing else polls `KIND_SESSIONS` during an isolated brain-trial\n        // window, so without this the wedge never drains: the promotion DRAINED gate\n        // (`old_gen_drained()` = `!any_local_controller_wedged()`, a PURE READ that never\n        // drives the evict) stays false for the full trial → \"alive but never ready\"\n        // rollback (the 2026-07-09 field freeze), AND the same wedge black-holes recovery\n        // (the late `BRAIN_SUBSCRIBER_STALL_EVICT`). Driven on the FIRST heartbeat and on\n        // BOTH the trial candidate AND the rollback/recovery brain (one loop covers both);\n        // best-effort — a transient sessions error must never fell the heartbeat (the\n        // net_status probe above already owns broker-liveness). Self-applies: it drives\n        // the CURRENT (older) broker's already-shipped reap verb, so a new brain un-strands\n        // its own trial against a live older broker — no coordinated broker restart.\n        // [impl->REQ-UPDATE-TRIAL-DRAIN-DRIVE]\n        let _ = brain.sessions();\n        let _ = write_ready(generation);\n    }\n…\n    // [impl->REQ-UPDATE-PROMOTE-DRAINED]\n    fn old_gen_drained(&self) -> bool {\n        // Drained ⟺ no hosted session has a wedged LOCAL (brain-own) controller conn.\n        // No broker (net-less / bind-failed) → no control plane → drained. A broker\n        // with no wedged local controller (the common case, incl. no daemon-hosted\n        // sessions) reads drained at once → the normal fast promote on `brain.ready`.\n        self.broker\n            .as_ref()\n            .is_none_or(|b| !b.any_local_controller_wedged())\n    }\n…\n    // Broker-owned generation custody (ADR-0018 V2, KH-2.4): the counter lives in\n    // the never-restarting supervisor and increments on EVERY spawn (cold or\n    // respawn), so the broker — which observes every respawn — is the\n    // authoritative source, not the outgoing brain. The first spawn is `Cold`;\n    // every supervised respawn here is a `Crash` (an unexpected exit the\n    // supervisor recovers from). A *planned* update respawn stamps `Update`\n    // instead — wired in D3-3, where the apply trigger drives this loop.\n    let mut generation: u64 = 0;\n    let mut reason = StartReason::Cold;\n    // Consecutive pre-ready exits of the CURRENT trial candidate (A8 budget K).\n    // In-memory: a broker restart resets it — acceptable, the trial restarts.\n    let mut pre_ready_exits: u32 = 0;\n    while !stop.load(Ordering::Relaxed) {\n        let started = Instant::now();\n        // Record-driven binary selection + trial latch (A8). The latch is the\n        // unpromoted `AppliedPending` record, NOT the start reason; a `RolledBack`\n        // record selects the last-known-good binary (and is never a trial — it is\n        // the recovery steady state). A missing/corrupt record degrades to \"no\n        // trial, default binary\" — the supervisor never panics on the record.\n        let record = env.applied_state();\n        let binary: Option<PathBuf> = match &record {\n            Some(AppliedRecord::RolledBack { rollback_binary, .. }) => {\n                Some(PathBuf::from(rollback_binary))\n            }\n            _ => None,\n        };\n        let is_trial = matches!(record, Some(AppliedRecord::AppliedPending { .. }));\n        // A9: clear `brain.ready` immediately before a trial spawn — the prior\n        // child is already waited-on dead (below), so the only stamp that can\n        // exist post-clear is THIS trial child's. Closes the cross-lifetime hole\n        // that same-generation equality alone cannot (a stale gen-N file from a\n        // pre-reboot lifetime matching a fresh gen-N trial).\n        if is_trial {\n            env.clear_ready();\n        }\n        // Did this cycle end because of a planned update restart (Q3) rather than\n        // an unexpected crash? The broker — observing the cycle it initiated — is\n        // the authoritative source of that distinction (V2).\n        let mut planned = false;\n        // Publish the generation about to be spawned BEFORE the child exists, so\n        // the coordinator-image announce (which races the spawn) always meets the\n        // current value — and so the PREVIOUS generation's recorded image stops\n        // answering the moment its child is replaced.\n        // [impl->REQ-UPDATE-RUNNING-IMAGE-SURFACE]\n        note_supervised_generation(generation);\n        match spawn_child(generation, reason, binary.as_deref()) {\n            Ok(mut child) => {\n                // The trial gate runs first for a candidate on trial; on\n                // promotion it falls through to the normal wait loop for the SAME\n                // child. A non-trial spawn skips straight to normal supervision.\n                let mut supervise = !is_trial;\n                if is_trial {\n                    match run_trial(&mut child, generation, window, env, stop) {\n                        TrialStep::Stopped => {\n                            let _ = child.kill();\n                            let _ = child.wait();\n…\n                        }\n…\n                        TrialStep::PreReadyExit => {\n…\n                            if pre_ready_exits >= TRIAL_PREREADY_BUDGET {\n…\n                            } else {\n…\n                            }\n                            generation += 1;\n                            reason = StartReason::Crash;\n                            sleep_backoff(backoff, stop);\n                            continue;\n                        }\n                        TrialStep::WindowElapsedAlive => {\n                            // A11: KILL the alive-but-never-ready candidate BEFORE\n                            // spawning the rollback binary — two live brains would\n                            // break every single-brain seed/socket/IPC assumption.\n                            spt_proto::emit_line_err!(\"BRAIN_TRIAL_TIMEOUT: candidate alive but never ready — killing before rollback\");\n                            let _ = child.kill();\n                            let _ = child.wait();\n                            rollback(env, &record);\n                            pre_ready_exits = 0;\n                            generation += 1;\n                            // The rollback binary is known-good — respawn promptly,\n                            // no crash-penalty backoff.\n                            reason = StartReason::Crash;\n                            backoff = base;\n                            continue;\n                        }\n                    }\n                }\n                if supervise {\n                    loop {\n                        if stop.load(Ordering::Relaxed) {\n                            let _ = child.kill();\n                            let _ = child.wait();\n                            return;\n                        }\n                        // A planned brain-restart request (apply swapped the\n                        // binary, ADR-0018 D3-3): drop this brain so the respawn\n                        // lands on the new executable. A deliberate kill, NOT a\n                        // crash — stamp it `update`.\n                        if restart.take() {\n                            spt_proto::emit_line_err!(\"BRAIN_UPDATE_RESTART: planned restart — cycling brain onto the swapped binary\");\n                            let _ = child.kill();\n                            let _ = child.wait();\n                            planned = true;\n                            break;\n                        }\n                        match child.try_wait() {\n                            Ok(Some(status)) => {\n                                spt_proto::emit_line_err!(\"BRAIN_EXIT: brain child exited ({status}) — respawning\");\n                                break;\n                            }\n                            Ok(None) => thread::sleep(TICK),\n                            Err(e) => {\n                                spt_proto::emit_line_err!(\"BRAIN_WAIT_FAIL: {e}\");\n                                break;\n                            }\n                        }\n                    }\n                }\n            }\n…\n        }\n…\n    }","startLine":265,"lineNumbers":[265,266,267,268,269,270,271,272,273,274,275,276,277,278,279,280,281,282,283,284,285,286,287,null,371,null,383,null,387,388,389,390,391,392,393,394,395,396,397,398,399,400,401,402,403,404,405,406,407,408,409,410,null,733,734,735,736,737,738,739,740,741,742,null,980,981,982,983,984,985,986,987,988,989,990,991,992,993,994,995,996,997,998,999,1000,1001,1002,1003,1004,1005,1006,1007,1008,1009,1010,1011,1012,1013,1014,1015,1016,1017,1018,1019,1020,1021,1022,1023,1024,1025,1026,1027,1028,1029,1030,1031,1032,1033,1034,1035,null,1037,null,1067,null,1070,null,1078,null,1080,1081,1082,1083,1084,1085,1086,1087,1088,1089,1090,1091,1092,1093,1094,1095,1096,1097,1098,1099,1100,1101,1102,1103,1104,1105,1106,1107,1108,1109,1110,1111,1112,1113,1114,1115,1116,1117,1118,1119,1120,1121,1122,1123,1124,1125,1126,1127,1128,1129,1130,1131,1132,1133,1134,1135,null,1137,null,1156]},"fileSize":122482,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\crates\\spt-daemon\\src\\brainproc.rs"}}},"isError":false,"timestamp":1789523993972}}
{"type":"message","id":"15f8bd78","parentId":"784d67da","timestamp":"2026-09-16T01:59:53.977Z","message":{"role":"toolResult","toolCallId":"call_0qi33lcdkZ6LzGTgqFCQx6bU|fc_0d104aba2b0f3edc016aa9f819e69c87d19ac73ad0f4539d16","toolName":"read","content":[{"type":"text","text":"fn viewer_writer(\n…\n) {\n…\n    // isolation is preserved). REQ-VIEWER-SKIP-TO-LIVE-ON-EVICT.\n    // [impl->REQ-VIEWER-SKIP-TO-LIVE-ON-EVICT]\n    if evicted.load(Ordering::Acquire) {\n        let marker = evicted_envelope(session_id);\n        let _ = send.write(&marker);\n    }\n}\n\n/// Round-4 observability bridge: a controller writer whose bounded conn write\n/// was DEADLINE-retired (`TimedOut` — the conn watchdog aborted a wedged\n/// write, or a gate-wait behind one exhausted the bound) IS the\n/// brain-subscriber stall-evict fact the W2 tally counts. Pre-round-4 that\n/// wedge aged until a logical sampler (`stall_evict_controller`) observed it;\n/// the physical bound now fires FIRST, so the tally is recorded here — same\n/// event class, same `daemon status` surface. An ORGANIC write failure (peer\n/// death, any non-TimedOut error) stays uncounted, exactly as before.\n// [impl->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\n// [impl->REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]\nfn note_controller_write_retired(sid: u64, conn: u64, e: &io::Error) {\n    if e.kind() != io::ErrorKind::TimedOut {\n        return;\n    }\n    // conn id + stamps (F-039 leg b). [impl->REQ-CONN-POISON-ATTRIBUTION]\n    spt_proto::emit_line_err!(\n        \"BRAIN_SUBSCRIBER_STALL_EVICT:{sid} conn={conn} {} — controller conn write \\\n         exceeded its bound; physical conn poisoned + retired (writer released, EOF \\\n         cleanup follows) [REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]\",\n        crate::conn::log_stamp()\n    );\n    record_stall_evict();\n}\n\n/// A message to the CONTROLLER writer thread (W1). `Output` carries a sequenced\n/// chunk and advances the authoritative `delivered_through` on a successful\n/// write. `Exit` is the session's terminal frame, enqueued by the exit waiter\n/// BEHIND every queued output frame (ADR-0043 decision 1,\n/// REQ-BROKER-OUTPUT-BEFORE-EXIT) — the writer is the sole FIFO sequencer, so\n/// Exit can never overtake output and strand the child's final frame. (The\n/// former `Control` out-of-band notice variant is gone — the terminal\n…\n        loop {\n…\n            match env.kind.as_str() {\n…\n                // stream also leaves this conn's sweep list, so the exit\n                // sweep does not re-detach a seat that is already gone.\n                // [impl->REQ-SEAT-LIFETIME-BOUNDED]\n                KIND_NET_STREAM_UNSUBSCRIBE => {\n                    match self.dispatch_net_stream_unsubscribe(env, &send) {\n                        Ok(id) => my_stream_subs.retain(|s| *s != id),\n                        Err(msg) => send_error(&send, &msg),\n                    }\n                }\n                KIND_NET_PRESENCE_SUBSCRIBE => {\n                    match self.dispatch_net_presence_subscribe(env, &send) {\n                        Ok(()) => my_presence_sub = true,\n                        Err(msg) => send_error(&send, &msg),\n                    }\n                }\n                KIND_BRAIN_RESTART => self.dispatch_brain_restart(&send),\n                // An unknown kind is tolerated (forward-compat): report, don't die.\n                other => send_error(&send, &format!(\"unknown command kind {other:?}\")),\n            }\n        }\n\n        // Brain gone: abort any seal ceremony this conn was a party to —\n        // requester (nobody left to answer) or controller (no human left at\n        // the overlay) — BEFORE the roles detach, so the other party is\n        // answered promptly rather than waiting on a corpse.\n        // [impl->REQ-SEAL-CEREMONY-ESC-CANCEL]\n        self.abort_ceremonies_for_conn(send.id());\n        // Brain gone: detach as subscriber but leave the sessions running.\n        // Clone the log Arcs out from under the sessions lock FIRST (ADR-0038\n        // Amendment fix 8, KH 7.12 kin): `detach_if` re-stamps info.json\n        // (`viewer_count`/`driven_by`) and emits a lifecycle record — file +\n        // conn I/O that must never run under the shared sessions table lock,\n        // or one slow perch write stalls every spawn/attach/kill dispatch.\n        // [impl->REQ-DISPATCH-HYGIENE-TELEMETRY]\n        let logs: Vec<Arc<Mutex<OutputLog>>> = {\n            let sessions = recover(&self.sessions);\n            my_subs\n                .iter()\n                .filter_map(|id| sessions.get(id).map(|h| Arc::clone(&h.log)))\n                .collect()\n        };\n        for log in &logs {\n            recover_log(log).detach_if(&send);\n        }\n        // Same for net streams: the streams + their rings persist (D4b).\n        if let Some(host) = self.net.get() {\n            for id in &my_stream_subs {\n                host.detach_stream_sub(*id, &send);\n            }\n            // ConnectionBound teardown (ADR-0040 decision 6): rows THIS conn\n            // opened as ConnectionBound die with it. FIN the stream first — the\n            // remote target's read pump sees EOF, `serve_attach` runs\n            // `detach_session` under its own serve generation (fix-6 validated:\n            // a stale opener's late close can never evict a newer controller),\n            // and the controller/viewer stamps clear — then terminal-retire the\n            // local row (eligibility + physical teardown, the W1 machinery).\n            // This is exactly the emphasys C2 hole: an rc viewport brain dying\n            // WITHOUT FIN left its attach stream restart-durable forever, the\n            // target never saw EOF, and CONTROLLED survived even a full broker\n            // restart. Durable rows (inter-brain) are untouched — brain\n            // disconnect must NEVER globally retire (brain-swap continuity).\n…\n        }"}],"details":{"displayContent":{"text":"fn viewer_writer(\n…\n) {\n…\n    // isolation is preserved). REQ-VIEWER-SKIP-TO-LIVE-ON-EVICT.\n    // [impl->REQ-VIEWER-SKIP-TO-LIVE-ON-EVICT]\n    if evicted.load(Ordering::Acquire) {\n        let marker = evicted_envelope(session_id);\n        let _ = send.write(&marker);\n    }\n}\n\n/// Round-4 observability bridge: a controller writer whose bounded conn write\n/// was DEADLINE-retired (`TimedOut` — the conn watchdog aborted a wedged\n/// write, or a gate-wait behind one exhausted the bound) IS the\n/// brain-subscriber stall-evict fact the W2 tally counts. Pre-round-4 that\n/// wedge aged until a logical sampler (`stall_evict_controller`) observed it;\n/// the physical bound now fires FIRST, so the tally is recorded here — same\n/// event class, same `daemon status` surface. An ORGANIC write failure (peer\n/// death, any non-TimedOut error) stays uncounted, exactly as before.\n// [impl->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]\n// [impl->REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]\nfn note_controller_write_retired(sid: u64, conn: u64, e: &io::Error) {\n    if e.kind() != io::ErrorKind::TimedOut {\n        return;\n    }\n    // conn id + stamps (F-039 leg b). [impl->REQ-CONN-POISON-ATTRIBUTION]\n    spt_proto::emit_line_err!(\n        \"BRAIN_SUBSCRIBER_STALL_EVICT:{sid} conn={conn} {} — controller conn write \\\n         exceeded its bound; physical conn poisoned + retired (writer released, EOF \\\n         cleanup follows) [REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]\",\n        crate::conn::log_stamp()\n    );\n    record_stall_evict();\n}\n\n/// A message to the CONTROLLER writer thread (W1). `Output` carries a sequenced\n/// chunk and advances the authoritative `delivered_through` on a successful\n/// write. `Exit` is the session's terminal frame, enqueued by the exit waiter\n/// BEHIND every queued output frame (ADR-0043 decision 1,\n/// REQ-BROKER-OUTPUT-BEFORE-EXIT) — the writer is the sole FIFO sequencer, so\n/// Exit can never overtake output and strand the child's final frame. (The\n/// former `Control` out-of-band notice variant is gone — the terminal\n…\n        loop {\n…\n            match env.kind.as_str() {\n…\n                // stream also leaves this conn's sweep list, so the exit\n                // sweep does not re-detach a seat that is already gone.\n                // [impl->REQ-SEAT-LIFETIME-BOUNDED]\n                KIND_NET_STREAM_UNSUBSCRIBE => {\n                    match self.dispatch_net_stream_unsubscribe(env, &send) {\n                        Ok(id) => my_stream_subs.retain(|s| *s != id),\n                        Err(msg) => send_error(&send, &msg),\n                    }\n                }\n                KIND_NET_PRESENCE_SUBSCRIBE => {\n                    match self.dispatch_net_presence_subscribe(env, &send) {\n                        Ok(()) => my_presence_sub = true,\n                        Err(msg) => send_error(&send, &msg),\n                    }\n                }\n                KIND_BRAIN_RESTART => self.dispatch_brain_restart(&send),\n                // An unknown kind is tolerated (forward-compat): report, don't die.\n                other => send_error(&send, &format!(\"unknown command kind {other:?}\")),\n            }\n        }\n\n        // Brain gone: abort any seal ceremony this conn was a party to —\n        // requester (nobody left to answer) or controller (no human left at\n        // the overlay) — BEFORE the roles detach, so the other party is\n        // answered promptly rather than waiting on a corpse.\n        // [impl->REQ-SEAL-CEREMONY-ESC-CANCEL]\n        self.abort_ceremonies_for_conn(send.id());\n        // Brain gone: detach as subscriber but leave the sessions running.\n        // Clone the log Arcs out from under the sessions lock FIRST (ADR-0038\n        // Amendment fix 8, KH 7.12 kin): `detach_if` re-stamps info.json\n        // (`viewer_count`/`driven_by`) and emits a lifecycle record — file +\n        // conn I/O that must never run under the shared sessions table lock,\n        // or one slow perch write stalls every spawn/attach/kill dispatch.\n        // [impl->REQ-DISPATCH-HYGIENE-TELEMETRY]\n        let logs: Vec<Arc<Mutex<OutputLog>>> = {\n            let sessions = recover(&self.sessions);\n            my_subs\n                .iter()\n                .filter_map(|id| sessions.get(id).map(|h| Arc::clone(&h.log)))\n                .collect()\n        };\n        for log in &logs {\n            recover_log(log).detach_if(&send);\n        }\n        // Same for net streams: the streams + their rings persist (D4b).\n        if let Some(host) = self.net.get() {\n            for id in &my_stream_subs {\n                host.detach_stream_sub(*id, &send);\n            }\n            // ConnectionBound teardown (ADR-0040 decision 6): rows THIS conn\n            // opened as ConnectionBound die with it. FIN the stream first — the\n            // remote target's read pump sees EOF, `serve_attach` runs\n            // `detach_session` under its own serve generation (fix-6 validated:\n            // a stale opener's late close can never evict a newer controller),\n            // and the controller/viewer stamps clear — then terminal-retire the\n            // local row (eligibility + physical teardown, the W1 machinery).\n            // This is exactly the emphasys C2 hole: an rc viewport brain dying\n            // WITHOUT FIN left its attach stream restart-durable forever, the\n            // target never saw EOF, and CONTROLLED survived even a full broker\n            // restart. Durable rows (inter-brain) are untouched — brain\n            // disconnect must NEVER globally retire (brain-swap continuity).\n…\n        }","startLine":3757,"lineNumbers":[3757,null,3763,null,3806,3807,3808,3809,3810,3811,3812,3813,3814,3815,3816,3817,3818,3819,3820,3821,3822,3823,3824,3825,3826,3827,3828,3829,3830,3831,3832,3833,3834,3835,3836,3837,3838,3839,3840,3841,3842,3843,3844,null,6085,null,6091,null,6405,6406,6407,6408,6409,6410,6411,6412,6413,6414,6415,6416,6417,6418,6419,6420,6421,6422,6423,6424,6425,6426,6427,6428,6429,6430,6431,6432,6433,6434,6435,6436,6437,6438,6439,6440,6441,6442,6443,6444,6445,6446,6447,6448,6449,6450,6451,6452,6453,6454,6455,6456,6457,6458,6459,6460,6461,6462,6463,6464,6465,null,6476]},"fileSize":877537,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\crates\\spt-daemon\\src\\broker.rs"}}},"isError":false,"timestamp":1789523993977}}
{"type":"message","id":"c506b717","parentId":"15f8bd78","timestamp":"2026-09-16T01:59:54.054Z","message":{"role":"toolResult","toolCallId":"call_3wSJKor4FarX2SnjE6hHDcOR|fc_0d104aba2b0f3edc016aa9f819e68887d1a185111102544b50","toolName":"read","content":[{"type":"text","text":"### 7.35 The cached ceremony-clock NTP offset must NOT survive an OS clock STEP — an offset measured against the pre-step clock strands every pairing for the TTL  `[REQ-HAZARD-CEREMONY-CLOCK-STEP]`\n- **Failure (paid-for, field evidence — the enlyzeam BIGNET-join RCA 2026-07-06):** enlyzeam's `w32time` was STOPPED, so its system clock drifted +3m30s; the pairing ceremony offset (REQ-PAIR-8, `ntp.rs`) correctly corrected the TOTP clock — until an operator `w32tm /resync` STEPPED the OS clock back −210s under the LIVE daemon. The cached offset had been measured against the OLD (drifted) clock and its snapshot was monotonic-only (`when: Option<Instant>`), so the step was invisible to the cache: for up to the 15-min TTL the ceremony clock stayed ~7 TOTP steps in the past, every `subnet join` returned NO_SEED_HOLDER while fresh-process probes met in <500ms, and only a daemon bounce (which forced a fresh query) healed it. Timeline-proven: refresh cadence 15:0…\n…\n- **spt-core mapping:** `spt_net::net::pairing::ntp` — `CachedOffset{when, wall_at}` pair + `refresh_needed`/`clock_stepped`/`signed_secs` (step-detect) + `force_refresh`/`invalidate` (the meet-exhaustion re-query seam); `spt_daemon::pairhost::meet_seed_holder` via `sweep_then_final_retry` (force-refresh + one final sweep on exhaustion). unit = `clock_step_forces_refresh_via_injected_reads` (injected mono/wall reads: a wall step forward OR back with mono barely moved re-queries; agreeing clocks within TTL do not; TTL-expiry and never-queried also refresh — no 15-min sleep) + `invalidate_clears_the_snapshot` + `meet_retries_once_after_refresh_on_exhaustion`/`meet_final_retry_failing_surrenders_once` (exactly one post-refresh retry, then surrender with the ric…\n- **Source:** JOIN-TRUTH W1/D3 (doyle /diagnose enlyzeam RCA; the offset defeated by a live clock step is the D3 of four defects, D1 multihome + D2 loud-fail + D4 verbose-clock the siblings).\n\n### 7.36 The broker control plane and PTY fan-out must NEVER block on a single subscriber connection — a suspended brain conn must not wedge control  `[REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]`\n<!-- [doc->REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE] -->\n- **Failure (paid-for, rig-CONFIRMED 2026-07-06/07 — `NtSuspendProcess` on the brain, no update involved):** a controller's writer thread does a BLOCKING socket write to its brain subscriber conn. When that brain is suspended (or black-holed) the write never returns. The output-driven eviction path (`append` → `mark_controller_gone`, bounded by 7.12's `CONTROLLER_WRITE_DEADLINE`) only fires on NEW output, and `reap_dead_controller` (7.29 companion) only caught a writer that had EXITED — so a controller BLOCKED (not exited) on an IDLE session was evicted by NEITHER. Within seconds: attached output froze, `detach` did not release the control stamp (release routed through the wedged conn), reattach was REFUSED (`controlled-by` latched), `rc --take` hung, and `d…\n- **Invariant:** no broker client ever observes a wedge because the BRAIN stalled. (1) The controller's writer publishes its IN-FLIGHT-write window (`write_blocked_since`: `Some(Instant)` set immediately before the blocking `write_frame`, `None` after — the mutex never held ACROSS the write); a reader sees a `Some` older than `BRAIN_WRITE_DEADLINE` (15 s = 3× the 7.12 controller bound, measured full-with-**zero writer progress**, so a slow-but-DRAINING brain never trips it) as a WEDGED writer and stall-evicts it. Eviction is triggered TIME-based, not output-driven: at the take/reattach path (`resolve_subscribe`, so control self-heals against broker truth even on an idle session) AND at the `KIND_SESSIONS` reap. (2) The stamp release is BROKER-side — `clear_c…\n- **spt-core mapping:** `broker.rs` — `ControllerSink::write_blocked_since` + `controller_writer` (marks the in-flight window around both the initial-batch and live-loop writes) + `controller_write_stalled` (pure predicate over injected `now`/`deadline`) + `stall_evict_controller` (Inline stamp at `resolve_subscribe`, Deferred stamp at `reap_dead_controller` so the reap closure does no I/O under the shared sessions lock) + `record_stall_evict`/`stall_evict_stats` (the tally) + `dispatch_stall_evicts` (`KIND_STALL_EVICTS` IPC); `brain.rs` `Brain::stall_evicts`; `cli.rs` `render_stall_evict_line` (daemon-status surface, public wording). unit = `controller_write_stall_predicate_distinguishes_wedged_from_idle` + `stalled_incumbent_is_evicted_on_reattach_and_cont…\n- **Source:** LIFECYCLE-TRUTH W2 (doyle rig `NtSuspendProcess` root + design ruling `docs/W2-DESIGN-RULING.md` @60ec0ed — minimal-plus scope, `BRAIN_WRITE_DEADLINE` 15 s, passive resubscribe + observability rider; todlando build 2026-07-07). The v0.27→v0.28 update-wedge night's flagship defect.\n\n### 7.37 The Layer-1 settle-gate must RE-ARM per delivery on an observable PTY — a mid-session reader reattach (`/clear`) re-creates the head-swallow race  `[REQ-HAZARD-INJECT-SETTLE-REARM]`\n<!-- [doc->REQ-HAZARD-INJECT-SETTLE-REARM] -->\n- **Failure (paid-for, field-CONFIRMED on 0.29.0 — doyle diagnosis + perri screenshot `WindowsTerminal_6iSjya8pMt.png`):** the shipped W5-A settle-gate (7-of-W5, `REQ-INJECT-MULTILINE-INTEGRITY`) gated Layer 1 behind a worker-local ONE-SHOT (`settled_once`) on the premise that the head-swallow race is a STARTUP condition (the input reader not yet attached after spawn). That premise is FALSE: a mid-session `/clear` re-enters the harness's raw-mode input reader, re-attaching it and RE-CREATING the pre-settle window — but the one-shot had already fired at spawn, so `settle_before_inject` was SKIPPED and the head was eaten again. A checkpoint-wake payload injected right after `/clear` arrived head-truncated (mid-path `spt/Cargo.toml)`); echo-verify (Layer 2) is …\n- **Invariant:** the settle-gate re-arms before EVERY delivery on an OBSERVABLE (echoing/interactive) PTY — the bug-prone class, where re-settling is cheap (a tick or two once the reader answers) and necessary (a mid-session reader reattach must be re-confirmed). The steady-state settle is latch-skipped ONLY where the probe is known UNOBSERVABLE on this PTY (a non-echoing ConPTY whose DSR answer never surfaces): there is no reader-reattach race to guard, and each settle would burn the full `INJECT_SETTLE` deadline. The class is discriminated by the settle's OWN return value — `settle_before_inject` returns `true` iff the session ring advanced (probe observed) — latched on the FIRST attempt only; a RE-DRIVE (`attempt > 1`, reached only on a swallowed head) AL…\n- **spt-core mapping:** `broker.rs` `run_inject_worker` — the `settled_once: bool` one-shot is replaced by a `probe_unobservable: bool` latch (init `false` = assume observable → settle first delivery); the gate is the pure `should_settle(attempt, probe_unobservable) = attempt > 1 || !probe_unobservable`, and the first-attempt settle's `observed` return sets `probe_unobservable = !observed`. `settle_before_inject` gains a `bool` return (observed vs timed-out). unit = `should_settle_rearms_on_observable_pty` (observable PTY re-settles each delivery; unobservable-probe PTY skips the steady-state settle; a re-drive settles in both classes — RED-first: restore the one-shot and the observable-PTY assert flips). No int edit: the `inject_control_wedge.rs` mocks (`fi…\n- **Source:** post-0.29.0 field fix (doyle diagnosis + design ruling `BUILD-SPEC-inject-settle-rearm`; todlando build 2026-07-08). An impl bug in the shipped W5-A Layer-1 fix, not a new invariant class — the settle-gate was correct but under-armed.\n\n### 7.38 EVERY write on a physical broker connection is bounded + cancelable + poison-on-failure — no writer holds the send gate across an UNBOUNDED OS write  `[REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]`\n<!-- [doc->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK] -->\n- **Failure (paid-for, field capture + deterministic Windows repro 2026-07-09 — hertz RCA, UPDATE-WEDGE round-4):** `controller_writer` held the `SharedSend = Arc<Mutex<SendHalf>>` guard ACROSS a blocking `write_frame`. On Windows, `interprocess 2.4.2` routes the send to `WriteFileEx` + `SleepEx(INFINITE, alertable)` — this local-socket send path exposes **no supported write timeout** (`set_timeout` → `Unsupported`). An `rc --take` controller consumer that stopped reading blocked the write **indefinitely** (~127.95 s in the capture, released only when a brain restart tore the connections down). The 7.36 logical `stall_evict_controller` removed the controller ROLE but neither canceled the in-flight pipe write, closed the physical connection, nor called `Cance…\n- **Invariant:** every write on a physical broker connection rides ONE broker-owned bounded/cancelable framed-write primitive with an **independent out-of-band aborter**. The `SendHalf` never leaves the conn object (a bypass cannot compile). An ABSOLUTE deadline stamped at write entry covers BOTH the serialized gate-wait AND OS write completion; a per-conn watchdog fires at that deadline and aborts the in-flight op — never relying on the write returning or on opportunistic stall-evict sampling. On deadline / partial write / cancellation / unknown completion: (1) poison the whole physical conn idempotently, (2) abort read+write so `handle_conn` reaches existing EOF cleanup, (3) wait for the canceled op to report completion before releasing its buffer, (4) NEV…\n- **spt-core mapping:** `spt-daemon/src/conn.rs` `BrokerConn` (Condvar-gated serialized write gate + poisoned state + per-in-flight op id + watchdog + `abort_physical()`: cfg(windows) `CancelIoEx` → completion handshake → `DisconnectNamedPipe` raw kernel32 externs; cfg(unix) `UnixStream::shutdown(Both)`); `broker.rs` `SharedSend = Arc<BrokerConn>` + `controller_writer`/`viewer_writer`/`send_frame`/`send_error` routed; `nethost.rs` stream-log/presence sends routed. int = `brain_decouple.rs` `non_draining_controller_stall_evict_releases_writer_and_connection` (`#[cfg(windows)]`, the hertz 9-step: real PTY burst, black-holed remote controller, exactly one logical stall-evict via an independent draining conn, the black-holed conn's `handle_conn` finishes ≤2 s of…\n- **Source:** UPDATE-WEDGE round-4 (hertz root doc `docs/UPDATE-WEDGE-ROOT-CAUSE.md` @f8596ca; doyle gate ruling `docs/UPDATE-WEDGE-RCA-DISPATCH-todlando.md`; todlando build 2026-07-09). The residual 7.36 left open: logical eviction without physical retirement.\n\n---\n\n…\n### 7.43 One wedged stream subscriber must NEVER stall stream serving — replay halts at the first failed write, a poisoned subscriber is removed, and recovery machinery must not manufacture new replay victims  `[REQ-HAZARD-REDISPATCH-STALL]`\n<!-- [doc->REQ-HAZARD-REDISPATCH-STALL] -->\n- **Failure (paid-for, hertz field RCA 2026-07-16 — live v0.34 boxes, recurrent 20–30s PTY/RC freezes, DISPATCH tails 17–62s):** a COMPOSITION, not one new timer. The dispatcher's retryable claims (500ms/1s ×3) re-drove the opener fallback whose guard arm is a broad `Err(_)` (`dispatch.rs:414` — comment intends old-broker-only; catches transport timeout/EOF/poison), each pass installing a throwaway peek subscriber. `StreamLog::attach` synchronously replays the ENTIRE retained ring under the per-stream mutex, `let _ = sub.write(&frame)` discards errors, iteration continues past failure, and the poisoned subscriber STAYS INSTALLED — one wedged subscriber conn = serial 15s bounded-write poison windows (33 observed, all 15,000–15,154 ms — the round-4 SharedSend …\n- **Invariant:** subscriber I/O never holds the `StreamLog` mutex (enqueue-under-lock, I/O outside, bounded per-subscriber queue; overflow = detach + resume-from-cursor); replay and fan-out halt at the FIRST failed subscriber write and the failed subscriber is REMOVED (attach + append + finish, all sites, `PresenceLog` mirrors included); the ring-peek fallback fires ONLY on the explicit unsupported-verb/old-broker answer; a deadline-poisoned replay is circuit-broken (global backoff, no reinstall until the prior subscriber is fully gone — an upgrade of bounded retry, never a revert to burn-the-claim abandonment); attach/detach validate ownership/generation so a stale worker never detaches a replacement controller (covers the unfinished-stale steal shape that …\n- **spt-core mapping:** `spt-daemon/src/nethost.rs` `StreamLog::attach`/`push`/`finish` + `PresenceLog` mirrors + serve/wire worker pairing; `spt-daemon/src/dispatch.rs` `first_line` fallback arm + claim retry/backoff + worker spawn; SharedSend bounded-write deadline (round-4, KH 7.7 kin).\n- **Source:** REDISPATCH-STALL W1 (hertz v0.34 field RCA 2026-07-16, doyle seam-verified same day; ADR-0038 Amendment 2026-07-16; `.claude/reports/2026-07-16-redispatch-stall/rca-summary.md`; distinct from `update --restart`'s bounded 30s rc reconnect loop).\n\n### 7.44 Streams and seats on a long-lived connection must have bounded lifetime — one-way rows terminal at FIN, seats released at serve completion, no per-chunk full-state rewrites in a drain loop  `[REQ-HAZARD-REGISTRY-STALL]`\n<!-- [doc->REQ-HAZARD-REGISTRY-STALL] -->\n- **Failure (paid-for, hertz post-close v0.36 field RCA 2026-07-17 — live prod box):** FOUR compounding consequences of \"physical teardown waits for conn close\" on connections that never close. The registry pump opens one fresh stream per feed (~30s) on the persistent pump conn; nothing retires the rows (the dispatcher — sole `retire_stream` caller — skips `initiated_locally`) → sender history unbounded. `stream_infos` filters `retired` only → every dispatcher poll serializes O(history) rows over IPC. `serve_registry_feed` ran `write_snapshots` (full rewrite of every subnet registry + heard.meta) up to TWICE PER 64KiB CHUNK, synchronously in the brain's event loop → brain stops draining IPC → broker seat writer blocks → 15s SharedSend bound → CONN_WRITE_POIS…\n- **Invariant:** a one-way (fire-and-forget) family's row is TERMINAL at successful FIN — the sender retires it; eligibility filtering happens server-side before IPC serialization; feed application is transactional per feed (snapshot writes O(feeds), never O(chunks × record-kinds) — no synchronous full-state rewrite inside a per-chunk drain iteration); a deadline-poisoned one-way replay retires terminal at a per-stream strike budget (safe: the next pump round re-advertises); every completed serve releases its SubscriberSeat (unsubscribe verb — thread joined, cursor dropped), all families; the exe self-hash is captured once per brain process (that IS the resident-bytes truthfulness contract).\n- **spt-core mapping:** `spt-daemon/src/pump/mod.rs` `push_feed` (sender retire); `spt-daemon/src/nethost.rs` `stream_infos` (server filter) + SubscriberSeat teardown + thread naming; `spt-daemon/src/dispatch.rs` `serve_registry_feed` (transactional apply) + worker completion unsubscribe + strike budget; `spt-daemon/src/registryhost.rs` `write_snapshots` call sites; `spt-daemon/src/brainproc.rs` `write_ready`/`current_exe_hash`.\n- **Source:** REGISTRY-LIFECYCLE W1 (hertz RCA `.claude/reports/2026-07-17-registry-stall-rca/README.md`, doyle-verified all legs 2026-07-17; ADR-0040; supersedes the v0.36 clean-window general-health ruling — the Attach-family fix stands, the class recurred for Registry).\n\n### 7.45 Endpoint lifecycle state converges to truth from every death path — no optimistic online without authority, no surviving control stamps, no immortal wake intent, no untruthful create  `[REQ-HAZARD-ENDPOINT-LIFECYCLE]`\n<!-- [doc->REQ-HAZARD-ENDPOINT-LIFECYCLE] -->\n- **Failure (paid-for, three hertz reports + operator field 2026-07-16):** four families, one root shape — lifecycle state written by multiple non-converging paths, optimistic stamps never verified. (a) `cmd_listen` stamped `status=online` from manifest capability alone → dead-PID hybrid rows survived EVERY restart (reconcile skipped them by state, cleanup gated on `controllable=true`, readers trusted `status=online` past dead-PID rejection). (b) Definitive session death stamped `status=offline` only, never `rest_state` → preserved `Active` intent = outstanding wake order → zombie WAKE_RESUME loops across generations (perri field). (c) `KIND_SESSIONS` poll snapshot raced the exit-waiter reap → stale poll relatched `controlled=true` on a dead session. (d) `--…\n- **Invariant:** online is stamped only from persisted state + hosting authority; legacy rows self-heal after a SUCCESSFUL broker query (broker failure ≠ empty set); control-stamp cleanup runs for every endpoint absent from session truth regardless of state/controllability, split from offline classification; definitive death = ONE atomic store mutation (offline + rest_state=suspended + clear dormant_since_ms, same info.json write) — never a rest-event edge, never reader-side inference; stamp writes are session/generation-validated so a pre-reap snapshot can't overwrite post-reap truth; fresh-create is a broker-atomic policy with a TYPED conflict (distinct wire kind, loud on N-1), never `Spawned(existing)`; cycle verbs share one liveness authority.\n- **spt-core mapping:** `spt/src/api/startup.rs` `cmd_listen` creator gate; `spt-daemon/src/livehost.rs` reconcile split + decide_resume; `spt-daemon/src/lifecycle.rs` `mark_offline` → terminal-normalize; `spt/src/cli.rs` `cmd_stop`/create dispatch; `spt-daemon/src/broker.rs` SpawnReq policy split + KIND_SPAWN_FRESH + dup-guard probe; KH 7.12/7.27/7.29 kin.\n- **Source:** REGISTRY-LIFECYCLE W2/W3 (hertz emphasys triple + rest-normalize + spawn-fresh reports 2026-07-16, doyle seam-verified same day; ADR-0041; operator deployah-wedge field evidence).\n\n### 7.46 An rc surface answers from live session authority, never a stale persisted projection — and a resuming perch is UNBOUND, not offline  `[REQ-HAZARD-RC-ATTACH-TRUTH]`\n…\n\n\n[Some lines truncated to 768 chars]"}],"details":{"displayContent":{"text":"### 7.35 The cached ceremony-clock NTP offset must NOT survive an OS clock STEP — an offset measured against the pre-step clock strands every pairing for the TTL  `[REQ-HAZARD-CEREMONY-CLOCK-STEP]`\n- **Failure (paid-for, field evidence — the enlyzeam BIGNET-join RCA 2026-07-06):** enlyzeam's `w32time` was STOPPED, so its system clock drifted +3m30s; the pairing ceremony offset (REQ-PAIR-8, `ntp.rs`) correctly corrected the TOTP clock — until an operator `w32tm /resync` STEPPED the OS clock back −210s under the LIVE daemon. The cached offset had been measured against the OLD (drifted) clock and its snapshot was monotonic-only (`when: Option<Instant>`), so the step was invisible to the cache: for up to the 15-min TTL the ceremony clock stayed ~7 TOTP steps in the past, every `subnet join` returned NO_SEED_HOLDER while fresh-process probes met in <500ms, and only a daemon bounce (which forced a fresh query) healed it. Timeline-proven: refresh cadence 15:0…\n…\n- **spt-core mapping:** `spt_net::net::pairing::ntp` — `CachedOffset{when, wall_at}` pair + `refresh_needed`/`clock_stepped`/`signed_secs` (step-detect) + `force_refresh`/`invalidate` (the meet-exhaustion re-query seam); `spt_daemon::pairhost::meet_seed_holder` via `sweep_then_final_retry` (force-refresh + one final sweep on exhaustion). unit = `clock_step_forces_refresh_via_injected_reads` (injected mono/wall reads: a wall step forward OR back with mono barely moved re-queries; agreeing clocks within TTL do not; TTL-expiry and never-queried also refresh — no 15-min sleep) + `invalidate_clears_the_snapshot` + `meet_retries_once_after_refresh_on_exhaustion`/`meet_final_retry_failing_surrenders_once` (exactly one post-refresh retry, then surrender with the ric…\n- **Source:** JOIN-TRUTH W1/D3 (doyle /diagnose enlyzeam RCA; the offset defeated by a live clock step is the D3 of four defects, D1 multihome + D2 loud-fail + D4 verbose-clock the siblings).\n\n### 7.36 The broker control plane and PTY fan-out must NEVER block on a single subscriber connection — a suspended brain conn must not wedge control  `[REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE]`\n<!-- [doc->REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE] -->\n- **Failure (paid-for, rig-CONFIRMED 2026-07-06/07 — `NtSuspendProcess` on the brain, no update involved):** a controller's writer thread does a BLOCKING socket write to its brain subscriber conn. When that brain is suspended (or black-holed) the write never returns. The output-driven eviction path (`append` → `mark_controller_gone`, bounded by 7.12's `CONTROLLER_WRITE_DEADLINE`) only fires on NEW output, and `reap_dead_controller` (7.29 companion) only caught a writer that had EXITED — so a controller BLOCKED (not exited) on an IDLE session was evicted by NEITHER. Within seconds: attached output froze, `detach` did not release the control stamp (release routed through the wedged conn), reattach was REFUSED (`controlled-by` latched), `rc --take` hung, and `d…\n- **Invariant:** no broker client ever observes a wedge because the BRAIN stalled. (1) The controller's writer publishes its IN-FLIGHT-write window (`write_blocked_since`: `Some(Instant)` set immediately before the blocking `write_frame`, `None` after — the mutex never held ACROSS the write); a reader sees a `Some` older than `BRAIN_WRITE_DEADLINE` (15 s = 3× the 7.12 controller bound, measured full-with-**zero writer progress**, so a slow-but-DRAINING brain never trips it) as a WEDGED writer and stall-evicts it. Eviction is triggered TIME-based, not output-driven: at the take/reattach path (`resolve_subscribe`, so control self-heals against broker truth even on an idle session) AND at the `KIND_SESSIONS` reap. (2) The stamp release is BROKER-side — `clear_c…\n- **spt-core mapping:** `broker.rs` — `ControllerSink::write_blocked_since` + `controller_writer` (marks the in-flight window around both the initial-batch and live-loop writes) + `controller_write_stalled` (pure predicate over injected `now`/`deadline`) + `stall_evict_controller` (Inline stamp at `resolve_subscribe`, Deferred stamp at `reap_dead_controller` so the reap closure does no I/O under the shared sessions lock) + `record_stall_evict`/`stall_evict_stats` (the tally) + `dispatch_stall_evicts` (`KIND_STALL_EVICTS` IPC); `brain.rs` `Brain::stall_evicts`; `cli.rs` `render_stall_evict_line` (daemon-status surface, public wording). unit = `controller_write_stall_predicate_distinguishes_wedged_from_idle` + `stalled_incumbent_is_evicted_on_reattach_and_cont…\n- **Source:** LIFECYCLE-TRUTH W2 (doyle rig `NtSuspendProcess` root + design ruling `docs/W2-DESIGN-RULING.md` @60ec0ed — minimal-plus scope, `BRAIN_WRITE_DEADLINE` 15 s, passive resubscribe + observability rider; todlando build 2026-07-07). The v0.27→v0.28 update-wedge night's flagship defect.\n\n### 7.37 The Layer-1 settle-gate must RE-ARM per delivery on an observable PTY — a mid-session reader reattach (`/clear`) re-creates the head-swallow race  `[REQ-HAZARD-INJECT-SETTLE-REARM]`\n<!-- [doc->REQ-HAZARD-INJECT-SETTLE-REARM] -->\n- **Failure (paid-for, field-CONFIRMED on 0.29.0 — doyle diagnosis + perri screenshot `WindowsTerminal_6iSjya8pMt.png`):** the shipped W5-A settle-gate (7-of-W5, `REQ-INJECT-MULTILINE-INTEGRITY`) gated Layer 1 behind a worker-local ONE-SHOT (`settled_once`) on the premise that the head-swallow race is a STARTUP condition (the input reader not yet attached after spawn). That premise is FALSE: a mid-session `/clear` re-enters the harness's raw-mode input reader, re-attaching it and RE-CREATING the pre-settle window — but the one-shot had already fired at spawn, so `settle_before_inject` was SKIPPED and the head was eaten again. A checkpoint-wake payload injected right after `/clear` arrived head-truncated (mid-path `spt/Cargo.toml)`); echo-verify (Layer 2) is …\n- **Invariant:** the settle-gate re-arms before EVERY delivery on an OBSERVABLE (echoing/interactive) PTY — the bug-prone class, where re-settling is cheap (a tick or two once the reader answers) and necessary (a mid-session reader reattach must be re-confirmed). The steady-state settle is latch-skipped ONLY where the probe is known UNOBSERVABLE on this PTY (a non-echoing ConPTY whose DSR answer never surfaces): there is no reader-reattach race to guard, and each settle would burn the full `INJECT_SETTLE` deadline. The class is discriminated by the settle's OWN return value — `settle_before_inject` returns `true` iff the session ring advanced (probe observed) — latched on the FIRST attempt only; a RE-DRIVE (`attempt > 1`, reached only on a swallowed head) AL…\n- **spt-core mapping:** `broker.rs` `run_inject_worker` — the `settled_once: bool` one-shot is replaced by a `probe_unobservable: bool` latch (init `false` = assume observable → settle first delivery); the gate is the pure `should_settle(attempt, probe_unobservable) = attempt > 1 || !probe_unobservable`, and the first-attempt settle's `observed` return sets `probe_unobservable = !observed`. `settle_before_inject` gains a `bool` return (observed vs timed-out). unit = `should_settle_rearms_on_observable_pty` (observable PTY re-settles each delivery; unobservable-probe PTY skips the steady-state settle; a re-drive settles in both classes — RED-first: restore the one-shot and the observable-PTY assert flips). No int edit: the `inject_control_wedge.rs` mocks (`fi…\n- **Source:** post-0.29.0 field fix (doyle diagnosis + design ruling `BUILD-SPEC-inject-settle-rearm`; todlando build 2026-07-08). An impl bug in the shipped W5-A Layer-1 fix, not a new invariant class — the settle-gate was correct but under-armed.\n\n### 7.38 EVERY write on a physical broker connection is bounded + cancelable + poison-on-failure — no writer holds the send gate across an UNBOUNDED OS write  `[REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK]`\n<!-- [doc->REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK] -->\n- **Failure (paid-for, field capture + deterministic Windows repro 2026-07-09 — hertz RCA, UPDATE-WEDGE round-4):** `controller_writer` held the `SharedSend = Arc<Mutex<SendHalf>>` guard ACROSS a blocking `write_frame`. On Windows, `interprocess 2.4.2` routes the send to `WriteFileEx` + `SleepEx(INFINITE, alertable)` — this local-socket send path exposes **no supported write timeout** (`set_timeout` → `Unsupported`). An `rc --take` controller consumer that stopped reading blocked the write **indefinitely** (~127.95 s in the capture, released only when a brain restart tore the connections down). The 7.36 logical `stall_evict_controller` removed the controller ROLE but neither canceled the in-flight pipe write, closed the physical connection, nor called `Cance…\n- **Invariant:** every write on a physical broker connection rides ONE broker-owned bounded/cancelable framed-write primitive with an **independent out-of-band aborter**. The `SendHalf` never leaves the conn object (a bypass cannot compile). An ABSOLUTE deadline stamped at write entry covers BOTH the serialized gate-wait AND OS write completion; a per-conn watchdog fires at that deadline and aborts the in-flight op — never relying on the write returning or on opportunistic stall-evict sampling. On deadline / partial write / cancellation / unknown completion: (1) poison the whole physical conn idempotently, (2) abort read+write so `handle_conn` reaches existing EOF cleanup, (3) wait for the canceled op to report completion before releasing its buffer, (4) NEV…\n- **spt-core mapping:** `spt-daemon/src/conn.rs` `BrokerConn` (Condvar-gated serialized write gate + poisoned state + per-in-flight op id + watchdog + `abort_physical()`: cfg(windows) `CancelIoEx` → completion handshake → `DisconnectNamedPipe` raw kernel32 externs; cfg(unix) `UnixStream::shutdown(Both)`); `broker.rs` `SharedSend = Arc<BrokerConn>` + `controller_writer`/`viewer_writer`/`send_frame`/`send_error` routed; `nethost.rs` stream-log/presence sends routed. int = `brain_decouple.rs` `non_draining_controller_stall_evict_releases_writer_and_connection` (`#[cfg(windows)]`, the hertz 9-step: real PTY burst, black-holed remote controller, exactly one logical stall-evict via an independent draining conn, the black-holed conn's `handle_conn` finishes ≤2 s of…\n- **Source:** UPDATE-WEDGE round-4 (hertz root doc `docs/UPDATE-WEDGE-ROOT-CAUSE.md` @f8596ca; doyle gate ruling `docs/UPDATE-WEDGE-RCA-DISPATCH-todlando.md`; todlando build 2026-07-09). The residual 7.36 left open: logical eviction without physical retirement.\n\n---\n\n…\n### 7.43 One wedged stream subscriber must NEVER stall stream serving — replay halts at the first failed write, a poisoned subscriber is removed, and recovery machinery must not manufacture new replay victims  `[REQ-HAZARD-REDISPATCH-STALL]`\n<!-- [doc->REQ-HAZARD-REDISPATCH-STALL] -->\n- **Failure (paid-for, hertz field RCA 2026-07-16 — live v0.34 boxes, recurrent 20–30s PTY/RC freezes, DISPATCH tails 17–62s):** a COMPOSITION, not one new timer. The dispatcher's retryable claims (500ms/1s ×3) re-drove the opener fallback whose guard arm is a broad `Err(_)` (`dispatch.rs:414` — comment intends old-broker-only; catches transport timeout/EOF/poison), each pass installing a throwaway peek subscriber. `StreamLog::attach` synchronously replays the ENTIRE retained ring under the per-stream mutex, `let _ = sub.write(&frame)` discards errors, iteration continues past failure, and the poisoned subscriber STAYS INSTALLED — one wedged subscriber conn = serial 15s bounded-write poison windows (33 observed, all 15,000–15,154 ms — the round-4 SharedSend …\n- **Invariant:** subscriber I/O never holds the `StreamLog` mutex (enqueue-under-lock, I/O outside, bounded per-subscriber queue; overflow = detach + resume-from-cursor); replay and fan-out halt at the FIRST failed subscriber write and the failed subscriber is REMOVED (attach + append + finish, all sites, `PresenceLog` mirrors included); the ring-peek fallback fires ONLY on the explicit unsupported-verb/old-broker answer; a deadline-poisoned replay is circuit-broken (global backoff, no reinstall until the prior subscriber is fully gone — an upgrade of bounded retry, never a revert to burn-the-claim abandonment); attach/detach validate ownership/generation so a stale worker never detaches a replacement controller (covers the unfinished-stale steal shape that …\n- **spt-core mapping:** `spt-daemon/src/nethost.rs` `StreamLog::attach`/`push`/`finish` + `PresenceLog` mirrors + serve/wire worker pairing; `spt-daemon/src/dispatch.rs` `first_line` fallback arm + claim retry/backoff + worker spawn; SharedSend bounded-write deadline (round-4, KH 7.7 kin).\n- **Source:** REDISPATCH-STALL W1 (hertz v0.34 field RCA 2026-07-16, doyle seam-verified same day; ADR-0038 Amendment 2026-07-16; `.claude/reports/2026-07-16-redispatch-stall/rca-summary.md`; distinct from `update --restart`'s bounded 30s rc reconnect loop).\n\n### 7.44 Streams and seats on a long-lived connection must have bounded lifetime — one-way rows terminal at FIN, seats released at serve completion, no per-chunk full-state rewrites in a drain loop  `[REQ-HAZARD-REGISTRY-STALL]`\n<!-- [doc->REQ-HAZARD-REGISTRY-STALL] -->\n- **Failure (paid-for, hertz post-close v0.36 field RCA 2026-07-17 — live prod box):** FOUR compounding consequences of \"physical teardown waits for conn close\" on connections that never close. The registry pump opens one fresh stream per feed (~30s) on the persistent pump conn; nothing retires the rows (the dispatcher — sole `retire_stream` caller — skips `initiated_locally`) → sender history unbounded. `stream_infos` filters `retired` only → every dispatcher poll serializes O(history) rows over IPC. `serve_registry_feed` ran `write_snapshots` (full rewrite of every subnet registry + heard.meta) up to TWICE PER 64KiB CHUNK, synchronously in the brain's event loop → brain stops draining IPC → broker seat writer blocks → 15s SharedSend bound → CONN_WRITE_POIS…\n- **Invariant:** a one-way (fire-and-forget) family's row is TERMINAL at successful FIN — the sender retires it; eligibility filtering happens server-side before IPC serialization; feed application is transactional per feed (snapshot writes O(feeds), never O(chunks × record-kinds) — no synchronous full-state rewrite inside a per-chunk drain iteration); a deadline-poisoned one-way replay retires terminal at a per-stream strike budget (safe: the next pump round re-advertises); every completed serve releases its SubscriberSeat (unsubscribe verb — thread joined, cursor dropped), all families; the exe self-hash is captured once per brain process (that IS the resident-bytes truthfulness contract).\n- **spt-core mapping:** `spt-daemon/src/pump/mod.rs` `push_feed` (sender retire); `spt-daemon/src/nethost.rs` `stream_infos` (server filter) + SubscriberSeat teardown + thread naming; `spt-daemon/src/dispatch.rs` `serve_registry_feed` (transactional apply) + worker completion unsubscribe + strike budget; `spt-daemon/src/registryhost.rs` `write_snapshots` call sites; `spt-daemon/src/brainproc.rs` `write_ready`/`current_exe_hash`.\n- **Source:** REGISTRY-LIFECYCLE W1 (hertz RCA `.claude/reports/2026-07-17-registry-stall-rca/README.md`, doyle-verified all legs 2026-07-17; ADR-0040; supersedes the v0.36 clean-window general-health ruling — the Attach-family fix stands, the class recurred for Registry).\n\n### 7.45 Endpoint lifecycle state converges to truth from every death path — no optimistic online without authority, no surviving control stamps, no immortal wake intent, no untruthful create  `[REQ-HAZARD-ENDPOINT-LIFECYCLE]`\n<!-- [doc->REQ-HAZARD-ENDPOINT-LIFECYCLE] -->\n- **Failure (paid-for, three hertz reports + operator field 2026-07-16):** four families, one root shape — lifecycle state written by multiple non-converging paths, optimistic stamps never verified. (a) `cmd_listen` stamped `status=online` from manifest capability alone → dead-PID hybrid rows survived EVERY restart (reconcile skipped them by state, cleanup gated on `controllable=true`, readers trusted `status=online` past dead-PID rejection). (b) Definitive session death stamped `status=offline` only, never `rest_state` → preserved `Active` intent = outstanding wake order → zombie WAKE_RESUME loops across generations (perri field). (c) `KIND_SESSIONS` poll snapshot raced the exit-waiter reap → stale poll relatched `controlled=true` on a dead session. (d) `--…\n- **Invariant:** online is stamped only from persisted state + hosting authority; legacy rows self-heal after a SUCCESSFUL broker query (broker failure ≠ empty set); control-stamp cleanup runs for every endpoint absent from session truth regardless of state/controllability, split from offline classification; definitive death = ONE atomic store mutation (offline + rest_state=suspended + clear dormant_since_ms, same info.json write) — never a rest-event edge, never reader-side inference; stamp writes are session/generation-validated so a pre-reap snapshot can't overwrite post-reap truth; fresh-create is a broker-atomic policy with a TYPED conflict (distinct wire kind, loud on N-1), never `Spawned(existing)`; cycle verbs share one liveness authority.\n- **spt-core mapping:** `spt/src/api/startup.rs` `cmd_listen` creator gate; `spt-daemon/src/livehost.rs` reconcile split + decide_resume; `spt-daemon/src/lifecycle.rs` `mark_offline` → terminal-normalize; `spt/src/cli.rs` `cmd_stop`/create dispatch; `spt-daemon/src/broker.rs` SpawnReq policy split + KIND_SPAWN_FRESH + dup-guard probe; KH 7.12/7.27/7.29 kin.\n- **Source:** REGISTRY-LIFECYCLE W2/W3 (hertz emphasys triple + rest-normalize + spawn-fresh reports 2026-07-16, doyle seam-verified same day; ADR-0041; operator deployah-wedge field evidence).\n\n### 7.46 An rc surface answers from live session authority, never a stale persisted projection — and a resuming perch is UNBOUND, not offline  `[REQ-HAZARD-RC-ATTACH-TRUTH]`\n…\n","startLine":742,"lineNumbers":[742,743,null,745,746,747,748,749,750,751,752,753,754,755,756,757,758,759,760,761,762,763,764,765,766,767,768,769,770,null,808,809,810,811,812,813,814,815,816,817,818,819,820,821,822,823,824,825,826,827,828,829,null,835]},"fileSize":360311,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\docs\\KNOWN-HAZARDS.md"},"limits":{"columnTruncated":{"maxColumn":768}}}},"isError":false,"timestamp":1789523994054}}
{"type":"message","id":"b03dcd6d","parentId":"c506b717","timestamp":"2026-09-16T01:59:54.069Z","message":{"role":"toolResult","toolCallId":"call_UexrlMKDmYJFeQfhKQER7iri|fc_0d104aba2b0f3edc016aa9f819e6ac87d1b0413bf87dbc6885","toolName":"grep","content":[{"type":"text","text":" 43|\n*44|use crate::broker::SharedSend;\n 45|\n 46|use crate::effect::Minter;\n 47|use crate::seedproofx::{prove_membership, MembershipSource, RosterExchange};\n...\n 60|/// brain-waiting QUIC await — `dial` / `open_stream` / `send_stream` — is bounded\n*61|/// by this so a black-holed or dead peer fails PROMPTLY with an ORDINARY error\n 62|/// the broker replies, never an unbounded await that wedges the single-threaded\n 63|/// pump (the B-half of REQ-HAZARD-PUMP-IPC-DEADLINE; the 2.2h hfenduleam wedge).\n 64|/// Chosen at **10s**: comfortably above any legitimate LAN/relay connect +\n...\n 71|/// frame's IPC transit + the brain's read scheduling. Test-overridable via\n*72|/// [`NetHost::set_quic_op_timeout`] so the hermetic black-hole tests fail fast.\n 73|const BROKER_QUIC_OP_TIMEOUT_MS: u64 = 10_000;\n 74|\n 75|/// Bounded concurrent-dial cap (PUMP-TRUTH W2, REQ-PUMP-PEER-ISOLATION, G4): a\n...\n 95|    /// Socket bind scope: [`BindScope::All`] in production;\n*96|    /// [`BindScope::Loopback`] in hermetic tests (no real interface touched,\n 97|    /// no per-test-binary Windows firewall prompt).\n 98|    pub scope: BindScope,\n 99|    /// Live source of this node's subnet credentials for the connect-time\n...\n 143|/// fixes 2+3, REQ-STREAMLOG-SUBSCRIBER-DISCIPLINE): frames ENQUEUE under the\n*144|/// log mutex; the conn I/O — the deadline-bounded `SharedSend::write` that can\n 145|/// take 15s against a wedged conn — runs on this seat's OWN writer thread,\n 146|/// OUTSIDE every lock. The field stall was exactly the old shape's serial\n 147|/// replay-under-mutex against a poisoned conn (33 windows, all ~15s).\n...\n 154|struct SubscriberSeat {\n*155|    sub: SharedSend,\n 156|    tx: std::sync::mpsc::SyncSender<Envelope>,\n 157|    /// Writer thread exited (failure or clean drain) — the seat is fully gone\n 158|    /// and a replacement may install (the no-reinstall-until-gone rule,\n...\n 169|impl SubscriberSeat {\n*170|    /// Install a seat: spawn its writer thread and hand back the handle.\n 171|    /// `replay` is the attach snapshot, moved in UNDER the log mutex (a\n 172|    /// memory move, never conn I/O): the writer drains it COMPLETELY before\n 173|    /// touching the live queue, so replay-then-live seq order holds by\n...\n 189|    fn install(\n*190|        sub: SharedSend,\n 191|        replay: Vec<Envelope>,\n 192|        lease: Option<Arc<StreamLease>>,\n 193|        label: &str,\n...\n 212|                    }\n*213|                    writer_sub.lifecycle_event(\"stream-sub-writer-poison\", halt_reason);\n 214|                };\n 215|                // Replay first — halt at the first failed write (Amendment\n 216|                // fix 2), never iterate past a poisoned conn.\n...\n 270|    /// Whether this seat belongs to `sub` (detach correlation).\n*271|    fn is(&self, sub: &SharedSend) -> bool {\n 272|        Arc::ptr_eq(&self.sub, sub)\n 273|    }\n 274|\n...\n 319|\n*320|/// Transport backing a held connection. A real `Quic` connection (the only\n*321|/// kind an inbound handshake can produce), or a broker-minted `Loopback` —\n 322|/// an in-process connection with NO iroh handshake whose identity IS this node,\n*323|/// established BY CONSTRUCTION. A loopback conn can ONLY be minted broker-side\n*324|/// via [`NetHost::dial_loopback`]; it can never arise from an inbound handshake\n 325|/// (a remote always presents its own key and cannot claim to be us). This is\n 326|/// what lets a LOCAL `spt rc` ride the byte-identical `serve_attach` pump (M12\n 327|/// W1.5, B1) — QUIC self-dial is refused by iroh (\"Connecting to ourself is not\n*328|/// supported\"), so local attach rides a broker loopback conn instead.\n 329|enum ConnKind {\n 330|    /// A real QUIC connection — kept alive by ownership here; dropping it closes\n 331|    /// the connection. Cloned (cheap handle) for stream opens.\n 332|    Quic(Connection),\n*333|    /// A broker-minted in-process loopback (no iroh handle, no handshake). Its\n 334|    /// streams are minted eagerly in cross-wired pairs by [`NetHost::open_stream`];\n 335|    /// there is no accept loop and no closed-watcher (it lives for the host).\n*336|    Loopback,\n 337|}\n 338|\n*339|/// One live connection the broker holds (QUIC or loopback).\n 340|struct ConnEntry {\n*341|    /// The transport backing this connection.\n 342|    kind: ConnKind,\n 343|    /// The remote node's stable name (hex pubkey). For QUIC, proven by the\n*344|    /// handshake; for loopback, THIS node's own id (the broker-minted identity).\n 345|    remote_id_hex: String,\n 346|    /// The subnets the peer proved current-epoch seed-knowledge of at connect\n 347|    /// (Mesh-D2, REQ-MESH-1). Empty when seed-proof is disabled (`membership:\n...\n 383|    finished: bool,\n*384|    /// **Retentive backpressure cap** (M11-W3, REQ-SHELL-4), in chunks; `0` ⇒ not\n 385|    /// retentive (ordinary bounded-overwrite stream). A shell-tunnel stream must\n 386|    /// NOT drop bytes — CONTEXT:262 \"tunneled protocols cannot drop frames, so\n 387|    /// congestion surfaces as lag, never loss.\" So a retentive log never cap-evicts\n...\n 431|\n*432|    /// Mark this log retentive with backpressure `cap` (chunks). Idempotent set.\n 433|    fn set_retentive(&mut self, cap: usize) {\n 434|        self.retentive_cap = cap;\n 435|    }\n...\n 466|        self.next_seq += 1;\n*467|        self.ring.push_back((seq, bytes.to_vec()));\n 468|        if self.retentive_cap == 0 {\n 469|            while self.ring.len() > self.cap_chunks {\n 470|                self.ring.pop_front();\n...\n 502|    /// Drain every undrained chunk in order (the shell-tunnel recv cursor, W3).\n*503|    /// Pops the whole ring front-to-back (FIFO by `seq`) and concatenates the\n 504|    /// bytes — chunk boundaries are arbitrary read-pump splits, irrelevant to an\n 505|    /// opaque byte stream. Lossless + ordered: a byte leaves the ring only here,\n 506|    /// once. Returns `(bytes, finished)`; `finished` reflects a clean/torn end so\n...\n 544|    // [impl->REQ-STREAMLOG-SUBSCRIBER-DISCIPLINE]\n*545|    fn begin_attach(&mut self, sub: SharedSend, from_seq: u64) -> io::Result<usize> {\n 546|        if let Some(prior) = &self.subscriber {\n 547|            // No-reinstall-until-gone applies to the POISON class only\n 548|            // (REQ-DISPATCH-FALLBACK-CIRCUIT): a wedged writer may sit in a\n...\n 610|    // [impl->REQ-SEAT-LIFETIME-BOUNDED]\n*611|    fn take_seat_if(&mut self, sub: &SharedSend) -> Option<SubscriberSeat> {\n 612|        if self.subscriber.as_ref().is_some_and(|cur| cur.is(sub)) {\n 613|            self.subscriber.take()\n 614|        } else {\n...\n 620|    /// not unsubscribe a newer brain).\n*621|    fn detach_if(&mut self, sub: &SharedSend) {\n 622|        if let Some(cur) = &self.subscriber {\n 623|            if cur.is(sub) {\n 624|                self.subscriber = None;\n...\n 672|    /// A `connected` event carrying the connection's OBSERVED remote address\n*673|    /// (REQ-CONV-1): the async-dial write-back seed the blocking `net-dialed`\n 674|    /// reply used to carry — now it rides the presence stream so the pump seeds\n 675|    /// `peer-addrs.json` from a non-blocking dial the same way.\n 676|    // [impl->REQ-CONV-1]\n...\n 720|        self.next_seq += 1;\n*721|        self.ring.push_back(ev.clone());\n 722|        while self.ring.len() > self.cap_events {\n 723|            self.ring.pop_front();\n 724|        }\n...\n 744|    // [impl->REQ-STREAMLOG-SUBSCRIBER-DISCIPLINE]\n*745|    fn begin_attach(&mut self, sub: SharedSend, from_seq: u64) -> io::Result<usize> {\n 746|        if let Some(prior) = &self.subscriber {\n 747|            // Poison-class only (see StreamLog::begin_attach): a healthy\n 748|            // prior presence subscriber is displaced by the newer brain.\n...\n 775|    /// not unsubscribe a newer brain).\n*776|    fn detach_if(&mut self, sub: &SharedSend) {\n 777|        if let Some(cur) = &self.subscriber {\n 778|            if cur.is(sub) {\n 779|                self.subscriber = None;\n...\n 788|/// write end of an in-process duplex). That is the M12 W1.5 one-pump invariant\n*789|/// made structural: a loopback attach and a cross-node attach diverge at this\n 790|/// leaf and nowhere above it.\n 791|enum SendHalf {\n 792|    Quic(SendStream),\n*793|    Loopback(WriteHalf<DuplexStream>),\n 794|}\n 795|\n 796|impl SendHalf {\n...\n 803|                .map_err(|e| io::Error::other(e.to_string())),\n*804|            SendHalf::Loopback(w) => w.write_all(bytes).await,\n 805|        }\n 806|    }\n 807|\n...\n 812|            SendHalf::Quic(s) => s.finish().map_err(|e| io::Error::other(e.to_string())),\n*813|            SendHalf::Loopback(w) => w.shutdown().await,\n 814|        }\n 815|    }\n 816|}\n...\n 818|/// A stream's read half — fed into the single read pump in [`register_stream`].\n*819|/// QUIC reads via quinn's chunk API; loopback via the read end of the duplex.\n 820|/// Both append to the same [`StreamLog`], so the subscriber (and everything\n 821|/// above it) cannot tell the transports apart.\n 822|enum RecvHalf {\n 823|    Quic(RecvStream),\n*824|    Loopback(ReadHalf<DuplexStream>),\n 825|}\n 826|\n*827|/// One live stream the broker holds (QUIC or loopback): the send half (brain\n 828|/// writes ride it over IPC) + the sequenced read log.\n 829|struct StreamEntry {\n 830|    conn_id: u64,\n...\n 842|    /// Wakes the read pump when a retentive (tunnel) stream's ring drains below\n*843|    /// its backpressure cap (M11-W3, REQ-SHELL-4) — see [`StreamLog::retentive_cap`].\n 844|    /// Unused by ordinary (non-retentive) streams.\n 845|    room: Arc<tokio::sync::Notify>,\n 846|    /// Retired = invisible to dispatch eligibility (ADR-0038 A), NOT torn down:\n...\n 904|/// the path address string, or `null` when id-only / path-less (nothing worth\n*905|/// seeding over discovery). Shared by [`register_conn`]'s presence write-back\n 906|/// and [`NetHost::conn_remote_addr_json`] so both seed identically.\n 907|// [impl->REQ-CONV-1]\n 908|fn observed_addr_json(conn: &Connection) -> serde_json::Value {\n...\n 1031|    // `connected` event can never race behind its own `disconnected` (D4c). The\n*1032|    // event carries the OBSERVED remote addr (REQ-CONV-1 seed write-back) so a\n 1033|    // non-blocking submit-dial seeds `peer-addrs.json` like the blocking reply\n 1034|    // did (id-only / path-less → `null`, nothing worth seeding).\n 1035|    let observed = observed_addr_json(&conn);\n...\n 1128|    // replayable either way. The transport differs only at this read leaf\n*1129|    // (QUIC chunk API vs the loopback duplex's `AsyncRead`); both append to the\n 1130|    // same `StreamLog`, so nothing above the pump can tell them apart.\n 1131|    //\n*1132|    // Backpressure (M11-W3, REQ-SHELL-4): a retentive (tunnel) log never evicts,\n 1133|    // so before pulling the next chunk the pump WAITS while the undrained ring is\n 1134|    // at its `retentive_cap`. The unread bytes stay in the transport's buffer, so\n*1135|    // QUIC stream flow-control (or the loopback duplex's bounded buffer) lags the\n 1136|    // SENDER — \"lag, never loss\" (CONTEXT:262). A `drain` wakes `room`. Ordinary\n 1137|    // streams have `retentive_cap == 0`, so this gate is a no-op for them.\n 1138|    tokio::spawn(async move {\n...\n 1149|            },\n*1150|            RecvHalf::Loopback(mut recv) => {\n 1151|                let mut buf = vec![0u8; STREAM_READ_CHUNK];\n 1152|                loop {\n 1153|                    await_ring_room(&log, &room).await;\n...\n 1168|\n*1169|/// Backpressure gate for a retentive (tunnel) read pump: park until the undrained\n 1170|/// ring has room below `retentive_cap`, woken by a `drain` on `room`. A no-op for\n 1171|/// an ordinary stream (`retentive_cap == 0`). Holds no lock across the await.\n 1172|// [impl->REQ-SHELL-4]\n...\n 1213|    roster: Arc<Mutex<Option<RosterExchange>>>,\n*1214|    /// The broker's single in-process loopback connection (M12 W1.5, B1),\n*1215|    /// minted lazily on the first [`Self::dial_loopback`] and reused thereafter\n 1216|    /// — a local `spt rc` needs exactly one. `Some(conn_id)` once minted; the\n 1217|    /// row lives for the host (no handshake to lose, no closed-watcher).\n*1218|    loopback_conn: Mutex<Option<u64>>,\n*1219|    /// Serializes [`Self::open_loopback_pair`] (M11-W3) so the peer-row discovery\n*1220|    /// (a before/after diff on the loopback conn's inbound rows) sees exactly the\n*1221|    /// one row this open minted — concurrent pair-opens on the shared loopback conn\n 1222|    /// would otherwise race the diff.\n*1223|    loopback_pair_lock: Mutex<()>,\n 1224|    /// The broker-side QUIC-op deadline in ms (REQ-HAZARD-BROKER-QUIC-DEADLINE),\n*1225|    /// default [`BROKER_QUIC_OP_TIMEOUT_MS`]; the hermetic black-hole tests lower\n 1226|    /// it via [`Self::set_quic_op_timeout`]. Interior-mutable so the setter takes\n 1227|    /// `&self`, like the live host's other post-`start` knobs (`set_roster_exchange`).\n 1228|    quic_op_timeout_ms: AtomicU64,\n...\n 1242|    /// runtime task stamps epoch-ms here every ~25ms. If the 2 workers are\n*1243|    /// MONOPOLIZED (a non-yielding dead-peer connect that never awaits back to\n 1244|    /// the scheduler) the stamp goes STALE — the DIRECT worker-starvation\n 1245|    /// discriminator the round-3 rig asserts on. A fresh stamp under a dead-peer\n 1246|    /// burst PROVES the QUIC awaits yield cleanly (⇒ the 2-worker-starvation\n...\n 1280|    pub fn start(cfg: NetConfig) -> io::Result<NetHost> {\n*1281|        // Hermeticity guard (REQ-PAIR-8): a loopback-scoped endpoint is the\n 1282|        // \"no network reached\" promise — the ceremony clock's lazy NTP query\n 1283|        // must honor it (tests must never probe a real NTP server). Real\n 1284|        // scopes keep the decision-18 offset correction.\n*1285|        spt_net::net::pairing::ntp::set_enabled(!matches!(cfg.scope, BindScope::Loopback));\n 1286|        let runtime = tokio::runtime::Builder::new_multi_thread()\n 1287|            .worker_threads(2)\n 1288|            .thread_name(\"spt-broker-net\")\n...\n 1353|                                    &self_addr,\n*1354|                                    &DialStage::new(), // accept side: stage untracked\n 1355|                                )\n 1356|                                .await\n 1357|                                {\n...\n 1406|            roster,\n*1407|            loopback_conn: Mutex::new(None),\n*1408|            loopback_pair_lock: Mutex::new(()),\n 1409|            quic_op_timeout_ms: AtomicU64::new(BROKER_QUIC_OP_TIMEOUT_MS),\n 1410|            dial_semaphore: Arc::new(tokio::sync::Semaphore::new(PUMP_DIAL_CONCURRENCY)),\n 1411|            active_dial_tasks: Arc::new(AtomicUsize::new(0)),\n...\n 1441|    /// Production keeps the [`BROKER_QUIC_OP_TIMEOUT_MS`] default; the hermetic\n*1442|    /// black-hole tests set a short bound so a dead-peer await fails fast instead\n 1443|    /// of pinning the test for ~10s. Off [`NetConfig`] by design (mirrors\n 1444|    /// [`Self::set_roster_exchange`]) — the many hermetic test configs need no change.\n 1445|    // [impl->REQ-HAZARD-BROKER-QUIC-DEADLINE]\n...\n 1452|    /// (REQ-HAZARD-BROKER-QUIC-DEADLINE) — the B-half of REQ-HAZARD-PUMP-IPC-\n*1453|    /// DEADLINE. A brain-waiting QUIC await against a black-holed/dead peer must\n 1454|    /// fail PROMPTLY with an ORDINARY error the broker replies, never the\n 1455|    /// unbounded await that wedged the single-threaded pump for 2.2h. On elapse\n 1456|    /// the future is DROPPED — cancelling the in-flight connect / stream op, so\n...\n 1539|    /// tasks currently on the 2-worker net runtime. The starvation rig reads this\n*1540|    /// to correlate a dead-peer burst with loopback-drain health.\n 1541|    pub fn active_dial_tasks(&self) -> usize {\n 1542|        self.active_dial_tasks.load(Ordering::Relaxed)\n 1543|    }\n...\n 1567|        // Bounded (REQ-HAZARD-BROKER-QUIC-DEADLINE): the `connect` + `prove_membership`\n*1568|        // awaits are the exact unbounded sites a dead peer black-holes — wrap\n 1569|        // them so a non-responsive peer fails as an ordinary error within the bound.\n 1570|        self.bounded_block_on(async move {\n 1571|            // Q2 caveat: the blocking dial acquires a permit from the SAME\n...\n 1638|                    // Ordinary per-peer failure rides the presence stream — the pump\n*1639|                    // backs THIS peer off and moves on; the round is never poisoned.\n 1640|                    // The reason is stage-attributed (REQ-PUMP-STAGE-TRUTH).\n 1641|                    shared.presence.lock().unwrap().append_dial_failed(\n 1642|                        &remote_id_hex,\n...\n 1713|    /// connection's current path addresses) as opaque JSON — the address-\n*1714|    /// seeding write-back the `net-dialed` reply carries (M8 decision 20,\n 1715|    /// REQ-CONV-1). `null` when the conn is gone, path-less, or fails to\n 1716|    /// serialize: the seed is an optimization, never an error path.\n 1717|    // [impl->REQ-CONV-1]\n...\n 1721|            match conns.get(&conn_id) {\n*1722|                // A loopback conn has no path to observe — nothing to seed.\n 1723|                Some(e) => match &e.kind {\n 1724|                    ConnKind::Quic(c) => c.clone(),\n*1725|                    ConnKind::Loopback => return serde_json::Value::Null,\n 1726|                },\n 1727|                None => return serde_json::Value::Null,\n 1728|            }\n...\n 1732|\n*1733|    /// Mint (or reuse) the broker's single in-process loopback connection\n 1734|    /// (M12 W1.5, B1) and return `(conn_id, remote_id_hex)` — `remote_id_hex` is\n 1735|    /// THIS node's own id, so a stream served over it gates same-node (access\n 1736|    /// short-circuit Allow). The conn carries NO iroh handshake: its identity is\n...\n 1742|    // [impl->REQ-RC-1]\n*1743|    pub fn dial_loopback(&self) -> (u64, String) {\n 1744|        let remote_id_hex = self.local_id.to_hex();\n*1745|        let mut slot = self.loopback_conn.lock().unwrap();\n 1746|        if let Some(id) = *slot {\n 1747|            return (id, remote_id_hex);\n 1748|        }\n...\n 1752|            ConnEntry {\n*1753|                kind: ConnKind::Loopback,\n 1754|                remote_id_hex: remote_id_hex.clone(),\n 1755|                proven_subnets: HashSet::new(),\n 1756|            },\n...\n 1765|    ///\n*1766|    /// On a [`ConnKind::Loopback`] conn there is no peer to accept the stream,\n 1767|    /// so the open mints a CROSS-WIRED PAIR in-process: an operator row\n 1768|    /// (`initiated_locally = true`, the returned id) and a peer row\n 1769|    /// (`initiated_locally = false`) the inbound dispatcher claims and serves via\n...\n 1781|    ) -> io::Result<u64> {\n*1782|        enum Backing {\n 1783|            Quic(Connection),\n*1784|            Loopback,\n 1785|        }\n*1786|        let (backing, remote_id_hex) = {\n 1787|            let conns = self.shared.conns.lock().unwrap();\n 1788|            let e = conns.get(&conn_id).ok_or_else(|| {\n 1789|                io::Error::new(io::ErrorKind::NotFound, format!(\"no such conn {conn_id}\"))\n 1790|            })?;\n*1791|            let backing = match &e.kind {\n*1792|                ConnKind::Quic(c) => Backing::Quic(c.clone()),\n*1793|                ConnKind::Loopback => Backing::Loopback,\n 1794|            };\n*1795|            (backing, e.remote_id_hex.clone())\n 1796|        };\n 1797|        let shared = Arc::clone(&self.shared);\n*1798|        match backing {\n*1799|            // Bounded (REQ-HAZARD-BROKER-QUIC-DEADLINE): `open_bi` on a black-holed\n 1800|            // peer can block on stream flow-control credit the dead peer never grants.\n*1801|            Backing::Quic(conn) => self.bounded_block_on(async move {\n 1802|                let (send, recv) = conn\n 1803|                    .open_bi()\n 1804|                    .await\n...\n 1816|            // Bounded like the QUIC arm (REQ-BROKER-ATTACH-JOURNAL-RESILIENT, bug\n*1817|            // #16): the loopback open is all in-memory and normally returns\n 1818|            // instantly, but a SICK NetHost runtime (saturated/wedged after a\n 1819|            // self-update brain-respawn) would otherwise hang this journaled op with\n 1820|            // no deadline — killing the per-conn reply thread so every attach\n...\n 1824|            // [impl->REQ-BROKER-ATTACH-JOURNAL-RESILIENT]\n*1825|            Backing::Loopback => self.bounded_block_on(async move {\n 1826|                // One duplex carries BOTH directions: operator writes reach the\n 1827|                // peer's read pump and vice-versa. Split each end into a read +\n 1828|                // write half so each row owns exactly its leg.\n...\n 1838|                    &remote_id_hex,\n*1839|                    SendHalf::Loopback(peer_w),\n*1840|                    RecvHalf::Loopback(peer_r),\n 1841|                    false,\n 1842|                    crate::msg::StreamLifetime::Durable,\n 1843|                );\n...\n 1847|                    &remote_id_hex,\n*1848|                    SendHalf::Loopback(op_w),\n*1849|                    RecvHalf::Loopback(op_r),\n 1850|                    true,\n 1851|                    lifetime,\n 1852|                ))\n...\n 1909|        // Bounded (REQ-HAZARD-BROKER-QUIC-DEADLINE): `write_all` / `finish` on a\n*1910|        // black-holed QUIC stream blocks on flow-control the dead peer never grants.\n*1911|        // Loopback (in-process) never hangs, so the bound is a no-op for it.\n 1912|        self.bounded_block_on(async move {\n 1913|            let mut send = entry.send.lock().await;\n 1914|            if !bytes.is_empty() {\n...\n 1948|    /// client-side guard (double-filter harmless, N-1 compatible both ways).\n*1949|    /// In-process introspection (tests, the loopback-pair diff) keeps the\n 1950|    /// unfiltered [`Self::stream_infos`].\n 1951|    // [impl->REQ-STREAM-INFOS-SERVER-FILTER]\n 1952|    pub fn dispatch_stream_infos(&self) -> Vec<NetStreamInfo> {\n...\n 2086|\n*2087|    /// Mark a held stream **retentive** with backpressure `cap` chunks (M11-W3,\n 2088|    /// REQ-SHELL-4): its read log stops evicting (no byte loss) and its read pump\n 2089|    /// lags the sender once the undrained ring reaches `cap` (no unbounded growth).\n 2090|    /// Used for shell-tunnel streams; ordinary streams stay bounded-overwrite.\n...\n 2100|    /// pops the whole ring in order, concatenated, and wakes the read pump (which\n*2101|    /// may have been lagging the sender at the backpressure cap). Returns\n 2102|    /// `(bytes, finished)`; `finished` ⇒ the read side ended (the poller closes).\n 2103|    /// `None` ⇒ no such stream. Empty bytes + `finished == false` ⇒ nothing pending.\n 2104|    // [impl->REQ-SHELL-4]\n...\n 2107|        let drained = entry.log.lock().unwrap().drain();\n*2108|        // The drain freed ring room — wake a pump that was backpressure-parked.\n 2109|        entry.room.notify_one();\n 2110|        Some(drained)\n 2111|    }\n...\n 2113|    /// A held stream's current undrained ring depth in chunks (test introspection\n*2114|    /// for the retentive backpressure bound). `None` ⇒ no such stream.\n 2115|    #[cfg(test)]\n 2116|    fn ring_len(&self, stream_id: u64) -> Option<usize> {\n 2117|        self.shared\n...\n 2125|    /// Inbound (peer-initiated) stream rows currently held on `conn_id` — the\n*2126|    /// candidate set [`Self::open_loopback_pair`] diffs to find the peer it just\n 2127|    /// minted.\n 2128|    fn inbound_rows_on(&self, conn_id: u64) -> HashSet<u64> {\n 2129|        self.shared\n...\n 2138|\n*2139|    /// Open a same-node loopback tunnel stream PAIR (M11-W3, REQ-SHELL-4) and\n 2140|    /// return `(owner_stream, shell_stream)` — the operator row the owner side\n*2141|    /// drives and the cross-wired peer row the shell side drives. The loopback\n 2142|    /// `open_stream` mints both rows in-process (one `duplex` wires them, so a\n 2143|    /// write to either send appears on the other's read log). Serialized so the\n 2144|    /// before/after diff sees exactly the one peer row this open minted.\n 2145|    // [impl->REQ-SHELL-4]\n*2146|    pub fn open_loopback_pair(&self) -> io::Result<(u64, u64)> {\n*2147|        let _guard = self.loopback_pair_lock.lock().unwrap();\n*2148|        let (conn_id, _) = self.dial_loopback();\n 2149|        let before = self.inbound_rows_on(conn_id);\n 2150|        let owner_stream = self.open_stream(conn_id, crate::msg::StreamLifetime::Durable)?;\n 2151|        let shell_stream = self\n...\n 2155|            .next()\n*2156|            .ok_or_else(|| io::Error::other(\"loopback peer row not registered\"))?;\n 2157|        Ok((owner_stream, shell_stream))\n 2158|    }\n 2159|\n...\n 2165|        stream_id: u64,\n*2166|        sub: SharedSend,\n 2167|        from_seq: u64,\n 2168|    ) -> io::Result<()> {\n 2169|        let log = {\n...\n 2199|    // [impl->REQ-SEAT-LIFETIME-BOUNDED]\n*2200|    pub(crate) fn unsubscribe_stream(&self, stream_id: u64, sub: &SharedSend) -> bool {\n 2201|        let log = {\n 2202|            let streams = self.shared.streams.lock().unwrap();\n 2203|            streams.get(&stream_id).map(|e| Arc::clone(&e.log))\n...\n 2220|    /// disconnect cleanup; the stream + its ring persist).\n*2221|    pub(crate) fn detach_stream_sub(&self, stream_id: u64, sub: &SharedSend) {\n 2222|        let log = {\n 2223|            let streams = self.shared.streams.lock().unwrap();\n 2224|            streams.get(&stream_id).map(|e| Arc::clone(&e.log))\n...\n 2233|    // [impl->REQ-EP-4]\n*2234|    pub(crate) fn subscribe_presence(&self, sub: SharedSend, from_seq: u64) {\n 2235|        // Same atomic install-with-replay discipline as subscribe_stream\n 2236|        // (the PresenceLog mirror): the seat's writer drains the replay\n 2237|        // before any live event, conn I/O outside the mutex. A replay write\n...\n 2250|    /// disconnect cleanup; the log + its ring persist).\n*2251|    pub(crate) fn detach_presence_sub(&self, sub: &SharedSend) {\n 2252|        self.shared.presence.lock().unwrap().detach_if(sub);\n 2253|    }\n 2254|\n...\n 2272|    /// closes), forcing the peer onto its QUIC idle timeout. That is the\n*2273|    /// sleeping-laptop / QUIC-stack-dead netsplit the transport-death rig\n 2274|    /// needs, and it stages it with no elevation, no firewall and no FFI.\n 2275|    ///\n 2276|    /// [`Self::start`] builds the runtime with `worker_threads(2)`, which is\n...\n 2344|            discovery: LocalDiscovery::Off,\n*2345|            scope: BindScope::Loopback,\n 2346|            membership: None,\n 2347|        }\n 2348|    }\n...\n 2418|    // [unit->REQ-DAEMON-2] a broker-owned host dials a second in-process host\n*2419|    // over hermetic loopback QUIC; both sides register the connection — the\n 2420|    // dialer in its table, the acceptor via its accept loop.\n 2421|    #[test]\n*2422|    fn two_hosts_dial_over_loopback() {\n 2423|        let a = NetHost::start(hermetic(&Identity::generate())).expect(\"host a\");\n 2424|        let b_id = Identity::generate();\n 2425|        let b = NetHost::start(hermetic(&b_id)).expect(\"host b\");\n...\n 2450|\n*2451|    // [unit->REQ-RC-1] the broker loopback conn (M12 W1.5, B1): `dial_loopback`\n 2452|    // mints a REUSED singleton whose origin is this node's OWN id, and\n 2453|    // `open_stream` on it cross-wires a PAIR in-process — an operator row and a\n 2454|    // peer row (the one the dispatcher serves) — with the operator's writes\n...\n 2456|    #[test]\n*2457|    fn loopback_conn_is_a_reused_singleton_that_cross_wires_a_stream_pair() {\n 2458|        let id = Identity::generate();\n 2459|        let host = NetHost::start(hermetic(&id)).expect(\"host start\");\n 2460|\n 2461|        // Singleton: reused, own-node origin (never a foreign/handshake id).\n*2462|        let (c1, hex1) = host.dial_loopback();\n*2463|        let (c2, hex2) = host.dial_loopback();\n*2464|        assert_eq!(c1, c2, \"the loopback conn is a reused singleton\");\n 2465|        assert_eq!(\n 2466|            hex1,\n 2467|            id.public_key().to_hex(),\n*2468|            \"loopback origin is THIS node's own id (broker-minted)\"\n 2469|        );\n 2470|        assert_eq!(hex2, hex1);\n*2471|        assert_eq!(host.conn_count(), 1, \"one loopback conn held\");\n 2472|\n 2473|        // open_stream mints the cross-wired pair: operator row + peer row.\n*2474|        let op_stream = host.open_stream(c1, crate::msg::StreamLifetime::Durable).expect(\"open loopback stream\");\n 2475|        let infos = host.stream_infos();\n 2476|        assert_eq!(infos.len(), 2, \"operator + peer rows\");\n 2477|        let op = infos\n...\n 2522|        let host = NetHost::start(hermetic(&Identity::generate())).expect(\"host\");\n*2523|        let (conn, _) = host.dial_loopback();\n 2524|\n 2525|        let cb_stream = host\n 2526|            .open_stream(conn, crate::msg::StreamLifetime::ConnectionBound)\n...\n 2590|\n*2591|    // [unit->REQ-SHELL-4] the loopback tunnel pair under BACKPRESSURE (M11-W3,\n 2592|    // doyle gate condition): a retentive shell-side stream that is not drained while\n 2593|    // the owner keeps writing must NOT lose bytes AND must NOT grow unbounded — the\n 2594|    // read pump lags the sender at the cap. After the reader catches up, every byte\n...\n 2596|    #[test]\n*2597|    fn loopback_tunnel_backpressures_without_loss() {\n 2598|        use std::sync::atomic::{AtomicUsize, Ordering as AtOrd};\n 2599|        let id = Identity::generate();\n 2600|        let host = Arc::new(NetHost::start(hermetic(&id)).expect(\"host start\"));\n*2601|        let (owner, shell) = host.open_loopback_pair().expect(\"open tunnel pair\");\n 2602|        const CAP: usize = 4;\n 2603|        host.mark_stream_retentive(owner, CAP);\n 2604|        host.mark_stream_retentive(shell, CAP);\n...\n 2624|\n*2625|        // Let the pump fill the shell ring to the cap and PARK (backpressure).\n 2626|        let mut peak = 0usize;\n 2627|        for _ in 0..200 {\n 2628|            if let Some(d) = host.ring_len(shell) {\n...\n 2632|            if sent.load(AtOrd::Relaxed) >= N {\n*2633|                break; // (shouldn't reach N while we never drain — backpressure)\n 2634|            }\n 2635|        }\n 2636|        assert!(peak > 0, \"the pump delivered some frames to the shell ring\");\n...\n 2638|            host.ring_len(shell).unwrap() <= CAP,\n*2639|            \"the retentive ring never exceeds the backpressure cap (bounded)\"\n 2640|        );\n 2641|        assert!(\n 2642|            sent.load(AtOrd::Relaxed) < N,\n*2643|            \"the un-drained consumer lags the sender (backpressure engaged, not all sent)\"\n 2644|        );\n 2645|\n 2646|        // Now drain to completion — the sender unblocks, every byte arrives in order.\n...\n 2662|        assert_eq!(got.len(), expected.len(), \"no byte lost or duplicated\");\n*2663|        assert_eq!(got, expected, \"byte-exact, in order, under backpressure\");\n 2664|    }\n 2665|\n 2666|    // [unit->REQ-HAZARD-BROKER-QUIC-DEADLINE] the broker-side QUIC-op bound (the\n...\n 2676|\n*2677|        // A never-completing await (the black-holed-peer shape) fails at the bound.\n 2678|        let started = std::time::Instant::now();\n 2679|        let err = host\n 2680|            .bounded_block_on(std::future::pending::<io::Result<()>>())\n...\n 2752|        let host = NetHost::start(hermetic(&Identity::generate())).expect(\"host start\");\n*2753|        let (owner, shell) = host.open_loopback_pair().expect(\"open pair\");\n 2754|        host.send_stream(owner, b\"{\\\"kind\\\":\\\"request\\\",\\\"session_id\\\":1}\\n\", false)\n 2755|            .expect(\"send opener\");\n 2756|        // The pump is async — wait for the shell row to hold the opener.\n...\n 2785|    // worker's reply send is fire-and-forget (brain returns before the broker\n*2786|    // processes the envelope), so the acked retire on its own conn can be\n 2787|    // processed FIRST. Pre-fix, retire removed the entry: the late send hit\n 2788|    // no-such-stream (or the drop reset the never-finished send half) and the\n 2789|    // requester got torn-EOF instead of its reply (wansend NoReply, CI red\n...\n 2793|        let host = NetHost::start(hermetic(&Identity::generate())).expect(\"host start\");\n*2794|        let (owner, shell) = host.open_loopback_pair().expect(\"open pair\");\n 2795|        // The requester's one-shot request: write + FIN (the rest/sync shape —\n 2796|        // the row is born finished on the serving side).\n 2797|        host.send_stream(\n...\n 2869|    /// A connected local-socket pair whose SERVER send half wraps as the\n*2870|    /// `SharedSend` a subscriber seat writes through — `controller_socket_pair`'s\n 2871|    /// idiom (broker.rs) with a caller-chosen write bound.\n 2872|    fn seat_socket_pair(\n 2873|        bound: Duration,\n 2874|    ) -> (\n*2875|        SharedSend,\n 2876|        interprocess::local_socket::Stream,\n 2877|        interprocess::local_socket::RecvHalf,\n 2878|    ) {\n...\n 3118|    #[test]\n*3119|    fn displaced_writer_failure_never_cancels_the_successor_lease() {\n 3120|        let (a, client_a, _ra) = seat_socket_pair(Duration::from_secs(5));\n 3121|        let (b, _cb, _rb) = seat_socket_pair(Duration::from_secs(5));\n 3122|        let mut log = StreamLog::new(11, 64);\n...\n 3209|        let host = NetHost::start(hermetic(&Identity::generate())).expect(\"host start\");\n*3210|        let (owner, _shell) = host.open_loopback_pair().expect(\"open pair\");\n 3211|        let lease = host.stream_lease(owner).expect(\"lease handle\");\n 3212|        assert!(!lease.is_canceled());\n 3213|        host.send_stream(owner, b\"ok\", false).expect(\"a live lease serves\");\n...\n 3231|        let host = NetHost::start(hermetic(&Identity::generate())).expect(\"host start\");\n*3232|        let (owner, shell) = host.open_loopback_pair().expect(\"open pair\");\n 3233|        host.send_stream(owner, b\"{\\\"subnet\\\":\\\"s\\\",\\\"instance\\\":{}}\\n\", false)\n 3234|            .expect(\"bytes onto the peer row\");\n 3235|\n...\n 3267|    #[test]\n*3268|    fn unsubscribe_releases_only_the_callers_seat_joins_writer_and_is_idempotent() {\n 3269|        let host = NetHost::start(hermetic(&Identity::generate())).expect(\"host start\");\n*3270|        let (owner, shell) = host.open_loopback_pair().expect(\"open pair\");\n 3271|        let _ = shell;\n 3272|\n 3273|        let (a, _ca, _ra) = seat_socket_pair(Duration::from_secs(5));\n...\n 3315|        // the retired_row_still_flushes_a_late_reply contract above).\n*3316|        let (owner, shell) = host.open_loopback_pair().expect(\"pair 1\");\n 3317|        let (sub, _c, _r) = seat_socket_pair(Duration::from_secs(5));\n 3318|        host.subscribe_stream(shell, Arc::clone(&sub), 0).expect(\"subscribe\");\n 3319|        let (rows_before, seats_before) = host.stream_counts();"}],"details":{"scopePath":"crates/spt-daemon/src/nethost.rs","searchPath":"C:\\Users\\decid\\Documents\\projects\\spt-core\\crates\\spt-daemon\\src\\nethost.rs","cwd":"C:\\Users\\decid\\Documents\\projects\\spt-core","matchCount":141,"fileCount":1,"files":["crates/spt-daemon/src/nethost.rs"],"fileMatches":[{"path":"crates/spt-daemon/src/nethost.rs","count":141}],"truncated":false,"displayContent":"   43│\n  *44│use crate::broker::SharedSend;\n   45│\n   46│use crate::effect::Minter;\n   47│use crate::seedproofx::{prove_membership, MembershipSource, RosterExchange};\n     │...\n   60│/// brain-waiting QUIC await — `dial` / `open_stream` / `send_stream` — is bounded\n  *61│/// by this so a black-holed or dead peer fails PROMPTLY with an ORDINARY error\n   62│/// the broker replies, never an unbounded await that wedges the single-threaded\n   63│/// pump (the B-half of REQ-HAZARD-PUMP-IPC-DEADLINE; the 2.2h hfenduleam wedge).\n   64│/// Chosen at **10s**: comfortably above any legitimate LAN/relay connect +\n     │...\n   71│/// frame's IPC transit + the brain's read scheduling. Test-overridable via\n  *72│/// [`NetHost::set_quic_op_timeout`] so the hermetic black-hole tests fail fast.\n   73│const BROKER_QUIC_OP_TIMEOUT_MS: u64 = 10_000;\n   74│\n   75│/// Bounded concurrent-dial cap (PUMP-TRUTH W2, REQ-PUMP-PEER-ISOLATION, G4): a\n     │...\n   95│    /// Socket bind scope: [`BindScope::All`] in production;\n  *96│    /// [`BindScope::Loopback`] in hermetic tests (no real interface touched,\n   97│    /// no per-test-binary Windows firewall prompt).\n   98│    pub scope: BindScope,\n   99│    /// Live source of this node's subnet credentials for the connect-time\n     │...\n  143│/// fixes 2+3, REQ-STREAMLOG-SUBSCRIBER-DISCIPLINE): frames ENQUEUE under the\n *144│/// log mutex; the conn I/O — the deadline-bounded `SharedSend::write` that can\n  145│/// take 15s against a wedged conn — runs on this seat's OWN writer thread,\n  146│/// OUTSIDE every lock. The field stall was exactly the old shape's serial\n  147│/// replay-under-mutex against a poisoned conn (33 windows, all ~15s).\n     │...\n  154│struct SubscriberSeat {\n *155│    sub: SharedSend,\n  156│    tx: std::sync::mpsc::SyncSender<Envelope>,\n  157│    /// Writer thread exited (failure or clean drain) — the seat is fully gone\n  158│    /// and a replacement may install (the no-reinstall-until-gone rule,\n     │...\n  169│impl SubscriberSeat {\n *170│    /// Install a seat: spawn its writer thread and hand back the handle.\n  171│    /// `replay` is the attach snapshot, moved in UNDER the log mutex (a\n  172│    /// memory move, never conn I/O): the writer drains it COMPLETELY before\n  173│    /// touching the live queue, so replay-then-live seq order holds by\n     │...\n  189│    fn install(\n *190│        sub: SharedSend,\n  191│        replay: Vec<Envelope>,\n  192│        lease: Option<Arc<StreamLease>>,\n  193│        label: &str,\n     │...\n  212│                    }\n *213│                    writer_sub.lifecycle_event(\"stream-sub-writer-poison\", halt_reason);\n  214│                };\n  215│                // Replay first — halt at the first failed write (Amendment\n  216│                // fix 2), never iterate past a poisoned conn.\n     │...\n  270│    /// Whether this seat belongs to `sub` (detach correlation).\n *271│    fn is(&self, sub: &SharedSend) -> bool {\n  272│        Arc::ptr_eq(&self.sub, sub)\n  273│    }\n  274│\n     │...\n  319│\n *320│/// Transport backing a held connection. A real `Quic` connection (the only\n *321│/// kind an inbound handshake can produce), or a broker-minted `Loopback` —\n  322│/// an in-process connection with NO iroh handshake whose identity IS this node,\n *323│/// established BY CONSTRUCTION. A loopback conn can ONLY be minted broker-side\n *324│/// via [`NetHost::dial_loopback`]; it can never arise from an inbound handshake\n  325│/// (a remote always presents its own key and cannot claim to be us). This is\n  326│/// what lets a LOCAL `spt rc` ride the byte-identical `serve_attach` pump (M12\n  327│/// W1.5, B1) — QUIC self-dial is refused by iroh (\"Connecting to ourself is not\n *328│/// supported\"), so local attach rides a broker loopback conn instead.\n  329│enum ConnKind {\n  330│    /// A real QUIC connection — kept alive by ownership here; dropping it closes\n  331│    /// the connection. Cloned (cheap handle) for stream opens.\n  332│    Quic(Connection),\n *333│    /// A broker-minted in-process loopback (no iroh handle, no handshake). Its\n  334│    /// streams are minted eagerly in cross-wired pairs by [`NetHost::open_stream`];\n  335│    /// there is no accept loop and no closed-watcher (it lives for the host).\n *336│    Loopback,\n  337│}\n  338│\n *339│/// One live connection the broker holds (QUIC or loopback).\n  340│struct ConnEntry {\n *341│    /// The transport backing this connection.\n  342│    kind: ConnKind,\n  343│    /// The remote node's stable name (hex pubkey). For QUIC, proven by the\n *344│    /// handshake; for loopback, THIS node's own id (the broker-minted identity).\n  345│    remote_id_hex: String,\n  346│    /// The subnets the peer proved current-epoch seed-knowledge of at connect\n  347│    /// (Mesh-D2, REQ-MESH-1). Empty when seed-proof is disabled (`membership:\n     │...\n  383│    finished: bool,\n *384│    /// **Retentive backpressure cap** (M11-W3, REQ-SHELL-4), in chunks; `0` ⇒ not\n  385│    /// retentive (ordinary bounded-overwrite stream). A shell-tunnel stream must\n  386│    /// NOT drop bytes — CONTEXT:262 \"tunneled protocols cannot drop frames, so\n  387│    /// congestion surfaces as lag, never loss.\" So a retentive log never cap-evicts\n     │...\n  431│\n *432│    /// Mark this log retentive with backpressure `cap` (chunks). Idempotent set.\n  433│    fn set_retentive(&mut self, cap: usize) {\n  434│        self.retentive_cap = cap;\n  435│    }\n     │...\n  466│        self.next_seq += 1;\n *467│        self.ring.push_back((seq, bytes.to_vec()));\n  468│        if self.retentive_cap == 0 {\n  469│            while self.ring.len() > self.cap_chunks {\n  470│                self.ring.pop_front();\n     │...\n  502│    /// Drain every undrained chunk in order (the shell-tunnel recv cursor, W3).\n *503│    /// Pops the whole ring front-to-back (FIFO by `seq`) and concatenates the\n  504│    /// bytes — chunk boundaries are arbitrary read-pump splits, irrelevant to an\n  505│    /// opaque byte stream. Lossless + ordered: a byte leaves the ring only here,\n  506│    /// once. Returns `(bytes, finished)`; `finished` reflects a clean/torn end so\n     │...\n  544│    // [impl->REQ-STREAMLOG-SUBSCRIBER-DISCIPLINE]\n *545│    fn begin_attach(&mut self, sub: SharedSend, from_seq: u64) -> io::Result<usize> {\n  546│        if let Some(prior) = &self.subscriber {\n  547│            // No-reinstall-until-gone applies to the POISON class only\n  548│            // (REQ-DISPATCH-FALLBACK-CIRCUIT): a wedged writer may sit in a\n     │...\n  610│    // [impl->REQ-SEAT-LIFETIME-BOUNDED]\n *611│    fn take_seat_if(&mut self, sub: &SharedSend) -> Option<SubscriberSeat> {\n  612│        if self.subscriber.as_ref().is_some_and(|cur| cur.is(sub)) {\n  613│            self.subscriber.take()\n  614│        } else {\n     │...\n  620│    /// not unsubscribe a newer brain).\n *621│    fn detach_if(&mut self, sub: &SharedSend) {\n  622│        if let Some(cur) = &self.subscriber {\n  623│            if cur.is(sub) {\n  624│                self.subscriber = None;\n     │...\n  672│    /// A `connected` event carrying the connection's OBSERVED remote address\n *673│    /// (REQ-CONV-1): the async-dial write-back seed the blocking `net-dialed`\n  674│    /// reply used to carry — now it rides the presence stream so the pump seeds\n  675│    /// `peer-addrs.json` from a non-blocking dial the same way.\n  676│    // [impl->REQ-CONV-1]\n     │...\n  720│        self.next_seq += 1;\n *721│        self.ring.push_back(ev.clone());\n  722│        while self.ring.len() > self.cap_events {\n  723│            self.ring.pop_front();\n  724│        }\n     │...\n  744│    // [impl->REQ-STREAMLOG-SUBSCRIBER-DISCIPLINE]\n *745│    fn begin_attach(&mut self, sub: SharedSend, from_seq: u64) -> io::Result<usize> {\n  746│        if let Some(prior) = &self.subscriber {\n  747│            // Poison-class only (see StreamLog::begin_attach): a healthy\n  748│            // prior presence subscriber is displaced by the newer brain.\n     │...\n  775│    /// not unsubscribe a newer brain).\n *776│    fn detach_if(&mut self, sub: &SharedSend) {\n  777│        if let Some(cur) = &self.subscriber {\n  778│            if cur.is(sub) {\n  779│                self.subscriber = None;\n     │...\n  788│/// write end of an in-process duplex). That is the M12 W1.5 one-pump invariant\n *789│/// made structural: a loopback attach and a cross-node attach diverge at this\n  790│/// leaf and nowhere above it.\n  791│enum SendHalf {\n  792│    Quic(SendStream),\n *793│    Loopback(WriteHalf<DuplexStream>),\n  794│}\n  795│\n  796│impl SendHalf {\n     │...\n  803│                .map_err(|e| io::Error::other(e.to_string())),\n *804│            SendHalf::Loopback(w) => w.write_all(bytes).await,\n  805│        }\n  806│    }\n  807│\n     │...\n  812│            SendHalf::Quic(s) => s.finish().map_err(|e| io::Error::other(e.to_string())),\n *813│            SendHalf::Loopback(w) => w.shutdown().await,\n  814│        }\n  815│    }\n  816│}\n     │...\n  818│/// A stream's read half — fed into the single read pump in [`register_stream`].\n *819│/// QUIC reads via quinn's chunk API; loopback via the read end of the duplex.\n  820│/// Both append to the same [`StreamLog`], so the subscriber (and everything\n  821│/// above it) cannot tell the transports apart.\n  822│enum RecvHalf {\n  823│    Quic(RecvStream),\n *824│    Loopback(ReadHalf<DuplexStream>),\n  825│}\n  826│\n *827│/// One live stream the broker holds (QUIC or loopback): the send half (brain\n  828│/// writes ride it over IPC) + the sequenced read log.\n  829│struct StreamEntry {\n  830│    conn_id: u64,\n     │...\n  842│    /// Wakes the read pump when a retentive (tunnel) stream's ring drains below\n *843│    /// its backpressure cap (M11-W3, REQ-SHELL-4) — see [`StreamLog::retentive_cap`].\n  844│    /// Unused by ordinary (non-retentive) streams.\n  845│    room: Arc<tokio::sync::Notify>,\n  846│    /// Retired = invisible to dispatch eligibility (ADR-0038 A), NOT torn down:\n     │...\n  904│/// the path address string, or `null` when id-only / path-less (nothing worth\n *905│/// seeding over discovery). Shared by [`register_conn`]'s presence write-back\n  906│/// and [`NetHost::conn_remote_addr_json`] so both seed identically.\n  907│// [impl->REQ-CONV-1]\n  908│fn observed_addr_json(conn: &Connection) -> serde_json::Value {\n     │...\n 1031│    // `connected` event can never race behind its own `disconnected` (D4c). The\n*1032│    // event carries the OBSERVED remote addr (REQ-CONV-1 seed write-back) so a\n 1033│    // non-blocking submit-dial seeds `peer-addrs.json` like the blocking reply\n 1034│    // did (id-only / path-less → `null`, nothing worth seeding).\n 1035│    let observed = observed_addr_json(&conn);\n     │...\n 1128│    // replayable either way. The transport differs only at this read leaf\n*1129│    // (QUIC chunk API vs the loopback duplex's `AsyncRead`); both append to the\n 1130│    // same `StreamLog`, so nothing above the pump can tell them apart.\n 1131│    //\n*1132│    // Backpressure (M11-W3, REQ-SHELL-4): a retentive (tunnel) log never evicts,\n 1133│    // so before pulling the next chunk the pump WAITS while the undrained ring is\n 1134│    // at its `retentive_cap`. The unread bytes stay in the transport's buffer, so\n*1135│    // QUIC stream flow-control (or the loopback duplex's bounded buffer) lags the\n 1136│    // SENDER — \"lag, never loss\" (CONTEXT:262). A `drain` wakes `room`. Ordinary\n 1137│    // streams have `retentive_cap == 0`, so this gate is a no-op for them.\n 1138│    tokio::spawn(async move {\n     │...\n 1149│            },\n*1150│            RecvHalf::Loopback(mut recv) => {\n 1151│                let mut buf = vec![0u8; STREAM_READ_CHUNK];\n 1152│                loop {\n 1153│                    await_ring_room(&log, &room).await;\n     │...\n 1168│\n*1169│/// Backpressure gate for a retentive (tunnel) read pump: park until the undrained\n 1170│/// ring has room below `retentive_cap`, woken by a `drain` on `room`. A no-op for\n 1171│/// an ordinary stream (`retentive_cap == 0`). Holds no lock across the await.\n 1172│// [impl->REQ-SHELL-4]\n     │...\n 1213│    roster: Arc<Mutex<Option<RosterExchange>>>,\n*1214│    /// The broker's single in-process loopback connection (M12 W1.5, B1),\n*1215│    /// minted lazily on the first [`Self::dial_loopback`] and reused thereafter\n 1216│    /// — a local `spt rc` needs exactly one. `Some(conn_id)` once minted; the\n 1217│    /// row lives for the host (no handshake to lose, no closed-watcher).\n*1218│    loopback_conn: Mutex<Option<u64>>,\n*1219│    /// Serializes [`Self::open_loopback_pair`] (M11-W3) so the peer-row discovery\n*1220│    /// (a before/after diff on the loopback conn's inbound rows) sees exactly the\n*1221│    /// one row this open minted — concurrent pair-opens on the shared loopback conn\n 1222│    /// would otherwise race the diff.\n*1223│    loopback_pair_lock: Mutex<()>,\n 1224│    /// The broker-side QUIC-op deadline in ms (REQ-HAZARD-BROKER-QUIC-DEADLINE),\n*1225│    /// default [`BROKER_QUIC_OP_TIMEOUT_MS`]; the hermetic black-hole tests lower\n 1226│    /// it via [`Self::set_quic_op_timeout`]. Interior-mutable so the setter takes\n 1227│    /// `&self`, like the live host's other post-`start` knobs (`set_roster_exchange`).\n 1228│    quic_op_timeout_ms: AtomicU64,\n     │...\n 1242│    /// runtime task stamps epoch-ms here every ~25ms. If the 2 workers are\n*1243│    /// MONOPOLIZED (a non-yielding dead-peer connect that never awaits back to\n 1244│    /// the scheduler) the stamp goes STALE — the DIRECT worker-starvation\n 1245│    /// discriminator the round-3 rig asserts on. A fresh stamp under a dead-peer\n 1246│    /// burst PROVES the QUIC awaits yield cleanly (⇒ the 2-worker-starvation\n     │...\n 1280│    pub fn start(cfg: NetConfig) -> io::Result<NetHost> {\n*1281│        // Hermeticity guard (REQ-PAIR-8): a loopback-scoped endpoint is the\n 1282│        // \"no network reached\" promise — the ceremony clock's lazy NTP query\n 1283│        // must honor it (tests must never probe a real NTP server). Real\n 1284│        // scopes keep the decision-18 offset correction.\n*1285│        spt_net::net::pairing::ntp::set_enabled(!matches!(cfg.scope, BindScope::Loopback));\n 1286│        let runtime = tokio::runtime::Builder::new_multi_thread()\n 1287│            .worker_threads(2)\n 1288│            .thread_name(\"spt-broker-net\")\n     │...\n 1353│                                    &self_addr,\n*1354│                                    &DialStage::new(), // accept side: stage untracked\n 1355│                                )\n 1356│                                .await\n 1357│                                {\n     │...\n 1406│            roster,\n*1407│            loopback_conn: Mutex::new(None),\n*1408│            loopback_pair_lock: Mutex::new(()),\n 1409│            quic_op_timeout_ms: AtomicU64::new(BROKER_QUIC_OP_TIMEOUT_MS),\n 1410│            dial_semaphore: Arc::new(tokio::sync::Semaphore::new(PUMP_DIAL_CONCURRENCY)),\n 1411│            active_dial_tasks: Arc::new(AtomicUsize::new(0)),\n     │...\n 1441│    /// Production keeps the [`BROKER_QUIC_OP_TIMEOUT_MS`] default; the hermetic\n*1442│    /// black-hole tests set a short bound so a dead-peer await fails fast instead\n 1443│    /// of pinning the test for ~10s. Off [`NetConfig`] by design (mirrors\n 1444│    /// [`Self::set_roster_exchange`]) — the many hermetic test configs need no change.\n 1445│    // [impl->REQ-HAZARD-BROKER-QUIC-DEADLINE]\n     │...\n 1452│    /// (REQ-HAZARD-BROKER-QUIC-DEADLINE) — the B-half of REQ-HAZARD-PUMP-IPC-\n*1453│    /// DEADLINE. A brain-waiting QUIC await against a black-holed/dead peer must\n 1454│    /// fail PROMPTLY with an ORDINARY error the broker replies, never the\n 1455│    /// unbounded await that wedged the single-threaded pump for 2.2h. On elapse\n 1456│    /// the future is DROPPED — cancelling the in-flight connect / stream op, so\n     │...\n 1539│    /// tasks currently on the 2-worker net runtime. The starvation rig reads this\n*1540│    /// to correlate a dead-peer burst with loopback-drain health.\n 1541│    pub fn active_dial_tasks(&self) -> usize {\n 1542│        self.active_dial_tasks.load(Ordering::Relaxed)\n 1543│    }\n     │...\n 1567│        // Bounded (REQ-HAZARD-BROKER-QUIC-DEADLINE): the `connect` + `prove_membership`\n*1568│        // awaits are the exact unbounded sites a dead peer black-holes — wrap\n 1569│        // them so a non-responsive peer fails as an ordinary error within the bound.\n 1570│        self.bounded_block_on(async move {\n 1571│            // Q2 caveat: the blocking dial acquires a permit from the SAME\n     │...\n 1638│                    // Ordinary per-peer failure rides the presence stream — the pump\n*1639│                    // backs THIS peer off and moves on; the round is never poisoned.\n 1640│                    // The reason is stage-attributed (REQ-PUMP-STAGE-TRUTH).\n 1641│                    shared.presence.lock().unwrap().append_dial_failed(\n 1642│                        &remote_id_hex,\n     │...\n 1713│    /// connection's current path addresses) as opaque JSON — the address-\n*1714│    /// seeding write-back the `net-dialed` reply carries (M8 decision 20,\n 1715│    /// REQ-CONV-1). `null` when the conn is gone, path-less, or fails to\n 1716│    /// serialize: the seed is an optimization, never an error path.\n 1717│    // [impl->REQ-CONV-1]\n     │...\n 1721│            match conns.get(&conn_id) {\n*1722│                // A loopback conn has no path to observe — nothing to seed.\n 1723│                Some(e) => match &e.kind {\n 1724│                    ConnKind::Quic(c) => c.clone(),\n*1725│                    ConnKind::Loopback => return serde_json::Value::Null,\n 1726│                },\n 1727│                None => return serde_json::Value::Null,\n 1728│            }\n     │...\n 1732│\n*1733│    /// Mint (or reuse) the broker's single in-process loopback connection\n 1734│    /// (M12 W1.5, B1) and return `(conn_id, remote_id_hex)` — `remote_id_hex` is\n 1735│    /// THIS node's own id, so a stream served over it gates same-node (access\n 1736│    /// short-circuit Allow). The conn carries NO iroh handshake: its identity is\n     │...\n 1742│    // [impl->REQ-RC-1]\n*1743│    pub fn dial_loopback(&self) -> (u64, String) {\n 1744│        let remote_id_hex = self.local_id.to_hex();\n*1745│        let mut slot = self.loopback_conn.lock().unwrap();\n 1746│        if let Some(id) = *slot {\n 1747│            return (id, remote_id_hex);\n 1748│        }\n     │...\n 1752│            ConnEntry {\n*1753│                kind: ConnKind::Loopback,\n 1754│                remote_id_hex: remote_id_hex.clone(),\n 1755│                proven_subnets: HashSet::new(),\n 1756│            },\n     │...\n 1765│    ///\n*1766│    /// On a [`ConnKind::Loopback`] conn there is no peer to accept the stream,\n 1767│    /// so the open mints a CROSS-WIRED PAIR in-process: an operator row\n 1768│    /// (`initiated_locally = true`, the returned id) and a peer row\n 1769│    /// (`initiated_locally = false`) the inbound dispatcher claims and serves via\n     │...\n 1781│    ) -> io::Result<u64> {\n*1782│        enum Backing {\n 1783│            Quic(Connection),\n*1784│            Loopback,\n 1785│        }\n*1786│        let (backing, remote_id_hex) = {\n 1787│            let conns = self.shared.conns.lock().unwrap();\n 1788│            let e = conns.get(&conn_id).ok_or_else(|| {\n 1789│                io::Error::new(io::ErrorKind::NotFound, format!(\"no such conn {conn_id}\"))\n 1790│            })?;\n*1791│            let backing = match &e.kind {\n*1792│                ConnKind::Quic(c) => Backing::Quic(c.clone()),\n*1793│                ConnKind::Loopback => Backing::Loopback,\n 1794│            };\n*1795│            (backing, e.remote_id_hex.clone())\n 1796│        };\n 1797│        let shared = Arc::clone(&self.shared);\n*1798│        match backing {\n*1799│            // Bounded (REQ-HAZARD-BROKER-QUIC-DEADLINE): `open_bi` on a black-holed\n 1800│            // peer can block on stream flow-control credit the dead peer never grants.\n*1801│            Backing::Quic(conn) => self.bounded_block_on(async move {\n 1802│                let (send, recv) = conn\n 1803│                    .open_bi()\n 1804│                    .await\n     │...\n 1816│            // Bounded like the QUIC arm (REQ-BROKER-ATTACH-JOURNAL-RESILIENT, bug\n*1817│            // #16): the loopback open is all in-memory and normally returns\n 1818│            // instantly, but a SICK NetHost runtime (saturated/wedged after a\n 1819│            // self-update brain-respawn) would otherwise hang this journaled op with\n 1820│            // no deadline — killing the per-conn reply thread so every attach\n     │...\n 1824│            // [impl->REQ-BROKER-ATTACH-JOURNAL-RESILIENT]\n*1825│            Backing::Loopback => self.bounded_block_on(async move {\n 1826│                // One duplex carries BOTH directions: operator writes reach the\n 1827│                // peer's read pump and vice-versa. Split each end into a read +\n 1828│                // write half so each row owns exactly its leg.\n     │...\n 1838│                    &remote_id_hex,\n*1839│                    SendHalf::Loopback(peer_w),\n*1840│                    RecvHalf::Loopback(peer_r),\n 1841│                    false,\n 1842│                    crate::msg::StreamLifetime::Durable,\n 1843│                );\n     │...\n 1847│                    &remote_id_hex,\n*1848│                    SendHalf::Loopback(op_w),\n*1849│                    RecvHalf::Loopback(op_r),\n 1850│                    true,\n 1851│                    lifetime,\n 1852│                ))\n     │...\n 1909│        // Bounded (REQ-HAZARD-BROKER-QUIC-DEADLINE): `write_all` / `finish` on a\n*1910│        // black-holed QUIC stream blocks on flow-control the dead peer never grants.\n*1911│        // Loopback (in-process) never hangs, so the bound is a no-op for it.\n 1912│        self.bounded_block_on(async move {\n 1913│            let mut send = entry.send.lock().await;\n 1914│            if !bytes.is_empty() {\n     │...\n 1948│    /// client-side guard (double-filter harmless, N-1 compatible both ways).\n*1949│    /// In-process introspection (tests, the loopback-pair diff) keeps the\n 1950│    /// unfiltered [`Self::stream_infos`].\n 1951│    // [impl->REQ-STREAM-INFOS-SERVER-FILTER]\n 1952│    pub fn dispatch_stream_infos(&self) -> Vec<NetStreamInfo> {\n     │...\n 2086│\n*2087│    /// Mark a held stream **retentive** with backpressure `cap` chunks (M11-W3,\n 2088│    /// REQ-SHELL-4): its read log stops evicting (no byte loss) and its read pump\n 2089│    /// lags the sender once the undrained ring reaches `cap` (no unbounded growth).\n 2090│    /// Used for shell-tunnel streams; ordinary streams stay bounded-overwrite.\n     │...\n 2100│    /// pops the whole ring in order, concatenated, and wakes the read pump (which\n*2101│    /// may have been lagging the sender at the backpressure cap). Returns\n 2102│    /// `(bytes, finished)`; `finished` ⇒ the read side ended (the poller closes).\n 2103│    /// `None` ⇒ no such stream. Empty bytes + `finished == false` ⇒ nothing pending.\n 2104│    // [impl->REQ-SHELL-4]\n     │...\n 2107│        let drained = entry.log.lock().unwrap().drain();\n*2108│        // The drain freed ring room — wake a pump that was backpressure-parked.\n 2109│        entry.room.notify_one();\n 2110│        Some(drained)\n 2111│    }\n     │...\n 2113│    /// A held stream's current undrained ring depth in chunks (test introspection\n*2114│    /// for the retentive backpressure bound). `None` ⇒ no such stream.\n 2115│    #[cfg(test)]\n 2116│    fn ring_len(&self, stream_id: u64) -> Option<usize> {\n 2117│        self.shared\n     │...\n 2125│    /// Inbound (peer-initiated) stream rows currently held on `conn_id` — the\n*2126│    /// candidate set [`Self::open_loopback_pair`] diffs to find the peer it just\n 2127│    /// minted.\n 2128│    fn inbound_rows_on(&self, conn_id: u64) -> HashSet<u64> {\n 2129│        self.shared\n     │...\n 2138│\n*2139│    /// Open a same-node loopback tunnel stream PAIR (M11-W3, REQ-SHELL-4) and\n 2140│    /// return `(owner_stream, shell_stream)` — the operator row the owner side\n*2141│    /// drives and the cross-wired peer row the shell side drives. The loopback\n 2142│    /// `open_stream` mints both rows in-process (one `duplex` wires them, so a\n 2143│    /// write to either send appears on the other's read log). Serialized so the\n 2144│    /// before/after diff sees exactly the one peer row this open minted.\n 2145│    // [impl->REQ-SHELL-4]\n*2146│    pub fn open_loopback_pair(&self) -> io::Result<(u64, u64)> {\n*2147│        let _guard = self.loopback_pair_lock.lock().unwrap();\n*2148│        let (conn_id, _) = self.dial_loopback();\n 2149│        let before = self.inbound_rows_on(conn_id);\n 2150│        let owner_stream = self.open_stream(conn_id, crate::msg::StreamLifetime::Durable)?;\n 2151│        let shell_stream = self\n     │...\n 2155│            .next()\n*2156│            .ok_or_else(|| io::Error::other(\"loopback peer row not registered\"))?;\n 2157│        Ok((owner_stream, shell_stream))\n 2158│    }\n 2159│\n     │...\n 2165│        stream_id: u64,\n*2166│        sub: SharedSend,\n 2167│        from_seq: u64,\n 2168│    ) -> io::Result<()> {\n 2169│        let log = {\n     │...\n 2199│    // [impl->REQ-SEAT-LIFETIME-BOUNDED]\n*2200│    pub(crate) fn unsubscribe_stream(&self, stream_id: u64, sub: &SharedSend) -> bool {\n 2201│        let log = {\n 2202│            let streams = self.shared.streams.lock().unwrap();\n 2203│            streams.get(&stream_id).map(|e| Arc::clone(&e.log))\n     │...\n 2220│    /// disconnect cleanup; the stream + its ring persist).\n*2221│    pub(crate) fn detach_stream_sub(&self, stream_id: u64, sub: &SharedSend) {\n 2222│        let log = {\n 2223│            let streams = self.shared.streams.lock().unwrap();\n 2224│            streams.get(&stream_id).map(|e| Arc::clone(&e.log))\n     │...\n 2233│    // [impl->REQ-EP-4]\n*2234│    pub(crate) fn subscribe_presence(&self, sub: SharedSend, from_seq: u64) {\n 2235│        // Same atomic install-with-replay discipline as subscribe_stream\n 2236│        // (the PresenceLog mirror): the seat's writer drains the replay\n 2237│        // before any live event, conn I/O outside the mutex. A replay write\n     │...\n 2250│    /// disconnect cleanup; the log + its ring persist).\n*2251│    pub(crate) fn detach_presence_sub(&self, sub: &SharedSend) {\n 2252│        self.shared.presence.lock().unwrap().detach_if(sub);\n 2253│    }\n 2254│\n     │...\n 2272│    /// closes), forcing the peer onto its QUIC idle timeout. That is the\n*2273│    /// sleeping-laptop / QUIC-stack-dead netsplit the transport-death rig\n 2274│    /// needs, and it stages it with no elevation, no firewall and no FFI.\n 2275│    ///\n 2276│    /// [`Self::start`] builds the runtime with `worker_threads(2)`, which is\n     │...\n 2344│            discovery: LocalDiscovery::Off,\n*2345│            scope: BindScope::Loopback,\n 2346│            membership: None,\n 2347│        }\n 2348│    }\n     │...\n 2418│    // [unit->REQ-DAEMON-2] a broker-owned host dials a second in-process host\n*2419│    // over hermetic loopback QUIC; both sides register the connection — the\n 2420│    // dialer in its table, the acceptor via its accept loop.\n 2421│    #[test]\n*2422│    fn two_hosts_dial_over_loopback() {\n 2423│        let a = NetHost::start(hermetic(&Identity::generate())).expect(\"host a\");\n 2424│        let b_id = Identity::generate();\n 2425│        let b = NetHost::start(hermetic(&b_id)).expect(\"host b\");\n     │...\n 2450│\n*2451│    // [unit->REQ-RC-1] the broker loopback conn (M12 W1.5, B1): `dial_loopback`\n 2452│    // mints a REUSED singleton whose origin is this node's OWN id, and\n 2453│    // `open_stream` on it cross-wires a PAIR in-process — an operator row and a\n 2454│    // peer row (the one the dispatcher serves) — with the operator's writes\n     │...\n 2456│    #[test]\n*2457│    fn loopback_conn_is_a_reused_singleton_that_cross_wires_a_stream_pair() {\n 2458│        let id = Identity::generate();\n 2459│        let host = NetHost::start(hermetic(&id)).expect(\"host start\");\n 2460│\n 2461│        // Singleton: reused, own-node origin (never a foreign/handshake id).\n*2462│        let (c1, hex1) = host.dial_loopback();\n*2463│        let (c2, hex2) = host.dial_loopback();\n*2464│        assert_eq!(c1, c2, \"the loopback conn is a reused singleton\");\n 2465│        assert_eq!(\n 2466│            hex1,\n 2467│            id.public_key().to_hex(),\n*2468│            \"loopback origin is THIS node's own id (broker-minted)\"\n 2469│        );\n 2470│        assert_eq!(hex2, hex1);\n*2471│        assert_eq!(host.conn_count(), 1, \"one loopback conn held\");\n 2472│\n 2473│        // open_stream mints the cross-wired pair: operator row + peer row.\n*2474│        let op_stream = host.open_stream(c1, crate::msg::StreamLifetime::Durable).expect(\"open loopback stream\");\n 2475│        let infos = host.stream_infos();\n 2476│        assert_eq!(infos.len(), 2, \"operator + peer rows\");\n 2477│        let op = infos\n     │...\n 2522│        let host = NetHost::start(hermetic(&Identity::generate())).expect(\"host\");\n*2523│        let (conn, _) = host.dial_loopback();\n 2524│\n 2525│        let cb_stream = host\n 2526│            .open_stream(conn, crate::msg::StreamLifetime::ConnectionBound)\n     │...\n 2590│\n*2591│    // [unit->REQ-SHELL-4] the loopback tunnel pair under BACKPRESSURE (M11-W3,\n 2592│    // doyle gate condition): a retentive shell-side stream that is not drained while\n 2593│    // the owner keeps writing must NOT lose bytes AND must NOT grow unbounded — the\n 2594│    // read pump lags the sender at the cap. After the reader catches up, every byte\n     │...\n 2596│    #[test]\n*2597│    fn loopback_tunnel_backpressures_without_loss() {\n 2598│        use std::sync::atomic::{AtomicUsize, Ordering as AtOrd};\n 2599│        let id = Identity::generate();\n 2600│        let host = Arc::new(NetHost::start(hermetic(&id)).expect(\"host start\"));\n*2601│        let (owner, shell) = host.open_loopback_pair().expect(\"open tunnel pair\");\n 2602│        const CAP: usize = 4;\n 2603│        host.mark_stream_retentive(owner, CAP);\n 2604│        host.mark_stream_retentive(shell, CAP);\n     │...\n 2624│\n*2625│        // Let the pump fill the shell ring to the cap and PARK (backpressure).\n 2626│        let mut peak = 0usize;\n 2627│        for _ in 0..200 {\n 2628│            if let Some(d) = host.ring_len(shell) {\n     │...\n 2632│            if sent.load(AtOrd::Relaxed) >= N {\n*2633│                break; // (shouldn't reach N while we never drain — backpressure)\n 2634│            }\n 2635│        }\n 2636│        assert!(peak > 0, \"the pump delivered some frames to the shell ring\");\n     │...\n 2638│            host.ring_len(shell).unwrap() <= CAP,\n*2639│            \"the retentive ring never exceeds the backpressure cap (bounded)\"\n 2640│        );\n 2641│        assert!(\n 2642│            sent.load(AtOrd::Relaxed) < N,\n*2643│            \"the un-drained consumer lags the sender (backpressure engaged, not all sent)\"\n 2644│        );\n 2645│\n 2646│        // Now drain to completion — the sender unblocks, every byte arrives in order.\n     │...\n 2662│        assert_eq!(got.len(), expected.len(), \"no byte lost or duplicated\");\n*2663│        assert_eq!(got, expected, \"byte-exact, in order, under backpressure\");\n 2664│    }\n 2665│\n 2666│    // [unit->REQ-HAZARD-BROKER-QUIC-DEADLINE] the broker-side QUIC-op bound (the\n     │...\n 2676│\n*2677│        // A never-completing await (the black-holed-peer shape) fails at the bound.\n 2678│        let started = std::time::Instant::now();\n 2679│        let err = host\n 2680│            .bounded_block_on(std::future::pending::<io::Result<()>>())\n     │...\n 2752│        let host = NetHost::start(hermetic(&Identity::generate())).expect(\"host start\");\n*2753│        let (owner, shell) = host.open_loopback_pair().expect(\"open pair\");\n 2754│        host.send_stream(owner, b\"{\\\"kind\\\":\\\"request\\\",\\\"session_id\\\":1}\\n\", false)\n 2755│            .expect(\"send opener\");\n 2756│        // The pump is async — wait for the shell row to hold the opener.\n     │...\n 2785│    // worker's reply send is fire-and-forget (brain returns before the broker\n*2786│    // processes the envelope), so the acked retire on its own conn can be\n 2787│    // processed FIRST. Pre-fix, retire removed the entry: the late send hit\n 2788│    // no-such-stream (or the drop reset the never-finished send half) and the\n 2789│    // requester got torn-EOF instead of its reply (wansend NoReply, CI red\n     │...\n 2793│        let host = NetHost::start(hermetic(&Identity::generate())).expect(\"host start\");\n*2794│        let (owner, shell) = host.open_loopback_pair().expect(\"open pair\");\n 2795│        // The requester's one-shot request: write + FIN (the rest/sync shape —\n 2796│        // the row is born finished on the serving side).\n 2797│        host.send_stream(\n     │...\n 2869│    /// A connected local-socket pair whose SERVER send half wraps as the\n*2870│    /// `SharedSend` a subscriber seat writes through — `controller_socket_pair`'s\n 2871│    /// idiom (broker.rs) with a caller-chosen write bound.\n 2872│    fn seat_socket_pair(\n 2873│        bound: Duration,\n 2874│    ) -> (\n*2875│        SharedSend,\n 2876│        interprocess::local_socket::Stream,\n 2877│        interprocess::local_socket::RecvHalf,\n 2878│    ) {\n     │...\n 3118│    #[test]\n*3119│    fn displaced_writer_failure_never_cancels_the_successor_lease() {\n 3120│        let (a, client_a, _ra) = seat_socket_pair(Duration::from_secs(5));\n 3121│        let (b, _cb, _rb) = seat_socket_pair(Duration::from_secs(5));\n 3122│        let mut log = StreamLog::new(11, 64);\n     │...\n 3209│        let host = NetHost::start(hermetic(&Identity::generate())).expect(\"host start\");\n*3210│        let (owner, _shell) = host.open_loopback_pair().expect(\"open pair\");\n 3211│        let lease = host.stream_lease(owner).expect(\"lease handle\");\n 3212│        assert!(!lease.is_canceled());\n 3213│        host.send_stream(owner, b\"ok\", false).expect(\"a live lease serves\");\n     │...\n 3231│        let host = NetHost::start(hermetic(&Identity::generate())).expect(\"host start\");\n*3232│        let (owner, shell) = host.open_loopback_pair().expect(\"open pair\");\n 3233│        host.send_stream(owner, b\"{\\\"subnet\\\":\\\"s\\\",\\\"instance\\\":{}}\\n\", false)\n 3234│            .expect(\"bytes onto the peer row\");\n 3235│\n     │...\n 3267│    #[test]\n*3268│    fn unsubscribe_releases_only_the_callers_seat_joins_writer_and_is_idempotent() {\n 3269│        let host = NetHost::start(hermetic(&Identity::generate())).expect(\"host start\");\n*3270│        let (owner, shell) = host.open_loopback_pair().expect(\"open pair\");\n 3271│        let _ = shell;\n 3272│\n 3273│        let (a, _ca, _ra) = seat_socket_pair(Duration::from_secs(5));\n     │...\n 3315│        // the retired_row_still_flushes_a_late_reply contract above).\n*3316│        let (owner, shell) = host.open_loopback_pair().expect(\"pair 1\");\n 3317│        let (sub, _c, _r) = seat_socket_pair(Duration::from_secs(5));\n 3318│        host.subscribe_stream(shell, Arc::clone(&sub), 0).expect(\"subscribe\");\n 3319│        let (rows_before, seats_before) = host.stream_counts();"},"isError":false,"timestamp":1789523994069}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_azmxnHAAueQRWexQOo5yFhI6|fc_0d104aba2b0f3edc016aa9f824173887d18c0d1a569ad5a341","toolName":"read","startedAt":"2026-09-16T02:00:04.166Z","args":{"path":"docs/KNOWN-HAZARDS.md:750-753,764-766,811-815:raw"},"intent":"Reading complete historical hazards"},"id":"0e16a9d9","parentId":"b03dcd6d","timestamp":"2026-09-16T02:00:04.166Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_HWXzUOeS6Y6h2UGMakJJYQfa|fc_0d104aba2b0f3edc016aa9f824174887d1ac7743d0abffbb2b","toolName":"read","startedAt":"2026-09-16T02:00:04.166Z","args":{"path":"traceable-reqs.toml:2376-2403:raw"},"intent":"Reading exact drain requirements"},"id":"78f88621","parentId":"0e16a9d9","timestamp":"2026-09-16T02:00:04.166Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_GvMjuagwuo4AXDFyZxwtaObd|fc_0d104aba2b0f3edc016aa9f824175087d1abcc610491b2c646","toolName":"hub","startedAt":"2026-09-16T02:00:04.167Z","intent":"Sharing controller attribution finding"},"id":"54a8a67f","parentId":"78f88621","timestamp":"2026-09-16T02:00:04.167Z"}
{"type":"message","id":"e5fe4564","parentId":"54a8a67f","timestamp":"2026-09-16T02:00:04.167Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":""},{"type":"toolCall","id":"call_azmxnHAAueQRWexQOo5yFhI6|fc_0d104aba2b0f3edc016aa9f824173887d18c0d1a569ad5a341","name":"read","arguments":{"i":"Reading complete historical hazards","path":"docs/KNOWN-HAZARDS.md:750-753,764-766,811-815:raw"},"intent":"Reading complete historical hazards"},{"type":"toolCall","id":"call_HWXzUOeS6Y6h2UGMakJJYQfa|fc_0d104aba2b0f3edc016aa9f824174887d1ac7743d0abffbb2b","name":"read","arguments":{"i":"Reading exact drain requirements","path":"traceable-reqs.toml:2376-2403:raw"},"intent":"Reading exact drain requirements"},{"type":"toolCall","id":"call_GvMjuagwuo4AXDFyZxwtaObd|fc_0d104aba2b0f3edc016aa9f824175087d1abcc610491b2c646","name":"hub","arguments":{"i":"Sharing controller attribution finding","op":"send","to":"AttachReplayRca","message":"Verified conn17/70/88/104 all session1 stream1 by=c96d… attach_gen1789523756051; all SLOT_CLOSED cause=detach. Eviction token for 70/88/104 comes note_controller_write_retired (broker3824), not logical stall_evict: paired stream-sub-writer-poison + TimedOut shared gate. old_gen_drained excludes all by:Some even local-loopback RC. Need your narrow wait mechanism/discriminators, no replay implementation investigation overlap."},"intent":"Sharing controller attribution finding"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":18611,"output":249,"cacheRead":63872,"cacheWrite":0,"totalTokens":82732,"reasoningTokens":15,"cost":{"input":0.18611000000000003,"output":0.012450000000000001,"cacheRead":0.063872,"cacheWrite":0,"total":0.262432}},"stopReason":"toolUse","timestamp":1789523994122,"responseId":"resp_0d104aba2b0f3edc016aa9f81a5d0487d1b568e6d01dcbf9d3","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"rs_0d104aba2b0f3edc016aa9f81c9e8087d180366434cb7a298f","type":"reasoning","content":[],"encrypted_content":"gAAAAABqqfgk0Uy1rxxTbt7X6B51AE9ayzlXWFE9NIR_lisgunfcCT9-bJNaRmq-Cl2lN_WVv6Wd9EBHwT-kSCw76d5j53fxW-ojgeRHsRfczkuBr2QKxVYVvb_XT2pNhRxyQy8N5iTGDishfq0F_DDWCH5AIDRkAcUgA3vQ-glJwq3t8G_gfCdUAWjBH-KDn8CKMOysDsQoanAGlivyI0yIVpCYh6N8f5fJB96Tx3NgiHeoVrg6qb-4HWLwT7vuVVPe_SSXPN6thZXj28GCVJuNCRM3I5j-ipiapbLmQA3tsrVOgH8Inx_7Iode5E6MWmKSo8kyGoLfMgCHop__JSL632qTDKSC3yXq5hww4cWeLyP7czLsWlg3HOIzjIJSqhHmlRZz9i1_hA9Z993tOuKBSTIPyfgGDb7Da8j68UEz1T0iXNQjYTAI-hFV3uR8FFnZpCcQQaoMU7iYqTYn9O_zKUn-HJDQlxY7rs6NheJQQge4woX3bA2BE2Oi6TzwMaJXfaQKZ8RWBTODdj3X6jJi3JLccI7FYz7S6DNWds_K-Wa6KqMPsNlTrCNxjVw0CHDZlMVTtD4RWpj5Eo0dxY3ZFc8oF4Kj-aFVROek1E7mIInTTU26buBONNTf5Z6yQEl0QRXr-Qj_e0ZdHLqzATTS612PCfbakx18Aw8Q3DdViFmwyW1hJkmDph0JRinD_UVZpzjHgm9-ShyP8D-4xKCGRQNAaGCWS310HFailmnWwGnNPFOCIMN3ct3pFqxjQWlqYjE9Jlp70vAgZZRaY-jhy8Y_pXBUrG6_9t0IzlR80OtvBA9rF4h1myqj5A4UYleOBEa3bTtoF_cURw_S4Fxv7jkazgakCp4lerzzs9F57PjK1iSbWQnGtGysU7E834QHNzMcOLUUQBxgO8mxFmkqgAvrBTlGK56ToeUyH3hQS42Cs2yEdUMINwrDFisTcpAgFInMQiwJnIwgDzjshGZ5T7NM9FbIbgpUfAP0MEAB7BiWTE033NO-9wA6ox1u8DwfJePolXLokntQ6Y42UdIH_Ua5iGnd9c-UDETy5iwj6-L9yIUVIUUR6zaf5itjnJs8VR2RusUiQwPUoTOW6qlmF2-Jv_-1az4ATiVePzzXtDB1sHiTt-2opBtYyyH6bmQy2dO45wF4Rgow904iId14B5QAjWYz4ofDSZ752Of_edFwwwewHYIjf9wQktM-Le2SZpGMYLkqyMBzPvPjjwtViPNCZrKMsCO4chs8FMllQewfsNGL0yuHKrBawUmVDdKOJUSbNV17LYJ57Y1ELHd4GKUasYAL6lsdPLkH2ILTErD0hCZC39mIc90A3ZWabv10ni245_eWzWCi9NktV-QSOpMn1AOhMg==","internal_chat_message_metadata_passthrough":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"},"summary":[],"metadata":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"}},{"id":"fc_0d104aba2b0f3edc016aa9f824173887d18c0d1a569ad5a341","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading complete historical hazards\",\"path\":\"docs/KNOWN-HAZARDS.md:750-753,764-766,811-815:raw\"}","call_id":"call_azmxnHAAueQRWexQOo5yFhI6","internal_chat_message_metadata_passthrough":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"},"name":"read","metadata":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"}},{"id":"fc_0d104aba2b0f3edc016aa9f824174887d1ac7743d0abffbb2b","type":"function_call","status":"completed","arguments":"{\"i\":\"Reading exact drain requirements\",\"path\":\"traceable-reqs.toml:2376-2403:raw\"}","call_id":"call_HWXzUOeS6Y6h2UGMakJJYQfa","internal_chat_message_metadata_passthrough":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"},"name":"read","metadata":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"}},{"id":"fc_0d104aba2b0f3edc016aa9f824175087d1abcc610491b2c646","type":"function_call","status":"completed","arguments":"{\"i\":\"Sharing controller attribution finding\",\"op\":\"send\",\"to\":\"AttachReplayRca\",\"message\":\"Verified conn17/70/88/104 all session1 stream1 by=c96d… attach_gen1789523756051; all SLOT_CLOSED cause=detach. Eviction token for 70/88/104 comes note_controller_write_retired (broker3824), not logical stall_evict: paired stream-sub-writer-poison + TimedOut shared gate. old_gen_drained excludes all by:Some even local-loopback RC. Need your narrow wait mechanism/discriminators, no replay implementation investigation overlap.\"}","call_id":"call_GvMjuagwuo4AXDFyZxwtaObd","internal_chat_message_metadata_passthrough":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"},"name":"hub","metadata":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"}}]},"duration":10041.482499999925,"ttft":2536.1646000000183,"completedAt":1789524004164,"contextSnapshot":{"promptTokens":82483,"nonMessageTokens":10480,"compactionEpoch":0}}}
{"type":"message","id":"b505b610","parentId":"e5fe4564","timestamp":"2026-09-16T02:00:04.168Z","message":{"role":"toolResult","toolCallId":"call_GvMjuagwuo4AXDFyZxwtaObd|fc_0d104aba2b0f3edc016aa9f824175087d1abcc610491b2c646","toolName":"hub","content":[{"type":"text","text":"Delivered to 1 peer(s):\n- AttachReplayRca: injected"}],"details":{"op":"send","from":"ControllerDrainRca","to":"AttachReplayRca","receipts":[{"to":"AttachReplayRca","outcome":"injected"}]},"isError":false,"timestamp":1789524004168}}
{"type":"message","id":"dc721943","parentId":"b505b610","timestamp":"2026-09-16T02:00:04.171Z","message":{"role":"toolResult","toolCallId":"call_azmxnHAAueQRWexQOo5yFhI6|fc_0d104aba2b0f3edc016aa9f824173887d18c0d1a569ad5a341","toolName":"read","content":[{"type":"text","text":"- **Failure (paid-for, rig-CONFIRMED 2026-07-06/07 — `NtSuspendProcess` on the brain, no update involved):** a controller's writer thread does a BLOCKING socket write to its brain subscriber conn. When that brain is suspended (or black-holed) the write never returns. The output-driven eviction path (`append` → `mark_controller_gone`, bounded by 7.12's `CONTROLLER_WRITE_DEADLINE`) only fires on NEW output, and `reap_dead_controller` (7.29 companion) only caught a writer that had EXITED — so a controller BLOCKED (not exited) on an IDLE session was evicted by NEITHER. Within seconds: attached output froze, `detach` did not release the control stamp (release routed through the wedged conn), reattach was REFUSED (`controlled-by` latched), `rc --take` hung, and `daemon status` still read healthy. Every brain cycle — including every `update apply` — has such a freeze window; a stalled or slow-draining NEW brain (6 sessions + psyches + WAN on the incident night) makes it a PERMANENT wedge until a bounce. `brain.ready ≠ subscribers drained` — why the 22:47 apply \"promoted\" while frozen.\n- **Invariant:** no broker client ever observes a wedge because the BRAIN stalled. (1) The controller's writer publishes its IN-FLIGHT-write window (`write_blocked_since`: `Some(Instant)` set immediately before the blocking `write_frame`, `None` after — the mutex never held ACROSS the write); a reader sees a `Some` older than `BRAIN_WRITE_DEADLINE` (15 s = 3× the 7.12 controller bound, measured full-with-**zero writer progress**, so a slow-but-DRAINING brain never trips it) as a WEDGED writer and stall-evicts it. Eviction is triggered TIME-based, not output-driven: at the take/reattach path (`resolve_subscribe`, so control self-heals against broker truth even on an idle session) AND at the `KIND_SESSIONS` reap. (2) The stamp release is BROKER-side — `clear_controller` re-stamps `driven_by`/`controlled` with no brain round-trip, so `controlled-by` can never latch behind a dead transit; `rc --take` and reattach-refuse consult live `has_controller()` truth. (3) The evict is OBSERVABLE, never a silent absence — a per-evict broker-log line plus a broker-global tally (`count` + `last_ms`) surfaced on `daemon status` (a brain that never returns is then a diagnosable fact). Idle ≠ wedged: an idle controller (writer parked on `rx.recv()`, `write_blocked_since` `None`) is NEVER evicted; only a genuinely-blocked writer is. Complements 7.12 (which bounds a merely SLOW controller's Full channel) and 7.29 (which converges stamps to broker truth) — this covers the BLOCKED-conn case both leave open. Deferred (Q2, minimal-plus): the WHOLESALE take/release verb-set migration off the brain stays out; the suspend-brain int rig is the arbiter — if it passes with the stall-evict shape, the wholesale move is scope creep. **Residual — BOUNDED by W2, CLOSED by W3 (`REQ-UPDATE-PROMOTE-DRAINED`):** the stall-evict shrank the update-apply FALSE-PROMOTE window from unbounded to `BRAIN_WRITE_DEADLINE` (15 s) — `brain.ready` could still arrive inside that window while blocked writes pended on the OLD-generation subscriber conn (`brain.ready ≠ subscribers drained`, the 22:47 apply). W3 CLOSES it: the ADR-0018 brain-trial promotion gate (`brainproc::run_trial`) now LATCHES ready-seen and promotes only on ready **AND** `TrialEnv::old_gen_drained()` — a broker-truth read (`Broker::any_local_controller_wedged`, the brain's own `by:None` conn wedged past the deadline; a remote conn is excluded, it is the W2 concern not an old-gen brain conn), no brain round-trip. Ready-but-never-drained elapses the trial window → kill + rollback (conservative; never a false-promote onto a wedge). The false-promote int rig (`tests/false_promote.rs`) exercises the promotion path itself, RED-first.\n- **spt-core mapping:** `broker.rs` — `ControllerSink::write_blocked_since` + `controller_writer` (marks the in-flight window around both the initial-batch and live-loop writes) + `controller_write_stalled` (pure predicate over injected `now`/`deadline`) + `stall_evict_controller` (Inline stamp at `resolve_subscribe`, Deferred stamp at `reap_dead_controller` so the reap closure does no I/O under the shared sessions lock) + `record_stall_evict`/`stall_evict_stats` (the tally) + `dispatch_stall_evicts` (`KIND_STALL_EVICTS` IPC); `brain.rs` `Brain::stall_evicts`; `cli.rs` `render_stall_evict_line` (daemon-status surface, public wording). unit = `controller_write_stall_predicate_distinguishes_wedged_from_idle` + `stalled_incumbent_is_evicted_on_reattach_and_control_released` + `reap_evicts_a_blocked_writer_not_only_an_exited_one` + `stall_evict_line_surfaces_only_a_real_tally`. int = `brain_decouple.rs` `suspended_brain_controller_is_stall_evicted_take_completes_viewer_ticks` — the deterministic in-process ANALOG of a suspended brain (a REMOTE controller that stops reading parks the broker writer in `write_frame`, the same primitive `NtSuspendProcess` produces), asserting EXACTLY four things: (1) an attached viewer keeps ticking during the wedge, (2) a second operator's `Control` subscribe COMPLETES via stall-evict (RED-first: remove the `resolve_subscribe` evict → `BusyControlled` + `evicts=0`), (3) it resolves promptly (well under the watchdog), (4) the evict is tallied (observability); Linux leg on kitsubito. It does NOT re-assert resume + cursor-replay byte-identity — W2's delta does not touch the replay path, so output integrity stays covered by the pre-existing ring/cursor units (effect-journal replay dedup, the `delivered_through` no-rewind CAS `advanced_cursor`, clean-repaint cold-attach). The live `NtSuspendProcess`/`SIGSTOP` run on a real brain child remains the MANUAL field rig from the incident night, not a CI test.\n- **Source:** LIFECYCLE-TRUTH W2 (doyle rig `NtSuspendProcess` root + design ruling `docs/W2-DESIGN-RULING.md` @60ec0ed — minimal-plus scope, `BRAIN_WRITE_DEADLINE` 15 s, passive resubscribe + observability rider; todlando build 2026-07-07). The v0.27→v0.28 update-wedge night's flagship defect.\n\n…\n\n- **Failure (paid-for, field capture + deterministic Windows repro 2026-07-09 — hertz RCA, UPDATE-WEDGE round-4):** `controller_writer` held the `SharedSend = Arc<Mutex<SendHalf>>` guard ACROSS a blocking `write_frame`. On Windows, `interprocess 2.4.2` routes the send to `WriteFileEx` + `SleepEx(INFINITE, alertable)` — this local-socket send path exposes **no supported write timeout** (`set_timeout` → `Unsupported`). An `rc --take` controller consumer that stopped reading blocked the write **indefinitely** (~127.95 s in the capture, released only when a brain restart tore the connections down). The 7.36 logical `stall_evict_controller` removed the controller ROLE but neither canceled the in-flight pipe write, closed the physical connection, nor called `CancelIoEx` — the detached writer kept its `SharedSend` clone and live stack-owned mutex guard, and every other write queued at that gate wedged behind it. Load-gated: needs a real `seq>0` frame plus a non-draining consumer (seq-0 boot conns and a quiescent update are clean — why v0.30.5's controlled apply passed). `CTRL_WRITE_LOCKED wait_us=0` on all four capture writers exonerates the mutex convoy: the block is INSIDE the OS write, after lock acquisition. Neither existing deadline bounds it: `CONTROLLER_WRITE_DEADLINE` is output-driven channel-full handling (a writer blocked on its first owned frame never fills the channel); `BRAIN_WRITE_DEADLINE` is an age predicate sampled opportunistically, not an I/O timer.\n- **Invariant:** every write on a physical broker connection rides ONE broker-owned bounded/cancelable framed-write primitive with an **independent out-of-band aborter**. The `SendHalf` never leaves the conn object (a bypass cannot compile). An ABSOLUTE deadline stamped at write entry covers BOTH the serialized gate-wait AND OS write completion; a per-conn watchdog fires at that deadline and aborts the in-flight op — never relying on the write returning or on opportunistic stall-evict sampling. On deadline / partial write / cancellation / unknown completion: (1) poison the whole physical conn idempotently, (2) abort read+write so `handle_conn` reaches existing EOF cleanup, (3) wait for the canceled op to report completion before releasing its buffer, (4) NEVER reuse the conn (a timed-out length-prefixed frame may be partially written — reuse would corrupt framing), (5) finish/join the retired writer before reporting physical cleanup. Controller (replay + live), viewer, dispatch-reply, and nethost stream/presence writes ALL route through the primitive — any raw unbounded write left behind the gate preserves the failure class. NO additional output queue (the bounded queue + isolated writer already exist; the block is BELOW them); NO `PIPE_NOWAIT` (recorded mid-frame corruption risk). The logical stall-evict (7.36) stays as the ROLE-release trigger; physical retirement authority is the bounded write itself. Unix keeps its existing semantics under the same poison/retire invariant.\n- **spt-core mapping:** `spt-daemon/src/conn.rs` `BrokerConn` (Condvar-gated serialized write gate + poisoned state + per-in-flight op id + watchdog + `abort_physical()`: cfg(windows) `CancelIoEx` → completion handshake → `DisconnectNamedPipe` raw kernel32 externs; cfg(unix) `UnixStream::shutdown(Both)`); `broker.rs` `SharedSend = Arc<BrokerConn>` + `controller_writer`/`viewer_writer`/`send_frame`/`send_error` routed; `nethost.rs` stream-log/presence sends routed. int = `brain_decouple.rs` `non_draining_controller_stall_evict_releases_writer_and_connection` (`#[cfg(windows)]`, the hertz 9-step: real PTY burst, black-holed remote controller, exactly one logical stall-evict via an independent draining conn, the black-holed conn's `handle_conn` finishes ≤2 s of logical release WITHOUT dropping the client, a write on the old client conn fails, a fresh controller resumes from the frozen cursor) — RED-first pre-fix; the cross-platform companion `suspended_brain_controller_is_stall_evicted_take_completes_viewer_ticks` stays green on both OSes.\n\n…\n\n- **Invariant:** subscriber I/O never holds the `StreamLog` mutex (enqueue-under-lock, I/O outside, bounded per-subscriber queue; overflow = detach + resume-from-cursor); replay and fan-out halt at the FIRST failed subscriber write and the failed subscriber is REMOVED (attach + append + finish, all sites, `PresenceLog` mirrors included); the ring-peek fallback fires ONLY on the explicit unsupported-verb/old-broker answer; a deadline-poisoned replay is circuit-broken (global backoff, no reinstall until the prior subscriber is fully gone — an upgrade of bounded retry, never a revert to burn-the-claim abandonment); attach/detach validate ownership/generation so a stale worker never detaches a replacement controller (covers the unfinished-stale steal shape that finished-row retirement cannot see).\n- **spt-core mapping:** `spt-daemon/src/nethost.rs` `StreamLog::attach`/`push`/`finish` + `PresenceLog` mirrors + serve/wire worker pairing; `spt-daemon/src/dispatch.rs` `first_line` fallback arm + claim retry/backoff + worker spawn; SharedSend bounded-write deadline (round-4, KH 7.7 kin).\n- **Source:** REDISPATCH-STALL W1 (hertz v0.34 field RCA 2026-07-16, doyle seam-verified same day; ADR-0038 Amendment 2026-07-16; `.claude/reports/2026-07-16-redispatch-stall/rca-summary.md`; distinct from `update --restart`'s bounded 30s rc reconnect loop).\n\n### 7.44 Streams and seats on a long-lived connection must have bounded lifetime — one-way rows terminal at FIN, seats released at serve completion, no per-chunk full-state rewrites in a drain loop  `[REQ-HAZARD-REGISTRY-STALL]`"}],"details":{"fileSize":360311,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\docs\\KNOWN-HAZARDS.md"}}},"isError":false,"timestamp":1789524004171}}
{"type":"message","id":"e5cc2b58","parentId":"dc721943","timestamp":"2026-09-16T02:00:04.181Z","message":{"role":"toolResult","toolCallId":"call_HWXzUOeS6Y6h2UGMakJJYQfa|fc_0d104aba2b0f3edc016aa9f824174887d1ac7743d0abffbb2b","toolName":"read","content":[{"type":"text","text":"id = \"REQ-UPDATE-PROMOTE-DRAINED\"\ntitle = \"W3 (LIFECYCLE-TRUTH, mechanic-d MOVED FROM W2 per doyle gate verdict @e5ae7a9 — binding): the update-apply brain-generation promotion completes only when the OLD generation's broker subscriber connection is CLOSED or stall-EVICTED — never while blocked writes still pend on it. ROOT: `brain.ready` != subscribers drained; W2's stall-evict (REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE) only BOUNDS the false-promote window to BRAIN_WRITE_DEADLINE (15s), it does NOT close it — a new brain can signal ready inside that window while the old gen's conn is still wedged, so the apply 'promotes' onto a still-frozen control plane (the 22:47 incident-night false-promote). FIX: the promotion gate (ADR-0018 brain-trial, brainproc.rs) adds an explicit DRAINED precondition — promote only on ready AND old-gen-subscriber-drained (conn closed OR stall-evicted); the drained signal reads broker truth (the W2 stall-evict tally / the old conn's liveness), no brain round-trip. The residual W2 left open, now closed. Int = a FALSE-PROMOTE rig that exercises the promotion path itself: an old-gen subscriber conn held wedged past ready must NOT promote until it drains (RED-first: ready-alone promotes).\"\nrequired_stages = [\"impl\", \"unit\", \"int\"]  # ACTIVATED LIFECYCLE-TRUTH W3 drained-condition slice (todlando 2026-07-07, work-start); brainproc.rs run_trial DRAINED gate + broker any_local_controller_wedged truth + false-promote int rig — closes the W2 15s residual (KH 7.36)\n\n[[requirements]]\nid = \"REQ-UPDATE-TRIAL-DRAIN-DRIVE\"\ntitle = \"UPDATE-WEDGE (counter-54, doyle-ruled 2026-07-09 — regression of the v0.29.0 seamless brain-swap): a brain generation DRIVES the broker's controller-liveness reap (a KIND_SESSIONS poll) each heartbeat throughout its boot/trial loop, so a hard-KILLED prior generation's black-holed LOCAL controller conn (by:None) is stall-evicted within the trial window and can never permanently strand the promotion DRAINED gate. ROOT (2026-07-09 field freeze, `spt update fetch --apply` v0.30.0->v0.30.2 froze all 7 live PTYs ~30s then rolled back): the promote gate (run_trial, brainproc.rs:657-661) needs BOTH `ready_generation==gen` AND `old_gen_drained()`; `old_gen_drained()` = `!any_local_controller_wedged()` (brainproc.rs:534) is a PURE READ of `write_blocked_since` (broker.rs:2703) — it never DRIVES the evict. The evict (`stall_evict_controller`, broker.rs:1039, same 15s `brain_write_deadline` the wedge-read uses) only runs via `reap_dead_controller` (broker.rs:967, severed->drop ELSE stall-evict) inside the KIND_SESSIONS snapshot closure (broker.rs:2879). During the isolated brain-trial window NOTHING polls KIND_SESSIONS: the old brain was hard-killed (`child.kill()`, brainproc.rs:851) so its local controller conn black-holes (live PTYs keep writing to a dead Windows named pipe -> writer BLOCKS, never EOFs) -> `write_blocked_since=Some` -> wedged=true for the full 30s -> gate false -> `WindowElapsedAlive` -> kill+rollback. The candidate DID reach ready (write_ready, brainproc.rs:211, runs before the loop; the v53 log's NET_FAMILY_GATE/PAIR_MEET_UP prove the loop was entered) — so `BRAIN_TRIAL_TIMEOUT: candidate alive but never ready` is the MISLEADING ready-stamped-but-never-DRAINED case, NOT a resume_sessions hang. SECOND LEG (recovery): post-rollback the gen-2 brain's KIND_SESSIONS poll finally reaps -> `BRAIN_SUBSCRIBER_STALL_EVICT:1` fires >15s late -> a session stayed black-holed through recovery -> continued freeze. ONE root, BOTH legs. FIX (BRAIN-SIDE, self-applying — doyle ruled brain-side to AVOID a broker-side coordinated-restart flag): the boot/trial heartbeat loop (brainproc.rs run_brain, currently only `net_status` at :244) also issues `Brain::sessions()` (KIND_SESSIONS, brain.rs:1397 — already exists) every heartbeat, driving the LIVE older broker's ALREADY-SHIPPED reap (>=v0.29.0 LIFECYCLE-TRUTH; the field-stuck broker is v0.30.0 so it HAS it). The old-gen wedged conn is stall-evicted ~15s < the 30s window -> `old_gen_drained()` flips true -> PROMOTE; the same poll on the rollback/recovery brain reaps promptly -> no >15s black-hole -> kills the STALL_EVICT recovery leg. Drive it on the FIRST heartbeat (no one-tick wait) and on BOTH the trial candidate AND the recovery brain (one loop covers both). SELF-APPLIES because it drives the current broker's existing reap verb — v54's brain fixes the v0.30.0->v54 update with NO coordinated broker restart. Forward-compat: brokers <v0.29.0 have neither the DRAINED gate nor the reap, so older-broker updates never hit this path — no regression. Composes with REQ-UPDATE-PROMOTE-DRAINED (the gate this un-strands) + REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE (the stall-evict it drives) + REQ-CONTROLLER-LIVENESS-REAP (the reap verb). Int = the brain-swap-under-live-sessions rig: (1) with a hard-killed prior gen holding a wedged by:None controller past 15s, the new brain PROMOTES within the window AND sessions stay served across the cycle (RED-first: without the drive, WindowElapsedAlive->rollback); (2) a deliberately-failing trial auto-rolls-back AND the restored brain re-drives EVERY session with NO BRAIN_SUBSCRIBER_STALL_EVICT.\"\nrequired_stages = [\"doc\", \"impl\", \"int\"]  # ACTIVATED counter-54 UPDATE-WEDGE fix (doyle ruled brain-side 2026-07-09; todlando builds). doc = a CONTEXT.md/LIFECYCLE note that the trial/boot brain drives controller-reap so a killed prior generation cannot strand old_gen_drained. impl = the brain.sessions() KIND_SESSIONS drive in run_brain's heartbeat loop (brainproc.rs, first-tick + every tick, trial + recovery brain). UNIT DELIBERATELY DROPPED (doyle FORK-2 ruling 2026-07-09): the impl is ONE unconditional branchless brain.sessions() call in the heartbeat loop — no pure branch/decision to unit-test; forcing a call-ordering unit would require a trait-seam refactor of the broker-critical brain heartbeat path purely to satisfy the stage (test theater, adds risk for no coverage). The behavior is load-bearing-guarded by the int rig (RED-first bites the PRODUCTION drive), and the broker-side reap the drive triggers is already unit-covered (broker.rs reap_evicts_a_blocked_writer_not_only_an_exited_one). int = the brain-swap-under-live-sessions rig (FORK-1 ruling A: a NEW crates/spt/tests real-SUBPROCESS sibling of brain_survive.rs — real broker + real PTY + real supervise_brain->run_trial + a real `spt daemon brain` child running production run_brain self-driving sessions(); RED-first bites the production drive), both assertions: (1) promote-within-window with a wedged old-gen local controller + sessions stay served; (2) failed trial rolls back with sessions re-driven, NO BRAIN_SUBSCRIBER_STALL_EVICT.\n\n[[requirements]]\nid = \"REQ-BRAIN-UPDATE-RESTART-CLEAN-CLOSE\"\ntitle = \"SEED (DEFERRED, doyle 2026-07-09 — post-counter-54 root-hardening for UPDATE-WEDGE; mint now, impl a FUTURE milestone): on a PLANNED brain-restart (`BRAIN_UPDATE_RESTART`, the seamless update-apply brain-cycle), the outgoing brain's LOCAL (by:None) controller conns are GRACEFULLY CLEAN-CLOSED as the brain is cycled, instead of hard-killed and left to black-hole. ROOT (field-pinned 2026-07-09, daemon.stderr.log L24277-24303): the update-restart path hard-kills the outgoing brain (`child.kill()`, brainproc.rs:851); its live-agent controller conns then block on dead pipes (never EOF) → the broker reads them WEDGED (broker.rs:2695-2700) → the new candidate's promotion DRAINED gate (`any_local_controller_wedged`, broker.rs:2704) stays true until the W2 stall-evict matures (~15s). REQ-UPDATE-TRIAL-DRAIN-DRIVE (counter-54) makes the candidate DRIVE that reap so it promotes within the 30s window — but at a ~15s wedge-maturity hitch (frozen PTYs during the swap). A CLEAN close makes the conn 'simply absent → drained=false AT ONCE → fast promote' (broker.rs:2699-2700), ELIMINATING the hitch = truly seamless (honors the paradigm the field freeze broke). SUPERSEDES the earlier livehost-reattach framing of 'Fix Y': livehost is SPAWN-FRESH (fresh session uuid/pid per boot, nothing to re-attach — wrong site, and it never ran in the trial window); the correct site is the brain-cycle / update-restart path (a bounded graceful-drain of the outgoing brain BEFORE the kill). Non-trivial: hard-kill → bounded graceful drain; a drain that hangs must NOT wedge the swap (timeout then kill anyway, never block the update). Composes with REQ-UPDATE-TRIAL-DRAIN-DRIVE (defense-in-depth reap-drive REMAINS for any conn that still black-holes — a peer/relay conn, a drain-timeout kill) + REQ-UPDATE-PROMOTE-DRAINED (the gate) + REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE (the stall-evict). Int: a planned update-restart under a live-agent controller conn → the new candidate promotes WITHOUT waiting the ~15s wedge-maturity (drained reads false immediately, no STALL_EVICT), RED-first vs the current hard-kill-then-reap-drive ~15s hitch.\"\nrequired_stages = []  # SEED minted DEFERRED (doyle 2026-07-09); activate post-counter-54 at the milestone that delivers HITCH-FREE seamless updates. Supersedes the livehost-reattach 'Fix Y' framing (wrong site). The counter-54 fix (REQ-UPDATE-TRIAL-DRAIN-DRIVE reap-drive) makes updates SELF-HEAL in ~15s; this seed removes the residual hitch.\n\n[[requirements]]\nid = \"REQ-LIVEHOST-RECONCILE-TRIAL-SILENT\"\ntitle = \"SEED (DEFERRED investigation, doyle 2026-07-09 — UPDATE-WEDGE follow-up): determine WHY the trial/rollback brain's livehost reconcile loop did NOT drive the broker controller-reap (nor re-host the live agents) during the ~30s field update-trial window, when livehost polls `query_live_session_endpoints()` → `brain.sessions()` (KIND_SESSIONS) UNCONDITIONALLY every `LIVE_RECONCILE_INTERVAL_MS`=5000ms (livehost.rs:1026). CONTEXT (surfaced building the counter-54 rig): livehost's 5s KIND_SESSIONS poll drives the SAME broker `reap_dead_controller` sweep the fix drives — so it would otherwise reap the 15s-matured wedge by ~T20 < the 30s trial and SELF-HEAL. It didn't (field froze 30s → rollback), so the field trial-brain livehost was silent/delayed (PIN Q2: no `DAEMON_RESTART_RESUME` under gen-1/gen-2; the 30s kill landed before/around livehost's first reconcile tick). The counter-54 fix (REQ-UPDATE-TRIAL-DRAIN-DRIVE) puts a RELIABLE 500ms reap-driver in run_brain's CORE heartbeat loop, making the wedge-reap INDEPENDENT of livehost — so this does NOT block counter-54. But the livehost silence is a latent anomaly with a SECOND consequence: live-agent HARNESS re-hosting was also delayed ~30s (a separate freeze contributor). Investigate: does `spawn_live_host`'s reconcile thread start promptly on a trial-brain boot, or is its first tick delayed past the trial window? Does its brain conn / `query_live_session_endpoints` block against the swap/wedge state? Register concrete REQ(s) once the mechanism is pinned. RELATED: [[REQ-BRAIN-UPDATE-RESTART-CLEAN-CLOSE]] (if the outgoing brain's black-holed conns perturb the new brain's livehost conn setup).\"\nrequired_stages = []  # SEED deferred investigation (doyle 2026-07-09); non-blocking for counter-54 (the core-loop reap-drive in REQ-UPDATE-TRIAL-DRAIN-DRIVE covers the wedge independently of livehost). Activate once the livehost-silence mechanism is pinned.\n\n[[requirements]]\nid = \"REQ-BRAIN-RESUME-NO-CONTROL-STEAL\"\ntitle = \"UPDATE-WEDGE round 2 (v0.30.4, doyle-ruled 2026-07-09 — field incident on the counter-54 fetch--apply): a brain-respawn must NEVER steal, then stall-evict, the controller of a broker PTY session the daemon brain does not DRIVE. ROOT (field-pinned + SME, docs/UPDATE-WEDGE-2-RCA.md + docs/UPDATE-WEDGE-2-SME-todlando.md): `resume_sessions` (brain.rs:985) re-attaches EVERY session `KIND_SESSIONS` returns via `subscribe` = `subscribe_with(AttachIntent::Control, by:None)` (brain.rs:1450-1455). The docstring's 'a None identity never displaces (falls back to viewer)' is a MYTH for FREE / SAME-LOCAL-IDENTITY slots: `resolve_subscribe` (broker.rs:1134-1153) stall-evicts FIRST, then `become_controller` if the slot is now free OR the incumbent is also local (None==None) — viewer-fallback fires ONLY when a DIFFERENT REMOTE controls. So on a box with N spt-hosted broker PTYs, a brain-respawn STEALS the by:None controller of every free/local-controlled session (incl. the operator's LOCAL `spt rc`), which the daemon brain never drains (it hosts no PTY sessions — brainproc.rs:184) → 15s later `stall_evict_controller` (broker.rs:1039) releases driven_by → the session is UNCONTROLLABLE (Failure A). It also head-of-line-blocks the shared brain↔broker conn on the N-session controller-replay burst → every journaled `spt rc` retake deadlines ('brain IPC read deadline', the REQ-BROKER-ATTACH-JOURNAL-RESILIENT / #16 shared-conn symptom) → global rc failure on ALL N (Failure B). Field 2026-07-09 (operator-confirmed): fetch--apply 0.30.2→0.30.3 PROMOTED CLEANLY (the counter-54 fix worked) but under 7 spt-hosted PTYs (ALL with LOCAL by:None controllers) the resume SILENTLY STOLE all 7 (become_controller same-local re-take, NO Displaced notice → orphaned, output froze immediately, no rc-detach splash) + blocked every rc retake. The 5-vs-2 stall-evict split is ACTIVE-vs-IDLE, not remote-vs-local: 5 producing output → stolen writer blocked >15s → stall-evict; 2 idle → writer parked → no evict, but still silently stolen+frozen. The counter-54 promotion fix did NOT cause this — pre-existing resume-steal latent bug, hidden until N broker PTYs were present at a respawn; the single-black-holed-session A'-rig never exercised N-live-controllers-under-replay. FIX (brain-side): `resume_sessions` re-attaches as **Viewer** (`AttachIntent::Viewer`), NOT Control — a viewer never touches driven_by and is never stall-evicted (broker.rs:1063+ bounded try_send + private eviction), so steal-then-drop vanishes and the operator keeps/regains control; and it relieves shared-conn pressure (a slow viewer is DROPPED, never a 15s controller block) so rc retake gets through. Control ONLY for sessions the daemon brain genuinely DRIVES (empty set today → all become Viewer; forward-correct for the live-agent-adapter future). SECONDARY (escalation, ONLY if the gate shows residual B): stagger the resume re-attach + bound the viewer replay so the respawn burst can't saturate the conn. Int = the multi-broker-PTY-session RESPAWN rig (the coverage the A'-rig lacked): N real broker-spawned sessions with controllers producing output → real brain respawn/promote → assert (1) EVERY session keeps its controller across the swap (no stall-evict of a session the brain doesn't drive), (2) `spt rc` attaches/retakes IMMEDIATELY post-promote (no shared-conn saturation), (3) promotion still succeeds. RED-first: the current Control re-attach steals+evicts + deadlines rc. Composes with REQ-UPDATE-TRIAL-DRAIN-DRIVE (the orthogonal counter-54 promote fix), REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE (the stall-evict it stops mis-firing on non-driven sessions), REQ-BROKER-ATTACH-JOURNAL-RESILIENT (the #16 shared-conn resilience). Distinct from REQ-BRAIN-UPDATE-RESTART-CLEAN-CLOSE (that = the OLD brain's outgoing black-hole; this = the NEW brain's resume-steal).\"\nrequired_stages = [\"doc\", \"impl\", \"int\"]  # ACTIVATED v0.30.4 UPDATE-WEDGE-2 fix (doyle ruled 2026-07-09; todlando builds). doc = brain.rs resume docstring corrected (kill the by:None-never-steals MYTH) + a CONTEXT/LIFECYCLE note that a brain-respawn re-attaches non-driven sessions as viewer. impl = resume path re-attaches AttachIntent::Viewer (Control only for genuinely-driven sessions, empty today). unit DROPPED if the impl is a branchless intent-flip (same FORK-2 reasoning as REQ-UPDATE-TRIAL-DRAIN-DRIVE — add a unit only if a real driven-vs-not predicate branch exists). int = the multi-broker-PTY-session respawn rig, all 3 assertions RED-first.\n\n[[requirements]]\nid = \"REQ-BRAIN-RESUME-NO-CONN-DEADLOCK\"\ntitle = \"UPDATE-WEDGE round 3 (v0.30.5, doyle-ruled Option A 2026-07-09 — the v0.30.4 field-verify re-wedge, root code-PROVEN + dead-peer-INDEPENDENT): the daemon brain must NOT subscribe broker PTY sessions onto its own request/reply IPC conn — it has no consumer for that output and the subscription DEADLOCKS the conn. ROOT (todlando code-read, docs/UPDATE-WEDGE-2-ROUND3-CODEREAD.md; the net-runtime AND the counter-54 reap-drive were both FALSIFIED first — docs/UPDATE-WEDGE-2-ROUND3-RIG-VERDICT.md): a conn's send half is a single `SharedSend = Arc<Mutex<SendHalf>>` (broker.rs:78). Subscriber writer threads (`viewer_writer` broker.rs:1333/1342, `controller_writer` :1451) hold `send.lock()` ACROSS a BLOCKING `write_frame`; the dispatch reply path (`send_frame` :4221 → KIND_SESSIONS_REPLY / KIND_NET_STATUS_REPLY) needs the SAME lock. `resume_sessions` (brain.rs:1031→1054) subscribes every session as a Viewer onto the brain's MAIN conn — which is ALSO the brain's request/reply channel. The daemon brain hosts no PTY sessions (brainproc.rs:184) so run_brain never drains that output; it reads the conn only during the 500ms-heartbeat net_status()/sessions() calls (drain-and-DISCARD, `_ => continue`). When an actively-streaming session backs the conn up, a subscriber writer BLOCKS in write_frame holding send.lock() → the dispatch thread can't send the heartbeat reply → net_status()/sessions() never return → the heartbeat loop stalls → the brain never drains → the writer stays blocked = SELF-DEADLOCK on the brain conn's send mutex → every subscriber writer on it wedges → BRAIN_SUBSCRIBER_STALL_EVICT (controller writer blocked >15s). BOTH severities, one mechanism: RESPAWN (resume subscribes N + the replay burst floods the conn before the loop drains) AND STEADY-STATE (an active streamer's output between heartbeats fills the socket buffer). Counter-54 (REQ-UPDATE-TRIAL-DRAIN-DRIVE) added a 2nd per-heartbeat reply round-trip (sessions()) through the contended mutex — WIDENED the window (regression-window-exact), did not create it. The round-2 Viewer fix (REQ-BRAIN-RESUME-NO-CONTROL-STEAL) removed the STEAL but kept the brain a SUBSCRIBER — viewer_writer has the same send.lock()-across-write pattern AND viewers have no stall-evict valve — so v0.30.4 field-verify wedged again. FIX (Option A, brain-side, SEAMLESS): resume_sessions does NOT subscribe (drop the subscribe_with call; §3 verification guard confirmed NO brain consumer — digest/relay/net-consumer/shellwake/presence — reads the resumed subs). The brain conn then carries only request/reply → no subscriber backpressure → no deadlock. Keep the session_cursors seed only if harmless. Rides the brain-swap ⇒ seamless (no daemon.rs:368 broker restart). Option B (a dedicated Split-reader drain of the brain conn, brain.rs:230 — the pump's carrier) is the FORWARD path for when genuinely daemon-DRIVEN sessions land (the live-agent adapter) — deferred, noted, not built. Option C (the broker-side durable CLASS fix) = REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK, deferred. Int = the confirmer rig (brain+broker+PTY, brain_decouple template): N sessions actively producing output, brain resume-subscribed onto its req/reply conn; RED-first = heartbeat stalls + an active-streaming controller stall-evicted >15s WITHOUT any dead peer present (proves dead-peer-independence); assert BOTH severities (respawn interleave + steady-state output backup); Option A turns both green. Composes with REQ-BRAIN-RESUME-NO-CONTROL-STEAL (the round-2 Viewer fix this supersedes as the wedge cure), REQ-UPDATE-TRIAL-DRAIN-DRIVE (the counter-54 reap-drive that widened the window), REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE (the stall-evict it stops triggering).\"\nrequired_stages = [\"doc\", \"impl\", \"int\"]  # ACTIVATED v0.30.5 UPDATE-WEDGE-3 fix (doyle ruled Option A 2026-07-09; todlando builds). doc = brain.rs resume_sessions docstring corrected (the brain no longer subscribes — D4-2/REQ-DAEMON-2 continuity intent) + a CONTEXT/LIFECYCLE note. impl = drop the subscribe_with call in resume_sessions (Option A). unit DROPPED (same FORK-2 reasoning as REQ-UPDATE-TRIAL-DRAIN-DRIVE: the impl is a deletion, no pure branch to unit-test; the int rig load-bearing-guards it RED-first). int = the confirmer rig, both severities RED-first, dead-peer-INDEPENDENT.\n\n[Showing lines 2376-2403 of 7747. Use :2404 to continue]"}],"details":{"truncation":{"content":"id = \"REQ-UPDATE-PROMOTE-DRAINED\"\r\ntitle = \"W3 (LIFECYCLE-TRUTH, mechanic-d MOVED FROM W2 per doyle gate verdict @e5ae7a9 — binding): the update-apply brain-generation promotion completes only when the OLD generation's broker subscriber connection is CLOSED or stall-EVICTED — never while blocked writes still pend on it. ROOT: `brain.ready` != subscribers drained; W2's stall-evict (REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE) only BOUNDS the false-promote window to BRAIN_WRITE_DEADLINE (15s), it does NOT close it — a new brain can signal ready inside that window while the old gen's conn is still wedged, so the apply 'promotes' onto a still-frozen control plane (the 22:47 incident-night false-promote). FIX: the promotion gate (ADR-0018 brain-trial, brainproc.rs) adds an explicit DRAINED precondition — promote only on ready AND old-gen-subscriber-drained (conn closed OR stall-evicted); the drained signal reads broker truth (the W2 stall-evict tally / the old conn's liveness), no brain round-trip. The residual W2 left open, now closed. Int = a FALSE-PROMOTE rig that exercises the promotion path itself: an old-gen subscriber conn held wedged past ready must NOT promote until it drains (RED-first: ready-alone promotes).\"\r\nrequired_stages = [\"impl\", \"unit\", \"int\"]  # ACTIVATED LIFECYCLE-TRUTH W3 drained-condition slice (todlando 2026-07-07, work-start); brainproc.rs run_trial DRAINED gate + broker any_local_controller_wedged truth + false-promote int rig — closes the W2 15s residual (KH 7.36)\r\n\r\n[[requirements]]\r\nid = \"REQ-UPDATE-TRIAL-DRAIN-DRIVE\"\r\ntitle = \"UPDATE-WEDGE (counter-54, doyle-ruled 2026-07-09 — regression of the v0.29.0 seamless brain-swap): a brain generation DRIVES the broker's controller-liveness reap (a KIND_SESSIONS poll) each heartbeat throughout its boot/trial loop, so a hard-KILLED prior generation's black-holed LOCAL controller conn (by:None) is stall-evicted within the trial window and can never permanently strand the promotion DRAINED gate. ROOT (2026-07-09 field freeze, `spt update fetch --apply` v0.30.0->v0.30.2 froze all 7 live PTYs ~30s then rolled back): the promote gate (run_trial, brainproc.rs:657-661) needs BOTH `ready_generation==gen` AND `old_gen_drained()`; `old_gen_drained()` = `!any_local_controller_wedged()` (brainproc.rs:534) is a PURE READ of `write_blocked_since` (broker.rs:2703) — it never DRIVES the evict. The evict (`stall_evict_controller`, broker.rs:1039, same 15s `brain_write_deadline` the wedge-read uses) only runs via `reap_dead_controller` (broker.rs:967, severed->drop ELSE stall-evict) inside the KIND_SESSIONS snapshot closure (broker.rs:2879). During the isolated brain-trial window NOTHING polls KIND_SESSIONS: the old brain was hard-killed (`child.kill()`, brainproc.rs:851) so its local controller conn black-holes (live PTYs keep writing to a dead Windows named pipe -> writer BLOCKS, never EOFs) -> `write_blocked_since=Some` -> wedged=true for the full 30s -> gate false -> `WindowElapsedAlive` -> kill+rollback. The candidate DID reach ready (write_ready, brainproc.rs:211, runs before the loop; the v53 log's NET_FAMILY_GATE/PAIR_MEET_UP prove the loop was entered) — so `BRAIN_TRIAL_TIMEOUT: candidate alive but never ready` is the MISLEADING ready-stamped-but-never-DRAINED case, NOT a resume_sessions hang. SECOND LEG (recovery): post-rollback the gen-2 brain's KIND_SESSIONS poll finally reaps -> `BRAIN_SUBSCRIBER_STALL_EVICT:1` fires >15s late -> a session stayed black-holed through recovery -> continued freeze. ONE root, BOTH legs. FIX (BRAIN-SIDE, self-applying — doyle ruled brain-side to AVOID a broker-side coordinated-restart flag): the boot/trial heartbeat loop (brainproc.rs run_brain, currently only `net_status` at :244) also issues `Brain::sessions()` (KIND_SESSIONS, brain.rs:1397 — already exists) every heartbeat, driving the LIVE older broker's ALREADY-SHIPPED reap (>=v0.29.0 LIFECYCLE-TRUTH; the field-stuck broker is v0.30.0 so it HAS it). The old-gen wedged conn is stall-evicted ~15s < the 30s window -> `old_gen_drained()` flips true -> PROMOTE; the same poll on the rollback/recovery brain reaps promptly -> no >15s black-hole -> kills the STALL_EVICT recovery leg. Drive it on the FIRST heartbeat (no one-tick wait) and on BOTH the trial candidate AND the recovery brain (one loop covers both). SELF-APPLIES because it drives the current broker's existing reap verb — v54's brain fixes the v0.30.0->v54 update with NO coordinated broker restart. Forward-compat: brokers <v0.29.0 have neither the DRAINED gate nor the reap, so older-broker updates never hit this path — no regression. Composes with REQ-UPDATE-PROMOTE-DRAINED (the gate this un-strands) + REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE (the stall-evict it drives) + REQ-CONTROLLER-LIVENESS-REAP (the reap verb). Int = the brain-swap-under-live-sessions rig: (1) with a hard-killed prior gen holding a wedged by:None controller past 15s, the new brain PROMOTES within the window AND sessions stay served across the cycle (RED-first: without the drive, WindowElapsedAlive->rollback); (2) a deliberately-failing trial auto-rolls-back AND the restored brain re-drives EVERY session with NO BRAIN_SUBSCRIBER_STALL_EVICT.\"\r\nrequired_stages = [\"doc\", \"impl\", \"int\"]  # ACTIVATED counter-54 UPDATE-WEDGE fix (doyle ruled brain-side 2026-07-09; todlando builds). doc = a CONTEXT.md/LIFECYCLE note that the trial/boot brain drives controller-reap so a killed prior generation cannot strand old_gen_drained. impl = the brain.sessions() KIND_SESSIONS drive in run_brain's heartbeat loop (brainproc.rs, first-tick + every tick, trial + recovery brain). UNIT DELIBERATELY DROPPED (doyle FORK-2 ruling 2026-07-09): the impl is ONE unconditional branchless brain.sessions() call in the heartbeat loop — no pure branch/decision to unit-test; forcing a call-ordering unit would require a trait-seam refactor of the broker-critical brain heartbeat path purely to satisfy the stage (test theater, adds risk for no coverage). The behavior is load-bearing-guarded by the int rig (RED-first bites the PRODUCTION drive), and the broker-side reap the drive triggers is already unit-covered (broker.rs reap_evicts_a_blocked_writer_not_only_an_exited_one). int = the brain-swap-under-live-sessions rig (FORK-1 ruling A: a NEW crates/spt/tests real-SUBPROCESS sibling of brain_survive.rs — real broker + real PTY + real supervise_brain->run_trial + a real `spt daemon brain` child running production run_brain self-driving sessions(); RED-first bites the production drive), both assertions: (1) promote-within-window with a wedged old-gen local controller + sessions stay served; (2) failed trial rolls back with sessions re-driven, NO BRAIN_SUBSCRIBER_STALL_EVICT.\r\n\r\n[[requirements]]\r\nid = \"REQ-BRAIN-UPDATE-RESTART-CLEAN-CLOSE\"\r\ntitle = \"SEED (DEFERRED, doyle 2026-07-09 — post-counter-54 root-hardening for UPDATE-WEDGE; mint now, impl a FUTURE milestone): on a PLANNED brain-restart (`BRAIN_UPDATE_RESTART`, the seamless update-apply brain-cycle), the outgoing brain's LOCAL (by:None) controller conns are GRACEFULLY CLEAN-CLOSED as the brain is cycled, instead of hard-killed and left to black-hole. ROOT (field-pinned 2026-07-09, daemon.stderr.log L24277-24303): the update-restart path hard-kills the outgoing brain (`child.kill()`, brainproc.rs:851); its live-agent controller conns then block on dead pipes (never EOF) → the broker reads them WEDGED (broker.rs:2695-2700) → the new candidate's promotion DRAINED gate (`any_local_controller_wedged`, broker.rs:2704) stays true until the W2 stall-evict matures (~15s). REQ-UPDATE-TRIAL-DRAIN-DRIVE (counter-54) makes the candidate DRIVE that reap so it promotes within the 30s window — but at a ~15s wedge-maturity hitch (frozen PTYs during the swap). A CLEAN close makes the conn 'simply absent → drained=false AT ONCE → fast promote' (broker.rs:2699-2700), ELIMINATING the hitch = truly seamless (honors the paradigm the field freeze broke). SUPERSEDES the earlier livehost-reattach framing of 'Fix Y': livehost is SPAWN-FRESH (fresh session uuid/pid per boot, nothing to re-attach — wrong site, and it never ran in the trial window); the correct site is the brain-cycle / update-restart path (a bounded graceful-drain of the outgoing brain BEFORE the kill). Non-trivial: hard-kill → bounded graceful drain; a drain that hangs must NOT wedge the swap (timeout then kill anyway, never block the update). Composes with REQ-UPDATE-TRIAL-DRAIN-DRIVE (defense-in-depth reap-drive REMAINS for any conn that still black-holes — a peer/relay conn, a drain-timeout kill) + REQ-UPDATE-PROMOTE-DRAINED (the gate) + REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE (the stall-evict). Int: a planned update-restart under a live-agent controller conn → the new candidate promotes WITHOUT waiting the ~15s wedge-maturity (drained reads false immediately, no STALL_EVICT), RED-first vs the current hard-kill-then-reap-drive ~15s hitch.\"\r\nrequired_stages = []  # SEED minted DEFERRED (doyle 2026-07-09); activate post-counter-54 at the milestone that delivers HITCH-FREE seamless updates. Supersedes the livehost-reattach 'Fix Y' framing (wrong site). The counter-54 fix (REQ-UPDATE-TRIAL-DRAIN-DRIVE reap-drive) makes updates SELF-HEAL in ~15s; this seed removes the residual hitch.\r\n\r\n[[requirements]]\r\nid = \"REQ-LIVEHOST-RECONCILE-TRIAL-SILENT\"\r\ntitle = \"SEED (DEFERRED investigation, doyle 2026-07-09 — UPDATE-WEDGE follow-up): determine WHY the trial/rollback brain's livehost reconcile loop did NOT drive the broker controller-reap (nor re-host the live agents) during the ~30s field update-trial window, when livehost polls `query_live_session_endpoints()` → `brain.sessions()` (KIND_SESSIONS) UNCONDITIONALLY every `LIVE_RECONCILE_INTERVAL_MS`=5000ms (livehost.rs:1026). CONTEXT (surfaced building the counter-54 rig): livehost's 5s KIND_SESSIONS poll drives the SAME broker `reap_dead_controller` sweep the fix drives — so it would otherwise reap the 15s-matured wedge by ~T20 < the 30s trial and SELF-HEAL. It didn't (field froze 30s → rollback), so the field trial-brain livehost was silent/delayed (PIN Q2: no `DAEMON_RESTART_RESUME` under gen-1/gen-2; the 30s kill landed before/around livehost's first reconcile tick). The counter-54 fix (REQ-UPDATE-TRIAL-DRAIN-DRIVE) puts a RELIABLE 500ms reap-driver in run_brain's CORE heartbeat loop, making the wedge-reap INDEPENDENT of livehost — so this does NOT block counter-54. But the livehost silence is a latent anomaly with a SECOND consequence: live-agent HARNESS re-hosting was also delayed ~30s (a separate freeze contributor). Investigate: does `spawn_live_host`'s reconcile thread start promptly on a trial-brain boot, or is its first tick delayed past the trial window? Does its brain conn / `query_live_session_endpoints` block against the swap/wedge state? Register concrete REQ(s) once the mechanism is pinned. RELATED: [[REQ-BRAIN-UPDATE-RESTART-CLEAN-CLOSE]] (if the outgoing brain's black-holed conns perturb the new brain's livehost conn setup).\"\r\nrequired_stages = []  # SEED deferred investigation (doyle 2026-07-09); non-blocking for counter-54 (the core-loop reap-drive in REQ-UPDATE-TRIAL-DRAIN-DRIVE covers the wedge independently of livehost). Activate once the livehost-silence mechanism is pinned.\r\n\r\n[[requirements]]\r\nid = \"REQ-BRAIN-RESUME-NO-CONTROL-STEAL\"\r\ntitle = \"UPDATE-WEDGE round 2 (v0.30.4, doyle-ruled 2026-07-09 — field incident on the counter-54 fetch--apply): a brain-respawn must NEVER steal, then stall-evict, the controller of a broker PTY session the daemon brain does not DRIVE. ROOT (field-pinned + SME, docs/UPDATE-WEDGE-2-RCA.md + docs/UPDATE-WEDGE-2-SME-todlando.md): `resume_sessions` (brain.rs:985) re-attaches EVERY session `KIND_SESSIONS` returns via `subscribe` = `subscribe_with(AttachIntent::Control, by:None)` (brain.rs:1450-1455). The docstring's 'a None identity never displaces (falls back to viewer)' is a MYTH for FREE / SAME-LOCAL-IDENTITY slots: `resolve_subscribe` (broker.rs:1134-1153) stall-evicts FIRST, then `become_controller` if the slot is now free OR the incumbent is also local (None==None) — viewer-fallback fires ONLY when a DIFFERENT REMOTE controls. So on a box with N spt-hosted broker PTYs, a brain-respawn STEALS the by:None controller of every free/local-controlled session (incl. the operator's LOCAL `spt rc`), which the daemon brain never drains (it hosts no PTY sessions — brainproc.rs:184) → 15s later `stall_evict_controller` (broker.rs:1039) releases driven_by → the session is UNCONTROLLABLE (Failure A). It also head-of-line-blocks the shared brain↔broker conn on the N-session controller-replay burst → every journaled `spt rc` retake deadlines ('brain IPC read deadline', the REQ-BROKER-ATTACH-JOURNAL-RESILIENT / #16 shared-conn symptom) → global rc failure on ALL N (Failure B). Field 2026-07-09 (operator-confirmed): fetch--apply 0.30.2→0.30.3 PROMOTED CLEANLY (the counter-54 fix worked) but under 7 spt-hosted PTYs (ALL with LOCAL by:None controllers) the resume SILENTLY STOLE all 7 (become_controller same-local re-take, NO Displaced notice → orphaned, output froze immediately, no rc-detach splash) + blocked every rc retake. The 5-vs-2 stall-evict split is ACTIVE-vs-IDLE, not remote-vs-local: 5 producing output → stolen writer blocked >15s → stall-evict; 2 idle → writer parked → no evict, but still silently stolen+frozen. The counter-54 promotion fix did NOT cause this — pre-existing resume-steal latent bug, hidden until N broker PTYs were present at a respawn; the single-black-holed-session A'-rig never exercised N-live-controllers-under-replay. FIX (brain-side): `resume_sessions` re-attaches as **Viewer** (`AttachIntent::Viewer`), NOT Control — a viewer never touches driven_by and is never stall-evicted (broker.rs:1063+ bounded try_send + private eviction), so steal-then-drop vanishes and the operator keeps/regains control; and it relieves shared-conn pressure (a slow viewer is DROPPED, never a 15s controller block) so rc retake gets through. Control ONLY for sessions the daemon brain genuinely DRIVES (empty set today → all become Viewer; forward-correct for the live-agent-adapter future). SECONDARY (escalation, ONLY if the gate shows residual B): stagger the resume re-attach + bound the viewer replay so the respawn burst can't saturate the conn. Int = the multi-broker-PTY-session RESPAWN rig (the coverage the A'-rig lacked): N real broker-spawned sessions with controllers producing output → real brain respawn/promote → assert (1) EVERY session keeps its controller across the swap (no stall-evict of a session the brain doesn't drive), (2) `spt rc` attaches/retakes IMMEDIATELY post-promote (no shared-conn saturation), (3) promotion still succeeds. RED-first: the current Control re-attach steals+evicts + deadlines rc. Composes with REQ-UPDATE-TRIAL-DRAIN-DRIVE (the orthogonal counter-54 promote fix), REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE (the stall-evict it stops mis-firing on non-driven sessions), REQ-BROKER-ATTACH-JOURNAL-RESILIENT (the #16 shared-conn resilience). Distinct from REQ-BRAIN-UPDATE-RESTART-CLEAN-CLOSE (that = the OLD brain's outgoing black-hole; this = the NEW brain's resume-steal).\"\r\nrequired_stages = [\"doc\", \"impl\", \"int\"]  # ACTIVATED v0.30.4 UPDATE-WEDGE-2 fix (doyle ruled 2026-07-09; todlando builds). doc = brain.rs resume docstring corrected (kill the by:None-never-steals MYTH) + a CONTEXT/LIFECYCLE note that a brain-respawn re-attaches non-driven sessions as viewer. impl = resume path re-attaches AttachIntent::Viewer (Control only for genuinely-driven sessions, empty today). unit DROPPED if the impl is a branchless intent-flip (same FORK-2 reasoning as REQ-UPDATE-TRIAL-DRAIN-DRIVE — add a unit only if a real driven-vs-not predicate branch exists). int = the multi-broker-PTY-session respawn rig, all 3 assertions RED-first.\r\n\r\n[[requirements]]\r\nid = \"REQ-BRAIN-RESUME-NO-CONN-DEADLOCK\"\r\ntitle = \"UPDATE-WEDGE round 3 (v0.30.5, doyle-ruled Option A 2026-07-09 — the v0.30.4 field-verify re-wedge, root code-PROVEN + dead-peer-INDEPENDENT): the daemon brain must NOT subscribe broker PTY sessions onto its own request/reply IPC conn — it has no consumer for that output and the subscription DEADLOCKS the conn. ROOT (todlando code-read, docs/UPDATE-WEDGE-2-ROUND3-CODEREAD.md; the net-runtime AND the counter-54 reap-drive were both FALSIFIED first — docs/UPDATE-WEDGE-2-ROUND3-RIG-VERDICT.md): a conn's send half is a single `SharedSend = Arc<Mutex<SendHalf>>` (broker.rs:78). Subscriber writer threads (`viewer_writer` broker.rs:1333/1342, `controller_writer` :1451) hold `send.lock()` ACROSS a BLOCKING `write_frame`; the dispatch reply path (`send_frame` :4221 → KIND_SESSIONS_REPLY / KIND_NET_STATUS_REPLY) needs the SAME lock. `resume_sessions` (brain.rs:1031→1054) subscribes every session as a Viewer onto the brain's MAIN conn — which is ALSO the brain's request/reply channel. The daemon brain hosts no PTY sessions (brainproc.rs:184) so run_brain never drains that output; it reads the conn only during the 500ms-heartbeat net_status()/sessions() calls (drain-and-DISCARD, `_ => continue`). When an actively-streaming session backs the conn up, a subscriber writer BLOCKS in write_frame holding send.lock() → the dispatch thread can't send the heartbeat reply → net_status()/sessions() never return → the heartbeat loop stalls → the brain never drains → the writer stays blocked = SELF-DEADLOCK on the brain conn's send mutex → every subscriber writer on it wedges → BRAIN_SUBSCRIBER_STALL_EVICT (controller writer blocked >15s). BOTH severities, one mechanism: RESPAWN (resume subscribes N + the replay burst floods the conn before the loop drains) AND STEADY-STATE (an active streamer's output between heartbeats fills the socket buffer). Counter-54 (REQ-UPDATE-TRIAL-DRAIN-DRIVE) added a 2nd per-heartbeat reply round-trip (sessions()) through the contended mutex — WIDENED the window (regression-window-exact), did not create it. The round-2 Viewer fix (REQ-BRAIN-RESUME-NO-CONTROL-STEAL) removed the STEAL but kept the brain a SUBSCRIBER — viewer_writer has the same send.lock()-across-write pattern AND viewers have no stall-evict valve — so v0.30.4 field-verify wedged again. FIX (Option A, brain-side, SEAMLESS): resume_sessions does NOT subscribe (drop the subscribe_with call; §3 verification guard confirmed NO brain consumer — digest/relay/net-consumer/shellwake/presence — reads the resumed subs). The brain conn then carries only request/reply → no subscriber backpressure → no deadlock. Keep the session_cursors seed only if harmless. Rides the brain-swap ⇒ seamless (no daemon.rs:368 broker restart). Option B (a dedicated Split-reader drain of the brain conn, brain.rs:230 — the pump's carrier) is the FORWARD path for when genuinely daemon-DRIVEN sessions land (the live-agent adapter) — deferred, noted, not built. Option C (the broker-side durable CLASS fix) = REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK, deferred. Int = the confirmer rig (brain+broker+PTY, brain_decouple template): N sessions actively producing output, brain resume-subscribed onto its req/reply conn; RED-first = heartbeat stalls + an active-streaming controller stall-evicted >15s WITHOUT any dead peer present (proves dead-peer-independence); assert BOTH severities (respawn interleave + steady-state output backup); Option A turns both green. Composes with REQ-BRAIN-RESUME-NO-CONTROL-STEAL (the round-2 Viewer fix this supersedes as the wedge cure), REQ-UPDATE-TRIAL-DRAIN-DRIVE (the counter-54 reap-drive that widened the window), REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE (the stall-evict it stops triggering).\"\r\nrequired_stages = [\"doc\", \"impl\", \"int\"]  # ACTIVATED v0.30.5 UPDATE-WEDGE-3 fix (doyle ruled Option A 2026-07-09; todlando builds). doc = brain.rs resume_sessions docstring corrected (the brain no longer subscribes — D4-2/REQ-DAEMON-2 continuity intent) + a CONTEXT/LIFECYCLE note. impl = drop the subscribe_with call in resume_sessions (Option A). unit DROPPED (same FORK-2 reasoning as REQ-UPDATE-TRIAL-DRAIN-DRIVE: the impl is a deletion, no pure branch to unit-test; the int rig load-bearing-guards it RED-first). int = the confirmer rig, both severities RED-first, dead-peer-INDEPENDENT.\r","truncated":true,"truncatedBy":"lines","totalLines":5372,"totalBytes":20391,"outputLines":28,"outputBytes":20391,"lastLinePartial":false,"firstLineExceedsLimit":false},"totalLines":7747,"displayContent":{"text":"id = \"REQ-UPDATE-PROMOTE-DRAINED\"\r\ntitle = \"W3 (LIFECYCLE-TRUTH, mechanic-d MOVED FROM W2 per doyle gate verdict @e5ae7a9 — binding): the update-apply brain-generation promotion completes only when the OLD generation's broker subscriber connection is CLOSED or stall-EVICTED — never while blocked writes still pend on it. ROOT: `brain.ready` != subscribers drained; W2's stall-evict (REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE) only BOUNDS the false-promote window to BRAIN_WRITE_DEADLINE (15s), it does NOT close it — a new brain can signal ready inside that window while the old gen's conn is still wedged, so the apply 'promotes' onto a still-frozen control plane (the 22:47 incident-night false-promote). FIX: the promotion gate (ADR-0018 brain-trial, brainproc.rs) adds an explicit DRAINED precondition — promote only on ready AND old-gen-subscriber-drained (conn closed OR stall-evicted); the drained signal reads broker truth (the W2 stall-evict tally / the old conn's liveness), no brain round-trip. The residual W2 left open, now closed. Int = a FALSE-PROMOTE rig that exercises the promotion path itself: an old-gen subscriber conn held wedged past ready must NOT promote until it drains (RED-first: ready-alone promotes).\"\r\nrequired_stages = [\"impl\", \"unit\", \"int\"]  # ACTIVATED LIFECYCLE-TRUTH W3 drained-condition slice (todlando 2026-07-07, work-start); brainproc.rs run_trial DRAINED gate + broker any_local_controller_wedged truth + false-promote int rig — closes the W2 15s residual (KH 7.36)\r\n\r\n[[requirements]]\r\nid = \"REQ-UPDATE-TRIAL-DRAIN-DRIVE\"\r\ntitle = \"UPDATE-WEDGE (counter-54, doyle-ruled 2026-07-09 — regression of the v0.29.0 seamless brain-swap): a brain generation DRIVES the broker's controller-liveness reap (a KIND_SESSIONS poll) each heartbeat throughout its boot/trial loop, so a hard-KILLED prior generation's black-holed LOCAL controller conn (by:None) is stall-evicted within the trial window and can never permanently strand the promotion DRAINED gate. ROOT (2026-07-09 field freeze, `spt update fetch --apply` v0.30.0->v0.30.2 froze all 7 live PTYs ~30s then rolled back): the promote gate (run_trial, brainproc.rs:657-661) needs BOTH `ready_generation==gen` AND `old_gen_drained()`; `old_gen_drained()` = `!any_local_controller_wedged()` (brainproc.rs:534) is a PURE READ of `write_blocked_since` (broker.rs:2703) — it never DRIVES the evict. The evict (`stall_evict_controller`, broker.rs:1039, same 15s `brain_write_deadline` the wedge-read uses) only runs via `reap_dead_controller` (broker.rs:967, severed->drop ELSE stall-evict) inside the KIND_SESSIONS snapshot closure (broker.rs:2879). During the isolated brain-trial window NOTHING polls KIND_SESSIONS: the old brain was hard-killed (`child.kill()`, brainproc.rs:851) so its local controller conn black-holes (live PTYs keep writing to a dead Windows named pipe -> writer BLOCKS, never EOFs) -> `write_blocked_since=Some` -> wedged=true for the full 30s -> gate false -> `WindowElapsedAlive` -> kill+rollback. The candidate DID reach ready (write_ready, brainproc.rs:211, runs before the loop; the v53 log's NET_FAMILY_GATE/PAIR_MEET_UP prove the loop was entered) — so `BRAIN_TRIAL_TIMEOUT: candidate alive but never ready` is the MISLEADING ready-stamped-but-never-DRAINED case, NOT a resume_sessions hang. SECOND LEG (recovery): post-rollback the gen-2 brain's KIND_SESSIONS poll finally reaps -> `BRAIN_SUBSCRIBER_STALL_EVICT:1` fires >15s late -> a session stayed black-holed through recovery -> continued freeze. ONE root, BOTH legs. FIX (BRAIN-SIDE, self-applying — doyle ruled brain-side to AVOID a broker-side coordinated-restart flag): the boot/trial heartbeat loop (brainproc.rs run_brain, currently only `net_status` at :244) also issues `Brain::sessions()` (KIND_SESSIONS, brain.rs:1397 — already exists) every heartbeat, driving the LIVE older broker's ALREADY-SHIPPED reap (>=v0.29.0 LIFECYCLE-TRUTH; the field-stuck broker is v0.30.0 so it HAS it). The old-gen wedged conn is stall-evicted ~15s < the 30s window -> `old_gen_drained()` flips true -> PROMOTE; the same poll on the rollback/recovery brain reaps promptly -> no >15s black-hole -> kills the STALL_EVICT recovery leg. Drive it on the FIRST heartbeat (no one-tick wait) and on BOTH the trial candidate AND the recovery brain (one loop covers both). SELF-APPLIES because it drives the current broker's existing reap verb — v54's brain fixes the v0.30.0->v54 update with NO coordinated broker restart. Forward-compat: brokers <v0.29.0 have neither the DRAINED gate nor the reap, so older-broker updates never hit this path — no regression. Composes with REQ-UPDATE-PROMOTE-DRAINED (the gate this un-strands) + REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE (the stall-evict it drives) + REQ-CONTROLLER-LIVENESS-REAP (the reap verb). Int = the brain-swap-under-live-sessions rig: (1) with a hard-killed prior gen holding a wedged by:None controller past 15s, the new brain PROMOTES within the window AND sessions stay served across the cycle (RED-first: without the drive, WindowElapsedAlive->rollback); (2) a deliberately-failing trial auto-rolls-back AND the restored brain re-drives EVERY session with NO BRAIN_SUBSCRIBER_STALL_EVICT.\"\r\nrequired_stages = [\"doc\", \"impl\", \"int\"]  # ACTIVATED counter-54 UPDATE-WEDGE fix (doyle ruled brain-side 2026-07-09; todlando builds). doc = a CONTEXT.md/LIFECYCLE note that the trial/boot brain drives controller-reap so a killed prior generation cannot strand old_gen_drained. impl = the brain.sessions() KIND_SESSIONS drive in run_brain's heartbeat loop (brainproc.rs, first-tick + every tick, trial + recovery brain). UNIT DELIBERATELY DROPPED (doyle FORK-2 ruling 2026-07-09): the impl is ONE unconditional branchless brain.sessions() call in the heartbeat loop — no pure branch/decision to unit-test; forcing a call-ordering unit would require a trait-seam refactor of the broker-critical brain heartbeat path purely to satisfy the stage (test theater, adds risk for no coverage). The behavior is load-bearing-guarded by the int rig (RED-first bites the PRODUCTION drive), and the broker-side reap the drive triggers is already unit-covered (broker.rs reap_evicts_a_blocked_writer_not_only_an_exited_one). int = the brain-swap-under-live-sessions rig (FORK-1 ruling A: a NEW crates/spt/tests real-SUBPROCESS sibling of brain_survive.rs — real broker + real PTY + real supervise_brain->run_trial + a real `spt daemon brain` child running production run_brain self-driving sessions(); RED-first bites the production drive), both assertions: (1) promote-within-window with a wedged old-gen local controller + sessions stay served; (2) failed trial rolls back with sessions re-driven, NO BRAIN_SUBSCRIBER_STALL_EVICT.\r\n\r\n[[requirements]]\r\nid = \"REQ-BRAIN-UPDATE-RESTART-CLEAN-CLOSE\"\r\ntitle = \"SEED (DEFERRED, doyle 2026-07-09 — post-counter-54 root-hardening for UPDATE-WEDGE; mint now, impl a FUTURE milestone): on a PLANNED brain-restart (`BRAIN_UPDATE_RESTART`, the seamless update-apply brain-cycle), the outgoing brain's LOCAL (by:None) controller conns are GRACEFULLY CLEAN-CLOSED as the brain is cycled, instead of hard-killed and left to black-hole. ROOT (field-pinned 2026-07-09, daemon.stderr.log L24277-24303): the update-restart path hard-kills the outgoing brain (`child.kill()`, brainproc.rs:851); its live-agent controller conns then block on dead pipes (never EOF) → the broker reads them WEDGED (broker.rs:2695-2700) → the new candidate's promotion DRAINED gate (`any_local_controller_wedged`, broker.rs:2704) stays true until the W2 stall-evict matures (~15s). REQ-UPDATE-TRIAL-DRAIN-DRIVE (counter-54) makes the candidate DRIVE that reap so it promotes within the 30s window — but at a ~15s wedge-maturity hitch (frozen PTYs during the swap). A CLEAN close makes the conn 'simply absent → drained=false AT ONCE → fast promote' (broker.rs:2699-2700), ELIMINATING the hitch = truly seamless (honors the paradigm the field freeze broke). SUPERSEDES the earlier livehost-reattach framing of 'Fix Y': livehost is SPAWN-FRESH (fresh session uuid/pid per boot, nothing to re-attach — wrong site, and it never ran in the trial window); the correct site is the brain-cycle / update-restart path (a bounded graceful-drain of the outgoing brain BEFORE the kill). Non-trivial: hard-kill → bounded graceful drain; a drain that hangs must NOT wedge the swap (timeout then kill anyway, never block the update). Composes with REQ-UPDATE-TRIAL-DRAIN-DRIVE (defense-in-depth reap-drive REMAINS for any conn that still black-holes — a peer/relay conn, a drain-timeout kill) + REQ-UPDATE-PROMOTE-DRAINED (the gate) + REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE (the stall-evict). Int: a planned update-restart under a live-agent controller conn → the new candidate promotes WITHOUT waiting the ~15s wedge-maturity (drained reads false immediately, no STALL_EVICT), RED-first vs the current hard-kill-then-reap-drive ~15s hitch.\"\r\nrequired_stages = []  # SEED minted DEFERRED (doyle 2026-07-09); activate post-counter-54 at the milestone that delivers HITCH-FREE seamless updates. Supersedes the livehost-reattach 'Fix Y' framing (wrong site). The counter-54 fix (REQ-UPDATE-TRIAL-DRAIN-DRIVE reap-drive) makes updates SELF-HEAL in ~15s; this seed removes the residual hitch.\r\n\r\n[[requirements]]\r\nid = \"REQ-LIVEHOST-RECONCILE-TRIAL-SILENT\"\r\ntitle = \"SEED (DEFERRED investigation, doyle 2026-07-09 — UPDATE-WEDGE follow-up): determine WHY the trial/rollback brain's livehost reconcile loop did NOT drive the broker controller-reap (nor re-host the live agents) during the ~30s field update-trial window, when livehost polls `query_live_session_endpoints()` → `brain.sessions()` (KIND_SESSIONS) UNCONDITIONALLY every `LIVE_RECONCILE_INTERVAL_MS`=5000ms (livehost.rs:1026). CONTEXT (surfaced building the counter-54 rig): livehost's 5s KIND_SESSIONS poll drives the SAME broker `reap_dead_controller` sweep the fix drives — so it would otherwise reap the 15s-matured wedge by ~T20 < the 30s trial and SELF-HEAL. It didn't (field froze 30s → rollback), so the field trial-brain livehost was silent/delayed (PIN Q2: no `DAEMON_RESTART_RESUME` under gen-1/gen-2; the 30s kill landed before/around livehost's first reconcile tick). The counter-54 fix (REQ-UPDATE-TRIAL-DRAIN-DRIVE) puts a RELIABLE 500ms reap-driver in run_brain's CORE heartbeat loop, making the wedge-reap INDEPENDENT of livehost — so this does NOT block counter-54. But the livehost silence is a latent anomaly with a SECOND consequence: live-agent HARNESS re-hosting was also delayed ~30s (a separate freeze contributor). Investigate: does `spawn_live_host`'s reconcile thread start promptly on a trial-brain boot, or is its first tick delayed past the trial window? Does its brain conn / `query_live_session_endpoints` block against the swap/wedge state? Register concrete REQ(s) once the mechanism is pinned. RELATED: [[REQ-BRAIN-UPDATE-RESTART-CLEAN-CLOSE]] (if the outgoing brain's black-holed conns perturb the new brain's livehost conn setup).\"\r\nrequired_stages = []  # SEED deferred investigation (doyle 2026-07-09); non-blocking for counter-54 (the core-loop reap-drive in REQ-UPDATE-TRIAL-DRAIN-DRIVE covers the wedge independently of livehost). Activate once the livehost-silence mechanism is pinned.\r\n\r\n[[requirements]]\r\nid = \"REQ-BRAIN-RESUME-NO-CONTROL-STEAL\"\r\ntitle = \"UPDATE-WEDGE round 2 (v0.30.4, doyle-ruled 2026-07-09 — field incident on the counter-54 fetch--apply): a brain-respawn must NEVER steal, then stall-evict, the controller of a broker PTY session the daemon brain does not DRIVE. ROOT (field-pinned + SME, docs/UPDATE-WEDGE-2-RCA.md + docs/UPDATE-WEDGE-2-SME-todlando.md): `resume_sessions` (brain.rs:985) re-attaches EVERY session `KIND_SESSIONS` returns via `subscribe` = `subscribe_with(AttachIntent::Control, by:None)` (brain.rs:1450-1455). The docstring's 'a None identity never displaces (falls back to viewer)' is a MYTH for FREE / SAME-LOCAL-IDENTITY slots: `resolve_subscribe` (broker.rs:1134-1153) stall-evicts FIRST, then `become_controller` if the slot is now free OR the incumbent is also local (None==None) — viewer-fallback fires ONLY when a DIFFERENT REMOTE controls. So on a box with N spt-hosted broker PTYs, a brain-respawn STEALS the by:None controller of every free/local-controlled session (incl. the operator's LOCAL `spt rc`), which the daemon brain never drains (it hosts no PTY sessions — brainproc.rs:184) → 15s later `stall_evict_controller` (broker.rs:1039) releases driven_by → the session is UNCONTROLLABLE (Failure A). It also head-of-line-blocks the shared brain↔broker conn on the N-session controller-replay burst → every journaled `spt rc` retake deadlines ('brain IPC read deadline', the REQ-BROKER-ATTACH-JOURNAL-RESILIENT / #16 shared-conn symptom) → global rc failure on ALL N (Failure B). Field 2026-07-09 (operator-confirmed): fetch--apply 0.30.2→0.30.3 PROMOTED CLEANLY (the counter-54 fix worked) but under 7 spt-hosted PTYs (ALL with LOCAL by:None controllers) the resume SILENTLY STOLE all 7 (become_controller same-local re-take, NO Displaced notice → orphaned, output froze immediately, no rc-detach splash) + blocked every rc retake. The 5-vs-2 stall-evict split is ACTIVE-vs-IDLE, not remote-vs-local: 5 producing output → stolen writer blocked >15s → stall-evict; 2 idle → writer parked → no evict, but still silently stolen+frozen. The counter-54 promotion fix did NOT cause this — pre-existing resume-steal latent bug, hidden until N broker PTYs were present at a respawn; the single-black-holed-session A'-rig never exercised N-live-controllers-under-replay. FIX (brain-side): `resume_sessions` re-attaches as **Viewer** (`AttachIntent::Viewer`), NOT Control — a viewer never touches driven_by and is never stall-evicted (broker.rs:1063+ bounded try_send + private eviction), so steal-then-drop vanishes and the operator keeps/regains control; and it relieves shared-conn pressure (a slow viewer is DROPPED, never a 15s controller block) so rc retake gets through. Control ONLY for sessions the daemon brain genuinely DRIVES (empty set today → all become Viewer; forward-correct for the live-agent-adapter future). SECONDARY (escalation, ONLY if the gate shows residual B): stagger the resume re-attach + bound the viewer replay so the respawn burst can't saturate the conn. Int = the multi-broker-PTY-session RESPAWN rig (the coverage the A'-rig lacked): N real broker-spawned sessions with controllers producing output → real brain respawn/promote → assert (1) EVERY session keeps its controller across the swap (no stall-evict of a session the brain doesn't drive), (2) `spt rc` attaches/retakes IMMEDIATELY post-promote (no shared-conn saturation), (3) promotion still succeeds. RED-first: the current Control re-attach steals+evicts + deadlines rc. Composes with REQ-UPDATE-TRIAL-DRAIN-DRIVE (the orthogonal counter-54 promote fix), REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE (the stall-evict it stops mis-firing on non-driven sessions), REQ-BROKER-ATTACH-JOURNAL-RESILIENT (the #16 shared-conn resilience). Distinct from REQ-BRAIN-UPDATE-RESTART-CLEAN-CLOSE (that = the OLD brain's outgoing black-hole; this = the NEW brain's resume-steal).\"\r\nrequired_stages = [\"doc\", \"impl\", \"int\"]  # ACTIVATED v0.30.4 UPDATE-WEDGE-2 fix (doyle ruled 2026-07-09; todlando builds). doc = brain.rs resume docstring corrected (kill the by:None-never-steals MYTH) + a CONTEXT/LIFECYCLE note that a brain-respawn re-attaches non-driven sessions as viewer. impl = resume path re-attaches AttachIntent::Viewer (Control only for genuinely-driven sessions, empty today). unit DROPPED if the impl is a branchless intent-flip (same FORK-2 reasoning as REQ-UPDATE-TRIAL-DRAIN-DRIVE — add a unit only if a real driven-vs-not predicate branch exists). int = the multi-broker-PTY-session respawn rig, all 3 assertions RED-first.\r\n\r\n[[requirements]]\r\nid = \"REQ-BRAIN-RESUME-NO-CONN-DEADLOCK\"\r\ntitle = \"UPDATE-WEDGE round 3 (v0.30.5, doyle-ruled Option A 2026-07-09 — the v0.30.4 field-verify re-wedge, root code-PROVEN + dead-peer-INDEPENDENT): the daemon brain must NOT subscribe broker PTY sessions onto its own request/reply IPC conn — it has no consumer for that output and the subscription DEADLOCKS the conn. ROOT (todlando code-read, docs/UPDATE-WEDGE-2-ROUND3-CODEREAD.md; the net-runtime AND the counter-54 reap-drive were both FALSIFIED first — docs/UPDATE-WEDGE-2-ROUND3-RIG-VERDICT.md): a conn's send half is a single `SharedSend = Arc<Mutex<SendHalf>>` (broker.rs:78). Subscriber writer threads (`viewer_writer` broker.rs:1333/1342, `controller_writer` :1451) hold `send.lock()` ACROSS a BLOCKING `write_frame`; the dispatch reply path (`send_frame` :4221 → KIND_SESSIONS_REPLY / KIND_NET_STATUS_REPLY) needs the SAME lock. `resume_sessions` (brain.rs:1031→1054) subscribes every session as a Viewer onto the brain's MAIN conn — which is ALSO the brain's request/reply channel. The daemon brain hosts no PTY sessions (brainproc.rs:184) so run_brain never drains that output; it reads the conn only during the 500ms-heartbeat net_status()/sessions() calls (drain-and-DISCARD, `_ => continue`). When an actively-streaming session backs the conn up, a subscriber writer BLOCKS in write_frame holding send.lock() → the dispatch thread can't send the heartbeat reply → net_status()/sessions() never return → the heartbeat loop stalls → the brain never drains → the writer stays blocked = SELF-DEADLOCK on the brain conn's send mutex → every subscriber writer on it wedges → BRAIN_SUBSCRIBER_STALL_EVICT (controller writer blocked >15s). BOTH severities, one mechanism: RESPAWN (resume subscribes N + the replay burst floods the conn before the loop drains) AND STEADY-STATE (an active streamer's output between heartbeats fills the socket buffer). Counter-54 (REQ-UPDATE-TRIAL-DRAIN-DRIVE) added a 2nd per-heartbeat reply round-trip (sessions()) through the contended mutex — WIDENED the window (regression-window-exact), did not create it. The round-2 Viewer fix (REQ-BRAIN-RESUME-NO-CONTROL-STEAL) removed the STEAL but kept the brain a SUBSCRIBER — viewer_writer has the same send.lock()-across-write pattern AND viewers have no stall-evict valve — so v0.30.4 field-verify wedged again. FIX (Option A, brain-side, SEAMLESS): resume_sessions does NOT subscribe (drop the subscribe_with call; §3 verification guard confirmed NO brain consumer — digest/relay/net-consumer/shellwake/presence — reads the resumed subs). The brain conn then carries only request/reply → no subscriber backpressure → no deadlock. Keep the session_cursors seed only if harmless. Rides the brain-swap ⇒ seamless (no daemon.rs:368 broker restart). Option B (a dedicated Split-reader drain of the brain conn, brain.rs:230 — the pump's carrier) is the FORWARD path for when genuinely daemon-DRIVEN sessions land (the live-agent adapter) — deferred, noted, not built. Option C (the broker-side durable CLASS fix) = REQ-HAZARD-SHAREDSEND-NO-BLOCKING-WRITE-UNDER-LOCK, deferred. Int = the confirmer rig (brain+broker+PTY, brain_decouple template): N sessions actively producing output, brain resume-subscribed onto its req/reply conn; RED-first = heartbeat stalls + an active-streaming controller stall-evicted >15s WITHOUT any dead peer present (proves dead-peer-independence); assert BOTH severities (respawn interleave + steady-state output backup); Option A turns both green. Composes with REQ-BRAIN-RESUME-NO-CONTROL-STEAL (the round-2 Viewer fix this supersedes as the wedge cure), REQ-UPDATE-TRIAL-DRAIN-DRIVE (the counter-54 reap-drive that widened the window), REQ-HAZARD-BROKER-VIEWER-BRAIN-DECOUPLE (the stall-evict it stops triggering).\"\r\nrequired_stages = [\"doc\", \"impl\", \"int\"]  # ACTIVATED v0.30.5 UPDATE-WEDGE-3 fix (doyle ruled Option A 2026-07-09; todlando builds). doc = brain.rs resume_sessions docstring corrected (the brain no longer subscribes — D4-2/REQ-DAEMON-2 continuity intent) + a CONTEXT/LIFECYCLE note. impl = drop the subscribe_with call in resume_sessions (Option A). unit DROPPED (same FORK-2 reasoning as REQ-UPDATE-TRIAL-DRAIN-DRIVE: the impl is a deletion, no pure branch to unit-test; the int rig load-bearing-guards it RED-first). int = the confirmer rig, both severities RED-first, dead-peer-INDEPENDENT.\r","startLine":2376,"lineNumbers":[2376,2377,2378,2379,2380,2381,2382,2383,2384,2385,2386,2387,2388,2389,2390,2391,2392,2393,2394,2395,2396,2397,2398,2399,2400,2401,2402,2403]},"fileSize":2076989,"meta":{"source":{"type":"path","value":"C:\\Users\\decid\\Documents\\projects\\spt-core\\traceable-reqs.toml"},"truncation":{"direction":"head","truncatedBy":"lines","totalLines":7747,"totalBytes":20391,"outputLines":28,"outputBytes":20391,"shownRange":{"start":2376,"end":2403},"nextOffset":2404}}},"isError":false,"timestamp":1789524004180}}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_DhuDSwrHxoSiQdwfvEZlJjZ1|fc_0d104aba2b0f3edc016aa9f828c5bc87d19460469aaa30b5d2","toolName":"yield","startedAt":"2026-09-16T02:01:37.536Z"},"id":"d6033a3e","parentId":"e5cc2b58","timestamp":"2026-09-16T02:01:37.536Z"}
{"type":"message","id":"bf6624d5","parentId":"d6033a3e","timestamp":"2026-09-16T02:01:37.537Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":"**Finalizing evidence report**"},{"type":"toolCall","id":"call_DhuDSwrHxoSiQdwfvEZlJjZ1|fc_0d104aba2b0f3edc016aa9f828c5bc87d19460469aaa30b5d2","name":"yield","arguments":{"type":null,"data":{"summary":"W1c proves the original controller and three replacement controllers actually vacated session 1, all via cause=detach; it does not prove graceful pre-kill close. Replacement IPC connections were separately deadline-retired, so continued viewport failure is not simply one immortal old connection. This local rc is by=Some(nodehex), explicitly outside old_gen_drained; ordinary refresh also does not by itself enter the AppliedPending promotion trial. RCA remains candidate-only: repeated failure of the reconstructed stream/consumer path despite functioning per-connection containment.","files":[{"path":"crates/spt-daemon/src/broker.rs","description":"Controller ownership/attach generation and epoch: 829-889, 2592-2694; closure causes, reap, wedge predicate: 2720-2917; detach identity: 3547-3588; timeout tally bridge: 3813-3836; controller writer: 3955-4124; aggregate drain predicate: 5989-6019; EOF cleanup: 6424-6465."},{"path":"crates/spt-daemon/src/conn.rs","description":"Physical-connection deadline versus organic retirement tokens: 331-360; serialized gate plus watchdog/retirement implementation: 609-710. Deadline includes gate waiting, not only OS write."},{"path":"crates/spt-daemon/src/brainproc.rs","description":"Cursor-only resume 268-286; heartbeat reap drive 391-409; ProductionTrialEnv drain predicate 733-742; ready-latched gate 884-918; AppliedPending-only trial selection 994-1033; planned restart kill/wait 1115-1121."},{"path":"crates/spt-daemon/src/nethost.rs","description":"Local rc uses broker-minted loopback with own-node identity, not by:None semantics: 320-345, 1733-1755. Stream subscriber has separate writer sharing SharedSend: 143-215."},{"path":".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/logs/daemon.stderr.log","description":"Primary observed lifecycle evidence: original conn17 at lines39-55; replacement conn70 lines65-80; conn88 lines81-95; conn104 lines96-110."},{"path":"docs/49-267-REFRESH-OBSERVATION-PLAN.md","description":"Q1/Q3 event interpretation limits, cause-attribution rule, no inference from absent lines; generation versus receipt and input echo versus acknowledgement/state distinctions."},{"path":"docs/KNOWN-HAZARDS.md","description":"Paid-for predecessor mechanisms: 7.36 logical release/promotion; 7.38 logical role eviction without physical cancellation; 7.43 composition of individually bounded writes and redispatch; 7.44 long-lived stream/seat cleanup."},{"path":"traceable-reqs.toml","description":"Historical refresh fix chain and deferred clean-close seed at2376-2403; distinguishes reap drive, resume ownership theft, cursor-only resume, and physical cancellation scopes."}],"architecture":"A session owns one ControllerSink keyed by SharedSend identity, by identity, attach generation, and a separate monotonically increased writer epoch. PTY output queues nonblocking to its writer; controller and stream-subscriber writers may share one physical BrokerConn serialized write gate. Physical retirement cancels connection I/O; EOF cleanup separately detaches matching session/stream roles. Promotion's drain predicate scans only currently seated by:None controllers whose write marker is older than the deadline; it is neither a physical-connection inventory nor end-to-end rc progress proof.","report":"## Observed identity and chronology\nAll four physical broker IPC connections carry session=1, endpoint=w1crig, stream-subscriber stream=1, and controller by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf. All subscribe records retain req_gen=1789523756051. That is the rc attachment generation, NOT brain generation 1 and NOT broker controller_epoch. Each replacement has decision=controller followed by decision=idempotent; no observed old-owner conflict is needed to explain these attempts.\n\n| conn | controller attach wall_ms | actual slot closure wall_ms | observed closure/failure |\n|---|---:|---:|---|\n|17|1789523756143|1789523803471|cause=detach; organic CONN_WRITE_RETIRED BrokenPipe, writer-exit BrokenPipe; transport-close poisoned=true timed_out=false at1789523803573|\n|70|1789523805976|1789523821012|cause=detach; CONN_WRITE_POISONED, stream-sub-writer-poison live write failed, controller writer-exit TimedOut; transport-close at1789523821028|\n|88|1789523821557|1789523836597|same deadline/stream/controller pattern, cause=detach; transport-close at1789523836612|\n|104|1789523837637|1789523852688|same pattern, cause=detach; transport-close at1789523852704|\n\nConn17 closure immediately follows the BRAIN_UPDATE_RESTART record. Source restart path calls child.kill() then child.wait(), not a graceful drain protocol. Generation1 reports BRAIN_RESUMED cursor establishment before replacement conn70 attaches. Replacement lifetimes from controller attach to poisoning are approximately15036ms,15040ms,15051ms. These are NOT blocked-write-onset measurements: first_write is the connection's first write, and no current record identifies the exact blocked operation's start.\n\n## Q1: narrow answer\nYES: occupied session1 controller slots became empty on all four named connections. The cause literal is detach, not send_deadline or stall_evict. clear_controller only emits after a real take of an occupied sink (broker2752-2775). detach_if requires SharedSend Arc identity equality before clearing (3562-3570). This establishes an actual role transition, not just an attempted eviction or already-empty slot.\n\nNO: it does not establish that the outgoing brain gracefully clean-closed conn17 before being killed, nor that rc recovered. conn17 shows BrokenPipe retirement; the supervised planned restart still hard-kills. Also cause=detach describes the closing call path, not a claim that the transport was healthy or that the detach was voluntary.\n\nFor70/88/104 the exact BRAIN_SUBSCRIBER_STALL_EVICT text is the timeout-observability bridge note_controller_write_retired (3824-3836): it tallies a TimedOut writer and says physical retirement/EOF cleanup follows. That helper does NOT clear the slot. The logical sampler stall_evict_controller (2879 onward) has different text and uses cause=stall_evict only in Inline mode; Deferred mode drops the sink directly without the clear_controller line. Therefore reading the token alone as cause=stall_evict is wrong. The observed cause=detach plus later matching session-detach/transport-close is consistent with retirement reaching existing EOF cleanup; line interleaving does not identify which concurrent writer initiated the physical timeout.\n\n## Ownership, epoch, and drain precision\nbecome_controller removes any incumbent, bumps controller_epoch, and installs a fresh writer/sink (2592-2694). Attach generation is separately checked by detach_if_gen; connection-cleanup detach uses pointer identity. A stale deadline eviction can clear only if its captured epoch is still current (2828-2831). Initial-batch writes check the epoch inside the serialized connection gate (controller_writer around4024); live backlog intentionally has no epoch check and drains the old channel until its senders disappear or a write fails. Thus a role replacement/drop is not itself proof a writer stack or OS operation has already terminated.\n\nThe observed writer-exit and transport-close records provide stronger termination evidence for these physical connections than a slot-close record alone. Repeated fresh IDs establish separate carrier failures rather than reuse of a poisoned conn17. The log does not print controller_epoch, so its exact numeric values must not be invented.\n\n## Q3: reachability limits in this setup\n1. Local geography is not local-controller identity. nethost320-345 and1733-1755 show same-node rc goes through broker-minted loopback carrying this node's hex identity. The logs explicitly show by=nodehex, i.e. Some, on17/70/88/104. local_controller_wedged (broker2858-2862) requires by.is_none(), so NONE of these rc controllers counts toward old_gen_drained, even while genuinely stalled. The word role=brain in the accumulated connection label does not override that predicate.\n2. old_gen_drained is !any_local_controller_wedged (brainproc733-742). It scans current session controller slots and aged write markers. It has no old-brain-generation parameter, does not inspect detached writers, excludes Some identities, and treats not-yet-deadline writes as not wedged. A missing broker or poisoned individual log also does not block promotion. Hence its literal contract is narrower than 'every old-generation write has completed'; do not promote historical prose into stronger runtime evidence.\n3. A planned restart/start-reason update is not itself a trial. supervise_brain selects run_trial only for an AppliedPending record (994-1033). The marker BRAIN_UPDATE_RESTART alone cannot prove Q3 was evaluated. Even in a genuine pending trial, these Some(nodehex) controllers cannot make it report NOT_DRAINED.\n4. BRAIN_PROMOTE_GATE_NOT_DRAINED proves exactly one ready-latched trial read returned false, with a trial generation only. It neither attributes a session nor proves heartbeat brain.sessions() completed. Conversely absence of the token is not proof of a drained old connection.\n\nA separate real signed AppliedPending trial remains required for Q3. To exercise the false branch, it additionally needs a qualifying by:None controller with an aged in-flight marker at the gate read. The present local rc alone cannot supply that condition. This is an evidence limitation, not authorization to manufacture pending state or change the predicate.\n\n## Why already-paid-for fixes can operate while the viewport remains stalled\nHistorical7.36 addressed role release and idle stalled-controller detection. Historical7.38 explicitly records why role release was insufficient: a detached writer could retain its SharedSend and remain inside unbounded Windows I/O. The current BrokerConn fixes that physical containment class with one absolute bound spanning gate acquisition and OS completion, poison/cancel/sever, and never-reuse (conn609-710). W1c shows that machinery firing on each replacement and those writers exiting. It does not show that physical containment has regressed to an immortal write.\n\nThe cursor-only resume fix removes the daemon heartbeat connection's undrained PTY subscription (brainproc268-286); BRAIN_RESUMED means cursor establishment, not rc live delivery. Heartbeat reap drive makes old by:None wedges clearable; neither feature validates the reconstructed rc stream pump. The promotion gate is narrower still and unavailable as an rc-progress guarantee here.\n\nHistorical7.43 already warns of composition: many independently bounded15s failures can create a long or persistent user stall, even though every individual connection eventually retires. Current W1c has that structural shape—replacement attaches on the same persistent stream and attachment generation, then distinct carrier failures. This is NOT proof that the old replay-under-mutex implementation returned: current SubscriberSeat uses its own writer outside log locks. The historical note supplies a hypothesis class, not a diagnosis by analogy.\n\n## Ranked RCA candidates and distinguishing evidence\n1. [INFERENCE / leading candidate] The reconstituted serve/consumer path ceases draining its shared IPC carrier after establishment, possibly while synchronously waiting on a stream/broker operation. Both stream-subscriber and controller writers target the SAME named physical connection; conn.write_gated serializes them. One write occupying the gate can make the other time out waiting. Prediction: a blocked capture should show the replacement serve-side reader not consuming conn70/88/104 while a specific request/stream wait remains outstanding, with a writer either in OS I/O or awaiting that same gate. Boundary evidence must distinguish gate-wait from OS-write ownership: the existing TimedOut/eviction pair cannot. A fresh independent query staying responsive is compatible with this per-carrier candidate and argues against a broker-global lock freeze, not against a per-stream wait.\n2. [INFERENCE / second candidate] Persistent stream/replay/ack state across brain restart reconstructs an immediately unhealthy consumer path. All replacements reattach stream1 from_seq=0 with unchanged rc attach_gen, while physical IDs change. Prediction: recovery differs between resuming this same stream and a genuinely new rc stream/attachment; trace boundary positions should show where fresh snapshot delivery stops and whether old stream control frames or acknowledgement expectations block subsequent reads. Snapshot receipt alone cannot decide this. AttachReplayRca owns the implementation-level investigation; this report does not claim its mechanism proven.\n3. [INFERENCE / lower candidate] The rig/rc output reader causes terminal backpressure after restart. Reader-enabled state does not logically prove continuous OS consumption, but the supplied raw capture/snapshot observations and independent child-input ledger make a simple parser-only failure insufficient. Prediction: raw terminal bytes continue but parsed PROG records disappear for a parser bug; a true paused underlying reader should instead show stalled raw reads/terminal write backpressure. A controlled independent fresh viewport would separate terminal reader behavior from stream-reconstruction behavior. No such intervention was performed by this read-only investigation.\n\nDisfavored narrow alternatives: (a) old conn17 merely never released—contradicted by its real slot closure, writer exit, transport close, and later successful replacement seats; (b) a stale detach unseats a new controller—current closure IDs match each incumbent, replacements occur afterward, and pointer/generation guards exist; (c) a stuck by:None promotion gate explains this rc freeze—the named controllers are explicitly excluded, and refresh alone does not establish a trial. Global dead-peer or runtime-starvation claims are not warranted by these lifecycle records; supplied active_dial_tasks=0 and fresh numeric canary are counterevidence to that broad story.\n\nScope compliance: only source, documentation, and preserved logs were read. No files changed; no builds, tests, lints, live probes, or service calls were run."},"error":null}}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":8791,"output":3064,"cacheRead":82304,"cacheWrite":0,"totalTokens":94159,"reasoningTokens":25,"cost":{"input":0.08791,"output":0.1532,"cacheRead":0.082304,"cacheWrite":0,"total":0.323414}},"stopReason":"toolUse","timestamp":1789524004195,"responseId":"resp_0d104aba2b0f3edc016aa9f8246d9c87d1b93bccb1d9cc1274","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"rs_0d104aba2b0f3edc016aa9f82595dc87d1b08ad86cad56148c","type":"reasoning","content":[],"encrypted_content":"gAAAAABqqfgo5D08gg-Sj0mfzUyx-Vbv1vS52S8CWpCx4r6rufxLsG9URRErxzkPczedXE5imrkBGc27nVnyAs2aSNzaS4j7J1q0HL0WsFxmSM_iMguTAarmKsDJgg3NgIgUkABXgqbMIOsB6BgYZgEwL8eFwfWDKRK-ZB6VkcZanP9F0kq6I1RfzCaW75xKae1ldV1z9dQVJ2cbLsxFHNku8upItV3ADeHe8yl7z9NiEqOW0I6XQaX1om3o6CwFgtGnlkMQCXk92uqukdnrgE57mvFmJS9g2qSX-54Dt4WpDn9Z74KlaTXhJPLBKgfPK5fpMdD0LMxh41iPEACf0XIWY3fl_UqGZT1BLt5jP0W44_SNYt61E8zPuBFWNeNlRC89sSYEGZS0XRgzsQ001qUkHCMF7kjJDmIxW6ohopI2xf0zNBd306soDTJ0zmiTgK2ihuYaromwrjFXwow5DWsUrZO69bhAzlaP8c2ZaFwU4twQzXPrW_PElU_fQTT7xf9FjwoMriTVXh-A33iKHIkOm7KKbsWQKbpr1y9sCgJZyPC_S_LyUgiyVz1YSb-bdGdtgMXliMQH6JKKCTKgRBKfBSlFqzg_l3vM80m4d5lYjs7LqTGxG4mPqT0zkGmAIi6kFT3G7acPbz2uF0C79egGPuMNulxgwqLSoYesGEwW4hlh0enBGSX_6rfC1cJmZFfeMzjdV_vVClVLof8kcsDeeojuj3hVmDx7x7UpIkV3z7p7g_ak9sieOkOcn2rJCC0QqR4YSktiQ_XHwI2B8yIGgd4_cF13MZ_B6WYdX_82CI43HgpMcBYFBFcgbJB754v0KdYhs3I9DfQQO16BVq4k1wVxpK4MQRowM3kOz-LNN666doygMaa90veJh0yEe_LYtaKYYu7mjCTsGE1z--dm4GH1PeSBu9GAFmbAGEreoIV2FdMnIsustKfML7zHOsvo8vl3BChmLVcOk2xglnAQGocH0UIXoNChS3t4G7sPu6rs0EbRIVzP49PKLGVl0qXgY1713ZH0Q6egIOmUcf7KE2B99vnlY68RQMowRb_rjMVjmrVpQ9QiEPnGBSNoXbz37dRmiUriLAshqE1WGYNOhFR2O7RuUcx39KqSMu-lxG2YKJSjAtqaRE0lgk1XACGGatIRAODNdw7Ya7zzkzG5t26cPwJ7romEnKhyZqbnMbCP3OKGpqGUUfDcx1P60R1Jrk36qGo_mIWRbrQAmC6Zrk85_dJDNUSlvKJgcJtQKJ-J8HGglKE8YnMsl8HzUE6qZVSz4qRLp39V4pE0NfVKfPW34IVCAyJ_GriVrt8BzPgWwJJEr9Ig9JIie8Xe4UXYZ46g4Ua_4RLznEk-BhKPFfVucRnX-k_px1cmRr7k1eA70LewkHqIPan7ukFWc-paTKIuxNIBggaxkl8wZEe3JyWmL9Hldh-UigOAiCBPGFh8MAZ_pUBLqZjnYv-H_rD-zZeira0g","internal_chat_message_metadata_passthrough":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"},"summary":[{"type":"summary_text","text":"**Finalizing evidence report**"}],"metadata":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"}},{"id":"fc_0d104aba2b0f3edc016aa9f828c5bc87d19460469aaa30b5d2","type":"function_call","status":"completed","arguments":"{\"type\":null,\"data\":{\"summary\":\"W1c proves the original controller and three replacement controllers actually vacated session 1, all via cause=detach; it does not prove graceful pre-kill close. Replacement IPC connections were separately deadline-retired, so continued viewport failure is not simply one immortal old connection. This local rc is by=Some(nodehex), explicitly outside old_gen_drained; ordinary refresh also does not by itself enter the AppliedPending promotion trial. RCA remains candidate-only: repeated failure of the reconstructed stream/consumer path despite functioning per-connection containment.\",\"files\":[{\"path\":\"crates/spt-daemon/src/broker.rs\",\"description\":\"Controller ownership/attach generation and epoch: 829-889, 2592-2694; closure causes, reap, wedge predicate: 2720-2917; detach identity: 3547-3588; timeout tally bridge: 3813-3836; controller writer: 3955-4124; aggregate drain predicate: 5989-6019; EOF cleanup: 6424-6465.\"},{\"path\":\"crates/spt-daemon/src/conn.rs\",\"description\":\"Physical-connection deadline versus organic retirement tokens: 331-360; serialized gate plus watchdog/retirement implementation: 609-710. Deadline includes gate waiting, not only OS write.\"},{\"path\":\"crates/spt-daemon/src/brainproc.rs\",\"description\":\"Cursor-only resume 268-286; heartbeat reap drive 391-409; ProductionTrialEnv drain predicate 733-742; ready-latched gate 884-918; AppliedPending-only trial selection 994-1033; planned restart kill/wait 1115-1121.\"},{\"path\":\"crates/spt-daemon/src/nethost.rs\",\"description\":\"Local rc uses broker-minted loopback with own-node identity, not by:None semantics: 320-345, 1733-1755. Stream subscriber has separate writer sharing SharedSend: 143-215.\"},{\"path\":\".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/logs/daemon.stderr.log\",\"description\":\"Primary observed lifecycle evidence: original conn17 at lines39-55; replacement conn70 lines65-80; conn88 lines81-95; conn104 lines96-110.\"},{\"path\":\"docs/49-267-REFRESH-OBSERVATION-PLAN.md\",\"description\":\"Q1/Q3 event interpretation limits, cause-attribution rule, no inference from absent lines; generation versus receipt and input echo versus acknowledgement/state distinctions.\"},{\"path\":\"docs/KNOWN-HAZARDS.md\",\"description\":\"Paid-for predecessor mechanisms: 7.36 logical release/promotion; 7.38 logical role eviction without physical cancellation; 7.43 composition of individually bounded writes and redispatch; 7.44 long-lived stream/seat cleanup.\"},{\"path\":\"traceable-reqs.toml\",\"description\":\"Historical refresh fix chain and deferred clean-close seed at2376-2403; distinguishes reap drive, resume ownership theft, cursor-only resume, and physical cancellation scopes.\"}],\"architecture\":\"A session owns one ControllerSink keyed by SharedSend identity, by identity, attach generation, and a separate monotonically increased writer epoch. PTY output queues nonblocking to its writer; controller and stream-subscriber writers may share one physical BrokerConn serialized write gate. Physical retirement cancels connection I/O; EOF cleanup separately detaches matching session/stream roles. Promotion's drain predicate scans only currently seated by:None controllers whose write marker is older than the deadline; it is neither a physical-connection inventory nor end-to-end rc progress proof.\",\"report\":\"## Observed identity and chronology\\nAll four physical broker IPC connections carry session=1, endpoint=w1crig, stream-subscriber stream=1, and controller by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf. All subscribe records retain req_gen=1789523756051. That is the rc attachment generation, NOT brain generation 1 and NOT broker controller_epoch. Each replacement has decision=controller followed by decision=idempotent; no observed old-owner conflict is needed to explain these attempts.\\n\\n| conn | controller attach wall_ms | actual slot closure wall_ms | observed closure/failure |\\n|---|---:|---:|---|\\n|17|1789523756143|1789523803471|cause=detach; organic CONN_WRITE_RETIRED BrokenPipe, writer-exit BrokenPipe; transport-close poisoned=true timed_out=false at1789523803573|\\n|70|1789523805976|1789523821012|cause=detach; CONN_WRITE_POISONED, stream-sub-writer-poison live write failed, controller writer-exit TimedOut; transport-close at1789523821028|\\n|88|1789523821557|1789523836597|same deadline/stream/controller pattern, cause=detach; transport-close at1789523836612|\\n|104|1789523837637|1789523852688|same pattern, cause=detach; transport-close at1789523852704|\\n\\nConn17 closure immediately follows the BRAIN_UPDATE_RESTART record. Source restart path calls child.kill() then child.wait(), not a graceful drain protocol. Generation1 reports BRAIN_RESUMED cursor establishment before replacement conn70 attaches. Replacement lifetimes from controller attach to poisoning are approximately15036ms,15040ms,15051ms. These are NOT blocked-write-onset measurements: first_write is the connection's first write, and no current record identifies the exact blocked operation's start.\\n\\n## Q1: narrow answer\\nYES: occupied session1 controller slots became empty on all four named connections. The cause literal is detach, not send_deadline or stall_evict. clear_controller only emits after a real take of an occupied sink (broker2752-2775). detach_if requires SharedSend Arc identity equality before clearing (3562-3570). This establishes an actual role transition, not just an attempted eviction or already-empty slot.\\n\\nNO: it does not establish that the outgoing brain gracefully clean-closed conn17 before being killed, nor that rc recovered. conn17 shows BrokenPipe retirement; the supervised planned restart still hard-kills. Also cause=detach describes the closing call path, not a claim that the transport was healthy or that the detach was voluntary.\\n\\nFor70/88/104 the exact BRAIN_SUBSCRIBER_STALL_EVICT text is the timeout-observability bridge note_controller_write_retired (3824-3836): it tallies a TimedOut writer and says physical retirement/EOF cleanup follows. That helper does NOT clear the slot. The logical sampler stall_evict_controller (2879 onward) has different text and uses cause=stall_evict only in Inline mode; Deferred mode drops the sink directly without the clear_controller line. Therefore reading the token alone as cause=stall_evict is wrong. The observed cause=detach plus later matching session-detach/transport-close is consistent with retirement reaching existing EOF cleanup; line interleaving does not identify which concurrent writer initiated the physical timeout.\\n\\n## Ownership, epoch, and drain precision\\nbecome_controller removes any incumbent, bumps controller_epoch, and installs a fresh writer/sink (2592-2694). Attach generation is separately checked by detach_if_gen; connection-cleanup detach uses pointer identity. A stale deadline eviction can clear only if its captured epoch is still current (2828-2831). Initial-batch writes check the epoch inside the serialized connection gate (controller_writer around4024); live backlog intentionally has no epoch check and drains the old channel until its senders disappear or a write fails. Thus a role replacement/drop is not itself proof a writer stack or OS operation has already terminated.\\n\\nThe observed writer-exit and transport-close records provide stronger termination evidence for these physical connections than a slot-close record alone. Repeated fresh IDs establish separate carrier failures rather than reuse of a poisoned conn17. The log does not print controller_epoch, so its exact numeric values must not be invented.\\n\\n## Q3: reachability limits in this setup\\n1. Local geography is not local-controller identity. nethost320-345 and1733-1755 show same-node rc goes through broker-minted loopback carrying this node's hex identity. The logs explicitly show by=nodehex, i.e. Some, on17/70/88/104. local_controller_wedged (broker2858-2862) requires by.is_none(), so NONE of these rc controllers counts toward old_gen_drained, even while genuinely stalled. The word role=brain in the accumulated connection label does not override that predicate.\\n2. old_gen_drained is !any_local_controller_wedged (brainproc733-742). It scans current session controller slots and aged write markers. It has no old-brain-generation parameter, does not inspect detached writers, excludes Some identities, and treats not-yet-deadline writes as not wedged. A missing broker or poisoned individual log also does not block promotion. Hence its literal contract is narrower than 'every old-generation write has completed'; do not promote historical prose into stronger runtime evidence.\\n3. A planned restart/start-reason update is not itself a trial. supervise_brain selects run_trial only for an AppliedPending record (994-1033). The marker BRAIN_UPDATE_RESTART alone cannot prove Q3 was evaluated. Even in a genuine pending trial, these Some(nodehex) controllers cannot make it report NOT_DRAINED.\\n4. BRAIN_PROMOTE_GATE_NOT_DRAINED proves exactly one ready-latched trial read returned false, with a trial generation only. It neither attributes a session nor proves heartbeat brain.sessions() completed. Conversely absence of the token is not proof of a drained old connection.\\n\\nA separate real signed AppliedPending trial remains required for Q3. To exercise the false branch, it additionally needs a qualifying by:None controller with an aged in-flight marker at the gate read. The present local rc alone cannot supply that condition. This is an evidence limitation, not authorization to manufacture pending state or change the predicate.\\n\\n## Why already-paid-for fixes can operate while the viewport remains stalled\\nHistorical7.36 addressed role release and idle stalled-controller detection. Historical7.38 explicitly records why role release was insufficient: a detached writer could retain its SharedSend and remain inside unbounded Windows I/O. The current BrokerConn fixes that physical containment class with one absolute bound spanning gate acquisition and OS completion, poison/cancel/sever, and never-reuse (conn609-710). W1c shows that machinery firing on each replacement and those writers exiting. It does not show that physical containment has regressed to an immortal write.\\n\\nThe cursor-only resume fix removes the daemon heartbeat connection's undrained PTY subscription (brainproc268-286); BRAIN_RESUMED means cursor establishment, not rc live delivery. Heartbeat reap drive makes old by:None wedges clearable; neither feature validates the reconstructed rc stream pump. The promotion gate is narrower still and unavailable as an rc-progress guarantee here.\\n\\nHistorical7.43 already warns of composition: many independently bounded15s failures can create a long or persistent user stall, even though every individual connection eventually retires. Current W1c has that structural shape—replacement attaches on the same persistent stream and attachment generation, then distinct carrier failures. This is NOT proof that the old replay-under-mutex implementation returned: current SubscriberSeat uses its own writer outside log locks. The historical note supplies a hypothesis class, not a diagnosis by analogy.\\n\\n## Ranked RCA candidates and distinguishing evidence\\n1. [INFERENCE / leading candidate] The reconstituted serve/consumer path ceases draining its shared IPC carrier after establishment, possibly while synchronously waiting on a stream/broker operation. Both stream-subscriber and controller writers target the SAME named physical connection; conn.write_gated serializes them. One write occupying the gate can make the other time out waiting. Prediction: a blocked capture should show the replacement serve-side reader not consuming conn70/88/104 while a specific request/stream wait remains outstanding, with a writer either in OS I/O or awaiting that same gate. Boundary evidence must distinguish gate-wait from OS-write ownership: the existing TimedOut/eviction pair cannot. A fresh independent query staying responsive is compatible with this per-carrier candidate and argues against a broker-global lock freeze, not against a per-stream wait.\\n2. [INFERENCE / second candidate] Persistent stream/replay/ack state across brain restart reconstructs an immediately unhealthy consumer path. All replacements reattach stream1 from_seq=0 with unchanged rc attach_gen, while physical IDs change. Prediction: recovery differs between resuming this same stream and a genuinely new rc stream/attachment; trace boundary positions should show where fresh snapshot delivery stops and whether old stream control frames or acknowledgement expectations block subsequent reads. Snapshot receipt alone cannot decide this. AttachReplayRca owns the implementation-level investigation; this report does not claim its mechanism proven.\\n3. [INFERENCE / lower candidate] The rig/rc output reader causes terminal backpressure after restart. Reader-enabled state does not logically prove continuous OS consumption, but the supplied raw capture/snapshot observations and independent child-input ledger make a simple parser-only failure insufficient. Prediction: raw terminal bytes continue but parsed PROG records disappear for a parser bug; a true paused underlying reader should instead show stalled raw reads/terminal write backpressure. A controlled independent fresh viewport would separate terminal reader behavior from stream-reconstruction behavior. No such intervention was performed by this read-only investigation.\\n\\nDisfavored narrow alternatives: (a) old conn17 merely never released—contradicted by its real slot closure, writer exit, transport close, and later successful replacement seats; (b) a stale detach unseats a new controller—current closure IDs match each incumbent, replacements occur afterward, and pointer/generation guards exist; (c) a stuck by:None promotion gate explains this rc freeze—the named controllers are explicitly excluded, and refresh alone does not establish a trial. Global dead-peer or runtime-starvation claims are not warranted by these lifecycle records; supplied active_dial_tasks=0 and fresh numeric canary are counterevidence to that broad story.\\n\\nScope compliance: only source, documentation, and preserved logs were read. No files changed; no builds, tests, lints, live probes, or service calls were run.\"},\"error\":null}","call_id":"call_DhuDSwrHxoSiQdwfvEZlJjZ1","internal_chat_message_metadata_passthrough":{"create_time":1789524004.625759,"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"},"name":"yield","metadata":{"turn_id":"b13ca808-0fcd-4fa1-8f3e-ed30551f6694"}}]},"duration":93339.08539999998,"ttft":1414.565600000089,"completedAt":1789524097536,"contextSnapshot":{"promptTokens":91095,"nonMessageTokens":10480,"compactionEpoch":0}}}
{"type":"message","id":"eac2a552","parentId":"bf6624d5","timestamp":"2026-09-16T02:01:37.541Z","message":{"role":"toolResult","toolCallId":"call_DhuDSwrHxoSiQdwfvEZlJjZ1|fc_0d104aba2b0f3edc016aa9f828c5bc87d19460469aaa30b5d2","toolName":"yield","content":[{"type":"text","text":"Result submitted."}],"details":{"data":{"summary":"W1c proves the original controller and three replacement controllers actually vacated session 1, all via cause=detach; it does not prove graceful pre-kill close. Replacement IPC connections were separately deadline-retired, so continued viewport failure is not simply one immortal old connection. This local rc is by=Some(nodehex), explicitly outside old_gen_drained; ordinary refresh also does not by itself enter the AppliedPending promotion trial. RCA remains candidate-only: repeated failure of the reconstructed stream/consumer path despite functioning per-connection containment.","files":[{"path":"crates/spt-daemon/src/broker.rs","description":"Controller ownership/attach generation and epoch: 829-889, 2592-2694; closure causes, reap, wedge predicate: 2720-2917; detach identity: 3547-3588; timeout tally bridge: 3813-3836; controller writer: 3955-4124; aggregate drain predicate: 5989-6019; EOF cleanup: 6424-6465."},{"path":"crates/spt-daemon/src/conn.rs","description":"Physical-connection deadline versus organic retirement tokens: 331-360; serialized gate plus watchdog/retirement implementation: 609-710. Deadline includes gate waiting, not only OS write."},{"path":"crates/spt-daemon/src/brainproc.rs","description":"Cursor-only resume 268-286; heartbeat reap drive 391-409; ProductionTrialEnv drain predicate 733-742; ready-latched gate 884-918; AppliedPending-only trial selection 994-1033; planned restart kill/wait 1115-1121."},{"path":"crates/spt-daemon/src/nethost.rs","description":"Local rc uses broker-minted loopback with own-node identity, not by:None semantics: 320-345, 1733-1755. Stream subscriber has separate writer sharing SharedSend: 143-215."},{"path":".spt/preserved/307/hertz-w1bc/w1c-run1/home-w1c/logs/daemon.stderr.log","description":"Primary observed lifecycle evidence: original conn17 at lines39-55; replacement conn70 lines65-80; conn88 lines81-95; conn104 lines96-110."},{"path":"docs/49-267-REFRESH-OBSERVATION-PLAN.md","description":"Q1/Q3 event interpretation limits, cause-attribution rule, no inference from absent lines; generation versus receipt and input echo versus acknowledgement/state distinctions."},{"path":"docs/KNOWN-HAZARDS.md","description":"Paid-for predecessor mechanisms: 7.36 logical release/promotion; 7.38 logical role eviction without physical cancellation; 7.43 composition of individually bounded writes and redispatch; 7.44 long-lived stream/seat cleanup."},{"path":"traceable-reqs.toml","description":"Historical refresh fix chain and deferred clean-close seed at2376-2403; distinguishes reap drive, resume ownership theft, cursor-only resume, and physical cancellation scopes."}],"architecture":"A session owns one ControllerSink keyed by SharedSend identity, by identity, attach generation, and a separate monotonically increased writer epoch. PTY output queues nonblocking to its writer; controller and stream-subscriber writers may share one physical BrokerConn serialized write gate. Physical retirement cancels connection I/O; EOF cleanup separately detaches matching session/stream roles. Promotion's drain predicate scans only currently seated by:None controllers whose write marker is older than the deadline; it is neither a physical-connection inventory nor end-to-end rc progress proof.","report":"## Observed identity and chronology\nAll four physical broker IPC connections carry session=1, endpoint=w1crig, stream-subscriber stream=1, and controller by=c96d58a2b4533a6b067bef2e90d5a47db9afcb0a4df1f71678531bed53e033cf. All subscribe records retain req_gen=1789523756051. That is the rc attachment generation, NOT brain generation 1 and NOT broker controller_epoch. Each replacement has decision=controller followed by decision=idempotent; no observed old-owner conflict is needed to explain these attempts.\n\n| conn | controller attach wall_ms | actual slot closure wall_ms | observed closure/failure |\n|---|---:|---:|---|\n|17|1789523756143|1789523803471|cause=detach; organic CONN_WRITE_RETIRED BrokenPipe, writer-exit BrokenPipe; transport-close poisoned=true timed_out=false at1789523803573|\n|70|1789523805976|1789523821012|cause=detach; CONN_WRITE_POISONED, stream-sub-writer-poison live write failed, controller writer-exit TimedOut; transport-close at1789523821028|\n|88|1789523821557|1789523836597|same deadline/stream/controller pattern, cause=detach; transport-close at1789523836612|\n|104|1789523837637|1789523852688|same pattern, cause=detach; transport-close at1789523852704|\n\nConn17 closure immediately follows the BRAIN_UPDATE_RESTART record. Source restart path calls child.kill() then child.wait(), not a graceful drain protocol. Generation1 reports BRAIN_RESUMED cursor establishment before replacement conn70 attaches. Replacement lifetimes from controller attach to poisoning are approximately15036ms,15040ms,15051ms. These are NOT blocked-write-onset measurements: first_write is the connection's first write, and no current record identifies the exact blocked operation's start.\n\n## Q1: narrow answer\nYES: occupied session1 controller slots became empty on all four named connections. The cause literal is detach, not send_deadline or stall_evict. clear_controller only emits after a real take of an occupied sink (broker2752-2775). detach_if requires SharedSend Arc identity equality before clearing (3562-3570). This establishes an actual role transition, not just an attempted eviction or already-empty slot.\n\nNO: it does not establish that the outgoing brain gracefully clean-closed conn17 before being killed, nor that rc recovered. conn17 shows BrokenPipe retirement; the supervised planned restart still hard-kills. Also cause=detach describes the closing call path, not a claim that the transport was healthy or that the detach was voluntary.\n\nFor70/88/104 the exact BRAIN_SUBSCRIBER_STALL_EVICT text is the timeout-observability bridge note_controller_write_retired (3824-3836): it tallies a TimedOut writer and says physical retirement/EOF cleanup follows. That helper does NOT clear the slot. The logical sampler stall_evict_controller (2879 onward) has different text and uses cause=stall_evict only in Inline mode; Deferred mode drops the sink directly without the clear_controller line. Therefore reading the token alone as cause=stall_evict is wrong. The observed cause=detach plus later matching session-detach/transport-close is consistent with retirement reaching existing EOF cleanup; line interleaving does not identify which concurrent writer initiated the physical timeout.\n\n## Ownership, epoch, and drain precision\nbecome_controller removes any incumbent, bumps controller_epoch, and installs a fresh writer/sink (2592-2694). Attach generation is separately checked by detach_if_gen; connection-cleanup detach uses pointer identity. A stale deadline eviction can clear only if its captured epoch is still current (2828-2831). Initial-batch writes check the epoch inside the serialized connection gate (controller_writer around4024); live backlog intentionally has no epoch check and drains the old channel until its senders disappear or a write fails. Thus a role replacement/drop is not itself proof a writer stack or OS operation has already terminated.\n\nThe observed writer-exit and transport-close records provide stronger termination evidence for these physical connections than a slot-close record alone. Repeated fresh IDs establish separate carrier failures rather than reuse of a poisoned conn17. The log does not print controller_epoch, so its exact numeric values must not be invented.\n\n## Q3: reachability limits in this setup\n1. Local geography is not local-controller identity. nethost320-345 and1733-1755 show same-node rc goes through broker-minted loopback carrying this node's hex identity. The logs explicitly show by=nodehex, i.e. Some, on17/70/88/104. local_controller_wedged (broker2858-2862) requires by.is_none(), so NONE of these rc controllers counts toward old_gen_drained, even while genuinely stalled. The word role=brain in the accumulated connection label does not override that predicate.\n2. old_gen_drained is !any_local_controller_wedged (brainproc733-742). It scans current session controller slots and aged write markers. It has no old-brain-generation parameter, does not inspect detached writers, excludes Some identities, and treats not-yet-deadline writes as not wedged. A missing broker or poisoned individual log also does not block promotion. Hence its literal contract is narrower than 'every old-generation write has completed'; do not promote historical prose into stronger runtime evidence.\n3. A planned restart/start-reason update is not itself a trial. supervise_brain selects run_trial only for an AppliedPending record (994-1033). The marker BRAIN_UPDATE_RESTART alone cannot prove Q3 was evaluated. Even in a genuine pending trial, these Some(nodehex) controllers cannot make it report NOT_DRAINED.\n4. BRAIN_PROMOTE_GATE_NOT_DRAINED proves exactly one ready-latched trial read returned false, with a trial generation only. It neither attributes a session nor proves heartbeat brain.sessions() completed. Conversely absence of the token is not proof of a drained old connection.\n\nA separate real signed AppliedPending trial remains required for Q3. To exercise the false branch, it additionally needs a qualifying by:None controller with an aged in-flight marker at the gate read. The present local rc alone cannot supply that condition. This is an evidence limitation, not authorization to manufacture pending state or change the predicate.\n\n## Why already-paid-for fixes can operate while the viewport remains stalled\nHistorical7.36 addressed role release and idle stalled-controller detection. Historical7.38 explicitly records why role release was insufficient: a detached writer could retain its SharedSend and remain inside unbounded Windows I/O. The current BrokerConn fixes that physical containment class with one absolute bound spanning gate acquisition and OS completion, poison/cancel/sever, and never-reuse (conn609-710). W1c shows that machinery firing on each replacement and those writers exiting. It does not show that physical containment has regressed to an immortal write.\n\nThe cursor-only resume fix removes the daemon heartbeat connection's undrained PTY subscription (brainproc268-286); BRAIN_RESUMED means cursor establishment, not rc live delivery. Heartbeat reap drive makes old by:None wedges clearable; neither feature validates the reconstructed rc stream pump. The promotion gate is narrower still and unavailable as an rc-progress guarantee here.\n\nHistorical7.43 already warns of composition: many independently bounded15s failures can create a long or persistent user stall, even though every individual connection eventually retires. Current W1c has that structural shape—replacement attaches on the same persistent stream and attachment generation, then distinct carrier failures. This is NOT proof that the old replay-under-mutex implementation returned: current SubscriberSeat uses its own writer outside log locks. The historical note supplies a hypothesis class, not a diagnosis by analogy.\n\n## Ranked RCA candidates and distinguishing evidence\n1. [INFERENCE / leading candidate] The reconstituted serve/consumer path ceases draining its shared IPC carrier after establishment, possibly while synchronously waiting on a stream/broker operation. Both stream-subscriber and controller writers target the SAME named physical connection; conn.write_gated serializes them. One write occupying the gate can make the other time out waiting. Prediction: a blocked capture should show the replacement serve-side reader not consuming conn70/88/104 while a specific request/stream wait remains outstanding, with a writer either in OS I/O or awaiting that same gate. Boundary evidence must distinguish gate-wait from OS-write ownership: the existing TimedOut/eviction pair cannot. A fresh independent query staying responsive is compatible with this per-carrier candidate and argues against a broker-global lock freeze, not against a per-stream wait.\n2. [INFERENCE / second candidate] Persistent stream/replay/ack state across brain restart reconstructs an immediately unhealthy consumer path. All replacements reattach stream1 from_seq=0 with unchanged rc attach_gen, while physical IDs change. Prediction: recovery differs between resuming this same stream and a genuinely new rc stream/attachment; trace boundary positions should show where fresh snapshot delivery stops and whether old stream control frames or acknowledgement expectations block subsequent reads. Snapshot receipt alone cannot decide this. AttachReplayRca owns the implementation-level investigation; this report does not claim its mechanism proven.\n3. [INFERENCE / lower candidate] The rig/rc output reader causes terminal backpressure after restart. Reader-enabled state does not logically prove continuous OS consumption, but the supplied raw capture/snapshot observations and independent child-input ledger make a simple parser-only failure insufficient. Prediction: raw terminal bytes continue but parsed PROG records disappear for a parser bug; a true paused underlying reader should instead show stalled raw reads/terminal write backpressure. A controlled independent fresh viewport would separate terminal reader behavior from stream-reconstruction behavior. No such intervention was performed by this read-only investigation.\n\nDisfavored narrow alternatives: (a) old conn17 merely never released—contradicted by its real slot closure, writer exit, transport close, and later successful replacement seats; (b) a stale detach unseats a new controller—current closure IDs match each incumbent, replacements occur afterward, and pointer/generation guards exist; (c) a stuck by:None promotion gate explains this rc freeze—the named controllers are explicitly excluded, and refresh alone does not establish a trial. Global dead-peer or runtime-starvation claims are not warranted by these lifecycle records; supplied active_dial_tasks=0 and fresh numeric canary are counterevidence to that broad story.\n\nScope compliance: only source, documentation, and preserved logs were read. No files changed; no builds, tests, lints, live probes, or service calls were run."},"status":"success"},"isError":false,"timestamp":1789524097537}}
{"type":"custom_message","customType":"irc:incoming","content":"<irc>\nIncoming IRC message from agent `AttachReplayRca`:\n\nConsumer has no ack-wait in normal attach: net_stream_subscribe + attach_as are send-only; confirmed input send_effect_no_ack, BUT pending_input flushed on Subscribed via send_effect(ack=true) attach.rs754-755. Replay from ring floor=0 replays all old Input, pre-role buffered; batch sends without reads can fill inbound pipe while per-conn broker blocks replying Applied on shared send. Snapshot may forward before Subscribed (attached gate, not role). Candidate only; actual replay counts/interleave not captured. Also stream-sub-writer-poison says LIVE write failed, so that writer already finished replay, excluding stuck replay-write itself. Failure os233 classified Transient, 3 attempts terminal, not SeatBlocked breaker (dispatch373-383,503-545).\n\nSent while waiting/working. Active interruptible wait stopped early for immediate reading.\n\nIf response expected, reply via `hub` (`op: \"send\"`, `to: \"AttachReplayRca\"`); may finish current step first. No one replies on your behalf.\n</irc>","display":true,"details":{"id":"158149cf166e71bc","from":"AttachReplayRca","message":"Consumer has no ack-wait in normal attach: net_stream_subscribe + attach_as are send-only; confirmed input send_effect_no_ack, BUT pending_input flushed on Subscribed via send_effect(ack=true) attach.rs754-755. Replay from ring floor=0 replays all old Input, pre-role buffered; batch sends without reads can fill inbound pipe while per-conn broker blocks replying Applied on shared send. Snapshot may forward before Subscribed (attached gate, not role). Candidate only; actual replay counts/interleave not captured. Also stream-sub-writer-poison says LIVE write failed, so that writer already finished replay, excluding stuck replay-write itself. Failure os233 classified Transient, 3 attempts terminal, not SeatBlocked breaker (dispatch373-383,503-545)."},"attribution":"agent","id":"e097caff","parentId":"eac2a552","timestamp":"2026-09-16T02:00:44.889Z"}
{"type":"custom","customType":"tool_execution_start","data":{"toolCallId":"call_fEQFfOi50znAXK2QTGGdxHQU|fc_0d104aba2b0f3edc016aa9f883b92887d1b11c820559ddb118","toolName":"hub","startedAt":"2026-09-16T02:01:44.241Z","intent":"Forwarding candidate discriminator refinement"},"id":"924d350e","parentId":"e097caff","timestamp":"2026-09-16T02:01:44.241Z"}
{"type":"message","id":"4f2ced06","parentId":"924d350e","timestamp":"2026-09-16T02:01:44.242Z","message":{"role":"assistant","content":[{"type":"toolCall","id":"call_fEQFfOi50znAXK2QTGGdxHQU|fc_0d104aba2b0f3edc016aa9f883b92887d1b11c820559ddb118","name":"hub","arguments":{"i":"Forwarding candidate discriminator refinement","op":"send","to":"Main","message":"Refinement from AttachReplayRca after my report: stream-sub-writer-poison says LIVE write failed, so subscriber writer finished its replay; do not call this a stuck replay write. Peer found pending_input flush on Subscribed uses send_effect(ack=true) attach.rs754-755, unlike confirmed input send_effect_no_ack; old Input replay can buffer before role, then batch sends without reads may deadlock against Applied replies on shared carrier. Candidate only; counts/interleave absent. Three os233 Transient attempts can terminate recovery, not trigger SeatBlocked breaker. My submitted lifecycle/drain report remains valid."},"intent":"Forwarding candidate discriminator refinement"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":3401,"output":156,"cacheRead":91008,"cacheWrite":0,"totalTokens":94565,"cost":{"input":0.034010000000000006,"output":0.0078000000000000005,"cacheRead":0.09100799999999999,"cacheWrite":0,"total":0.132818}},"stopReason":"toolUse","timestamp":1789524097553,"responseId":"resp_0d104aba2b0f3edc016aa9f881cdac87d1a3e61f74cbac5900","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"fc_0d104aba2b0f3edc016aa9f883b92887d1b11c820559ddb118","type":"function_call","status":"completed","arguments":"{\"i\":\"Forwarding candidate discriminator refinement\",\"op\":\"send\",\"to\":\"Main\",\"message\":\"Refinement from AttachReplayRca after my report: stream-sub-writer-poison says LIVE write failed, so subscriber writer finished its replay; do not call this a stuck replay write. Peer found pending_input flush on Subscribed uses send_effect(ack=true) attach.rs754-755, unlike confirmed input send_effect_no_ack; old Input replay can buffer before role, then batch sends without reads may deadlock against Applied replies on shared carrier. Candidate only; counts/interleave absent. Three os233 Transient attempts can terminate recovery, not trigger SeatBlocked breaker. My submitted lifecycle/drain report remains valid.\"}","call_id":"call_fEQFfOi50znAXK2QTGGdxHQU","internal_chat_message_metadata_passthrough":{"create_time":1789524098.123113,"turn_id":"53b29861-63d2-4a9c-951a-bcf54e7d463d"},"name":"hub","metadata":{"turn_id":"53b29861-63d2-4a9c-951a-bcf54e7d463d"}}]},"duration":6686.223600000143,"ttft":2199.9904000000097,"completedAt":1789524104240,"contextSnapshot":{"promptTokens":94409,"nonMessageTokens":10480,"compactionEpoch":0}}}
{"type":"message","id":"9e11e3b5","parentId":"4f2ced06","timestamp":"2026-09-16T02:01:44.243Z","message":{"role":"toolResult","toolCallId":"call_fEQFfOi50znAXK2QTGGdxHQU|fc_0d104aba2b0f3edc016aa9f883b92887d1b11c820559ddb118","toolName":"hub","content":[{"type":"text","text":"Delivered to 1 peer(s):\n- Main: injected"}],"details":{"op":"send","from":"ControllerDrainRca","to":"Main","receipts":[{"to":"Main","outcome":"injected"}]},"isError":false,"timestamp":1789524104243}}
{"type":"message","id":"26c6fec4","parentId":"9e11e3b5","timestamp":"2026-09-16T02:01:50.286Z","message":{"role":"assistant","content":[{"type":"text","text":"Additional discriminator forwarded to Main: the stream writer reports **live write failed**, so it had completed its initial replay. A pending-input acknowledgement interaction remains a candidate, not a proven mechanism.","textSignature":"{\"v\":1,\"id\":\"msg_0d104aba2b0f3edc016aa9f88cf85887d1956293e23415843e\",\"phase\":\"final_answer\"}"}],"api":"openai-codex-responses","provider":"openai-codex","model":"gpt-6-astra","usage":{"input":379,"output":41,"cacheRead":94208,"cacheWrite":0,"totalTokens":94628,"cost":{"input":0.0037900000000000004,"output":0.00205,"cacheRead":0.094208,"cacheWrite":0,"total":0.100048}},"stopReason":"stop","timestamp":1789524104297,"responseId":"resp_0d104aba2b0f3edc016aa9f888855887d1b368c626cf803546","providerPayload":{"type":"openaiResponsesHistory","provider":"openai-codex","dt":true,"items":[{"id":"msg_0d104aba2b0f3edc016aa9f88cf85887d1956293e23415843e","type":"message","status":"completed","content":[{"type":"output_text","annotations":[],"logprobs":[],"text":"Additional discriminator forwarded to Main: the stream writer reports **live write failed**, so it had completed its initial replay. A pending-input acknowledgement interaction remains a candidate, not a proven mechanism."}],"internal_chat_message_metadata_passthrough":{"create_time":1789524104.694155,"turn_id":"53b29861-63d2-4a9c-951a-bcf54e7d463d"},"phase":"final_answer","role":"assistant","metadata":{"turn_id":"53b29861-63d2-4a9c-951a-bcf54e7d463d"}}]},"duration":5988.802300000098,"ttft":4699.602399999974,"completedAt":1789524110286,"contextSnapshot":{"promptTokens":94587,"nonMessageTokens":10480,"compactionEpoch":0}}}
{"type":"custom","customType":"session_exit","data":{"reason":"dispose","kind":"normal","recordedAt":"2026-09-16T02:08:37.562Z"},"id":"3b856275","parentId":"26c6fec4","timestamp":"2026-09-16T02:08:37.562Z"}
