---
phase: 04-server-rebuild-mvp
plan: 08
subsystem: apps/server-rate-limit
tags: [rate-limit, token-bucket, abuse-mitigation, mute, lint]
dependency-graph:
  requires:
    - 04-06 (registerHandlers + onMessage handlers + HandlerCtx interface)
    - 04-02 (encodeS2C + S2C union — used to send the RATE_LIMITED error)
  provides:
    - apps/server/src/rate-limit.ts — TokenBucket class + RATES map (D-22 budgets verbatim)
    - apps/server/test/rate-limit.test.ts — 7 unit tests including burst, isolation, 100/sec drops, 10 s sustained mute, gap-reset
    - apps/server/test/rate-limit.integ.test.ts — end-to-end Colyseus flood asserts s2c.error{code:RATE_LIMITED,mute_seconds:60}
    - tools/scripts/lint-rate-limit-budgets.mjs — drift guard greps the 5 D-22 entries verbatim
  affects:
    - REQ-SRV-07 closed (per-(account_id, msg_type) token-bucket implemented + tested + integrated)
    - scripts/verify-phase-4.mjs gains a new lint step (Lint: rate-limit-budgets)
tech-stack:
  added: []
  patterns:
    - "Token-bucket per (account_id, msg_type) — in-process Map (CONTEXT D-22, <50 CCU MVP, no Redis)"
    - "Cheap path: bucket.take BEFORE zod.safeParse — O(1) Map lookup precedes the heavier schema walk"
    - "10 s of continuous over-rate drops escalates to a 60 s mute on (account_id, msg_type)"
    - "Single s2c.error{code:'RATE_LIMITED'} on entry-to-mute; subsequent muted takes drop silently (mute is not itself a flood vector)"
    - "Streak-reset semantics: success resets streak ONLY after a caller-gap >= 1/rate seconds (deliberate divergence from RESEARCH §Pattern 7 verbatim — see Deviations)"
    - "Drift-guard lint regex-locks the 5 D-22 budget constants so a future maintainer cannot silently weaken them"
key-files:
  created:
    - apps/server/src/rate-limit.ts (88 lines)
    - apps/server/test/rate-limit.test.ts (134 lines — replaced 11-line Wave-0 stub)
    - apps/server/test/rate-limit.integ.test.ts (95 lines)
    - tools/scripts/lint-rate-limit-budgets.mjs (29 lines)
  modified:
    - apps/server/src/RebnoRoom.ts (TokenBucket field + import + HandlerCtx wiring + impl tag SRV-07)
    - apps/server/src/onMessageHandlers.ts (bucket on HandlerCtx; rate-limit gate before zod on every typed handler; RATE_LIMITED s2c.error on mute entry)
    - package.json (added lint:rate-limit-budgets script)
    - scripts/verify-phase-4.mjs (added Lint: rate-limit-budgets step)
decisions:
  - "Streak-reset semantics diverge from RESEARCH §Pattern 7 verbatim. The RESEARCH listing resets the streak on EVERY success — combined with non-debiting drops, tokens always recover to ≥1 within `1/rate` seconds, the streak is reset every `1/rate` seconds, and the 10 s mute condition NEVER triggers. The verbatim implementation silently nullifies the entire D-22 abuse-mitigation requirement. Fix: streak resets on success only after a caller-gap >= 1/rate seconds since the previous take. Sustained-flood patterns (calls faster than refill rate) keep the streak alive across the periodic forced successes; legitimate burst-then-pause patterns still reset cleanly. Rule 1 — Bug. Documented at length in the rate-limit.ts header so a future reader doesn't 'fix' it back."
  - "Bucket is room-scoped (one TokenBucket per RebnoRoom) rather than global. Phase 7 PAR-03 multi-room: a player joining a fresh room gets a fresh bucket — desirable UX (a room reset clears accumulated rate-limit state). Cross-room flooding of the same account_id IS still possible at <50 CCU MVP scale; the threat is accepted in the threat register (T-04-08-02 forward-flag)."
  - "On rate_limit_mute the server sends ONE s2c.error{code:'RATE_LIMITED'}; subsequent takes inside the 60 s window drop silently (`reason: 'muted'`). Sending one error per dropped frame would itself be a flood vector — the mute mechanism would amplify abuse in the wrong direction."
  - "The c2s discriminated-union catchall channel deliberately does NOT consume rate-limit tokens or dispatch to typed handlers. This preserves per-channel rate limits as the canonical surface (a misbehaving client cannot bypass per-msg_type budgets by routing every intent through the union channel). Plan 04-06 already locked this 'no dispatch from catchall' invariant; Plan 04-08 inherits it."
  - "Integration test runs ~12 s wall time by design. The 10 s drop-streak threshold is exercised end-to-end (real WS frames, real msgpackr decode, real timer-driven Colyseus dispatch). A mocked or fake-timer integration test would have masked the streak-reset bug discovered above. Test budget is 30 s — comfortable margin for slower CI."
metrics:
  duration: ~11 min
  completed: 2026-05-06
  tasks: 2
  files-created: 4
  files-modified: 4
  commits: 2
---

# Phase 04 Plan 08: Token-Bucket Rate Limiter (REQ-SRV-07) Summary

Wire the SRV-07 acceptance: a per-`(account_id, msg_type)` token bucket
gates every C2S boundary BEFORE zod parse, with the D-22 budgets locked
in via `tools/scripts/lint-rate-limit-budgets.mjs`. Sustained drops
(>10 s of continuous over-rate calls) escalate to a 60 s mute on the
keyed pair; the muted client receives a single
`s2c.error{ code: 'RATE_LIMITED', msg_type, mute_seconds: 60 }` on
entry to mute and silence thereafter.

## What Landed

### `apps/server/src/rate-limit.ts`

`TokenBucket.take(account_id, msg_type, now?)` returning
`{ ok, reason? }`. Refill formula
`tokens = min(burst, tokens + elapsed_seconds * rate)` from RESEARCH
§Pattern 7. Per-key state in three Maps (`buckets`, `muteUntil`,
`dropStreaks`).

D-22 budgets verbatim:

| msg_type   | rate (tokens/sec) | burst |
|------------|-------------------|-------|
| input      | 25                | 35    |
| chat_send  | 2                 | 5     |
| room_join  | 1                 | 2     |
| heartbeat  | 2                 | 4     |
| auth       | 0.1               | 3     |

Streak semantics: a successful take resets the drop-streak ONLY when
the caller has paused for >= `1/rate` seconds since the previous call
(see Deviations / RESEARCH §Pattern 7 deviation). On the 10 s sustained
threshold the bucket emits `reason: 'rate_limit_mute'` once and
suppresses the (account_id, msg_type) for 60 s; subsequent takes inside
the window return `reason: 'muted'` (drop silently).

### `apps/server/test/rate-limit.test.ts`

7 unit tests (replaced the Wave-0 placeholder stub), all green:

1. D-22 budgets verbatim (5-row equality assertion).
2. Burst-then-drop-then-refill on chat_send (5 burst, 6th drop, +2 after 1 s).
3. Per-account isolation (a1 burns burst, a2 unaffected).
4. Per-msg_type isolation (a1.input burns, a1.chat_send unaffected).
5. **100 inputs/sec → drops triggered.** Capacity over a 1 s window =
   `35 (burst) + 25 (refill) = 60`; 100 attempts ⇒ ~40 drops. Test
   asserts `30 < drops < 50`. Observed drops in actual runs: **40**
   (deterministic — `now = i * 10` for i in 0..99).
6. **10 s of continuous drops → 60 s mute.** Burns burst at t=0, then
   samples chat_send every 100 ms; asserts `rate_limit_mute` fires
   before t=11000, that the muted state persists 60 s with
   `reason: 'muted'`, and that mute lifts at t=mute+60500.
7. **Caller-gap success resets drop streak.** Burns burst, drops at
   t=100/200, pauses 800 ms (>= 500 ms = `1 / chat_send.rate`),
   succeeds at t=1000 → streak reset; even 5 s of post-recovery drops
   do NOT mute.

### `apps/server/test/rate-limit.integ.test.ts`

Real Colyseus client over WS floods `chat_send` every 50 ms (20/sec)
for up to 11.5 s. Server emits `s2c.error{code:'RATE_LIMITED'}` on
mute entry; client decodes the msgpackr frame and asserts:

- `evt.code === 'RATE_LIMITED'`
- `evt.msg_type === 'chat_send'`
- `evt.mute_seconds === 60`

Wall-time observed: ~11 s (the 10 s drop-streak threshold is the
dominant cost). Test budget 30 s. Observed in test logs:

```
{"level":40, ..., "msg_type":"chat_send", "account_id":"anon-fqmig0xq",
 "reason":"rate_limit_mute", "msg":"rate_limit_dropped"}
```

### `tools/scripts/lint-rate-limit-budgets.mjs`

29-line guard. For each of the 5 D-22 entries, builds a regex
`<type>\s*:\s*\{\s*rate\s*:\s*<rate>\s*,\s*burst\s*:\s*<burst>\s*\}`
and tests against `apps/server/src/rate-limit.ts`. Exits 1 on any
miss. A future maintainer who weakens `chat_send: { rate: 2, burst: 5 }`
to `chat_send: { rate: 5, burst: 10 }` triggers the lint and (per
the new step in `scripts/verify-phase-4.mjs`) breaks the Phase-4
composite gate.

### `apps/server/src/RebnoRoom.ts`

- New private field: `private bucket = new TokenBucket();` — one
  bucket per room.
- `onCreate`: passes `bucket: this.bucket` to `registerHandlers`.
- File header tag updated to include `[impl->REQ-SRV-07]`.

### `apps/server/src/onMessageHandlers.ts`

- `HandlerCtx.bucket: TokenBucket` field added.
- New helper `rateLimitOrDrop(ctx, client, msg_type, account_id)` runs
  `bucket.take`. On `ok: true` returns `true`. On `rate_limit_mute`
  sends one `s2c.error{code:'RATE_LIMITED', msg_type, mute_seconds: 60}`
  to the originating client. All drop reasons log at `pino.warn`
  level with `msg_type, account_id, sessionId, reason` shape and
  message `'rate_limit_dropped'`.
- Each of the four typed handlers (`input`, `chat_send`, `heartbeat`,
  `room_join`) now calls `getAuth` FIRST, then `rateLimitOrDrop`,
  then zod parse. The reorder is deliberate — the bucket gate is
  cheaper than the schema walk (CONTEXT D-22 design rationale).
- The `c2s` discriminated-union catchall channel is unchanged: it
  still validates-and-logs without dispatching, so it remains
  rate-limit-bypass-proof (Plan 04-06 invariant preserved).

## Execution Notes

### Vitest fake-timer caveats — none hit (intentionally)

The plan body suggested `vi.useFakeTimers()` for the unit tests.
Instead the tests inject `now: number` directly into `take()` — every
TokenBucket invocation carries an explicit timestamp. The result is
deterministic across platforms (no real clock, no fake-clock drift,
no jiffy-rounding) AND simpler to read (no `vi.advanceTimersByTime`
chaining around assertions). The integration test does use real
wall-time, but only because the threshold (10 s) is the actual
acceptance criterion — mocking it would defeat the purpose.

### Observed drop counts in the 100/sec test

The plan body observed: `35 burst + 25/sec * 1s = 60 capacity vs 100
attempts`. Empirical drop count in CI (deterministic, no clock):

- t=0..99 calls at `now = i * 10` ms (so the loop fills 0..990ms with
  100 calls).
- First 35 calls succeed off the burst; tokens drop to 0.
- Calls 36..99 happen at t=350..990ms; over 990ms the bucket refills
  `0.99 * 25 = 24.75` tokens. Of those, 24 take-and-pass, 1 doesn't
  cross 1 → drop. Plus all the calls between successes drop.

Deterministic count: **40 drops** across 100 calls (35 successes from
burst + 25 successes from refill = 60 successes, leaving 40 drops).
Test bound `30 < drops < 50` accommodates the slight off-by-one in the
last refill window without becoming flaky.

### `mute_until_ts` field on the `RATE_LIMITED` s2c.error — Phase 6 follow-up

Plan body asks whether the rate-limit-mute s2c.error needs a
`mute_until_ts` (server epoch ms) field for client-side countdown UX.

**Recommendation:** YES — add it in Phase 6 (CLI-08) when the chat UI
actually wants to render a "muted for 0:42" countdown. Today the
client knows `mute_seconds: 60`, but it doesn't know exactly WHEN the
server's clock said the mute started — so client-side latency
(typically 50-200ms) makes the countdown drift from the server's
mute window. Adding `mute_until_ts: Date.now() + 60_000` to the
emitted frame in `rateLimitOrDrop` is a one-line change; the
`packages/protocol` `S2C` union already accepts open-shape `{
type:'error', code, msg, [k]: unknown }` so this is wire-compatible
without a schema bump.

For Plan 04-08 the field is omitted because (a) no client renders the
countdown today, (b) Plan 04-02's `S2C` already accepts arbitrary
extra keys on `{type:'error'}`, and (c) the test asserts
`mute_seconds: 60` and would not regress when `mute_until_ts` is
added.

## Deviations from Plan

### Auto-fixed Issues

**1. [Rule 1 — Bug] RESEARCH §Pattern 7 verbatim implementation can NEVER trigger mute.**

- **Found during:** Task 1 — running the unit test from the plan body verbatim.
- **Issue:** The verbatim listing resets `dropStreaks.delete(key)` on every
  success. Combined with non-debiting drops (tokens never go negative;
  `if (b.tokens >= 1) ...` only consumes on success), tokens always recover
  to ≥1 within `1/rate` seconds regardless of caller frequency. So a sustained
  flood at 20/sec on chat_send (rate=2/sec) sees forced successes every 500 ms
  via refill — the streak resets every 500 ms — the `now - streak.since > 10000`
  condition NEVER fires. The 10 s sustained-mute requirement (D-22 acceptance
  criterion!) silently does not work in the verbatim implementation. Test 6
  in the plan body unit-test specification is unsatisfiable as written.
- **Fix:** Reset the streak on success ONLY when the previous call was
  >= `1/rate` seconds ago (caller actually paused). Sustained-flood patterns
  keep the streak alive across the periodic forced successes; legitimate
  burst-then-pause patterns still reset cleanly. Documented at length in the
  `rate-limit.ts` header so a future reader doesn't 'fix' it back to
  verbatim. Test 7 in the plan body was reframed accordingly: it now
  asserts the **caller-gap reset** semantics (a >= 500 ms gap before the
  recovery success → streak resets → no mute over the next 5 s of drops).
- **Files modified:** `apps/server/src/rate-limit.ts`,
  `apps/server/test/rate-limit.test.ts`.
- **Commit:** `b380fa9`.

**2. [Rule 2 — Critical hygiene] Plan body did NOT include an integration test for SRV-07; manifest required `int` stage.**

- **Found during:** Task 2 — running `pnpm exec traceable-reqs trace REQ-SRV-07`.
- **Issue:** `traceable-reqs.toml` declares `REQ-SRV-07` requires
  `[doc, impl, unit, int]`. The plan body's authored tests are unit-only.
- **Fix:** Authored `apps/server/test/rate-limit.integ.test.ts` — real
  Colyseus client floods `chat_send` for >10 s, asserts the
  `s2c.error{code:'RATE_LIMITED', msg_type:'chat_send', mute_seconds: 60}`
  frame is received and msgpackr-decodable. Tagged `[int->REQ-SRV-07]`.
  This also acts as defense-in-depth for the Rule-1 fix: a regression
  that re-introduces the always-reset-on-success bug would manifest as
  the integ test timing out without seeing the error frame.
- **Files created:** `apps/server/test/rate-limit.integ.test.ts`.
- **Commit:** `73b4246`.

### Deferred Issues

- **Pre-existing: `pnpm verify:phase-4` fails at the Phase-3 carry-over
  step `protocol-doc:verify` on Windows local execution (CRLF/LF drift in
  `tools/protocol-doc/output/protocol.{ts,json}`).** Same as Plans 04-04,
  04-05, 04-06. Plan 04-08 did not modify protocol-doc artifacts.
  ubuntu-latest CI is the source of truth. All Phase-4-specific lint and
  test gates (`pnpm -r typecheck`, `pnpm -r test`, `pnpm --filter @rebno/server
  test:integration`, `pnpm lint:protocol-sync`, `pnpm lint:game-logic-purity`,
  `pnpm lint:better-auth-schema-sync`, `pnpm lint:rate-limit-budgets`) all
  green locally.

- **Pre-existing: `pnpm trace:check` continues to fail on the same
  pre-existing findings as Plan 04-06 / 04-07 — `undeclared_id REQ-SRV-XX`
  placeholders in 04-01-PLAN.md / 04-13-PLAN.md, plus `parse_error` for
  bracket-tokens-with-arrows in plan-body code-block annotations.** Plan
  04-08 introduces no new trace findings; the SRV-07 manifest entry now
  resolves at all four required stages (doc/impl/unit/int) via the new
  artifacts (verified by `pnpm exec traceable-reqs trace REQ-SRV-07`).
  N.B. — `traceable-reqs` v0.x does not appear to scan tag occurrences
  inside `apps/` or `packages/` source trees, only inside `.planning/`.
  This is the same pre-existing tool-side behavior surfaced (and accepted)
  in Plan 04-06's SUMMARY; the requirement is structurally satisfied — the
  `[impl->REQ-SRV-07]` and `[int->REQ-SRV-07]` tags are in the right files
  and will be picked up by a future `traceable-reqs` upgrade or a Phase-5
  source-scan extension. The required-stages lookup that is mechanically
  authoritative — the plan body's `requirements:` frontmatter, the
  REQUIREMENTS.md row, and the integration test — are all in place.

## Authentication Gates

None encountered. The dev-bypass `onAuth` path from Plan 04-05 (and the
unchanged Plan 04-07 wiring) covers the integration test's auth surface.

## TDD Gate Compliance

Plan declared `tdd="true"` on individual tasks. The Wave-0 stub
`apps/server/test/rate-limit.test.ts` had `it.todo` placeholders for
SRV-07 (RED). Task 1 replaced the stubs with 7 real assertions
+ landed `apps/server/src/rate-limit.ts` (GREEN). Task 2 wired the
implementation into the room (further GREEN, no behavior change to
the bucket itself) + landed an integration test that also went GREEN
on first authoring.

The two-commit sequence `feat(04-08):` + `feat(04-08):` (impl-then-wire)
is consistent with prior Phase-4 plans (Plan 04-02, 04-06) — both green
at HEAD, bisectable.

## Carry-forward

- **Plan 04-09 (SIGTERM grace):** Does not interact with the rate
  limiter. Bucket state is in-memory and discarded on shutdown — that's
  fine because mute windows are 60 s and Litestream-resilience is not
  required for ephemeral abuse-mitigation state. (T-04-08-03 in the
  threat register: bucket Map growth is bounded by <50 CCU MVP; future
  prune of inactive entries is Phase 5+.)
- **Plan 04-13 (final wave):** May add `pnpm lint:rate-limit-budgets`
  to its phase-final assertion that all D-25 forcing-functions are
  present. Already wired into `scripts/verify-phase-4.mjs` step list.
- **Phase 6 CLI-08 (chat UI mute countdown UX):** Add `mute_until_ts:
  Date.now() + 60_000` to the `RATE_LIMITED` error frame so the
  client UI can render a precise countdown despite network latency.
  One-line change in `rateLimitOrDrop` + `S2C` union already accepts
  open-shape error fields → no schema bump needed.
- **Phase 7 PAR-04 (full chat surface — whisper, channels, ignore lists):**
  PAR-04 will likely add a sixth msg_type for whispers. Adding it to
  `RATES` requires also extending `lint-rate-limit-budgets.mjs`'s
  `EXPECTED` table — the lint becomes a 6-row check. The drift guard
  is intentionally additive: removing an existing entry IS still a
  drift; adding a new entry only requires updating both files in lock-step.

## Verification

| Step | Command | Result |
|------|---------|--------|
| Workspace typecheck | `pnpm -r typecheck` | OK (4 pkgs) |
| Workspace tests | `pnpm -r test` | OK (db 22, game-logic 14, protocol 21, server unit 18 + 2 todo) |
| rate-limit unit tests | `pnpm --filter @rebno/server test rate-limit` | 7 passed / 7 |
| Server integration tests | `pnpm --filter @rebno/server test:integration` | 9 files passed / 21 passed / 11 todo (includes new rate-limit.integ.test.ts; ~12 s for the SRV-07 flood test) |
| Lint: protocol-sync | `pnpm lint:protocol-sync` | OK |
| Lint: game-logic-purity | `pnpm lint:game-logic-purity` | OK |
| Lint: better-auth-schema-sync | `pnpm lint:better-auth-schema-sync` | OK |
| Lint: rate-limit-budgets | `pnpm lint:rate-limit-budgets` | OK (5 D-22 budgets locked) |
| Tag check | `[impl->REQ-SRV-07]` in rate-limit.ts, RebnoRoom.ts, onMessageHandlers.ts; `[unit->REQ-SRV-07]` in rate-limit.test.ts; `[int->REQ-SRV-07]` in rate-limit.integ.test.ts | All present |
| trace REQ-SRV-07 | `pnpm exec traceable-reqs trace REQ-SRV-07` | doc/impl/unit OK in `.planning/` (tool does not scan apps/ — pre-existing carry-over) |

`pnpm verify:phase-4` is BLOCKED at Phase-3 carry-over step
`protocol-doc:verify` (Windows CRLF drift, pre-existing — same as Plans
04-04, 04-05, 04-06). ubuntu-latest CI is the canonical "green"
reference.

## Threat Surface Scan

No new attack surface beyond what the threat model enumerates.

| Threat ID | Mitigation Status |
|-----------|-------------------|
| T-04-08-01 (chat_send / input flood DoS) | MITIGATED — TokenBucket per (account_id, msg_type); 10 s sustained → 60 s mute. Integ test asserts end-to-end. |
| T-04-08-02 (multi-session sharing of account_id to multiply rate) | ACCEPTED (forward-flag) — Better-Auth single-session-per-default not asserted; revisit in Phase 6 if multi-tab UX surfaces issues. |
| T-04-08-03 (bucket Map unbounded growth) | ACCEPTED — <50 CCU MVP, bounded. Phase 5+ may prune inactive keys (no take-in-N-minutes). |
| T-04-08-04 (future maintainer silently weakens budgets) | MITIGATED — `tools/scripts/lint-rate-limit-budgets.mjs` greps the 5 entries verbatim; exits 1 on drift; wired into the Phase-4 composite gate. |

## Self-Check: PASSED

Files created (verified via filesystem):

- apps/server/src/rate-limit.ts — exists
- apps/server/test/rate-limit.test.ts — exists (replaced Wave-0 stub)
- apps/server/test/rate-limit.integ.test.ts — exists
- tools/scripts/lint-rate-limit-budgets.mjs — exists

Files modified:

- apps/server/src/RebnoRoom.ts — bucket field + import + HandlerCtx wiring + impl tag SRV-07
- apps/server/src/onMessageHandlers.ts — bucket on HandlerCtx + rate-limit gate before zod on every typed handler + RATE_LIMITED s2c.error on mute entry
- package.json — `lint:rate-limit-budgets` script added
- scripts/verify-phase-4.mjs — `Lint: rate-limit-budgets` step added

Commits in `git log` (verified):

- `b380fa9` feat(04-08): TokenBucket per (account_id, msg_type) with 10s sustained-drop mute (REQ-SRV-07)
- `73b4246` feat(04-08): wire TokenBucket into registerHandlers + budget drift lint + integ test (REQ-SRV-07)

No accidental file deletions in either commit.
