* docs(plans): phased remediation plan for the 2026-08-19 audit Executes the audit's §8 MUST-fix verdict and §9.1 fix order: one phase per finding group, statuses updated in place as phases land. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HtkxwdqE4pUv82GQPRsTeQ * test(client): give the renderWindow-breaker test its own timeout (audit F-5) 30 synchronous 100-row jsdom rebuilds can exceed vitest's default 5s on a loaded runner; the test timed out once under CI-like load and passes in isolation, so it now carries an explicit 20s budget. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HtkxwdqE4pUv82GQPRsTeQ * docs: fix the ten wrong reference-doc statements from audit 2026-08-19 (B-01..B-10) schema.md: migrations 030/031 documented, attachments ON DELETE SET NULL (matching 030's rebuild), index inventory rewritten from cumulative migration state, writer/reader pool split described, default-roles table made a consistent post-migration snapshot, dbgen preamble updated. protocol.md: DM chat events documented as sequenced/ring-buffered/replayable (they are), plugin_broadcast seq flipped to Yes, retry_after claim removed (no WS error carries it), the five enforced-but-documented-as-None rate limits added (channel_focus, mark_read, call_decline, chat_command, ping), E2EE announce/offer budgets corrected incl. the per-target inner cap, BAD_PAYLOAD and NOT_KEY_HOLDER added to the error table, ready voice_states/ roles field lists completed, member_join top-level status documented. api.md: diagnostics endpoint is ADMINISTRATOR-only (H-8) with a per-IP limiter and host:port livekit_url, error-code table now matches emitted codes (INTERNAL_ERROR, STORAGE_ERROR 507; oversize upload is 400), body-cap exemptions listed, identity_public_key documented on PATCH /users/me, plugin endpoints' plain-text errors + X-Plugin-Runtime header documented, /health 503 degraded state documented, metrics/LiveKit CIDR keys named, updates/apply restart-conflict 409s added. Also folds in the audit's D-04/D-05 comment and plan-header staleness fixes (buildReady comment, e2e spec-count comments, logctx stray word, three plan status headers). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HtkxwdqE4pUv82GQPRsTeQ * fix(server): log the five silently-discarded persistence errors (audit F-3/F-4/D-16) Lockout Upsert/Delete/Cleanup failures (auth/ratelimit.go), the H-6 session-cap eviction failure in CreateSession (db/auth_queries.go), and the channel_focus read-state write failure (service/channel.go) all discarded their errors with no trace — a brute-force lockout could silently fail to survive a restart. In-memory behavior is unchanged (warn-and-continue); the lockout write paths are pinned by tests mirroring OC-0061's load-path test. The session-cap and read-state sites are log-only additions on seams the existing suites already exercise on the success path. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HtkxwdqE4pUv82GQPRsTeQ * fix(dm): blocking a user evicts them from the pair's live 1:1 DM voice call (audit F-1) The block gate ran only at voice_join and voluntary voice_token_refresh, so a blocked user already in the shared 1:1 DM call kept their session indefinitely — the same guard-asymmetry family as A-2026-08-03. handleBlockUser now severs the call through the dmVoiceEvictor capability handleCloseDM already exercises, using a new find-only FindDMChannelIDBetween lookup (sqlc-generated; mirrors GetOrCreateDMChannel's is_group=0 clause so group DM calls stay exempt, matching requireDMNotBlocked). Pinned by three handler tests: shared-DM eviction, no-DM no-op, group-only no-op. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HtkxwdqE4pUv82GQPRsTeQ * fix(ws): close the role-reassignment/handshake race (audit F-2) A role reassignment landing mid-handshake was invisible for the socket's whole life: both handshake paths resolved permissions from the auth-time c.user snapshot, revokeUnreadableChannels early-returns for a user not yet in h.clients, and its Unsubscribe no-ops on the pubsub identity guard once a reconnect replaced the client. Three coordinated fixes: (1) refreshUserSnapshot re-reads the user row (and role name) in reconnectPrecheck and handleFreshConnect, fail-closed; (2) the resume-fallback path re-reads the role once more after registerNow and runs the revocation pass when it moved, so the reassignment-vs-registration orderings meet in the middle; (3) revokeUnreadableChannels re-resolves the live client immediately before acting, mirroring RefreshChannelVisibility. Pinned by four tests driving real WS handshakes through the existing race hooks plus a new pre-register/pre-act hook pair; ws suite green under the default and deadlock builds. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HtkxwdqE4pUv82GQPRsTeQ * refactor(client): remove the inert replay-dedup machinery (audit F-6) The server writes auth_ok before the replay burst, so replayDedup — created on socket-open and cleared when auth_ok is processed — could never be active for a real replayed frame, and the dispatcher's isReplaying() unread gates never fired. Their no-op behavior is the correct behavior (a buffer/db resume has no ready payload, so replayed frames must count as unread), so the machinery, the gates, and the misleading comments are removed rather than repaired. The pinning tests injected replay frames in an order a spec-compliant server never produces; they are replaced by a test pinning the real contract (frames after auth_ok are dispatched verbatim; duplicate handling belongs to the stores). Client suite green: 5036/5036. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HtkxwdqE4pUv82GQPRsTeQ * docs(plans): mark remediation phases 1-6 done Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HtkxwdqE4pUv82GQPRsTeQ * fix(ws): nolint the context-less revoke call golangci-lint flags revokeUnreadableChannels takes no context by design (admin HubBroadcaster interface); annotate the one call site inside a ctx-taking function, matching the RefreshChannelVisibility precedent. golangci-lint v2.11.3: 0 issues. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HtkxwdqE4pUv82GQPRsTeQ --------- Co-authored-by: Claude <noreply@anthropic.com>
12 KiB
Bug-detection improvements — design
Date: 2026-08-08
Status: partially implemented (verified 2026-08-19) — Tier 1a's make fuzz
target exists (Server/Makefile) and Tier 2's five custom ESLint rules
shipped 2026-08-08 (Client/tauri-client/eslint-rules.js), so the gap table
below is stale for those two rows; Tiers 1b/1c are on-demand npm scripts;
Tiers 3–4 remain unimplemented.
Problem
The multi-agent bug hunt finds real defects at a high rate — the 2026-08-08 client hunt confirmed 88 bugs and the follow-up sweeps fixed 101 — but it is the only mechanism doing so, it costs a large token budget per run, and it has never converged. Meanwhile several bug-catching tools are already installed, configured, and committed to this repository, and none of them execute.
This design adds mechanical detection alongside the agentic hunt, prioritised by yield per token spent.
What already exists and does not run
| Asset | State | Gap |
|---|---|---|
14 Fuzz* harnesses under Server/**/*_fuzz_test.go |
Committed | go test ./... runs a Fuzz* function against its seed corpus only — one pass per seed, zero generated inputs. -fuzz appears nowhere in the repo. |
| Stryker mutation testing | stryker.config.mjs + npm run test:mutate |
Referenced in ci.yml only inside an npm-audit comment. Has never run. |
| Browser-mode vitest | vitest.config.browser.ts + npm run test:browser |
CI runs jsdom only. |
| Cross-package coverage | make cover-all prints every 0.0%-covered function |
Output is not fed to anything. |
Separately, three of the codebase's sharpest invariants are documented in
CLAUDE.md files as prose and asserted nowhere:
ws: "a frame that skips the queue, or a seq allocated for a frame that is then dropped, is silently unrecoverable"- voice: cleanup in an aborted attempt "must be scoped to that attempt's own
room — a global
leaveVoice()there kills the live session" - E2EE: "must never report an unverified peer as verified"
Prose fails no build.
Locked decisions
Everything in this design runs locally, on demand. Nothing is added to GitHub Actions.
Rationale: go test -fuzz writes each crashing input to
testdata/fuzz/<Target>/<hash>, and that file is a working reproducer. The
root CLAUDE.md states: "This repo is public — unfixed defects do not belong
in commits, issues, or PR descriptions." Actions artifacts on a public repo are
downloadable by anyone, and a red scheduled job is itself a public signal that
something is broken. A Stryker surviving-mutant report is a milder version of
the same disclosure: a precise map of which behaviour nobody tests.
Local-only also means zero new workflow files and zero CI minutes.
Corpus discipline. A crasher stays uncommitted until its fix exists. The
testdata/fuzz/ corpus entry and the fix are committed together, as one
regression test. This is the same shape as the existing test-first rule.
Always replay a crasher before believing it. Go runs fuzz targets in
separate worker processes. When a worker dies without reporting, the
coordinator cannot tell "crashed on this input" from "was killed externally",
so it saves the in-flight input to testdata/fuzz/ as a suspected crasher.
Interrupting a fuzz run therefore manufactures a fake reproducer that is
indistinguishable at a glance from a real security finding. Confirm with
go test ./<pkg> -run='<FuzzTarget>' — a real crasher fails there. Observed
2026-08-08: a 1666-byte malformed JPEG appeared under
api/testdata/fuzz/FuzzImageDimensions/ purely because the run was killed.
Never rm -r a testdata/fuzz/<Target>/ directory to clear a false
crasher. Committed seed corpus files live in the same directory — deleting the
directory takes them with it. Remove the single offending file by name.
Superseded 2026-08-08: Tier 2 ships as ESLint rules, not semgrep. Semgrep
has no native Windows support (WSL or Docker only), so on this machine it would
join make as tooling that cannot be run locally. ESLint flat config supports
an inline plugin, so custom rules cost no new dependency — and npx eslint src/ is already a blocking CI gate, which removes the promotion step entirely.
Rules live in Client/tauri-client/eslint-rules.js, tested with RuleTester
in tests/unit/eslint-rules.test.ts. See "Tier 2 — delivered" below.
Tier 1 — Turn on what already exists
1a. make fuzz
Go fuzzes one target per package per invocation, so this cannot be a single
go test -fuzz ./.... The target enumerates fuzz functions and runs each with
a time budget.
Add to Server/Makefile, and to its header comment block:
# fuzz Actually fuzz. CI only replays the seed corpus; this generates inputs.
# Override the per-target budget: FUZZTIME=2m make fuzz
FUZZTIME ?= 30s
fuzz:
@for pkg in $$(go list ./...); do \
for fn in $$(go test -list='^Fuzz' $$pkg 2>/dev/null | grep '^Fuzz'); do \
echo "── $$pkg $$fn"; \
go test $$pkg -run='^$$' -fuzz="^$$fn$$" -fuzztime=$(FUZZTIME) || exit 1; \
done; \
done
Add fuzz to the .PHONY list.
Default budget 30s per target — a full sweep of 14 targets is about 10 minutes
unattended. FUZZTIME=2m for a deep run.
1b. Scoped Stryker runs
stryker.config.mjs already scopes mutation to src/lib/** and src/stores/**
with thresholds.break: 50. A full run over that scope is expensive; a
hotspot run is not:
npx stryker run --mutate "src/lib/dispatcher.ts,src/lib/ws.ts,src/lib/livekitE2EE.ts"
Roughly 25 minutes for three files. Surviving mutants identify lines whose behaviour can be changed with the entire 4800-test suite still green.
Treat the result as advisory. Do not gate on thresholds.break — the threshold
in the config file applies to a full-scope run and is meaningless for a
three-file subset.
Target the files the hunt keeps returning to: dispatcher.ts, ws.ts,
livekitE2EE.ts, identity.ts, and the voice session module.
1c. Browser-mode vitest
Run npm run test:browser locally. The client CLAUDE.md already documents
jsdom diverging from native Web Storage semantics; browser mode is the only
configured surface that observes that class.
1d. Prerequisite
Confirm Client/tauri-client/reports/, Client/tauri-client/.stryker-tmp/,
Server/coverage-all.out, and Server/**/testdata/fuzz/ interim output are
covered by .gitignore before running any of the above. Add entries where
they are missing.
Tier 2 — Bugs to permanent detectors
Roughly 200 confirmed real bugs have been fixed across the hunt and harvest runs. Each one currently bought exactly one fix. Encoding the recurring classes converts them into permanent detectors.
Sources to mine: bughunt commit history on fix/bughunt-* and
fix/bughunt-harvest-* branches, .superpowers/harvest-med-low-checklist.md,
and docs/audit-*.md.
Method: cluster findings by class, not by symptom. Use the installed
semgrep-rule-creator skill, which is test-first — each rule ships with a
positive fixture that must match and a negative fixture that must not.
Tier 2 — delivered 2026-08-08
Five rules, all scoped to the modules their invariant governs, all proven to fire by reintroducing the historical bug shape into real source and reverting:
| Rule | Encodes |
|---|---|
no-leave-voice-when-superseded |
A global leaveVoice() inside a branch that already confirmed supersession tears down the newer live session |
e2ee-epoch-needs-keypair-check |
A non-key-holder never bumps the epoch, so an epoch-only staleness guard cannot see a restarted session |
e2ee-verified-status-literal |
Keeps "verified" tied to a hand-written call site that earned it, never a computed status |
no-identity-scope-fallback |
A ?? 0 placeholder scope mints a keypair under the wrong account |
no-store-write-in-ws-on |
Page-local ws.on handlers may read stores, not write them |
Declined: await-then-stale-snapshot. Not AST-expressible. Whether an
await needs a guard — and whether the guard present is sufficient and correctly
placed — is intent, not shape. livekitSession.ts alone expresses supersession
guards in at least four different forms, and several awaits legitimately need
no guard. Any rule here would be too narrow to catch real bugs or broad enough
to flag most of the file's already-correct guard code. A rule that misfires on
correct code gets disabled and trains people to ignore the linter.
Found while writing these: the dispatcher invariant in the client
CLAUDE.md was factually wrong. It claimed ws.on(...) appears only in
dispatcher.ts; eight handlers across main.ts, MainPage.ts and
ChannelController.ts say otherwise. The true invariant — dispatcher is the
single path by which server events write to stores — is what the rule
encodes, and the doc has been corrected to match.
Still open: the server-side ws seq/FIFO invariant, which needs a Go
runtime assertion rather than a lint rule.
Not every fixed bug becomes a rule. A class earns one when it has recurred at
least twice, or when it corresponds to an invariant already written down in a
CLAUDE.md.
Tier 3 — Stateful and chaos testing
Fuzzing and property tests find bad functions. Every recurring bug in this
codebase's history is a bad ordering: registerNow reconnect-transfer,
superseded voice sessions, duplicate-message reconciliation, resync corruption,
the auth-race deep link, the logout/auto-login race. Nothing in the repo
generates orderings.
3a. Client model-based tests
fast-check v4 is already a dependency, and
tests/unit/*.property.test.ts establish the house pattern. Use fc.commands.
- Commands:
Connect,Disconnect,RegisterNow,Receive(seq),Supersede,Resync,Logout. - Model: a minimal reference implementation of expected store state — not a second copy of the real logic.
- Invariants: message ids never duplicate; per-client seq is monotonic; a verified peer never flips to unverified and back; an aborted voice attempt never tears down a live session owned by a newer attempt.
The shrinking is the point: fast-check reduces a 40-step failure to the minimal 3-step reproducer, which is what makes an ordering bug fixable at all.
3b. Server hub simulation
A ws package test driving random interleavings of subscribe, broadcast, ack
and disconnect under -race, asserting the FIFO and seq property already
stated in Server/CLAUDE.md. Seeded and therefore replayable.
3c. Fault-injected transport
A test-only wrapper that drops, reorders, duplicates and delays frames from a seed. Shared by 3a and 3b. Deterministic: a failing seed reproduces exactly.
Tier 4 — Sharpen the hunt
The 2026-08-08 client hunt fixed 101 bugs and still did not converge. Four changes, cheapest first:
- Persistent seen-ledger. Key on
(file, symbol, class)and persist across runs, not only within one. Each run currently starts cold and re-derives ground already covered — the most likely reason convergence never arrives. - Sibling-sweep lens. For every confirmed bug, enumerate the other callers of the touched function. This is the root-cause rule turned into a lens: it converts one finding into its whole family, which is also what stops the same class reappearing in the next round.
- Coverage-guided targeting.
make cover-allalready prints every function at 0.0%. Feed that list to the finders as a priority surface. - Anti-pattern priming. Supply the fixed-bug corpus as "confirmed-real classes in this codebase — hunt siblings" rather than starting each finder from a cold read.
Order and effort
| Step | Effort | Runs in |
|---|---|---|
1a make fuzz |
15 min to write | 10 min/sweep unattended |
| 1b Stryker hotspots | 0 (already configured) | ~25 min for 3 files |
| 1d gitignore check | 5 min | — |
| 2 first four semgrep rules | ~1 afternoon | seconds |
| 4.1 + 4.2 ledger and sibling lens | ~2 hours | within existing hunt |
| 1c browser-mode vitest | 0 | minutes |
| 3 model-based and chaos harnesses | ~1 day | minutes |
Tier 1a is first because 14 harnesses — the expensive part — are already written and produce nothing today.
Non-goals
- No new GitHub Actions workflows, jobs, or scheduled runs.
- No gating of any existing CI check on mutation score or fuzz results.
- No change to the existing test suites' assertions. The client suite is green and stays green.
- No promotion of semgrep to CI in this scope.
- No replacement of the agentic bug hunt. Tier 4 sharpens it; Tiers 1 to 3 run beside it.