mirror of
https://github.com/J3vb/OwnCord.git
synced 2026-09-03 03:50:00 +03:00
* chore(format): one Prettier config at the repository root Every formatting rule in this repository lived under Client/ and covered exactly two globs: Client/src/**/*.ts and Client/tests/**/*.ts. Root Markdown, all of docs/, every YAML and JSON, all CSS, the root scripts and tools/mcp-introspect were formatted by nothing. There was no .editorconfig. The obvious fix -- a second Prettier config at the root for "everything else" -- gives two configs and two ignore files that can silently disagree about the same file. So the root takes ownership instead: config, ignore file and gate move up, and Client/ folds in. Client's inline "prettier" block, its .prettierignore, its format/format:check scripts and its now-unused prettier devDependency are all deleted; knip would have failed client-check on that last one. The .prettierrc.json values are lifted byte-for-byte from Client/package.json, which is what keeps the reformat commit free of client TypeScript churn: 87 tracked files need reformatting and not one of them is under Client/src or Client/tests. .prettierignore carries only what .gitignore does not. Prettier 3 reads the root .gitignore by default, so node_modules/, dist/, coverage/, Client/src/generated/ and docs/security-findings/ need no entry. It does NOT read nested .gitignore files, which is why .remember/ is listed explicitly -- 38 untracked per-machine scratch files were otherwise able to turn a shared gate red. graphify-out/ is listed because its seven files are tracked and .graphify_labels.json is signed byte-for-byte by its .sig, so formatting it would silently invalidate the signature. check:hygiene is registered in scripts/run.mjs and folded into check and release:preflight. It deliberately contains no `gofmt -l` step: gofmt -l prints offenders and still exits 0, so it cannot fail a build. Go formatting is enforced separately. shellcheck and actionlint take their file lists from `git ls-files`, never a filesystem glob -- .claude/worktrees/ holds a gitignored pre-flatten copy of the tree with three .sh files a glob would happily lint. This commit leaves the tree non-conformant on purpose. The reformat is the next commit, so the 87-file diff is reviewable separately from the rule that caused it. Not included: editorconfig-checker. .editorconfig is the editor baseline the audit asked for; Prettier, gofmt and rustfmt already fail CI on the same indentation and newline rules, so a fourth tool checking them again is a gate with no failure mode of its own. Verified: `npx prettier --check .` names 87 tracked files and zero untracked ones; the same command listed 38 .remember/ scratch files before the ignore entry and none after. `node scripts/run.mjs --list` resolves check:hygiene to 8 shell targets and 4 workflow targets. Both package.json files parse. Refs RL-19 / L-13, S-05. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(format): reformat the tree to the repository Prettier rules Mechanical. This commit is `npx prettier --write .` and nothing else -- the rule that caused it landed in the previous commit so this diff can be reviewed as a transformation rather than as 84 files of hunks. 84 tracked files: 54 Markdown, 7 .mjs, 7 JSON, 6 YAML, 4 .js, 3 CSS, 2 TypeScript (the two Playwright configs at Client's root, which the old Client/src + Client/tests globs never covered). No file under Client/src or Client/tests moves, because .prettierrc.json carries Client's former inline values byte-for-byte. Prettier rewrote 87 files, not 84. The three in .github/ISSUE_TEMPLATE/ had CRLF on disk and differ only in line endings, which .gitattributes (`* text=auto eol=lf`) already normalises, so their committed blobs are unchanged. Worth knowing before someone reconciles the two numbers. The largest single diff is .superpowers/findings-ledger.json at 7976 lines rewritten. That is safe to format: nothing writes the ledger programmatically -- render-ledger.mjs reads it and writes only FINDINGS.md -- so no tool will fight Prettier over its style on the next hunt. FINDINGS.md itself is ignored as generated. Verified: `npx prettier --check .` reports "All matched files use Prettier code style", so the pass is both complete and idempotent. All 7 reformatted JSON files were parsed before and after and compared as values: semantically identical, zero content changes. `node .superpowers/render-ledger.mjs --check` still reports 348 valid findings and leaves FINDINGS.md untouched. `node scripts/check-doc-counts.mjs` still passes its selftest and still agrees on 27 claims across 9 watched documents -- table realignment did not break the patterns it matches on. `node scripts/run.mjs --list` still parses. Refs RL-19 / L-13, S-05. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(lint): enforce Go formatting in the Server linter S-05: repository-wide Go formatting was not a required gate. The only gofmt enforcement anywhere was .githooks/pre-commit, which is opt-in per clone (`npm run hooks:install`), only sees staged files, and warns-and-skips when gofmt is off PATH. The obvious fix -- a `gofmt -l` step in CI -- does not work: `gofmt -l` prints its offenders and still exits 0, so the step passes no matter what it finds. scripts/run.mjs has the same problem, which is why check:hygiene has no Go step either. So gofmt goes where it can actually fail something: Server/.golangci.yml. The file was already `version: "2"` but had no `formatters:` block at all, so the 19 enabled linters ran with zero formatters. In v2 gofmt/gofumpt/goimports moved out of `linters.enable` into their own section with its own exclusions. Adding it there means the gate reports through the Lint step of "Server Build & Test", which is already pinned as required on dev -- no new job and no new pin. Every tracked .go file is under Server/ (551 of them, one go.mod), so Server-scoped is repository-wide here. One file was genuinely misformatted: a one-space struct field alignment in Server/admin/handlers_users_broadcast_test.go, fixed in the same commit because a single line does not need its own reformat commit. Trap worth recording: `gofmt -l .` on a Windows working tree lists every file that has CRLF on disk, because gofmt normalises line endings. That reported 18 offenders here, 17 of them ghosts. The blobs are all LF -- .gitattributes forces `eol=lf` -- so CI never saw them, and the honest test is to run gofmt over `git show HEAD:<file>` rather than the working copy. Doing that across all 551 tracked Go files found exactly the one real offender above. Verified both directions with golangci-lint v2 locally: `golangci-lint run ./...` reports 0 issues on the formatted tree; appending a misformatted function to Server/auth/constants.go produces 2 gofmt findings; appending the same misformatted function to Server/db/dbgen/admin.sql.go produces 0, so the exclusion holds. Both files restored and verified clean afterwards. Refs RL-19 / L-13, S-05. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(scripts): escape the NUL separator instead of embedding one The `tracked()` helper added earlier in this branch splits `git ls-files -z` output on NUL. The separator was written as a literal NUL byte rather than the two-character JavaScript escape, so scripts/run.mjs became a binary file: `git diff` refused to show it, `grep` reported "Binary file matches" instead of the line, and `* text=auto` in .gitattributes stops normalising line endings for a blob it detects as binary. The code worked -- splitting on a raw NUL and splitting on "\0" are the same operation -- which is exactly why this is worth fixing before it is inherited. A source file that tooling classifies as binary is a file nobody can review. Verified: zero NUL bytes remain, `grep -n "split("` now prints line 50 instead of "Binary file matches", `node scripts/run.mjs --list` still resolves the same 8 shell and 4 workflow targets, and prettier still reports the file clean. * chore(lint): enforce Rust formatting Rust had no formatting gate of any kind: no rustfmt.toml, no `cargo fmt` anywhere in CI, in scripts/run.mjs, in the Makefile or in the git hooks. Clippy was the only Rust gate, and clippy does not check layout. `cargo fmt --all -- --check` now runs in the rust-tests job, ahead of clippy: a formatting failure is cheap to produce and cheap to fix, and there is no reason to spend a clippy pass to surface one. The stable toolchain in that job requested `components: clippy` only, so rustfmt is added there. Only that job. ci.yml has a second, byte-identical `Install Rust` block in tauri-build; it stays clippy-only, because a full desktop build is the wrong place to discover a misplaced brace. No rustfmt.toml. The default profile is the point of a baseline -- a config file here would be a second opinion about style with nothing to say. Client/src-tauri is a single `[package]`, not a workspace, so `--all` is a safeguard against a future member rather than a fan-out today. Verified: `node scripts/run.mjs --list` resolves check:rust to three steps with `cargo fmt --all -- --check` first, `npm run format` now also runs `cargo fmt --all`, and prettier reports ci.yml, run.mjs and the ci-check skill clean. `cargo fmt --all -- --check` currently fails on 13 files -- that is the reformat, and it is the next commit. Refs RL-19 / L-13. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(format): reformat the Rust crate to rustfmt defaults Mechanical. This commit is `cargo fmt --all` and nothing else; the gate that demands it landed in the previous commit so this diff is reviewable on its own. 13 of the 17 tracked .rs files, +343/-164. The crate had never been formatted, so the changes are the usual first-run set: aligned trailing comments collapsed to single spaces, single-element slice literals folded onto one line, long method chains broken across lines, closure bodies expanded into blocks. Verified: `cargo fmt --all -- --check` is clean, so the pass is complete and idempotent. `cargo clippy --all-targets -- -D warnings` finishes with no warnings, and `cargo test --lib` reports 115 passed / 0 failed -- identical to before the reformat, which is what "mechanical" has to mean for a commit that touches this much of the crate. Refs RL-19 / L-13. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(scripts): make the root facade actually run on Windows Adding the first gate that a contributor would run from the repository root exposed two bugs in the facade, both of which made it silently wrong on the platform this project is developed on. 1. Every npm and npx step failed. bin() appends `.cmd` on Windows, but Node refuses to spawn a .cmd or .bat with shell:false -- the CVE-2024-27980 mitigation -- and fails with EINVAL and a *null* exit status. run.mjs only special-cased ENOENT, so the result was `FAILED: npx prettier --check . exited null` with nothing to explain it. check:client has three npm steps and has never been able to run here. Fixed by spawning only the npm shims through a shell. They are concatenated into a single command string rather than passed as an args array, because shell:true plus a separate array is deprecated (DEP0190) and prints a warning on every invocation; no argument in this file contains a space. 2. Every optional() step was skipped, always. onPath() shelled out to `where` on Windows, but where.exe lives in C:\WINDOWS\System32, which a Git Bash PATH does not necessarily contain -- on this machine PATH carries System32\Wbem, System32\WindowsPowerShell\v1.0 and System32\OpenSSH but not System32 itself. The probe could not start, `probe.status === 0` was false, and golangci-lint and sqlc reported as "not installed" while installed. Fixed by resolving against PATH and PATHEXT directly. No subprocess, and no dependency on which directories happen to be on PATH. A spawn error other than ENOENT now reports its code instead of surfacing as a null exit status. Verified: before, `node scripts/run.mjs check:hygiene` died with "exited null" and both optional steps printed SKIP with the tools present on PATH. After, the same command runs prettier, shellcheck and actionlint and prints "check:hygiene: passed", with no deprecation warning. `golangci-lint` is detected by the new onPath where the old one missed it. Refs RL-20 / L-14. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(format): ignore build output that nested gitignores hide Prettier honours the root .gitignore and no other. Every build and scratch directory in this repository is ignored by a *nested* one -- Client/.gitignore, .serena/.gitignore, .superpowers/sdd/.gitignore -- so none of them were excluded from the new repository-wide gate. The effect is not subtle. Running `cargo test` once drops roughly 850 formattable files into Client/src-tauri/target/, and the hygiene gate goes from clean to "Code style issues found in 939 files". CI never sees it, because a fresh checkout has no build output; every contributor sees it on their second command. Mirrors the three nested files rather than inventing a list: dist, coverage, playwright-report, test-results, .vite, src-tauri/target and src-tauri/gen from Client/.gitignore, plus .serena/ and .superpowers/sdd/. node_modules needs no entry -- Prettier ignores it by default. Verified: `npx prettier --check .` reports "All matched files use Prettier code style" with a fully populated Client/src-tauri/target/ present on disk, and still names README.md when a misformatted table is appended to it. Refs RL-19 / L-13. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(ci): shellcheck, actionlint, and a repository hygiene job The last two gates RL-19 asks for. Neither existed: the shell scripts were never linted, the workflows were never syntax-checked, and .githooks/pre-commit carried hand-written `# shellcheck disable=` directives that nothing had ever read. New `hygiene` job, ubuntu-only and root-scoped, modelled on docs-consistency for the same reason: every gate in it is platform-independent text analysis, and .gitattributes pins eol=lf so a second OS would only re-prove line endings. It runs `npm run check:hygiene` -- the same entry point a contributor runs, not a parallel copy of the commands. shellcheck ships in the runner image. actionlint does not, so it is pinned by version and verified by sha256: an installer script piped from a branch would be the one unverified download in a workflow file that pins every action by commit SHA. Prettier's step moves here from client-check, where it no longer belongs. Both linters found real defects. shellcheck, 3 findings in 8 scripts. Two are SC1125 errors in .githooks/pre-commit: `# shellcheck disable=SC2086 - repo paths contain no spaces` is not a valid directive. Trailing prose makes shellcheck discard the rest of the line, so neither suppression was ever in effect -- and one of the two was written earlier in this same branch, which is a fair demonstration of why the gate is worth having. The prose moves to its own line above. The third is SC2015 in start-server.sh, rewritten as an explicit if. actionlint, 5 findings, all inside `run:` blocks it shellchecks once shellcheck is on PATH. Three SC2015 in load-baseline.yml, rewritten as explicit ifs. Two SC2035 in release.yml, where `sha256sum *` should not become `sha256sum ./*`: the comment four lines above records that ParseChecksumFile exact-matches the last field, so a "./" prefix would strand every deployed server exactly as a "windows/" prefix would. `sha256sum -- *` satisfies the linter and leaves the output bytes identical. Verified all three gates in both directions with shellcheck 0.10.0 and actionlint 1.7.7 on PATH. Passing: `node scripts/run.mjs check:hygiene` prints "check:hygiene: passed" with all three steps run, not skipped. Failing: appending `bait_fn() { cat $1; }` to Server/scripts/voice-test.sh fails on SC2086; changing a runs-on to `ubunt-latest` fails on runner-label; appending a misformatted table to README.md fails prettier. All three files restored and confirmed clean afterwards. actionlint validates the new job in ci.yml itself. Refs RL-19 / L-13, S-05. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(plans): record B1 progress through B1-3 The header still read "B1-0 is complete; B1-1 is the next step" three merged phases later. A plan that misstates where it is costs a reader the same confusion whether it is stale by one phase or three. B1-0 (#1410), B1-1 (#1411), B1-2 (#1412) and B1-3 (this branch) are done; B1-4, dependency automation, is next. Verified: `node scripts/check-doc-counts.mjs` still agrees on 27 claims across 9 watched documents -- this file is one of them -- and prettier reports it clean. * chore(ci): pin Repository Hygiene as a required check on dev The second half of S-05. Its acceptance criterion is "tree is formatted AND a fast required gate fails future drift" -- a check that runs but is not pinned lets a formatting regression merge, so the gate is not a gate until this lands. The name was read off PR #1414 with `gh pr checks` after the job reported `pass` in 26s, not copied out of ci.yml. That order matters: the B0 script records that three pinned names exist in no workflow file at all, and that a required check which never reports blocks every PR forever. Extends the existing script rather than adding a second one, per the B1 plan. Also records, in the "deliberately NOT pinned" list, that Docs & Ledger Consistency reports and passes on a dev PR yet is unpinned. That reads as an oversight from the 2026-08-25 pass rather than a decision, but it belongs to G-04, so it is documented here and not changed. NOT APPLIED YET. Running this script now would pin a check that PR #1413 cannot report -- its branch predates the hygiene job, so the job does not exist in its workflow file and the check would never arrive. Run it after #1414 merges; #1413 needs a rebase onto dev regardless. Verified: shellcheck clean, the embedded JSON parses, and `check:hygiene` passes with prettier, shellcheck and actionlint all running. Refs S-05, RL-14 / G-03. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
270 lines
14 KiB
Markdown
270 lines
14 KiB
Markdown
# Bug-detection improvements — design
|
||
|
||
Date: 2026-08-08
|
||
Status: partially implemented (verified 2026-08-19) — Tier 1a's `make fuzz`
|
||
target exists (`Server/Makefile`) and Tier 2's five custom ESLint rules
|
||
shipped 2026-08-08 (`Client/eslint-rules.js`), so the gap table
|
||
below is stale for those two rows; Tiers 1b/1c are on-demand npm scripts;
|
||
Tiers 3–4 remain unimplemented.
|
||
|
||
## Problem
|
||
|
||
The multi-agent bug hunt finds real defects at a high rate — the 2026-08-08
|
||
client hunt confirmed 88 bugs and the follow-up sweeps fixed 101 — but it is
|
||
the only mechanism doing so, it costs a large token budget per run, and it has
|
||
never converged. Meanwhile several bug-catching tools are already installed,
|
||
configured, and committed to this repository, and none of them execute.
|
||
|
||
This design adds mechanical detection alongside the agentic hunt, prioritised
|
||
by yield per token spent.
|
||
|
||
## What already exists and does not run
|
||
|
||
| Asset | State | Gap |
|
||
| ----------------------------------------------------- | --------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||
| 14 `Fuzz*` harnesses under `Server/**/*_fuzz_test.go` | Committed | `go test ./...` runs a `Fuzz*` function against its **seed corpus only** — one pass per seed, zero generated inputs. `-fuzz` appears nowhere in the repo. |
|
||
| Stryker mutation testing | `stryker.config.mjs` + `npm run test:mutate` | Referenced in `ci.yml` only inside an npm-audit comment. Has never run. |
|
||
| Browser-mode vitest | `vitest.config.browser.ts` + `npm run test:browser` | CI runs jsdom only. |
|
||
| Cross-package coverage | `make cover-all` prints every 0.0%-covered function | Output is not fed to anything. |
|
||
|
||
Separately, three of the codebase's sharpest invariants are documented in
|
||
`CLAUDE.md` files as prose and asserted nowhere:
|
||
|
||
- `ws`: "a frame that skips the queue, or a seq allocated for a frame that is
|
||
then dropped, is silently unrecoverable"
|
||
- voice: cleanup in an aborted attempt "must be scoped to that attempt's own
|
||
room — a global `leaveVoice()` there kills the live session"
|
||
- E2EE: "must never report an unverified peer as verified"
|
||
|
||
Prose fails no build.
|
||
|
||
## Locked decisions
|
||
|
||
**Everything in this design runs locally, on demand. Nothing is added to
|
||
GitHub Actions.**
|
||
|
||
Rationale: `go test -fuzz` writes each crashing input to
|
||
`testdata/fuzz/<Target>/<hash>`, and that file _is_ a working reproducer. The
|
||
root `CLAUDE.md` states: "This repo is public — unfixed defects do not belong
|
||
in commits, issues, or PR descriptions." Actions artifacts on a public repo are
|
||
downloadable by anyone, and a red scheduled job is itself a public signal that
|
||
something is broken. A Stryker surviving-mutant report is a milder version of
|
||
the same disclosure: a precise map of which behaviour nobody tests.
|
||
|
||
Local-only also means zero new workflow files and zero CI minutes.
|
||
|
||
**Corpus discipline.** A crasher stays uncommitted until its fix exists. The
|
||
`testdata/fuzz/` corpus entry and the fix are committed together, as one
|
||
regression test. This is the same shape as the existing test-first rule.
|
||
|
||
**Always replay a crasher before believing it.** Go runs fuzz targets in
|
||
separate worker processes. When a worker dies without reporting, the
|
||
coordinator cannot tell "crashed on this input" from "was killed externally",
|
||
so it saves the in-flight input to `testdata/fuzz/` as a suspected crasher.
|
||
Interrupting a fuzz run therefore manufactures a fake reproducer that is
|
||
indistinguishable at a glance from a real security finding. Confirm with
|
||
`go test ./<pkg> -run='<FuzzTarget>'` — a real crasher fails there. Observed
|
||
2026-08-08: a 1666-byte malformed JPEG appeared under
|
||
`api/testdata/fuzz/FuzzImageDimensions/` purely because the run was killed.
|
||
|
||
**Never `rm -r` a `testdata/fuzz/<Target>/` directory** to clear a false
|
||
crasher. Committed seed corpus files live in the same directory — deleting the
|
||
directory takes them with it. Remove the single offending file by name.
|
||
|
||
**Superseded 2026-08-08: Tier 2 ships as ESLint rules, not semgrep.** Semgrep
|
||
has no native Windows support (WSL or Docker only), so on this machine it would
|
||
join `make` as tooling that cannot be run locally. ESLint flat config supports
|
||
an inline plugin, so custom rules cost no new dependency — and `npx eslint
|
||
src/` is already a blocking CI gate, which removes the promotion step entirely.
|
||
Rules live in `Client/eslint-rules.js`, tested with `RuleTester`
|
||
in `tests/unit/eslint-rules.test.ts`. See "Tier 2 — delivered" below.
|
||
|
||
## Tier 1 — Turn on what already exists
|
||
|
||
### 1a. `make fuzz`
|
||
|
||
Go fuzzes **one target per package per invocation**, so this cannot be a single
|
||
`go test -fuzz ./...`. The target enumerates fuzz functions and runs each with
|
||
a time budget.
|
||
|
||
Add to `Server/Makefile`, and to its header comment block:
|
||
|
||
```make
|
||
# fuzz Actually fuzz. CI only replays the seed corpus; this generates inputs.
|
||
# Override the per-target budget: FUZZTIME=2m make fuzz
|
||
FUZZTIME ?= 30s
|
||
fuzz:
|
||
@for pkg in $$(go list ./...); do \
|
||
for fn in $$(go test -list='^Fuzz' $$pkg 2>/dev/null | grep '^Fuzz'); do \
|
||
echo "── $$pkg $$fn"; \
|
||
go test $$pkg -run='^$$' -fuzz="^$$fn$$" -fuzztime=$(FUZZTIME) || exit 1; \
|
||
done; \
|
||
done
|
||
```
|
||
|
||
Add `fuzz` to the `.PHONY` list.
|
||
|
||
Default budget 30s per target — a full sweep of 14 targets is about 10 minutes
|
||
unattended. `FUZZTIME=2m` for a deep run.
|
||
|
||
### 1b. Scoped Stryker runs
|
||
|
||
`stryker.config.mjs` already scopes mutation to `src/lib/**` and `src/stores/**`
|
||
with `thresholds.break: 50`. A full run over that scope is expensive; a
|
||
hotspot run is not:
|
||
|
||
```bash
|
||
npx stryker run --mutate "src/lib/dispatcher.ts,src/lib/ws.ts,src/lib/livekitE2EE.ts"
|
||
```
|
||
|
||
Roughly 25 minutes for three files. Surviving mutants identify lines whose
|
||
behaviour can be changed with the entire 4800-test suite still green.
|
||
|
||
Treat the result as advisory. Do not gate on `thresholds.break` — the threshold
|
||
in the config file applies to a full-scope run and is meaningless for a
|
||
three-file subset.
|
||
|
||
Target the files the hunt keeps returning to: `dispatcher.ts`, `ws.ts`,
|
||
`livekitE2EE.ts`, `identity.ts`, and the voice session module.
|
||
|
||
### 1c. Browser-mode vitest
|
||
|
||
Run `npm run test:browser` locally. The client `CLAUDE.md` already documents
|
||
jsdom diverging from native Web Storage semantics; browser mode is the only
|
||
configured surface that observes that class.
|
||
|
||
### 1d. Prerequisite
|
||
|
||
Confirm `Client/reports/`, `Client/.stryker-tmp/`,
|
||
`Server/coverage-all.out`, and `Server/**/testdata/fuzz/` interim output are
|
||
covered by `.gitignore` before running any of the above. Add entries where
|
||
they are missing.
|
||
|
||
## Tier 2 — Bugs to permanent detectors
|
||
|
||
Roughly 200 confirmed real bugs have been fixed across the hunt and harvest
|
||
runs. Each one currently bought exactly one fix. Encoding the recurring
|
||
_classes_ converts them into permanent detectors.
|
||
|
||
**Sources to mine:** bughunt commit history on `fix/bughunt-*` and
|
||
`fix/bughunt-harvest-*` branches, `.superpowers/harvest-med-low-checklist.md`,
|
||
and `docs/audit-*.md`.
|
||
|
||
**Method:** cluster findings by class, not by symptom. Use the installed
|
||
`semgrep-rule-creator` skill, which is test-first — each rule ships with a
|
||
positive fixture that must match and a negative fixture that must not.
|
||
|
||
### Tier 2 — delivered 2026-08-08
|
||
|
||
Five rules, all scoped to the modules their invariant governs, all proven to
|
||
fire by reintroducing the historical bug shape into real source and reverting:
|
||
|
||
| Rule | Encodes |
|
||
| -------------------------------- | ------------------------------------------------------------------------------------------------------------- |
|
||
| `no-leave-voice-when-superseded` | A global `leaveVoice()` inside a branch that already confirmed supersession tears down the newer live session |
|
||
| `e2ee-epoch-needs-keypair-check` | A non-key-holder never bumps the epoch, so an epoch-only staleness guard cannot see a restarted session |
|
||
| `e2ee-verified-status-literal` | Keeps `"verified"` tied to a hand-written call site that earned it, never a computed status |
|
||
| `no-identity-scope-fallback` | A `?? 0` placeholder scope mints a keypair under the wrong account |
|
||
| `no-store-write-in-ws-on` | Page-local `ws.on` handlers may read stores, not write them |
|
||
|
||
**Declined: `await`-then-stale-snapshot.** Not AST-expressible. Whether an
|
||
await needs a guard — and whether the guard present is sufficient and correctly
|
||
placed — is intent, not shape. `livekitSession.ts` alone expresses supersession
|
||
guards in at least four different forms, and several awaits legitimately need
|
||
no guard. Any rule here would be too narrow to catch real bugs or broad enough
|
||
to flag most of the file's already-correct guard code. A rule that misfires on
|
||
correct code gets disabled and trains people to ignore the linter.
|
||
|
||
**Found while writing these:** the dispatcher invariant in the client
|
||
`CLAUDE.md` was factually wrong. It claimed `ws.on(...)` appears only in
|
||
`dispatcher.ts`; eight handlers across `main.ts`, `MainPage.ts` and
|
||
`ChannelController.ts` say otherwise. The true invariant — dispatcher is the
|
||
single path by which server events _write to stores_ — is what the rule
|
||
encodes, and the doc has been corrected to match.
|
||
|
||
**Still open:** the server-side `ws` seq/FIFO invariant, which needs a Go
|
||
runtime assertion rather than a lint rule.
|
||
|
||
Not every fixed bug becomes a rule. A class earns one when it has recurred at
|
||
least twice, or when it corresponds to an invariant already written down in a
|
||
`CLAUDE.md`.
|
||
|
||
## Tier 3 — Stateful and chaos testing
|
||
|
||
Fuzzing and property tests find bad **functions**. Every recurring bug in this
|
||
codebase's history is a bad **ordering**: `registerNow` reconnect-transfer,
|
||
superseded voice sessions, duplicate-message reconciliation, resync corruption,
|
||
the auth-race deep link, the logout/auto-login race. Nothing in the repo
|
||
generates orderings.
|
||
|
||
### 3a. Client model-based tests
|
||
|
||
`fast-check` v4 is already a dependency, and
|
||
`tests/unit/*.property.test.ts` establish the house pattern. Use `fc.commands`.
|
||
|
||
- **Commands:** `Connect`, `Disconnect`, `RegisterNow`, `Receive(seq)`,
|
||
`Supersede`, `Resync`, `Logout`.
|
||
- **Model:** a minimal reference implementation of expected store state — not
|
||
a second copy of the real logic.
|
||
- **Invariants:** message ids never duplicate; per-client seq is monotonic; a
|
||
verified peer never flips to unverified and back; an aborted voice attempt
|
||
never tears down a live session owned by a newer attempt.
|
||
|
||
The shrinking is the point: fast-check reduces a 40-step failure to the minimal
|
||
3-step reproducer, which is what makes an ordering bug fixable at all.
|
||
|
||
### 3b. Server hub simulation
|
||
|
||
A `ws` package test driving random interleavings of subscribe, broadcast, ack
|
||
and disconnect under `-race`, asserting the FIFO and seq property already
|
||
stated in `Server/CLAUDE.md`. Seeded and therefore replayable.
|
||
|
||
### 3c. Fault-injected transport
|
||
|
||
A test-only wrapper that drops, reorders, duplicates and delays frames from a
|
||
seed. Shared by 3a and 3b. Deterministic: a failing seed reproduces exactly.
|
||
|
||
## Tier 4 — Sharpen the hunt
|
||
|
||
The 2026-08-08 client hunt fixed 101 bugs and still did not converge. Four
|
||
changes, cheapest first:
|
||
|
||
1. **Persistent seen-ledger.** Key on `(file, symbol, class)` and persist
|
||
_across_ runs, not only within one. Each run currently starts cold and
|
||
re-derives ground already covered — the most likely reason convergence never
|
||
arrives.
|
||
2. **Sibling-sweep lens.** For every confirmed bug, enumerate the other callers
|
||
of the touched function. This is the root-cause rule turned into a lens: it
|
||
converts one finding into its whole family, which is also what stops the
|
||
same class reappearing in the next round.
|
||
3. **Coverage-guided targeting.** `make cover-all` already prints every
|
||
function at 0.0%. Feed that list to the finders as a priority surface.
|
||
4. **Anti-pattern priming.** Supply the fixed-bug corpus as "confirmed-real
|
||
classes in this codebase — hunt siblings" rather than starting each finder
|
||
from a cold read.
|
||
|
||
## Order and effort
|
||
|
||
| Step | Effort | Runs in |
|
||
| --------------------------------- | ---------------------- | ----------------------- |
|
||
| 1a `make fuzz` | 15 min to write | 10 min/sweep unattended |
|
||
| 1b Stryker hotspots | 0 (already configured) | ~25 min for 3 files |
|
||
| 1d gitignore check | 5 min | — |
|
||
| 2 first four semgrep rules | ~1 afternoon | seconds |
|
||
| 4.1 + 4.2 ledger and sibling lens | ~2 hours | within existing hunt |
|
||
| 1c browser-mode vitest | 0 | minutes |
|
||
| 3 model-based and chaos harnesses | ~1 day | minutes |
|
||
|
||
Tier 1a is first because 14 harnesses — the expensive part — are already
|
||
written and produce nothing today.
|
||
|
||
## Non-goals
|
||
|
||
- No new GitHub Actions workflows, jobs, or scheduled runs.
|
||
- No gating of any existing CI check on mutation score or fuzz results.
|
||
- No change to the existing test suites' assertions. The client suite is green
|
||
and stays green.
|
||
- No promotion of semgrep to CI in this scope.
|
||
- No replacement of the agentic bug hunt. Tier 4 sharpens it; Tiers 1 to 3 run
|
||
beside it.
|