infra: observability, backups, guardrails, and deployment hardening (#1376)

* docs: add infrastructure roadmap plan

Records the verified recommendations from an infrastructure review in three
tracks: raising the single-instance ceiling, cheap seams for a possible
multi-instance future, and ops hygiene. Includes explicit anti-recommendations
and sequencing. Security-sensitive detail is intentionally excluded per
docs/security.md.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj

* feat(server): real health checks and saturation metrics

/api/v1/metrics now exposes signals that were already computed in memory but
never surfaced: reconnect replay tier hits, event-persister counters, SQLite
writer-pool wait stats, aggregate per-client backpressure counters (including
previously invisible low-priority drops), and permission-cache hit/miss.

/health now returns a real verdict: hub dispatch-loop liveness, a bounded
database ping, and a free-disk check, returning 503 with a subsystem reason
when degraded. Checks are cached so the unauthenticated endpoint cannot
amplify load. The hub's panic breaker now exits the process so a supervisor
can restart it, instead of leaving broadcast delivery silently dead while
clients still appear online.

OTel instruments that were declared but never recorded are now wired
(ws_active_connections, ws_broadcast_latency_seconds, ws_messages_total,
ws_events_dropped_total, voice gauges) or removed (db_query_duration_seconds).
Also corrects the docs/api.md description of broadcast_drops, which counts
hub-queue overflow, not client send-queue overflow.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj

* feat(server): implement scheduled backups, retention, and backup verification

The backup_schedule and backup_retention settings have existed in the admin
panel and API since the initial schema but were never read by any code. The
15-minute maintenance loop now enforces them: a scheduled backup is taken
when the newest backup on disk is older than the schedule interval (manual
backups reset the clock), and retention prunes backups older than the
configured days while always keeping the newest one.

Backups are now verified with PRAGMA integrity_check immediately after
VACUUM INTO (a failed backup is removed rather than listed as restorable)
and again before a restore may overwrite the live database. A failed VACUUM
INTO also cleans up its partial output file — but never a pre-existing one.

The backup directory is configurable via a new backup.dir key (default
data/backups) so operators can point backups at another disk or an off-host
mount, mirroring the SetDatabasePath plumb.

Restore-handler tests now use real SQLite fixtures (the integrity gate
correctly refuses text files) with the mid-copy failure injected through a
test-only copy hook. Also adds audited gosec suppressions to the Windows
disk-free syscall added in the previous commit, which the Windows lint leg
flagged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj

* feat(server): capacity and failure-mode guardrails

- server.max_ws_connections: optional cap on concurrent WebSocket clients,
  checked before the upgrade with a 503 + Retry-After; rejections are counted
  and exposed as ws_conn_rejects in /api/v1/metrics.
- Single-process database lock: an OS-level advisory lock (flock / exclusive
  handle) beside the SQLite file makes a second server process fail fast with
  a clear message instead of silently fighting the first over process-local
  state. A bounded retry covers the self-update/restore restart handoff, and
  the lock mechanism failing (e.g. network filesystems) only warns.
- Disk-space awareness: boot-time warnings for the data and backup volumes,
  plus a disk_free_mb metrics field, via a small cross-platform diskutil
  package (already used by /health).
- Upload storage failures: storage.Save now marks server-side filesystem
  failures with a sentinel (storage.ErrIO); handlers return 507 for those
  instead of blaming the client with a 400, and the emoji route stops echoing
  raw storage errors (which embed absolute paths) into responses.
- Unknown config keys now warn at startup — a typo like admin_alowed_cidrs
  previously kept the default silently while the operator believed the
  setting changed. Never fatal: newer servers tolerate older configs.
- Admin settings honesty: the three stored-but-inert settings (server_icon,
  max_upload_bytes, voice_quality) are shown read-only with a note pointing
  at the real config.yaml keys, instead of pretending to apply.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj

* perf(db): write-path efficiency and capacity knobs

- channel_focus/mark_read now skip the read-state UPSERT when the stored row
  already matches (same last_message_id, no mentions) — refocus events fire
  at up to 10/s/user and every no-op write still occupied the single SQLite
  writer connection. The extra existence check runs on the reader pool, which
  doesn't serialize. Same shape as the session-touch throttle.
- DeleteExpiredSessions is now sargable: migration 031 normalizes legacy
  expiry formats to the RFC3339-Z layout the server writes and indexes
  expires_at, replacing the strftime full-table scan that ran on the writer
  every 15 minutes.
- Boot-time ANALYZE runs only when a migration actually applied; unchanged
  schemas get the cheap PRAGMA optimize instead (which also covers
  crash-restarts that never reached the shutdown optimize).
- The read/write SQL router gets a table-driven test with explicit expected
  values (INSERT ... RETURNING must hit the writer despite being :one).
- New knobs, all defaulting to current behavior: database.max_readers,
  security.auth_rate_limit_multiplier (for shared-NAT communities),
  event_persistence.replay_ring_size and replay_cold_limit.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj

* fix(server): shutdown lifecycle ordering

- The event pruner and maintenance loop are now joined (bounded) before the
  database closes: bgCtx cancellation used to run AFTER database.Close via
  LIFO defers, contradicting its own comment, and neither goroutine was ever
  waited on — a mid-tick scheduled backup or prune could still hold the
  writer while the pool tore down. StartEventPruner returns a done channel
  with the same join contract EventPersister.Stop already had.
- srv.Shutdown now runs before hub.GracefulStop, so in-flight HTTP handlers'
  broadcasts still reach a live hub and the event persister instead of
  vanishing from the replay/event store across a restart. Shutdown does not
  wait on hijacked WebSocket connections, so the swap adds no delay.
- GracefulStopContext threads the 30s shutdown budget into the hub: the 5s
  client-notice window (matching the countdown clients are shown) ends early
  when the budget expires, and is skipped entirely when nobody is connected —
  early-return startup paths and idle servers no longer sleep 5s for an
  audience of zero.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj

* build(deploy): systemd unit, compose hardening, boot-smoked releases, CI polish

- deploy/owncord.service: hardened systemd unit template with the two
  verified caveats encoded (install dir stays writable for self-update under
  ProtectSystem=strict; CAP_NET_BIND_SERVICE for ACME's :80), plus a
  'Linux (systemd)' deployment docs section — the Linux service story was
  previously 'Docker or nothing'.
- New 'Reverse Proxy Topology' docs section with a working nginx snippet and
  the correct signaling-vs-media distinction: /livekit/* is already proxied
  by the server, only WebRTC media ports must be directly reachable.
- docker-compose: log rotation, commented resource limits, and a healthcheck
  backed by a new 'chatserver healthcheck' subcommand (the distroless image
  has no shell) that probes /health without config side effects.
- release.yml: a concurrency group (queue, never cancel), and boot-smoke
  gates — the freshly built server binaries and the Docker image are cold
  booted and probed healthy BEFORE anything is signed or pushed. The release
  feed drives signed self-updates, so a binary that compiles but dies on
  boot previously would have shipped itself to every auto-updating instance.
- ci.yml: client-check/client-tests move to ubuntu with the reasoning
  recorded (no win32 code paths, LF enforced repo-wide); admin-e2e gets a
  written graduation criterion instead of an open-ended non-blocking status.
- docs: Tailscale guide notes the CGNAT range vs the default admin CIDRs;
  architecture overview records presence/voice state as the fifth
  single-instance blocker and the macOS client scope decision.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj

* perf(server): measured load tooling, narrowed invalidation, presence coalescing, storage and CIDR seams

- Fix scripts/k6/ws-load.js against the real wire protocol: envelope-wrapped
  frames, correct message types (typing_start, presence_update), the correct
  /api/v1/ws path, and thresholds that fail a run where nobody authenticated
  or went ready — the script had drifted to pre-envelope framing and reported
  100% green while every auth failed on the first frame. A new
  workflow_dispatch-only load-baseline workflow boots a real server, seeds
  users through the setup/invite APIs, runs the script, and uploads the k6
  summary plus a metrics snapshot for before/after comparison.
- Role-scoped channel-override changes now evict only the affected role's
  members from the permission cache (fail-safe: unreadable member list still
  flushes everything). InvalidateAll here repopulated every connected user —
  two reads each — synchronously inside the admin request via
  RefreshChannelVisibility, a stampede that scaled with total population
  rather than the role's size. Same pattern the per-user override endpoints
  already used.
- Connect/disconnect presence broadcasts now pass through a 300ms latest-wins
  coalescer (QueuePresence): each un-coalesced presence change is a sequenced
  global broadcast (an O(clients) fan-out under seqMu), so a reconnect storm
  fired O(users) of them from the connect critical path. A flap inside the
  window collapses to its final state; the wire format, seq ordering, and
  replay behaviour are unchanged, and the delivery path (BroadcastPresence)
  is untouched.
- Storage seam: api handlers now consume a FileStore interface (consumer-side,
  same pattern as service.Store) with Open returning a seekable storage.File —
  writing down the contract (range-request seeks included) an alternative
  backend would have to meet, without building one.
- The metrics surfaces and the LiveKit webhook/health endpoints get their own
  allowlist keys (metrics_allowed_cidrs, livekit_webhook_allowed_cidrs, both
  defaulting to admin_allowed_cidrs), so a central Prometheus scraper or an
  externally-hosted LiveKit no longer requires widening the admin panel's
  perimeter. Startup now also warns when admin_allowed_cidrs is customized
  while trusted_proxies is empty — behind a proxy or container network the
  check would otherwise compare the proxy's private address, not the client's.
- The container healthcheck probe now PINS the server's own certificate from
  disk (VerifyConnection, exact-match) instead of skipping TLS verification,
  addressing the CodeQL finding on the previous commit; WebPKI verification
  is used when no local cert exists (ACME).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj

* fix(server): address self-review findings on the hardening branch

Seven fixes from a high-effort review of the full branch diff:

- healthcheck CLI now works under tls.mode acme: it overrides ServerName
  with the configured domain for WebPKI verification instead of pinning a
  cert that doesn't exist (or is stale) in that mode. Previously an ACME
  deployment's container healthcheck failed forever.
- /health pings the READER pool (new db.PingRead): the writer ping queued
  behind a scheduled backup's VACUUM INTO and reported the server degraded
  for the whole backup — which an autoheal watchdog would turn into a
  nightly mid-backup restart.
- /health runs its cached checks under context.WithoutCancel so a probe
  that disconnects mid-request cannot poison the shared cache with a false
  degraded verdict for the next 5 seconds.
- The token CLI uses a new db.OpenShared that skips the single-process
  lock: minting a token against a running server is safe under WAL and was
  a documented workflow the lock had broken.
- The per-user TOTP failure cap is no longer scaled by
  security.auth_rate_limit_multiplier — that knob exists for per-IP limits;
  scaling the only cross-IP brute-force defence multiplied an attacker's
  distributed guess budget. Mirrors the unscaled per-user login threshold.
- A direct presence_update now drops the user's queued entry in the
  connect/disconnect coalescer, so a stale connect-time presence can no
  longer flush 300ms later over the user's fresher chosen status.
- The scheduled-backup filename collision loop breaks on any stat error
  and bounds its suffix probing, instead of spinning the maintenance
  goroutine forever on a persistent EACCES.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj

* test(admin): real SQLite fixture for the merged Close-failure restore test

TestHandleRestoreBackup_RestartsWhenCloseFails arrived from main (#1375)
with a plain-text backup fixture; this branch's restore handler verifies
backups with integrity_check before touching the live database, so the text
fixture was (correctly) refused with 400 before the Close-failure branch
under test was reached. Use a real backup via BackupToSafe, matching the
other restore tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj

---------

Co-authored-by: Claude <noreply@anthropic.com>
This commit is contained in:
J3vb
2026-08-15 20:50:47 +02:00
committed by GitHub
co-authored by Claude Fable 5
parent ea0430c5b0
commit f5faf82a60
97 changed files with 3946 additions and 315 deletions
+189
View File
@@ -0,0 +1,189 @@
package admin
import (
"context"
"errors"
"fmt"
"log/slog"
"os"
"path/filepath"
"strconv"
"strings"
"time"
"github.com/owncord/server/db"
)
// Scheduled-backup intervals for the backup_schedule setting values the admin
// UI offers. "off" (or anything unrecognised) disables scheduling.
const (
backupIntervalDaily = 24 * time.Hour
backupIntervalWeekly = 7 * 24 * time.Hour
)
// MaintainBackups implements the backup_schedule / backup_retention settings
// the admin panel has always offered. It is driven by main.go's 15-minute
// maintenance loop, mirroring the expired-session sweep: read the settings
// each tick, take a scheduled backup when the newest backup on disk is older
// than the schedule interval, and prune backups past the retention window.
//
// Freshness is judged by the newest *.db file's mtime, manual backups
// included — an operator who clicked "Backup now" this morning does not need
// a second copy tonight. Retention prunes by mtime in whole days, but never
// removes the newest backup, so a long-dead schedule cannot delete the last
// copy in the directory.
//
// The returned error feeds the maintenance loop's circuit breaker; settings
// simply not existing (fresh DB mid-migration) is not an error.
func MaintainBackups(ctx context.Context, database *db.DB) error {
schedule, err := database.GetSetting(ctx, "backup_schedule")
if err != nil {
if errors.Is(err, db.ErrNotFound) {
return nil
}
return fmt.Errorf("MaintainBackups: reading backup_schedule: %w", err)
}
var interval time.Duration
switch strings.ToLower(strings.TrimSpace(schedule)) {
case "daily":
interval = backupIntervalDaily
case "weekly":
interval = backupIntervalWeekly
}
var firstErr error
if interval > 0 {
if err := runScheduledBackup(ctx, database, interval); err != nil {
slog.Warn("scheduled backup failed", "error", err)
firstErr = err
}
}
if err := pruneExpiredBackups(ctx, database); err != nil {
slog.Warn("backup retention pruning failed", "error", err)
if firstErr == nil {
firstErr = err
}
}
return firstErr
}
// runScheduledBackup takes a backup when the newest existing one is older
// than interval (or none exists).
func runScheduledBackup(ctx context.Context, database *db.DB, interval time.Duration) error {
newest, _, err := scanBackups()
if err != nil {
return err
}
if !newest.IsZero() && time.Since(newest) < interval {
return nil
}
if err := os.MkdirAll(backupBaseDir, 0o750); err != nil {
return fmt.Errorf("creating backup dir: %w", err)
}
// VACUUM INTO refuses an existing destination, and the timestamp only has
// second resolution — suffix on collision instead of failing the tick.
base := "scheduled_" + time.Now().UTC().Format("20060102_150405")
name := base + ".db"
path := filepath.Join(backupBaseDir, name)
for i := 2; ; i++ {
if _, err := os.Stat(path); err != nil {
// ENOENT is the free-slot case. Any OTHER stat error (EACCES on
// the dir, an unreadable mount) cannot be fixed by trying more
// suffixes — stop probing and let VACUUM INTO surface the real
// failure with a legible error instead of spinning this loop.
break
}
if i > 100 {
return fmt.Errorf("scheduled backup: no free filename after %s (tried 100 suffixes)", base)
}
name = fmt.Sprintf("%s_%d.db", base, i)
path = filepath.Join(backupBaseDir, name)
}
if err := database.BackupToSafe(ctx, path, backupBaseDir); err != nil {
return err
}
if err := db.CheckBackupIntegrity(ctx, path); err != nil {
_ = os.Remove(path)
return fmt.Errorf("scheduled backup failed verification: %w", err)
}
slog.Info("scheduled backup created", "name", name)
// Actor 0 = system, same audit action the manual handler writes.
db.WriteAudit(ctx, database, 0, "backup_create", "server", 0, "scheduled backup saved: "+name)
return nil
}
// pruneExpiredBackups deletes *.db backups whose mtime is older than the
// backup_retention window (in days), always keeping the newest one.
func pruneExpiredBackups(ctx context.Context, database *db.DB) error {
retStr, err := database.GetSetting(ctx, "backup_retention")
if err != nil {
if errors.Is(err, db.ErrNotFound) {
return nil
}
return fmt.Errorf("reading backup_retention: %w", err)
}
// Malformed values parse to 0; zero-or-below means retention is off, so a
// typo disables pruning rather than failing the maintenance tick.
days, _ := strconv.Atoi(strings.TrimSpace(retStr))
if days <= 0 {
return nil
}
newest, entries, err := scanBackups()
if err != nil || len(entries) == 0 {
return err
}
cutoff := time.Now().Add(-time.Duration(days) * 24 * time.Hour)
pruned := 0
for _, e := range entries {
if e.mtime.Before(cutoff) && !e.mtime.Equal(newest) {
if rmErr := os.Remove(e.path); rmErr != nil {
slog.Warn("backup retention: failed to remove expired backup", "path", e.path, "error", rmErr)
continue
}
pruned++
}
}
if pruned > 0 {
slog.Info("backup retention: pruned expired backups", "count", pruned, "retention_days", days)
db.WriteAudit(ctx, database, 0, "backup_delete", "server", 0,
fmt.Sprintf("retention pruned %d backup(s) older than %d days", pruned, days))
}
return nil
}
type backupFile struct {
path string
mtime time.Time
}
// scanBackups lists *.db files in the backup dir, returning the newest mtime
// and all entries. A missing directory is "no backups", not an error.
func scanBackups() (newest time.Time, files []backupFile, err error) {
entries, err := os.ReadDir(backupBaseDir)
if err != nil {
if os.IsNotExist(err) {
return time.Time{}, nil, nil
}
return time.Time{}, nil, fmt.Errorf("reading backup dir: %w", err)
}
for _, e := range entries {
if e.IsDir() || filepath.Ext(e.Name()) != ".db" {
continue
}
info, infoErr := e.Info()
if infoErr != nil {
continue
}
mt := info.ModTime()
files = append(files, backupFile{path: filepath.Join(backupBaseDir, e.Name()), mtime: mt})
if mt.After(newest) {
newest = mt
}
}
return newest, files, nil
}
+152
View File
@@ -0,0 +1,152 @@
package admin_test
import (
"context"
"os"
"path/filepath"
"strings"
"testing"
"time"
"github.com/owncord/server/admin"
)
func listBackupFiles(t *testing.T, dir string) []string {
t.Helper()
entries, err := os.ReadDir(dir)
if err != nil {
if os.IsNotExist(err) {
return nil
}
t.Fatalf("ReadDir: %v", err)
}
var names []string
for _, e := range entries {
if filepath.Ext(e.Name()) == ".db" {
names = append(names, e.Name())
}
}
return names
}
func backdate(t *testing.T, path string, age time.Duration) {
t.Helper()
old := time.Now().Add(-age)
if err := os.Chtimes(path, old, old); err != nil {
t.Fatalf("Chtimes: %v", err)
}
}
// TestMaintainBackups_ScheduleAndRetention exercises the full settings-driven
// lifecycle: off is a no-op, daily creates one backup and only one, staleness
// triggers the next, and retention prunes expired backups while always
// keeping the newest.
func TestMaintainBackups_ScheduleAndRetention(t *testing.T) {
database := openAdminTestDB(t)
dir := t.TempDir()
admin.SetBackupBaseDir(dir)
t.Cleanup(func() { admin.SetBackupBaseDir(filepath.Join("data", "backups")) })
ctx := context.Background()
// Settings absent → no-op, no error.
if err := admin.MaintainBackups(ctx, database); err != nil {
t.Fatalf("MaintainBackups with no settings: %v", err)
}
if got := listBackupFiles(t, dir); len(got) != 0 {
t.Fatalf("no-settings tick created files: %v", got)
}
// Schedule off → still a no-op.
mustSetSetting(t, database, "backup_schedule", "off")
mustSetSetting(t, database, "backup_retention", "7")
if err := admin.MaintainBackups(ctx, database); err != nil {
t.Fatalf("MaintainBackups with schedule=off: %v", err)
}
if got := listBackupFiles(t, dir); len(got) != 0 {
t.Fatalf("schedule=off created files: %v", got)
}
// Daily → first tick creates exactly one scheduled backup.
mustSetSetting(t, database, "backup_schedule", "daily")
if err := admin.MaintainBackups(ctx, database); err != nil {
t.Fatalf("MaintainBackups daily #1: %v", err)
}
files := listBackupFiles(t, dir)
if len(files) != 1 || !strings.HasPrefix(files[0], "scheduled_") {
t.Fatalf("after first daily tick files = %v, want one scheduled_*.db", files)
}
first := filepath.Join(dir, files[0])
// Fresh backup on disk → next tick is a no-op.
if err := admin.MaintainBackups(ctx, database); err != nil {
t.Fatalf("MaintainBackups daily #2: %v", err)
}
if got := listBackupFiles(t, dir); len(got) != 1 {
t.Fatalf("fresh-backup tick changed files: %v", got)
}
// Backup older than a day (but inside retention) → a new one is taken and
// the old one is kept.
backdate(t, first, 25*time.Hour)
if err := admin.MaintainBackups(ctx, database); err != nil {
t.Fatalf("MaintainBackups daily #3: %v", err)
}
if got := listBackupFiles(t, dir); len(got) != 2 {
t.Fatalf("stale-backup tick files = %v, want 2", got)
}
// Old backup past the 7-day retention window → pruned; the fresh one stays.
backdate(t, first, 8*24*time.Hour)
if err := admin.MaintainBackups(ctx, database); err != nil {
t.Fatalf("MaintainBackups daily #4: %v", err)
}
got := listBackupFiles(t, dir)
if len(got) != 1 {
t.Fatalf("retention tick files = %v, want 1", got)
}
if filepath.Join(dir, got[0]) == first {
t.Fatalf("retention pruned the newest backup instead of the expired one")
}
}
// TestMaintainBackups_RetentionNeverDeletesNewest locks the safety rule: even
// when every backup is past retention, the newest survives.
func TestMaintainBackups_RetentionNeverDeletesNewest(t *testing.T) {
database := openAdminTestDB(t)
dir := t.TempDir()
admin.SetBackupBaseDir(dir)
t.Cleanup(func() { admin.SetBackupBaseDir(filepath.Join("data", "backups")) })
ctx := context.Background()
mustSetSetting(t, database, "backup_schedule", "off")
mustSetSetting(t, database, "backup_retention", "7")
// Two ancient backups, one slightly newer than the other.
older := filepath.Join(dir, "chatserver_a.db")
newer := filepath.Join(dir, "chatserver_b.db")
for _, p := range []string{older, newer} {
if err := os.WriteFile(p, []byte("x"), 0o600); err != nil {
t.Fatal(err)
}
}
backdate(t, older, 30*24*time.Hour)
backdate(t, newer, 20*24*time.Hour)
if err := admin.MaintainBackups(ctx, database); err != nil {
t.Fatalf("MaintainBackups: %v", err)
}
got := listBackupFiles(t, dir)
if len(got) != 1 || got[0] != "chatserver_b.db" {
t.Fatalf("files = %v, want only chatserver_b.db (newest kept)", got)
}
}
func mustSetSetting(t *testing.T, database interface {
SetSetting(ctx context.Context, key, value string) error
}, key, value string,
) {
t.Helper()
if err := database.SetSetting(context.Background(), key, value); err != nil {
t.Fatalf("SetSetting(%s): %v", key, err)
}
}
+12
View File
@@ -36,6 +36,18 @@ func SetSetupLimiterReapTiming(interval, maxWindow time.Duration) (restore func(
// at a temp dir. Lives here so it stays out of the production binary.
func SetBackupBaseDir(dir string) { backupBaseDir = dir }
// StubCopyBackup swaps the restore path's file-copy hook so tests can inject
// mid-copy failures that pass the pre-copy integrity gate. CopyBackupForTest
// is the real implementation, for stubs that only want to fail once.
func StubCopyBackup(fn func(src, dst string) error) (restore func()) {
prev := copyBackupFile
copyBackupFile = fn
return func() { copyBackupFile = prev }
}
// CopyBackupForTest exposes the real copyFile for StubCopyBackup delegates.
var CopyBackupForTest = copyFile
// StubCloseError makes the next handleRestoreBackup call's database.Close()
// return err instead of actually closing the pools, so tests can exercise the
// Close-failure branch without a genuine driver-level close error (see
+50 -10
View File
@@ -30,15 +30,30 @@ const (
// backupBaseDir is the directory for backup files, resolved to an absolute
// path at package init time so handlers don't depend on the process CWD (L14).
// Overridden at startup via SetBackupDir with cfg.Backup.Dir.
var backupBaseDir string
func init() {
abs, err := filepath.Abs(filepath.Join("data", "backups"))
if err == nil {
backupBaseDir = abs
} else {
backupBaseDir = filepath.Join("data", "backups")
backupBaseDir = absOrRaw(filepath.Join("data", "backups"))
}
// SetBackupDir points every backup handler and the scheduled-backup
// maintenance at the operator-configured directory. Call once at startup with
// cfg.Backup.Dir (main.go, next to SetDatabasePath); tests use it to isolate
// a temp dir. Mirrors SetDatabasePath: without it, a configured backup.dir
// would be ignored while backups keep landing in the default location.
func SetBackupDir(dir string) {
if dir == "" {
return
}
backupBaseDir = absOrRaw(dir)
}
func absOrRaw(p string) string {
if abs, err := filepath.Abs(p); err == nil {
return abs
}
return p
}
// dbFilePath is the live SQLite database file that "Restore backup"
@@ -73,12 +88,22 @@ func handleBackup(database *db.DB) http.Handler {
// Detached like the restore path's safety backup: an interrupted
// VACUUM INTO leaves a truncated .db that handleListBackups would
// present as restorable.
if err := database.BackupTo(context.WithoutCancel(r.Context()), backupPath); err != nil {
// present as restorable. BackupToSafe is rooted at the configured
// backup dir (SetBackupDir), not the historical hardcoded default.
if err := database.BackupToSafe(context.WithoutCancel(r.Context()), backupPath, backupDir); err != nil {
writeErr(w, http.StatusInternalServerError, "INTERNAL_ERROR", "backup failed")
return
}
// Verify before reporting success: a backup that fails integrity_check
// is worse than no backup, because the operator believes they have one.
if err := db.CheckBackupIntegrity(context.WithoutCancel(r.Context()), backupPath); err != nil {
slog.Error("backup failed integrity check — removing", "path", backupPath, "err", err)
_ = os.Remove(backupPath)
writeErr(w, http.StatusInternalServerError, "INTERNAL_ERROR", "backup failed verification")
return
}
actor := actorFromContext(r)
backupName := filepath.Base(backupPath)
slog.Info("database backup created", "actor_id", actor, "name", backupName)
@@ -191,6 +216,16 @@ func handleRestoreBackup(database *db.DB, hub HubBroadcaster) http.Handler {
return
}
// Refuse to overwrite the live database with a file SQLite itself
// rejects — a truncated pre-crash backup, a stray non-database .db.
// The pre-restore safety copy would make this survivable, but "restore
// succeeded" followed by a broken server is still the worst UX here.
if err := db.CheckBackupIntegrity(context.WithoutCancel(r.Context()), target); err != nil {
slog.Error("restore refused: backup failed integrity check", "backup", name, "err", err)
writeErr(w, http.StatusBadRequest, "BAD_REQUEST", "backup file failed integrity verification")
return
}
dbPath := dbFilePath
actor := actorFromContext(r)
@@ -215,7 +250,7 @@ func handleRestoreBackup(database *db.DB, hub HubBroadcaster) http.Handler {
// a server started from another working directory writes it somewhere
// the operator will never find it.
preRestore := filepath.Join(backupBaseDir, "pre_restore_"+time.Now().UTC().Format("20060102_150405")+".db")
if err := database.BackupTo(context.WithoutCancel(r.Context()), preRestore); err != nil {
if err := database.BackupToSafe(context.WithoutCancel(r.Context()), preRestore, backupBaseDir); err != nil {
// Fail closed. The admin panel promises "a pre-restore backup will
// be created" before an irreversible overwrite; proceeding without
// one takes away the safety net the operator was shown, exactly
@@ -254,7 +289,7 @@ func handleRestoreBackup(database *db.DB, hub HubBroadcaster) http.Handler {
// Stream the backup file over the (now closed) database to avoid loading
// the entire DB into memory (could be hundreds of MiB).
if err := copyFile(target, dbPath); err != nil {
if err := copyBackupFile(target, dbPath); err != nil {
// copyFile truncates the destination with os.Create before it can know
// whether the read will succeed, so the live database file is already
// destroyed by the time we get here — and the DB is closed, so nothing
@@ -262,7 +297,7 @@ func handleRestoreBackup(database *db.DB, hub HubBroadcaster) http.Handler {
// leaving the operator with a zero-byte database.
slog.Error("restore copy failed — rolling back to the pre-restore safety copy", "backup", name, "err", err)
msg := "failed to restore database file — the pre-restore safety copy was put back, server restarting"
if rbErr := copyFile(preRestore, dbPath); rbErr != nil {
if rbErr := copyBackupFile(preRestore, dbPath); rbErr != nil {
slog.Error("rollback from the pre-restore safety copy failed — recover manually",
"safety_copy", preRestore, "err", rbErr)
msg = "failed to restore database file AND failed to roll back — recover manually from " + filepath.Base(preRestore)
@@ -356,6 +391,11 @@ func restartProcess(reason string) {
os.Exit(0) //nolint:gocritic // backstop if the SIGTERM handler didn't exit
}
// copyBackupFile is the restore path's file-copy hook. It exists as a var so
// tests can inject the hard-to-simulate mid-copy failure (truncate-then-fail)
// the rollback branch exists for; production never swaps it.
var copyBackupFile = copyFile
// copyFile streams src to dst without loading the entire file into memory.
func copyFile(src, dst string) error {
in, err := os.Open(src) //nolint:gosec // G703: src is from sanitized backup path
+43 -13
View File
@@ -4,6 +4,7 @@ import (
"bytes"
"context"
"encoding/json"
"fmt"
"net/http"
"os"
"path/filepath"
@@ -267,12 +268,13 @@ func TestHandleRestoreBackup_Success(t *testing.T) {
t.Fatalf("MkdirAll data: %v", err)
}
// Write content as the "backup" to restore from.
// A real SQLite backup to restore from — the handler now verifies backups
// with integrity_check before touching the live database, so a text
// fixture would be (correctly) refused.
backupName := "chatserver_20240101_120000.db"
backupPath := filepath.Join(backupDir, backupName)
fakeContent := []byte("fake sqlite db content")
if err := os.WriteFile(backupPath, fakeContent, 0o644); err != nil {
t.Fatalf("WriteFile backup: %v", err)
if err := database.BackupToSafe(context.Background(), backupPath, backupDir); err != nil {
t.Fatalf("BackupToSafe fixture: %v", err)
}
restarted, restoreHook := admin.StubRestart()
@@ -369,11 +371,31 @@ func TestHandleRestoreBackup_RollsBackWhenCopyFails(t *testing.T) {
t.Fatalf("WriteFile live db: %v", err)
}
// A valid backup (it must pass the pre-copy integrity gate); the mid-copy
// failure is injected through the copy hook below, reproducing the exact
// failure mode the rollback exists for: os.Create truncates the live DB,
// then the copy dies.
backupName := "chatserver_20240102_120000.db"
if err := os.MkdirAll(filepath.Join(backupDir, backupName), 0o750); err != nil {
t.Fatalf("MkdirAll fake backup: %v", err)
if err := database.BackupToSafe(context.Background(), filepath.Join(backupDir, backupName), backupDir); err != nil {
t.Fatalf("BackupToSafe fixture: %v", err)
}
failedOnce := false
restoreCopy := admin.StubCopyBackup(func(src, dst string) error {
if !failedOnce {
failedOnce = true
// Truncate the destination the way the real copy's os.Create
// does, then fail — the state the rollback must repair.
f, createErr := os.Create(dst)
if createErr == nil {
_ = f.Close()
}
return fmt.Errorf("injected copy failure")
}
return admin.CopyBackupForTest(src, dst)
})
defer restoreCopy()
restarted, restoreHook := admin.StubRestart()
defer restoreHook()
@@ -442,9 +464,13 @@ func TestHandleRestoreBackup_RestartsWhenCloseFails(t *testing.T) {
if err := os.WriteFile(dbPath, []byte("original live contents"), 0o600); err != nil {
t.Fatalf("WriteFile live db: %v", err)
}
// A real SQLite backup — the restore handler verifies backups with
// integrity_check before touching the live database, so a text fixture
// would be (correctly) refused with 400 before the Close-failure branch
// under test is ever reached.
backupName := "chatserver_20240103_120000.db"
if err := os.WriteFile(filepath.Join(backupDir, backupName), []byte("replacement contents"), 0o644); err != nil {
t.Fatalf("WriteFile backup: %v", err)
if err := database.BackupToSafe(context.Background(), filepath.Join(backupDir, backupName), backupDir); err != nil {
t.Fatalf("BackupToSafe fixture: %v", err)
}
restarted, restoreRestartHook := admin.StubRestart()
@@ -483,8 +509,8 @@ func TestHandleRestoreBackup_AbortsWithoutSafetyBackup(t *testing.T) {
}
backupName := "chatserver_20240101_120000.db"
dbFile := filepath.Join(tmpDir, "data", "chatserver.db")
if err := os.WriteFile(filepath.Join(backupDir, backupName), []byte("replacement"), 0o644); err != nil {
t.Fatalf("WriteFile backup: %v", err)
if err := database.BackupToSafe(context.Background(), filepath.Join(backupDir, backupName), backupDir); err != nil {
t.Fatalf("BackupToSafe fixture: %v", err)
}
if err := os.WriteFile(dbFile, []byte("original"), 0o644); err != nil {
t.Fatalf("WriteFile db: %v", err)
@@ -551,9 +577,13 @@ func TestHandleRestoreBackup_UsesConfiguredDatabasePath(t *testing.T) {
t.Cleanup(func() { admin.SetDatabasePath(filepath.Join("data", "chatserver.db")) })
backupName := "chatserver_20240101_120000.db"
backupContent := []byte("restored contents")
if err := os.WriteFile(filepath.Join(backupDir, backupName), backupContent, 0o644); err != nil {
t.Fatalf("WriteFile backup: %v", err)
backupPath := filepath.Join(backupDir, backupName)
if err := database.BackupToSafe(context.Background(), backupPath, backupDir); err != nil {
t.Fatalf("BackupToSafe fixture: %v", err)
}
backupContent, err := os.ReadFile(backupPath)
if err != nil {
t.Fatalf("ReadFile fixture: %v", err)
}
restarted, restoreHook := admin.StubRestart()
+26 -2
View File
@@ -160,8 +160,20 @@ func handlePutChannelPermission(database *db.DB, hub HubBroadcaster, permInvalid
db.WriteAudit(context.WithoutCancel(r.Context()), database, actor, "channel_perms_update", "channel", ch.ID,
fmt.Sprintf("set overrides for role %s on #%s (allow=%#x deny=%#x)", role.Name, ch.Name, allow, deny))
// Narrow the eviction to users actually holding this role: a
// role-scoped override cannot change any other user's verdict, and
// InvalidateAll here made every connected user repopulate (2 reads
// each) synchronously inside RefreshChannelVisibility below — a
// whole-cache stampede that grows with total population, not with the
// role's size. Same rationale (and same fail-safe) as the role-perms
// handler: an unreadable member list falls back to the full flush,
// because a missed eviction is a stale grant.
if permInvalidator != nil {
permInvalidator.InvalidateAll()
if affected, listErr := database.ListUserIDsByRole(r.Context(), roleID); listErr == nil {
invalidateUsers(permInvalidator, affected)
} else {
permInvalidator.InvalidateAll()
}
}
if hub != nil {
hub.RefreshChannelVisibility(ch)
@@ -222,8 +234,20 @@ func handleDeleteChannelPermission(database *db.DB, hub HubBroadcaster, permInva
db.WriteAudit(context.WithoutCancel(r.Context()), database, actor, "channel_perms_clear", "channel", ch.ID,
fmt.Sprintf("cleared overrides for role %s on #%s", role.Name, ch.Name))
// Narrow the eviction to users actually holding this role: a
// role-scoped override cannot change any other user's verdict, and
// InvalidateAll here made every connected user repopulate (2 reads
// each) synchronously inside RefreshChannelVisibility below — a
// whole-cache stampede that grows with total population, not with the
// role's size. Same rationale (and same fail-safe) as the role-perms
// handler: an unreadable member list falls back to the full flush,
// because a missed eviction is a stale grant.
if permInvalidator != nil {
permInvalidator.InvalidateAll()
if affected, listErr := database.ListUserIDsByRole(r.Context(), roleID); listErr == nil {
invalidateUsers(permInvalidator, affected)
} else {
permInvalidator.InvalidateAll()
}
}
if hub != nil {
hub.RefreshChannelVisibility(ch)
+25 -4
View File
@@ -107,6 +107,13 @@ func TestPutChannelPermission_PersistsAndPropagates(t *testing.T) {
t.Fatalf("CreateChannel: %v", err)
}
// A member of the targeted role, so the narrowed invalidation has someone
// to evict. Users of other roles must NOT be evicted.
memberID, err := database.CreateUser(context.Background(), "role3member", "hash", 3)
if err != nil {
t.Fatalf("CreateUser: %v", err)
}
denyPrivate := permissions.ReadMessages | permissions.ConnectVoice
body := map[string]any{"allow": 0, "deny": denyPrivate}
w := doRequest(t, handler, http.MethodPut,
@@ -123,8 +130,14 @@ func TestPutChannelPermission_PersistsAndPropagates(t *testing.T) {
t.Errorf("persisted override = (%#x, %#x), want (0, %#x)", allow, deny, denyPrivate)
}
if inv.invalidateAllN != 1 {
t.Errorf("InvalidateAll calls = %d, want 1", inv.invalidateAllN)
// The eviction is narrowed to the targeted role's members — a role-scoped
// override cannot change any other user's verdict, so the whole-cache
// flush (and its repopulate stampede) is reserved for the fail-safe path.
if inv.invalidateAllN != 0 {
t.Errorf("InvalidateAll calls = %d, want 0 (narrowed invalidation)", inv.invalidateAllN)
}
if len(inv.invalidateUserIDs) != 1 || inv.invalidateUserIDs[0] != memberID {
t.Errorf("InvalidateUser calls = %v, want exactly [%d]", inv.invalidateUserIDs, memberID)
}
if len(hub.visibilityRefreshes) != 1 || hub.visibilityRefreshes[0].ID != chID {
t.Errorf("RefreshChannelVisibility not called for channel %d", chID)
@@ -317,6 +330,10 @@ func TestDeleteChannelPermission_ClearsOverride(t *testing.T) {
if err := database.UpsertChannelOverride(context.Background(), chID, 3, 0, permissions.ReadMessages); err != nil {
t.Fatalf("UpsertChannelOverride: %v", err)
}
memberID, err := database.CreateUser(context.Background(), "role3clear", "hash", 3)
if err != nil {
t.Fatalf("CreateUser: %v", err)
}
w := doRequest(t, handler, http.MethodDelete,
"/channels/"+itoa(chID)+"/permissions/3", token, nil)
@@ -331,8 +348,12 @@ func TestDeleteChannelPermission_ClearsOverride(t *testing.T) {
if allow != 0 || deny != 0 {
t.Errorf("override still present: (%#x, %#x)", allow, deny)
}
if inv.invalidateAllN != 1 {
t.Errorf("InvalidateAll calls = %d, want 1", inv.invalidateAllN)
// Narrowed invalidation: only the targeted role's members are evicted.
if inv.invalidateAllN != 0 {
t.Errorf("InvalidateAll calls = %d, want 0 (narrowed invalidation)", inv.invalidateAllN)
}
if len(inv.invalidateUserIDs) != 1 || inv.invalidateUserIDs[0] != memberID {
t.Errorf("InvalidateUser calls = %v, want exactly [%d]", inv.invalidateUserIDs, memberID)
}
if len(hub.visibilityRefreshes) != 1 {
t.Errorf("RefreshChannelVisibility calls = %d, want 1", len(hub.visibilityRefreshes))
+3 -3
View File
@@ -1568,12 +1568,12 @@ async function renderSettings(){
let html='<div class="page-title">Server Settings</div><div class="page-desc">Configure your OwnCord server</div>';
html+='<div class="section-card"><div class="section-card-header"><h3>General</h3></div><div class="section-card-body">';
html+='<div class="setting-row"><div class="setting-info"><div class="setting-name">Server Name</div></div><div class="setting-ctrl"><input class="form-input" id="s-server_name" value="'+esc(v('server_name'))+'" style="width:240px" oninput="markSettingsChanged()"></div></div>';
html+='<div class="setting-row"><div class="setting-info"><div class="setting-name">Server Icon URL</div></div><div class="setting-ctrl"><input class="form-input" id="s-server_icon" value="'+esc(v('server_icon'))+'" style="width:240px" oninput="markSettingsChanged()"></div></div>';
html+='<div class="setting-row"><div class="setting-info"><div class="setting-name">Server Icon URL</div><div class="setting-desc">Not used by the server or client yet &mdash; stored for a future release</div></div><div class="setting-ctrl"><input class="form-input" id="s-server_icon" value="'+esc(v('server_icon'))+'" style="width:240px" disabled title="Not implemented yet"></div></div>';
html+='<div class="setting-row"><div class="setting-info"><div class="setting-name">Message of the Day</div><div class="setting-desc">Shown to users when they connect</div></div><div class="setting-ctrl"><input class="form-input" id="s-motd" value="'+esc(v('motd'))+'" style="width:300px" oninput="markSettingsChanged()"></div></div>';
html+='</div></div>';
html+='<div class="section-card"><div class="section-card-header"><h3>Limits</h3></div><div class="section-card-body">';
html+='<div class="setting-row"><div class="setting-info"><div class="setting-name">Max Upload Size (bytes)</div></div><div class="setting-ctrl"><input class="form-input" id="s-max_upload_bytes" value="'+esc(v('max_upload_bytes'))+'" style="width:160px" type="number" oninput="markSettingsChanged()"></div></div>';
html+='<div class="setting-row"><div class="setting-info"><div class="setting-name">Voice Quality</div></div><div class="setting-ctrl"><select class="filter-select" id="s-voice_quality" onchange="markSettingsChanged()"><option value="low" '+(v('voice_quality')==='low'?'selected':'')+'>Low</option><option value="medium" '+(v('voice_quality')==='medium'?'selected':'')+'>Medium</option><option value="high" '+(v('voice_quality')==='high'?'selected':'')+'>High</option></select></div></div>';
html+='<div class="setting-row"><div class="setting-info"><div class="setting-name">Max Upload Size (bytes)</div><div class="setting-desc">Controlled by upload.max_size_mb in config.yaml (requires restart) &mdash; this display value has no effect</div></div><div class="setting-ctrl"><input class="form-input" id="s-max_upload_bytes" value="'+esc(v('max_upload_bytes'))+'" style="width:160px" type="number" disabled title="Set upload.max_size_mb in config.yaml and restart"></div></div>';
html+='<div class="setting-row"><div class="setting-info"><div class="setting-name">Voice Quality</div><div class="setting-desc">Controlled by voice.quality in config.yaml (requires restart) &mdash; this display value has no effect</div></div><div class="setting-ctrl"><select class="filter-select" id="s-voice_quality" disabled title="Set voice.quality in config.yaml and restart"><option value="low" '+(v('voice_quality')==='low'?'selected':'')+'>Low</option><option value="medium" '+(v('voice_quality')==='medium'?'selected':'')+'>Medium</option><option value="high" '+(v('voice_quality')==='high'?'selected':'')+'>High</option></select></div></div>';
html+='</div></div>';
html+='<div class="section-card"><div class="section-card-header"><h3>Security</h3></div><div class="section-card-body">';
html+='<div class="setting-row"><div class="setting-info"><div class="setting-name">Require 2FA</div><div class="setting-desc">Require all users to enable two-factor authentication</div></div><div class="setting-ctrl"><button class="toggle '+(isOn('require_2fa')?'on':'')+'" id="s-require_2fa" onclick="this.classList.toggle(\'on\');markSettingsChanged()"></button></div></div>';