mirror of
https://github.com/J3vb/OwnCord.git
synced 2026-09-03 03:50:00 +03:00
* docs: add infrastructure roadmap plan Records the verified recommendations from an infrastructure review in three tracks: raising the single-instance ceiling, cheap seams for a possible multi-instance future, and ops hygiene. Includes explicit anti-recommendations and sequencing. Security-sensitive detail is intentionally excluded per docs/security.md. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj * feat(server): real health checks and saturation metrics /api/v1/metrics now exposes signals that were already computed in memory but never surfaced: reconnect replay tier hits, event-persister counters, SQLite writer-pool wait stats, aggregate per-client backpressure counters (including previously invisible low-priority drops), and permission-cache hit/miss. /health now returns a real verdict: hub dispatch-loop liveness, a bounded database ping, and a free-disk check, returning 503 with a subsystem reason when degraded. Checks are cached so the unauthenticated endpoint cannot amplify load. The hub's panic breaker now exits the process so a supervisor can restart it, instead of leaving broadcast delivery silently dead while clients still appear online. OTel instruments that were declared but never recorded are now wired (ws_active_connections, ws_broadcast_latency_seconds, ws_messages_total, ws_events_dropped_total, voice gauges) or removed (db_query_duration_seconds). Also corrects the docs/api.md description of broadcast_drops, which counts hub-queue overflow, not client send-queue overflow. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj * feat(server): implement scheduled backups, retention, and backup verification The backup_schedule and backup_retention settings have existed in the admin panel and API since the initial schema but were never read by any code. The 15-minute maintenance loop now enforces them: a scheduled backup is taken when the newest backup on disk is older than the schedule interval (manual backups reset the clock), and retention prunes backups older than the configured days while always keeping the newest one. Backups are now verified with PRAGMA integrity_check immediately after VACUUM INTO (a failed backup is removed rather than listed as restorable) and again before a restore may overwrite the live database. A failed VACUUM INTO also cleans up its partial output file — but never a pre-existing one. The backup directory is configurable via a new backup.dir key (default data/backups) so operators can point backups at another disk or an off-host mount, mirroring the SetDatabasePath plumb. Restore-handler tests now use real SQLite fixtures (the integrity gate correctly refuses text files) with the mid-copy failure injected through a test-only copy hook. Also adds audited gosec suppressions to the Windows disk-free syscall added in the previous commit, which the Windows lint leg flagged. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj * feat(server): capacity and failure-mode guardrails - server.max_ws_connections: optional cap on concurrent WebSocket clients, checked before the upgrade with a 503 + Retry-After; rejections are counted and exposed as ws_conn_rejects in /api/v1/metrics. - Single-process database lock: an OS-level advisory lock (flock / exclusive handle) beside the SQLite file makes a second server process fail fast with a clear message instead of silently fighting the first over process-local state. A bounded retry covers the self-update/restore restart handoff, and the lock mechanism failing (e.g. network filesystems) only warns. - Disk-space awareness: boot-time warnings for the data and backup volumes, plus a disk_free_mb metrics field, via a small cross-platform diskutil package (already used by /health). - Upload storage failures: storage.Save now marks server-side filesystem failures with a sentinel (storage.ErrIO); handlers return 507 for those instead of blaming the client with a 400, and the emoji route stops echoing raw storage errors (which embed absolute paths) into responses. - Unknown config keys now warn at startup — a typo like admin_alowed_cidrs previously kept the default silently while the operator believed the setting changed. Never fatal: newer servers tolerate older configs. - Admin settings honesty: the three stored-but-inert settings (server_icon, max_upload_bytes, voice_quality) are shown read-only with a note pointing at the real config.yaml keys, instead of pretending to apply. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj * perf(db): write-path efficiency and capacity knobs - channel_focus/mark_read now skip the read-state UPSERT when the stored row already matches (same last_message_id, no mentions) — refocus events fire at up to 10/s/user and every no-op write still occupied the single SQLite writer connection. The extra existence check runs on the reader pool, which doesn't serialize. Same shape as the session-touch throttle. - DeleteExpiredSessions is now sargable: migration 031 normalizes legacy expiry formats to the RFC3339-Z layout the server writes and indexes expires_at, replacing the strftime full-table scan that ran on the writer every 15 minutes. - Boot-time ANALYZE runs only when a migration actually applied; unchanged schemas get the cheap PRAGMA optimize instead (which also covers crash-restarts that never reached the shutdown optimize). - The read/write SQL router gets a table-driven test with explicit expected values (INSERT ... RETURNING must hit the writer despite being :one). - New knobs, all defaulting to current behavior: database.max_readers, security.auth_rate_limit_multiplier (for shared-NAT communities), event_persistence.replay_ring_size and replay_cold_limit. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj * fix(server): shutdown lifecycle ordering - The event pruner and maintenance loop are now joined (bounded) before the database closes: bgCtx cancellation used to run AFTER database.Close via LIFO defers, contradicting its own comment, and neither goroutine was ever waited on — a mid-tick scheduled backup or prune could still hold the writer while the pool tore down. StartEventPruner returns a done channel with the same join contract EventPersister.Stop already had. - srv.Shutdown now runs before hub.GracefulStop, so in-flight HTTP handlers' broadcasts still reach a live hub and the event persister instead of vanishing from the replay/event store across a restart. Shutdown does not wait on hijacked WebSocket connections, so the swap adds no delay. - GracefulStopContext threads the 30s shutdown budget into the hub: the 5s client-notice window (matching the countdown clients are shown) ends early when the budget expires, and is skipped entirely when nobody is connected — early-return startup paths and idle servers no longer sleep 5s for an audience of zero. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj * build(deploy): systemd unit, compose hardening, boot-smoked releases, CI polish - deploy/owncord.service: hardened systemd unit template with the two verified caveats encoded (install dir stays writable for self-update under ProtectSystem=strict; CAP_NET_BIND_SERVICE for ACME's :80), plus a 'Linux (systemd)' deployment docs section — the Linux service story was previously 'Docker or nothing'. - New 'Reverse Proxy Topology' docs section with a working nginx snippet and the correct signaling-vs-media distinction: /livekit/* is already proxied by the server, only WebRTC media ports must be directly reachable. - docker-compose: log rotation, commented resource limits, and a healthcheck backed by a new 'chatserver healthcheck' subcommand (the distroless image has no shell) that probes /health without config side effects. - release.yml: a concurrency group (queue, never cancel), and boot-smoke gates — the freshly built server binaries and the Docker image are cold booted and probed healthy BEFORE anything is signed or pushed. The release feed drives signed self-updates, so a binary that compiles but dies on boot previously would have shipped itself to every auto-updating instance. - ci.yml: client-check/client-tests move to ubuntu with the reasoning recorded (no win32 code paths, LF enforced repo-wide); admin-e2e gets a written graduation criterion instead of an open-ended non-blocking status. - docs: Tailscale guide notes the CGNAT range vs the default admin CIDRs; architecture overview records presence/voice state as the fifth single-instance blocker and the macOS client scope decision. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj * perf(server): measured load tooling, narrowed invalidation, presence coalescing, storage and CIDR seams - Fix scripts/k6/ws-load.js against the real wire protocol: envelope-wrapped frames, correct message types (typing_start, presence_update), the correct /api/v1/ws path, and thresholds that fail a run where nobody authenticated or went ready — the script had drifted to pre-envelope framing and reported 100% green while every auth failed on the first frame. A new workflow_dispatch-only load-baseline workflow boots a real server, seeds users through the setup/invite APIs, runs the script, and uploads the k6 summary plus a metrics snapshot for before/after comparison. - Role-scoped channel-override changes now evict only the affected role's members from the permission cache (fail-safe: unreadable member list still flushes everything). InvalidateAll here repopulated every connected user — two reads each — synchronously inside the admin request via RefreshChannelVisibility, a stampede that scaled with total population rather than the role's size. Same pattern the per-user override endpoints already used. - Connect/disconnect presence broadcasts now pass through a 300ms latest-wins coalescer (QueuePresence): each un-coalesced presence change is a sequenced global broadcast (an O(clients) fan-out under seqMu), so a reconnect storm fired O(users) of them from the connect critical path. A flap inside the window collapses to its final state; the wire format, seq ordering, and replay behaviour are unchanged, and the delivery path (BroadcastPresence) is untouched. - Storage seam: api handlers now consume a FileStore interface (consumer-side, same pattern as service.Store) with Open returning a seekable storage.File — writing down the contract (range-request seeks included) an alternative backend would have to meet, without building one. - The metrics surfaces and the LiveKit webhook/health endpoints get their own allowlist keys (metrics_allowed_cidrs, livekit_webhook_allowed_cidrs, both defaulting to admin_allowed_cidrs), so a central Prometheus scraper or an externally-hosted LiveKit no longer requires widening the admin panel's perimeter. Startup now also warns when admin_allowed_cidrs is customized while trusted_proxies is empty — behind a proxy or container network the check would otherwise compare the proxy's private address, not the client's. - The container healthcheck probe now PINS the server's own certificate from disk (VerifyConnection, exact-match) instead of skipping TLS verification, addressing the CodeQL finding on the previous commit; WebPKI verification is used when no local cert exists (ACME). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj * fix(server): address self-review findings on the hardening branch Seven fixes from a high-effort review of the full branch diff: - healthcheck CLI now works under tls.mode acme: it overrides ServerName with the configured domain for WebPKI verification instead of pinning a cert that doesn't exist (or is stale) in that mode. Previously an ACME deployment's container healthcheck failed forever. - /health pings the READER pool (new db.PingRead): the writer ping queued behind a scheduled backup's VACUUM INTO and reported the server degraded for the whole backup — which an autoheal watchdog would turn into a nightly mid-backup restart. - /health runs its cached checks under context.WithoutCancel so a probe that disconnects mid-request cannot poison the shared cache with a false degraded verdict for the next 5 seconds. - The token CLI uses a new db.OpenShared that skips the single-process lock: minting a token against a running server is safe under WAL and was a documented workflow the lock had broken. - The per-user TOTP failure cap is no longer scaled by security.auth_rate_limit_multiplier — that knob exists for per-IP limits; scaling the only cross-IP brute-force defence multiplied an attacker's distributed guess budget. Mirrors the unscaled per-user login threshold. - A direct presence_update now drops the user's queued entry in the connect/disconnect coalescer, so a stale connect-time presence can no longer flush 300ms later over the user's fresher chosen status. - The scheduled-backup filename collision loop breaks on any stat error and bounds its suffix probing, instead of spinning the maintenance goroutine forever on a persistent EACCES. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj * test(admin): real SQLite fixture for the merged Close-failure restore test TestHandleRestoreBackup_RestartsWhenCloseFails arrived from main (#1375) with a plain-text backup fixture; this branch's restore handler verifies backups with integrity_check before touching the live database, so the text fixture was (correctly) refused with 400 before the Close-failure branch under test was reached. Use a real backup via BackupToSafe, matching the other restore tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj --------- Co-authored-by: Claude <noreply@anthropic.com>
449 lines
18 KiB
Go
449 lines
18 KiB
Go
// Package db provides database access for the OwnCord server.
|
|
// It uses modernc.org/sqlite — a pure-Go SQLite driver requiring no CGO.
|
|
package db
|
|
|
|
import (
|
|
"context"
|
|
"database/sql"
|
|
"errors"
|
|
"fmt"
|
|
"log/slog"
|
|
"runtime"
|
|
"strings"
|
|
"sync/atomic"
|
|
|
|
"github.com/owncord/server/db/dbgen"
|
|
"github.com/owncord/server/migrations"
|
|
_ "modernc.org/sqlite" // register the sqlite3 driver
|
|
)
|
|
|
|
// DB wraps the underlying SQLite pools and exposes the subset of methods
|
|
// needed by the server.
|
|
//
|
|
// q is the sqlc-generated query layer (db/dbgen). Query method bodies delegate
|
|
// to it — sqlc is the source of truth for the SQL text and parameter binding
|
|
// (verified in CI by `make sqlc-verify`), while this package keeps the stable
|
|
// public API and the domain model types the rest of the server consumes.
|
|
// Migration is incremental (decision D2); methods not yet delegated still run
|
|
// their raw SQL directly against writer/reader.
|
|
type DB struct {
|
|
// writer is a single-connection pool that owns every statement that can
|
|
// mutate the database: INSERT/UPDATE/DELETE (including RETURNING forms),
|
|
// transactions, migrations, ANALYZE/VACUUM and PRAGMA writes. Pinning
|
|
// writes to one connection makes concurrent writers queue on the Go side
|
|
// instead of colliding on SQLite's single write lock.
|
|
writer *sql.DB
|
|
|
|
// reader is a multi-connection pool serving read-only statements
|
|
// (SELECT / PRAGMA reads). Under WAL, readers run concurrently with each
|
|
// other and with the writer, which is the point of the split. For
|
|
// in-memory databases reader and writer are the same handle.
|
|
reader *sql.DB
|
|
|
|
q *dbgen.Queries
|
|
|
|
// auditWriter, when installed via SetAuditWriter (main.go server
|
|
// startup only), turns WriteAudit calls backed by this DB into
|
|
// non-blocking enqueues. Nil (the default) keeps audit writes
|
|
// synchronous — the token CLI and tests rely on that.
|
|
auditWriter atomic.Pointer[AuditWriter]
|
|
|
|
// lockRelease drops the single-process advisory lock taken by openFile.
|
|
// Nil for in-memory databases and when the lock mechanism is unavailable.
|
|
lockRelease func()
|
|
}
|
|
|
|
// filePragmas are the per-connection PRAGMAs applied to every file-backed
|
|
// connection via `_pragma=` DSN parameters (modernc.org/sqlite executes each
|
|
// one in newConn, busy_timeout first). They MUST be in the DSN rather than
|
|
// Exec'd after Open: with a pool larger than one connection an Exec'd PRAGMA
|
|
// lands on one arbitrary connection and every other connection would silently
|
|
// run with foreign_keys=OFF.
|
|
const filePragmas = "_pragma=busy_timeout(5000)" + // wait up to 5s for the write lock instead of failing instantly
|
|
"&_pragma=journal_mode(WAL)" + // WAL: readers don't block the writer and vice versa
|
|
"&_pragma=foreign_keys(1)" + // enforce foreign key constraints
|
|
"&_pragma=synchronous(NORMAL)" + // performance tuning, safe with WAL
|
|
"&_pragma=temp_store(MEMORY)" +
|
|
"&_pragma=mmap_size(268435456)" +
|
|
"&_pragma=cache_size(-64000)"
|
|
|
|
// sqliteTimeLayout is how SQLite's own datetime('now') writes a timestamp, and
|
|
// therefore the only shape a Go-side cutoff may take when it is compared against
|
|
// such a column: the comparison is bytewise TEXT, so RFC3339's 'T' separator
|
|
// sorts after the space and quietly turns "older than X" into "any earlier date".
|
|
const sqliteTimeLayout = "2006-01-02 15:04:05"
|
|
|
|
// isMemoryPath reports whether path names an in-memory database
|
|
// (":memory:", "file::memory:" or any URI carrying mode=memory).
|
|
func isMemoryPath(path string) bool {
|
|
return strings.Contains(path, ":memory:") || strings.Contains(path, "mode=memory")
|
|
}
|
|
|
|
// Open opens (or creates) a SQLite database at path, enables WAL mode and
|
|
// foreign key enforcement, and returns a ready-to-use DB.
|
|
//
|
|
// Two modes:
|
|
//
|
|
// - In-memory databases keep the historical single-handle behavior: one
|
|
// *sql.DB pinned to a single connection with the PRAGMAs Exec'd once.
|
|
// A one-connection pool makes DSN PRAGMAs unnecessary, all callers share
|
|
// the same in-memory state, and connection-scoped PRAGMA toggles in tests
|
|
// (e.g. temporarily disabling foreign_keys) behave deterministically.
|
|
// In this mode reader and writer are the same handle.
|
|
//
|
|
// - File-backed databases get a reader/writer pool split. Both pools carry
|
|
// the PRAGMAs in the DSN so every physical connection is configured
|
|
// identically. The writer is additionally opened with _txlock=immediate
|
|
// so explicit transactions take the write lock up front instead of
|
|
// failing with SQLITE_BUSY on upgrade.
|
|
//
|
|
// Path assumptions (file mode): path is either a plain filesystem path or an
|
|
// existing file: URI. It must not contain '?', '#' or '%' characters — the
|
|
// path is embedded in a file: URI without escaping. cfg.Database.Path is a
|
|
// plain path (default "data/chatserver.db"), which satisfies this.
|
|
func Open(path string) (*DB, error) {
|
|
return OpenWithMaxReaders(path, 0)
|
|
}
|
|
|
|
// OpenWithMaxReaders is Open with an explicit reader-pool bound
|
|
// (database.max_readers). maxReaders <= 0 keeps the automatic
|
|
// max(4, NumCPU) sizing; values are clamped to [1, 64]. Ignored for
|
|
// in-memory databases, which use a single shared connection.
|
|
func OpenWithMaxReaders(path string, maxReaders int) (*DB, error) {
|
|
if isMemoryPath(path) {
|
|
return openMemory(path)
|
|
}
|
|
return openFile(path, maxReaders, true)
|
|
}
|
|
|
|
// OpenShared opens the database WITHOUT taking the single-process lock. It
|
|
// exists for short-lived tooling — the `server token` CLI — that must work
|
|
// while the server is running. SQLite's own WAL locking makes the concurrent
|
|
// access safe at the file level; the process lock only protects the SERVER's
|
|
// process-local state (presence, replay ring, rate-limit windows), which a
|
|
// CLI does not touch. Long-lived processes must use Open.
|
|
func OpenShared(path string) (*DB, error) {
|
|
if isMemoryPath(path) {
|
|
return openMemory(path)
|
|
}
|
|
return openFile(path, 0, false)
|
|
}
|
|
|
|
// openMemory preserves the pre-split behavior exactly: a single connection
|
|
// with PRAGMAs applied by Exec. reader == writer.
|
|
func openMemory(path string) (*DB, error) {
|
|
sqlDB, err := sql.Open("sqlite", path)
|
|
if err != nil {
|
|
return nil, fmt.Errorf("opening sqlite db: %w", err)
|
|
}
|
|
|
|
// Verify the connection is actually usable.
|
|
if err := sqlDB.Ping(); err != nil {
|
|
_ = sqlDB.Close()
|
|
return nil, fmt.Errorf("pinging sqlite db: %w", err)
|
|
}
|
|
|
|
// A single connection ensures all callers share the same in-memory state
|
|
// and makes the Exec'd PRAGMAs below apply to every future statement.
|
|
sqlDB.SetMaxOpenConns(1)
|
|
|
|
// The same PRAGMA set as filePragmas, Exec'd because there is exactly one
|
|
// connection to configure. journal_mode is a no-op for :memory: (SQLite
|
|
// reports "memory") but is kept for symmetry with file mode.
|
|
for _, p := range []struct{ name, stmt string }{
|
|
{"enabling WAL mode", "PRAGMA journal_mode=WAL;"},
|
|
{"setting busy_timeout", "PRAGMA busy_timeout=5000;"},
|
|
{"enabling foreign keys", "PRAGMA foreign_keys=ON;"},
|
|
{"setting synchronous mode", "PRAGMA synchronous=NORMAL;"},
|
|
{"setting temp_store", "PRAGMA temp_store=MEMORY;"},
|
|
{"setting mmap_size", "PRAGMA mmap_size=268435456;"},
|
|
{"setting cache_size", "PRAGMA cache_size=-64000;"},
|
|
} {
|
|
if _, err := sqlDB.Exec(p.stmt); err != nil {
|
|
_ = sqlDB.Close()
|
|
return nil, fmt.Errorf("%s: %w", p.name, err)
|
|
}
|
|
}
|
|
|
|
return newDB(sqlDB, sqlDB), nil
|
|
}
|
|
|
|
// openFile opens the writer and reader pools for a file-backed database.
|
|
// takeLock is false only for OpenShared (short-lived CLI tooling).
|
|
func openFile(path string, maxReaders int, takeLock bool) (*DB, error) {
|
|
// Single-process guard: SQLite's own locking prevents file corruption,
|
|
// but everything built above it — presence derived from hub membership,
|
|
// the replay ring, rate-limit windows, the boot-time status reset — is
|
|
// process-local and assumes exactly one server owns this database file.
|
|
// A second process starting is almost always an accident (double systemd
|
|
// unit, container + binary); fail fast with a clear message rather than
|
|
// letting two instances silently fight over shared state.
|
|
var release func()
|
|
if takeLock {
|
|
var lockErr error
|
|
release, lockErr = acquireProcessLock(path)
|
|
if lockErr != nil {
|
|
if errors.Is(lockErr, errAlreadyLocked) {
|
|
return nil, fmt.Errorf(
|
|
"database %s is in use by another running OwnCord process — stop that process first (the lock is released automatically when it exits)",
|
|
path)
|
|
}
|
|
// The lock mechanism itself failed (e.g. a network filesystem that
|
|
// rejects advisory locks). Warn and continue — refusing to start on
|
|
// an NFS data dir would be a regression, and SQLite still protects
|
|
// the file itself.
|
|
slog.Warn("db: could not take the single-process lock; continuing unprotected",
|
|
"path", lockFilePath(path), "error", lockErr)
|
|
release = nil
|
|
}
|
|
}
|
|
ok := false
|
|
defer func() {
|
|
if !ok && release != nil {
|
|
release()
|
|
}
|
|
}()
|
|
|
|
base := path
|
|
if !strings.HasPrefix(base, "file:") {
|
|
base = "file:" + base
|
|
}
|
|
sep := "?"
|
|
if strings.Contains(base, "?") {
|
|
sep = "&"
|
|
}
|
|
dsn := base + sep + filePragmas
|
|
|
|
// Writer: one connection so concurrent writes queue on the Go side, with
|
|
// BEGIN IMMEDIATE transactions (see Open doc).
|
|
writer, err := sql.Open("sqlite", dsn+"&_txlock=immediate")
|
|
if err != nil {
|
|
return nil, fmt.Errorf("opening sqlite db: %w", err)
|
|
}
|
|
writer.SetMaxOpenConns(1)
|
|
|
|
// Verify the connection is actually usable (this also creates the file,
|
|
// so the reader below never races file creation).
|
|
if err := writer.Ping(); err != nil {
|
|
_ = writer.Close()
|
|
return nil, fmt.Errorf("pinging sqlite db: %w", err)
|
|
}
|
|
|
|
// Reader: sized for concurrent request handling. Idle == open so warm
|
|
// connections (and their page caches) are kept rather than churned.
|
|
reader, err := sql.Open("sqlite", dsn)
|
|
if err != nil {
|
|
_ = writer.Close()
|
|
return nil, fmt.Errorf("opening sqlite reader pool: %w", err)
|
|
}
|
|
readConns := max(4, runtime.NumCPU())
|
|
if maxReaders > 0 {
|
|
readConns = min(max(maxReaders, 1), 64)
|
|
}
|
|
reader.SetMaxOpenConns(readConns)
|
|
reader.SetMaxIdleConns(readConns)
|
|
if err := reader.Ping(); err != nil {
|
|
_ = reader.Close()
|
|
_ = writer.Close()
|
|
return nil, fmt.Errorf("pinging sqlite reader pool: %w", err)
|
|
}
|
|
|
|
d := newDB(writer, reader)
|
|
d.lockRelease = release
|
|
ok = true
|
|
return d, nil
|
|
}
|
|
|
|
// newDB assembles a DB whose sqlc query layer routes through dbtx.
|
|
func newDB(writer, reader *sql.DB) *DB {
|
|
return &DB{
|
|
writer: writer,
|
|
reader: reader,
|
|
q: dbgen.New(&dbtx{writer: writer, reader: reader}),
|
|
}
|
|
}
|
|
|
|
// dbtx routes statements between the reader and writer pools. It satisfies
|
|
// sqlc's DBTX interface (db/dbgen/db.go) so the generated query layer picks
|
|
// the correct pool per statement without touching generated code, and it
|
|
// backs the DB.QueryContext/QueryRowContext wrappers for the same reason.
|
|
//
|
|
// Routing table:
|
|
//
|
|
// - ExecContext → writer.
|
|
// - PrepareContext → writer. sqlc's generated code in this repo prepares
|
|
// nothing through DBTX (only *sql.Tx.PrepareContext inside the batch
|
|
// persisters, which already run on writer transactions); this method
|
|
// exists for interface completeness. Tradeoff: if a future caller
|
|
// prepared a hot SELECT here it would run serialized on the writer.
|
|
// - QueryContext / QueryRowContext → reader, but only when the statement
|
|
// is provably read-only (isReadOnlySQL). sqlc routes
|
|
// INSERT/UPDATE/DELETE ... RETURNING statements through
|
|
// QueryRowContext/QueryContext (messages, attachments), and those writes
|
|
// must stay on the single writer connection; anything not provably
|
|
// read-only conservatively falls back to the writer.
|
|
//
|
|
// For in-memory databases writer == reader, so routing degenerates to the
|
|
// historical single-handle behavior.
|
|
type dbtx struct {
|
|
writer *sql.DB
|
|
reader *sql.DB
|
|
}
|
|
|
|
func (t *dbtx) ExecContext(ctx context.Context, query string, args ...any) (sql.Result, error) {
|
|
return t.writer.ExecContext(ctx, query, args...)
|
|
}
|
|
|
|
func (t *dbtx) PrepareContext(ctx context.Context, query string) (*sql.Stmt, error) {
|
|
return t.writer.PrepareContext(ctx, query)
|
|
}
|
|
|
|
func (t *dbtx) QueryContext(ctx context.Context, query string, args ...any) (*sql.Rows, error) {
|
|
return t.pool(query).QueryContext(ctx, query, args...)
|
|
}
|
|
|
|
func (t *dbtx) QueryRowContext(ctx context.Context, query string, args ...any) *sql.Row {
|
|
return t.pool(query).QueryRowContext(ctx, query, args...)
|
|
}
|
|
|
|
// pool selects the reader for read-only statements and the writer otherwise.
|
|
func (t *dbtx) pool(query string) *sql.DB {
|
|
if isReadOnlySQL(query) {
|
|
return t.reader
|
|
}
|
|
return t.writer
|
|
}
|
|
|
|
// isReadOnlySQL reports whether the statement's leading keyword — after
|
|
// skipping whitespace and `--` line comments — is SELECT or PRAGMA. Comment
|
|
// skipping matters because sqlc-generated SQL starts with a `-- name: ...`
|
|
// line. Anything unrecognized (including `/* */` block comments, which this
|
|
// package does not use) is conservatively treated as a write.
|
|
func isReadOnlySQL(query string) bool {
|
|
s := query
|
|
for {
|
|
s = strings.TrimLeft(s, " \t\r\n")
|
|
if !strings.HasPrefix(s, "--") {
|
|
break
|
|
}
|
|
nl := strings.IndexByte(s, '\n')
|
|
if nl < 0 {
|
|
return false // comment-only "statement" — let the writer reject it
|
|
}
|
|
s = s[nl+1:]
|
|
}
|
|
return hasKeywordPrefix(s, "SELECT") || hasKeywordPrefix(s, "PRAGMA")
|
|
}
|
|
|
|
// hasKeywordPrefix reports whether s starts with the keyword (ASCII
|
|
// case-insensitive) followed by a non-identifier character or end of input.
|
|
func hasKeywordPrefix(s, keyword string) bool {
|
|
if len(s) < len(keyword) || !strings.EqualFold(s[:len(keyword)], keyword) {
|
|
return false
|
|
}
|
|
if len(s) == len(keyword) {
|
|
return true
|
|
}
|
|
c := s[len(keyword)]
|
|
return c != '_' && (c < '0' || c > '9') && (c < 'a' || c > 'z') && (c < 'A' || c > 'Z')
|
|
}
|
|
|
|
// Migrate runs all SQL migration files from the embedded migrations FS in
|
|
// lexicographic order, applying each file exactly once. It delegates to
|
|
// MigrateFS (defined in migrate.go) which maintains the schema_versions
|
|
// tracking table.
|
|
func Migrate(database *DB) error {
|
|
applied, err := migrateFSCount(database, migrations.FS)
|
|
if err != nil {
|
|
return err
|
|
}
|
|
// Refresh the query planner's statistics when the schema changed, so newly
|
|
// created indexes (e.g. migration 019) are actually chosen. A full ANALYZE
|
|
// grows with row count and blocks startup, so unchanged schemas get the
|
|
// cheap PRAGMA optimize instead — which also covers crash-restarts that
|
|
// never reached Close()'s optimize. Both write sqlite_stat rows, so they
|
|
// run on the writer.
|
|
if applied > 0 {
|
|
if _, err := database.writer.Exec("ANALYZE;"); err != nil {
|
|
return fmt.Errorf("running ANALYZE after migrations: %w", err)
|
|
}
|
|
return nil
|
|
}
|
|
if _, err := database.writer.Exec("PRAGMA optimize;"); err != nil {
|
|
return fmt.Errorf("running PRAGMA optimize at startup: %w", err)
|
|
}
|
|
return nil
|
|
}
|
|
|
|
// Close releases the underlying database connections (both pools).
|
|
func (d *DB) Close() error {
|
|
// Run PRAGMA optimize to analyze and update query planner statistics.
|
|
// It may write statistics, so it runs on the writer.
|
|
_, _ = d.writer.Exec("PRAGMA optimize;")
|
|
var readerErr error
|
|
if d.reader != d.writer {
|
|
readerErr = d.reader.Close()
|
|
}
|
|
err := errors.Join(d.writer.Close(), readerErr)
|
|
// Release the single-process lock only after the pools are closed, so a
|
|
// successor process acquiring it can rely on this one being done writing.
|
|
if d.lockRelease != nil {
|
|
d.lockRelease()
|
|
d.lockRelease = nil
|
|
}
|
|
return err
|
|
}
|
|
|
|
// QueryRowContext executes a query that returns at most one row, with context.
|
|
// Read-only statements run on the reader pool; anything else on the writer.
|
|
func (d *DB) QueryRowContext(ctx context.Context, query string, args ...any) *sql.Row {
|
|
return d.routePool(query).QueryRowContext(ctx, query, args...)
|
|
}
|
|
|
|
// ExecContext executes a query that doesn't return rows, with context.
|
|
// Always runs on the writer.
|
|
func (d *DB) ExecContext(ctx context.Context, query string, args ...any) (sql.Result, error) {
|
|
return d.writer.ExecContext(ctx, query, args...)
|
|
}
|
|
|
|
// QueryContext executes a query that returns multiple rows, with context.
|
|
// Read-only statements run on the reader pool; anything else on the writer.
|
|
func (d *DB) QueryContext(ctx context.Context, query string, args ...any) (*sql.Rows, error) {
|
|
return d.routePool(query).QueryContext(ctx, query, args...)
|
|
}
|
|
|
|
// routePool mirrors dbtx.pool for the public wrapper methods.
|
|
func (d *DB) routePool(query string) *sql.DB {
|
|
if isReadOnlySQL(query) {
|
|
return d.reader
|
|
}
|
|
return d.writer
|
|
}
|
|
|
|
// BeginTx starts a database transaction with context and options.
|
|
// Transactions always run on the writer (file-backed writers BEGIN IMMEDIATE
|
|
// via _txlock, so the write lock is taken up front).
|
|
func (d *DB) BeginTx(ctx context.Context, opts *sql.TxOptions) (*sql.Tx, error) {
|
|
return d.writer.BeginTx(ctx, opts)
|
|
}
|
|
|
|
// SQLDb returns the underlying writer *sql.DB for cases requiring direct
|
|
// access. It is the escape hatch for statements the wrappers can't route —
|
|
// notably PRAGMA wal_checkpoint(TRUNCATE) in the admin backup handler, which
|
|
// must run on the writer to checkpoint the WAL it just stopped appending to.
|
|
func (d *DB) SQLDb() *sql.DB {
|
|
return d.writer
|
|
}
|
|
|
|
// PingRead answers whether the database can serve reads, via a bounded
|
|
// SELECT 1 on the READER pool. The health endpoint uses it deliberately:
|
|
// pinging the single-connection writer would queue behind any long write —
|
|
// most notably a scheduled backup's VACUUM INTO — and report a healthy,
|
|
// read-serving server as degraded for the backup's whole duration. Writer
|
|
// saturation is reported separately (SQLDb().Stats() in /api/v1/metrics),
|
|
// where it is a capacity signal rather than a liveness verdict.
|
|
func (d *DB) PingRead(ctx context.Context) error {
|
|
var one int
|
|
return d.reader.QueryRowContext(ctx, "SELECT 1").Scan(&one)
|
|
}
|