mirror of
https://github.com/J3vb/OwnCord.git
synced 2026-09-03 03:50:00 +03:00
* docs: add infrastructure roadmap plan Records the verified recommendations from an infrastructure review in three tracks: raising the single-instance ceiling, cheap seams for a possible multi-instance future, and ops hygiene. Includes explicit anti-recommendations and sequencing. Security-sensitive detail is intentionally excluded per docs/security.md. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj * feat(server): real health checks and saturation metrics /api/v1/metrics now exposes signals that were already computed in memory but never surfaced: reconnect replay tier hits, event-persister counters, SQLite writer-pool wait stats, aggregate per-client backpressure counters (including previously invisible low-priority drops), and permission-cache hit/miss. /health now returns a real verdict: hub dispatch-loop liveness, a bounded database ping, and a free-disk check, returning 503 with a subsystem reason when degraded. Checks are cached so the unauthenticated endpoint cannot amplify load. The hub's panic breaker now exits the process so a supervisor can restart it, instead of leaving broadcast delivery silently dead while clients still appear online. OTel instruments that were declared but never recorded are now wired (ws_active_connections, ws_broadcast_latency_seconds, ws_messages_total, ws_events_dropped_total, voice gauges) or removed (db_query_duration_seconds). Also corrects the docs/api.md description of broadcast_drops, which counts hub-queue overflow, not client send-queue overflow. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj * feat(server): implement scheduled backups, retention, and backup verification The backup_schedule and backup_retention settings have existed in the admin panel and API since the initial schema but were never read by any code. The 15-minute maintenance loop now enforces them: a scheduled backup is taken when the newest backup on disk is older than the schedule interval (manual backups reset the clock), and retention prunes backups older than the configured days while always keeping the newest one. Backups are now verified with PRAGMA integrity_check immediately after VACUUM INTO (a failed backup is removed rather than listed as restorable) and again before a restore may overwrite the live database. A failed VACUUM INTO also cleans up its partial output file — but never a pre-existing one. The backup directory is configurable via a new backup.dir key (default data/backups) so operators can point backups at another disk or an off-host mount, mirroring the SetDatabasePath plumb. Restore-handler tests now use real SQLite fixtures (the integrity gate correctly refuses text files) with the mid-copy failure injected through a test-only copy hook. Also adds audited gosec suppressions to the Windows disk-free syscall added in the previous commit, which the Windows lint leg flagged. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj * feat(server): capacity and failure-mode guardrails - server.max_ws_connections: optional cap on concurrent WebSocket clients, checked before the upgrade with a 503 + Retry-After; rejections are counted and exposed as ws_conn_rejects in /api/v1/metrics. - Single-process database lock: an OS-level advisory lock (flock / exclusive handle) beside the SQLite file makes a second server process fail fast with a clear message instead of silently fighting the first over process-local state. A bounded retry covers the self-update/restore restart handoff, and the lock mechanism failing (e.g. network filesystems) only warns. - Disk-space awareness: boot-time warnings for the data and backup volumes, plus a disk_free_mb metrics field, via a small cross-platform diskutil package (already used by /health). - Upload storage failures: storage.Save now marks server-side filesystem failures with a sentinel (storage.ErrIO); handlers return 507 for those instead of blaming the client with a 400, and the emoji route stops echoing raw storage errors (which embed absolute paths) into responses. - Unknown config keys now warn at startup — a typo like admin_alowed_cidrs previously kept the default silently while the operator believed the setting changed. Never fatal: newer servers tolerate older configs. - Admin settings honesty: the three stored-but-inert settings (server_icon, max_upload_bytes, voice_quality) are shown read-only with a note pointing at the real config.yaml keys, instead of pretending to apply. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj * perf(db): write-path efficiency and capacity knobs - channel_focus/mark_read now skip the read-state UPSERT when the stored row already matches (same last_message_id, no mentions) — refocus events fire at up to 10/s/user and every no-op write still occupied the single SQLite writer connection. The extra existence check runs on the reader pool, which doesn't serialize. Same shape as the session-touch throttle. - DeleteExpiredSessions is now sargable: migration 031 normalizes legacy expiry formats to the RFC3339-Z layout the server writes and indexes expires_at, replacing the strftime full-table scan that ran on the writer every 15 minutes. - Boot-time ANALYZE runs only when a migration actually applied; unchanged schemas get the cheap PRAGMA optimize instead (which also covers crash-restarts that never reached the shutdown optimize). - The read/write SQL router gets a table-driven test with explicit expected values (INSERT ... RETURNING must hit the writer despite being :one). - New knobs, all defaulting to current behavior: database.max_readers, security.auth_rate_limit_multiplier (for shared-NAT communities), event_persistence.replay_ring_size and replay_cold_limit. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj * fix(server): shutdown lifecycle ordering - The event pruner and maintenance loop are now joined (bounded) before the database closes: bgCtx cancellation used to run AFTER database.Close via LIFO defers, contradicting its own comment, and neither goroutine was ever waited on — a mid-tick scheduled backup or prune could still hold the writer while the pool tore down. StartEventPruner returns a done channel with the same join contract EventPersister.Stop already had. - srv.Shutdown now runs before hub.GracefulStop, so in-flight HTTP handlers' broadcasts still reach a live hub and the event persister instead of vanishing from the replay/event store across a restart. Shutdown does not wait on hijacked WebSocket connections, so the swap adds no delay. - GracefulStopContext threads the 30s shutdown budget into the hub: the 5s client-notice window (matching the countdown clients are shown) ends early when the budget expires, and is skipped entirely when nobody is connected — early-return startup paths and idle servers no longer sleep 5s for an audience of zero. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj * build(deploy): systemd unit, compose hardening, boot-smoked releases, CI polish - deploy/owncord.service: hardened systemd unit template with the two verified caveats encoded (install dir stays writable for self-update under ProtectSystem=strict; CAP_NET_BIND_SERVICE for ACME's :80), plus a 'Linux (systemd)' deployment docs section — the Linux service story was previously 'Docker or nothing'. - New 'Reverse Proxy Topology' docs section with a working nginx snippet and the correct signaling-vs-media distinction: /livekit/* is already proxied by the server, only WebRTC media ports must be directly reachable. - docker-compose: log rotation, commented resource limits, and a healthcheck backed by a new 'chatserver healthcheck' subcommand (the distroless image has no shell) that probes /health without config side effects. - release.yml: a concurrency group (queue, never cancel), and boot-smoke gates — the freshly built server binaries and the Docker image are cold booted and probed healthy BEFORE anything is signed or pushed. The release feed drives signed self-updates, so a binary that compiles but dies on boot previously would have shipped itself to every auto-updating instance. - ci.yml: client-check/client-tests move to ubuntu with the reasoning recorded (no win32 code paths, LF enforced repo-wide); admin-e2e gets a written graduation criterion instead of an open-ended non-blocking status. - docs: Tailscale guide notes the CGNAT range vs the default admin CIDRs; architecture overview records presence/voice state as the fifth single-instance blocker and the macOS client scope decision. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj * perf(server): measured load tooling, narrowed invalidation, presence coalescing, storage and CIDR seams - Fix scripts/k6/ws-load.js against the real wire protocol: envelope-wrapped frames, correct message types (typing_start, presence_update), the correct /api/v1/ws path, and thresholds that fail a run where nobody authenticated or went ready — the script had drifted to pre-envelope framing and reported 100% green while every auth failed on the first frame. A new workflow_dispatch-only load-baseline workflow boots a real server, seeds users through the setup/invite APIs, runs the script, and uploads the k6 summary plus a metrics snapshot for before/after comparison. - Role-scoped channel-override changes now evict only the affected role's members from the permission cache (fail-safe: unreadable member list still flushes everything). InvalidateAll here repopulated every connected user — two reads each — synchronously inside the admin request via RefreshChannelVisibility, a stampede that scaled with total population rather than the role's size. Same pattern the per-user override endpoints already used. - Connect/disconnect presence broadcasts now pass through a 300ms latest-wins coalescer (QueuePresence): each un-coalesced presence change is a sequenced global broadcast (an O(clients) fan-out under seqMu), so a reconnect storm fired O(users) of them from the connect critical path. A flap inside the window collapses to its final state; the wire format, seq ordering, and replay behaviour are unchanged, and the delivery path (BroadcastPresence) is untouched. - Storage seam: api handlers now consume a FileStore interface (consumer-side, same pattern as service.Store) with Open returning a seekable storage.File — writing down the contract (range-request seeks included) an alternative backend would have to meet, without building one. - The metrics surfaces and the LiveKit webhook/health endpoints get their own allowlist keys (metrics_allowed_cidrs, livekit_webhook_allowed_cidrs, both defaulting to admin_allowed_cidrs), so a central Prometheus scraper or an externally-hosted LiveKit no longer requires widening the admin panel's perimeter. Startup now also warns when admin_allowed_cidrs is customized while trusted_proxies is empty — behind a proxy or container network the check would otherwise compare the proxy's private address, not the client's. - The container healthcheck probe now PINS the server's own certificate from disk (VerifyConnection, exact-match) instead of skipping TLS verification, addressing the CodeQL finding on the previous commit; WebPKI verification is used when no local cert exists (ACME). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj * fix(server): address self-review findings on the hardening branch Seven fixes from a high-effort review of the full branch diff: - healthcheck CLI now works under tls.mode acme: it overrides ServerName with the configured domain for WebPKI verification instead of pinning a cert that doesn't exist (or is stale) in that mode. Previously an ACME deployment's container healthcheck failed forever. - /health pings the READER pool (new db.PingRead): the writer ping queued behind a scheduled backup's VACUUM INTO and reported the server degraded for the whole backup — which an autoheal watchdog would turn into a nightly mid-backup restart. - /health runs its cached checks under context.WithoutCancel so a probe that disconnects mid-request cannot poison the shared cache with a false degraded verdict for the next 5 seconds. - The token CLI uses a new db.OpenShared that skips the single-process lock: minting a token against a running server is safe under WAL and was a documented workflow the lock had broken. - The per-user TOTP failure cap is no longer scaled by security.auth_rate_limit_multiplier — that knob exists for per-IP limits; scaling the only cross-IP brute-force defence multiplied an attacker's distributed guess budget. Mirrors the unscaled per-user login threshold. - A direct presence_update now drops the user's queued entry in the connect/disconnect coalescer, so a stale connect-time presence can no longer flush 300ms later over the user's fresher chosen status. - The scheduled-backup filename collision loop breaks on any stat error and bounds its suffix probing, instead of spinning the maintenance goroutine forever on a persistent EACCES. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj * test(admin): real SQLite fixture for the merged Close-failure restore test TestHandleRestoreBackup_RestartsWhenCloseFails arrived from main (#1375) with a plain-text backup fixture; this branch's restore handler verifies backups with integrity_check before touching the live database, so the text fixture was (correctly) refused with 400 before the Close-failure branch under test was reached. Use a real backup via BackupToSafe, matching the other restore tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj --------- Co-authored-by: Claude <noreply@anthropic.com>
734 lines
26 KiB
Go
734 lines
26 KiB
Go
package api
|
||
|
||
import (
|
||
"context"
|
||
"encoding/json"
|
||
"errors"
|
||
"fmt"
|
||
"log/slog"
|
||
"net/http"
|
||
"strings"
|
||
"time"
|
||
"unicode/utf8"
|
||
|
||
"github.com/go-chi/chi/v5"
|
||
"github.com/microcosm-cc/bluemonday"
|
||
"github.com/owncord/server/auth"
|
||
"github.com/owncord/server/db"
|
||
"github.com/owncord/server/permissions"
|
||
"github.com/owncord/server/service"
|
||
)
|
||
|
||
// sanitizer strips all HTML from user-supplied strings before storage.
|
||
var sanitizer = bluemonday.StrictPolicy()
|
||
|
||
// maxLoginUsernameLen bounds the username accepted by handleLogin, mirroring
|
||
// auth.ValidateUsername's 32-rune cap on registered usernames. Enforced
|
||
// before the value is ever used to build a RateLimiter map key — see the
|
||
// check in handleLogin for why.
|
||
const maxLoginUsernameLen = 32
|
||
|
||
// genericAuthError is returned for all login/register failures to avoid
|
||
// revealing whether a username exists.
|
||
var genericAuthError = errorResponse{
|
||
Error: "INVALID_CREDENTIALS",
|
||
Message: "invalid invite or credentials",
|
||
}
|
||
|
||
// registerRequest is the JSON body for POST /api/v1/auth/register.
|
||
type registerRequest struct {
|
||
Username string `json:"username"`
|
||
Password string `json:"password"`
|
||
InviteCode string `json:"invite_code"`
|
||
}
|
||
|
||
// loginRequest is the JSON body for POST /api/v1/auth/login.
|
||
type loginRequest struct {
|
||
Username string `json:"username"`
|
||
Password string `json:"password"`
|
||
}
|
||
|
||
// userResponse is the user shape included in auth responses.
|
||
type userResponse struct {
|
||
ID int64 `json:"id"`
|
||
Username string `json:"username"`
|
||
Avatar string `json:"avatar,omitempty"`
|
||
// DisplayName and About are always present (null = unset) so the settings
|
||
// form can tell "cleared" from "the server does not know this field".
|
||
DisplayName *string `json:"display_name"`
|
||
About *string `json:"about"`
|
||
// CustomStatus is the user's own free-text status line.
|
||
CustomStatus *string `json:"custom_status"`
|
||
// Status is the user's own true status, invisible included. This response
|
||
// only ever describes the caller, so there is nothing to hide from them.
|
||
Status string `json:"status"`
|
||
RoleID int64 `json:"role_id"`
|
||
TOTPEnabled bool `json:"totp_enabled"`
|
||
CreatedAt string `json:"created_at"`
|
||
}
|
||
|
||
// authSuccessResponse is returned on successful login/register.
|
||
type authSuccessResponse struct {
|
||
Token string `json:"token,omitempty"`
|
||
PartialToken string `json:"partial_token,omitempty"`
|
||
Requires2FA bool `json:"requires_2fa"`
|
||
User *userResponse `json:"user,omitempty"`
|
||
}
|
||
|
||
// AuthBroadcaster is the interface handleDeleteAccount uses to notify
|
||
// connected WebSocket clients that an account is gone. Satisfied by *ws.Hub
|
||
// (which already implements BroadcastMemberBan for the admin ban path this
|
||
// mirrors).
|
||
type AuthBroadcaster interface {
|
||
BroadcastMemberBan(userID int64)
|
||
}
|
||
|
||
// MountAuthRoutes registers all auth endpoints on the given router.
|
||
// Rate limiters are applied per-endpoint as specified. trustedProxies is the
|
||
// list of CIDRs whose X-Forwarded-For / X-Real-IP headers are honoured for
|
||
// rate-limiting IP resolution. totpKey is the AES-256 key used to encrypt
|
||
// TOTP secrets at rest (M1 security hardening).
|
||
//
|
||
// broadcaster is variadic and optional: MountAuthRoutes is called before the
|
||
// hub exists (router.go mounts auth routes first, and the hub needs the
|
||
// router to register its own webhook route), so a caller that cannot supply
|
||
// one yet may omit it entirely and self-deletion simply sends no event,
|
||
// exactly like today. A caller mounted after hub creation should pass it so
|
||
// DELETE /api/v1/auth/account can broadcast the same member_ban event the
|
||
// admin ban path already sends for the identical anonymise-and-ban DB state.
|
||
func MountAuthRoutes(r chi.Router, database *db.DB, limiter *auth.RateLimiter, trustedProxies []string, totpKey []byte, broadcaster ...AuthBroadcaster) {
|
||
var ab AuthBroadcaster
|
||
if len(broadcaster) > 0 {
|
||
ab = broadcaster[0]
|
||
}
|
||
registerLimiter := limiter
|
||
loginLimiter := limiter
|
||
partialStore := auth.NewPartialAuthStore(partialAuthStoreTTL)
|
||
pendingTOTPStore := auth.NewPendingTOTPStore(pendingTOTPStoreTTL)
|
||
usedTOTPCodes := auth.NewUsedTOTPCodeStore()
|
||
|
||
r.Route("/api/v1/auth", func(r chi.Router) {
|
||
r.With(RateLimitMiddleware(registerLimiter, "register:", scaledAuthLimit(registerRateLimitPerMinute), time.Minute, trustedProxies)).
|
||
Post("/register", handleRegister(database, trustedProxies))
|
||
|
||
r.With(RateLimitMiddleware(loginLimiter, "login:", scaledAuthLimit(loginRateLimitPerMinute), time.Minute, trustedProxies)).
|
||
Post("/login", handleLogin(database, limiter, partialStore, trustedProxies))
|
||
|
||
r.With(RateLimitMiddleware(limiter, "totp_verify:", scaledAuthLimit(verifyTOTPRateLimitPerMinute), time.Minute, trustedProxies)).
|
||
Post("/verify-totp", handleVerifyTOTP(database, partialStore, limiter, usedTOTPCodes, totpKey))
|
||
|
||
r.With(AuthMiddleware(database)).
|
||
Post("/logout", handleLogout(database))
|
||
|
||
r.With(AuthMiddleware(database)).
|
||
Get("/me", handleMe())
|
||
|
||
r.With(AuthMiddleware(database),
|
||
RateLimitMiddleware(limiter, "del_account:", scaledAuthLimit(sensitiveEndpointRateLimitPerMinute), time.Minute, trustedProxies)).
|
||
Delete("/account", handleDeleteAccount(database, limiter, ab))
|
||
})
|
||
|
||
r.With(AuthMiddleware(database),
|
||
RateLimitMiddleware(limiter, "totp:", scaledAuthLimit(sensitiveEndpointRateLimitPerMinute), time.Minute, trustedProxies)).
|
||
Post("/api/v1/users/me/totp/enable", handleEnableTOTP(pendingTOTPStore, limiter))
|
||
|
||
r.With(AuthMiddleware(database),
|
||
RateLimitMiddleware(limiter, "totp:", scaledAuthLimit(sensitiveEndpointRateLimitPerMinute), time.Minute, trustedProxies)).
|
||
Post("/api/v1/users/me/totp/confirm", handleConfirmTOTP(database, pendingTOTPStore, usedTOTPCodes, limiter, totpKey))
|
||
|
||
r.With(AuthMiddleware(database),
|
||
RateLimitMiddleware(limiter, "totp:", scaledAuthLimit(sensitiveEndpointRateLimitPerMinute), time.Minute, trustedProxies)).
|
||
Delete("/api/v1/users/me/totp", handleDisableTOTP(database, pendingTOTPStore, limiter))
|
||
}
|
||
|
||
// handleRegister processes POST /api/v1/auth/register.
|
||
func handleRegister(database *db.DB, trustedProxies []string) http.HandlerFunc {
|
||
proxyNets := parseCIDRList(trustedProxies) // W3-3a: parse once at construction
|
||
return func(w http.ResponseWriter, r *http.Request) {
|
||
registrationOpen, err := isRegistrationOpen(r.Context(), database)
|
||
if err != nil {
|
||
writeJSON(w, http.StatusInternalServerError, errorResponse{
|
||
Error: "INTERNAL_ERROR",
|
||
Message: "failed to load registration policy",
|
||
})
|
||
return
|
||
}
|
||
if !registrationOpen {
|
||
writeJSON(w, http.StatusForbidden, errorResponse{
|
||
Error: "FORBIDDEN",
|
||
Message: "registration is currently closed",
|
||
})
|
||
return
|
||
}
|
||
|
||
require2FA, err := isRequire2FAEnabled(r.Context(), database)
|
||
if err != nil {
|
||
writeJSON(w, http.StatusInternalServerError, errorResponse{
|
||
Error: "INTERNAL_ERROR",
|
||
Message: "failed to load registration policy",
|
||
})
|
||
return
|
||
}
|
||
if require2FA {
|
||
writeJSON(w, http.StatusForbidden, errorResponse{
|
||
Error: "FORBIDDEN",
|
||
Message: "registration is unavailable while two-factor authentication is required",
|
||
})
|
||
return
|
||
}
|
||
|
||
var req registerRequest
|
||
if err := json.NewDecoder(r.Body).Decode(&req); err != nil {
|
||
writeJSON(w, http.StatusBadRequest, errorResponse{
|
||
Error: "INVALID_INPUT",
|
||
Message: "malformed request body",
|
||
})
|
||
return
|
||
}
|
||
|
||
// F: use the fixpoint sanitizer (service.SanitizeText), not the bare
|
||
// sanitizer.Sanitize below — Sanitize's output is always HTML-escaped
|
||
// (' -> ', & -> &, " -> "), so a plain call here would store
|
||
// a different string than what handleLogin looks up (which only
|
||
// trims), permanently locking out any username containing one of
|
||
// those characters. See service.SanitizeText's doc comment.
|
||
req.Username = strings.TrimSpace(service.SanitizeText(req.Username))
|
||
req.InviteCode = strings.TrimSpace(req.InviteCode)
|
||
|
||
if req.Username == "" || req.Password == "" || req.InviteCode == "" {
|
||
writeJSON(w, http.StatusBadRequest, errorResponse{
|
||
Error: "INVALID_INPUT",
|
||
Message: "username, password, and invite_code are required",
|
||
})
|
||
return
|
||
}
|
||
|
||
// Validate username format (length, no control/invisible chars).
|
||
if err := auth.ValidateUsername(req.Username); err != nil {
|
||
writeJSON(w, http.StatusBadRequest, errorResponse{
|
||
Error: "INVALID_INPUT",
|
||
Message: err.Error(),
|
||
})
|
||
return
|
||
}
|
||
|
||
// Validate password strength before anything else.
|
||
if err := auth.ValidatePasswordStrength(req.Password); err != nil {
|
||
writeJSON(w, http.StatusBadRequest, errorResponse{
|
||
Error: "INVALID_INPUT",
|
||
Message: err.Error(),
|
||
})
|
||
return
|
||
}
|
||
|
||
// Hash password before consuming the invite so that a hashing failure
|
||
// does not burn a valid invite code.
|
||
hash, err := auth.HashPassword(req.Password)
|
||
if err != nil {
|
||
writeJSON(w, http.StatusInternalServerError, errorResponse{
|
||
Error: "INTERNAL_ERROR",
|
||
Message: "failed to process registration",
|
||
})
|
||
return
|
||
}
|
||
|
||
// Atomically consume the invite and create the user so failed
|
||
// registrations do not burn a valid invite code.
|
||
uid, err := database.CreateUserWithInvite(r.Context(), req.Username, hash, int(permissions.MemberRoleID), req.InviteCode)
|
||
if err != nil {
|
||
// UNIQUE constraint violation → duplicate username → 400.
|
||
// Any other DB error → 500.
|
||
switch {
|
||
case db.IsUniqueConstraintError(err):
|
||
writeJSON(w, http.StatusBadRequest, genericAuthError)
|
||
case errors.Is(err, db.ErrNotFound):
|
||
writeJSON(w, http.StatusBadRequest, genericAuthError)
|
||
default:
|
||
slog.Error("CreateUserWithInvite failed", "err", err, "username", req.Username)
|
||
writeJSON(w, http.StatusInternalServerError, errorResponse{
|
||
Error: "INTERNAL_ERROR",
|
||
Message: "registration failed — please try again",
|
||
})
|
||
}
|
||
return
|
||
}
|
||
|
||
ip := clientIPWithProxies(r, proxyNets)
|
||
slog.Info("user registered", "username", req.Username, "user_id", uid, "ip", ip)
|
||
db.WriteAudit(context.WithoutCancel(r.Context()), database, uid, "user_register", "user", uid,
|
||
"new account created via invite")
|
||
|
||
// Issue session.
|
||
token, err := auth.GenerateToken()
|
||
if err != nil {
|
||
writeJSON(w, http.StatusInternalServerError, errorResponse{
|
||
Error: "INTERNAL_ERROR",
|
||
Message: "failed to create session",
|
||
})
|
||
return
|
||
}
|
||
|
||
device := truncateDevice(r.Header.Get("User-Agent"))
|
||
if _, err := database.CreateSession(r.Context(), uid, auth.HashToken(token), device, ip); err != nil {
|
||
writeJSON(w, http.StatusInternalServerError, errorResponse{
|
||
Error: "INTERNAL_ERROR",
|
||
Message: "failed to create session",
|
||
})
|
||
return
|
||
}
|
||
|
||
user, err := database.GetUserByID(r.Context(), uid)
|
||
if err != nil || user == nil {
|
||
slog.Error("failed to fetch user after registration", "user_id", uid, "error", err)
|
||
writeJSON(w, http.StatusInternalServerError, errorResponse{
|
||
Error: "INTERNAL_ERROR",
|
||
Message: "registration succeeded but user fetch failed",
|
||
})
|
||
return
|
||
}
|
||
writeJSON(w, http.StatusCreated, authSuccessResponse{
|
||
Token: token,
|
||
Requires2FA: false,
|
||
User: toUserResponse(user),
|
||
})
|
||
}
|
||
}
|
||
|
||
// handleLogin processes POST /api/v1/auth/login.
|
||
func handleLogin(database *db.DB, limiter *auth.RateLimiter, partialStore *auth.PartialAuthStore, trustedProxies []string) http.HandlerFunc {
|
||
proxyNets := parseCIDRList(trustedProxies) // W3-3a: parse once at construction
|
||
return func(w http.ResponseWriter, r *http.Request) {
|
||
var req loginRequest
|
||
if err := json.NewDecoder(r.Body).Decode(&req); err != nil {
|
||
writeJSON(w, http.StatusBadRequest, errorResponse{
|
||
Error: "INVALID_INPUT",
|
||
Message: "malformed request body",
|
||
})
|
||
return
|
||
}
|
||
|
||
req.Username = strings.TrimSpace(req.Username)
|
||
// Do NOT trim req.Password — passwords may intentionally contain
|
||
// leading/trailing whitespace. Bcrypt handles arbitrary bytes.
|
||
|
||
if req.Username == "" || req.Password == "" {
|
||
writeJSON(w, http.StatusBadRequest, errorResponse{
|
||
Error: "INVALID_INPUT",
|
||
Message: "username and password are required",
|
||
})
|
||
return
|
||
}
|
||
|
||
// F: reject an over-long username before it is ever used to build a
|
||
// RateLimiter map key below (unameKey, failKey, userFailKey, lockout
|
||
// keys). Unlike registration, login has no account to validate
|
||
// against yet, so nothing else bounds this value — an unauthenticated
|
||
// caller could otherwise pin an arbitrarily large, body-sized string
|
||
// as a retained key (Cleanup only evicts it after hours). Mirrors the
|
||
// same 32-rune cap auth.ValidateUsername enforces at registration.
|
||
if utf8.RuneCountInString(req.Username) > maxLoginUsernameLen {
|
||
writeJSON(w, http.StatusBadRequest, errorResponse{
|
||
Error: "INVALID_INPUT",
|
||
Message: "username is too long",
|
||
})
|
||
return
|
||
}
|
||
|
||
ip := clientIPWithProxies(r, proxyNets)
|
||
|
||
// Check per-IP lockout first.
|
||
lockKey := "login_lock:" + ip
|
||
if limiter.IsLockedOut(lockKey) {
|
||
writeJSON(w, http.StatusTooManyRequests, errorResponse{
|
||
Error: "RATE_LIMITED",
|
||
Message: "account temporarily locked due to too many failed attempts",
|
||
})
|
||
return
|
||
}
|
||
|
||
// BUG-110: Also check per-username lockout to prevent distributed brute force.
|
||
// F1: canonicalize the username the same way GetUserByUsername does (COLLATE
|
||
// NOCASE) before keying the lockout, so case variants of one account
|
||
// (admin/Admin/ADMIN) share a single bucket instead of each getting its own.
|
||
unameKey := strings.ToLower(req.Username)
|
||
userLockKey := "login_user_lock:" + unameKey
|
||
if limiter.IsLockedOut(userLockKey) {
|
||
writeJSON(w, http.StatusTooManyRequests, errorResponse{
|
||
Error: "RATE_LIMITED",
|
||
Message: "account temporarily locked due to too many failed attempts",
|
||
})
|
||
return
|
||
}
|
||
|
||
// Constant-time lookup: always attempt bcrypt compare even when user
|
||
// does not exist to prevent timing-based username enumeration.
|
||
user, err := database.GetUserByUsername(r.Context(), req.Username)
|
||
|
||
// Distinguish DB errors from authentication failures. DB errors
|
||
// should NOT increment the rate limiter — otherwise a transient
|
||
// DB outage would lock out legitimate users.
|
||
if err != nil && user == nil {
|
||
// Could be a real DB error or simply "user not found".
|
||
// GetUserByUsername returns (nil, nil) for not-found, so a
|
||
// non-nil error here is a genuine DB failure.
|
||
slog.Error("login: GetUserByUsername failed", "err", err, "ip", ip)
|
||
writeJSON(w, http.StatusInternalServerError, errorResponse{
|
||
Error: "INTERNAL_ERROR",
|
||
Message: "login temporarily unavailable",
|
||
})
|
||
return
|
||
}
|
||
|
||
failKey := "login_fail:" + ip
|
||
userFailKey := "login_user_fail:" + unameKey
|
||
// F3: atomically reserve this attempt BEFORE the bcrypt compare. The
|
||
// read-only IsLockedOut gates above are check-then-act: N concurrent
|
||
// requests all pass them before any failure is recorded below, so the
|
||
// per-username cap — the only cross-IP brute-force defence — bound
|
||
// only sequential attackers. Allow records the attempt under the
|
||
// limiter's lock, capping a concurrent burst at the same budget a
|
||
// sequential attacker gets. Sized at threshold+1 so the sequential
|
||
// accepted-input set is unchanged: failures 1–10 still land, the 10th
|
||
// still trips the lockout (via the Check below), and a correct
|
||
// password on attempt 10 still succeeds — successful logins reset
|
||
// both counters. The reservation sits after the DB-error return above
|
||
// so a transient DB outage still does not consume attempts.
|
||
if !limiter.Allow(failKey, scaledAuthLimit(loginFailureThreshold)+1, loginFailureWindow) ||
|
||
!limiter.Allow(userFailKey, loginUserFailureThreshold+1, loginUserFailureWindow) {
|
||
writeJSON(w, http.StatusTooManyRequests, errorResponse{
|
||
Error: "RATE_LIMITED",
|
||
Message: "account temporarily locked due to too many failed attempts",
|
||
})
|
||
return
|
||
}
|
||
// Always run the password check — with an empty hash when the user does
|
||
// not exist. auth.CheckPassword performs a dummy bcrypt comparison for an
|
||
// empty hash, so bcrypt executes on every path and response time stays
|
||
// constant, preventing timing-based username enumeration. (A `user == nil
|
||
// || CheckPassword(...)` short-circuit would skip bcrypt entirely for
|
||
// unknown usernames, reintroducing the timing side-channel.)
|
||
storedHash := ""
|
||
if user != nil {
|
||
storedHash = user.PasswordHash
|
||
}
|
||
if !auth.CheckPassword(storedHash, req.Password) {
|
||
// The attempt was already recorded atomically up-front (F3); here
|
||
// only decide the lockouts, at the same boundary as before: the
|
||
// 10th in-window failure locks the key. Check is read-only, so
|
||
// the reservation is not double-counted.
|
||
if !limiter.Check(failKey, scaledAuthLimit(loginFailureThreshold)+1, loginFailureWindow) {
|
||
limiter.Lockout(r.Context(), lockKey, loginLockoutDuration)
|
||
}
|
||
// BUG-110: per-username lockout on threshold.
|
||
if !limiter.Check(userFailKey, loginUserFailureThreshold+1, loginUserFailureWindow) {
|
||
limiter.Lockout(r.Context(), userLockKey, loginUserLockoutDuration)
|
||
}
|
||
slog.Info("login failed", "ip", ip, "username_len", len(req.Username))
|
||
writeJSON(w, http.StatusUnauthorized, errorResponse{
|
||
Error: "UNAUTHORIZED",
|
||
Message: "invalid credentials",
|
||
})
|
||
return
|
||
}
|
||
|
||
// Reset failure counters on success.
|
||
limiter.Reset(r.Context(), failKey)
|
||
limiter.Reset(r.Context(), userFailKey)
|
||
|
||
if auth.IsEffectivelyBanned(user) {
|
||
slog.Warn("banned user login attempt", "username", user.Username, "user_id", user.ID, "ip", ip)
|
||
db.WriteAudit(context.WithoutCancel(r.Context()), database, user.ID, "login_blocked_banned", "user", user.ID,
|
||
"banned user attempted login from "+ip)
|
||
writeJSON(w, http.StatusForbidden, errorResponse{
|
||
Error: "FORBIDDEN",
|
||
Message: "your account has been suspended",
|
||
})
|
||
return
|
||
}
|
||
|
||
require2FA, err := isRequire2FAEnabled(r.Context(), database)
|
||
if err != nil {
|
||
writeJSON(w, http.StatusInternalServerError, errorResponse{
|
||
Error: "INTERNAL_ERROR",
|
||
Message: "failed to load authentication policy",
|
||
})
|
||
return
|
||
}
|
||
if user.TOTPSecret != nil {
|
||
partialToken, err := partialStore.Issue(user.ID, truncateDevice(r.Header.Get("User-Agent")), ip)
|
||
if err != nil {
|
||
writeJSON(w, http.StatusInternalServerError, errorResponse{
|
||
Error: "INTERNAL_ERROR",
|
||
Message: "failed to start two-factor challenge",
|
||
})
|
||
return
|
||
}
|
||
writeJSON(w, http.StatusOK, authSuccessResponse{
|
||
PartialToken: partialToken,
|
||
Requires2FA: true,
|
||
})
|
||
return
|
||
}
|
||
if require2FA {
|
||
writeJSON(w, http.StatusForbidden, errorResponse{
|
||
Error: "FORBIDDEN",
|
||
Message: "two-factor authentication must be enabled on this account before login",
|
||
})
|
||
return
|
||
}
|
||
|
||
// Issue session.
|
||
token, err := issueSession(r.Context(), database, user.ID, truncateDevice(r.Header.Get("User-Agent")), ip)
|
||
if err != nil {
|
||
writeJSON(w, http.StatusInternalServerError, errorResponse{
|
||
Error: "INTERNAL_ERROR",
|
||
Message: "failed to create session",
|
||
})
|
||
return
|
||
}
|
||
|
||
// Don't set status to "online" here — the WebSocket connection in
|
||
// serve.go does that when the user actually connects. Setting it here
|
||
// would leave the user permanently "online" if they never open a WS
|
||
// connection or if the client crashes before connecting.
|
||
slog.Info("user logged in", "username", user.Username, "user_id", user.ID, "ip", ip)
|
||
db.WriteAudit(context.WithoutCancel(r.Context()), database, user.ID, "user_login", "user", user.ID,
|
||
"logged in from "+ip)
|
||
writeJSON(w, http.StatusOK, authSuccessResponse{
|
||
Token: token,
|
||
Requires2FA: false,
|
||
User: toUserResponse(user),
|
||
})
|
||
}
|
||
}
|
||
|
||
// handleLogout processes POST /api/v1/auth/logout.
|
||
func handleLogout(database *db.DB) http.HandlerFunc {
|
||
return func(w http.ResponseWriter, r *http.Request) {
|
||
sess, ok := r.Context().Value(SessionKey).(*db.Session)
|
||
if !ok || sess == nil {
|
||
writeJSON(w, http.StatusUnauthorized, errorResponse{
|
||
Error: "UNAUTHORIZED",
|
||
Message: "not authenticated",
|
||
})
|
||
return
|
||
}
|
||
|
||
// The client clears its token optimistically — once logout reaches the
|
||
// server, the revocation must not die with a dropped connection.
|
||
if err := database.DeleteSession(context.WithoutCancel(r.Context()), sess.TokenHash); err != nil {
|
||
writeJSON(w, http.StatusInternalServerError, errorResponse{
|
||
Error: "INTERNAL_ERROR",
|
||
Message: "failed to logout",
|
||
})
|
||
return
|
||
}
|
||
|
||
// A custom status is a "what I am doing right now" note. Leaving it
|
||
// standing after the user signed out states something about them that
|
||
// is no longer true, so logout clears it — unlike the chosen presence
|
||
// status, which is a preference and deliberately survives.
|
||
if err := database.UpdateUserCustomStatus(context.WithoutCancel(r.Context()), sess.UserID, nil); err != nil {
|
||
slog.Warn("failed to clear custom status on logout", "user_id", sess.UserID, "err", err)
|
||
}
|
||
|
||
slog.Info("user logged out", "user_id", sess.UserID)
|
||
db.WriteAudit(context.WithoutCancel(r.Context()), database, sess.UserID, "user_logout", "user", sess.UserID, "")
|
||
|
||
w.WriteHeader(http.StatusNoContent)
|
||
}
|
||
}
|
||
|
||
// handleMe processes GET /api/v1/auth/me.
|
||
func handleMe() http.HandlerFunc {
|
||
return func(w http.ResponseWriter, r *http.Request) {
|
||
user, ok := r.Context().Value(UserKey).(*db.User)
|
||
if !ok || user == nil {
|
||
writeJSON(w, http.StatusUnauthorized, errorResponse{
|
||
Error: "UNAUTHORIZED",
|
||
Message: "not authenticated",
|
||
})
|
||
return
|
||
}
|
||
writeJSON(w, http.StatusOK, toUserResponse(user))
|
||
}
|
||
}
|
||
|
||
// deleteAccountRequest is the JSON body for DELETE /api/v1/auth/account.
|
||
type deleteAccountRequest struct {
|
||
Password string `json:"password"`
|
||
}
|
||
|
||
// handleDeleteAccount processes DELETE /api/v1/auth/account.
|
||
// The caller must supply their current password for confirmation.
|
||
// Progressive lockout mirrors the login handler: 3 failures → 15-min lock.
|
||
// broadcaster may be nil, in which case no event is sent and other connected
|
||
// clients converge on their next reconnect instead (same fallback every
|
||
// other broadcaster-optional handler in this package uses).
|
||
func handleDeleteAccount(database *db.DB, limiter *auth.RateLimiter, broadcaster AuthBroadcaster) http.HandlerFunc {
|
||
return func(w http.ResponseWriter, r *http.Request) {
|
||
user, ok := r.Context().Value(UserKey).(*db.User)
|
||
if !ok || user == nil {
|
||
writeJSON(w, http.StatusUnauthorized, errorResponse{
|
||
Error: "UNAUTHORIZED",
|
||
Message: "not authenticated",
|
||
})
|
||
return
|
||
}
|
||
|
||
// Per-user lockout to prevent password brute-force on this destructive endpoint.
|
||
lockKey := auth.Key("delete_lock", user.ID)
|
||
if limiter.IsLockedOut(lockKey) {
|
||
writeJSON(w, http.StatusTooManyRequests, errorResponse{
|
||
Error: "RATE_LIMITED",
|
||
Message: "too many failed attempts, try again later",
|
||
})
|
||
return
|
||
}
|
||
|
||
var req deleteAccountRequest
|
||
if err := json.NewDecoder(r.Body).Decode(&req); err != nil {
|
||
writeJSON(w, http.StatusBadRequest, errorResponse{
|
||
Error: "INVALID_INPUT",
|
||
Message: "malformed request body",
|
||
})
|
||
return
|
||
}
|
||
|
||
if req.Password == "" {
|
||
writeJSON(w, http.StatusBadRequest, errorResponse{
|
||
Error: "INVALID_INPUT",
|
||
Message: "password is required",
|
||
})
|
||
return
|
||
}
|
||
|
||
// Verify the supplied password matches the stored hash.
|
||
failKey := auth.Key("delete_fail", user.ID)
|
||
if !auth.CheckPassword(user.PasswordHash, req.Password) {
|
||
if !limiter.Allow(failKey, deleteAccountFailureThreshold, deleteAccountFailureWindow) {
|
||
limiter.Lockout(r.Context(), lockKey, deleteAccountLockoutDuration)
|
||
}
|
||
writeJSON(w, http.StatusBadRequest, errorResponse{
|
||
Error: "INVALID_INPUT",
|
||
Message: "incorrect password",
|
||
})
|
||
return
|
||
}
|
||
limiter.Reset(r.Context(), failKey)
|
||
|
||
if err := database.DeleteAccount(r.Context(), user.ID); err != nil {
|
||
if errors.Is(err, db.ErrLastAdmin) {
|
||
writeJSON(w, http.StatusForbidden, errorResponse{
|
||
Error: "FORBIDDEN",
|
||
Message: "cannot delete the last admin account",
|
||
})
|
||
return
|
||
}
|
||
slog.Error("DeleteAccount failed", "err", err, "user_id", user.ID)
|
||
writeJSON(w, http.StatusInternalServerError, errorResponse{
|
||
Error: "INTERNAL_ERROR",
|
||
Message: "failed to delete account",
|
||
})
|
||
return
|
||
}
|
||
|
||
ip := clientIP(r)
|
||
slog.Info("account deleted", "username", user.Username, "user_id", user.ID, "ip", ip)
|
||
db.WriteAudit(context.WithoutCancel(r.Context()), database, user.ID, "account_deleted", "user", user.ID,
|
||
"account self-deleted from "+ip)
|
||
|
||
// DeleteAccount left the row in exactly the state an admin ban does
|
||
// (anonymised, banned, sessions revoked) — broadcast the same event so
|
||
// every other connected client drops the deleted user immediately
|
||
// instead of keeping their pre-deletion username until it reconnects.
|
||
if broadcaster != nil {
|
||
broadcaster.BroadcastMemberBan(user.ID)
|
||
}
|
||
|
||
w.WriteHeader(http.StatusNoContent)
|
||
}
|
||
}
|
||
|
||
// toUserResponse converts a db.User to the API response shape.
|
||
func toUserResponse(u *db.User) *userResponse {
|
||
avatar := ""
|
||
if u.Avatar != nil {
|
||
avatar = *u.Avatar
|
||
}
|
||
resp := &userResponse{
|
||
ID: u.ID,
|
||
Username: u.Username,
|
||
Avatar: avatar,
|
||
DisplayName: u.DisplayName,
|
||
About: u.About,
|
||
CustomStatus: u.CustomStatus,
|
||
Status: u.Status,
|
||
RoleID: u.RoleID,
|
||
TOTPEnabled: u.TOTPSecret != nil,
|
||
CreatedAt: u.CreatedAt,
|
||
}
|
||
return resp
|
||
}
|
||
|
||
// truncateDevice truncates the User-Agent to prevent oversized session records.
|
||
const maxDeviceLen = 512
|
||
|
||
func truncateDevice(ua string) string {
|
||
if len(ua) > maxDeviceLen {
|
||
return ua[:maxDeviceLen]
|
||
}
|
||
return ua
|
||
}
|
||
|
||
func issueSession(ctx context.Context, database *db.DB, userID int64, device, ip string) (string, error) {
|
||
token, err := auth.GenerateToken()
|
||
if err != nil {
|
||
return "", err
|
||
}
|
||
if _, err := database.CreateSession(ctx, userID, auth.HashToken(token), device, ip); err != nil {
|
||
return "", err
|
||
}
|
||
return token, nil
|
||
}
|
||
|
||
func isRequire2FAEnabled(ctx context.Context, database *db.DB) (bool, error) {
|
||
return getBooleanSetting(ctx, database, "require_2fa", false)
|
||
}
|
||
|
||
func isRegistrationOpen(ctx context.Context, database *db.DB) (bool, error) {
|
||
return getBooleanSetting(ctx, database, "registration_open", true)
|
||
}
|
||
|
||
func getBooleanSetting(ctx context.Context, database *db.DB, key string, defaultValue bool) (bool, error) {
|
||
value, err := database.GetSetting(ctx, key)
|
||
if err != nil {
|
||
if errors.Is(err, db.ErrNotFound) {
|
||
return defaultValue, nil
|
||
}
|
||
return false, err
|
||
}
|
||
return parseBooleanSettingValue(value)
|
||
}
|
||
|
||
func parseBooleanSettingValue(value string) (bool, error) {
|
||
switch strings.ToLower(strings.TrimSpace(value)) {
|
||
case "1", "true":
|
||
return true, nil
|
||
case "0", "false":
|
||
return false, nil
|
||
default:
|
||
return false, fmt.Errorf("invalid boolean setting value %q", value)
|
||
}
|
||
}
|
||
|
||
func requirePasswordConfirmation(user *db.User, password string) error {
|
||
if password == "" {
|
||
return fmt.Errorf("password is required")
|
||
}
|
||
if !auth.CheckPassword(user.PasswordHash, password) {
|
||
return fmt.Errorf("password confirmation failed")
|
||
}
|
||
return nil
|
||
}
|