Files
OwnCord/docs/server-configuration.md
T
J3vbandClaude Fable 5 f5faf82a60 infra: observability, backups, guardrails, and deployment hardening (#1376)
* docs: add infrastructure roadmap plan

Records the verified recommendations from an infrastructure review in three
tracks: raising the single-instance ceiling, cheap seams for a possible
multi-instance future, and ops hygiene. Includes explicit anti-recommendations
and sequencing. Security-sensitive detail is intentionally excluded per
docs/security.md.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj

* feat(server): real health checks and saturation metrics

/api/v1/metrics now exposes signals that were already computed in memory but
never surfaced: reconnect replay tier hits, event-persister counters, SQLite
writer-pool wait stats, aggregate per-client backpressure counters (including
previously invisible low-priority drops), and permission-cache hit/miss.

/health now returns a real verdict: hub dispatch-loop liveness, a bounded
database ping, and a free-disk check, returning 503 with a subsystem reason
when degraded. Checks are cached so the unauthenticated endpoint cannot
amplify load. The hub's panic breaker now exits the process so a supervisor
can restart it, instead of leaving broadcast delivery silently dead while
clients still appear online.

OTel instruments that were declared but never recorded are now wired
(ws_active_connections, ws_broadcast_latency_seconds, ws_messages_total,
ws_events_dropped_total, voice gauges) or removed (db_query_duration_seconds).
Also corrects the docs/api.md description of broadcast_drops, which counts
hub-queue overflow, not client send-queue overflow.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj

* feat(server): implement scheduled backups, retention, and backup verification

The backup_schedule and backup_retention settings have existed in the admin
panel and API since the initial schema but were never read by any code. The
15-minute maintenance loop now enforces them: a scheduled backup is taken
when the newest backup on disk is older than the schedule interval (manual
backups reset the clock), and retention prunes backups older than the
configured days while always keeping the newest one.

Backups are now verified with PRAGMA integrity_check immediately after
VACUUM INTO (a failed backup is removed rather than listed as restorable)
and again before a restore may overwrite the live database. A failed VACUUM
INTO also cleans up its partial output file — but never a pre-existing one.

The backup directory is configurable via a new backup.dir key (default
data/backups) so operators can point backups at another disk or an off-host
mount, mirroring the SetDatabasePath plumb.

Restore-handler tests now use real SQLite fixtures (the integrity gate
correctly refuses text files) with the mid-copy failure injected through a
test-only copy hook. Also adds audited gosec suppressions to the Windows
disk-free syscall added in the previous commit, which the Windows lint leg
flagged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj

* feat(server): capacity and failure-mode guardrails

- server.max_ws_connections: optional cap on concurrent WebSocket clients,
  checked before the upgrade with a 503 + Retry-After; rejections are counted
  and exposed as ws_conn_rejects in /api/v1/metrics.
- Single-process database lock: an OS-level advisory lock (flock / exclusive
  handle) beside the SQLite file makes a second server process fail fast with
  a clear message instead of silently fighting the first over process-local
  state. A bounded retry covers the self-update/restore restart handoff, and
  the lock mechanism failing (e.g. network filesystems) only warns.
- Disk-space awareness: boot-time warnings for the data and backup volumes,
  plus a disk_free_mb metrics field, via a small cross-platform diskutil
  package (already used by /health).
- Upload storage failures: storage.Save now marks server-side filesystem
  failures with a sentinel (storage.ErrIO); handlers return 507 for those
  instead of blaming the client with a 400, and the emoji route stops echoing
  raw storage errors (which embed absolute paths) into responses.
- Unknown config keys now warn at startup — a typo like admin_alowed_cidrs
  previously kept the default silently while the operator believed the
  setting changed. Never fatal: newer servers tolerate older configs.
- Admin settings honesty: the three stored-but-inert settings (server_icon,
  max_upload_bytes, voice_quality) are shown read-only with a note pointing
  at the real config.yaml keys, instead of pretending to apply.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj

* perf(db): write-path efficiency and capacity knobs

- channel_focus/mark_read now skip the read-state UPSERT when the stored row
  already matches (same last_message_id, no mentions) — refocus events fire
  at up to 10/s/user and every no-op write still occupied the single SQLite
  writer connection. The extra existence check runs on the reader pool, which
  doesn't serialize. Same shape as the session-touch throttle.
- DeleteExpiredSessions is now sargable: migration 031 normalizes legacy
  expiry formats to the RFC3339-Z layout the server writes and indexes
  expires_at, replacing the strftime full-table scan that ran on the writer
  every 15 minutes.
- Boot-time ANALYZE runs only when a migration actually applied; unchanged
  schemas get the cheap PRAGMA optimize instead (which also covers
  crash-restarts that never reached the shutdown optimize).
- The read/write SQL router gets a table-driven test with explicit expected
  values (INSERT ... RETURNING must hit the writer despite being :one).
- New knobs, all defaulting to current behavior: database.max_readers,
  security.auth_rate_limit_multiplier (for shared-NAT communities),
  event_persistence.replay_ring_size and replay_cold_limit.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj

* fix(server): shutdown lifecycle ordering

- The event pruner and maintenance loop are now joined (bounded) before the
  database closes: bgCtx cancellation used to run AFTER database.Close via
  LIFO defers, contradicting its own comment, and neither goroutine was ever
  waited on — a mid-tick scheduled backup or prune could still hold the
  writer while the pool tore down. StartEventPruner returns a done channel
  with the same join contract EventPersister.Stop already had.
- srv.Shutdown now runs before hub.GracefulStop, so in-flight HTTP handlers'
  broadcasts still reach a live hub and the event persister instead of
  vanishing from the replay/event store across a restart. Shutdown does not
  wait on hijacked WebSocket connections, so the swap adds no delay.
- GracefulStopContext threads the 30s shutdown budget into the hub: the 5s
  client-notice window (matching the countdown clients are shown) ends early
  when the budget expires, and is skipped entirely when nobody is connected —
  early-return startup paths and idle servers no longer sleep 5s for an
  audience of zero.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj

* build(deploy): systemd unit, compose hardening, boot-smoked releases, CI polish

- deploy/owncord.service: hardened systemd unit template with the two
  verified caveats encoded (install dir stays writable for self-update under
  ProtectSystem=strict; CAP_NET_BIND_SERVICE for ACME's :80), plus a
  'Linux (systemd)' deployment docs section — the Linux service story was
  previously 'Docker or nothing'.
- New 'Reverse Proxy Topology' docs section with a working nginx snippet and
  the correct signaling-vs-media distinction: /livekit/* is already proxied
  by the server, only WebRTC media ports must be directly reachable.
- docker-compose: log rotation, commented resource limits, and a healthcheck
  backed by a new 'chatserver healthcheck' subcommand (the distroless image
  has no shell) that probes /health without config side effects.
- release.yml: a concurrency group (queue, never cancel), and boot-smoke
  gates — the freshly built server binaries and the Docker image are cold
  booted and probed healthy BEFORE anything is signed or pushed. The release
  feed drives signed self-updates, so a binary that compiles but dies on
  boot previously would have shipped itself to every auto-updating instance.
- ci.yml: client-check/client-tests move to ubuntu with the reasoning
  recorded (no win32 code paths, LF enforced repo-wide); admin-e2e gets a
  written graduation criterion instead of an open-ended non-blocking status.
- docs: Tailscale guide notes the CGNAT range vs the default admin CIDRs;
  architecture overview records presence/voice state as the fifth
  single-instance blocker and the macOS client scope decision.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj

* perf(server): measured load tooling, narrowed invalidation, presence coalescing, storage and CIDR seams

- Fix scripts/k6/ws-load.js against the real wire protocol: envelope-wrapped
  frames, correct message types (typing_start, presence_update), the correct
  /api/v1/ws path, and thresholds that fail a run where nobody authenticated
  or went ready — the script had drifted to pre-envelope framing and reported
  100% green while every auth failed on the first frame. A new
  workflow_dispatch-only load-baseline workflow boots a real server, seeds
  users through the setup/invite APIs, runs the script, and uploads the k6
  summary plus a metrics snapshot for before/after comparison.
- Role-scoped channel-override changes now evict only the affected role's
  members from the permission cache (fail-safe: unreadable member list still
  flushes everything). InvalidateAll here repopulated every connected user —
  two reads each — synchronously inside the admin request via
  RefreshChannelVisibility, a stampede that scaled with total population
  rather than the role's size. Same pattern the per-user override endpoints
  already used.
- Connect/disconnect presence broadcasts now pass through a 300ms latest-wins
  coalescer (QueuePresence): each un-coalesced presence change is a sequenced
  global broadcast (an O(clients) fan-out under seqMu), so a reconnect storm
  fired O(users) of them from the connect critical path. A flap inside the
  window collapses to its final state; the wire format, seq ordering, and
  replay behaviour are unchanged, and the delivery path (BroadcastPresence)
  is untouched.
- Storage seam: api handlers now consume a FileStore interface (consumer-side,
  same pattern as service.Store) with Open returning a seekable storage.File —
  writing down the contract (range-request seeks included) an alternative
  backend would have to meet, without building one.
- The metrics surfaces and the LiveKit webhook/health endpoints get their own
  allowlist keys (metrics_allowed_cidrs, livekit_webhook_allowed_cidrs, both
  defaulting to admin_allowed_cidrs), so a central Prometheus scraper or an
  externally-hosted LiveKit no longer requires widening the admin panel's
  perimeter. Startup now also warns when admin_allowed_cidrs is customized
  while trusted_proxies is empty — behind a proxy or container network the
  check would otherwise compare the proxy's private address, not the client's.
- The container healthcheck probe now PINS the server's own certificate from
  disk (VerifyConnection, exact-match) instead of skipping TLS verification,
  addressing the CodeQL finding on the previous commit; WebPKI verification
  is used when no local cert exists (ACME).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj

* fix(server): address self-review findings on the hardening branch

Seven fixes from a high-effort review of the full branch diff:

- healthcheck CLI now works under tls.mode acme: it overrides ServerName
  with the configured domain for WebPKI verification instead of pinning a
  cert that doesn't exist (or is stale) in that mode. Previously an ACME
  deployment's container healthcheck failed forever.
- /health pings the READER pool (new db.PingRead): the writer ping queued
  behind a scheduled backup's VACUUM INTO and reported the server degraded
  for the whole backup — which an autoheal watchdog would turn into a
  nightly mid-backup restart.
- /health runs its cached checks under context.WithoutCancel so a probe
  that disconnects mid-request cannot poison the shared cache with a false
  degraded verdict for the next 5 seconds.
- The token CLI uses a new db.OpenShared that skips the single-process
  lock: minting a token against a running server is safe under WAL and was
  a documented workflow the lock had broken.
- The per-user TOTP failure cap is no longer scaled by
  security.auth_rate_limit_multiplier — that knob exists for per-IP limits;
  scaling the only cross-IP brute-force defence multiplied an attacker's
  distributed guess budget. Mirrors the unscaled per-user login threshold.
- A direct presence_update now drops the user's queued entry in the
  connect/disconnect coalescer, so a stale connect-time presence can no
  longer flush 300ms later over the user's fresher chosen status.
- The scheduled-backup filename collision loop breaks on any stat error
  and bounds its suffix probing, instead of spinning the maintenance
  goroutine forever on a persistent EACCES.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj

* test(admin): real SQLite fixture for the merged Close-failure restore test

TestHandleRestoreBackup_RestartsWhenCloseFails arrived from main (#1375)
with a plain-text backup fixture; this branch's restore handler verifies
backups with integrity_check before touching the live database, so the text
fixture was (correctly) refused with 400 before the Close-failure branch
under test was reached. Use a real backup via BackupToSafe, matching the
other restore tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017RtDNHSYWwPKArL8MsRdbj

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-15 20:50:47 +02:00

18 KiB
Raw Blame History

Server Configuration Reference

Complete reference for all OwnCord server configuration options.

Overview

OwnCord server reads configuration from config.yaml in the working directory. On first run, if the file does not exist, a default config.yaml is created automatically.

Configuration is loaded in three layers (later layers override earlier ones):

  1. Built-in defaults (compiled into the binary)
  2. YAML file (config.yaml)
  3. Environment variables (prefix: OWNCORD_)

First-run setup wizard

The admin panel's first-time setup wizard (shown at /admin until the Owner account exists) writes a subset of these keys into config.yaml for you: server.port, server.name, tls.mode, tls.domain, upload.max_size_mb, voice.quality and voice.auto_download_livekit. It also persists the auto-generated voice.livekit_api_key / voice.livekit_api_secret (only when the file has none) so voice tokens survive restarts. The wizard patches the file in place — comments and hand-edited values it doesn't manage are preserved — and restarts the server automatically when a startup-only value changed. Note that OWNCORD_* environment variables still override anything the wizard writes.

Config Key Reference

Server (server)

Key Type Default Description
server.port int 8443 HTTP(S) listen port
server.name string "OwnCord Server" Server display name (shown in /api/v1/info and admin panel)
server.data_dir string "data" Directory for database, certs, uploads, backups
server.max_ws_connections int 0 Cap on concurrently connected WebSocket clients; further upgrades get 503 until connections free up. 0 = unlimited. Every connection costs goroutines and buffered send queues — set a ceiling that matches the host's memory before opening the server to a large community.
server.metrics_allowed_cidrs []string [] Separate allowlist for /api/v1/metrics and the Prometheus /metrics exporter, so a central scraper can be admitted without widening /admin to its network. Empty = falls back to admin_allowed_cidrs.
server.livekit_webhook_allowed_cidrs []string [] Separate allowlist for the LiveKit webhook/health endpoints (which also authenticate cryptographically) — an externally-hosted LiveKit's IP goes here, not in the admin allowlist. Empty = falls back to admin_allowed_cidrs.
server.allowed_origins string[] [] WebSocket CORS allowed origins for web/browser clients; empty list DENIES all cross-origin (set to ["*"] to allow any origin). The OwnCord desktop client needs no entry here — its webview origins (http(s)://tauri.localhost, tauri://localhost) are always accepted.
server.trusted_proxies string[] [] CIDRs of trusted reverse proxies (for X-Forwarded-For)
server.admin_allowed_cidrs string[] private networks CIDRs allowed to access /admin routes. Default: 127.0.0.0/8, ::1/128, 10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16, fc00::/7
server.waf_enabled bool false Enable the Coraza WAF middleware (inline rules + OWASP Core Rule Set)
server.waf_paranoia_level int 2 OWASP CRS paranoia level 14; values outside that range fall back to 2
server.waf_crs_mode string "detect" CRS layer mode: off (inline rules only), detect (matches logged, never blocks), block (anomaly-scoring blocking). Unknown values fall back to detect.

TLS (tls)

Key Type Default Description
tls.mode string "self_signed" TLS mode: self_signed, acme, manual, off
tls.cert_file string "data/cert.pem" Path to TLS certificate (used by manual and self_signed)
tls.key_file string "data/key.pem" Path to TLS private key
tls.domain string "" Domain for ACME/Let's Encrypt (required when mode: acme)
tls.acme_cache_dir string "data/acme_certs" Directory for cached Let's Encrypt certificates

Database (database)

Key Type Default Description
database.type string "sqlite" Database backend. sqlite (or empty) is the only supported value — any other value makes the server refuse to start.
database.path string "data/chatserver.db" Path to SQLite database file
database.max_readers int 0 Bound on the read-only connection pool. 0 = automatic (max(4, CPU count)); clamped to 164. Readers beyond the CPU count mostly buy queueing, not throughput.

Backups (backup)

Key Type Default Description
backup.dir string "data/backups" Directory where database backups are written and pruned. Point it at another disk or an off-host mount so backups don't share a single point of failure with the live database. The admin panel's Backup Schedule and Retention settings operate on this directory.

Security (security)

Key Type Default Description
security.auth_rate_limit_multiplier float 1.0 Scales the per-IP auth rate limits and failure thresholds (registration, login, TOTP, sensitive endpoints). The defaults assume roughly one person per IP; raise this for communities behind a shared NAT (office, school). Clamped to 0.1100.

Uploads (upload)

Key Type Default Description
upload.max_size_mb int 100 Maximum file upload size in megabytes
upload.storage_dir string "data/uploads" Directory where uploaded files are stored

Voice / LiveKit (voice)

Key Type Default Description
voice.livekit_api_key string (random per run) LiveKit API key. Set a stable value for persistent voice tokens.
voice.livekit_api_secret string (random per run) LiveKit API secret (min 32 chars). Set a stable value for persistent tokens.
voice.livekit_url string "ws://localhost:7880" LiveKit server WebSocket URL
voice.livekit_binary string "" Path to an existing livekit-server binary; set to skip auto-download and run your own build
voice.auto_download_livekit bool false (compiled) / true in the generated config When no livekit_binary is set, download a pinned livekit-server release from the official LiveKit GitHub releases (verified against the release checksums.txt) into data/livekit/ and run it automatically
voice.livekit_version string "" Override the pinned livekit-server version used by auto-download (e.g. "1.13.5"); empty = built-in pin
voice.node_ip string "" Public IP for WebRTC ICE candidates; empty = auto-detect. Required for remote users behind NAT.
voice.advertise_internal_ip bool false Also advertise internal (LAN) IPs as ICE candidates. Enable when the server is reachable via both a LAN IP and a public IP so local-network clients can connect to voice.
voice.quality string "medium" Voice quality preset: low, medium, high

Warning: If livekit_api_key or livekit_api_secret are left empty, random credentials are generated on each startup. This means voice tokens break on restart. Always set stable credentials in production. See LiveKit Setup for details.

Server with both a LAN and a public IP

If your server is dual-homed (e.g. 192.168.1.10 on the LAN and 47.x.x.x public), set voice.node_ip to the public IP and voice.advertise_internal_ip: true. LiveKit then advertises the LAN address in addition to the public one, so clients on the local network connect directly while remote clients use the public IP.

For LiveKit options OwnCord does not model, you can take ownership of the auto-started server's config: edit data/livekit.yaml and delete the header line containing the auto-generated marker — OwnCord will stop regenerating the file on startup (your keys: entry must still match voice.livekit_api_key / voice.livekit_api_secret).

GitHub / Updates (github)

Key Type Default Description
github.token string "" Optional GitHub API token for higher rate limits on update checks (5000 req/hr vs 60)
github.owner string "J3vb" Owner of the GitHub repository server and client updates are fetched from
github.repo string "OwnCord" Repository whose releases carry update assets. Must stay publicly readable — both the server self-update and the client auto-update chain fetch release assets from it

Event Persistence (event_persistence)

Controls the tiered event log used for WebSocket reconnection replay. When enabled, missed events are stored in the database so clients that reconnect after the in-memory ring buffer window (replay_ring_size events) can still replay missed events from the DB tier.

Key Type Default Description
event_persistence.enabled bool true Enable cold-storage event persistence. When false, only the in-memory ring buffer is used (lower durability).
event_persistence.retention_hours int 24 How long persisted events are kept before the pruner deletes them
event_persistence.batch_size int 50 Maximum events per database flush
event_persistence.batch_flush_ms int 100 Maximum delay between flushes (milliseconds)
event_persistence.pruner_interval_minutes int 60 How often the pruner goroutine wakes up to delete expired events
event_persistence.replay_ring_size int 1000 Capacity of the in-memory reconnect replay ring. Larger rings bridge longer disconnects without touching the database, at ~1 message payload of memory per slot.
event_persistence.replay_cold_limit int 5000 Maximum persisted events a single reconnect may replay; a larger gap falls back to a full resync. Watch the reconnect_tier_full metric before raising it.

Telemetry / OpenTelemetry (telemetry)

Controls the OpenTelemetry SDK. Requires building with -tags otel (see Contributing). When disabled, the server uses no-op tracer/meter providers; the legacy JSON /api/v1/metrics endpoint exists regardless of this setting (it is admin-IP-restricted, like all metrics surfaces).

Key Type Default Description
telemetry.enabled bool false Enable the OTel SDK
telemetry.exporter string "none" Exporter backend: none, prometheus, otlp
telemetry.otlp_endpoint string "" gRPC endpoint for the OTLP exporter (e.g. localhost:4317). Only used when exporter: otlp.
telemetry.otlp_insecure bool false Disable TLS for the OTLP gRPC connection. Only set true in development / private-network deployments.
telemetry.service_name string "owncord-server" OTel service.name resource attribute

Local development: Run make otel-up (from Server/) to start Jaeger + Prometheus via Docker. Jaeger UI: http://localhost:16686 — Prometheus UI: http://localhost:9090

Plugins (plugins)

Controls the Wazero WASM plugin runtime. Requires building with -tags wazero. When disabled, no plugins are loaded; plugin admin lifecycle endpoints return 503 Service Unavailable and the plugin list endpoint returns an empty list.

Key Type Default Description
plugins.enabled bool false Enable plugin loading at startup
plugins.directory string "data/plugins" Directory scanned for plugin packages on startup
plugins.max_memory_mb int 64 Maximum WASM linear memory per plugin (megabytes)
plugins.cpu_budget_ms int 100 Maximum CPU time per plugin invocation (milliseconds)
plugins.http_allowlist string[] [] Host suffixes plugins may reach via the host_http capability (e.g. ["api.steampowered.com"]). Empty = no outbound HTTP.

GIF Picker (gif)

Powers the client's GIF picker. The server proxies the Klipy API so the key never ships in the desktop bundle — the client only ever calls /api/v1/gif/* on its own server.

Disabled by default. With no gif.api_key set, /api/v1/gif/* returns 503 GIF_DISABLED and clients hide their GIF button. Nothing else changes.

Key Type Default Description
gif.api_key string "" Klipy API key. Get one at partner.klipy.com. Empty = feature off.

Treat this as a credential. Prefer OWNCORD_GIF_API_KEY (or a secrets manager) over writing it into config.yaml, and rotate it if it has ever been exposed to a client build.

Logging (logging)

Controls server log verbosity. The level gates both stdout and the in-memory ring buffer that backs the admin panel's live log view.

Key Type Default Description
logging.level string "info" Minimum level logged: debug, info, warn, error. Empty = info; an unrecognised value falls back to info with a startup warning.

Environment Variable Overrides

Every config key can be overridden via environment variables using the prefix OWNCORD_.

Format: OWNCORD_<SECTION>_<KEY> — the first _ after the prefix maps to the section/key dot; the scheme covers every key in the file, including ones absent from the table below (it is a representative subset, not the full list).

Environment Variable Config Path
OWNCORD_SERVER_PORT server.port
OWNCORD_SERVER_NAME server.name
OWNCORD_SERVER_DATA_DIR server.data_dir
OWNCORD_DATABASE_PATH database.path
OWNCORD_TLS_MODE tls.mode
OWNCORD_TLS_CERT_FILE tls.cert_file
OWNCORD_TLS_DOMAIN tls.domain
OWNCORD_UPLOAD_MAX_SIZE_MB upload.max_size_mb
OWNCORD_UPLOAD_STORAGE_DIR upload.storage_dir
OWNCORD_VOICE_LIVEKIT_API_KEY voice.livekit_api_key
OWNCORD_VOICE_LIVEKIT_API_SECRET voice.livekit_api_secret
OWNCORD_VOICE_LIVEKIT_URL voice.livekit_url
OWNCORD_VOICE_NODE_IP voice.node_ip
OWNCORD_VOICE_ADVERTISE_INTERNAL_IP voice.advertise_internal_ip
OWNCORD_VOICE_QUALITY voice.quality
OWNCORD_GITHUB_TOKEN github.token
OWNCORD_EVENT_PERSISTENCE_ENABLED event_persistence.enabled
OWNCORD_EVENT_PERSISTENCE_RETENTION_HOURS event_persistence.retention_hours
OWNCORD_TELEMETRY_ENABLED telemetry.enabled
OWNCORD_TELEMETRY_EXPORTER telemetry.exporter
OWNCORD_TELEMETRY_OTLP_ENDPOINT telemetry.otlp_endpoint
OWNCORD_TELEMETRY_SERVICE_NAME telemetry.service_name
OWNCORD_PLUGINS_ENABLED plugins.enabled
OWNCORD_PLUGINS_DIRECTORY plugins.directory
OWNCORD_GIF_API_KEY gif.api_key
OWNCORD_SERVER_WAF_ENABLED server.waf_enabled
OWNCORD_DATABASE_TYPE database.type
OWNCORD_TELEMETRY_OTLP_INSECURE telemetry.otlp_insecure
OWNCORD_LOGGING_LEVEL logging.level

Example config.yaml

# OwnCord Server Configuration
server:
  port: 8443
  name: "OwnCord Server"
  data_dir: "data"
  allowed_origins: []             # empty = deny all cross-origin; set to ["*"] to allow any
  trusted_proxies: []              # e.g. ["10.0.0.0/8"] if behind a reverse proxy
  admin_allowed_cidrs:
    - "127.0.0.0/8"
    - "::1/128"
    - "10.0.0.0/8"
    - "172.16.0.0/12"
    - "192.168.0.0/16"

database:
  path: "data/chatserver.db"

tls:
  mode: "self_signed"              # self_signed | acme | manual | off
  cert_file: "data/cert.pem"
  key_file: "data/key.pem"
  domain: ""                       # required for acme mode
  acme_cache_dir: "data/acme_certs"

upload:
  max_size_mb: 100
  storage_dir: "data/uploads"

voice:
  livekit_api_key: "your-api-key"
  livekit_api_secret: "your-secret-at-least-32-characters-long"
  livekit_url: "ws://localhost:7880"
  livekit_binary: ""               # path to livekit-server binary
  node_ip: ""                      # public IP for remote users behind NAT
  advertise_internal_ip: false     # also advertise LAN IPs (dual-homed servers)
  quality: "medium"                # low | medium | high

github:
  token: ""                        # optional GitHub PAT for update check rate limits
  owner: "J3vb"                    # update source repo owner
  repo: "OwnCord"                  # repo holding release assets (binaries + source snapshots)

# Event persistence (tiered reconnect replay)
event_persistence:
  enabled: true
  retention_hours: 24
  batch_size: 50
  batch_flush_ms: 100
  pruner_interval_minutes: 60

# OpenTelemetry (requires build tag: -tags otel)
telemetry:
  enabled: false
  exporter: "none"                 # none | prometheus | otlp
  otlp_endpoint: ""                # e.g. "localhost:4317" for OTLP gRPC
  service_name: "owncord-server"

# Plugin runtime (requires build tag: -tags wazero)
plugins:
  enabled: false
  directory: "data/plugins"
  max_memory_mb: 64
  cpu_budget_ms: 100
  http_allowlist: []               # host suffixes plugins may reach, e.g. ["api.steampowered.com"]

# GIF picker (server-side Klipy proxy). Empty key = feature off.
# Prefer OWNCORD_GIF_API_KEY over storing the key in this file.
gif:
  api_key: ""

# Logging. "level" gates what is logged, to stdout and the admin panel's live
# log view alike. Override without editing this file via OWNCORD_LOGGING_LEVEL.
logging:
  level: "info"                    # debug | info | warn | error

See Also