mirror of
https://github.com/J3vb/OwnCord.git
synced 2026-09-02 19:43:10 +03:00
* feat(server): supervisor detection and server.restart_mode config key RunningUnderSupervisor detects systemd (INVOCATION_ID) and, best-effort, NSSM (NSSM_SERVICE_NAME — 2.24 does not set it, so NSSM deployments set the mode explicitly). server.restart_mode (auto|spawn|supervised, default auto, env OWNCORD_SERVER_RESTART_MODE) selects how a self-restart hands off after the server drains: exit for the supervisor to relaunch, or spawn the replacement directly. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ngzj2Rx9UGC35uLHAfErMp * fix(server): make the self-restart handoff drain fully before starting the successor The update/restore/wizard restart previously spawned the replacement while the old server was still serving, then SIGTERMed itself and hard-exited after 10s. That design failed in every documented deployment mode: under the shipped systemd unit the spawned child (same cgroup) was killed when the old main process exited and Restart=on-failure never relaunched a clean exit; on Windows the self-SIGTERM is unsupported and silently dropped, so graceful shutdown never ran — hub.GracefulStop (the only caller of LiveKitProcess.Stop) was skipped, orphaning livekit-server on TCP 7880/UDP 50000-60000 and dropping queued event/audit batches; and NSSM's relaunch raced the self-spawned replacement for the database lock. Admin handlers now perform only the on-disk swap and request a restart through an injected hook (admin.SetRestartHandoff). The main package's restart coordinator cancels the parent of run()'s signal.NotifyContext — the exact drain a SIGTERM triggers, on every platform — and after run() has fully torn down (listeners closed, hub and LiveKit stopped, queues flushed, DB closed and its lock released) main() performs the handoff: spawn the replacement in spawn mode, or exit 0 for the supervisor in supervised mode. A 90s backstop force-exits a wedged teardown; the DB-lock and bind retries demote to safety nets. A three-state guard (idle/busy/restart-pending) serializes update apply, backup restore, and setup-wizard restarts against each other: concurrent applies no longer race the same staged .new file or broadcast a spurious update_aborted, and conflicting requests get 409 UPDATE_IN_PROGRESS / RESTART_PENDING. The swap being free of process side effects also makes the apply success path unit-testable for the first time. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ngzj2Rx9UGC35uLHAfErMp * fix(server): errno-based bind-conflict detection, ACME bind retry, LiveKit Pdeathsig isAddrInUse now unwraps to the platform errno (EADDRINUSE; WSAEADDRINUSE 10048 on Windows) with the English strings kept only as fallback — the string-only match never fired on localized Windows, silently disabling the bind retry. The retry loop is extracted into serveWithBindRetry and now also covers the ACME :80 challenge server, which previously gave up on first conflict and stayed dead (breaking HTTP-01 renewals) until the next restart. The .old-binary boot cleanup retries briefly for the window where a spawn-mode predecessor has not fully exited. The companion livekit-server gets Pdeathsig SIGKILL on Linux so a parent killed without teardown (kill -9, OOM, backstop exit) cannot orphan it with the voice ports held. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ngzj2Rx9UGC35uLHAfErMp * docs(deploy): Restart=always unit and per-supervisor restart-mode guidance Restart=always is what lets the deliberate clean exit after a self-update/restore relaunch under systemd (systemctl stop is never auto-restarted; failure exits behave as before). Deployment docs gain the required NSSM AppEnvironmentExtra line, the Task Scheduler and Docker restart-policy notes, and the new drain-then-handoff update flow. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ngzj2Rx9UGC35uLHAfErMp --------- Co-authored-by: Claude <noreply@anthropic.com>
57 lines
2.3 KiB
Desktop File
57 lines
2.3 KiB
Desktop File
# OwnCord server — systemd unit template.
|
|
#
|
|
# Install:
|
|
# 1. Create a service user and install directory:
|
|
# sudo useradd --system --home /opt/owncord --shell /usr/sbin/nologin owncord
|
|
# sudo mkdir -p /opt/owncord && sudo chown owncord:owncord /opt/owncord
|
|
# 2. Place the chatserver binary (from GitHub Releases) in /opt/owncord/.
|
|
# 3. sudo cp deploy/owncord.service /etc/systemd/system/owncord.service
|
|
# 4. sudo systemctl daemon-reload && sudo systemctl enable --now owncord
|
|
#
|
|
# Logs: journalctl -u owncord -f
|
|
|
|
[Unit]
|
|
Description=OwnCord chat server
|
|
After=network-online.target
|
|
Wants=network-online.target
|
|
|
|
[Service]
|
|
Type=simple
|
|
User=owncord
|
|
Group=owncord
|
|
WorkingDirectory=/opt/owncord
|
|
ExecStart=/opt/owncord/chatserver
|
|
# always, not on-failure: the server exits 0 ON PURPOSE after an admin-panel
|
|
# self-update, backup restore, or setup-wizard change, expecting systemd to
|
|
# relaunch it (now running the swapped binary). `systemctl stop` is unaffected
|
|
# — systemd never auto-restarts an explicitly stopped unit. Failure exits
|
|
# (e.g. the WebSocket dispatch panic breaker, exit 1) restart the same as
|
|
# they did under on-failure.
|
|
Restart=always
|
|
RestartSec=3
|
|
# A crash-looping binary still trips systemd's start limit (default 5 starts
|
|
# in 10s) and parks the unit as failed; raise StartLimitIntervalSec /
|
|
# StartLimitBurst in [Unit] if you want a longer leash.
|
|
|
|
# The server drains gracefully on SIGTERM with a 30s budget; give it a little
|
|
# headroom before systemd escalates to SIGKILL.
|
|
TimeoutStopSec=35
|
|
|
|
# ── Hardening ───────────────────────────────────────────────────────────────
|
|
NoNewPrivileges=true
|
|
ProtectSystem=strict
|
|
ProtectHome=true
|
|
PrivateTmp=true
|
|
# The install directory must stay WRITABLE: the admin panel's self-update
|
|
# renames the new binary into place (chatserver -> chatserver.old swap), and
|
|
# the data dir (SQLite, uploads, certs, backups) lives beneath it by default.
|
|
# If you disable self-update and move data_dir elsewhere, narrow this.
|
|
ReadWritePaths=/opt/owncord
|
|
|
|
# Only needed for tls.mode: acme (binds :80 for HTTP-01 challenges) as a
|
|
# non-root user. Harmless otherwise; remove if you prefer.
|
|
AmbientCapabilities=CAP_NET_BIND_SERVICE
|
|
|
|
[Install]
|
|
WantedBy=multi-user.target
|