fix(live-ws): reject on bind failure instead of crashing the process (closes #6324) - #6332
Conversation
…loses diegosouzapw#6324) startLiveDashboardServer() called server.listen() with no error handler. Since 3.8.45 wired the live-WS daemon into the standalone runtime (instrumentation-node.ts), the daemon auto-starts on port 20129 — the same port the API bridge binds when API_PORT=20129 (the documented value). The bind then fails with EADDRINUSE, and with no error handler the failure surfaced as an uncaughtException that crash-looped the whole container (ai gateway returned 503). Attach a one-shot error handler so the bind failure rejects the returned promise instead. The module-level auto-start already has a .catch that logs it non-fatally, so the process now survives — matching the graceful degradation already implemented in apiBridgeServer, embedWsProxy, and startHttpProxyServer. The handler must be on the WebSocketServer (wss), not the underlying http.Server: ws re-emits the server's error via wss.emit(), which throws synchronously when wss has no listener, before any server.on('error') handler could run. Adds a regression test that occupies the port first and asserts startLiveDashboardServer rejects with EADDRINUSE (fails without the fix).
There was a problem hiding this comment.
Code Review
This pull request addresses issue #6324 by modifying startLiveDashboardServer in src/server/ws/liveServer.ts to gracefully handle port bind failures (such as EADDRINUSE) instead of crashing the process. It now returns a Promise that rejects upon encountering a WebSocket server error, cleaning up event subscriptions and clients in the process. Additionally, a new unit test file tests/unit/live-ws-eaddrinuse-6324.test.ts has been added to verify this behavior by simulating a port conflict. No review comments were provided, so I have no feedback to offer.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
|
Merged — thank you, @vinayakkulkarni! Great catch on attaching the error listener to |
…loses diegosouzapw#6324) (diegosouzapw#6332) LiveWS server rejects on bind failure instead of crash-looping the process (closes diegosouzapw#6324). Integrated into release/v3.8.46.
…loses diegosouzapw#6324) (diegosouzapw#6332) LiveWS server rejects on bind failure instead of crash-looping the process (closes diegosouzapw#6324). Integrated into release/v3.8.46.
Summary
The published
omniroutestandalone image (3.8.45) crash-loops on startup whenAPI_PORT=20129is set (the documented value), taking the whole gateway down with a 503. This makesstartLiveDashboardServer()degrade gracefully on a bind failure instead of crashing the process.Root cause
3.8.45 wired the live-dashboard WebSocket daemon into the standalone runtime (the
await import("@/server/ws/liveServer")added toinstrumentation-node.ts, #6202). That import fires the module-level auto-start, which bindsLIVE_WS_PORT(default 20129). But 20129 is also the API bridge's port whenAPI_PORT=20129— so the second listener fails withEADDRINUSE:startLiveDashboardServer()calledserver.listen()with no error handler, so the bind failure surfaced as anuncaughtExceptionand crash-looped the container. This is the opposite of the graceful degradation already implemented inapiBridgeServer,embedWsProxy, andstartHttpProxyServer, which all catchEADDRINUSEand disable just that one aux server.Verified empirically against the published
3.8.45-webimage onlinux/arm64: withOMNIROUTE_ENABLE_LIVE_WS=0the container comes up healthy and serves200/307, confirming the LiveWS bind is the sole fatal fault.Fix
Attach a one-shot error handler that rejects the returned promise on bind failure. The module-level auto-start already has a
.catchthat logs it non-fatally, so the process now survives.The listener is on the
WebSocketServer(wss), not the underlyinghttp.Server:wsre-emits the server'serrorviawss.emit("error", …), which throws synchronously whenwsshas no listener — before anyserver.on("error")handler could run. (Aserver.on("error")handler looks correct but does not catch it; confirmed with a minimal repro.)Related Issues
Validation
npm run lint(eslint on changed files — clean)npm run test:unit(new test + existing live-ws suite: 34/34 pass)npm run test:coverage(not run locally — full-suite coverage; CI will run it)Typecheck (
tsc -p tsconfig.typecheck-core.json) is also clean.Red/green proof: with the fix reverted, the new test fails with a raw
EADDRINUSE; with the fix, it passes.Tests Added Or Updated
tests/unit/live-ws-eaddrinuse-6324.test.ts(new) — occupies the target port first, then assertsstartLiveDashboardServer()rejects withEADDRINUSEinstead of crashing.Coverage Notes
src/server/ws/liveServer.ts. The new test exercises the added bind-failure rejection path; the happy path is already covered bytests/integration/live-ws-startup.test.ts.Reviewer Notes
3.8.45-webimage. A second, non-fatal log-only error also appears in that image on startup —Error: file data stream has unexpected number of bytes at .build/next/server/chunks/_099ro22._.js— but it does not stop the server (verified: container serves200with LiveWS disabled). It looks like a build-artifact issue rather than a source defect, so it's intentionally out of scope for this PR; happy to open a separate issue if useful.