Skip to content

fix(desktop): bound reconnect awaits so a stuck IPC round-trip can't latch the UI frozen - #93476

Closed
chelsealong wants to merge 2 commits into
NousResearch:mainfrom
chelsealong:fix/desktop-bound-reconnect-awaits-93454
Closed

chelsealong wants to merge 2 commits into
NousResearch:mainfrom
chelsealong:fix/desktop-bound-reconnect-awaits-93454

Conversation

@chelsealong

Copy link
Copy Markdown

What does this PR do?

Bounds the three IPC-round-trip awaits in the primary reconnect loop
(attemptReconnect() in use-gateway-boot.ts) with a 20s timeout, so a
stuck main-process call can no longer permanently freeze Desktop's
reconnect backoff.

Problem

Fixes #93454.

Desktop connected to a remote gateway over an SSH tunnel periodically
freezes on reconnect after a transient liveness-probe failure. The UI is
left stuck ("gateway needs setup" / stuck reconnecting) with no error
surfaced, even though the backend itself is confirmed healthy the whole
time (curl .../api/status → 200 OK). Only killing and relaunching the
app recovers it.

attemptReconnect() sets a reconnecting guard, then does:

await desktop.revalidateConnection?.().catch(() => undefined)
const conn = await desktop.getConnection()
...
const wsUrl = await resolveGatewayWsUrl(desktop, conn)
await gateway.connect(wsUrl)

revalidateConnection(), getConnection(), and resolveGatewayWsUrl() are
all IPC calls into the Electron main process with no timeout of their own. If
any stalls (a wedged revalidation after the liveness probe trips, even
though the remote backend answers fine), the await never settles. While it's
pending, reconnecting never clears, so every later scheduleReconnect() /
attemptReconnect() early-returns forever — the backoff loop is latched
and the UI never self-heals.

An earlier version of this PR bounded only getConnection() and
resolveGatewayWsUrl() — review correctly flagged that the PR's own
diagnosis names a "wedged revalidation" as the trigger, and that call
(revalidateConnection()) sat one line above the first bounded await,
untouched, with only a .catch(() => undefined) that swallows a rejection
but does nothing for a promise that never settles. This revision bounds
that call too.

This is the same defect described (and fixed, but never merged) in closed
PR #40008 for a different trigger (no-network suspend); the code has
since grown several more getConnection()/resolveGatewayWsUrl() call
sites, so that patch no longer applies, but the primary automatic
reconnect loop this issue reports against still has no timeout.

Fix

Add a small withTimeout() helper and wrap all three awaits in
attemptReconnect() with a 20s bound — including revalidateConnection(),
which keeps its original best-effort .catch(() => undefined) (a stall or
a rejection there is still non-fatal to the reconnect attempt; only
getConnection()/resolveGatewayWsUrl() failing aborts it). On timeout
each await rejects, the existing catch/finally clears the
reconnecting guard, and scheduleReconnect() resumes the backoff — so the
UI recovers on its own once the gateway is reachable again, matching the
reporter's expected behavior. gateway.connect() already has its own
separate 15s connect timeout and is unchanged.

How to Test

  1. Connect Desktop to a remote gateway (SSH local port-forward).
  2. Force the reconnect path's revalidateConnection()/getConnection()/
    resolveGatewayWsUrl() IPC call to hang (e.g. a wedged main-process
    revalidation) while the remote backend itself keeps answering
    /api/status with 200 OK.
  3. Before this fix: the UI stays stuck "reconnecting" indefinitely.
  4. After this fix: within ~20s the stalled call times out, the backoff
    loop resumes, and the app reconnects on its own.

Automated regression tests in use-gateway-boot.test.tsx:

  • one hangs desktop.getConnection() forever (new Promise(() => undefined)) on every reconnect attempt after a drop, and asserts a
    further reconnect attempt still fires after the internal timeout elapses.
  • a second hangs desktop.revalidateConnection() specifically (the call
    named in the bug report and this file's own comment as the trigger,
    getConnection() left fast/unmocked) and asserts getConnection() is
    still reached and the socket reopens once the timeout elapses — this is
    the case review found missing from the first version of this PR.
$ npx vitest run --project ui src/app/gateway/hooks/use-gateway-boot.test.tsx
 Test Files  1 passed (1)
      Tests  21 passed (21)

Confirmed the new revalidateConnection() test fails without the fix
(reverted only the source file via git stash push -- <file>, kept the
test):

AssertionError: expected 1 to be greater than 1
 ❯ src/app/gateway/hooks/use-gateway-boot.test.tsx:520:53
 Tests  1 failed | 20 passed (21)
$ npm run typecheck
> tsc -p . --noEmit && tsc -p tsconfig.electron.json --noEmit && tsc -p tsconfig.e2e.json --noEmit
(clean)

$ npx eslint src/app/gateway/hooks/use-gateway-boot.ts src/app/gateway/hooks/use-gateway-boot.test.tsx
(clean)

$ npx vitest run --project ui src/app/gateway
 Test Files  3 passed (3)
      Tests  30 passed (30)

Checklist

  • I've read the Contribution Guide and repository AGENTS.md.
  • My commit message follows Conventional Commits.
  • The PR contains only the reconnect-timeout fix and its regression test.
  • I've added a behavior-level regression test and confirmed it fails
    without the fix.

AI assistance disclosure

This change was drafted with AI assistance (Claude) and reviewed before
submission.

…latch the UI frozen

After a liveness-probe-triggered reconnect on a remote gateway,
attemptReconnect() awaits desktop.getConnection() and resolveGatewayWsUrl()
with no timeout. If either stalls (e.g. main process wedged mid-revalidation
even though the backend itself is reachable), the `reconnecting` guard never
clears, so every later scheduleReconnect()/attemptReconnect() early-returns
forever and the UI stays stuck in "reconnecting" until the app is restarted.

Bound both awaits with a 20s timeout so a stall rejects instead of hanging;
the existing catch/finally already clears the guard and resumes backoff on
rejection. gateway.connect() keeps its own separate connect timeout.

Fixes NousResearch#93454
…h#93454)

attemptReconnect() awaited desktop.revalidateConnection?.() unbounded,
immediately before the two IPC calls the previous commit wrapped in
withTimeout(). A wedged revalidation after a liveness-probe trip -
the exact trigger NousResearch#93454 and this file's own comment describe - hung
that await forever, so the reconnecting guard never cleared and the
prior fix never got reached.

Wrap it in the same 20s withTimeout() (still swallowing the result via
.catch, matching its existing best-effort semantics) and extend the
regression test to hang revalidateConnection() specifically, proving
getConnection() and the socket still proceed once the stall times out.
@teknium1

Copy link
Copy Markdown
Collaborator

Merged via #93662 with your commits cherry-picked (authorship preserved); we bounded the sibling boot/gateway-switch awaits on top. Prior art credit to @victorftrdba's #40008. Thanks @chelsealong!

@teknium1 teknium1 closed this Aug 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/desktop Electron desktop app (apps/desktop/*) P3 Low — cosmetic, nice to have type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Desktop: remote gateway reconnect hangs indefinitely after liveness-probe trip, requires manual app restart

3 participants