Skip to content

fix(chatCore): stop executeWithUpstreamStartTimeout leaking its abortPromise listener (hedge-cancelled process exit) - #12406

Merged
diegosouzapw merged 7 commits into
diegosouzapw:release/v3.8.51from
Beexly:fix/upstream-timeouts-abort-listener-leak
Sep 18, 2026
Merged

diegosouzapw merged 7 commits into
diegosouzapw:release/v3.8.51from
Beexly:fix/upstream-timeouts-abort-listener-leak

Conversation

@Beexly

@Beexly Beexly commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

Summary

Fixes the Error [AbortError]: hedge-cancelled process exit (exit code 7) seen on a production instance on 2026-08-31, and hardens the process crash guard around it.

Root cause (mapped from the built chunks back to source and reproduced): executeWithUpstreamStartTimeout in open-sse/handlers/chatCore/upstreamTimeouts.ts registered a { once: true } abort listener for abortPromise on the long-lived client/stream signal and never removed it in finally, so every executor attempt (and every retry) leaked one listener. Promise.race only subscribes once the array literal has been evaluated; when execute() threw synchronously the race never ran, abortPromise was orphaned, and the next hedge cancellation / client disconnect aborted the signal with the string reason that streamHandler.ts forwards. createAbortError() rebuilt it as an AbortError-named Error and rejected a promise nothing awaited. That unhandledRejection reached src/shared/utils/httpClientAbortGuard.mjs, whose handler re-threw it as an uncaughtException, and the process exited.

Commits

  1. fix(sse): stop mergeAbortSignals from leaking abort listeners - the listener-hygiene change from fix(sse): stop mergeAbortSignals from leaking abort listeners #12391 (open-sse/executors/base.ts), included so this branch is self-contained. Note: that change is correct but is not on the crash path above.
  2. fix(server): stop the crash guard re-throwing combo abort reasons - isClientAbortError() now absorbs AbortErrors and the combo abort reasons from comboAbortReasons.ts (hedge-cancelled, combo-per-model-timeout). Ports the upstream guard tests and makes the child-process tests Windows-portable (a file:// URL instead of a bare path in dynamic import()).
  3. fix(chatCore): stop executeWithUpstreamStartTimeout leaking its abortPromise listener - the root-cause fix: keep a handle to the listener and remove it with the others in finally; mark the two race-loser promises as handled so a synchronous throw from execute() can never orphan them. Race semantics are unchanged.
  4. fix(server): absorb raw string abort reasons in the crash guard; document the verified crash path - streamHandler.ts aborts with raw string reasons and undici rejects with signal.reason verbatim, so a cancellation can reach process level as a bare string; the guard now absorbs those too. Comments corrected to describe the verified mechanism.

Tests

  • New regression tests in tests/unit/chatcore-upstream-timeouts.test.ts (+2; both fail against the previous implementation: one asserts no listener growth after a resolving execute, one asserts that a synchronously throwing execute followed by a late hedge-cancelled abort produces no unhandledRejection).
  • tests/unit/httpClientAbortGuard.test.mjs (+9), including a real child-process replay of the production shape on both the uncaughtException and unhandledRejection routes, and a genuine-error case that must still crash.
  • Results: httpClientAbortGuard 18/18, chatcore-upstream-timeouts 7/7, executor-base-utils 23/23; 127 related unit test files 968/970 (the 2 failures reproduce identically on the base commit: a stale G13 golden and a Windows temp-dir EPERM on cleanup).
  • tsc --noEmit -p open-sse/tsconfig.json clean; eslint clean on the changed files.

Related

  • Extends fix(sse): stop mergeAbortSignals from leaking abort listeners #12391, which on its own does not fix the exit.
  • Separately observed on Windows and not addressed here: scripts/build/bootstrap-env.mjs resolves DATA_DIR to %APPDATA%\omniroute while src/lib/dataPaths.ts prefers a legacy ~/.omniroute when it exists, so the first npm start of a source checkout can auto-generate a new STORAGE_ENCRYPTION_KEY and orphan every previously encrypted credential. Happy to file it as an issue.

🤖 Generated with Claude Code

Beexly and others added 4 commits September 1, 2026 21:25
mergeAbortSignals() attached "abort" listeners to its primary/secondary
signals but never removed them once the merged signal settled. Every
executor fetch attempt calls this (fetchWithStartTimeout, once per
URL/retry), so a busy combo request accumulated one live listener per
call on the long-lived combo/client signal. A leaked listener still
fires when that signal is later aborted (e.g. a hedge cancellation
arriving after this merge's own caller already finished), for a merged
output nothing is watching anymore.

Mirrors the already-correct self-cleaning pattern in
open-sse/utils/directResponseStartTimeout.ts's local mergeAbortSignals.

Regression test measures listener growth across repeated merges of the
same long-lived signal: 25 merges leaked exactly 25 listeners pre-fix,
0 post-fix.

(cherry picked from commit 0796965)
Production crash 2026-08-31 (omniroute.log): on a client disconnect,
handleDisconnect aborted the combo controller and a late abort listener
threw the abort reason on an empty stack:

    Error [AbortError]: hedge-cancelled
        at ... AbortController.abort ... handleDisconnect
    file:///.../src/shared/utils/httpClientAbortGuard.mjs:130  throw err;

isClientAbortError() only knew Node's stream codes and "aborted", so
shouldSwallowUncaught() said false and the guard re-threw, taking the
whole server down.

- Port upstream's AbortError line (name "AbortError" + abort-flavoured
  message) so request_signal_aborted / DOMException aborts are absorbed.
- Add an exact-message match for the combo abort reasons from
  open-sse/services/combo/comboAbortReasons.ts ("hedge-cancelled",
  "combo-per-model-timeout"), name-agnostic because the raw reason is a
  plain Error that only gets name="AbortError" stamped on the way out.
  A losing hedge / stalled target is never a server fault. Inlined so
  this .mjs stays dependency-free for scripts/dev/run-next.mjs.

Tests: port upstream's guard tests, add the exact crash shape, a
child-process replay of the crash (dies pre-fix, survives post-fix), a
genuine-error case that must still crash, and a sync check against
comboAbortReasons.ts. The child-process helper passes a file:// URL, not
a bare path, so the tests run on Windows.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 90c9bce)
…Promise listener

Root cause of the 2026-08-31 production exit (Error [AbortError]:
hedge-cancelled), verified by mapping the crash frames in
.build/next/server/chunks/13721.js back to this file:

- The abortPromise abort listener registered on the long-lived client /
  stream signal was never removed in the finally block (only abortListener
  and timeoutAbortListener were), so every executor attempt (and every
  retry) leaked one listener onto that signal.
- Promise.race only subscribes to abortPromise/timeoutPromise once the
  array literal has been evaluated. When execute() threw synchronously the
  race never ran, abortPromise was orphaned, and the next hedge
  cancellation / client disconnect aborted the signal with the string
  reason streamHandler.ts forwards; createAbortError() rebuilt it as an
  AbortError-named Error and rejected a promise nothing awaited. That
  unhandledRejection reached the process crash guard, which re-threw it as
  an uncaughtException and exited with code 7.

Keep a handle to the listener and remove it with the others, and mark the
two race-loser promises as handled so a synchronous throw from execute()
can never orphan them. Race semantics are unchanged (the race still
observes their rejections).

Regression tests: (1) a resolving execute leaves the listener count on
the client signal unchanged; (2) a synchronously throwing execute leaks
no listener and a later abort with the string "hedge-cancelled" produces
no unhandledRejection. Both fail against the previous implementation.

Note: commit 0796965 (mergeAbortSignals cleanup) is correct listener
hygiene but is not on this crash path; this is the fix for the incident.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit e68a50a)
…ment the verified crash path

Follow-ups from the adversarial review of 90c9bce:

- open-sse/utils/streamHandler.ts aborts the stream controller with a raw
  string reason (getClientAbortReason / handleDisconnect) and undici
  rejects with signal.reason verbatim, so a cancellation can reach
  process level as a bare string. isClientAbortError() returned false for
  every non-object, which would still have exited the process. Absorb the
  combo abort reasons and the stream-handler disconnect reasons when they
  arrive as strings.
- Correct the mechanism comment: the 2026-08-31 exit was a leaked
  upstreamTimeouts.ts abortPromise listener rejecting a promise nothing
  awaited (unhandledRejection), escalated by this guard, not a listener
  throwing synchronously. The leak is fixed at the source in the previous
  commit; this guard remains the last-resort net.
- Reword the inlining rationale (plain node launcher, no reliance on
  type-stripping for the .ts constants module).
- Tests: the child-process replay now also exercises the
  unhandledRejection route with the exact production error shape and with
  raw string reasons; add unit coverage for string reasons and non-object
  look-alikes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 696fcc8)
@Beexly
Beexly requested a review from diegosouzapw as a code owner September 2, 2026 02:58
…am-timeouts-abort-listener-leak

Keep the abort-listener leak fix (executeWithUpstreamStartTimeout +
mergeAbortSignals) and the crash-guard combo/string abort absorption on
top of release/v3.8.51. Extract mergeAbortSignals out of frozen
open-sse/executors/base.ts so the file-size cap is not raised.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
@diegosouzapw

Copy link
Copy Markdown
Owner

Solid fix for the abort-listener leak, and the extraction of mergeAbortSignals into its own
module with an explicit cleanup() is a real architectural improvement over the inline version.
Verified 47/47 tests pass across the three touched test files. This looks merge-ready. Since this
supersedes your own earlier #12391 with the same fix in a cleaner shape, we'll close that one in
favor of this PR — let us know if you'd rather we merge #12391 instead, but this version is the
more complete one.

Resolves the conflict in open-sse/executors/base.ts — kept the base's
single-line `cliFingerprints` import with its `// prettier-ignore` marker
(the branch had only reformatted that line); the mergeAbortSignals
extraction into ./base/mergeAbortSignals.ts merged cleanly.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
@diegosouzapw

Copy link
Copy Markdown
Owner

Thanks @Beexly — merging via the release merge-train. Validated in local merge-train /tmp/mt-train3b.log on 192.168.0.113 @ train tip 408e2128791696a18966681953136fc96aeb99b0 (32 PRs boarded): static gates green; full test:unit 40694 tests, 17 failing — every one reproduces on the pure release tip (base-red sweep list), zero new reds. Merged --admin per merge-gates §4/§7.

Re-sync onto the current tip — both sides appended tests at the end of tests/unit/httpClientAbortGuard.test.mjs (this PR's raw-string abort reasons, the base's diegosouzapw#14064 full-error logging); kept both, and restored the fileURLToPath import the auto-merge dropped.

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
@diegosouzapw

Copy link
Copy Markdown
Owner

Thanks @Beexly — merging via the release merge-train. Validated in local merge-train (merge-train-20260918-130526-suite.log) on the devbox @ train tip 28fb420c9ad860ba275294ebb4ebbecb9da83318 with the sibling PRs of this batch: static gates green; changed-area node:test 293/294 (0 failing) + vitest 482/482 (fast parity — full suite ran today on the tip via the base-red and 3b trains). Merged --admin per merge-gates §7.

@diegosouzapw
diegosouzapw merged commit 706dc75 into diegosouzapw:release/v3.8.51 Sep 18, 2026
3 checks passed
SCys pushed a commit to SCys/OmniRoute that referenced this pull request Sep 25, 2026
…no live body

Root cause: diegosouzapw#14342 changed executeWithUpstreamStartTimeout to keep the
client-signal `abortListener` (the link that aborts combinedController)
whenever the start-timeout race settled successfully, so a client abort
after headers could still reach the upstream fetch. It kept the link for
every resolved result, including results that carry no streaming body, so
one listener stayed on the client signal after the race settled. That
broke the diegosouzapw#12406 guard ("every listener registered for the race must be
removed once it settles", 1 !== 0).

Fix: keep the link only when the settled result (a Response or the
executor's `{ response }` wrapper) still has an unconsumed body that will
stream on the combined signal. Otherwise remove it in the finally, as for
failed attempts. The abortPromise listener is still always removed. The
diegosouzapw#14342 post-headers propagation is unchanged for streaming responses.

Tests: the diegosouzapw#12406 guard is green again; two new cases pin both sides
(a body-less Response releases the link; a streaming body keeps exactly
one link, propagates the abort, and it self-removes on abort).

Refs diegosouzapw#14547
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…Promise listener (hedge-cancelled process exit) (diegosouzapw#12406)

* fix(sse): stop mergeAbortSignals from leaking abort listeners

mergeAbortSignals() attached "abort" listeners to its primary/secondary
signals but never removed them once the merged signal settled. Every
executor fetch attempt calls this (fetchWithStartTimeout, once per
URL/retry), so a busy combo request accumulated one live listener per
call on the long-lived combo/client signal. A leaked listener still
fires when that signal is later aborted (e.g. a hedge cancellation
arriving after this merge's own caller already finished), for a merged
output nothing is watching anymore.

Mirrors the already-correct self-cleaning pattern in
open-sse/utils/directResponseStartTimeout.ts's local mergeAbortSignals.

Regression test measures listener growth across repeated merges of the
same long-lived signal: 25 merges leaked exactly 25 listeners pre-fix,
0 post-fix.

(cherry picked from commit 0796965)

* fix(server): stop the crash guard re-throwing combo abort reasons

Production crash 2026-08-31 (omniroute.log): on a client disconnect,
handleDisconnect aborted the combo controller and a late abort listener
threw the abort reason on an empty stack:

    Error [AbortError]: hedge-cancelled
        at ... AbortController.abort ... handleDisconnect
    file:///.../src/shared/utils/httpClientAbortGuard.mjs:130  throw err;

isClientAbortError() only knew Node's stream codes and "aborted", so
shouldSwallowUncaught() said false and the guard re-threw, taking the
whole server down.

- Port upstream's AbortError line (name "AbortError" + abort-flavoured
  message) so request_signal_aborted / DOMException aborts are absorbed.
- Add an exact-message match for the combo abort reasons from
  open-sse/services/combo/comboAbortReasons.ts ("hedge-cancelled",
  "combo-per-model-timeout"), name-agnostic because the raw reason is a
  plain Error that only gets name="AbortError" stamped on the way out.
  A losing hedge / stalled target is never a server fault. Inlined so
  this .mjs stays dependency-free for scripts/dev/run-next.mjs.

Tests: port upstream's guard tests, add the exact crash shape, a
child-process replay of the crash (dies pre-fix, survives post-fix), a
genuine-error case that must still crash, and a sync check against
comboAbortReasons.ts. The child-process helper passes a file:// URL, not
a bare path, so the tests run on Windows.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 90c9bce)

* fix(chatCore): stop executeWithUpstreamStartTimeout leaking its abortPromise listener

Root cause of the 2026-08-31 production exit (Error [AbortError]:
hedge-cancelled), verified by mapping the crash frames in
.build/next/server/chunks/13721.js back to this file:

- The abortPromise abort listener registered on the long-lived client /
  stream signal was never removed in the finally block (only abortListener
  and timeoutAbortListener were), so every executor attempt (and every
  retry) leaked one listener onto that signal.
- Promise.race only subscribes to abortPromise/timeoutPromise once the
  array literal has been evaluated. When execute() threw synchronously the
  race never ran, abortPromise was orphaned, and the next hedge
  cancellation / client disconnect aborted the signal with the string
  reason streamHandler.ts forwards; createAbortError() rebuilt it as an
  AbortError-named Error and rejected a promise nothing awaited. That
  unhandledRejection reached the process crash guard, which re-threw it as
  an uncaughtException and exited with code 7.

Keep a handle to the listener and remove it with the others, and mark the
two race-loser promises as handled so a synchronous throw from execute()
can never orphan them. Race semantics are unchanged (the race still
observes their rejections).

Regression tests: (1) a resolving execute leaves the listener count on
the client signal unchanged; (2) a synchronously throwing execute leaks
no listener and a later abort with the string "hedge-cancelled" produces
no unhandledRejection. Both fail against the previous implementation.

Note: commit 0796965 (mergeAbortSignals cleanup) is correct listener
hygiene but is not on this crash path; this is the fix for the incident.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit e68a50a)

* fix(server): absorb raw string abort reasons in the crash guard; document the verified crash path

Follow-ups from the adversarial review of 90c9bce:

- open-sse/utils/streamHandler.ts aborts the stream controller with a raw
  string reason (getClientAbortReason / handleDisconnect) and undici
  rejects with signal.reason verbatim, so a cancellation can reach
  process level as a bare string. isClientAbortError() returned false for
  every non-object, which would still have exited the process. Absorb the
  combo abort reasons and the stream-handler disconnect reasons when they
  arrive as strings.
- Correct the mechanism comment: the 2026-08-31 exit was a leaked
  upstreamTimeouts.ts abortPromise listener rejecting a promise nothing
  awaited (unhandledRejection), escalated by this guard, not a listener
  throwing synchronously. The leak is fixed at the source in the previous
  commit; this guard remains the last-resort net.
- Reword the inlining rationale (plain node launcher, no reliance on
  type-stripping for the .ts constants module).
- Tests: the child-process replay now also exercises the
  unhandledRejection route with the exact production error shape and with
  raw string reasons; add unit coverage for string reasons and non-object
  look-alikes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 696fcc8)

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
Co-authored-by: Beexly <Beexly@users.noreply.github.com>
diegosouzapw added a commit that referenced this pull request Sep 29, 2026
…leak, MCP bundle deadlocks, sidebar keys, flush-empty-retry, proxy-status and pack-policy tests (#14820)

Fixes the reds still present on the tip: client-abort listener leak (#14342 vs #12406), MCP bundle deadlock (quotaCache init cycle via the quotaCacheState leaf, plus the new #14892 proxyLogs -> upstreamStatusCapture -> usage/migrations cycle via the zero-import timingMs leaf), sidebar Model catalog keys, and the stale flush-empty-retry / proxy-status / pack-policy tests. Combo pre-content retry and zh-TW glossary were already fixed on the tip and dropped. Merged tree: 102/102 focused incl. mcp-bundle-startup (red on the tip), typecheck:core clean, open-sse typecheck 0, check:cycles OK, file-size OK.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants