Skip to content

Drain the microtask checkpoint after unhandledRejection handlers run - #33356

Closed
robobun wants to merge 5 commits into
mainfrom
farm/184ecb12/unhandled-rejection-microtask-checkpoint
Closed

robobun wants to merge 5 commits into
mainfrom
farm/184ecb12/unhandled-rejection-microtask-checkpoint

Conversation

@robobun

@robobun robobun commented Jul 5, 2026 •

Copy link
Copy Markdown
Collaborator

Repro

process.on("unhandledRejection", () => {
  // any microtask hop — which is what `await` compiles to
  Promise.resolve().then(() => setTimeout(() => console.log("recovered"), 10));
});
Promise.reject(new Error("x"));
node: recovered
bun:  (nothing, exits 0)

process.nextTick and queueMicrotask are dropped the same way. Scheduling the timer directly in the handler works, so only the hop is affected.

It matters because unhandledRejection handlers are where apps flush a logger, write a crash report, or start a graceful shutdown, and almost all of that code awaits something before scheduling a continuation.

Cause

An unhandledRejection handler is user JS that can queue microtasks and nextTicks. Node loops back into its checkpoint afterwards:

do {
  while ((tock = queue.shift()) !== null) { /* nextTick */ }
  runMicrotasks();
} while (!queue.isEmpty() || processPromiseRejections());

Bun's handleRejectedPromises() emits the event and returns straight to the is_event_loop_alive() check. The handler's microtask is still sitting in the queue, which is_event_loop_alive() does not count, so the loop exits and the work never runs.

A second gap in the same family: auto_tick_active (the timer + poll phase that bun run's main loop and on_before_exit drive) had no rejection handling at its tail at all. A rejection raised by a timer callback was therefore notified only after beforeExit, from the exit path, with no loop iteration left. That dropped even a directly scheduled setTimeout, and emitted unhandledRejection after beforeExit instead of before it.

Fix

EventLoop::drain_rejected_promises(): notify pending rejections, then drain the nextTick + microtask checkpoint they filled, repeating while the drain produces new rejections. It returns Result<(), Stopped> like the rest of the loop-level code on main, so a VM termination met inside a handler propagates instead of being swallowed. Used at the tail of EventLoop::tick_turn, auto_tick, auto_tick_active, and in on_before_exit's drain loop, so whatever a handler schedules is visible before liveness is re-checked.

The new hasPendingRejectedPromises() binding keeps the common path to a single isEmpty() read: when nothing rejected, no drain runs.

Emission order within a batch is unchanged (all handlers, then one checkpoint), matching node.

Verification

Differential against node v26. All of these produce byte-identical output with the fix and were wrong before it:

rejection raised handler schedules before after / node
module top level setTimeout behind .then() nothing recovered
module top level setTimeout behind queueMicrotask nothing queueMicrotask
module top level setTimeout behind nextTick nothing nextTick
setTimeout callback setTimeout behind .then() nothing handler, recovered
setTimeout callback setTimeout directly beforeExit only handler, timer-from-handler, beforeExit
setTimeout callback nothing beforeExit, handler handler, beforeExit
beforeExit listener setTimeout behind .then() nothing handler, recovered, beforeExit#2

Three of those are the new cases in test/js/node/process/process-on.test.ts; they fail on main and pass with the fix.

Suites run against the debug build

No new failures in:

  • test/js/node/process/process-on.test.ts
  • test/js/node/process/process.test.js (including delivers many unhandledRejections in order, which pins batch ordering)
  • test/js/node/process/process-nexttick.test.js
  • test/js/node/test/parallel/test-promises-unhandled-rejections.js and the seven other test-promise-unhandled-* files
  • test/js/web/workers/worker.test.ts
  • test/js/web/timers/{setTimeout,setInterval,setImmediate}.test.js
  • test/js/node/timers/node-timers.test.ts
  • test/js/bun/cron/in-process-cron.test.ts
  • test/js/bun/http/serve.test.ts, test/js/web/streams/streams.test.js, test/js/bun/spawn/spawn-stdin-readable-stream.test.ts

bun run rust:check-all passes on all 10 targets.

Pre-existing failures in this container, unrelated and reproducible without the change: the timer RSS-leak fixtures gate on process.execPath.includes("bun-asan"), which is false for bun-debug, so the 192 MB ASAN threshold never applies (a release build measures a 2.0 MB delta); a few serve.test.ts cases need IPv6 / a non-root uid.

Second fix: captureRejections

Fixing the event loop made test/js/node/test/parallel/test-event-capture-rejections.js fail. That test is a chain of twelve stages, each handing off with process.nextTick from inside an unhandledRejection handler, so on main it silently stopped after its third stage and still exited 0. With the chain repaired the remaining nine stages run for the first time and hit three pre-existing captureRejections gaps, all reproducible on released bun:

bun 1.4.0 node
handler returns a thenable (not a real Promise) rejection dropped silently error event
then getter throws rejection dropped silently error event
inherits(X, EventEmitter) with no super(), global captureRejections = true unhandledRejection error event

In src/js/node/events.ts:

  • addCatch only attached to real promises ($isPromise(result)). Node duck-types then. Reading then is observable (Promises/A+ allows a getter), so it is read once and a throwing getter becomes an error event.
  • emit is swapped per instance in the constructor, so an emitter built by inherits() without a super() call keeps the non-capturing emit regardless of the global flag. The static captureRejections setter now swaps the prototype's emit too.
  • That swap alone would over-capture: an emitter constructed while the flag was off owns kCapture = false and has no own emit, so flipping the flag later would start capturing it, which node does not do. addCatch now re-checks kCapture, as node does.

That prototype swap only ever replaces one of the two internal emit functions. Node's setter leaves emit alone, so a userland patch of EventEmitter.prototype.emit (what APM and tracing libraries install) has to survive a toggle of the flag.

Five tests in test/js/node/events/event-emitter.test.ts; three fail on main, and two are negative tests for the guards above (each fails if its guard is deleted). test-event-capture-rejections.js now runs all twelve stages and exits 0.

Rebase notes

Rebased across a large stretch of main. Two resolutions changed code rather than just context:

  • Stopped model. Main's tick() became tick_turn() -> Result<(), Stopped>, and handle_rejected_promises() / drain_microtasks() now return Results. drain_rejected_promises() adopts the same signature and uses ? throughout; the auto_tick / auto_tick_active / Run::start call sites discard it with let _ = exactly as main does for the call it replaces, and on_before_exit returns on Err, matching the script_allowed() returns main added around it (process: skip 'beforeExit' after a fatal uncaught exception #34639).
  • on_before_exit guards. Main now skips beforeExit after a fatal throw and stops the drain on a termination request. The rejection drain is inserted at the point the loop goes idle, ahead of those guards, so they still gate the re-dispatch.
  • Two dispatch sites in events.ts. node:events: single listeners stored bare like node; copy-on-write arrays #34519 split emitWithRejectionCapture into a single-listener fast path and an array path, each with its own $isPromise check. Auto-merge only updated one; the other is updated by hand so a thenable is captured regardless of listener count. test-event-capture-rejections.js covers both shapes.

Everything else merged cleanly. Verified on the new base: process-on.test.ts + event-emitter.test.ts 86/86, test-event-capture-rejections.js all 12 stages, 56 upstream test-event* / test-promise* files, the 7 node-differential repros byte-identical, rust:check-all 12/12 targets, and the 6 expected fail-before failures against released bun.

Note on #33354

#33354 changes when a rejection is emitted (moving the scan into EventLoop::exit() so it lands before anything the rejecting callback scheduled). This PR changes what happens after the handler runs. They are independent and do not conflict textually, but whichever lands second should have handle_rejected_promises_after_tick() call drain_rejected_promises() rather than global_ref().handle_rejected_promises(), otherwise the checkpoint is skipped for rejections reported from exit().


[review] gate passed · iteration 9 · 11 files touched

fails on main (without fix)
ASAN without fix: 6 FAILED
$ BUN_DEBUG_QUIET_LOGS=1 bun scripts/build.ts --profile=debug --quiet test "--reporter=junit" "--reporter-outfile=/tmp/mechgate.xml" test/js/node/events/event-emitter.test.ts test/js/node/process/process-on.test.ts
bun test v1.4.1 (861e9ae04)

test/js/node/process/process-on.test.ts:
(pass) process.on > when called from the main thread [380.02ms]
 [245ms]  bundle  1 modules
[2.77s] compile  ./out
(pass) process.on > should work inside --compile [3413.51ms]
[macro] call initialize
Bundled 1 module in 451ms

  out.ts  8 bytes  (entry point)

(pass) process.on > should work inside a macro [813.28ms]
144 |       env: bunEnv,
145 |       stdout: "pipe",
146 |       stderr: "pipe",
147 |     });
148 |     const [stdout, stderr, exitCode] = await Promise.all([proc.stdout.text(), proc.stderr.text(), proc.exited]);
149 |     expect({ stdout: stdout.trim(), stderr, exitCode }).toEqual({
                                                              ^
error: expect(received).toEqual(expected)

  {
    "exitCode": 0,
    "stderr": "",
-   "stdout": "handler,recovered,beforeExit",
+   "stdout": "beforeExit",
  }

- Expected  - 1
+ Received  + 1

      at <anonymous> (/workspace/b
... (truncated)

release without fix: all passed
bun test v1.4.1-canary.1 (792f4a752)

test/js/node/process/process-on.test.ts:
(pass) process.on > when called from the main thread [7.85ms]
  [28ms]  bundle  1 modules
 [419ms] compile  ./out
(pass) process.on > should work inside --compile [502.78ms]
Bundled 1 module in 57ms

  out.ts  8 bytes  (entry point)

(pass) process.on > should work inside a macro [109.31ms]
(pass) process.on('unhandledRejection') > runs for a rejection raised by a beforeExit listener [8.44ms]
(pass) process.on('unhandledRejection') > runs before beforeExit when the rejection comes from a timer [9.45ms]
(pass) process.on('unhandledRejection') > keeps the process alive for work the handler schedules after a microtask hop [11.33ms]

test/js/node/events/event-emitter.test.ts:
(pass) node:events > captureRejectionSymbol [0.04ms]
(pass) node:events > once [1.07ms]
(pass) node:events > once (abort) [0.32ms]
(pass) node:events > once (two events in same tick) [10.39ms]
(pass) node:events > once removes the listener afterwards [0.87ms]
(pass) node:events > once is an async function [0.03ms]
(pass) node:events > once with already-aborted signal rejects (not a synchronous throw) [0.25ms]
(pass) node
... (truncated)
passes on PR (with fix)
ASAN with fix: all passed
$ BUN_DEBUG_QUIET_LOGS=1 bun scripts/build.ts --profile=debug --quiet test "--reporter=junit" "--reporter-outfile=/tmp/mechgate.xml" test/js/node/events/event-emitter.test.ts test/js/node/process/process-on.test.ts
bun test v1.4.1 (861e9ae04)

test/js/node/process/process-on.test.ts:
(pass) process.on > when called from the main thread [397.51ms]
 [303ms]  bundle  1 modules
[1.929s] compile  ./out
(pass) process.on > should work inside --compile [2591.54ms]
[macro] call initialize
Bundled 1 module in 349ms

  out.ts  8 bytes  (entry point)

(pass) process.on > should work inside a macro [705.56ms]
(pass) process.on('unhandledRejection') > runs for a rejection raised by a beforeExit listener [285.98ms]
(pass) process.on('unhandledRejection') > runs before beforeExit when the rejection comes from a timer [304.45ms]
(pass) process.on('unhandledRejection') > keeps the process alive for work the handler schedules after a microtask hop [462.41ms]

test/js/node/events/event-emitter.test.ts:
(pass) node:events > captureRejectionSymbol [2.21ms]
(pass) node:events > once [38.89ms]
(pass) node:events > once (abort) [11.78ms]
(pass) node:events > once (two events in same tick) 
... (truncated)

release with fix: all passed
$ bun scripts/build.ts --profile=release
[configured] bun-profile → bun (stripped) in 926ms (unchanged)
ninja: Entering directory `/workspace/bun/build/release'
[1/125] gen generated_host_exports.rs
generated_host_exports.rs: 116 exports (host=5, lazy=10, generic=101, rust=0); 242 extern-C blocks audited
[2/125] gen cpp.rs (cppbind)
[3/125] gen JS modules (bundle-modules)
Preprocess modules (7641ms)
Bundle modules (174ms)
Postprocesss modules (1096ms)
Bundle Functions (496ms)
Generate Code (62ms)

[9.48s] Bundled "src/js" for production
  2595 kb
  197 internal modules
  13 native modules
  50 internal functions across 16 files
[3/124] cargo bun_runtime → libbun_runtime.a (--target x86_64-unknown-linux-gnu)

  nightly-2026-07-20-x86_64-unknown-linux-gnu unchanged - rustc 1.99.0-nightly (9f36de775 2026-07-19)

�[1m�[92m   Compiling�[0m bun_jsc v0.0.0 (/workspace/bun/src/jsc)
�[1m�[92m   Compiling�[0m bun_js_parser_jsc v0.0.0 (/workspace/bun/src/js_parser_jsc)
�[1m�[92m   Compiling�[0m bun_semver_jsc v0.0.0 (/workspace/bun/src/semver_jsc)
�[1m�[92m   Compiling�[0m bun_bundler_jsc v0.0.0 (/workspace/bun/src/bundler_jsc)
�[1m�[92m   Compiling�[0m bun_ast_jsc v0.0.0 
... (truncated)
diff hotspot
src/js/node/events.ts                     |  30 +++++--
 src/jsc/JSGlobalObject.rs                 |   8 ++
 src/jsc/VirtualMachine.rs                 |  20 +++--
 src/jsc/bindings/ZigGlobalObject.h        |   1 +
 src/jsc/bindings/bindings.cpp             |   5 ++
 src/jsc/bindings/headers.h                |   1 +
 src/jsc/event_loop.rs                     |  18 ++++-
 src/runtime/cli/run_command.rs            |   2 +-
 src/runtime/jsc_hooks.rs                  |  21 +++--
 test/js/node/events/event-emitter.test.ts | 129 ++++++++++++++++++++++++++++++
 test/js/node/process/process-on.test.ts   |  98 +++++++++++++++++++++++
 11 files changed, 307 insertions(+), 26 deletions(-)

gate history · 4 passed · 0 rejected · iteration 9

evidence per changed file
file                                       reads  edits  tests
src/js/node/events.ts                          9     10      0
src/jsc/JSGlobalObject.rs                      1      2      0
src/jsc/VirtualMachine.rs                      4      4      0
src/jsc/bindings/ZigGlobalObject.h             1      1      0
src/jsc/bindings/bindings.cpp                  1      1      0
src/jsc/bindings/headers.h                     1      1      0
src/jsc/event_loop.rs                          3      3      0
src/runtime/cli/run_command.rs                 2      1      0
src/runtime/jsc_hooks.rs                       5      5      0
test/js/node/events/event-emitter.test.ts      2      3      0
test/js/node/process/process-on.test.ts        1      1      0

@github-actions github-actions Bot added the claude label Jul 5, 2026
@robobun

robobun commented Jul 5, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 2:09 AM PT - Aug 25th, 2026

❌ @robobun, your commit 6da9bca has 1 failures in Build #105519 (All Failures):


🧪   To try this PR locally:

bunx bun-pr 33356

That installs a local version of the PR into your bun-33356 executable, so you can run:

bun-33356 --bun

@github-actions

github-actions Bot commented Jul 5, 2026

Copy link
Copy Markdown
Contributor

Found 1 issue this PR may fix:

  1. Top level await causes promise rejections inside global.setTimeout to be ignored, allowing execution to continue #22546 - The PR's auto_tick_active fix adds rejection handling at the tail of the timer/poll phase, ensuring rejections from timer callbacks are properly processed rather than silently ignored when top-level await keeps the loop alive.

If this is helpful, copy the block below into the PR description to auto-close this issue on merge.

Fixes #22546

🤖 Generated with Claude Code

@coderabbitai

coderabbitai Bot commented Jul 5, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Walkthrough

This PR adds a pending-rejected-promises query and event-loop draining helper, wires that draining into shutdown and auto-tick paths, and expands EventEmitter.captureRejections to thenables. New tests cover thenable rejection capture and unhandledRejection ordering.

Changes

Rejected promise draining

Layer / File(s) Summary
Pending-rejected-promises query
src/jsc/bindings/ZigGlobalObject.h, src/jsc/bindings/bindings.cpp, src/jsc/bindings/headers.h, src/jsc/JSGlobalObject.rs
Adds a C++ accessor, exported binding, Rust wrapper, and FFI declaration for checking pending rejected promises.
Event loop rejection draining
src/jsc/event_loop.rs
Adds EventLoop::drain_rejected_promises() and changes tick() to call it before decrementing the entered-event-loop count.
Shutdown and auto-tick wiring
src/jsc/VirtualMachine.rs, src/runtime/cli/run_command.rs, src/runtime/jsc_hooks.rs
Updates on_before_exit, Run::start, auto_tick, and auto_tick_active to drain rejected promises from the event loop and adjusts the before-exit dispatch flow and related comments.

EventEmitter captureRejections thenables

Layer / File(s) Summary
Thenable rejection capture
src/js/node/events.ts
Changes listener rejection capture to duck-type thenables, preserve throwing then getter behavior, and keep the prototype emit function in sync with captureRejections toggles.
captureRejections tests
test/js/node/events/event-emitter.test.ts
Adds tests for thenable rejection capture, throwing then getters, and subprocess cases covering capture toggling and prototype emit patching.

Possibly related PRs

  • oven-sh/bun#32554: Rewrites the rejected-promise draining logic that this PR’s event-loop draining path depends on.
🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly identifies the primary change: draining the microtask checkpoint after unhandledRejection handlers run.
Description check ✅ Passed The description explains the problem, cause, implementation, scope, follow-up captureRejections fixes, and verification results. It does not use the exact template headings, but it provides the requir…
Full details: Description check

Explanation

The description explains the problem, cause, implementation, scope, follow-up captureRejections fixes, and verification results. It does not use the exact template headings, but it provides the required information in equivalent sections.


Comment @coderabbitai help to get the list of available commands.

@robobun

robobun commented Jul 5, 2026

Copy link
Copy Markdown
Collaborator Author

Checked the #22546 repro against this branch: it behaves exactly like main, so this PR does not fix it and Fixes #22546 should not go in the description.

import { setTimeout as sleep } from "node:timers/promises";
let x = 1;
async function main() {
  global.setTimeout(() => Promise.reject("oops"), 300);
  while (true) { x++; await sleep(100); console.log(x); }
}
await main();
output
node v26 2, 3, then ERR_UNHANDLED_REJECTION, exit 1
bun main 2, 3, error: oops, 4, 5, 6, … keeps running
bun this branch 2, 3, error: oops, 4, 5, 6, … keeps running

The rejection is already reported on main there, which is the part auto_tick_active is responsible for. What does not happen is the loop stopping, and that is a separate mechanism: under top-level await the module promise is driven by EventLoop::wait_for_promise, which loops on promise status alone and never consults unhandled_error_counter / is_event_loop_alive(). So the usual "an unhandled rejection ends the loop" path is bypassed, and --unhandled-rejections=throw keeps running too.

The one visible difference this PR makes on a variant of that repro without top-level await is that the process stops one loop iteration sooner, because the rejection is now notified in the timer phase rather than on the following tick. That is a side effect of the ordering fix, not a fix for the issue.

Comment thread src/runtime/jsc_hooks.rs
@robobun

robobun commented Jul 5, 2026

Copy link
Copy Markdown
Collaborator Author

Pushed fa94543.

Doc comment: fixed. The sibling comment on the RuntimeHooks::auto_tick_active field still listed handleRejectedPromises among the things the body skips. Good catch.

CI: test/js/node/test/parallel/test-event-capture-rejections.js failed on every platform. It is a real regression from this PR, and tracking it down turned up three more bugs, so here is the full story.

That test is a chain of twelve stages, each handing off to the next with process.nextTick from inside an unhandledRejection handler. On main that chain is exactly the thing this PR fixes, so the test silently stopped after its third stage and still exited 0:

main:  captureRejections, captureRejectionsTwoHandlers, defaultValue          (then nothing, exit 0)
node:  captureRejections, ... all twelve stages ...,    argValidation

With the event-loop fix the chain keeps going, the remaining nine stages run for the first time, and they land on three pre-existing captureRejections gaps. All three reproduce on released bun with no part of this PR:

bun 1.4.0 node
handler returns a thenable (not a real Promise) rejection dropped silently error event
then getter throws rejection dropped silently error event
inherits(X, EventEmitter) with no super() call, global captureRejections = true unhandledRejection error event

Causes, in src/js/node/events.ts:

  • addCatch only attached to real promises ($isPromise(result)). Node duck-types then. Reading then is observable (Promises/A+ allows a getter), so it is read once and a throwing getter becomes an error event.
  • emit is swapped per instance in the constructor (this.emit = emitWithRejectionCapture). An emitter built by inherits() without a super() call never runs that constructor, so it keeps the non-capturing emit no matter what the global flag says. The static captureRejections setter now swaps the prototype's emit as well.

That second fix alone would over-capture: an emitter constructed while the flag was off has its own kCapture = false and no own emit, so flipping the flag later would have started capturing it, which node does not do. So addCatch now re-checks kCapture, exactly as node does. I verified that clause is load-bearing by deleting it and watching the new negative test fail.

Four tests added to test/js/node/events/event-emitter.test.ts; three fail on main, and the fourth is the negative test for the guard above. test-event-capture-rejections.js now runs all twelve stages and exits 0, matching node.

Sweeps, no new failures: all 39 test-event*.js, 259 test-stream*.js + test-promise*.js (4 failures, all reproducing on released bun), event-emitter.test.ts, process-on.test.ts, and process.test.js's delivers many unhandledRejections in order. The seven node-differential repros in the description still match node byte for byte.

Comment thread src/js/node/events.ts Outdated

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the guard on the prototype .emit swap (and the test) — that closes out my last note. No further findings from me, but the event-loop / on_before_exit control-flow changes are the kind of shutdown-semantics work a human should sign off on.

Extended reasoning...

Overview

This PR does two things: (1) adds EventLoop::drain_rejected_promises() and wires it into tick, auto_tick, auto_tick_active, on_before_exit, and Run::start so microtasks/nextTicks scheduled inside an unhandledRejection handler are drained before the loop re-checks liveness; and (2) fixes three captureRejections gaps in src/js/node/events.ts (thenables, throwing then getter, inherits()-without-super) that the first fix exposed by making test-event-capture-rejections.js run all twelve stages. A new hasPendingRejectedPromises C++/FFI binding backs the fast-path check.

Since my last review, commit 070f31c addressed the .emit-clobbering concern exactly as suggested (guard the prototype swap on the current value being one of Bun's two internal functions) and added a dedicated test. Both earlier findings are now resolved.

Security risks

None identified. No auth/crypto/network-boundary code; the changes are event-loop scheduling and a Node-compat builtin. The drain_rejected_promises loop terminates on TerminationException, and has_pending_rejected_promises is a plain isEmpty() read.

Level of scrutiny

High. The on_before_exit restructuring (moving drain_rejected_promises() + is_event_loop_alive() above the dispatch branch, and unconditionally continue-ing after dispatch_on_before_exit) changes shutdown control flow. auto_tick_active gaining a rejection tail and tick's tail switching from a single handle_rejected_promises to a looping drain both touch the hottest runtime path. These are exactly the areas where subtle ordering has bitten before, and the PR description itself flags an interaction with #33354 that whichever lands second must reconcile.

Other factors

The PR is unusually well-tested (differential vs Node across seven scenarios, three new process-on tests, five new event-emitter tests including the negative guard, and broad suite sweeps), and the author has been responsive to bot feedback. That raises confidence, but the scope — core event-loop tick/shutdown semantics plus Node-compat builtin behavior — still puts it firmly in "human maintainer should review" territory rather than bot-approvable.

@robobun

robobun commented Jul 5, 2026 •

Copy link
Copy Markdown
Collaborator Author

Status

Rebased onto current main (2c10950) and ready for review.

Build 105519 on the current head (6da9bca): finished 180 passed, 1 failed. The failed job is annotated step failed outside runner: a darwin any aarch64 shard on agent darwin-aarch64-26.6.2-1 hit buildkite-agent artifact download timed out after 120s for step 'darwin-aarch64-build-bun' and refused to run with a partial binary. Zero tests ran on it and no bun binary of this branch was executed. Every lane that has actually run is green, binary-size included.

The named darwin agents are the only thing that has been red on this PR for three consecutive builds, each time failing to fetch bits over the network before any test runs: a git clone at ~13 KiB/s dying at 14% (105308, 26.6.1-1), and artifact-download timeouts (105519, 26.6.2-1; and the 26.5.1-1 / 26.3.1-1 generation before them). Across all branches in the last 60 builds, darwin-aarch64-26.6.x test jobs fail 17% of the time against 0–2% for every other darwin agent class (tart VMs and the x64 minis). That pool looks worth a look independent of this PR.

This fourth rebase was trivial: one context-only conflict in JSGlobalObject.rs, where #40410's exception-check lint made the neighbouring readable_stream_to_array_buffer fallible right where this PR's has_pending_rejected_promises accessor sits. Both kept. Since #40410 adds an exception-discipline lint, I also ran its runtime form (BUN_JSC_validateExceptionChecks=1 BUN_JSC_dumpSimulatedThrows=1) over all 7 node-differential repros and both of this PR's test files: zero findings. (The new C++ here is a const isEmpty() accessor and a one-line trampoline; neither declares a ThrowScope nor calls anything that throws.)

The previous head's build, 105308, finished 180 passed, 1 failed with no test-failure annotation: the one failed job never checked out the repository (three git clone attempts on darwin-aarch64-26.6.1-1 crawled at ~13 KiB/s and died at 14%). Zero tests ran on it. Every lane that actually ran was green.

This third rebase was trivial: the only conflict was in headers.h, where main's dead-code pass (#39581) deleted the putCachedObject declaration this PR's new hasPendingRejectedPromises declaration had been inserted next to. The new declaration is kept; the dead neighbour stays deleted. No other file conflicted. Re-verified on the new base: the C++ definition, header, Rust extern and wrapper all agree; both addCatch dispatch sites in events.ts are still thenable-shaped; on_before_exit is untouched by main's VirtualMachine.rs rework.

The previous head's build, 103223, finished 180 passed, 1 failed. The one failed job was a single darwin any aarch64 shard on agent darwin-aarch64-26.6.2-1 carrying three upstream HTTP tests (test-http-should-keep-alive.js, test-http-pipeline-requests-connection-leak.js, test-http-full-response.js). Not this diff: none of the three uses captureRejections or raises an unhandled rejection, so with no pending rejections drain_rejected_promises() is the same single isEmpty() read handleRejectedPromises() already performed and the events.ts changes never run; all three pass 3/3 locally on the exact binary; and the agent had a 0-for-1 record with a 1-for-3 sibling. All three were reported for main-break triage. binary-size is green, and the earlier Windows deinitialization.test.ts segfault (also reported) did not recur.

The previous head's build, 100373, finished 178 passed, 1 failed. The failure was test/bake/deinitialization.test.ts segfaulting at process exit on Windows 2019 x64: a pre-existing crash on main (identical signature on unrelated branches 100346 / 100372), reported for main-break triage separately. binary-size is green on this base.

The previous build on the rebased branch, 100358, finished 178 passed, 1 failed. The one failure was the binary-size check reporting +545–637 KB on every target. That was a baseline artifact, not this diff: the check compares against the latest canary, which at that point was 11 commits ahead of the branch's base and included #39484 (drops unused libarchive formats) and #39486. The branch is now rebased past both, so the comparison is like-for-like; this diff itself is ~80 lines and adds one isEmpty() accessor plus one small loop.

Every lane that was flaky on the old base (the two darwin agents, tls-syscall-fault.test.ts, terminal.test.ts) passed in 100358.

Review: all threads resolved. The comment-bot pass did turn up one real thing, now in 7f1ac49: drain_rejected_promises() no longer hoists the global/VM into locals (it uses global_ref() / drain_microtasks() directly), which removed the borrowck note along with it. Remaining comments are the rustdoc on the new fn, the SAFETY: lines, and two one/two-line notes on non-obvious guards.

Verified on the current head: process-on.test.ts + event-emitter.test.ts 98/98 (main added cases to that file), plus the new test/js/bun/jsc/exception-checks.test.ts 4/4; test-event-capture-rejections.js all 12 stages; 56 upstream test-event* / test-promise* files; 7 node-differential repros byte-identical to node v26; rust:check-all 12/12; the 6 expected fail-before failures against released bun.

Earlier CI history on the pre-rebase base (superseded)

Builds 68503 and 68530 on the original base were green on every lane except: two darwin 26 aarch64 shards whose agent (darwin-aarch64-26.5.1-1) timed out downloading the build artifact and never ran a test (the same failure hit 27 of 40 contemporaneous builds on unrelated branches); tls-syscall-fault.test.ts on x64-asan, which also failed on two unrelated branches and whose assertion greps for a different OOM string than the one the crash handler printed; and one terminal.test.ts timeout on a macOS agent that the identical code had passed on the previous build and that did not reproduce locally (8/8 in 3–4 s against a 90 s budget, 60-iteration PTY stress clean). All four passed on the rebased base.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/js/node/events.ts`:
- Around line 842-849: The comment block in EventEmitterPrototype.emit handling
is too long and exceeds the 3-line limit; trim the explanatory text while
keeping the same meaning. Update the nearby comment in the emit swap logic so it
stays concise, and preserve the actual assignment logic that toggles between
emitWithRejectionCapture and emitWithoutRejectionCapture.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: d4e61d41-1022-4b9e-a639-4289a5ac5311

📥 Commits

Reviewing files that changed from the base of the PR and between 15a0952 and 16290bc.

📒 Files selected for processing (3)
  • src/js/node/events.ts
  • src/jsc/VirtualMachine.rs
  • test/js/node/events/event-emitter.test.ts

Comment thread src/js/node/events.ts Outdated
@robobun

robobun commented Jul 5, 2026

Copy link
Copy Markdown
Collaborator Author

Heads up: this overlaps with #33354 (opened a few hours earlier). Both are real and neither subsumes the other, but they add the same report-then-drain loop two different ways and touch the same five files, so they will conflict on merge.

They fix different bugs

#33354 this PR
bug ordering — the event is emitted after macrotasks the rejecting callback scheduled liveness — work the handler defers behind a microtask hop is dropped entirely
hook end of each timer / immediate callback the existing scan sites + auto_tick_active's missing tail

I checked this PR's repro against #33354's build, and it is not fixed there:

node            : recovered
bun main        : (nothing)
bun with #33354 : (nothing)

A top-level rejection is still scanned at auto_tick's tail, which #33354 doesn't touch. So both changes are needed.

Where they collide

Both add while (pending) { handle; drain }, just plumbed differently:

Same five files either way: src/jsc/event_loop.rs, src/jsc/JSGlobalObject.rs, src/jsc/bindings/ZigGlobalObject.h, src/jsc/bindings/bindings.cpp, src/jsc/bindings/headers.h.

Suggestion

Whichever lands first, the other rebases and reuses the loop that's already there rather than adding a second one. The two mechanisms are interchangeable; the signal just needs to exist once. #33354 is the older and currently the further-along of the two (green except for the darwin artifact-download infra failure that's hitting every build right now), but I have no stake in which way it goes.

@robobun

robobun commented Aug 18, 2026

Copy link
Copy Markdown
Collaborator Author

Another report of the same bug, in the await shape: a suspended async function whose promise is resolved from inside the unhandledRejection listener never resumes, and the process exits 0.

let res; const p = new Promise(r => res = r);
process.on("unhandledRejection", () => { console.log("handler"); res(); });
(async () => { await p; console.log("continuation"); })();
Promise.reject(new Error("x"));

bun 1.4.0 prints handler only; node prints handler, continuation. queueMicrotask, process.nextTick and a fresh Promise.resolve().then() inside the listener are dropped the same way when nothing else keeps the loop alive.

Two data points that may help whoever picks this up:

  • The default mode is the only one affected. bun --unhandled-rejections=none|warn|warn-with-error-code|strict|throw all print continuation, because every arm of VirtualMachine::unhandled_rejection except Mode::Bun drains microtasks after Bun__handleUnhandledRejection returns (src/jsc/VirtualMachine.rs, the drain(self) calls). The .bun arm has been the odd one out since the Zig version.
  • This branch is 1281 commits behind main and conflicts in src/js/node/events.ts, src/jsc/VirtualMachine.rs, src/jsc/event_loop.rs, src/runtime/cli/run_command.rs and src/runtime/jsc_hooks.rs. The auto_tick_active half of the change has since been redone on its own in event loop: report promise rejections left by the turn that let the loop go idle #37981, which applies cleanly to current main; the post-listener drain (the part that fixes the repro above) and the captureRejections fixes in events.ts are still only here.

Not opening a separate PR for this since the fix is already here.

@robobun
robobun force-pushed the farm/184ecb12/unhandled-rejection-microtask-checkpoint branch from ec1737f to 6a99346 Compare August 18, 2026 03:51
Comment thread src/js/node/events.ts Outdated
Comment thread src/js/node/events.ts Outdated
Comment thread src/js/node/events.ts Outdated
Comment thread src/js/node/events.ts Outdated
Comment thread src/jsc/JSGlobalObject.rs
Comment thread src/jsc/VirtualMachine.rs
Comment thread src/jsc/VirtualMachine.rs Outdated
Comment thread src/jsc/VirtualMachine.rs
Comment thread src/jsc/event_loop.rs Outdated
Comment thread src/jsc/event_loop.rs Outdated
Comment thread src/runtime/jsc_hooks.rs
@robobun
robobun force-pushed the farm/184ecb12/unhandled-rejection-microtask-checkpoint branch from 6a99346 to 7f1ac49 Compare August 18, 2026 04:17
Comment thread src/js/node/events.ts
Comment thread src/jsc/event_loop.rs
@robobun
robobun force-pushed the farm/184ecb12/unhandled-rejection-microtask-checkpoint branch 2 times, most recently from 7578c54 to 792f4a7 Compare August 25, 2026 03:09

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this PR and the current bug-hunting pass found no new issues; both earlier findings (the stale auto_tick_active doc comment and the captureRejections setter clobbering a userland emit patch) are resolved with tests. Because this restructures the on_before_exit drain loop and adds a rejection drain at the tail of every auto_tick/auto_tick_active exit — core shutdown-ordering paths — a human look is still worthwhile, and there is the noted #33354 overlap to reconcile on merge.

What was reviewed:

  • drain_rejected_promises loop and its four call sites; Stopped propagation matches the surrounding let _ = / early-return pattern on main.
  • on_before_exit reflow: the unconditional continue after dispatch_on_before_exit re-enters the loop so rejections from a beforeExit listener drain before the next liveness check; dispatch = false still terminates on the following pass.
  • addCatch thenable duck-typing: then is read once inside a try, primitives are filtered by the !== null / !== undefined guards, and the kCapture re-check prevents over-capture when the prototype holds the capturing variant.
Extended reasoning...

Overview

Two related fixes: (1) a new EventLoop::drain_rejected_promises() that runs the nextTick + microtask checkpoint after unhandledRejection handlers, wired into tick_turn, both auto_tick exits, both auto_tick_active exits, Run::start's post-loop scan, and the on_before_exit drain loop (which is also restructured so rejections from a beforeExit listener are notified before the next liveness check); and (2) three captureRejections gaps in src/js/node/events.ts — thenable duck-typing, a throwing then getter, and emitters whose constructor never ran — plus a guard so the prototype-emit swap does not clobber a userland monkey-patch. New C++/Rust binding hasPendingRejectedPromises keeps the common path to one isEmpty() read. Eight new tests across two files; all subprocess-based with piped stdout/stderr and combined-object assertions.

Security risks

None identified. No auth/crypto/permissions surface. The addCatch change now reads .then off arbitrary listener return values, but only under captureRejections (opt-in), inside a try/catch, and matches Node's reference implementation.

Level of scrutiny

High. This touches the event-loop tick tail, the beforeExit re-dispatch loop, and EventEmitter.prototype.emit — all load-bearing for shutdown ordering and process-exit semantics. The on_before_exit reflow moves the is_event_loop_alive() check from after dispatch_on_before_exit to before it and makes the post-dispatch continue unconditional; I traced it and believe it terminates correctly (dispatch = false on the second pass), but this is exactly the kind of subtle control-flow change a maintainer should confirm.

Other factors

The PR is thoroughly tested (differential against node v26, fail-before verified, upstream test suites swept) and both prior review findings are addressed with negative tests. The description notes a semantic overlap with #33354 that whoever lands second must reconcile (handle_rejected_promises_after_tick should call drain_rejected_promises). CI on the current head is 180/1 with the one failure triaged as an unrelated darwin agent issue. Given the scope and the merge-order coordination needed, deferring to a human.

An unhandledRejection handler is user JS that can queue microtasks and
nextTicks. Node loops back into processTicksAndRejections after
processPromiseRejections(), so anything the handler schedules behind an
await still runs. Bun returned straight to the liveness check, dropping
that work and exiting 0.

Add EventLoop::drain_rejected_promises(): notify pending rejections, then
re-enter the nextTick + microtask checkpoint they filled, repeating while
the drain produces new rejections. Use it at the tail of EventLoop::tick,
auto_tick and auto_tick_active, and in the before-exit drain loop.

auto_tick_active had no rejection handling at all, so a rejection raised
in the timer or poll phase was only notified after beforeExit, with no
loop iteration left to run what the handler scheduled.
…nstructor never ran

Unmasked by the unhandledRejection fix: test-event-capture-rejections.js
stopped after its third stage on main because the nextTick chain inside the
rejection handler was dropped, so the later stages never ran.

Three gaps in captureRejections, all reproducible on released bun:

- addCatch only attached to real promises, so a thenable returned by a
  handler was never captured. Duck-type `then` like node, read it once, and
  route a throwing getter to an 'error' event.
- emit is swapped per instance in the constructor, so an emitter built by
  inherits() without a super call kept the non-capturing emit. Swap the
  prototype's emit from the static captureRejections setter too.
- addCatch now re-checks kCapture, so an emitter that opted out does not
  start capturing when the global flag flips after it was constructed.

Also drop handleRejectedPromises from the auto_tick_active hook doc, which
no longer describes the body.
The captureRejections setter swapped the prototype's emit unconditionally,
so toggling the flag discarded a monkey-patched emit (what APM and tracing
libraries install). Node's setter only flips the flag. Swap only when the
current value is one of the two internal emit functions.
…eeds

Use global_ref() and drain_microtasks() directly instead of hoisting the
global and VM into locals, which removes the borrowck note explaining the
hoist. Drop two comments that restate what the adjacent code already says
and shorten the rest.
@robobun
robobun force-pushed the farm/184ecb12/unhandled-rejection-microtask-checkpoint branch from 792f4a7 to 6da9bca Compare August 25, 2026 08:51

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this PR and the bug-hunting pass found no issues. Both concerns I raised earlier (the stale auto_tick_active doc comment and the emit prototype-patch clobbering) have been addressed. Given the scope — event-loop tick/shutdown semantics, the on_before_exit control-flow restructuring, and the noted coordination with #33354 — a human look is still warranted.

Checked: the new drain_rejected_promises loop and its four call sites propagate Stopped consistently with neighbouring loop-level code; the on_before_exit restructuring terminates (unconditional continue after re-dispatch falls through to break on the next pass with dispatch = false); both emitWithRejectionCapture dispatch sites got the thenable change; the addCatch kCapture guard prevents over-capture when the prototype holds the capturing variant.

Extended reasoning...

Overview

Two related fixes: (1) drain the nextTick + microtask checkpoint after unhandledRejection handlers run so work they schedule behind an await isn't dropped, wired into tick_turn, auto_tick, auto_tick_active, on_before_exit, and Run::start; (2) three captureRejections Node-compat gaps in events.ts (thenable duck-typing, throwing then getter, inherits() without super()) uncovered when fix (1) let test-event-capture-rejections.js run past its third stage. Adds one C++ accessor + Rust FFI wrapper for hasPendingRejectedPromises. Eight new tests across two files, all differentially verified against Node.

Security risks

None identified. No untrusted-input parsing, no auth/crypto, no path handling. The events.ts change reads user-controlled .then under a try/catch as Node does.

Level of scrutiny

High. This touches the core event-loop tick and shutdown/drain paths (EventLoop::tick_turn, VirtualMachine::on_before_exit, auto_tick/auto_tick_active), where ordering mistakes surface as hangs, dropped work, or wrong beforeExit semantics rather than clean failures. The on_before_exit loop was restructured (the post-dispatch liveness check moved and the continue became unconditional). events.ts is a hot-path built-in, and the setter now writes EventEmitter.prototype.emit (guarded).

Other factors

The PR has been through 8 iterations with extensive verification (CI green on 180 lanes, node-differential repros byte-identical, rust:check-all on all targets, fail-before checks against released Bun). Both issues I raised previously are resolved with tests. There is a stated interaction with #33354 (same five files, overlapping report-then-drain loop) that whoever merges will need to reconcile — the PR description documents exactly what the second-to-land should do. That coordination and the event-loop-semantics scope are why this warrants a human reviewer rather than auto-approval.

robobun added a commit that referenced this pull request Sep 8, 2026
… await pending; fold in the cases from #37981, #34193 and #33356

A worker whose 'beforeExit' listener calls process.exit(0) while the
entry module's top-level await is still pending exited 13. The exit-13
stamp now also checks that no stop was requested, as the enclosing
'beforeExit' condition does. Node exits 0 here.

Tests carried over from the PRs this one supersedes:
- #37981: repl -e last-timer rejection, two main-thread 'beforeExit'
  orderings, the top-level-await worker cases, and the
  reported-before-queued-task ordering.
- #34193: worker-error-exit-code.test.ts (message-listener and timer
  sites, the worker's own 'exit' argument, listener suppression).
- #33356: a main-thread listener that recovers from a 'beforeExit'
  rejection after a microtask hop.
@robobun

robobun commented Sep 8, 2026

Copy link
Copy Markdown
Collaborator Author

Closing in favor of #42031 and #42032, which are the same fixes on current main.

I ran this PR's 8 new tests on #42032's branch (it includes #42031): the 5 event-emitter cases and 2 of the 3 process-on cases pass, and the 'beforeExit' case was added to #42032 in bb390cb. The remaining one, "keeps the process alive for work the handler schedules after a microtask hop", passes on substance (all three timers run) but not on order: #42031 drains per listener rather than per batch, so it prints then, queueMicrotask, nextTick where Node prints nextTick, then, queueMicrotask. That is noted on #42031 with the repro.

@robobun robobun closed this Sep 8, 2026
robobun added a commit that referenced this pull request Sep 26, 2026
… await pending; fold in the cases from #37981, #34193 and #33356

A worker whose 'beforeExit' listener calls process.exit(0) while the
entry module's top-level await is still pending exited 13. The exit-13
stamp now also checks that no stop was requested, as the enclosing
'beforeExit' condition does. Node exits 0 here.

Tests carried over from the PRs this one supersedes:
- #37981: repl -e last-timer rejection, two main-thread 'beforeExit'
  orderings, the top-level-await worker cases, and the
  reported-before-queued-task ordering.
- #34193: worker-error-exit-code.test.ts (message-listener and timer
  sites, the worker's own 'exit' argument, listener suppression).
- #33356: a main-thread listener that recovers from a 'beforeExit'
  rejection after a microtask hop.
robobun added a commit that referenced this pull request Sep 26, 2026
… await pending; fold in the cases from #37981, #34193 and #33356

A worker whose 'beforeExit' listener calls process.exit(0) while the
entry module's top-level await is still pending exited 13. The exit-13
stamp now also checks that no stop was requested, as the enclosing
'beforeExit' condition does. Node exits 0 here.

Tests carried over from the PRs this one supersedes:
- #37981: repl -e last-timer rejection, two main-thread 'beforeExit'
  orderings, the top-level-await worker cases, and the
  reported-before-queued-task ordering.
- #34193: worker-error-exit-code.test.ts (message-listener and timer
  sites, the worker's own 'exit' argument, listener suppression).
- #33356: a main-thread listener that recovers from a 'beforeExit'
  rejection after a microtask hop.
robobun added a commit that referenced this pull request Sep 26, 2026
… await pending; fold in the cases from #37981, #34193 and #33356

A worker whose 'beforeExit' listener calls process.exit(0) while the
entry module's top-level await is still pending exited 13. The exit-13
stamp now also checks that no stop was requested, as the enclosing
'beforeExit' condition does. Node exits 0 here.

Tests carried over from the PRs this one supersedes:
- #37981: repl -e last-timer rejection, two main-thread 'beforeExit'
  orderings, the top-level-await worker cases, and the
  reported-before-queued-task ordering.
- #34193: worker-error-exit-code.test.ts (message-listener and timer
  sites, the worker's own 'exit' argument, listener suppression).
- #33356: a main-thread listener that recovers from a 'beforeExit'
  rejection after a microtask hop.
robobun added a commit that referenced this pull request Sep 26, 2026
… await pending; fold in the cases from #37981, #34193 and #33356

A worker whose 'beforeExit' listener calls process.exit(0) while the
entry module's top-level await is still pending exited 13. The exit-13
stamp now also checks that no stop was requested, as the enclosing
'beforeExit' condition does. Node exits 0 here.

Tests carried over from the PRs this one supersedes:
- #37981: repl -e last-timer rejection, two main-thread 'beforeExit'
  orderings, the top-level-await worker cases, and the
  reported-before-queued-task ordering.
- #34193: worker-error-exit-code.test.ts (message-listener and timer
  sites, the worker's own 'exit' argument, listener suppression).
- #33356: a main-thread listener that recovers from a 'beforeExit'
  rejection after a microtask hop.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant