Skip to content

hot/watch: dispatch beforeExit once per drain, not once per event loop wakeup - #38658

Open
robobun wants to merge 1 commit into
mainfrom
farm/7b9bd7eb/hot-before-exit-once-per-drain
Open

robobun wants to merge 1 commit into
mainfrom
farm/7b9bd7eb/hot-before-exit-once-per-drain

Conversation

@robobun

@robobun robobun commented Aug 14, 2026

Copy link
Copy Markdown
Collaborator

Problem

  • Under bun --hot and bun --watch, process.on("beforeExit", ...) fires again about once per second while the script is idle, and in a tight loop whenever some fd the loop polls is ready. Plain bun and node (with or without --watch) emit it once when the loop drains. Reproduces on 1.4.0 and main (probes below).
  • Cause: the watcher arm of the run loop in Run::start (src/runtime/cli/run_command.rs, the loop after // core run-loop) calls vm.on_before_exit() on every iteration, and one iteration is one return from tick_possibly_forever() (src/jsc/event_loop.rs), which returns on every wakeup: its one second bounded wait, a signal, a ready one-shot poll. Nothing checked whether the loop had done any work since the previous dispatch.
  • Not new: the Zig version of this loop made the same unconditional call, but its park was only interrupted by a 4 minute keep-alive timer (and fd wakeups); the one second bounded wait the Rust tick_possibly_forever() uses is what makes the idle case visible.

Fix

  • The watcher arm dispatches beforeExit only when something drained: once for the initial evaluation, again after its while vm.is_event_loop_alive() body ran (the loop had work and ran out of it), and again when vm.hot_reload_counter moved (a reload happened). A wakeup that did none of these goes straight back to parking.
  • The counter check is needed, not just the "body ran" flag: the entry file is transpiled on the JS thread (transpile_file in src/runtime/jsc_hooks.rs skips the concurrent store for is_main), so a reload of an entry without imports is evaluated entirely inside the tick of the wakeup that delivered the reload task, and the loop is never alive for it. Verified by building without the counter check: the --hot test below then fails because the reloaded generation never gets its beforeExit.
  • Why this is correct: it is the rule node and the non-watcher arm right below already follow (emit when the loop drains; emit again only if a listener re-armed it), applied to a process that is kept alive by the watcher instead of exiting. The re-arm case itself is unchanged, it is handled inside VirtualMachine::on_before_exit(). A listener that schedules a timer once still gets start, beforeExit, timer, beforeExit under --hot, the same as plain bun. Two reloads that land in the same tick drain together and are announced once.
  • Everything else the loop did per iteration (report_exception_in_hot_reloaded_module_if_needed(), which reports a reloaded generation's rejection, retries a deferred reload and re-arms the watcher on the entry) still runs on every iteration; only the dispatch is gated.
  • Test: test/cli/hot/hot.test.ts, --hot and --watch variants of "emits beforeExit once per generation, not once per event loop wakeup". The wake source is SIGUSR2: a signal listener runs without keeping the loop alive, so after each signal line nothing else may be printed. The test then saves the entry and expects the new generation to print start and exactly one more beforeExit. Signals are what make the test POSIX-only; the run loop is shared by all platforms.
    • Unfixed (1.4.0 via USE_SYSTEM_BUN=1, and the previous debug build): both variants fail within 100ms with expected "signal", received "beforeExit". Fixed: pass, 40/40 with --rerun-each=20 on the debug (ASAN) build.
    • Also green on the debug build: the rest of test/cli/hot/hot.test.ts, test/cli/hot/watch.test.ts, test/cli/watch/watch.test.ts, test/cli/watch/watcher-trace.test.ts, test/cli/hot/watch-many-dirs.test.ts, test/js/bun/cron/in-process-cron.test.ts, test/js/bun/resolve/bun-main-entry-point.test.ts, test/regression/issue/29524.test.ts, and the beforeExit tests in test/js/node/process/process.test.js.
  • Not in scope: console.write() (a FileSink write + flush) inside a beforeExit listener still re-fires forever, with or without a watcher, because the flushed write leaves its poll registered and ref'd, so the loop reads as alive and on_before_exit()'s own re-dispatch rule fires; that is a FileSink keep-alive bug being fixed separately. Once it is, this change is what stops the same scenario from looping under --hot (the poll still wakes the loop, but no longer counts as a drain).
  • Nearby open PRs touch the same line and compose with this one: hot: report a reloaded generation's error while a beforeExit listener keeps the loop alive #38638 replaces the on_before_exit() call with on_before_exit_with(...) (the call simply moves inside the new if), hot/watch: keep driving timers and surviving errors after an unhandled error #38206 and hot: reset unhandled-error state on reload so timers recover after a failed reload #34657 change the error counters the dispatch is already guarded on, hot: replace a generation that is parked on a top-level await #38613 changes when a reload defers.

Background

  • beforeExit is node's "the event loop is empty" event: emitted when the loop drains; if a listener schedules more work, the loop runs again and the event is emitted again when it drains; if not, the process exits. VirtualMachine::on_before_exit() implements exactly that (dispatch, drain, re-dispatch only if the drain did work).
  • With --hot or --watch the process must not exit, so Run::start takes a different arm: drive the loop while is_event_loop_alive(), dispatch beforeExit, then park in tick_possibly_forever() until something wakes the loop. --hot handles a file change by posting a task that re-evaluates the entry in the same process (one "generation" per evaluation; hot_reload_counter is incremented by each reload, and globalThis and process listeners survive it); --watch re-executes the process, so each generation starts from the top of this function.
  • is_event_loop_alive() counts ref'd handles, queued tasks and in-flight module loads. Signal listeners and unref'd handles do not count, so their callbacks run on a wakeup without the loop ever being alive.
Probes (1.4.0, linux x64)
$ cat log.js
process.on("beforeExit", () => console.log("beforeExit"));
$ timeout 3 bun log.js | wc -l          # 1
$ timeout 3 bun --hot log.js | wc -l    # 3, one per second of idling
$ timeout 3 bun --watch log.js | wc -l  # 3

$ cat entry.js
console.log("start");
if (!globalThis.listening) {
  globalThis.listening = true;
  process.on("beforeExit", () => console.log("beforeExit"));
  process.on("SIGUSR2", () => console.log("signal"));
}
# bun --hot entry.js, then three kill -USR2, then save the file:
start, beforeExit, signal, beforeExit, signal, beforeExit, signal, beforeExit, start, beforeExit
# same with --watch. With this change (debug build), both modes:
start, beforeExit, signal, signal, signal, start, beforeExit

…p wakeup

The watcher arm of the run loop called on_before_exit() after every
return from tick_possibly_forever(), which returns on every wakeup:
its one second bounded wait, a signal, a ready one-shot poll. Dispatch
only after the loop actually ran work and drained, or after a reload
bumped hot_reload_counter (an entry with nothing to transpile off-thread
is re-evaluated inside the wakeup's own tick without the loop ever
being alive), plus once for the initial evaluation.
@coderabbitai

coderabbitai Bot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

@robobun, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 30 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: bb2f5e09-6353-4367-93e8-00802ef87965

📥 Commits

Reviewing files that changed from the base of the PR and between 97a4363 and db963ab.

📒 Files selected for processing (2)
  • src/runtime/cli/run_command.rs
  • test/cli/hot/hot.test.ts

Comment @coderabbitai help to get the list of available commands.

@robobun

robobun commented Aug 14, 2026

Copy link
Copy Markdown
Collaborator Author
Updated 12:37 PM PT - Aug 14th, 2026

@robobun, your commit db963ab is building: #96492

@robobun

robobun commented Aug 14, 2026

Copy link
Copy Markdown
Collaborator Author

Status: ready for review.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this PR and didn't find any bugs. Because it changes the semantics of beforeExit dispatch in the core --hot/--watch run loop and interacts with several nearby open PRs (#38638, #38206, #34657, #38613), a human look would still be worthwhile.

What was reviewed:

  • The gating logic in run_command.rs: initial-eval / drained-work / hot_reload_counter-moved each set the flag; spurious wakeups (bounded wait, signal, unref'd poll) do not.
  • hot_reload_counter is a plain u32 read and written only on the JS thread (incremented in VirtualMachine::reload), so no atomics needed.
  • Confirmed the re-arm case (a beforeExit listener that schedules work) still re-dispatches: the scheduled work makes is_event_loop_alive() true on the next iteration, which re-enters the while body and sets the flag.
  • Test uses forEachLine/tempDir/writeHotFileAtomicSync from harness, awaits each line rather than sleeping, and the while (line === "start") loop absorbs coalesced watcher events.
Extended reasoning...

Overview

This PR gates vm.on_before_exit() in the watcher arm of Run::start (src/runtime/cli/run_command.rs) so that beforeExit fires once per drain of the event loop rather than once per return from tick_possibly_forever(). Two new flags track whether a drain occurred: dispatch_before_exit (set on initial entry and whenever the while is_event_loop_alive() body runs) and a snapshot of hot_reload_counter (to catch reloads that complete entirely inside a single wakeup without the loop ever being alive). A test in test/cli/hot/hot.test.ts covers both --hot and --watch, using SIGUSR2 as a wakeup that must not re-emit and a file save as a reload that must emit exactly once more.

Security risks

None. This is control flow around event dispatch in the run loop; no untrusted input, parsing, allocation, or privilege boundaries are involved.

Level of scrutiny

High. The watcher arm of Run::start is the core run loop for --hot and --watch — every long-running dev session goes through it. The change alters when a Node-compat lifecycle event fires, which is user-observable and something applications may depend on (cleanup handlers, graceful-shutdown hooks). The PR description is unusually thorough (mechanism traced, USE_SYSTEM_BUN=1 failure confirmed, --rerun-each=20 pass on ASAN, the counter-check verified load-bearing by removing it), and the logic reads correctly to me: signal/bounded-wait/unref'd-poll wakeups skip straight to tick_possibly_forever(), while any real work (loop-alive body ran, or the reload counter moved) re-arms the dispatch. The re-arm-from-listener case is preserved because scheduled work makes the loop alive on the next iteration.

Other factors

The PR itself notes four nearby open PRs touching the same line or the counters it reads, and describes how they compose. That merge-ordering judgment, plus the fact that this is a semantic change to a Node lifecycle event in the hottest dev-mode code path, is exactly what a maintainer should weigh in on. The bug hunting system found nothing; I checked the thread-safety of hot_reload_counter (single-threaded, JS-thread-only), the re-arm path, and the test's handling of coalesced watcher events, and found no issues. Deferring so a maintainer can confirm the intended beforeExit semantics under watch modes and the composition with the adjacent PRs.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant