Skip to content

FileSink: release the event loop keep-alive when flush() drains the buffer - #38641

Merged
Jarred-Sumner merged 1 commit into
mainfrom
farm/c8d1672c/filesink-flush-keepalive
Aug 14, 2026
Merged

Jarred-Sumner merged 1 commit into
mainfrom
farm/c8d1672c/filesink-flush-keepalive

Conversation

@robobun

@robobun robobun commented Aug 14, 2026

Copy link
Copy Markdown
Collaborator

Problem

  • process.on("beforeExit", () => console.write("x\n")) with stdout on a pipe never exits: beforeExit is re-emitted over and over and x is printed until the process is killed (timeout 2 bun -e '...' | wc -l prints several hundred thousand lines). process.stdout.write and console.log in the same listener print once and the process exits, as in Node.
  • Same for user code doing write() + flush() on Bun.stdout.writer(), Bun.stderr.writer() or any pipe/socket-backed Bun.file(fd).writer() from the listener. console.write() is exactly that (write in src/js/builtins/ConsoleObject.ts).
  • Cause, in src/runtime/webcore/FileSink.rs: on POSIX a write() below the coalescing threshold is only buffered. on_write then turns the loop keep-alive on because bytes are pending (update_ref(evtloop, has_pending_data)), and the deferred on_auto_flush turns it off after draining them.
  • flush_from_js (the JS flush()) drains the same bytes but leaves the keep-alive on, so until the next deferred-task drain the sink still tells the loop it has pending work.
  • VirtualMachine::on_before_exit (src/jsc/VirtualMachine.rs) therefore sees the loop alive after the listener returns, ticks (the auto-flush turns the keep-alive off), and, as it must when a listener really did schedule work, emits beforeExit again. The listener writes again and the cycle repeats.
  • Only pipes and sockets are affected: a regular file has no poll to keep alive, and a TTY forces the sink synchronous so nothing is ever buffered. Piped stdout is what every test harness gives a child, which is how this was noticed.

Fix

  • flush_from_js turns the keep-alive off when its flush emptied a non-empty buffer, on the success and the error arm alike. This is what on_auto_flush already does after its own drain.
  • Why this is right: the keep-alive only exists to cover bytes sitting in the buffer. Once a flush has drained them the loop has nothing to wait for, so the sink's liveness now matches a sink whose bytes went straight to the fd (the process.stdout.write case that already behaved).
  • A flush that could not drain (EAGAIN) still has bytes buffered, so the keep-alive and the armed writable poll stay and the process still waits for them. Covered by a test.
  • Gating on "had buffered data" keeps a flush() with nothing buffered a no-op, so it does not touch the explicit ref()/unref() state, which shares the same flag.
  • Windows is unaffected: has_pending_data() stays true while a uv_write is in flight (the release keeps happening in on_write on completion), and its stdout/stderr sinks are synchronous.
  • The one-shot writable poll armed by the buffering write is left in place: it is the writer's backstop for a later EAGAIN, it does not count toward loop liveness, and removing it would add an epoll_ctl to every console.write(). Under --hot that poll still wakes the watcher loop once, and that loop re-dispatches beforeExit on every wakeup (once a second even with an empty listener); that is a separate pre-existing run-loop issue, tracked separately.
  • Tests: test/js/bun/console/console-write.test.ts (the reported console.write() case) and test/js/bun/util/filesink.test.ts, "FileSink flush() from a 'beforeExit' listener": Bun.stdout.writer() and Bun.stderr.writer() on pipes, the EPIPE error arm (child's stdout is a socket whose peer is closed), and the could-not-drain case above (child's stdout is a socket whose send buffer the test filled first). Each fixture writes from the first beforeExit only and reports how many it saw: the first four report 2 on the released binary and 1 with this change; the last passes before and after, as intended.
  • Also ran against the debug build: the rest of filesink.test.ts, test/js/bun/console/, test/js/node/process/process.test.js (existing beforeExit keep-alive tests), test/js/bun/io/bun-write.test.js, test/js/web/fetch/blob-write.test.ts, test/js/bun/spawn/spawn-stdin-readable-stream.test.ts.

Background

  • FileSink: the native sink behind Bun.file(...).writer(), Bun.stdout.writer() and console.write(). On POSIX it wraps a PosixStreamingWriter (src/io/PipeWriter.rs) that coalesces writes smaller than a page into a buffer instead of issuing one write(2) per call; flush() pushes that buffer out.
  • Keep-alive: for a pipe or socket the sink owns a FilePoll. enable_keeping_process_alive / disable_keeping_process_alive add to or remove from the uSockets loop's active count, and VirtualMachine::is_event_loop_alive is true while that count (or a queue of pending tasks) is non-zero. The poll being registered with epoll/kqueue is separate from this count.
  • Deferred tasks / AutoFlusher: a queue run right after the microtask queue. A buffering write registers the sink there so the buffer goes out at the end of the current JS turn without an explicit flush(); on_auto_flush is that callback.
  • beforeExit (same in Node): emitted when the loop drains; if a listener makes the loop alive again, the runtime runs the loop and emits beforeExit again when it drains again. The bug was the sink claiming to have made the loop alive while it had nothing pending.

…uffer

A write() below the coalescing threshold buffers its bytes; on_write then
marks the event loop alive until they are flushed, and on_auto_flush
releases that once it drains them. flush_from_js drained them too but left
the mark in place until the next deferred-task drain, so console.write()
(write()+flush() on a stdout FileSink) called from a 'beforeExit' listener
with stdout on a pipe counted as newly scheduled work and 'beforeExit' was
emitted again, forever.

Release the keep-alive in flush_from_js when the flush emptied a non-empty
buffer, on the success and on the error path alike.
@coderabbitai

coderabbitai Bot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

@robobun, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 16 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: c2fce621-6a42-4930-9e7a-b2663103fadb

📥 Commits

Reviewing files that changed from the base of the PR and between 97a4363 and b2beceb.

📒 Files selected for processing (3)
  • src/runtime/webcore/FileSink.rs
  • test/js/bun/console/console-write.test.ts
  • test/js/bun/util/filesink.test.ts

Comment @coderabbitai help to get the list of available commands.

@robobun

robobun commented Aug 14, 2026

Copy link
Copy Markdown
Collaborator Author
Updated 11:50 AM PT - Aug 14th, 2026

@robobun, your commit b2beceb is building: #96431

@robobun

robobun commented Aug 14, 2026

Copy link
Copy Markdown
Collaborator Author

Status: fix pushed, waiting on CI.

Reproduced on the released binary (linux x64) with timeout 2 bun -e 'process.on("beforeExit", () => console.write("x\n"))' | wc -l printing several hundred thousand lines; the same listener with process.stdout.write or console.log prints 1. The new tests in test/js/bun/console/console-write.test.ts and test/js/bun/util/filesink.test.ts report 2 beforeExit emits on the released binary and 1 with this branch.

The --hot variant of the same listener still re-emits beforeExit after this change because the watcher run loop re-dispatches it on every wakeup (it does so once a second even with an empty listener); that is a separate pre-existing issue, tracked separately from this PR.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this PR and didn't find any bugs. The 10-line Rust change mirrors what on_auto_flush already does after a drain, is gated conservatively, and disable_keeping_process_alive is flag-idempotent so a redundant release is harmless. Because it touches FileSink's event-loop keep-alive accounting — subtle runtime lifecycle code — a human look would still be worthwhile.

What was reviewed:

  • flush_from_js gate: only releases when the buffer transitioned from non-empty to empty; the Pending/partial-drain path leaves has_pending_data() true so the keep-alive stays (covered by the could-not-drain test).
  • Error arm: POSIX flush() resets outgoing on Err, so the release fires there too — matches the EPIPE test.
  • Idempotency: FilePoll::disable_keeping_process_alive decrements by HasIncrementedActiveCount as u32, so a later on_auto_flush's update_ref(false) on the same sink is a no-op.
  • Windows test concern (Bun.stdout.writer() on a pipe being async) was raised and refuted.
Extended reasoning...

Overview

The PR fixes an infinite beforeExit loop when console.write() (or any Bun.file(fd).writer() write()+flush()) is called from a beforeExit listener with stdout/stderr on a pipe. The fix is ~10 lines in src/runtime/webcore/FileSink.rs: after flush_from_js's w.flush() returns, if the buffer went from non-empty to empty, call self.update_ref(false) to release the event-loop keep-alive that on_write had set. This is exactly what on_auto_flush already does after its own drain, so the change makes the explicit-flush path consistent with the deferred-flush path. Two test files add five subprocess-based regression tests (console.write, Bun.stdout/stderr.writer(), EPIPE error arm, and a could-not-drain case that verifies the keep-alive is not released when the flush left bytes buffered).

Security risks

None. This is process-lifecycle bookkeeping (uSockets loop active count) with no untrusted-input parsing, no auth/crypto, no memory-safety change. The added tests use local subprocesses and createSocketPair() from bun:internal-for-testing; no network access.

Level of scrutiny

Higher than a mechanical change: FileSink's keep-alive accounting is subtle and has been the source of leak/hang fixes before (the file already carries several such fixes with detailed comments). Getting it wrong causes either premature process exit (releasing too eagerly) or hangs (never releasing). I traced the change against the underlying FilePoll::disable_keeping_process_alive (idempotent via HasIncrementedActiveCount flag) and the POSIX flush() implementation (resets outgoing on both full-drain and error, leaves it non-empty on Pending/partial), and the gate is correct on every arm. The interaction with an explicit user sink.ref() is unchanged relative to the existing on_auto_flush path (both call update_ref(false) after a drain), so no new inconsistency is introduced.

Other factors

The PR description is unusually thorough — root cause traced through on_write → update_ref(evtloop, has_pending_data) → VirtualMachine::on_before_exit, Windows behavior explained (has_pending_data() stays true while a uv_write is in flight so the release still happens in on_write on completion), and the --hot watcher-loop wakeup called out as a separate pre-existing issue. Test coverage is strong: the could-not-drain test specifically guards against over-release (an unref'd timer only gets to run because pending bytes hold the loop open), and each fixture is designed so the pre-fix binary produces count: 2 and the fixed one produces count: 1. The bug-hunting system found no issues; one candidate (Windows async Bun.stdout.writer()) was raised and refuted. Given the change touches process-lifecycle-critical code, I'm deferring rather than approving, but I have no specific concerns to flag.

@Jarred-Sumner
Jarred-Sumner merged commit 9cff2a1 into main Aug 14, 2026
10 of 11 checks passed
@Jarred-Sumner
Jarred-Sumner deleted the farm/c8d1672c/filesink-flush-keepalive branch August 14, 2026 19:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants