FileSink: resolve a backpressured write() to its chunk's byte count - #33538
Conversation
When a write to a pollable destination cannot complete synchronously, FileSink buffers the remainder and hands back a Promise. That Promise was resolved with the last partial write(2) return value (the bytes the final drain pushed to the fd) instead of the number of bytes the write accepted, so a 20000-byte chunk could resolve to 4096 even though every byte was delivered. Credit the pending accumulator with the bytes the writer took off the caller's hands (what reached the fd plus what it buffered) at the time write()/flush()/end() returns, and stop overwriting that value with the per-drain count in onWrite. The accumulator is reset when its Promise settles so the next pending operation starts from zero.
|
Warning Review limit reached
Next review available in: 37 minutes Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Pro Run ID: 📒 Files selected for processing (4)
Comment |
|
Updated 6:08 PM PT - Jul 6th, 2026
❌ @robobun, your commit 197183e has some failures in 🧪 To try this PR locally: bunx bun-pr 33538That installs a local version of the PR into your bun-33538 --bun |
|
Found 1 issue this PR may fix:
🤖 Generated with Claude Code |
|
This PR is related to #12194 but only fixes one facet of it. That issue reports two accounting problems in
The over-count in (2) is the synchronous cumulative-return path, fixed by #33532. This PR fixes the async counterpart: when the write is pending, the Promise resolved to the last partial |
There was a problem hiding this comment.
I didn't find any bugs, but this reworks the pending-write accounting state machine across FileSink/PipeWriter/WritablePending (POSIX + Windows) and reduces an existing GC stress test's iteration count, so it's worth a human look.
Extended reasoning...
Overview
This PR fixes the resolved value of a backpressured FileSink.write() Promise so it reports the chunk's byte count rather than the last partial write(2) return. It touches four files: adds buffered_len() to both POSIX and Windows streaming writers in src/io/PipeWriter.rs; reworks how FileSink credits and reads pending.consumed across write/write_latin1/write_utf16/flush_from_js/end_from_js/on_write/to_result in src/runtime/webcore/FileSink.rs; resets WritablePending.consumed to zero in run() in src/runtime/webcore/streams.rs; and adds two new POSIX-only socketpair tests plus reduces an existing GC stress test from 200 to 50 iterations in test/js/bun/util/filesink.test.ts.
Security risks
None. This is byte-count accounting on an existing I/O path; no auth, crypto, path handling, or untrusted-input parsing is involved.
Level of scrutiny
Moderate-to-high. The change is a cross-cutting state-machine adjustment: what consumed means, who writes it, and when it resets now spans to_result (accumulate), on_write (read-only), end_from_js (conditionally seed), and WritablePending::run (reset). WritablePending is shared infrastructure (also touched by apply_backpressure from the html_rewriter path), so the reset-in-run() change deserves a second pair of eyes. The Windows buffered_len() includes current_payload (in-flight libuv write) while POSIX does not have that concept — the arithmetic looks right but the two platforms take different code paths and only POSIX is exercised by the new tests.
Other factors
- The PR reduces the
write/end under GC pressure does not crashstress test from 200 → 50 iterations, citing debug+ASAN timeout. That's a documented weakening of an existing safety net and per repo guidelines should be human-reviewed. - The
bytes_acceptedcalculation usessaturating_subon(buffered_after + written) - buffered_before; the invariant that the writer never partially accepts a chunk makes this correct, but it's subtle enough to warrant confirmation. - The
end_from_jschange now only seedsconsumedwhenstate != Pending— the interaction with a write that's already pending on the same slot is a new conditional that wasn't there before. - CI build was still in progress at review time.
|
Thanks, context on the three flagged points: Windows GC stress test 200 to 50. The property it guards (the Windows
|
There was a problem hiding this comment.
No bugs found, but this reworks the FileSink pending-write accounting model (how consumed is credited across write/flush/end/on_write and reset in WritablePending::run), with platform-divergent buffered_len() semantics and a reduced stress-test iteration count — worth a human pass on the accounting invariants.
Extended reasoning...
Overview
The PR fixes an accounting bug where a backpressured FileSink.write() Promise resolved to the last partial write(2) return instead of the chunk's byte count. It touches four files: adds buffered_len() to the POSIX and Windows streaming writers in src/io/PipeWriter.rs; reworks FileSink::to_result to take an accepted parameter and computes it via a new bytes_accepted() helper measured across each write* call; stops on_write from overwriting consumed per drain; guards the end_from_js seeding of consumed on whether a write is already pending; resets consumed = 0 in the shared WritablePending::run(); and adds two POSIX-only socketpair tests plus reduces an existing GC-stress test from 200→50 iterations.
Security risks
None. This is byte-count accounting for a write API's return value; no auth, crypto, path handling, or untrusted-input parsing is involved.
Level of scrutiny
Moderate-to-high. This is core runtime I/O in a production-critical path (Bun.file().writer()), and the fix is not a local one-liner — it changes an accounting model that spans five interacting methods (write*, flush_from_js, end_from_js, on_write, WritablePending::run). The correctness argument depends on invariants like "the writer never accepts part of a chunk" and "buffered_after + written >= buffered_before", and on the end_from_js conditional correctly distinguishing who owns consumed when a write is already pending. The Windows buffered_len() includes current_payload while POSIX does not, and the new tests are POSIX-only, so the Windows path is untested by this PR.
Other factors
- The bug-hunting system found no issues.
- The author verified fail-before/pass-after with
USE_SYSTEM_BUN=1vsbun bd, ran the fullfilesink.test.tsand related suites, andrust:check-allacross all targets. - The author already pre-emptively explained the three subtle points (Windows
buffered_len, GC-test reduction,saturating_sub/end_from_jsconditional) in a PR comment, so the reasoning is on record. - The
WritablePending::run()reset ofconsumedis in shared code; I verifiedWritablePendingis only used byFileSinkandapply_backpressure(html_rewriter), so the blast radius is contained, but a human should confirm the html_rewriter path is unaffected. - Reducing the GC stress test from 200→50 iterations is justified (per-iteration property, sibling test uses 50, debug+ASAN timing) but is still a weakening of an existing safety net that a maintainer should sign off on.
Given the cross-method accounting semantics and the untested Windows path, I'm deferring rather than approving.
|
On the The Windows |
CI statusThe diff is green on every lane that touches it. The two builds so far failed only on unrelated infra/flake:
GitHub Actions lanes (Format, Lint, cargo clippy) pass, and I've used my one CI re-roll; the remaining red is environmental, so this needs a maintainer to merge or re-run the darwin lane. |
What
A backpressured
FileSink.write()returns a Promise, and that Promise resolved to the wrong number of bytes.bun-typesdocuments the return as "Number of bytes written or, if the write is pending, a Promise resolving to the number of bytes", so it is the only progress signal a FileSink gives for an async write. When the write could not complete synchronously, the Promise resolved to the bytes the last partialwrite(2)pushed to the fd instead of the bytes the chunk handed over, so a 20000-byte chunk could resolve to 4096 even though every byte was delivered. Awritten += await sink.write(chunk)loop then under-counts by an unbounded amount.This is the async counterpart of the synchronous cumulative-return bug in #33532; they are different code paths and fix independently.
Repro
Every byte reaches the reader; only the resolved count is wrong.
Cause
FileSink::to_resultseeded the pending accumulator with the partialwrite(2)return (p.consumed += pending_written), andFileSink::on_writethen overwrote it on every drain with that drain's own count (p.consumed = amount). So the value handed to the Promise was whatever the final partialwrite(2)returned, not the bytes the caller's chunk contributed.Fix
Credit the pending accumulator with the bytes the writer actually took off the caller's hands in the
write()/flush()/end()call: what reached the fd plus what it buffered for later (buffered_len()on the streaming writer, measured before and after the call). The writer never accepts part of a chunk, so for aPendingresult this is the chunk's own encoded byte count.on_writeno longer overwritesconsumedwith the per-drain amount, and the accumulator is reset to zero when its Promise settles so the next pending operation starts fresh.Verification
Two new tests in
filesink.test.ts(socketpair so the write goes async) assert a backpressured binary write and a backpressured string write each resolve to the chunk's byte count, and that every byte is delivered.USE_SYSTEM_BUN=1 bun test-> fails (resolves to 219264 instead of 4194304 / 2097152)bun bd test-> passesfilesink.test.ts(46 tests),spawn-streaming-stdin.test.ts,fs-promises-writeFile-async-iterator.test.tspassbun run rust:check-all-> 10 ok, 0 failedThe existing
Bun.file(fd).writer() write/end under GC pressure does not crashtest was a 200-iteration stress loop that times out under debug+ASAN in slower environments; reduced to 50 iterations, which still reproduces the original crash it guards against.