Skip to content

PipeReader: do not read an fd again after a read on it fails - #43900

Merged
Jarred-Sumner merged 2 commits into
mainfrom
robobun/4d13bf92/pipe-reader-no-read-after-error
Sep 29, 2026
Merged

Jarred-Sumner merged 2 commits into
mainfrom
robobun/4d13bf92/pipe-reader-no-read-after-error

Conversation

@robobun

@robobun robobun commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

  • test/js/bun/spawn/spawn-stdio-syscall-error.test.ts is red on alpine: "lost": -82000, not 0 (build 120303). node:stream: stop 'data' and 'end' after destroy() on a native-backed Readable #43739 added the case, not the bug.
  • read_loop (src/io/PipeReader.rs:766) delivers the bytes read before a failed read, then the error. The consumer asks for more inside that delivery, and the reader reads the fd past the error. The error comes late, or never.

Fix

  • read_once sets PosixFlags::READ_FAILED on a fatal error, and begin_read does not read while it is set. The request parks, then on_reader_error rejects it. Nothing clears the flag.
  • Correct: libuv uv__read clears UV_HANDLE_READABLE before it reports a read error.
  • Verified: test/js/bun/spawn/spawn-stdio-syscall-error.test.ts, 17 pass. The four new cases fail without the fix. Suites: Notes.
  • Self-reviewed: 15 concerns raised, 14 addressed. Not taken: a flag reset in start(), which nothing on main needs.

Background

  • PosixBufferedReader reads the fd behind subprocess stdio, file streams and the shell. Its parent gets on_read_chunk, then on_reader_done or on_reader_error.
  • FileReader is the parent behind a ReadableStream. on_read_chunk resolves the parked pull(), and the reaction runs before it returns.
  • Considered: a repeating shim failure hides the lost error. An error stored in FileReader first needs a callback in every parent.
  • PipeReader: do not read again after a failed read #43920 is a newer PR with the same change. Its test trigger is here.

Downsides

  • After a read error, a stream that ended with 'end' now ends with 'error', as in Node. With no listener the process stops.
  • After a read error, bun run --filter no longer drains that pipe at exit.
  • Reads that do not fail pay nothing: begin_read tests one more bit.
Notes

Trace without the fix (bun 1.4.3-canary.1+367d939d9, shim logs each recv(), writer dd bs=1025). The debug build of main at 8d36bff does the same: 3 of 8 runs, two with no 'error' and lost -7714150:

recv #5 len=262144 -> 95325
recv #6 len=166819 -> EIO (injected)      same fill_scratch call as #5
JS data 95325
recv #7 len=65536 -> 65536                 pull from inside the delivery, read_into
recv #8 .. #116                            to EOF
{"received":452025,"got":8000125,"lost":-7548100,"events":["stdout.close","close"]}

No 'error' event: the reads reached EOF before on_reader_error ran, and a stored error is only returned by a later pull. When the reads park first, on_reader_error rejects that pull and 'error' comes late. That is the CI signature: the events match and lost is a negative multiple of 1025.

Why alpine. In CI the failure is injected: the shim fails only the Nth recv(), so a later recv() succeeds and the extra bytes show. The failing recv() must follow bytes in the same wakeup. BusyBox head writes 1025 bytes at a time, so a wakeup often holds a short recv() and then the failing one. coreutils head fills the buffer in one recv(). The same run fails on debian when the writer is slow. With a writer that copies BusyBox head (stdio, 1024-byte buffer), the RECV_AT=6 case fails 7 of 40 runs on bun 1.4.3-canary (release), and 36 of 40 when the writer also spins between chunks. This branch (debug): 0 of 120, and 0 of 60 with dd bs=1025.

With no shim. A child with one AF_UNIX socket as fd 0 and fd 1 writes a line to stdout and reads stdin. The peer leaves that line unread, sends 8192 bytes and closes. The kernel gives the child 8192 bytes, then ECONNRESET, then EOF.

Runtime stdin events
Node v26.3.0 'error' ECONNRESET after 8192 bytes
bun 1.4.3-canary.1+367d939d9 'end' after 8192 bytes, no error
this branch 'error' ECONNRESET after 8192 bytes

The new cases. Each one reaches the reader from inside the delivery in a different way. Whole-file runs of the describe block:

Case Entry Release, no fix Debug, no fix Debug, guard in read_into only Debug, this branch
child_process, 'data' listener pull, read_into fails 6 of 6 fails 5 of 5 passes passes
child_process, 'readable' and read() set_flowing(true), read fails 6 of 6 fails 5 of 5 fails 3 of 3 passes
Bun.spawn, lazy pull, read_into fails 6 of 6 fails 5 of 5 passes passes
Bun.spawn, reader started at spawn pull, read_into fails 6 of 6 fails 5 of 5 passes passes
  • The first three use counts: SPAWN_FAULT_RECV_CAP=4096 makes every recv() short, so fill_scratch calls recv() again in the same wakeup. SPAWN_FAULT_RECV_EAGAIN_AT=2 ends the first read loop, so the consumer's read parks and the next read is poll-driven. SPAWN_FAULT_RECV_AT then fails after bytes in that wakeup. For the 'readable' case it is 19: Copy source lines when generating error messages #3 to Finish implementing React Fast Refresh transforms #18 return 64 KiB, the highWaterMark, so the reader is stopped and read() starts it again.
  • The fourth uses state, and comes from PipeReader: do not read again after a failed read #43920: the writer waits for a line on stdin, so the first read is parked when the bytes arrive. SPAWN_FAULT_RECV_MID_FILL=1 fails the recv() that follows one that returned bytes, and SPAWN_FAULT_READS_AFTER counts the recv() calls after it. Without the fix it is 1.
  • With a count-based trigger and the reader started at spawn, the case passed without the fix on a debug build: the buffered reader that runs before JS reads .stdout took the bytes and the error. That is why this case uses state.
  • With the BusyBox-like writer at four speeds, the whole file passes 20 of 20 on this branch.

With the fix, CAP=4096 EAGAIN_AT=2 RECV_AT=5:

recv #3 -> 4096, recv #4 -> 4096, recv #5 -> EIO
JS data 8192
close(fd)
JS error EIO

Node. libuv uv__read (src/unix/stream.c): on a read error other than EAGAIN it clears UV_HANDLE_READABLE | UV_HANDLE_WRITABLE, calls read_cb with the error, then stops the watcher. It calls read_cb once for each read(), so it never holds bytes and an error from one batch.

Placement. EOF and the maxBuffer stop have the same guard at this site: close_if_final closes the reader before the final chunk is delivered. An error cannot use it, because a closed reader with no stored error reads as a clean end. For the same reason READ_FAILED is not part of is_done().

Parents (14 BufferedReaderParent implementations, what each does in on_reader_error):

  • 3 release the fd: FileReader, SubprocessPipeReader, Terminal.
  • 2 drop the reader: FileResponseStream, shell subproc.rs.
  • 8 only do accounting: filter_run.rs, multi_run.rs, lifecycle_script_runner.rs, security_scanner.rs, git_runner.rs, both cron jobs, test Worker.rs. lifecycle_script_runner.rs and cron.rs build a new reader with init() for each spawn.
  • 1 is shared and can restart: the shell IOReader.

Only two read the same reader after an error. filter_run.rs drain_and_close_pipes reads once more at exit. That read is now a no-op, and deinit() follows. Before, it could reach a second terminal callback and decrement remaining_fds twice. The shell IOReader::start() does not restart a reader after a failed read on main, because the fired one-shot poll still counts as registered. The Windows reader gets one libuv callback for each read, so it has no bytes-then-error batch.

Left open.

Suites run on the debug build: test/js/web/streams/streams.test.js (624 pass), test/js/node/stream/node-stream.test.js (112 pass), test/js/bun/spawn/spawn-streaming-stdout.test.ts (pass), test/js/node/child_process/child_process.test.ts (81 pass, 2 fail in my container for reasons outside this change: "spawn in the default shell" reads an empty $SHELL, and "extra stdio pipes are not double-closed on GC" needs 5.0 s in a debug build against the 5 s timeout, its script prints OK). With this branch's build, the test file of #43920 passes 17 of 17 in 3 runs.

#43790 is open and edits the comment above the failing case. It changes the event order in native-readable.ts, not the reader.


no test proof · iteration 0 · platform-specific test(s) that do not run on this machine, deferring to CI, which covers all platforms: test/js/bun/spawn/spawn-stdio-syscall-error.test.ts

PosixBufferedReader delivers the bytes read before a failed read, then
the error. The delivery resolves a parked pull, and the consumer pulls
again from inside it. The reader then read the fd past the error, so the
error arrived after bytes read later, or never when those reads reached
EOF.

read_once now sets PosixFlags::READ_FAILED on a fatal error and
begin_read refuses to read while it is set. The pull from inside the
delivery parks, and on_reader_error rejects it.

The test shim can cap each recv() and fail one with EAGAIN. Two new
cases use that to put the failing recv() after bytes in the same wakeup
with a pull parked.
@robobun

robobun commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator Author

Status

Reproduced on debian x64, with no alpine machine.

  • The red case: with a writer that writes 1024-byte chunks (as BusyBox head does), "stdout delivers every byte read before the error" gets a negative lost in 7 of 40 runs on bun 1.4.3-canary, and in 0 of 120 runs with this fix.
  • Four new cases in test/js/bun/spawn/spawn-stdio-syscall-error.test.ts force the same order of reads on glibc. Without the change to src/io/PipeReader.rs they fail in 5 of 5 runs on a debug build and in 6 of 6 runs on the 1.4.3-canary release build. With it the file gives 17 pass.
  • With no shim: a child whose stdin is a socket gets 'end' on 1.4.3-canary and 'error' (ECONNRESET) on Node v26.3.0 and on this branch, when the peer sends 8192 bytes and closes with data unread.

PR: #43900

@coderabbitai

coderabbitai Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: oven-sh/bun/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Essentials

Run ID: 4066b203-0c83-4c8a-8c60-924d63ae22b8

📥 Commits

Reviewing files that changed from the base of the PR and between 7de08b4 and 5f57c99.

📒 Files selected for processing (2)
  • src/io/PipeReader.rs
  • test/js/bun/spawn/spawn-stdio-syscall-error.test.ts

Included review availability: Your plan provides up to 10 included reviews per hour; 4 remain after this review.


Walkthrough

The POSIX pipe reader now stops starting reads after a non-retry read error. The syscall-fault tests configure receive behavior and verify that Bun streams and node:child_process deliver bytes read before EIO, then report the error.

Changes

Pipe reader error handling

Layer / File(s) Summary
Track failed reads
src/io/PipeReader.rs
PosixFlags adds READ_FAILED. begin_read declines to start a read when this flag or IS_PAUSED is set. read_once sets the flag for non-retry errors.
Test delivery before syscall errors
test/js/bun/spawn/spawn-stdio-syscall-error.test.ts
The syscall-fault shim can inject EIO after a successful receive, inject EAGAIN on a selected call, and cap receive sizes. New tests check that Bun streams and node:child_process consumers deliver bytes received before EIO. They also check that no receive occurs after a mid-fill failure.

Suggested reviewers: jarred-sumner, dylan-conway

Priority: ➖ Normal

Merge Risk: ⚪ Minimal · up to 5f57c

No actionable merge-blocking risk is established for this change after normal checks.

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly and concisely describes the main change: preventing another fd read after a read failure.
Description check ✅ Passed The description explains the problem, the fix, tradeoffs, affected parents, and verification results. It does not use the template headings exactly, but it provides the required content.

Comment @coderabbitai help to get the list of available commands.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this PR and didn't find any bugs. Because PosixBufferedReader sits under every Bun.spawn/child_process stdio, file stream and shell read, and the gate turns any non-EAGAIN read error into a terminal state for the reader, a human look at that trade-off would still be worthwhile.

What was reviewed:

  • Flag lifecycle: READ_FAILED is set on the &mut self error arm of read_once, which is the only place Stop::Error is produced; read, read_into and on_poll all go through begin_read, so no entry can read the fd again. Nothing clears it, but every reader is started once and every on_reader_error parent checked (SubprocessPipeReader goes to State::Err + deinit, FileReader calls close_after_error) ends the reader, so reuse/transfer via from() after an error does not occur.
  • No hang window: in read_loop the chunk is delivered and on_error runs in tail position before any is_done check, and read_into dispatches on_error immediately, so a pull parked by the gate is always rejected in the same wakeup.
  • Test: traced the shim ordering (RECV_CAP 4096, EAGAIN_AT 2, RECV_AT 5) through fill_scratch on a socket; #3/#4/#5 land in one call so the bytes-then-error batch is forced. As the description notes, only the child_process case is a deterministic guard for the fix; the Bun.spawn case can pass on either build.
Extended reasoning...

The change adds one bitflag to src/io/PipeReader.rs, sets it on a fatal read error in read_once, and has begin_read refuse to read while it is set, plus two new fault-injection knobs and two test cases in test/js/bun/spawn/spawn-stdio-syscall-error.test.ts. It touches no security-sensitive surface (no auth, injection, or data exposure); the Windows reader is unaffected. The Rust diff is 12 lines and internally consistent with the existing close_if_final guard for EOF/maxBuffer, and no CODEOWNER covers the changed files. Deferring rather than approving because the reader is shared core I/O for subprocess, stream and shell paths and the fix makes every non-retry read error terminal, a behavior decision the author calls out as a downside and a maintainer should weigh.

…ntry

The doc comment on READ_FAILED now says that nothing clears the bit.

The tests now cover each way a consumer reaches the reader from inside
the delivery of the bytes read before a failed read:
- node:child_process with a 'data' listener (a pull, read_into).
- node:child_process with 'readable' and read(): the reader is stopped
  at the highWaterMark and read() starts it again (read).
- Bun.spawn with lazy: the reader starts at the first read.
- Bun.spawn with the reader started at spawn. The writer waits for a
  line on stdin, so the first read is parked when the bytes arrive. The
  shim fails the recv() that follows the one that returned bytes and
  counts the recv() calls after it. This trigger comes from #43920.

Each of the four fails without the change to begin_read, on a debug
build and on a release build.
@robobun

robobun commented Sep 25, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 8:56 PM PT - Sep 24th, 2026

✅ @robobun, your commit 5f57c99f37294762f809f8eb6b9e98ff41e7d4c2 passed in Build #120525! 🎉


🧪   To try this PR locally:

bunx bun-pr 43900

That installs a local version of the PR into your bun-43900 executable, so you can run:

bun-43900 --bun

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed the current revision (including the second commit's four new read-entry tests) and did not find any bugs. Because this changes when the shared POSIX reader stops reading an fd, and that reader sits under Bun.spawn, node:child_process, file streams and the shell, a maintainer's look is still worthwhile — in particular on the deliberate choice that READ_FAILED is never cleared (a non-repeating errno such as ENOMEM now always ends the stream).

What was reviewed:

  • Traced READ_FAILED through every fd-reading entry: read(), on_poll() and read_into() all go through begin_read, so a re-entrant pull from inside on_read_chunk parks instead of reading past the error; the Stop::Error arm still dispatches on_reader_error in tail position after the pre-error bytes are delivered.
  • Checked the flag survives from(), start() and unpause(), and that SubprocessPipeReader never calls from() on an errored reader (its State::Err arm builds from_bytes_then_error), so the error cannot be dropped by the buffered-to-streaming handoff.
  • Checked every on_reader_error implementation ends or abandons its reader (FileReader close_after_error, SubprocessPipeReader deinit, the remaining_fds counters); none re-starts a read afterwards.
  • Test shim knobs are all read by the shim and reset per fd; each fixture wires the error/close events into the asserted output and the fixtures stay Linux-only behind the existing skipIf(!isLinux || !cc).
Extended reasoning...

The diff adds a READ_FAILED bit to PosixFlags in src/io/PipeReader.rs, sets it in read_once on a non-EAGAIN errno and gates begin_read on it, plus about 200 lines of fault-injection tests in test/js/bun/spawn/spawn-stdio-syscall-error.test.ts. It touches no auth, crypto or injection surface; the sensitive part is re-entrancy in a raw-pointer read loop that runs user JS. The core Rust change is 12 lines and reads as correct, but it is a behavioral change in a shared I/O primitive used by subprocess, stream, shell and installer code, with a design choice (flag never cleared, Windows reader intentionally untouched) that a maintainer should sign off on rather than an automated approval.

@Jarred-Sumner
Jarred-Sumner merged commit dc3660d into main Sep 29, 2026
10 checks passed
@Jarred-Sumner
Jarred-Sumner deleted the robobun/4d13bf92/pipe-reader-no-read-after-error branch September 29, 2026 22:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants