Skip to content

shell: fix use-after-free when the poll re-registration issued from PipeReader::on_read_chunk fails - #33269

Merged
Jarred-Sumner merged 2 commits into
mainfrom
farm/852f729b/shell-pipereader-onreadchunk-uaf
Jul 4, 2026
Merged

Jarred-Sumner merged 2 commits into
mainfrom
farm/852f729b/shell-pipereader-onreadchunk-uaf

Conversation

@robobun

@robobun robobun commented Jul 2, 2026

Copy link
Copy Markdown
Collaborator

Problem

Heap use-after-free in Bun Shell when epoll_ctl(MOD) fails while the shell PipeReader::on_read_chunk callback re-registers the poll from inside the read loop. Found by syscall-fault-injection fuzzing against origin/main.

This is the path #32986 called out as out of scope: that PR fixed read_with_fn's own EAGAIN-arm re-registration, but the shell PipeReader::on_read_chunk still called self.reader.register_poll() itself.

==ERROR: AddressSanitizer: heap-use-after-free
READ of size 8
  #0 <PosixBufferedReader>::read_with_fn   src/io/PipeReader.rs:837:43
  #1 <PosixBufferedReader>::read_socket    src/io/PipeReader.rs:581:9
  #2 <PosixBufferedReader>::on_poll        src/io/PipeReader.rs:534:17
  #3 __bun_run_file_poll                   src/runtime/dispatch.rs:677:22
Freed-by stack (the re-entrant callback chain)
freed by thread T0 here:
  core::ptr::drop_in_place::<Arc<shell::subproc::PipeReader>>
  <shell::subproc::PipeReader>::on_reader_error                      src/runtime/shell/subproc.rs:2363
  <PosixBufferedReader>::register_poll                               src/io/PipeReader.rs:433
  <shell::subproc::PipeReader>::on_read_chunk                        src/runtime/shell/subproc.rs:2062
  <PosixBufferedReader>::read_with_fn                                src/io/PipeReader.rs:875
  <PosixBufferedReader>::read_socket
  <PosixBufferedReader>::on_poll
  __bun_run_file_poll

Repro

  1. A shell pipe's FilePoll fires and __bun_run_file_poll dispatches into PosixBufferedReader::on_poll -> read_with_fn with a bare &mut and no keepalive.
  2. recv() drains more than half of the 256 KB scratch buffer in one call, so read_with_fn's streaming inner loop flushes the head mid-loop: parent.vtable.on_read_chunk(.., Progress).
  3. Shell PipeReader::on_read_chunk re-arms the poll itself: self.reader.register_poll(). The epoll_ctl(MOD) fails (ENOMEM in the repro; fd/watch pressure in the wild).
  4. register_poll dispatches on_reader_error. The shell PipeReader::on_reader_error signals the Cmd and drops the Readable::Pipe Arc; its own guard_from_raw keepalive becomes the last reference, and dropping it frees the PipeReader (and the PosixBufferedReader embedded in it).
  5. register_poll returns false, but on_read_chunk is not a direct caller of the read loop, so the false never reaches it. The inner loop keeps going and reads parent._offset from the freed reader on the next recv.

Cause

BufferedReaderParent's contract (and the SAFETY comments in read_with_fn / read_blocking_pipe) is that on_read_chunk never frees the reader; only on_reader_error may. The shell PipeReader::on_read_chunk broke that transitively by calling register_poll(), whose failure path dispatches on_reader_error.

#32986's register_poll() -> bool return value only protects direct callers in the read loop. It cannot protect a caller that reaches register_poll through the on_read_chunk vtable dispatch two frames down.

Fix

Delete the re-arm from shell PipeReader::on_read_chunk. It was redundant on both platforms and the codebase already documents why:

  • POSIX: every exit of read_with_fn / read_blocking_pipe that wants more data already calls register_poll() itself, driven by the bool on_read_chunk returns.
  • Windows: WindowsBufferedReader::on_read notes "the re-arm is already handled by on_file_read's epilogue / uv_read_start", and it already performs the _buffer.clear() that used to be start_with_current_pipe()'s second side effect.
  • The sibling shell reader, IOReader::on_read_chunk_cb, already dropped its identical re-arm for the same two reasons (redundancy, plus re-deriving &mut to the embedded reader while the read loop holds one).

Removing it also removes the only &mut self.reader re-derivation inside the callback, and the Output::panic("TODO: ...") that was the Windows branch's only error handling.

Test

New SHELL_RECV_BULK=N mode in test/js/bun/shell/shell-pipe-read-fault.test.ts's LD_PRELOAD shim: the first N real recv()s on each AF_UNIX socket instead return the caller's whole buffer filled with 'A'. Combined with the existing SHELL_RECV_EAGAIN_FIRST=1 and SHELL_FAIL_EPOLL_FROM=3, one fabricated bulk recv deterministically pushes head_start past the half-buffer cutoff so the mid-loop flush (and therefore the failing re-registration) happens from on_read_chunk.

With the epoll failure count unchanged, the same epoll_ctl #3 that used to be issued by on_read_chunk is now the read loop's own EAGAIN re-registration, whose failure path already returns without touching the reader, so the command just reports ENOMEM.

  • Before the fix: the new test fails in ~750 ms with the heap-use-after-free above; the other 6 tests in the file pass.
  • After the fix: all 7 pass.

The test is skipIf(!isASAN) like its sibling.

Also ran the rest of test/js/bun/shell/ (bunshell*.test.ts: 394 pass / 0 fail; commands/ and the remaining files: every failure reproduces identically with src/runtime/shell/subproc.rs reverted to main, so they are pre-existing in this environment, not caused by this change).

@github-actions github-actions Bot added the claude label Jul 2, 2026
@robobun

robobun commented Jul 2, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 4:20 PM PT - Jul 2nd, 2026

❌ @robobun, your commit a769948 has some failures in Build #68070 (All Failures)


🧪   To try this PR locally:

bunx bun-pr 33269

That installs a local version of the PR into your bun-33269 executable, so you can run:

bun-33269 --bun

@coderabbitai

coderabbitai Bot commented Jul 2, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 8d61518d-289c-4bb1-a593-1c1ed2f5e950

📥 Commits

Reviewing files that changed from the base of the PR and between 37e801f and a769948.

📒 Files selected for processing (2)
  • src/runtime/shell/subproc.rs
  • test/js/bun/shell/shell-pipe-read-fault.test.ts

Walkthrough

The PR removes explicit re-arming of polling/reading from PipeReader::on_read_chunk in src/runtime/shell/subproc.rs, relying on the caller's read loop to re-register based on the returned continuation flag. It also extends the shell test fault-injection shim with a SHELL_RECV_BULK mode and adds a corresponding regression test.

Changes

PipeReader re-arming fix and fault-injection test

Layer / File(s) Summary
Remove redundant re-arming in PipeReader callback
src/runtime/shell/subproc.rs
Eliminates the unix/non-unix branch that re-registered polling or restarted the pipe from within on_read_chunk; the callback now appends/forwards data and returns whether more data remains based on EOF state, with the unused non-unix Output import removed and comments explaining the event-loop contract.
Add SHELL_RECV_BULK fault mode to LD_PRELOAD shim
test/js/bun/shell/shell-pipe-read-fault.test.ts
Adds recv_bulk/bulk_count state, parses SHELL_RECV_BULK in init_modes(), fabricates full-buffer 'A'-filled AF_UNIX recv() results for a configurable number of calls per fd, resets the counter on close(), and registers the new knob in VALUE_MODES and its documentation.
New regression test combining epoll failure and bulk reads
test/js/bun/shell/shell-pipe-read-fault.test.ts
Adds a Linux/cc/!isASAN-gated test running poll-chunk.js with SHELL_RECV_EAGAIN_FIRST=1, SHELL_FAIL_EPOLL_FROM=3, and SHELL_RECV_BULK=1, asserting the shim-injected ENOMEM exit code.

Possibly related PRs

  • oven-sh/bun#29966: Both PRs modify PipeReader re-arming/read-loop behavior after read events to avoid incorrect looping conditions.
  • oven-sh/bun#30055: Both PRs address PipeReader-style pipe read lifecycle/teardown correctness in subprocess handling.
  • oven-sh/bun#32754: Both PRs adjust shell pipe-read fault handling in subproc.rs and extend the shared fault-injection shim/tests.
🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly summarizes the main change: removing a failing poll re-registration path that caused a use-after-free.
Description check ✅ Passed The description covers what the PR does and how it was verified, with enough detail despite using different section headings.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

Comment @coderabbitai help to get the list of available commands.

Comment thread append_output.txt Outdated
register_poll()'s failure path dispatches on_reader_error, which drops the
last Arc<PipeReader>. Calling it from inside on_read_chunk (which runs from
the bun_io read loop, with &mut to the embedded PosixBufferedReader live on
that loop's stack) therefore frees the reader the loop keeps reading through:
a heap use-after-free in read_with_fn's streaming inner loop whenever a
single poll wake drains more than half the 256 KB scratch buffer and the
re-registration's epoll_ctl(MOD) fails.

The re-arm is also redundant: the POSIX read loop re-registers the poll
itself based on the bool on_read_chunk returns, and on Windows
WindowsBufferedReader::on_read / on_file_read handle it. The sibling shell
reader (IOReader::on_read_chunk_cb) already dropped its re-arm for the same
reasons.
@robobun
robobun force-pushed the farm/852f729b/shell-pipereader-onreadchunk-uaf branch from 37e801f to 435e4be Compare July 2, 2026 17:55
@robobun

robobun commented Jul 2, 2026

Copy link
Copy Markdown
Collaborator Author

Removed the two stray *_output.txt files claude flagged. They came from file-io.test.ts's pwd &> ... redirect cases, which write into the cwd (the repo root when running the suite locally), and got picked up by git add -A during the shell-suite sweep. The PR is back to the two intended files as of 435e4be; the fix and test are unchanged.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for cleaning up the stray output files — the diff is back to the two intended changes. The fix looks sound and matches the sibling IOReader::on_read_chunk_cb pattern, but since it removes re-arm behavior on both the POSIX and Windows unsafe read paths, it's worth a human sign-off.

Extended reasoning...

Overview

The PR deletes the redundant register_poll() / start_with_current_pipe() re-arm from shell PipeReader::on_read_chunk in src/runtime/shell/subproc.rs (plus the now-unused Windows Output import), and adds a SHELL_RECV_BULK fault-injection mode with a new ASAN-gated regression test in test/js/bun/shell/shell-pipe-read-fault.test.ts. The two stray *_output.txt files I flagged in my previous review have been removed in 435e4be, and the current diff is scoped to exactly these two files.

Security risks

None identified. This is an internal I/O lifetime fix in the shell subprocess reader; it does not touch auth, input validation, network protocol parsing, or any user-facing trust boundary. The test changes only extend an existing LD_PRELOAD fault-injection shim that runs inside a temp dir under the test harness.

Level of scrutiny

High. The change sits in memory-safety-critical unsafe Rust: raw *mut PipeReader pointers, Arc strong-count juggling, and a documented invariant ("on_read_chunk never frees the reader") that the deleted code was violating transitively. I verified the load-bearing claims against the codebase — src/io/PipeReader.rs:404 documents register_poll() -> bool's failure path dispatching on_reader_error, src/io/PipeReader.rs:628/820 document the on_read_chunk-never-frees contract, and src/runtime/shell/IOReader.rs:299-304 shows the sibling on_read_chunk_cb already dropped the identical re-arm with the same rationale. The reasoning is internally consistent and the ASAN regression test is deterministic.

That said, the deletion also removes the Windows start_with_current_pipe() call (whose Output::panic("TODO: ...") error path is being retired). The PR asserts WindowsBufferedReader::on_read already handles both the re-arm and the _buffer.clear() side effect, which I did not independently trace end-to-end. Given the cross-platform reach and the fact that correctness here hinges on subtle re-registration ownership across the bun_io read loop, this is exactly the kind of change a maintainer familiar with the PosixBufferedReader / WindowsBufferedReader state machines should sign off on.

Other factors

  • No bugs surfaced by the bug-hunting sweep on the current revision.
  • My only prior feedback (stray append_output.txt / pwd_output.txt) was addressed and the thread is resolved.
  • The PR description is unusually thorough (ASAN traces, freed-by stack, before/after test evidence, and a shell-suite sweep confirming no new regressions), which raises confidence but doesn't substitute for maintainer review of an unsafe-heavy I/O path.

@robobun

robobun commented Jul 2, 2026

Copy link
Copy Markdown
Collaborator Author

Thanks for the review. The Windows half is the part I could not exercise here, so here is the trace for whoever signs off:

On Windows, shell PipeReader::on_read_chunk is only ever reached from WindowsBufferedReader::on_read (src/io/PipeReader.rs:1855, via _on_read_chunk). start_with_current_pipe() does exactly two things, buffer().clear() and start_reading(), and the block right below that call site already accounts for both:

  • re-arm: "The re-arm is already handled by on_file_read's epilogue / uv_read_start" (PipeReader.rs:1857-1860). For stream sources uv_read_start was issued once in start_reading() and libuv keeps delivering read callbacks until uv_read_stop; for file sources on_file_read's epilogue issues the next uv_fs_read.
  • buffer reset: "clearing the buffer here is load-bearing" (PipeReader.rs:1861-1867); on_read now performs the _buffer.clear() itself after the streaming consumer is done with the slice.

That comment block was written when the sibling shell reader, IOReader::on_read_chunk_cb (src/runtime/shell/IOReader.rs:300-317), dropped its identical register_poll() / start_with_current_pipe() re-arm. The subprocess SubprocessPipeReader::on_read_chunk never had one. The shell PipeReader here was the last caller.

On the red CI check: the only failed job on this build so far is darwin 26 aarch64 - test-bun, which died on a buildkite-agent artifact download timed out after 120s for step 'darwin-aarch64-build-bun' before running any tests, so it is unrelated to the diff.

@robobun

robobun commented Jul 2, 2026

Copy link
Copy Markdown
Collaborator Author

CI status for 435e4be

282 jobs passed, including the debian 13 x64-asan lane, which is the one that actually exercises this PR's regression test (it is skipIf(!isASAN)). That shard reports:

--- [3/227] test/js/bun/shell/shell-pipe-read-fault.test.ts
bun test v1.4.0-canary.1 (435e4be3d)
.......
7 pass
0 fail
Ran 7 tests across 1 file. [5.84s]

Two jobs are red, both on darwin, and neither is reachable from this diff:

1. darwin 26 aarch64 - test-bun — never ran a test:

Error: buildkite-agent artifact download timed out after 120s for step 'darwin-aarch64-build-bun'.
Refusing to continue with a partial download (would silently fall back to the wrong binary).

2. darwin 14 x64 - test-bun — exactly one failing test, on all 4 attempts:

✗ Bun.spawn with terminal option > creates subprocess with terminal attached [90000.00ms]
  ^ this test timed out after 90000ms.
88 pass, 1 todo, 1 fail

That test drives Bun.spawn({ terminal }), i.e. the Terminal reader, which has its own on_read_chunk (src/runtime/api/bun/Terminal.rs:1766). This diff only changes shell::subproc::PipeReader::on_read_chunk (the ShellPipeReader variant, constructed solely by the shell interpreter). After this change, register_poll() has exactly one caller outside src/io/: shell/IOWriter.rs:807, the writer, untouched. Neither Terminal nor SubprocessPipeReader ever call it. There is no path from this change to a PTY read.

I am not pushing a retrigger commit: that terminal timeout is deterministic (identical 90 s timeout on all four attempts), so a re-run would reproduce it rather than clear it. It also looks latent rather than new, which is probably why it is not showing up elsewhere: on the four most recent finished PR builds I checked (68037, 68040, 68041, 68042), the darwin build-cpp step was canceled, so their darwin test lanes never executed at all. This build is one of the few where they did.

Happy to rebase or dig into the terminal lane separately if a maintainer would like.

@robobun

robobun commented Jul 2, 2026

Copy link
Copy Markdown
Collaborator Author

Correction to my previous comment, and a retrigger.

Build 68021 has now finished, and the darwin lanes it did not get to earlier changed the picture. darwin 14 aarch64 - test-bun ran the same binary against the same file and passed:

--- [299/2493] test/js/bun/terminal/terminal.test.ts
bun test v1.4.0-canary.1 (435e4be3d)
89 pass
1 todo
0 fail
Ran 90 tests across 1 file. [874.00ms]

So creates subprocess with terminal attached passes in 874 ms on darwin 14 aarch64 and hangs for the full 90 s timeout on darwin 14 x64, from the same build of the same commit. This change is arch-independent, which rules it out as the cause and points at the x64 agent instead.

It also means I was wrong to call the failure deterministic: all four retries ran inside the same job on the same agent, so they tell us nothing about a fresh one. I have pushed a single ci: retrigger (a769948) to re-roll on new agents. The code and test are untouched.

Final tally for 435e4be was 284 passed, 2 failed:

  • darwin 26 aarch64 - test-bun: buildkite-agent artifact download timed out after 120s, ran zero tests.
  • darwin 14 x64 - test-bun: the terminal timeout above.

Both darwin 14 aarch64 shards and the sibling shard of each failing lane passed, and the debian 13 x64-asan lane (the only one that can run this PR's skipIf(!isASAN) regression test) reported 7 pass, 0 fail.

@robobun

robobun commented Jul 2, 2026

Copy link
Copy Markdown
Collaborator Author

Status: diff is green, CI is red on unrelated darwin infra. Needs a maintainer.

I have used my one retrigger (a769948) and will not push another. Summary of where this stands, so nobody has to re-derive it.

The change is verified

  • debian 13 x64-asan is the only lane that can run this PR's regression test (it is skipIf(!isASAN)). On build 68021 it reported 7 pass, 0 fail for test/js/bun/shell/shell-pipe-read-fault.test.ts.
  • Locally: the new test fails in ~750 ms with the heap-use-after-free in the PR description against unpatched src/, and passes with the fix. The rest of test/js/bun/shell/ was swept; every other failure there reproduces identically with src/runtime/shell/subproc.rs reverted to main.
  • Build 68021 finished 284 passed / 2 failed.

The two red jobs, neither reachable from this diff

darwin 26 aarch64 - test-bun — recurring infra, 2 builds out of 2. Identical on 68021 and 68070, and it runs zero tests before dying:

Searching for artifacts: "**" within step: "darwin-aarch64-build-bun"
Found 2 artifacts. Starting to download to: .../bun/bun/release
Error: buildkite-agent artifact download timed out after 120s for step 'darwin-aarch64-build-bun'.
Refusing to continue with a partial download (would silently fall back to the wrong binary).
    at getExecPathFromBuildKite (scripts/runner.node.mjs:2182:13)

This is scripts/runner.node.mjs's own 120 s guard tripping on a slow artifact fetch. It is unrelated to any diff and will red any PR whose darwin aarch64 test lane happens to run.

darwin 14 x64 - test-bun (on 68021) — agent-specific hang. One test, test/js/bun/terminal/terminal.test.ts > Bun.spawn with terminal option > creates subprocess with terminal attached, hit its 90 s timeout. The same binary ran the same file on darwin 14 aarch64 and passed in 874 ms:

--- [299/2493] test/js/bun/terminal/terminal.test.ts
bun test v1.4.0-canary.1 (435e4be3d)
89 pass, 1 todo, 0 fail
Ran 90 tests across 1 file. [874.00ms]

This change is arch-independent, so a test that passes on aarch64 and hangs on x64 from the same build is not caused by it. Mechanically it also cannot be: that test drives Bun.spawn({ terminal }), i.e. the Terminal reader, which has its own on_read_chunk (src/runtime/api/bun/Terminal.rs:1766). This PR only touches shell::subproc::PipeReader::on_read_chunk (the ShellPipeReader variant, constructed solely by the shell interpreter). After this change register_poll() has exactly one caller outside src/io/: shell/IOWriter.rs:807, the writer, untouched. Neither Terminal nor SubprocessPipeReader ever call it.

Ask

The code is unchanged since 435e4be and mergeable: MERGEABLE. Happy to rebase whenever, and happy to look at the darwin aarch64 artifact-download timeout separately if that is useful, but it is not something this PR can fix.

@robobun

robobun commented Jul 2, 2026

Copy link
Copy Markdown
Collaborator Author

Retrigger result: everything that ran tests is green. The only red is one broken CI agent.

Build 68070 (on a769948) finished 284 passed, 2 failed, and the two failures are both the same Buildkite agent failing to download the build artifacts before it ever starts a test.

The failure is an agent, not a lane and not this diff

Comparing the darwin 26 aarch64 - test-bun jobs across both builds by agent:

Build Agent Result
68021 darwin-aarch64-26.5.1-1 failed, artifact download timeout
68021 darwin-aarch64-26.4-1 passed
68070 darwin-aarch64-26.5.1-1 failed, artifact download timeout
68070 darwin-aarch64-26.5.1-1 failed, artifact download timeout

Every job that lands on darwin-aarch64-26.5.1-1 dies the same way, with zero tests executed:

Searching for artifacts: "**" within step: "darwin-aarch64-build-bun"
Found 2 artifacts. Starting to download to: .../darwin-aarch64-26.5.1-1/bun/bun/release
Error: buildkite-agent artifact download timed out after 120s for step 'darwin-aarch64-build-bun'.
Refusing to continue with a partial download (would silently fall back to the wrong binary).
    at getExecPathFromBuildKite (scripts/runner.node.mjs:2182:13)

Every job on any other agent passes. On 68070 both shards of that lane happened to land on the bad agent, which is why the lane is fully red this time. This looks worth a separate look by someone with infra access; it is not something a PR can fix.

The previous darwin 14 x64 failure did not recur

test/js/bun/terminal/terminal.test.ts > creates subprocess with terminal attached, which hit its 90 s timeout on 68021, passed on both darwin 14 x64 shards this build. That settles it as an agent-level flake, consistent with the same binary passing the same file on darwin 14 aarch64 in 874 ms on the earlier build.

The change is verified

All 20 debian 13 x64-asan shards passed. That is the only lane that can execute this PR's regression test (it is skipIf(!isASAN)), and it reported 7 pass, 0 fail for test/js/bun/shell/shell-pipe-read-fault.test.ts.

Ask

The code is unchanged since 435e4be, mergeable: MERGEABLE, and I have used my one retrigger so I will not push again. Every lane that actually ran tests is green on both builds. This is ready for a maintainer.

@Jarred-Sumner
Jarred-Sumner merged commit b28cf5b into main Jul 4, 2026
76 of 78 checks passed
@Jarred-Sumner
Jarred-Sumner deleted the farm/852f729b/shell-pipereader-onreadchunk-uaf branch July 4, 2026 02:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants