Skip to content

spawn(windows): stop closing an exposed extra stdio pipe handle twice - #39966

Merged
Jarred-Sumner merged 1 commit into
mainfrom
farm/1d92b523/win-stdio-pipe-double-close
Aug 21, 2026
Merged

Jarred-Sumner merged 1 commit into
mainfrom
farm/1d92b523/win-stdio-pipe-double-close

Conversation

@robobun

@robobun robobun commented Aug 21, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

  • On Windows, Subprocess.stdio[i] for i >= 3 exposes the HANDLE of the parent-side uv_pipe_t (src/runtime/api/bun/subprocess.rs, get_stdio). node:child_process passes that value to net.connect({ fd }) (src/js/node/child_process.ts:1281), which adopts it into a second uv_pipe_t (src/runtime/socket/Listener.rs:1195) and closes it with the socket. The Subprocess kept the first uv_pipe_t and closed the same HANDLE value again in finalize_streams. Windows reuses a closed handle value at once, so the second close destroyed whatever owned the value by then. child_process.spawn(cmd, { stdio: ["pipe", "pipe", "pipe", "pipe"] }) is enough to hit it (Playwright and ffmpeg launchers use such pipes).
  • Found while investigating the crash failed to join on thread: ... (os error 6) in WebWorker::join (Sentry BUN-4NEC, Windows, 1.4.0): the std thread handle of a Worker had been closed by someone else in the process. This stale close is one Bun-side way for that to happen. The tags of the report (spawn, workers_spawned) fit, but one report cannot prove it was this one.
  • Bun.spawn: don't double-close extra stdio fds exposed via .stdio #33828 fixed the same ownership problem on POSIX. The Windows arm was left out, and its regression tests skip on Windows.

Fix

  • get_stdio on Windows duplicates each Buffer slot's handle with bun_sys::dup, closes its own pipe right away, and stores the duplicate as the new WindowsStdioResult::UnownedFd. Reads return the duplicate's value, and finalize_streams and the VM teardown path leave it alone. This is the Windows form of the POSIX downgrade to ExtraPipe::UnownedFd.
  • The pipe is closed at once, not at finalize: while Bun's handle stayed open, the child would not see EOF after the caller closed its copy. The duplicate refers to the same pipe, so net.connect reads it as before.
  • If dup fails, the slot stays Unavailable and reads as null. The three exhaustive matches on WindowsStdioResult gain the new arm.
  • Verified on a Windows x64 debug build: the two new tests in test/js/bun/spawn/spawn.test.ts and test/js/node/child_process/child_process.test.ts fail 6 of 6 runs on the current canary and pass on this build. spawn.test.ts, child_process.test.ts and the spawn IPC tests pass on that build.

Background

  • A uv_pipe_t owns its OS handle: uv_close on it calls CloseHandle. Bun.connect({ fd }) on Windows turns a HANDLE value into a CRT fd with _open_osfhandle and gives it to a new uv_pipe_t, so after that call two uv_pipe_t owned one handle.
  • On Windows files, pipes, events, processes and threads share one handle table per process, and a freed value is reissued at once. A stale CloseHandle therefore closes an unrelated object. Fd::close asserts on this only in debug builds, and this path did not go through it at all.
  • stdio_pipes holds the slots for indexes 3 and up. finalize_streams (GC) and stop_for_vm_teardown (a Worker exiting) close every slot that still holds a pipe.
Notes
  • The spawn.test.ts test pins the mechanism down exactly. After the socket closed the exposed value, it allocates events until the kernel reissues that value, drops the Subprocess, runs the GC and checks the event is still signaled. On the unfixed build every checked value is closed again (3 of 3 in each of 6 runs). An iteration is skipped when something else reissues the value first. On the debug build that is usually the second of the three iterations, the other two were conclusive in 8 of 8 runs. On the release canary all three were conclusive in 6 of 6 runs.
  • The child_process.test.ts test covers the wiring on top: the three sockets deliver the child's data, and the files opened after the sockets closed survive the GC of the ChildProcess. On the unfixed build 3 or 6 of them die with EBADF in every run.
  • Debug timings on the Windows machine used here: about 1.9 s and 2.1 s per test, most of it three debug-build child starts each.
  • Not changed: Listener.rs interprets a { fd } number as a CRT fd first and only then as a HANDLE, so a HANDLE value that happens to match an open CRT fd number is still adopted as the wrong object. Separate issue.
  • The related stale close in Dir::cwd() (src/sys/dir.rs, Windows only, needs a concurrent process.chdir()) is handled separately.

no test proof · iteration 0 · Platform-specific test(s) that do not run on this machine. Deferring to CI, which covers all platforms: test/js/bun/spawn/spawn.test.ts test/js/node/child_process/child_process.test.ts

On Windows, Subprocess.stdio[i >= 3] exposed the HANDLE of the parent-side
uv_pipe_t. node:child_process hands that value to net.connect({ fd }), which
adopts it into a second uv_pipe_t and closes it with the socket. The
Subprocess still closed its own uv_pipe_t, and with it the same HANDLE
value, when it was finalized. By then the value usually belonged to another
object, which the second close destroyed.

The getter now hands out a duplicate of the handle, closes its own pipe at
once, and keeps the duplicate in the slot so repeated reads return the same
value. The duplicate is the caller's to close, like ExtraPipe::UnownedFd on
POSIX.
@coderabbitai

coderabbitai Bot commented Aug 21, 2026

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

Your included review limit has been reached.

You’re in a promotional period — use the checkbox below to run this review for free:

  • Run review for free

On-demand reviews are free for the next 30 days. After that, they cost $0.25 per reviewed file.

How can I continue?

Run this review now using the option above, or comment @coderabbitai review --use-credits.

You can also wait for the limit to reset (next review available in 9 minutes), then comment @coderabbitai review or push new commits to the PR.

An organization admin can change what happens after included review limits in Billing.

How do review limits work?

CodeRabbit enforces per-developer PR review limits within each organization.

For paid Pro and Pro+ reviews, CodeRabbit uses a developer's included PR review attempts over the past 7 days to set the current hourly allowance. At typical activity levels, the full plan allowance applies. Higher sustained activity can lower the allowance until earlier attempts leave the 7-day window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 15f20e41-179d-400d-b20c-743220a16cfb

📥 Commits

Reviewing files that changed from the base of the PR and between 4448a2e and be9e7ba.

📒 Files selected for processing (7)
  • src/runtime/api/bun/js_bun_spawn_bindings.rs
  • src/runtime/api/bun/subprocess.rs
  • src/runtime/shell/subproc.rs
  • src/spawn/process.rs
  • src/spawn/static_pipe_writer.rs
  • test/js/bun/spawn/spawn.test.ts
  • test/js/node/child_process/child_process.test.ts

Comment @coderabbitai help to get the list of available commands.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this PR and didn't find any bugs. Because it changes Windows HANDLE ownership in the spawn stdio path (dup + early uv_close, with EOF-visibility implications for the child), a human look would still be worthwhile.

What was reviewed:

  • The new UnownedFd arm is covered in every exhaustive match on WindowsStdioResult (spawn bindings sync-cleanup, shell PipeReader::create, StaticPipeWriter::create, finalize_streams, stop_for_vm_teardown).
  • close_and_destroy calls open_handles::remove first, so the entry record_stdio_pipe_ownership registered is cleaned up and stop_for_vm_teardown cannot reach the freed pipe.
  • core::mem::take on the slot is sound (WindowsStdioResult: Default → Unavailable); repeat .stdio reads see the stored UnownedFd and return the same value.
  • bun_sys::dup on Windows uses DuplicateHandle(bInheritHandle=FALSE), so the exposed duplicate is not inherited by later spawns.
Extended reasoning...

Overview

Adds WindowsStdioResult::UnownedFd(Fd) and, in Subprocess::get_stdio on Windows, converts each extra-stdio Buffer(Box<uv::Pipe>) slot into an UnownedFd holding a DuplicateHandle of the pipe's OS handle, then closes the original uv_pipe_t immediately via close_and_destroy. The duplicate's numeric value is what JS receives; finalize_streams, stop_for_vm_teardown, and the sync-spawn error-cleanup path all skip the new variant. Three other match sites gain the arm mechanically. Two Windows-only regression tests are added.

Security risks

None identified. This is a resource-ownership fix: it removes a stale CloseHandle on a value the kernel may have reissued to an unrelated object. bun_sys::dup creates the duplicate non-inheritable, so it does not leak into children spawned while it is open.

Level of scrutiny

High. This is Windows-specific handle-lifetime code in the spawn subsystem. The reasoning is subtle: closing Bun's uv_pipe_t immediately (rather than at finalize) is required so the child sees EOF once the caller closes the duplicate; the duplicate keeps the kernel object alive for net.connect({fd}). The exhaustive-match updates and the open_handles interaction check out, but the EOF/ordering argument and the intentional "caller owns the duplicate, we never close it" leak-on-abandon semantics (matching the POSIX UnownedFd from #33828) deserve a human confirmation.

Other factors

  • Maybe<Fd> is core::result::Result, so the if let Ok(dup) pattern is correct; on dup failure the slot stays Unavailable and reads as null.
  • close_and_destroy removes the pipe from the thread's open-handles list before uv_close, so the ownership record set by record_stdio_pipe_ownership does not go stale.
  • The two new tests are precise: the spawn.test.ts one reoccupies the freed HANDLE value with a signaled event and asserts GC does not close it; the child_process.test.ts one checks the net.Socket wiring end-to-end and that files opened afterward survive GC. Both were verified failing on canary per the description.
  • No prior human reviews or outstanding comments on the PR.

@Jarred-Sumner
Jarred-Sumner merged commit 6e10a65 into main Aug 21, 2026
11 of 12 checks passed
@Jarred-Sumner
Jarred-Sumner deleted the farm/1d92b523/win-stdio-pipe-double-close branch August 21, 2026 21:23
Jarred-Sumner pushed a commit that referenced this pull request Aug 22, 2026
…dle (#39969)

### Problem
- On Windows `Fd::cwd()` returned the handle ntdll keeps in the PEB
(`src/bun_core/util.rs:1047`). `process.chdir()` on any thread closes
that handle and installs a new one, and Windows gives the closed value
to the next object created. So an `Fd::cwd()` taken before a chdir is a
stale snapshot afterwards: it no longer compares equal to `Fd::cwd()`,
and closing it closes an unrelated object.
- `Dir::drop` (`src/sys/dir.rs:17`) relies on that comparison to skip
the cwd. `fs.rm` / `fs.rmdir` with `recursive` hold a `Dir::cwd()` for
the whole walk on a thread pool thread
(`src/runtime/node/node_fs.rs:7664`, `:7702`), and the transpiler cache
holds one across every cache write on worker and pool threads
(`src/jsc/RuntimeTranspilerCache.rs:871`). On the current canary,
`fs.promises.rm` of 1500 files with a `process.chdir()` during the walk
closes one of 1024 events created after the chdir, in 6 of 6 runs.
- Found while investigating `failed to join on thread: ... (os error 6)`
in `WebWorker::join` (Sentry BUN-4NEC, Windows, 1.4.0): a Worker's
thread handle had been closed by other code in the process. A thread
created after a chdir is one of the objects this close can hit. The
report's tags (`transpiler_cache`, `workers_spawned`) fit, but one
report cannot prove it was this path. #39966 fixed a second stale close
found in the same investigation.

### Fix
- `Fd::cwd()` on Windows is now a fixed value, the counterpart of
`AT_FDCWD`. The value is bit 62 alone: handle values fit in 32 bits, bit
63 is the uv tag, and `INVALID_HANDLE_VALUE` masks to all of bits 0..63,
so nothing else can produce it. `decode_windows()` maps it to the PEB
handle at the time of each call, so every syscall still resolves against
the live cwd, and `Display` prints it as `[cwd]`.
- Because the value is constant, `Dir::drop`'s comparison is exact on
every platform, and `Dir::cwd()` and its callers need no change.
`Fd::close()` refuses the sentinel with EBADF, as `close(AT_FDCWD)` does
on POSIX, instead of closing ntdll's handle.
- Verified on a Windows x64 debug build: the new test in
`test/js/node/fs/fs.test.ts` fails on the current canary (`rm closed 1
handle(s) it did not own`) and passes 4 of 4 runs here. On the same
build: `fs.test.ts` (one disk-bound 4.9 GB write times out here,
unrelated), `bun-write.test.js`, `transpiler-cache.test.ts`,
`bun-link.test.ts`, `bun-pm.test.ts`, `bundler/cli.test.ts` pass, and
relative fs operations after a `process.chdir()` resolve against the new
cwd.

### Background
- `bun_sys::Dir` (#30878) closes its fd on `Drop` unless the fd is
`Fd::INVALID` or `Fd::cwd()`. On POSIX `Fd::cwd()` is the `AT_FDCWD`
constant, so that check was exact there and only broke on Windows.
- `Fd` on Windows packs a kind bit and a value. `decode_windows()` is
the one place that turns a system-kind value into a HANDLE for
`native()`, `close()` and friends, which is why the mapping lives there.
The Windows `*at` wrappers already read the PEB handle per call for an
invalid dirfd; this makes the cwd fd itself behave that way.
- On Windows all kernel objects share one handle table per process and a
freed value is reused at once, so a stale close destroys an unrelated
object. `Fd::close` notices that only in debug builds.

<details><summary>Notes</summary>

- The first version of this PR made `Dir::cwd()` return
`ManuallyDrop<Dir>` and rewrote two callers. Review asked for the cwd
abstraction itself to be fixed so callers stay platform-uniform, which
is what the sentinel does.
- The test creates 1500 files, starts `fs.promises.rm` (thread pool),
polls until the walk has started deleting, calls `process.chdir()`,
creates 1024 signaled events (one of them receives the retired handle
value), awaits the rm and checks that every event is still signaled. On
the debug build the walk is detected with about 1480 files to go and
takes about 140 ms, so the chdir lands inside it. About 0.85 s per run
on the debug build.
- The same comparison exists once more in
`RuntimeTranspilerCache.rs:884-888` on a raw `Fd`; it is exact now as
well.
- `Fd::cwd()` is passed as a dirfd in about 230 places. All of them
reach the handle through `native()`, so they see the live cwd as before.
The only behavior changes are the stable comparison, the refused close
and the `Display` output.
</details>

<!-- robobun:evidence:begin -->

---

**no test proof** · iteration 2 · platform-specific test(s) that do not
run on this machine, deferring to CI, which covers all platforms:
test/js/node/fs/fs.test.ts

<!-- robobun:evidence:end -->
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants