Skip to content

spawn: hold the Subprocess as a RefPtr across spawn_maybe_sync - #37700

Closed
robobun wants to merge 1 commit into
mainfrom
farm/c83f5856/spawn-shared-subprocess-in-spawn
Closed

robobun wants to merge 1 commit into
mainfrom
farm/c83f5856/spawn-shared-subprocess-in-spawn

Conversation

@robobun

@robobun robobun commented Aug 12, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

  • Nothing user-visible breaks. This is a lifetime cleanup of spawn_maybe_sync (src/runtime/api/bun/js_bun_spawn_bindings.rs), which builds a Subprocess for Bun.spawn and Bun.spawnSync.
  • It held the new Subprocess as &mut for its whole length, but stdin setup, an already aborted signal, and the spawnSync loop all re-enter the object before it returns. Each collision had its own raw-pointer workaround. The ref count started at 2 for two holders named in a comment, and every exit balanced it by hand.

Fix

  • The allocation is a RefPtr<Subprocess>. The function's ref is a local, the exit handler and the JS wrapper each clone() one where they are installed, and a release is a drop. finalize turns the wrapper's Box back into its RefPtr. The spawnSync tail releases its own through a scope guard.
  • The body binds &Subprocess (every field is Cell-style), so the workarounds become plain field access.
  • Correct because the count after every exit is the same as before. Two exits leaked the Subprocess before and still do, now with an explicit into_raw(): a release there would free the object without its teardown while pipes still point at it.
  • Verified: test/js/bun/spawn/ (35 files) and test/js/node/child_process/ under the debug build. Windows cargo check and clippy are clean.

Background

  • Subprocess is intrusively ref-counted: the last holder to release frees it. Holders are the JS wrapper, the exit handler, and the stdin pipe while open.
  • RefPtr<T> (bun_ptr, since bun_ptr: RefPtr releases on Drop; remove ScopedRef/IntrusiveRc/DestructorCtx #40478) is one counted ref held as a value: clone takes a ref, drop releases it, into_raw/from_raw move it across an FFI boundary.
  • The generated finalizer hands the wrapper's m_ctx back as a Box<Subprocess>. Other holders may still be alive then, so finalize must not drop it as a Box.
  • spawnSync never creates a JS wrapper. It runs an isolated event loop until the child exits, then tears the Subprocess down itself.
Notes

This PR was first stacked on #37665, which added an OwnedRef<T> type for this purpose. #40478 gave RefPtr<T> that role (release on drop, Clone, into_raw/from_raw), so the PR was redone on top of main with RefPtr in place of OwnedRef.

The leak fix the earlier revision carried (the spawnSync tail returned without finalize when building the result threw) landed on main in #39564. The scope guard here covers the same exits.

to_js_from_ptr(*mut) becomes to_js_from_ref(RefPtr), which moves the wrapper's ref into m_ctx. Writable::init takes &Subprocess and passes as_ctx_ptr() to the four StaticPipeWriter::create sites. The stdin source handle is a BackRef::new, the abort signal is stored through the Cell, and the failed-watch() exit notification holds its own RefPtr clone instead of a raw pointer behind a flag. The subprocess_ipc_owner helper (a null check on a pointer that cannot be null) is gone.

Ref count per exit, before and after. Writable::init error: two derefs vs one drop, both reach 0. Handler installed, then the stdin ReadableStream setup throws: leaked before, leaked now (into_raw()). Buffer-stdin start() fails in spawnSync: same. Windows IPC setup error: one deref vs one drop. Normal Bun.spawn return: wrapper + handler both times. Bun.spawnSync: the function's ref is consumed by finalize_owned where finalize(Box::from_raw(..)) ran before. The earlier OwnedRef revision dropped the function's ref on the two leaking exits as well. That is the one place this revision differs from it. On Windows the uv exit callback fires without watch(), so the handler's release could then free the object with finalize_streams never run, while the readers and the extra stdio uv_pipe_t handles still point at it.

Test results in the container used here: two tests in spawnsync-isolated-event-loop.test.ts (the GC keep-alive ones from #40508) hit the 5 s timeout under the debug ASAN build. Their fixtures print OK in about 7 s with and without this change. spawn_waiter_thread.test.ts fails its CPU-time bound the same way on main's build. child_process.test.ts "default shell" needs $SHELL. Everything else passes: spawn.test.ts 140 pass, the other 34 spawn files 207 pass, child_process 130 pass.

Original description

Stacked on #37665 (the base of this PR is that branch). Two commits: the original one, which binds the Subprocess as &Subprocess instead of &mut across spawn_maybe_sync, and, following review ("Can we use a better RAII container type instead?"), a second one that holds it as an OwnedRef.

spawn_maybe_sync materialised the freshly boxed Subprocess as &mut and kept that borrow live to the end of the function, although the code in between hands out back-pointers to the same allocation and re-enters it. Each collision had its own workaround. The binding became the shared form every field of Subprocess is built for, Writable::init takes &Subprocess, and the workarounds became plain field access. The allocation became an OwnedRef: the function's ref is a local, the exit handler and the JS wrapper each take a clone where they are installed, and a release is a drop.

@robobun
robobun requested a review from alii August 12, 2026 01:12
@coderabbitai

coderabbitai Bot commented Aug 12, 2026

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

@robobun, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 2 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 9251236d-7228-47e6-abdc-73ccb208aeaa

📥 Commits

Reviewing files that changed from the base of the PR and between e7abdf7 and 6f866ac.

📒 Files selected for processing (2)
  • src/runtime/api/bun/js_bun_spawn_bindings.rs
  • src/runtime/api/bun/subprocess/Writable.rs

Comment @coderabbitai help to get the list of available commands.

@robobun

robobun commented Aug 12, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 5:53 AM PT - Aug 12th, 2026

✅ @robobun, your commit 835077dd3c0304aa2149c6854d122869fe0b086a passed in Build #92935! 🎉


🧪   To try this PR locally:

bunx bun-pr 37700

That installs a local version of the PR into your bun-37700 executable, so you can run:

bun-37700 --bun

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this PR and didn't find any bugs. Because it reworks unsafe pointer/aliasing invariants in spawn_maybe_sync (provenance of subprocess_nn, the exit-notification guard's liveness reasoning), a human look would still be worthwhile.

Checked: NonNull::from(&Subprocess) provenance is sufficient for all consumers (R-2 interior mutability, BackRef<_, Shared> only yields &T); as_ctx_ptr() contract matches what StaticPipeWriter::create needs; the defer! guard only dereferences subprocess_nn when send_exit_notification is set, which happens only on the !IS_SYNC path where the JS wrapper owns the box past function return; the let-chain reordering in the stdin Pipe block is semantically equivalent; abort_signal.set ordering vs add_listener is unchanged.

Extended reasoning...

Overview

This PR changes spawn_maybe_sync to bind the freshly-boxed Subprocess as &SubprocessT<'static> instead of &mut, collapsing several per-site unsafe workarounds (raw (*subprocess_ptr) writes, a subprocess_ptr_exit copy, lifetime-erasing casts) into plain shared-field access. Writable::init drops its &mut Subprocess parameter to &Subprocess and routes the four StaticPipeWriter::create sites through the existing Subprocess::as_ctx_ptr() helper. Writable::pipe_sink is widened to pub(in crate::api) to match its siblings. Net: −1 unsafe block, −4 &mut→*mut casts, −2 lifetime-erasing casts.

Security risks

None. No user-facing behavior, validation, or resource-handling changes; this is a type-level refactor of how an already-boxed allocation is borrowed.

Level of scrutiny

High. spawn_maybe_sync is memory-safety-critical (the most-blocked category per REVIEW.md): it wires back-pointers that re-enter the same allocation, arms a defer! guard that runs after finalize on the sync path, and hands the box to a JS wrapper on the async path. The change is a strict tightening (&mut → &), but the soundness of NonNull::from(&T) provenance for downstream consumers, and the exit-guard's flag-gated dereference, deserve a maintainer's confirmation.

Other factors

I verified: as_ctx_ptr()'s documented contract ("callbacks deref it as &*const shared") matches the new call sites; BackRef::new and BackRef<_, Shared> only ever yield &T, so shared-read provenance from NonNull::from(subprocess) suffices given all Subprocess mutation goes through Cell/JsCell; send_exit_notification is only set on the !IS_SYNC branch (after to_js_from_ptr), so the guard's subprocess_nn.as_ref() never runs on the spawnSync path where finalize has already freed the box; the stdin Writable::Pipe let-chain checks the same two conditions as the old matches! && cond { if let } shape; the two abort_signal.set calls still happen after add_listener. 208 spawn tests pass per the description. No prior review comments to address.

@Jarred-Sumner Jarred-Sumner left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we use a better RAII container type instead?

@robobun robobun changed the title spawn: hold &Subprocess, not &mut, across spawn_maybe_sync spawn: hold the Subprocess as an OwnedRef across spawn_maybe_sync Aug 12, 2026
@robobun
robobun changed the base branch from main to farm/c83f5856/ptr-owned-ref August 12, 2026 06:08
@robobun
robobun force-pushed the farm/c83f5856/spawn-shared-subprocess-in-spawn branch from 6f866ac to 8e980dc Compare August 12, 2026 06:09
Comment thread src/ptr/lib.rs Outdated
Comment thread src/ptr/ref_count.rs Outdated
Comment thread src/ptr/ref_count.rs Outdated
Comment thread src/ptr/ref_count.rs Outdated
Comment thread src/ptr/ref_count.rs Outdated
Comment thread src/ptr/ref_count.rs Outdated
Comment thread src/ptr/ref_count.rs Outdated
Comment thread src/ptr/ref_count.rs Outdated
Comment thread src/ptr/ref_count.rs Outdated
Comment thread src/ptr/ref_count.rs Outdated
Comment thread src/ptr/ref_count.rs Outdated
Comment thread src/ptr/ref_count.rs Outdated
Comment thread src/ptr/ref_count.rs Outdated
Comment thread src/ptr/ref_count.rs Outdated
Comment thread src/ptr/ref_count.rs Outdated
Comment thread src/ptr/ref_count.rs Outdated
Comment thread src/ptr/ref_count.rs Outdated
Comment thread src/ptr/ref_count.rs Outdated
Comment thread src/ptr/ref_count.rs Outdated
Comment thread src/ptr/ref_count.rs Outdated
Comment thread src/bun_core/lib.rs
Comment thread src/bun_core/lib.rs
Comment thread src/bun_core/lib.rs
Comment thread src/bun_core/lib.rs
Comment thread src/bun_core/lib.rs
Comment thread src/bun_core/lib.rs
Comment thread src/bun_core/lib.rs
Comment thread src/bun_core/lib.rs
Comment thread src/bun_core/lib.rs
Comment thread src/bun_core/lib.rs
Comment thread src/bun_core/lib.rs
Comment thread src/bun_core/lib.rs
Comment thread src/bun_core/lib.rs
Comment thread src/bun_core/lib.rs
Comment thread src/bun_core/lib.rs
Comment thread src/bun_core/lib.rs
Comment thread src/bun_core/output.rs
Comment thread src/bun_core/output.rs
Comment thread src/bun_core/string/immutable.rs
Comment thread src/bun_core/string/wtf.rs
Comment thread src/bun_core/string/wtf.rs
Comment thread src/bun_core/string/wtf.rs
Comment thread src/bun_core/util.rs
@robobun robobun changed the title spawn: hold the Subprocess as an OwnedRef across spawn_maybe_sync spawn: hold the Subprocess as a RefPtr across spawn_maybe_sync Aug 26, 2026
@robobun
robobun changed the base branch from farm/c83f5856/ptr-owned-ref to main August 26, 2026 08:44
@robobun

robobun commented Aug 26, 2026

Copy link
Copy Markdown
Collaborator Author

Retargeted onto main after #40478. That PR gave RefPtr<T> the role OwnedRef<T> (#37665) was written for, so this PR was redone on top of main with RefPtr in place of OwnedRef, as one commit (ccb40c7).

What changed against the earlier revision:

  • The leak fix it carried (the spawnSync tail returned without finalize when building the result threw) is already on main via Throw instead of aborting when child output does not fit in a Buffer #39564. The scope guard here covers the same exits, so this is now a pure lifetime cleanup.
  • The two exits that leaked the Subprocess before (a throwing stdin ReadableStream setup, a failing buffer-stdin start() in spawnSync) keep leaking it, now with an explicit into_raw(). The earlier revision dropped the function's ref there. On Windows the uv exit callback fires without watch(), so that release could free the object with finalize_streams never run while the readers and the extra stdio pipe handles still point at it. The ref count after every exit is now the same as on main.

Rebuilt and ran test/js/bun/spawn/ and test/js/node/child_process/ under the debug build. The PR body has the details.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Additional findings (outside the current diff — GitHub can't attach inline comments there):

  • 🟡 packages/bun-usockets/src/bsd.c:171 — nit: bsd_recvmmsg gained a max_packets parameter (used by loop.c to cap shared cluster UDP sockets to 1 packet per dispatch via u->shared_fd ? 1 : LIBUS_UDP_RECV_COUNT), but the _WIN32 branch was not updated: the for loop at line 150 and the return at line 171 still hard-code LIBUS_UDP_RECV_COUNT, so the new argument (and the clamp at line 148) is dead on Windows.

    Extended reasoning...

    On Windows, a node:cluster worker adopts a shared UDP socket (via the new bsd_socket_import/us_create_udp_socket_from_fd(..., shared=1, ...) path added in this diff — src/js/node/dgram.ts sets $sharedFd, udp_socket.rs:670 forwards it, udp.c:209 records udp->shared_fd = 1). When the poll fires readable, loop.c:991 calls bsd_recvmmsg(fd, &recvbuf, MSG_DONTWAIT, 1); on Linux/macOS this reads at most one datagram so other workers sharing the duplicated kernel socket get a fair share, but on Windows the loop still drains up to LIBUS_UDP_RECV_COUNT packets per event. Relative to the base branch (which had no max_packets and no shared-UDP path) this is not a regression, but it defeats the intent of the change on one platform: a single Windows worker can monopolise the shared receive queue while its siblings starve. Fix: replace both LIBUS_UDP_RECV_COUNT occurrences in the _WIN32 arm with max_packets.

    Verification: nit — The _WIN32 branch of bsd_recvmmsg was not updated to use the new max_packets argument. At packages/bun-usockets/src/bsd.c:150 the loop is for (int i = 0; i < LIBUS_UDP_RECV_COUNT; i++) and at line 171 it returns LIBUS_UDP_RECV_COUNT, while the Apple branch (lines 175, 180, 192) and Linux branch (line 195) were changed to use max_packets. The clamp at line 148 (`if (max_packets

@robobun

robobun commented Aug 26, 2026

Copy link
Copy Markdown
Collaborator Author

Closing this as superseded by #40204 and #40516.

#40204 converts the same function the same way, in a larger scope: spawn_maybe_sync builds the Subprocess with RefPtr::new and RefCount::init(), the exit handler's ref is a value (exit_ref: Cell<Option<RefPtr<Self>>>, installed with set_exit_handler(subprocess_this) and released at the tail of on_process_exit), and the wrapper's ref is subprocess_js_ref, moved into to_js_from_ptr. That covers every structural hunk here. #40516 covers the finalize half generator-wide: refCounted: true on the class, finalize(&self), and the generated thunk releases the wrapper's ref, which makes the finalize(Box<Self>) to RefPtr adapter here unnecessary.

Merging this first would only give both of them conflicts in js_bun_spawn_bindings.rs, subprocess.rs and Writable.rs.

One finding from this revision carries over to #40204 once it is rebased onto #40478 (where RefPtr releases on drop), posted there: two exits of spawn_maybe_sync leak the Subprocess on main, and a function-held RefPtr that drops on those exits releases a ref main keeps. On Windows the uv exit callback fires without watch(), so the handler's release can then free the object with finalize_streams never run, while the extra stdio uv_pipe_t handles and the readers still point at it.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants