Skip to content

spawnSync: keep event-loop ref counts on the loop they were taken on - #37754

Closed
robobun wants to merge 3 commits into
mainfrom
farm/c728b41c/spawnsync-loop-counts
Closed

robobun wants to merge 3 commits into
mainfrom
farm/c728b41c/spawnsync-loop-counts

Conversation

@robobun

@robobun robobun commented Aug 12, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

  • bun test --parallel workers intermittently hang inside Bun.spawnSync of a local bun child: CI prints Interrupted while still running: <file> after 4 minutes, or a run of tests in one worker fails with killed 1 dangling process and this test timed out after 90000ms with empty child stdout. The files pass alone. Also reported on macOS (worker at 100% CPU, child a zombie) and reproduced locally in 3 of about 73 CI-shaped batches.
  • Cause: while spawnSync waits, the VM's loop handle points at spawnSync's private loop, and every keep-alive counter released on whichever loop was current. So a ref taken on the main loop and released during the wait (GC sweeping dead objects, or the runner moving on after a timeout) came off the private loop.
  • Once the private loop's count reads 0 it is never polled again, so spawnSync spins without seeing the exit. The private loop is reused, so the worker stays broken for the rest of its life.
  • A second bug in the same code: the saved "previous loop handle" was one slot shared by nested spawnSync calls, so after a nested call returned the VM stayed on the private loop for good.
  • Fixes Bug: spawnSync never returns: child exit lost (child stays zombie), wait loop busy-spins at 100% CPU re-registering a finished pipe reader (macOS ARM64) #34069

Fix

  • Each of the four counters (file polls, KeepAlive refs, the timer and immediate counts, and each JS event loop's keep-alive counter) records the loop it counted on and releases on that same loop, instead of resolving the loop at release time.
  • The previous loop handle is returned to the caller and lives on its stack frame, so nested spawnSync calls restore in the right order.
  • The property to check: every increment and its matching decrement now hit the same loop. The handle swap itself stays, since spawnSync's own polls rely on it, and each recorded loop outlives what was counted on it (the thread's loop is freed at VM teardown after timers and the heap, the private loop after that).
  • Verification: five new test cases fail on the unfixed build (the inner bun test spins until killed) and pass with the fix; four release one kind of ref each during a wait, and the nesting case alone still fails with only the counter fix applied. Related suites were also run and the remaining failures judged unrelated. The new tests are skipped on Windows, where three of the counters do not go through the swapped handle.

Background

  • Bun.spawnSync waits on a private uws loop with its own epoll/kqueue fd, not the main loop, so main-loop work such as JS timers does not run during the wait. While it waits, vm.event_loop_handle is repointed at that private loop so the child's pidfd and pipe polls register there.
  • A uws loop tracks how many polls and refs are keeping it alive; us_loop_run_bun_tick returns at once without polling when that count is 0. That is why an off-by-one shows up as a busy spin, not a blocked wait.
  • The four counters the diff touches: FilePoll (one per fd registered on a loop: pipes, pidfds, sockets), KeepAlive (the on/off ref used by servers, sockets, fetch, fs, dns and similar), timer::All (refs the loop while any JS timer or immediate is pending), and jsc::EventLoop's concurrent_ref (refs from MessagePort, BroadcastChannel and similar, folded into a uws loop).
  • JS runs inside a spawnSync wait through one path: when a bun:test test times out inside spawnSync after its callback has left the stack, the runner continues with the following tests inside the wait. The new tests use that path to release refs during a wait.

no test proof · iteration 0 · Platform-specific test(s) that do not run on this machine. Deferring to CI, which covers all platforms: test/js/bun/spawn/spawnsync-isolated-event-loop.test.ts

Original description

Problem

bun test --parallel workers intermittently stop inside Bun.spawnSync of a local bun child. In CI this is the flaky annotation "failed in the parallel batch ... passed alone": the batch goes quiet, the runner kills it after 4 minutes and prints Interrupted while still running: <file> (24xs), and the file passes in ~100ms on its own. The victims are whatever spawn-heavy files land in the batch (napi do.test.ts files, resolve.test.ts, child_process.test.ts, ctrl-c.test.ts, text-loader.test.ts, malformed-integrity-base64.test.ts, ...), across ubuntu, debian and alpine lanes (for example builds 92307, 92340, 92426, 92160). A second shape of the same failure is a run of consecutive tests in one worker each failing with killed 1 dangling process + this test timed out after 90000ms, with the spawned child's stdout coming back empty (build 92426, resolve.test.ts, six in a row).

#34069 describes the same state outside our CI, on macOS (kqueue): a bun test --parallel (or plain --isolate) worker pegged at 100% CPU inside spawnMaybeSync, the child a zombie, the pipes already closed, tickWithTimeout returning immediately, and the same nested-spawnSync escalation through the bun:test timeout path once the per-test timeout fires inside the spin. The accounting below is shared by the epoll and kqueue backends, and the new tests run on both.

Reproduced locally by running CI-shaped batches (bun test --parallel=3 --timeout=90000, CI's BUN_GARBAGE_COLLECTOR_LEVEL=1 environment, pinned to 4 CPUs): 3 of ~73 batches stalled or timed out the same way, and the stalled worker's main thread was busy-spinning inside BunObject_callback_spawnSync with its child already a zombie.

Cause

spawn_maybe_sync points vm.event_loop_handle at spawnSync's private uws loop for the duration of the call, so that the child's pidfd and pipe polls register on that loop. Everything that maintains a loop's keep-alive bookkeeping resolved that handle at the moment it ran, so it acted on whichever loop happened to be current:

  • FilePoll::activate/deactivate (num_polls and active, plus the epoll/kqueue fd used to register and unregister),
  • KeepAlive::ref_/unref (Loop::ref_/unref), used by servers, sockets, fetch, node:fs, dns, zlib, napi async work, ...
  • timer::All::increment_timer_ref / increment_immediate_ref, which ref the loop when the first JS timer or immediate appears and unref it when the last one goes away,
  • jsc::EventLoop::apply_concurrent_ref_delta, which folds the refKeepAlive counter used by MessagePort, BroadcastChannel and ScriptExecutionContext into vm.platform_loop_opt(), immediately on every JS-thread ref/unref.

So a ref taken on the main loop and released while a spawnSync is waiting was subtracted from the private loop instead. That happens whenever teardown runs during the wait: JSC sweeping a dead Timeout, port or pipe reader while spawnSync allocates its result (very frequent under --isolate + BUN_GARBAGE_COLLECTOR_LEVEL=1, where the previous file's leftovers are swept during the next file), or the test runner carrying on with the following tests after a test times out inside spawnSync. The private loop is reused for every spawnSync in the process, so the error persists for the rest of the worker's life.

us_loop_run_bun_tick returns immediately when loop->num_polls == 0. With the private loop's count off by one, a pipeless spawnSync (the napi harness) starts at 0 and spins from the first tick; a piped one reaches 0 as soon as its pipes hit EOF, before the pidfd event is dispatched. Either way the wait loop spins without polling and never observes the exit. The per-test timeout still fires inside the spin: it prints killed 1 dangling process into the worker's captured output, and if nothing is left to close the spin continues forever (the silent 240s variant); if pipe readers are still open, closing them makes the count non-zero, the loop polls once, and spawnSync returns the long-exited child with empty stdout (the timed-out variant), leaving the count wrong for the next call, hence the consecutive failures. The main loop is left with counts that are too high, which in a plain script also keeps the process from exiting.

A second problem in the same code: SpawnSyncEventLoop kept the saved handle in the struct that both levels of a nested spawnSync share (the timeout path above nests), so after the inner call returned, the outer call restored the private loop's handle and the VM stayed on the private loop for good.

Fix

Each counter records the loop it counted on and uncounts that one:

  • FilePoll gets a counted_loop, used for unregister/deactivate, re-registration and the active-count adjustments (this also makes the unregister hit the epoll/kqueue instance the fd is actually registered with). Because the poll keeps the pointer, FilePoll::register/unregister and friends now take *mut Loop instead of &mut Loop and the callers pass the loop handle they already hold; the &'static mut accessor EventLoopCtx::platform_event_loop goes away with that.
  • KeepAlive records the loop it ref'd. unref_on_next_tick keeps deferring through the pending counter when that is the thread's loop and unrefs directly otherwise, since only the thread's loop drains that counter.
  • timer::All records the loop for each of its two counts.
  • Every jsc::EventLoop records the uws loop it drives (uws_loop, previously a Windows-only field; the VM's loops get the thread's loop in ensure_waker, a spawnSync loop gets the private one) and folds its keep-alive counter into that.
  • The handle returned by SpawnSyncEventLoop::prepare lives on the calling frame and is passed back to cleanup, so nesting restores correctly.

Why this is the right shape: the invariant that broke is "an increment and its matching decrement hit the same loop", and the handle swap itself is what spawnSync's own polls rely on to find the private loop, so the swap stays and each counter remembers its loop instead. That is independent of when or why the release happens (GC, the runner, user code), and the recorded pointers are always live when used: the thread's loop is freed by VirtualMachine::teardown only after timers are cancelled and the heap is destroyed (which is also when the last fold happens), and the private loop, owned by RareData, is freed after that. On Windows the first three counters go through the thread's loop regardless of the handle, so their new fields are cfg(not(windows)); the EventLoop counter did follow the swapped handle there too and is fixed on every platform.

Tests

test/js/bun/spawn/spawnsync-isolated-event-loop.test.ts gets a block of five cases. Each runs a small file under bun test in which test a enters a spawnSync after leaving its synchronous section and times out, so the runner executes the following tests while the private loop is current. a's spawnSync holds exactly one poll on the private loop, so one misdirected release is enough to make it spin forever; on the fixed build the release comes off the main loop, a's (killed) child is observed, and the file finishes with everything but a passing.

  • four cases release one ref each of a different kind inside the wait: a JS timer (clearTimeout), a KeepAlive (Bun.serve().unref()), a registered FilePoll (cancelling a subprocess's stdout), and an event-loop keep-alive ref (BroadcastChannel.unref());
  • the fifth nests a spawnSync inside the wait and then checks that the main loop is polled again afterwards (a child whose pidfd is registered on the main loop has its exit observed).

On the unfixed build all five fail (the inner bun test spins, the outer test times out and bun test kills it). On a build with the counter fixes but without the nested-handle fix, the four release cases pass and only the nesting case fails, so the cases are independent. Each case takes ~1.4s on a debug build; the block is deliberately sequential because bun test only kills a timed-out test's dangling processes when the test is not in a concurrent group, and that is what cleans up a spinning inner process on a regressed build.

Also ran the spawn, spawnSync, maxBuffer, pidfd, child_process, BroadcastChannel, MessagePort/worker, dns, macro, fs.watch, timers and setImmediate suites on the debug build, plus bun test --parallel=3 over the CI victim files with BUN_GARBAGE_COLLECTOR_LEVEL=1. The failures left in those runs are unrelated to this change: network-dependent dns lookups, RSS and throughput tests over their budget under ASAN, a pre-existing LSan report in worker-terminate-lifetime that reproduces with a plain require("fs"), and container-specific process/uid cases; the ones that were checked against an unfixed debug build fail identically there.

Fixes #34069

While Bun.spawnSync waits for its child, vm.event_loop_handle points at
spawnSync's private uws loop. FilePoll, KeepAlive and the JS timer /
immediate ref counts all adjusted whichever loop that handle named at the
moment they ran, so a ref taken on the main loop and released during the
wait (a timer or reader swept by GC, or tests run by the test runner's
timeout path) was subtracted from the private loop instead. Once the
private loop's num_polls read 0 with the child's pidfd still registered,
us_loop_run_bun_tick returned without polling and spawnSync spun forever
without reading stdout or observing the exit.

FilePoll now records the loop it was counted on (and registered with) and
uses it for unregister/deactivate and the active-count adjustments;
KeepAlive records the loop it ref'd; timer::All records the loop each of
its two ref counts ref'd. The handle saved by SpawnSyncEventLoop::prepare
moves to the caller's frame so a spawnSync nested inside another one's
wait restores the correct handle at each level instead of leaving the VM
on the private loop.
@coderabbitai

coderabbitai Bot commented Aug 12, 2026 •

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

@robobun, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 4 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 266d4e08-60ec-418d-b48e-7db5f1ad58b0

📥 Commits

Reviewing files that changed from the base of the PR and between 2ffb8d4 and 1bb1627.

📒 Files selected for processing (13)
  • src/event_loop/SpawnSyncEventLoop.rs
  • src/io/ParentDeathWatchdog.rs
  • src/io/keep_alive.rs
  • src/io/lib.rs
  • src/io/posix_event_loop.rs
  • src/jsc/event_loop.rs
  • src/runtime/api/bun/js_bun_spawn_bindings.rs
  • src/runtime/dns_jsc/dns.rs
  • src/runtime/dns_jsc/dns_sd.rs
  • src/runtime/node/memory_pressure.rs
  • src/runtime/timer/mod.rs
  • src/spawn/process.rs
  • test/js/bun/spawn/spawnsync-isolated-event-loop.test.ts

Comment @coderabbitai help to get the list of available commands.

@robobun

robobun commented Aug 12, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 3:19 AM PT - Aug 12th, 2026

✅ @robobun, your commit 1bb1627f21b5e7e19e589d206f7b26c923212a99 passed in Build #93058! 🎉


🧪   To try this PR locally:

bunx bun-pr 37754

That installs a local version of the PR into your bun-37754 executable, so you can run:

bun-37754 --bun

@robobun

robobun commented Aug 12, 2026 •

Copy link
Copy Markdown
Collaborator Author

Status: ready for review

Reproduced locally by looping CI-shaped batches (bun test --parallel=3 --timeout=90000 over ~200 allowlisted files with CI's BUN_GARBAGE_COLLECTOR_LEVEL=1 environment, pinned to 4 CPUs): 3 of ~73 batches stalled or timed out the way the flaky annotations describe, and the stuck worker's main thread was spinning inside BunObject_callback_spawnSync with the child already a zombie. Also the state described in #34069 (macOS).

The fix is in three commits: the FilePoll / KeepAlive / timer counters plus the nested-handle restore (c1e1ed4); the EventLoop keep-alive counter used by MessagePort and BroadcastChannel, plus FilePoll taking raw loop pointers, which is the provenance point raised in review (d94f890); and a comment trim (1bb1627). The tests are one case per kind of ref plus one for the nested restore: all five fail on an unfixed build, and a build with the counters fixed but without the nested restore fails only the nesting case.

CI passed on all three commits (builds 92859, 93047, 93058). The only items in those builds are retry-passed flakes unrelated to this change (napi_get_value_string output ordering, inspect-error-leak exceeding its timeout in an ASAN batch, a watch-mode signal test on alpine aarch64). For what it is worth, the three main builds this was diagnosed from each had 3 to 5 "failed in the parallel batch, passed alone" entries; the three builds of this branch have 0, 0 and 1, and that one is the CPU-bound leak test, which does not spawn anything.

Review: the provenance thread is addressed in d94f890 and the re-review found nothing further; the comment-length bot's threads are resolved (see the comment below). A separate self-review pass over the final diff turned up nothing either.

@github-actions

Copy link
Copy Markdown
Contributor

Found 1 issue this PR may fix:

  1. Bug: spawnSync never returns: child exit lost (child stays zombie), wait loop busy-spins at 100% CPU re-registering a finished pipe reader (macOS ARM64) #34069 - Reports the exact failure this PR fixes: spawnSync never returning with the child left a zombie and the wait loop busy-spinning at 100% CPU because tickWithTimeout returns immediately (the private loop's num_polls hit 0 from cross-loop ref accounting), including the nested-spawnSync escalation from the bun:test timeout path.

If this is helpful, copy the block below into the PR description to auto-close these issues on merge.

Fixes #34069

🤖 Generated with Claude Code

@robobun

robobun commented Aug 12, 2026

Copy link
Copy Markdown
Collaborator Author

Checked #34069 against this: it is the same state. That report (macOS, so the kqueue side of the same FilePoll accounting) has the worker spinning inside spawnMaybeSync with the child a zombie and the pipes already closed, tickWithTimeout returning immediately, and the nested spawnSync through the bun:test timeout path once the per-test timeout fires inside the spin. Those are the two things this branch changes (the loop counts reaching 0 with the proc poll still registered, and the saved handle being clobbered by the nested call). Added Fixes #34069 to the description; the new test runs on macOS as well.

Comment thread src/io/posix_event_loop.rs
…p; FilePoll takes raw loop pointers

EventLoop::apply_concurrent_ref_delta() applied MessagePort / BroadcastChannel /
ScriptExecutionContext keep-alive refs to vm.event_loop_handle, which
Bun.spawnSync repoints at its private loop while it waits, so a port or channel
released during that wait (GC sweep, or tests run by the test runner's timeout
path) was subtracted from the private loop. Every EventLoop now records the uws
loop it drives (uws_loop, previously Windows-only) and folds its refs into that.

FilePoll::register/unregister and friends now take *mut Loop instead of
&mut Loop: the poll keeps the pointer in counted_loop until it is unregistered,
and a pointer derived from a caller's reborrow would not stay valid for that
long. Callers pass the loop handle they already hold, which also removes the
&'static mut accessor EventLoopCtx::platform_event_loop.

The test is now one fixture per kind of ref (JS timer, KeepAlive, FilePoll,
EventLoop keep-alive ref), each released on its own inside a spawnSync wait,
plus one for a spawnSync nested inside the wait restoring the main loop.
Comment thread src/event_loop/SpawnSyncEventLoop.rs
Comment thread src/event_loop/SpawnSyncEventLoop.rs Outdated
Comment thread src/event_loop/SpawnSyncEventLoop.rs
Comment thread src/io/keep_alive.rs Outdated
Comment thread src/io/keep_alive.rs Outdated
Comment thread src/io/lib.rs
Comment thread src/io/lib.rs
Comment thread src/io/lib.rs
Comment thread src/io/posix_event_loop.rs Outdated
Comment thread src/io/posix_event_loop.rs Outdated
Comment thread src/io/posix_event_loop.rs
Comment thread src/io/posix_event_loop.rs Outdated
Comment thread src/io/posix_event_loop.rs
Comment thread src/jsc/event_loop.rs Outdated
Comment thread src/jsc/event_loop.rs Outdated
Comment thread src/jsc/event_loop.rs
Comment thread src/runtime/api/bun/js_bun_spawn_bindings.rs Outdated
Comment thread src/runtime/timer/mod.rs Outdated
Comment thread src/runtime/timer/mod.rs Outdated
Comment thread src/event_loop/SpawnSyncEventLoop.rs
Comment thread src/io/keep_alive.rs
Comment thread src/io/posix_event_loop.rs
Comment thread src/io/posix_event_loop.rs
Comment thread src/jsc/event_loop.rs
Comment thread src/jsc/event_loop.rs
Comment thread src/runtime/timer/mod.rs
Comment thread src/runtime/timer/mod.rs
@robobun

robobun commented Aug 12, 2026

Copy link
Copy Markdown
Collaborator Author

Pushed two more commits on top of the original fix:

  • d94f890 adds the fourth counter that had the same problem, jsc::EventLoop::apply_concurrent_ref_delta (the refKeepAlive path used by MessagePort / BroadcastChannel / ScriptExecutionContext), by having every EventLoop record the uws loop it drives, and takes the review point about counted_loop's provenance: FilePoll::register/unregister now take the loop as *mut Loop, callers pass the handle they already hold, and EventLoopCtx::platform_event_loop is gone. The tests are restructured into one case per kind of ref plus one for the nested-spawnSync restore; all five fail on an unfixed build, and a build with the counters fixed but not the nested restore fails only the nesting case.
  • 1bb1627 shortens the comments (the reasoning is on FilePoll::counted_loop, the other sites point at it) and drops an unreachable fallback in unref_refd_loop.

The coderabbit comments are rate-limit notices. The comment-length bot flags every doc comment in the diff, including the one- and two-line ones left after the trim; those are the docs for the new fields and functions rather than workaround justifications, so I have resolved those threads. The PR description is updated to match.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I re-reviewed after d94f890 and didn't find any bugs. The earlier Stacked Borrows note is addressed — activate/register/unregister now take *mut Loop and counted_loop is stored from the raw handle, matching the KeepAlive::loop_ and timer_refd_loop paths. Given this reworks ref-count bookkeeping across FilePoll, KeepAlive, timer::All, and EventLoop::uws_loop on every POSIX target, a human look is still worthwhile.

Checked: counted_or routing for re-register/unregister and the keep-alive toggles; unref_on_next_tick's Loop::get() compare falling back to a direct unref; own_uws_loop vs. usockets_loop so apply_concurrent_ref_delta folds into the isolated loop; the prepare→frame-local→cleanup handle for nested spawnSync; that removed helpers (platform_event_loop, loop_unref) have no remaining callers.

Extended reasoning...

Overview

The PR fixes cross-loop ref-count leakage between the main uws loop and spawnSync's private loop by recording, at increment time, which loop each of three counters (FilePoll::counted_loop, KeepAlive::loop_, timer::All::{timer,immediate}_refd_loop) was taken on, and decrementing that same loop later. It also makes jsc::EventLoop::uws_loop a real field on POSIX so apply_concurrent_ref_delta folds into the loop this EventLoop actually drives, and moves SpawnSyncEventLoop's saved handle onto the caller's stack frame so nested spawnSync restores correctly. Signature changes ripple *mut Loop (instead of &mut Loop) through FilePoll::register*/unregister* and their call sites in process.rs, dns.rs, dns_sd.rs, memory_pressure.rs, ParentDeathWatchdog.rs, and io/lib.rs.

Security risks

None identified. No untrusted input parsing, no auth/crypto, no path handling. The new raw-pointer fields point at per-thread uws loops that outlive everything counted on them; SAFETY comments state the lifetime and single-thread invariants.

Level of scrutiny

High. This is core event-loop accounting on the hot path for every fd poll, timer, and keep-alive on POSIX, with new raw-pointer state whose lifetime is enforced only by convention. A miscount here is exactly the kind of bug the PR is fixing (spins/hangs, process-won't-exit). The change is coherent and well-tested, but it is not mechanical.

Other factors

  • My earlier inline comment about activate() storing a pointer derived from a transient &mut Loop reborrow was addressed in d94f890: FilePoll now takes and stores *mut Loop end-to-end, and counted_or/{enable,disable}_keeping_process_alive deref through loop_mut(*mut Loop) with provenance from the long-lived handle.
  • The new test exercises each ref kind (timer, KeepAlive via Bun.serve, FilePoll via a piped child, concurrent_ref via BroadcastChannel) plus the nested-spawnSync restore, and the PR description confirms it fails on the unfixed build.
  • unref_refd_loop panics if the recorded loop is null when the count crosses back to ≤0 — that's a hard invariant check on a state that should be unreachable, which seems appropriate.
  • The comment-cop bot has open style notes about comment length; those are not correctness issues and 1bb1627 already shortened several.
  • Windows is explicitly untouched (all new fields are cfg(not(windows)); EventLoopCtx::loop_add_active/loop_sub_active are now Windows-only, which matches their only remaining caller in windows_event_loop).

@robobun

robobun commented Aug 12, 2026

Copy link
Copy Markdown
Collaborator Author

Thanks for the re-review. One correction to its summary so nobody is misled: after d94f890 Windows is not entirely untouched. The FilePoll / KeepAlive / timer fields are cfg(not(windows)) because those counters already resolve to the thread's loop there, but EventLoop::uws_loop is now set on every platform and apply_concurrent_ref_delta folds into it everywhere, since prepare() swaps event_loop_handle on Windows as well and that counter followed it. The Windows, macOS and FreeBSD targets type-check; the behaviour change there is only for that one counter.

A separate self-review pass over the diff did not turn up anything further either. Waiting on CI (build 93058).

@robobun

robobun commented Aug 22, 2026

Copy link
Copy Markdown
Collaborator Author

Superseded by #40078, which applies the same invariant (every poll, keep-alive and timer ref is released against the loop it was counted on) to current main.

This branch conflicts with main in src/io/posix_event_loop.rs and src/runtime/api/bun/js_bun_spawn_bindings.rs, and its tests rely on the bun:test timeout callback running inside the spawnSync wait, which #38883 removed. #40078 carries a debug-only hook (BUN_INTERNAL_SPAWN_SYNC_GC) so the finalizer path can still be exercised deterministically.

@robobun robobun closed this Aug 22, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Bug: spawnSync never returns: child exit lost (child stays zombie), wait loop busy-spins at 100% CPU re-registering a finished pipe reader (macOS ARM64)

1 participant