Repository navigation
Conversation
|
Updated 4:23 PM PT - Jul 17th, 2026
✅ @autofix-ci[bot], your commit 3bf4bbccd1a739fbbf5aed68dadc42ed782b4a3f passed in 🧪 To try this PR locally: bunx bun-pr 31951That installs a local version of the PR into your bun-31951 --bun |
|
Found 3 issues this PR may fix:
🤖 Generated with Claude Code |
|
Checked all three against this PR's builds before adding any Fixes lines. None of them verify, so I am not claiming them:
|
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
WalkthroughStops and waits for nested child workers during parent shutdown (step 3.5), refactors termination into a filtered terminate-and-wait helper, changes futex wake semantics to wake-all, clears stale JSC termination exceptions during shutdown/log flush, and adds two regression tests for termination lifetimes. ChangesNested worker termination and shutdown guarantees
Possibly related issues
🚥 Pre-merge checks | ✅ 4✅ Passed checks (4 passed)
Comment |
There was a problem hiding this comment.
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
src/jsc/web_worker.rs (1)
419-436:⚠️ Potential issue | 🔴 Critical | 🏗️ Heavy liftWait on a monotonic wake sequence, not
OUTSTANDINGitself.Between Line 420 and Line 436,
OUTSTANDINGcan gon → n-1 → nbefore this thread actually entersFutex::wait(): one worker exits, a mid-WebWorker__createworker registers, bothwake()calls fire, and the counter is back at the originaln. At that pointFutex::wait(&OUTSTANDING, n, ...)will sleep on a stale value, so the newly registered worker is never re-swept and never getsrequested_terminate = true. If the wait then times out, the caller can still tear down the parent/main VM while that worker is instart_vm().Possible direction
pub(super) static OUTSTANDING: AtomicU32 = AtomicU32::new(0); +pub(super) static WAKE_SEQ: AtomicU32 = AtomicU32::new(0); pub(super) fn register(worker: *mut WebWorker) { MUTEX.lock(); ... OUTSTANDING.fetch_add(1, Ordering::Release); - Futex::wake(&OUTSTANDING, u32::MAX); + WAKE_SEQ.fetch_add(1, Ordering::AcqRel); + Futex::wake(&WAKE_SEQ, u32::MAX); MUTEX.unlock(); } pub(super) fn mark_exited() { OUTSTANDING.fetch_sub(1, Ordering::Release); - Futex::wake(&OUTSTANDING, u32::MAX); + WAKE_SEQ.fetch_add(1, Ordering::AcqRel); + Futex::wake(&WAKE_SEQ, u32::MAX); } - let n = live_workers::OUTSTANDING.load(Ordering::Acquire); + let n = live_workers::OUTSTANDING.load(Ordering::Acquire); + let seq = live_workers::WAKE_SEQ.load(Ordering::Acquire); live_workers::MUTEX.unlock(); ... - let _ = Futex::wait(&live_workers::OUTSTANDING, n, Some(deadline_ns - elapsed)); + let _ = Futex::wait(&live_workers::WAKE_SEQ, seq, Some(deadline_ns - elapsed));As per coding guidelines, "Rust code: fix the whole bug class in the same PR."
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/jsc/web_worker.rs` around lines 419 - 436, The wait is currently sleeping on the mutable counter live_workers::OUTSTANDING (via Futex::wait(&live_workers::OUTSTANDING, n, ...)), which can change n→n-1→n and cause a stale-wait race; instead introduce and wait on a monotonic sequence/version (e.g., live_workers::WAKE_SEQ or similar) that you increment whenever workers are added/removed or when requested_terminate is toggled, read that sequence under live_workers::MUTEX, compute the done condition the same way (using parent_filter or outstanding==0), then call Futex::wait(&live_workers::WAKE_SEQ, seq, Some(...)) so wakes are not lost (also update all corresponding wake() calls to futex::wake(&live_workers::WAKE_SEQ) when registering/unregistering workers or setting requested_terminate, and ensure WebWorker__create and any code that toggles requested_terminate increments the sequence while holding live_workers::MUTEX).Source: Coding guidelines
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Outside diff comments:
In `@src/jsc/web_worker.rs`:
- Around line 419-436: The wait is currently sleeping on the mutable counter
live_workers::OUTSTANDING (via Futex::wait(&live_workers::OUTSTANDING, n, ...)),
which can change n→n-1→n and cause a stale-wait race; instead introduce and wait
on a monotonic sequence/version (e.g., live_workers::WAKE_SEQ or similar) that
you increment whenever workers are added/removed or when requested_terminate is
toggled, read that sequence under live_workers::MUTEX, compute the done
condition the same way (using parent_filter or outstanding==0), then call
Futex::wait(&live_workers::WAKE_SEQ, seq, Some(...)) so wakes are not lost (also
update all corresponding wake() calls to futex::wake(&live_workers::WAKE_SEQ)
when registering/unregistering workers or setting requested_terminate, and
ensure WebWorker__create and any code that toggles requested_terminate
increments the sequence while holding live_workers::MUTEX).
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: 7892402e-59aa-4913-a414-fb0fff0d8faf
📒 Files selected for processing (1)
src/jsc/web_worker.rs
There was a problem hiding this comment.
Additional findings (outside current diff — PR may have been updated during review):
-
🟡
src/jsc/web_worker.rs:420-428— nit: this timeout log still saysterminateAllAndWaitand prints the globaln = OUTSTANDING, but the function is now also reached viaterminate_children_and_waitfromshutdown()step 3.5 — in that path the relevant figure ismatching(the global count always includes the caller itself plus unrelated workers). Consider branching onparent_filterto emit a distinct message withmatching, or at least genericizing the prefix.Extended reasoning...
What this is
This PR refactored
terminate_all_and_waitinto a sharedterminate_and_wait(parent_filter, timeout_ms)with a new filtered caller,terminate_children_and_wait, invoked from every exiting worker'sshutdown()step 3.5. The timeout-path debug log was not updated to reflect the generalization:if elapsed >= deadline_ns { log!("terminateAllAndWait: timed out with {} outstanding", n); return; }
When reached via the filtered path, this is misleading on two axes:
- Prefix names the wrong entry point — it says
terminateAllAndWaiteven when the caller isterminate_children_and_waitfrom a worker's shutdown. - Reports the wrong count —
nis the globalOUTSTANDINGcount, which in the filtered path always includes the caller worker itself (it unlinks after step 3.5 returns) plus any unrelated workers in the process. The figure the filtered wait actually conditions on ismatching.
Concrete example
Take Main → M → G, with M terminated and G stuck somewhere past the cooperative checkpoints for 10 s:
- M's step 3.5 sweep finds G:
matching = 1. n = OUTSTANDING.load()under the mutex →n = 2(M itself + G; M is still registered until after step 3.5).- 10 s elapses → log emits
terminateAllAndWait: timed out with 2 outstanding.
A developer with
BUN_DEBUG_Worker=1enabled who is chasing a per-worker child-wait stall sees a message that (a) names the main-thread global-exit sweep and (b) reports 2 outstanding when only 1 child of M is actually blocking the wait. If there were sibling worker trees in the process,nwould be larger still whilematchingstayed at 1.Why this isn't entirely academic
This is a
hiddenscoped debug log (define_scoped_log!(log, Worker, hidden)— only emits when the Worker debug scope is explicitly enabled via env var) on a 10 s timeout path, so there is zero functional impact and it is invisible to users. The counter-argument that this is below the nit threshold is reasonable.That said, this PR's own review history went through two rounds of debugging timeout stalls in exactly this code path (the lost-wakeup snapshot race fixed in b6b37a7, and the wake-one-vs-multiple-waiters issue fixed in 9e0980f — the latter explicitly cited a 10 s deadline stall on the debian-13 ASAN job). Anyone debugging the next stall here with the Worker scope enabled will land on this line, and "terminateAllAndWait timed out with N" pointing at the wrong sweep with an inflated count is a small but real speed bump. Since the PR is what generalized the function, updating the one diagnostic it contains to match seems worth a one-line follow-up.
Suggested fix
if elapsed >= deadline_ns { match parent_filter { Some(_) => log!("terminateChildrenAndWait: timed out with {} matching ({} outstanding)", matching, n), None => log!("terminateAllAndWait: timed out with {} outstanding", n), } return; }
(or just genericize the prefix to
terminate_and_waitand print both counts). - Prefix names the wrong entry point — it says
There was a problem hiding this comment.
Additional findings (outside current diff — PR may have been updated during review):
-
🟡
src/jsc/web_worker.rs:710-714— The parallel nested-worker comment increate()(line 631) still says "parent is itself a worker, not joined on exit" — the PR updated this rationale at the file header, theparentfield doc, andset_ref(), but missed this 4th site. With step 3.5 in place, worker parents do terminate-and-wait for children, so this wording now contradicts the rest of the file; it should get the same updateset_ref()did.Extended reasoning...
What the issue is
This PR updates the nested-worker keepalive rationale to reflect that worker parents now terminate-and-wait for their children via the new
shutdown()step 3.5. It updates that rationale at three sites:- the file header (lines 54-60): "Nested workers ARE stopped when their WORKER parent tears down:
shutdown()step 3.5 terminates and waits for them…" - the
parentfield doc (lines 88-94): "When the parent is itself a worker, itsshutdown()step 3.5 terminates us and waits for ourunlink()before freeing its VM…" - the
set_ref()comment (lines 710-714): changed from "worker parents aren't joined on exit" to "keeping the parent's loop alive avoids the child being terminated by the parent's natural exit (shutdown()step 3.5)"
But the parallel comment in
create()was missed:// src/jsc/web_worker.rs:628-632 // Keep the parent's event loop alive until the close task releases this. // If the user passed `{ ref: false }` we skip — they've opted out of the // worker keeping the process alive. Exception: a nested worker (parent is // itself a worker, not joined on exit) must hold the parent-loop keepalive // regardless, because the child holds a non-owning `BackRef` to the parent VM.
Why it's now wrong
The phrase "not joined on exit" was the old rationale: before this PR, a worker parent did not wait for its children, so the keepalive was the only thing standing between the child's
BackRefand a freed parent VM. After this PR, step 3.5 does wait for children before freeing the parent VM — so "not joined on exit" is no longer true, and the implied UAF-prevention rationale is no longer the operative one. The keepalive's purpose for nested workers is now to keep the parent's loop alive so the child isn't terminated by the parent's natural exit (exactly what the updatedset_ref()comment says).Step-by-step
- Pre-PR: file header documented "nested workers are not stopped when their WORKER parent's context tears down" → both
create()andset_ref()said "not joined on exit" → consistent. - PR adds step 3.5 (
terminate_children_and_wait) → worker parents now DO wait for children. - PR updates file header,
parentfield doc, andset_ref()to match → 3 of 4 sites consistent. create()at line 631 still says "not joined on exit" → contradicts the other 3 sites and the actual behavior.
Impact
Pure documentation inconsistency — no runtime effect. But these comments document a memory-safety invariant (the parent-VM
BackReflifetime), and a future reader who lands oncreate()first would form the wrong model of why the nested-worker keepalive exists and what guarantees it provides.Fix
Give
create()the same wordingset_ref()got, e.g.:// … Exception: a nested worker (parent is itself a worker) must hold the // parent-loop keepalive regardless: the child holds a non-owning `BackRef` // to the parent VM, and keeping the parent's loop alive avoids the child // being terminated by the parent's natural exit (`shutdown()` step 3.5).
- the file header (lines 54-60): "Nested workers ARE stopped when their WORKER parent tears down:
There was a problem hiding this comment.
All four of my earlier findings (snapshot race, wake-one vs wake-all, timeout log, stale create() comment) are addressed and the current WAKE_SEQ approach looks sound, but this reworks the cross-thread worker shutdown/futex protocol enough that a human should sign off.
Extended reasoning...
Overview
This PR fixes a use-after-free where a terminated worker frees its VirtualMachine while nested children are still inside start_vm() reading it via BackRef. The fix adds step 3.5 to WebWorker::shutdown(): terminate_children_and_wait(vm_ptr, 10_000) terminates direct children and futex-waits for each to unlink from the live-workers registry before the parent VM is freed. Supporting changes: terminate_all_and_wait is refactored into a generic terminate_and_wait(parent_filter, timeout); a new monotonic WAKE_SEQ futex word replaces waiting on OUTSTANDING (avoids ABA where exit+register restores the expected value); both register() and mark_exited() now wake-all since multiple threads can wait concurrently; flush_logs and shutdown() step 2 now clear stale TerminationExceptions and handle JsError::Terminated/Thrown instead of panic!. Two regression tests are added, and the test file is added to the LeakSanitizer exclusion list.
Security risks
None identified. This is internal lifetime/synchronization machinery for worker teardown; no auth, crypto, untrusted input parsing, or external surface.
Level of scrutiny
High. The change sits at the intersection of cross-thread raw-pointer lifetime management (extensive unsafe, BackRef validity invariants documented in load-bearing comments), a futex wait/wake protocol that this PR substantially redesigns (new WAKE_SEQ word, wake-one → wake-all, snapshot-under-mutex ordering), and the worker shutdown() sequence whose step ordering is a documented hard constraint. Earlier review rounds on this PR surfaced two real concurrency bugs (lost-wakeup race; wrong-waiter with wake-one) that were fixed, which underscores the subtlety. The 10s timeout fallback bounds any remaining protocol error to a stall rather than a UAF, but that's still user-visible.
Other factors
All four of my prior inline findings are resolved in the current diff. Since my last review the author landed 50fff74 (the WAKE_SEQ monotonic-sequence redesign — a stronger fix than the one I suggested) and c02cb19 (LeakSanitizer exclusion for the new test, with a documented rationale matching the file-header's intentional-leak note). The robobun CI comment for c99f30c shows the new test failing on x64-asan and an unrelated bunx.test.ts failure across all platforms; c02cb19 should address the former, but I haven't seen green CI for HEAD. Given the scope — new synchronization primitive, modified shutdown ordering, and a leaksan exclusion — this warrants human review rather than auto-approval.
There was a problem hiding this comment.
All four of my earlier findings have been addressed and I didn't spot anything new in the WAKE_SEQ rework, but this is cross-thread futex synchronization on the worker VM teardown path — worth a human sign-off.
Extended reasoning...
Overview
This PR fixes a use-after-free in nested Web Worker teardown by adding a terminate_children_and_wait step (3.5) to WebWorker::shutdown(), so a worker parent terminates and waits for its own children before freeing the VirtualMachine they back-reference. To support this it refactors terminate_all_and_wait into a shared terminate_and_wait(parent_filter, timeout) helper, replaces the futex word with a monotonic WAKE_SEQ counter (avoiding ABA on OUTSTANDING), switches register()/mark_exited() to wake-all, and hardens flush_logs() / the shutdown path against stale JSC TerminationExceptions. Two regression tests and a LeakSan exclusion are added; the VirtualMachine.rs change is doc-only.
Security risks
None in the conventional sense (no auth/crypto/input-handling). The change is memory-safety-relevant though: it edits the cross-thread lifetime contract that justifies every BackRef<VirtualMachine> deref in start_vm(), and the ordering of steps in shutdown() is load-bearing for soundness.
Level of scrutiny
High. This is core runtime concurrency code: an intrusive linked list walked under a process-global mutex, futex wait/wake protocol with a newly-introduced sequence word, and a multi-step teardown where each step is a barrier for the next. The PR already went through four review iterations here (lost-wakeup snapshot race, wake-one vs wake-all with multiple waiters, timeout-log accuracy, stale doc comment) — all fixed — which itself signals the subtlety. The final WAKE_SEQ design looks correct to me (snapshot taken under MUTEX before the sweep; every state change bumps it and wakes all), but a human reviewer familiar with this file's history should confirm the step-3.5 placement (after JSC teardown, before unlink()/dispatchExit) and the 10-second bounded-wait fallback are acceptable.
Other factors
- All four of my prior inline comments are resolved by commits on this branch, and the current bug-hunting pass found nothing new.
- Regression tests reproduce the original UAF and the
flush_logspanic on unfixed builds per the PR description. - The
no-validate-leaksan.txtaddition is justified (intentional leak documented in the file header) and matches existing exclusions forworker.test.ts/worker_blob.test.ts. - I'm not approving because the change is neither simple nor mechanical and sits squarely on a critical, hard-to-test code path.
|
CI status for reviewers: everything this PR touches is green. The worker test suites pass on all lanes, including debian-13 x64-asan after the LeakSanitizer exclusion (that lane's earlier failures were a flaky 32-byte leak report from the intentional worker-teardown leak, reproduced and verified locally on a release-asan build). The remaining red is unrelated: bunx.test.ts fails on every OS in every recent build of this branch, including builds whose diff was a text file or an empty commit, and the two single-platform failures in build 61283 (sql-mysql-bind-blob-borrow on alpine, install migration complex-workspace on ubuntu 25.04) are in suites this PR does not touch. |
|
Follow-up from a second, independent investigation of 1. Deterministic, race-free public-API repros for both armsThe missing ingredient this PR's test supplies via nested-worker/blob-URL races is simply that Terminated arm, one run, no race, no nesting: // main.ts
const w = new Worker(new URL("./pkg/worker.ts", import.meta.url).href);
w.addEventListener("error", () => {});
const closed = new Promise(r => w.addEventListener("close", r, { once: true }));
await new Promise(r => w.addEventListener("message", r, { once: true }));
w.terminate();
await closed;
console.log("done");The busy loop is what removes the race: the worker's only way out of JS is the termination trap, so by the time the event-loop break is observable the JSC side is guaranteed to be terminating. On a release build of main this aborts the whole process ( Thrown arm, one run, no Two warnings make Both are checked in under ~1s each in 2.
|
|
The // app/package.json -> "{invalid" (any non-fatal resolver diagnostic works)
// app/w.ts -> postMessage("up"); setInterval(() => {}, 1000);
const w = new Worker("./app/w.ts");
w.addEventListener("error", () => {});
await new Promise(resolve => w.addEventListener("message", resolve, { once: true }));
w.terminate();Resolving the worker's entry point writes the package.json diagnostics into Every run, on the 1.4.0 canary (942c222) and on a debug build of current main. Two notes on the
I pushed that variant, with a deterministic regression test (fails on unpatched main, passes with either fix), to e61f92eb714784a0722a64a4f5f04ae6d9cc0519 ( |
|
This still reproduces on today's I have the complementary half of the fix on The two approaches compose and each covers something the other does not. The terminate-and-wait here also fixes the orphaned-nested-worker behavior gap, which the snapshot does not touch; but its 10s timeout is a last resort, and a child that cannot reach a safepoint in time would still read freed memory. The snapshot removes the cross-thread read structurally with no timeout, but leaves the orphan behavior as is. One integration note if both land: I did not open a second PR since this one already covers the crash. I can fold the snapshot into this branch if that is preferred, or it can land separately after this one with the small filter-key adjustment. |
ab6ba82 to
300094d
Compare
|
Rebased onto main (3eaadbe). One conflict in |
|
CI triage for build 74278 after the rebase, plus the flush_logs review follow-up: Fixed in 1dddfe2 (just pushed): the The remaining
|
|
The
which the DevShell loader promotes to a real environment variable (the job log's env dump shows it verbatim). That var flows through Being fixed in #34355; re-running CI after that lands (or amending the commit message so no body line matches |
There was a problem hiding this comment.
No bugs found this pass, but this is cross-thread memory-safety code (worker VM lifetime, futex wait/wake redesign, JSC termination-exception handling across three re-entry sites) that has iterated enough times that a maintainer should sign off — in particular on the 10s terminate_children_and_wait fallback (still a UAF if a child can't reach a safepoint in time, per the July 2 note) and whether the complementary snapshot-the-parent-state approach should land alongside it.
What was reviewed:
WAKE_SEQfutex vs. the oldOUTSTANDINGwait — snapshot is taken underMUTEXbefore the sweep, and everyregister/mark_exitedbumps+wake-alls, so no lost-wake/ABA; a child'sunlink()doesn't bump, but any subsequent registry event (or the deadline) wakes the parent to re-sweep and observematching == 0.- Step 3.5 placement after JSC teardown — no JS can run to register a new child, and the filter can't match
self(self.parentis the grandparent, nevervm_ptr). flush_logs: all three JS-entry sites (top gate,Thrownarm,dispatcherror) now re-checkhas_requested_terminate(); the top gate is keyed on the JSC request soconfigure_defines-failure diagnostics still dispatch.
Extended reasoning...
Overview
The PR fixes a Fuzzilli-found heap-use-after-free where a terminated worker's shutdown() freed its VirtualMachine while nested children were still inside start_vm() reading it via a BackRef. The fix adds terminate_children_and_wait() (step 3.5 in shutdown()), which reuses the existing terminate_all_and_wait sweep with a parent-VM filter and futex-waits until each matching child has unlinked from the live-workers registry. Supporting changes: the futex word moved from OUTSTANDING (ABA-prone with concurrent waiters) to a monotonic WAKE_SEQ, wake-one became wake-all (multiple waiters are now possible), and flush_logs()'s panic!("unhandled exception") was replaced with termination-aware handling gated on has_requested_terminate() && has_termination_request(). Three regression tests added; the file is added to no-validate-leaksan.txt for the documented intentional teardown leak.
Security risks
None user-facing. This is internal runtime lifetime management. The change reduces a memory-safety hazard (UAF reachable from public Worker API on release builds).
Level of scrutiny
High. This is exactly the category REVIEW.md flags as most-blocked: cross-thread lifetime, futex synchronization, refcounts, and JSC exception-state handling in unsafe Rust. The PR went through 13 iterations; four separate review findings (wake-one → wake-all, stale doc comments, unreachable JsError::Terminated arm, over-broad terminate gate suppressing configure_defines errors, missing re-check in the dispatch error block) were each real and fixed in follow-up commits. That history alone argues for a human maintainer's sign-off rather than bot approval.
Other factors
- Design alternative on the table: the July 2 thread proposes a complementary structural fix (snapshot everything
start_vm()needs by value increate(), delete theparentBackRefentirely) that removes the cross-thread read without a timeout. The author noted the two compose and each covers something the other doesn't; whether to fold that in is a maintainer call. - Timeout is a last resort, not a guarantee: if a child cannot reach a cooperative safepoint within 10s, step 5 still frees the VM out from under it. The PR description doesn't claim otherwise, but it's worth a maintainer confirming that's an acceptable residual.
- All prior inline findings from my earlier passes are addressed and their threads resolved; the current diff matches the fixes described (5acc128, 5a5ff01, 1dddfe2).
- Tests are spawned fixtures with pipes drained concurrently, assert on
{exitCode, signalCode}, and were shown to fail on the unfixed ASAN build; the deterministicflush_logsrepro (malformed package.json → resolver diagnostic) was added per the June 26 follow-up.
5acc128 to
8532647
Compare
There was a problem hiding this comment.
All four rounds of my prior findings are now addressed (3ebddde handles the last two — stderr assertions and the timeout-fallback comment). No new issues found. Deferring to a human reviewer given this is ~220 lines of cross-thread memory-safety work in the worker VM lifecycle (futex/WAKE_SEQ redesign, JSC termination-exception handling at multiple JS re-entry points, shutdown ordering), plus the acknowledged residual 10s-timeout gap with a planned structural follow-up on farm/81bc4778.
Extended reasoning...
Overview
Fixes a Fuzzilli-found heap-use-after-free where a terminated worker frees its VirtualMachine while nested children are still inside start_vm() reading it via BackRef. The fix adds terminate_children_and_wait (step 3.5 in shutdown()), refactors the terminate_all_and_wait futex loop to a shared terminate_and_wait with a parent-VM filter, replaces the futex word with a monotonic WAKE_SEQ (ABA-safe, wake-all for concurrent waiters), and reworks flush_logs()'s error handling around TerminationException re-arming. Three new regression tests plus a LeakSanitizer exclusion. Touches src/jsc/web_worker.rs (~180 net lines), src/jsc/VirtualMachine.rs (comment only), the test file, and no-validate-leaksan.txt.
Security risks
None user-facing. This is internal lifecycle/teardown ordering; no untrusted-input parsing, auth, or crypto surface changes.
Level of scrutiny
High. This is squarely in REVIEW.md's most-blocked category (native memory safety, cross-thread lifetime, refcount/keepalive balance, JSC exception-scope discipline). The change redesigns a futex protocol used by concurrent waiters, introduces a new lock-ordering interaction (live_workers::MUTEX → per-worker vm_lock while a worker parent waits), changes what clear_termination_exception() is called on and where, and has an explicitly documented residual UAF on the timeout fallthrough. These are exactly the areas where a subtle mistake produces a rare production crash — a human with worker-lifecycle context should sign off.
Other factors
- This PR has been through five iterations of review-and-fix on the
flush_logstermination gate alone; each fix was correct but the density of edge cases (atomic-only vs JSC-trap-armed termination, re-arming from the trap bit, three separate JS-entry points needing the same re-check) argues for a human pass over the final shape. - The
WAKE_SEQ/ wake-all change to the pre-existingterminate_all_and_waitpath is well-reasoned (ABA + multi-waiter), but it's a behavior change to a pre-existing process-exit path, not just an addition. - The 2026-07-02 thread proposes a complementary structural fix (snapshot parent state in
create()) that composes with this one; a maintainer should decide whether to land this first or fold both together. - Tests look solid (spawned fixtures, ASAN-verified fail-before/pass-after, deterministic repro for the flush_logs case), and CI triage in the thread attributes remaining reds to unrelated/pre-existing issues.
3ebddde to
dcbb4fd
Compare
A worker's children hold a BackRef to its VirtualMachine and read it throughout start_vm() (transform options, env clone, standalone module graph). When the parent worker was terminated, its shutdown() freed that VM without waiting for children still starting up, a use-after-free caught by ASAN in VirtualMachine::init_worker. shutdown() now terminates its own children and waits for each to unlink from the live-workers registry before freeing the VM, the per-worker analogue of the main thread's terminate_all_and_wait() and of Node's Environment::stop_sub_worker_contexts(). Also fix two crashes on the same termination race: - flush_logs() panicked with "unhandled exception" when the log-to-JS conversion hit a pending TerminationException; skip dispatching the error event instead. - A TerminationException left pending after JSC clears the termination request at VM-entry-scope exit tripped ASSERT(vm.hasTerminationRequest()) in VMTraps::deferTerminationSlow when the death path re-entered JS; clear the stale exception at the re-entry points (flush_logs and the exit-handler step of shutdown).
The filtered wait derives its done-condition from the list sweep (matching) but passed a futex expected value loaded after the mutex was released. If the last matching child ran unlink() and mark_exited() in that gap, matching stayed stale at nonzero while OUTSTANDING already equaled the loaded value, so Futex::wait slept through the already-fired wake until the timeout. Loading OUTSTANDING before releasing the mutex guarantees every counted child's decrement is still ahead of the expected value, so the wait either observes the change or is woken.
terminate_children_and_wait means the main thread and any number of worker parents can wait on OUTSTANDING at the same time. A futex waiter queued with a stale expected value is only released by a wake or its timeout, so waking a single waiter can hand the event to a thread whose own condition is unmet while the one that could make progress stays asleep until its deadline (e.g. a three-level tree where the grandchild's exit wakes the grandparent instead of the middle worker). Wake all waiters in register() and mark_exited(); each re-sweeps under the mutex and re-checks its own condition, and the waiter count is bounded by the worker-tree depth.
OUTSTANDING as the futex word is ABA-prone: one worker exit plus one registration in the gap between a waiter's snapshot and its Futex::wait restores the exact expected value while both wakes fire with no waiter queued, so the wait sleeps through events it needed to re-sweep for (a newly registered worker then misses requested_terminate until the deadline forces a re-sweep). Add WAKE_SEQ, bumped on register() and mark_exited(), and wait on that; it only increments, so it cannot alias. The sequence is snapshotted under the registry mutex before the sweep and the OUTSTANDING load, so any event invalidating those inputs changes the futex word afterwards.
The timeout log always named terminateAllAndWait and printed the global OUTSTANDING count, which on the filtered path includes the calling worker itself plus unrelated workers; matching is the number the wait actually conditions on.
Each child is a full JSC VM; 12 children x 4 rounds per test overloaded saturated ASAN runners into the test timeout. 6 children x 3 rounds still crashed the unfixed build on every verification run (the race window spans the whole child VM startup), and the fixed suite drops to about 2.3s per test locally. Also update the one remaining create() comment that still described worker parents as not waiting for their children on exit.
The debian-13 x64-asan lane runs tests with detect_leaks=1 and BUN_DESTRUCT_VM_ON_EXIT=1 unless the file is listed in test/no-validate-leaksan.txt. Worker teardown intentionally leaks when a parent context is gone before the close task runs (the thread-held Worker ref and detached thread bookkeeping documented in src/jsc/web_worker.rs), and the new regression tests terminate workers while nested children are starting, which is exactly that window. Reproduced on a local release-asan build with the lane environment: a flaky 32-byte direct leak report in the spawned fixture's stderr failed the test about one run in five. worker.test.ts and worker_blob.test.ts are already listed for the same reason. Leak exclusion does not affect the tests' crash coverage; AddressSanitizer UAF detection stays on.
The JsError::Terminated match arm was unreachable: to_bun_string (via bun_string_jsc::from_js) and Log::to_js's aggregate-error path both map every failure, including a re-armed TerminationException, to JsError::Thrown. So a plain worker.terminate() that raced this flush fell into the Thrown arm and called report_uncaught_exception on JSC's termination sentinel with the exception still pending, the opposite of the skip-dispatch intent. Gate on self.has_requested_terminate() before entering any JS: a worker that is being terminated should not dispatch a late error event, and this avoids the JS re-entry entirely. Also re-check in the Thrown arm in case terminate() lands mid-call. Adds a deterministic regression test (malformed package.json next to the worker entry puts a non-fatal resolver diagnostic in vm.log; the final flush_logs on the way out then always has work to do, and terminate() is called after the worker is running). Fails on the unfixed build with ASSERT(vm.hasTerminationRequest()) in VMTraps::deferTerminationSlow; passes with the gate.
…tomic has_requested_terminate() over-captures: start_vm's configure_defines() failure path sets only the atomic (its comment says "vm.log carries the error for flushLogs") without arming the JSC trap, so JS entry would succeed and the error should dispatch. Key the gate on both the atomic and vm.jsc_vm().has_termination_request(), which is only set by notify_need_termination() for external terminate(). The re-check in the Thrown arm stays on the atomic (by that point JS has already been entered, so if the atomic flipped mid-call it was an external terminate).
… path too WebWorker__dispatchError re-enters JS (worker-side 'error' listeners, SerializedScriptValue::create), and from_js_host_call_generic reports a TerminationException that lands during that call as JsError::Thrown, not Terminated. The dispatch error handler lacked the same has_requested_ terminate() re-check as the to_js/to_bun_string Thrown arm, so a terminate() landing mid-dispatchError would reach take_exception and report_uncaught_exception on the termination sentinel with the exception still pending.
- Drop the strict stderr assertions in the two nested-worker tests and switch to the combined-object form, matching the third new test and REVIEW.md's subprocess-test rule (ASAN/debug lanes may emit benign warnings; the regression signal is fully covered by exit code 0). - Document at the step 3.5 call site that the 10s timeout is a should-never-happen fallback whose fallthrough reintroduces the UAF, and point at the complementary snapshot-in-create() structural fix (farm/81bc4778) that removes the cross-thread parent-VM read.
dcbb4fd to
1a9b9be
Compare
There was a problem hiding this comment.
No new findings this pass — the earlier flush_logs termination-gate issues and the stderr-assertion / timeout-doc nits are all addressed. Deferring to a human reviewer given the scope: this reworks the live-workers futex protocol (adds WAKE_SEQ, wake-all), inserts a blocking wait on the worker-thread teardown path, and changes TerminationException handling at three JS re-entry points.
What was reviewed:
terminate_and_waitrefactor:WAKE_SEQsnapshot is taken under the mutex before the sweep, so a child'sunlink+mark_exitedcannot land between the condition check and the futex wait; the filtered exit keys on list membership (matching == 0), and the caller's own entry is skipped becauseself.parent≠vm_ptr.flush_logs: top gate now requireshas_termination_request()so theconfigure_definesshort-circuit (atomic-only) still dispatches; both theThrownarm and thedispatchErrorerror handler re-check beforetake_exception.- Step 3.5 placement: after
teardownJSCVM(no JS → no new child canregister), beforeunlink/dealloc; timeout fallthrough is documented as a residual gap with the follow-up branch named. - Tests: pipes drained concurrently, stderr not asserted empty,
signalCode: nullcovered.
Extended reasoning...
Overview
Fixes a nested-worker use-after-free: a worker parent's shutdown() freed its VirtualMachine while children were still inside start_vm() reading it via BackRef. The fix adds terminate_children_and_wait (step 3.5 in shutdown()), which reuses the existing terminate_all_and_wait sweep parameterized by a parent_filter. Supporting changes: the futex word moves from OUTSTANDING to a monotonic WAKE_SEQ (avoiding an ABA between one exit + one register, and because multiple waiters — main thread + N worker parents — now sleep on it and need wake-all); flush_logs and shutdown() step 2 clear the stale TerminationException before re-entering JS and gate/re-check on termination so the sentinel is never fed to report_uncaught_exception. Three regression tests added; the file is added to no-validate-leaksan.txt alongside the other worker tests.
Security risks
None identified. This is internal lifetime/concurrency plumbing with no user-facing input parsing, auth, or network surface. The change reduces a memory-safety hazard (heap UAF reachable from public API) rather than introducing one.
Level of scrutiny
High. web_worker.rs is production-critical concurrency code — every field carries a threading contract, and the diff touches (a) a process-global futex protocol shared by the main thread and every worker, (b) the ordered teardown sequence whose step numbering exists precisely because reordering is unsafe, and (c) JSC TerminationException semantics that already tripped two follow-up findings during this review. The 10s timeout at step 3.5 is a documented best-effort fallback whose fallthrough reintroduces the UAF; that trade-off (vs. the structural snapshot-in-create() alternative on farm/81bc4778) is a design call a maintainer should sign off on.
Other factors
The PR has been through many iterations; every prior inline finding from this reviewer is resolved and the current bug-hunting pass found nothing. Test coverage is solid (three new spawned-fixture regressions verified to fail on the unfixed ASAN build). CI triage in the thread attributes remaining red to pre-existing/unrelated flakes. Still, ~230 net lines in a file whose header is a hand-maintained lifetime proof is beyond what should land without a human maintainer's eyes.
On the unfixed build the subprocess panics and dumps core, which under ASAN takes about 7s (the panic path prints a full stack then aborts); the default 5s test timeout fired before the assertion could run, producing a hang+dangling-process cleanup instead of a clean assertion failure. Passing the file's timeout (60s on slow lanes) lets the fail-before assertion land.
There was a problem hiding this comment.
No further issues found — all four rounds of earlier findings (unreachable JsError::Terminated arm, over-broad has_requested_terminate() gate, missing re-check in the dispatch error handler, stderr assertions / timeout-fallback doc) have been addressed and are visible in the current diff. Deferring to a human for the merge call: this is ~230 lines of cross-thread lifetime and futex-synchronization changes in worker VM teardown, with an acknowledged residual gap (the 10s timeout fallthrough, now documented at the step 3.5 call site) and a competing/complementary structural approach on farm/81bc4778 that a maintainer should weigh in on.
What was reviewed:
terminate_and_waitrefactor:WAKE_SEQmonotonic futex word (ABA avoidance), wake-all, snapshot-before-sweep ordering, and the parent-filter exit condition (matching == 0vsOUTSTANDING == 0) — the caller's own registry entry is correctly excluded.shutdown()step 3.5 placement: after JSC teardown (no new children can register), beforeunlink/dealloc; the filter never dereferencesw.parent, onlyptr::eq.flush_logstermination handling: gate keyed on the JSC request (soconfigure_definesfailure still dispatches),clear_termination_exception()before JS re-entry, and re-checks in both theThrownarm and thedispatcherror handler.- Tests: subprocess pipes drained concurrently, stderr not asserted empty, combined
{stdout, exitCode, signalCode}assertions;no-validate-leaksan.txtentry scoped to LSan only.
Extended reasoning...
Overview
The PR fixes a Fuzzilli-found heap-use-after-free where a terminated worker frees its VirtualMachine while nested children are still inside start_vm() reading it via BackRef. The fix adds terminate_children_and_wait (a parent-filtered variant of the existing terminate_all_and_wait) as step 3.5 of shutdown(), refactors the futex wait to a monotonic WAKE_SEQ word with wake-all semantics (needed now that multiple threads can wait concurrently), and hardens flush_logs/shutdown against stale/re-armed TerminationExceptions. Four files touched: web_worker.rs (~230 lines net), a 4-line comment update in VirtualMachine.rs, three new regression tests in worker-terminate-lifetime.test.ts, and a LeakSanitizer exclusion.
Security risks
None in the injection/auth/data-exposure sense. This is memory-safety-critical concurrent native code: the class of risk is UAF/race, which is exactly what the PR is fixing. The change is a strict improvement over the pre-PR state (every nested-termination raced to UAF; now only a >10s stall inside start_vm() — which runs no user JS/FFI — can hit the documented fallback).
Level of scrutiny
High. This touches cross-thread worker VM lifetime, futex synchronization semantics, JSC TerminationException / VMTraps interaction, and shutdown() step ordering — all of which are subtle and have already produced four rounds of real findings during review. Per REVIEW.md's memory-safety guidance ("know the thread affinity of every line", "never let a pointer outlive the memory it points into"), this warrants a maintainer's eyes even though the current state looks correct to me and the bug-hunting system found nothing new.
Other factors
- Prior review history: I raised four issues over the course of today's review (unreachable
Terminatedarm; gate catching the internalconfigure_definesshort-circuit; missing re-check afterWebWorker__dispatchError; stderr-empty assertions + timeout doc). All are fixed in the current diff and marked resolved. - Architectural decision pending: the 2026-07-02 thread proposes a complementary snapshot-in-
create()approach onfarm/81bc4778that removes the cross-thread parent-VM read structurally. The step 3.5 comment now references it as follow-up. Whether to land this first and follow up, or fold both together, is a maintainer call. - Residual gap acknowledged: the 10s timeout fallthrough is now documented at the call site as reintroducing the UAF under pathological load — honest, but a maintainer should sign off on shipping a bounded-wait-then-free rather than a fail-safe.
- Test coverage: three new spawned-fixture regression tests, each verified to crash the unfixed ASAN build; existing worker suites reported green on the rebased debug build. Evidence block in the PR description shows the fails-without/passes-with gate.
- CI: prior triage attributed remaining reds to unrelated pre-existing flakes (Windows
worker.test.ts0xC0000409 reproduces on main;bunx.test.ts; a commit-message env-var leak fixed separately in #34355).
…t_thrown - The blanket has_requested_terminate() early-return swallowed the configure_defines() failure dispatch (start_vm self-signals via set_requested_terminate() alone, without arming the JSC trap, and relies on spin()'s first checkpoint to flush_logs the error). Gate on has_termination_request() too so only an external terminate (which arms the trap via notify_need_termination) skips the dispatch. - Extract the terminate-check/take_exception/report sequence into a report_thrown closure used by both error arms (was duplicated with inconsistent None handling). - Drop the dead flush_logs call in the new spin() checkpoint. - Add VM::has_termination_request() (the FFI already existed). - Restructure the stress test to two levels (main spawns failing workers directly, no middle worker) so it does not also trip the nested-worker parent-VM UAF that #31951 addresses. The earlier three-level fixture caught that UAF on the debian-13 x64-asan lane in build 81404.
|
superseded by #37075, parents now join their child workers before teardown. nested worker + process.exit no longer crashes on main |
What does this PR do?
Fixes a use-after-free found by Fuzzilli (fingerprint
Address:heap-use-after-free, flaky, on a Worker thread).Root cause: a worker's children hold a
BackRefto the parent'sVirtualMachineand read it throughoutstart_vm()(transform options clone, env map clone,standalone_module_graphinVirtualMachine::init_worker). When the parent is itself a worker and gets terminated, itsshutdown()freed that VM immediately. Children still starting up then read freed memory:This was the "Known gap" documented in the
web_worker.rsheader: the main thread waits for all workers at exit (terminate_all_and_wait), but a worker parent did not wait for its own children. The first regression test below crashes release builds too, so this is reachable in production, not only under ASAN.The fix: the fixing line is the
terminate_children_and_wait(vm_ptr, 10_000)call inshutdown()(step 3.5): before freeing its VM, an exiting worker terminates the workers whose parent VM is its own and futex-waits until each has unlinked from the live-workers registry, which is past every parent-VM read. It reuses the existingterminate_all_and_waitsweep, parameterized with a parent filter. This mirrors Node'sEnvironment::stop_sub_worker_contexts(). The placement after JSC teardown means no JS can run to register new children, so the sweep is complete.Two more crashes on the same termination race, surfaced by the same repro once the UAF was fixed:
flush_logs()hitpanic!("unhandled exception")when a child's entry-point resolution failed (revoked blob URL) while its termination was in flight: the pending TerminationException makes the log-to-JS conversion returnJsError::Terminated. Termination is a normal event, not an invariant violation, so skip dispatching the error event. A genuineJsError::Thrownfrom the conversion is reported throughreport_uncaught_exceptioninstead of crashing.handleTrapsre-arms from the trap bit). Re-entering JS in that state tripsASSERT(vm.hasTerminationRequest())inVMTraps::deferTerminationSlow(debug builds). The worker death path now clears the stale exception at its JS re-entry points, using the existingclear_termination_exception()(same pattern as the test runner and repl).How did you verify your code works?
Two regression tests in
test/js/web/workers/worker-terminate-lifetime.test.ts, each a spawned fixture where a worker creates children and is terminated while they start:flush_logspanic and the VMTraps assertion path; crashes 5/5 on the unfixed ASAN build.Also ran the existing worker suites (
worker.test.ts,worker_blob.test.ts,worker-terminate-lifetime.test.ts,worker_threads.test.ts,worker-async-dispose.test.ts,message-port-context-destroy-leak.test.ts) with the debug+ASAN build. The pre-existing failures inworker.test.ts("worker with event listeners doesn't close event loop", 1000ms budget) andworker_threads.test.ts("eval does not leak source code") reproduce identically on an unfixed build on this runner; they are unrelated timing flakes on slow debug machines.The file is also added to
test/no-validate-leaksan.txt: the debian-13 x64-asan lane runs tests withdetect_leaks=1, and worker teardown intentionally leaks when a parent context is gone before the close task runs (the thread-held Worker ref documented inweb_worker.rs), which these tests exercise by design. Reproduced on a local release-asan build (flaky 32-byte leak report in the fixture's stderr, about one run in five);worker.test.tsandworker_blob.test.tsare already excluded for the same reason. AddressSanitizer's use-after-free detection is unaffected by the exclusion.[review] gate passed · iteration 23 · 4 files touched
fails on main (without fix)
passes on PR (with fix)
diff hotspot
gate history · 8 passed · 1 rejected · iteration 23
evidence per changed file