Skip to content

JSCTaskScheduler: drop deferred-work tasks scheduled after the event loop's last tick - #34293

Merged
Jarred-Sumner merged 4 commits into
mainfrom
farm/161a219b/jsc-deferred-work-leak
Jul 17, 2026
Merged

Jarred-Sumner merged 4 commits into
mainfrom
farm/161a219b/jsc-deferred-work-leak

Conversation

@robobun

@robobun robobun commented Jul 16, 2026 •

Copy link
Copy Markdown
Collaborator

test/js/web/timers/timer-heap-race.test.ts went red on the x64-asan lane after #33131 landed (build 73570):

SUMMARY: AddressSanitizer: 1280 byte(s) leaked in 40 allocation(s).
  ConcurrentTask::create src/event_loop/ConcurrentTask.rs:319
  Bun__queueJSCDeferredWorkTaskConcurrently src/jsc/JSCScheduler.rs:64
  Bun::JSCTaskScheduler::onScheduleWorkSoon JSCTaskScheduler.cpp:54
  JSC::DeferredWorkTimer::scheduleWorkSoon DeferredWorkTimer.cpp:235
  JSC::Waiter::cancelAndClear WaiterListManager.cpp:298
  JSC::WaiterListManager::unregister(JSC::VM*) WaiterListManager.cpp:310
  JSC::VM::~VM() VM.cpp:591
  WebWorker__teardownJSCVM Worker.cpp:676

The leak is pre-existing; #33131's new cross-thread Atomics.waitAsync fixture is the first test that terminates a worker with pending async waiters on a SharedArrayBuffer the parent keeps alive.

Cause

Worker shutdown() drains the concurrent task queue once (release_queued_tasks_for_shutdown), then calls WebWorker__teardownJSCVM, which ends in ~VM(). ~VM() runs WaiterListManager::unregister(this), and for every pending Atomics.waitAsync ticket that reaches Waiter::cancelAndClear → DeferredWorkTimer::scheduleWorkSoon → our onScheduleWorkSoon hook. The hook allocates a JSCDeferredWorkTask and a ConcurrentTask and enqueues them into the worker's concurrent queue, which was just drained for the last time. When the worker's VirtualMachine box is raw-dealloc'd, both become unreachable. The same path is reachable from the final collectNow via JSFinalizationRegistry::finalizeUnconditionally.

A second, narrower leak: a cross-thread Atomics.notify that lands between the worker's last tick and teardownJSCVM enqueues a JSCDeferredWorkTask that release_queued_tasks_for_shutdown forwards into self.tasks. __bun_release_task_at_shutdown had no arm for that tag, so it was re-queued, and EventLoop::deinit re-queued it once more into a freshly allocated LinearFifo buffer that leaked on worker dealloc.

Fix

  • JSCTaskScheduler gets an std::atomic<bool> m_isShuttingDown, set at the start of WebWorker__teardownJSCVM and Zig__GlobalObject__destructOnExit (mirroring the existing ctx->markTerminating()). onScheduleWorkSoon and onAddPendingWork drop the work once it's set; onScheduleWorkSoon also balances the onAddPendingWork ref via onCancelPendingWork.
  • __bun_release_task_at_shutdown gains a JSCDeferredWorkTask arm that deletes the job via a new Bun__deleteDeferredWorkTask FFI. This runs before JSC teardown, so ~Ref<TicketData> and the captured Task lambda release against a live VM.

Test

timer-heap-atomics-teardown-fixture.ts terminates a worker with 32 pending Atomics.waitAsync tickets on a parent-owned SAB, a few of them notified cross-thread first, under detect_leaks=1. Without the fix LSan reports ~29 ConcurrentTask allocations from WaiterListManager::unregister and SIGABRTs; with it the fixture exits clean. The original race fixture is also 0/10 failures under the CI env (was ~1/5 on a debug build and 1/1 on release-asan).

Overlap with #34270

#34270 adds the same m_isShuttingDown/onScheduleWorkSoon gate (there named m_isTerminating) while fixing a separate FinalizationRegistry assert, but without the __bun_release_task_at_shutdown arm the race fixture still fails ~1/10 on that branch. Whichever lands first, the other is a small rebase over JSCTaskScheduler.{h,cpp}.


[review] gate passed · iteration 3 · 9 files touched

fails on main (without fix)
ASAN without fix: 1 FAILED
$ BUN_DEBUG_QUIET_LOGS=1 bun scripts/build.ts --profile=debug --quiet test "--reporter=junit" "--reporter-outfile=/tmp/mechgate.xml" test/js/web/timers/timer-heap-race.test.ts
info: syncing channel updates for nightly-2026-05-06-x86_64-unknown-linux-gnu
info: latest update on 2026-05-06 for version 1.97.0-nightly (e95e73209 2026-05-05)
info: component rust-src is up to date
info: checking for self-update (current version: 1.29.0)
bun test v1.4.0 (4d9b926bb)

test/js/web/timers/timer-heap-race.test.ts:
(pass) timer heap survives cross-thread Atomics.waitAsync timeout cancellation [3806.16ms]
(pass) timer heap stays consistent while GC re-arms the RunLoop timer [2408.11ms]
51 |       ASAN_OPTIONS: "allow_user_segv_handler=1:disable_coredump=0:detect_leaks=1:abort_on_error=1",
52 |       LSAN_OPTIONS: `malloc_context_size=30:print_suppressions=0:suppressions=${path.join(import.meta.dir, "..", "..", "..", "leaksan.supp")}`,
53 |     });
54 |     // LSan writes its leak report to stderr and SIGABRTs; stdout holds the
55 |     // fixture's own OK line either way, so assert exitCode/signal explicitly.
56 |     expect({ stdout, stderr, signal, exitCode }).toEqual({
        
... (truncated)

release without fix: 2 skipped
bun test v1.4.0-canary.1 (1498d7b77)

test/js/web/timers/timer-heap-race.test.ts:
(pass) timer heap survives cross-thread Atomics.waitAsync timeout cancellation [3042.35ms]
(skip) timer heap stays consistent while GC re-arms the RunLoop timer
(skip) terminating a worker with pending Atomics.waitAsync tickets does not leak deferred-work tasks

 1 pass
 2 skip
 0 fail
 1 expect() calls
Ran 3 tests across 1 file. [3.20s]
__F:0:S:2
passes on PR (with fix)
ASAN with fix: all passed
$ BUN_DEBUG_QUIET_LOGS=1 bun scripts/build.ts --profile=debug --quiet test "--reporter=junit" "--reporter-outfile=/tmp/mechgate.xml" test/js/web/timers/timer-heap-race.test.ts
info: syncing channel updates for nightly-2026-05-06-x86_64-unknown-linux-gnu
info: latest update on 2026-05-06 for version 1.97.0-nightly (e95e73209 2026-05-05)
info: component rust-src is up to date
info: checking for self-update (current version: 1.29.0)
bun test v1.4.0 (4d9b926bb)

test/js/web/timers/timer-heap-race.test.ts:
(pass) timer heap survives cross-thread Atomics.waitAsync timeout cancellation [3805.27ms]
(pass) timer heap stays consistent while GC re-arms the RunLoop timer [2481.63ms]
(pass) terminating a worker with pending Atomics.waitAsync tickets does not leak deferred-work tasks [7403.00ms]

 3 pass
 0 fail
 3 expect() calls
Ran 3 tests across 1 file. [15.73s]
__F:0:S:0

release with fix: 2 skipped
$ bun scripts/build.ts --profile=release
info: syncing channel updates for nightly-2026-05-06-x86_64-unknown-linux-gnu
info: latest update on 2026-05-06 for version 1.97.0-nightly (e95e73209 2026-05-05)
info: component rust-src is up to date
info: checking for self-update (current version: 1.29.0)
[configured] bun-profile → bun (stripped)
  target       linux-x64-gnu
  build type   Release
  build dir    ./build/release
  revision     4d9b926bb1
  features     (none)

22 deps, 106 codegen, 1168 objects in 891ms

ninja: Entering directory `/workspace/bun/build/release'
[1/1231] install /workspace/bun
bun install v1.4.0-canary.1 (1498d7b77)

Checked 124 installs across 170 packages (no changes) [13.00ms]
[2/1231] install /workspace/bun/packages/bun-error
bun install v1.4.0-canary.1 (1498d7b77)

Checked 1 install across 2 packages (no changes) [3.00ms]
[3/1231] install /workspace/bun/src/node-fallbacks
bun install v1.4.0-canary.1 (1498d7b77)

Checked 129 installs across 147 packages (no changes) [11.00ms]
[4/1231] gen ErrorCode+*.h
[5/1231] fetch zlib
[zlib] up to date
[6/1231] gen bindgenv2
[7/1231] fetch picohttpparser
[picohttpparser] up to date
[8/1231] gen .bind
... (truncated)
diff hotspot
src/jsc/VirtualMachine.rs                          |  8 +++
 src/jsc/bindings/JSCTaskScheduler.cpp              | 77 ++++++++++++++++++----
 src/jsc/bindings/JSCTaskScheduler.h                | 13 ++++
 src/jsc/bindings/ZigGlobalObject.cpp               |  2 +
 src/jsc/bindings/webcore/Worker.cpp                |  5 ++
 src/jsc/web_worker.rs                              |  8 +++
 src/runtime/dispatch.rs                            | 15 +++++
 .../timers/timer-heap-atomics-teardown-fixture.ts  | 32 +++++++++
 test/js/web/timers/timer-heap-race.test.ts         | 25 ++++++-
 9 files changed, 171 insertions(+), 14 deletions(-)

gate history · 2 passed · 0 rejected · iteration 3

evidence per changed file
file                                                      reads  edits  tests
src/jsc/VirtualMachine.rs                                     5      2      0
src/jsc/bindings/JSCTaskScheduler.cpp                         7     10      0
src/jsc/bindings/JSCTaskScheduler.h                           5      4      0
src/jsc/bindings/ZigGlobalObject.cpp                          1      1      0
src/jsc/bindings/webcore/Worker.cpp                           1      1      0
src/jsc/web_worker.rs                                         5      4      0
src/runtime/dispatch.rs                                       2      1      0
…st/js/web/timers/timer-heap-atomics-teardown-fixture.ts      0      2      0
test/js/web/timers/timer-heap-race.test.ts                    2      3      0

…loop's last tick

Worker teardown runs release_queued_tasks_for_shutdown, then
WebWorker__teardownJSCVM which ends in ~VM(). ~VM() calls
WaiterListManager::unregister for every Atomics.waitAsync ticket still
pending on that VM, and each one reaches
DeferredWorkTimer::scheduleWorkSoon -> JSCTaskScheduler::onScheduleWorkSoon,
which allocates a JSCDeferredWorkTask and a ConcurrentTask and enqueues
into the worker's concurrent queue. That queue was already drained for
the last time, so the nodes become unreachable when the worker's
VirtualMachine box is dealloc'd and LSan reports them.

Gate onScheduleWorkSoon and onAddPendingWork on a new m_isShuttingDown
atomic, set at the start of WebWorker__teardownJSCVM and
Zig__GlobalObject__destructOnExit so work scheduled from the final
collectNow or ~VM() is dropped instead of enqueued.

A JSCDeferredWorkTask can also land in the queue before the flag is set
(cross-thread Atomics.notify from another VM while this one is between
its last tick and teardown). release_queued_tasks_for_shutdown forwards
it into self.tasks where __bun_release_task_at_shutdown didn't recognise
the tag, so it was re-queued into a freshly allocated LinearFifo buffer
that leaked on worker dealloc. Add a JSCDeferredWorkTask arm that
deletes the job via a new Bun__deleteDeferredWorkTask FFI.

Surfaced by test/js/web/timers/timer-heap-race.test.ts on the x64-asan
lane (build 73570) after #33131 added the cross-thread Atomics.waitAsync
fixture.
@robobun

robobun commented Jul 16, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 3:19 PM PT - Jul 16th, 2026

❌ @robobun, your commit 4d9b926 has 1 failures in Build #74072 (All Failures):


🧪   To try this PR locally:

bunx bun-pr 34293

That installs a local version of the PR into your bun-34293 executable, so you can run:

bun-34293 --bun

@coderabbitai

coderabbitai Bot commented Jul 16, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Walkthrough

Changes

Deferred work scheduling now observes VM shutdown, rejects new work, and reclaims queued deferred-work tasks during teardown. VM and worker shutdown paths signal the scheduler, and an ASAN-gated Atomics.waitAsync regression test verifies cleanup.

Deferred work shutdown

Layer / File(s) Summary
Shutdown state and teardown signaling
src/jsc/bindings/JSCTaskScheduler.h, src/jsc/bindings/ZigGlobalObject.cpp, src/jsc/bindings/webcore/Worker.cpp, src/jsc/web_worker.rs, src/jsc/VirtualMachine.rs
JSCTaskScheduler stores a lock-protected shutdown state, while VM and worker teardown signal scheduler and deferred-work shutdown.
Scheduler guards and task cleanup
src/jsc/bindings/JSCTaskScheduler.cpp, src/runtime/dispatch.rs
Pending tickets use shared removal logic, new work is rejected after shutdown, and queued JSCDeferredWorkTask payloads are deleted.
Atomics teardown regression coverage
test/js/web/timers/timer-heap-atomics-teardown-fixture.ts, test/js/web/timers/timer-heap-race.test.ts
A worker fixture creates pending Atomics.waitAsync tickets, and an ASAN-gated test checks clean termination without LeakSanitizer output.

Possibly related issues

Possibly related PRs

  • oven-sh/bun#34270 — Changes overlapping JSCTaskScheduler shutdown and deferred-work scheduling paths.
  • oven-sh/bun#34278 — Changes worker shutdown ordering with an early termination marker before queued-task draining.

Suggested reviewers: jarred-sumner, cirospaciari

🚥 Pre-merge checks | ✅ 3 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Description check ⚠️ Warning The description is detailed, but it does not use the required template headings for the PR summary and verification steps. Reformat the description to match the template with '### What does this PR do?' and '### How did you verify your code works?' sections.
✅ Passed checks (3 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly summarizes the main fix: dropping deferred-work tasks scheduled after the event loop's last tick.

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/jsc/bindings/JSCTaskScheduler.cpp`:
- Around line 118-124: Release the pending ticket before reclaiming queued jobs:
in src/jsc/bindings/JSCTaskScheduler.cpp lines 118-124, add a cancel-and-delete
FFI entry point that invokes onCancelPendingWork() for job->ticket before
deleting the job; in src/runtime/dispatch.rs lines 1206-1214, bind and call this
new entry point instead of Bun__deleteDeferredWorkTask.

In `@src/jsc/bindings/JSCTaskScheduler.h`:
- Around line 23-31: Serialize shutdown with deferred-work enqueueing: in
src/jsc/bindings/JSCTaskScheduler.h#L23-L31, update
JSCTaskScheduler::markShuttingDown to use the scheduler’s synchronization
protocol; in src/jsc/bindings/JSCTaskScheduler.cpp#L55-L68, hold the same lock
across the m_isShuttingDown check and the ownership handoff to
Bun__queueJSCDeferredWorkTaskConcurrently, so no enqueue can occur after
teardown releases queued tasks.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 7a6e08f6-4bce-49e8-9e29-c8c73d3fe101

📥 Commits

Reviewing files that changed from the base of the PR and between d29b050 and dccb8b9.

📒 Files selected for processing (7)
  • src/jsc/bindings/JSCTaskScheduler.cpp
  • src/jsc/bindings/JSCTaskScheduler.h
  • src/jsc/bindings/ZigGlobalObject.cpp
  • src/jsc/bindings/webcore/Worker.cpp
  • src/runtime/dispatch.rs
  • test/js/web/timers/timer-heap-atomics-teardown-fixture.ts
  • test/js/web/timers/timer-heap-race.test.ts

Comment thread src/jsc/bindings/JSCTaskScheduler.cpp
Comment thread src/jsc/bindings/JSCTaskScheduler.h Outdated
@github-actions

Copy link
Copy Markdown
Contributor

This PR may be a duplicate of:

  1. JSCTaskScheduler: skip DeferredWorkTimer tasks once a worker is terminating #34270 - Both add an atomic shutdown flag to JSCTaskScheduler (m_isShuttingDown vs m_isTerminating) with the same early-return guard in onScheduleWorkSoon to prevent ConcurrentTask leaks from DeferredWorkTimer during worker teardown

🤖 Generated with Claude Code

Comment thread src/jsc/bindings/webcore/Worker.cpp
Comment thread test/js/web/timers/timer-heap-race.test.ts
Address review: the atomic flag alone leaves a gap where a cross-thread
Atomics.notify that reads m_isShuttingDown == false before the worker
thread sets it can still enqueue after release_queued_tasks_for_shutdown
has drained.

Hold m_lock across the check and the enqueue in onScheduleWorkSoon (and
across the check in onAddPendingWork, matching the existing lock there),
and take the same lock in markShuttingDown. Introduce
Bun__JSCTaskScheduler__markShuttingDown so worker shutdown can flip the
flag before the drain: a notifier that enqueues under the lock before
the flip is visible to the drain via the queue's own release/acquire,
and a notifier that acquires the lock after the flip sees the flag and
drops.

Factor the pending-ticket removal into dropPendingTicketLocked so the
shutdown paths (onScheduleWorkSoon's early return and
Bun__deleteDeferredWorkTask) balance the event-loop ref and ticket-set
entry the same way runPendingWork would have.

Relax the new test's stderr assertion to not.stringContaining per the
file convention; abort_on_error=1 already turns a leak into SIGABRT.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Beyond the inline nit, I checked a7c0da0's lock change for ordering hazards — onScheduleWorkSoon now holds m_lock across Bun__queueJSCDeferredWorkTaskConcurrently, but that's a lock-free MPSC push (enqueue_task_concurrent), so no nesting concern. Also verified the onCancelPendingWork → dropPendingTicketLocked refactor is behavior-preserving (else-branch just unlocks slightly earlier). My two earlier nits (drain→flag window on the worker path, stderr: "" assertion) are both addressed.

Extended reasoning...

a7c0da0 addressed all four prior review threads (two CodeRabbit majors, my two nits). The remaining inline finding is the main-thread sibling of the drain→flag ordering fix — nit-level and BUN_DESTRUCT_VM_ON_EXIT-gated. Not approving: this is cross-thread VM-teardown code with a known overlap to coordinate against #34270, so a human should sign off.

Comment thread src/jsc/bindings/ZigGlobalObject.cpp
Mirror the worker shutdown path so a cross-thread scheduleWorkSoon that
races the main-thread BUN_DESTRUCT_VM_ON_EXIT teardown is either caught
by the drain or observes m_isShuttingDown under m_lock. The main VM box
is static-rooted so LSan would not flag it, but the sibling paths should
match.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

bf3fe64 addresses my last note — global_exit() now fences JSCTaskScheduler before both concurrent-queue drains, symmetric with web_worker.rs::shutdown(). No further issues found; deferring to a human for the cross-thread teardown / m_lock-scope changes and the merge-order call vs #34270.

What was reviewed:

  • m_lock now held across Bun__queueJSCDeferredWorkTaskConcurrently — verified that path is a lock-free MPSC push + wakeup (JSCScheduler.rs:56-65), no lock-order hazard.
  • Bun__deleteDeferredWorkTask mirrors runPendingWork's ticket take + event-loop-ref balance under the shared dropPendingTicketLocked helper.
  • global_exit() placement of markShuttingDown before drop_concurrent_cpp_tasks vs. web_worker.rs placing it after close_all_socket_groups — ruled out as a problem (onAddPendingWork also gates on the flag, so JS run during shutdown that creates new deferred work just drops it).
Extended reasoning...

Overview

This PR gates JSCTaskScheduler::onScheduleWorkSoon/onAddPendingWork on a new lock-protected m_isShuttingDown flag so deferred-work tasks scheduled during VM teardown (~VM → WaiterListManager::unregister, collectNow → JSFinalizationRegistry::finalizeUnconditionally) are dropped instead of enqueued into a concurrent queue that will never drain again. It also adds a JSCDeferredWorkTask arm to __bun_release_task_at_shutdown so tasks already queued at shutdown are deleted (with pending-ticket / event-loop-ref balance) rather than re-queued and leaked. Touches JSCTaskScheduler.{h,cpp}, Worker.cpp, ZigGlobalObject.cpp, web_worker.rs, VirtualMachine.rs, dispatch.rs, plus a new ASAN-gated fixture and test.

Review history

I left three prior inline comments across two review passes; all were addressed:

  • a7c0da0 serialized the shutdown transition under m_lock (closing the check-then-enqueue TOCTOU CodeRabbit and I both flagged), moved the worker-side markShuttingDown before the drain, made Bun__deleteDeferredWorkTask release the pending ticket, and switched the test's stderr assertion to expect.not.stringContaining("LeakSanitizer").
  • bf3fe64 mirrored the pre-drain fence into VirtualMachine::global_exit().

The bug-hunting system found no issues on the current head. One finder candidate — that global_exit() sets the flag before any close_all_socket_groups-equivalent JS runs, unlike web_worker.rs — was verified not to be a bug: onAddPendingWork also checks m_isShuttingDown, so new tickets created by JS during shutdown are dropped rather than half-tracked.

Security risks

None. This is internal VM/worker teardown lifecycle management with no user-controlled input, auth, or data-exposure surface.

Level of scrutiny

High. This is exactly the category CLAUDE.md flags as most-blocked: cross-thread lifetime management, a lock-scope change (m_lock now held across an FFI enqueue), and refcount balancing on error/shutdown paths. I verified the enqueue under lock is a lock-free MPSC push (no re-entry into m_lock), that dropPendingTicketLocked is a behavior-preserving extraction of the pre-existing onCancelPendingWork body, and that the three markShuttingDown call sites are all idempotent. The reasoning looks sound, but VM-teardown concurrency is subtle enough that a maintainer should sign off.

Other factors

  • The PR description explicitly calls out overlap with #34270 (same m_isShuttingDown gate under a different name). Whichever lands first, the other needs a rebase over JSCTaskScheduler.{h,cpp} — that's a human coordination decision.
  • Test coverage: the new ASAN-gated fixture terminates a worker with 32 pending Atomics.waitAsync tickets (some notified cross-thread first) under detect_leaks=1:abort_on_error=1; asserts signal: null / exitCode: 0 so an LSan SIGABRT fails the test independent of stderr noise.

@robobun

robobun commented Jul 16, 2026

Copy link
Copy Markdown
Collaborator Author

CI on bf3fe64 (build 73608): 277 passed, 2 failed. timer-heap-race.test.ts passes on the x64-asan lane (the one that was red in build 73570).

The only red test is test/js/node/test/parallel/test-net-connect-memleak.js on alpine 3.23 x64 and x64-baseline, which is pre-existing (#33045 quarantines it on linux-x64-musl; #20627 tracks the underlying fix) and unrelated to this change; it passes 5/5 locally under this build.

Everything else is yellow (passed on retry): complex-workspace, fetch-backpressure, test-repl-close, test-fs-promises-file-handle-readFile, no-orphans, npmrc, net-mongodb-pattern-leak, serve-protocols. The darwin-aarch64 test shards are still waiting on a scheduled build-rust job.

Ready for review.

@Jarred-Sumner

Copy link
Copy Markdown
Collaborator

@robobun fix conflicts

Resolve conflict in src/jsc/web_worker.rs: both ScriptExecutionContext
markTerminating (postTaskTo fence, #34278) and JSCTaskScheduler
markShuttingDown (scheduleWorkSoon fence, this PR) are called before the
final concurrent-queue drain.
@robobun

robobun commented Jul 16, 2026 •

Copy link
Copy Markdown
Collaborator Author

CI on 4d9b926 (build 74072): 285/286 passed. timer-heap-race.test.ts passes on every lane including x64-asan.

The one red is test/js/node/test/parallel/test-worker-message-port-transfer-terminate.js SIGABRT on debian 13 x64-asan, which is a documented pre-existing intermittent flake on that lane (see the header comment in test/js/node/worker_threads/worker-transfer-terminate-stress.test.ts added by #32488: "SIGABRTs intermittently on the x64-asan lane only ... It does not reproduce locally (0/115 loaded runs)"). The assertion is in JSObject::getOwnPropertyDescriptor during MessagePort structured-clone serialization, unrelated to the JSCTaskScheduler path this PR touches; 0/30 locally under this build.

Everything else is yellow (passed on retry). Ready for review.

@Jarred-Sumner
Jarred-Sumner merged commit 73411ad into main Jul 17, 2026
78 of 79 checks passed
@Jarred-Sumner
Jarred-Sumner deleted the farm/161a219b/jsc-deferred-work-leak branch July 17, 2026 02:02
robobun added a commit that referenced this pull request Jul 17, 2026
…hutdown handling

WebKit: merged 8be995561a (AsyncContextSwapScope RAII refactor),
ae5110d307 (reifyStaticProperty termination), 365cb02471 (Dockerfile.windows)
into the upgrade branch. JSMicrotask.cpp resolved by keeping upstream's
microtaskCallCache threading and asyncFunctionGeneratorBodyCall extraction on
top of the new RAII helper; AsyncGeneratorDriverResume now uses
wrapWithCurrent / unwrapContextTuple like its siblings.

Bun: merged origin/main. JSCTaskScheduler.{h,cpp} resolved by keeping the
new m_isShuttingDown gate, dropPendingTicketLocked, markShuttingDown and
Bun__deleteDeferredWorkTask from #34293, updated to the new
DeferredWorkTimer::Ticket / Ref<Ticket>&& / Ticket& API.

WEBKIT_VERSION bumped to the preview tag for the new WebKit head.
cirospaciari added a commit that referenced this pull request Jul 17, 2026
Only src/js/node/worker_threads.ts conflicted, in three hunks where #34338
("don't hang when captured stdout/stderr is never consumed") and this branch
touch the same lines. #34338 removed the #stdoutAutoPipe/#stderrAutoPipe fields
and moved the stdio port ref/unref out of ref()/unref() — ports now manage their
own ref via makePortReadable's incrementsPortRef. This branch only added #hasRef
bookkeeping there, so main's structure is taken wholesale and only the two
`if (!this.#exited) this.#hasRef = ...` lines and the field are kept.

async_hooks.ts (#31825) and VirtualMachine.rs (#34293, #32498) auto-merged.

`git diff origin/main -- src/js/node/worker_threads.ts` is a pure addition:
zero deleted lines, so nothing from #34338 or #31825 is reverted.

Verified on the merge result: test-worker-hasref, test-worker-error-stack-
getter-throws, test-perf-hooks-worker-timeorigin, test-diagnostics-channel-
worker-threads and the new "online fires before the entry point finishes" all
pass; #34338's own repro still exits 0 like node; BroadcastChannel ref()/unref()
and the 'online' timing fix both still match node v26.3.0.
Jarred-Sumner pushed a commit that referenced this pull request Sep 6, 2026
… atomics fixture by iterations (#41428)

### Problem
- `test/js/web/timers/timer-heap-race.test.ts` is in the slowest 5
percent of test files in CI: 12s on debian 13 x64-asan (build 110300).
Locally on the debug ASAN build it takes 11.3s.
- The three tests run serially (3.5s + 2.1s + 3.7s) although each one
spawns one independent child. The atomics fixture runs for a fixed 3s of
wall time, so a release build does 25x more rounds than a debug build
for the same 3s.

### Fix
- The three tests run with `it.concurrent`. The file takes as long as
its slowest child.
- `timer-heap-atomics-fixture.ts` runs a fixed number of pump rounds per
worker instead of a clock deadline. The test passes the count (100
rounds, 3 workers). 100 rounds is what the old 3s gave on the debug ASAN
build, the lane where the heap assertion lives, so that lane keeps the
same number of cross-thread collisions. The main thread hammers until
every worker reports done.
- Assertions: the atomics fixture reports `OK 3 workers 300 pumps` and
the test asserts the exact line from the shared constants. The GC
fixture takes its tick count from the test (`GC_TICKS = 30`) and the
test asserts `ok 30` from the same constant. The LeakSanitizer stderr
check on the teardown test is unchanged.
- Verified: `bun bd test test/js/web/timers/timer-heap-race.test.ts`
went from 11.28s to 5.6s (three runs: 5.63s, 5.55s, 5.56s).
`USE_SYSTEM_BUN=1 bun test` on the same file went from 3.2s to 0.2s.
Also ran `test/js/web/timers/setTimeout.test.js` and
`timer-gc-roots.test.ts`: the same four pre-existing debug ASAN failures
as on main (the three RSS leak tests and one 5s timeout), none in files
this PR touches.

### Background
- The atomics fixture guards #33131. `Atomics.waitAsync` timeouts are
`WTF::RunLoop` timers that Bun keeps in its own heap (`All.wtf_timers`
in `src/runtime/timer/mod.rs`). Other threads can touch that heap, so it
has its own mutex, separate from the `setTimeout` heap.
- The teardown fixture guards #34293: a worker terminated with live
`waitAsync` tickets must drop its deferred-work tasks, not enqueue them
into the dead loop. It runs only under ASAN with `detect_leaks=1`.
- The 20s ceilings stay. Each fixture still runs for seconds on a debug
build and the default is 5s.

<details><summary>Notes</summary>

- About 3s of the teardown test is LeakSanitizer's exit scan. A trivial
`bun -e 'console.log(1)'` with `ASAN_OPTIONS=detect_leaks=1` takes 3.3s
on the debug build. The fixture itself is 0.5s. That cost is not
reachable from the test and is now overlapped with the other two tests.
- I tried to calibrate the round count against the original race by
building a debug binary with the `wtf_timers` pop moved outside its
lock. The fixture ran 30s without tripping. The vendored JSC no longer
stops a `waitAsync` timer from the notifying thread
(`Waiter::clearTimer` in `WaiterListManager.h` only drops its
reference), and `DeferredWorkTimer::scheduleWorkSoonIfActive` goes
through Bun's `onScheduleWorkSoon` hook instead of a RunLoop timer. So
the cross-thread cancel the fixture was written against does not exist
in this tree, and the count cannot be proven against it. The count keeps
the debug ASAN lane at its previous iteration budget.
- Release builds did about 2500 rounds per worker in the old 3s. They
now do 100. The release binary compiles out the heap assertion, so the
debug lane is the one that detects a heap corruption.
- GC fixture per-tick cost on the debug build is dominated by
`Bun.gc(true)`, not by the 16ms delay, so the delay is unchanged.
</details>

<!-- robobun:evidence:begin -->

---

**[auto-merge]** gate passed · iteration 1 · 3 files touched

<details><summary>passes on PR (with fix)</summary>

```console
Test-only change.

Debug/ASAN (expected pass):
$ bun bd test 'test/js/web/timers/timer-heap-race.test.ts'
$ BUN_DEBUG_QUIET_LOGS=1 bun scripts/build.ts --profile=debug --quiet test test/js/web/timers/timer-heap-race.test.ts
bun test v1.4.3 (e0a2b82)

test/js/web/timers/timer-heap-race.test.ts:
(pass) timer heap stays consistent while GC re-arms the RunLoop timer [1945.78ms]
(pass) timer heap survives cross-thread Atomics.waitAsync timeout cancellation [3519.14ms]
(pass) terminating a worker with pending Atomics.waitAsync tickets does not leak deferred-work tasks [3624.34ms]

 3 pass
 0 fail
 3 expect() calls
Ran 3 tests across 1 file. [5.55s]
Exit: 0
```

</details>

<details><summary>diff hotspot</summary>

```
test/js/web/timers/timer-heap-atomics-fixture.ts | 27 +++++++++------
 test/js/web/timers/timer-heap-gc-fixture.ts      |  2 +-
 test/js/web/timers/timer-heap-race.test.ts       | 44 ++++++++++++++++--------
 3 files changed, 47 insertions(+), 26 deletions(-)
```

</details>

**gate history** · 2 passed · 0 rejected · iteration 1

<details><summary>evidence per changed file</summary>

```
file                                              reads  edits  tests
test/js/web/timers/timer-heap-atomics-fixture.ts      1      2     10
test/js/web/timers/timer-heap-gc-fixture.ts           1      2      7
test/js/web/timers/timer-heap-race.test.ts            1      2      6
```

</details>

**root cause** · written by the author bot

The file was slow because its three independent subprocess tests ran
serially and the atomics fixture spun against a fixed wall-clock
deadline rather than stopping once it had done a bounded amount of
race-provoking work, so every run paid the full sleep regardless of how
quickly the race reproduced. The fix switches the tests to it.concurrent
so the file's wall time collapses to the slowest single fixture, and
replaces the clock-based loop with a fixed per-worker pump budget that
each worker reports back, calibrated to match the previous duration on
the debug ASAN lane where the heap asser…

<!-- robobun:evidence:end -->
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants