Skip to content

test: clear the stale stack slot that keeps the last abandoned fetch body alive - #41607

Open
robobun wants to merge 2 commits into
mainfrom
robobun/77f13b0b/fetch-stream-cancel-leak-stale-slots
Open

robobun wants to merge 2 commits into
mainfrom
robobun/77f13b0b/fetch-stream-cancel-leak-stale-slots

Conversation

@robobun

@robobun robobun commented Sep 6, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

  • test/js/web/fetch/fetch-stream-cancel-leak.test.ts > "one read() after it parked" went red on darwin aarch64 in builds 110683 and 110902: expect(N - aborted).toBeLessThan(N / 4), Expected: < 5, Received: 5. Up to 4 of the 20 abandoned bodies were allowed to survive the 3 s wait, and 5 did.
  • The survivors are not rooted by bun. After the last body's final read(), its ReadableStreamDefaultController stays in a stale slot of the JSC::runInternalMicrotask frame that ran that read. The wait loop resumes through that same frame, so the conservative stack scan marks the controller, and through it the stream, on every Bun.gc(true). Nothing in the loop writes that slot again, so the stream lives until unrelated work does, about one second later on an idle box, and past the deadline on the CI lane.

Fix

  • Before each collection the loop reads one chunk of a throwaway body from a /scrub route of the same origin and cancels it. That runs the same fetch-body code path, so the stale slot then holds a stream nobody cares about. /scrub aborts do not count.
  • With the slot cleared, the assertion is exact: expect(aborted).toBe(N). Unfixed (fetch: park an unread body stream instead of buffering it without bound #39590 reverted), none of the bodies are aborted, so the test still fails.
  • The test still exercises the same paths: park on 256 KiB unread, read() after the park, unpark, re-park, collect, abort from the sweep.
  • Verified: the file 15 of 15 times on darwin aarch64 with the CI build of 110902 (release and profile), 6 of 6 with the release build on linux x64, 3 of 3 with bun bd test (ASAN). The stall is gone in all of them.

Background

  • A fetch body that nothing reads is parked once it holds 256 KiB (fetch: park an unread body stream instead of buffering it without bound #39590): the transport pauses and the tasklet stops rooting the stream, so the GC can collect it and the sweep aborts the fetch. The test counts those aborts at the origin.
  • JSC scans the native stack conservatively: every word between the stack pointer and the stack base that looks like a cell pointer keeps that cell alive. A stale word inside a live frame is not sanitized.
  • runInternalMicrotask is JSC's dispatcher for internal microtasks. Promise reactions of the stream machinery and async-function resumes both run through it, at the same depth, but they spill different values into the same frame area.
Notes

How the retainer was found. A standalone copy of the "dropped while a reader holds the lock" shape stalls for about 1 s in 25 to 45 % of runs on darwin and linux release builds. During the stall: the origin's socket has a full send queue, the client socket a full receive queue (netstat, ss -tnie: rwnd_limited), the HTTP thread does no recvfrom on that socket (fs_usage), and the JS thread sends no wakeup to the HTTP thread (a DYLD_INSERT_LIBRARIES interposer on mach_msg). So the stream is parked and the transport paused, and the only thing missing is the collection.

generateHeapSnapshotForDebugging() taken after Bun.gc(true) during the stall: the stream, its controller, reader, BytesInternalReadableStreamSource and NativeStreamSourceAdapter form a cluster with no incoming edge from outside it and no roots entry.

lldb on the profile build of 110902, stopped from inside the stall: the controller's cell address 0x45b98648e70 is on the main thread stack at 0x16fdfdc38, inside frame #10 JSC::runInternalMicrotask (fp 0x16fdfdd20) of the chain drainWithUseCallOnEachMicrotask > runInternalMicrotask > asyncModuleExecutionResume > the script's continuation > Bun.gc. The fp chain was walked by hand so JIT frames do not stop it. /proc/self/mem on linux finds the same controller address in the [stack] mapping.

Why a throwaway fetch and not a throwaway ReadableStream. A JS-source stream or a Bun.file().stream() read before each collection did not help (stalls in 16 to 19 of 20 runs, they spill into other slots). A throwaway fetch body read made 30 of 30 runs collect everything within 100 ms on darwin and 20 of 20 on linux. One throwaway fetch read in the middle of a stall collects the retained stream within 20 ms in 12 of 12 stalled runs.

Why the 1 s on an idle box. Not measured to the instruction. The slot is overwritten only by the same microtask path, and the origin's own stream microtasks stop while it is backpressured. The next unrelated work on the JS thread clears it.

The CI count. 5 survivors did not reproduce locally, with or without CPU load, under the release or the profile build. The mechanism is the same for every survivor that the heap snapshot cannot explain, and the scrub re-runs the whole path (tasklet progress task, ByteStream::on_data, the pull microtasks, the read request), so it overwrites the slots of each of those frames.

Related. #41111 rewrites this file and waits on WeakRefs with an exact count. It notes the same ~970 ms stall. The scrub applies there too.

Suites run: this file under the darwin release and profile builds of 110902 (15 runs), the linux release build (6 runs), and bun bd test ASAN (3 runs).


[auto-merge] gate passed · iteration 1 · 1 files touched

passes on PR (with fix)
Test-only change.

Debug/ASAN (expected pass):
$ bun bd test 'test/js/web/fetch/fetch-stream-cancel-leak.test.ts'
$ BUN_DEBUG_QUIET_LOGS=1 bun scripts/build.ts --profile=debug --quiet test test/js/web/fetch/fetch-stream-cancel-leak.test.ts
bun test v1.4.3 (f42e98025)

test/js/web/fetch/fetch-stream-cancel-leak.test.ts:
(pass) ReadableStream from fetch should be GC'd after reader.cancel() [915.97ms]
(pass) ReadableStream from fetch should be GC'd after body.cancel() [831.24ms]
(pass) response.body.cancel() on a never-read body aborts the underlying fetch [418.98ms]
(pass) an abandoned fetch body stream is collected and its fetch is aborted > res.body touched [535.51ms]
(pass) an abandoned fetch body stream is collected and its fetch is aborted > one read(), then releaseLock() [677.89ms]
(pass) an abandoned fetch body stream is collected and its fetch is aborted > dropped while a reader holds the lock [578.03ms]
(pass) an abandoned fetch body stream is collected and its fetch is aborted > one read() after it parked [3602.14ms]
(pass) a proxied fetch body is aborted once the response's client is gone > new Response(upstream.body), client leaves while reading [140.51ms]
(pass) a proxied fetch body is aborted once the response's client is gone > new Response(upstream.body), client leaves without reading [231.04ms]
(pass) a proxied fetch body is aborted once the response's client is gone > the upstream Response itself, client leaves while reading [80.22ms]
(pass) a proxied fetch body is aborted once the response's client is gone > the upstream Response itself, client leaves without reading [199.19ms]

 11 pass
 0 fail
 15 expect() calls
Ran 11 tests across 1 file. [12.39s]
Exit: 0
diff hotspot
test/js/web/fetch/fetch-stream-cancel-leak.test.ts | 17 +++++++++++++----
 1 file changed, 13 insertions(+), 4 deletions(-)

gate history · 2 passed · 0 rejected · iteration 1

evidence per changed file
file                                                reads  edits  tests
test/js/web/fetch/fetch-stream-cancel-leak.test.ts      2      3     11

…body alive

The abandoned-body tests collect 20 unread fetch bodies and count the
aborts the origin sees. The last body's controller stays in a stale slot
of the JSC::runInternalMicrotask frame that ran its final read, and the
conservative stack scan keeps it alive through every Bun.gc(true) in the
wait loop. The loop now reads one chunk of a throwaway body and cancels
it before each collection. That runs the same code path, so the slot
holds a stream nobody cares about. The assertion becomes exact.
@coderabbitai

coderabbitai Bot commented Sep 6, 2026 •

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

  • Run on-demand review

On-demand reviews are free for the next 14 days. After that, they cost $0.25 per reviewed file.

Or wait 10 minutes for your next included review.

Check out review usage here.

View limit details

Limit details: You’ve used all 10 included reviews currently available.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Essentials

Run ID: e68641f2-5b33-4e4e-ba91-c21afe8f2e80

📥 Commits

Reviewing files that changed from the base of the PR and between d316760 and e00fa6d.

📒 Files selected for processing (1)
  • test/js/web/fetch/fetch-stream-cancel-leak.test.ts

Comment @coderabbitai help to get the list of available commands.

@robobun

robobun commented Sep 6, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 9:00 AM PT - Sep 6th, 2026

❌ @robobun, your commit e00fa6d has 3 failures in Build #111285 (All Failures):

  • test/js/web/fetch/fetch-backpressure.test.ts - code 1 on 🪟 11 aarch64
  • test/js/node/test/parallel/test-crypto-dh-leak.js - code 1 on 🐧 13 x64-asan
  • ❌ CPU instruction violation on Linux x64 — 1 check(s) failed
  • The baseline build contains instructions not available on Nehalem (SSE4.2, no AVX/AVX2/AVX512).

    • Static instruction scan

    Static scan violations

    llint_op_wide16  [INVLPGB]  (1 insns)
    

    If these are runtime-dispatched behind a CPUID gate: add each symbol to scripts/verify-baseline-static/allowlist-x64.txt with a comment pointing at the gate.

    If there's no gate: this is a real bug — a -march leaked into a subbuild.

  • ❌ CPU instruction violation on Linux x64 — 1 check(s) failed
  • The baseline build contains instructions not available on Nehalem (SSE4.2, no AVX/AVX2/AVX512).

    • Static instruction scan

    Static scan violations

    llint_op_wide16  [INVLPGB]  (1 insns)
    

    If these are runtime-dispatched behind a CPUID gate: add each symbol to scripts/verify-baseline-static/allowlist-x64.txt with a comment pointing at the gate.

    If there's no gate: this is a real bug — a -march leaked into a subbuild.


🧪   To try this PR locally:

bunx bun-pr 41607

That installs a local version of the PR into your bun-41607 executable, so you can run:

bun-41607 --bun

@robobun

robobun commented Sep 6, 2026

Copy link
Copy Markdown
Collaborator Author

Status: the diff is ready for review. Reproduced the retention with a standalone copy of the abandoned-body shapes on darwin aarch64 (CI build of 110902) and linux x64 release: the last body stalls about 1 s in 25 to 45 % of runs. The heap snapshot and the lldb stack search in the PR body show the retainer is a stale slot in the JSC::runInternalMicrotask frame. With the scrub, 15 of 15 runs on darwin, 6 of 6 on linux and 3 of 3 under the ASAN build collect all 20 bodies. The coderabbit comment above is a rate-limit notice, nothing to address.

Comment thread test/js/web/fetch/fetch-stream-cancel-leak.test.ts Outdated
@robobun

robobun commented Sep 6, 2026

Copy link
Copy Markdown
Collaborator Author

Evidence for the retainer, as CLAUDE.md rule 15 asks for. Both artifacts come from the release build of CI build 110902 (the profile variant for lldb, since it carries symbols) and a standalone copy of the "dropped while a reader holds the lock" shape: 20 fetches, one read() each, then a Bun.gc(true) / Bun.sleep(10) loop. The last body is the one that stalls.

1. Heap snapshot. generateHeapSnapshotForDebugging() after Bun.gc(true) during the stall, on linux x64. The cluster of the retained stream has no incoming edge from outside the cluster and no entry in roots:

node ReadableStream#55 size=96 addr=0x20fdd240b80
   -> Property:bunNativePtr BytesInternalReadableStreamSource#79
   -> Property:controller ReadableStreamDefaultController#50
   -> Property:reader ReadableStreamDefaultReader#1052
incoming edges, transitively:
  ReadableStreamDefaultController#50 --Property:stream--> #55
  ReadableStreamDefaultReader#1052 --Property:stream--> #55
    NativeStreamSourceAdapter#60 --Internal--> #50
    ReadableStream#55 --Property:controller--> #50
    ReadableStream#55 --Property:reader--> #1052
      ReadableStreamDefaultController#50 --Property:algorithmContext--> #60
      Function#128 --Internal--> #60
      Function#973 --Internal--> #60
        BytesInternalReadableStreamSource#79 --Property:onDrainCallback--> #128
        BytesInternalReadableStreamSource#79 --Property:onCloseCallback--> #973
          NativeStreamSourceAdapter#60 --Internal--> #79
          ReadableStream#55 --Property:bunNativePtr--> #79
roots entries for #55, #50, #1052, #60, #79, #128, #973: none

ReadableStream#55 is the only stream node in the snapshot with a bunNativePtr that is a BytesInternalReadableStreamSource. A WeakRef to it still dereferences at that point.

2. Debugger session. lldb on bun-profile (darwin aarch64), the process stopped by process.kill(process.pid, "SIGSTOP") from inside the stall, right after the snapshot above was taken in the same process and the cluster addresses were written to a file. A script walks the main thread's frame-pointer chain by hand (the unwinder stops at JIT frames) and searches every word from sp to the stack top for the cluster addresses:

cluster addrs: 0x45b98644760(stream) 0x45b98650980(ReadableStreamDefaultReader)
  0x45b98648e70(ReadableStreamDefaultController) 0x45b98670270(BytesInternalReadableStreamSource)
  0x45b98678270(NativeStreamSourceAdapter)
sp0 0x16fdfd4d0 frames in fp chain: 29
  #0 fp=0x16fdfd510 libsystem_kernel.dylib`__kill
  #1 fp=0x16fdfd610 bun-profile`jsc_llint_0_doVMEntry__copyArgsDone_LowLevelInterpreter64_asm_177
  #2 fp=0x16fdfd780 bun-profile`JSC::Interpreter::executeCall(...)
  #3 fp=0x16fdfd860 bun-profile`Bun::Process_functionKill(JSC::JSGlobalObject*, JSC::CallFrame*)
  #4 fp=0x16fdfd870 ?`0x104ffdf7c            (JIT frame)
  #5 fp=0x16fdfd990 ?`0x10500b0b4            (JIT frame)
  #6 fp=0x16fdfdaa0 bun-profile`llint_call_javascript
  #7 fp=0x16fdfdb90 bun-profile`JSC::Interpreter::executeModuleProgram(...)
  #8 fp=0x16fdfdbc0 bun-profile`JSC::JSModuleRecord::evaluate(...)
  #9 fp=0x16fdfdc00 bun-profile`JSC::asyncModuleExecutionResume(...)
  #10 fp=0x16fdfdd20 bun-profile`JSC::runInternalMicrotask(JSC::JSGlobalObject*, JSC::VM&, JSC::InternalMicrotask, unsigned char, std::span<JSC::JSValue const, 4ul>, JSC::MicrotaskCallCache*)
  #11 fp=0x16fdfdfc0 bun-profile`JSC::MicrotaskQueue::drainWithUseCallOnEachMicrotask(...)
  #12 fp=0x16fdfe070 bun-profile`JSC::VM::drainMicrotasks()
  #13 fp=0x16fdfe090 bun-profile`Zig::GlobalObject::drainMicrotasks()
  #14 fp=0x16fdfe100 bun-profile`<bun_jsc::event_loop::EventLoop>::exit
  #15 fp=0x16fdfe1a0 bun-profile`<bun_runtime::timer::timer_object_internals::TimerObjectInternals>::fire
  #16 fp=0x16fdfe310 bun-profile`__bun_fire_timer
  #17 fp=0x16fdfe390 bun-profile`<bun_runtime::timer::All>::drain_timers
  #18 fp=0x16fdfe480 bun-profile`bun_runtime::jsc_hooks::auto_tick
  #19 fp=0x16fdfe4e0 bun-profile`<bun_jsc::event_loop::EventLoop>::wait_for_promise
  #20 fp=0x16fdfe530 bun-profile`<bun_jsc::virtual_machine::VirtualMachine>::load_entry_point
  ... #27 main, #28 dyld`start
scanned 0x16fdfd4d0 .. 0x16fe00000 (11056 bytes)
HIT 0x45b98648e70 at 0x16fdfdc38 (sp+0x768): frame #10 fp=0x16fdfdd20 fn=bun-profile`JSC::runInternalMicrotask(...)

The only cluster address on the stack is the controller's, at 0x16fdfdc38, between asyncModuleExecutionResume's fp (0x16fdfdc00) and runInternalMicrotask's fp (0x16fdfdd20), so in the locals of runInternalMicrotask. That frame is live: it is the one resuming the script after await Bun.sleep(10), and Bun.gc(true) runs under it. The microtask it is running (the async module resume) does not take a controller; the stream's own pull and enqueue microtasks ran through the same frame earlier. On linux the same search through /proc/self/mem finds the controller address in the [stack] mapping at a fixed offset.

3. The slot is what holds it. With the process stalled, one throwaway fetch body read (the scrub in this PR) is followed by the collection and the abort within 20 ms, in 12 of 12 stalled runs. A throwaway JS-source ReadableStream or a Bun.file().stream() read in the same place does not release it (16 to 19 of 20 runs still stall), because they spill into other slots.

The transport was also checked during the stall: the client socket's receive queue is full, the HTTP thread does no recvfrom on it (fs_usage), the JS thread sends it no wakeup (mach_msg interposed), and the server has a full send queue. That is the parked state #39590 intends. The only missing step is the collection.

@robobun

robobun commented Sep 6, 2026 •

Copy link
Copy Markdown
Collaborator Author

Status: ready for a maintainer. Build 111285 ran this file on every test lane, including both darwin aarch64 jobs (the lane that went red in 110683 and 110902), darwin x64, windows x64 and aarch64, and the ASAN lane: all passed. The red jobs in that build are not touched by this diff and are red on main too: verify-baseline on x64 and x64-musl (LLInt misdecode in the static instruction scan, tracked separately), test/js/node/test/parallel/test-crypto-dh-leak.js on x64-asan, and test/js/web/fetch/fetch-backpressure.test.ts on windows 11 aarch64 (both reported as main breaks). A retrigger would hit the same three, so I am not pushing one.

Reproduction and evidence: #41607 (comment) (heap snapshot with no retainer, lldb stack word at 0x16fdfdc38 in the live JSC::runInternalMicrotask frame).

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code review found no issues

No high-confidence issues detected in this change.

robobun added a commit to oven-sh/WebKit that referenced this pull request Sep 16, 2026
…ng to occupy

A microtask checkpoint is where a turn of the embedder's event loop resumes an async function: a
timer or an I/O callback settles a promise, and the queue is drained. Every turn does that from the
same stack depth, so the frames of the checkpoint (MicrotaskQueue::drain*(), runInternalMicrotask(),
the job's, the entry to JS) are at the same addresses in every turn. They stay for as long as the
turn's JS runs, and they lie over whatever ran at that depth since the last checkpoint: that
checkpoint's own frames, or the embedder's. What they do not write is in reach of the conservative
scan of every collection made from under them, and sanitizeStackForVM(), which such a collection
calls, clears only what is below the frame that calls it.

#673 took the words that ASan poisons (redzones, locals out of scope) out of the scan. This is about
the words it cannot take out: spill slots and locals of a live frame that this job's path through
the function does not write. Builds without ASan have only those, and ASan builds have them too.
runInternalMicrotask() is one switch over every kind of job, with callMicrotask() inlined into
several arms, and a job uses one arm.

Seen as: what a finished async function held is not collected by collections made from later turns.

    async function makeGarbage() { /* two objects in a list, one await per object */ }
    await makeGarbage();
    for (;;) { await turn(); fullGC(); }    // turn(): a promise that a timer resolves

jsc shell of c281568 (linux amd64, lto): 32 turns until the objects are collected (the loop
tiers up and the frames change), 60 of 60 turns not collected with --useJIT=0. Bun at that WebKit
(release, x86_64): not collected in 8 of 8 turns when the timer is Bun.sleep(1); Bun's debug ASan
build, which has #673: not collected in 8 of 8 turns for the module shape with any timer. A heap
snapshot has the objects held by `list`, held by the lexical environment of the finished
function's JSAsyncFunctionGenerator, which has no incoming edge and no root. In a debugger, the one
word of the scanned span that holds the generator's address is at rbp-64 of
runInternalMicrotask()'s frame (272 bytes), under drainWithUseCallOnEachMicrotask() (624 bytes).
The job that runs in the later turns (AsyncModuleExecutionResume for the module,
AsyncFunctionResume when the loop is in an async function of its own) does not write that slot.
Which shape shows it moves with the compiler: the same script with setTimeout() kept the objects
with clang 21 and does not with clang 23. Earlier sightings of this frame in Bun's leak tests:
oven-sh/bun#37853 (the same generator word), oven-sh/bun#41607 (darwin arm64, release).

So a checkpoint that has jobs to run first clears MicrotaskQueue::stackBytesClearedForCheckpoint
bytes of the stack below performMicrotaskCheckpoint(), before the first of those frames exists:
2 KB in a release build (from there to the JS frame is 1.5 KB for x86_64), 32 KB with assertions or
ASan (12.6 KB with both). It does that with a function that is never inlined and whose frame is an
array of that size, which it zeroes with memset(): that frame lies exactly where the caller's next
callee has its own, on every ABI, and the compiler probes the stack for it where a platform needs
that. The array's address goes through an empty asm statement. Without it the stores are dead to
the compiler (zeroBytes() and secureZeroBytes() both compile to a bare `ret` here: the memory
clobber of secureZeroSpan() does not keep stores to a local whose address does not escape).
callMicrotask() asserts that it runs within seven eighths of the window below the checkpoint's
MicrotaskCallCache, which is a local of drainImpl().

Not sanitizeStackForVM() at the checkpoint: it clears from where it was last called, and nothing
need have called it from below the checkpoint since. The jsc shell does not change with it
(JSLock::didAcquireLock() resets VM::m_lastStackTop for every task). Not sanitizeStackForVMImpl()
with m_lastStackTop lowered to the bottom of the window, which was the first version: its loop for
x86_64 stores 8 bytes per iteration, and a checkpoint with one job went from 106 ns to 195 ns.

Cost, Xeon 8375C, clang 23: `promise.then(noop); drainMicrotasks()` 5 M times in the jsc shell,
101 ns per iteration without the clear and 122 ns with it. Bun, 1 M turns of
`await new Promise(r => setImmediate(r))`: 2531 ms without and 2547 ms with (medians of 11, same
binary, the option). 20 M awaits in one checkpoint: 602 ms and 602 ms. An empty checkpoint clears
nothing.

Once per checkpoint and not once per job, which would cost every job that much: what one job
leaves is still there for the later jobs of the same checkpoint, and is gone at the next one. Not
when the window would reach below the soft stack limit.

An embedder that holds the API lock for the life of its thread can get much of this from a
sanitizeStackForVM() per event loop turn instead (oven-sh/bun#37853): in Bun that collects the
same six cases at the same cost. It depends on something having sanitized from below the stale
word since it was written, clears per turn and not per checkpoint, and covers the embedder's own
frames as well. The two do not exclude each other.

Options::clearStackForMicrotaskCheckpoint (default on) is for comparisons.

Tests: JSTests/modules/microtask-checkpoint-clears-its-stack.js (an async function, then the
module) and JSTests/stress/microtask-checkpoint-clears-its-stack.js (a promise reaction, then an
async generator, each followed by an async function). Both fail in the shell of c281568 and with
--clearStackForMicrotaskCheckpoint=0, and pass in the 28 configurations that
run-javascriptcore-tests runs them in (release, x86_64).
robobun added a commit to oven-sh/WebKit that referenced this pull request Sep 16, 2026
…ng to occupy

A microtask checkpoint is where a turn of the embedder's event loop resumes an async function: a
timer or an I/O callback settles a promise, and the queue is drained. Every turn does that from the
same stack depth, so the frames of the checkpoint (MicrotaskQueue::drain*(), runInternalMicrotask(),
the job's, the entry to JS) are at the same addresses in every turn. They stay for as long as the
turn's JS runs, and they lie over whatever ran at that depth since the last checkpoint: that
checkpoint's own frames, or the embedder's. What they do not write is in reach of the conservative
scan of every collection made from under them, and sanitizeStackForVM(), which such a collection
calls, clears only what is below the frame that calls it.

#673 took the words that ASan poisons (redzones, locals out of scope) out of the scan. This is about
the words it cannot take out: spill slots and locals of a live frame that this job's path through
the function does not write. Builds without ASan have only those, and ASan builds have them too.
runInternalMicrotask() is one switch over every kind of job, with callMicrotask() inlined into
several arms, and a job uses one arm.

Seen as: what a finished async function held is not collected by collections made from later turns.

    async function makeGarbage() { /* two objects in a list, one await per object */ }
    await makeGarbage();
    for (;;) { await turn(); fullGC(); }    // turn(): a promise that a timer resolves

jsc shell of c281568 (linux amd64, lto): 32 turns until the objects are collected (the loop
tiers up and the frames change), 60 of 60 turns not collected with --useJIT=0. Bun at that WebKit
(release, x86_64): not collected in 8 of 8 turns when the timer is Bun.sleep(1); Bun's debug ASan
build, which has #673: not collected in 8 of 8 turns for the module shape with any timer. A heap
snapshot has the objects held by `list`, held by the lexical environment of the finished
function's JSAsyncFunctionGenerator, which has no incoming edge and no root. In a debugger, the one
word of the scanned span that holds the generator's address is at rbp-64 of
runInternalMicrotask()'s frame (272 bytes), under drainWithUseCallOnEachMicrotask() (624 bytes).
The job that runs in the later turns (AsyncModuleExecutionResume for the module,
AsyncFunctionResume when the loop is in an async function of its own) does not write that slot.
Which shape shows it moves with the compiler: the same script with setTimeout() kept the objects
with clang 21 and does not with clang 23. Earlier sightings of this frame in Bun's leak tests:
oven-sh/bun#37853 (the same generator word), oven-sh/bun#41607 (darwin arm64, release).

So a checkpoint that has jobs to run first clears MicrotaskQueue::stackBytesClearedForCheckpoint
bytes of the stack below performMicrotaskCheckpoint(), before the first of those frames exists:
2 KB in a release build (from there to the JS frame is 1.5 KB for x86_64), 32 KB with assertions or
ASan (12 KB with both). It does that with a function that is never inlined and whose frame is an
array of that size, which it zeroes with memset(): that frame lies exactly where the caller's next
callee has its own, on every ABI, and the compiler probes the stack for it where a platform needs
that. The array's address goes through an empty asm statement. Without it the stores are dead to
the compiler (zeroBytes() and secureZeroBytes() both compile to a bare `ret` here: the memory
clobber of secureZeroSpan() does not keep stores to a local whose address does not escape).
callMicrotask() asserts that it runs within seven eighths of the window below the checkpoint's
MicrotaskCallCache, which is a local of drainImpl().

Not sanitizeStackForVM() at the checkpoint: it clears from where it was last called, and nothing
need have called it from below the checkpoint since. The jsc shell does not change with it
(JSLock::didAcquireLock() resets VM::m_lastStackTop for every task). Not sanitizeStackForVMImpl()
with m_lastStackTop lowered to the bottom of the window, which was the first version: its loop for
x86_64 stores 8 bytes per iteration, and a checkpoint with one job went from 106 ns to 195 ns.

Cost, Xeon 8375C, clang 23: `promise.then(noop); drainMicrotasks()` 5 M times in the jsc shell,
101 ns per iteration without the clear and 122 ns with it. Bun, 1 M turns of
`await new Promise(r => setImmediate(r))`: 2531 ms without and 2547 ms with (medians of 11, same
binary, the option). 20 M awaits in one checkpoint: 602 ms and 602 ms. An empty checkpoint clears
nothing.

Once per checkpoint and not once per job, which would cost every job that much: what one job
leaves is still there for the later jobs of the same checkpoint, and is gone at the next one. Not
when the window would reach below the soft stack limit.

An embedder that holds the API lock for the life of its thread can get much of this from a
sanitizeStackForVM() per event loop turn instead (oven-sh/bun#37853): in Bun that collects the
same six cases at the same cost. It depends on something having sanitized from below the stale
word since it was written, clears per turn and not per checkpoint, and covers the embedder's own
frames as well. The two do not exclude each other.

Options::clearStackForMicrotaskCheckpoint (default on) is for comparisons.

Tests: JSTests/modules/microtask-checkpoint-clears-its-stack.js (an async function, then the
module) and JSTests/stress/microtask-checkpoint-clears-its-stack.js (a promise reaction, then an
async generator, each followed by an async function). Both fail in the shell of c281568 and with
--clearStackForMicrotaskCheckpoint=0, and pass in the 28 configurations that
run-javascriptcore-tests runs them in (release, x86_64).

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant