Conversation
…ntion The upgradeTLS test counted live TLSSocket cells after a GC loop. JSC scans the machine stack conservatively, so stale words that NewSocket::on_data left in stack memory kept the last [raw, tls] pair alive for as long as the live JSC::runInternalMicrotask frame that resumes the loop covered them. On some ASAN binaries the count stayed at 3 for the whole loop. Assert the retention model directly. The protected object counts show the Strong move from the TCP wrapper to the two TLS wrappers on upgrade and its release on close. A debugging heap snapshot shows that no recorded GC root reaches a wrapper after close. Neither depends on the conservative scan. The test now also runs on Windows.
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Essentials Run ID: 📒 Files selected for processing (1)
Included review availability: Your plan provides up to 10 included reviews per hour; 2 remain after this review. WalkthroughThe ChangesTLS socket retention validation
Priority: ⬇️ Low Merge Risk: ⚪ Minimal · up to The updated regression test has no supported merge-blocking issue. 🚥 Pre-merge checks | ✅ 4✅ Passed checks (4 passed)
Comment |
|
Status: ready for a maintainer to merge. This is a test-only change, and it is green on every lane. Build 116387 ran The build is red for one test that this diff does not touch: How I reproduced the red test:
A release-asan build of main at the merge base (b841a68) keeps one wrapper alive the same way and passes the old test. The heap snapshot and the gdb stack search that show what holds the survivors are in the Notes of the PR body. |
A heap snapshot of the test runner's own heap takes seconds in a debug build, and the test needed a timeout for it. In a fresh process the heap is small: the test takes about 2 s in a debug build and 60 ms in a release build, so the timeout is gone.
|
Updated 12:14 AM PT - Sep 16th, 2026
❌ @robobun, your commit fdf1aa0 has 1 failures in 🧪 To try this PR locally: bunx bun-pr 42877That installs a local version of the PR into your bun-42877 --bun |
Problem
test/js/bun/net/socket-retention.test.tsis red on x64-asan for some PR builds, on every retry (116201, 116220, 116272):Expected: <= 2,Received: 3atexpect(count).toBeLessThanOrEqual(baseline + 2).NewSocket::on_datapoint at the last[raw, tls]pair. The liveJSC::runInternalMicrotaskframe that resumes the test's GC loop covers them, so everyBun.gc(true)marks both wrappers.baseline + 2(Migrate TCPSocket/TLSSocket from hasPendingActivity to jsc.JSRef #29451) allowed the one wrapper that was always pinned. build: compile ASAN builds with -fsanitize-address-use-after-return=never #42775 changed the ASAN frame layout, and some binaries now pin two.Fix
heapStats().protectedObjectTypeCounts: the Strong moves on upgrade (TLSSocket +2, TCPSocket -1) and is gone after close. These counts are exact.bun bd, release, Windows x64 and aarch64.Background
TCPSocketandTLSSockethold their JS wrapper through aJsRef: a Strong while the socket is active, a weak pointer after close.upgradeTLSreturns two wrappers for one connection.generateHeapSnapshotForDebugging()runs a full GC and records its roots and edges, except conservative roots.Notes
Heap snapshot. Release-asan build of 71c3a0a (the commit of build 116272), after the five connections and two
Bun.gc(true):objectTypeCounts.TLSSocketis 3 andprotectedObjectTypeCounts.TLSSocketis 0.generateHeapSnapshotForDebugging()lists two TLSSocket wrappers, cells0x712f75c24180and0x712f75c24200(adjacent, the last pair). Neither has an incoming edge or arootsentry. The third cell is the prototype.Debugger. gdb on the same binary, breakpoint on
JSC::ConservativeRoots::add(void*, void*, JITStubRoutineSet&, CodeBlockSet&)during the nextBun.gc(true). The scanned span is the main thread stack, 42,328 bytes. A search of that span for the two cell addresses finds 6 words. I walked the frame-pointer chain by hand, so JIT frames do not stop it:The chain at that point is
auto_tick>drain_timers>__bun_fire_timer>EventLoop::exit>Zig::GlobalObject::drainMicrotasks>VM::drainMicrotasks>runInternalMicrotask>asyncFunctionGeneratorBodyCall> the test function >Bun.gc. This is the resume ofawait Bun.sleep(10)ingcUntilCountAtMost. Each pass of the loop runs through the same frames at the same addresses.With
bun bdthe same search finds 3 words for one wrapper, all inrunInternalMicrotask(a 6320-byte frame there). Hardware watchpoints on those 3 words show that the last writer of each isNewSocket<true>::on_data(+4037, +4045, +4510), called throughus_internal_ssl_on_data>us_dispatch_data. That frame is popped long before the GC loop starts.runInternalMicrotaskis one large switch. On the async-function resume path it does not write the slots of its other cases.Why it depends on the binary. Release-asan builds (
--profile=release-asan --ci=on),BUN_GC_TIMER_DISABLE=1, 60 passes ofBun.gc(true)+Bun.sleep(10)after the five connections:Expected: <= 2,Received: 3The stale words are always in
runInternalMicrotask. Which of them line up with a wrapper address depends on the frame sizes of the two call chains, and a change to any Rust code inbun_runtimemoves them. #42775 is correct. It made the ASAN frames static, and that moved the words to where a second wrapper can be pinned. The ASAN lane runs on PR builds only, so main's own builds never showed it.Why every retry fails in CI and a local run can pass. The survivors stay until something writes over those words. The GC controller's 1 s repeating timer does: its callback runs at the same stack depth. The old loop is 50 passes. In CI the whole test took 778 ms, so the loop ended before the timer fired. On my machine a pass takes about 20 ms, and in most runs the timer fires inside the loop (the count drops from 3 to 1 at pass 33 to 45). With
BUN_GC_TIMER_DISABLE=1orBUN_GC_TIMER_INTERVAL=60000the old test fails every time.Same frame as #41607. That PR found a stale
runInternalMicrotaskslot behind a fetch-body test on a darwin release build, so this is not specific to ASAN.Checks on the new assertions.
downgrade()removed inmark_inactive: the test fails withprotectedAfterClose: 5andwrappers: 5(expected 0 and 0).wrappers: 5. Same result on Windows.await donecan run inside the close dispatch, beforeCloseTeardowndowngrades the tls wrapper. The script waits onesetImmediateturn before it reads the counts.Subprocess. The first version of this PR ran the same assertions inside the test runner. On the red binary, where the two pinned wrappers exist, that version passed 10 of 10 with the CI environment and 10 of 10 with
BUN_GC_TIMER_DISABLE=1, so the snapshot walk does ignore them. The snapshot of the test runner's heap (about 12,000 cells) took 2.5 s in a debug build, and the test needed a timeout. A freshbun -eprocess has about 2,800 cells. In that process the red binary pins no wrapper at all (other frames are on the stack), which shows again that the old count depended on the stack layout and not on Bun's references.Windows. The test was skipped there because the residual count varied. With canary 1.4.3-canary.1+f5649a7ed the new test passes 20 of 20 on Windows Server 2019 x64 and 20 of 20 on Windows 11 aarch64. The whole file passes 3 of 3 on each.
Time. The test takes 1.8 to 2.2 s in a debug build, 0.3 s on release-asan and 60 ms on release, with no timeout of its own. The old loop took 0.8 to 0.9 s on release-asan.
Runs.
bun bd test test/js/bun/net/socket-retention.test.ts: 5 pass, 1 skip. Release-asan build of 71c3a0a with the CI environment (BUN_JSC_validateExceptionChecks=1,BUN_DESTRUCT_VM_ON_EXIT=1, LSAN options): 10 of 10. Same binary withBUN_GC_TIMER_DISABLE=1: 10 of 10. Release: 3 of 3.no test proof · iteration 0 · platform-specific test(s) that do not run on this machine, deferring to CI, which covers all platforms: test/js/bun/net/socket-retention.test.ts