Repository navigation
Conversation
|
Updated 6:05 PM PT - Aug 12th, 2026
❌ @robobun, your commit c46dde9 has 3 failures in
🧪 To try this PR locally: bunx bun-pr 33018That installs a local version of the PR into your bun-33018 --bun |
|
Found 3 issues this PR may fix:
🤖 Generated with Claude Code |
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
WalkthroughAdjusts libuv poll teardown and nested loop handling, adds a Windows libuv poll patch and wires it into dependency setup, hardens dependency patch application, and adds a subprocess regression test for socket close during event-loop re-entry. Changeslibuv poll lifecycle and re-entrancy fix
Suggested reviewers
🚥 Pre-merge checks | ✅ 4✅ Passed checks (4 passed)
Comment |
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@packages/bun-usockets/src/eventing/libuv.c`:
- Around line 219-226: In us_create_poll, the newly allocated p and p->uv_p are
used without checking for allocation failure, so add OOM handling before
dereferencing either pointer. Use the existing native OOM path in this
function’s flow to validate the malloc/calloc results and return or abort
consistently before assigning p->uv_p->data, keeping the fix localized to
us_create_poll and its allocation setup.
In `@test/js/bun/net/socket.test.ts`:
- Around line 1655-1657: The nested HTMLRewriter test only checks that it
doesn’t crash, so it doesn’t prove the re-entry path actually ran. Update the
HTMLRewriter::transform() repro in socket.test.ts to make the nested tick
observable by flipping a reentered flag inside the setImmediate callback and
asserting it immediately after transform() returns, and remove the setTimeout(r,
5) fallback so the test stays deterministic. Keep the focus on the nested
HTMLRewriter re-entry case rather than relying on sock.terminate() alone.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: c800994f-8229-44c2-8464-7e07a1f458bc
📒 Files selected for processing (3)
packages/bun-usockets/src/eventing/libuv.cpackages/bun-usockets/src/internal/loop_data.htest/js/bun/net/socket.test.ts
|
This PR may be a duplicate of:
🤖 Generated with Claude Code |
|
Good flag. #30028 (@Jarred-Sumner) did not surface in my pre-open duplicate search (I had scoped it to robobun-authored PRs). I've now read it and tested the relationship both ways on a Windows debug build. Summary: #30028 is one of this PR's two changes, this PR is a strict superset, and neither makes the other redundant. The overlap. Both PRs move Not a duplicate: #30028 alone does not fix the heap corruption. Verified, not just traced. I applied only #30028's Mechanism: the nested tick's The converse holds. #30028's regression test, run verbatim against this branch: so this PR covers #30028's symptom too. Remaining delta over #30028's @Jarred-Sumner this is your call on how the two should land:
I did not want to step on an open PR without flagging the overlap. |
|
Evaluated all three. I'm not adding any #22444 (crash in #32585 (Windows segfault in a long-running HTTP proxy, fault address #27692 (Windows BSOD, If a maintainer wants #22444 linked on the strength of the signature match, no objection from me, but I would rather not auto-close an issue I cannot prove. |
|
Converting to draft: CI on this branch has a Windows regression I introduced, and the verification section in the description is wrong until it's fixed. Six node-compat tests time out on all three Windows lanes ( None of them is in I'm reproducing @Jarred-Sumner heads up since I pinged you above: please hold off on the #30028 / #33018 question until this is resolved. |
84b2342 to
eb014e8
Compare
|
Un-drafting. CI was right and this PR's first revision was wrong. The correction changes the diff substantially and also explains #30028's two-month-old CI failures, so here is the complete story. What CI caught. Six node http tests ( Root cause of the hang. The first revision moved The fix now. A correction to my earlier comments on this PR. I wrote that this PR is a strict superset of #30028 and that #30028's regression test passes on this branch. Both were true of the first revision and are no longer true: the (The earlier review note about One build-system finding, for whoever owns |
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@scripts/build/deps/libuv.ts`:
- Around line 51-64: The patch application flow in applyPatch is allowing a
skipped git apply to look successful, which can leave vendor/libuv unpatched
while still writing the fetch stamp. Update applyPatch in
scripts/build/fetch-cli.ts to detect when git apply skips or applies nothing and
treat that as a hard failure before the .ref file is written. Use the
applyPatch/fetch stamping path to locate the fix and ensure malformed patch
formats are rejected instead of being marked done.
In `@test/js/bun/net/socket.test.ts`:
- Around line 1685-1687: The subprocess regression assertion in socket.test.ts
is dropping native crash diagnostics by discarding stderr; update the combined
expectation around the Promise.all result so stderr remains part of the asserted
object while staying unconstrained (for example using an any-string matcher),
and remove the separate stderr discard so failures still surface diagnostics
from proc.stderr.text().
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: 5c5e6eab-0499-46fe-a4b8-99a5021d2b5f
📒 Files selected for processing (5)
packages/bun-usockets/src/eventing/libuv.cpackages/bun-usockets/src/internal/loop_data.hpatches/libuv/win-poll-no-reendgame-after-close.patchscripts/build/deps/libuv.tstest/js/bun/net/socket.test.ts
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@scripts/build/fetch-cli.ts`:
- Around line 245-254: Condense the explanatory block comment in fetch-cli.ts to
fit the repository’s 3-line limit while preserving the core rationale. Keep only
the essential points about GIT_CEILING_DIRECTORIES preventing skipped git-format
patches from being treated as successfully applied and the fail-loud backstop
using the skipped-patch check. Use the existing comment block near the
patch-application logic and shorten the wording without changing behavior.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: 90936f46-7f94-4661-84c0-3cfa451f43cb
📒 Files selected for processing (3)
scripts/build/deps/libuv.tsscripts/build/fetch-cli.tstest/js/bun/net/socket.test.ts
💤 Files with no reviewable changes (1)
- scripts/build/deps/libuv.ts
|
CI status, for whoever reviews this: the red check is infrastructure, not the diff. Across the last two builds on this branch (66519 and 66525), the only hard-failed job each time is It never ran a single test either time, and both failures were on the same agent ( For the lanes this PR can actually affect, both builds are clean: no failures in, and no annotation mentions of, |
|
One correction to my comment above. I wrote that no annotation mentions It is in the flaky (retried and passed) bucket, not a failure, and it is not my test. Two reasons it cannot be this PR:
The hard-failure picture is unchanged from my comment above: 280 jobs passed, and the only 2 failures are both |
|
The same failure is back on build 66564 (the push for the review nit above), so here is the consolidated CI picture across this branch's last three builds, for whoever can unblock it. Every lane that can exercise this change is green. On 66525, all 24 Windows The only failure is one Buildkite host. Across 66519, 66525, and now 66564, every hard-failed job is a That host cannot reach the artifact store. Re-pushing just re-dispatches onto it (66564 hit it again within minutes of my last commit), so I am going to stop pushing. My Buildkite token is read-scoped (the single-job retry endpoint returns 403), so I cannot retry it from here either. The red What this needs: someone with Buildkite access to retry this job, or to read the darwin failure as the infrastructure noise it is. The Windows fail-before and pass-after for the actual change are in the description. |
|
Build 66564 finished: 284 passed, 2 failed. Both failures are noise, but one of them disproves a claim in my previous comment, so I am correcting it here and explaining both. I was wrong that a re-push can only land on the sick agent. There are two The other failure is They cannot be caused by this diff, and that is a preprocessor fact rather than a judgment call. #if defined(_WIN32)
#define LIBUS_USE_LIBUV
#elif defined(__APPLE__) || defined(__FreeBSD__)
#define LIBUS_USE_KQUEUE
Since a re-push is not futile after all, I have pushed a single empty |
|
The retrigger (build 66646) went straight back to the sick agent. That is structural, not a run of bad luck. As long as that agent is online in the What this needs, in order of preference:
|
|
Build 66646 settled with a second failure beyond the darwin agent, and this one is on a Windows lane, so it deserves real scrutiny rather than a hand-wave: this is a Windows PR. The job is
So 66646's two failures are one persistently sick darwin build agent and one known flaky Windows install test that the retry happened not to save on one of three lanes. My one retrigger is spent and I am not pushing again. Everything in my previous comment stands: retry the darwin job, or merge on the evidence. |
fb8e250 to
c5c4804
Compare
|
Rebased onto The conflict was a pure append-vs-append in Since #31859 rewrote the Rust socket dispatch layer, I re-checked that the test's mechanism is intact: Full-file run on the rebased branch: |
|
Post-rebase CI picture (build 70743, finished: 285 passed, 1 failed; the two The darwin agent blocker is gone. Both The one failure is the current repo-wide Windows flake, not this PR. This PR's own surface is clean. 23 of 24 Windows I have already pushed this conversation's one retrigger (on the pre-rebase builds), so I will not push again for the napi flake. Anyone with Buildkite access can retry the one job; I cannot (read-scoped token). |
c5c4804 to
9d58b55
Compare
|
Rebased again for the same append-vs-append conflict in |
|
Build 71517 finished: 268 passed, 18 failed. All 18 are accounted for by three tests that are broken on main and present on other branches; none of the three is related to this PR, and none of them mentions
16 + 1 + 1 = the 18 failed shards. The other-branch hits are a TLS ALPN change, an mTLS close-notify fix, a blob deserialize fix, an ASan overflow sweep, a worker-threads ELU change, an RSA JWK validation fix, and so on: none related to each other, to uSockets, or to this PR. I am not going to fold an HTTP-client redirect fix or a darwin TLS-proxy stress-test fix into a uSockets |
|
Build 72796 finished: 283 passed, 3 failed; all four
The |
|
Independent confirmation of this diagnosis from a different symptom, plus a rebase onto current main.
Repro numbers on a Windows Server 2019 x64 box, running that bake test's child fixture in a loop:
This PR no longer applies to main ( Overlap note: #37105 (opened yesterday) contains the |
|
A third independent sighting of the same bug, with a trigger shape that differs from the two already described here, for whoever reviews this. Symptom. Trigger shape. Unlike the Read against |
|
@robobun rebase conflicts |
2662a98 to
c46dde9
Compare
|
Rebased onto current main (
Rebuilt on the rebased tree: the test passes, and the full file goes from |
|
Build 93542 (the rebase push) is red because of a CI-wide outage, not the diff: three |
alii
left a comment
There was a problem hiding this comment.
The tick_depth bracket, the closed sentinel and the libuv patch hold up against current main, and the rebase you pushed this afternoon took care of the us_free and active_handles conflicts. Two things left, plus one note.
- us_loop_pump: nothing exercises the pump bracket; add an unref'd variant of the new test.
- The fetch-cli.ts hardening is unrelated to this fix and the body still says it was left for a separate PR; split it out or update the body.
- The c-ares uv_poll_t in src/runtime/dns_jsc/dns.rs has the same shape (on_dns_socket_state uv_closes it from inside its own poll_cb, on_close_uv frees it, and the microtask drain at the end of on_dns_poll_uv can re-enter the loop). A follow-up is fine, but say so in the body.
Windows CI on the rebase died in the github.com outage, so it still needs a green Windows lane.
…e stack (Windows) On Windows each c-ares socket is driven by a uv_poll_t embedded in a heap UvDnsPoll. c-ares closes a UDP socket as soon as its answer is processed, which happens inside on_dns_poll_uv, so uv_close() is issued from within the handle's own poll callback. The callback then drains microtasks; a promise reaction that spins the event loop (Bun.build() with a plugin whose setup() returns a pending promise, bun:test's .resolves, ...) runs a nested uv_run, whose endgame phase invoked on_close_uv and freed the UvDnsPoll. libuv's uv__fast_poll_process_poll_req frame for that handle was still suspended underneath and reads the handle again once the callback returns; with the stale bytes it re-queued the endgame, close_cb ran a second time and the struct was freed twice. A script that re-enters the loop from resolve4() reactions segfaults after about ten queries on the release build. Count the on_dns_poll_uv frames active for a handle. on_close_uv only frees the struct when none is active; otherwise it marks the handle and the outermost frame, once its microtask drain is over, hands the struct to the event loop's task queue, which is drained only after the libuv callback has returned. The common path (close callback delivered after the poll callback returned) is unchanged. The libuv side of the same scenario is the win-poll-no-reendgame-after- close patch from #33018 (identical copy): with the struct kept alive, libuv would otherwise re-queue the endgame of the already CLOSED handle from the post-callback check and invoke close_cb twice (uv__poll_endgame asserts on this in debug builds).
|
Two more CI sightings of this bug in Sightings. Both on 2026-08-13, on unrelated branches that had just merged main:
Reproduction. On a Windows Server 2019 box where With the Confirmation that this is the whole story for that test. I tried a variant of the fix that avoids the libuv patch: keep a per-poll dispatch depth in Not opening a separate PR; this one covers it. |
|
One more trigger shape for this bug, with a different crash site than the three already described here. Posting it because it is the smallest repro so far and because the nested tick enters through Symptom. Trigger. A test body resumes inside TLS socket A's Numbers (Windows Server 2019 x64, canary
I have not built this branch on Windows; the nested tick here goes through Repro (bun test file; keys from test/js/node/test/fixtures/keys)import { expect, test } from "bun:test";
import { once } from "node:events";
import { readFileSync } from "node:fs";
import https from "node:https";
import tls from "node:tls";
const keys = "test/js/node/test/fixtures/keys/";
async function peerCN(port: number, extra = {}) {
const socket = tls.connect({ host: "127.0.0.1", port, rejectUnauthorized: false, ...extra });
const errored = once(socket, "error");
await Promise.race([once(socket, "secureConnect"), errored.then(([e]) => Promise.reject(e))]);
const cn = socket.getPeerCertificate().subject?.CN;
socket.destroy(); // closes socket A from inside its own on_data
return cn;
}
test("rejects after a socket event", async () => {
const server = https.createServer({
key: readFileSync(keys + "agent1-key.pem", "utf8"),
cert: readFileSync(keys + "agent1-cert.pem", "utf8"),
minVersion: "TLSv1.3",
});
server.listen(0);
await once(server, "listening");
const port = (server.address() as any).port;
try {
expect(await peerCN(port)).toBe("agent1"); // resumes inside A's us_internal_ssl_on_data
await expect(peerCN(port, { maxVersion: "TLSv1.2" })).rejects.toThrow(); // nested us_loop_run ticks
} finally {
server.close();
}
}); |
|
One more data point for Verification of this PR's exact diff (the
That reduced fixture is on |
|
This also fixes the intermittent Windows failure of CI symptomThe child How the fixture gets there
An instrumented build confirmed the free: Verification of this branch (Windows x64, debug builds)
For what it is worth, I independently tried the variant that moves the |
|
Arrived at the same three changes independently from build 98537 (
|
|
This is also the cause of the top hard failure on the windows-x64 CI lane: How that fixture reaches this bug: the case before it awaits the WebSocket Independent reproduction on a windows-x64 box with a For reference, branch |
c46dde9 to
ea2a10c
Compare
|
All three addressed in ea2a10c (rebased onto current main at the same time, which picked up the libuv pin bump in #39472; the three-patch stack still applies at the new pin, re-verified against
Windows CI: the push will give it a fresh run; the outage build is superseded. One unrelated thing I hit while building the unref'd variant, filed separately as well: on Windows, if the only things pending are unref'd immediates queued from inside a uv callback and there is no I/O in flight, the process neither runs them nor exits, it blocks (no |
us_internal_loop_post defers us_internal_free_closed_sockets to the outermost loop tick via loop->data.tick_depth, because a nested tick (a poll callback re-entering the loop, e.g. waitForPromise) must not free a socket the suspended outer dispatch frame still reads after its callback returns. That counter was only maintained by the epoll/kqueue backend; on libuv (Windows) it stayed 0 forever, so every nested uv_run ran the sweep. The result is a deterministic heap corruption (STATUS_HEAP_CORRUPTION, 0xC0000374 on bun 1.3.14) when a socket's data callback closes it and then synchronously re-enters the event loop. Three pieces, all of which are no-ops for a non-nested close: - eventing/libuv.c: us_loop_run and us_loop_pump bracket uv_run with tick_depth++ and --, the direct mirror of us_loop_run and us_loop_run_bun_tick in epoll_kqueue.c. us_internal_loop_post runs from the uv_check_t registered in us_create_loop, so it fires inside uv_run and now sees the nesting depth. - eventing/libuv.c: deferring the sweep means the nested tick's uv__process_endgames can now run close_cb_free_poll before us_poll_free has re-pointed uv_p->data at the us_poll_t. That branch used to silently do nothing, which would leak both allocations: the later us_poll_free sees uv_is_closing() still true (uv_is_closing reports CLOSED as well as CLOSING) and arms a callback that already fired. close_cb_free_poll now marks the handle (h->data = h) and us_poll_free frees both when it sees the mark. - patches/libuv/win-poll-no-reendgame-after-close.patch: the post-poll_cb tail of uv__fast_poll_process_poll_req (and its slow-poll sibling) re-queues an endgame for any CLOSING handle with no outstanding AFD requests. When a nested uv_run already ran that handle's endgame, and so cleared UV_HANDLE_ENDGAME_QUEUED and set UV_HANDLE_CLOSED, the re-queue runs uv__poll_endgame a second time, which it asserts against, and invokes close_cb_free_poll twice. Guard the check on !(flags & UV_HANDLE_CLOSED). us_poll_stop, us_create_poll, and the common path of us_poll_free are unchanged from main. In particular, us_poll_stop keeps issuing the uv_close: that uv_close is what keeps uv__loop_alive() true for a closed socket (uv__handle_closing calls uv__active_handle_add even for an unref'd handle), which is what makes uv_run execute the next iteration, which is the only thing that runs the sweep. An earlier revision of this branch moved the uv_close into us_poll_free; that cycle hung six node http tests on every Windows CI lane, because a socket closed from the timer phase (after the check phase) dropped active_handles to zero and the very next uv_run returned without running a single phase. PR 30028 makes the same move for a different reason and its CI shows the same failures. Repro: a Bun.listen data handler that calls sock.terminate() and then synchronously re-enters the loop (an HTMLRewriter element handler that returns a pending promise) crashes bun 1.3.14 on Windows with STATUS_HEAP_CORRUPTION on the first hit. Either half alone does not crash. Not reproducible on Linux or macOS, where the epoll/kqueue backend already maintains tick_depth.
ea2a10c to
7f1dd1a
Compare
|
This PR has been closed because it was flagged as AI slop. Many AI-generated PRs are fine, but this one was identified as having one or more of the following issues:
If you believe this was done in error, please leave a comment explaining why. |
There was a problem hiding this comment.
I re-reviewed after 7f1dd1a and found no issues — the stdout && JSON.parse(stdout) guard is in place, so a crashed child now reaches the toEqual diff. Given this is a memory-safety fix in the Windows event-loop backend plus a vendored libuv patch, and a human reviewer's earlier change requests were just addressed today, a human re-review is still warranted before merge.
What was reviewed:
tick_depthbrackets inus_loop_run/us_loop_pumpmirror the epoll/kqueue backend's placement; both are the onlyuv_runentry points on the uSockets loop.close_cb_free_poll/us_poll_freeself-marker handshake — the newh->data = hwrite only fires in the previously-no-opdata == 0branch, and the only reader is the newus_poll_freecase, so no regression for the existing stop-then-free ordering or theus_poll_start_rcinit-failure path.- The libuv patch adds a
UV_HANDLE_CLOSEDguard to both fast/slow poll siblings; it only narrows whenuv__want_endgameis queued. - Test: both ref'd/unref'd variants drain stdout/stderr/exited concurrently, wire every failure event to
resolve, and assert the observedactiveHandleson Windows only.
Extended reasoning...
Overview
This PR fixes a Windows-only heap corruption (STATUS_HEAP_CORRUPTION, exit 0xC0000374) triggered when a socket's data callback closes the socket and then synchronously re-enters the event loop (e.g. via waitForPromise, expect().rejects, or HTMLRewriter.transform). It touches five files: packages/bun-usockets/src/eventing/libuv.c (tick_depth brackets + close_cb marker handshake), packages/bun-usockets/src/internal/loop_data.h (comment update), a new patches/libuv/win-poll-no-reendgame-after-close.patch, scripts/build/deps/libuv.ts (registers the patch), and a two-variant regression test in test/js/bun/net/socket.test.ts.
Security risks
None identified. This is a use-after-free / double-free fix in event-loop infrastructure; it does not touch auth, crypto, permissions, or any user-input parsing surface. The libuv patch strictly narrows an existing condition (adds !(flags & UV_HANDLE_CLOSED)).
Level of scrutiny
High. This is exactly the category REVIEW.md calls out as most-blocked: native memory safety, handle lifecycle, and re-entrancy in the event loop. The fix has three interacting parts (tick_depth deferral, a self-marker handshake between close_cb_free_poll and us_poll_free, and a vendored-libuv patch), and the PR description honestly notes that with the libuv patch and marker present, deleting either tick_depth bracket degrades to stale reads that a non-ASAN Windows build cannot observe — so the brackets' correctness rests on the fail-before against the fully-unfixed tree plus reasoning by symmetry with the epoll/kqueue backend, not on per-bracket fail-on-delete. A previous revision of this PR broke six Windows node-http tests by moving uv_close to the wrong place, which confirms the area is delicate. The change is also entirely inside #ifdef LIBUS_USE_LIBUV, so POSIX binaries are byte-identical.
Other factors
- A human reviewer left three substantive change requests on 2026-08-13 (pump-path test coverage, splitting out unrelated
fetch-cli.tshardening, and thedns.rssibling). All three were addressed in ea2a10c on 2026-08-19, but that reviewer has not yet re-reviewed. - My previous run's one finding (the
JSON.parse("")failure-diagnostic issue) was addressed in 7f1dd1a and is confirmed present in the current diff. - The PR has extensive independent corroboration in the timeline: at least five separate CI sightings (
test/bake/deinitialization.test.ts, postgres listen/notify, node-https-server-context) root-caused to the same three defects, plus three independent branches that arrived at the same fix. - The test wires every failure event (
close/end/error/connectErrorand the connect promise rejection) to the awaited resolver, drains all three subprocess streams concurrently, and assertsactiveHandleson Windows to prove which entry point dispatched — the pump variant addresses the reviewer's coverage request.
Given the memory-safety complexity, the vendored-dependency patch, and the pending human re-review of freshly-addressed change requests, this is not a candidate for automated approval.
…libuv are done with it (#39643) ### Problem - Windows crashes after a handler returns: `us_socket_is_closed` or `us_internal_socket_follow_adopted` at `0xFFFFFFFFFFFFFFFF` (BUN-43AH, BUN-4NC8), `us_internal_socket_close_raw` on a freed socket (BUN-442H, BUN-442Y, BUN-4NP3), a garbage `poll_cb` called from `uv__fast_poll_process_poll_req` (BUN-442Z), mimalloc free-list crashes, and the `test/bake/deinitialization.test.ts` failures on the Windows lanes. The 1.4.0 release (`34cbb9a40`) reports this family from `bun test` on Windows (as of 2026-08-21: BUN-442H 12 events, BUN-4NC8 6, BUN-442Y 5, BUN-4NP3 2, BUN-442Z 1) and the TLS forms BUN-4MNY, BUN-4NFZ, BUN-4NJG, BUN-4NKV and BUN-4QEY, one each. BUN-4QEV, BUN-4QHN and BUN-4QQ6 are the reused-handle arm described in the notes. - `loop.c:450` frees closed sockets only at `tick_depth <= 1`. The libuv backend never counts it, so a handler that waits for a promise (a nested `uv_run`) frees the socket the outer dispatch still reads (`loop.c:746`). - The nested run also finishes closing that socket's `uv_poll_t` under libuv's outer `uv__fast_poll_process_poll_req` frame. That frame then queues the endgame again, and `close_cb_free_poll` runs twice. `uv_run` is documented as not reentrant, so the misuse is ours. ### Fix - `us_loop_run` and `us_loop_pump` count `tick_depth`, as the POSIX backend does. - While a `poll_cb` frame is on the stack (`poll_cb_depth`), `us_poll_stop` only disarms the handle. The outermost frame calls `uv_close` on its way back into libuv, which libuv supports, and then closes the socket (`us_internal_poll_close_fd`), so libuv sees the same order as before. libuv is unchanged. - `us_poll_free` and `close_cb_free_poll` record who ran first (`released`, `uv_closed`). The second one frees both blocks. The old `data = NULL` handshake leaked when the callback ran first (#37105). A stopped poll is never re-armed or re-initialized. - Verified: the new case in `test/js/bun/net/socket.test.ts`. On a Windows debug build of main its child segfaults at `0xFFFFFFFFFFFFFFFF` (the BUN-43AH signature), and the release canary crashes too. It passes with the fix (Windows x64 and arm64 debug, Linux ASAN). `appcontainer.test.ts` gets a step for the socket order. Notes have the rest. ### Background - On Windows a socket is two blocks: the `us_socket_t` (it starts with the `us_poll_t`) and a libuv `uv_poll_t`. libuv's in-flight AFD requests live inside the `uv_poll_t`, so it has to live until the close callback. A closed socket waits on `loop->data.closed_head` until `us_internal_loop_post` frees it, and `loop.c` reads it after its handlers return. - Endgame: `uv_close` cancels the requests. When the last one completes, libuv queues the endgame, which unlinks the handle and runs the close callback. `uv__fast_poll_process_poll_req` checks for it right after `poll_cb` returns. - Nested tick: `wait_for_promise` (`expect().resolves`, auto-install) runs `uv_run` inside the handler. <details><summary>Notes</summary> **Symbolized fail-before** (Windows x64 debug build of main `a35696478d`, a standalone copy of the new test's fixture. The test itself fails on that build with the same fault address): ``` panic(main thread): Segmentation fault at address 0xFFFFFFFFFFFFFFFF us_internal_socket_follow_adopted packages/bun-usockets/src/internal/internal.h:357 us_internal_dispatch_ready_poll packages/bun-usockets/src/loop.c:746 (after us_dispatch_data returned) poll_cb packages/bun-usockets/src/eventing/libuv.c:165 uv__fast_poll_process_poll_req vendor/libuv/src/win/poll.c:233 uv__process_reqs / uv_run vendor/libuv/src/win/core.c us_loop_run packages/bun-usockets/src/eventing/libuv.c:417 ``` The freed socket is filled with mimalloc's debug poison, so `flags.adopted` reads as set and `prev` is followed. The `open()` variant crashes the same way at `loop.c:556`. The same fixture without the nested wait passes on that build. The release canary (`1.4.0-canary.1+32e87032b`) passes the first case and crashes in the second, at `0xFFFFFFFFFFFFFFFF` in one run and `0x1AF0000002A` in another, which is the heap corruption from the double free. **How the three Sentry shapes follow.** In release, `mi_free` overwrites the first word of the freed `uv_poll_t`, which is `data`. The rest stays intact, so the outer frame sees `events == 0`, `CLOSING`, and no requests in flight, and queues the endgame again. The second `close_cb_free_poll` frees `h->data`, now the free-list link to another freed block, and `h` itself again. Later allocations alias (a garbage `poll_cb`, BUN-442Z, or `us_poll_start_rc` faulting at 0 in `deinitialization.test.ts`), or the freed socket is reused and the outer dispatch runs the error close on it (BUN-442Y), or `follow_adopted` reads a reused block (BUN-43AH). **Socket order, found by CI.** The first push closed the socket in `close_raw` before the deferred `uv_close`. `appcontainer.test.ts` failed on the Windows 11 arm64 lane with exit `0xC0000008` (`STATUS_INVALID_HANDLE`), and failed 10 of 10 runs on an arm64 machine with that build against 5 of 5 passes with main's `libuv.c` on the same machine. `GetProcessMitigationPolicy(ProcessStrictHandleCheckPolicy)` inside the container reports `0x3`, so a call on a closed handle raises instead of failing, and `uv__poll_close` cancels the in-flight request with an ioctl on the socket. A probe with four steps (listener only, serve + fetch, terminate from `data()`, terminate from a timer with the peer closed by its own dispatch) died in every step that closes a socket from a dispatch. With `us_internal_poll_close_fd` the four steps and the test pass (10 of 10 runs on arm64, and on x64). The test now also closes a socket from its own `data()` handler, which is the shortest path to this order. Outside a nested tick, the later socket close is not observable from JS: the loop does not run again before the handler's dispatch ends. Inside one, a handler that closes its socket and then waits for the peer to notice now waits until the handler returns. That wait was already unreliable on Windows (the peer's completion is often in the outer run's batch) and then corrupted the heap. **The reused-handle arm.** Between the inner run's endgame and the outer frame's return, `us_create_poll` can hand the freed `uv_poll_t` block to a new poll (same size class, most recent free). Then the second `close_cb_free_poll` runs on a live handle: `h->data` is the new poll's `us_socket_t`, so that socket is freed with no close path, and the new `uv_poll_t` is freed under its pending request. The JS wrapper of that socket still reports open. BUN-4QHN (`ws.send()`: `us_socket_group_ext` on the freed socket's `group`), BUN-4QQ6 (`socket.readyState`: `SSL_get_shutdown` on its `ssl`) and BUN-4QEV (`uv__process_poll_req` reading the freed handle for a completed request, `poll.c:574`) are that arm. With one `uv_close` per handle, issued after its last callback frame, the second callback no longer exists. **Why the hand-off is needed together with `tick_depth`.** With `tick_depth` alone, every socket closed during a nested tick with no live frame finishes closing in the inner run, and the deferred `us_poll_free` then met a handle whose close callback had already run: the old code handed it `data` and leaked both blocks. #37105 fixes that case on its own with a marker in `data`. This PR covers it with the two bits. **Defensive branches with no current caller:** `us_poll_free` on a poll that was never stopped, `us_poll_start_rc` on a registered poll (mask change, or `UV_EBADF` once stopped), `us_poll_change` after stop. Every current caller creates, starts, and later stops a poll exactly once. The `uv_poll_init_socket` failure path now goes through `us_poll_stop`. `uv_poll_stop` on a half-initialized handle only clears `events` (checked against `uv__poll_set`). That path was reasoned about, not run: it needs a `--socketFaultInjection=on` build, which rebuilds every Rust crate. **Suites run on the Windows x64 debug build with the fix**, all green: `test/js/bun/net/socket.test.ts` (86 pass), `tcp-server`, `socket-retention`, `socket-syscall-fault`, `test/js/bun/udp/udp_socket.test.ts`, `test/js/bun/websocket/websocket-server.test.ts`, `test/js/node/net/node-net-server.test.ts`, `test/js/node/tls/node-tls-server.test.ts`, `test/js/node/http/node-http.test.ts`, `test/js/bun/http/serve.test.ts` (294 pass), `test/bake/deinitialization.test.ts` 3 of 3 (the unfixed release canary hangs for 60 s and is killed on the same machine). `test/js/web/fetch/fetch.test.ts` had 14 timeouts of `(with gc)` body tests on this slow debug box. The two socket tests among them pass in isolation. Per the analysis in #33018, `deinitialization.test.ts` closes the socket from a nested `poll_cb` on the same handle, so it also covers `poll_cb_depth > 1`. **Sentry groups on the 1.4.0 release** (`34cbb9a40`, all Windows): BUN-4NC8 is `follow_adopted` (`internal.h:357`) under `loop.c:746`, the frames of the fail-before above. BUN-442H is `close_raw` at `socket.c:292`, the low-prio unlink writing through `s->prev` of a reused block, reached from the eof or error close at the end of the dispatch (BUN-442Y and BUN-4NP3 are the same call: the unlink at `context.c:235` and `:240` writing through the reused block's `prev` and `next`, and `us_internal_disable_sweep_timer` through its `group`). BUN-4MNY is `ssl_retry_parked_write` (`openssl.c:2103`, `s->group` read as NULL) in the tail of `us_internal_ssl_on_data` after `us_dispatch_data` returned: the TLS form of the same read of a freed socket. BUN-4NFZ is the other exit of that tail: `us_internal_ssl_close` (`openssl.c:1927`, again `s->group` NULL) from the `ssl_close` call at `openssl.c:2324` after the data callback. BUN-4NJG and BUN-4NKV are the keylog and session flushes in the same tail (`openssl.c:2430` to `:351`, and `:2429` to `:426`): `s->ssl` read from the reused block, so `SSL_get_ex_data` returns garbage. Both flushes are guarded by `!s->ssl || is_closed`, which only a reused block gets past. The four TLS groups are the four helpers of that tail, in whichever order a given reused block fails. With the closed list left alone until the outermost tick, `ssl_gone`, `is_closed` and `s->ssl` read the real state there and those tails return. **Limit.** `tick_depth` counts the ticks Bun itself runs (`us_loop_run`, `us_loop_pump`). A native addon that calls `uv_run()` on this loop from inside a handler is not counted. The socket whose own callback is on the stack is still safe there, because its `uv_close` and frees are tied to `poll_cb_depth`, but a socket that another frame holds (an accepted socket closed from its own `open()` while the accept loop still reads it, or a block retired by adoption) is not. Counting the frames of the callbacks libuv runs for us (`poll_cb`, `prepare_cb`, `timer_cb`, `async_cb`) would cover that entry as well. None of the crashes above needed it, so it is left for a follow-up rather than re-verifying this PR for it. **Release-build numbers** for `test/bake/deinitialization.test.ts` are in the comments below: on one Windows Server 2019 machine the fixture fails 8 of 8 runs on main (6 hangs, 2 segfaults) and passes 20 of 20 with this diff. **Related.** #33018 was an earlier attempt at this bug. It patched libuv's post-callback check instead of moving the `uv_close`. #38024 has the same bug class in the c-ares poll and still carries that patch. Deferring its `uv_close` past the outermost callback frame would remove the need for it there too. The test fixture waits on a timer on purpose: a nested run cannot wait for an event of the peer socket, because that completion may sit in the outer run's IOCP batch. </details> <!-- robobun:evidence:begin --> --- **no test proof** · iteration 4 · platform-specific test(s) that do not run on this machine, deferring to CI, which covers all platforms: test/js/bun/windows/appcontainer.test.ts, test/js/bun/net/socket.test.ts <!-- robobun:evidence:end --> --------- Co-authored-by: Ciro Spaciari <ciro.spaciari@gmail.com>
This PR has been marked as AI slop and the description has been updated to avoid confusion or misleading reviewers.
Many AI PRs are fine, but sometimes they submit a PR too early, fail to test if the problem is real, fail to reproduce the problem, or fail to test that the problem is fixed. If you think this PR is not AI slop, please leave a comment.