Repository navigation
usockets: treat ENOBUFS/ENOMEM from send() as would-block, not fatal - #33803
Conversation
us_socket_write_check_error classified every non-EWOULDBLOCK send() errno as a fatal write error. For the transient resource-exhaustion errnos (ENOBUFS/ENOMEM on POSIX, WSAENOBUFS on Windows) the connection is still healthy and a later retry can succeed, so the node:net drain path (internal_flush in socket_body.rs) must keep the buffered bytes and re-arm the writable poll instead of dropping them. Dropping them fired 'drain' on a socket that had silently lost a contiguous span of the application's byte stream: the next write went straight to the kernel on the same live connection, so the peer received the stream with bytes missing from the middle, with no 'error' event, no failed write callback, and no close.
|
Warning Review limit reached
Next review available in: 31 minutes Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Pro Run ID: 📒 Files selected for processing (4)
Comment |
|
Updated 11:27 PM PT - Jul 8th, 2026
✅ @robobun, your commit cd450a6155289194324ca1aa55501492a7d062fd passed in 🧪 To try this PR locally: bunx bun-pr 33803That installs a local version of the PR into your bun-33803 --bun |
There was a problem hiding this comment.
No issues found, but this changes error classification in usockets' core TCP write path (which underpins all networking in Bun), so it's worth a human sanity check.
Extended reasoning...
Overview
Adds bsd_send_is_transient_error() (ENOBUFS/ENOMEM on POSIX, WSAENOBUFS on Windows) and treats it as would-block rather than fatal in us_socket_write_check_error(). Includes a fault-injection test in net-syscall-fault.test.ts covering both errnos. Touches packages/bun-usockets/src/{bsd.c,socket.c,internal/networking/bsd.h}.
Security risks
None identified — this widens the retry set for send() errnos, not the accept set for any input. No auth/crypto/parsing surface.
Level of scrutiny
Moderate-to-high. The diff is small and the reasoning is sound (matches libuv's uv__write retry set), and the new helper is deliberately kept separate from bsd_would_block() so recv-side callers are unaffected. But us_socket_write_check_error sits in the write path shared by every TCP connection (node:net, fetch, Bun.serve, WebSocket), and errno classification here decides whether buffered bytes are retried or silently dropped. A subtle mistake would manifest as rare, hard-to-reproduce data loss or hangs across the whole runtime, so a human should sign off.
Other factors
The bug-hunting system found no issues. The test uses the in-tree fault-injection harness and follows the existing patterns in the file. The only caller of us_socket_write_check_error is the node:net flush path via src/uws_sys/us_socket_t.rs, so the blast radius of the behavior change itself is contained, but the new helper lives in shared bsd.c and the reviewer should confirm the recv-side isolation reasoning holds.
…ted sockets too
SocketBody (and its buffered_data_for_node_net / internal_flush path) is
shared between connecting and net.createServer-accepted sockets, so the
fix in the previous commit already covers both. Parametrise the
drain-path transient-errno test over {client writer, server writer} so
the server surface (every raw-TCP response writer built on node:net) is
proven by the same fail-before/pass-after.
|
Pushed cd450a6: the drain-path test now runs against both a
The same sweep confirmed The complementary "report fatal to JS" half (honouring |
There was a problem hiding this comment.
No issues found — the fix looks correct and matches libuv's ENOBUFS handling, but this changes error classification in the core usockets write path (an area with prior Windows regressions per the socket_body.rs comments), so worth a human sign-off.
Extended reasoning...
Overview
Adds bsd_send_is_transient_error() (ENOBUFS/ENOMEM on POSIX, WSAENOBUFS on Windows) and ORs it with bsd_would_block() in us_socket_write_check_error() so transient kernel resource exhaustion is treated as would-block rather than fatal. Prevents internal_flush() in socket_body.rs from silently dropping buffered_data_for_node_net on a healthy connection. ~15 lines of production code across bsd.c/bsd.h/socket.c plus a fault-injection test covering {ENOBUFS, ENOMEM} × {client, server} writers.
Security risks
None identified. This narrows the set of errnos classified as fatal (moving two transient ones to the retry path); it does not relax any validation of untrusted input or expose new surface. EPIPE/ECONNRESET still reach the fatal branch.
Level of scrutiny
High. This is the core usockets TCP write error-classification path that every node:net / Bun.connect socket flows through, and misclassification here causes silent mid-stream data loss. The existing comment at src/runtime/socket/socket_body.rs:904-911 documents a prior regression (f0325bd) where changes to fatal-detection semantics broke Windows FIN-terminated responses, so this area has subtle cross-platform interactions that merit a maintainer's eye.
Other factors
The change is conservative and mirrors libuv's uv__write (EAGAIN || EWOULDBLOCK || ENOBUFS). The new helper is deliberately kept separate from bsd_would_block() so recv() callers are unaffected. Tests use the in-tree fault-injection harness and assert byte-exact delivery. The PR description notes tests were deferred to CI (didn't run locally), and the CI build was still in progress at review time. No CODEOWNERS match for packages/bun-usockets/.
|
CI: 284/284 jobs that ran passed on build #70819, including every Linux, Windows, darwin-x64 and darwin-26-aarch64 test lane. The two Ready for review. |
us_socket_write_check_errorclassified every non-EWOULDBLOCKsend()errno as a fatal write error. For transient kernel resource exhaustion (ENOBUFS/ENOMEMon POSIX,WSAENOBUFSon Windows) the connection is still healthy and a later retry can succeed, so thenode:netdrain path must keep the buffered bytes and re-arm the writable poll instead of dropping them.Repro
Deterministic via the in-tree usockets fault injection (debug/ASan builds):
Without the fix the peer receives
head[0]followed by whatever is written next; the 199 buffered bytes are gone, no'error'event, no failed write callback, no close. The short send that leaves data buffered is ordinary kernel backpressure on any real-network upload; a one-offENOBUFS/ENOMEMfromsend(2)under memory pressure is ordinary kernel behaviour.Cause
packages/bun-usockets/src/socket.cus_socket_write_check_error():src/runtime/socket/socket_body.rsinternal_flush()then clearsbuffered_data_for_node_netonfatal, andon_writabledispatches'drain'regardless, so JS keeps writing on a connection that has silently lost a span of its byte stream.Fix
Add
bsd_send_is_transient_error()(ENOBUFS/ENOMEM on POSIX, WSAENOBUFS on Windows) and treat it the same as would-block inus_socket_write_check_error: setlast_write_failed, re-arm the writable poll, return 0, do not setfatal_write_error. This matches libuv'suv__write, which retries onEAGAIN || EWOULDBLOCK || ENOBUFS.EPIPE/ECONNRESETstill reach the fatal branch; those are connection-fatal and surfaced on the read side. Making the write side itself fail those pending writes is the complementary fix in #33506.Verification
New
test.each(["ENOBUFS", "ENOMEM"])case intest/js/node/net/net-syscall-fault.test.ts.Without the fix (
git stash -- packages/ && bun bd test ...):With the fix:
The rest of
net-syscall-fault.test.ts,tls-syscall-fault.test.ts, andsocket-syscall-fault.test.tspass unchanged.no test proof · iteration 0 · Platform-specific test(s) that do not run on this machine. Deferring to CI, which covers all platforms: test/js/node/net/net-syscall-fault.test.ts