Conversation
us_udp_socket_send() subtracted each pass's batch size from num before testing it, so the loop ended after the first pass and the drain guard compared the pass's result against the remaining count instead of the pass's own size. sendMany(300) on an idle socket returned 204, left the other 96 unsent, and never armed the writable poll, so the documented "if fewer were sent, wait for drain and resend the rest" protocol had nothing to wait for; sendMany(500) returned 408 and armed a drain with no backlog. Walk the batch with an offset instead, and key backpressure off sendmmsg's own report: -1 with EWOULDBLOCK arms the writable poll and returns the count that did go out, any other errno goes to the caller. sendmmsg(2) turns a datagram that fails mid-batch into a short count and drops its errno, so a short count now retries from the failed datagram, which brings the errno back as -1. An oversized payload therefore throws EMSGSIZE wherever it sits in the batch rather than only at index 0.
|
Updated 4:54 PM PT - Jul 5th, 2026
✅ @robobun, your commit 2d2cd41bf289b7250a210479c377260598ac8ce0 passed in 🧪 To try this PR locally: bunx bun-pr 33391That installs a local version of the PR into your bun-33391 --bun |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Pro Run ID: 📒 Files selected for processing (2)
WalkthroughThis PR reworks the UDP batch-send loop in ChangesUDP send retry logic and tests
Sequence Diagram(s)sequenceDiagram
participant Caller
participant us_udp_socket_send
participant bsd_udp_setup_sendbuf
participant bsd_sendmmsg
Caller->>us_udp_socket_send: send(payloads, lengths, addresses, num)
loop while total_sent < num
us_udp_socket_send->>bsd_udp_setup_sendbuf: setup(payloads+total_sent, num-total_sent)
us_udp_socket_send->>bsd_sendmmsg: sendmmsg()
alt sent < 0 and not would_block
bsd_sendmmsg-->>us_udp_socket_send: error
us_udp_socket_send-->>Caller: return error
else would_block
us_udp_socket_send->>us_udp_socket_send: register writable, break
else sent >= 0
bsd_sendmmsg-->>us_udp_socket_send: sent count
us_udp_socket_send->>us_udp_socket_send: total_sent += sent
end
end
Estimated code review effort: 3/5 (High complexity C-level retry logic change requiring careful review of socket error semantics; tests are straightforward.) Related issues: None found in provided context. Related PRs: None found in provided context. Suggested labels: bun:usockets, needs-tests-verified Suggested reviewers: None found in provided context. 🥁 A drumroll for datagrams stuck in line, 🚥 Pre-merge checks | ✅ 4✅ Passed checks (4 passed)
Comment |
|
Closing as part of a cleanup of stale pull requests. This PR has had no new commits since 2026-07-05 and it conflicts with main. This is not a judgment on the fix itself. If the problem still reproduces on a current build, reopen this PR after a rebase or open a new one against main. |
|
Follow-up for anyone who lands here. #32625 fixed the main finding of this PR (the loop that stopped after the first batch of about 204 datagrams), so this PR stays closed. The second half of this PR made a mid-batch errno throw at any position. That is not revived: #32602 chose the opposite contract on purpose (a count once one datagram is out, and the errno on the next call), which matches |
Bun.udpSocket().sendMany()silently drops packets once the batch outgrows the internalsendmmsgsend buffer (~204 slots), and the documented backpressure protocol has nothing to resume it.Repro
On 1.4.0:
N=300returns204,N=500returns408plus onedrainwith no backlog behind it. Anything at or below 204 works, which is why the existing 100-packet tests never caught it.Cause
us_udp_socket_send()decrementednumby each pass's batch size before the loop condition and the drain guard read it:For
num = 300: pass one sends 204,numbecomes 96,total_sent < numis204 < 96so the loop exits with 96 packets never handed to the kernel.sent < numis204 < 96too, so the writable poll is never armed anddrainnever fires. Atnum = 500the same comparison fires the other way in pass one (204 < 296) and arms adrainfor a send that was not actually blocked.The same guard is what a real short send lands in, so genuine kernel backpressure could not arm
draineither.Fix
Walk the batch with an offset and leave
numalone, then key backpressure off whatsendmmsgactually reports:-1withEWOULDBLOCKis the only recoverable outcome: arm the writable poll, return the count that went out.sentfailed.sendmmsg(2)drops that datagram's errno ("If an error occurs after at least one message has been sent... the error code is lost"), so the next pass retries from it, and the errno comes back as-1.That last part also fixes a sibling bug in the same function: an oversized datagram threw
EMSGSIZEat index 0 but came back as a short count ("1 of 3 sent") anywhere after it, so an un-sendable packet was indistinguishable from backpressure. It now throws wherever it sits in the batch.Two behavior changes worth naming:
EAGAINout ofsend()/sendMany(). It now returnsfalse/ a short count and armsdrain, which is what the docs describe.drainlater) is what made oversized packets look like backpressure.Verification
test/js/bun/udp/udp_socket.test.ts: 8 of the 10 new cases fail on the unfixed build (the two "index 0" cases pass both ways on purpose, they are the control for the position dependence).All 188 pass with the fix, as do
test/js/bun/udp/dgram.test.ts,udp_socket_recv_flags.test.ts, and everytest/js/node/test/parallel/test-dgram-*.js. The 500-packet send was run 30 times without a short count.