Skip to content

udp: take every queued ICMP report before the first error handler runs - #44461

Open
robobun wants to merge 3 commits into
mainfrom
robobun/0480ed13/udp-errqueue-take-then-deliver
Open

robobun wants to merge 3 commits into
mainfrom
robobun/0480ed13/udp-errqueue-take-then-deliver

Conversation

@robobun

@robobun robobun commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Fixes #44218

Problem

  • On Linux a UDP error handler that sends again to a closed port on the same host never leaves one poll dispatch: 100000 'error' events; setImmediate ran 0 times; the 1 ms timer fired: false.
  • The error pass in us_internal_dispatch_ready_poll (packages/bun-usockets/src/loop.c:1004) calls the handler between two recvmsg(MSG_ERRQUEUE). Over loopback the handler's datagram queues its report before send() returns.

Fix

  • The pass takes every report off the queue, then calls the handler for each. A handler's own report belongs to the next event.
  • Then it reads SO_ERROR once: the pending errno of a handler's datagram would fail the next recvmmsg. The pass reports it only if no report is queued.
  • Verified: test/js/bun/udp/udp_socket_recv_flags.test.ts (15 new tests, 6 fail on main). Also the UDP, dgram, QUIC and HTTP/3 suites.
  • Self-reviewed: 21 concerns raised, 20 addressed. Declined: code for gVisor (Notes).

Background

Downsides

  • A handler that sends 3 datagrams for each error still gets up to 221 calls in one loop iteration (node: 1).
  • An event with reports costs 1 more syscall (getsockopt), 2 after a handler sent. Release text: +512 bytes.
  • Reports of one event arrive grouped by errno, not in queue order.
Notes

Who can hit it. All of these must hold:

  • Linux.
  • The peer is on the same host and has no listener. Only then the ICMP error arrives before send() returns.
  • The socket is a Bun.udpSocket, or a node:dgram socket that is connected. dgram.ts drops the reports of a socket that is not connected.
  • The error handler sends again before the loop runs again: in the handler, in process.nextTick, or after an await. A send from setImmediate does not hang on main.

No user reported it. A review of the UDP dispatch found it. The drain loop is in releases since v1.3.14 (#29768).

Why the kernel does not pace this loop. A maintainer said of the TCP read loop (#44220) that the kernel paces the reads. That loop needs a peer to make its input. This pass does not: the send() of the handler queues the next report before it returns, so each pass of the loop made the input of the next pass.

The class, and what this PR leaves. us_internal_dispatch_ready_poll has four loops of this shape: this error pass, the UDP data loop, the accept loop and the TCP read loop. A handler can feed the first three with no peer. This PR changes the error pass only. With this PR the other two still keep the loop in one event:

Why the error pass is a unit of its own. Only the error queue has these two properties:

  • A report reduces to an errno. So the pass can take the whole queue before the first handler runs. It needs no constant, and an event with no EPOLLERR pays nothing.
  • A report that stays queued costs inbound datagrams (80 of 1024 with a limit of 32). So a count is not an option here.

The rule. One event handles what its queue held when the event began. What a handler queues belongs to the next event, after the timers and the other polls. This PR applies the rule to the error queue, where it is exact. The data loop and the accept loop cannot take their queue before the handler runs, so #37103 and #44198 use a count of 32. If a maintainer wants the same rule there, those counts become a bound from a capacity (the receive buffer, the listen backlog).

The 221 calls. A handler that sends 3 datagrams queues more than 1 report for the next event (about half of its sends fail with the pending error). So the number of reports grows from event to event until the receive buffer is full (221 reports). Node gets 1 error for each loop turn because it never sets IP_RECVERR: the kernel then keeps one pending error. node:dgram without IP_RECVERR is the change that ends this. It is not built yet, and it is not in this PR.

The SO_ERROR read. It is not part of the fix for the hang. It keeps two behaviours of main. On main the loop took the report of the handler's datagram in the same event, and the kernel clears the pending error when the last report leaves the queue.

  • A datagram that waits behind a report is delivered in the same event. The errno of the handler's datagram is not reported a second time.
  • A send after the event does not fail with the errno of the handler's datagram.

poll runs only if SO_ERROR was not zero, that is after a handler sent a datagram that was refused. If no report is queued then, the kernel had no room for one, and the pass reports the errno once with errqueue not set. Main reported that errno from the recvmmsg that failed with it.

Measurements. Release builds of main (f4d755a) and of this PR on that commit, Linux x64, kernel 7.0.0. An LD_PRELOAD library counted the calls. gdb single steps counted the instructions that us_internal_dispatch_ready_poll itself executes.

main this PR
100 readable UDP events, no error 200 recvmmsg, 0 error-queue calls the same
Instructions for an event with no EPOLLERR: UDP readable, TCP readable, TCP writable 122, 150, 108 118, 149, 108
One report 2 recvmsg 2 recvmsg, 1 getsockopt
Retry chain of 100 100 errors in 1 loop iteration, 101 recvmsg 1 error in each of 100 iterations, 200 recvmsg, 100 getsockopt, 99 poll
100 reports in one event 101 recvmsg 101 recvmsg, 1 getsockopt
64 iterations of 64 reports and 16 datagrams 4096 errors, 1024 datagrams, 4160 recvmsg the same, and 64 getsockopt
32 reports, then an errno with no room for a report 33 errors, 33 recvmsg, 3 recvmmsg 33 errors, 33 recvmsg, 2 recvmmsg, 1 getsockopt, 1 poll
The handler sends 3 datagrams for each error, until 2000 errors 2001 errors in 1 iteration at most 221 errors in one iteration
8 reports in one event, the handler sends once for each 1 send reaches the wire, 7 fail with ECONNREFUSED 4 reach the wire, 4 fail
Release binary, text 80731481 bytes 80731993 bytes

The row with 8 reports is a change in behaviour. At each dequeue the kernel sets the pending error to the errno of the next queued report, so on main a send from the handler fails while reports wait. With this PR the queue is empty when the handlers run, and only the send after a refused datagram fails.

What stays as on main.

  • The receive path and send() still report the pending error while its report is queued (udp: stop IP_RECVERR on unconnected sockets; never fatal-by-default for recv errors #35922 describes it).
  • A queued report of another errno hides an ICMP errno that the kernel had no room to queue. On main the dequeue of the queued report clears the pending error in the same way.
  • A node:quic endpoint has an error callback that does nothing, so udp.c:238 gives it IP_RECVERR. The pass takes its reports and drops them.
  • gVisor: not run. Its netstack raises EPOLLERR only for the pending error, and a read from its error queue does not clear the pending error (pkg/tcpip/transport/udp/endpoint.go, Readiness and onICMPError). So main reports each ICMP error two times there, in two events. This PR reports it two times in one event. The review asked for one more read of the error queue for that case. It is not here: the number of reports is the same as on main, and no gVisor run exists to test it.

Tests. The block error queue (IP_RECVERR) has 15 new tests. They count the error calls by getEventLoopStats().iteration.

  • On main the 6 rows of an error handler that sends again gets the next report in the next loop iteration fail: 100 errors in one iteration, and the immediate did not run.
  • The other 9 pass on main. They guard what the change must keep: every datagram of a mixed load arrives, each report has its errno (IPv4, IPv6, IPv4 from an IPv6 socket), no handler runs after a close, no error stays on the socket after its event (CPU 100% on UDP error !!!! #29436), an errno with no room for a report is reported once, and the two behaviours of the SO_ERROR read.
  • 9 mutations of the change each fail a test: no SO_ERROR read, the read never reports, the read without the poll, the read only without a datagram, the read only with a datagram, a handler call after the close, one report for each event, one errno slot, no IPV6_RECVERR branch.
  • Suites on the debug build: udp_socket_recv_flags (17), udp_socket (230), dgram (62), node's test-dgram-* (78 files), test-quic-endpoint-* and test-quic-sni* (15 files), quic-endpoint (7), quic-stream (7), quic-sni (6), fetch-http3-client (65), serve-http3 (73).

Other open PRs in this code.

This PR, #35922 and #37103 change the same switch case. They need one decision: does Bun.udpSocket keep IP_RECVERR on every socket, and does node:dgram drop it as node does.


no test proof · iteration 0 · platform-specific test(s) that do not run on this machine, deferring to CI, which covers all platforms: test/js/bun/udp/udp_socket_recv_flags.test.ts

The Linux error pass of the UDP case in us_internal_dispatch_ready_poll
called the error handler between two recvmsg(MSG_ERRQUEUE). A handler
that sends again to a closed port on the same host queues the next
report before send() returns, so the pass never found the queue empty.

The pass now takes every report off the queue first, counted by errno,
and then calls the handler once for each report. After the handlers it
reads SO_ERROR once and reports that errno when no queued report
carries it.
@github-actions github-actions Bot added the claude label Oct 2, 2026
@robobun

robobun commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator Author

Status: ready for review.

How to reproduce: run the script of #44218 on Linux.

Build Output
main (f4d755a, release, Linux x64, kernel 7.0.0) 100000 'error' events; setImmediate ran 0 times; the 1 ms timer fired: false
this PR 200 loop turns; 198 'error' events; the 1 ms timer fired: true
Node v26.3.0 200 loop turns; 199 'error' events; the 1 ms timer fired: true

The test test/js/bun/udp/udp_socket_recv_flags.test.ts shows the same. On main the handler gets 100 errors in one loop iteration. With this PR it gets 1 in each iteration.

CI: build 122940 (the first commit) is green. Build 123162 (the current head, which adds a test fix and restores two comments) has one failed job: test/js/bun/spawn/spawn.test.ts on debian 13 x64-asan. That test also fails on main, and this PR does not touch it. udp_socket_recv_flags.test.ts passes on every lane.

@coderabbitai

coderabbitai Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository: oven-sh/bun/.coderabbit.yaml
  • Review profile: ASSERTIVE
  • Plan: Essentials
  • Run ID: 288024af-e35d-44be-9be7-f915adc894f3
📥 Commits

Reviewing files that changed from the base of the PR and between 5e6a7f5 and 985e4ae.

📒 Files selected for processing (1)
  • test/js/bun/udp/udp_socket_recv_flags.test.ts

Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 2 remain after this review.


Walkthrough

Linux UDP polling now collects queued error reports before invoking callbacks, then checks for a pending socket error. Linux-only helpers expose pending and queued error state. New tests cover repeated errors, datagram delivery, and event-loop ordering.

Changes

Linux UDP error handling

Layer / File(s) Summary
Pending and queued error primitives
packages/bun-usockets/src/internal/networking/bsd.h, packages/bun-usockets/src/bsd.c, packages/bun-usockets/src/internal/internal.h
Linux-only helpers retrieve pending SO_ERROR and check for queued error reports. The callback comment describes queued reports and pending errors.
Batch and dispatch UDP errors
packages/bun-usockets/src/loop.c, src/runtime/socket/udp_socket.rs, src/js/node/dgram.ts
The poll handler batches queued reports before callbacks, then checks for a pending error. Runtime comments describe the callback paths, and the node:dgram comment describes pending socket errors.
Linux UDP error and event-loop tests
test/js/bun/udp/udp_socket_recv_flags.test.ts
Linux tests cover repeated error callbacks, queued reports, datagram delivery under receive-buffer pressure, socket closure, and event ordering across Bun and Node UDP sockets.

Priority: ➖ Normal

Severity of issue fixed: Medium

Merge Risk: ⚪ Minimal · up to d46ef

The reviewed change has no identified merge-blocking issue. The added Linux UDP tests still need to complete in CI.

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly and specifically describes the main change: draining queued ICMP reports before invoking error handlers.
Description check ✅ Passed The description explains the problem, the fix, and verification. It uses different headings from the template, but it includes the information required by both template sections.
Linked Issues check ✅ Passed Issue #44218 requires retries from Linux UDP error handlers to leave the current poll dispatch without leaving queued reports that consume receive-buffer space. The reviewed change drains queued repor…
Out of Scope Changes check ✅ Passed The changes stay within issue #44218. The Linux helpers and loop.c changes support error-queue draining and pending-error handling. The node:dgram and callback-comment updates clarify error delive…
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Comment @coderabbitai help to get the list of available commands.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Beyond the inline findings, I also checked two things: the error queue is still drained when on_recv_error is NULL (the take runs unconditionally and recv_error_surfaced keeps its old meaning, so the quic no-op-callback case is unchanged), and the new code is entirely __linux__-gated with no sibling kqueue/Windows UDP path that uses IP_RECVERR, so there is no parallel backend left unfixed.

Extended reasoning...

The change reworks the Linux-only EPOLLERR branch of the UDP poll dispatch in packages/bun-usockets/src/loop.c into a two-phase drain-then-dispatch, adds two getsockopt/poll helpers in bsd.c, updates comments in internal.h, udp_socket.rs and dgram.ts, and adds ~495 lines of Linux-only tests. It touches no auth, crypto or input-parsing surface; the one kernel-reported length (cmsg_len) gains a bounds check. Three confirmed findings (test race, errno reordering within one event, gVisor double-report) are posted inline, so a human look is already signaled; this note only records the extra items ruled out.

Findings marked 🟡 are optional suggestions and need no follow-up push.

Comment thread test/js/bun/udp/udp_socket_recv_flags.test.ts Outdated
Comment thread packages/bun-usockets/src/loop.c
Comment thread packages/bun-usockets/src/loop.c
refusedSends() detected a late ICMP error only when the second send of a
pair did not fail. The pending error of that late ICMP error then failed
the first send of the next pair, and the exception failed the test. Every
send that a late ICMP error can fail now counts as late, so the scenario
is set up again.
Comment thread src/js/node/dgram.ts Outdated
Comment thread src/runtime/socket/udp_socket.rs Outdated
Comment thread src/runtime/socket/udp_socket.rs Outdated
@robobun

robobun commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 11:34 PM PT - Oct 2nd, 2026

❌ @robobun, your commit d46efc7 has 1 failures in Build #123162 (All Failures):


🧪   To try this PR locally:

bunx bun-pr 44461

That installs a local version of the PR into your bun-44461 executable, so you can run:

bun-44461 --bun

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I re-reviewed after the test-only push and found no bugs; because the usockets UDP error-dispatch change in loop.c is unchanged since my last pass and alters when and in what order Linux ICMP reports reach handlers, a human look there is still worthwhile.

What was reviewed:

  • The new refused() helper and its use in refusedSends and the in-handler send("retry") sites (test/js/bun/udp/udp_socket_recv_flags.test.ts:212-237, :532, :575): a late-ICMP ECONNREFUSED on any of those sends now counts as late and re-runs the scenario via firstUsable, which closes the uncaught-throw path I raised earlier.
  • sendTooLarge now returns a late count (:275-285), so the otherErrno matrix feeds refusedSends the same way.
  • loop.c, bsd.c, bsd.h, internal.h: no changes since the prior review; the drain-then-dispatch loop, the !u->closed guard between handler calls, and the SO_ERROR/poll pair were re-read with nothing new found.
Extended reasoning...

The diff reworks the Linux-only MSG_ERRQUEUE drain in packages/bun-usockets/src/loop.c for us_udp_socket_t (take all reports into a per-errno tally, then dispatch, then read SO_ERROR and poll for POLLERR), adds two getsockopt/poll helpers in bsd.c/bsd.h, updates a comment in internal.h, dgram.ts and udp_socket.rs, and adds ~500 lines of Linux-gated tests. It touches no auth, crypto, or injection surface; the sensitive part is event-loop dispatch order and syscall-count per UDP error event. The only change since the prior review is the test commit 985e4ae, which addresses my earlier inline finding at the first send of refusedSends. The native change is non-trivial and changes observable errno ordering and per-event syscalls, so it is not a simple/mechanical change suitable for approve without a human.

Still open from earlier reviews (2):

  • Unresolved: 2 minor or pre-existing.

internal.h holds the contract of is_errqueue. The two callers keep the
wording of main.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code review found no issues

No high-confidence issues detected in this change.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

udp: an error handler that sends again to a closed local port holds the event loop (Linux)

1 participant