Skip to content

Bun.listen / Bun.connect: report a write error when a close loses the queued tail of end(data) - #44313

Open
robobun wants to merge 7 commits into
robobun/3bf73240/socket-end-queue-tailfrom
robobun/1e6803b5/socket-end-tail-lost-error
Open

robobun wants to merge 7 commits into
robobun/3bf73240/socket-end-queue-tailfrom
robobun/1e6803b5/socket-end-tail-lost-error

Conversation

@robobun

@robobun robobun commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Stacked on #41785.

Problem

  • socket.end(data) queues what one send does not take. If the connection closes with no read error, nothing reports the loss. TLS on a unix socket: end() returned 8388608, the peer got less, close got no error.
  • NewSocket::on_close (src/runtime/socket/socket_body.rs) reports only a read error or a failed send. A hangup gives the clean code, and TLS hides the errno of a failed send.

Fix

  • on_close reports a clean close over a queued tail as a write error: EPIPE, syscall write.
  • The error handler gets it, or close without one: the rule of Bun.listen / Bun.connect: send the whole end(data) chunk before the socket closes #41785 for a failed send.
  • A close by the application reports nothing: terminate(), close(), and close_all (listener.stop(true)), which marks its group.
  • Verified: 33 new tests in test/js/bun/net/socket.test.ts, all fail on the base. Self-reviewed: 17 concerns raised, 14 addressed.

Background

  • uSockets gives on_close a code: 0 (clean), 1 (reset), 2 (fast shutdown), or a read errno. On Linux a failed send() takes the pending error, so the later hangup gives 0.
  • Considered the uWS HTTP rule (close when a writable event sends nothing after the peer's FIN): it misses a peer that sends no FIN.

Downsides

  • Behaviour change: error or close gets an error where close got undefined, also for shutdown() after end(data).
  • Still silent: a read error with no close handler, unref() with a queued tail, a tail that only TLS holds, Windows named pipes.
  • Cost: release binary 80832072 B, .text 58146549 B, NewSocket 504 B, us_socket_group_t 88 B and syscalls per connection are unchanged.
Notes

Order and scope

A close with a read error

  • A reset that the loop reads closes the socket with the read errno. close gets ECONNRESET with syscall read, and error is not called, with or without a queued tail. This PR does not change that.
  • Probe (debug builds): a peer resets 300 ms after the socket called end(16 MiB). With error and close handlers: close ECONNRESET read, 15 of 15. With close only: the same, 15 of 15. With error only: no handler runs, 15 of 15. The base gives the same 45 results.
  • When a send sees the reset first, fail_write (Bun.listen / Bun.connect: send the whole end(data) chunk before the socket closes #41785) gives ECONNRESET with syscall write to error: 25 of 25 on this branch with a peer that reads 64 KiB and then resets.
  • So one event has two shapes, and a socket with only an error handler learns of the loss in the second shape only. tcp.mdx and the JSDoc of end(), close and error in bun.d.ts state both cases.
  • Not built: with no close handler, give the error handler the error of a read-error close over a queued tail. It is one more clause in on_close, and it changes what a read error does, so it needs a maintainer's decision.

shutdown() after end(data), and #44300

unref()

  • end(16 MiB), then unref(), with nothing else alive: the process exits with the tail queued, and the peer sees a clean close. No handler runs. Release builds, 40 runs each: the base lost data in 31 runs, this branch in 33. In the other runs the sends finished before the exit.
  • Node 26 delivers all 16 MiB in 3 of 3 runs: a pending write keeps its loop alive. node:net in Bun loses data the same way (3 of 3 on both builds).
  • This PR does not change it: there is no close to report. socket: shutdown() sends the FIN after queued bytes; a close frees the queue and keeps bytesWritten #44300 names the event-loop hold for queued bytes as the next step of its stack.

Not in this PR

  • The Windows part of the design: a flush that reports its fatal errno on Windows, and a kernel probe after a flush that sent nothing. The review found two defects in the draft. For a TLS socket that is stalled on its own ciphertext, the close only arms a deferred close that waits for that ciphertext to drain. And the probe repeats a decision that usockets already makes in three places (eventing/libuv.c, crypto/openssl.c, HttpContext.h). No Windows machine was available, and a comment in the code keeps that path: "Until that detection is verified on Windows, keep the legacy contract there". So Windows keeps the base behaviour: a flush that gets a fatal errno drops the tail without a report.
  • The errno that usockets drops. us_socket_raw_write folds a failed send of TLS ciphertext to 0, so the wrapper never sees its errno. The report here says EPIPE, also where the kernel said ECONNRESET (the two flush() rows of the table).
  • A tail that only the TLS ciphertext spill holds. The wrapper queue is empty and the wrapper is detached, so on_close sees nothing. us_socket_ssl_spill_pending at close is one way to cover it. The JSDoc of error and close names this limit.
  • upgradeTLS() after end(data): the wrapper that is retired keeps its queue and gets no close.
  • node:net and node:tls. They queue with $write, never call end(data), and fail a queued write themselves (ERR_SOCKET_CLOSED on a clean close, where Node reports write EPIPE).

Repro (linux-x64, release builds)

end(8 MiB) from a Bun.listen socket with error and close handlers. The peer reads one chunk and calls terminate(). Events after end(). The table is from the first revision of this PR (base 6927ed3):

Transport Peer got Base This PR
TLS over unix 229376 close error EPIPE write, close
TLS over TCP, the socket calls flush() after the reset 81920 close error EPIPE write, close
TLS over TCP 81920 close ECONNRESET read same
TCP 128000 close ECONNRESET read same
TCP, the socket calls flush() after the reset 128000 error ECONNRESET write, close same
unix 219264 error EPIPE write, close same
TLS over unix, end(250000) 229376 close close (the limit above)

The first row again on this revision (base fe18122), 20 runs for each set of handlers. error and close: the base gives close, this PR gives error EPIPE write, close. close only: close, and close EPIPE write. error only: no handler runs, and error EPIPE write.

When on_close reports

The close has code 0, the socket is a usockets socket, end() was called, the handshake did not fail, the queue is not empty, and the group is not in close_all.

Close site (clean code) Origin Report
loop.c:916 the peer's FIN after this side's FIN: shutdown() after end(data), or a TLS socket after a failed send yes
loop.c:947 hangup yes
openssl.c:2413, :2538, :2547 a TLS failure on the read side yes
openssl.c:2528 the peer's close_notify after this side's yes
openssl.c:2162, :2305, :2316, :1914 the handshake fails no: the handshake handler reports it
openssl.c:2157, :2358, :2470, :2771 a deferred close that resumes with its stored code as the close that was deferred
context.c close_all walk and low-priority drain the owner: listener.stop(true), listener finalize, test isolation, VM teardown no: group mark

Closes that the socket starts: terminate(), a rejected handshake and the end of end() go through close_and_detach, which empties the queue. close() uses code 2. A reconnect goes through detach_for_reconnect. fail_write closes after the flush emptied the queue.

Tests

  • 33 new tests in one describe block. With the src/ and packages/ of the base: 33 fail (one of them by its timeout: it waits for an error event that the base never sends). With this branch: 33 pass, in 9 runs of the block on the debug build and 5 on the release build.
  • Lost by a TLS record that cannot be read, the same on every event backend: with and without an error handler, an error handler without a close handler, and the two reload() rows. This is also the control of every row that expects no report. These rows run over AF_UNIX on POSIX and over TCP on Windows. Over TCP on macOS they closed with no error in some CI runs: the kernel took the whole queued tail before the record arrived. The kernel buffers of an AF_UNIX socket are small and do not grow.
  • Lost by shutdown() after end(data), the same on every backend: TCP and TLS, both sides, with and without an error handler, and the reconnect row.
  • Lost by a peer that resets, per backend: epoll gives the write error. kqueue and libuv close with a read error (close ECONNRESET read), which the base gives too, so those rows do not tell this change from the base there.
  • No report: terminate(), close(), listener.stop(true), a failed handshake in five configurations, a slow TLS peer that gets every byte, a node:tls socket. Each of these rows also checks the control, so it fails on the base.
  • Not run by me: macOS and Windows. The expectations for kqueue and libuv come from the code and from CI.
  • Each clause of the guard, removed in a debug build (on the first revision; these rows did not change): no code clause, the TCP close() row fails. No END_AFTER_FLUSH clause, the node:tls row fails. No REJECTED clause, three of the five handshake rows get EPIPE. No group mark, the TCP stop(true) row fails (the TLS one ends with code 1). No is_usockets_backed clause, no row fails: it decides for named pipes and upgraded duplexes, whose close has code 0 for every origin.
  • The whole file on a debug build of this revision: 271 pass, 6 skip, 2 fail. should not call drain before handshake needs the public internet. end(data) without an end handler keeps the process alive until the queued tail is sent is a test of Bun.listen / Bun.connect: send the whole end(data) chunk before the socket closes #41785 that timed out at 5 s. Both fail on a build of the base too. In the full run on the base build, two more tests of the base timed out at that load.
  • That second test fails alone, 5 of 5 runs on this branch and 4 of 4 on the base, while the load average of my machine is 700 to 900. It builds a 4 MiB payload with a callback that runs 1,048,576 times, once in the parent and once in the child. One build of the payload takes 2.5 to 3.4 s on the debug build at that load. It passed in the runs of the first revision, at a lower load.
  • Two tests of Bun.listen / Bun.connect: send the whole end(data) chunk before the socket closes #41785 failed once each in the CI of this PR and passed on the retry. On Windows, end(data) without an end handler keeps the process alive until the queued tail is sent printed its close line before its fin read line. On macOS, end(data) whose tail is still queued when the peer's FIN is read > tls 1.2 > Bun.connect socket > allowHalfOpen: false > end(data) in data() with an end handler got end() returning -2 with no request read.
  • Also run on the debug build of the first revision: test/js/bun/net/{tcp-server,socket-syscall-fault}.test.ts, test/js/node/net/{node-net-allowHalfOpen.test.js,node-net-server.test.ts,node-net.test.ts}, test/js/node/tls/{node-tls-socket-allow-half-open-option,node-tls-raw-end,node-tls-server,node-tls-connect,tls-syscall-fault}.test.ts, test/js/node/http/{node-http-server-socket-end-drain,node-http-backpressure}.test.ts, test/js/node/http2/node-http2.test.js, the nine test-net-half-open-peer-reset-*.mjs, and three node parallel half-open scripts. node-net.test.ts fails the same 11 tests on a build of the base. On this revision: test/integration/bun-types/bun-types.test.ts, 21 pass.
  • cargo check -p bun_runtime for x86_64-pc-windows-msvc (this revision) and aarch64-apple-darwin (the first revision).

Measurements (release builds, base = fe18122)

  • File size 80832072 B for both. size -A: every section has the same size on both builds (.text 58146549 B, .rodata 19804908 B, .data 59648 B, .bss 1810040 B).
  • Symbol sizes (nm -S on the unstripped binaries): the sum grows by 268 B over 23 symbols, while .text keeps its size. on_close +279 B (TCP) and +327 B (TLS). The new report_write_error is 477 B. fail_write is 461 B smaller for each of TCP and TLS, because it calls report_write_error. us_socket_group_close_all_ex +2 B. The other 17 symbols change by 1 to 13 B.
  • size_of from the debug info: NewSocket<false> and NewSocket<true> 504 B, SocketGroup 88 B, before and after. SocketGroup is the Rust mirror of us_socket_group_t, and a compile-time check holds the two to one size. The new byte takes padding.
  • Syscalls per connection, counted with an LD_PRELOAD shim on the libc calls, (count at N=300 - count at N=100) / 200, sequential connections from a node peer. Nothing queued, TCP: send 1, recv 1, close 1, setsockopt 2, accept4 2, epoll_ctl ADD 1, MOD 0, DEL 1, on both builds, three runs each. Nothing queued, TLS: send 3, recv 3, the rest as TCP, on both builds, three runs each.
  • With a 4 MiB tail that the peer reads (TCP, (count at N=60 - count at N=20) / 40): send 2.13 to 2.65 on the base and 2.15 to 2.90 on this branch, four runs each. recv 1, close 1, epoll_ctl ADD 1, MOD 1, DEL 1 on both. The difference in send is inside the spread between runs.
  • Per close, from the diff: up to six tests of a field or a flag in loses_end_tail, three in on_close, and a second read of the close callback. No allocation and no syscall. Per close_all: one byte load and two byte stores.
  • No perf, valgrind, bloaty or strace on the machine, so there are no instruction counts.

Self-review


no test proof · iteration 0 · platform-specific test(s) that do not run on this machine, deferring to CI, which covers all platforms: test/js/bun/net/socket.test.ts

@robobun
robobun requested a review from alii as a code owner September 30, 2026 17:52
@robobun

robobun commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator Author

Status: ready for review. Stacked on #41785 (the base of this PR is that PR's branch, at fe18122).

How it was reproduced (linux-x64, release builds, 20 runs for each set of handlers):

  • A Bun.listen TLS server on a unix socket answers a request with socket.end(Buffer.alloc(8 << 20)). The peer reads one chunk and calls terminate().
  • Base: end() returned 8388608, the peer got less, and close(socket, undefined) ran. No error event. With only an error handler, no handler ran.
  • This branch: the error handler gets EPIPE with syscall write, then close runs. Without an error handler, close gets that error.

The regression tests are in test/js/bun/net/socket.test.ts: 33 new tests. With the src/ and packages/ of the base, all 33 fail. With this branch, all 33 pass.

A connection that a read error closes (a reset that the loop reads) still gives that error to close only. The docs in this PR say so.

Not run by me: macOS and Windows.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

Beyond the inline finding, I also checked the re-entrancy ordering in on_close: loses_end_tail reads flags, the queue and the group mark before any JS entry, and handlers.on_close() is re-read after the error handler runs, so a reload() from that handler is respected. The closing_all byte lands in existing tail padding of us_socket_group_t, so the Rust size assertion and ABI stay intact, and the save/restore in us_socket_group_close_all_ex keeps the mark across re-entrant calls.

Extended reasoning...

The change touches usockets (us_socket_group_t gains a closing_all byte set during close_all), its Rust mirror, NewSocket::on_close/fail_write in socket_body.rs, the .d.ts/docs contract, and adds 33 tests. No security-sensitive surface (no auth, crypto, or parsing); it is a behavior change in close/error reporting for Bun.listen/Bun.connect. A confirmed inline finding (TLS ciphertext spill tail still unreported) plus the behavior change and stacked-PR base mean a human should weigh the contract.

Comment thread src/runtime/socket/socket_body.rs
us_socket_group_close_all_ex sets closing_all on the group for the time
of the call. An on_close handler can then tell a close that the owner of
the group started from a close that the loop or the peer started. The
field takes one byte of the padding of us_socket_group_t.
end(data) queues what one send does not take and returns the size of the
chunk. When the connection closed with the clean code while that tail
was still queued, nothing reported it.

NewSocket::on_close now reports that close as a write error, EPIPE with
syscall write. The report uses the rule that fail_write already has for
a failed send: the error handler gets it, and without an error handler
the close handler gets it. report_write_error holds that rule for both.

A close that the application starts reports nothing: close_and_detach
and detach_for_reconnect empty the queue, close() uses another close
code, and close_all marks its group. A failed TLS handshake stays with
the handshake handler. Named pipes and upgraded duplexes are left out:
their close has code 0 for every origin.
…te error

The rows lose the tail by a shutdown() after end(data), by a TLS record
that cannot be read, and by a peer that resets. They check the error
handler first, the close handler without one, and an error handler
without a close handler.

terminate(), close(), listener.stop(true), a failed handshake, a slow
TLS peer and a node:tls socket report nothing new. Each of those rows
also checks that a lost tail is reported, so it fails on a build that
never reports.
tcp.mdx and the SocketHandler comments in bun.d.ts say which handler
gets the error, which calls discard the queued part, and what the
report does not cover.
@robobun

robobun commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator Author

A note on how this PR meets #44300 (shutdown() sends the FIN after queued bytes, a close frees the queue). Both edit the same place at the top of NewSocket::on_close.

The two merge with a text conflict on that line. The right order is: compute write_errno with loses_end_tail first, then free the queue. With the other order the queue is already empty and loses_end_tail is always false. The tests of this PR catch that.

One more effect: with #44300, shutdown() after end(data) no longer cuts the tail, so the Downsides line here about that case no longer applies when both are in.

…r a queued tail

The sentences said that every close over a queued end(data) tail gives a write
error to the error handler. A connection that a read error closes (a reset that
the loop reads) gives that error to the close handler and calls no error
handler. State both cases in tcp.mdx and in the JSDoc of end(), close and error.
… TLS record

The control of the rows that expect no report, the row with an error handler
and no close handler, and the two reload() rows lost the tail by shutdown()
after end(data). They now lose it by a TLS record that the peer makes
unreadable, which no call of the socket under test starts. Only the eight
shutdown() rows and the reconnect row depend on what shutdown() does with a
queued tail.
@robobun
robobun force-pushed the robobun/1e6803b5/socket-end-tail-lost-error branch from 015d112 to f17532a Compare September 30, 2026 21:12

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code review found no issues

No high-confidence issues detected in this change.

Still open from earlier reviews (1):

  • Unresolved: 1 minor or pre-existing.

On macOS the rows that lose the tail by an unreadable TLS record closed with
no error in some runs: over TCP the kernel took the whole queued tail before
the record arrived, so the close had nothing queued. An AF_UNIX socket has
small kernel buffers that do not grow, so the tail is still queued at the
close. Windows keeps TCP.
@robobun

robobun commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 4:15 PM PT - Sep 30th, 2026

✅ @robobun, your commit 938bd74ec6ec6ddb922357ca56a1918b8969e894 passed in Build #122017! 🎉


🧪   To try this PR locally:

bunx bun-pr 44313

That installs a local version of the PR into your bun-44313 executable, so you can run:

bun-44313 --bun

@robobun

robobun commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator Author

The one open item is the tail that only the TLS ciphertext spill holds. It stays a named limit of this PR, and the thread has the reason.

Changes since that review:

  • The branch is rebased on the head of Bun.listen / Bun.connect: send the whole end(data) chunk before the socket closes #41785 (fe18122).
  • tcp.mdx and the JSDoc of end(), close and error now state both cases: a connection that a read error closes gives that error to close only, and a connection that closes with no read error gives the write error to error (or to close without an error handler).
  • Tests: the rows that do not test shutdown() lose the tail by a TLS record that cannot be read. Those rows run over AF_UNIX on POSIX, because over TCP on macOS the kernel took the whole queued tail before the record arrived in some runs.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code review found no issues

No high-confidence issues detected in this change.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant