socket: shutdown() sends the FIN after queued bytes; a close frees the queue and keeps bytesWritten - #44300
Conversation
…e queue and keeps bytesWritten shutdown() over a queue with bytes records the FIN, and internal_flush sends it after the last byte. Every discard of the queue goes through one helper that frees the buffer and adds the dropped count to bytes_written, so bytesWritten keeps the accepted total. on_close frees the queue of a usockets-backed socket at once.
|
Status: ready for review. The diff is green. How I reproduced it (loopback, debug build unless noted):
Tests: Self-review: 17 concerns raised, 15 addressed, 2 rejected. The Notes in the body list them. CI at |
…3bf73240/fin-after-queue-release-at-close
…3bf73240/fin-after-queue-release-at-close
…3bf73240/fin-after-queue-release-at-close
|
Updated 2:14 AM PT - Oct 1st, 2026
✅ @robobun, your commit daba962dd377db4c2a57640f2eddc4e3740feeff passed in 🧪 To try this PR locally: bunx bun-pr 44300That installs a local version of the PR into your bun-44300 --bun |
… the raw half keeps its immediate FIN close() detaches the wrapper before a TLS close that waits for its ciphertext spill, so on_close read a detached wrapper and kept the queue. It now reads the socket of the dispatch. shutdown() on the raw half of an upgradeTLS pair sends the FIN at once, as before: no writable event reaches that half, so a deferred FIN was never sent.
…minate(), and the raw half
|
What comes after this PR. It is not built yet, and it needs a maintainer decision first. The defect that is left.
Where the bytes are. Accepted bytes sit in two stores: the queue of The plan, one PR each:
The decision.
|
…the queue The queue lends its bytes to a send through lend(). A release that runs while the send is on the stack (the socket closed from inside the write) takes effect when the send returns, so the slice stays valid for the whole call. With that, on_close frees the queue for every transport, and close() frees it at once, also when a TLS close waits for its peer.
There was a problem hiding this comment.
Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.
Still open from earlier reviews (2):
- Unresolved: 2 minor or pre-existing.
Additional findings (outside the current diff — GitHub can't attach inline comments there):
-
🟣
src/runtime/socket/socket_body.rs— pre-existing: a node:net write on the raw half of atls.connect({ socket })pair that the kernel takes only partly is queued into a buffer nothing ever flushes, so its tail is lost and its write callback never fires. socket_body.rs:2807 passesbuffer_unwritten_data = truefor every socket, whileendat socket_body.rs:3382 passes!BYPASS_TLSbecause no writable event reaches that half (the PR's own comment at 3306). Fix: queue no tail on aBYPASS_TLSsocket inwrite_or_end_bufferedtoo, and report the short write to JS as the raw half'sendalready does, so the caller can retry fromdraininstead of waiting forever.Why this was flagged
After
tls.connect({ socket: raw })net.ts setsconnection._handle = raw(src/js/node/net.ts:2243), the BYPASS_TLS twin built at socket_body.rs:3816-3842. Araw.write(chunk)whose send is short reacheswrite_buffered->write_or_end_bufferedwith an empty queue, which at socket_body.rs:2807 callswrite_or_end::<false>(global, &mut values, true); at 3132-3136 the unsent remainder is appended tobuffered_data_for_node_net. Dispatch for thatus_socket_tgoes to the tls half via the ext slot; the raw half is reached only byus_dispatch_ssl_raw_tap(openssl.c:2393-2394), which is a data hook, soNewSocket::on_writableandinternal_flushnever run for it and the tail is never sent. net.ts stores the write callback at src/js/node/net.ts:3050 and waits for a drain that never comes, soraw.end()never finishes. The base branch behaves the same; the PR'sendbinding already avoids queuing on BYPASS_TLS (socket_body.rs:3382) butwrite_or_end_buffereddoes not.Verification: pre-existing (the base takes the identical route). Trigger: user code writes to the net.Socket that
tls.connect({ socket })adopted and the kernel takes the send only partly. With an empty queue socket_body.rs:2807 passesbuffer_unwritten_data = truefor every socket, BYPASS_TLS included. Nothing ever flushes that queue on the raw half; the siblingendbinding avoids this at 3381-3382.
|
Three commits since the review above.
Why the test needed that: it found a defect in the TLS write path that is on On the additional finding (a node:net write on the raw half of a |
Stacked on #44291.
Problem
socket.shutdown()sends the FIN at once, before bytes thatend(data)or a node:net write left in the queue. Of 8,388,608 bytes, 2,633,835 arrive, and the peer sees a clean close.NewSocket::shutdown(src/runtime/socket/socket_body.rs) does not read the queue.Fix
shutdown()over a queue with bytes records the FIN inPendingWrites.internal_flushsends it after the last byte.discard_pending_writes. It frees the buffer and keeps the dropped count inbytesWritten, as node does.on_closeandclose()call it too: 344 B remain.test/js/bun/net/socket.test.tsandtest/js/node/net/node-net.test.ts, 6 fail on the base. Notes list the suites.Background
buffered_data_for_node_net) holds bytes that a write accepted and the socket did not take.internal_flushsends them later.Flagsbit for the waiting FIN.Flagsis a fullu16.Downsides
size_of::<NewSocket>grows from 240 B to 248 B. The mimalloc block stays 256 B.bytesWrittenafterterminate()or a failed send now includes the queued bytes that were dropped.Notes
Delivery, debug build, base to this PR
end(data)thenshutdown(), tcpend(data)thenshutdown(), tlssocket._handle.shutdown()Memory of one socket whose peer sent FIN, never read a 16 MiB
end(data), then reset (estimateShallowMemoryUsageOf, release builds)bytesWrittenafternode v26.3.0 also reads 16,777,216 from
bytesWrittenafter a queued write and a reset by the peer.Costs, release builds of the base (
04d1f0645b) and this PR (ca01d4c216).textis 80,673,561 B and 80,672,793 B (768 B less).size_of::<NewSocket<false>>and<true>: 240 B to 248 B (the queue type goes from 32 B to 40 B for one flag byte).mi_usable_size(mi_malloc(n))is 256 for n = 232, 240, 248 and 256, so the heap block does not grow.end()), counted with anLD_PRELOADshim:send1,000,recv1,997,epoll_ctlADD 2,002 / MOD 1,001 / DEL 2,001,close2,006,setsockopt4,001 on both builds. 2,000 connections give twice these counts on both.write_or_endis not changed, and the new check inwrite_or_end_bufferedis after its empty-queue return.shutdown(), one byte load and branch ininternal_flushwhen the socket does not end after the flush, and onereleaseper close (a flag test, a length read, a buffer reset and an add). A send of queued bytes runs throughlend: two byte stores and one load more.lend, the free inclose(), the raw-half test inshutdown()). I did not build release binaries again for them. The struct sizes did not change: the debug-assertions layout reads 520 B forNewSocketand 40 B for the queue type at the final head, as atca01d4c216.perfandvalgrindare not on my machine.Which close frees the queue. Every one:
on_closefor a close that the peer or an error starts, andclose(),terminate()and the other detach sites for a close that bun starts.close()frees it at once, also when a TLS close still waits for ciphertext that the kernel did not take.A release under a send. A send reads the queue through
PendingWrites::lend. The socket can close from inside that send: usockets closes a TLS socket from insideus_internal_ssl_writev("closed from inside the call"), and a Duplex transport runs user JS inside its write. Arelease()in that window only records itself, and the caller frees the queue when the send returned, after its ownconsume. So the slice stays valid for the whole call. The review of #44291 found the missing guard onmainforterminate()under a Duplex send. No test reaches a release under a send: a Duplex whosewrite()destroys the socket during the flush after the handshake gave no ASAN report on the build before this guard either.The raw half of an
upgradeTLS()pair keeps the FIN of the base:shutdown()sends it at once. No writable event reaches that half, so nothing would send a FIN that waits.Other discard sites. Before this PR,
handle_connect_error,close_and_detach,detach_for_reconnect,write_bufferedandend_bufferedon a detached socket, and the two fatal-send paths freed the queue, andbytesWrittenthen lost those bytes. They now keep the count.upgrade_tls_implnow frees the queue of the TCP wrapper that it retires.Tests. Windows takes a first send of any size whole, so the tests fill the kernel with writes first and then queue the chunk under test. I cannot run Windows here, so CI is the check for that. The tail of the
shutdown()test continues the byte stream of its writes, like the retry of a short write does: with other bytes a TLS defect onmainshows (see the comments on this PR).socket.test.ts:shutdown() after end(data)(tcp, tls),a close by the peer frees a queued end(data) tail and keeps bytesWritten,close() on a TLS socket frees a queued end(data) tail at once, while the close waits for the peer,terminate() over a queued end(data) tail keeps the tail in bytesWritten. All 5 fail on the base.node-net.test.ts,a write that is still queued natively:is sent before the FIN when the handle shuts downfails on the base.stays in bytesWritten when the peer resets the connectionpasses on the base and fails ifon_closefrees the queue and does not keep the count.does not hold back the FIN of a shutdown on the raw half ...passes on the base and fails (times out) if the raw half defers its FIN.is_fin_deferred()check inwrite_or_end_buffered: test code cannot call the private$writewhile a FIN waits.Suites on a debug build with this PR:
socket.test.ts244 pass, 6 skip, 1 fail (should not call drain before handshakeneedswww.example.com).node-net.test.ts112 pass, 10 fail, all 10 also on the base (8net.Socket readtests,#13126,unref should exit when no more work pending: no internet). All pass:socket-pending-writes,socket-syscall-fault,tcp-server,net-syscall-fault,node-net-server,node-net-allowHalfOpen,node-tls-connect,node-tls-server,node-tls-raw-end,node-tls-upgrade,tls-syscall-fault,node-tls-socket-allow-half-open-option,node-tls-wrapped-socket-close, and twonode-tls-duplex-*files.Self-review. Eight reviewers read the diff, one per area (FIN state, writes after a shutdown, the byte count, the free in
on_close, node:net, the transports, the tests, the text). They raised 17 concerns, each with a probe. I checked each against the code.close()or node:tlsdestroy()on a TLS socket with a queued tail kept the tail after the close (14,197,112 B), becauseon_closeread the detached wrapper. Fixed:close()frees the queue itself, andon_closefrees it for every transport. New test.shutdown()over a queued tail on the raw half of anupgradeTLS()pair never sent its FIN. Fixed: the raw half sends it at once, as on the base. New test.stays in bytesWritten when destroy()could not fail, and no test read the count after a local close. Fixed: that test is gone, and theterminate()test fails on the base.shutdown()test did not prove that a tail was queued. Fixed: it fills the kernel and asserts it.write()aftershutdown()described the effect ofend(data). Removed.on_closecould run under a live borrow of the queue during a TLS flush, and the fill loop of one test stopped at 16 steps. Fixed:lend, and 64 steps.write()orend(data)after ashutdown()that waits for a node:net write are accepted, where the base returned-1.write_or_endalready asserts in debug builds that the public calls do not run over a node:net tail, and a check for it would add a load and a branch to everywrite().flush()after ashutdown()that waits sends the last byte and the FIN, and nodrainfollows, so the node:net write callback does not run.net.tsnever callsflush(), and the base lost the data in the same call order.Meets #44313. That PR reports a write error when a close loses the queued tail. It reads the queue length at the top of
on_close, where this PR frees the queue. The two merge with a text conflict there. The right order is the check of #44313 first, then the free.Not in this PR
unref()with a queued tail). It is the next step of this stack and needs a maintainer decision on its shape first.timeouthandler.mainthat this work found: a short TLSwrite()can leave one sealed record that it did not count (handed off), and a node:net write on the raw half of atls.connect({ socket })pair queues a tail that nothing sends.bun.d.tssaysshutdown(true)shuts the write side. The code shuts the read side fortrueand the write side for no argument. This PR does not touch that text.no test proof · iteration 3 · platform-specific test(s) that do not run on this machine, deferring to CI, which covers all platforms: test/js/node/net/node-net.test.ts, test/js/bun/net/socket.test.ts