Skip to content

node:http2: destroy the accepted socket when an injected connection's TLS proxy is destroyed - #38154

Closed
robobun wants to merge 1 commit into
mainfrom
farm/42ce50d0/http2-upgrade-destroy-raw-socket
Closed

robobun wants to merge 1 commit into
mainfrom
farm/42ce50d0/http2-upgrade-destroy-raw-socket

Conversation

@robobun

@robobun robobun commented Aug 13, 2026

Copy link
Copy Markdown
Collaborator

Problem

  • A connection handed to an http2.createSecureServer() instance with h2Server.emit("connection", rawSocket) (the pattern used by grpc-js's connection injector, http2-wrapper and similar) is never released once the server side gives up on it: after a failed TLS handshake 'tlsClientError'/'clientError' fire and the peer gets a FIN, but rawSocket.destroyed stays false, rawSocket never emits 'close', the net.Server that accepted it keeps reporting it in getConnections() and netServer.close(cb) never calls back. Node v26.3.0 destroys the accepted socket right away (probe below).
  • Same leak after a successful handshake whenever the server side tears the connection down: session.destroy(), a rejected client certificate (requestCert + rejectUnauthorized), or an 'unknownProtocol' client being destroyed after unknownProtocolTimeout. A peer that keeps its side of the TCP connection open keeps the accepted socket (and the connection count) alive indefinitely.
  • Cause: upgradeRawSocketToH2 (src/js/node/_http2_upgrade.ts) runs the TLS engine over a proxy Duplex. When the proxy is destroyed, tlsSocketDestroy (_http2_upgrade.ts:106) only closed the native TLS handle, and the native side of a duplex upgrade only end()s the underlying stream (UpgradedDuplex::on_close in src/runtime/socket/UpgradedDuplex.rs, call_write_or_end(None)), so the raw socket was half-closed and left waiting for a FIN the peer never has to send. Nothing ever called rawSocket.destroy().

Fix

  • tlsSocketDestroy destroys ctx.rawSocket (if not already destroyed) after closing the native handle. Every teardown of the injected connection ends in the proxy's _destroy: the handshake-failure and certificate-rejection branches of socketHandshake, socketError, Http2SecureServer's own 'tlsClientError' handler, the session's socket teardown, the unknownProtocol timer, and the native close callback (socketClose) that runs when the peer's close_notify or EOF arrives. So the raw socket is released in all of them from one place.
  • This is node's behavior: TLSWrap.prototype.close destroys _parentWrap (the socket the TLSSocket was built on) whenever the TLSSocket's handle closes, however it went down. The raw socket is destroyed without an error, so its 'close' reports hadError === false, as node's _parentWrap.destroy() does.
  • Timing is unchanged, only the final step differs: the proxy is destroyed at the same points as before (and the engine's writes to rawSocket.write() are plain TCP sends that complete before it), and on the graceful paths (session.close(), HTTP/1 fallback res.end(), proxy.end()) the engine only reports close once the peer's close_notify or EOF has arrived, which is also when node destroys its TLSSocket and the socket under it. Checked that the HTTP/1 fallback (allowHTTP1: true) over an injected connection still delivers its response with the change.
  • 'tlsClientError'/'clientError' are still emitted before the proxy is destroyed, so their listeners see the connection intact, as before.
  • Companion to node:tls: destroy the wrapped net.Socket when a TLS wrap is destroyed #38028, which fixes the same leak for the wraps node:net/node:tls create (new TLSSocket(raw), tls.connect({ socket }), tls.Server#emit('connection')); those go through a different mechanism and that change does not reach this path.
  • Tests: test/js/node/http2/node-http2-upgrade.test.mts, new block "the accepted socket is released when the server side goes down": failed handshake with and without a 'clientError' listener (the two destroy sites on that path), session.destroy(err) and session.destroy(), rejected client certificate, unknownProtocol teardown, plus a control that a live session does not release the connection. In each case the peer holds its TCP side open (for the TLS cases the client runs over an in-memory carrier on top of a held-open net.Socket), and the test waits for the accepted socket's 'close', checks getConnections() === 0 and that netServer.close(cb) calls back. The file also re-runs itself under node (existing "tests should run on node.js" case); all new cases pass on node v26.3.0.
  • Without the src/ change (main debug build and released bun 1.4.0) the six teardown cases time out waiting for the accepted socket's 'close'; with it the whole file passes. Also run with the change: the vendored test-http2-socket-close, -client-connection-tunnelling, -autoselect-protocol, -client-proxy-over-http2, -backpressure, -generic-streams(-sendfile), -sensitive-headers, -session-unref, -write-finishes-after-stream-destroy (all exit 0) and grpc-js's (currently describe.todo) connection-injector test, which passes when enabled.

Background

  • Injected connections: Http2SecureServer extends tls.Server, and node lets you feed it already-accepted plain sockets via server.emit("connection", socket); the server then performs the TLS handshake itself. In bun, Http2SecureServer#emit routes a non-TLS socket to upgradeRawSocketToH2, which creates a proxy Duplex that the HTTP/2 session (or the HTTP/1 fallback) treats as its TLS socket, and starts the native TLS engine over the raw socket with upgradeDuplexToTLS. Ciphertext flows raw socket -> engine -> proxy push(), and proxy _write() -> engine -> rawSocket.write(). The proxy's _destroy (tlsSocketDestroy) is the one place every teardown of such a connection passes through.
  • net.Server connection accounting: an accepted socket is counted in server._connections until its _destroy runs, and server.close(cb) emits 'close' (calling cb) only when that count reaches zero. A socket that is only end()ed stays counted.
  • end() vs destroy() on a socket: end() sends a FIN and leaves the socket open to keep reading until the peer also closes; destroy() closes the descriptor and emits 'close' regardless of the peer. Node's TLSWrap.close uses the latter on the wrapped socket.
Probe from the report (peer keeps its side open after the server's FIN; state sampled after 'tlsClientError')
node v26.3.0: ["clientError","tlsClientError","raw-close","client-got-fin","connections=0","raw.destroyed=true","server-close-cb"]
bun 1.4.0:    ["clientError","tlsClientError","client-got-fin","connections=1","raw.destroyed=false","TIMEOUT"]
bun + fix:    ["clientError","tlsClientError","raw-close","client-got-fin","connections=0","raw.destroyed=true","server-close-cb"]

The same probe with a TLS peer that completes the handshake (carrier over a held-open TCP socket) and the server calling session.destroy(err) / session.destroy(), rejecting the client certificate, or turning away a client without ALPN: node and bun + fix release the connection; bun 1.4.0 and main time out in every case. (The error code bun reports for the non-TLS peer, UNABLE_TO_GET_ISSUER_CERT where node reports ERR_SSL_HTTP_REQUEST, is unchanged here and tracked separately.)

… TLS proxy is destroyed

Http2SecureServer.emit("connection", rawSocket) runs the TLS engine over a
proxy Duplex (_http2_upgrade.ts). When that proxy was destroyed (failed
handshake, rejected client certificate, session destroy, unknownProtocol
timeout) the native close only end()ed the raw net.Socket, so it stayed
open, and counted by the net.Server that accepted it, until the peer
happened to close its side. Destroy it from the proxy's _destroy, like
node's TLSWrap destroys the socket it wraps.
@coderabbitai

coderabbitai Bot commented Aug 13, 2026

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

@robobun, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 29 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 70a6fcd2-ee28-4da0-a0bb-e4ef03368b51

📥 Commits

Reviewing files that changed from the base of the PR and between b7a0431 and 74d59f8.

📒 Files selected for processing (2)
  • src/js/node/_http2_upgrade.ts
  • test/js/node/http2/node-http2-upgrade.test.mts

Comment @coderabbitai help to get the list of available commands.

@robobun

robobun commented Aug 13, 2026

Copy link
Copy Markdown
Collaborator Author

Status

  • Reproduced on released bun 1.4.0 and on a main debug build with the probe from the report (h2.emit("connection", raw) + a non-TLS peer that keeps its side open): connections=1, raw.destroyed=false, netServer.close(cb) never fires. Node v26.3.0 releases the connection. The same holds after a successful handshake for session.destroy(), a rejected client certificate and an unknownProtocol teardown.
  • Fix: src/js/node/_http2_upgrade.ts, the proxy's _destroy now destroys the raw socket (this PR).
  • Tests: test/js/node/http2/node-http2-upgrade.test.mts, six teardown cases time out without the fix and pass with it; the file also passes under node.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this PR and didn't find any bugs. The source change is small and well-justified against Node's TLSWrap behavior, and the new tests cover the teardown matrix thoroughly (including a Node cross-check). Given it changes socket-lifecycle ordering in the TLS/HTTP2 injection path, a quick human look would still be worthwhile.

What was reviewed:

  • tlsSocketDestroy now destroys ctx.rawSocket after h.close(); checked the double-destroy guard covers the existing catch block and the socketClose → destroy() re-entry.
  • Verified onTlsClose listener cleanup and the events[3] raw-socket 'close' path don't race into a second destroy — both hit .destroyed guards.
  • Tests: await real conditions (raw socket 'close', getConnections, close(cb)), port: 0, finally cleanup, control case for a live session, and re-run under Node.
Extended reasoning...

Overview

This PR touches two files: a ~9-line addition to tlsSocketDestroy in src/js/node/_http2_upgrade.ts that destroys the underlying raw socket when the TLS proxy Duplex is torn down, and a new 170-line describe block in test/js/node/http2/node-http2-upgrade.test.mts covering seven teardown variants (failed handshake ×2, session.destroy() ×2, rejected client cert, unknownProtocol timeout, and a live-session control). The fix addresses a leak where an injected connection's accepted net.Socket is only end()ed (never destroy()ed) when the server side gives up, so net.Server keeps counting it and close(cb) never fires.

Security risks

None identified. The change tightens resource release rather than loosening any check; no auth, crypto, or validation logic is altered. The only new behavior is calling .destroy() on a socket the server already owns and has decided to tear down.

Level of scrutiny

Medium-high. The source diff is tiny and the mechanism (mirror Node's TLSWrap.close → _parentWrap.destroy()) is clearly argued, but it lives in the TLS/HTTP2 teardown path where ordering matters: h.close() triggers the native side to end() the raw socket (close_notify + FIN), and immediately after we now destroy() it. The PR description asserts the engine's outbound writes are plain TCP sends that complete before the destroy, and the vendored Node http2 tests plus the file's Node re-run all pass, so the risk of truncating a final GOAWAY/close_notify appears to have been considered and empirically checked. Still, this is exactly the kind of lifecycle change a maintainer familiar with UpgradedDuplex should glance at.

Other factors

  • No CODEOWNERS entry covers these paths.
  • No prior human or bot review comments beyond a CodeRabbit rate-limit notice.
  • I traced re-entrancy: the pre-existing catch block calls rawSocket.destroy(e) before tlsSocket.destroy(e), and socketClose (native close, possibly triggered by the raw socket's own 'close' via events[3]) calls this.destroy() — in both cases the new !rawSocket.destroyed guard or the proxy's own destroyed flag prevents a second destroy.
  • Tests meet the repo's bar: they await observable conditions rather than sleeping, use port: 0, hold the peer's TCP side open via allowHalfOpen + an in-memory carrier so the fix is the only way the accepted socket can close, include a negative control, and are verified to time out on the unfixed build and pass under Node v26.3.0.
  • The PR is a companion to #38028 (same leak in the node:tls wrap path), which is a design-level detail a human should be aware of when landing.

Deferring rather than approving because socket-lifecycle changes in TLS/HTTP2 warrant a human sign-off even when the diff is small and the tests are strong.

@robobun

robobun commented Aug 13, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 5:05 AM PT - Aug 13th, 2026

❌ @robobun, your commit 74d59f8 has 1 failures in Build #94562 (All Failures):


🧪   To try this PR locally:

bunx bun-pr 38154

That installs a local version of the PR into your bun-38154 executable, so you can run:

bun-38154 --bun

@Jarred-Sumner

Copy link
Copy Markdown
Collaborator

Superseded by #44618, which consolidates the open TLS pull requests. This fix and its tests are in there, either as written, rewritten smaller, or merged with the other PRs that patched the same cause (see the "By area" list in that PR). Closing in favor of it.

Jarred-Sumner added a commit that referenced this pull request Oct 6, 2026
…cted connection (#38154)

Closing the proxy's native handle only end()s the raw socket, so a peer that
never closes kept the fd and a slot in the net.Server after a failed handshake,
a destroyed session, a rejected certificate or an unknown protocol.
Jarred-Sumner added a commit that referenced this pull request Oct 7, 2026
…cted connection (#38154)

Closing the proxy's native handle only end()s the raw socket, so a peer that
never closes kept the fd and a slot in the net.Server after a failed handshake,
a destroyed session, a rejected certificate or an unknown protocol.
Jarred-Sumner added a commit that referenced this pull request Oct 8, 2026
…cted connection (#38154)

Closing the proxy's native handle only end()s the raw socket, so a peer that
never closes kept the fd and a slot in the net.Server after a failed handshake,
a destroyed session, a rejected certificate or an unknown protocol.
Jarred-Sumner added a commit that referenced this pull request Oct 8, 2026
… it still has to send

The close of a TLS socket destroyed the stream it wraps, as in node. There every
write has completed by then. Here a write over a stream completes once the
stream has taken the ciphertext, so the destroy dropped what the stream had yet
to send: when the peer closed its side first, 4 of 8 MiB arrived with TLS in
TLS and 4 of 32 MiB on the http2 emit('connection') path.

Until the verdict on the peer lets the session through, the application cannot
have written over it. There the stream is still destroyed, and so are the
sessions below it, which releases the connection of a failed handshake, of a
rejected certificate and of an early destroy(). The same holds for an http2
socket that the application never got, and for resetAndDestroy().

After that, as before this branch:
- a net.Socket only gets the end() of the engine. What left it at 'finish' is
  still the kernel's to send, and a close over unread input drops that. It
  closes at the FIN of its peer. Until then it keeps its own timeout, which the
  TLS socket cleared through _parent, and reports its own errors, which went to
  the TLS socket that is now gone.
- any other stream is destroyed with the TLS socket.

A wrapped socket whose handle is gone hands the error to its TLS socket and
closes by itself. The engine resumes the stream at its close only if it paused
it, so that a paused net.Socket reads the FIN.

Two http2 tests of #38154 need a write that completes as in node and are todo:
a client that never closes its side holds the accepted socket of a session that
the server destroyed, or of a client it turned away with a 403.
Jarred-Sumner added a commit that referenced this pull request Oct 8, 2026
… it still has to send

The close of a TLS socket destroyed the stream it wraps, as in node. There every
write has completed by then. Here a write over a stream completes once the
stream has taken the ciphertext, so the destroy dropped what the stream had yet
to send: when the peer closed its side first, 4 of 8 MiB arrived with TLS in
TLS and 4 of 32 MiB on the http2 emit('connection') path.

Until the verdict on the peer lets the session through, the application cannot
have written over it. There the stream is still destroyed, and so are the
sessions below it, which releases the connection of a failed handshake, of a
rejected certificate and of an early destroy(). The same holds for an http2
socket that the application never got, and for resetAndDestroy().

After that, as before this branch:
- a net.Socket only gets the end() of the engine. What left it at 'finish' is
  still the kernel's to send, and a close over unread input drops that. It
  closes at the FIN of its peer. Until then it keeps its own timeout, which the
  TLS socket cleared through _parent, and reports its own errors, which went to
  the TLS socket that is now gone.
- any other stream is destroyed with the TLS socket.

A wrapped socket whose handle is gone hands the error to its TLS socket and
closes by itself. The engine resumes the stream at its close only if it paused
it, so that a paused net.Socket reads the FIN.

Two http2 tests of #38154 need a write that completes as in node and are todo:
a client that never closes its side holds the accepted socket of a session that
the server destroyed, or of a client it turned away with a 403.
steipete added a commit to openclaw/bun that referenced this pull request Oct 8, 2026
Match Node 24 destruction for Duplex-backed TLS while keeping graceful
shutdown separate. Retain adopted-fd close ownership and release HTTP/2
injected transports when their TLS proxy is destroyed.

Adapt HTTP/2 lifecycle coverage from oven-sh#38154.

Co-authored-by: robobun <117481402+robobun@users.noreply.github.com>
steipete added a commit to openclaw/bun that referenced this pull request Oct 8, 2026
Match Node 24 destruction for Duplex-backed TLS while keeping graceful
shutdown separate. Retain adopted-fd close ownership and release HTTP/2
injected transports when their TLS proxy is destroyed.

Adapt HTTP/2 lifecycle coverage from oven-sh#38154.

Co-authored-by: robobun <117481402+robobun@users.noreply.github.com>
steipete added a commit to openclaw/bun that referenced this pull request Oct 8, 2026
* fix(tls): destroy wrapped transports without ending them

Match Node 24 destruction for Duplex-backed TLS while keeping graceful
shutdown separate. Retain adopted-fd close ownership and release HTTP/2
injected transports when their TLS proxy is destroyed.

Adapt HTTP/2 lifecycle coverage from oven-sh#38154.

Co-authored-by: robobun <117481402+robobun@users.noreply.github.com>

* fix(tls): preserve wrapped transport ownership and half-close ordering

Keep adopted-fd ownership on the TLS socket after its raw handle detaches.
Separate peer EOF, graceful writable shutdown, and full TLS destruction;
wait for transport completion without forwarding its error a second time.
Use the inherited tls.Server connection path for injected HTTP/2 sockets
instead of maintaining a second TLS transport adapter.

Port HTTP/2 destroy-versus-close and final event ordering from
oven-sh#38195, and the destroyed-socket EOF guard from oven-sh#43392.
Retain the six-case destruction regression and add Node 24 controls for
half-open replies, raw EOF, renegotiation shutdown, and transport ownership.
Synchronize conformance assertions with the sessionError event itself.

Co-authored-by: robobun <117481402+robobun@users.noreply.github.com>

* fix(tls): flush half-open writes after peer shutdown

Continue draining encrypted output after close_notify while the transport remains open. Add a Node 24 parity case that writes outside the receive callback and waits for peer receipt before ending, so shutdown cannot mask a missing flush.

---------

Co-authored-by: robobun <117481402+robobun@users.noreply.github.com>
Jarred-Sumner added a commit that referenced this pull request Oct 10, 2026
…cted connection (#38154)

Closing the proxy's native handle only end()s the raw socket, so a peer that
never closes kept the fd and a slot in the net.Server after a failed handshake,
a destroyed session, a rejected certificate or an unknown protocol.
Jarred-Sumner added a commit that referenced this pull request Oct 10, 2026
… it still has to send

The close of a TLS socket destroyed the stream it wraps, as in node. There every
write has completed by then. Here a write over a stream completes once the
stream has taken the ciphertext, so the destroy dropped what the stream had yet
to send: when the peer closed its side first, 4 of 8 MiB arrived with TLS in
TLS and 4 of 32 MiB on the http2 emit('connection') path.

Until the verdict on the peer lets the session through, the application cannot
have written over it. There the stream is still destroyed, and so are the
sessions below it, which releases the connection of a failed handshake, of a
rejected certificate and of an early destroy(). The same holds for an http2
socket that the application never got, and for resetAndDestroy().

After that, as before this branch:
- a net.Socket only gets the end() of the engine. What left it at 'finish' is
  still the kernel's to send, and a close over unread input drops that. It
  closes at the FIN of its peer. Until then it keeps its own timeout, which the
  TLS socket cleared through _parent, and reports its own errors, which went to
  the TLS socket that is now gone.
- any other stream is destroyed with the TLS socket.

A wrapped socket whose handle is gone hands the error to its TLS socket and
closes by itself. The engine resumes the stream at its close only if it paused
it, so that a paused net.Socket reads the FIN.

Two http2 tests of #38154 need a write that completes as in node and are todo:
a client that never closes its side holds the accepted socket of a session that
the server destroyed, or of a client it turned away with a 403.
Jarred-Sumner added a commit that referenced this pull request Oct 10, 2026
…ps, WebSocket, SQL) (#44618)

### What does this PR do?

Consolidates the open TLS pull requests into one. Each was reproduced on
`main` and, for `node:*` behavior, on Node v26.3.0 first. About a third
are ported as written, the rest are rewritten smaller or merged into one
fix where several PRs patched the same cause. One commit per fix, so it
can be read commit by commit.

Fixes #43520, fixes #31396, fixes #43635, fixes #37193, fixes #43846,
fixes #17932, fixes #41061, fixes #36887, fixes #31810, fixes #35240,
fixes #32234, fixes #44365, fixes #43807, fixes #42280, fixes #44517.
Addresses #41856 (SNI and `servername`; not `checkServerIdentity` for
SQL), #24845 (the spin is gone, shown with fault injection on Linux; not
run on macOS), #19754 (node-fetch forwards the agent's TLS options; the
Kubernetes client itself was not run).

#### The ones that matter most

| | On `main` | PRs |
|---|---|---|
| Client certificate disclosure | `https.request()` with a client
certificate sends it to a server it then refuses (wrong name,
`checkServerIdentity`, `destroy()` in `'secureConnect'`, `terminate()`
in `handshake`). A server can force it with a junk record behind its
Finished | #43946 |
| False `authorized` | Over a Duplex, `secureConnect` with `authorized
=== true` for a peer that failed the key proof; `secureConnect` for a
plaintext peer with `rejectUnauthorized: false` | #44422, #32929 |
| Cleartext https | `https.createServer()` without a usable key/cert
answers plain HTTP | #41672, #33539 |
| Revoked client certificates | An https mTLS server never sees `crl`,
so a revoked client is `authorized` | #41641 |
| Pooled sockets | Requests with different client certificates or CAs
share an `https.Agent` socket and session | #42498 |
| Silent plaintext | `tls: [...]` given to `Bun.listen` / `Bun.connect`
is plain TCP | #41490 |
| Server weakened by a client knob | `NODE_TLS_REJECT_UNAUTHORIZED=0`
turns off a server's client-certificate enforcement | #35245 |
| Pins never checked | `WebSocket` never calls `tls.checkServerIdentity`
and ignores `tls.serverName` | #41648 |
| `verify-full` dropped | `PGSSLMODE=verify-*` is lost next to a `TLS_*`
URL variable; `tls: true` sends no SNI | #44498 |
| Crashes | use-after-free from `destroy()` in `ALPNCallback` over a
Duplex; `abort()` on a late `setSession()`; SIGABRT in `fetch` with an
https proxy from the environment and a `Bun.file()` body | #44462,
#41671, #44458 |
| Stream corruption | A TLS `write()` can lose 16 KiB it reported as
written while another socket on the loop is stalled | #44529 |
| Hangs and spins | 100% CPU on a failing `send()`; a fatal `SSL_write`
leaves the socket open forever; `idleTimeout` never sheds a TLS client
that ignores `close_notify` | #34510, #38176, #42336 |
| Wrong certificate (regression since 1.3.14) | Connections accepted
before `stop()` / `close()` get the default certificate and skip their
entry's `requestCert` / `ca` | #42355 |
| Quadratic Duplex / proxy tunnel | Reading one chunk over a Duplex,
CPU: 8 MB 0.88 s → 0.14 s, 16 MB 3.18 s → 0.23 s, 32 MB 11.75 s → 0.39
s; `fetch` upload through CONNECT: 1.7 s → 0.18 s (debug build) | #44464
|

#### By area

- **fd engine, write path** (`openssl.c`, `socket.c`): #42352, #34510 +
#38176 + #42336 as one change, #44529, #44458, #44192. A rejected
`send()` ends the write side only and closes at the next writable event
unless the peer's bytes are still queued (a 413 sent before a reset is
still read). No new per-socket state. Also, on kqueue, **a FIN no longer
ends a socket that waits in the low-priority queue** (`loop.c`): with
more than 5 TLS handshakes at once, a client that ended right after its
handshake could be reset and its server socket report `socket hang up`,
because the eof that the sentinel read knote reports was acted on ahead
of the unread Finished. That is on `main` too (the macOS entry for
`node-tls-server.test.ts` in `test/flaky-tests.txt`: 7 of 48 recent
builds of other branches), and this branch made it likelier (6 of 8
builds), since Finished now leaves in one segment with the close_notify.
- **Error reporting, both engines**: #44422, #32929, #44516, #37094,
#41272 + #42324 + #44223 as one change, #44021, #37472, #43946, #33630.
One channel: a fatal error on an established session is reported, then
**the engine closes the connection itself**, whatever the owner does
with the report. `test/js/bun/net/tls-fatal-error-closes.test.ts`
asserts closed-and-nothing-delivered for every owner (node:tls,
`Bun.connect`, `Bun.listen`, `fetch` direct and through CONNECT,
`Bun.serve`, `WebSocket` direct and through a proxy, Postgres, MySQL,
Valkey, Duplex).
- **Duplex engine** (`SSLWrapper`, `UpgradedDuplex`): #44462, #43529,
#42332, #44464. #43877 + #44394 were in and are **out again**, see
"Worth a look" 5.
- **node:tls wrap lifecycle** (`net.ts`, `tls.ts`): #38007, #38058,
#38028 + #38122 + #38076 as one change (six copies of the attach code
become two helpers), #38311, #39008, #38154, #42340 + #42343 + #42339 +
#42453 as one change, #43791, #42425, #44085, #42683, #39088, #39040,
#40375, and what was still real of #36534.
- **SNI, ALPN, server contexts**: #43080, #42050, #37195 + #43849 as one
change (**one** SNI matcher for TCP and HTTP/3), #42355, #42285, #33253,
part of #37896, part of #37013. A `tls.Server` has one `SSL_CTX`.
- **Verification and options**: #44738, #41490, #37005 + the cwd pin of
#40984, #31811, #43982, #33483 + #35245, #41810, #32235, #44441, #38092.
- **node:tls API and CA store**: #41671, #38145, #32824, #43594, #39997,
#41696, #33534, #34748, #42991, #42996, #42970.
- **node:https, Agent, `ws`, node-fetch**: #41672 (https half), #41641,
#38261, #42498, #44346, #35609, #31397, #42325.
- **WebSocket client**: #41648, #37487 + #43048 as one change.
- **SQL, Redis**: #33666, #41711, #44498, part of #42054.
- **Tests only**: #41426, #40040, #44395, #44016, #37860, #40591,
#44440, #41424.

Found on the way and fixed here: an upload that a TLS 1.2 server
interrupts with a renegotiation never completes on `main` (0 of 32 runs
over `https.request`, `fetch`, `node:tls` and `Bun.connect`: the
renegotiation ClientHello lands inside an application record that is
still unsent, or the socket gets no `drain` again) and completes here,
with two tests from robobun; the fix for #40653 (final flight and first
write in one segment) stopped working whenever another TLS socket on the
loop was stalled, on `main` too; the `tls.Server` prototype pinned the
last server constructed and every `SSL_CTX` it owned;
`Object.create(process.env).NODE_TLS_REJECT_UNAUTHORIZED = "0"` turned
verification off process-wide once a `SHARE_ENV` worker existed; two
debug panics when wrapping a shut-down or still-connecting socket; a
`fetch` POST through a proxy sent its headers twice when the origin
renegotiated; `BlockList` ignored IPv6 zone ids; a test now ties
`root_certs.der` to `certdata.txt`.

#### Behavior changes

- **A server's `ca` without `requestCert` no longer asks for a client
certificate** (`Bun.serve`, `Bun.listen`, HTTP/3, node:tls). It matches
the docs and Node. On `main` such a server refused clients with no
certificate but served any unrelated self-signed one, so it was never
authentication. **Set `requestCert: true` to require a certificate.** A
matrix test pins that `requestCert: true` still refuses no certificate
and an untrusted one on 8 kinds of server, TLS 1.2 and 1.3, with
`NODE_TLS_REJECT_UNAUTHORIZED` unset and `0`.
- `NODE_TLS_REJECT_UNAUTHORIZED=0` no longer relaxes a server.
- `Bun.connect` / `Bun.listen` hear of a fatal TLS error after the
handshake through `error(socket, err)`. With no `error` handler the
socket just closes.
- HTTP/3 server names match like TCP: `*.` covers exactly one label,
case is ignored, a trailing dot is ignored, the last registration of a
name wins.
- `requestCert` on node:https is `=== true`, as in Node.
- An array where a generated options dictionary is expected throws
(`tls: []`, `jest.useFakeTimers([])`).
- `key` / `cert` arrays serve every identity. A client that can use both
gets ECDSA, where `main` served whichever pair came last.
- `ecdhCurve` is forwarded by node:https, `ws` and node-fetch now, so a
group BoringSSL lacks (`X448`) throws there as it already does in
`tls.createServer`.
- A wrapped socket's error is re-emitted on the TLS socket as in Node,
so `raw.destroy(err)` with a listener on `raw` only is uncaught, as in
Node.
- `sql.options.tls` is always an object, never `true`. `RedisClient`
sends SNI.
- `tls: { secureContext }` alone asks for TLS on `Bun.listen` /
`Bun.connect` (it was plain TCP), and a value that is not a
`SecureContext` throws. The context is served as it is: the
`requestCert` / `rejectUnauthorized` it was created with hold whatever
the options next to it say, and `requestCert` in the options over a
context that does not ask throws at `listen()`.
- `tls.DEFAULT_CIPHERS` reaches every client once assigned (`fetch`,
`WebSocket`, `Bun.connect`, `RedisClient`, `Bun.SQL`, `S3Client`, proxy
tunnels) and servers again. A list that selects no cipher throws
`ERR_SSL_NO_CIPHER_MATCH` at the assignment. `fetch.preconnect()` dials
nothing after an assignment.
- The warning for an unreadable `NODE_EXTRA_CA_CERTS` is Node's one
line, without the `warn:` prefix.
- `BUN_CONFIG_WS_CLOSE_TIMEOUT` (default 30 s): how long a `WebSocket`
client waits for the server to close the connection after the closing
handshake.

#### Worth a look in review

1. **#44529**: the kernel-refused remainder of a TLS write moves from
the loop's one slot onto the connection (in the existing rare struct),
so the write BIO never refuses a sealed record. Nothing is allocated on
an unstalled path (200 writes: 0 appends, same `send()` count as
`main`), memory with 16 stalled writers is lower than on `main` (276 KB
vs 340 KB, which `main` holds inside BoringSSL's buffers), `us_socket_t`
stays 80 bytes. It needs a bound on how long a deferred close waits, or
a peer that stops reading pins the fd past `destroy()`:
`US_SSL_CLOSE_AFTER_SPILL_TIMEOUT` is a fixed 10 s, not re-armed on
progress. Separate commits, but the fix that keeps the client
certificate off the wire beside a stalled socket builds on them.
2. **The default name check of node:tls also runs inside the
handshake**, so a wrong-name server gets no client certificate on TLS
1.2 either. JS still runs it after every successful handshake, so a
difference between the two matchers can only refuse. Error objects are
byte-identical.
3. **#44441** widens trust by design: a self-issued leaf whose
`keyUsage` lacks `keyCertSign` (`dotnet dev-certs`) is its own anchor
when the store holds a byte-identical copy. No BoringSSL change. Expired
pin, same subject with another key, wrong EKU and a pinned intermediate
are tested to fail.
4. **#32235** only adds Ed25519 and ECDSA P-521 to the verify list. A
captured ClientHello shows `main`'s list with the two inserted;
`rsa_pkcs1_sha1` stays.

5. **A stream that a TLS socket wraps, when that TLS socket closes.** An
earlier state of this branch lost data here while CI was green (found by
#44709's report): with the peer closing first, 4 of 8 MiB arrived with
TLS in TLS, 4 of 32 MiB on the http2 `emit("connection")` path, and a
`write()` with no `'error'` listener ended the process. Three
Node-parity changes only hold together: destroying the wrapped stream at
the close (#38028 + #38122 + #38076, #38154) is safe only if every write
has really completed (#43877), which in turn needs Node's handling of
the peer's close_notify, which needs half-open sockets that the GC can
collect. So:
- #43877 + #44394 are reverted and reopened. A write over a stream
completes once the stream has taken the ciphertext, as on `main`.
- Until the verdict on the peer lets the session through, the
application cannot have written over it. There the wrapped stream is
destroyed as in Node, with the sessions below it. That keeps the release
of the connection after a failed handshake, a rejected certificate and
an early `destroy()`. The same for an http2 socket the application never
got, and for `resetAndDestroy()`.
- After that it is `main`'s teardown: a `net.Socket` only gets the
engine's `end()`, closes at its peer's FIN, keeps its own timeout and
reports its own errors. Any other stream is destroyed with the TLS
socket.

The regular suites cannot see any of this (999 files were green on every
broken variant), so it was steered by eleven seeded differential fuzzers
run on this build, `main`, Node v26.3.0 and the earlier state: close,
`end()`, `destroy()`, `destroySoon()`, resets, hung and half-open peers,
paused writers, timeouts, two and three sessions deep, over TCP and over
Duplexes, before, at and after the handshake, and http2 requests. See
"How did you verify".

#### Known limits

- `fetch` with a `checkServerIdentity` function still sends the client
certificate (not the request) to a server the function refuses. On TLS
1.2 any verdict a JS callback gives is too late, as in Node.
- `addContext()` / `SNICallback` still do not apply to a server-side
socket on the stream engine (`emit("connection", duplex)`, TLS in TLS,
unflushed writes, named pipes), as on `main`.
- A CA bundled in a pfx extends an explicit `ca` only, for `ws` /
node-fetch / `WebSocket`: the native `ca` can only replace the default
store, and that store keeps `SSL_CERT_FILE` / `SSL_CERT_DIR`.
- P-521 leaves work on TLS 1.3 only. TLS 1.2 needs secp521r1 in every
ClientHello (`it.todo`).
- Once `tls.DEFAULT_CIPHERS` is assigned, `fetch(url, { protocol:
"http3" })` is `HTTP3Unsupported`, as with an explicit `ciphers`.
- `addCACert()` by hand does not extend the chains of a context with
several identities.
- A throwing `ALPNCallback` sends `no_application_protocol` on both
engines. Node sends nothing and its client sees `ECONNRESET`.
- TLS in TLS, peer FIN while the outer handshake runs: the inner socket
gets one `write EPIPE`, where Node gives `ECONNRESET` (`main` gives it
no error at all).
- `@SECLEVEL` in `ciphers` is dropped by the `ws` / node-fetch shims,
which used to ignore `ciphers`. node:tls keeps throwing
`ERR_SSL_INVALID_COMMAND`.
- Beside a stalled TLS socket only the first record (16 KiB) of the
first write leaves with the handshake flight. The rest goes record by
record, which is what bounds the memory of stalled writers.
- After a fatal error on an established session the socket emits
`'error'` and then `'close'`. Node emits `'error'` and leaves the socket
open.
- A paused reader whose own write the kernel rejects loses what it had
not read yet, with an `EPIPE`, as on Node. `main` reports no error there
and delivers it.
- On `main` too: a `Bun.listen` socket without `allowHalfOpen` that has
unsent ciphertext when the client's `shutdown()` arrives loses that
ciphertext (32 KiB), and over plain TCP `end()` with the peer still
sending is a close over unread input, so a reset.
- Differences from both `main` and Node that the differential runs below
found and that stay, all with a peer that aborts: `ECONNRESET` instead
of a clean `'end'` after the socket's own `'finish'` when the peer
destroyed with unread data; under TLS 1.2, a zero-length `write()`
followed by `destroy()` in `'secureConnection'` leaves the client
without `'secureConnect'` (a plain `destroy()` there matches Node); a
TLS 1.2 client that destroys in `'secureConnect'` gets no `'session'`; a
`ClientRequest` whose handshake fails with an alert emits `'error'` and
`'close'` but no `'finish'` (`writableFinished` is true).
- Once `tls.DEFAULT_CIPHERS` is assigned, `fetch.preconnect()` opens
nothing: `fetch()` then uses a context of its own, and a socket warmed
under the default one would never be picked up.
- A TLS `send()` that the kernel refuses outside a `write()` call (the
drain of unsent ciphertext) is reported with the close, as `read EPIPE`
/ `read ECONNRESET`. Node says `write EPIPE`. `main` does not report it
at all.
- Once the application has a session over a `net.Socket` (TLS in TLS,
http2 `emit("connection")`), a peer that never sends its FIN holds that
socket after the TLS socket closed, as on `main`. Node destroys it. Two
tests of #38154 are `todo` for this. Closing it any earlier (at its
`'finish'`, say) makes the kernel drop what it has not sent yet as soon
as the peer's close_notify arrives.
- Plaintext that was queued on a socket before it was wrapped (STARTTLS
with a backlog) is dropped when the TLS socket is destroyed, or its
handshake fails, before the session is accepted. Node drops it too,
except on `destroySoon()`. `main` sends it.
- Over a stream that is no `net.Socket`, `end()` can still cut what that
stream has buffered, and there is no backpressure, both as on `main`
(#43877).
- `tls.secureContext` (the undocumented door node:tls uses) is not read
by a Windows named pipe listener, which builds its context from the
options. On `upgradeTLS({ isServer: true })` the options next to it are
the policy, as with Node's `SetVerifyMode`.
- `selectServerName()` rebuilds the name tree per ClientHello for
injected sockets of a server with `addContext()` entries: 0.4 µs for 1
entry, 3.7 µs for 10, 41 µs for 100, against 631–1111 µs for a
handshake.

#### Not included

Left open, because they need a decision or are not TLS: #43877 + #44394
(see "Worth a look" 5; #43874 stays open with them), #38548, #38591
(both shrink who is trusted), #41589 (`verify-full` vs
`NODE_TLS_REJECT_UNAUTHORIZED=0`), #37197, #41706, #43216, #33487,
#33545, #36707, #32435, #37255, #28691, #40275, #30314 (features),
#38120 (needs the BoringSSL fork, as did #33517, which the stale bot has
closed since), #38529 (needs a Windows measurement), #34342, #38232,
#43089, #44454, #40451, #42710, #44527, #38088, #38093, #41898. #37896,
#42054 and #37013 stay open for the halves not taken.
`http.createServer({ key, cert })` keeps serving TLS on purpose.

One open question: `tls: {}` (an object that names no TLS option) is
plain TCP on `Bun.listen` / `Bun.connect`, here and on `main`. It is the
same trap as `tls: []`, but changing it changes a Bun default, so it is
left alone.

### How did you verify your code works?

- Every new test fails on `main` for the stated reason and passes here,
except guards that pin existing behavior, each shown to fail when its
clause is removed. `node:*` tests also pass on Node v26.3.0; the few
that cannot say which Node version has the behavior.
- 212 test files that touch TLS, sockets, http, http2, fetch, WebSocket,
SQL, Valkey and workers: 6154 pass, 2 fail. Both are seen on `main` too:
`serve.test.ts` "root range port" (the box runs as root), and
`worker_threads.test.ts` "terminate(): nothing of the worker's runs
after the request", which is flaky there and passed in the run below.
- 58 of those files the way the ASAN lane runs them (LeakSanitizer +
`BUN_JSC_validateExceptionChecks`): 58 files, 48 of them with leak
checking, 4132 pass, 3 fail. All three also fail on `main`:
`serve.test.ts` "root range port", `node-net.test.ts` "should not leak
when connect({path}) fails synchronously on a reused handle" (times out
under this environment), `worker_threads.test.ts` "process.exit() with a
shell cp in flight" (a `ShellCpTask` leak).
- 647 vendored `test-tls-*`, `test-https-*`, `test-net-*`,
`test-http2-*`: the only two failures also fail on `main`.
- The SNI matcher was diffed against both old matchers: 3 seeds × 1.23 M
lookups × 3 registration flavours, every difference in one of the
intended classes, TCP and HTTP/3 identical on every lookup.
- The headline rows were also driven by hand with scripts against this
build, `main` and Node v26.3.0: cleartext https, `crl`, `tls: []` / `{
secureContext }`, the `ca` / `requestCert` matrix, the client
certificate on a wrong-name server, late `setSession()`, `destroy()` in
`ALPNCallback`, `[rsa, ec]` identities with an intermediate from `ca`,
`WebSocket` `checkServerIdentity`, a corrupted record, the Duplex read
above, `tls.DEFAULT_CIPHERS`.
- The `setSession()` guard was checked against the real `abort()` at 43
handshake states.
- `bun run rust:check-all`: 12 of 12 targets. `tsc`, oxlint, source
lints, prettier, rustfmt, mordant clean.
- usockets' `_Nonnull` is compiled out of debug builds, so 105 of those
files were also run on a local release ASAN build with the CI runner's
environment (92 with leak checking): 4595 pass, 1 fail,
`child_process.test.ts` "spawn reports EPERM after dropping privileges",
which cannot pass as root and fails on `main` too.
- The close of a TLS socket over another stream ("Worth a look" 5):
eleven seeded differential fuzzers, 8,424 scenarios compared, each run
on a release ASAN build of this branch, on `main`, on Node v26.3.0 and
on the earlier state of the branch. Against `main`:
- Data that `main` delivers in full is cut in 5 scenarios, and about 150
that `main` cuts arrive in full. Of the 5, in 2 `main` never notices the
peer's close and keeps the socket for good, 2 call `end()` on the middle
one of three sessions over an in-memory Duplex, and 1 does the same on
Node.
- No dead timeout, no silent reset and no uncaught error that `main`
does not have (4 uncaught errors fewer).
- A socket stays open where `main` closes it in 109, and closes where
`main` keeps it in 295. 92 of the 109 do the same on Node or on the
earlier state (a `destroy()` that an in-memory Duplex does not show its
peer, half-open peers). 14 wait for a peer that paused reading and so
does not read the FIN (#42332's backpressure, as in Node); the socket's
own timeout fires there. 3 are left: one on a 5 ms timer, two with three
sessions over an in-memory Duplex.
- The earlier state of the branch cut data in 173 of the 400 scenarios
of one of them, where `main` cuts none and this cuts none.
- 23 new tests pin what they found. Each earlier attempt at this fix
fails the ones that describe it, the earlier state of the branch fails
7, and all pass on Node.
- After that change: 999 test files on the release ASAN build (20,246
pass; the 11 files that fail need a database, Docker, DNS or a non-root
user, or share a temp directory with a parallel run and pass alone), 61
on the debug build.
- TLS over a file descriptor (`openssl.c`, the path of `fetch`,
`Bun.serve`, `tls.connect`, `Bun.connect`) got the same treatment after
the rebase: seeded differential fuzzers on CI's release build of this
branch, on `main` and, for `node:*`, on Node v26.3.0. Every runtime also
against itself for the noise floor, injected faults and known bugs of
`main` as positive controls, and a difference counts only if it shows in
5 of 5 fresh processes.
- `node:tls` over TCP: 11,500 scenarios (one connection with Node as the
oracle line by line; 2 to 60 connections beside stalled neighbours; raw
peers that break the handshake). HTTPS: about 136,000 runs over
`Bun.serve` + `fetch`, `node:https`, `node:http2` and `wss://`, also
with the two ends in different runtimes. `Bun.connect` / `Bun.listen` /
`upgradeTLS`: 11,500 scenarios and 720 slow connections, with writers
driven by what `write()` returns, beside up to 6 stalled, dripping,
closing or resetting neighbours, and plain TCP as a second oracle. No
crash, hang, duplication, reordering or silent truncation, and no change
in time or in connection reuse.
- They found six things that `main` does better, none of which any test
showed. All are fixed, each with a test that fails on the build before:
what the peer sent lost behind a rejected `send()` (23 scenarios, and an
early HTTPS response lost with only `EPIPE`), the same silently for a
paused reader, `server.close()` never calling back on a half-open server
after a ClientHello and a reset (17), `closeAllConnections()` taking 12
s with a stalled client, `end()` losing up to 1.3 of 4 MiB that
`write()` had reported while the peer still uploads, and `end()` a
little after a stall never closing beside other stalled TLS sockets. The
last two fixes also deliver the 1 to 2 MiB that `main` loses there, and
close the socket that `main` keeps for good without such neighbours.
- All of them again after every fix, on CI's release build of it. That
caught one regression of a fix itself (a reader stopped for backpressure
lost 86,385 bytes, 1 of 6,000 scenarios), fixed too. On the last build:
scenarios that lose data where `main` does not 23 → 2, and Node loses it
in both, with the same `EPIPE`; `server.close()` that never calls back
17 → 0; connections held 4 → 0; requests that end in an error only where
`main` has a response 6 → 0. With a Node server in another process, a
request ends in an error only in 8 and 10 of 1,500 scenarios here, 5 and
3 on `main`, 10 with Node as the client.
- `Bun.connect` / `Bun.listen` on the last build against `main`, in
scenarios: hangs 0 against 1,031, sockets and fds never released 0
against 965, corrupted data 0 against 345, `abort()` 0 against 26
(`setSession()` after the handshake), writers that never close 0 against
101 of 720 connections. No kind of failure shows here and not on `main`.
About a third of the slow connections close later than on `main`, in 1
to 16 s instead of at once, waiting for unsent ciphertext or for the
peer's close_notify, and 79 more of them deliver all that `write()`
reported. RSS and time with 16 to 256 stalled writers are the same.
- What they found that `main` does worse: a `WebSocket` that calls
`close()` with sends pending loses messages in 81 of 999 scenarios (0
here), 37 server sockets left open, 10 `server.close()` that never call
back, 20 write callbacks that never run.
- The kqueue fix cannot be run on Linux. The `connectionListener` count
test now says what became of a missing connection, which is how the
cause was found (`'tlsClientError'` "socket hang up", then `read
ECONNRESET` at the client of the same port, after its
`'secureConnect'`). On macOS x64 it failed every attempt of the three
builds before the fix and passed at the first attempt of the build with
it.
- Windows and macOS were only run by CI. Four new tests asserted what
only the Linux kernel does (a FIN read ahead of a reset, unread bytes
surviving a reset, loopback buffer sizes, `fstat()` on a socket) and now
say so per platform.

---------

Co-authored-by: autofix-ci[bot] <114827586+autofix-ci[bot]@users.noreply.github.com>
Co-authored-by: robobun <117481402+robobun@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants