Skip to content

tls: queue unsent ciphertext per connection so write() counts every sealed record - #44529

Closed
robobun wants to merge 2 commits into
mainfrom
robobun/abc61b77/tls-write-honest-count
Closed

robobun wants to merge 2 commits into
mainfrom
robobun/abc61b77/tls-write-honest-count

Conversation

@robobun

@robobun robobun commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

  • A Bun.connect / Bun.listen TLS write() can report N bytes while one more sealed record stays inside BoringSSL, not counted. The next write of other data sends that record and loses as many accepted bytes. A smaller next write fails with SSL_R_BAD_WRITE_RETRY and never drains.
  • Cause: while another TLS socket owns the loop's spill slot, BIO_s_custom_write (packages/bun-usockets/src/crypto/openssl.c) refuses a record the kernel does not take. BoringSSL already sealed it, and us_internal_ssl_writev does not count it.

Fix

  • Each TLS connection has a queue for ciphertext the kernel did not take. The write BIO takes every record, so BoringSSL never waits for a retry.
  • Every sealed record is counted and reaches the peer before any later write. The spill slot is gone.
  • Verified: 6 new tests in test/js/bun/net/socket.test.ts (5 fail without the change), and the TLS fault-injection and backpressure suites.

Background

  • usockets sends the records of one TLS write with one send(). The spill slot held the part the kernel refused, one socket per event loop.
  • While the slot was taken, other TLS sockets sent record by record. That rule and its byte bound stay.
  • Considered: count the refused record and retry it from a saved plaintext copy. That covers one of the five SSL calls that use this BIO.

Downsides

  • us_ssl_rare_t grows from 56 to 64 bytes per accepted TLS connection.
  • A close that waits for unsent ciphertext ends after 8 to 12 s without progress. Before, the slot owner waited forever.
  • Not run on macOS or Windows.
Notes

Reproduction (linux x64 loopback, release 1.4.3-canary.1+367d939d9 and a debug build of main)

Two TLS clients on one loop. The first fills the kernel and stays stalled. The second calls write(1 MiB) until a write is short, then writes other data.

Next write main this PR
256 KiB of other bytes the peer gets 16384 bytes that write() did not report, and 16384 accepted bytes never arrive exact
100 other bytes write() returns 0, no drain, the socket is marked fatal with no event sent
same stream after setMaxSendFragment(512) same as the row above sent

Mechanism

  • do_tls_write (BoringSSL ssl/s3_pkt.cc) seals the record before it flushes. When the BIO refuses, it keeps the record and pending_write, and expects a retry with the same bytes.
  • The context sets SSL_MODE_ACCEPT_MOVING_WRITE_BUFFER, so a retry with another buffer of the same size or more flushes the old record and reports that many bytes of the new buffer as written. A smaller buffer fails with SSL_R_BAD_WRITE_RETRY.
  • node:net, node:tls and Bun.serve keep the unsent remainder and retry exactly it, which is what BoringSSL expects.

What changes besides the count

  • SSL_shutdown, SSL_do_handshake and SSL_read write through the same BIO. Their records (close_notify, handshake flights, alerts) are queued too. The dropped close_notify branch in ssl_handle_shutdown is gone: the alert goes out when the kernel takes it, then the FIN.
  • A socket that sent its close_notify keeps its writable poll while the alert is queued (loop.c, us_socket_resume).
  • A close waits for the queue only after the handshake completed. A socket closed in its handshake while it held unsent ciphertext now reports the failed handshake like every other socket (the low-prio fixture counted 128 reports for 130 such sockets, now 130).
  • A close that waits, for the queue or for the peer's close_notify, ends when the socket's timeout fires. Before, that timeout went to a holder that had already closed the socket, and a Bun.connect socket then stayed open. When the holder armed no timeout and bytes are still unsent, the close arms 10 s, the default idle timeout of Bun.serve. The timeout wheel ticks every 4 s, so that is 8 to 12 s, and each drain progress starts it again.
  • Removed as dead: the SSL_ERROR_WANT_WRITE arms of the handshake and read drivers with ssl_read_wants_write, the foreign-slot kill in ssl_flush_write_batch, the spill re-point in us_internal_ssl_socket_relocated.

Worst case held in the queue: one flush unit (at most about 147 KB) for the one connection per loop that holds the rest of a batch, and the rest of one record (16 KiB of plaintext) plus a handshake flight or an alert for any other. Main held the same: the record sat inside BoringSSL. No record of the other drivers is unbounded: BoringSSL sends one KeyUpdate reply per write of ours, and client renegotiation is limited to 3 per 600 s.

Measurements (main 519963edc8 against this branch)

  • Sizes, one translation unit each with the release flags without ThinLTO (llvm-dwarfdump, llvm-size): us_socket_t 80 to 80, us_ssl_rare_t 56 to 64, loop_ssl_data 384 to 368 bytes. .text of openssl.o 22423 to 23121, loop.o 5983 to 6079, socket.o 4882 to 4896 bytes. A release binary was not built.
  • send() calls, gdb breakpoint hit counts (release bun-profile of 367d939 against the debug build of this branch): 200 writes of 64 KiB with nothing stalled 204 to 204. Writer beside a stalled socket 193 to 193. Single writer that stalls 27 to 27.
  • Queue appends (gdb hit count on ssl_out_queue_append): 0 for the 200 unstalled writes. In the new test (one run): 42256 bytes for the stalled socket (rest of a batch) and 7462 bytes for the socket under test (rest of one record).
  • 4 TLS clients that stop reading a 64 MiB response of Bun.serve({ tls, idleTimeout: 2 }): connections still open 30 s after the timeout, 1 of 4 to 0 of 4 (runs of 40 s: 1 on the release build, 2 on this branch).
  • Not measured: instructions per write (no perf or valgrind in the container) and server CPU beside a stalled socket (no release build of the branch).

Open question for a maintainer

The usockets rewrite (#34037) records a directive to keep one loop-shared spill slot so that memory stays O(1), and it rejects a spill per socket. This PR keeps the byte bound (the same gate), but the bytes live on the connection. If the buffer itself must stay loop-shared, the alternative is to keep the slot and add a place for one refused record per connection.

Other PRs

Not covered

Tests run on the debug build (linux x64): socket.test.ts, socket-syscall-fault.test.ts, tls-syscall-fault.test.ts, node-tls-server.test.ts, node-http-backpressure.test.ts, node-tls-connect.test.ts, node-tls-duplex-end-verify.test.ts. Before the last rebase also node-tls-cert, node-tls-context, node-tls-upgrade, renegotiation, node-tls-raw-end, node-tls-socket-allow-half-open-option, node-tls-wrapped-socket-close, tls-connect-socket-churn, tcp-server, socket-retention, tls-reject-before-client-cert, bun-serve-ssl, tls-keepalive, node-http-pinned-write, node-https-checkServerIdentity, fetch.tls, node-http2. Failures seen were 5 s timeouts of tests that pass when run alone on the loaded machine, and one test that needs the public internet.

The low-prio queue fixture (tls-low-prio-queue-fixture.ts) failed sockets through the refused-flight close. Its clients now send a fatal alert instead. With the double-park guard in loop.c removed it still aborts on group->low_prio_count == 0.

…ealed record

The write BIO of a usockets TLS socket refused a record that the kernel
did not take while another socket owned the loop's spill slot. BoringSSL
had already sealed that record and kept it for a retry with the same
bytes, and us_internal_ssl_writev returned a count without it. The next
write of other data then sent the stale record and lost as many accepted
bytes, and a smaller next write failed with SSL_R_BAD_WRITE_RETRY.

Each TLS connection now has its own queue for ciphertext that the kernel
did not take, and the BIO takes every record. The spill slot of the loop
is removed. The batching rule stays: while one connection holds the rest
of a batch, the other sockets send record by record.

A close_notify that the kernel does not take is queued, and the FIN
follows it. A close that waits for queued ciphertext ends at the
socket's timeout.
@robobun

robobun commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 6:08 AM PT - Oct 3rd, 2026

❌ @robobun, your commit 4f60a4f has 2 failures in Build #123306 (All Failures):


🧪   To try this PR locally:

bunx bun-pr 44529

That installs a local version of the PR into your bun-44529 executable, so you can run:

bun-44529 --bun

@robobun

robobun commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator Author

Status of #44529: draft. The self-review has not finished yet.

How the problem was reproduced:

  • Two TLS clients on one event loop. The first fills the kernel and stays stalled. The second calls write(1 MiB) until a write is short, then writes other data.
  • On main the peer receives 16384 bytes that write() did not report, and 16384 accepted bytes never arrive. A smaller next write returns 0 and the socket never drains.
  • The tests are in test/js/bun/net/socket.test.ts (TLS write() that ends short while another TLS socket is stalled). USE_SYSTEM_BUN=1 bun test test/js/bun/net/socket.test.ts -t "ends short" fails 3 of 4. bun bd test passes 4 of 4.

@Jarred-Sumner

Copy link
Copy Markdown
Collaborator

Superseded by #44618, which consolidates the open TLS pull requests. This fix and its tests are in there, either as written, rewritten smaller, or merged with the other PRs that patched the same cause (see the "By area" list in that PR). Closing in favor of it.

Jarred-Sumner added a commit that referenced this pull request Oct 6, 2026
…s every record it sealed (#44529)

While another TLS socket owned the loop's spill slot, the write BIO refused what
the kernel did not take. BoringSSL had already sealed that record and kept it for
a retry, and us_internal_ssl_writev did not count it. The next write of other
data flushed the old record and dropped as many accepted bytes (silent stream
corruption); a shorter one, or any write after setMaxSendFragment(), failed with
BAD_WRITE_RETRY and the socket went fatal with no event.

The BIO now takes every record. Only what the kernel refused is copied, to the
connection (us_ssl_rare_t), and freed when it drains. The loop-wide rule stays:
while one connection holds the rest of a batch flush nobody batches, so userspace
holds one flush unit per loop plus at most the rest of one record per other
stalled connection. That record used to sit, whole, in BoringSSL's own
per-connection write buffer. 16 stalled writers on one loop: 340378 B before
(94288 slot + 15 x 16406), 275968 B + 76 B per connection now.

Nothing is allocated or copied while the kernel keeps up: 200 writes of 64 KiB,
203 send() calls and 0 spills, as before. us_ssl_rare_t 56 -> 64 bytes,
loop_ssl_data 392 -> 368, us_socket_t 80 -> 80 (one bit of its padding).

Gone with the slot: its owner, the relocation hook for it, and the "slot is
another socket's, so this connection dies" branch.

A close now waits behind the spill of every stalled connection, not only the slot
owner's (16 peers that stop reading, destroy() on all: 9 fds stay open, 1 before).
The next commit bounds that wait.

tls-low-prio-queue-fixture.ts: a refused flight no longer leaves the next byte
unread, so the burst that fails the parked handshakes is a fatal alert.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants