Skip to content

socket: drain node:net's pending-write queue with a cursor - #44219

Open
robobun wants to merge 1 commit into
mainfrom
robobun/0fba4e9e/socket-write-queue-linear-drain
Open

robobun wants to merge 1 commit into
mainfrom
robobun/0fba4e9e/socket-write-queue-linear-drain

Conversation

@robobun

@robobun robobun commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

  • A node:net write that one send does not take whole costs CPU by the square of its size. socket.end(Buffer.alloc(64 << 20)) over a unix socket takes 1,107 ms (Node v26.3.0: 18 ms). No user reported it.
  • After each partial send, internal_flush (src/runtime/socket/socket_body.rs:3156) and write_or_end_buffered (:2822, :2864) move all unsent bytes to the front of buffered_data_for_node_net.

Fix

  • The queue is PendingWrites, a newtype over bun_io::StreamBuffer. consume advances a cursor and moves no bytes.
  • The same bytes leave in the same order. Only this queue changes.
  • Verified: test/internal/socket-pending-writes.test.ts, test/js/node/net/net-syscall-fault.test.ts. On main the first cannot load, the second fails 2 of 8 tests.
  • Self-reviewed: 18 concerns raised, 15 addressed. Open: fan-out and slow-link numbers, issues for the sites in Notes.

Background

Downsides

  • NewSocket: 232 to 240 bytes (same allocator class). A whole write: 954 to 956 instructions. The binary: +4,419 bytes.
  • A write behind a tail that is under half sent doubles the allocation: 2,088,960 bytes for 525,312 unsent, main 1,044,480.
  • At 64 MiB the gain is 16x over a unix socket, about 2x over loopback TCP or TLS, and none for a sender that waits for drain.
Notes

Sender CPU, release builds

socket.end(Buffer.alloc(N)), linux-x64, loopback, reader is Node v26.3.0 in a second process. The numbers are the CPU time of the sender (user plus system) from end() to its callback, in ms. Median of 7 interleaved rounds. main is 9f70da0, this PR is bd4967d.

Transport Build 16 MiB 32 MiB 64 MiB Per doubling
unix socket bun 1.4.2 75.9 252.1 1,182.8 x3.3, x4.7
unix socket main 63.4 263.4 1,107.1 x4.2, x4.2
unix socket this PR 15.9 33.9 66.9 x2.1, x2.0
unix socket Node 4.8 9.4 17.8
TCP bun 1.4.2 22.3 46.4 131.9 x2.1, x2.8
TCP main 24.3 58.0 146.7 x2.4, x2.5
TCP this PR 17.4 34.6 63.8 x2.0, x1.8
TCP Node 9.9 11.6 23.8
TLS bun 1.4.2 24.7 65.0 152.4 x2.6, x2.3
TLS main 26.9 68.7 147.4 x2.6, x2.1
TLS this PR 21.2 43.0 89.0 x2.0, x2.1
TLS Node 26.6 54.3 110.5

A sender that writes 64 MiB in 64 KiB chunks and waits for drain (15 rounds): unix socket 42.1 ms on main, 42.2 ms on main again, 42.7 ms with this PR. TCP 34.8, 33.3 and 33.2 ms.

The peer decides the size of a send. end() of 64 MiB over TCP to a reader that sets SO_RCVBUF to 2,048 and reads as fast as it can (3 rounds): 528.1 ms on main, 110.5 ms with this PR, 65.5 ms on Node.

Not measured: many connections that drain at the same time, and a link with a small congestion window.

The machine had a load average near 380 on 16 cores, so these are CPU times and not wall times.

Scripts for the sender CPU table

sender.js <unix|tcp|tls> <path or port> <bytes> <end|drain>, run with each build:

// usage: sender.js <transport> <target> <size> <mode>
//   transport: unix | tcp | tls      target: socket path or port
//   mode: end (one chunk) | drain (64 KiB chunks, waits for 'drain')
const net = require("node:net");
const tls = require("node:tls");
const [transport, target, sizeArg, mode] = process.argv.slice(2);
const size = Number(sizeArg);
const payload = Buffer.alloc(size, 0x61);
const onConnect = () => {
  const c0 = process.cpuUsage();
  const done = () => {
    const c = process.cpuUsage(c0);
    console.log(JSON.stringify({ cpu: (c.user + c.system) / 1000 }));
  };
  if (mode === "drain") {
    const chunk = 64 * 1024;
    let at = 0;
    const pump = () => {
      while (at < size) {
        const ok = sock.write(payload.subarray(at, at + chunk));
        at += chunk;
        if (!ok) return sock.once("drain", pump);
      }
      sock.end(done);
    };
    pump();
  } else {
    sock.end(payload, done);
  }
};
const sock =
  transport === "unix"
    ? net.connect({ path: target }, onConnect)
    : transport === "tcp"
      ? net.connect({ port: Number(target), host: "127.0.0.1" }, onConnect)
      : tls.connect({ port: Number(target), host: "127.0.0.1", rejectUnauthorized: false }, onConnect);
sock.resume();

reader.js <unix|tcp|tls> <path>, run with Node:

// usage: reader.js <transport> <path-or-0>   prints "ready <port-or-path>", then the byte total on stderr
const net = require("node:net");
const tls = require("node:tls");
const fs = require("node:fs");
const [transport, target] = process.argv.slice(2);
let total = 0;
const onSocket = s => {
  s.on("data", c => (total += c.length));
  s.on("end", () => {
    console.error(total);
    s.end();
    server.close();
  });
  s.on("error", () => {});
};
let server;
if (transport === "tls") {
  const dir = process.env.CERT_DIR;
  server = tls.createServer({ key: fs.readFileSync(dir + "/key.pem"), cert: fs.readFileSync(dir + "/cert.pem") }, onSocket);
} else {
  server = net.createServer(onSocket);
}
if (transport === "unix") {
  try { fs.unlinkSync(target); } catch {}
  server.listen(target, () => console.log("ready " + target));
} else {
  server.listen(0, "127.0.0.1", () => console.log("ready " + server.address().port));
}

Bytes moved

Debug builds, TCP loopback, sender net.connect(port, () => sock.end(payload)), every send clamped with socketFaultInjection.set({ syscall: "send", action: "short", bytes, repeat: -1 }).

Payload, clamp main: flushes main: bytes moved this PR: moves
4 MiB, 16 KiB 254 526,434,304 0
4 MiB, 4 KiB 1,022 2,137,010,176 0
32 MiB, 4 KiB 8,190 137,355,079,680 0
  • main: sum of the onWritable buffered_data_for_node_net <len> lines of BUN_DEBUG_Socket=1, measured at 36cd151. The three sites are the same on 9f70da0. On main that length is what the flush just moved. With this PR the line still prints the queue length, and it no longer means bytes moved.
  • This PR: hits of a gdb breakpoint on the one line that moves bytes (self.list.drain(..self.cursor) in StreamBuffer::compact).
  • Per doubling, main moves x4.0 (4 MiB to 32 MiB at 4 KiB is x64.3).
  • CPU time of the sender process for 32 MiB at 4 KiB, debug builds: 16.7 s on main, 1.0 s with this PR.

Cost for a socket that never queues

  • size_of::<NewSocket>(): release 232 to 240 bytes (both in mimalloc's 256-byte class), debug 504 to 512. estimateShallowMemoryUsageOf(socket._handle) of an idle socket: 329 to 337.
  • One 1 KiB write() that one send takes whole, counted with gdb stepi on the release builds: TCPSocketPrototype__writeBuffered and its callees 954 to 956 instructions, 20 to 20 calls. The bytesWritten getter 36 to 37 instructions, 0 calls. The queue length is now two loads and a compare.
  • Release binary (size): text 80,666,326 to 80,670,745 bytes, data and bss the same. 3,090 of the 4,419 bytes are the test probe.
  • Host functions: 1 added, pendingWritesReplayProbe in bun:internal-for-testing. No new rows in sockets.classes.ts.
  • Allocator calls of the queue, counted with gdb on debug builds. A write that the queue drains: 1 allocation and 1 free, as on main. A writev that takes the whole tail and part of the next chunk: 1 free and 1 allocation, where main has 1 realloc.
  • Tools that the machine does not have: perf, valgrind, strace, bloaty.

Tests

  • test/internal/socket-pending-writes.test.ts runs on every lane. It replays schedules of appends and drains on a PendingWrites and reads how many unsent bytes changed address.
  • The new block in net-syscall-fault.test.ts needs a build that can clamp a send, so it runs where socketFaultInjection.available() is true. It is skipped on Windows like the rest of that file.
  • On main, 2 of its 8 tests fail: "writev clamp below the tail" (80,896 bytes allocated expected, 79,872 received) and "a write behind a tail that writable events partly sent" (1,309,696 expected, 1,307,648 received). main has dropped the sent bytes from the allocation by then, and that takes the move. The other 6 tests pass on main: they pin bytes, order and bytesWritten.
  • Each of these changes to the fix breaks at least one new test: a move in consume, no clamp in consume, reset() in place of a free, a length that counts the sent prefix, and a count that is off by one at each of the three drain sites.
  • Also ran with the debug build: every *.test.* file in test/js/node/net/, test/js/node/tls/tls-syscall-fault.test.ts, node-tls-connect, node-tls-server, node-tls-upgrade, node-tls-raw-end, node-tls-duplex-end-verify, test/js/bun/net/socket.test.ts, socket-syscall-fault.test.ts, node-http-syscall-fault.test.ts, and from test/js/node/test/parallel: test-net-bytes-written-large, test-net-write-slow, test-net-throttle, test-net-write-fully-async-buffer, test-net-write-fully-async-hex-string, test-tls-buffersize, test-tls-fast-writing, test-tls-connect-stream-writes. The machine was under heavy load, so tests that spawn processes hit their 5 s timeout at random. Every failure also shows with main's src/, or goes away when the test runs alone.
  • bun run rust:check-all: 12 of 12 targets.

Windows

Not measured. The queue is the same type on every platform and transport. The writev arm is #[cfg(unix)], so on Windows a write behind a tail takes the append-then-send arm.

Not changed


no test proof · iteration 0 · platform-specific test(s) that do not run on this machine, deferring to CI, which covers all platforms: test/js/node/net/net-syscall-fault.test.ts

A short send leaves the rest of a node:net write in
`NewSocket::buffered_data_for_node_net`. The three places that take
sent bytes out of that queue moved every unsent byte to the front of
the Vec after each partial send. For a tail of C bytes that leaves in
steps of s bytes that is C * C / (2 * s) bytes moved.

The queue is now `PendingWrites`, a newtype over `bun_io::StreamBuffer`.
`consume` advances the cursor. The queue frees its allocation when it
empties, and an append drops the sent prefix once it is at least as
large as the unsent bytes (the rule StreamBuffer got in #42731).
@robobun

robobun commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

Status of #44219

How the cost was reproduced:

  • Release builds, linux-x64. The sender runs net.connect({ path }, () => sock.end(Buffer.alloc(64 << 20))) over a unix socket. The reader is Node v26.3.0 in a second process. The CPU time of the sender from end() to its callback is 1,107 ms on main (9f70da0) and 67 ms with this PR.
  • Debug builds, TCP loopback, every send clamped to 4 KiB with socketFaultInjection.set({ syscall: "send", action: "short", bytes: 4096, repeat: -1 }). A 32 MiB end() takes 8,190 flushes. On main they move 137,355,079,680 bytes (the sum of the onWritable buffered_data_for_node_net <len> lines of BUN_DEBUG_Socket=1). With this PR a breakpoint on the one line that moves bytes has 0 hits.

@coderabbitai

coderabbitai Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: oven-sh/bun/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Essentials

Run ID: 3bc98e6d-2eee-4223-b2d6-dcebf0ace8c3

📥 Commits

Reviewing files that changed from the base of the PR and between 9f70da0 and bd4967d.

📒 Files selected for processing (7)
  • src/codegen/generate-js2native.ts
  • src/js/internal-for-testing.ts
  • src/runtime/socket/mod.rs
  • src/runtime/socket/pending_writes.rs
  • src/runtime/socket/socket_body.rs
  • test/internal/socket-pending-writes.test.ts
  • test/js/node/net/net-syscall-fault.test.ts

Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 6 remain after this review.


Walkthrough

The socket runtime adds PendingWrites to track unsent bytes and consume successfully written prefixes. A Rust-backed replay probe reports queue integrity and buffer metrics. Tests cover replay schedules and short writes across TCP and TLS sockets.

Changes

Pending-write queue

Layer / File(s) Summary
Queue semantics and replay probe
src/runtime/socket/pending_writes.rs, src/runtime/socket/mod.rs, src/codegen/generate-js2native.ts, src/js/internal-for-testing.ts, test/internal/socket-pending-writes.test.ts
PendingWrites wraps a StreamBuffer and supports appending, inspecting, and consuming unsent bytes. The replay probe reports integrity, byte counts, capacity, and buffer movement. Tests exercise append and drain schedules.
Socket pending-write integration
src/runtime/socket/socket_body.rs
NewSocket stores pending data in PendingWrites. Partial-write paths consume the written prefix, and TCP and TLS construction paths initialize the buffer.
TCP and TLS queue validation
test/js/node/net/net-syscall-fault.test.ts
Tests inject short sends and check payload order and integrity, callback behavior, and queue allocation as data drains across TCP and TLS scenarios.

Suggested reviewers: cirospaciari

Priority: ➖ Normal

Merge Risk: ⚪ Minimal · up to bd496

This change makes node:net partial-write draining cheaper without changing byte ordering. Tests cover short sends over TCP and TLS. No concrete merge-blocking risk was identified.

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly and concisely describes the main change: replacing repeated pending-write buffer moves with cursor-based draining in node:net.
Description check ✅ Passed The description thoroughly explains the problem, implementation, tradeoffs, verification results, test coverage, and limitations. It does not use the template headings exactly, but it provides the req…

Comment @coderabbitai help to get the list of available commands.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

Beyond the inline finding, I also checked the three collapsed call sites in src/runtime/socket/socket_body.rs: consume(n) reproduces each old branch (n >= len frees, n < len keeps the tail) including the writev case where written exceeds the queue length, and at the TLS site (:2819) remaining_in_input_data borrows the caller's input buffer rather than the queue, so advancing the cursor cannot shift bytes under it.

Extended reasoning...

The change swaps the Vec node:net pending-write queue for a StreamBuffer-backed cursor newtype and rewires three drain sites plus a test-only replay probe; it touches no auth, crypto, or input-parsing surface. One inline finding (the 2x allocation growth behind a partly sent tail) is already posted, so this note only records the semantic-equivalence checks on the call sites that were ruled out.

Comment thread src/runtime/socket/pending_writes.rs

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant