Skip to content

ipc: postpone disconnect() while a handle send is still queued - #43081

Open
robobun wants to merge 2 commits into
mainfrom
robobun/24cbda27/ipc-disconnect-flush-queued-handle
Open

robobun wants to merge 2 commits into
mainfrom
robobun/24cbda27/ipc-disconnect-flush-queued-handle

Conversation

@robobun

@robobun robobun commented Sep 17, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

Fix

  • disconnect() also sets close_after_flush when an item in queue carries a handle. continue_send already closes the socket once the queue is empty and no ACK is pending.
  • Correct because "an item in queue has a handle, or waiting_for_ack is set" is the same state as node's _handleQueue !== null. kill(), child exit, and peer EOF still close at once through close_socket_next_tick.
  • Verified: two new tests in test/js/node/child_process/child_process_ipc_handle.test.ts, one per direction. Both fail on bun 1.4.3-canary. Other suites: see Notes.

Background

  • SendQueue is the outgoing side of an IPC channel. queue holds the messages that are not fully written. A write larger than the socket buffer is partial.
  • A message with a handle (a socket or a server) needs an ACK from the peer. After the write, the item moves from queue to waiting_for_ack. No other message goes out until the ACK arrives.
  • close_after_flush marks a disconnect() that waits. connected reports false at once. The socket closes after the queue drains.
Notes

Repro (bun parent.js and node parent.js):

// parent.js
const { fork } = require("node:child_process");
const net = require("node:net");
const child = fork(__dirname + "/child.js", { stdio: ["ignore", "pipe", "inherit", "ipc"] });
let childOut = "";
child.stdout.on("data", d => (childOut += d));
const server = net.createServer();
server.listen(0, "127.0.0.1", () => {
  net.connect(server.address().port, "127.0.0.1", function () {
    const r = { big: "never", handle: "never", after: "never" };
    const pad = Buffer.alloc(1 << 21, "d").toString(); // larger than the IPC socket buffer
    child.send({ pad }, e => (r.big = e));
    child.send("handle", this, e => (r.handle = e));
    child.send("after", e => (r.after = e));
    child.disconnect();
    child.on("exit", () =>
      setImmediate(() => {
        console.log(JSON.stringify({ callbacks: r, childReceived: childOut.trim().split("\n").filter(Boolean) }));
        process.exit(0);
      }),
    );
  });
});
// child.js
process.on("message", (m, h) => {
  console.log(typeof m === "string" ? m + (h ? "+handle" : "") : "big");
  if (h) h.destroy();
});
process.on("disconnect", () => process.exit(0));
runtime callbacks child received
node v26.3.0 big: null, handle: null, after: null big, handle+handle, after
bun 1.4.3-canary all "never" nothing
this PR big: null, handle: null, after: null big, handle+handle, after

The child to parent direction (process.send of a net.Server, then process.disconnect()) gives the same three rows.

One difference from node that this PR does not change. After the last handle is acknowledged, bun flushes the whole remaining queue before it closes. node closes on the next tick and cancels a trailing write that is still pending. With big1, handleA, mid, handleB, big2, last and then disconnect(), node delivers up to handleB. bun with this PR delivers all six. That is how close_after_flush already worked for a written handle (#31829).

Paths that stay as they are. kill(), child exit, and peer EOF call close_socket_next_tick, which clears close_after_flush and closes. The existing test "channel close: written handle callback fires null; unsent queued handle callback never fires" pins that and still passes.

Sibling site left as it is. serialize_and_send decides the false return of send() from waiting_for_ack alone (indicate_backoff, src/runtime/ipc.rs:1626). It is bun's form of node's return this._handleQueue.length === 1, and it has the same gap: while the handle is still in queue, every send() returns true. In the repro state, five sends (big, handle, then three plain messages) return false, false, true, false, false on node v26.3.0 and true five times on bun, with and without this PR. It is excluded on purpose. It changes only the return value of send(), not what is delivered. The first two false come from node's writeQueueSize threshold, which bun does not have. That is #30569 (open; #32925 was a fix for it and was closed unmerged). The return value of send() should be fixed in one change, under #30569.

Relation to other PRs.

Provenance and reach.

  • Found in review of ipc: run the callbacks of sends still queued when the channel closes #43000, by reading node's source against SendQueue::disconnect. A tracker search (disconnect, send, handle, ipc and cluster terms) found no report of the symptom.
  • Measured on bun 1.4.3-canary, Linux x64, with the repro above and a smaller pad, three runs per size. At 200 KB and 220 KB everything arrives. 240 KB is mixed (2 of 3 runs deliver). From 256 KiB up nothing arrives. The limit is the socket buffer (net.core.wmem_default is 208 KiB here) while the forked child is still starting and does not read yet. A busy peer would build the same backlog from smaller messages (not measured).
  • Who passes handles: code that calls send(message, handle) itself, and cluster for node:net servers under SCHED_RR, one handle per accepted connection from the primary to the worker. node:http workers listen with reusePort and pass no handles.

Windows and macOS. The change is in shared code. The tests in this file are POSIX-only (describe.skipIf(isWindows)), so CI does not run the new tests on Windows. On Windows the pipe write completes asynchronously (write() in ipc.rs), so a handle is still in queue whenever disconnect() runs in the same tick as the send. The new branch is the usual path there, not a backlog case. Run by hand on Windows Server 2019 x64 with the two test fixtures: bun 1.4.3-canary.1+b64b63069 delivers big and the handle and drops after in both directions. A debug build with this change delivers all three and calls every callback with null, 3 of 3 runs per direction. One caveat there: an ACK that is read before the write callback runs is dropped (#37815, open), and a postponed disconnect() then waits until the child exits. test/js/node/test/parallel/test-cluster-shared-leak.js can reach that state. In two CI runs of this PR it timed out once, on Windows 11 aarch64, and passed on retry (build 117039). It did not time out in build 117060. main shows the same timeout on the same lane (build 116425, 1 of the last 12 main builds), so two runs do not show whether this change moves the rate. The fixtures were not run by hand on macOS. The two new tests pass on the macOS CI lanes.

Suites run on the debug build, all pass:

  • test/js/node/child_process/child_process_ipc_handle.test.ts (16 tests, 6 full runs, and the 2 new tests 15 more times)
  • test/js/node/child_process/child_process_ipc.test.js, child_process_send_cb.test.js, child_process_ipc_large_disconnect.test.js
  • test/js/bun/spawn/spawn.ipc.test.ts, spawn.ipc.bun-node.test.ts, spawn.ipc.node-bun.test.ts, bun-ipc-inherit.test.ts, spawn-ipc-gc.test.ts
  • test/js/node/cluster.test.ts (35 tests)
  • 111 node tests: test/js/node/test/parallel/test-child-process-{fork,ipc,send,disconnect,recv,net,exit,stdio-ipc,constructor,internal}* and test-cluster-*
  • test/js/node/test/sequential/test-cluster-send-handle-large-payload.js, test-child-process-pass-fd.js, test-child-process-exit.js, test-cluster-port-reuse-between-workers.js, test-cluster-net-listen-ipv6only-{rr,none}.js

no test proof · iteration 0 · platform-specific test(s) that do not run on this machine, deferring to CI, which covers all platforms: test/js/node/child_process/child_process_ipc_handle.test.ts

disconnect() postponed the close only when a handle was already written
and waited for its ACK. A handle that was still queued behind an
unfinished write did not count. The channel closed on the next tick and
dropped the handle and every message behind it.

node postpones the disconnect from the moment it submits the handle
write. Treat a queued item that carries a handle the same as a handle
that waits for its ACK.
@coderabbitai

coderabbitai Bot commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Review Change StackReview Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Essentials

Run ID: be9f4521-e6c8-4eef-b617-1b0b3141462e

📥 Commits

Reviewing files that changed from the base of the PR and between b52d513 and a55cf2f.

📒 Files selected for processing (2)
  • src/runtime/ipc.rs
  • test/js/node/child_process/child_process_ipc_handle.test.ts

Included review availability: Your plan provides up to 10 included reviews per hour; 0 remain after this review.


Walkthrough

SendQueue::disconnect now defers closure when queued messages contain handles. Parent-side and child-side tests verify ordered delivery, immediate disconnected state, and null send callbacks.

Changes

IPC disconnect behavior

Layer / File(s) Summary
Defer disconnect for queued handles
src/runtime/ipc.rs
SendQueue::disconnect now treats queued handle-bearing messages as outstanding work and keeps them eligible for flushing.
Validate parent and child disconnect flows
test/js/node/child_process/child_process_ipc_handle.test.ts
Tests verify that large messages, handles, and following messages arrive in order after disconnect. They also verify immediate disconnected state and null send callbacks.

Suggested reviewers: cirospaciari

Priority: ⬇️ Low

Merge Risk: ⚪ Minimal · up to a55cf

Queued handle transfers and following messages retain orderly delivery during disconnect, with no remaining concrete merge-blocking risk identified.

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly and concisely describes the main change: delaying disconnect() while a handle send remains queued.
Description check ✅ Passed The description clearly explains the problem, root cause, fix, scope, compatibility behavior, platform considerations, and extensive verification results. Although it does not use the template heading…

Comment @coderabbitai help to get the list of available commands.

@robobun

robobun commented Sep 17, 2026 •

Copy link
Copy Markdown
Collaborator Author

Status: fix is up in this PR (#43081).

How I reproduced it: the two-file repro in the PR Notes (child.send of a 2 MiB message, then a net.Socket handle, then 'after', then child.disconnect()).

  • node v26.3.0: all three callbacks get null, and the child receives big, the handle, and after.
  • bun 1.4.3-canary: no callback runs and the child receives nothing.
  • This branch: same result as node, in both directions (subprocess.disconnect() and process.disconnect()).

The two new tests in test/js/node/child_process/child_process_ipc_handle.test.ts fail with USE_SYSTEM_BUN=1 bun test and pass with bun bd test.

Self-review: the verdict was to ship after changes. State of its requests:

One more check, run by hand: a child that holds nothing else on the event loop (server unref'd, no listeners) stays alive until the postponed process.disconnect() has flushed. The parent receives all three messages and every callback gets null.

CI: build 117060 passes on all 181 jobs, Windows and macOS lanes included. No test in the cluster, child_process, or IPC suites failed or needed a retry in that build.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this PR and didn't find any bugs. The Rust change is three lines and mirrors the existing close_after_flush mechanism, but it changes when disconnect() actually closes the channel and overlaps with #43001 in the same function, so a maintainer look at the intended node-parity tradeoff would still be worthwhile.

What was reviewed:

  • SendQueue::disconnect in src/runtime/ipc.rs: the queue.get().iter() borrow is only entered from the JS host functions (Subprocess.disconnect, Bun__closeChildIPC), never inside continue_send's with_mut closure, so no re-entrant JsCell borrow; continue_send already closes once the queue is empty and no ack is pending, and close_socket_next_tick still clears close_after_flush on kill/exit/EOF.
  • Considered whether a peer that never reads/acks now makes disconnect() hang: that matches node's _handleQueue behavior and kill()/exit still close immediately, so it is intentional rather than a regression.
  • Both new tests are inside the existing describe.skipIf(isWindows) block, drain pipes with Promise.all, and assert a combined object; the 2 MiB pad is needed to exceed the socket buffer so the handle is still queued when disconnect() runs. I did not build and run them locally.
Extended reasoning...

Overview

The PR touches SendQueue::disconnect() in /home/claude/bun/src/runtime/ipc.rs (lines 1264-1273) and adds two test.concurrent cases to /home/claude/bun/test/js/node/child_process/child_process_ipc_handle.test.ts. Previously disconnect() only deferred the close (close_after_flush = true) when waiting_for_ack was set, i.e. the handle had already been written. Now it also defers when any item still in queue carries a handle. Everything else (continue_send closing once the queue drains, close_socket_next_tick clearing the flag on kill/exit/EOF, is_connected() reporting false while close_after_flush is set) is unchanged.

Security risks

None identified. The change only affects the timing of a local IPC socket close between parent and child processes; it does not parse untrusted data, touch credentials, or alter permission checks.

Level of scrutiny

Moderate. The diff is tiny and reuses an existing mechanism, but it is a semantic change in a process-lifecycle path: a disconnect() that used to close on the next tick now waits for the peer to read and ack a handle that may sit behind an arbitrarily large backlog. That matches node (the PR cites _handleQueue handling in lib/internal/child_process.js, which I could not fetch to verify from this sandbox), and kill(), child exit, and peer EOF still force the close, so I do not consider it a bug. However, the PR description itself notes a conflict with #43001 in the same function and a remaining divergence from node (bun flushes the entire trailing queue after the last ack rather than closing next tick), which are design tradeoffs a maintainer should weigh rather than an automated approval.

Other factors

I checked the JsCell borrow rule from src/CLAUDE.md: self.queue.get() in disconnect() is reached only from Subprocess::disconnect (src/runtime/api/bun/subprocess.rs) and close_child_ipc (src/runtime/hw_exports.rs), both JS host functions, and continue_send's with_mut closure does not run user JS, so no overlapping borrow is possible. The new tests cover both directions (parent child.disconnect() and child process.disconnect()), assert the peer received all three messages plus 'disconnect', and assert each callback got null; they live inside the existing describe.skipIf(isWindows) block. No debug build was available and I did not build one, so I did not confirm locally that the tests fail on the base branch and pass with the change; the author's claim that they fail on 1.4.3-canary is plausible given that on the base branch the handle item would be dropped at next-tick close.

@robobun

robobun commented Sep 17, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 6:47 AM PT - Sep 17th, 2026

@robobun, your commit 69f26f2 is building: #117060

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant