Skip to content

worker_threads: settle terminate() from the exit handler - #44252

Open
robobun wants to merge 2 commits into
mainfrom
robobun/6ecdf277/worker-terminate-in-exit-drain
Open

robobun wants to merge 2 commits into
mainfrom
robobun/6ecdf277/worker-terminate-in-exit-drain

Conversation

@robobun

@robobun robobun commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

  • await worker.terminate() never continues when code that the Worker's exit handler runs calls it first: a 'message' listener, a stdio 'readable' listener at EOF, or a stdin write callback.
  • terminate() added a native 'close' listener per call (src/js/node/worker_threads.ts:1074). The exit handler #onClose is the 'close' listener and stored the exit state (:1209) after that user code. A listener added during a dispatch does not run.

Fix

  • #onClose settles the pending terminate() after 'exit'. terminate() adds no listener.
  • #onClose follows node's kOnExit order: last messages, exit state, stdio EOF, 'exit'. A terminate() from a 'message' listener resolves the exit code, one from a stdio EOF resolves undefined, as in node v26.3.0.
  • Verified: 25 new tests in test/js/node/worker_threads/worker_threads.test.ts (18 fail on main), that whole file, vendored test-worker-* files.
  • Self-reviewed: 11 concerns raised, 9 addressed. Rejected: a second 'close' listener (each Worker pays), native reports for listener throws (Notes).

Background

  • Node's kOnExit delivers the worker's last messages, then disposes the handle: threadId becomes -1 and terminate() resolves undefined.
  • Weighed: an arm in terminate() for a close in progress (keeps a listener per call), node's once('exit') (removeAllListeners() loses the promise).

Downsides

  • Bytecodes: Worker exit 121 (main 99), new Worker() +4, threadId read 15 (main 7). Builtin source +586 bytes, binary text +613 bytes.
  • Inside the exit handler, a 'message' listener that throws no longer stops it, and it reads a running Worker (threadId, threadName, postMessage()). Node does both.
Notes

Repro with the public API only, on each run

const { Worker } = require("node:worker_threads");
const w = new Worker("", { eval: true, stdout: true });
w.on("exit", code => console.log("exit", code));
w.stdout.on("readable", async () => {
  if (w.stdout.read() !== null) return;
  console.log("terminate() resolved", await w.terminate());
});
output
node v26.3.0 exit 0, terminate() resolved undefined
bun 1.4.2, main (ad60a9b) exit 0
this PR exit 0, terminate() resolved undefined

bun 1.3.x has no worker.stdout, so this form starts with 1.4.0.

The reported form

The exit handler delivers a 'message' only when the worker posted and ended before the parent started its port. The parent starts the port at the end of new Worker(), after it started the thread. Once the port is started, the port itself delivers what is left when the worker's side closes (MessagePort::peerClosed), before the exit handler runs. The script holds the parent inside the constructor to reach that state on each run.

import { Worker } from "node:worker_threads";
import { EventEmitter } from "node:events";
const on = EventEmitter.prototype.on; let once = true;
EventEmitter.prototype.on = function (name, fn) { if (once && name === "data") { once = false; const end = Date.now() + 3000; while (Date.now() < end); } return on.call(this, name, fn); };
const w = new Worker("require('node:worker_threads').parentPort.postMessage(1)", { eval: true });
w.on("exit", code => console.log("exit", code));
await new Promise(done => w.on("message", () => w.terminate().then(done)));
console.log("terminated");
  • main prints exit 0 and stops there, 3 of 3. Under top-level await the process then never ends on a release build (the spin is the separate matter in Detect unsettled top-level await and exit like Node (exit 13 + warning) #33283).
  • This PR and node v26.3.0 print exit 0, then terminated, 3 of 3.
  • This form is a regression from 1.3.14. It came with Worker / worker_threads: WebCore-shaped lifetimes, joined threads, one ordered VM teardown #37075, which made the exit handler deliver the worker's last messages.
  • The report measured 1 to 7 of 1,000 process starts on a busy machine with no device. I saw it with no device on a machine with a load average of about 600 (16 cpus). The worker posts 3,000 messages, and the parent calls terminate() from the listener of message 1,500. On main the promise stayed pending in 2 of 190 runs. With this PR the exit handler delivered that message in 2 of 190 runs, and the promise settled in both. With one message: 0 of 400 on main, 0 of 400 on this PR.

Value per place and form

main leaves the promise pending in all 12 cells.

first terminate() from promise terminate(callback) [Symbol.asyncDispose]()
a 'message' that the exit handler delivers the exit code, after 'exit' callback(null, code) after 'exit', then the promise undefined
stdout or stderr 'readable' at EOF undefined undefined, callback not called undefined
a parked worker.stdin.write() callback undefined undefined, callback not called undefined

node v26.3.0 gives the same values in the first two rows, except that it ignores the callback. It does not run the stdin callback at all, because its exit handler does not destroy stdin.

A listener that throws inside the exit handler

throw from main this PR node v26.3.0
a 'message' that the exit handler delivers the handler stops: no more messages, no 'exit', promise pending reported, then the other messages, 'exit', and the promise the same, but it reports the errors after 'exit'
stdout or stderr 'readable' at EOF the handler stops: no 'exit'. The promise resolves. as main the handler stops: no 'exit'. The promise stays pending.
an 'exit' listener, a terminate() callback reported, the promise resolves as main the promise stays pending
  • The first row is a behaviour change. With 3,000 messages in the port and a listener that throws on 3 of them, main delivers 1 and this PR delivers 3,000.
  • The throw is caught in JavaScript and passed to the uncaught exception path, as guardCallback does for fs and dns callbacks. For a thrown value that is not an Error, the report then has no code frame. A native report needs a change to jsFunctionReportUncaughtException or a second native listener. I did not take either.

Other behaviour changes

  • A 'message' that the exit handler delivers now comes before the stdio EOF. main ends stdio first. Node's order is messages first.
  • Inside such a listener the Worker reads as one that runs: threadId, threadName, postMessage() (it throws for a value that cannot be cloned), and worker.stdin is not destroyed. main reads -1 and null, and postMessage() returns.
  • Outside the exit handler threadId is the native value, as on main. So a Worker of a disposed Bun.ModuleGraph reads -1 after its thread ended. A listener for a 'message' from the worker's global postMessage() that the native close task delivers reads -1 too.
  • A terminate() before the exit, then a second one from a stdio EOF listener: the second resolves undefined (main: the exit code). Node resolves undefined.
  • startHeapProfile() now reads the native state, so it rejects ERR_WORKER_NOT_RUNNING in each of these places, as before.
  • Bun.ModuleGraph: when a graph is disposed from inside the exit handler of its Worker, a pending terminate() now resolves. On main it stays pending, because the native dispatch skips the listeners of a disposed graph. test/js/bun/module-graph/module-graph-workers.test.ts:706 accepts both.

Differences from node that stay as on main

  • terminate(callback) is honoured, with DEP0132. Node v26 ignores the callback.
  • Two calls before the exit return one promise. Node makes one per call.
  • terminate() adds no 'exit' listener, so removeAllListeners() and an 'exit' listener that throws do not lose the promise.
  • The listeners stay on the Worker after 'exit', and worker.stdin is destroyed at exit.

Measurements

Release builds of ad60a9b with and without this change, unless the line says another build.

  • Native listener registrations: new Worker() 11 (main 11), terminate() 0 (main 1), terminate(callback) 0 more (main 1 more). listenerCount('exit') does not change in either.
  • GC cells per terminate() (heapStats, 20 Workers, 2 runs): first call 4 (main 8), repeated call 0 (main 2), call after exit 1 (main 3). Retained while pending: 4 (main 5).
  • Per first terminate() (gdb hit counters on debug builds, 20 and 40 Workers): addEventListener 0 (main 1), JSEventListener 0 (main 1), Weak handles 0 (main 2), fastMalloc about 3 (main about 6).
  • Bytecodes on the straight path (BUN_JSC_dumpGeneratedBytecodes): terminate() first, cached, after exit: 35, 17, 14 (main 48, 21, 22). One Worker exit: 121 and 10 branches (main 99 and 9), no allocation. new Worker(): 4 more, for one more private field. threadId on a Worker that runs: 15 (main 7). #onMessage is unchanged.
  • Messages that the exit handler delivers in ordinary runs: 0 of 180,600 (main 0). So the guarded loop body does not run there.
  • Builtin source: 29,419 to 30,005 bytes. Binary text (size): 80,667,862 to 80,668,475.
  • No native file changes. The node:worker_threads binding keeps 17 entries.
  • perf and valgrind are not available in my environment, so there is no instruction count for bench/postMessage.

Tests

  • 18 of the 25 new tests fail on main. The other 7 pin what must not move: terminate() on a running Worker in the three forms, the stdio listener that throws, and the threadId of a Worker whose Bun.ModuleGraph was disposed.
  • The 'message' tests use no timer. They replace EventTarget.prototype.addEventListener during new Worker() so that the public port never starts. Then only the exit handler can deliver what the worker posted.
  • I removed or changed 13 clauses of the change one at a time (the finally, each guard, the copy of the thread id, the place where the exit state is stored). Each change fails at least one new test.
  • Suites run on a debug build: worker_threads.test.ts (167 pass), worker_destruction, worker-async-dispose, worker-transfer-terminate-stress, worker-top-level-await, worker-transfer-list, worker-shutdown-post-leak, 15787, message-channel. Vendored test-worker-*: 107 of 108 pass (debug build at ef8e000, release build at 225acf1). test-worker-arraybuffer-zerofill.js fails on main in the same way (it calls describe from node:test outside the test runner).
  • module-graph-workers.test.ts, rows with a node Worker: 253 pass. On a machine under heavy load one row at times reports fdsAboveBaseline: 2. It did so in 2 of 12 runs on main and in 2 of 12 runs on this PR.

Not in this PR

  • Output that a worker writes to stdout or stderr right before process.exit() is lost when the worker ends before the parent starts its stdio ports. Node's kOnExit drains that port too. This is on main and this PR does not change it.
  • MessagePort.close(callback) waits with this.once('close', callback). It has the same shape as the old terminate(). Its output is the same on main and here.
  • After terminate(), the native Worker drops the messages that the worker posted through the global postMessage() (m_wasTerminated in Worker::dispatchEvent). So a terminate() from the listener of the first of 3,000 such messages leaves 1 message received, on main and here. Messages through parentPort.postMessage() do not pass that gate: main, this PR and node v26.3.0 deliver all 3,000.

Prior work

#19940 had #onClose resolve a deferred that the Worker owns. This PR keeps that idea. Here the first terminate() makes the deferred, and #onClose resolves it after 'exit'.

terminate() waited for the native 'close' event with a listener that it
added on each call. The parent-side exit handler is itself the 'close'
listener, and it runs user code before it stores the exit state: the
'message' listeners for what the worker posted before it ended, the
'readable' listeners at the stdio EOF, and parked stdin write callbacks.
A terminate() from one of them added its listener during the only
dispatch of 'close', so its promise never settled.

The exit handler now settles the pending terminate() after it emits
'exit', and terminate() adds no listener. The handler runs in the order
of node's kOnExit: deliver the worker's last messages, store the exit
state, end stdio, emit 'exit'. A 'message' or 'exit' listener that
throws is reported as an uncaught exception and the handler continues.
A stdio listener that throws ends the handler, as before, and
terminate() still settles.

Co-authored-by: Alistair Smith <hi@alistair.sh>
@robobun

robobun commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator Author

Status: ready for review.

How I reproduced it

  • new Worker("", { eval: true, stdout: true }) with await worker.terminate() in the worker.stdout 'readable' listener at EOF (script in the PR body). bun 1.4.2 and main print exit 0 and stop. node v26.3.0 and this PR also print terminate() resolved undefined.
  • The reported form, terminate() in a 'message' listener, needs a parent that is held inside new Worker() until the worker has ended. With a 3 s hold, main stops after exit 0 in 3 of 3 runs.
  • The reported form with no hold inside the constructor, on a machine with a load average of about 600: the worker posts 3,000 messages and the parent calls terminate() from the listener of message 1,500. On main the promise stayed pending in 2 of 190 runs. With this PR it settled in 190 of 190.
  • test/js/node/worker_threads/worker_threads.test.ts: 18 of the 25 new tests fail on main. All 167 tests of the file pass with this change (bun bd test).

@coderabbitai

coderabbitai Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: oven-sh/bun/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Essentials

Run ID: 9e4e7d52-e999-42af-836f-3d03377531a0

📥 Commits

Reviewing files that changed from the base of the PR and between ef8e000 and 225acf1.

📒 Files selected for processing (2)
  • src/js/node/worker_threads.ts
  • test/js/node/worker_threads/worker_threads.test.ts

Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 4 remain after this review.


Walkthrough

Worker termination now shares a pending promise across calls and retains the thread ID for exit-time operations. Close handling drains queued messages before closing the port and emitting exit. Guarded listeners report exceptions. Tests cover event ordering, callback forms, thread state, and settlement after exceptions.

Changes

Worker termination lifecycle

Layer / File(s) Summary
Worker identity and shared termination state
src/js/node/worker_threads.ts, test/js/node/worker_threads/worker_threads.test.ts
The Worker retains its thread ID and shares a pending promise across repeated terminate() calls. Tests cover termination forms, return values, warnings, callbacks, and thread state.
Exit draining and guarded settlement
src/js/internal/shared.ts, src/js/node/worker_threads.ts, test/js/node/worker_threads/worker_threads.test.ts
Close handling drains queued messages and settles pending termination in a finally path. Guarded message, exit, and callback exceptions are reported. Tests cover exit handling and settlement when listeners throw or close handling is interrupted.

Suggested reviewers: cirospaciari

Priority: ⬇️ Low

Merge Risk: ⚪ Minimal · up to 225ac

The change addresses termination settlement during worker exit handling. No actionable merge-blocking issue remains in the supplied evidence; merge after normal checks pass.

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly and concisely describes the primary change: settling terminate() when called from the exit handler.
Description check ✅ Passed The description is detailed and directly related to the change. It explains the problem, fix, behavior differences, verification results, tests, performance impact, and limitations. It does not use th…

Comment @coderabbitai help to get the list of available commands.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

Comment thread src/js/node/worker_threads.ts Outdated
Comment thread src/js/node/worker_threads.ts
The exit handler does not run for a Worker whose Bun.ModuleGraph was
disposed, so the copy of the thread id stayed live after the thread
ended. threadId now reads the copy only while the exit handler runs,
before it stores the exit state. At every other time it reads the
native id, as before.
@robobun

robobun commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 4:29 PM PT - Sep 29th, 2026

✅ @robobun, your commit 225acf1da5d7ed30254092228d9e6aa9a762a748 passed in Build #121700! 🎉


🧪   To try this PR locally:

bunx bun-pr 44252

That installs a local version of the PR into your bun-44252 executable, so you can run:

bun-44252 --bun

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I re-reviewed after the follow-up commit and found no new bugs; the second commit addresses the disposed-graph threadId concern (the getter now reads the native id outside the exit handler, with a test for the disposed-graph case). Because this reworks the exit/terminate lifecycle where user listeners re-enter mid-handler, a human look is still worthwhile.

What was reviewed:

  • #onClose/#exit ordering and the finally settle: #exited, terminate callbacks and the pending promise are settled even when a stdio listener throws out of #exit.
  • terminate() called while the exit handler runs (#exiting, not yet #exited): the extra #worker.ref() is a no-op there since WorkerMessagingProxy::setKeepAlive returns early once keep-alive is released, so it cannot pin the event loop.
  • getHeapSnapshot/startHeapProfile switching to #worker.threadId === -1 keeps ERR_WORKER_NOT_RUNNING for both the exited and mid-exit cases.
Extended reasoning...

The change touches src/js/node/worker_threads.ts (Worker exit handler, terminate(), threadId getter) and exports reportUncaughtException from src/js/internal/shared.ts, with ~330 lines of new tests. It touches no auth, crypto or injection surface; the sensitive area is re-entrancy of user JS during the exit handler and node-compat ordering of message/exit/stdio events. The follow-up commit resolved the one non-pre-existing issue I raised earlier and added a regression test for it. I did not approve because the lifecycle rework changes observable ordering versus main in several cells (messages before stdio EOF, threadId/postMessage semantics inside late message listeners), which is a judgment call for a maintainer rather than a mechanical change.

Still open from earlier reviews (1):

  • Unresolved: 1 minor or pre-existing.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants