Skip to content

node:fs: release the { signal } of readFile/writeFile on the JS thread, not on the pool - #37420

Open
robobun wants to merge 2 commits into
mainfrom
farm/c33c1c51/fs-signal-ref-js-side
Open

robobun wants to merge 2 commits into
mainfrom
farm/c33c1c51/fs-signal-ref-js-side

Conversation

@robobun

@robobun robobun commented Aug 11, 2026 •

Copy link
Copy Markdown
Collaborator

Symptom

fs.readFile / fs.promises.readFile / writeFile / appendFile called with { signal } inside a worker that is torn down (terminate(), process.exit(), uncaught error) while those operations are still queued on the thread pool crash the process once the pool gets to them. Same thing on the main thread under BUN_DESTRUCT_VM_ON_EXIT=1. Affects every platform (these ops use AsyncFSTask everywhere).

Repro: test/js/node/fs/abort-signal-read-write-file-worker-teardown-fixture.ts, run with UV_THREADPOOL_SIZE=2. It parks both pool threads in a readFile() of a FIFO (the open("w") handshake proves they are inside those reads), has a worker queue readFile/writeFile/appendFile calls with three kinds of signals behind them, terminates the worker, then closes the FIFOs. On main, a debug build dies with both of these at once (one per pool thread):

ASSERTION FAILED: Thread::mayBeGCThread()
src/jsc/bindings/webcore/EventListenerMap.h(77) : void WebCore::EventListenerMap::releaseAssertOrSetThreadUID()

panic: VirtualMachine.get() called with no VM on this thread

with the stack <bun_runtime::node::fs::args::ReadFile as Drop>::drop -> ExternalShared<AbortSignal>::drop -> WebCore__AbortSignal__unref -> WebCore::AbortSignal::~AbortSignal() on a pool thread. The first is a RELEASE_ASSERT, so release builds die the same way for a signal that has abort listeners.

Cause

args::ReadFile / args::WriteFile owned the AbortSignalRef (and the pending-activity count), and the arguments are part of the job's off-thread half. job.rs releases a job's JS side on the VM's own thread at teardown and leaves the off-thread half to whoever holds it; for a job still queued on the pool that is a pool thread, after the VM, its heap and the signal's wrapper are gone. The job's ref is the last one by then, so ~AbortSignal() runs on the pool thread against a dead VM:

  • a signal with abort listeners: ~EventTarget trips EventListenerMap's thread check (release assert), and each JSEventListener releases a JSC::Weak into the freed heap;
  • an aborted signal: m_reason releases a JSC::Weak into the freed heap (heap corruption in release);
  • AbortSignal.timeout(): cancelTimer() -> AbortSignal__Timeout__deinit resolves the VM through the calling thread's thread-local, which a pool thread does not have;
  • in every case the non-atomic refcount is decremented off its thread.

Fix

The arguments keep only a pointer to the signal, and the only thing the operation does with it is poll aborted(). The flags byte behind aborted() is std::atomic now (relaxed; GC marker threads were already reading it through JSAbortSignalOwner), so the pool's poll is a plain atomic load instead of the racy byte read it was before.

Keeping the signal alive while the operation is pending is args::SignalHold: GC protection on the wrapper plus the pending-activity count the code already took (it is what marks a timeout signal as observed). It is taken when the arguments are parsed and released on the JS thread in every case: a synchronous call drops it with the arguments, and AsyncFSTask::create moves it onto the job's JS side (AsyncFSJs.signal), which the completion releases, or JobList::release_all_js at teardown, same as the promise. The off-thread half of a job freed after its VM is gone holds nothing of the signal any more. No allocation, no listener.

Why the retention is a protection on the wrapper rather than a ref of our own held on the JS side: teardown releases the JS sides (release_all_js) before VmHandle::close() waits for bodies still running on the pool. Unprotecting only makes the wrapper collectible, and nothing collects it before the VM is destroyed, which happens after those bodies have finished, so a body still polling keeps reading a live signal; that is the same thing every other job's off-thread pointers rely on. Dropping what might be the last ref at that point would free the signal under a running body (a signal whose wrapper was already collected mid-operation, e.g. a throwaway controller).

ReadFile/WriteFile no longer need a Drop impl; the three ReadFile::default()-then-assign sites that existed because of it are struct literals now (clippy's field_reassign_with_default fires on them otherwise).

Tests

In test/js/node/fs/promises.test.js:

  • "signals of operations still queued on the pool when their worker is torn down" (POSIX; the FIFO setup makes the queued state deterministic, and the fixture reports how many of the worker's writes ran, which must be 0 since released jobs never run: without UV_THREADPOOL_SIZE=2 it reports 12, so a run that did not actually park the pool fails instead of passing vacuously). Fails on main as above (a release build of main dies in ~170ms), passes with the fix. Runs the child with Malloc=1 so the heap side is visible to ASAN as well.
  • "readFile aborted while in flight keeps its reason when the caller holds nothing" (all platforms): a readFile whose controller, signal and reason are referenced by nothing else. While it is in flight exactly one AbortSignal is protected (heapStats().protectedObjectTypeCounts), after two GCs the rejection still carries its cause, and afterwards the protection is gone. Fails on main: nothing is protected, and on a release build of main the cause also comes back undefined (20/20; the reason is only reachable through the wrapper, which nothing kept alive).
  • "readFile/writeFile/appendFile release their signal once they have finished": plain, listener-bearing and timeout signals through all seven signal-taking entry points (promise, callback and sync forms); afterwards the protected count is back at its baseline and a GC plus a wrapper count checks the pending activity (a count that is never released pins two of the three per iteration; verified by temporarily removing the unref: 51 live against a bound of 25 at the time). Passes before and after; it guards the release side of the new ownership.

Also green locally on the debug/ASAN build: fs.test.ts, promises.test.js, test/js/web/abort/, the fetch abort tests, worker-refused-completion.test.ts, the Node test-fs-{promises-,}readfile* / writefile* files; cargo clippy -p bun_runtime, clang-format on the header, and a cargo check of the Windows target.

Related but different: #36259 and #35021 add more aborted() polls between chunks and compose with this. #36983 / #37170 / #34154 predate #37075's job model and address in-flight or blocked work; #37170's refused-post disposal also leaks the signal ref with ManuallyDrop instead of releasing it, which this supersedes (its signal arm can be dropped on rebase).

@robobun

robobun commented Aug 11, 2026 •

Copy link
Copy Markdown
Collaborator Author

Status: reproduced on main with the fixture in this PR (UV_THREADPOOL_SIZE=2, debug build): ASSERTION FAILED: Thread::mayBeGCThread() in EventListenerMap::releaseAssertOrSetThreadUID plus panic: VirtualMachine.get() called with no VM on this thread, both from ~AbortSignal() running on a pool thread out of args::ReadFile's Drop. Fix is 18401a3 (atomic flags byte, pointer in the arguments, protection + pending activity held on the JS thread); 1655472 adds a cross-platform in-flight test and makes the teardown fixture self-checking. CI is green on 1655472 (build 92264, 190/190). Open question above is whether to keep the protected field; the probe showing the in-flight teardown use-after-free without it is in the thread.

@coderabbitai

coderabbitai Bot commented Aug 11, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Walkthrough

Summary

Asynchronous Node filesystem operations now retain abort signals on the JS thread and share atomic cancellation flags with worker threads. Tests cover signal cleanup and worker teardown. Related ReadFile construction sites use direct struct initialization.

Changes

Filesystem abort lifecycle

Layer / File(s) Summary
Signal ownership and cancellation state
src/runtime/node/node_fs.rs
SignalHold owns signal listeners and activity. AbortFlag shares cancellation state with workers. ReadFile, WriteFile, and AppendFile transfer signals and poll atomic flags.
Async task signal transfer
src/runtime/node/node_fs.rs, src/runtime/api/js_bundle_completion_task.rs
AsyncFSTask moves signal state into AsyncFSJs. Completion checks the stored signal. Sourcemap writes pass no abort flag.
Call-site updates and lifecycle tests
src/runtime/socket/SSLConfig.rs, src/runtime/webcore/Blob.rs, src/runtime/webcore/fetch.rs, test/js/node/fs/*
ReadFile call sites use struct initialization. Tests verify signal cleanup and queued-operation safety during worker teardown.

Possibly related PRs

Suggested reviewers: jarred-sumner, dylan-conway

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly and concisely describes the main fix: releasing readFile and writeFile signals on the JavaScript thread instead of the thread pool.
Description check ✅ Passed The description clearly explains the symptom, cause, fix, and verification tests, although it uses different headings from the repository template.

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@test/js/node/fs/abort-signal-read-write-file-worker-teardown-fixture.ts`:
- Around line 51-52: Update the queued readiness promise around
worker.once("message") so it rejects when the worker emits "error" or exits
early with a nonzero status before posting the ready message. Preserve
resolution on the expected "message" event and ensure all failure events are
wired to rejection so the fixture cannot hang.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 5ac79998-5497-4609-9103-549ee4365904

📥 Commits

Reviewing files that changed from the base of the PR and between 9fcdea8 and 85e7ef3.

📒 Files selected for processing (7)
  • src/runtime/api/js_bundle_completion_task.rs
  • src/runtime/node/node_fs.rs
  • src/runtime/socket/SSLConfig.rs
  • src/runtime/webcore/Blob.rs
  • src/runtime/webcore/fetch.rs
  • test/js/node/fs/abort-signal-read-write-file-worker-teardown-fixture.ts
  • test/js/node/fs/promises.test.js

Comment thread test/js/node/fs/abort-signal-read-write-file-worker-teardown-fixture.ts Outdated
Comment thread src/runtime/node/node_fs.rs Outdated
Comment thread src/runtime/node/node_fs.rs Outdated
Comment thread src/runtime/node/node_fs.rs Outdated
Comment thread src/runtime/node/node_fs.rs Outdated
Comment thread src/runtime/node/node_fs.rs Outdated
Comment thread src/runtime/node/node_fs.rs Outdated
@robobun

robobun commented Aug 11, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 6:18 AM PT - Aug 11th, 2026

✅ @robobun, your commit 16554722e4bd09dcfe31adf5aeee4f609fbc4d3f passed in Build #92264! 🎉


🧪   To try this PR locally:

bunx bun-pr 37420

That installs a local version of the PR into your bun-37420 executable, so you can run:

bun-37420 --bun

@robobun

robobun commented Aug 11, 2026

Copy link
Copy Markdown
Collaborator Author

Pushed two follow-ups to the review comments: the fixture now rejects its readiness promise if the worker errors or exits before queueing (a worker-side failure exits the fixture with that error instead of sitting there until the test times out), and the new doc comments in node_fs.rs are down to one line each; the ownership rationale stays in the commit message and the description. The flagged comment threads are resolved since those blocks no longer exist. No change to the fix itself; promises.test.js re-run on the debug build.

@github-actions

Copy link
Copy Markdown
Contributor

This PR may be a duplicate of:

  1. Don't block worker/VM teardown on fs thread-pool ops that never complete #37170 - Also disarms the AbortSignal held by args::ReadFile/args::WriteFile in node_fs.rs so it is not deref'd on a pool thread against a dead VM, using a conflicting mechanism (ManuallyDrop on the refused path vs. moving the hold to AsyncFSJs.signal).

🤖 Generated with Claude Code

Comment thread test/js/node/fs/abort-signal-read-write-file-worker-teardown-fixture.ts Outdated
@robobun

robobun commented Aug 11, 2026

Copy link
Copy Markdown
Collaborator Author

Not a duplicate of #37170. That PR is about teardown hanging on fs ops that never complete, and it is written against the pre-#37075 task model (ConcurrentPosterGate / fs_discard; it no longer merges). Its signal handling is a side effect of its refused-post disposal: disarm_for_dead_vm takes the AbortSignalRef out of the arguments and leaks it with ManuallyDrop, so the ref and the pending-activity count are never released for a job the pool frees after the VM is gone, and a job whose body is still running at teardown keeps reading the signal from the pool thread.

This PR changes who owns the signal under the current job.rs model: the hold lives on the job's JS side and is released by the VM on its own thread (completion or release_all_js), nothing is leaked, and the pool thread only ever touches an atomic flag. If #37170 is rebased it can drop its signal arm; the rest of it is independent of this change.

@Jarred-Sumner Jarred-Sumner left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This adds a heap allocation and adds more state to track. Not necessary.

Feels like the better design here is:

  • Continue to ref/unref on main thread, only before entering/exiting JS
  • Make it so an AbortSignal can read the aborted boolean atomically across threads

@robobun

robobun commented Aug 11, 2026

Copy link
Copy Markdown
Collaborator Author

Agreed, reworking it that way: no allocation and no listener. The pool will poll aborted() on the signal itself (making the flags byte atomic on the C++ side, which also covers the reads JSAbortSignalOwner already does from GC threads), and the arguments will carry only the pointer; the retention plus the pending-activity ref/unref stay on the JS thread, taken when the operation is created and released in its completion (or by the VM's teardown). Will push shortly.

…d, not on the pool

readFile, writeFile and appendFile with { signal } kept the AbortSignal ref
(and its pending-activity count) in the arguments, which travel with the
off-thread half of the pool job. A job still queued on the pool when its
worker is torn down is freed later by a pool thread, so that thread released
what was by then the last ref and ~AbortSignal ran against a destroyed VM:
EventListenerMap's thread check fires for a signal with abort listeners,
AbortSignal.timeout's timer deinit looks the VM up on the pool thread, and an
aborted signal's reason releases a JSC::Weak into the freed heap.

The arguments now carry only a pointer to the signal, and all the operation
does with it is poll aborted(), whose flags byte is atomic on the C++ side
(GC marker threads were already reading it through JSAbortSignalOwner). What
keeps the signal alive while the operation is pending (GC protection on the
wrapper, plus the pending-activity count as before) is args::SignalHold:
a synchronous call drops it with the arguments, and AsyncFSTask::create moves
it onto the job's JS side, which the completion or the VM's teardown releases
on the JS thread like the job's other JS state. The retention is a protection
rather than a ref of our own because teardown releases JS sides before it
waits for bodies still running on the pool: unprotecting only makes the
wrapper collectible, and nothing collects it before the VM is destroyed,
after those bodies are done, whereas dropping a last ref would free the
signal under a body still polling it.

ReadFile/WriteFile no longer need Drop impls; the three default-then-assign
construction sites that existed because of them are struct literals now.
@robobun
robobun force-pushed the farm/c33c1c51/fs-signal-ref-js-side branch from 1a12b87 to 18401a3 Compare August 11, 2026 06:06
@robobun

robobun commented Aug 11, 2026

Copy link
Copy Markdown
Collaborator Author

Reworked as suggested and force-pushed as one commit (18401a3), description updated.

  • AbortSignal::m_flags is std::atomic<uint8_t> (relaxed), so aborted() is readable from the pool; the Rust-side doc on aborted() says so. Same accessors, so the reads JSAbortSignalOwner already did from GC threads are covered too.
  • The arguments carry a pointer and poll aborted() through it, as before. No Arc, no listener.
  • What stays on the JS thread is args::SignalHold: the pending-activity ref/unref that was already there, plus GC protection on the wrapper to keep the signal alive while the operation is out. Taken when the arguments are parsed, dropped with them for sync calls, moved onto the job's JS side by AsyncFSTask::create for async ones, so it is released in the completion or by the VM's teardown, never on the pool.

One deliberate detail: the retention is a protection rather than a C++ ref held on the JS side, because teardown releases the JS sides before it waits for bodies still running on the pool. Unprotecting leaves the wrapper alive until the VM is destroyed (after those bodies are done), while dropping a last ref there would free the signal under a body still polling it. Details in the description.

Both tests unchanged in substance; the teardown one got an explicit timeout since a worker start plus teardown alone is ~3s on a debug build and it timed out once on a loaded machine.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this and didn't find any bugs. Given it reworks cross-thread ownership of AbortSignal (JS-side release via AsyncFSJs, a JsPtr polled from the pool, and m_flags made atomic), a human look at the lifetime/ordering argument would still be worthwhile.

What was reviewed:

  • The Signal/SignalHold split and take_signal plumbing into AsyncFSJs; pending_activity_ref/unref pairing on all paths (sync drop, completion, release_all_js).
  • The JsPtr<AbortSignal> SAFETY claim vs. teardown ordering — checked that Protected release only unprotects the wrapper (finalization is a later GC/VM-destruction step, after bodies under a borrow finish); ruled out as a bug.
  • AbortSignal.h: m_flags → std::atomic<uint8_t> with fetch_or/fetch_and; the consolidated setFlag preserves the previous per-flag setters' behavior, and existing GC-thread reads (JSAbortSignalOwner) now go through the same atomic load.
  • The three ReadFile { .. } FRU call sites and the fixture (error/exit now reject the readiness promise — earlier comment addressed).
Extended reasoning...

Overview

The PR fixes a crash where ~AbortSignal() runs on a thread-pool thread against a dead VM when a worker with queued readFile/writeFile/appendFile({ signal }) is torn down. The fix moves the signal's retention (GC protection + pending-activity count) from the off-thread args::ReadFile/WriteFile into the job's JS side (AsyncFSJs.signal), so it is released on the JS thread by completion or JobList::release_all_js. The pool-thread half keeps only a JsPtr<AbortSignal> and polls aborted(), which is now an atomic load (m_flags becomes std::atomic<uint8_t>). ReadFile/WriteFile lose their Drop impls, so three internal call sites switch back to functional-record-update. Two new tests: a FIFO-parked worker-teardown fixture (POSIX) and a heap-stats leak guard over all seven signal-taking entry points.

Security risks

None identified. This is internal lifetime management; no user-controlled input reaches new parsing or allocation.

Level of scrutiny

High. Per REVIEW.md this sits squarely in the most-blocked category — cross-thread ownership of a refcounted C++ object with GC handles, a new JsPtr whose validity depends on the ordering between release_all_js, wrapper finalization, and pool bodies still holding a Borrow, and a change to C++ atomics/memory ordering. The approach was also reworked mid-review (from Arc<AtomicBool> + native listener to atomic m_flags + JsPtr + Protected), so the PR description is one revision behind the code. That is exactly the kind of subtle invariant a maintainer familiar with job.rs teardown should sign off on.

Other factors

The fixture-readiness comment I left earlier is addressed (worker error/exit now reject). The comment-cop threads are resolved (the flagged blocks are gone). The bug-hunting verifier examined the JsPtr SAFETY claim against teardown and refuted it as a bug. CI is still building per the last robobun status. This supersedes the signal-leak arm of #37170 (noted in the description); a maintainer may want to weigh the interaction.

@robobun

robobun commented Aug 11, 2026

Copy link
Copy Markdown
Collaborator Author

The description was refreshed right after the force-push and matches 18401a3 now (atomic flags byte, pointer in the arguments, SignalHold on the JS side); the earlier Arc + listener version is gone from both the code and the text. CI on this head is down to its last few jobs with nothing red so far; the three yellow entries (inspect-error-leak timing out in the ASAN parallel batch, webview-chrome on alpine aarch64, fetch-leak's RSS threshold on Windows aarch64) passed on retry and are unrelated to this change.

@Jarred-Sumner

Copy link
Copy Markdown
Collaborator

The GC protection on the wrapper is not needed. The wrapper is kept alive while there are pending native callbacks due to has pending activity.

@robobun

robobun commented Aug 11, 2026

Copy link
Copy Markdown
Collaborator Author

I tried exactly that: dropped the protection and kept only the ref + pending activity on the JS side (main's hold, relocated). It is not enough in one case, and it reproduces.

Pending activity keeps the wrapper alive only while the signal has an abort listener (the nesting in JSAbortSignalOwner::isReachableFromOpaqueRoots from #13687), so for a listener-less signal the wrapper can be collected while the operation is out, and the job's ref is then the last one. At teardown release_all_js drops it before close() waits for bodies still running on the pool, so a body's next aborted() poll reads a freed signal. With the ref-only variant, ASAN (Malloc=1), 3 of 3 runs:

ERROR: AddressSanitizer: heap-use-after-free ... READ of size 1 ... thread T14 (Bun Pool 0)
    WebCore::AbortSignal::aborted() const   webcore/AbortSignal.h:90
    WebCore__AbortSignal__aborted           bindings.cpp
    <bun_runtime::node::fs::NodeFS>::read_file_with_options   node_fs.rs:7281
freed by thread T11 (Worker) here: ...

Same script against the current head (with the protection): clean, 3 of 3. That is what the protection is for: it makes the pointer the body polls valid until the VM is destroyed, which is the lifetime the job model already assumes for everything else a body reads off-thread (buffers, wrappers). Main does not have this particular read because main keeps the ref in the off-thread half, which is the original bug.

So the choices as I see them: keep the one protected field (current head), or make teardown wait for in-flight bodies before it releases JS sides, which fixes this class generically but changes the teardown sequence from #37075. I have left the PR on the first; glad to switch to the second if you would rather have that. The probe below is not in the PR as a test because it needs a sleep to order teardown ahead of the poll.

probe (mkfifo a FIFO, then Malloc=1 bun probe.js /path/to/fifo on an ASAN build)
// Probe for the residual teardown window: a body still running on the pool
// (blocked reading a FIFO) whose signal's wrapper was already collected, while
// the worker is torn down. release_all_js drops the last ref; the body then
// polls aborted() once its 256 KiB pre-stat buffer fills. Run with Malloc=1
// under ASAN. (A probe, so it uses a sleep to order teardown before the poll.)
const fs = require("node:fs");
const { Worker } = require("node:worker_threads");

const fifo = process.argv[2];
const CHUNK = 256 * 1024;

const worker = new Worker(
  `
  const fs = require("node:fs");
  const { parentPort, workerData: fifo } = require("node:worker_threads");
  function start() {
    // Nothing references the controller or the signal after this returns.
    const ac = new AbortController();
    fs.promises.readFile(fifo, { signal: ac.signal }).then(() => {}, () => {});
  }
  start();
  function churn(n) { let a = []; for (let i = 0; i < n; i++) a.push({ i }); return a.length; }
  churn(1000); Bun.gc(true); churn(1000); Bun.gc(true);
  const { heapStats } = require("bun:jsc");
  parentPort.postMessage(heapStats().objectTypeCounts.AbortSignal ?? 0);
  `,
  { eval: true, workerData: fifo },
);

worker.once("message", async liveSignalWrappers => {
  console.log("signal wrappers alive in worker after gc:", liveSignalWrappers);
  const wfd = fs.openSync(fifo, "w");
  // All but the last byte: the body consumes these and blocks in read().
  fs.writeSync(wfd, Buffer.alloc(CHUNK - 1, 0x61));
  await Bun.sleep(200);
  console.log("body parked in read() with the buffer one byte short");
  const terminated = worker.terminate();
  // Let teardown release the job's JS side and start waiting for the body.
  await Bun.sleep(1000);
  // Now the body fills its buffer and polls the signal.
  fs.writeSync(wfd, Buffer.from("b"));
  fs.closeSync(wfd);
  await terminated;
  console.log("worker torn down");
});

…n fixture self-checking

- A readFile whose controller, signal and reason are referenced by nothing
  else: the operation shows up as exactly one protected AbortSignal while
  in flight, the rejection keeps its cause across a GC, and the protection
  is gone afterwards. Fails on main (nothing is protected; release builds
  also lose the cause).
- The release test also checks the protected count returns to its baseline.
- The teardown fixture counts the files the worker's writes would have
  produced: jobs released by the teardown never run, so the count must be
  zero, and a run where the pool was not actually parked fails instead of
  passing vacuously (without UV_THREADPOOL_SIZE=2 it reports 12).
@robobun

robobun commented Aug 11, 2026

Copy link
Copy Markdown
Collaborator Author

Pushed 1655472 (tests only) after going over the coverage once more:

  • New cross-platform test next to the in-flight abort tests: a readFile whose controller, signal and reason nothing else references. It checks that exactly one AbortSignal is protected while the operation is out, that the rejection still has its cause after GCs, and that the protection is gone afterwards. On main nothing is protected; on a release build of main the cause also comes back undefined (20/20 runs), since the reason is only reachable through the wrapper and nothing kept the wrapper alive. That is the user-visible side of the same retention discussed above.
  • The release test also checks the protected count returns to its baseline.
  • The teardown fixture now prints how many of the worker's writes actually ran, pinned to 0 in the expectation (released jobs never run). Without the UV_THREADPOOL_SIZE=2 parking it reports 12, so the test can no longer pass vacuously if the parking ever stops working.

Fix commit unchanged.

@robobun

robobun commented Aug 14, 2026

Copy link
Copy Markdown
Collaborator Author

Re-checked on current main (97a4363) after the teardown rework in #38299. Part of this is covered now, part is not, so this stays open but should shrink on rebase.

Covered by #38299: the crash in the title. A job's off-thread half, the args with their signal ref included, is now always dropped on the worker's own thread with the heap alive, either in then or in Job::release_unrun during the teardown wait, so nothing releases the signal from a pool thread any more. Observed on a debug+ASAN build of main:

  • The fixture as written no longer crashes but deadlocks: terminate() now waits for the worker's queued jobs (the debug build reports 60 ticket(s) still held off-thread, taken at src/runtime/node/node_fs.rs:1324), and the fixture only releases the parked pool after terminate() resolves. Re-sequenced to close the FIFOs while terminate() is pending, it passes with Malloc=1 in 3 of 3 runs: worker torn down with code 1, 0 of its writes ran, no assertion failure or panic.
  • The same 15 { signal } operations (fs.promises readFile/writeFile/appendFile plus callback readFile/writeFile, each with an aborted signal, one with a listener, and an AbortSignal.timeout()) whose jobs did run and whose completions came back during the wait (BUN_DEBUG_TEST_WORKER_TEARDOWN_GATE) are released cleanly as well.
  • "readFile/writeFile/appendFile release their signal once they have finished" passes on main.

Not covered: the in-flight liveness half. Nothing keeps the signal's wrapper alive while the operation is pending, so "readFile aborted while in flight keeps its reason when the caller holds nothing" still fails on main (the protected AbortSignal count stays at 0), and when the stale stack slots are overwritten before the GC runs, the rejection comes back as a bare ABORT_ERR with no cause in 50 of 50 runs on a debug build of main (181 of 200 on a release canary). The SignalHold on the JS side and the atomic flags are still needed for that; the teardown motivation, the worker fixture and its test can be dropped when rebasing.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants