Conversation
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Pro Run ID: 📒 Files selected for processing (14)
💤 Files with no reviewable changes (2)
Included review availability: Your plan provides up to 5 included reviews per hour; 1 remains after this review. WalkthroughChangesThe change routes poll lifecycle operations through Suggested reviewers: 🚥 Pre-merge checks | ✅ 4✅ Passed checks (4 passed)
Comment |
|
Updated 9:31 PM PT - Sep 29th, 2026
✅ @robobun, your commit 2d3f147ea0f7afcd1bf6e61a8a140d2d7d1d4bd6 passed in 🧪 To try this PR locally: bunx bun-pr 40078That installs a local version of the PR into your bun-40078 --bun |
|
Status: ready for review. 1-bit main/isolated flag per review. The tests run in CI on release builds. Head 2d3f147, merged with main ba3f27d. CI build 121775 passes (181 of 181 jobs). What changed on 2026-09-30:
How it reproduces on current main:
In CI: the two builds of this branch since the merge (121730, 121775) have no annotation for Supersedes #37754. Complements #40508, which fixed the |
1967316 to
a9e1382
Compare
|
Review sweep:
No code change since a9e1382. The test file still passes on the debug build (5 tests). |
|
This bug now also fails Signature (build 116040, alpine 3.23 x64, shard 16): Other runs print Why this file
The sinks die in the GC epilogue, so no synchronous GC is necessary Reproduction on the canary, without CI
Both CI variants follow from the drift. I set it by hand at the first isolated tick, with
A This branch conflicts with |
|
The wrong-epoll Symptom in CI. On the glibc Linux lanes, the 78-file Local repro. 1.4.3-canary.1+b52d51348, the same 78 files, the environment that Trace. An
One instance, worker pid 17509: So the mechanism in this PR explains that flake in full. #38008 removes the same path for the test runner from the other side. It ends each file's stdio sinks at the global swap. Then no sink is left for a finalizer to release inside spawnSync. A fallback at registration is not a substitute. I tried the libuv approach: when const a = Bun.file(0).stream().getReader();
const b = Bun.file(0).stream().getReader();
console.log(await Promise.allSettled([a.read(), b.read()]));On main Repro recipe and tracerFile list: the T=$(mktemp -d)
env -i PATH="$PATH" HOME="$HOME" TMPDIR="$T" BUN_TMPDIR="$T" TEST_TMPDIR="$T" BUN_INSTALL_CACHE_DIR="$T" \
FORCE_COLOR=1 CI=true BUN_FEATURE_FLAG_INTERNAL_FOR_TESTING=1 BUN_DEBUG_QUIET_LOGS=1 \
BUN_GARBAGE_COLLECTOR_LEVEL=1 BUN_JSC_randomIntegrityAuditRate=1.0 BUN_RUNTIME_TRANSPILER_CACHE_PATH=0 \
BUN_ENABLE_CRASH_REPORTING=0 LD_PRELOAD=$PWD/eptrace.so EPTRACE_OUT=/tmp/eptrace.log \
taskset -c 0-2 bun test --parallel=3 --timeout=45000 --dots $FILES
grep -v dup_of=-1 /tmp/eptrace.logNode child processes also report EEXIST on ADD ( // cc -O1 -fPIC -shared -o eptrace.so eptrace.c -ldl -lpthread
// LD_PRELOAD=./eptrace.so EPTRACE_OUT=/tmp/eptrace.log bun test --parallel=3 ...
// Reports, per process: an EPOLL_CTL_DEL that misses while another epoll still
// holds the fd, a close() of a dup of fd 0/1/2 that still has an epoll entry,
// and an EPOLL_CTL_ADD that fails with EEXIST.
#define _GNU_SOURCE
#include <dlfcn.h>
#include <errno.h>
#include <fcntl.h>
#include <pthread.h>
#include <stdarg.h>
#include <stdio.h>
#include <stdlib.h>
#include <sys/epoll.h>
#include <sys/stat.h>
#include <sys/syscall.h>
#include <unistd.h>
#define MAXFD 8192
static int registered_in[MAXFD]; // epoll fd that holds an entry for this fd, 0 = none
static pthread_mutex_t lock = PTHREAD_MUTEX_INITIALIZER;
static long (*real_syscall)(long, ...);
static int (*real_close)(int);
static int (*real_epoll_ctl)(int, int, int, struct epoll_event *);
static int self_pid;
static void init(void) {
if (real_syscall) return;
real_syscall = dlsym(RTLD_NEXT, "syscall");
real_close = dlsym(RTLD_NEXT, "close");
real_epoll_ctl = dlsym(RTLD_NEXT, "epoll_ctl");
}
// A vfork child shares this table with its parent. Ignore it.
static int in_vfork_child(void) {
int pid = (int)real_syscall(SYS_getpid);
if (!self_pid) self_pid = pid;
return pid != self_pid;
}
static void at_fork_child(void) { self_pid = 0; }
__attribute__((constructor)) static void setup(void) { pthread_atfork(NULL, NULL, at_fork_child); }
static int stdio_dup(int fd) {
struct stat a, b;
if (fd <= 2 || fstat(fd, &a)) return -1;
for (int i = 0; i <= 2; i++)
if (!fstat(i, &b) && a.st_dev == b.st_dev && a.st_ino == b.st_ino) return i;
return -1;
}
static void report(const char *what, int fd, int epfd) {
const char *path = getenv("EPTRACE_OUT");
if (!path) return;
char line[256];
int n = snprintf(line, sizeof line, "%s pid=%d tid=%d fd=%d dup_of=%d epfd=%d entry_in_epfd=%d\n", what, self_pid,
(int)real_syscall(SYS_gettid), fd, stdio_dup(fd), epfd, registered_in[fd]);
int out = (int)real_syscall(SYS_openat, AT_FDCWD, path, O_WRONLY | O_APPEND | O_CREAT | O_CLOEXEC, 0644);
if (out < 0) return;
if (write(out, line, n) < 0) {}
real_syscall(SYS_close, out);
}
static void on_close(int fd) {
if (fd < 0 || fd >= MAXFD || in_vfork_child()) return;
int saved = errno;
pthread_mutex_lock(&lock);
if (registered_in[fd] && stdio_dup(fd) >= 0) report("CLOSE_WITH_ENTRY", fd, -1);
registered_in[fd] = 0;
pthread_mutex_unlock(&lock);
errno = saved;
}
static int on_epoll_ctl(int epfd, int op, int fd, struct epoll_event *ev, int raw) {
int rc = raw ? (int)real_syscall(SYS_epoll_ctl, epfd, op, fd, ev) : real_epoll_ctl(epfd, op, fd, ev);
int err = rc ? errno : 0;
if (fd >= 0 && fd < MAXFD && !in_vfork_child()) {
pthread_mutex_lock(&lock);
if (!rc && op != EPOLL_CTL_DEL) registered_in[fd] = epfd;
else if (!rc) registered_in[fd] = 0;
else if (op == EPOLL_CTL_DEL && err == ENOENT && registered_in[fd] && registered_in[fd] != epfd)
report("DEL_ON_OTHER_EPOLL", fd, epfd);
else if (op == EPOLL_CTL_ADD && err == EEXIST) report("ADD_EEXIST", fd, epfd);
pthread_mutex_unlock(&lock);
}
errno = err;
return rc;
}
int close(int fd) {
init();
on_close(fd);
return real_close(fd);
}
int epoll_ctl(int epfd, int op, int fd, struct epoll_event *ev) {
init();
return on_epoll_ctl(epfd, op, fd, ev, 0);
}
// Bun calls close(2) and epoll_ctl(2) through syscall(2).
long syscall(long number, ...) {
init();
va_list ap;
va_start(ap, number);
long a = va_arg(ap, long), b = va_arg(ap, long), c = va_arg(ap, long);
long d = va_arg(ap, long), e = va_arg(ap, long), f = va_arg(ap, long);
va_end(ap);
if (number == SYS_close) on_close((int)a);
if (number == SYS_epoll_ctl) return on_epoll_ctl((int)a, (int)b, (int)c, (struct epoll_event *)d, 1);
return real_syscall(number, a, b, c, d, e, f);
} |
main's #38354 made a failed POSIX `start()` leave the fd with the caller. Take that contract here: - `FileSink::setup` and `PosixStreamingWriter::start` keep main's bodies. This branch's fd-ownership change is gone. `start()` still passes the `EventLoopCtx` to `register_with_fd`. - Drop this branch's filesink.test.ts case. main's "whose registration fails closes the dup exactly once" covers the double close. - `unregister_with_fd_impl` takes main's `disarmed_only` logic and resolves the watcher fd through `loop_mut(vm)`.
|
Merged main (367d939) into this branch at 2cd5c71. The PR body and the status comment describe the result. One part of the merge is not mechanical: main's #38354 fixed the One note for whoever triages a future failure of |
|
Independent confirmation on Bun 1.4.2 ( We initially observed intermittent Save as const { spawnSync } = require('node:child_process');
const closeExplicitly = process.argv[2] === 'end';
(async () => {
for (let round = 0; round < 8; round++) {
for (let i = 0; i < 3; i++) {
const sink = Bun.stdout.writer();
sink.write('test');
if (closeExplicitly) await sink.end();
}
const first = spawnSync('/usr/bin/sleep', ['0.05'], { timeout: 1000 });
const started = performance.now();
const second = spawnSync('/usr/bin/printf', ['expected'], {
timeout: 1000, encoding: 'utf8',
});
if (first.error || second.error || second.status !== 0 || second.stdout !== 'expected') {
console.error(JSON.stringify({
round,
firstError: first.error?.code,
secondError: second.error?.code,
secondStatus: second.status,
stdout: second.stdout,
secondMs: performance.now() - started,
}));
process.exit(1);
}
await new Promise(resolve => setImmediate(resolve));
}
console.error(JSON.stringify({ passed: 8 }));
process.exit(0);
})();Run with stdout connected to a pipe and const { spawn } = require('node:child_process');
const child = spawn('bun', ['repro.cjs', ...process.argv.slice(2)], {
env: { PATH: process.env.PATH, BUN_JSC_collectContinuously: '1' },
stdio: ['ignore', 'pipe', 'pipe'],
});
child.stdout.resume();
child.stderr.pipe(process.stderr);
const guard = setTimeout(() => child.kill('SIGKILL'), 12000);
child.once('exit', code => {
clearTimeout(guard);
process.exitCode = code ?? 1;
});In three runs of the reduced program, two failed with The explicit-close control ( An instrumented version with the same sink/spawn sequence captured this ownership mismatch:
The counter reads used the matching release's Linux x86-64 layout and checked the epoll FD and wakeup back-reference. No memory writes or binary patches were used in this reduced reproduction. This gives an independent stdout-based reproducer for the |
|
CI data for this PR, from a scan of the annotations of the last 400 finished builds (118369 to 119356, all branches). The spin shows up in the
Files with these signatures, by number of builds:
Each of these files calls I could not reproduce the spin on my machine with the same file lists (0 of 420 batch runs), so I have no independent check of the fix. I did not read the diff. |
…c exits On the hosted runners the parallel pass failed once on macOS (4 tests) and once on Ubuntu (5 tests) in the first three runs of #178, never on Windows. Every failure was a test that ran into its own budget, followed by Bun killing a dangling child. Reproduced under WSL Ubuntu with Bun 1.4.0. The full suite at 4 workers, stopped early, had 40 timeouts and 73 dangling children. --parallel=4 over apps and packages/mcp-server hit it in 2 of 5 runs: 3 tests ran into their budgets in one, another hung past 25 minutes. The same files with --isolate and no --parallel passed 441/441. A watcher caught the state: the git child spawned by spawnSync had exited and stayed a zombie, the worker held no pipe any more, and its main thread spun at a full core. This is oven-sh/bun#34069 (open, reproduced on 1.4.0 and 1.4.2; fix proposed in oven-sh/bun#40078): spawnSync loses a child's exit on the posix event loops, and parallel workers multiply the exposure. Only Windows now runs the parallel pass. Linux and macOS run one sequential pass over the same roots, the canonical command plus a JUnit report, and the same completeness check covers both plans: every discovered file reported by exactly one pass, each pass reporting exactly its expected files. Windows is the critical path of the required verdict, so the gain stays; re-enable the other platforms once a fixed Bun is pinned and five hosted runs per OS are green. Refs: HOK-822 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
The spin has a third CI signature, and it is the most frequent one I found: a test that times out at the File. CI data. The flaky annotation names this test in 35 of the 85 builds from 119960 to 120061 that have annotations: 33 on The history follows allocation history, as the Notes of this PR predict:
Reproduction with the unmodified test. Put the fixture of Release canary 6808912 (linux x64 musl), release canary 367d939 (linux x64 glibc) and a debug build give this output. The dangling process is the #43862 makes that test create its FIFO with the |
CI impact of this bugFour flaky test files in CI match the two symptoms in this PR. At least one of the four fails in 279 of the last 400 finished builds (Buildkite builds 120072 to 120698, 2026-09-23 to 2026-09-25). All four fail only in the
How the signatures were matchedThe fixtures below ran on
These fixtures did not run against this branch. State of this PR's CIThe last CI run of this branch (build 117738) is red on one file: Fixtures
const n = Number(process.argv[2]);
const shape = process.argv[3];
globalThis.writers = [];
const refs = [];
for (let i = 0; i < n; i++) {
const writer = Bun.stderr.writer();
writer.write("");
writer.flush();
globalThis.writers.push(writer);
refs.push(new WeakRef(writer));
}
await Bun.sleep(0);
const first = { cmd: ["echo", "first"], stdout: "pipe", stderr: "ignore", timeout: 2000 };
Bun.spawnSync(first);
globalThis.writers = null;
Bun.spawnSync(first);
let collectedDuringCall = 0;
for (let i = 0; i < refs.length; i++) if (refs[i].deref() === undefined) collectedDuringCall++;
const t0 = performance.now();
const result =
shape === "piped"
? Bun.spawnSync({ cmd: ["sh", "-c", "sleep 0.3; echo 'clang version 21.1.8'"], stdin: "ignore", stdout: "pipe", stderr: "pipe", timeout: 5000 })
: Bun.spawnSync({ cmd: ["sh", "-c", "sleep 0.3"], stdin: "inherit", stdout: "ignore", stderr: "inherit", timeout: 5000 });
const ms = Math.round(performance.now() - t0);
console.log(JSON.stringify({ n, shape, collectedDuringCall, ms, exitedDueToTimeout: result.exitedDueToTimeout, stdout: result.stdout?.toString() }));
import { fstatSync, openSync, readdirSync } from "node:fs";
import { $ } from "bun";
const stderrInode = fstatSync(2).ino;
const stderrDups = () =>
readdirSync("/proc/self/fd")
.map(Number)
.filter(fd => {
if (fd === 2) return false;
try {
return fstatSync(fd).ino === stderrInode;
} catch {
return false;
}
});
for (let i = 0; i < 4; i++) Bun.file(2).writer().write("");
const before = stderrDups();
Bun.gc(false);
Bun.spawnSync({ cmd: ["sleep", "0.2"] });
const after = stderrDups();
const occupied = [];
if (process.argv[2] === "occupy") for (let i = 0; i < 4; i++) occupied.push(openSync("/dev/null", "r"));
const shell = [];
for (let i = 0; i < 2; i++) {
const res = await $`seq`.nothrow();
shell.push({ exitCode: res.exitCode, stderr: res.stderr.toString() });
}
let writerError = null;
try {
const w = Bun.file(2).writer();
w.write("writer after\n");
await w.flush();
} catch (e) {
writerError = String(e?.code ?? e);
}
console.log(JSON.stringify({ before, after, occupied, shell, writerError })); |
|
What CI shows. In the last 400 finished builds (121144 to 121668, 2026-09-27 to 2026-09-29) the file fails in the
Probe on main. Release canary
The child exits with status 0 at once. The parent does not poll until the timeout closes the pipes. Not verified. I did not identify which objects supply the count of three in the CI batch. The batch matters: the x64-asan lane runs the file in a different bucket and has 0 failures in the same 400 builds. This branch has a merge conflict with main at this time. |
Conflicts: - src/io/ParentDeathWatchdog.rs: main renamed `Pollable` to `Flags`. The `register` call keeps the `EventLoopCtx` argument of this branch. - src/runtime/socket/WindowsNamedPipe.rs: main replaced the `impl_streaming_writer_parent!` invocation with hand-written Windows impls. Take main's file; the `uws_loop` arm this branch removed has no user left there. - src/runtime/timer/mod.rs: main dropped a clippy allow on `increment_timer_ref`/`increment_immediate_ref`. Keep this branch's signatures, which take the `EventLoopCtx`.
Conflict: src/io/posix_event_loop.rs. main added `Flags::Tty` at the end of the enum where this branch added `Flags::SpawnSyncLoop`. Keep both.
|
Merged main into the branch (0dedb9f). The PR was in conflict with main.
|
A release build of the previous commit had 66,560 more bytes of .text than its merge base on linux-x64. The loop lookup that `KeepAlive::ref_` and `unref` gained was inlined into each of their several hundred callers. `EventLoopCtx::loop_ref` now returns the spawnSync-loop bit itself, and it and `loop_unref_for` are `#[inline(never)]`. A caller keeps a status check and one call. `FilePoll::register_with_fd` and `unregister_with_fd` resolve the loop once and hand it to the private helpers, which take `&mut Loop` again as on main.
|
Follow-up on the binary size, which my last comment left open.
The two builds of this branch since the merge (121730 and 121775) have no annotation for |
Fixes #34069
Problem
Bun.spawnSyncwaits,vm.event_loop_handlepoints at its isolated uws loop.FilePollandKeepAliveresolved their loop through it, so a GC finalizer in that window released main-loop polls against the isolated loop.spawnSyncspins until its timeout: the isolated loop'snum_pollsreads 0, and the tick returns withoutepoll_wait. CI:killed 1 dangling process(10887.test.ts),Cannot tell which /…(build-toolchain-identity.test.ts, 47 of the last 60 builds).EPOLL_CTL_DELalso leaves a stale epoll entry:EEXIST: file already exists, epoll_ctl(bun test --isolate (Linux): intermittent unhandledEEXIST: file already exists, epoll_ctlwhile initializing process.stderr fails test files with no named tests #37968).Fix
Flags::SpawnSyncLoop,KeepAlive::spawn_sync_loop, two bools intimer::All).EventLoopCtx::loop_forresolves the loop from that bit at release.FilePoll::register/unregisterand the timer ref functions take theEventLoopCtx, not a loop pointer.spawnsync-isolated-event-loop.test.tsandspawnSync.test.ts, on every Linux lane. Both fail on main dc3660d (spin, EEXIST).Background
Bun.spawnSyncblocks on a private uws loop, so main-loop timers do not run. The thread's own loop isuws::Loop::get().jsc::EventLooppath.Downsides
KeepAlivegrows from 1 to 2 bytes.FilePollstays 40 bytes.FilePoll::initandKeepAlive::ref_do 2 more loop-pointer lookups each.ref_/unrefmake one call that main inlined. No allocation or syscall..texton linux-x64: 6,656 bytes smaller than main.Notes
Merge with main (ba3f27d, 2026-09-30), in two merge commits (bb97ecd, 0dedb9f), followed by 87f1190 (see Costs). Four files conflicted.
ParentDeathWatchdog.rs: main renamedfile_poll::Pollabletofile_poll::Flags, and theregistercall keeps this branch'sEventLoopCtxargument.WindowsNamedPipe.rs: main replaced theimpl_streaming_writer_parent!invocation with hand-written Windows impls, so main's file is taken whole and this PR no longer touches it.timer/mod.rs: main removed a clippy allow aboveincrement_timer_refandincrement_immediate_ref, and this branch's signatures stay.posix_event_loop.rs: #43868 addedFlags::Ttyat the end of the enum where this branch addedFlags::SpawnSyncLoop. Both stay, and the flag set (25 variants) is still 4 bytes. main's other changes toPipeReader.rs(#43868, #43900, #44086),PipeWriter.rsandFileSink.rsmerged without conflict.Fail-before on current main: a debug build with
src/andpackages/from main dc3660d and this branch's tests.finalizers that run inside spawnSync do not stall the next spawnSyncfails with{"collectedDuringCall":3,"stdout":"","exitedDueToTimeout":true,"exitCode":0}.a writer finalized during spawnSync does not break the next writer on the same fdfails withEEXIST: file already exists, epoll_ctl. The other six cases of the first file pass there. On this branch the two files pass (7 pass with--timeout 120000, and 18 pass with 2 skip).build-toolchain-identity.test.ts: the flaky annotation for the casea file changes when its tool's path resolves to another tool, and only then, with the messageCannot tell which /…, is on 47 of the 60 most recent finished builds (121638 to 121728). The case fails in the parallel batch and passes alone.toolIdentity(scripts/build/tools.ts:217) runschild_process.spawnSync(exe, ["--version"], { timeout: 30_000, stdio: ["ignore", "pipe", "pipe"] })on a two-line/bin/shscript, and the call ends at its own timeout. Stand-in: threeBun.stderr.writer()sinks held onglobalThiswith stderr on a pipe, a warm-upBun.spawnSync, the drop, one moreBun.spawnSyncunderBUN_JSC_slowPathAllocsBetweenGCs=5, then thatchild_process.spawnSyncwith a 3s timeout. Debug build of main dc3660d:ETIMEDOUTafter 3.4s to 3.7s in 3 of 3 runs. Release 1.4.3-canary 367d939:ETIMEDOUTin 3 of 3. This branch:clang version 21.1.8with no error in 0.3s to 0.7s in 9 of 9. With zero writers main returns normally. The test file alone passes 15 of 15 on this branch.Costs, how they were measured. Struct sizes: a compile-time
size_ofprobe on both trees (KeepAlive1 and 2 bytes,FilePoll40 and 40,FlagsSet4 and 4).timer::Allgains two bools, one instance per VM. Lookups, counted from the diff:FilePoll::initreads the ctx's loop pointer anduws_get_loop()(a thread-local read) once each to set the bit, where main did no lookup.register_with_fdandunregister_with_fdresolve the loop once, as their callers did on main.EventLoopCtx::loop_refdoes the same two reads asinitbefore the ref, andloop_unref_fordoes one lookup, asloop_unrefdid. Binary size:sizeand a per-symbolnm -Sdiff on linux-x64 release builds of the merge base ba3f27d and of this branch, same checkout, onlysrc/swapped. The build is reproducible (two builds of 0dedb9f are byte-identical). The first merged revision, 0dedb9f, had 66,560 more bytes of.textthan the base (58,217,461 against 58,150,901). The lookups thatKeepAlive::ref_andunrefgained were inlined with them into 572 symbols: thenode:fsasync tasks, work-pool jobs, dns, sockets and the other users. CI's binary-size annotation on build 121730 agrees (+64.0 KB on linux-x64 and on linux-x64-musl). In 87f1190EventLoopCtx::loop_refreturns the bit itself, and it andloop_unref_forare#[inline(never)], so a caller keeps a status check and one call.FilePoll::register_with_fdandunregister_with_fdresolve the loop once and pass&mut Loopto the private helpers, as on main. Result for 2d3f147:.text58,144,245 bytes, 6,656 fewer than the base, and the file is 8,192 bytes smaller. CI's binary-size annotation on build 121775 compares against main #121761. No target is larger. linux-x64 is 12.0 KB smaller and linux-aarch64 is 64.0 KB smaller.strace,perf,valgrindandbloatyare not installed in this container, so there is no instruction or syscall count from a tool.Suites on the merged debug build (0dedb9f), on a host under heavy load:
spawnsync-isolated-event-loop.test.ts(7 pass, where the two #40508 cases take 10s to 13s undercollectContinuouslyon this build and on main, above the 5s default),spawnSync.test.ts(18 pass, 2 skip),process-stdio.test.ts(18 pass), thenspawn.test.ts,10887.test.ts,filesink.test.ts,node-timers.test.ts,spawnsync-no-microtask-drain.test.tsandbuild-toolchain-identity.test.tsin one run with--timeout 120000(278 pass, 8 skip, 1 fail). The one failure is thewith BUN_FEATURE_FLAG_FORCE_WAITER_THREADwrapper inspawn.test.ts, which hit its own 192s cap. The same inner run, started directly with a 60s per-test timeout, passes (177 pass, 9 skip, 210s). The time goes to the 200 debugbun -echildren that several of its cases start.cargo checkpasses forx86_64-pc-windows-msvc,x86_64-apple-darwinandaarch64-unknown-linux-musl. Again on 87f1190: the two test files, the stand-in above (3 of 3 return normally),10887.test.ts,filesink.test.ts,node-timers.test.ts,spawnsync-no-microtask-drain.test.ts,build-toolchain-identity.test.ts,process-stdio.test.tsandspawn-streaming-stdin.test.ts(120 pass in one run),spawn.test.tsin waiter-thread mode (177 pass, 9 skip) and in the default mode (177 pass, 8 skip, and the same wrapper at its 192s cap), and the same threecargo checktargets. One run ofspawnSync.test.tsfailedshould not set a timeout if timeout is 0 and inherit is used for stdouton itstoBeLessThan(1000)wall-clock check (1190ms for a debug child). AspawnSyncofsleep 0.3takes 300ms to 1600ms on this host with the release builds of main and of this branch alike.Earlier summary, moved here from the visible part. The user-visible face (#34069): under
bun test --isolate, a file that touchesprocess.stderrand callsBun.spawnSynctimes out after 5000ms, or the run hangs until it is killed.spawnSyncnever returns, the child stays a zombie, and the parent spins at 100% CPU. Other signatures of the spin and of the stale epoll entry:✗ does not segfault [90001.51ms](10887.test.ts), and in the shellShellError: Failed with exit code 17(#40611). An epoll entry is keyed by the open file and the fd number. The kernel drops it only when the last fd for that open file closes. A dup of stdio never reaches that point while the process runs, so a missedEPOLL_CTL_DELleaves the entry behind for the next dup that lands on the same number.FilePollis one registration of an fd on a loop.KeepAliveis the on/off loop ref used by servers, sockets, fetch, fs and dns.FilePoll::activateanddeactivatetake theEventLoopCtxtoo, andSpawnSyncEventLoop::prepareasserts in debug builds that the isolated loop's counters are zero before each spawn. This PR supersedes #37754 (same cause, predates #38883) and #40611 (sameFilePollcause, stored a loop pointer instead of the bit, and its test is folded in here). main fixed the double close that #40611 also reported in #38354. This PR's tests still fail on main with #40508, becauseFilePoll,KeepAliveand the timer counters resolve their loop throughEventLoopCtx, not through theEventLoop.The notes below predate the merge of 2026-09-30.
How the bucket was identified: every flaking build ran the same 65-file batch, and in that batch
10887.test.tswas the tenth file of the third worker, right afterweb-globals.test.js(fourstdin: "pipe"children), the napitest_objectsuite (four pipeless spawnSyncs), 013880, 04893 and 05828. 11 of the last ~110 PR builds flaked on at least one of debian 13 x64/aarch64 and ubuntu 25.04 x64/aarch64, no flake on any build with a different bucket. #38883 saw the same state on 07261.Repro of the mechanism: with a debug build instrumented to record, per
FilePoll, the loop it was activated on, a forcedcollectNow(Sync)at the end of spawnSync loggedFileSink__finalize → PosixStreamingWriter::drop → FilePoll::deinit_force_unregister → deactivateagainst the isolated loop for polls activated on the main loop, thenprepare: num_polls=-1on the next spawnSync and a pipeless napi spawnSync spinning withnum_polls=0. Without the forced GC the drift did not reproduce locally in ~100 batch runs (release canary and debug, with CI'sBUN_GARBAGE_COLLECTOR_LEVEL=1), so the CI trigger is GC timing. The isolated loop is reused for every spawnSync in the process, which is why the drift accumulates.The asan abort (#40611): 15 of 16 asan builds, always the worker that ran
test/regression/issue/20753.test.jsright after19652.test.ts(aBun.spawnSync). The GC epilogue runs inside spawnSync on the first JS allocation after the wait (create_resource_usage_object) and finalizes the previous test file'sprocess.stderrFileSink. Timeline from an instrumented asan build:epoll ADD ok (default loop),epoll DEL ENOENT (other loop)fromFileSink::finalize,Closer::closeon the work pool,epoll ADD EEXIST (default loop)for the nextBun.file(2).writer(),FileSink::setupcloses the fd, then the writer's Drop schedules a second close that fails with EBADF. The panic was visible in the worker log but not in annotations: the runner re-ran the crashed files alone and marked the job flaky. On macOS kqueue knotes are per fd, so the EEXIST symptom cannot appear there; thespawnSync.test.tscase is Linux-only for that reason. The wrong-loopEV_DELETEand the counter drift applied on macOS too and are fixed by the same change.Merge with main (367d939): main's #38354 fixed the double close with the opposite fd contract. A failed POSIX
start()now leaves the fd with the caller, andFileSink::setupcloses it. This branch first made the writer own the fd. The merge takes main's contract:FileSink::setupandPosixStreamingWriter::startkeep main's bodies, and this branch'sfilesink.test.tscase is gone because main's "whose registration fails closes the dup exactly once" covers it. The EEXIST half stays here, with itsspawnSync.test.tscase.unregister_with_fd_impltakes main'sdisarmed_onlylogic and resolves the watcher fd throughloop_mut(vm). On the merged debug build:spawnsync-isolated-event-loop.test.ts(the new case passes; the two #40508 cases still need more than the 5s default on a local debug build),spawnSync.test.ts(18 pass, 2 skip),filesink.test.ts(68 pass),spawn-pipe-start-error.test.ts(9 pass, 1 skip),spawn.test.ts(144 pass, 8 skip),isolation.test.ts(41 pass),node-timers.test.ts(24 pass),10887.test.ts,spawn-streaming-stdin.test.tsandspawnsync-no-microtask-drain.test.ts.child_process.test.tshas 65 pass and 2 fail:extra stdio pipes are not double-closed on GCneeds 6.4s on this debug build and passes with a longer timeout, andshould allow us to spawn in the default shellfails because$SHELLis empty in the test environment of this container.cargo check --workspacepasses forx86_64-pc-windows-msvcandaarch64-apple-darwin.Stand-in for the user-visible face: 30 files under
bun test --isolate --smolwith stderr on a pipe. Each file touchesprocess.stderrandprocess.stdout, allocates 20000 small objects, then callsBun.spawnSync(["/bin/true"])30 times and counts the fds on the stderr pipe around each call. Release build of main b52d513: 11 of 20 runs were killed at the 25s limit after(fail) t04 [5019.48ms]or(fail) t08. Release build of this branch merged with main: 0 of 20 hung and 0 failed. In 4 of those 20 runs, 13 sinks of earlier files were finalized inside aBun.spawnSynccall: the fd count and the main loop'snumPollsboth dropped across the call. A second variant reads, at the end of every file, the main loop'snumPollsminus the fds on the stderr pipe. On main the value moves from -2 to 0 and 1 in the runs that hang. On this branch it is -2 in all 600 files of 20 runs, and in 8 of those runs sinks were finalized inside a call. The same stand-in on the debug build never finalizes a sink inside a call, so it shows nothing there.The script from #34069:
spawnsync-spin.test.mjsfrom the comment of 2026-09-19, unchanged, in the shape its runner uses (8 copies of the file,bun test --isolate --timeout 0, stdout on a pipe, 3 such processes at once, a process that is silent for 40s counts as wedged). Release builds, blocks of 6 process-runs that alternate between the two binaries, Linux x64. main 367d939: 15 of 72 wedged, each in state R with its CPU time rising 5s per 5s and oneshchild a zombie. This branch (2cd5c71): 0 of 72. bun 1.4.2: 5 of 6.Why
BUN_JSC_slowPathAllocsBetweenGCs: a finalizer runs inside spawnSync only when JSC sweeps at one of the allocations spawnSync makes after the wait.LocalAllocator::doTestCollectionsIfNeededrunscollectNow(Sync, Full), which sweeps synchronously, every N slow-path allocations. Two details make the window deterministic. The first spawnSync call of a process creates lazy structures andBun.spawnSyncitself, about 11 slow-path allocations before the loop swap, so the fixture warms spawnSync up while the writers are still reachable; the next call then makes no allocation before the swap (measured with markers inprepare/cleanupandBUN_JSC_logGC=1: 0 GCs before the swap, 3 inside at N=5) and the first forced GC after the drop lands inside it. The probe child sleeps 300ms so the parent has to poll for its exit; withechothe child could exit before the polls were registered and the unfixed build returned without spinning. N=5 spins on the unfixed release, canary and main (with #40508) builds in 105 of 105 runs with the allocation phase shifted by 0 to 150 objects of five kinds. N=1 also works but makes a debug build spend 12s in startup GCs; N=5 costs about 1s. An earlier revision used a debug-onlyrun_gc(true)hook behindBUN_INTERNAL_SPAWN_SYNC_GC, which no CI lane exercised; the hook is gone.Why a bit and not a pointer: an earlier revision stored the loop pointer on every
FilePoll,KeepAliveand timer counter and dropped the ctx argument fromunref. Review asked for a 1-bit identifier instead, since only two loops exist per thread. The bit also leaves the existingref_(ctx)/unref(ctx)call sites unchanged.Isolated-loop lifetime: every poll that records
SpawnSyncLoopis created and freed inside the samespawnSynccall. The stdio readers and the process poller start afterprepare()swaps the handle in,SubprocessT::finalizeruns before thecleanup()scopeguard swaps it back, and the syncSubprocessnever gets a JS wrapper.loop_forasserts this in debug builds, andKeepAlive::unref_on_next_tick(deferred throughpending_unref_counter, drained on the thread's loop) asserts its ref was not taken on the isolated loop. Timer refs are never taken during spawnSync, since no JS runs there, but a finalizer can release one, so the timer counters record the bit too.Windows:
VirtualMachine::uws_loop()returnsLoop::get()unconditionally there, sois_spawn_sync_loop()is always false and every release resolves to the same loop as before. The WindowsFilePollis unchanged.Why the timer takes the
EventLoopCtxtoo:timer::Allfirst carried its own copy of the spawnSync predicate and release path over a raw*mut Loop, without the release-after-return debug assert. It now records the bit fromEventLoopCtx::is_spawn_sync_loopand releases throughloop_unref_for, likeFilePollandKeepAlive. The four callers (timer_object_internals.rs,subprocess.rs, two indns.rs) already had the per-thread ctx in reach.Why the fixture uses WeakRefs and not
bun:internal-for-testingornode:fs:fileSinkInternals.liveCount()or an fd listing would give exact counts, but loading either module underslowPathAllocsBetweenGCs=5costs a debug build 5s to 15s (1s without). AWeakRefmade before a job boundary (await Bun.sleep(0)) does not keep its target alive afterwards (JSCcurrentWeakRefVersion), andderef()after the probe call returnsundefinedexactly when the writer was collected during the call. The writers hang offglobalThisuntil the drop: a module binding that is only written after the await is dead at the suspension point, and the writers were collected in the warm-up call instead (measured: 0 of 3 alive after the warm-up). With the global, an unfixed build spins withcollectedDuringCall: 3for every padding tried before the warm-up (0, 2, 5 and 40 objects). What the check cannot see is a GC that runs after the drop but before the loop swap: an allocation between the drop and the call (the comment forbids code there), or a slow-path allocation in spawnSync's argument parsing, of which there is none today (N=1 also spins).Scope: the drift requires a teardown under the isolated loop. GC finalizers of
FileSink,Subprocess,FetchTasklet,MessagePortand similar are the only code that runs there since #38883 moved the bun:test timeout callback out of the wait.EventLoopCtx::loop_unref,loop_add_activeandloop_sub_activeare gone on POSIX (the last two remain for the WindowsFilePoll).EventLoopCtx::platform_event_loopandPosixStreamingWriterParent::loop_(with theuws_looparm ofimpl_streaming_writer_parent!) lost their last callers and are removed.Fail-before was checked against a debug build of the merge base (6d13a3a) in a second worktree, the same shape as the gate's stash run: all three tests this branch had then fail there (spin, EEXIST,
err.is_none()abort) and pass on this branch. The third one left with the merge above.Suites run on the debug build:
spawn.test.ts(140 pass, 8 skip),spawnSync.test.ts(17 pass, 2 skip),filesink.test.ts,spawn-streaming-stdin.test.ts,spawn-stdin-readable-stream-sync.test.ts,spawn-stdin-pipe-fd-leak.test.ts,node-timers.test.ts,dns-resolver-concurrent-timeout.test.ts,dns-lookup-keepalive.test.ts,dns-interleave.test.ts,spawn-kill-signal.test.ts,spawn-maxbuf.test.ts,spawn-signal.test.ts,spawnsync-isolated-event-loop.test.ts,spawnsync-no-microtask-drain.test.ts,spawn-streaming-stdout.test.ts,spawn-streaming-stdin.test.ts,spawn-stdin-readable-stream-sync.test.ts,exit-code.test.ts,spawn-kill-signal.test.ts,spawn-many-teardown.test.ts,10887.test.ts,filesink.test.tsandbun-write.test.js(123 pass),child_process.test.tsandchild_process-node.test.js,node-timers.test.ts.resolve-dns.test.tsneeds network access this container lacks (getaddrinfo ETIMEOUT).cargo checkpasses for linux-gnu x64, darwin x64 and windows x64. Rebased onto main after #40508: the only conflicts were the test file's imports. #40508's two new tests in the same file pass on this branch (OK, exit 0) but need more than the 5s default timeout on a local debug build (8s and 11s undercollectContinuously). The same applies toextra stdio pipes are not double-closed on GCinchild_process.test.ts(passes alone in 9s).no test proof · iteration 7 · platform-specific test(s) that do not run on this machine, deferring to CI, which covers all platforms: test/js/bun/spawn/spawnsync-isolated-event-loop.test.ts, test/js/bun/spawn/spawnSync.test.ts