ci: detect stalled test batches by silence; fix worker panic under --bail - #36232
Conversation
|
Warning Review limit reachedYou’ve reached a temporary PR review limit under our Fair Usage Limits Policy. Next review available in: 19 minutes Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Pro Run ID: 📒 Files selected for processing (1)
WalkthroughChangesThe runner now applies job-budget-aware and output-idle timeouts to parallel Bun processes. Worker reaping preserves crash accounting during bail handling, with regression coverage for parallel worker crashes and abort reporting. Parallel execution reliability
Possibly related PRs
Suggested reviewers: 🚥 Pre-merge checks | ✅ 4✅ Passed checks (4 passed)
Comment |
There was a problem hiding this comment.
Beyond the inline nit, I also checked: jobBudgetMs() with BUILDKITE_TIMEOUT unset or already exceeded — the Infinity fallback and outer Math.max(60_000, …) floor keep the batch timeout sane; and the Coordinator change still routes SIGTERM'd siblings (from a prior panic's kill loop) through account_unfinished, since SIGTERM isn't in is_panic_status.
Extended reasoning...
The reported finding is a harmless orphaned timer; the runner exits via process.exit() so it can't hold anything open. I traced a few adjacent edges: (1) when BUILDKITE_TIMEOUT is unset, parseInt("") → NaN → jobBudgetMs() returns Infinity, and Math.min(X, Infinity - 300000) = X, so behavior matches the old cap; when the budget is already negative the outer 60 s floor applies. (2) On the spawnSafe EBUSY retry path, the "spawn" event never fired so idleTimer was never armed — no leak across the recursive retry. (3) In Coordinator.rs, workers killed by abort_on_worker_panic's SIGTERM still hit was_bailed && !panicked → account_unfinished, matching the old behavior; only genuinely panicked workers now go through account_crash + banner even when bailed was already set, which is the stated intent.
| timer = setTimeout(expire, timeout); | ||
| if (options.idleTimeout) { | ||
| armIdleTimer = () => { | ||
| clearTimeout(idleTimer); | ||
| idleTimer = setTimeout(() => { | ||
| idledOut = true; | ||
| expire(); | ||
| }, options.idleTimeout); | ||
| }; | ||
| armIdleTimer(); | ||
| } |
There was a problem hiding this comment.
🟡 When idleTimer fires and calls expire(), the graceful branch reassigns timer = setTimeout(SIGKILL, 15s) without first clearing the original timer = setTimeout(expire, timeout) armed just below — that wall-clock timer (potentially tens of minutes for a large batch) is orphaned and fires later against the already-dead subprocess. Harmless in practice since the runner ends with an explicit process.exit() and kill() on a dead child is a no-op, but a clearTimeout(timer) at the top of expire() (alongside clearTimeout(idleTimer)) makes the invariant hold for both entry points.
Extended reasoning...
What changed the invariant
Before this PR, the timeout callback body was only ever reached as the wall-clock timer's own callback — by the time it ran, timer had already fired, so reassigning timer = setTimeout(SIGKILL, 15s) in the graceful branch was safe: there was nothing armed under the old handle.
This PR extracts that body into expire() and adds a second entry point: idleTimer's callback now also invokes expire(). On that path, the original timer = setTimeout(expire, timeout) at scripts/runner.node.mjs:1521 is still armed when expire() runs. expire() clears idleTimer but not timer, then overwrites the timer variable with the 15 s SIGKILL escalation handle — orphaning the original wall-clock timeout.
Step-by-step
- Batch spawns;
timer = setTimeout(expire, timeout)arms with, say, a 40-minute cap.idleTimerarms with 4 minutes. - The batch wedges silently. After 4 minutes
idleTimerfires →idledOut = true→expire(). expire()runsclearTimeout(idleTimer), then (since the batch always setsgracefulTimeout: trueand this is POSIX) sends SIGTERM and doestimer = setTimeout(SIGKILL, 15_000). The 40-minute handle is now unreferenced but still scheduled.- Subprocess exits on SIGTERM →
done()runsclearTimeout(timer)— buttimernow points at the 15 s escalation timer, so only that is cleared. The 40-minute timer remains armed. - ~36 minutes later (if the runner is still alive) the orphaned timer fires →
expire()again → SIGTERM to a dead PID (no-op,kill()returnsfalse), arms yet another 15 s timer, andresolve()is a no-op on the already-settled promise.
Why nothing else catches it
done() is the only place that clears timer, and it clears whatever handle the variable currently holds. Once expire() reassigns it, nothing retains a reference to the original timeout. The retry loop and the outer promise are already resolved by the time the orphan fires, so it can't corrupt results — it's purely a leaked timer.
Impact
Low. runner.node.mjs ends with an explicit process.exit(), so the orphan can't hold the process open, and the late expire() call has no observable effect (dead subprocess, settled promise). The only downside is a leaked timer per stalled batch and potentially confusing noise if someone is debugging with per-timer logging — a spurious SIGTERM against a long-dead subprocess minutes after the batch was reported.
Fix
One line — add clearTimeout(timer); at the top of expire() next to clearTimeout(idleTimer):
const expire = () => {
timedOut = true;
clearTimeout(timer);
clearTimeout(idleTimer);
if (options.gracefulTimeout && !isWindows) {
...That makes expire() idempotent regardless of which timer invoked it.
There was a problem hiding this comment.
Actionable comments posted: 3
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
scripts/runner.node.mjs (1)
1508-1516: 🩺 Stability & Availability | 🟠 Major | ⚡ Quick winClear the original wall-clock timer before installing the grace timer.
An idle expiry calls
expire()before the wall-clock timeout. The original timer is then overwritten, sodone()only clears the 15-second grace timer; the original timer keeps the runner alive and later callsexpire()again.Proposed fix
const expire = () => { timedOut = true; + clearTimeout(timer); clearTimeout(idleTimer); if (options.gracefulTimeout && !isWindows) {🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@scripts/runner.node.mjs` around lines 1508 - 1516, Update the expire function in the timeout handling flow to clear the original wall-clock timer before replacing it with the 15-second graceful-shutdown timer. Preserve the existing idle-timer cleanup and graceful termination behavior while ensuring the original timer cannot fire again after expire() runs.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@scripts/runner.node.mjs`:
- Line 989: Update the idleTimeout configuration in the runner options to parse
BUN_RUNNER_BATCH_IDLE_MS with Number rather than parseInt, rejecting empty,
non-integer, and negative values while intentionally preserving an explicit
zero; only use the four-minute default when the environment value is absent or
invalid.
- Around line 985-988: Update the timeout calculation in the batch runner around
jobBudgetMs() so it never exceeds the remaining Buildkite budget after applying
the reserve. Remove or adjust the outer 60-second minimum, and skip the batch
when the remaining budget cannot support that minimum timeout.
In `@test/cli/test/parallel.test.ts`:
- Around line 203-216: Update the parallel bail panic fixture around the
Bun.spawn command to disable core dumps before launching the test process on
POSIX, using the existing issue-30205 approach where applicable. Preserve the
panic and sibling-worker behavior, and skip the core-dump suppression or provide
an equivalent Windows-specific setup so the test remains hermetic across
platforms.
---
Outside diff comments:
In `@scripts/runner.node.mjs`:
- Around line 1508-1516: Update the expire function in the timeout handling flow
to clear the original wall-clock timer before replacing it with the 15-second
graceful-shutdown timer. Preserve the existing idle-timer cleanup and graceful
termination behavior while ensuring the original timer cannot fire again after
expire() runs.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: 8da39f67-8122-491d-8a5f-361622a20f07
📒 Files selected for processing (3)
scripts/runner.node.mjssrc/runtime/cli/test/parallel/Coordinator.rstest/cli/test/parallel.test.ts
| timeout: Math.max( | ||
| 60_000, | ||
| Math.min(Math.max(10 * 60_000, bucketFiles.length * 5_000), jobBudgetMs() - 5 * 60_000), | ||
| ), |
There was a problem hiding this comment.
🩺 Stability & Availability | 🟠 Major | ⚡ Quick win
Do not exceed the remaining Buildkite budget.
With 30 seconds remaining, this computes a 60-second timeout because the outer Math.max(60_000, ...) applies after the budget cap. Skip the batch when the reserve is exhausted, or cap the final timeout by the remaining budget.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@scripts/runner.node.mjs` around lines 985 - 988, Update the timeout
calculation in the batch runner around jobBudgetMs() so it never exceeds the
remaining Buildkite budget after applying the reserve. Remove or adjust the
outer 60-second minimum, and skip the batch when the remaining budget cannot
support that minimum timeout.
| 60_000, | ||
| Math.min(Math.max(10 * 60_000, bucketFiles.length * 5_000), jobBudgetMs() - 5 * 60_000), | ||
| ), | ||
| idleTimeout: parseInt(process.env.BUN_RUNNER_BATCH_IDLE_MS || "", 10) || 4 * 60_000, |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win
Parse BUN_RUNNER_BATCH_IDLE_MS strictly.
"0" silently becomes the four-minute default, while values such as "100ms" are accepted as 100. Parse the present value with Number, reject empty/non-integer/negative values, and handle an explicit zero intentionally.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@scripts/runner.node.mjs` at line 989, Update the idleTimeout configuration in
the runner options to parse BUN_RUNNER_BATCH_IDLE_MS with Number rather than
parseInt, rejecting empty, non-integer, and negative values while intentionally
preserving an explicit zero; only use the four-minute default when the
environment value is absent or invalid.
| test( | ||
| "--parallel --bail: a worker panic still prints the panic banner and stops sibling workers", | ||
| async () => { | ||
| using dir = tempDir("parallel-bail-panic", { | ||
| "a-hang.test.js": `import {test} from "bun:test"; test("hang", async () => { await new Promise(() => {}); }, 999999);`, | ||
| "b-panic.test.js": `import {test} from "bun:test"; test("panic", () => { process.kill(process.pid, "SIGSEGV"); });`, | ||
| }); | ||
| await using proc = Bun.spawn({ | ||
| cmd: [bunExe(), "test", "--parallel=2", "--bail=1"], | ||
| env: { ...bunEnv, BUN_TEST_PARALLEL_SCALE_MS: "0" }, | ||
| cwd: String(dir), | ||
| stderr: "pipe", | ||
| stdout: "pipe", | ||
| }); |
There was a problem hiding this comment.
🩺 Stability & Availability | 🟠 Major | ⚡ Quick win
Suppress deliberate crash core dumps in this fixture.
Unlike test/regression/issue/30205.test.ts:153-181, this SIGSEGV fixture leaves core dumps enabled. In --coredump-upload lanes, scripts/runner.node.mjs will detect the worker-produced core and fail this enclosing test. On POSIX, wrap the command with ulimit -c 0 && exec "$@"; skip or provide a Windows-specific fixture.
As per coding guidelines, tests must be hermetic.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@test/cli/test/parallel.test.ts` around lines 203 - 216, Update the parallel
bail panic fixture around the Bun.spawn command to disable core dumps before
launching the test process on POSIX, using the existing issue-30205 approach
where applicable. Preserve the panic and sibling-worker behavior, and skip the
core-dump suppression or provide an equivalent Windows-specific setup so the
test remains hermetic across platforms.
Source: Coding guidelines
| if (timedOut && (!error || error === signalCode || /^code \d+$/.test(error))) | ||
| error = idledOut ? "stalled" : "timeout"; |
There was a problem hiding this comment.
🟡 The new "stalled" error label never reaches any output: the only caller passing idleTimeout (the parallel batch in runParallelBucket) destructures error at line 973 and never references it again, and spawnBun only reads result.error to decide whether to overwrite it with "core dumped". So the PR description's "a silence stop reads stalled, a wall-clock one timeout" isn't actually delivered — either surface error in the retry log / annotation when !ok, or drop the idledOut distinction.
Extended reasoning...
What the PR added vs. what is consumed
spawnSafe now sets error = idledOut ? "stalled" : "timeout" at scripts/runner.node.mjs:1700-1701, and the PR description promises a "distinct annotation title for the new case: a silence stop reads stalled, a wall-clock one timeout."
idledOut can only become true when options.idleTimeout is set. Grepping the file, idleTimeout is passed at exactly one call site — the parallel batch spawn inside runParallelBucket (via spawnBun → spawnSafe). No other spawnSafe caller sets it, so the "stalled" value can only flow back to that one destructure.
The consumer drops it on the floor
At scripts/runner.node.mjs:973 the batch caller does:
const { ok, error, stdout, crashes } = await startGroup(...)Reading runParallelBucket end to end (lines 973–1120), error is never referenced again after that destructure. The recovery path branches on ok, suites.size, and the parsed stdout/junit; the retry log at ~line 1096 ("parallel bucket: retrying … / no junit and no streamed evidence …") does not include error; the per-file BuildKite annotation at ~line 1111 does not include it either. startGroup (utils.mjs) simply returns the callback's promise value without printing it, and spawnBun (lines 1800–1821) reads result.error only to decide whether to overwrite it with "core dumped" in the coredump branch — it never emits it.
Step-by-step trace
- Batch wedges silently → after 4 min
idleTimerfires →idledOut = true→expire()→ SIGTERM → subprocess exits. spawnSafecomputeserror = "stalled"and returns{ok: false, error: "stalled", stdout, …}.spawnBunreceives it, checksif (isLinux && result.exitCode…)for coredumps (not taken), returns the result unchanged.startGroupreturns it unchanged.runParallelBucketdestructures{ok, error, stdout, crashes}and then executes the!okrecovery: parses junit / streamed interrupt report, computesfailed/incomplete, logs the yellow retry line, re-runs solo.erroris not read at any of those steps.
The wall-clock-timeout path (error = "timeout") reaches the exact same code with the exact same observable output. There is no annotation, log line, or group title that distinguishes the two.
Why this matters (and why it's only a nit)
Per REVIEW.md: "trace new state to an actual consumer — parsed-but-never-read is a red flag" and "Delete dead code in the same PR that makes it dead." The idledOut boolean and the "stalled" branch are new state with no consumer, and the PR description overstates what shipped.
That said, the stall-detection mechanism itself works: the idle timer fires, SIGTERM triggers the coordinator's interrupt report (which names the hung file), and runParallelBucket retries it solo. Nothing breaks if this merges as-is; the only casualty is the promised observability distinction. Hence nit, not blocking.
Fix
Either wire it up — e.g. include error in the yellow retry log when !ok:
console.log(`${getAnsi("yellow")}parallel bucket ${error ? `(${error})` : ""}: ${evidence ? …}`);— or, if the distinction isn't worth keeping, drop idledOut and the "stalled" branch and update the PR description.
| subprocess.stdout.on("data", chunk => { | ||
| armIdleTimer?.(); | ||
| const text = chunk.toString("utf-8"); | ||
| stdout?.(text); | ||
| buffer += text; | ||
| }); | ||
| subprocess.stderr.on("data", chunk => { | ||
| armIdleTimer?.(); | ||
| const text = chunk.toString("utf-8"); | ||
| stderr?.(text); | ||
| buffer += text; |
There was a problem hiding this comment.
🟡 Addendum to the orphaned-timer nit above: on the exitCode === 0 path, done() clears idleTimer but leaves armIdleTimer callable and does not destroy the streams — since Node's "exit" can fire before stdio drains, a trailing "data" event then re-arms a fresh 4-minute idleTimer that nothing ever clears (the clean-exit branch bypassed beforeDone's counter, so the subsequent "end" won't call done() again). Same harmless class (runner ends via process.exit(), late expire() on a dead subprocess is a no-op), but the clearTimeout(timer)-in-expire() fix from the earlier comment doesn't cover this path — also set armIdleTimer = undefined in done().
Extended reasoning...
Relation to the earlier comment
The earlier nit on this PR is about the wall-clock timer handle being orphaned when idleTimer → expire() reassigns timer without clearing the original. This is a distinct leak on a different path: here it's idleTimer itself being re-armed after done() on the clean-exit path, and the fix suggested there (clearTimeout(timer) at the top of expire()) does not address it.
The path
On exitCode === 0, the "exit" handler calls done(resolve) directly (runner.node.mjs:1544), bypassing beforeDone(). done():
- clears
idleTimer(line 1466) ✅ - unrefs stdout/stderr but does not destroy them (the destroy branch at 1470–1476 requires
exitCode === undefined) - does not null out
armIdleTimer
Per Node's child_process contract, "exit" can fire while stdio streams still have buffered data — this file's own beforeDone() machinery exists precisely for that ordering, but the exitCode === 0 branch bypasses it. So a trailing stdout/stderr "data" event can arrive after done() and each one calls armIdleTimer?.() (lines 1551/1557), arming a fresh 4-minute idleTimer.
The subsequent stdout.on("end") → beforeDone() only increments doneCalls from 0 to 1 (the clean-exit path called done() directly, so the counter is still 0; 0 === 1 is false), so done() does not run again and nothing ever clears the re-armed idleTimer.
Step-by-step
- Parallel batch spawns with
idleTimeout: 4 * 60_000.armIdleTimeris defined and armsidleTimer. - Batch runs to completion, coordinator exits with code 0. Node fires
"exit"with(0, null). "exit"handler takes the else branch →done(resolve).done()clearstimerandidleTimer, unrefs (but does not destroy) stdout/stderr, resolves the promise.doneCallsis still 0.- A last buffered stdout chunk (e.g. the tail of the summary) is delivered →
"data"handler runsarmIdleTimer?.()→idleTimer = setTimeout(expire, 240_000). stdoutemits"end"→beforeDone()→doneCalls++ === 1?0 === 1→ false.done()is not called;idleTimeris not cleared.- 4 minutes later (if the runner is still alive) the orphaned
idleTimerfires →idledOut = true→expire()→ SIGTERM to a dead PID (no-op), arms a 15 s SIGKILL timer, andresolve()on the settled promise is a no-op.
The parallel-batch spawn is the only spawnSafe caller that sets idleTimeout, and it produces voluminous output ending on stdout, so this ordering is realistic there.
Why the earlier fix doesn't cover it
Adding clearTimeout(timer) at the top of expire() makes expire() idempotent for whichever timer invoked it. It does nothing here: the leak is idleTimer being re-armed after done() on a path that never reaches expire() again until the orphan fires. The re-arm happens in the "data" handler via armIdleTimer, which done() leaves callable.
Impact
Same harmless class as the earlier nit:
runner.node.mjsends with an explicitprocess.exit(), so an orphaned timer cannot hold the process open.- The late
expire()mutatestimedOut/idledOuton a closure whosespawnSafehas already returned, andkill()on a dead child returnsfalse.
No observable failure — just a leaked timer per clean-exit batch and a spurious late SIGTERM against a dead PID.
Fix
One line in done(), alongside the existing clearTimeout(idleTimer):
const done = resolve => {
if (timer) clearTimeout(timer);
clearTimeout(idleTimer);
armIdleTimer = undefined; // <-- so post-teardown "data" events no-op
...(Optionally the same at the top of expire(), so the SIGTERM-window "Interrupted while still running:" chunks don't redundantly re-arm either — though done() does clear that one on the way out.)
| const runnerStartedAt = Date.now(); | ||
| const jobBudgetMs = () => { | ||
| const minutes = parseInt(process.env.BUILDKITE_TIMEOUT || "", 10); | ||
| if (!Number.isFinite(minutes) || minutes <= 0) return Infinity; | ||
| return minutes * 60_000 - (Date.now() - runnerStartedAt); | ||
| }; |
There was a problem hiding this comment.
🟡 BUILDKITE_TIMEOUT is measured from when the agent starts the job (bootstrap → checkout → artifact download), but runnerStartedAt is captured at module load — so jobBudgetMs() overestimates the remaining budget by the pre-runner setup time. If that setup exceeds the 5-minute headroom, a still-producing-output batch's wall-clock cap can outlive the job, so "can never outlive the job" doesn't strictly hold. Still a strict improvement (and the 4-min idleTimeout covers the actual incident); consider widening the headroom constant or noting that it must absorb pre-runner setup.
Extended reasoning...
What the calculation misses
jobBudgetMs() computes the remaining Buildkite job budget as BUILDKITE_TIMEOUT × 60_000 − (Date.now() − runnerStartedAt), where runnerStartedAt = Date.now() at module load (scripts/runner.node.mjs:77). But BUILDKITE_TIMEOUT is the step's timeout_in_minutes (e.g. 45 for darwin/windows at .buildkite/ci.mjs:849), and Buildkite counts that from when the agent accepts the job — which includes bootstrap, git checkout, pre-command hooks, and artifact download of the built binary. All of that happens before runner.node.mjs is even loaded, so none of it is reflected in Date.now() − runnerStartedAt.
The result is that jobBudgetMs() systematically overestimates the true remaining budget by the pre-runner setup duration S. The batch cap at scripts/runner.node.mjs:985-988 is min(5s × N, jobBudgetMs() − 5min), so if S > 5min the cap can exceed the job's actual remaining time and Buildkite kills the job before expire() fires.
Step-by-step example
Take a 45-minute darwin shard where checkout + artifact download takes S = 7 min, and the batch is dispatched T = 2 min after runner.node.mjs loads (so 9 min into the job's real clock):
jobBudgetMs()returns45 − 2 = 43min (it thinks only 2 min have elapsed).- Actual job remaining is
45 − 7 − 2 = 36min. - For ~800 files,
5s × 800 ≈ 67min, so the batch cap ismin(67, 43 − 5) = 38min. - 38 min > 36 min actual remaining, so if the batch is actively producing output (idle timer keeps re-arming) and simply runs long, Buildkite kills the job at the 36-min mark — 2 minutes before
expire()would have fired to name the hung file and trigger the solo retry.
Why the existing safeguards don't fully cover it
The 4-minute idleTimeout catches a silent batch regardless of jobBudgetMs() accuracy — that's the primary defense and it handles the actual incident from build 84092. But the wall-clock cap is the backstop for a batch that keeps emitting output while running long (e.g. one file with many slow tests), and that backstop is off by S. There's no standard Buildkite env var exposing the job-start wallclock (only BUILDKITE_TIMEOUT and BUILDKITE_JOB_ID are available), so the runner can't measure S directly.
Impact
Low — this is a nit, not a blocker:
- Strict improvement over the pre-PR state: the old cap was
max(10min, 5s × N) ≈ 68min on a 45-min job; the new cap is bounded to ~40 + S min. Even with the overestimate, the window where Buildkite can kill the job beforeexpire()fires shrank from ~23 min tomax(0, S − 5)min. - Narrow trigger: requires pre-runner setup > 5 min and a batch that continuously produces output up to the cap. The reported failure mode (a wedged/silent batch) is caught by
idleTimeoutindependently. - The PR description's "can never outlive the job" is an overstatement of the guarantee, not a code bug per se.
Fix
Either widen the headroom constant (e.g. jobBudgetMs() − 10 * 60_000 to absorb typical checkout + artifact-download time), or add a comment on the − 5 * 60_000 noting it must cover pre-runner setup. A harder fix would inject the job-start timestamp from the pipeline step (RUNNER_JOB_STARTED_AT) so jobBudgetMs() can subtract from the real clock.
Two failure modes from the first day of parallel batches in CI: - A darwin shard died at the Buildkite job timeout inside an 813-file batch: the batch cap (5 s per file) was ~68 min on a 45-minute job, so the interrupt/name-the-hung-file/retry net never fired. A wedged batch is silent, so watch for that instead of wall clock: spawnSafe gains an idleTimeout re-armed on every output chunk; the batch aborts gracefully after 4 minutes of silence, naming the hung file, and the wall-clock cap is now bounded by the job's remaining time (BUILDKITE_TIMEOUT) so the backstop can't outlive the job either. - On a worker panic, whether it was the panic path or --bail that set `bailed` was checked after account_crash — which can itself flip `bailed` via bail_out() — so the Nth failure being a panic skipped the panic banner and left in-flight siblings running. Snapshot `bailed` first, and run abort_on_worker_panic on every panic (its contract: runs even when already bailed).
--parallel --bail=1 where the first failure is a worker killed by a fatal signal, with a sibling worker inside a hanging file: the panic banner must print and the sibling must be stopped. With the previous gate the sibling ran on and the run never finished (fails the test's timeout); with the fix it exits promptly non-zero.
c426d1d to
3983787
Compare
| await using proc = Bun.spawn({ | ||
| cmd: [bunExe(), "test", "--parallel=2", "--bail=1"], | ||
| env: { ...bunEnv, BUN_TEST_PARALLEL_SCALE_MS: "0" }, | ||
| cwd: String(dir), | ||
| stderr: "pipe", | ||
| stdout: "pipe", | ||
| }); |
There was a problem hiding this comment.
🔴 This SIGSEGV fixture omits the ulimit -c 0 wrapper and BUN_CRASH_REPORT_URL: "" / BUN_ENABLE_CRASH_REPORTING: "0" env overrides that the near-identical precedent at test/regression/issue/30205.test.ts:153-181 uses for the same abort_on_worker_panic path. On Linux Buildkite (where --coredump-upload defaults on), spawnBun in runner.node.mjs scans coresDir before/after and marks the enclosing test error: "core dumped" when the worker's core appears — so parallel.test.ts fails on that lane. This is orthogonal to the Windows-skip comment above; even with test.skipIf(isWindows) added, POSIX coredump lanes still red. Copy the ["/bin/sh", "-c", 'ulimit -c 0 && exec "$@"', "--", …] cmd wrapper and the two env keys from 30205.
Extended reasoning...
What the fixture does vs. what the precedent does
The new test at test/cli/test/parallel.test.ts:203-223 has a worker execute process.kill(process.pid, "SIGSEGV") under --parallel=2 --bail=1 and asserts the panic banner. That is the same abort_on_worker_panic code path already covered by test/regression/issue/30205.test.ts:153-181, whose comment (lines 160-163) says exactly why it wraps the spawn:
CI lanes with coredump-upload flag any new core file in coresDir as a test failure — including the one the worker deliberately produces here.
ulimit -c 0on the coordinator is inherited by the workers; the test is POSIX-only so /bin/sh is available.
The 30205 test therefore spawns as ["/bin/sh", "-c", 'ulimit -c 0 && exec "$@"', "--", bunExe(), "test", …] and passes env: { ...bunEnv, …, BUN_CRASH_REPORT_URL: "", BUN_ENABLE_CRASH_REPORTING: "0" }. The new test does neither: it spawns bunExe() directly and its env is { ...bunEnv, BUN_TEST_PARALLEL_SCALE_MS: "0" }. test/cli/run/run-crash-handler.test.ts:396-402 uses the identical wrapper for the identical reason, so this is an established harness convention.
Why the core is actually written
process.kill(pid, "SIGSEGV") delivers a real signal — it does not go through the internal js_segfault/js_panic test hooks in crash_handler_jsc.rs that call suppress_core_dumps_if_necessary(). Bun's crash handler catches SIGSEGV, prints the trace, then re-raises via SIG_DFL, and the kernel writes a core when RLIMIT_CORE > 0. Nothing in this test lowers that limit; bunEnv (harness.ts:64-90) even sets ASAN_OPTIONS with disable_coredump=0, so ASAN doesn't suppress it either.
Why the runner then fails parallel.test.ts
scripts/runner.node.mjs on Linux Buildkite defaults --coredump-upload on (line ~181: isBuildkite && isLinux). Inside spawnBun (lines ~1804-1826) it does readdirSync(coresDir) before and after the child, and if a new core appeared sets result.ok = false and result.error = "core dumped"; "core dumped" is in isAlwaysFailure so the retry gate does not save it. The scan is directory-wide, so the grandchild worker's core is attributed to the enclosing parallel.test.ts invocation.
Step-by-step on a Linux Buildkite lane
runner.node.mjsrecordsexistingCores = readdirSync(coresDir)and spawnsbun test test/cli/test/parallel.test.ts.- The new test spawns
bun test --parallel=2 --bail=1; worker B loadsb-panic.test.jsand executesprocess.kill(process.pid, "SIGSEGV"). - Bun's crash handler runs, then re-raises SIGSEGV via
SIG_DFL; kernel writes/var/bun-cores-…/core.<pid>(ulimit unchanged). - Coordinator sees
is_panic_status(Signaled(SIGSEGV)) == true, prints the banner, kills the sibling — the assertions inside the test pass. parallel.test.tsfinishes;spawnBunre-scanscoresDir, finds the new core, and overwrites the result to{ ok: false, error: "core dumped" }.parallel.test.tsis reported as failed with errorcore dumped;isAlwaysFailure("core dumped")is true so it stays red on retry.
Crash-reporter env
Separately, bunEnv does not set BUN_CRASH_REPORT_URL / BUN_ENABLE_CRASH_REPORTING (harness.ts:64-85). It spreads ...process.env, so on Buildkite it may inherit the runner's remap-server URL rather than being disabled — and locally it inherits nothing, so the crash handler is free to hit the default reporting host, which violates REVIEW.md's "Never contact external network hosts" hermeticity rule. The 30205 precedent explicitly sets both to "" / "0", and REVIEW.md says "Copy harness conventions exactly." The core-dump failure alone is sufficient to block; this is the second reason to copy the precedent env verbatim.
Relation to the earlier Windows comment
The prior 🔴 comment on this test says to add test.skipIf(isWindows) because is_panic_status() is POSIX-only. That is correct but does not fix this: with the Windows skip in place, the test still runs on Linux Buildkite and still leaks the core. Both fixes are needed. Conveniently, once skipIf(isWindows) is added the /bin/sh wrapper is safe to use unconditionally — exactly as 30205 does.
Fix
Match the precedent one-for-one:
await using proc = Bun.spawn({
cmd: ["/bin/sh", "-c", `ulimit -c 0 && exec "$@"`, "--", bunExe(), "test", "--parallel=2", "--bail=1"],
env: { ...bunEnv, BUN_TEST_PARALLEL_SCALE_MS: "0", BUN_CRASH_REPORT_URL: "", BUN_ENABLE_CRASH_REPORTING: "0" },
cwd: String(dir),
stderr: "pipe",
stdout: "pipe",
});(combined with test.skipIf(isWindows) from the earlier comment).
There was a problem hiding this comment.
The switch to crash_handler.segfault() fixes the core-dump half — that hook calls suppress_core_dumps_if_necessary() (crash_handler_jsc.rs:75) which zeroes RLIMIT_CORE before faulting, so the ulimit -c 0 wrapper is no longer needed. 👍
The two env keys are still missing though, and js_segfault() does not call suppress_reporting(). On CI runner.node.mjs sets BUN_CRASH_REPORT_URL=http://localhost:<remapPort> (line 1780), which bunEnv inherits via ...process.env — so the worker's deliberate crash still POSTs to the remap server. Per the comment at test/cli/run/run-crash-handler.test.ts:7-10, that makes the runner pin the crash on the next unrelated failing test as "crash reported" and blocks its retry. Add the env overrides:
env: { ...bunEnv, BUN_TEST_PARALLEL_SCALE_MS: "0", BUN_CRASH_REPORT_URL: "", BUN_ENABLE_CRASH_REPORTING: "0" },|
Updated 4:22 PM PT - Jul 28th, 2026
@Jarred-Sumner, your commit f4dc8d3 is building: |
process.kill(pid, "SIGSEGV") throws "Unknown signal" on Windows instead of crashing the worker, so there the file was an ordinary failing test and no panic ever occurred. Trigger a genuine segfault with the crash-handler test primitive, which every platform delivers as a fatal signal.
| 60_000, | ||
| Math.min(Math.max(10 * 60_000, bucketFiles.length * 5_000), jobBudgetMs() - 5 * 60_000), | ||
| ), | ||
| idleTimeout: parseInt(process.env.BUN_RUNNER_BATCH_IDLE_MS || "", 10) || 4 * 60_000, |
There was a problem hiding this comment.
🟡 On ASAN lanes perTestTimeout is Math.ceil(180_000/2) * 3 = 270s (line 957), which exceeds the hard-coded 4-minute idleTimeout here — so at the batch tail, when only one worker is left and no sibling dots are keeping the idle timer re-armed, a single test legitimately running 240–270s is SIGTERM'd as "stalled" even though it's within its per-test budget. Not a hard failure (the file lands in incomplete and passes on solo retry), just ~4 min of wasted CI and a false "stalled" classification. Consider idleTimeout: … || Math.max(4 * 60_000, perTestTimeout + 30_000) so the invariant idleTimeout > perTestTimeout holds regardless of the ASAN multiplier.
Extended reasoning...
The invariant that should hold, and why it does not on ASAN
The new idleTimeout at scripts/runner.node.mjs:994 defaults to 4 * 60_000 = 240_000 ms. The same batch spawn passes --timeout=${perTestTimeout} (line 983), where perTestTimeout = Math.ceil(testTimeout / 2) * (isAsan ? 3 : 1) (line 957) and testTimeout = 3 * 60_000 (line 85). On an ASAN lane that is Math.ceil(180_000 / 2) * 3 = 270_000 ms. So on ASAN, idleTimeout (240 s) < perTestTimeout (270 s) by 30 s. The idle timer is meant to detect a wedged batch, but on ASAN it can fire while a single test is still legitimately inside its allotted per-test window.
Why silence is expected at the batch tail
The batch runs with --dots (line 984). In --dots mode the coordinator only writes to stderr on TestDone (a dot per completed test) and flushes captured worker output at TestDone/FileDone; FileStart writes nothing (Coordinator.rs::on_frame). armIdleTimer is re-armed on every stdout/stderr chunk (lines 1556/1562), so as long as any worker is finishing tests, dots keep the idle timer alive. But once the batch drains to its last inflight file — every other worker has sent Shutdown and exited — the only source of output is that one worker's next TestDone. If its current test runs longer than 240 s, nothing is written and idleTimer fires.
Step-by-step on an ASAN lane
runParallelBucketspawnsbun test --parallel=N --timeout=270000 --dots …withidleTimeout: 240000.- The batch runs down until worker 0 is the only live worker, running the last file's last test — say a 250 s integration-ish case (allowed: 250 s < 270 s per-test timeout).
- All other workers have already emitted their final dots and exited. From this point the coordinator writes nothing until worker 0 sends
TestDone. - At t + 240 s since the last dot,
idleTimerfires →idledOut = true→expire()→subprocess.kill("SIGTERM"). - The coordinator's SIGTERM handler runs
abort_all(), printsInterrupted while still running: <file>, and exits.spawnSafereturns{ok: false, error: "stalled", …}. runParallelBucket's!okrecovery parses the interrupt report, puts the file inincomplete, and re-runs it solo — where the per-file timeout applies and the 250 s test passes.
At t + 270 s the worker's own --timeout would have killed the test, printed a (fail) … [timeout] line (a TestDone frame → a chunk → idle timer re-armed), and the batch would have carried on. The idle timer pre-empts that by 30 s.
Why nothing else prevents it
isBucketCandidate (lines 1131–1138) filters by directory allowlist and excludes /stress/i, but does not filter by expected test duration, and ASAN lanes do run runParallelBucket (line 956/966 handle isAsan inside it — there is no ASAN gate on the bucketable path). The BUN_RUNNER_BATCH_IDLE_MS override is not set anywhere in .buildkite/, so the 4-minute default applies.
Impact — why this is a nit, not blocking
No hard failure: the SIGTERM path names the hung file, the file is retried solo, and (since it was within its per-test timeout) it passes. The cost is up to ~4 min of wasted wall-clock plus a misleading "stalled" classification for a healthy test. The trigger is narrow — batch tail on an ASAN lane with a single test in the 240–270 s band — but ASAN is exactly where slow tests cluster. It is also worth noting the non-ASAN perTestTimeout is 90 s, comfortably below 240 s, so this only bites the ASAN multiplier.
Fix
Derive the idle window from the per-test timeout so the invariant holds regardless of the ASAN multiplier:
idleTimeout: parseInt(process.env.BUN_RUNNER_BATCH_IDLE_MS || "", 10) || Math.max(4 * 60_000, perTestTimeout + 30_000),(or apply the same isAsan ? 3 : 1 multiplier to the 4-minute constant).
Worker crashes are recognized as panics by fatal signal, which Windows never surfaces: there Bun's abort() is ExitProcess(3), indistinguishable from process.exit(3), so is_panic_status is documented as POSIX-only and Windows worker crashes deliberately take the per-file-failure path. The banner-and-kill contract this test asserts therefore only exists on POSIX; skip it on Windows rather than assert behavior the platform can't have.
Follow-up to #36175, from the first day of runs.
What went wrong
A darwin shard died at the Buildkite job timeout inside its batch (build 84092,
darwin-14-x64,019fa900): darwin buckets ~800 files into onebun test --parallelbatch, and the runner's batch cap wasmax(10 min, 5 s × files)≈ 68 min — longer than the 45-minute job. So when the batch wedged, the interrupt → name-the-hung-file → solo-retry net never fired; the job was simply killed with no report.Panic under
--bail=N: the coordinator decided whether the panic path or--bailhad setbailedby reading it afteraccount_crash— which can itself set it viabail_out(). When the Nth failure was a worker panic, the panic banner was skipped and in-flight siblings kept running.Fixes
spawnSafegains anidleTimeout, re-armed on every stdout/stderr chunk. A wedged batch is silent, so the batch now aborts gracefully after 4 minutes of no output, naming the hung file — a bound that doesn't grow with batch size.min(5 s × files, job remaining − 5 min)viaBUILDKITE_TIMEOUT, so it can never outlive the job either.--bailbailedbefore accounting; runabort_on_worker_panicon every panic (its own contract: "runs even if--bailalready setbailed").Distinct annotation title for the new case: a silence stop reads
stalled, a wall-clock onetimeout.Verified
Drove the runner over a bucket with one worker-only silent hang and a 10 s idle window: the guard fired at 13 s, the coordinator named the hung file, the seven finished files kept their results, and the hanger was retried alone — 22 s total instead of running to the cap.
parallel.test.tscrash/bail/junit cases pass.