Skip to content

runner: kill the crash remap server on exit instead of beforeExit, which never fires - #38981

Closed
robobun wants to merge 1 commit into
mainfrom
farm/f1d23af5/runner-kill-remap-server-on-exit
Closed

robobun wants to merge 1 commit into
mainfrom
farm/f1d23af5/runner-kill-remap-server-on-exit

Conversation

@robobun

@robobun robobun commented Aug 15, 2026

Copy link
Copy Markdown
Collaborator

Problem

  • Every CI-mode run of scripts/runner.node.mjs (Linux and macOS lanes) leaves its crash-report remap server behind: a bun run --silent ci-remap-server ... process plus the bun node_modules/.bin/ci-remap-server ... script it wraps, both reparented to PID 1. On persistent agents these pile up one pair per job.
  • The orphan inherits the runner's stderr (stdio: ["ignore", "pipe", "inherit"]), so anything capturing the runner's output through a pipe keeps waiting after the runner has exited. Locally, node scripts/runner.node.mjs ... 2>&1 | cat was still blocked 5s after the runner exited; killing the orphan released it.
  • Cause: server.kill() was only called from a process.once("beforeExit", ...) handler (scripts/runner.node.mjs:809-815), and the runner never exits in a way that emits beforeExit. main() ends in process.exit() (:3431), so do --bail (:753) and the SIGINT/SIGTERM/SIGHUP handler (:3076), and an uncaught exception does not emit it either.
  • Even if beforeExit had fired, the handler called server.off("error") / server.off("exit") without a listener, which throws ERR_INVALID_ARG_TYPE before reaching server.kill().
  • The docker coordinator started a few lines earlier already used process.once("exit", () => coordinator.kill()) and was cleaned up on the same runs, which is what pointed at the event.

Fix

  • Add killOnExit(child) to scripts/utils.mjs: process.once("exit", () => child.kill()). Use it for the remap server and for the docker coordinator (same behavior as its inline hook).
  • Why exit is the right event: Node emits exit both from process.exit() and after an uncaught exception, which together are every way the runner ends; beforeExit is emitted only when the event loop drains, which never happens here. exit listeners must be synchronous, and kill() is.
  • Why kill()'s default SIGTERM, rather than SIGKILL, is enough for the two-process tree: on Linux bun run stays alive as the parent of the script and forwards SIGTERM to it, then exits itself; SIGKILL is not forwarded and would orphan the script. Checked against the real tree: SIGTERM to the wrapper took both processes down and released the pipe. On macOS bun run --silent execs the script in place, so there is only one process. The comment at the call site records this so nobody "hardens" it to SIGKILL.
  • The exiting flag and the off() calls existed only to serve the removed handler; the did not start branch no longer unregisters anything since a second kill() on a dead child is a no-op.
  • Test: test/internal/runner-kill-on-exit.test.ts. A fixture spawns a helper that inherits the fixture's stdout (as the remap server inherits the runner's stderr), hooks it with killOnExit, and leaves via process.exit(), an uncaught exception from await main(), or a SIGTERM handler that calls process.exit(); stdout reaching EOF is the proof the helper died. A control case without the hook checks the helper survives and keeps the pipe open, so the EOF signal is not vacuous. Runs under node (what the runner uses) and under bun.
  • bun bd test test/internal/runner-kill-on-exit.test.ts: 8 pass.
  • With scripts/ stashed: 8 fail (no export). With killOnExit temporarily hooked on beforeExit instead: the 6 positive cases time out waiting for the pipe, the 2 controls pass.
  • Real runner, before and after, each of the three exit paths (normal, uncaught exception, SIGTERM): before, 2 ci-remap-server processes left behind every time; after, 0, and a reader on the runner's output got EOF 4ms after the runner exited (log below).

Background

  • The remap server: in CI the runner starts ci-remap-server (from the bun-tracestrings package) once per shard and points every test process at it with BUN_CRASH_REPORT_URL, so crash trace strings from tests get symbolicated and attached to the failure output. It is meant to live exactly as long as the runner.
  • beforeExit vs exit in Node: beforeExit fires when the event loop has no more work and the process would exit on its own; it is not emitted for process.exit() or for a fatal error. exit is emitted in both of those cases as well, and its listeners run synchronously right before the process ends.
  • bun run <bin> on Linux spawns the bin as a child and waits for it, forwarding catchable signals it receives (SIGTERM, SIGINT, SIGHUP, ...) to that child; with --silent on macOS it execs the bin in place instead.
Runner probes (this container, GITHUB_ACTIONS=true, one test file)

Before, normal exit:

runner exited with 0 at 1786786303.289
reader: still blocked 5s after the runner exited
   6918       1  /workspace/bun/build/release/bun run --silent ci-remap-server /workspace/bun/build/release/bun /workspace/bun c1ae5ca1...
   6920    6918  bun /workspace/bun/node_modules/.bin/ci-remap-server /workspace/bun/build/release/bun /workspace/bun c1ae5ca1...
(SIGTERM sent to 6918)
reader: got EOF
no ci-remap-server processes left

Before, SIGTERM to the runner:

::group::Received SIGTERM, exiting...
remap processes after exit: 2
coordinator processes after exit: 0

After, normal exit:

runner exited with 0 at 1786786902.4616
reader got EOF at 1786786902.4655

After, uncaught exception (GITHUB_REPOSITORY unset makes the runner throw after starting the server) and SIGTERM:

server started: 1
Error: Environment variable is missing: GITHUB_REPOSITORY
remap processes after exit: 0

::group::Received SIGTERM, exiting...
remap processes after exit: 0

The runner always leaves through process.exit() (main(), --bail, the
SIGINT/SIGTERM/SIGHUP handlers), or through an uncaught exception. Neither
emits beforeExit, which was the only place the ci-remap-server child was
killed, so every CI-mode run left the `bun run --silent ci-remap-server`
wrapper and its script alive under PID 1, holding the runner's inherited
stderr open. The beforeExit handler also called server.off() without a
listener, which throws before reaching kill().

Add killOnExit() to scripts/utils.mjs, hooked on the exit event (which
process.exit() and uncaught exceptions both emit), and use it for the remap
server and the docker coordinator, which already used that event inline.
@coderabbitai

coderabbitai Bot commented Aug 15, 2026

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

@robobun, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 15 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 08d1eb8e-aae2-426a-a8b9-a69b09f10abc

📥 Commits

Reviewing files that changed from the base of the PR and between 732491c and 88f5eb1.

📒 Files selected for processing (3)
  • scripts/runner.node.mjs
  • scripts/utils.mjs
  • test/internal/runner-kill-on-exit.test.ts

Comment @coderabbitai help to get the list of available commands.

@robobun

robobun commented Aug 15, 2026

Copy link
Copy Markdown
Collaborator Author

Status: ready for review.

Reproduced on main with the real runner in CI mode (GITHUB_ACTIONS=true node scripts/runner.node.mjs --exec-path=... --quiet --retries=0 --vendor=false internal/linear-fifo.test.ts): after each run, on all three exit paths (normal process.exit, uncaught exception, SIGTERM handler), two ci-remap-server processes were left under PID 1 and a pipe reading the runner's output was still open 5s after the runner had exited. With this branch the same runs leave nothing behind and the pipe closes a few ms after exit.

Fix and test are described in the PR body; the regression test is test/internal/runner-kill-on-exit.test.ts.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM — straightforward fix moving the remap-server cleanup hook from beforeExit (never fires here) to exit.

What was reviewed:

  • Verified removing the exiting guard is safe: during process exit no further event-loop ticks run, so the child's exit handler cannot fire; in the did not start branch errorResolve was already called even before this change (guard was still false there).
  • Checked that a second kill() on the already-dead child (from the exit hook after the timeout branch killed it) is a no-op in Node.
  • Test design: control case proves stdout-EOF is a real signal, BUN_FEATURE_FLAG_NO_ORPHANS is unset so bun's own reaping doesn't mask a missing hook, leaked helpers are SIGKILLed in afterAll, Windows is skipped (matches the isCI && !isWindows gate at the call site).
Extended reasoning...

Overview

This PR fixes an orphaned-process leak in the CI test runner (scripts/runner.node.mjs). The crash-report remap server was hooked on process.once("beforeExit", ...), but the runner always leaves via process.exit() (main path, --bail, signal handlers) or an uncaught exception — none of which emit beforeExit. The fix adds a tiny killOnExit(child) helper to scripts/utils.mjs that hooks process.once("exit", ...) instead, applies it to both the remap server and the docker coordinator (which was already using an inline exit hook), and removes the now-dead exiting flag and broken server.off("error") calls (which would have thrown ERR_INVALID_ARG_TYPE anyway). A new test at test/internal/runner-kill-on-exit.test.ts covers all three exit paths under both node and bun, with a control case proving the assertion is not vacuous.

Security risks

None. This touches CI orchestration scripts only — no runtime code, no user-facing surface, no auth/crypto/permissions.

Level of scrutiny

Low-to-medium. This is internal CI tooling (scripts/), not shipped runtime code, and the change is small (~15 net lines of production diff). The diagnosis is unusually thorough — the PR description traces every exit path in the runner to a process.exit() call, explains why beforeExit never fires, documents why SIGTERM (not SIGKILL) is correct for the bun run wrapper tree on Linux vs macOS, and includes before/after probes on the real runner showing 2→0 leaked processes and pipe-EOF timing. No CODEOWNERS entry covers scripts/.

Other factors

  • The removed exiting guard was only ever set inside onBeforeExit, so removing it doesn't change behavior on the surviving code path (the did not start branch already called errorResolve post-kill before this change, harmlessly).
  • The test follows repo conventions well: tempDir/bunEnv/bunExe/nodeExe from harness, test.concurrent for independent subprocess cases, describe.skipIf(isWindows || !exe) matching the !isWindows gate at the call site, afterAll cleanup of any helpers a failing test would leak, and explicit unset of BUN_FEATURE_FLAG_NO_ORPHANS so bun's own orphan-killing doesn't hide a missing hook.
  • The docker coordinator change is a pure refactor to the shared helper — identical semantics to its previous inline process.once("exit", () => coordinator.kill()).
  • Placement in test/internal/ matches the existing pattern for tests of build/CI tooling (alongside parallel-allowlist.test.ts, build-*.test.ts, etc.).

Jarred-Sumner pushed a commit that referenced this pull request Aug 17, 2026
…ls never download from GitHub (#39446)

### Problem
- The "Lint JavaScript" (lint.yml) and "Format" (format.yml) checks go
red on PRs that did not touch anything they check. Their `bun install`
step fails with:
  ```
error: failed to download
bun-tracestrings@github:oven-sh/bun.report#912ca63: HTTP 5xx
  Failed to install 1 package
  ```
Seen on #29642 at 5309742 (a C++ comment change), both attempts of
runs 32040959983 and 32040960113, while other PRs' Lint runs flapped
red/green in the same minutes.
- Cause: `package.json:13` pins `bun-tracestrings` as a `github:`
dependency, so every root `bun install` downloads a tarball from GitHub.
That install runs on a fresh runner (no cache) in lint.yml, format.yml,
rust-lints.yml (4 jobs), bun-types.yml and packages-ci.yml, and in every
build (`scripts/build/codegen.ts` `bun_install`). bun retries a tarball
5 times back to back, which does not cover an outage of a few minutes.
- The only user of the package is `scripts/runner.node.mjs`, which runs
its `ci-remap-server` bin on Buildkite test shards
(`runner.node.mjs:803` on main). Nothing in lint, format, types,
rust-lints or the build uses it. Removing it outright was tried in
#25425 and closed for that reason.

### Fix
- Move the dependency to a new `scripts/ci-remap-server/package.json` (+
`bun.lock`) and drop it from the root `package.json` / `bun.lock` (a
pure removal: 94 lockfile entries, the package and its transitive
closure; the lockfile stays `lockfileVersion: 1`).
- `runner.node.mjs` installs that directory right before starting the
server, through the same `spawnBunInstall` as root and test/ so it uses
the agent's baked install cache, and runs the bin from there. The
install is best-effort like the server start already is: on failure it
warns and the tests run without crash remapping instead of failing the
shard.
- `bootstrap.sh` warms the new directory into the image's install cache
next to root and test/. No image version bump needed: the new lockfile
carries over the exact resolutions the root lockfile had (checked entry
by entry), so the caches baked from the old root lockfile already
contain everything except `@types/bun@1.3.14` and `bun-types@1.3.14`,
which resolved to the workspace packages before and now come from npm.
- New source lint
`test/internal/source-lints/lockfile-registry-only.test.ts`: the root
and test/ lockfiles (the two that every PR's checks and every shard
install) may not contain `github:`, `git+` or tarball-URL resolutions.
`source-lints.yml` now also triggers on `bun.lock` / `test/bun.lock`.
- Why this is the right layer: the failing jobs had no use for the
package, and a retry or cache in lint.yml/format.yml would still leave
GitHub on the install path of rust-lints, bun-types, the build and every
dev's `bun install`, and would still fail during a multi-minute blip.
After this, root installs only talk to the npm registry; the one job
that needs GitHub gets it from a baked cache and degrades gracefully
when it is unavailable.
- Verified:
- `test/internal/source-lints/lockfile-registry-only.test.ts`: fails on
main's `bun.lock` (`bun-tracestrings ->
bun-tracestrings@github:oven-sh/bun.report#912ca63`), passes here, under
`bun bd test`, the system bun and bun 1.3.14 (the version the workflows
pin). Whole `test/internal/source-lints/` passes.
- Root `bun install --frozen-lockfile` (lint.yml's command) with bun
1.3.14 and the debug build: no changes. Fresh checkout with GitHub
unreachable (`GITHUB_API_URL=http://127.0.0.1:1`, empty cache): main
fails with `failed to download bun-tracestrings@github:...
ConnectionRefused`, this branch installs.
- `bun install --frozen-lockfile` in `scripts/ci-remap-server` with bun
1.3.14, canary and the debug build: no changes; `bun run --silent
ci-remap-server` from that directory prints a port and serves `/traces`.
- This PR's own Buildkite build (#100049): every non-Windows shard logs
`scripts/ci-remap-server/package.json` / `86 packages installed` in 0.4
to 1s (the baked cache, as predicted; the root install went from 102
packages to 21), followed by `crash reports parsed on port ...`. The
server came up on 17 of 40 debian+ubuntu x64 shards against 14 of 40 on
main's build #100080: the remaining shards hit the runner's pre-existing
5s startup timeout (`ci-remap server did not start: timeout`), which
this PR does not change and is worth a follow-up of its own.
- Locally, `CI=true node scripts/runner.node.mjs ...` with a cold cache
got a 429 from codeload.github.com for the tarball; the runner warned
`ci-remap server not installed (...), crash reports will not be
remapped`, ran the test and exited 0. Same with the directory removed
(`spawn error`).
- `bun lint`, prettier `--check` on the touched files, `sh -n
scripts/bootstrap.sh`, `bun bd` reconfigure after the root package.json
change.
- Overlaps with #38981 on neighbouring lines of the same block in
`runner.node.mjs` (it changes how the server is killed, this changes
where it is installed and started); either rebases trivially onto the
other.

### Background
- `bun-tracestrings` is the npm name of github.com/oven-sh/bun.report,
the service that turns the trace strings in bun's crash reports back
into stack traces. Its `ci-remap-server` bin is a local copy of that:
`runner.node.mjs` starts it, points every spawned bun at it via
`BUN_CRASH_REPORT_URL`, and prints the remapped traces of any test that
crashed. It is purely diagnostic output; tests still fail on their exit
code without it.
- A `github:` dependency has no registry tarball: bun downloads it from
GitHub on install (and caches it under the same name@resolution key as
registry packages).
- `bootstrap.sh` builds the CI agent images. It clones the repo and runs
`bun install` in root and test/ with `BUN_INSTALL_CACHE_DIR` set, so
test shards install from disk; a package only downloads when its
resolution is not in the baked cache. The `# Version:` header is only
bumped when the image itself must change, which this does not require.
- `test/internal/source-lints/` holds repo-invariant tests that run on a
bare checkout in source-lints.yml (no `bun install`), which is why the
lint lives there and why that workflow's path filter had to learn about
the lockfiles.
@robobun

robobun commented Aug 21, 2026

Copy link
Copy Markdown
Collaborator Author

#39851 replaces the remap server block that this PR patches. It starts the server through a helper in utils.mjs that hooks the exit event for the reason given here, sets BUN_FEATURE_FLAG_NO_ORPHANS=1 on the server for the case where the runner is killed, and ports the three exit path cases of this test. If #39851 lands first, what remains of this PR is the coordinator call site, which already has an exit hook.

@robobun

robobun commented Aug 21, 2026

Copy link
Copy Markdown
Collaborator Author

Closing in favor of #39851.

#39851 rewrites the remap server block of the runner. Its spawnBackgroundServer() in scripts/utils.mjs hooks exit for the reason described here, and it also sets BUN_FEATURE_FLAG_NO_ORPHANS=1 on the server for the case where the runner is killed. Its test, test/internal/spawn-background-server.test.ts, carries the cases of test/internal/runner-kill-on-exit.test.ts: process.exit(), an uncaught exception, a signal handler that calls process.exit(), and the control case, under node and under bun. I ran that test at a4d48b7: all 20 cases pass, and with the exit hook removed exactly those three exit path cases fail under both runtimes.

The only part of this PR that #39851 does not carry is the coordinator call site, which already has an exit hook on main. This branch also conflicts with main now. Nothing remains to land here.

@robobun robobun closed this Aug 21, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant