Conversation
ShellSubprocess::on_process_exit ignored Status::Err, so Cmd::on_exit never ran and the Bun.$ promise never settled. The command now exits with 1 and reports the wait error on its stderr, like a failed spawn.
Each case spawns bun. On a loaded machine the debug build needs more than the 700 ms these cases allowed, so the file failed without a hang.
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository: oven-sh/bun/.coderabbit.yaml Review profile: ASSERTIVE Plan: Essentials Run ID: 📒 Files selected for processing (4)
Included review availability: Your plan provides up to 10 included reviews per hour; 0 remain after this review. WalkthroughThe shell now stores subprocess wait errors, reports them during command completion, and returns exit status 1. Shell tests cover quiet and loud execution, propagated ChangesShell wait-error handling
Suggested reviewers: Priority: ⬇️ Low 🚥 Pre-merge checks | ✅ 4✅ Passed checks (4 passed)
Comment |
|
Status: the fix and a test are in this PR. How I reproduced it, on release 1.4.3-canary.1 (Linux x64): The command never settles, so The new case in CI on bf4ee42 (build 118400) is final. The build is red because of one file outside this diff: |
There was a problem hiding this comment.
Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.
Beyond the inline findings, I also traced the new wait_error field from on_process_exit to its consumer: finish_if_done is the only producer of CmdState::Done for subprocess-backed Cmds and always yields Next(this_id), so the Done arm in Cmd::next runs on both the exit-first and stdio-closes-last orderings, and the take() plus WaitingWriteErr -> child_done(.., 1) path means the error is reported exactly once. The Cmd::new constructor is the only initializer site and was updated.
Extended reasoning...
Findings were reported inline (the waiter-thread test race and the sibling unnamed-signal arm in on_process_exit), so this body only records what else was examined. I read src/runtime/shell/states/Cmd.rs around on_exit/finish_if_done (lines 1022-1048) and the Done arm of next (lines 299-320): finish_if_done sets Done only when the exit code is present and every piped stdio has closed, then yields into the state machine, so the new wait_error check is reached regardless of whether the wait failure or the last pipe close arrives first. cmd_write_failing_error transitions to WaitingWriteErr and its completion calls child_done(parent, this, 1) (line 344), matching the existing spawn-error path. The only constructor site (Cmd::new, line 222) initializes the field. Nothing here is a guarantee of correctness; the inline findings still warrant a human look.
Without pidfd_open, bun on Linux waits for children on a second thread. That thread reaps the command before the waitpid(-1) in the test, so the wait never fails. The child probes pidfd_open and skips when it is not available, and the test drops BUN_FEATURE_FLAG_FORCE_WAITER_THREAD.
The list was already incomplete (it left out Cmd::child_done) and goes stale with each new call site.
There was a problem hiding this comment.
Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.
Beyond the inline finding, I also checked the new Done branch in src/runtime/shell/states/Cmd.rs for a re-entry loop or double finish: wait_error is take()n before cmd_write_failing_error, the Fd path's completion (on_io_writer_chunk) goes straight to child_done(1) without re-entering Done, and the Pipe (.quiet()) path appends to the captured stderr and finishes synchronously — so each failed wait reports exactly once.
Extended reasoning...
The Status::Err arm in subproc.rs now yields exit code 1 and stores the error on the Cmd using the same handle.cmd_mut() write the existing interrupted flag already uses, before on_exit, so no new pointer-lifetime shape is introduced. In Cmd.rs, the Done arm consumes wait_error once via take(); Builtin::cmd_write_failing_error either enqueues to the IOWriter and moves to WaitingWriteErr (whose completion callback calls child_done(parent, this, 1) directly) or, for captured stderr, appends and calls child_done synchronously. Neither path re-enters the Done match, so there is no loop or duplicate report. The argv0 slice uses saturating_sub(1) to strip the NUL and falls back to <unknown> for empty args. The remaining concern is the test's silent early return on hosts without pidfd_open, which is already posted inline.
|
Updated 12:20 PM PT - Sep 19th, 2026
❌ @robobun, your commit bf4ee42 has 1 failures in
🧪 To try this PR locally: bunx bun-pr 43469That installs a local version of the PR into your bun-43469 --bun |
The case returned early when the child found no pidfd_open, which the runner counts as a pass. The test process now probes pidfd_open once and gates the case with test.skipIf, so the runner reports a skip.
When the case times out its child bun never exits. The runner kills the child of a timed-out serial test and leaves the child of a concurrent one.
Problem
await Bun.$cmd`` never settles when the wait for the command's process fails. bun stays alive and prints no error.ShellSubprocess::on_process_exit(src/runtime/shell/subproc.rs:944). ItsStatus::Errarm is an empty// TODO: handle errorblock, so the function returns beforeCmd::on_exitand theCmdnever finishes.Fix
on_process_exitstores the wait error on theCmdand callsCmd::on_exit(1). TheCmdstill waits for its piped stdio to close, so the output is kept. Then it writesbun: failed to wait for <argv0>: <error>to its stderr and exits with 1.Status::Erris final:Process::on_exitremoves the exit handler. The exit status is lost, so the PR does not guess one.test/js/bun/shell/shell-hang.test.ts(new POSIX case, release 1.4.3-canary.1 times out). Alsobunshell,execandthrow.Signaledarm of the same function, which spawn: name exit signals and send named signals with the platform's numbers #39970 fixes.Background
Status(src/spawn/process.rs) is the result of a wait:Exited,Signaled, orErr. It isErrwhenwait4fails (ECHILD here).waitpid(-1)first, or SIGCHLD isSIG_IGNand the kernel reaped the child.Cmd(src/runtime/shell/states/Cmd.rs) is the shell state for one command. It finishes when it has an exit code and every piped stdio is closed.Notes
Repro on release 1.4.3-canary.1.
repro.jsisconst r = await Bun.$/bin/echo hi.nothrow().quiet(); console.log(r.exitCode, r.stderr.toString()).With this change:
The same start with
bun run --shell=bun <script>hangs on the release build. With this change it prints the message, and the script exits with 1.What this PR does not do. Started with SIGCHLD ignored, bash, dash and Node all report the real exit status (I ran
sh -c 'exit 3'under each: 3 every time). Each replaces the inherited disposition. bun keeps it, so the kernel discards the status. This PR turns the hang into a reported failure. #38022 resets an inheritedSIG_IGNbefore the first spawn, which keeps the real status in that setup. The two changes do not depend on each other. After #38022,Status::Errstill arrives when other code reaps the child, when the poll for the process cannot be registered again (Process::on_wait_pid), and on Windows when libuv reports a negative exit status (on_exit_uv). #37293 makes the second of these rarer on Linux. It does not change what the shell does when the status still arrives.Self-review. A review of the first version of this diff ran, and the run was cut off several times. I read its verdict (keep the PR) and its arguments for and against, but not its final list of concerns, so there may be points I have not seen. From what I read, I changed these: the body claimed that every
Statusnow reachesCmd::on_exit, which is false while the unnamed-signal arm remains. The first test ignored SIGCHLD, which pinned a result that bash, dash and Node do not give and depended on #38022 resetting the disposition only once. It was also Linux only with no reason given. I kept the message wording: it prints the resolved path and the errno text, like the otherbun:lines inCmd.rs.Sibling site left alone on purpose. In the same function, a
Signaledstatus whose signal has no name (for examplekill -40 $$) also returns beforeCmd::on_exit. #39970 fixes that arm and adds its own case to the same test file. This PR does not touch it, so the second of the two PRs to land needs a small rebase (adjacent lines insubproc.rs, the import line and the end ofshell-hang.test.ts).Why exit 1 and a message, not a rejected promise. The closest case in the shell is a spawn that fails:
Cmd::transition_to_execreports it throughBuiltin::cmd_write_failing_error, which writes to the stderr of the command and finishes it with 1. The wait error uses the same function.||,;, pipelines and.nothrow()then behave as they do for every other failed command, and the captured stdout is kept. Without.nothrow()the promise rejects with aShellError(exit code 1, the message instderr).Why the message is written in
Done.cmd_write_failing_errorfinishes theCmdwhen the write completes. It does not wait for the subprocess pipes.Cmd::finish_if_donesetsDoneonly after the exit code is in and every piped stdio is closed, soDoneis the first point where it is safe to finish through that function. The message therefore comes after the stderr of the command itself.The test. A child bun starts
sh -c "echo out; echo err >&2; exit 3"withBun.$, then calls a blockingwaitpid(-1)throughbun:ffibefore it returns to the event loop. That call always reapsshfirst (it reads exit status 3), so the wait in bun fails with ECHILD. The child does this with.quiet()(stderr is captured only) and without it (stderr also goes through theIOWriterto the real stderr, which is theWaitingWriteErrstate). Both must give exit code 1,out\non stdout, anderr\nplus the message on stderr. The test does not depend on the SIGCHLD disposition, so #38022 does not change it. It does depend on bun waiting on the JS thread. Withoutpidfd_open(a seccomp profile can block it), bun on Linux sets its waiter-thread flag itself (src/spawn_sys/spawn_process.rs:544) and waits on a second thread, which reapsshfirst. With that mode forced, the first version of this test failed in 30 of 30 runs (reaped: false, exit code 3). So on Linux the test process callspidfd_openonce throughbun:ffi, and the case usestest.skipIfwhen it fails. The runner then reports a skip, not a pass. The test also removesBUN_FEATURE_FLAG_FORCE_WAITER_THREADfrom the child env. I checked this on a host withoutpidfd_open: a seccomp filter that returns EPERM for syscall 434, installed before exec (x86_64 only, EPERM only). With the filter the file reports 6 pass, 1 skip, 0 fail. With the filter and no force flag, the first version of the test body givesreaped: falseand exit code 3 in 10 of 10 runs, so bun does fall back by itself. Without the filter the case runs and passes. The case is serial on purpose, unlike the other cases in the file. On a regression its child bun never exits. On the release build, the runner kills that child when the timed-out test is serial (killed 1 dangling process, no child left) and leaves it alive when the test is concurrent. On the release build the command never settles and the test times out. I ran it on Linux only. In CI build 118353 every Linux test lane passed it (glibc, musl, x64, aarch64, ASAN). On macOS both lanes passed in CI build 118400, which is the head (bf4ee422cd): aarch64 and x64, 2 of 2 test shards each. The case is not skipped on macOS, so it ran and passed there. macOS aarch64 also passed in build 118385. macOS x64 finished only in build 118400, so that is one run. macOS never uses the waiter thread.Builtin.rs. The only change there removes a sentence from thecmd_write_failing_errordoc that listed its callers. The list was already incomplete before this PR (it left outCmd::child_done).The 700 ms timeouts. The six existing cases in
shell-hang.test.tseach spawn bun and had a 700 ms timeout. On a loaded machine the debug build needs 0.5 to 1.9 s per case, so the file failed with no hang. The second commit removes the explicit timeouts. The default timeout still catches a hang.Suites run on the debug (ASAN) build.
shell-hang,bunshell,exec,throw,lazy,yield,shell-pipe-read-fault,shell-write-fault. One of fivebunshell.test.tsruns had a single failure that did not repeat. I did not capture which test it was.Also run on the debug (ASAN) build. A script with the wait error in a pipeline,
||,&&,;, a subshell, anifcondition,< ${buffer},> ${buffer},> file,.text(), and a command that throwsShellError. Every case completes with the command output kept. Twenty iterations withdetect_leaks=1report no leak. Forty failing waits leave the open fd count unchanged (9 before, 9 after).cargo clippy -p bun_runtimereports nothing for the changed files.no test proof · iteration 0 · platform-specific test(s) that do not run on this machine, deferring to CI, which covers all platforms: test/js/bun/shell/shell-hang.test.ts