Skip to content

crash_handler(windows): report abort() and int3 crashes, exit with the crash's own status - #38860

Open
robobun wants to merge 4 commits into
mainfrom
farm/ccc9bd59/windows-abort-crash-report
Open

robobun wants to merge 4 commits into
mainfrom
farm/ccc9bd59/windows-abort-crash-report

Conversation

@robobun

@robobun robobun commented Aug 15, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

  • On Windows, every abort() inside bun.exe kills the process with exit code 0xC0000409 and prints nothing: no Bun has crashed banner, no bun.report trace string. WTF RELEASE_ASSERT / CRASH() compiles to std::abort() in release builds off Darwin (vendor/WebKit/Source/WTF/wtf/Assertions.h:353), so this covers every JSC and Bun C++ assertion, plus aborts in mimalloc, BoringSSL and libuv. POSIX has reported these since crash_handler: catch SIGABRT and SIGTRAP so native aborts and traps get reported #34771.
  • Repro on canary 1.4.0-canary.1+a5c86aec7, Windows Server 2019 x64: the process.env snippet from process.env: make a failed env build throw instead of aborting (Windows 0xC0000409 on Bun.$ / Bun.sql / process.env near the stack limit) #38821 exits 0xC0000409; stderr holds only JSC's own dataLog line. With the abort test hook calling a real abort() (below), stderr is empty.
  • Cause: UCRT abort() (ucrt/startup/abort.cpp) calls raise(SIGABRT) only if a CRT-level SIGABRT handler is installed, otherwise it __fastfails. A fast-fail is not an exception, so the vectored handler and unhandled-exception filter that init() installs in src/crash_handler/lib.rs never run. Bun never installed a CRT SIGABRT handler.
  • Two adjacent gaps found while fixing it:
    • x64 int3 (JSC JIT abortWithReason(), LLInt break, __debugbreak) raises EXCEPTION_BREAKPOINT, which classify_exception_windows did not know, so those died silently too (exit 0x80000003). The POSIX fix caught SIGABRT and SIGTRAP together; this is its Windows twin.
    • After printing a report, crash() ended every Windows crash in ExitProcess(3), which a parent cannot tell from process.exit(3). bun test --parallel classifies worker deaths by exit status (Coordinator.rs, is_fatal_windows_exit_code), so reporting an abort would have turned a run-aborting 0xC0000409 worker death (bun test --parallel: abort the run when a Windows worker dies with a fatal NTSTATUS #37129) into a per-file failure.

Fix

  • init() installs a CRT SIGABRT handler (signal(SIGABRT, handle_abort_windows)) that enters crash_handler(CrashReason::Abort, ..), the same entry point the POSIX SIGABRT handler uses. It is removed wherever the VEH is: reset_segfault_handler() after a report, and raise_ignoring_panic_handler_raw() (bun run re-raising a child's signal; UCRT raise(6) maps onto the same slot, so without that reset a forwarded SIGABRT would print a bogus Bun report). _set_abort_behavior(0, _WRITE_ABORT_MSG) stops the debug UCRT from showing its own abort() has been called message first; the release UCRT never sets that flag.
  • classify_exception_windows maps EXCEPTION_BREAKPOINT to CrashReason::Trap. It goes through the same out-of-image rule as every other code in the VEH (JIT-pool traps reach the handler through JSC's unwind-info route from crash_handler(windows): let foreign first-chance AVs reach SEH via JSC unwind info #35083), and an attached debugger consumes breakpoints before any handler runs. WTF's own VEH (SignalsWin.cpp) only recovers access violations, illegal instructions and FP exceptions, never breakpoints.
  • crash() exits with CrashReason::terminal_exit_code(), the Windows counterpart of terminal_signal(): the exception's own status for faults (0xC0000005, 0xC000001D, 0xC00000FD, 0x80000002, 0x80000003), STATUS_STACK_BUFFER_OVERRUN (0xC0000409, what an unreported abort or Rust abort exits with) for the abort/panic/OOM bucket. Parents therefore see the same status with or without a report; the coordinator's allowlist gains 0x80000002, the one classified code it lacked. The debug-build int3 before ExitProcess now only runs when a debugger is attached; unconditionally it terminated the process with 0x80000003 before the exit code was reached.
  • Why the hook is correct: the CRT signal slot is the only hook UCRT abort() has and it is consulted before the fast-fail (verified against the UCRT source in the Windows SDK and with a /MT probe, below); UCRT raise() resets the slot to SIG_DFL before calling the handler, so an abort during the report fast-fails instead of recursing (one-shot, like SA_RESETHAND); nothing else in bun uses the slot (process.on("SIGABRT") is libuv, process.abort() is _exit(134) on Windows, Rust aborts and the fastfail hook are bare __fastfails), and the tests pin all three staying unreported. bun.exe links the UCRT statically, so only abort() calls that resolve to bun.exe's CRT (JSC/WTF, Bun, vendored deps) are caught; an addon linked against ucrtbase.dll still fast-fails and still classifies as before.
  • arm64 Windows finding (probed on a Windows 11 arm64 machine): the kernel reports a brk with any immediate it does not define, including WTF's 0xbb08 that JSC emits everywhere, as STATUS_ILLEGAL_INSTRUCTION; only brk #0xF000 (__debugbreak) is a breakpoint. So JSC traps on arm64 were already reported as illegal instructions; the silent-trap gap was x64. Documented on CrashReason::Trap, and the Windows test expects per architecture.
  • Test hooks: crash_handler.abort() and .trap() now execute the real abort() / int3 / brk #0xbb08 on Windows instead of calling the handler directly (which is what hid this); raiseIgnoringPanicHandler() takes an optional signal.
  • Tests:
    • test/cli/run/run-crash-handler.test.ts: a describe.if(isWindows) matrix asserting banner text and exit status for abort, panic, outOfMemory (9, the low byte of 0xC0000409 as Bun.spawn reports it), segfault (5) and trap (3 on x64, 0x1D on arm64, with a real fault address); process.abort() (134) and fastfail() (9) print nothing; abort added to the upload loop; the raise ignoring panic handler test now covers SIGSEGV and SIGABRT on every platform.
    • test/cli/test/parallel.test.ts: the worker-segfault case runs on Windows too and asserts the coordinator sees exit code 0xC0000005 and aborts the run.
    • test/napi/node-napi-tests/test/common/index.js: the Windows abort exit-code list moves from 3 to 9 (verified: spawnSync of a panicking bun reports status: 9 on Windows).
  • Verification:
    • Windows Server 2019 x64 debug build with src/crash_handler/lib.rs and src/bun_core/Global.rs reverted: the abort test fails with Received: "", everything else passes. Full branch: run-crash-handler.test.ts 19 pass / 0 fail, parallel.test.ts crash cases 6 pass, crash-report-command-char.test.ts 3 pass.
    • Linux x64 bun bd (ASAN): run-crash-handler.test.ts 19 pass / 16 skip, parallel.test.ts crash cases 6 pass, 30205.test.ts 4 pass.
    • cargo check of bun_crash_handler, bun_core, bun_sys for x86_64- and aarch64-pc-windows-msvc clean; cargo fmt --check, prettier clean. Deliberately left out, as follow-up material: classifying IN_PAGE_ERROR / integer-divide / privileged-instruction exceptions (the SIGBUS/SIGFPE analogs).
    • The discriminating tests are Windows-only, so the automated fail-before check cannot observe them on Linux; the reverted-files run above is that proof, and the Windows CI lanes run them with the fix.

Background

  • CRT signals: the Microsoft C runtime emulates a few POSIX signals in user space. signal() stores a function pointer in a CRT-global slot; raise() resets the slot and calls it synchronously on the calling thread; SIG_DFL is _exit(3). Unrelated to libuv's uv_signal, which is what process.on(signal) uses on Windows.
  • __fastfail: an int 0x29 that makes the kernel terminate the process at once with 0xC0000409, skipping all user-mode exception dispatch (VEH, SEH, unhandled-exception filter). It is how UCRT abort() and Rust's std::process::abort() end a process on Windows.
  • VEH / UEF: AddVectoredExceptionHandler and SetUnhandledExceptionFilter, the hooks the crash handler already had. They only see exceptions, so faults were reported and aborts were not; EXCEPTION_BREAKPOINT is an exception they did see but did not classify.
  • Exit status on Windows: a process killed by an unhandled exception exits with that exception's NTSTATUS as its exit code; that is what bun test --parallel and bun's CI runner classify on, and what terminal_exit_code() reproduces. Bun's Bun.spawn / child_process currently expose only the low byte of it (hence 9, 5, 3 in the tests); the coordinator reads the full value.
UCRT abort(), probes, and fail-before output

ucrt/startup/abort.cpp (Windows SDK 10.0.26100):

__crt_signal_handler_t const sigabrt_action = __acrt_get_sigabrt_handler();
if (sigabrt_action != SIG_DFL)
    raise(SIGABRT);
if (__abort_behavior & _CALL_REPORTFAULT)      // release default
    __fastfail(FAST_FAIL_FATAL_APP_EXIT);
_exit(3);                                      // debug default (_WRITE_ABORT_MSG only)

ucrt/misc/signal.cpp raise(): SIGABRT and SIGABRT_COMPAT (6) share one action; SIG_DFL is _exit(3); the action is set back to SIG_DFL before the user handler is called.

x64 probe (clang-cl, abort() / raise(6) with and without a signal(SIGABRT, h) that _exit(77)s):

release /MT   abort(), no handler      -> exit 0xC0000409, nothing printed
release /MT   abort(), handler         -> handler ran (sig=22), exit 77
release /MT   raise(6), SIG_DFL        -> exit 3
release /MT   raise(6), handler        -> handler ran (sig=6), exit 77
debug   /MTd  abort(), no handler      -> exit 3, nothing printed
debug   /MTd  abort(), handler         -> handler ran (sig=22), exit 77

arm64 probe (Windows 11 arm64, VEH printing the exception code):

brk #0        -> 0xC000001D STATUS_ILLEGAL_INSTRUCTION
brk #1        -> 0xC000001D
brk #0xbb08   -> 0xC000001D   (WTF_FATAL_CRASH_CODE: JIT abortWithReason, LLInt break, WTF_FATAL_CRASH_INST)
brk #0xF000   -> 0x80000003 STATUS_BREAKPOINT   (__debugbreak)
brk #0xF001   -> 0xC0000420 STATUS_ASSERTION_FAILURE

Canary 1.4.0-canary.1+a5c86aec7 on Windows Server 2019 running the #38821 snippet:

exit: 0xC0000409
stdout: (empty)
stderr: Static hashtable initialiation for env did not produce a property.

Fail-before (this branch minus lib.rs / Global.rs), debug build, Windows Server 2019:

error: expect(received).toContain(expected)
Expected to contain: "panic(main thread): abort() called"
Received: ""
(fail) ... abort() produces a crash report

Full branch, same machine: run-crash-handler.test.ts 19 pass / 16 skip / 0 fail; parallel.test.ts prints a test worker process crashed with exit code 0xC0000005 and Aborting for a worker that went through the crash handler; spawnSync of a panicking child reports status: 9.


no test proof · iteration 0 · Platform-specific test(s) that do not run on this machine. Deferring to CI, which covers all platforms: test/cli/run/run-crash-handler.test.ts test/cli/test/parallel.test.ts test/regression/issue/30205.test.ts

UCRT abort() never raises an exception: it calls the CRT-level SIGABRT
handler if one is installed and otherwise __fastfail()s, so the VEH and
unhandled-exception filter never saw it and a WTF RELEASE_ASSERT (which is
std::abort() in release builds off Darwin), a mimalloc or BoringSSL abort,
or any plain C abort() exited 0xC0000409 with no crash report. Install a
CRT SIGABRT handler in init() that routes into crash_handler(Abort), and
reset it wherever the VEH is torn down (reset_segfault_handler and
raise_ignoring_panic_handler). Clear _WRITE_ABORT_MSG so the debug UCRT
does not put up its own abort() message first.

The abort test hook now calls the real abort() on Windows too, and
raiseIgnoringPanicHandler takes an optional signal so the SIGABRT reset
is covered.
@coderabbitai

coderabbitai Bot commented Aug 15, 2026 •

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

@robobun, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 2 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: d1611ff8-4da6-46a5-85ea-7aa237795fd3

📥 Commits

Reviewing files that changed from the base of the PR and between 39fb3c1 and a83d314.

📒 Files selected for processing (9)
  • src/bun_core/Global.rs
  • src/crash_handler/lib.rs
  • src/runtime/api/crash_handler_jsc.rs
  • src/runtime/cli/test/parallel/Coordinator.rs
  • src/sys/windows/mod.rs
  • test/cli/run/run-crash-handler.test.ts
  • test/cli/test/parallel.test.ts
  • test/napi/node-napi-tests/test/common/index.js
  • test/regression/issue/30205.test.ts

Comment @coderabbitai help to get the list of available commands.

@robobun

robobun commented Aug 15, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 4:41 AM PT - Aug 15th, 2026

@robobun, your commit a83d314 is building: #97989

Status: reproduced and fixed in this PR (#38860); self-review follow-ups folded in.

  • Reproduced how: on Windows Server 2019 x64, canary 1.4.0-canary.1+a5c86aec7 running the process.env RELEASE_ASSERT snippet from process.env: make a failed env build throw instead of aborting (Windows 0xC0000409 on Bun.$ / Bun.sql / process.env near the stack limit) #38821 exits 0xC0000409 with nothing from Bun on stderr. With this branch's test hook calling a real abort() but src/crash_handler/lib.rs and src/bun_core/Global.rs reverted, the new Windows abort test in test/cli/run/run-crash-handler.test.ts fails with Received: ""; with the full branch the file passes (19 pass on that machine, 19 pass / 16 skip on Linux x64 ASAN, where the Windows block is skipped).
  • Review follow-ups (d73ad69): EXCEPTION_BREAKPOINT is now classified as a trap (x64 int3 died silently the same way), and a reported crash exits with the status the unreported crash would have had instead of 3, so bun test --parallel keeps classifying a reported abort as fatal (the un-skipped Windows case in test/cli/test/parallel.test.ts covers it). Probing on Windows 11 arm64 showed the kernel reports JSC's brk #0xbb08 as an illegal instruction, so arm64 traps were already reported; the test expects per architecture.
  • The discriminating tests are Windows-only, so the automated fail-before check cannot observe them on Linux; the reverted-files run above is that proof, and the Windows CI lanes run them with the fix.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this PR and didn't find any bugs. Because it changes process-global crash-handler state on Windows (CRT SIGABRT slot, _set_abort_behavior) and touches bun run's signal-forwarding path, a human look would still be worthwhile.

What was reviewed:

  • The CRT SIGABRT install/teardown pairing: init() installs, reset_segfault_handler() and raise_ignoring_panic_handler_raw() both reset to SIG_DFL — the SIGABRT_COMPAT (6) aliasing means a forwarded child SIGABRT would otherwise re-enter the handler.
  • handle_abort_windows → crash_handler → Windows crash() ends via ExitProcess(3), so no re-entry into UCRT abort(); UCRT raise() resetting the slot before dispatch gives one-shot semantics.
  • js_abort now calls real libc::abort() on Windows (previously masked by direct handler invocation); js_raise_ignoring_panic_handler signal arg goes through signal_code_jsc::from_js → Linux numbering, matching what raise_ignoring_panic_handler_raw expects.
  • New tests pin process.abort() (exit 134) and fastfail (0xC0000409) as non-reporting; the reworked raise ignoring panic handler test now asserts stderr is clean rather than sleeping-then-checking.
Extended reasoning...

Overview

This PR installs a CRT-level SIGABRT handler on Windows so that abort() calls (WTF RELEASE_ASSERT/CRASH(), mimalloc, BoringSSL, etc.) route into Bun's crash reporter instead of silently dying with 0xC0000409. It touches: src/crash_handler/lib.rs (install signal(SIGABRT, handle_abort_windows) + _set_abort_behavior in init(), reset in reset_segfault_handler()), src/bun_core/Global.rs (reset the slot in raise_ignoring_panic_handler_raw so bun run's forwarded SIGABRT doesn't produce a bogus report), src/runtime/api/crash_handler_jsc.rs (js_abort now calls real abort() on Windows; raiseIgnoringPanicHandler accepts a signal arg), plus comment-only updates in Coordinator.rs/parallel.test.ts and new/reworked tests in run-crash-handler.test.ts.

Security risks

None identified. The crash reporter is diagnostic-only; the new hook does not process untrusted input and does not weaken any existing check. _set_abort_behavior(0, _WRITE_ABORT_MSG) only suppresses the debug-CRT dialog.

Level of scrutiny

High. The crash handler is the last-resort diagnostic path and this change installs a process-global CRT signal handler on Windows. It also intersects bun run's signal-forwarding: without the SIG_DFL reset in raise_ignoring_panic_handler_raw, a child's SIGABRT (POSIX number 6, aliased to the CRT SIGABRT slot as SIGABRT_COMPAT) would print a spurious Bun crash report. The author covered that, and I verified the teardown runs on the outer CRASH_HANDLER_INSTALLED gate (not the inner VEH-handle-non-null check). The primary new test is Windows-only, so Linux CI won't exercise it; the PR provides a manual fail-before/pass-after run on Windows Server 2019.

Other factors

  • Recursion: UCRT raise() resets the slot to SIG_DFL before invoking the handler, and crash_handler's Windows crash() path terminates via ExitProcess(3) (the module-local abort()), so an abort() during report writing fast-fails rather than recursing.
  • The reworked raise ignoring panic handler test is stronger than before (asserts clean stderr and exact exit code / signal instead of .not.toBe(0) after a 2s sleep), and now covers SIGABRT on both platforms.
  • js_raise_ignoring_panic_handler now uses signal_code_jsc::from_js, which returns the Linux-numbered bun_sys::SignalCode(u8) — consistent with the existing raise_ignoring_panic_handler_raw callers in run_command.rs/bunx_command.rs.
  • I did not find a case where process.on('SIGABRT') (libuv) would collide with the CRT slot; they are separate tables as the PR states.

Given the crash handler is critical infrastructure and the correctness hinges on UCRT internals I can only verify from documentation, deferring to a human reviewer.

…sh's own status

EXCEPTION_BREAKPOINT (x64 int3 from JSC JIT abortWithReason and LLInt
break, brk #0xF000 on arm64) was not classified, so those traps died
silently like abort() did. arm64 Windows reports JSC's brk #0xbb08 as an
illegal instruction, which was already classified.

crash() used to ExitProcess(3) for every reported crash, which a parent
cannot tell apart from process.exit(3); bun test --parallel therefore
stopped classifying a worker's abort() as fatal once the new SIGABRT hook
reported it. Exit with the status the unreported crash would have had
instead (the exception code, or STATUS_STACK_BUFFER_OVERRUN for the
abort/panic/OOM bucket, mirroring terminal_signal on POSIX), and only
execute the debug-build breakpoint when a debugger is attached so that
status is actually reached. The trap test hook now executes the real
instruction on Windows too.
Comment thread src/bun_core/Global.rs Outdated
Comment thread src/crash_handler/lib.rs Outdated
Comment thread src/crash_handler/lib.rs Outdated
Comment thread src/crash_handler/lib.rs Outdated
Comment thread src/crash_handler/lib.rs Outdated
Comment thread src/crash_handler/lib.rs Outdated
Comment thread src/crash_handler/lib.rs Outdated
Comment thread src/crash_handler/lib.rs Outdated
Comment thread src/crash_handler/lib.rs Outdated
Comment thread src/crash_handler/lib.rs Outdated
Comment thread src/crash_handler/lib.rs Outdated
Comment thread src/runtime/api/crash_handler_jsc.rs Outdated
Comment thread src/runtime/api/crash_handler_jsc.rs Outdated
Comment thread src/runtime/api/crash_handler_jsc.rs Outdated
Comment thread src/runtime/api/crash_handler_jsc.rs Outdated
Comment thread src/runtime/cli/test/parallel/Coordinator.rs Outdated
Comment thread src/runtime/cli/test/parallel/Coordinator.rs Outdated
Comment thread src/sys/windows/mod.rs Outdated
Comment thread src/crash_handler/lib.rs
Comment thread src/crash_handler/lib.rs
Comment thread src/crash_handler/lib.rs
Comment thread src/crash_handler/lib.rs
Comment thread src/crash_handler/lib.rs
Comment thread src/crash_handler/lib.rs
Comment thread src/runtime/api/crash_handler_jsc.rs
Comment thread src/runtime/api/crash_handler_jsc.rs
Comment thread src/runtime/cli/test/parallel/Coordinator.rs
Comment thread src/sys/windows/mod.rs
@robobun robobun changed the title crash_handler(windows): report abort() crashes instead of exiting 0xC0000409 silently crash_handler(windows): report abort() and int3 crashes, exit with the crash's own status Aug 15, 2026

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this PR and didn't find any bugs. Because it changes process-wide crash-handler installation and exit-code semantics on Windows (CRT SIGABRT hook, VEH breakpoint classification, terminal_exit_code), a human look would still be worthwhile.

Checked: the SIGABRT hook is torn down everywhere the VEH is (reset_segfault_handler, raise_ignoring_panic_handler_raw); terminate_windows gates the int3 on IsDebuggerPresent so it can't recurse into the newly-classified EXCEPTION_BREAKPOINT path; terminal_exit_code's statuses line up with is_fatal_windows_exit_code so a reported crash still classifies as a panic in the coordinator; the aarch64 brk #0xbb08 change is POSIX-neutral (still SIGTRAP).

Extended reasoning...

Overview

Windows crash-handler coverage: installs a CRT signal(SIGABRT, ...) hook so UCRT abort() (WTF RELEASE_ASSERT, mimalloc, BoringSSL) reaches the crash reporter instead of __fastfailing silently; classifies EXCEPTION_BREAKPOINT in the VEH so x64 int3 (JSC JIT abortWithReason, LLInt) is reported; replaces the fixed ExitProcess(3) after a report with terminal_exit_code() so parents see the fault's own NTSTATUS. Test hooks (crash_handler_jsc.rs) now execute the real abort()/int3/brk #0xbb08 on Windows instead of calling the handler directly. raiseIgnoringPanicHandler gains an optional signal argument. Coordinator's fatal-exit-code allowlist gains STATUS_DATATYPE_MISALIGNMENT; a comment-only update elsewhere. Tests: new Windows describe.if matrix in run-crash-handler.test.ts, un-skipped Windows case in parallel.test.ts, updated NAPI nodeProcessAborted exit-code list.

Security risks

None identified. This is diagnostic/termination-path code; no untrusted input is parsed. _set_abort_behavior and IsDebuggerPresent are precondition-free by-value calls. The new libc::signal(SIGABRT, ...) slot is process-global but scoped to bun.exe's statically-linked CRT (addon CRTs are unaffected, as documented and pinned by the fastfail test).

Level of scrutiny

High. This is process-wide crash-handling infrastructure: the SIGABRT hook, VEH classification, and exit-code contract all affect how every C++/Rust assertion, JSC trap, and vendored-library abort surfaces to users, CI, and bun test --parallel. The exit-code change (3 → NTSTATUS) is user-visible to any parent classifying by status. The discriminating tests are Windows-only and can't be exercised on this Linux review machine — Windows CI is the actual verification. The PR description is unusually thorough (UCRT source citations, x64/arm64 probes, fail-before output), but the interaction surface (VEH ↔ CRT signal ↔ JSC's own SEH ↔ debugger-attached) is the kind a maintainer familiar with the Windows crash-handler history should sign off on.

Other factors

  • The comment-cop bot flagged over-long comments; all were addressed in a83d314 and the threads are resolved.
  • Confirmed the SIGABRT reset in raise_ignoring_panic_handler_raw and reset_segfault_handler mirrors the existing VEH/UEF teardown, so bun run re-raising a child's SIGABRT won't loop into the reporter.
  • The debug-build int3 in terminate_windows is now guarded by IsDebuggerPresent(), avoiding re-entry into the newly-classified EXCEPTION_BREAKPOINT handler when no debugger is attached.
  • terminal_exit_code() maps every reason to a code already in is_fatal_windows_exit_code, so reported crashes still abort a --parallel run rather than being downgraded to per-file failures.
  • The aarch64 brk #0 → brk #0xbb08 change in js_trap is fine on POSIX (both deliver SIGTRAP) and matches the arm64-Windows probe result documented in the PR.

@robobun

robobun commented Aug 21, 2026

Copy link
Copy Markdown
Collaborator Author

#39985 is stacked on this branch. It adds the out of memory case of a reported abort() during JSC::initialize(), so it should land after this one.

@robobun

robobun commented Aug 21, 2026

Copy link
Copy Markdown
Collaborator Author

A data point from the CI build of #39985, which is built on this branch (https://buildkite.com/bun/bun/builds/102724): on the darwin aarch64 test lane, SIGABRT/SIGTRAP are caught by the crash handler > trap produces a crash report fails with an empty stderr. That is the brk #0xbb08 this branch puts into the trap hook on every aarch64 platform: macOS kills the process on that instruction before a SIGTRAP handler runs (#39967 and #39978 describe the same thing for WTF's own brk #0xbb08). main uses brk #0 there, which macOS does deliver as SIGTRAP. The build of this PR itself (#97989) had no darwin aarch64 test lane, so it did not show up here.

Unrelated to the test: the binary-size step of that build reports every target 1 to 3.7 MB over the current canary, but the sizes are identical to the ones in #97989, so that is only the age of the base of this branch.

@robobun

robobun commented Sep 17, 2026

Copy link
Copy Markdown
Collaborator Author

The docs work in #43039 found the third gap of this PR again. I reproduced it on canary 1.4.3-canary.1+630e921db, Windows Server 2019 x64.

The repro has five test files and runs bun test --parallel=2. One file calls a crash_handler helper from bun:internal-for-testing.

Helper Coordinator output on Windows Run on Windows
segfault(), panic(), abort(), trap(), outOfMemory() ✗ b.test.js (worker crashed: exit code 3) Continues: 4 pass, 1 fail
fastfail() a test worker process crashed with exit code 0xC0000409 while running b.test.js. Aborts: 1 pass, 4 fail

On Linux x64 (canary 1.4.3-canary.1+c6b7fcb5b) the run aborts for each of the first five helpers, with SIGSEGV, SIGABRT or SIGTRAP. The terminal_exit_code() change in this PR removes that difference, so there is no second PR for it.

Docs coupling: #43039 describes the current behavior in docs/test/parallel.mdx, section "Crashes and interrupts":

A crash that Bun reports with its Bun has crashed banner exits the worker with code 3 on Windows, so the coordinator counts that file as failed and the run continues.

The PR that merges second must update or remove that sentence.

This branch has merge conflicts with main at this time.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant