crash_handler(windows): report a JSC initialization that ran out of memory as an error, not a crash - #39985
Conversation
…emory as an error, not a crash JSC::initialize() commits 10.5 MiB of mimalloc arena metadata for the Structure heap. When the commit limit of the machine or the memory limit of a job object is used up, that commit fails and JSC aborts. With the abort() now reported, that would be a crash report for the machine's condition. bun_jsc::initialize marks the call with Action::InitializeJsc. A crash under that action in a process that cannot commit 16 MiB anymore prints an error that names the commit charge and the job object limits, and exits with code 1.
|
Status: ready for review. Stacked on #38860, which this depends on: without it an Reproduced on Windows Server 2019 x64 with the current canary ( CI build 102724 finished with 177 of 179 jobs green. The two red ones come from the base branch, not from this diff. The |
There was a problem hiding this comment.
I reviewed this PR and didn't find any bugs. Because it changes the crash handler's early path and encodes a policy decision (JSC-init OOM → error+exit(1) instead of crash report) with a heuristic 16 MiB threshold, a human look would still be worthwhile.
What was reviewed:
- Verified
scoped_action/current_actionmechanics and that the new check is a no-op on non-Windows. - Checked Win32 constants (
MEM_RESERVE,MEM_RELEASE,JOB_OBJECT_LIMIT_*) andMEMORYSTATUSEXlayout against the headers. - Cross-checked the test's
JOBOBJECT_EXTENDED_LIMIT_INFORMATIONoffsets (16/112/128, size 144) against the Rust struct — they match. - Confirmed
bun_crash_handleris already abun_jscdep andGlobal::exit(1)matches sibling exit sites.
Extended reasoning...
Overview
This PR makes the Windows crash handler recognize when a crash during JSC::initialize() was caused by commit-charge exhaustion (machine or job-object memory limit) and turn that into a user-facing error + exit(1) instead of a crash report. It adds an Action::InitializeJsc variant, a scoped-action guard around JSCInitialize, an early check in crash_handler() that probes whether 16 MiB can still be committed, and three Win32 helpers (can_commit, commit_charge, job_memory_limits) in bun_sys::windows. Two Windows-only tests exercise both the OOM path (via a job-object memory-limit walk driven through cmd.exe + FFI) and the non-OOM path (a RELEASE_ASSERT triggered by BUN_JSC_structureHeapSizeInKB=3072).
Security risks
None identified. The new code runs only after a crash has already occurred, calls read-only Win32 introspection APIs (GlobalMemoryStatusEx, QueryInformationJobObject, K32GetProcessMemoryInfo) plus a VirtualAlloc/VirtualFree probe on a fresh region, and prints diagnostic values to stderr. No untrusted input is parsed.
Level of scrutiny
Medium-high. The crash handler is safety-critical infrastructure: the new branch runs before PANIC_STAGE is bumped, so a bug there could mask real crashes. The change is narrow (gated on Action::InitializeJsc, no-op on non-Windows) and the Win32 wrappers are straightforward, but the 16 MiB threshold is a heuristic that could drift with JSC changes, and the choice to treat this as an error rather than a crash report is a UX/policy call. The test harness (calibrated job-object memory walk via cmd.exe intermediary and hand-laid struct offsets) is unusual enough that a maintainer familiar with the Windows CI lanes should sign off on its stability.
Other factors
The PR description is exceptionally thorough — it derives the 10.5 MiB commit from mimalloc's arena metadata, tabulates canary behavior at each limit, and documents fail-before verification against the base branch. I verified the FFI struct offsets in the test against the workspace's JOBOBJECT_EXTENDED_LIMIT_INFORMATION definition, checked that bun_crash_handler is already in bun_jsc's deps, and confirmed the ActionGuard drop semantics restore the previous action after JSCInitialize returns. No prior review comments exist on this PR.
|
On the two points of the review above: The 16 MiB probe is a lower bound, not a measurement that has to stay exact. The question it answers is "could JSC have started in this process at all", and JSC's initialization commits at least the 10.5 MiB of Structure heap metadata, so a process that cannot commit 16 MiB could not have finished it. If JSC ever needs more, the probe still fails whenever initialization failed for memory, so the error is still right. If mimalloc's metadata ever shrinks, the cost is that a real bug during initialization on a machine with less than 16 MiB of commit left gets blamed on memory, and such a process would have died of memory a moment later anyway. The tests pin the current numbers: the second one only passes if the error is produced by a limit inside the span that mimalloc's commit creates. On the stability of the second test: the limits under which the error must appear span 10.5 MiB (17.4 to 27.9 MiB on the current release canary), the walk starts from a measured value ( |
|
Updated 10:23 AM PT - Aug 21st, 2026
❌ @robobun, your commit a6f490d has 2 failures in
Add 🧪 To try this PR locally: bunx bun-pr 39985That installs a local version of the PR into your bun-39985 --bun |
||||||||||||||||||||||||||||||||||||||||||||||||||||||
Stacked on #38860 (its branch is the base, so the diff is this commit only). Windows counterpart of #39967.
Problem
JSC::StructureMemoryManager::StructureMemoryManager(). ItsRELEASE_ASSERTs areabort(), which until crash_handler(windows): report abort() and int3 crashes, exit with the crash's own status #38860 died silently with 0xC0000409, so they are not these events (Notes).StructureAlignedMemoryAllocator.cpp:147does fire: mimalloc commits 10.5 MiB of metadata for the Structure heap, which fails once the commit limit of the machine or of a job object is used up (canary under a 20 to 24 MB job limit). With crash_handler(windows): report abort() and int3 crashes, exit with the crash's own status #38860 that is a crash report for a condition of the machine.Fix
bun_jsc::initializesetsAction::InitializeJscaroundJSCInitialize(the lines crash_handler: report an exhausted address space during JSC initialization as an error, not a crash #39967 adds too, the second PR to land drops them).test/cli/run/run-crash-handler.test.ts. Both fail on crash_handler(windows): report abort() and int3 crashes, exit with the crash's own status #38860 alone. The file passes with this branch on Windows Server 2019. Alsocargo checkfor both Windows targets.Background
Structures at startup and, since 1.4, gives it to mimalloc as an arena, whose metadata mimalloc commits up front.CURRENT_ACTIONis the thread local printed as "Crashed while ...". The check runs first incrash_handler(), for every reason, so it covers the releaseabort()and the trap of a debug build.Notes
Why the Sentry events are not these asserts: on Windows bun classifies access violations, illegal instructions, misalignment and stack overflow, and
abort()raised none of them. "Segmentation fault at address X" is an access violation with X as the data address (0x7FF7B9289011, 0x1 and 0xFFFFFFFFFFFFFFFF in the three issues). Nothing in the constructor dereferences such addresses. Sentry was not reachable from this session, so the events themselves were not examined. Two related observations: the "(pre-init)" command of the 1.4.0 events means abun build --compileexecutable (boot_standalonenever sets the command character), so these come from end user machines. And under a 32 to 56 MB job limit the canary produces a real access violation report (a null dereference inJSRunLoopTimer::Manager::scheduleTimerduringVM::VM, symbolicated with bun.report), so commit exhaustion does produce access violation reports in other frames.Canary
1.4.0-canary.1+4448a2e21,bun -e 1under a per process job memory limit (JOB_OBJECT_LIMIT_PROCESS_MEMORY):MIMALLOC_VERBOSE=1at 20 and 24 MB:cannot commit OS memory (error: 1455 (0x5AF), address: 0x01E300010000, size: 0xA80000 bytes), thenunable to commit meta-data for OS memory. 0xA80000 is 65536 slices times 168 bytes ofmi_page_t. The address is the 4 GiB aligned heap plus one 64 KiB slice. The peak commit of the canary at this point is 17.4 MiB (job objectPeakProcessMemoryUsed), so the limits that hit the constructor are 17.4 to 27.9 MiB, which matches the table.GlobalMemoryStatusExunder a 96 MB job limit still reports the machine's 75 GB, which is whyQueryInformationJobObjectis printed too. With a null handle it returns the innermost job, verified through a nested job.Output of this branch under a 30 MB job limit (debug build):
Tests. The second test measures the peak commit of
bun --version(12.2 MiB on the release canary, 18.3 MiB on the debug build) and walks the limit up from 8 MiB above it in 4 MiB steps, which is below the 10.5 MiB span of limits that hit JSC's initialization. JSC's initialization starts 5.2 MiB above that peak on the release canary (17.4 MiB) and about 8 MiB above it on the debug build, so the first or second step lands in the span. A step below the span dies of a Rust allocation failure, a__fastfail, which Windows Error Reporting holds for about 1.4 seconds: the test takes 2.5 seconds on the debug build (three runs) and should take one step on a release build. On #38860 alone, the debug build reports "Illegal instruction" (the trap a debug build dies of) at every limit from 30 to 90 MiB and the test fails after listing them. The release canary dies silently there. The first test usesstructureHeapSizeInKB=3072, which trips the assert at line 110 with memory to spare: on #38860 alone it is reported without the "Crashed while" line.Both tests are Windows only, so the fail-before check is deferred to the Windows CI lanes.
Also run:
cargo check -p bun_crash_handler -p bun_jsc -p bun_sysnatively and forx86_64-pc-windows-msvcandaarch64-pc-windows-msvc,cargo fmt --check, prettier,test/internal/source-lints/,test/cli/run/crash-report-command-char.test.ts, and the crash cases oftest/cli/test/parallel.test.tson Windows.