Repository navigation
arena: the idle pull-forward uses the CAS paths, and a two arena test for the purge schedule - #42
Merged
Jarred-Sumner merged 3 commits intoSep 14, 2026
Conversation
…ena test for the purge schedule `_mi_arenas_purge_now` pulled the expire of an arena forward with a load and a store, and did the same for `subproc->purge_expire`. A pass resets both with a CAS or before a walk, so a store that lands after that reset arms an arena that the pass has just taken care of, and a value of `subproc->purge_expire` that was read before the reset says nothing about what the pass leaves behind. In the worst case an arena is armed and the subprocess is not until the scavenger's 30 s look at the arenas. The arena expire is now moved with a CAS from the value that was read, and the subprocess expire goes through `mi_subproc_schedule_purge`, the CAS loop that the passes use. test-purge-behind-pass gets a third part for more than one arena. Both arenas get a batch, and one more block goes into the main arena while the pass is in the other one. A pass that only reads the expire of each arena as it leaves it and installs the earliest one at its end misses that free: 8565de2 with the one line fix for `all_visited` ends with `258 of 264 MiB`, 3 of 3 runs. The first part alone passes there once a second arena exists.
…ks on a 32-bit target A block of 8 MiB and a bit starts on a boundary in an arena: each one takes 32 MiB of it on a 32-bit target, where a 96 MiB arena held 3 of the 4 blocks and the part failed with 'out of memory' (Debug, Win32). The reservation is four times what goes in now, and an arena that still does not take all of the blocks ends the part with a message and does not fail it.
…e sanitizer and light runs The thread sanitizer job has a time limit for each test. With MI_TSAN or MI_TEST_LIGHT the first part runs two rounds and the third part works with a quarter of the blocks, which is less memory traffic in total than the test had before the third part. The smaller part still fails on the one line fix for all_visited (129 of 136 MiB, 2 of 3 runs).
Jarred-Sumner
added a commit
to oven-sh/bun
that referenced
this pull request
Sep 14, 2026
Pins 3aa0ac9bea44, the merge of oven-sh/mimalloc#42 on bun-dev3-v2, on top of #40 and #43. #42 hardens the purge pull-forward that runs when a thread goes idle (mi_on_thread_idle -> _mi_arenas_purge_now). It moved each armed arena's expire, and the sub-process expire that the purge thread sleeps on, to "now" with a load and a store. A purge pass resets both first thing, so a store based on a value read before that reset could undo it. The arena expire is now moved with a compare-and-swap, and the sub-process expire through the CAS loop that the passes use. Nothing was observed; it closes the window. 12 lines in src/arena.c, no header changes. Allocator micro benchmark: 2.2384 G instructions before and after.
Jarred-Sumner
added a commit
to oven-sh/bun
that referenced
this pull request
Sep 15, 2026
Pins 3aa0ac9bea44, the merge of oven-sh/mimalloc#42 on bun-dev3-v2, on top of #40 and #43. #42 hardens the purge pull-forward that runs when a thread goes idle (mi_on_thread_idle -> _mi_arenas_purge_now). It moved each armed arena's expire, and the sub-process expire that the purge thread sleeps on, to "now" with a load and a store. A purge pass resets both first thing, so a store based on a value read before that reset could undo it. The arena expire is now moved with a compare-and-swap, and the sub-process expire through the CAS loop that the passes use. Nothing was observed; it closes the window. 12 lines in src/arena.c, no header changes. Allocator micro benchmark: 2.2384 G instructions before and after.
Jarred-Sumner
added a commit
to oven-sh/bun
that referenced
this pull request
Sep 15, 2026
Pins 3aa0ac9bea44, the merge of oven-sh/mimalloc#42 on bun-dev3-v2, on top of #40 and #43. #42 hardens the purge pull-forward that runs when a thread goes idle (mi_on_thread_idle -> _mi_arenas_purge_now). It moved each armed arena's expire, and the sub-process expire that the purge thread sleeps on, to "now" with a load and a store. A purge pass resets both first thing, so a store based on a value read before that reset could undo it. The arena expire is now moved with a compare-and-swap, and the sub-process expire through the CAS loop that the passes use. Nothing was observed; it closes the window. 12 lines in src/arena.c, no header changes. Allocator micro benchmark: 2.2384 G instructions before and after.
Jarred-Sumner
added a commit
to oven-sh/bun
that referenced
this pull request
Sep 15, 2026
, #44, #45, #46) (#42668) ### What does this PR do? Moves `MIMALLOC_COMMIT` from `ab13501334a8` to `69050cf3ec14be4847ca3ba6c840da821fa283b0`, the head of `bun-dev3-v2`: main's pin plus six merged PRs of the fork and nothing else, and changes the two callers in bun that the last of them asks to change. `MI_FREE_USE_PAGEMAP=1` stays. - oven-sh/mimalloc#40: the sampled profiler hooks (`mi_profiler_t`) attribute a sample to the small allocation that refilled its page, keep `bytes_since_last_sample` right when `on_alloc` raises the rate, sample a new thread's heap from its first allocation, report the frees of `mi_heap_destroy`, and give `on_free` of an aligned block the pointer that `on_alloc` got. What the heap profiler in #42626 needs. - oven-sh/mimalloc#43: an aligned request only becomes a sample through the over-allocating path (with #40 alone `mi_malloc_aligned(64, 64)` could free a sampled block inside the allocator; `MI_DEBUG` asserted). - oven-sh/mimalloc#42: the purge pull-forward of an idle thread (`mi_on_thread_idle`) goes through compare-and-swap, so it cannot undo the reset that a purge pass does first thing. 12 lines in `src/arena.c`. - oven-sh/mimalloc#44: a thread that does not sample pays for the profiler hooks in the generic path only: +0.2% of instructions on the allocator micro benchmark where they cost +1.5%. - oven-sh/mimalloc#45: a debug assertion that raced with `mi_profiler_stop` is gone. - oven-sh/mimalloc#46: **free blocks in the 4 MiB pages (blocks of 96 to 512 KiB) are handed back to the OS.** Hole purging skipped these pages (their 1024 OS pages do not fit the 256-bit purge bitmap), so what a process freed in that range stayed resident for as long as one block of the page was alive: 97% of it by `mincore`. The bitmap's unit is now per page (16 KiB for a 4 MiB page), and the free blocks of a large page stay while the page is in use, which the allocator tells without a clock on the allocating side: the idle sweep counts epochs of 100 ms or more, an allocation from a large page leaves the current epoch in the page (a load, a compare, and a store once per epoch), and the blocks go when that is two epochs ago. Without that hold a `Bun.serve` with 256 KiB bodies discarded and refaulted 3 GB in 20 s. Nothing in bun attaches a profiler yet. The fast paths of `mi_malloc` and `mi_free` are untouched by all six. `include/mimalloc.h` gains one function, `mi_on_thread_idle_pending()`, and comments; `mimalloc-stats.h` is identical, `mimalloc-profile.h` differs in comments, `mimalloc/types.h` gains fields of the internal `mi_theap_t` and `mi_tld_t`. No struct or enum that bun binds changed (`src/mimalloc_sys/mimalloc.rs`, the declarations in `BunJSCEventLoop.cpp` and `epoll_kqueue.c`). #### The callers that sweep once and then block With oven-sh/mimalloc#46 ONE call of `mi_on_thread_idle()` never takes the free blocks of a large page that the thread allocated from since its previous sweep: a sweep ends one epoch at the most. The event loop is fine: `us_loop_run_bun_tick` hands the thread's heaps to mimalloc's scavenger thread across the poll (`mi_on_thread_idle_start()` / `_end()`), and the scavenger comes back for them by itself, two or three sweeps an interval apart. Two callers swept inline once and then blocked for good, and would have kept such blocks, as they do on main: - **`src/threading/ThreadPool.rs`**: a worker swept once, after its futex wait timed out at 100 ms, and then waited without a timeout. It now brackets that second wait with `mi_on_thread_idle_start()` / `mi_on_thread_idle_end()`, and sweeps inline (once, as before) only when `_start` says there is no scavenger to hand off to. Nothing runs between the two calls but `Futex::wait`, which allocates nothing; `drain_idle_events` runs before `_start`. Only the wait that follows a timed-out one is bracketed. Bracketing every park was measured and is slower (table below: `_start` is two read-modify-writes on cache lines that every parking thread shares, and it wakes the scavenger, which costs half a context switch per task), and sweeping on every park is what the comment there warns about (13% of `vite preview` requests per second). As it is, a worker that parks between tasks never reaches `_start`, and `notify()` is what it was. - **The libuv loop on Windows** (`Bun__JSC_onBeforeWait` in `BunJSCEventLoop.cpp`, from `us_loop_run` in `libuv.c`) sweeps inline at most every 100 ms and then `uv_run(UV_RUN_ONCE)` blocks. It cannot hand off: `uv_run` dispatches the completions right after its poll inside the same call (`uv__poll` at `src/win/core.c:732`, `uv__process_reqs` at `:737`), and libuv allocates from mimalloc there (`uv_replace_allocator`), so there is no point between "about to block" and "awake" that bun's code sees without a patch to libuv. It keeps the inline sweep, and while `mi_on_thread_idle_pending()` says blocks were left, `Bun__JSC_onBeforeWait` returns 100 (ms) and `us_loop_run` arms an unref'd timer for that long after the sweep, so the hook runs again: three times at the most after JS last ran, which is the bound the fork gives. The timer does not keep the loop alive. Not changed: when no scavenger runs (`MIMALLOC_SCAVENGER=0` or a purge delay of 0; bun sets neither) the fallbacks in `us_loop_run_bun_tick` and in the pool sweep inline once, as on main, and the free blocks of large pages stay. ### How did you verify your code works? Release builds of one tree (this branch on `04823d8fad67`), Linux x64: **base** = main's pin, **pin** = this pin with the callers unchanged (the first commit), **this PR** = both commits. **Test.** `test/js/bun/jsc/heapStats-mimalloc.test.ts` gains "an idle work pool thread hands back the free blocks of its large pages": a subprocess with 4 pool threads reads 480 files of 96 to 512 KiB with `fs.promises.readFile` (the buffers are allocated on the pool threads), keeps 1 in 12, collects the rest, and polls `RssAnon` while nothing gives the pool work. Anonymous memory beyond what it started with and what it keeps has to come down to 12 MB within 2.5 s. | | free and resident at the end | | --- | --- | | base | 54 to 67 MB: fails | | pin, callers unchanged | 38 to 47 MB: fails (also with `BUN_GARBAGE_COLLECTOR_LEVEL=1` and `2`, and with the system bun 1.4.1: 50 MB) | | this PR | -1 MB, reached 350 ms into the idle: passes, 10 of 10 runs | **The park and notify path of the pool** (what the `has_swept` comment protects). 100000 `await fs.promises.stat("/")` one after the other (every call notifies a worker, which runs one task and parks), the same in batches of 8, and `await randomBytes(16)`; `perf stat` of the whole process, 8 cores, 5 interleaved runs, medians (range): | user instructions per call | base | pin | this PR | | --- | --- | --- | --- | | `stat`, sequential | 8794 (8784..8828) | 8792 (8716..8827) | 8797 (8780..8816) | | `stat`, 8 at a time | 6027 (6014..6131) | 6025 (6004..6080) | 6034 (6008..6102) | | `randomBytes(16)`, sequential | 12608 (12513..12629) | 12581 (12522..12609) | 12589 (12582..12627) | Context switches per call are 3.96 to 3.97, 0.51 to 0.53 and 3.97 in all three, kernel instructions per call overlap (41.6 k..49.9 k sequential), wall time is within the spread of the machine. Bracketing EVERY park instead (a build with a switch for it, 200000 calls, 4 interleaved runs), against the same build with the handoff after a timed-out wait only: | | after a timed-out wait (this PR) | every park | | --- | --- | --- | | `stat` sequential: user / kernel instructions, context switches per call | 8565..8601 / 45.6 k..46.7 k / 3.95..3.99 | 9193..9335 (+7%) / 49.6 k..50.7 k (+9%) / 4.2..4.55 | | `stat` 8 at a time | 5635..5659 / 8.8 k..9.3 k / 0.53..0.55 | 5771..5960 (+3%) / 10.1 k..11.6 k (+13%) / 0.63..0.75 | | `randomBytes(16)` sequential | 12145..12283 / 43.2 k..45.5 k / 3.97..3.99 | 12815..12893 (+5%) / 48.2 k..51.2 k (+11%) / 4.41..4.5 | | wall time, `stat` sequential | 11.6 s..12.7 s | 12.4 s..13.5 s (+6 to 8%) | **Allocator micro benchmarks** (mimalloc alone with bun's flags: clang 21, `-O3 -march=nehalem`, C++, `MI_FREE_USE_PAGEMAP`; user-mode instructions, the runs of one build equal to within 100): | 25 M `mi_malloc` + 25 M `mi_free` of 16..1039 bytes | instructions | | --- | --- | | before the upstream sync (`707d90be`) | 2.0731 G | | main's pin (`ab135013`) | 2.0568 G | | with oven-sh/mimalloc#40, #43, #42 (`3aa0ac9b`) | 2.0556 G (-0.06%) | | with #44, #45 (`58460467`) | 2.0305 G (-1.28%) | | this pin (`69050cf3`) | **2.0305 G (-1.28%)**, the same as `58460467` to within 400 instructions | | this pin with `MI_PROFILE=0` (not what bun builds) | 2.0244 G | (The absolute numbers are 8% below the ones this description had before: the benchmark's own object file is built differently here. All rows are from one recipe.) | 20 M `mi_malloc(128 KiB)` + `mi_free`, the generic path into a large page | instructions per pair | | --- | --- | | main's pin | 230.6 | | `58460467` | 193.6 | | this pin | 206.6 (+13.0 for the epoch stamp and its gate; -10% against main's pin) | **bun** (5 interleaved runs pinned to 8 cores; user-mode instructions and peak RSS, medians): | | base | pin | this PR | | --- | --- | --- | --- | | `bun -e 1` | 10.405 M, 27.0 MB | 10.410 M (+0.04%), 27.0 MB | 10.412 M (+0.06%), 27.0 MB | | `Bun.serve` hello, per request (20000 sequential `fetch`) | 52071, 49.0 MB | 51129 (-1.8%), 50.6 MB | 51361 (-1.4%; the spread between runs is 3%), 50.1 MB | | `Bun.serve` start and stop alone | 12.180 M | 12.179 M | 12.182 M (+0.01%) | | `JSON.stringify` + `JSON.parse` of a 200-user object, 3000 times | 5105.20 M, 44.0 MB | 5104.58 M, 44.0 MB | 5104.00 M (-0.02%), 43.5 MB | | string concat + `Buffer.from` + `split`, 300 x 5000 | 1900 M, 59.6 MB | 1912 M (+0.6%; the spread between its runs is 6.5%) , 60.5 MB | 1897 M (-0.2%), 59.7 MB | | idle after start (`Bun.sleep`), RSS / anonymous from `smaps_rollup`, binaries read into the page cache first | 28.42 MB / 3.14 | 28.45 MB / 3.05 | 28.45 MB / 3.05 | Nothing is slower by more than 0.1% outside the spread of its own runs. **The large-buffer server of oven-sh/mimalloc#46 at bun level** (`Bun.serve` that reads 256 KiB request bodies and answers 90 to 270 KB of JSON; 16 connections for 15 s, then idle; RSS / anonymous from `smaps_rollup`): | at idle | base | pin | this PR | | --- | --- | --- | --- | | 40 bodies kept, 2 s idle | 74.7 / 36.2 MB | 54.3 / 15.9 MB | 54.3 / 15.9 MB (**-20.4 MB**) | | 40 bodies kept, 15 s idle | 74.7 / 36.2 MB | 54.2 / 15.8 MB | 54.2 / 15.9 MB | | nothing kept, 15 s idle | 47.0 / 8.6 MB | 43.6 / 5.3 MB | 43.7 / 5.5 MB (-3.3 MB) | | at the end of the load, before the idle (40 kept) | 74.4 / 36.4 MB | 73.7 / 35.7 MB | 73.9 / 35.9 MB | Under sustained load (the same server, 20 s, 3 interleaved rounds; instructions and page faults of the server process from `perf stat -p`): | | requests | user instr / request | kernel instr / request | page faults / request | | --- | --- | --- | --- | --- | | base | 83.6 k .. 92.7 k | 283.9 k .. 284.5 k | 119.4 k .. 126.6 k | 0.34 .. 0.53 | | pin | 87.0 k .. 91.4 k | 284.0 k .. 284.1 k | 118.8 k .. 127.4 k | 0.50 .. 0.64 | | this PR | 80.4 k .. 91.4 k | 284.0 k .. 284.7 k | 120.9 k .. 125.3 k | 0.51 .. 0.86 | Requests and instructions per request are the same; the buffers stay while the server is busy and go when it is not. **Tests with the release build of this PR** - `test/js/bun/spawn/spawn-pipe-leak.test.ts` 25 of 25 runs, `test/js/bun/jsc/heapStats-mimalloc.test.ts` (with the new test) and `test/js/bun/gc/gc-controller-cadence.test.ts` 10 of 10 each. - 66 test files, 2165 tests pass: `require-cache`, `serve.test.ts`, all of `test/js/web/workers`, `test/js/node/worker_threads`, `test/js/node/child_process` and `test/js/node/fs` (the pool-heavy ones), `spawn-noread-leak`, `bun-file`, `bun-write`, `bundler_edgecase`, `esbuild/splitting`. The two failures need a non-root user (the root-range port in `serve.test.ts`, EPERM after dropping privileges in `child_process.test.ts`) and fail the same way with the base build. - The wider list of 113 files (bake dev, bundler, fetch, http, streams, sqlite, install, resolve, shell and others): 2485 tests pass; the two failures are in `resolve.test.ts`, want a directory that root cannot read, and fail the same way with the base build. - The new test fails with `USE_SYSTEM_BUN=1` (bun 1.4.1: 50 MB) and with the first commit alone (38 to 47 MB), and passes with both. - No debug or ASAN build was made for this revision of the PR (the earlier revisions ran `heapStats-mimalloc`, `gc-controller-cadence` and `worker.test.ts` on one at `58460467`); the new test skips itself under ASAN, where malloc is not mimalloc. **Not run here:** macOS and Windows. The Windows half of the second commit (`BunJSCEventLoop.cpp` under `OS(WINDOWS)`, `libuv.c`, `libuv.h`) was compiled by nothing on this machine: the C++ function was syntax-checked with `OS(WINDOWS)` forced on against stubs, `cargo check --target x86_64-pc-windows-msvc` passes for `bun_uws_sys`, `bun_threading` and `bun_mimalloc_sys` (the `WindowsLoop` mirror gained the field), and the rest is for CI. The fork's CI passed 23 of 23 jobs at each of the six merges (Linux x64/arm64 with ASAN, UBSAN, TSAN and guarded builds, the `MI_FREE_USE_PAGEMAP` build that matches bun's configuration, macOS, Windows x64/arm64, FreeBSD).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up to #39 from a review after its merge.
Problem
_mi_arenas_purge_now(src/arena.c) pulls an arena expire forward with a load and a store, and does the same forsubproc->purge_expire. A purge pass resets both. A store that lands after the reset arms an arena that the pass just handled.test-purge-behind-passcovers one arena. With two, a pass that only fixesall_visitedstill loses a free into an arena that the pass has left, and the test does not see it.Fix
mi_subproc_schedule_purge, as in a pass.all_visitedfix (258 of 264 MiB, 3 of 3 runs) and passes here. Full ctest passes.Background
arena->purge_expireis when an arena is due. A free arms it (sets it from 0).subproc->purge_expireis the earliest of these. The scavenger sleeps until then.subproc->purge_expirefirst and puts back the earliest arena expire still pending at its end (arena: what is freed during a purge pass is purged by the next pass #39)._mi_arenas_purge_nowruns when a thread goes idle. It makes every scheduled purge due at once.Notes
The third part needs no thread of its own.
arena_purgescounts the arenas a pass went into, so the test waits for it to move by two (the main arena first, then the arena of the test) and frees the last block then. If the pass is over by then, the free takes the normal path and the part passes as well, so it cannot fail on a build that has the fix.The reservation for the test arena is four times what goes into it. A block of 8 MiB and a bit starts on a boundary in an arena: three share 32 MiB on a 64-bit target, and each one takes 32 MiB on a 32-bit one. The first push reserved less, and the part failed with
out of memoryin the Win32 Debug jobs (3 of the 4 blocks fit). An arena that still does not take all of the blocks ends the part with a message and does not fail it. On Win32 (Debug and Release, MSVC) the part now runs and ends with72 of 72 MiB.With
MI_TSANorMI_TEST_LIGHTthe first part runs two rounds and the third part uses a quarter of the blocks. That is less memory traffic in total than the test had before the third part, so the thread sanitizer job has no more to do per test than before. The first push was red there as well. I cannot read the job logs from where I work, so I do not know which test it was. A Debug build spends 20 s in this test (17 s of it in the kernel), which makes the per test time limit my first suspect. The smaller part still fails on the one lineall_visitedfix (129 of 136 MiB, 2 of 3 runs).Not done, and why. The review also proposed a debug check under the purge guard that an armed arena implies an armed subprocess. A free arms the arena first and the subprocess second, so the check can see the state between the two and fail without a defect.
Local runs on Linux x64: Release 31 of 31, Debug (
MI_DEBUG_FULL) 32 of 32. The purge, park and fork tests also pass with ASAN,MI_GUARDED,MI_SECUREandMI_USE_CXX. My stress harness (8 busy threads, pause without a park): 0 stuck rounds of 230.No pin bump in bun is needed for this by itself. It can ride with the next one.