Skip to content

arena: the idle pull-forward uses the CAS paths, and a two arena test for the purge schedule - #42

Merged
Jarred-Sumner merged 3 commits into
oven-sh:bun-dev3-v2from
robobun:robobun/39824388/purge-now-schedule
Sep 14, 2026
Merged

Jarred-Sumner merged 3 commits into
oven-sh:bun-dev3-v2from
robobun:robobun/39824388/purge-now-schedule

Conversation

@robobun

@robobun robobun commented Sep 13, 2026 •

Copy link
Copy Markdown
Collaborator

Follow-up to #39 from a review after its merge.

Problem

  • _mi_arenas_purge_now (src/arena.c) pulls an arena expire forward with a load and a store, and does the same for subproc->purge_expire. A purge pass resets both. A store that lands after the reset arms an arena that the pass just handled.
  • In the worst case an arena is armed and the subprocess is not, until the scavenger's 30 s look at the arenas. It needs a thread preempted between its load and its store for a whole pass: a hardening, nothing was seen.
  • test-purge-behind-pass covers one arena. With two, a pass that only fixes all_visited still loses a free into an arena that the pass has left, and the test does not see it.

Fix

  • The arena expire moves with a CAS from the value read. The subprocess expire goes through mi_subproc_schedule_purge, as in a pass.
  • New test part: both arenas get a batch, then one more block goes into the main arena while the pass is in the other.
  • Verified: the new part fails on 8565de2 plus the one line all_visited fix (258 of 264 MiB, 3 of 3 runs) and passes here. Full ctest passes.

Background

  • arena->purge_expire is when an arena is due. A free arms it (sets it from 0). subproc->purge_expire is the earliest of these. The scavenger sleeps until then.
  • A pass resets subproc->purge_expire first and puts back the earliest arena expire still pending at its end (arena: what is freed during a purge pass is purged by the next pass #39).
  • _mi_arenas_purge_now runs when a thread goes idle. It makes every scheduled purge due at once.
Notes

The third part needs no thread of its own. arena_purges counts the arenas a pass went into, so the test waits for it to move by two (the main arena first, then the arena of the test) and frees the last block then. If the pass is over by then, the free takes the normal path and the part passes as well, so it cannot fail on a build that has the fix.

The reservation for the test arena is four times what goes into it. A block of 8 MiB and a bit starts on a boundary in an arena: three share 32 MiB on a 64-bit target, and each one takes 32 MiB on a 32-bit one. The first push reserved less, and the part failed with out of memory in the Win32 Debug jobs (3 of the 4 blocks fit). An arena that still does not take all of the blocks ends the part with a message and does not fail it. On Win32 (Debug and Release, MSVC) the part now runs and ends with 72 of 72 MiB.

With MI_TSAN or MI_TEST_LIGHT the first part runs two rounds and the third part uses a quarter of the blocks. That is less memory traffic in total than the test had before the third part, so the thread sanitizer job has no more to do per test than before. The first push was red there as well. I cannot read the job logs from where I work, so I do not know which test it was. A Debug build spends 20 s in this test (17 s of it in the kernel), which makes the per test time limit my first suspect. The smaller part still fails on the one line all_visited fix (129 of 136 MiB, 2 of 3 runs).

Not done, and why. The review also proposed a debug check under the purge guard that an armed arena implies an armed subprocess. A free arms the arena first and the subprocess second, so the check can see the state between the two and fail without a defect.

Local runs on Linux x64: Release 31 of 31, Debug (MI_DEBUG_FULL) 32 of 32. The purge, park and fork tests also pass with ASAN, MI_GUARDED, MI_SECURE and MI_USE_CXX. My stress harness (8 busy threads, pause without a park): 0 stuck rounds of 230.

No pin bump in bun is needed for this by itself. It can ride with the next one.

…ena test for the purge schedule

`_mi_arenas_purge_now` pulled the expire of an arena forward with a load and a
store, and did the same for `subproc->purge_expire`. A pass resets both with a
CAS or before a walk, so a store that lands after that reset arms an arena that
the pass has just taken care of, and a value of `subproc->purge_expire` that was
read before the reset says nothing about what the pass leaves behind. In the
worst case an arena is armed and the subprocess is not until the scavenger's
30 s look at the arenas. The arena expire is now moved with a CAS from the value
that was read, and the subprocess expire goes through
`mi_subproc_schedule_purge`, the CAS loop that the passes use.

test-purge-behind-pass gets a third part for more than one arena. Both arenas
get a batch, and one more block goes into the main arena while the pass is in
the other one. A pass that only reads the expire of each arena as it leaves it
and installs the earliest one at its end misses that free: 8565de2 with the
one line fix for `all_visited` ends with `258 of 264 MiB`, 3 of 3 runs. The
first part alone passes there once a second arena exists.
…ks on a 32-bit target

A block of 8 MiB and a bit starts on a boundary in an arena: each one takes
32 MiB of it on a 32-bit target, where a 96 MiB arena held 3 of the 4 blocks
and the part failed with 'out of memory' (Debug, Win32). The reservation is
four times what goes in now, and an arena that still does not take all of the
blocks ends the part with a message and does not fail it.
…e sanitizer and light runs

The thread sanitizer job has a time limit for each test. With MI_TSAN or
MI_TEST_LIGHT the first part runs two rounds and the third part works with a
quarter of the blocks, which is less memory traffic in total than the test had
before the third part. The smaller part still fails on the one line fix for
all_visited (129 of 136 MiB, 2 of 3 runs).
@Jarred-Sumner
Jarred-Sumner merged commit 3aa0ac9 into oven-sh:bun-dev3-v2 Sep 14, 2026
23 checks passed
Jarred-Sumner added a commit to oven-sh/bun that referenced this pull request Sep 14, 2026
Pins 3aa0ac9bea44, the merge of oven-sh/mimalloc#42 on bun-dev3-v2, on top of #40 and #43.

#42 hardens the purge pull-forward that runs when a thread goes idle (mi_on_thread_idle -> _mi_arenas_purge_now). It moved each
armed arena's expire, and the sub-process expire that the purge thread sleeps on, to "now" with a load and a store. A purge pass
resets both first thing, so a store based on a value read before that reset could undo it. The arena expire is now moved with a
compare-and-swap, and the sub-process expire through the CAS loop that the passes use. Nothing was observed; it closes the window.

12 lines in src/arena.c, no header changes. Allocator micro benchmark: 2.2384 G instructions before and after.
Jarred-Sumner added a commit to oven-sh/bun that referenced this pull request Sep 15, 2026
Pins 3aa0ac9bea44, the merge of oven-sh/mimalloc#42 on bun-dev3-v2, on top of #40 and #43.

#42 hardens the purge pull-forward that runs when a thread goes idle (mi_on_thread_idle -> _mi_arenas_purge_now). It moved each
armed arena's expire, and the sub-process expire that the purge thread sleeps on, to "now" with a load and a store. A purge pass
resets both first thing, so a store based on a value read before that reset could undo it. The arena expire is now moved with a
compare-and-swap, and the sub-process expire through the CAS loop that the passes use. Nothing was observed; it closes the window.

12 lines in src/arena.c, no header changes. Allocator micro benchmark: 2.2384 G instructions before and after.
Jarred-Sumner added a commit to oven-sh/bun that referenced this pull request Sep 15, 2026
Pins 3aa0ac9bea44, the merge of oven-sh/mimalloc#42 on bun-dev3-v2, on top of #40 and #43.

#42 hardens the purge pull-forward that runs when a thread goes idle (mi_on_thread_idle -> _mi_arenas_purge_now). It moved each
armed arena's expire, and the sub-process expire that the purge thread sleeps on, to "now" with a load and a store. A purge pass
resets both first thing, so a store based on a value read before that reset could undo it. The arena expire is now moved with a
compare-and-swap, and the sub-process expire through the CAS loop that the passes use. Nothing was observed; it closes the window.

12 lines in src/arena.c, no header changes. Allocator micro benchmark: 2.2384 G instructions before and after.
Jarred-Sumner added a commit to oven-sh/bun that referenced this pull request Sep 15, 2026
, #44, #45, #46) (#42668)

### What does this PR do?

Moves `MIMALLOC_COMMIT` from `ab13501334a8` to
`69050cf3ec14be4847ca3ba6c840da821fa283b0`, the head of `bun-dev3-v2`:
main's pin plus six merged PRs of the fork and nothing else, and changes
the two callers in bun that the last of them asks to change.
`MI_FREE_USE_PAGEMAP=1` stays.

- oven-sh/mimalloc#40: the sampled profiler hooks (`mi_profiler_t`)
attribute a sample to the small allocation that refilled its page, keep
`bytes_since_last_sample` right when `on_alloc` raises the rate, sample
a new thread's heap from its first allocation, report the frees of
`mi_heap_destroy`, and give `on_free` of an aligned block the pointer
that `on_alloc` got. What the heap profiler in #42626 needs.
- oven-sh/mimalloc#43: an aligned request only becomes a sample through
the over-allocating path (with #40 alone `mi_malloc_aligned(64, 64)`
could free a sampled block inside the allocator; `MI_DEBUG` asserted).
- oven-sh/mimalloc#42: the purge pull-forward of an idle thread
(`mi_on_thread_idle`) goes through compare-and-swap, so it cannot undo
the reset that a purge pass does first thing. 12 lines in `src/arena.c`.
- oven-sh/mimalloc#44: a thread that does not sample pays for the
profiler hooks in the generic path only: +0.2% of instructions on the
allocator micro benchmark where they cost +1.5%.
- oven-sh/mimalloc#45: a debug assertion that raced with
`mi_profiler_stop` is gone.
- oven-sh/mimalloc#46: **free blocks in the 4 MiB pages (blocks of 96 to
512 KiB) are handed back to the OS.** Hole purging skipped these pages
(their 1024 OS pages do not fit the 256-bit purge bitmap), so what a
process freed in that range stayed resident for as long as one block of
the page was alive: 97% of it by `mincore`. The bitmap's unit is now per
page (16 KiB for a 4 MiB page), and the free blocks of a large page stay
while the page is in use, which the allocator tells without a clock on
the allocating side: the idle sweep counts epochs of 100 ms or more, an
allocation from a large page leaves the current epoch in the page (a
load, a compare, and a store once per epoch), and the blocks go when
that is two epochs ago. Without that hold a `Bun.serve` with 256 KiB
bodies discarded and refaulted 3 GB in 20 s.

Nothing in bun attaches a profiler yet. The fast paths of `mi_malloc`
and `mi_free` are untouched by all six.

`include/mimalloc.h` gains one function, `mi_on_thread_idle_pending()`,
and comments; `mimalloc-stats.h` is identical, `mimalloc-profile.h`
differs in comments, `mimalloc/types.h` gains fields of the internal
`mi_theap_t` and `mi_tld_t`. No struct or enum that bun binds changed
(`src/mimalloc_sys/mimalloc.rs`, the declarations in
`BunJSCEventLoop.cpp` and `epoll_kqueue.c`).

#### The callers that sweep once and then block

With oven-sh/mimalloc#46 ONE call of `mi_on_thread_idle()` never takes
the free blocks of a large page that the thread allocated from since its
previous sweep: a sweep ends one epoch at the most. The event loop is
fine: `us_loop_run_bun_tick` hands the thread's heaps to mimalloc's
scavenger thread across the poll (`mi_on_thread_idle_start()` /
`_end()`), and the scavenger comes back for them by itself, two or three
sweeps an interval apart. Two callers swept inline once and then blocked
for good, and would have kept such blocks, as they do on main:

- **`src/threading/ThreadPool.rs`**: a worker swept once, after its
futex wait timed out at 100 ms, and then waited without a timeout. It
now brackets that second wait with `mi_on_thread_idle_start()` /
`mi_on_thread_idle_end()`, and sweeps inline (once, as before) only when
`_start` says there is no scavenger to hand off to. Nothing runs between
the two calls but `Futex::wait`, which allocates nothing;
`drain_idle_events` runs before `_start`.
Only the wait that follows a timed-out one is bracketed. Bracketing
every park was measured and is slower (table below: `_start` is two
read-modify-writes on cache lines that every parking thread shares, and
it wakes the scavenger, which costs half a context switch per task), and
sweeping on every park is what the comment there warns about (13% of
`vite preview` requests per second). As it is, a worker that parks
between tasks never reaches `_start`, and `notify()` is what it was.
- **The libuv loop on Windows** (`Bun__JSC_onBeforeWait` in
`BunJSCEventLoop.cpp`, from `us_loop_run` in `libuv.c`) sweeps inline at
most every 100 ms and then `uv_run(UV_RUN_ONCE)` blocks. It cannot hand
off: `uv_run` dispatches the completions right after its poll inside the
same call (`uv__poll` at `src/win/core.c:732`, `uv__process_reqs` at
`:737`), and libuv allocates from mimalloc there
(`uv_replace_allocator`), so there is no point between "about to block"
and "awake" that bun's code sees without a patch to libuv. It keeps the
inline sweep, and while `mi_on_thread_idle_pending()` says blocks were
left, `Bun__JSC_onBeforeWait` returns 100 (ms) and `us_loop_run` arms an
unref'd timer for that long after the sweep, so the hook runs again:
three times at the most after JS last ran, which is the bound the fork
gives. The timer does not keep the loop alive.

Not changed: when no scavenger runs (`MIMALLOC_SCAVENGER=0` or a purge
delay of 0; bun sets neither) the fallbacks in `us_loop_run_bun_tick`
and in the pool sweep inline once, as on main, and the free blocks of
large pages stay.

### How did you verify your code works?

Release builds of one tree (this branch on `04823d8fad67`), Linux x64:
**base** = main's pin, **pin** = this pin with the callers unchanged
(the first commit), **this PR** = both commits.

**Test.** `test/js/bun/jsc/heapStats-mimalloc.test.ts` gains "an idle
work pool thread hands back the free blocks of its large pages": a
subprocess with 4 pool threads reads 480 files of 96 to 512 KiB with
`fs.promises.readFile` (the buffers are allocated on the pool threads),
keeps 1 in 12, collects the rest, and polls `RssAnon` while nothing
gives the pool work. Anonymous memory beyond what it started with and
what it keeps has to come down to 12 MB within 2.5 s.

| | free and resident at the end |
| --- | --- |
| base | 54 to 67 MB: fails |
| pin, callers unchanged | 38 to 47 MB: fails (also with
`BUN_GARBAGE_COLLECTOR_LEVEL=1` and `2`, and with the system bun 1.4.1:
50 MB) |
| this PR | -1 MB, reached 350 ms into the idle: passes, 10 of 10 runs |

**The park and notify path of the pool** (what the `has_swept` comment
protects). 100000 `await fs.promises.stat("/")` one after the other
(every call notifies a worker, which runs one task and parks), the same
in batches of 8, and `await randomBytes(16)`; `perf stat` of the whole
process, 8 cores, 5 interleaved runs, medians (range):

| user instructions per call | base | pin | this PR |
| --- | --- | --- | --- |
| `stat`, sequential | 8794 (8784..8828) | 8792 (8716..8827) | 8797
(8780..8816) |
| `stat`, 8 at a time | 6027 (6014..6131) | 6025 (6004..6080) | 6034
(6008..6102) |
| `randomBytes(16)`, sequential | 12608 (12513..12629) | 12581
(12522..12609) | 12589 (12582..12627) |

Context switches per call are 3.96 to 3.97, 0.51 to 0.53 and 3.97 in all
three, kernel instructions per call overlap (41.6 k..49.9 k sequential),
wall time is within the spread of the machine.

Bracketing EVERY park instead (a build with a switch for it, 200000
calls, 4 interleaved runs), against the same build with the handoff
after a timed-out wait only:

| | after a timed-out wait (this PR) | every park |
| --- | --- | --- |
| `stat` sequential: user / kernel instructions, context switches per
call | 8565..8601 / 45.6 k..46.7 k / 3.95..3.99 | 9193..9335 (+7%) /
49.6 k..50.7 k (+9%) / 4.2..4.55 |
| `stat` 8 at a time | 5635..5659 / 8.8 k..9.3 k / 0.53..0.55 |
5771..5960 (+3%) / 10.1 k..11.6 k (+13%) / 0.63..0.75 |
| `randomBytes(16)` sequential | 12145..12283 / 43.2 k..45.5 k /
3.97..3.99 | 12815..12893 (+5%) / 48.2 k..51.2 k (+11%) / 4.41..4.5 |
| wall time, `stat` sequential | 11.6 s..12.7 s | 12.4 s..13.5 s (+6 to
8%) |

**Allocator micro benchmarks** (mimalloc alone with bun's flags: clang
21, `-O3 -march=nehalem`, C++, `MI_FREE_USE_PAGEMAP`; user-mode
instructions, the runs of one build equal to within 100):

| 25 M `mi_malloc` + 25 M `mi_free` of 16..1039 bytes | instructions |
| --- | --- |
| before the upstream sync (`707d90be`) | 2.0731 G |
| main's pin (`ab135013`) | 2.0568 G |
| with oven-sh/mimalloc#40, #43, #42 (`3aa0ac9b`) | 2.0556 G (-0.06%) |
| with #44, #45 (`58460467`) | 2.0305 G (-1.28%) |
| this pin (`69050cf3`) | **2.0305 G (-1.28%)**, the same as `58460467`
to within 400 instructions |
| this pin with `MI_PROFILE=0` (not what bun builds) | 2.0244 G |

(The absolute numbers are 8% below the ones this description had before:
the benchmark's own object file is built differently here. All rows are
from one recipe.)

| 20 M `mi_malloc(128 KiB)` + `mi_free`, the generic path into a large
page | instructions per pair |
| --- | --- |
| main's pin | 230.6 |
| `58460467` | 193.6 |
| this pin | 206.6 (+13.0 for the epoch stamp and its gate; -10% against
main's pin) |

**bun** (5 interleaved runs pinned to 8 cores; user-mode instructions
and peak RSS, medians):

| | base | pin | this PR |
| --- | --- | --- | --- |
| `bun -e 1` | 10.405 M, 27.0 MB | 10.410 M (+0.04%), 27.0 MB | 10.412 M
(+0.06%), 27.0 MB |
| `Bun.serve` hello, per request (20000 sequential `fetch`) | 52071,
49.0 MB | 51129 (-1.8%), 50.6 MB | 51361 (-1.4%; the spread between runs
is 3%), 50.1 MB |
| `Bun.serve` start and stop alone | 12.180 M | 12.179 M | 12.182 M
(+0.01%) |
| `JSON.stringify` + `JSON.parse` of a 200-user object, 3000 times |
5105.20 M, 44.0 MB | 5104.58 M, 44.0 MB | 5104.00 M (-0.02%), 43.5 MB |
| string concat + `Buffer.from` + `split`, 300 x 5000 | 1900 M, 59.6 MB
| 1912 M (+0.6%; the spread between its runs is 6.5%) , 60.5 MB | 1897 M
(-0.2%), 59.7 MB |
| idle after start (`Bun.sleep`), RSS / anonymous from `smaps_rollup`,
binaries read into the page cache first | 28.42 MB / 3.14 | 28.45 MB /
3.05 | 28.45 MB / 3.05 |

Nothing is slower by more than 0.1% outside the spread of its own runs.

**The large-buffer server of oven-sh/mimalloc#46 at bun level**
(`Bun.serve` that reads 256 KiB request bodies and answers 90 to 270 KB
of JSON; 16 connections for 15 s, then idle; RSS / anonymous from
`smaps_rollup`):

| at idle | base | pin | this PR |
| --- | --- | --- | --- |
| 40 bodies kept, 2 s idle | 74.7 / 36.2 MB | 54.3 / 15.9 MB | 54.3 /
15.9 MB (**-20.4 MB**) |
| 40 bodies kept, 15 s idle | 74.7 / 36.2 MB | 54.2 / 15.8 MB | 54.2 /
15.9 MB |
| nothing kept, 15 s idle | 47.0 / 8.6 MB | 43.6 / 5.3 MB | 43.7 / 5.5
MB (-3.3 MB) |
| at the end of the load, before the idle (40 kept) | 74.4 / 36.4 MB |
73.7 / 35.7 MB | 73.9 / 35.9 MB |

Under sustained load (the same server, 20 s, 3 interleaved rounds;
instructions and page faults of the server process from `perf stat -p`):

| | requests | user instr / request | kernel instr / request | page
faults / request |
| --- | --- | --- | --- | --- |
| base | 83.6 k .. 92.7 k | 283.9 k .. 284.5 k | 119.4 k .. 126.6 k |
0.34 .. 0.53 |
| pin | 87.0 k .. 91.4 k | 284.0 k .. 284.1 k | 118.8 k .. 127.4 k |
0.50 .. 0.64 |
| this PR | 80.4 k .. 91.4 k | 284.0 k .. 284.7 k | 120.9 k .. 125.3 k |
0.51 .. 0.86 |

Requests and instructions per request are the same; the buffers stay
while the server is busy and go when it is not.

**Tests with the release build of this PR**
- `test/js/bun/spawn/spawn-pipe-leak.test.ts` 25 of 25 runs,
`test/js/bun/jsc/heapStats-mimalloc.test.ts` (with the new test) and
`test/js/bun/gc/gc-controller-cadence.test.ts` 10 of 10 each.
- 66 test files, 2165 tests pass: `require-cache`, `serve.test.ts`, all
of `test/js/web/workers`, `test/js/node/worker_threads`,
`test/js/node/child_process` and `test/js/node/fs` (the pool-heavy
ones), `spawn-noread-leak`, `bun-file`, `bun-write`, `bundler_edgecase`,
`esbuild/splitting`. The two failures need a non-root user (the
root-range port in `serve.test.ts`, EPERM after dropping privileges in
`child_process.test.ts`) and fail the same way with the base build.
- The wider list of 113 files (bake dev, bundler, fetch, http, streams,
sqlite, install, resolve, shell and others): 2485 tests pass; the two
failures are in `resolve.test.ts`, want a directory that root cannot
read, and fail the same way with the base build.
- The new test fails with `USE_SYSTEM_BUN=1` (bun 1.4.1: 50 MB) and with
the first commit alone (38 to 47 MB), and passes with both.
- No debug or ASAN build was made for this revision of the PR (the
earlier revisions ran `heapStats-mimalloc`, `gc-controller-cadence` and
`worker.test.ts` on one at `58460467`); the new test skips itself under
ASAN, where malloc is not mimalloc.

**Not run here:** macOS and Windows. The Windows half of the second
commit (`BunJSCEventLoop.cpp` under `OS(WINDOWS)`, `libuv.c`, `libuv.h`)
was compiled by nothing on this machine: the C++ function was
syntax-checked with `OS(WINDOWS)` forced on against stubs, `cargo check
--target x86_64-pc-windows-msvc` passes for `bun_uws_sys`,
`bun_threading` and `bun_mimalloc_sys` (the `WindowsLoop` mirror gained
the field), and the rest is for CI. The fork's CI passed 23 of 23 jobs
at each of the six merges (Linux x64/arm64 with ASAN, UBSAN, TSAN and
guarded builds, the `MI_FREE_USE_PAGEMAP` build that matches bun's
configuration, macOS, Windows x64/arm64, FreeBSD).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants