Repository navigation
profile: sample the small allocation that refilled its page, new threads, and the blocks mi_heap_destroy frees - #40
Merged
Conversation
…ads, and the blocks mi_heap_destroy frees - page.c: with coarse sampling the blocks the fast path took from a page are counted when its free list is refilled. If that uses up the countdown, the allocation in progress is the sample, not the next one through the generic path (which is an allocation above MI_SMALL_SIZE_MAX far more often than a small one). mi_page_extend_free counts the previous extension too. - sample-profile.c: a rate that on_alloc raised starts its period in full, so bytes_since_last_sample stays what was requested. - theap.c: a theap created while a profiler runs is sampled from its first allocation, and starts with a full period. - sample-profile.c, arena.c: mi_heap_destroy calls on_free for the sampled blocks still in use. A profiled block carries a check word (address xor a random key) so the walk does not take a regular block that starts with the tag's value for one; the word is cleared when the block is freed. No instruction is added to the mi_malloc/mi_free fast paths.
…ned allocation mi_malloc_aligned aligns inside the block that _mi_theap_malloc_profiled returned, so mi_free of a sampled aligned block called on_free with the aligned pointer. test-profile's on_free asserts they are the same; it met such a block with guarded sampling on in the same process (the CI's guarded jobs set a rate for every test): samples then land on mimalloc's own mi_heap_zalloc_aligned of the per-heap arena page table, which mi_heap_destroy frees. _mi_page_profile_on_free computes the pointer from the block's layout, and checks the block's check word before it calls on_free. test-profile gets MIMALLOC_GUARDED_SAMPLE_RATE=0 explicitly in guarded builds (CMakeLists.txt meant it to have none: what the samples add up to is off with guarded samples in between, also without this branch). The new case turns guarded sampling on for itself: sampled, guarded and aligned blocks through mi_free, mi_free on another thread, mi_heap_realloc and mi_heap_destroy.
This was referenced Sep 14, 2026
Jarred-Sumner
added a commit
to oven-sh/bun
that referenced
this pull request
Sep 15, 2026
Pins 64d2625acfed, the merge of oven-sh/mimalloc#40 on bun-dev3-v2. Five fixes to the sampled profiler hooks (mi_profiler_t) that came in with the upstream sync: - the small allocation that refills its page is the one that gets sampled (before, the sample went to the next allocation above 1 KiB far more often than to the size class that used up the countdown) - bytes_since_last_sample stays right when on_alloc raises the rate; a new thread's heap is sampled from its first allocation - mi_heap_destroy reports the frees of the sampled blocks it releases, and on_free of an aligned sampled block gets the pointer that on_alloc got They are what the heap profiler in #42626 needs for unbiased attribution and correct in-use accounting. Nothing in bun attaches a profiler yet: without one the change is one load and a branch on the allocator's generic path. 25 M malloc + 25 M free with bun's flags: 2.2403 G -> 2.2427 G instructions (+0.11%; 2.2463 G before the upstream sync). bun -e 1, Bun.serve, JSON and string workloads: within 0.3% of the old pin, inside the run to run spread.
Jarred-Sumner
added a commit
to oven-sh/bun
that referenced
this pull request
Sep 15, 2026
Pins 64d2625acfed, the merge of oven-sh/mimalloc#40 on bun-dev3-v2. Five fixes to the sampled profiler hooks (mi_profiler_t) that came in with the upstream sync: - the small allocation that refills its page is the one that gets sampled (before, the sample went to the next allocation above 1 KiB far more often than to the size class that used up the countdown) - bytes_since_last_sample stays right when on_alloc raises the rate; a new thread's heap is sampled from its first allocation - mi_heap_destroy reports the frees of the sampled blocks it releases, and on_free of an aligned sampled block gets the pointer that on_alloc got They are what the heap profiler in #42626 needs for unbiased attribution and correct in-use accounting. Nothing in bun attaches a profiler yet: without one the change is one load and a branch on the allocator's generic path. 25 M malloc + 25 M free with bun's flags: 2.2403 G -> 2.2427 G instructions (+0.11%; 2.2463 G before the upstream sync). bun -e 1, Bun.serve, JSON and string workloads: within 0.3% of the old pin, inside the run to run spread.
Jarred-Sumner
added a commit
to oven-sh/bun
that referenced
this pull request
Sep 15, 2026
, #44, #45, #46) (#42668) ### What does this PR do? Moves `MIMALLOC_COMMIT` from `ab13501334a8` to `69050cf3ec14be4847ca3ba6c840da821fa283b0`, the head of `bun-dev3-v2`: main's pin plus six merged PRs of the fork and nothing else, and changes the two callers in bun that the last of them asks to change. `MI_FREE_USE_PAGEMAP=1` stays. - oven-sh/mimalloc#40: the sampled profiler hooks (`mi_profiler_t`) attribute a sample to the small allocation that refilled its page, keep `bytes_since_last_sample` right when `on_alloc` raises the rate, sample a new thread's heap from its first allocation, report the frees of `mi_heap_destroy`, and give `on_free` of an aligned block the pointer that `on_alloc` got. What the heap profiler in #42626 needs. - oven-sh/mimalloc#43: an aligned request only becomes a sample through the over-allocating path (with #40 alone `mi_malloc_aligned(64, 64)` could free a sampled block inside the allocator; `MI_DEBUG` asserted). - oven-sh/mimalloc#42: the purge pull-forward of an idle thread (`mi_on_thread_idle`) goes through compare-and-swap, so it cannot undo the reset that a purge pass does first thing. 12 lines in `src/arena.c`. - oven-sh/mimalloc#44: a thread that does not sample pays for the profiler hooks in the generic path only: +0.2% of instructions on the allocator micro benchmark where they cost +1.5%. - oven-sh/mimalloc#45: a debug assertion that raced with `mi_profiler_stop` is gone. - oven-sh/mimalloc#46: **free blocks in the 4 MiB pages (blocks of 96 to 512 KiB) are handed back to the OS.** Hole purging skipped these pages (their 1024 OS pages do not fit the 256-bit purge bitmap), so what a process freed in that range stayed resident for as long as one block of the page was alive: 97% of it by `mincore`. The bitmap's unit is now per page (16 KiB for a 4 MiB page), and the free blocks of a large page stay while the page is in use, which the allocator tells without a clock on the allocating side: the idle sweep counts epochs of 100 ms or more, an allocation from a large page leaves the current epoch in the page (a load, a compare, and a store once per epoch), and the blocks go when that is two epochs ago. Without that hold a `Bun.serve` with 256 KiB bodies discarded and refaulted 3 GB in 20 s. Nothing in bun attaches a profiler yet. The fast paths of `mi_malloc` and `mi_free` are untouched by all six. `include/mimalloc.h` gains one function, `mi_on_thread_idle_pending()`, and comments; `mimalloc-stats.h` is identical, `mimalloc-profile.h` differs in comments, `mimalloc/types.h` gains fields of the internal `mi_theap_t` and `mi_tld_t`. No struct or enum that bun binds changed (`src/mimalloc_sys/mimalloc.rs`, the declarations in `BunJSCEventLoop.cpp` and `epoll_kqueue.c`). #### The callers that sweep once and then block With oven-sh/mimalloc#46 ONE call of `mi_on_thread_idle()` never takes the free blocks of a large page that the thread allocated from since its previous sweep: a sweep ends one epoch at the most. The event loop is fine: `us_loop_run_bun_tick` hands the thread's heaps to mimalloc's scavenger thread across the poll (`mi_on_thread_idle_start()` / `_end()`), and the scavenger comes back for them by itself, two or three sweeps an interval apart. Two callers swept inline once and then blocked for good, and would have kept such blocks, as they do on main: - **`src/threading/ThreadPool.rs`**: a worker swept once, after its futex wait timed out at 100 ms, and then waited without a timeout. It now brackets that second wait with `mi_on_thread_idle_start()` / `mi_on_thread_idle_end()`, and sweeps inline (once, as before) only when `_start` says there is no scavenger to hand off to. Nothing runs between the two calls but `Futex::wait`, which allocates nothing; `drain_idle_events` runs before `_start`. Only the wait that follows a timed-out one is bracketed. Bracketing every park was measured and is slower (table below: `_start` is two read-modify-writes on cache lines that every parking thread shares, and it wakes the scavenger, which costs half a context switch per task), and sweeping on every park is what the comment there warns about (13% of `vite preview` requests per second). As it is, a worker that parks between tasks never reaches `_start`, and `notify()` is what it was. - **The libuv loop on Windows** (`Bun__JSC_onBeforeWait` in `BunJSCEventLoop.cpp`, from `us_loop_run` in `libuv.c`) sweeps inline at most every 100 ms and then `uv_run(UV_RUN_ONCE)` blocks. It cannot hand off: `uv_run` dispatches the completions right after its poll inside the same call (`uv__poll` at `src/win/core.c:732`, `uv__process_reqs` at `:737`), and libuv allocates from mimalloc there (`uv_replace_allocator`), so there is no point between "about to block" and "awake" that bun's code sees without a patch to libuv. It keeps the inline sweep, and while `mi_on_thread_idle_pending()` says blocks were left, `Bun__JSC_onBeforeWait` returns 100 (ms) and `us_loop_run` arms an unref'd timer for that long after the sweep, so the hook runs again: three times at the most after JS last ran, which is the bound the fork gives. The timer does not keep the loop alive. Not changed: when no scavenger runs (`MIMALLOC_SCAVENGER=0` or a purge delay of 0; bun sets neither) the fallbacks in `us_loop_run_bun_tick` and in the pool sweep inline once, as on main, and the free blocks of large pages stay. ### How did you verify your code works? Release builds of one tree (this branch on `04823d8fad67`), Linux x64: **base** = main's pin, **pin** = this pin with the callers unchanged (the first commit), **this PR** = both commits. **Test.** `test/js/bun/jsc/heapStats-mimalloc.test.ts` gains "an idle work pool thread hands back the free blocks of its large pages": a subprocess with 4 pool threads reads 480 files of 96 to 512 KiB with `fs.promises.readFile` (the buffers are allocated on the pool threads), keeps 1 in 12, collects the rest, and polls `RssAnon` while nothing gives the pool work. Anonymous memory beyond what it started with and what it keeps has to come down to 12 MB within 2.5 s. | | free and resident at the end | | --- | --- | | base | 54 to 67 MB: fails | | pin, callers unchanged | 38 to 47 MB: fails (also with `BUN_GARBAGE_COLLECTOR_LEVEL=1` and `2`, and with the system bun 1.4.1: 50 MB) | | this PR | -1 MB, reached 350 ms into the idle: passes, 10 of 10 runs | **The park and notify path of the pool** (what the `has_swept` comment protects). 100000 `await fs.promises.stat("/")` one after the other (every call notifies a worker, which runs one task and parks), the same in batches of 8, and `await randomBytes(16)`; `perf stat` of the whole process, 8 cores, 5 interleaved runs, medians (range): | user instructions per call | base | pin | this PR | | --- | --- | --- | --- | | `stat`, sequential | 8794 (8784..8828) | 8792 (8716..8827) | 8797 (8780..8816) | | `stat`, 8 at a time | 6027 (6014..6131) | 6025 (6004..6080) | 6034 (6008..6102) | | `randomBytes(16)`, sequential | 12608 (12513..12629) | 12581 (12522..12609) | 12589 (12582..12627) | Context switches per call are 3.96 to 3.97, 0.51 to 0.53 and 3.97 in all three, kernel instructions per call overlap (41.6 k..49.9 k sequential), wall time is within the spread of the machine. Bracketing EVERY park instead (a build with a switch for it, 200000 calls, 4 interleaved runs), against the same build with the handoff after a timed-out wait only: | | after a timed-out wait (this PR) | every park | | --- | --- | --- | | `stat` sequential: user / kernel instructions, context switches per call | 8565..8601 / 45.6 k..46.7 k / 3.95..3.99 | 9193..9335 (+7%) / 49.6 k..50.7 k (+9%) / 4.2..4.55 | | `stat` 8 at a time | 5635..5659 / 8.8 k..9.3 k / 0.53..0.55 | 5771..5960 (+3%) / 10.1 k..11.6 k (+13%) / 0.63..0.75 | | `randomBytes(16)` sequential | 12145..12283 / 43.2 k..45.5 k / 3.97..3.99 | 12815..12893 (+5%) / 48.2 k..51.2 k (+11%) / 4.41..4.5 | | wall time, `stat` sequential | 11.6 s..12.7 s | 12.4 s..13.5 s (+6 to 8%) | **Allocator micro benchmarks** (mimalloc alone with bun's flags: clang 21, `-O3 -march=nehalem`, C++, `MI_FREE_USE_PAGEMAP`; user-mode instructions, the runs of one build equal to within 100): | 25 M `mi_malloc` + 25 M `mi_free` of 16..1039 bytes | instructions | | --- | --- | | before the upstream sync (`707d90be`) | 2.0731 G | | main's pin (`ab135013`) | 2.0568 G | | with oven-sh/mimalloc#40, #43, #42 (`3aa0ac9b`) | 2.0556 G (-0.06%) | | with #44, #45 (`58460467`) | 2.0305 G (-1.28%) | | this pin (`69050cf3`) | **2.0305 G (-1.28%)**, the same as `58460467` to within 400 instructions | | this pin with `MI_PROFILE=0` (not what bun builds) | 2.0244 G | (The absolute numbers are 8% below the ones this description had before: the benchmark's own object file is built differently here. All rows are from one recipe.) | 20 M `mi_malloc(128 KiB)` + `mi_free`, the generic path into a large page | instructions per pair | | --- | --- | | main's pin | 230.6 | | `58460467` | 193.6 | | this pin | 206.6 (+13.0 for the epoch stamp and its gate; -10% against main's pin) | **bun** (5 interleaved runs pinned to 8 cores; user-mode instructions and peak RSS, medians): | | base | pin | this PR | | --- | --- | --- | --- | | `bun -e 1` | 10.405 M, 27.0 MB | 10.410 M (+0.04%), 27.0 MB | 10.412 M (+0.06%), 27.0 MB | | `Bun.serve` hello, per request (20000 sequential `fetch`) | 52071, 49.0 MB | 51129 (-1.8%), 50.6 MB | 51361 (-1.4%; the spread between runs is 3%), 50.1 MB | | `Bun.serve` start and stop alone | 12.180 M | 12.179 M | 12.182 M (+0.01%) | | `JSON.stringify` + `JSON.parse` of a 200-user object, 3000 times | 5105.20 M, 44.0 MB | 5104.58 M, 44.0 MB | 5104.00 M (-0.02%), 43.5 MB | | string concat + `Buffer.from` + `split`, 300 x 5000 | 1900 M, 59.6 MB | 1912 M (+0.6%; the spread between its runs is 6.5%) , 60.5 MB | 1897 M (-0.2%), 59.7 MB | | idle after start (`Bun.sleep`), RSS / anonymous from `smaps_rollup`, binaries read into the page cache first | 28.42 MB / 3.14 | 28.45 MB / 3.05 | 28.45 MB / 3.05 | Nothing is slower by more than 0.1% outside the spread of its own runs. **The large-buffer server of oven-sh/mimalloc#46 at bun level** (`Bun.serve` that reads 256 KiB request bodies and answers 90 to 270 KB of JSON; 16 connections for 15 s, then idle; RSS / anonymous from `smaps_rollup`): | at idle | base | pin | this PR | | --- | --- | --- | --- | | 40 bodies kept, 2 s idle | 74.7 / 36.2 MB | 54.3 / 15.9 MB | 54.3 / 15.9 MB (**-20.4 MB**) | | 40 bodies kept, 15 s idle | 74.7 / 36.2 MB | 54.2 / 15.8 MB | 54.2 / 15.9 MB | | nothing kept, 15 s idle | 47.0 / 8.6 MB | 43.6 / 5.3 MB | 43.7 / 5.5 MB (-3.3 MB) | | at the end of the load, before the idle (40 kept) | 74.4 / 36.4 MB | 73.7 / 35.7 MB | 73.9 / 35.9 MB | Under sustained load (the same server, 20 s, 3 interleaved rounds; instructions and page faults of the server process from `perf stat -p`): | | requests | user instr / request | kernel instr / request | page faults / request | | --- | --- | --- | --- | --- | | base | 83.6 k .. 92.7 k | 283.9 k .. 284.5 k | 119.4 k .. 126.6 k | 0.34 .. 0.53 | | pin | 87.0 k .. 91.4 k | 284.0 k .. 284.1 k | 118.8 k .. 127.4 k | 0.50 .. 0.64 | | this PR | 80.4 k .. 91.4 k | 284.0 k .. 284.7 k | 120.9 k .. 125.3 k | 0.51 .. 0.86 | Requests and instructions per request are the same; the buffers stay while the server is busy and go when it is not. **Tests with the release build of this PR** - `test/js/bun/spawn/spawn-pipe-leak.test.ts` 25 of 25 runs, `test/js/bun/jsc/heapStats-mimalloc.test.ts` (with the new test) and `test/js/bun/gc/gc-controller-cadence.test.ts` 10 of 10 each. - 66 test files, 2165 tests pass: `require-cache`, `serve.test.ts`, all of `test/js/web/workers`, `test/js/node/worker_threads`, `test/js/node/child_process` and `test/js/node/fs` (the pool-heavy ones), `spawn-noread-leak`, `bun-file`, `bun-write`, `bundler_edgecase`, `esbuild/splitting`. The two failures need a non-root user (the root-range port in `serve.test.ts`, EPERM after dropping privileges in `child_process.test.ts`) and fail the same way with the base build. - The wider list of 113 files (bake dev, bundler, fetch, http, streams, sqlite, install, resolve, shell and others): 2485 tests pass; the two failures are in `resolve.test.ts`, want a directory that root cannot read, and fail the same way with the base build. - The new test fails with `USE_SYSTEM_BUN=1` (bun 1.4.1: 50 MB) and with the first commit alone (38 to 47 MB), and passes with both. - No debug or ASAN build was made for this revision of the PR (the earlier revisions ran `heapStats-mimalloc`, `gc-controller-cadence` and `worker.test.ts` on one at `58460467`); the new test skips itself under ASAN, where malloc is not mimalloc. **Not run here:** macOS and Windows. The Windows half of the second commit (`BunJSCEventLoop.cpp` under `OS(WINDOWS)`, `libuv.c`, `libuv.h`) was compiled by nothing on this machine: the C++ function was syntax-checked with `OS(WINDOWS)` forced on against stubs, `cargo check --target x86_64-pc-windows-msvc` passes for `bun_uws_sys`, `bun_threading` and `bun_mimalloc_sys` (the `WindowsLoop` mirror gained the field), and the rest is for CI. The fork's CI passed 23 of 23 jobs at each of the six merges (Linux x64/arm64 with ASAN, UBSAN, TSAN and guarded builds, the `MI_FREE_USE_PAGEMAP` build that matches bun's configuration, macOS, Windows x64/arm64, FreeBSD).
Jarred-Sumner
added a commit
to oven-sh/bun
that referenced
this pull request
Sep 15, 2026
…ollow-ups Stacked on the MIMALLOC_COMMIT bump to 64d2625acfed (oven-sh/mimalloc#40). - mimalloc reports the sampled blocks that mi_heap_destroy frees: the live-sample table and bun_alloc's heap_destroy_hook go, and a sampled block carries its bucket and weight itself. - The interval between samples is drawn from an exponential distribution. - sampleInterval can go down to 64 KiB.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Five changes to the sampled-profiling hooks (
mi_profiler_t) that Bun's heap profiler needs. None of them adds an instruction to themi_malloc/mi_freefast paths. Each comes with a test intest/test-profile.cthat fails atab135013.1. A small allocation is sampled when its own size class used up the countdown (
page.c)With coarse sampling (
MI_SAMPLE==1, whatMI_PROFILE=1gives) the fast path does not touch the sample countdown. The blocks it took from a page are counted when the page's free list is refilled in the generic path._mi_malloc_genericlooked at the countdown beforemi_page_queue_find_freedid that refill, so when the refill used up the countdown the sample went to the next allocation to come through the generic path. Every allocation aboveMI_SMALL_SIZE_MAXcomes through it, so those took nearly all the samples that small allocations had earned.test_profiler_small_attributionallocates 40 blocks of 64 bytes for each block of 1100 bytes (70% of the bytes are in the 64-byte blocks) and keeps them all. Share of the sampled bytes that went to the 64-byte blocks:ab135013,MI_PROFILE=1MI_PROFILE=1MI_PROFILE=2(exact countdown in the fast path)The change: after
mi_page_queue_find_free, if sampling is on and the countdown is used up, the allocation in progress is the sample (it has the size class of the blocks that were just counted).mi_page_extend_freenow counts the blocks of the previous extension too; before, a page that only grew (nothing freed) was first counted when it became full.What remains of the bias comes from the counting being late by up to one refill (4 KiB); it goes either way depending on the mix and is bounded by that.
Cost when no profiler runs: a load and a branch on
theap->sample_rateper call of_mi_malloc_genericthat finds a page. On the malloc/free benchmark of #37 (25 M mallocs, 25 M frees; clang 21-O3 -march=nehalem, as C++): 2.26451 G -> 2.26696 G user instructions (+0.11%).mi_mallocandmi_freeare byte-identical. (I tried folding it into the existing check in front ofmi_page_queue_find_free, which is one check instead of two: clang then keepspagelive across it and that costs more, +0.29%.)2.
bytes_since_last_samplestays exact whenon_allocchanges the rate (sample-profile.c)_mi_theap_malloc_sampledtakes the period that just ended to have beensample_ratebytes long. Whenon_allocreturned a higher rate,_mi_theap_set_profile_sample_rateleft the countdown at the old rate ("todo: adjust difference?"), so the next period was the old length but was reported with the new one: each sample reported the difference on top. With a rate that varies around a mean (which is how a profiler avoids sampling a periodic pattern always at the same allocation) the samples added up to 18% more than was allocated in the test, and up to 50% with an exponential distribution. Now the new period starts in full at a rate change. The same test checks that the samples add up to what was requested.3. A thread that starts while a profiler runs is sampled from its first allocation (
theap.c)A theap only noticed an enabled profiler in
mi_malloc_generic_admin, every 1000 generic allocations. A thread that makes a hundred large allocations and ends was never sampled; a new heap (an arena) on a thread lost its first ~4 MiB._mi_theap_initnow sets the profile rate when the heap's profiler is already enabled, and starts with a full period (the countdowns are 0 at that point: without that the first allocation of each new theap is a sample that reportsinitial_sample_ratebytes nobody requested).test_profiler_new_thread: 8 allocations of 1 MiB on a new thread, 8 samples or more (was 1).test_profiler_new_thread_no_phantom_period: 4 threads that allocate 32 bytes each report 0 bytes (2 MiB without the full period, withinitial_sample_rate= 512 KiB).4.
mi_heap_destroyreports the sampled blocks it frees (sample-profile.c,arena.c)mi_heap_destroyfrees pages without ami_freeper block, soon_freewas never called for the sampled blocks in them and a profiler counted them as in use forever. There is_mi_page_unguard_allat that spot for guarded blocks;_mi_page_profile_free_allnext to it walks the used blocks of a page that has the interior-pointer flag (only pages that ever held a sampled block do) and callson_freefor the profiled ones.mi_freetells a profiled block by its tag and by being handed an interior pointer. The walk has no pointer to go by, and the first word of a regular block in use is program data that can hold the value of the tag. So a profiled block now carries a check word behind the tag (its address xor a random key), which_mi_page_profile_on_freeclears (the free list only overwrites the tag, and the block can come back as a regular block).test_profiler_heap_destroyfills every block with{1, ...}and checks that exactly the sampled blocks are reported, once (those freed withmi_freebefore the destroy are not reported again).test_profiler_heap_destroy_recycledfrees 400 sampled blocks, takes the same blocks again as regular ones with only the first word written (1), and destroys the heap: nothing is reported for them (12 of them were, in a release build, before the word was cleared; a debug build fills freed memory and hides it).Layout of a profiled block:
[tag][check][mi_profiler_sample_data_t ...][user block]; one word more than before.Testing
Linux x64, the configurations of
.github/workflows/test.yamlwith their flags and environment: Release, DebugMI_DEBUG=FULL, Secure, Debug and Release as C++, ASAN, UBSAN, TSAN, Debug and ReleaseMI_GUARDED=FULL(with the job'sMIMALLOC_GUARDED_SAMPLE_RATE), Debug and ReleaseMI_FREE_USE_PAGEMAP=ON(what Bun builds with), Debug and ReleaseMI_SECURE=FULL, TLS pthreads, SIMD, and ReleaseMI_PROFILE=OFF.ctestis 100% in each excepttest-emulated-tlsin the secure and guarded ones (see below; the same without this branch).src/static.ccompiles as C++20 with Bun's flags. Atab135013the new tests fail (shares and counts above).5.
on_freegets the pointer thaton_allocgot, also for an aligned allocation (sample-profile.c)mi_malloc_alignedaligns inside the block that_mi_theap_malloc_profiledreturned, somi_freeof a sampled aligned block calledon_freewith the aligned pointer and not the oneon_allocwas given.test-profile.c's ownon_freeasserts they are the same; it only never met a sampled aligned block. It does with guarded sampling on in the same process (the CI'sDebug, Guardedjob setsMIMALLOC_GUARDED_SAMPLE_RATEfor the whole job): samples then land on mimalloc's ownmi_heap_zalloc_alignedof the per-heap arena page table, whichmi_heap_destroyfrees._mi_page_profile_on_freenow computes the pointer from the block's layout. It also checks the block's check word before it callson_free(it did only in themi_heap_destroywalk), so a block that merely starts with the tag's value and is freed through an interior pointer is not reported.test_profiler_guarded_mixed: 6000 blocks of five sizes in a new heap, a third of them with an alignment (64, 1 KiB, 16 KiB), with guarded sampling on in guarded builds (it sets the rate itself; 1333 of the 6000 are guarded, 2801 sampled); a quarter is freed on this thread, a quarter on another thread, a quarter moved bymi_heap_realloc, the rest destroyed with the heap. Every sampled block is reported once with the pointeron_allocgot. Without this change it fails (the assert in debug builds, a wrong pointer in release builds).Tests and guarded sampling
CMakeLists.txtdeliberately gavetest-profileno guarded sample rate, but the CI's guarded jobs set one in the environment for every test. With guarded samples in between, what the profiler's samples add up to is off atab135013already (34 of 70 MiB, share 0.88 for 0.70), sotest-profilenow getsMIMALLOC_GUARDED_SAMPLE_RATE=0explicitly, liketest-purge-holes; the case above turns guarded sampling on for itself.Seen on the way, not changed
_mi_theap_malloc_sampledsubtracts the bytes since the last profile sample from both derived countdowns at every combined sample, so a profile sample comes early and the reported bytes do not add up (the numbers above). Bun does not build withMI_GUARDED.test-emulated-tlsfails here in theMI_SECUREandMI_GUARDEDconfigurations with and without this branch ("unexpected block layout": it expects two consecutive small allocations to be adjacent); it passes in CI.