Evict code a parked process does not use: idle by allocation, module graph at 2 min, the executable's own pages and rebuildable bytecode at 10 min - #42383
Conversation
|
Updated 7:45 AM PT - Sep 16th, 2026
❌ @Jarred-Sumner, your commit 3dbaba3 has 1 failures in
🧪 To try this PR locally: bunx bun-pr 42383That installs a local version of the PR into your bun-42383 --bun |
c0022f8 to
fc1d3b3
Compare
9e77b38 to
2faeb57
Compare
fc1d3b3 to
ef12a83
Compare
ef12a83 to
f3b5196
Compare
8d35e8a to
83602bb
Compare
f3b5196 to
945f762
Compare
There was a problem hiding this comment.
Findings marked 🟡 are optional suggestions and need no follow-up push.
Still open from earlier reviews (1):
- Unresolved: 1 blocking on lines changed since (possibly already fixed).
If you have decided not to act on one of these findings, resolve its thread (a reply alone leaves it open) and the next review stops counting it. To review this commit again now, use Re-run on its "Claude Code Review" check.
…s everything at once The ladder is now: when a tick last saw the program busy (`last_busy_at`), and how many rungs have run since (`idle_rungs_done`). A tick that sees the program busy stamps the time and starts over; any other tick runs the next rung if the program has been idle that long, at most one per tick, and the timer is re-armed for the earlier of its normal interval and the next rung. The tick uses the clock reading the timer dispatcher already took. - Busy is what it was (the heap grew by more than the slack), or a tick that comes more than two seconds after it was due: the thread was in a synchronous call, or the process was stopped or asleep, and that time was not spent idle. That is main's definition and it cannot tell a small server from a parked program: a server's heap does not grow, Bun.serve answering 20 requests a second walks the whole ladder. The definition that can (what the program allocates) needs a counter that the WebKit upgrade of #42383 brings, which this should land with. - Rung 1: an idle full collection, requested. Rung 2 of three: the same, and the page-out of a standalone executable's module graph. - The last rung (the only one of a list of one): the collection, and the page-out of the module graph and of the executable's own code and constants right after it (Linux). The collection is synchronous when a page-out follows, because the collector is the last thing that would read the executable back in, and requested like the others when none does. The rung waits for a tick without JS on the stack. - Nothing is paged out while a Worker is alive: the page-outs are for the whole process and the ladder only watches the main thread. And a Worker's own VM no longer runs the ladder: it was initialised before it knew it was a Worker's. - An entry of BUN_IDLE_GC_SECONDS that is not a positive decimal number turns the ladder off instead of ending the list (which made "1,x" a list of one); a number too large to read is the hour that is the maximum. Gone: the quiet-time accumulator, the last rung coming round again, the tick after the last rung that paged out if the process had used under 3% of a CPU (quiet_since_cpu_ms, page_out_due, process_cpu_time_ms). What that gives up: what a parked program's own later jobs read back in stays until it has been busy and idle again, and a process that is working in native code on its main thread at the last rung has the pages it is executing read back from the page cache, once. Tests: the child runs from a file (-e runs with one GC marker thread, with which a rung's concurrent collection stays open and the next rung's request is folded into it), stamps its stderr with its own clock, and exits when the parent has seen the rungs it waits for or at its own deadline, after an EXIT stamp that keeps the teardown's collection (BUN_DESTRUCT_VM_ON_EXIT) out of the count; it must exit cleanly. Lists of one, two and three run that many collections a second apart, and none after the last; malformed lists and "0" run none; a program whose heap keeps growing gets none; 2.5 s in a synchronous call are not idle time; a 10 ms timer does not keep a program from going idle. One standalone executable is built for the page-out tests and they run it one at a time (with a copy each, side by side, the children that had started did not get to run while the next copies were written, on a slow disk): what counts is the resident size of the executable's own mappings in /proc/self/smaps, not RssFile, which the kernel can lower by itself. The module graph goes with the second rung of three and the code does not; the code goes with the last rung only; neither goes after "1,99999999999", while a Worker is alive, or with BUN_FEATURE_FLAG_DISABLE_STANDALONE_MADVISE=1. The FTL-aging test reports as soon as the code is gone. The file takes 19 s (was 23 s). The file is taken out of the parallel test batch (test/parallel-denylist.txt, and excludeFiles in the allowlist as the generator would write it): it measures seconds of idleness, which a loaded batch stretches.
83602bb to
3f655c5
Compare
e728444 to
e8c023a
Compare
There was a problem hiding this comment.
Code review found no new issues
No new issues were found in this update; 1 finding from earlier reviews is still open above.
Still open from earlier reviews (1):
- Unresolved: 1 blocking on lines changed since (possibly already fixed).
If you have decided not to act on one of these findings, resolve its thread (a reply alone leaves it open) and the next review stops counting it. To review this commit again now, use Re-run on its "Claude Code Review" check.
… run it The idle full collections (BUN_IDLE_GC_SECONDS) become a ladder that keeps two facts: when a tick last saw the program busy (`last_busy_at`), and how many rungs have run since (`idle_rungs_done`). A tick that sees the program busy stamps the time and starts over; any other tick runs the next rung if the program has been idle that long, at most one per tick, and the timer is re-armed for the earlier of its normal interval and the next rung. The tick uses the clock reading the timer dispatcher already took. What runs is what ran before: an idle full collection per rung, and the page-out of a standalone executable's module graph with the second rung, or the only one. This changes when they run, not what they do. - Busy is what it was (the heap grew by more than the slack), or a tick that comes more than two seconds after it was due: the thread was in a synchronous call, or the process was stopped or asleep, and that time was not spent idle (Bun.spawnSync(["sleep", "8"]) used to run all three rungs the moment it returned). As on main, that cannot tell a small server from a parked program: a server's heap does not grow. #42383 replaces the heap-growth term by what the program allocates, and only then adds the step that would hurt a working server, the page-out of the executable image. - The defaults go from "10,65,65" (10 s, 75 s, 140 s) to "30,90,480" (30 s, 2 min, 10 min): someone who reads output for a few minutes should come back to a process that is still warm. - A Worker's own VM no longer runs the ladder (its idle full collections and the module-graph page-out): init asked is_main_thread(), which is "has no worker", before the Worker's VM was given one; on main too. And the module graph is not paged out while a Worker is alive: the page-out is for the whole process and the ladder only watches the main thread. - An entry of BUN_IDLE_GC_SECONDS that is not a positive decimal number turns the ladder off instead of ending the list (which made "1,x" a list of one); a number too large to read is the hour that is the maximum. Gone: the quiet-time accumulator that counted nominal intervals rather than time. Tests: the child runs from a file (-e runs with one GC marker thread, with which a rung's concurrent collection stays open and the next rung's request is folded into it), stamps its stderr with its own clock, and exits when the parent has seen the rungs it waits for or at its own deadline, after an EXIT stamp that keeps the teardown's collection (BUN_DESTRUCT_VM_ON_EXIT) out of the count; it must exit cleanly. Lists of one, two and three run that many collections a second apart, and none after the last; malformed lists and "0" run none; a program whose heap keeps growing gets none; 2.5 s in a synchronous call are not idle time; a 10 ms timer does not keep a program from going idle. One standalone executable is built for the page-out tests and they run it one at a time (with a copy each, side by side, the children that had started did not get to run while the next copies were written, on a slow disk): what counts is the resident size of the executable's own mappings in /proc/self/smaps, not RssFile, which the kernel can lower by itself. The module graph goes with the second rung of three and of two and with the only rung of one, the code never; the graph stays after "1,99999999999", while a Worker is alive, and with BUN_FEATURE_FLAG_DISABLE_STANDALONE_MADVISE=1. The FTL-aging test reports as soon as the code is gone. The file takes 19 s (was 23 s). The file is taken out of the parallel test batch (test/parallel-denylist.txt, and excludeFiles in the allowlist as the generator would write it): it measures seconds of idleness, which a loaded batch stretches.
3f655c5 to
6e6434f
Compare
… run it The idle full collections (BUN_IDLE_GC_SECONDS) become a ladder that keeps two facts: when a tick last saw the program busy (`last_busy_at`), and how many rungs have run since (`idle_rungs_done`). A tick that sees the program busy stamps the time and starts over; any other tick runs the next rung if the program has been idle that long, at most one per tick, and the timer is re-armed for the earlier of its normal interval and the next rung. The tick uses the clock reading the timer dispatcher already took. What runs is what ran before: an idle full collection per rung, and the page-out of a standalone executable's module graph with the second rung, or the only one. This changes when they run, not what they do. - Busy is what it was (the heap grew by more than the slack), or a tick that comes more than two seconds after it was due: the thread was in a synchronous call, or the process was stopped or asleep, and that time was not spent idle (Bun.spawnSync(["sleep", "8"]) used to run all three rungs the moment it returned). As on main, that cannot tell a small server from a parked program: a server's heap does not grow. #42383 replaces the heap-growth term by what the program allocates, and only then adds the step that would hurt a working server, the page-out of the executable image. - The defaults go from "10,65,65" (10 s, 75 s, 140 s) to "30,90,480" (30 s, 2 min, 10 min): someone who reads output for a few minutes should come back to a process that is still warm. - A Worker's own VM no longer runs the ladder (its idle full collections and the module-graph page-out): init asked is_main_thread(), which is "has no worker", before the Worker's VM was given one; on main too. And the module graph is not paged out while a Worker is alive: the page-out is for the whole process and the ladder only watches the main thread. - An entry of BUN_IDLE_GC_SECONDS that is not a positive decimal number turns the ladder off instead of ending the list (which made "1,x" a list of one); a number too large to read is the hour that is the maximum. Gone: the quiet-time accumulator that counted nominal intervals rather than time. Tests: the child runs from a file (-e runs with one GC marker thread, with which a rung's concurrent collection stays open and the next rung's request is folded into it), stamps its stderr with its own clock, and exits when the parent has seen the rungs it waits for or at its own deadline, after an EXIT stamp that keeps the teardown's collection (BUN_DESTRUCT_VM_ON_EXIT) out of the count; it must exit cleanly. Lists of one, two and three run that many collections a second apart, and none after the last; malformed lists and "0" run none; a program whose heap keeps growing gets none; 2.5 s in a synchronous call are not idle time; a 10 ms timer does not keep a program from going idle. One standalone executable is built for the page-out tests and they run it one at a time (with a copy each, side by side, the children that had started did not get to run while the next copies were written, on a slow disk): what counts is the resident size of the executable's own mappings in /proc/self/smaps, not RssFile, which the kernel can lower by itself. The module graph goes with the second rung of three and of two and with the only rung of one, the code never; the graph stays after "1,99999999999", while a Worker is alive, and with BUN_FEATURE_FLAG_DISABLE_STANDALONE_MADVISE=1. The FTL-aging test reports as soon as the code is gone. The file takes 19 s (was 23 s). The file is taken out of the parallel test batch (test/parallel-denylist.txt, and excludeFiles in the allowlist as the generator would write it): it measures seconds of idleness, which a loaded batch stretches.
6e6434f to
9d1dcfa
Compare
e8c023a to
29c333a
Compare
29c333a to
c471002
Compare
There was a problem hiding this comment.
This pull request has been reviewed before and this review found new issues. Where they share a root cause, one fix may close them together.
Still open from earlier reviews (1):
- Unresolved: 1 blocking on lines changed since (possibly already fixed).
If you have decided not to act on one of these findings, resolve its thread (a reply alone leaves it open) and the next review stops counting it. To review this commit again now, use Re-run on its "Claude Code Review" check.
c471002 to
35fff22
Compare
There was a problem hiding this comment.
Code review found no new issues
No new issues were found in this update; 2 findings from earlier reviews are still open above.
2 optional suggestions (nits or notes on pre-existing code) were found and not posted.
Still open from earlier reviews (2):
- 🔴
src/jsc/GarbageCollectionController.rs:267—Users of a parked standalone app that does one short job get the last rung's sync full GC, bytecode drop and executable… - Also unresolved: 1 blocking on lines changed since (possibly already fixed).
If you have decided not to act on one of these findings, resolve its thread (a reply alone leaves it open) and the next review stops counting it. To review this commit again now, use Re-run on its "Claude Code Review" check.
…at rest After the second idle full collection the GC controller already pages out a standalone executable's embedded module graph. An idle process keeps tens of MB of its own text and read-only data resident as well, almost none of which it touches while parked. On Linux, once the heap has stopped growing for the second BUN_IDLE_GC_SECONDS threshold, the controller now also asks the kernel to reclaim the resident pages of the program's read-only PT_LOAD segments (MADV_PAGEOUT: clean file-backed pages, read back from the file when touched again). Because a flat heap does not prove a program is idle, this only happens when the process used at most 3% of one CPU (all threads) both since the heap last grew and over the last timer tick; it is evaluated on ticks that run no idle collection, retried every tick until it holds, and not repeated within 10 minutes. Heap growth resets it. bun_sys::page_out_range is the one place that issues the madvise (4 MB per call so mmap_lock is not held across a whole segment) and honours BUN_FEATURE_FLAG_DISABLE_STANDALONE_MADVISE; the module graph page-out uses it too. bun_sys::elf::page_out_program_image finds the image with the same dl_iterate_phdr walk as find_loaded_module and issues the calls after the walk returns, outside the loader lock. bun_core::time::process_cpu_time_ms is new. Idle `bun -e` on Linux x64: RssFile 27 MB -> 10 MB three seconds in with BUN_IDLE_GC_SECONDS=1,1; a compiled standalone app 24 MB -> 13 MB. The kernel skips pages that another process maps too, and everything if the caller neither owns nor can write the file, so those cases are unchanged.
… second With the default BUN_IDLE_GC_SECONDS=10,65,65 the image page-out was armed when the second threshold was crossed, ran on the next quiet tick (~105 s), and the third idle collection at ~140 s read the collector's code straight back in with nothing left to page it out again. Arm it when the last threshold is crossed; the module graph's page-out stays with the second collection as before. Schedules on a release build, time at which file-backed resident memory drops: "1,1" 3.0 s, "1,1,1" 4.1 s, "1,2,2" (collections at 1/3/5 s) 6.1 s, "1,3,3" (1/4/7 s) 8.0 s. New test with collections at 1, 2 and 4 s: the drop must come after 4 s. With the arming reverted it comes at 3.1 s and the test fails. The test environment also pins BUN_FEATURE_FLAG_DISABLE_STANDALONE_MADVISE to unset unless a case sets it.
The follow-up that split the module graph's page-out from the image's left both
on the second threshold; the image's is the last one ("1,1,2": armed at 4 s, not
2 s), which is what its test expects.
clippy::double_ended_iterator_last: the filtered thresholds are a DoubleEndedIterator, so last() walks all of them for the same element.
The idle work is a ladder of rungs from BUN_IDLE_GC_SECONDS, with one quiet clock: every rung is an idle full collection; the second (or only) one also pages out a standalone executable's module graph; a quiet tick after the last one pages out the executable's own code and constants, if the process (all threads) has used under 3% of a CPU since the heap went quiet. The last rung comes round again every time the quiet has lasted that long once more, which releases what a parked program's occasional jobs read back in. This replaces IdleImagePageOut (CpuSample, quiet_since, pending, the 10-minute minimum interval) with two fields. The default list goes from "10,65,65" to "30,90,480": a collection after 30 s, the module graph after 2 minutes, the executable after 10. Someone who reads output or thinks for a few minutes comes back to a process that is still warm; the sessions that matter for memory sit idle for tens of minutes to hours.
… run it The idle full collections (BUN_IDLE_GC_SECONDS) become a ladder that keeps two facts: when a tick last saw the program busy (`last_busy_at`), and how many rungs have run since (`idle_rungs_done`). A tick that sees the program busy stamps the time and starts over; any other tick runs the next rung if the program has been idle that long, at most one per tick, and the timer is re-armed for the earlier of its normal interval and the next rung. The tick uses the clock reading the timer dispatcher already took. What runs is what ran before: an idle full collection per rung, and the page-out of a standalone executable's module graph with the second rung, or the only one. This changes when they run, not what they do. - Busy is what it was (the heap grew by more than the slack), or a tick that comes more than two seconds after it was due: the thread was in a synchronous call, or the process was stopped or asleep, and that time was not spent idle (Bun.spawnSync(["sleep", "8"]) used to run all three rungs the moment it returned). As on main, that cannot tell a small server from a parked program: a server's heap does not grow. #42383 replaces the heap-growth term by what the program allocates, and only then adds the step that would hurt a working server, the page-out of the executable image. - The defaults go from "10,65,65" (10 s, 75 s, 140 s) to "30,90,480" (30 s, 2 min, 10 min): someone who reads output for a few minutes should come back to a process that is still warm. - A Worker's own VM no longer runs the ladder (its idle full collections and the module-graph page-out): init asked is_main_thread(), which is "has no worker", before the Worker's VM was given one; on main too. And the module graph is not paged out while a Worker is alive: the page-out is for the whole process and the ladder only watches the main thread. - An entry of BUN_IDLE_GC_SECONDS that is not a positive decimal number turns the ladder off instead of ending the list (which made "1,x" a list of one); a number too large to read is the hour that is the maximum. Gone: the quiet-time accumulator that counted nominal intervals rather than time. Tests: the child runs from a file (-e runs with one GC marker thread, with which a rung's concurrent collection stays open and the next rung's request is folded into it), stamps its stderr with its own clock, and exits when the parent has seen the rungs it waits for or at its own deadline, after an EXIT stamp that keeps the teardown's collection (BUN_DESTRUCT_VM_ON_EXIT) out of the count; it must exit cleanly. Lists of one, two and three run that many collections a second apart, and none after the last; malformed lists and "0" run none; a program whose heap keeps growing gets none; 2.5 s in a synchronous call are not idle time; a 10 ms timer does not keep a program from going idle. One standalone executable is built for the page-out tests and they run it one at a time (with a copy each, side by side, the children that had started did not get to run while the next copies were written, on a slow disk): what counts is the resident size of the executable's own mappings in /proc/self/smaps, not RssFile, which the kernel can lower by itself. The module graph goes with the second rung of three and of two and with the only rung of one, the code never; the graph stays after "1,99999999999", while a Worker is alive, and with BUN_FEATURE_FLAG_DISABLE_STANDALONE_MADVISE=1. The FTL-aging test reports as soon as the code is gone. The file takes 19 s (was 23 s). The file is taken out of the parallel test batch (test/parallel-denylist.txt, and excludeFiles in the allowlist as the generator would write it): it measures seconds of idleness, which a loaded batch stretches.
…e-out, not the rung The rung that pages out a standalone executable's module graph counted as run without paging anything out while a Worker was alive, so once the Worker had gone the graph stayed resident until the program had been busy and idle again. The rung still runs when it is due (its collection, and the ones after it, are for this thread's heap and have nothing to do with Workers), and what it could not do is owed: the first tick that finds no Worker pages the graph out. A busy tick forgets the debt with the rest of the ladder. Tests: with a Worker alive for the whole ladder all three collections run and the graph stays; with one that is terminated after 2 s the graph is resident until then, although the rung was due after 1 s, and gone right after. The second fails without the change.
…he first rung is back at 10 s A rung requests its collection; the collection then proceeds at the mutator's safepoints, and in a program that runs no JS those are this timer's ticks. The first rung had moved from 10 s to 30 s, which is also when the timer goes to its 30 s tick (30 ticks without heap growth): the collection started on time, ended on the next tick half a minute later, and what it freed went back to the system on the tick after that. A server with 800 MB live that had just served 2,000 requests a second held 2.36 GB until 92 s after the traffic stopped; main, whose first rung is at 10 s, is back at 0.8 GB after 13 s. - After a rung the timer is on its fast tick for the next 30 ticks, whatever the heap does. - The default list is "10,110,480": 10 s, 2 min, 10 min. The reason for 30 s was that heap growth is a coarse signal for a rung that early; main has always had it at 10 s, and a server that has gone quiet should not keep a burst's garbage for half a minute. The same server on this change: 2.39 GB at +10 s, 0.81 GB at +15 s (first below 1.2 GB at +13 s and +14 s in two runs; main at +13 s). Test: a child with 50 MB alive and 300 MB that have survived a full collection and that nothing refers to any more, on a 20 ms tick (slow after 0.6 s) with the rung 2 s in: its RSS is down by 250 MB within 6 s. Without the change it still holds 405 MB then.
…when its heap has stopped growing The idle ladder (BUN_IDLE_GC_SECONDS) starts over whenever a tick sees the program busy, and busy was "JSC's block bytes plus extra memory are more than 2 MB above the previous tick's sample". Block bytes do not say what the program is doing: - A small server's heap does not grow. Bun.serve answering 20 requests a second walks the whole ladder. - A program that works hard but recycles the same blocks has a flat footprint. An interactive CLI app allocates 10-50 MB a second while it works on a request and read as quiet on 5 of those 13 ticks. - The fast/slow switch compared two samples for equality, so a few KB of timer chatter kept the timer at one tick a second for ever. A tick is now loud when the program allocated, since the last tick (Heap::totalBytesAllocated(), the cycle in progress included, so there is no lag of a collection), more than 1/64 per second of what the collector lets it allocate before it collects by itself (Heap::allocationBudgetThisCycle(), which follows the size of the heap: 128 KB a second with the 8 MB of a new heap, 1.8 MB a second with the 117 MB of a large CLI application at its prompt), or when it came more than two seconds late (#42329's rule). That is the one signal: the ladder's rungs run their entries after the last loud tick, a loud tick puts the timer on its 1 s tick, and 30 ticks in a row that are not loud put it on the 30 s one. The first tick sees what starting up allocated and is loud. What runs, and the collection the timer requests on every tick, are what they were. Measured: that application parked at its prompt allocates under 1 KB a second, and 60 KB a second in the half minute of its own job ten minutes after a session; a server with a 40 MB working set answering 50 requests a second, 27 MB a second against a budget of 153 MB. A server answering "hello" five times a second allocates 4 KB a second and is not loud: it goes down the ladder like a parked program. On the 30 s tick the event loop puts the timer back on the fast one as soon as the program has allocated 8 MB since the last tick and the tick so far would be a loud one (so a large program that allocates steadily under its loud line stays on the 30 s tick), and that counts as a loud tick of its own (the ladder starts over, the next tick measures from there): work that starts then would otherwise run for up to half a minute without the collections requested every second, and the application above peaked 50-65 MB higher after a pause. process.memoryUsage().heapUsed reads Heap::sizeAfterLastCollection() directly; the HeapObserver that reconstructed it, gc_last_heap_size and JSC__VM__blockBytesAllocated are gone. Tests, on the default 1 s ticks (on 100 ms ones a single free list of 16 KB is a loud tick): lists of one, two and three run that many collections a second apart and none after; programs that allocate 5 MB and 40 MB a second on a flat heap get none; a 10 ms timer and a timer that builds a spinner frame every 80 ms do not keep a program from going idle; 4 s in a synchronous call are not idle time; a burst on the 30 s tick brings the fast tick back, after half a second in which nothing was requested. The page-out tests look at /proc/<pid>/smaps from the test (reading it in the child is a loud tick), tell the child to exit before its deadline (the mappings of an exiting process are no measure), and run a child again when the mapping was gone before any rung could have run. A Worker that is alive puts off the module graph's page-out, not the collections (#42329). After a rung the timer is on its fast tick for 30 ticks, so that the requested collection has safepoints to finish on (#42329).
…t its own code and drops what JSC can decode again
An idle Bun process keeps tens of MB of its own code and read-only data
resident, almost none of which it touches while parked. In a standalone
executable (bun build --compile: the applications this is for) the last rung
of the idle ladder (10 minutes after the last loud tick by default; the only
one of a list of one) now does three things in one step; any other process
(bun file.js) gets the requested collection it got before and keeps its code:
- JSC lets go of its parser caches and of the unlinked bytecode of functions
that have no linked code any more (the earlier collections unlinked what had
not run) and that can be decoded again from a --compile --bytecode executable
(VM::shrinkFootprintNow, KeepCodeInUse, LeaveCollectionToCaller). Nothing that
would have to be parsed again is dropped. It declines with JS on the stack
(the timer fired from a nested event loop): the rung then waits for a tick
that has none.
- The rung's idle full collection, which frees that. It is synchronous when a
page-out follows, because the collector is the last thing that would read the
executable back in, and requested like the others when none does (not Linux,
BUN_FEATURE_FLAG_DISABLE_STANDALONE_MADVISE).
- The page-out of the module graph and of the executable image (Linux):
MADV_PAGEOUT over the program's read-only PT_LOAD segments, clean file-backed
pages that are read back from the file when touched.
This comes after the change of what busy means because it is the step that
hurts a program that is working: under the heap-growth rule a server under load
was idle, and would have had the code it serves from paged out. A server that
allocates less than a loud program ("hello" five times a second) still is:
measured, the ten requests after the rung take the same user instructions as
the ten before and 55-113 major plus 130-190 minor faults, once, 3.5-4 MB of
text read back; the slowest request of the run is 70 ms against a median of
0.3 ms.
bun_sys::page_out_range is the one place that issues the madvise;
BUN_FEATURE_FLAG_DISABLE_STANDALONE_MADVISE is read by the controller, once;
StandaloneModuleGraph::page_out
uses it too. bun_sys::elf::page_out_program_image finds the image with
find_loaded_module, which now also returns the program headers, and issues the
calls after the walk has returned, outside the loader lock. The kernel skips pages that
another process maps too (a second instance of the same executable), and
everything if the caller neither owns nor can write the file. While a Worker is
alive the rung runs and its page-outs are owed, as the module graph's is (#42329);
its collection is only requested then.
What the rung drops is decoded again when it next runs (0.7 MB of allocation in
the half minute after it in a large CLI application, far from loud). What a parked program's own
later jobs read back in stays until it has been busy and parked again, and a
process that is working in native code on its main thread at that moment has
the pages it is executing read back from the page cache, once.
Tests (Linux release lanes, disk-backed temp dir): the executable run as
plain bun (BUN_BE_BUN=1, so that nothing else maps it) keeps its code; the
module graph goes with
the second rung of three and the code does not; the code goes with the last
rung of two and with a list of one; neither goes after "1,99999999999", while a
Worker is alive (both go once it has gone, and the collections run meanwhile), or with
BUN_FEATURE_FLAG_DISABLE_STANDALONE_MADVISE=1. On a
--compile --bytecode executable the last rung of "1,1" drops the re-decodable
unlinked code and the results are the same afterwards; the first of "1,30" does
not.
6600f22 to
10b7d52
Compare
10b7d52 to
0102f81
Compare
There was a problem hiding this comment.
1 optional suggestion (a nit or a note on pre-existing code) was found and not posted.
2 verified lower-impact observations (convention, logging or cleanup points) were not posted.
Still open from earlier reviews (2):
- 🔴
src/jsc/GarbageCollectionController.rs:340—A compiled server that had a busy hour and then serves modest steady traffic still gets the last rung's stop-the-world… - 🔴
src/jsc/GarbageCollectionController.rs:368—A parked standalone server whose JS heap is big enough that the last rung's synchronous full collection takes over 3 s…
If you have decided not to act on one of these findings, resolve its thread (a reply alone leaves it open) and the next review stops counting it. To review this commit again now, use Re-run on its "Claude Code Review" check.
0102f81 to
075686e
Compare
1281124 to
621bacb
Compare
…ates when it works A tick was loud when the program allocated more than 1/64 per second of the collector's budget for the cycle. The budget follows the size of the heap, so a server with a big heap and a modest allocation rate (300 MB live, 2 MB a second, all the time) read as parked and would have had a synchronous collection, its bytecode dropped and its executable paged out while it served. No fixed line helps either: a server answering "hello" five times a second allocates 4 KB a second, less than the 60 KB a second of the background job of a large CLI application that is parked. Only the program's own history tells them apart. - working_rate is the highest rate seen between two looks at the program, fading over the ladder's length (the last rung's time, 10 minutes by default) until the ladder has run to its end: what a program went all the way idle against stays its yardstick until something is loud against it. (Left to fade, the application's job would have crossed the line 17 minutes in and the first rung would have repeated every minute or so from there.) - The program is loud when its rate is more than working_rate / 16, and never for less than 2 KB a second, or when a tick came more than two seconds late. The collector's budget is no longer read: with the floor at the 2 KB a second it came to for a new heap, every measured program is classified as before. - A look is a tick. On the 30 s tick the next one is brought forward once the program has allocated 8 MB, to a second after the last one at the earliest: a burst in the middle of half a minute is judged by its own rate, whatever the line, also when the program parks right after it (the look comes from the timer, not from an event loop that may have nothing to turn for), and a steady rate reads the same on either tick. If the look is not loud the timer goes back to where it was. This replaces judging a window as if it had been 5 s long, under which a steady rate read six times higher on the slow tick than on the fast one. - The first tick and the tick after a rung do not judge what was allocated (starting up; the rung's doing: freed blocks handed out again, and the code the last rung dropped decoded again, 0.7 MB in that application) and run no rung. One flag instead of two sentinel values of the byte count. - A tick reads the clock itself. The dispatcher's reading is from before the callbacks it ran in the same drain: a timer of the program's that sat in a synchronous call for seconds ahead of the tick left it "on time". The tick is late by that reading like any other when it comes late, the first one and the one after a rung included; the timer is armed from the clock again at the end, so the time a rung took is not time the next tick is late by. - The time a tick is late by (a synchronous call, the process stopped, the machine asleep) does not fade working_rate: it fades up to when the tick was due. A pause longer than the ladder took the whole yardstick away, and the program's background job was its normal rate from then on: loud, for good. - With a debugger attached nothing is due, as without a ladder: the timer goes to its 30 s tick instead of firing every second for a rung that never runs. The consequence, by design: a program with a modest steady allocation at its own normal rate is running, not idle. Measured (release build): hello at 5 and at 50 requests a second (3.7 and 36 KB/s) and 300 MB live at 2 MB/s, loud on every tick under traffic; a server with 800 MB live under 2,000 requests a second, 2.38 GB -> 0.80 GB 13 s after the load (main: 14 s in the same session); the application parked at its prompt, rungs at 12 s, 2 min and 10 min and nothing after the last; after a session, rungs 12 s, 2 min and 10 min after it, its own job 64 KB/s against a line of 2.0 MB/s, and in a 40-minute run not one loud tick after the session, the timer on its 30 s tick from 30 s after the last rung. Tests: the children are observed without allocating (stamps written from a buffer they keep; the FTL and bytecode children look at the heap only once their rungs have had their time). On a 20 ms tick with lists of seconds: after a second at 80 MB a second, 0.5 MB a second goes down the ladder and 16 MB a second does not; a job that comes and goes after the last rung does not start the ladder over (it does if the yardstick goes on fading); 2 MB a second on the slow tick, looked at every 8 MB, gets its rung when it is due; a second of work in the middle of a slow tick starts the ladder over. On the default tick: rungs a second apart in the list run two seconds apart (the tick between them only looks; they are a second apart if it runs rungs); a server answering five requests a second reaches no rung, and the same server without requests does; a standalone executable with 300 MB live that allocates 2 MB a second gets no collection from the controller and keeps its code; 3 s in a synchronous call that make the tick after a rung late start the ladder over, and so do 3 s in a call made from a timer in the same turn of the event loop as the tick; a process stopped for 3 s, longer than its ladder, still has both rungs with 1/200 of its peak rate afterwards; a burst that is over within a second of the last tick, in a program that parks, brings the fast tick back; with --inspect the timer reaches its slow tick. And on the 20 ms tick: work at 320 MB a second that comes back while the timer is slow is loud within a second against a line of 20 MB a second.
621bacb to
3dbaba3
Compare
| let now = &Timespec::now(TimespecMockMode::ForceRealTime); | ||
| let late = now.duration(&due).ms_unsigned() > Self::LATE_TICK_MS; | ||
| let judged = !this.skip_next_tick.replace(false); | ||
| let rate = this.sample(jsc, now, &due, judged); | ||
| let loud = late || (judged && this.is_loud(rate)); |
There was a problem hiding this comment.
🔴 Operators of a compiled server whose traffic is served natively (static routes, HTML bundles, Bun.file responses) get its whole executable paged out and a sync full GC under live load, which the base never does. Loudness at GarbageCollectionController.rs:345 comes only from total_bytes_allocated() (:266), and StaticRoute.rs:318-333 answers requests without touching the JS heap, so a busy static server reads as parked and reaches the last rung (:186-201) after 600 s. Fix: count native work on the JS thread as loud too (e.g. requests completed via on_pending_request, or I/O completions since the last tick), so a thread that is serving never runs the eviction rung. [also at: src/jsc/GarbageCollectionController.rs:284 - Operators of a compiled server whose traffic never enters JS (static routes, Bun.file/directory routes, HTML bundles) get its code paged out and a synchronous full GC while it serves, which the base never did.]
Extended reasoning...
The PR description says a program that runs no JS is parked and goes down the ladder on schedule; a bun build --compile server with routes: { "/": new Response(...) }, an HTML import route, or Bun.file() responses serves thousands of requests a second with zero JS allocation, so that note does not hold for it.
StaticRoute::on (src/runtime/server/StaticRoute.rs:318-333) does on_pending_request, do_render_blob, and the uWS write; no JS runs and Heap::totalBytesAllocated() does not move.
Every tick: sample at :263-268 reads 0 bytes, is_loud at :281 is false, late is false because the loop is never blocked; silent_ticks reaches 30 and the timer goes to 30 s ticks (:98-103).
next_rung_in_ms at :210-220 hits Some(0) at 10 s, 120 s and 600 s of "idle" since last_loud_at.
At 600 s idle_tick: evict is true (:186, standalone graph present), shrink_footprint_now (:187) drops re-decodable bytecode, collect_idle(true)…
Verification: normal (bounded: one synchronous full collection, bytecode drop and MADV_PAGEOUT of the executable's own read-only segments on a serving process, then major-fault re-reads of its hot native HTTP code under load; once per ladder, repeated only if a tick is later than 2 s) — triggered by any Linux bun build --compile server whose traffic is answered natively (StaticRoute / FileRoute / HTML…
| fn is_loud(&self, rate: u64) -> bool { | ||
| rate > (self.working_rate.get() / Self::LOUD_SHARE_OF_WORKING_RATE) | ||
| .max(Self::QUIET_BYTES_PER_SECOND) | ||
| } |
There was a problem hiding this comment.
🔴 Operators of a parked program with a small steady churn of plain JS objects (a spinner, a 50-100 ms poll) get an idle full collection every ~70 s forever and never the second or last rung, where the base walked the ladder. On the 1 s tick perform_gc() at :359 collects every second, so JSC re-sweeps the same block and total_bytes_allocated charges nothing; on the 30 s tick blocks fill and each charges ~16 KB, so is_loud at :282 sees a few KB/s over the 2 KB/s floor and restarts the ladder. Fix: judge both ticks the same, e.g. give the floor a block-granularity term over the window (loud only when allocated exceeds line*elapsed plus a few MarkedBlock sizes) or keep the quiet fast ticks' collections from masking the reading.
Extended reasoning...
JSC's Heap::didAllocate, which is what totalBytesAllocated() counts, is only called when a LocalAllocator adopts a block (LocalAllocator::tryAllocateIn charges m_freeList.originalSize(), about 16 KB) or through reportExtraMemoryAllocated; after a collection resumeAllocating re-sweeps the allocator's last block into a fresh free list without a charge. So plain cells (strings, arrays, closures) are counted only when a block fills. This differs from the resolved :253 entry, whose cause was the 5 s window; that fix does not touch this asymmetry and a different fix is needed. Trigger: a main-thread program that allocates roughly 100-300 cells a second in a size class (a 10 fps spinner building a few strings per frame, a poll loop building a handful of objects) whose start-up fit in the first tick, which :343 skips, so working_rate is near 0 and the line is QUIET_BYTES_PER_SECOND (:72) = 2 KB/s. Fast ticks: each tick's collect_async(false) (:321, :359) frees the cells, the block never fills, allocated at :267 is 0, the tick is quiet; rung 1 runs at 10 s (:180-190); after 30 silent ticks…
Verification: normal (bounded: no crash; a parked program with a small steady churn of small cells gets rung 1's idle full collection every ~70 s and never reaches the module-graph / executable page-out rungs, which the base ran for it) — triggered by a main-thread program that allocates plain small cells at more than 2 KB/s but under roughly one free list (~15 KB) per size class per second, with a judged…
|
Closing in favour of one smaller PR (number to follow here). Measuring what each idle step saves next to what the next request or turn pays for it showed two things. Releasing anonymous memory at idle (compiled code, re-decodable bytecode, heap garbage) is a real saving for one re-warmed turn. Paging out clean file-backed pages (the module graph, the executable's own code) costs the next turn as much again or more, all of it in major page faults, for memory the kernel reclaims for free under pressure. Without the page-outs a wrong guess about "idle" is cheap, so the classifier this PR grew is not needed either. The replacement keeps the ladder of idle collections, the Worker fix and the fast ticks after a rung, drops re-decodable bytecode at the last rung, removes the module-graph page-out that is on main today, and has no page-out and no adaptive rule. |
|
Replaced by #43174: the ladder of idle collections at 10 s / 2 min / 10 min, Worker VMs do not run it, fast ticks after a collection, the last one drops re-decodable bytecode; no page-outs of anything and no adaptive rule. |
…ged out; the last drops re-decodable bytecode (#43174) Replaces #42329, #42383 and #43167 (one PR instead of three). Two commits: the first removes the module-graph page-out from the idle collection, the second is the ladder. The idle full collections (`BUN_IDLE_GC_SECONDS`) let JSC age out code that no longer runs and give its memory back. That memory is anonymous, so only the runtime can release it, and the price is a one-off re-warm: the next thing the program does compiles again what a collection aged out. This PR keeps that, makes it cheaper for a program whose user comes back, and fixes what is around it. `GarbageCollectionController.rs` is 248 lines (main: 260). 1. **The default list is `10,110,480`**: collections 10 s, 2 min and 10 min after the heap stopped growing (main: 10 s, 75 s, 140 s). What is freed is the same; a program that is used again within two minutes no longer pays for the second collection. 2. **The last collection also drops what JSC gets back cheaply**: before it, `VM::shrinkFootprintNow(LeaveCollectionToCaller | KeepCodeInUse)` lets go of the unlinked bytecode of functions that have no linked code any more (the earlier collections unlinked what had not run) and that a bytecode cache can hand back (a `--compile --bytecode` executable's embedded bytecode), and of the parser's caches. Nothing that would have to be parsed again, nothing in use. Deleting code waits for a collection that is under way (`Heap::preventCollection`), which on a big heap in the middle of a concurrent full collection is hundreds of milliseconds of the event loop, and JSC declines with JS on the stack (a timer fired from a nested event loop): the binding declines in the first case and JSC in the second, nothing is dropped, that tick's quiet is not counted, and the next tick tries again. What the binding cannot see is a collection that has been requested and has not started: the drop then waits for it, which is at most the collection the previous tick requested, an eden collection of a heap that has not grown. The collection that follows the drop is the same requested, concurrent one. It is not gated on standalone executables: without embedded bytecode there is little to drop and nothing that costs anything to get back. 3. **Every JS thread runs them for its own heap, as on main, and the code now says so.** The controller was written for the main thread only (`if vm.is_main_thread()`), but that asks whether the VM has a Worker, and a Worker's VM is initialised before it is given one: the test was always true and Workers have always run the ladder. It is removed rather than fixed. Measured: a pool of 8 Workers, each with 50 MB of old-generation garbage after a burst, 452 MB resident; 12 s later 40 MB, and 452 MB for good when only the main thread ran the ladder (each Worker's collections are its own; nothing the ladder does is process-wide any more, and the code drop is per VM). 4. **After an idle collection the timer stays on its fast tick for 30 ticks.** The collection is requested, not run: it proceeds at the mutator's safepoints, which in a program that runs no JS are this timer's ticks. The second and third were requested on the 30 s tick, where a server held on to a burst's garbage for a minute and more. 5. **Nothing is paged out.** Main's second idle collection also asked the kernel to reclaim the pages of a standalone executable's embedded module graph (`MADV_PAGEOUT`). Those pages are clean and file-backed: they are not ours to evict. The kernel drops them by itself, at no cost, as soon as it needs the memory, and until then they are a cache that makes the program's next action fast; paging them out by hand only lowers the RSS column, and the next thing the program does reads them back from the disk, one major fault at a time (numbers below). `StandaloneModuleGraph::page_out`, the trait method and the call are gone. `BUN_FEATURE_FLAG_DISABLE_STANDALONE_MADVISE` stays, because two other hints read it, both one-off at start-up and unchanged: the read-ahead of the module graph's start-up pages (`MADV_WILLNEED` / `F_RDADVISE`) and the `MADV_DONTNEED` hint for the embedded source text once the entry point has been evaluated. After this PR it gates only those two. "Busy" is what it is on main: the heap grew by more than 2 MB since the last tick. A wrong "idle" now costs a concurrent collection and a re-warm, which does not justify a finer rule. ## What each collection saves, and what the next request pays A ~200 MB compiled command-line program: a scripted run of 20 requests against a stub API, a pause, one more request (release builds; every run on its own copy of the executable; faults from `/proc/<pid>/stat`, instructions from `perf stat`; medians of 2-4 runs on a busy machine, so wall times are noisy and instructions are not). In steady state a request is 224 ms, 0.9 G instructions, 4 major faults. "main, no page-out" is main with `BUN_FEATURE_FLAG_DISABLE_STANDALONE_MADVISE=1`; "the page-out alone" is main with `BUN_JSC_forceCodeBlockLiveness=1` (no code aging, so the collection costs the next request nothing). | collection | anonymous memory freed | the next request pays | |---|---|---| | first (10 s) | 56-60 MB (the run's garbage; with code aging switched off it is the same) | nothing measurable (+4 to +35 ms, no extra instructions) | | second (main 75 s, here 2 min) | 55-57 MB (none of it with `BUN_JSC_forceCodeBlockLiveness=1`: it all hangs off aged-out code) | +200 to +340 ms wall, +1.3 to +2.2 G instructions, +15 K minor faults, no major faults; the request after it +20 to +60 ms | | third (main 140 s, here 10 min) | 24-34 MB (34 with the code drop) | +0.4 to +0.9 G instructions on top of the second's when both have run | | all of it | 90 MB of 257 idle after the scripted run; 12 MB of 104 idle after start-up | | | main's module-graph page-out (with the second collection) | none: 57 MB of file-backed RSS (104 -> 46 MB) | **+240 to +290 ms wall and 210-220 major faults** (the page-out alone, after 90 s and 3 min) | | the request after a pause of | main | main, no page-out | this PR | |---|---|---|---| | 30 s | +19 ms | -4 ms | not measured (nothing differs before 2 min) | | 90 s | **+600 ms**, +1.55 G instr., 262 major faults | +332 ms, +1.98 G, 0 | **+6 ms**, +0.1 G, 3 | | 3 min | +671 ms, +2.51 G, 262 | +209 ms, +2.53 G, 0 | +205 ms, +2.18 G, 0 | | 11 min | +768 ms, +2.67 G, 188 | +348 ms, +2.35 G, 0 | +197 ms, +1.75 G, 0 | Memory, anonymous | file-backed MB, seconds after the scripted run's last request: | | +5 s | +30 s | +80 s | +135 s | +150 s | +300 s | +610 s | |---|---|---|---|---|---|---|---| | main | 315 \| 104 | 259 \| 105 | 199 \| 46 | 199 \| 47 | 167 \| 47 | 165 \| 47 | 168 \| 54 | | main, no page-out | 324 \| 106 | 264 \| 104 | 201 \| 104 | 200 \| 104 | 167 \| 104 | 166 \| 104 | 169 \| 104 | | main, no code aging | 321 \| 105 | 265 \| 105 | 258 \| 48 | 257 \| 49 | 257 \| 49 | 257 \| 49 | 259 \| 55 | | this PR | 317 \| 105 | 265 \| 105 | 254 \| 105 | 197 \| 105 | 197 \| 105 | 196 \| 105 | **162** \| 105 | Idle after start-up: 115 | 97 at +5 s, 100 | 97 at +135 s, 89 | 97 at +610 s (main without the page-out: 108, 91, 86). The first request of a program that was idle for 10 minutes after start-up: 1402 ms and 402 major faults on main, 774 ms and 40 without the page-out. So between 75 s and 2 min after it was last used the program holds 55 MB more than on main, and from 10 minutes on a few MB less; in exchange the second collection's re-warm is only paid by a program that really was left alone. ## A server at scale Half an LRU Map of objects, half retained 16-512 KiB buffers; 2,000 requests a second for 60 s, then silence. Anonymous RSS in MB: | live | | under load avg / max | end | +10 s | +14 s | +15 s | +70 s | p50 / p99 | |---|---|---|---|---|---|---|---|---| | 800 MB | main | 1680 / 2470 | 2470 | 2385 | 803 | 803 | 803 | 0.22 / 14.5 ms | | 800 MB | this PR | 1681 / 2470 | 2470 | 2385 | 1159 | 803 | 803 | 0.22 / 13.8 ms | | 80 MB | main | 215 / 337 | 196 | 104 | 103 | 103 | 103 | 0.20 / 6.8 ms | | 80 MB | this PR | 213 / 355 | 185 | 125 | 108 | 108 | 107 | 0.20 / 6.9 ms | The first idle collection comes 10 s after the load as on main and what it frees is back within 5 s. (With that collection requested on the 30 s tick, as the second and third are on main, an earlier state held 2.36 GB until 92 s after the traffic had stopped; that is what the fast ticks after a collection are for.) ## What was tried and dropped #42329 and #42383 also paged out the module graph and, on the last rung, the executable's own code and constants: measured, that bought 57 + 36 MB of file-backed RSS, which the kernel reclaims for free under pressure, for +240 to +530 ms and 210-560 major faults on the next request. Because a wrong "idle" was then expensive, #42383 grew a classifier (the program's allocation rate against its own history, several clocks, a wake from the slow tick) that four rounds of review kept finding edge cases in, and that no allocation-only signal can get right for a server whose traffic never touches the JS heap. Without page-outs and without a synchronous collection a wrong "idle" is cheap, so all of that is gone. ## Tests `gc-controller-cadence.test.ts`, 16 tests, 6 s for the file, as on main (the new ones run alongside the existing ones; `test/expected-durations.json` still said 2 s and has estimates now, until it is regenerated): a Worker's 100 MB of old-generation garbage are given back by the Worker's own idle collection (nothing else asks for a full collection of its heap; this pins what main does); a `--compile --bytecode` executable (one, shared) loses its re-decodable unlinked code with the second collection of `1,1` and gives the same results afterwards, and still has it right after the first of `1,30` has been logged; 300 MB of old-generation garbage next to 50 MB of live data are back within seconds of the idle collection on a 20 ms tick. The last two behaviours fail on main. No test for the page-out: it would assert the absence of code that is deleted. Release build, 3 runs and twice with `BUN_DESTRUCT_VM_ON_EXIT=1`; clippy on `bun_jsc`, `bun_standalone_graph` and `bun_resolver`; `cargo check` for `x86_64-pc-windows-msvc` and `aarch64-apple-darwin` (before the last small change to `idle_tick`).
Contains #42329's commits until that lands (it makes main's idle collections one ladder and stops Worker VMs from running
it; the three commits here change its controller). The base is main; once #42329 is merged this
rebases down to its own three commits, which are the last three on the branch. The WebKit it needs (
Heap::totalBytesAllocated(),VM::shrinkFootprintNow(flags)) is on main since #42822 (Heap::allocationBudgetThisCycle(), which the first commit reads, isnot read any more after the third).
GC controller: a program is idle when it has stopped allocating, not when its heap has stopped growing.GC controller: in a standalone executable the last idle rung pages out its own code and drops what JSC can decode again(the page-out of the executable image used to be in GC controller: one idle ladder of full collections; Worker VMs do not run it #42329; it is here because it needs commit 1 to be safe for a program that
is serving).
GC controller: loud is measured against what the program itself allocates when it works(the line the first commit drew wasa share of the collector's budget; this replaces it).
The GC controller: one signal, and what the last rung evicts
#42329 makes main's idle work one ladder:
BUN_IDLE_GC_SECONDS(default10,110,480) lists the rungs, the controllerremembers when a tick last saw the program busy and how many rungs have run since, and a busy tick starts over. Busy there is
main's "the heap grew by more than 2 MB", which cannot tell a server from a parked program: a server's heap does not grow, and
a program that works hard on a flat footprint reads as quiet. This PR replaces that one term and then, on top of it, adds what
would hurt a working program if it were mistaken for a parked one. When collections run is otherwise what main and #42329
do: the ladder's requested idle collections and the collection the timer requests on every tick. An earlier state of this PR
also decided a "burst has ended" collection from the collector's budget and kept a byte bucket for the rungs; measured against
main on a pausing server, a spike and WebSocket bursts it made no difference worth its code (the same collections inside
bursts, the same RSS in the gaps, a collection 5 s after a burst where main has one after about 10 s), and the bucket made the
last rung slip to 20 minutes after a session. Both are gone.
GarbageCollectionController.rsis 376 lines (main 260, #42329290).
Loud (first and third commit). The controller reads
Heap::totalBytesAllocated()(the cycle in progress included, so thereis no lag of a collection) and keeps one number,
working_rate: the highest allocation rate it has seen between two looks atthe program, fading over the ladder's length (the last rung's time: 10 minutes by default) until the ladder has run to its
end. The program is loud when its rate is more than
working_rate / 16, and never for less than 2 KB a second, or when atick came more than two seconds late (#42329's rule). The rungs run their entries after the last loud look; a loud one puts the
timer on its 1 s tick and 30 ticks in a row that are not loud put it on the 30 s one. Four details:
the last one at the earliest (
process_gc_timerre-arms the timer; it judges nothing itself). The tick judges what wasallocated over the time it took. Loud: the timer is on its fast tick. Not loud: it goes back to where it was. So a burst in the
middle of half a minute is judged by its own rate, whatever the line (the server with 800 MB live, whose line is 22 MB a
second, is back on its fast tick within a second of its traffic coming back after a minute of silence), also when the
program parks right after the burst (the look comes from the timer, not from a turn of an event loop that may have nothing to
turn for), and a steady rate reads the same on either tick. (An earlier state judged every window as if it had
been 5 s long, which made a steady rate read six times higher on the slow tick than on the fast one: a trickle between 1/96
and 1/16 of the working rate flipped between loud and quiet every 70 s and never got past the first rung. The application's
job, 1.9 MB in one 30 s window, never reaches 8 MB and is judged as the 64 KB a second it is.)
until something is loud against it (which starts the ladder, and the fading, over). Left to fade, the application's job
would have been over the line 17 minutes after a session, and from 34 minutes on every tick would have been loud.
starting up, or the rung's doing: freed blocks handed out again, and after the last rung the code it dropped decoded again,
0.7 MB in the application, which at its first prompt is more than a sixteenth of anything it has allocated since starting
up. Rungs a second apart in the list therefore run two ticks apart.
drain, before the callbacks it runs: a timer of the program's that fires in the same drain ahead of the tick and sits in a
synchronous call for seconds left the tick "on time", those seconds idle ones and their allocation one tick's worth. Every
tick is late when it comes late, the two that do not judge included (a synchronous call is not idle time whichever tick
it delays); the timer is armed from the clock again at the end of the tick, so the time a rung's collection took is not time
the next tick is late by.
stopped process or a machine asleep for longer than the ladder took the whole of it away, and what the program allocated
afterwards, its background job, was its normal rate from then on: loud for good, never idle again.
second for a rung that never ran.
300 MB live that allocates 2 MB a second all the time read as parked, and would have had a synchronous collection, its bytecode
dropped and its code paged out while it served. And no fixed line works either, because a server answering "hello" five
times a second allocates 4 KB a second, less than the 60 KB a second of the background job of a parked application. Only the
program's own history tells them apart. The budget is not read any more: the floor was 1/4096 of it, 2 KB a second for a
new heap, and with a constant 2 KB a second every program below is classified as before, so
JSC__VM__allocationBudgetThisCycleis gone.The consequence, by design: a program with a modest steady allocation at its own normal rate is running, not idle. A small
server at a few requests a second is never evicted; neither is a program that redraws something every second at the rate it has
always allocated at. A program that runs no JS allocates nothing and goes down the ladder on schedule.
What this rule cannot tell apart. "A server that was busy and now serves modest steady traffic" and "an application that
ran a turn and now runs a modest steady job" are the same shape and differ only in the ratio to the peak; 1/16 is where the
line is drawn. The application's job is 1/300 to 1/1500 of a turn's rate. A server whose modest traffic is above a sixteenth of
its peak stays busy; one below it goes down the ladder once (bounded: three collections, the last one synchronous, and its
code read back in from the page cache as it is used) and then, because the yardstick no longer fades, nothing more happens to
it until something is loud against that line. (The fade is
working *= 1 - dt/Tper look, so e^-1 = 37 % of the peak is leftat the last rung: the line there is 1/43 of the peak. Fading linearly to zero by the last rung's time instead would make the
application's job, 57-63 KB a second against a line of 2 MB a second x (1 - t/600 s), loud at about 582 s, just before the
last rung: it would start the ladder over and the application would never be evicted, which is what this PR is for.)
Starting up is counted, except for what fits in the first tick. A real server's start-up spans several ticks, so its
working rate starts at its start-up peak; low steady traffic after that looks idle for up to ln(peak / 16 / rate) x T, and such
a server goes down the ladder once in its first ten minutes and then stays classified idle until something is loud against the
line it stopped at. The 5-requests-a-second rows below are for a server whose whole start-up fits in the first tick. Not
counting start-up does not work as it stands: at its first prompt the application has nothing else to be measured against, and
its ticks there (3.0, 1.4 and 0.7 KB a second at the three rungs) are no quieter than that server's whole load (3.7 KB a second).
How the measured programs are classified (release build; the 5 / 50 / 2 MB rows re-measured on this state for 66 s each with the default list, the others unchanged from the previous state):
The last rung (second commit; the only one of a list of one). All of this is for standalone executables (
bun build --compile,the applications it was built for): a process started as
bun file.jsgets the requested collection it gets on #42329, nothingis dropped and nothing is paged out. Before its collection,
JSC__VM__shrinkFootprintNowhas JSC dropits parser caches and the unlinked bytecode of functions that have no linked code any more (the earlier rungs' collections
unlinked what had not run) and that it can decode again from the executable's embedded bytecode; the rung's collection frees
it. Nothing that would need a re-parse, and nothing at all with JS on the stack: the binding returns whether it did anything,
and if it did not the rung is not counted and comes again on the next tick. The collection is synchronous when a page-out
follows, because the collector is the last thing that would read the executable back in (0.04 ms for the code drop plus 68 ms
for the collection on a 237 MB JS heap, 348 ms on a 1.0 GB one; measured on the previous state, the code is the same). The
executable is paged out with
MADV_PAGEOUTover the program's read-onlyPT_LOADsegments (bun_sys::page_out_range;page_out_program_imagefinds the image withfind_loaded_module, which now also returns the program headers, and issuesthe calls after that walk has returned, outside the loader lock). While a Worker is
alive the rungs run when they are due (their collections are for this thread's heap; the last one's is only requested then) and
what they could not page out is owed: the first tick that finds no Worker does it, and a loud tick forgets the debt with the rest
of the ladder;
BUN_FEATURE_FLAG_DISABLE_STANDALONE_MADVISEturns the page-outs off. The page-out of theexecutable used to be in #42329; it is here because it needs the first commit: under the heap-growth rule a server under load was idle.
process.memoryUsage().heapUsedreadsHeap::sizeAfterLastCollection(); theHeapObserverthat reconstructed it andJSC__VM__blockBytesAllocatedare gone.A server at scale: what is given back after a burst, and when
Half an LRU Map of objects, half retained 16-512 KiB buffers; 2,000 requests a second for 60 s (JSON parse, LRU churn and a
16-512 KiB scratch buffer per request), then silence. Anonymous RSS in MB, main (f937cf4) and this PR side by side, one run each (this state):
The server is loud on every tick of the load (351 MB a second; its working rate peaks at 406) and quiet from the first tick
after it, so the first rung comes 10 s after the load, as on main. An earlier state of both PRs had the first rung at 30 s,
which is when the timer goes to its 30 s tick: the rung's requested collection then started on time and ended a tick later, and
both held 2.36 GB until 92 s after the traffic had stopped. After any rung the timer now stays on its fast tick for 30 ticks
(#42329 has the details and the test).
The application (release build,
--lto=off; RSS in MB, anonymous + file-backed)Rungs at 12 s, 2 min and 10 min at the first prompt and 12 s, 2 min and 10 min after the session, in all three runs; nothing
starts the ladder over. The application's own job ten minutes after the session is 1.9 MB in one 30 s window, 64 KB a second
against a line of 2.05 MB a second, and what the last rung dropped being decoded again falls in the tick that only looks. One
of the runs went on for 41 minutes after the session: not one loud tick, no rung after the third, the timer on its 30 s tick
from 30 s after the last rung to the end. (With the first rung at 30 s and the budget-relative line these rows were 213 ->
185 -> 128 -> 93 and 416 / 427 -> 334 / 355 -> 228 / 233 -> 171 / 172.)
Numbers
A large bundled CLI application built as a standalone executable with embedded bytecode, scripted 20-request session against a
local fake API, release builds (
--lto=off). Total RSS in MB (anonymous + file-backed in brackets).(The two tables that follow were measured on an earlier state of this PR with the same rungs at the same times; the table above is this state.)
What each rung gives, default thresholds (2 runs each; the before/30 s/2 min/10 min columns are within a few MB of the
previous state of this PR, 4-5 runs):
After the session the application runs a job of its own ten minutes after the last request, a few seconds before the last rung;
the code that job ran is in use and stays, and over the following minutes it reads 7-13 MB of the executable back in. The last rung
does not come round again (#42329), so that stays until the application has worked and gone idle again: with the repeat the
previous state of this PR was at 161-173 (152-156 + 9-17) after 20 minutes. With the previous thresholds (
10,65,65) that stateended at 168-176 (156-160 + 12-16) 8 minutes after the session, and the state before it (shrink on its own clock, page-out
repeated on demand) at 171-174 (154-157 + 14-18) after the session and 105-108 (89-90 + 16-18) at the prompt. With a preloaded job
that allocates 6 MB every 5 s the application parked at its prompt still went down the ladder (107-109 (96-98 + 11) after 10
minutes; not re-measured on this state, the busy predicate is unchanged).
What coming back costs: the first request after the pause, a fresh process each time, user-mode instructions of that request and
of all 20 (
perf stat -Ilined up with the request's frames; wall time on a shared machine):(1-2 runs each on a machine at load 80-130, so the wall times say little. The previous state of this PR, measured alongside:
614 ms / 2.97 G / 19.4 G with no rung, 1215 ms / 3.29 G / 20.9 G after two, 1458 ms / 3.35 G / 19.9 G after all three: the first
request after the last rung costs 0.2-0.3 G more now, the session as a whole the same.)
The second rung is the one that costs most: about +0.5 G instructions and half a second on the first request. The last one adds
0.2-0.3 G to that.
What the WebKit upgrade gives on its own (fresh prompt, idle, peak,
--help) is measured in #42822.Tests
Release build (
bun scripts/build.ts --profile=release --lto=off) on main after #42822.test/js/bun/gc/gc-controller-cadence.test.ts: 50 tests, all pass on the release build (3 runs, 67 s, and once withBUN_DESTRUCT_VM_ON_EXIT=1; on the ASAN build CI made of this head, with the lane's environment, 23 pass / 28 skip twice, on the previous head; the three tests added since are for release lanes, one of them also runs on ASAN). The threetests that expect the executable to stay resident now run their child again whenever its code is gone, not only when it is
gone before a rung could have run: on a loaded machine the kernel takes freshly read pages now and then (one in six runs
here), a page-out that should not be there does it every time. That is far over a file's budget, because what it measures is seconds
of idleness: the "idle release" block (20 s of wall clock) runs its children side by side; the standalone-executable block
(45 s) is one at a time because
MADV_PAGEOUTskips pages another process maps, and every child there maps the same 100 MBexecutable. The file is out of the parallel batch (GC controller: one idle ladder of full collections; Worker VMs do not run it #42329 put it in
parallel-allowlist.json'sexcludeFiles, which iswhat the runner reads, and in
parallel-denylist.txt, which is whatscripts/update-parallel-allowlist.mjsregenerates theformer from: both are needed for the entry to survive a regeneration). The children are observed
without allocating: a child that allocated a little on every stamp would be a program at its normal rate, so the stamps are
written with
fs.writeSyncfrom a buffer the child keeps, and the FTL and bytecode children look at the heap (8 KB a look)only once their rungs have had their time. No test waits for a time: when the first rung comes depends on how many ticks
starting up is spread over (two on a release build, four on an ASAN one), so the positive ones end a moment after the rungs they
wait for have been seen. Lists of one, two and three (
3,2,2,2,1,1: the time the yardstick fades over is secondsthere, not a tick or two) run that many collections and none in the 2.5 s after (a tick that only looks and one that would
run a rung); rungs a second apart in the list run two seconds apart (a second apart when the tick after a rung runs
rungs); malformed lists and
0run none; programs that allocate 5 MB and 40 MB a second on a flat heap get none; a 10 ms timerand a timer that builds a spinner frame every 80 ms do not keep a program from going idle; 4 s in a synchronous call are not
idle time, nor are 3 s that make the tick after a rung late (the ladder starts over; taking that tick for an idle one had
the second rung run on the next), nor 3 s in a call made from a timer that fires in the same turn of the event loop as the
tick, ahead of it (the rung of
2ran 7 s in, a second after the call; it is due two seconds after); a process stopped(
SIGSTOP) for 3 s, longer than the ladder of1,1, has both rungs after it goes on, with 1/200 of its peak rate (nonebefore: the 3 s had taken the whole yardstick away); with
--inspectthe timer reaches its slow tick (no collection after the first4 s; once a second before; on ASAN lanes that child runs with
detect_leaks=0, because LeakSanitizer reports 376 bytes ofthe debugger's own start-up at exit and aborts); a burst on the 30 s tick brings the fast tick back, also when it is
over within a second of the last tick and the program parks (one collection in the 1.5 s after it before, a hundred now). The rule, on a 20 ms tick with lists of seconds (release
builds only: the rates mean nothing at a tenth of the speed): after a second at 80 MB a second, 0.5 MB a second goes down
1,1,1and 16 MB a second does not (a 20 ms tick sees a timer's allocation in lumps, twice the rate or none, hence ratiosof 1/160 and 1/5 and not 1/40 and 1/8; and since 37 % of the peak is left at the last rung, nothing above 1/43 can finish a
ladder); after half a second at 320 MB a second, a job of 4 MB a second that runs for 2.5 s out of every 5 does not start
1,1over once it has run to its end (it does, after every job, when the yardstick goes on fading: checked by takingthe condition out); 2 MB a second after half a second at 160, on the slow tick, where it is looked at every 8 MB, gets the
rung of
6when it is due; a second of work (6 MB, under what makes the controller look early) in the middle of a slowtick starts the ladder over; work at 320 MB a second that comes back on the slow tick is loud against a line of 20 MB a
second (looking every 8 MB over at least a second could report 8 MB a second at most: the rung came in the middle of the
work). On the default tick: a server answering five requests a second reaches no rung and the same
server without requests does; a standalone executable with 300 MB live that allocates 2 MB a second gets no collection from
the controller and keeps its code (the large-heap case). A burst's garbage (50 MB alive, 300 MB old-generation garbage) is
given back within seconds of the rung
(GC controller: one idle ladder of full collections; Worker VMs do not run it #42329). Page-out (Linux release lanes): the module graph goes with the second rung of three and the code does not; the code
goes with the last rung of two and with a list of one; neither goes after
1,99999999999, withBUN_FEATURE_FLAG_DISABLE_STANDALONE_MADVISE=1, or while a Worker is alive (both go right after it has; with a Worker alivefor the whole of a list of three all three collections run); the same executable run as plain
bun(BUN_BE_BUN=1) keeps itscode. These stay one at a time:
MADV_PAGEOUTskips pages another process maps. It is the test that looks, at/proc/<pid>/smaps; it tells the child to exit before its deadline, counts a mapping as gone when it has less than a quarterof what it had, runs a child again when the mapping was gone before the rung in question could have run, and says so when a
child exits before it is ready. On a
--compile --bytecodeexecutable the last rung of1,1drops the re-decodable unlinkedcode and the results are the same afterwards, the first of
1,30does not.Also on this build:
require-cache18/18,bundler_bytecode_portable22/22 andcompile-bytecode-tooling4/4 (both as inUpgrade WebKit to c775a5dc527d: less memory for code that does not run #42822),
html-rewriter-leak13 pass / 1 skip,heapStats-mimalloc; clippy onbun_jsc,bun_sysandbun_standalone_graph,cargo checkforx86_64-pc-windows-msvcandaarch64-apple-darwin(on this state).