Skip to content

[Perf][mem_cache] Persistent lazy eviction heap for the UnifiedRadixCache Full component (bit-identical order) - #38371

Open
zsun6 wants to merge 13 commits into
sgl-project:mainfrom
zsun6:perf/unified-radix-lazy-eviction-heap
Open

zsun6 wants to merge 13 commits into
sgl-project:mainfrom
zsun6:perf/unified-radix-lazy-eviction-heap

Conversation

@zsun6

@zsun6 zsun6 commented Sep 7, 2026 •

Copy link
Copy Markdown

Motivation

FullComponent._evict_device_start rebuilds a heap over every evictable device leaf on every eviction call
([(key(n), n) for n in evictable_device_leaves] + heapify), and _evict_device_end throws it away;
drive_host_eviction does the same over evictable_host_leaves. The call sits on the decode path: once the KV pool is full,
check_decode_capacity -> evict_from_tree_cache -> evict_for_alloc -> _evict -> _evict_components runs on every step that
needs to free the shortfall, so the per-call cost is O(#evictable leaves) of pure Python plus O(#leaves) tuple garbage,
independent of how few tokens are actually evicted. The Python UnifiedTreeCore is the default backend
(SGLANG_UNIFIED_RADIX_TREE_CORE_BACKEND="python"); the Rust core (rust/sglang-radix-tree/src/components/full.rs)
has the same per-call rebuild and is left untouched here (see Follow-ups).

Measured on CPU (Apple M3 Pro) with the real UnifiedRadixCache (Python core, decode-shaped loop: pool full,
evict(64 tokens) then insert one 64-token sequence per step, 300 steps, gc disabled, machine idle, min-of-3 p50;
--chunk-len 128 --kv-size 4000000, --num-seqs 5000/20000/50000; "heap rebuild" = time inside _evict_device_start,
the O(L) _evict_device_end list drop is the rest of the removed share):

evictable leaves evict() p50 before of which heap rebuild + drop evict() p50 after speedup
~3.5K 326 us 273 us (84%) 30 us 10.9x
~14K 1,426 us 1,262 us (89%) 32 us 44.6x
~35K 4,907 us 4,300 us (88%) 33 us 149x

With gc enabled the per-call tuple churn also drives garbage collections: at ~14K leaves, 300 steps triggered
4,471 / 407 / 11 (gen0/gen1/gen2) collections before vs 0 / 0 / 0 after; mean per call 8,582 → 35 us and the
p99 goes from 179 ms (a gen2 pass landing inside the call) to 82 us.
We claim only the collection frequency; gen2 pause length is a property of the live heap (#28067).

Amortization: with tails of 1-2K tokens an eviction call happens every ~15-30 decode steps at batch 64; with 100-300-token
agentic tails (the setting of #34012) every 1-3 steps. Under the overlap scheduler the scheduler-side CPU time is hidden
only when the scheduler is not CPU-bound; GC stalls never are.

Design

_LazyLeafHeap (unified_tree_core.py, ~180 lines): a persistent min-heap per leaf set with lazy invalidation.

  • Entries keep the legacy (key, node) shape (so UnifiedTreeNode.__lt__ stays the tie-break). _live[node] is the key a
    member was last pushed with; an entry is valid iff _live[node] == key; everything else is stale and skipped on pop.
  • refresh(node) (membership add), touch(node) (key input changed; one dict lookup for non-members), forget(node)
    (membership removal), promote(node) (the walk-time explicit parent push). Stale entries are bounded by compaction
    (len(heap) <= 2*len(live) + 64).
  • A walk (begin_walk / pop_next / end_walk) sees exactly what the rebuilt heap saw: keys frozen at begin_walk
    (the stored key is authoritative for the walk), nodes entering the set mid-walk stay invisible unless promoted,
    a yielded node loses its live entry until end_walk re-keys it if it is still a member (declined write-back victims).
  • Two instances owned by UnifiedTreeCore (full_device_heap, full_host_heap), created in reset(); the key function
    is resolved lazily on first use (session-ref tuple when --enable-session-radix-cache, else the strategy priority) so
    component overrides that are not a FullComponent keep working.
  • Hooks at every key-mutation site (last_access_time in _match_post_processor/_touch_node/Mamba commit,
    hit_count, priority, Full session_ref in the three coverage helpers) and at every membership change
    (_update_evictable_leaf_sets, the direct discards in _release_all_component_layers, _evict_host_leaf, tombstone
    cleanup, acquire_component_lock). sanity_check() verifies the heap invariants (live set == leaf set, every live key
    fresh, every live entry present, compaction bound) at all its existing call sites.
  • FullComponent._evict_device_start/_next_node/_end and drive_host_eviction become begin_walk / pop_next +
    promote(parent) / end_walk; import heapq leaves the component.
  • Kill switch: SGLANG_UNIFIED_RADIX_LAZY_EVICTION_HEAP=0 re-keys every member at begin_walk through the same code path
    (legacy cost, identical order). Happy to drop it if you prefer no flag.

Eviction order is identical to today for every strategy (lru/lfu/fifo/mru/filo/priority/slru) and the session-ref
tuple: same snapshot at walk start (I2: every member live with a fresh key), same explicit parent push at the same
instant, same yield-once semantics for declined victims, no ties among simultaneous members (last_access_time is
unique per node), and the same tuple shape.

Complexity: O(log H) per refresh / pop instead of O(L) per call; per-call garbage no longer proportional to L.
Hot-path cost of the hooks: one dict lookup per touched node (touch on a non-member), measured below.

Tests

  • New test/registered/unit/mem_cache/test_unified_radix_eviction_heap.py (register_cpu_ci, est_time=25):
    • _LazyLeafHeap unit tests (order, key updates, forget, walk snapshot / entrants / promote, mid-walk touch deferral,
      nested-walk assert, compaction bound, rebuild-each-walk mode, invariant reporting).
    • Order parity against the legacy per-call rebuild: the pre-change _evict_device_* / drive_host_eviction bodies
      are installed as a LegacyFullComponent via component_registry_override; both caches replay the same seeded op
      stream (insert with shared prefixes and priorities, match, lock/unlock, evict, session register/release) and every
      eviction call must yield the same victims (root-to-node token paths), the same evicted-token counts and the same
      final leaf set; sanity_check() every 25 ops. 7 policies × session on/off × 12 seeds, page_size 4, and the kill
      switch vs the lazy heap: 19 tests / 176 sub-tests (seed 8 caught a compaction-bound miss in end_walk, fixed in the head commit).
    • Targeted: MRU touched-leaf-first, LRU touched-leaf-protected, session release returns a leaf to the unreferenced
      band, parent promotion within one call, exception mid-walk leaves the heap consistent, 5k matches keep the heap
      compact, reset() empties both heaps.
  • Existing suites on CPU: test_unified_radix_allocation_eviction.py, test_evict_policy.py,
    test_session_unified_radix_cache.py green; test_unified_radix_cache_unittest.py produces the identical pass/fail
    set as main on this macOS box (481 passed; the 705 failures are the pre-existing CUDA/HiCache-host-pool environment
    failures, sanity_check — which now also checks the heaps — runs in all passing classes).
  • Hook coverage (each hook line commented out in turn, CPU suites re-run): 8 of the 19 hook sites are individually
    load-bearing on the Full-only CPU paths and are caught by the new file; of the rest, two are redundant with a sibling
    hook on the same insert walk and nine are only reachable through a HiCache host pool, the write_back host walk, or a
    FULL+MAMBA tree, which the existing unittest classes cover in CI where every sanity_check() call also validates the
    heap invariants.

Benchmarks

test_unified_radix_cache_bench.py gains --benchmarks evict_step (decode-shaped: pool full, evict B tokens, insert one
B-token sequence per step; prints the evictable-leaf count) and mean in the report line.

Regression guard (--benchmarks match insert lock cache_finished evict --num-seqs 20000 --chunk-len 128 --kv-size 4000000 --components full, p50 us, before → after):

bench (20K seqs, p50 us) before (main) after delta
match_prefix 34-35 36 +1-2 us (one dict lookup per matched ancestor)
insert 101-104 110-112 +7 us (~7%; 3-5 heap upserts per insert: new leaf + split/parent re-keys)
lock_unlock (pair) 21 23 +2 us (forget on lock, re-key + push on unlock)
cache_finished 127-129 137-139 +9 us (~7%; insert path + lock release)
evict (prefill-shaped, 20K tokens/call) p50 13 / p99 4,294 p50 3 / p99 4,395 p50 -10 us; p99 unchanged (dominated by the actual frees)

These are per-request costs (one insert/lock/finish per request) against savings of 300-4,900 us per eviction call
(one call per decode step under memory pressure); the hooks are the price of not rebuilding the heap. If the insert
delta matters to you I can trim the split path to a single upsert.

Follow-ups (not in this PR)

  • Rust core: full.rs evict_device_start/drive_host_eviction rebuild the heap per call the same way; a port of this
    design (BinaryHeap<Reverse<(PriorityKey, NodeIdx_)>> + live map, ~120 lines) is straightforward — I can open it, or
    it can fold into the Rust roadmap work.
  • Legacy RadixCache.evict / HiRadixCache.evict_host share the pattern (LMCache / flexkv / cpp tree paths).
  • Revealing tombstone-collapsed leaves within the same walk would yield more victims per call (behaviour change, kept out
    to preserve parity).

Relates to #20415 ("LRU Optimization"), #24072, and #34012. T-LRU (#34012) changes the Full eviction key inputs; whichever lands second should wire those mutations into the lazy-heap key refresh path.


AI assistance: the implementation and benchmarks in this PR were developed with AI assistance. I reviewed the changes, ran the tests and benchmarks myself, and take responsibility for the contribution.


CI States

Latest PR Test (Base): 🚫 Run #37561065667
Latest PR Test (Extra): 🚫 Run #37561065312
Latest PR Test (AMD ROCm 10): ❌ Run #37561065582

…ache Full component

FullComponent._evict_device_start rebuilt a heap over every evictable device
leaf on every eviction call ([(key(n), n) for n in evictable_device_leaves] +
heapify, dropped again in _evict_device_end); drive_host_eviction did the same
over the host leaves. Once the KV pool is full this runs on the decode path
(check_decode_capacity -> evict_from_tree_cache -> evict_for_alloc), so every
eviction call cost O(#evictable leaves) of Python plus O(#leaves) tuple garbage
regardless of how few tokens were freed.

Replace the per-call rebuild with a persistent lazy-deletion heap per leaf set
(_LazyLeafHeap, owned by UnifiedTreeCore): entries keep the legacy (key, node)
shape, a node's live key is tracked in a dict, stale entries are skipped on
pop and bounded by compaction, every key-mutation site (last_access_time,
hit_count, priority, Full session_ref) and every membership change refreshes /
forgets the node, and a walk keeps the legacy snapshot semantics (keys frozen
at walk start, entrants invisible unless promoted, yielded nodes re-keyed at
walk end). Eviction order is therefore identical for all strategies and the
session-ref tuple; sanity_check() now verifies the heap invariants.

Decode-shaped CPU benchmark (pool full, evict 64 tokens + insert one 64-token
sequence per step, gc off, min-of-3 p50): evict() 326 / 1,426 / 4,907 us ->
30 / 32 / 33 us at ~3.5K / 14K / 35K evictable leaves; with gc on at ~14K
leaves, 300 steps triggered 4,471 / 407 / 11 gen0 / gen1 / gen2 collections
before and none after (mean per call 8,582 -> 35 us).

Tests: new test_unified_radix_eviction_heap.py (heap unit tests; randomized
order-parity replay against the legacy component bodies installed through
component_registry_override, 7 policies x session on/off x seeds, page_size 4,
kill switch; targeted cases). Bench: new --benchmarks evict_step mode.
Kill switch: SGLANG_UNIFIED_RADIX_LAZY_EVICTION_HEAP=0 re-keys every leaf at
each walk through the same code path.

Signed-off-by: Zhongrui Sun <sunzhongrui66@gmail.com>
zsun6 and others added 5 commits September 7, 2026 13:33
Signed-off-by: Zhongrui Sun <sunzhongrui66@gmail.com>
…payloads

A walk yields victims through pop_next, which drops each victim's live entry
without a bound check, and forget() short-circuits on an already-yielded node.
A K-victim walk therefore left up to K stale heap entries while _live shrank by
K, so check_invariants reported an uncompacted heap and sanity_check raised.
Reproduced with the parity replay at mru/seed=8.

Compact once at the end of the walk and route all four bound checks through a
single _maybe_compact()/_compact_threshold() pair so the assertion cannot drift
from the enforcement again.

Tests: widen the parity seeds to 12 (seed 8 exposed the bug), add an
eviction-path compaction regression test, and pin the lazy side of replay_pair
so the suite still exercises the persistent heap when the kill switch is
exported. Bench: evict_step sliced shared-prefix sequences and re-inserted ~10
distinct payloads; synthesize a distinct tail per step instead.

Signed-off-by: Zhongrui Sun <sunzhongrui66@gmail.com>
@alphabetc1 alphabetc1 self-assigned this Sep 11, 2026
@zsun6

zsun6 commented Sep 16, 2026

Copy link
Copy Markdown
Author

@alphabetc1 could you run /tag-and-rerun-ci on this one? All the red checks are the pr-gate "Require run-ci label" step and the *-finish jobs downstream of it; no test job has run on the branch yet. I'm not in CI_PERMISSIONS.json, so I can't add the label from my side.

@alphabetc1 alphabetc1 added run-ci CI: run the baseline test suite on this PR run-ci-extra CI: also run the extra suite (requires run-ci) labels Sep 17, 2026
@zsun6

zsun6 commented Sep 18, 2026

Copy link
Copy Markdown
Author

/rerun-failed-ci

Resolve environ.py against the rust TreeCore default (sgl-project#39627) and follow
the new register_session_ref / IncLockRefResult signatures in the
eviction-heap test.
@zsun6

zsun6 commented Sep 30, 2026 •

Copy link
Copy Markdown
Author

Rebased on main after #39627. Python is the fallback core now rather than the default, but _evict_device_start there still clears and rebuilds the heap over all evictable leaves on every eviction call. That path is still the only core on non-Linux platforms, and it's what sessions and no-Rust-extension installs resolve to, so the fix still applies there.

CI defaults to the Rust core after sgl-project#39627, so the suite built a Rust-backed
cache and hit node_by_id, which is not ported yet. The heap under test lives
in the Python core, so the fixture selects it explicitly.
@zsun6

zsun6 commented Sep 30, 2026

Copy link
Copy Markdown
Author

/rerun-failed-ci

sgl-project#39627 removed tree_cls from _make_env and the other bench functions; the
tree core is now picked by SGLANG_UNIFIED_RADIX_TREE_CORE_BACKEND. The new
evict_step bench still passed it, so base-b failed with "_make_env() takes
from 4 to 5 positional arguments but 6 were given".
@zsun6

zsun6 commented Oct 3, 2026

Copy link
Copy Markdown
Author

/rerun-failed-ci

Preserve the persistent Python eviction heap toggle alongside the upstream sliding-window release setting and its legacy environment alias.

Assisted-by: OpenAI Codex
Signed-off-by: Zhongrui Sun <sunzhongrui66@gmail.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

run-ci CI: run the baseline test suite on this PR run-ci-extra CI: also run the extra suite (requires run-ci) unified-radix-cache

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants