Repository navigation
Conversation
…ache Full component FullComponent._evict_device_start rebuilt a heap over every evictable device leaf on every eviction call ([(key(n), n) for n in evictable_device_leaves] + heapify, dropped again in _evict_device_end); drive_host_eviction did the same over the host leaves. Once the KV pool is full this runs on the decode path (check_decode_capacity -> evict_from_tree_cache -> evict_for_alloc), so every eviction call cost O(#evictable leaves) of Python plus O(#leaves) tuple garbage regardless of how few tokens were freed. Replace the per-call rebuild with a persistent lazy-deletion heap per leaf set (_LazyLeafHeap, owned by UnifiedTreeCore): entries keep the legacy (key, node) shape, a node's live key is tracked in a dict, stale entries are skipped on pop and bounded by compaction, every key-mutation site (last_access_time, hit_count, priority, Full session_ref) and every membership change refreshes / forgets the node, and a walk keeps the legacy snapshot semantics (keys frozen at walk start, entrants invisible unless promoted, yielded nodes re-keyed at walk end). Eviction order is therefore identical for all strategies and the session-ref tuple; sanity_check() now verifies the heap invariants. Decode-shaped CPU benchmark (pool full, evict 64 tokens + insert one 64-token sequence per step, gc off, min-of-3 p50): evict() 326 / 1,426 / 4,907 us -> 30 / 32 / 33 us at ~3.5K / 14K / 35K evictable leaves; with gc on at ~14K leaves, 300 steps triggered 4,471 / 407 / 11 gen0 / gen1 / gen2 collections before and none after (mean per call 8,582 -> 35 us). Tests: new test_unified_radix_eviction_heap.py (heap unit tests; randomized order-parity replay against the legacy component bodies installed through component_registry_override, 7 policies x session on/off x seeds, page_size 4, kill switch; targeted cases). Bench: new --benchmarks evict_step mode. Kill switch: SGLANG_UNIFIED_RADIX_LAZY_EVICTION_HEAP=0 re-keys every leaf at each walk through the same code path. Signed-off-by: Zhongrui Sun <sunzhongrui66@gmail.com>
zsun6
requested review from
Ying1123,
alphabetc1,
hanming-lu,
hnyls2002,
huangtingwei9988,
hzh0425,
ispobock,
merrymercy,
xiezhq-hermann and
yizhang2077
as code owners
September 7, 2026 20:33
Signed-off-by: Zhongrui Sun <sunzhongrui66@gmail.com>
…payloads A walk yields victims through pop_next, which drops each victim's live entry without a bound check, and forget() short-circuits on an already-yielded node. A K-victim walk therefore left up to K stale heap entries while _live shrank by K, so check_invariants reported an uncompacted heap and sanity_check raised. Reproduced with the parity replay at mru/seed=8. Compact once at the end of the walk and route all four bound checks through a single _maybe_compact()/_compact_threshold() pair so the assertion cannot drift from the enforcement again. Tests: widen the parity seeds to 12 (seed 8 exposed the bug), add an eviction-path compaction regression test, and pin the lazy side of replay_pair so the suite still exercises the persistent heap when the kill switch is exported. Bench: evict_step sliced shared-prefix sequences and re-inserted ~10 distinct payloads; synthesize a distinct tail per step instead. Signed-off-by: Zhongrui Sun <sunzhongrui66@gmail.com>
Author
|
@alphabetc1 could you run |
Author
|
/rerun-failed-ci |
Resolve environ.py against the rust TreeCore default (sgl-project#39627) and follow the new register_session_ref / IncLockRefResult signatures in the eviction-heap test.
Author
|
Rebased on main after #39627. Python is the fallback core now rather than the default, but |
CI defaults to the Rust core after sgl-project#39627, so the suite built a Rust-backed cache and hit node_by_id, which is not ported yet. The heap under test lives in the Python core, so the fixture selects it explicitly.
Author
|
/rerun-failed-ci |
sgl-project#39627 removed tree_cls from _make_env and the other bench functions; the tree core is now picked by SGLANG_UNIFIED_RADIX_TREE_CORE_BACKEND. The new evict_step bench still passed it, so base-b failed with "_make_env() takes from 4 to 5 positional arguments but 6 were given".
Author
|
/rerun-failed-ci |
Preserve the persistent Python eviction heap toggle alongside the upstream sliding-window release setting and its legacy environment alias. Assisted-by: OpenAI Codex Signed-off-by: Zhongrui Sun <sunzhongrui66@gmail.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
FullComponent._evict_device_startrebuilds a heap over every evictable device leaf on every eviction call(
[(key(n), n) for n in evictable_device_leaves]+heapify), and_evict_device_endthrows it away;drive_host_evictiondoes the same overevictable_host_leaves. The call sits on the decode path: once the KV pool is full,check_decode_capacity -> evict_from_tree_cache -> evict_for_alloc -> _evict -> _evict_componentsruns on every step thatneeds to free the shortfall, so the per-call cost is O(#evictable leaves) of pure Python plus O(#leaves) tuple garbage,
independent of how few tokens are actually evicted. The Python
UnifiedTreeCoreis the default backend(
SGLANG_UNIFIED_RADIX_TREE_CORE_BACKEND="python"); the Rust core (rust/sglang-radix-tree/src/components/full.rs)has the same per-call rebuild and is left untouched here (see Follow-ups).
Measured on CPU (Apple M3 Pro) with the real
UnifiedRadixCache(Python core, decode-shaped loop: pool full,evict(64 tokens)then insert one 64-token sequence per step, 300 steps, gc disabled, machine idle, min-of-3 p50;--chunk-len 128 --kv-size 4000000,--num-seqs 5000/20000/50000; "heap rebuild" = time inside_evict_device_start,the O(L)
_evict_device_endlist drop is the rest of the removed share):evict()p50 beforeevict()p50 afterWith gc enabled the per-call tuple churn also drives garbage collections: at ~14K leaves, 300 steps triggered
4,471 / 407 / 11 (gen0/gen1/gen2) collections before vs 0 / 0 / 0 after; mean per call 8,582 → 35 us and the
p99 goes from 179 ms (a gen2 pass landing inside the call) to 82 us.
We claim only the collection frequency; gen2 pause length is a property of the live heap (#28067).
Amortization: with tails of 1-2K tokens an eviction call happens every ~15-30 decode steps at batch 64; with 100-300-token
agentic tails (the setting of #34012) every 1-3 steps. Under the overlap scheduler the scheduler-side CPU time is hidden
only when the scheduler is not CPU-bound; GC stalls never are.
Design
_LazyLeafHeap(unified_tree_core.py, ~180 lines): a persistent min-heap per leaf set with lazy invalidation.(key, node)shape (soUnifiedTreeNode.__lt__stays the tie-break)._live[node]is the key amember was last pushed with; an entry is valid iff
_live[node] == key; everything else is stale and skipped on pop.refresh(node)(membership add),touch(node)(key input changed; one dict lookup for non-members),forget(node)(membership removal),
promote(node)(the walk-time explicit parent push). Stale entries are bounded by compaction(
len(heap) <= 2*len(live) + 64).begin_walk/pop_next/end_walk) sees exactly what the rebuilt heap saw: keys frozen atbegin_walk(the stored key is authoritative for the walk), nodes entering the set mid-walk stay invisible unless
promoted,a yielded node loses its live entry until
end_walkre-keys it if it is still a member (declined write-back victims).UnifiedTreeCore(full_device_heap,full_host_heap), created inreset(); the key functionis resolved lazily on first use (session-ref tuple when
--enable-session-radix-cache, else the strategy priority) socomponent overrides that are not a
FullComponentkeep working.last_access_timein_match_post_processor/_touch_node/Mamba commit,hit_count,priority, Fullsession_refin the three coverage helpers) and at every membership change(
_update_evictable_leaf_sets, the direct discards in_release_all_component_layers,_evict_host_leaf, tombstonecleanup,
acquire_component_lock).sanity_check()verifies the heap invariants (live set == leaf set, every live keyfresh, every live entry present, compaction bound) at all its existing call sites.
FullComponent._evict_device_start/_next_node/_endanddrive_host_evictionbecomebegin_walk/pop_next+promote(parent)/end_walk;import heapqleaves the component.SGLANG_UNIFIED_RADIX_LAZY_EVICTION_HEAP=0re-keys every member atbegin_walkthrough the same code path(legacy cost, identical order). Happy to drop it if you prefer no flag.
Eviction order is identical to today for every strategy (
lru/lfu/fifo/mru/filo/priority/slru) and the session-reftuple: same snapshot at walk start (I2: every member live with a fresh key), same explicit parent push at the same
instant, same yield-once semantics for declined victims, no ties among simultaneous members (
last_access_timeisunique per node), and the same tuple shape.
Complexity: O(log H) per refresh / pop instead of O(L) per call; per-call garbage no longer proportional to L.
Hot-path cost of the hooks: one dict lookup per touched node (
touchon a non-member), measured below.Tests
test/registered/unit/mem_cache/test_unified_radix_eviction_heap.py(register_cpu_ci,est_time=25):_LazyLeafHeapunit tests (order, key updates, forget, walk snapshot / entrants / promote, mid-walk touch deferral,nested-walk assert, compaction bound, rebuild-each-walk mode, invariant reporting).
_evict_device_*/drive_host_evictionbodiesare installed as a
LegacyFullComponentviacomponent_registry_override; both caches replay the same seeded opstream (insert with shared prefixes and priorities, match, lock/unlock, evict, session register/release) and every
eviction call must yield the same victims (root-to-node token paths), the same evicted-token counts and the same
final leaf set;
sanity_check()every 25 ops. 7 policies × session on/off × 12 seeds, page_size 4, and the killswitch vs the lazy heap: 19 tests / 176 sub-tests (seed 8 caught a compaction-bound miss in
end_walk, fixed in the head commit).band, parent promotion within one call, exception mid-walk leaves the heap consistent, 5k matches keep the heap
compact,
reset()empties both heaps.test_unified_radix_allocation_eviction.py,test_evict_policy.py,test_session_unified_radix_cache.pygreen;test_unified_radix_cache_unittest.pyproduces the identical pass/failset as
mainon this macOS box (481 passed; the 705 failures are the pre-existing CUDA/HiCache-host-pool environmentfailures,
sanity_check— which now also checks the heaps — runs in all passing classes).load-bearing on the Full-only CPU paths and are caught by the new file; of the rest, two are redundant with a sibling
hook on the same insert walk and nine are only reachable through a HiCache host pool, the write_back host walk, or a
FULL+MAMBA tree, which the existing unittest classes cover in CI where every
sanity_check()call also validates theheap invariants.
Benchmarks
test_unified_radix_cache_bench.pygains--benchmarks evict_step(decode-shaped: pool full, evict B tokens, insert oneB-token sequence per step; prints the evictable-leaf count) and
meanin the report line.Regression guard (
--benchmarks match insert lock cache_finished evict --num-seqs 20000 --chunk-len 128 --kv-size 4000000 --components full, p50 us, before → after):These are per-request costs (one insert/lock/finish per request) against savings of 300-4,900 us per eviction call
(one call per decode step under memory pressure); the hooks are the price of not rebuilding the heap. If the insert
delta matters to you I can trim the split path to a single upsert.
Follow-ups (not in this PR)
full.rsevict_device_start/drive_host_evictionrebuild the heap per call the same way; a port of thisdesign (
BinaryHeap<Reverse<(PriorityKey, NodeIdx_)>>+ live map, ~120 lines) is straightforward — I can open it, orit can fold into the Rust roadmap work.
RadixCache.evict/HiRadixCache.evict_hostshare the pattern (LMCache / flexkv / cpp tree paths).to preserve parity).
Relates to #20415 ("LRU Optimization"), #24072, and #34012. T-LRU (#34012) changes the Full eviction key inputs; whichever lands second should wire those mutations into the lazy-heap key refresh path.
AI assistance: the implementation and benchmarks in this PR were developed with AI assistance. I reviewed the changes, ran the tests and benchmarks myself, and take responsibility for the contribution.
CI States
Latest PR Test (Base): 🚫 Run #37561065667
Latest PR Test (Extra): 🚫 Run #37561065312
Latest PR Test (AMD ROCm 10): ❌ Run #37561065582