Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
41 commits
Select commit Hold shift + click to select a range
8cd4e81
[Radix Cache] Sync Rust TreeCore and use it by default
Jialin Sep 15, 2026
5893507
[Radix Cache] Centralize rollout fallback and run unsupported test cases
Jialin Sep 15, 2026
7689fcc
Initialize pinned Rust toolchain for AMD CI
Jialin Sep 15, 2026
9848fd6
Build Xeon CI image from the tested checkout revision
Jialin Sep 15, 2026
d150590
Use shared TreeCore inspectors in scripted runtime tests
Jialin Sep 15, 2026
030f999
Fallback to Python for SWA buffer-mode window repair
Jialin Sep 15, 2026
1678578
Fix import grouping in TreeCore inspection test
Jialin Sep 15, 2026
f7b90b2
Allow available UCX transports in AMD NIXL CI tests
Jialin Sep 15, 2026
ef110a3
Port SWA buffer-mode window repair from #39283 to Rust
Jialin Sep 16, 2026
8f9c654
fix: register HIP weight allocations at their base address
Jialin Sep 16, 2026
c070212
Keep tree-core sync focused on NVIDIA validation
Jialin Sep 16, 2026
58743cf
Fix concurrent mock 3FS IO with positional reads and writes
Jialin Sep 16, 2026
f39cc49
Separate independent CI fixes from tree core sync
Jialin Sep 16, 2026
49a1918
Finalize Rust tree core default and document Python fallbacks
Jialin Sep 16, 2026
fe6d7ba
List Rust tree core fallback cases as separate bullets
Jialin Sep 16, 2026
c229619
Address TreeCore review naming and documentation nits
Jialin Sep 17, 2026
bb687f7
Fall back to Python for T-LRU eviction
Jialin Sep 18, 2026
5eb19dd
Document T-LRU in Rust backend fallback comment
Jialin Sep 18, 2026
3994569
Sync unified-memory SWA HiCache transfers from #37507
Jialin Sep 22, 2026
e58edbc
Sync internal Mamba write-back eviction from #40680
Jialin Sep 22, 2026
8f69b19
fix(cache): align Rust SWA reuse and host eviction with Python
Jialin Sep 23, 2026
41433e0
fix(cache): align Rust transfer eviction and acknowledgment ordering
Jialin Sep 23, 2026
0027895
feat(cache): support T-LRU in the Rust tree core
Jialin Sep 23, 2026
4d4e76f
refactor(cache): use one insertion-ordered node set
Jialin Sep 23, 2026
7a5024f
docs(cache): explain NodeSet links and insertion order
Jialin Sep 23, 2026
19ceefb
feat(cache): support floating-point T-LRU parameters
Jialin Sep 23, 2026
e60d2a6
refactor(cache): evaluate large T-LRU estimates directly
Jialin Sep 23, 2026
4ef5b64
refactor(cache): use bounded native T-LRU arithmetic
Jialin Sep 23, 2026
115d674
fix(cache): preserve component order in hybrid transfer plans
Jialin Sep 23, 2026
24f4d3c
Fix deferred Mamba eviction under unrelated SWA backup
Jialin Sep 24, 2026
0eb572b
Fall back to Python when the Rust toolchain is unavailable
Jialin Sep 24, 2026
c9530eb
Preserve internal SWA write-back and surface native tree-core panics
Jialin Sep 24, 2026
e402b92
Make native panic regressions executable by the CI runner
Jialin Sep 24, 2026
26697b3
Correct hybrid Mamba write-back test prefix accounting
Jialin Sep 24, 2026
2aafff9
Adapt tree-core parity coverage to current cache APIs
Jialin Sep 26, 2026
f705c05
Preserve rebased cache APIs and return native Python errors
Jialin Sep 26, 2026
56c564e
Unify Python and Rust internal write-back eviction
Jialin Sep 27, 2026
85e244c
Return internal eviction backup requests from Python components
Jialin Sep 27, 2026
4244b72
Merge branch 'main' into codex/tree-core-sync-latest-20260915
hzh0425 Sep 28, 2026
8950ef4
Merge branch 'main' into codex/tree-core-sync-latest-20260915
ispobock Sep 28, 2026
ba5b81c
test: bound optimistic prefill TCP transfer batches
Jialin Sep 28, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 0 additions & 4 deletions docs/docs/advanced_features/radix_eviction_policy.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -50,10 +50,6 @@ Omit the flag to accept every default. An unrecognized key fails at startup rath
TypeError: SLRUStrategy.__init__() got an unexpected keyword argument 'protected_treshold'
```

<Note>
`--radix-eviction-policy-config` is not supported by the experimental Rust tree core (`SGLANG_UNIFIED_RADIX_TREE_CORE_BACKEND=rust`), which builds its strategy from the policy name alone. Passing both fails at startup.
Comment thread
Jialin marked this conversation as resolved.
</Note>

### Policy parameters

`slru` and `tlru` take parameters. `lru`, `lfu`, and `priority` take none, so `--radix-eviction-policy-config` has no effect with them and any key is an error.
Expand Down
13 changes: 11 additions & 2 deletions python/sglang/srt/environ.py
Original file line number Diff line number Diff line change
Expand Up @@ -677,8 +677,17 @@ class Envs:
SGLANG_OPT_UNIFIED_CACHE_FREE_OUT_OF_WINDOW_SLOTS = EnvBool(True)
# Decode batches between SWA out-of-window evictions.
SGLANG_SWA_EVICTION_INTERVAL = EnvInt(128)
# Registered TreeCore backend serving the unified radix cache.
SGLANG_UNIFIED_RADIX_TREE_CORE_BACKEND = EnvStr("python")
# The tree-core registry falls back to Python for:
# - Session-aware caching.
# - C128 or other unsupported components.
# - Custom component overrides.
# - Non-Linux platforms.
# - Unsupported PyTorch versions.
# - Devices other than CPU or CUDA.
# - Installs with neither the Rust extension nor its sources.
# - Source builds with a missing or unusable Rust toolchain.
# This also applies when Rust is explicitly selected.
SGLANG_UNIFIED_RADIX_TREE_CORE_BACKEND = EnvStr("rust")
SGLANG_OPT_SWA_RELEASE_LEAF_LOCK_AFTER_WINDOW = EnvBool(False)

# ===================================================================
Expand Down
199 changes: 176 additions & 23 deletions python/sglang/srt/mem_cache/rust_tree_core/adapter.py

Large diffs are not rendered by default.

Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@
ComponentType,
EvictLayer,
ExternalLinkerLoadPhase,
InternalStateBackup,
LinkerTransferPhase,
LRURefreshPhase,
PrepareLoadBackResult,
Expand All @@ -23,6 +24,7 @@
"ComponentData",
"ComponentType",
"ExternalLinkerLoadPhase",
"InternalStateBackup",
"LinkerTransferPhase",
"EvictLayer",
"FullComponent",
Expand Down
18 changes: 14 additions & 4 deletions python/sglang/srt/mem_cache/unified_cache/components/base.py
Original file line number Diff line number Diff line change
Expand Up @@ -63,6 +63,15 @@ class EvictLayer(IntFlag):
ALL = DEVICE | HOST


@dataclasses.dataclass(frozen=True)
class InternalStateBackup:
"""Pause eviction for a host backup before freeing internal device state."""

node_id: NodeId
# Host capacity required by the backup, in this component's pool units.
num_tokens: int


@dataclasses.dataclass(frozen=True)
class PrepareLoadBackResult:
"""Outcome of prepare_load_back; default = nothing to prepare."""
Expand Down Expand Up @@ -522,11 +531,12 @@ def evict_device_next_node(
tracker: dict[ComponentType, int],
device_frees: dict[ComponentType, list[torch.Tensor]],
host_frees: dict[ComponentType, list[torch.Tensor]],
) -> Optional[NodeId]:
"""Advance one eviction step and return a device leaf, if selected.
) -> NodeId | InternalStateBackup | None:
"""Return a device leaf, an internal backup request, or no selection.

Implementations must return after one allocator-relevant internal
mutation so the caller can drain pending frees before continuing.
Backup requests leave device state intact until the core resumes eviction.
"""
assert self.is_evict_device_ongoing, (
f"{self.component_type} device eviction not started"
Expand All @@ -552,8 +562,8 @@ def _evict_device_next_node(
tracker: dict[ComponentType, int],
device_frees: dict[ComponentType, list[torch.Tensor]],
host_frees: dict[ComponentType, list[torch.Tensor]],
) -> Optional[NodeId]:
"""Advance the walk by at most one allocator-relevant mutation."""
) -> NodeId | InternalStateBackup | None:
"""Select a leaf, request backup, or perform at most one internal mutation."""
...

@abstractmethod
Expand Down
50 changes: 16 additions & 34 deletions python/sglang/srt/mem_cache/unified_cache/components/mamba.py
Original file line number Diff line number Diff line change
Expand Up @@ -30,6 +30,7 @@
CacheTransferPhase,
ComponentType,
EvictLayer,
InternalStateBackup,
LinkerTransferPhase,
LRURefreshPhase,
PrepareLoadBackResult,
Expand Down Expand Up @@ -369,7 +370,7 @@ def _evict_device_next_node(
tracker: dict[ComponentType, int],
device_frees: dict[ComponentType, list[torch.Tensor]],
host_frees: dict[ComponentType, list[torch.Tensor]],
) -> Optional[NodeId]:
) -> NodeId | InternalStateBackup | None:
"""Advance one device-eviction step and return a leaf, if selected.

An internal tombstone is one complete step so the caller can apply its
Expand Down Expand Up @@ -404,9 +405,18 @@ def _evict_device_next_node(
)
return x.id
if not enabled:
x_next = lru.get_prev_no_lock(x)
# write_back: demote the state to host before the internal tombstone.
self._maybe_backup_node_before_state_tombstone(x)
self._evict_device_cursor = lru.get_prev_no_lock(x)
cd = x.component_data[ct]
if (
self.tree_core.enable_hicache
and self.tree_core.is_write_back
and cd.host_value is None
and not x.backuped
and x.component_data[BASE_COMPONENT_TYPE].value is not None
):
# Keep the state live until the controller has attempted its host
# backup. Session cursors advance after the resumed tombstone.
return InternalStateBackup(node_id=x.id, num_tokens=1)
self.tree_core._evict_component_and_detach_lru(
x,
self,
Expand All @@ -418,38 +428,10 @@ def _evict_device_next_node(
self.tree_core._cascade_evict(
x, self, tracker, device_frees=device_frees, host_frees=host_frees
)
self._evict_device_cursor = lru.cursor_next() if enabled else x_next
if enabled:
self._evict_device_cursor = lru.cursor_next()
return None

def _maybe_backup_node_before_state_tombstone(self, node: UnifiedTreeNode) -> None:
"""Demote an internal node's mamba state to host before its tombstone
(write_back only), mirroring the leaf deferred-demote path.

The match validator passes only nodes holding the state on some
layer, so a dropped internal state caps the match frontier at this
node forever, leaving the subtree's still-resident KV unservable.
Best-effort: this walk must make progress (it satisfies an imminent
slot allocation), so any failure falls back to the legacy drop.
"""
cache = self.cache
cd = node.component_data[self.component_type]
if (
cache.cache_controller is None
or not cache.is_write_back
or cd.host_value is not None
or node.backuped
or node.component_data[BASE_COMPONENT_TYPE].value is None
):
return
# The backup executor pre-evicts only the KV host pool; make room
# for the state slot the way the PREFETCH hook does.
if (
self._mamba_pool_host is not None
and self._mamba_pool_host.available_size() < 1
):
cache.evict_host(1, self.component_type)
cache.backup_node_for_write_back(node.id)

def _evict_device_end(self) -> None:
"""Clear the device-eviction walk cursor state."""
if self.tree_core.enable_session_radix_cache:
Expand Down
75 changes: 31 additions & 44 deletions python/sglang/srt/mem_cache/unified_cache/components/swa.py
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,7 @@
ComponentType,
EvictLayer,
ExternalLinkerLoadPhase,
InternalStateBackup,
LinkerTransferPhase,
LRURefreshPhase,
PreparePrefetchResult,
Expand Down Expand Up @@ -711,7 +712,7 @@ def _evict_device_next_node(
tracker: dict[ComponentType, int],
device_frees: dict[ComponentType, list[torch.Tensor]],
host_frees: dict[ComponentType, list[torch.Tensor]],
) -> Optional[NodeId]:
) -> NodeId | InternalStateBackup | None:
"""Advance one device-eviction step and return a leaf, if selected.

An internal tombstone is one complete step so the caller can apply its
Expand Down Expand Up @@ -746,9 +747,33 @@ def _evict_device_next_node(
)
return x.id
if not enabled:
x_next = lru.get_prev_no_lock(x)
# write_back: demote the SWA KV to host before the internal tombstone.
self._maybe_backup_node_before_swa_tombstone(x)
self._evict_device_cursor = lru.get_prev_no_lock(x)
cd = x.component_data[ct]
if (
self.tree_core.enable_hicache
and self.tree_core.is_write_back
and self.tree_core.has_swa_host_pool
and cd.host_value is None
and not x.backuped
and x.component_data[BASE_COMPONENT_TYPE].value is not None
):
# Reserve the whole unbacked window, not just this victim, before
# its best-effort host backup and resumed internal tombstone.
needed = sum(
len(node.component_data[ct].value)
for node in self._collect_unbacked_swa_nodes(x)
)
if needed:
if ct == ComponentType.SWA:
return InternalStateBackup(node_id=x.id, num_tokens=needed)
# Custom SWA component types use Python's inherited inline
# path; the shared finish hook is for built-in SWA only.
if (
self._swa_kv_pool_host is not None
and self._swa_kv_pool_host.available_size() < needed
):
self.cache.evict_host(needed, ct)
self.cache.backup_node_for_write_back(x.id)
self.tree_core._evict_component_and_detach_lru(
x,
self,
Expand All @@ -760,48 +785,10 @@ def _evict_device_next_node(
self.tree_core._cascade_evict(
x, self, tracker, device_frees=device_frees, host_frees=host_frees
)
self._evict_device_cursor = lru.cursor_next() if enabled else x_next
if enabled:
self._evict_device_cursor = lru.cursor_next()
return None

def _maybe_backup_node_before_swa_tombstone(self, node: UnifiedTreeNode) -> None:
"""Demote an internal node's SWA KV to host before its tombstone
(write_back only), mirroring the leaf deferred-demote path.

The match validator treats an unbacked tombstone as a window reset,
so a dropped internal SWA segment caps the match frontier until a
full sliding window re-accumulates below it, leaving up to one
window of still-resident KV unservable. The leaf backup walk covers
ancestors only within one window of the evicted leaf, so an
internal node whose child spans the window arrives here unbacked.
Best-effort: this walk must make progress (it satisfies an imminent
allocation), so any failure falls back to the legacy drop.
"""
cache = self.cache
cd = node.component_data[self.component_type]
if (
cache.cache_controller is None
or not cache.is_write_back
or not self.tree_core.has_swa_host_pool
or cd.host_value is not None
or node.backuped
or node.component_data[BASE_COMPONENT_TYPE].value is None
):
return
# The backup executor pre-evicts only the KV host pool; make room
# in the SWA host pool the way the PREFETCH hook does.
needed = sum(
len(n.component_data[self.component_type].value)
for n in self._collect_unbacked_swa_nodes(node)
)
if needed == 0:
return
if (
self._swa_kv_pool_host is not None
and self._swa_kv_pool_host.available_size() < needed
):
cache.evict_host(needed, self.component_type)
cache.backup_node_for_write_back(node.id)

def _evict_device_end(self) -> None:
"""Clear the device-eviction walk cursor state."""
if self.tree_core.enable_session_radix_cache:
Expand Down
Loading
Loading