Skip to content

Fix unified HiCache physical transfers - #39479

Merged
merrymercy merged 173 commits into
mainfrom
yonghao/ump-hicache-physical-transfers
Oct 4, 2026
Merged

merrymercy merged 173 commits into
mainfrom
yonghao/ump-hicache-physical-transfers

Conversation

@ZYHowell

@ZYHowell ZYHowell commented Sep 14, 2026 •

Copy link
Copy Markdown
Collaborator

Description corrected after merge to match the merged code, following review #39479 (review). Struck text is what this description originally said; Corrected is the behavior as merged; Follow-up is what #42708 changes (not part of this PR).

Original PR: #37496 — previous reviews and comments.

Motivation

Unified HiCache must preserve index ownership and address lifetimes while device or shared host pages can move. Main now provides per-pool L2 translation, SWA load binding, device move gates, and Python SWA write-back demotion. This PR retains the remaining shared-host buffer and storage support while using those upstream contracts.

The preceding #39478 has merged. This PR now targets main; the conflict resolution preserves the existing head history and incorporates main f404db94d425b1a10dd4954e9f9a70d6ecb65402.

Modifications

  • Reuse main's L2 engine translation, virtual-owner SWA allocation/rollback, independent host-transfer move gate, deferred D2H batching, and sparse layer-ID span. Remove duplicate eager physical translation and allocator reservation/event bookkeeping.
    • Corrected: SWA H2D targets are physical reservations (alloc_physical / cancel_physical_reservation / free_physical), not main's virtual-owner SWA binding. The allocator keeps that bookkeeping: set_hicache_transfer_done_event records the transfer-done events, and _pending_hicache_load_pages blocks page moves until the load is queued. That count covered only lazy pools. In the FULL/SWA/Mamba tri-pool the SWA pool floats and is not lazy, so its reservations were not counted. A Mamba allocation could then move a reserved SWA page before its load was queued, while the queued copy kept the old address. Main's binding path stayed in the code but became unreachable.
    • Corrected: as merged, free-page allocation (_alloc_bind_fast_or_slow, alloc_physical) also waited on the registered transfer-done events whenever a lazy hole existed.
    • Follow-up: removes the unreachable binding path, and the allocation-time wait. Compaction and frees keep their waits. It also counts the floating SWA pool's reservations, so they block page moves until the load is queued or cancelled.
  • Preserve main's unified pool-shape support and host-pool selection, removing the original extra storage-backend allowlist {None, file, sim, mori, shm} along with the superseded broad model/PP/I/O restrictions. Retain the direct external-cache-linker guard because that path bypasses the L2 translation/lifetime contract.
    • Corrected: host-pool selection changed. Only no storage backend, or mori, selects the shared page envelope, in every host memory mode. file, sim, shm, mooncake and the other backends use separate host pools, in cache mode too.
    • Corrected: for unified memory with HiCache, this PR added two argument-layer asserts: --pp-size 1, and lazy compaction (SGLANG_DISABLE_LAZY_COMPACTION unset).
    • Follow-up: removes both asserts. This is not a claim that pipeline-parallel serving is validated. The allocator's own lazy-compaction checks predate this PR and are unchanged.
  • Support shared FULL/SWA host arenas in buffer_only for eligible two-ended layouts except Mooncake, with byte-based staging admission, load reserve, occupancy, atomic prefetch allocation, rank consensus, and rollback.
    • Corrected: only a backend that selects the envelope gets the shared arena; in buffer_only that is mori. The shared-host allocation consensus collectives were removed before merge. Other consensus checks are unchanged.
  • Keep Mooncake buffer-only startup on main's separate K/V host pools. Reject runtime Mooncake attachment to an already-created shared buffer arena before storage threads/distributed groups are created; restarting with --hicache-storage-backend mooncake selects compatible pools.
    • Corrected: the runtime rejection applies to every backend that needs separate host pools, not only Mooncake, and in either host memory mode. Attaching one of them at runtime to an existing envelope arena is rejected; main accepted that attach in cache mode.
  • Store each unified page envelope as one UMBP object; ordinary MHA K/V and split-head object cardinality remain unchanged.
  • Keep Python's upstream best-effort SWA demotion. Retain the Rust backup-action bridge, align failed-backup behavior to preserve FULL/descendants while dropping only SWA, and provide the resident/new FULL anchor segments required by main's SWA load-back controller.
    • Follow-up: the Rust backup-action bridge (a SWA write-back eviction barrier) fired only after the pending-node barrier from PR 39627 had already backed up any unbacked SWA window, and then ran a FULL write-back that SWA preservation does not need. The resident/new FULL anchor segments fed only the unreachable binding path. The follow-up removes both.
  • Preserve bounded decode prefix matching across L1/L2/L3 and unified-storage error-to-miss cleanup. Remove an identical test left at the obsolete singular kernel/ path.

Known inherited limitation: main already permits Mooncake in cache mode and selects a shared envelope for eligible FULL/SWA pools, but Mooncake's MHA metadata still expects two K/V pointers per page while that envelope supplies one. This raises even with page_first; an uncaught backup-worker exception can prevent ACK/release completion. This PR preserves that upstream cache-mode behavior and does not claim it is supported. The buffer-only fallback prevents the newly added shared-buffer support from extending the mismatch to main's working separate-pool path. Switching an existing shared buffer to Mooncake at runtime requires restart. Other untested backend/layout combinations are not claimed compatible.

Corrected: the limitation above does not apply to the merged code. In cache mode, Mooncake selects separate host pools, not the shared envelope. Other untested backend/layout combinations are still not claimed compatible.

Corrected: these merged behavior changes were not listed above:

  • Write-through host refill. Under write-through, insert_host keeps host slots already filled by an L3 prefetch when the matching node has device KV but no host copy. This is a cache-mode change and starts no extra D2H copy. Buffer-only completion takes the staging path before insert_host. See Fix unified HiCache physical transfers #39479 (comment).
  • Load-back admission. PrefillAdder takes a host lock on the selected host match from before prepare_load_back's reclaim until init_load_back holds its own locks, and builds a FULL load-back spec. Follow-up: the FULL spec is built only for SharedSWAPrefillBudget, the budget that reads it. The host lock stays.
  • Joint reclaim quotas. evict_to_free_tokens passes the FULL/SWA reclaim quotas together to one evict(EvictParams(...)) call instead of evict_for_alloc. This also applies to unified memory without HiCache. See the comment linked above.
  • Host-pool retraction. It is allowed for unified hybrid-SWA and for the decode radix cache.
  • Write-policy attach check. Attaching storage with write_through over a write_back cache first runs is_write_through_compatible. Host trees that write-through cannot hold are rejected before any policy change. Buffer-only mode is not checked.

Accuracy Tests

No model accuracy or serving runs were performed. Focused CPU contract tests passed, including the actual compiled Rust extension. Existing host-pressure/index-domain expectations were updated to the main contracts; no new permanent test cases were added.

Follow-up: #42708 adds CPU regression tests for these mechanisms: physical reservations (including the tri-pool's floating SWA pool), allocation and transfer order, the transfer lifecycles, load-back admission, the write-policy attach check, envelope selection for every storage backend choice, host-pool retraction (including decode radix retraction with a shared prefix) and SWA write-back eviction. CPU only; no GPU, serving or performance runs.

Speed Tests and Profiling

No performance runs or speedup claims. Main's move gate can defer compaction until transfers are acknowledged; main's D2H batching is retained.

Test plan

  • Focused existing host-pool, shared-buffer, index-domain, staged-transfer, and decode-lock tests: 63 passed, 1 CUDA skipped.
  • Tree/prefetch selection: 163 passed, 164 CUDA skipped; the changed host-pressure and retraction contract tests subsequently passed (2 tests, 4 subtests).
  • Compiled Rust SWA/load-back/backup/eviction selection: 30 passed.
  • Temporary CPU probes passed for mixed resident/new FULL anchors into the SWA controller, native component-only fallback, and UMBP key cardinality; these are not backend end-to-end tests.
  • Existing tombstone positive-write/completeness and buffer snapshot sanity cases: 4 passed, 20 subtests; all substantive assertions retained. Empty backup staging explicitly supplies zero occupied units. Existing pool/buffer/prefetch modules: 26 passed, 20 subtests.
  • Temporary CPU probe passed for actual Mooncake metadata cardinality (separate FULL/SWA: two keys/two pointers; shared envelope: confirmed two-key/one-pointer mismatch), pool selection versus main, and rejection before attachment side effects. No Mooncake client/IO ran.
  • Changed Python pre-commit hooks, targeted Rust crate Clippy/rustfmt, and git diff --check against main passed.
  • UMBP's test module could not collect because mori is absent in the existing environment. GPU/serving, storage backend IO, multi-rank/multi-node PD, and performance/accuracy tests were not run.

Checklist

  • Format changed code with applicable pre-commit and Rust checks.
  • Run focused existing unit tests and preserve their coverage under the upstream contracts.
  • Update the PR description for its current main-based implementation.
  • Provide model accuracy and performance results (not run).
  • Follow existing pool/controller/tree ownership boundaries.

Review and Merge Process

Prior discussion remains on the original PR linked above. This conflict fix preserves the current PR and branch history; review and CI must be evaluated against its new head.


CI States

Latest PR Test (Base): 🚫 Run #37174694531
Latest PR Test (Extra): ❌ Run #37174694422
Latest PR Test (AMD ROCm 10): ❌ Run #37174694602

yhzhuang and others added 30 commits September 1, 2026 15:04
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Co-authored-by: yhzhuang <yhzhuang@fb.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Minimize both component reclaim quotas under the joint capacity predicate so a blocked compaction path cannot turn one allocation shortfall into an all-SWA eviction.

Original prod_inference commit: 56c7082a92bc0c9024585387460e64b065652918

Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Co-authored-by: yhzhuang <yhzhuang@fb.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Co-authored-by: yhzhuang <yhzhuang@fb.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
# Conflicts:
#	python/sglang/srt/mem_cache/common.py
# Conflicts:
#	python/sglang/srt/mem_cache/common.py
@metamergebot

Copy link
Copy Markdown
Collaborator

This port clears our bar on e8364fe732f6b51423bd519dc7c216b7c809c645: every base-a and base-b job is green.
base-c and the accelerator suites are reported but not gated, because they run
on scarce hardware and upstream merges through them routinely.

gate 20✓✗6 lint✓ build✓ A✓ B✓ C 0/10⏳ other✓ · accel ✗amd,mlx,musa

@merrymercy @ZYHowell — could one of you take a look and merge it if you are happy? Nothing else is waiting on our side.

@metamergebot

Copy link
Copy Markdown
Collaborator

@ZYHowell this branch reads CONFLICTING against main since #42295 landed. It is one hunk: python/sglang/srt/managers/schedule_policy.py, PrefillAdder._lock_node. Main dropped the is_tree_cache() branch and now always calls inc_lock_ref(...).to_dec_params(); this PR wraps the same body with the lock_host host lock. A three-way probe of the other five files both sides touched since this branch's last main merge (89f2167) is clean.

Suggested resolution: keep this PR's signature and host-lock wrapper, take main's body.

    @contextmanager
    def _lock_node(self, last_node: TreeNode, *, lock_host: bool = False):
        host_lock_params = (
            self.tree_cache.inc_host_lock_ref(last_node).to_dec_params()
            if lock_host
            else None
        )
        try:
            # Replay the acquire's receipt (SWA boundary uuid, mamba flag) so the
            # release takes back exactly what this temporary lock took.
            dec_lock_params = self.tree_cache.inc_lock_ref(last_node).to_dec_params()
            try:
                yield None
            finally:
                self.tree_cache.dec_lock_ref(last_node, dec_lock_params)
        finally:
            if host_lock_params is not None:
                self.tree_cache.dec_host_lock_ref(last_node, host_lock_params)

base-a and base-b were fully green on e8364fe; the merge request above stands once the branch is mergeable again.

Preserve host locks while replaying cache lock receipts unconditionally.
Keep the logical host pool envelope and storage-format defaults.
@metamergebot metamergebot added the bypass-fail-fast CI: a failing job no longer aborts its siblings (lint still gates) label Oct 3, 2026
@metamergebot

Copy link
Copy Markdown
Collaborator

This port clears our bar on 2ce5783e68f2979048c92a928e97a926c8de6d40: every base-a and base-b job is green.
base-c and the accelerator suites are reported but not gated, because they run
on scarce hardware and upstream merges through them routinely.

gate✓ lint✓ build✓ A✓ B✓ C 0/10⏳ other✓ · accel ✗amd,npu

@merrymercy @ZYHowell — could one of you take a look and merge it if you are happy? Nothing else is waiting on our side.

@metamergebot

Copy link
Copy Markdown
Collaborator

@ZYHowell this branch reads CONFLICTING against main again since #41961 landed. It is one add/add hunk: python/sglang/srt/mem_cache/allocator/unified_sub_pool.py, end of MultiEndedAllocator, right after free_group_end. Main appended _set_capacity(...) there; this PR appends its HiCache physical-reservation block (_pending_hicache_load_pages, set_hicache_transfer_done_event, _wait_hicache_transfers, _expand_pages_to_tokens, alloc_physical, cancel_physical_reservation, free_physical) at the same spot. A three-way probe of the other ten files both sides touched since this branch's last main merge (00bcc25) is clean, and the #42362 is_chunk_cache → supports_prefix_sharing rename resolves to main's side without any edit on this branch.

Suggested resolution: keep both, main's method first, then this PR's block unchanged. _set_capacity ends in self.clear(), and this PR's clear() already resets _hicache_transfer_done_events and _pending_hicache_load_pages, so post-capture resizing and the reservation fence stay consistent.

    def free_group_end(self) -> None:
        pending, self.free_page_reps_group = self.free_page_reps_group, None
        super().free_group_end()
        if pending:
            reps = torch.cat(pending)
            self.free(reps, _pages=reps // self.page_size)

    def _set_capacity(
        self, max_slots: int, *, virtual_num_pages: Optional[int] = None
    ) -> None:
        """Set active ranges while retaining the captured v2p/p2v storage."""
        num_pages = int(max_slots) // self.page_size
        num_virtual_ids = (
            num_pages if virtual_num_pages is None else int(virtual_num_pages)
        )
        self.num_pages = num_pages
        self.num_virtual_ids = num_virtual_ids
        self.max_slots = num_pages * self.page_size
        self.size = self.max_slots
        self.clear()

    _pending_hicache_load_pages: _CapacityField[int] = _CapacityField()

    def set_hicache_transfer_done_event(self, transfer_key: Hashable, event) -> None:
        ...  # this PR's block, unchanged through free_physical

base-a and base-b were fully green on 2ce5783 (run 37103097720, with base-c 14/14 as well); the merge request above stands once the branch is mergeable again.

@metamergebot

Copy link
Copy Markdown
Collaborator

@ZYHowell gentle ping: #39479 still reads CONFLICTING against main, and a fresh three-way probe against the current main tip finds the same single add/add hunk in python/sglang/srt/mem_cache/allocator/unified_sub_pool.py described above, nothing new. The branch sits in sgl-project itself, so if it is easier, anyone with write access could apply that resolution directly. base-a, base-b and base-c were fully green on 2ce5783, and the merge request above stands once the branch is mergeable again.

…ty add/add)

Merges upstream main 4ab720e into yonghao/ump-hicache-physical-transfers.
The only conflict was the end of MultiEndedAllocator in
python/sglang/srt/mem_cache/allocator/unified_sub_pool.py, where main added
_set_capacity (#41961) next to this branch's HiCache transfer block; both are
kept. Pushed by the OSS sync bot on maintainer merrymercy's instruction.
@metamergebot

Copy link
Copy Markdown
Collaborator

@ZYHowell on maintainer @merrymercy's instruction I merged main (4ab720e) into this branch as 90e9dcb so the PR is mergeable again. The only conflict was the end of MultiEndedAllocator in python/sglang/srt/mem_cache/allocator/unified_sub_pool.py, where #41961 added _set_capacity next to this branch's HiCache transfer block; both are kept, exactly as in the resolution posted above. No other file needed a hand edit. CI is running on the new head; if you would rather resolve it differently, feel free to force-push over it.

… fix)

Merges upstream main 169e622, which carries #42454, the fix for the
base-a-test-cpu (11) breakage (MooncakeKVSender args in
test_prefill_complete_integration.py). Pushed by the OSS sync bot on
maintainer merrymercy's instruction.
@metamergebot

Copy link
Copy Markdown
Collaborator

@ZYHowell merged main (169e622) into this branch again as e6487f4 to pick up #42454, which fixes the base-a-test-cpu (11) breakage that failed the previous run; no conflicts this time. CI is running on the new head.

@metamergebot

Copy link
Copy Markdown
Collaborator

This port clears our bar on e6487f4e2d89546368e9312b4b9a4037dc3ef6c5: every base-a and base-b job is green.
base-c and the accelerator suites are reported but not gated, because they run
on scarce hardware and upstream merges through them routinely.

gate 20✓✗1 lint✓ build 9✓✗1✂3 A✓ B✓ C 0/10⏳ other✓ · accel ✗amd,npu,xpu

@merrymercy @ZYHowell — could one of you take a look and merge it if you are happy? Nothing else is waiting on our side.

@metamergebot

Copy link
Copy Markdown
Collaborator

/rerun-test test/registered/e2e/hicache/test_hicache_unified_memory.py test/registered/e2e/disaggregation/test_disaggregation_unified_memory_swa.py test/registered/e2e/disaggregation/test_disaggregation_unified_memory_tri.py

@github-actions

github-actions Bot commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test test/registered/e2e/hicache/test_hicache_unified_memory.py test/registered/e2e/disaggregation/test_disaggregation_unified_memory_swa.py test/registered/e2e/disaggregation/test_disaggregation_unified_memory_tri.py:

🚀 2-gpu-h100 (3 tests): ✅ View workflow run

cd test/ && python3 registered/e2e/hicache/test_hicache_unified_memory.py
cd test/ && python3 registered/e2e/disaggregation/test_disaggregation_unified_memory_swa.py
cd test/ && python3 registered/e2e/disaggregation/test_disaggregation_unified_memory_tri.py

@merrymercy
merrymercy merged commit 937c0a6 into main Oct 4, 2026
145 of 173 checks passed
@merrymercy
merrymercy deleted the yonghao/ump-hicache-physical-transfers branch October 4, 2026 07:38

@ch-wan ch-wan left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-reviewed at e6487f4 against main 169e622.

Everything from the last round is addressed. Envelope selection is now opt-in (no backend or mori) both at startup and on runtime attach, the buffer-only check uses the resolved threshold, the shared-host votes are gone, and the unreachable handlers and dead code are removed. The commits added since then need another pass: SWA reload through physical reservations, transfer-done events, the write-through host refill and the load-back host lock.

Should fix

  • Allocation now waits on the latest HiCache transfer-done events whenever a lazy hole exists. As far as I can tell, this makes a load-back batch's forward wait for its whole H2D. See the inline comment on _alloc_bind_fast_or_slow.
  • Two new startup rejections for unified memory + HiCache: --pp-size > 1 and SGLANG_DISABLE_LAZY_COMPACTION=1. See the inline comment in kv_cache_hook.py.
  • _swa_allocation_callbacks turns off main's SWA load binding for every unified stack. That leaves bind_swa_for_loaded_rows/free_swa and the new binds_swa_to_full branch dead. See the inline comment.
  • The PR description no longer matches the code:
    • It says the PR reuses main's virtual-owner SWA allocation and removes the allocator reservation/event bookkeeping. b55790f adds alloc_physical/cancel_physical_reservation/free_physical, _pending_hicache_load_pages and set_hicache_transfer_done_event, and bypasses bind_swa_for_loaded_rows.
    • It says buffer_only uses shared arenas for every backend except Mooncake, with rank consensus. Now only no backend or mori qualifies, and the consensus collectives were removed.
    • The "known inherited limitation" for cache-mode Mooncake no longer applies. _uses_unified_page_envelope_host now requires supports_page_envelope_host in every mode. So in cache mode, file/sim/shm/mooncake move from the envelope to separate pools, and attaching one of them at runtime to an envelope arena is rejected. Main accepts that attach.
    • It says the PP restrictions were removed, but kv_cache_hook.py now adds one.
    • These changes are not mentioned:
      • Under write-through, insert_host now refills host rows on device-only nodes. This is on the default HiCache + storage path.
      • PrefillAdder now takes a host lock and builds a FULL load-back spec for every load-back, including non-unified HiCache.
      • evict_to_free_tokens switches from evict_for_alloc to evict. This also affects unified memory without HiCache.
      • Host-pool retraction is now allowed for unified hybrid-SWA and for the decode radix cache.
      • Attach now runs is_write_through_compatible before switching from write_back to write_through.
  • The new mechanisms have no permanent tests. That covers physical reserve/cancel with the _pending_hicache_load_pages move block, the transfer-done event wait, the load-back host lock and prepare_load_back, the write-policy switch check, and the per-backend envelope selection from the last round. The changed tests mostly adapt mocks.

Cleanup / nits

  • The PR's backup_kv SWA barrier overlaps the barrier main added in #39627, which now runs first. See the inline comment in swa.rs.

Comment on lines +941 to +942
if self.lazy_compaction and self._free_phys_pages.numel() > 0:
self._wait_hicache_transfers()

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

start_loading records the load's finish_event through set_hicache_transfer_done_event (cache_controller.py:993). In _get_new_batch_prefill_raw, ready_to_load_host_cache() runs before prepare_for_extend(). So the batch's own extend allocation reaches this branch whenever either sub-pool has a lazy hole. _wait_hicache_transfers then makes the schedule stream wait on that event, and run_batch calls forward_stream.wait_stream(self.schedule_stream). As far as I can tell, the forward of every load-back batch then waits for the whole H2D, instead of waiting layer by layer through layer_done_counter. The next allocation waits on the write event from start_writing in the same way.

Scenario: --enable-unified-memory --enable-hierarchical-cache (lazy compaction is the default), with some freed pages not yet compacted. Every step with a load-back loses the H2D/compute overlap.

Could you confirm this, or share numbers showing it doesn't matter? If some path really needs the wait, for example reusing a page that is freed while a transfer still targets it, could the event be recorded only for pages freed while a transfer is in flight? _pending_reuse already does this for forward events. That would avoid waiting on every outstanding transfer.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Confirmed, your reading is right. start_loading registers the load's finish event before the batch's own extend allocation. When either sub-pool had a lazy hole, _alloc_bind_fast_or_slow / alloc_physical then called _wait_hicache_transfers(). That made the schedule stream, and through forward_stream.wait_stream(schedule_stream) the forward too, wait for the whole H2D and for any outstanding D2H.

#42708 removes that wait from both allocation paths instead of narrowing it to pages freed during a transfer. No page that an in-flight copy uses can be on the free list:

  • Write-through and buffer-only D2H sources keep a device lock until the ack has synchronized the copy's finish event. Write-back drains its D2H (writing_check(write_back=True)) before the demote frees anything.
  • An H2D target stays owned until its ack:
    • cache mode: by the load-back's locks, until loading_check;
    • buffer-only: by the request lock on the published span; redundant SWA destinations are released by free_physical at the ack.
  • Page moves stay blocked:
    • for an uncommitted reservation, by _pending_hicache_load_pages;
    • for a queued or unacknowledged transfer, by the host-transfer move gate.
  • _flush, the eager free path and free_physical still wait on the registered events.

One gap in that protection is fixed too. #39479 counted reservations in _pending_hicache_load_pages only on lazy pools. The SWA pool of the FULL/SWA/Mamba tri-pool floats and is not lazy, so a Mamba allocation in the same load could move the reserved SWA page before the load was queued, and the queued copy kept the old address. The follow-up counts the floating pool's reservations as well, until the load is queued or cancelled. Such a load now fails and rolls back unless evicting a checkpoint frees a slot.

test/registered/unit/mem_cache/test_unified_hicache_transfer_lifecycle.py checks this on CPU, with a stand-in that records the stream calls:

  • hole reuse and alloc_physical add no wait (these tests fail on current main);
  • moves and frees wait first;
  • the rows a copy uses are neither reused nor moved, including a floating SWA reservation in the tri-pool;
  • each layer's attention read still waits on its own load event.

These are CPU ordering checks, not GPU measurements. We have not measured the overlap or any speedup on a GPU yet. One note for that measurement, about the shared page-envelope host pools (no backend, or mori):

  • Each component's envelope is copied once, at that component's first mapped layer. FULL and SWA are copied separately, not all before layer 0.
  • So nothing within an envelope is copied layer by layer, but a later component's envelope copy can still overlap earlier layers.
  • With separate host pools, layer-by-layer overlap is possible within a component too.

Comment on lines +579 to +587
if cfg.enable_hierarchical_cache:
assert cfg.pp_size == 1, (
"--enable-unified-memory with hierarchical cache does not support "
"pipeline parallelism (--pp-size > 1)."
)
assert not envs.SGLANG_DISABLE_LAZY_COMPACTION.get(), (
"--enable-unified-memory with hierarchical cache requires lazy "
"compaction so pending H2D physical reservations remain stable."
)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On main, handle_unified_memory_pool requires pp_size == 1 and lazy compaction only under PD disaggregation. Here both become startup assertions for any unified-memory + HiCache run. --enable-unified-memory --enable-hierarchical-cache --pp-size 2 now fails to launch, and so does the same command with SGLANG_DISABLE_LAZY_COMPACTION=1, the A/B and rollback escape hatch. Main starts both.

The lazy-compaction assert comes from the new unbound alloc_physical reservations. Main's bind-at-allocation path has no equivalent guard. The PP assert has no stated reason, and the description says the PP restrictions were removed.

Could you either keep main's binding path for these configurations, or give the reason and list both as trade-offs in the description?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed. #42708 removes both asserts for unified memory with HiCache; the PD disaggregation asserts stay. test/registered/unit/server_args/test_unified_hicache_startup_args.py checks that --pp-size 2 and SGLANG_DISABLE_LAZY_COMPACTION=1 pass the argument layer. That is only an argument-layer check: pipeline-parallel serving with unified memory and HiCache has not been validated.

One correction on the lazy-compaction case. Main before #39479 also stopped at startup with SGLANG_DISABLE_LAZY_COMPACTION=1, just later:

  • The scheduler installs the host-transfer move gate for unified memory with HiCache (python/sglang/srt/managers/scheduler.py L598–L609 at 6cc661f, the parent of the Fix unified HiCache physical transfers #39479 merge).
  • install_move_gate asserts lazy compaction (python/sglang/srt/mem_cache/allocator/unified_sub_pool.py L242–L261, same commit).
  • SGLANG_DISABLE_LAZY_COMPACTION=1 turns lazy compaction off (python/sglang/srt/mem_cache/kv_cache_configurator.py L149–L152).

That allocator-level check predates #39479, and the follow-up leaves it unchanged.

Comment on lines +278 to +282
if isinstance(allocator, MultiEndedAllocator):
return dict(
device_alloc_fn=allocator.alloc_physical,
device_free_fn=allocator.cancel_physical_reservation,
)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Every unified SWA stack builds its SWA sub-allocator as a MultiEndedAllocator or FloatMultiEndedAllocator, so this branch always wins. As a result:

  • bind/free_bound (bind_swa_for_loaded_rows/free_swa, passed at :504 and :1344) are never installed, and no entry gets device_indices_from_anchor_fn.
  • The anchor loop in _resolve_device_transfers is unreachable.
  • The binds_swa_to_full branch that 6027473 just added to BufferModePipeline (pipeline.py:1295-1330 and :1367) is unreachable too, so buffer-only staged loads quietly take the elif swa_dev is not None path instead.

Could you keep one mechanism? Either drop the bind path, the swa_indices_from_anchor_fn/swa_free_from_anchor_fn plumbing and the binds_swa_to_full branch, or keep main's binding and drop the physical reservation.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in #42708. There is one mechanism now: physical reservation. The follow-up removes:

  • bind_swa_for_loaded_rows;
  • the device_indices_from_anchor_fn / swa_indices_from_anchor_fn / swa_free_from_anchor_fn plumbing;
  • anchor_index_parts and its producers (Python core and Rust adapter);
  • the anchor loop in _resolve_device_transfers;
  • the binds_swa_to_full branch in BufferModePipeline.

free_swa stays, because free_swa_segment still uses it when page_size == 1.

New tests cover the physical path (test/registered/unit/mem_cache/test_unified_hicache_transfer_lifecycle.py):

  • cache-mode load-back reserves only the missing window rows and binds them;
  • a failed reservation rolls back the whole load;
  • a later-pool failure cancels the reservation;
  • a buffer-only staged load reserves the whole window and frees the rows of resident pages at the ack.

Comment on lines +805 to +816
if tree_core.is_write_back
&& tree_core.swa_write_back_eviction_barrier_enabled
&& tree_core.component_state(SWA).evict_device_last_backup != Some(x)
&& !tree_core.arena.node(x).backuped()
{
// A later Full backup cannot recover SWA data after this
// internal node is tombstoned. Pause on the same cursor so
// the Controller can preserve the dirty path first.
tree_core.component_state_mut(SWA).evict_device_backup_node = Some(x);
tree_core.component_state_mut(SWA).evict_device_last_backup = Some(x);
cursor = Some(x);
break 'step None;

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

After the merge with #39627, main's SWA write-back barrier just above runs first. It already pauses on every internal node that has FULL on device and neither FULL nor SWA on host. As far as I can tell, that leaves this barrier firing only when SWA already has a host copy and FULL does not. Tombstoning device SWA loses nothing in that case, yet the controller still runs a FULL write-back for the node through backup_kv (_evict_device_next_node in unified_radix_cache.py).

Could you drop this barrier and its plumbing, or say which case it covers that main's barrier doesn't? The plumbing is swa_write_back_eviction_barrier_enabled, evict_device_backup_node/evict_device_last_backup, EvictDeviceNextNodeResult.backup_kv and the retry loop in _evict_device_next_node.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed. The barrier fired only when SWA already had a host copy, or had nothing to back up, and then ran a FULL write-back that nothing needed. #42708 drops the barrier and its plumbing:

  • swa_write_back_eviction_barrier_enabled and its enable call;
  • evict_device_backup_node / evict_device_last_backup;
  • EvictDeviceNextNodeResult.backup_kv;
  • the retry loop in _evict_device_next_node.

The barrier from PR 39627 still backs up an unbacked window first. The host-pressure test keeps its four cases (Rust/session × pinned/unpinned). New tests check three cases:

  • a hosted window is dropped with no D2H. This test is pinned to the Rust TreeCore, where the barrier lived, and reports a skip when that core cannot load; it fails on current main. A twin runs the same check on the Python TreeCore;
  • an unbacked window is backed up first;
  • under host pressure only the window is dropped.

@ZYHowell

ZYHowell commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator Author

Re: review, the bullets on evict_to_free_tokens switching from evict_for_alloc to evict, and on the write-through insert_host refill.

Eviction: The FULL/SWA reclaim plan returns joint eviction quotas, which are passed together in one EvictParams request. The current evict_for_alloc implementation uses per-component capacity targets; for a shared arena, those targets can count the same free bytes independently and stop before the combined FULL/SWA allocation fits. We are retaining the joint-quota call here. The distinction is the stopping condition, not whether compaction is allowed. This path is the unified-memory hybrid-SWA allocator's, so it also applies to unified memory without HiCache, as you noted. A joint allocation-goal API could consolidate this later, but that API refactor is outside this follow-up.

Write-through host refill: This is a separate cache-mode behavior change, not a requirement for unified-memory buffer-only operation. insert_host retains slots already populated by L3 prefetch when the matching node has device KV but no host copy. This can preserve the host prefix needed to retain a prefetched suffix, at the cost of additional L2 occupancy; it does not initiate another D2H copy. Buffer-only completion takes the staging path before insert_host. We are not adding another refill change or reverting this behavior in this follow-up. The refill is present in public commit 4154c31; it should be documented as an additional cache-mode change rather than presented as a prerequisite for buffer-only support. This explanation does not claim a measured performance benefit.

@ZYHowell

ZYHowell commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator Author

Following up on the remaining points of the review #39479 (review):

The follow-up is #42708. It is open and not merged.

"The new mechanisms have no permanent tests." The follow-up adds CPU-registered tests that run the real allocator, cache, controller and page-envelope host pool, and check row contents as well as counts:

  • physical reserve, bind, cancel and free with the _pending_hicache_load_pages move block, and the state after commit, completion and failure, including the tri-pool's floating SWA pool and a reservation that succeeds on retry after device eviction;
  • the transfer-done events:
    • allocation adds no wait;
    • compaction and frees wait first;
    • a closed transfer gate keeps rows in place;
    • copy rows are not overwritten;
  • the load-back host lock and prepare_load_back (below);
  • the write-policy switch check (below);
  • envelope selection for every --hicache-storage-backend choice, dynamic included, in cache and buffer-only modes (only no backend or mori selects the envelope), plus the rejection of a runtime attach to an envelope arena.

The files are:

  • test/registered/unit/mem_cache/test_unified_hicache_transfer_lifecycle.py;
  • test/registered/unit/mem_cache/test_hybrid_pool_assembler.py;
  • test/registered/unit/server_args/test_unified_hicache_startup_args.py.

PrefillAdder host lock and FULL load-back spec. The FULL spec is a load plan. It is now built only for SharedSWAPrefillBudget, the one budget whose prepare_load_back reads full_tokens; other budgets use the host hit length.

The host lock stays. It covers a different interval from load-back's own locks: from prepare_load_back's reclaim, which can write back device victims and evict host leaves, until init_load_back takes its own locks. A test runs admission under host pressure. It checks that:

  • the matched node stays pinned while reclaim evicts other host leaves;
  • its data loads intact;
  • the pins are released on success, failure and exception.

With the lock removed, reclaim evicts the matched SWA host copy and the load-back fails.

Host-pool retraction for unified hybrid-SWA and the decode radix cache. The behavior is unchanged.

  • One new test round-trips FULL and SWA rows through host-pool retraction on the unified pool, where the request shares a cached prefix. It checks that the restored rows and the shared prefix are intact.
  • A second test runs PD decode with the decode radix cache. Two requests share a prefix. One is retracted through release_req, restored with restore_kv_cache and finally released. At each step the test checks the restored rows and the host pool, and that the other request keeps its row and its locks on the shared prefix.
  • Other combinations are not validated beyond these.

The write_back→write_through attach check. It stays. It stops a write-back host tree that write-through cannot hold from being switched to write-through: auxiliary host rows without their FULL rows, or a host suffix below a parent whose FULL rows are not on host. It does not apply to buffer-only mode. The tests check that:

  • both shapes are rejected with no policy, thread or backend side effects;
  • a compatible tree switches;
  • buffer-only mode is not checked.

Validation. CPU only: torch 2.8 with the Python TreeCore, and torch 2.11 with the Rust TreeCore built from the branch. No GPU, serving or performance runs.

@ZYHowell

ZYHowell commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator Author

I've corrected this PR's description so that it matches the merged code, as the review asked. Each correction is marked inline. Struck text is what the description originally said, Corrected is how #39479 behaved as merged, and Follow-up is what #42708 changes. In short:

  • SWA H2D targets use physical reservations. The reservation and transfer-event bookkeeping still exists. The reservation count did not cover the tri-pool's floating SWA pool, which the follow-up fixes.
  • Only no backend or mori selects the shared page envelope, in every mode. A runtime attach of another backend to an envelope arena is rejected.
  • Only the shared-host allocation consensus collectives were removed.
  • The Mooncake "known inherited limitation" no longer applies.
  • Fix unified HiCache physical transfers #39479 added the PP and lazy-compaction startup asserts, and the follow-up removes them.
  • The behavior changes that were not listed (host refill, the admission host lock and FULL spec, the eviction quotas, host-pool retraction, the attach check) are now listed, together with the validation that actually ran.

ZYHowell added a commit that referenced this pull request Oct 6, 2026
The review of #39479 asked for permanent tests of five mechanisms:
physical reserve/cancel with the _pending_hicache_load_pages move block,
the transfer-done event wait, the load-back host lock and
prepare_load_back, the write-policy switch check, and the per-backend
envelope selection. Keep a short behavior test or two for each, and the
tests that replace base coverage the bind path took with it. Delete the
rest, and the support code only they used.

test_unified_hicache_transfer_lifecycle.py keeps 11 of its 37 tests:
- reserve/cancel and the move block:
  test_reservation_is_counted_unbound_and_released_by_cancel,
  test_pending_reservation_blocks_moves_until_its_load_is_queued, and,
  for the floating SWA pool whose reservations did not block moves,
  test_state_allocation_cannot_move_a_reservation_before_its_load_is_queued;
- the transfer-done event wait:
  test_hole_reuse_does_not_wait_for_unrelated_copies and
  test_moves_and_frees_wait_for_every_copy_first, which no longer fixes
  the order of the two waits, only that both come before the move;
- the load-back host lock and prepare_load_back:
  test_host_match_stays_pinned_while_reclaim_evicts_host_leaves and
  test_shared_budget_reclaims_for_the_load_and_pending_demand;
- the write-policy switch check:
  test_incompatible_host_trees_are_rejected_without_side_effects;
- SWA load-back into physical reservations, in place of the removed
  TestSwaLoadAllocation cases:
  test_only_the_missing_window_rows_are_reserved_and_bound,
  test_reservation_retries_once_after_device_eviction (now only the
  retry itself) and test_later_pool_failure_cancels_the_swa_reservation.

TestUnifiedPageEnvelopeSelection keeps
test_only_no_backend_and_mori_select_the_envelope, which now reads the
backend choices from the parser. The startup-args file keeps its two
argument-layer tests and drops the PD one.
wangwenmingaa pushed a commit to wangwenmingaa/sglang that referenced this pull request Oct 8, 2026
Keep the eviction-headroom reserve from this PR while adopting
upstream sgl-project#39479, which switched the unified hybrid-SWA call site from
evict_for_alloc() to evict() so cumulative reclaim quotas are fully
honored. Update the unit-test assertion accordingly and refresh the
mocked tree-cache guard renamed by sgl-project#42362 (is_chunk_cache ->
supports_prefix_sharing).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bypass-fail-fast CI: a failing job no longer aborts its siblings (lint still gates) hicache Hierarchical Caching for SGLang memory-pool run-ci CI: run the baseline test suite on this PR unified-radix-cache

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants