Repository navigation
[unified memory pool] Follow-up to #39479: allocate without transfer waits, pin floating SWA reservations, drop the startup asserts, add lifecycle tests - #42708
Closed
ZYHowell wants to merge 17 commits into
Conversation
Every unified SWA sub-allocator is a MultiEndedAllocator, so _swa_allocation_callbacks always returns alloc_physical and cancel_physical_reservation and never sets device_indices_from_anchor_fn. That left dead in every configuration: the anchor loop in _resolve_device_transfers, bind_swa_for_loaded_rows, the anchor_index_parts producers (the Python core and the Rust adapter) and BufferModePipeline's binds_swa_to_full branch. Remove them. Buffer-only staged loads keep their single path, which loads into reserved slots. free_swa stays because free_swa_segment still uses it for page_size == 1. Tests that exercised only the removed path are removed. The SWA load-back spec test keeps its spec assertions. The tri-pool assembly test now checks the physical callbacks.
PrefillAdder walked the tree for a FULL transfer spec on every UnifiedRadixCache load-back, including plain HiCache. The spec is used only to count new FULL slots for prepare_load_back, and that count is read only by SharedSWAPrefillBudget; the base budget's prepare_load_back is a no-op. Budgets now declare load_back_needs_full_tokens. The other budgets get the host hit length, an upper bound on the new slots. The host lock around load-back admission is unchanged.
The pending-node barrier from PR 39627 runs first and backs up an internal node's unbacked SWA window. So the barrier added with #39479 fired only when SWA already had a host copy, or had nothing to back up. It then started a FULL write-back that SWA preservation does not need. Remove that barrier and its plumbing: - the Rust walk state and its enable switch; - EvictionStepResult.backup_kv and the matching binding field; - the Python enable call; - the backup_kv branch of _evict_device_next_node, which returns to its pre-#39479 form. test_swa_write_back_host_full_preserves_full_and_continues keeps all four cases (Rust/session x pinned/unpinned) and their assertions. Each case now initializes HiCache through set_hicache_enabled() and a host pool group instead of the removed switch.
_alloc_bind_fast_or_slow and alloc_physical take pages that no copy uses, so they no longer wait on the transfer-done events when a lazy hole exists. A page under a copy never becomes a hole while the copy runs: - write-through and buffer-only D2H sources keep a device lock until the write ack has synchronized the copy's finish event; - write-back drains the D2H (writing_check(write_back=True)) before it demotes or frees the node, and retraction backups synchronize too; - an H2D target stays locked by its load-back until loading_check has synchronized the finish event (buffer-only: the request lock pins the staged span), and a physical reservation is cancelled only before its copy is submitted. Compaction keeps its protection: _flush, the eager free path and free_physical still wait on every transfer event, pending reservations block moves, and the host-transfer move gate stays installed. An allocation that has to compact goes through _flush.
…ction #39479 made two startup assertions unconditional for unified memory with hierarchical cache: --pp-size must be 1, and SGLANG_DISABLE_LAZY_COMPACTION must be unset. Before that, main had these checks only under PD disaggregation, where they stay. Remove both for HiCache; the allocator's own lazy-compaction checks are unchanged.
handle_unified_memory_pool accepts --enable-hierarchical-cache with --pp-size > 1 and with SGLANG_DISABLE_LAZY_COMPACTION set. The test checks only the argument layer, not pipeline-parallel serving. PD disaggregation keeps its own pipeline-parallel and lazy-compaction asserts, which the test also covers.
The page-envelope host pool copies with index_copy_ on CPU, so the real allocator, cache, controller and host pool run here. Stream waits are recorded through a stand-in for torch.cuda.current_stream(). The tests check row contents as well as counts. Allocator: - physical reserve, bind, cancel and free: counts and ownership; failed reservations; moves blocked until the load is queued. - hole reuse and alloc_physical do not wait on HiCache copies. - _flush, free_physical and eager free wait before any page moves. - a closed host-transfer gate keeps rows in place. Cache: - write-through and buffer-only D2H sources stay locked until their acks. - write-back frees rows only after the copy is synchronized. - load-back targets stay owned until the ack. - cache-mode SWA load-back reserves only missing window rows, and a failed reservation rolls back the whole load. - a buffer-only staged load reserves the whole window and frees the rows of resident pages at the ack. - host-pool retraction restores FULL and SWA rows and keeps a shared prefix. - each attention read waits on its own layer's load event. Admission: - the host match stays pinned through reclaim. - pins are released on success, failure and exception. - only the shared FULL/SWA budget builds the FULL load-back spec. Attach and eviction: - write_back to write_through attaches reject both incompatible host-tree shapes with no side effects, allow compatible trees and skip buffer-only. - a runtime attach of a separate-pool backend onto the envelope arena fails before any backend exists. - a hosted SWA window is dropped without a FULL write-back; an unbacked one is backed up first; host pressure drops only the window. The assembler tests cover the envelope choice per backend in both host modes.
In the FULL/SWA/Mamba tri-pool the SWA middle is a float, which never runs lazy compaction. alloc_physical counted a reservation in _pending_hicache_load_pages only on lazy pools, so a float reservation left moves_blocked() open. The controller allocates the Mamba checkpoint after the SWA reservation and before it queues the load. When the Mamba end needed room, make_room moved the float, including the unbound reserved page, and wrote through that page's -1 owner into the v2p sentinel. The transfer kept the old address, which then lay in Mamba state. A float now counts its reservations the way lazy end pools do. The count blocks make_room and compact_holes until the load event is registered or the reservation is cancelled; from the queueing on, the host-transfer gate also holds until the ack. End pools keep their behavior. A tri-pool load whose Mamba slot needs the float to move now fails and rolls back, unless evicting a checkpoint frees a slot. Tests, through HybridCacheController.load on the real CPU tri-pool: the state allocation cannot move the reservation and rollback returns the count to zero; with an eviction the load succeeds, the reserved rows stay put through bind and submission until the ack, and the modeled H2D leaves the Mamba state bytes unchanged.
The selection test covered only None, mori, file, sim, shm, mooncake and nixl. It now covers every --hicache-storage-backend choice in both host memory modes: npu_memcache, hf3fs, aibrix, eic, simm, tensorcast and dynamic too. A new case checks the list against the parser's choices. Another case runs the dynamic path with an extra config, under a custom name and under "mori": a dynamic backend keeps separate pools, because its class, and therefore its object layout, comes from the config. The selector is unchanged and no backend is created.
R7 deleted test_binding_retries_after_eviction along with the virtual SWA binding path it mocked. This restores the coverage on the physical path, with the real cache, controller and allocator. Tree-owned rows fill the SWA end, and their backups are acked. Then a SWA load transfer resolves through alloc_physical and the cache's own device eviction callback. Only recorders wrap the two callbacks. The test checks one eviction of the requested size and two reservation attempts, the first of which fails. The second attempt's physical rows become the transfer's destinations. The reclaimed SWA page goes into the reservation, and the evicted node stays on host. Survivors keep their rows. The reservation is counted and unbound, and allocation and byte accounting balance before and after the cancel.
The removed SWA write-back eviction barrier existed only in the Rust TreeCore. The test used whichever core the environment selected, so a passing run under the Python core said nothing about the barrier. The cache fixture takes an optional tree_core_backend. When the selection falls back from the requested core, the test is skipped and the skip names that core. The hosted-window test now pins Rust. A twin pins Python, so each run names the core each result came from.
The FULL/SWA retraction round trip pins a prefix by hand and calls the cache's backup and restore directly. It never runs PD decode with a radix cache, and no other request holds the prefix. The new test runs on the CPU unified FULL/SWA cache in decode mode, with the decode radix cache and host-pool retraction backups. Two real Req objects share a two-page prefix. Each joins at the root as a row owner and is inserted with checkpoint. The second request's insert reuses the first one's prefix rows. The scheduler's functions then drive it: release_req retracts it into the host pool; after its leaf leaves the device, a fresh row and restore_kv_cache bring the KV back; rejoining re-shares the prefix. At the end, release_req without a backup releases both requests. It is used instead of release_kv_cache, whose keyword differs between this tree and OSS main. At each step the test checks the restored FULL rows and SWA window and the host pool usage. It also checks the survivor's row, its protected length and its leaf lock, the shared node's lock count and the cache's protected size. Rows come from the composite alloc, because the decode preallocation's SWA-tail allocation is a Triton kernel.
test_shared_budget_counts_only_the_full_slots_a_load_adds checked the FULL argument that admission passes to prepare_load_back, on an idle pool. With no pending demand and no capacity pressure, dropping total_offset or swa_offset from load-back preparation would not change its result. The test, renamed to test_shared_budget_reclaims_for_the_load_and_pending_demand, uses the same resident-FULL fixture on a constrained pool. A running request locks a 40-token path, so only its SWA past the window can be reclaimed. The adder carries 24 mixed decode tokens as pending demand on both bands. The FULL argument check stays, and the test now also checks the outcome: - admission continues and the load is queued; - while the queued load blocks page moves, the batch's own extend and the pending decode tokens still allocate, with no move; - that room came from reclaiming the running request's out-of-window SWA, and its FULL rows and window keep their data; - after the ack the loaded rows hold the original data, only the request's device lock remains, and byte accounting balances. With total_offset or swa_offset dropped from SharedSWAPrefillBudget.prepare_load_back, nothing is reclaimed and the pending decode tokens no longer fit, so the test fails.
ZYHowell
requested review from
Jialin,
Ying1123,
alphabetc1,
hanming-lu,
hnyls2002,
huangtingwei9988,
hzh0425,
ispobock,
merrymercy,
xiezhq-hermann and
yizhang2077
as code owners
October 6, 2026 04:26
4 of 5 tasks
The review of #39479 asked for permanent tests of five mechanisms: physical reserve/cancel with the _pending_hicache_load_pages move block, the transfer-done event wait, the load-back host lock and prepare_load_back, the write-policy switch check, and the per-backend envelope selection. Keep a short behavior test or two for each, and the tests that replace base coverage the bind path took with it. Delete the rest, and the support code only they used. test_unified_hicache_transfer_lifecycle.py keeps 11 of its 37 tests: - reserve/cancel and the move block: test_reservation_is_counted_unbound_and_released_by_cancel, test_pending_reservation_blocks_moves_until_its_load_is_queued, and, for the floating SWA pool whose reservations did not block moves, test_state_allocation_cannot_move_a_reservation_before_its_load_is_queued; - the transfer-done event wait: test_hole_reuse_does_not_wait_for_unrelated_copies and test_moves_and_frees_wait_for_every_copy_first, which no longer fixes the order of the two waits, only that both come before the move; - the load-back host lock and prepare_load_back: test_host_match_stays_pinned_while_reclaim_evicts_host_leaves and test_shared_budget_reclaims_for_the_load_and_pending_demand; - the write-policy switch check: test_incompatible_host_trees_are_rejected_without_side_effects; - SWA load-back into physical reservations, in place of the removed TestSwaLoadAllocation cases: test_only_the_missing_window_rows_are_reserved_and_bound, test_reservation_retries_once_after_device_eviction (now only the retry itself) and test_later_pool_failure_cancels_the_swa_reservation. TestUnifiedPageEnvelopeSelection keeps test_only_no_backend_and_mori_select_the_envelope, which now reads the backend choices from the parser. The startup-args file keeps its two argument-layer tests and drops the PD one.
_pending_hicache_load_pages is read only by moves_blocked(). The reservations it counts come only from the HiCache SWA load callbacks (_swa_allocation_callbacks), which run on unified memory with HiCache. install_move_gate requires lazy compaction there, so the SWA allocator is a lazy end pool or the tri-pool's floating pool, and both already counted their reservations. Eager end pools never hand one out. So alloc_physical and cancel_physical_reservation now count in every pool, and _reservations_block_moves() and its FloatMultiEndedAllocator override are gone. An uncommitted reservation still blocks moves, in the floating pool too, until its load event is registered or it is cancelled.
Keep one test per behavior the review's test bullet names, plus the SWA reservation's retry after device eviction, which replaces the base test that went with the bind path. - Lifecycle file: 11 -> 7 tests. Drop the reserve/cancel bookkeeping test, the moves-and-frees wait test, and the missing-window and later-pool rollback tests; the move-block, hole-reuse and floating-pool tests and the adapted spec test cover them. Check one incompatible host tree, not two. The kept tests no longer assert private counters, page maps or addresses. - Assembler: the envelope selection is checked for no backend, mori, mooncake and dynamic, once; the host memory mode plays no part. - Remove the startup-args test file.
Comment and docstring changes only: the HiCache wait docstring, the SWA load-back and allocation-callback comments, the prefill-budget flag comment, and two in the lifecycle test. Each file's AST is unchanged.
Collaborator
Author
|
hand off to merge bot |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
This follows up #39479 (unified HiCache physical transfers) and addresses review #39479 (review). The review's R6 (joint reclaim quotas) and R7 (write-through host refill) need no code change; the explanation is at #39479 (comment).
How #39479 behaved as merged, and what this PR changes:
_alloc_bind_fast_or_slowandalloc_physicalwaited on every registered HiCache transfer event whenever a lazy hole existed.alloc_physicalcounted reservations in_pending_hicache_load_pagesonly on lazy pools. The tri-pool's SWA pool floats and is not lazy, so a Mamba allocation in the same load could move the reserved page (make_room) before the load was queued. The queued copy kept the old address, which then lay in Mamba state.handle_unified_memory_poolasserted--pp-size 1and lazy compaction for unified memory with HiCache.alloc_physical/cancel_physical_reservation), plus an unreachable bind-to-FULL path (bind_swa_for_loaded_rows,device_indices_from_anchor_fn,anchor_index_parts).PrefillAdderUnifiedRadixCacheload-back.SharedSWAPrefillBudget, the one budget that reads it.Modifications
Allocate free pages without waiting on HiCache transfers (
unified_sub_pool.py).Allocation takes pages that no copy uses. A page under a copy never becomes free while the copy runs:
writing_check(write_back=True)) before it demotes or frees the node. Retraction backup and restore synchronize too.loading_check;free_physicalat the ack._pending_hicache_load_pages); the next change extends this to the tri-pool's floating SWA pool. A queued or unacknowledged transfer blocks them through the host-transfer move gate._flush, the eagerfreepath andfree_physicalstill wait on the registered events. An allocation that must move pages goes through_flush, or a float'smake_room; both refuse to move while either protection is active.Count floating SWA reservations as pending HiCache loads (
unified_sub_pool.py).alloc_physicalandcancel_physical_reservationcount every reservation, in every pool. Only the HiCache SWA load callbacks reserve, and HiCache requires lazy compaction (install_move_gate), so this keeps the lazy end pools as before and adds the floating pool; eager end pools never reserve.make_roomandcompact_holesuntil the load event is registered (set_hicache_transfer_done_event, which the tri-pool forwards to the float) or the reservation is cancelled. From the queueing on, the host-transfer gate, installed on all three tri-pool members, holds until the ack.Drop the unified-memory HiCache startup asserts for PP and lazy compaction (
kv_cache_hook.py).--pp-size > 1and withSGLANG_DISABLE_LAZY_COMPACTIONset.install_move_gate,_compact_pending_impl) predate Fix unified HiCache physical transfers #39479 and are unchanged.Remove the unreachable SWA bind-to-FULL load path.
MultiEndedAllocator, so_swa_allocation_callbacksalways selectsalloc_physicalandcancel_physical_reservation._resolve_device_transfers;bind_swa_for_loaded_rows;anchor_index_partsproducers (Python core and Rust adapter);BufferModePipeline'sbinds_swa_to_fullbranch.free_swastays, becausefree_swa_segmentuses it whenpage_size == 1.Build the FULL load-back spec only for budgets that use it.
load_back_needs_full_tokens.init_load_backtakes its own locks.Remove the redundant SWA write-back eviction barrier. This removes:
EvictionStepResult.backup_kv, its binding field andEvictDeviceNextNodeResult.backup_kv;backup_kvbranch of_evict_device_next_node.The host-pressure test keeps all four cases (Rust/session × pinned/unpinned).
Tests. All CPU-registered. The review asked for permanent tests of five mechanisms; each has one or two short behavior tests. One more keeps the retry-after-eviction scenario of the removed binding-path tests.
test/registered/unit/mem_cache/test_unified_hicache_transfer_lifecycle.py(7 tests). These run the real CPU allocator, cache, controller and page-envelope host pool, and record CUDA stream waits through a stand-in fortorch.cuda.current_stream()._pending_hicache_load_pagesmove block:HybridCacheController.loadwith the production SWA and Mamba callbacks: a Mamba allocation that needs the floating SWA pool to move cannot move the reservation; the load rolls back, the reservation is cancelled, and the float can move again.alloc_physicaladd no wait, and the rows copies use keep their data.prepare_load_back:prepare_load_backreclaims a running request's out-of-window SWA before the queued load blocks page moves. The batch's own extend and the pending decode tokens then allocate with no move, and the loaded and running rows keep their data.test/registered/unit/mem_cache/test_hybrid_pool_assembler.py(1 new test): no backend andmoriselect the shared page envelope;mooncakeanddynamicdo not.Accuracy Tests
No model accuracy or serving runs were performed. No GPU was used for this PR.
CPU validation on this branch. Of the last four commits, two trim tests, one simplifies the reservation counting in
unified_sub_pool.py, and one shortens comments without changing any file's AST. The changed test files were rerun at the head (with the Rust TreeCore, on the same tests before two final edits: comments and one unused test line), and the seven test files that exercise the reservation counter at the counting change. The final lifecycle file was also run againstmainand against the commit before the floating-reservation fix. The other results below are fromcf45b4354b(fromdfa28af4b2for the joint-budget test), whose production code differs from the head's only in that counting and in comments.test_unified_hicache_transfer_lifecycle.py7 passed;test_hybrid_pool_assembler.py42;main, 2 fail: the floating-reservation test and the allocation-wait check. The floating-reservation test stops early there:main's resolver still readsdevice_indices_from_anchor_fn, which this branch removes, before it allocates.total_offset, orswa_offset, dropped fromSharedSWAPrefillBudget.prepare_load_back(a temporary change), the joint-budget test fails: the pending decode tokens no longer fit.mainand on this branch, apart from the added tests. There are 28 files: the assembler, HiCache regressions, storage-prefetch lifecycle, prefill budget and adder, multi-ended allocator, buffer-mode sidecar, staged write-back, tri-pool, physical write loc, allocation eviction, tree-core registry, PP prefetch ticket, rust-TreeCore integration, unified radix cache and server args. Because the allocator changed, they also include full-loc fast path, KV index translator, byte accounting, byte-budget sizing, capacity memo, dynamic gate capacity, float move gate, free without host sync, strided state, MHA and MLA views, and pending event batches.mainand on this branch. The failures are CPU-environment ones: CUDA fixtures, a child process, and the rust-TreeCore integration file's collection under torch 2.8.main). The HiCache regressions file loses the sixTestSwaLoadAllocationtests, which exercised only the removed binding path; their retry-after-eviction and later-pool rollback scenarios are now covered on the physical path.test_rust_tree_core_integration.py444 passed, 1 skipped, andtest_rust_tree_core_errors.py25 passed (the same counts as onmain).SGLANG_USE_CPU_ENGINE=1and its CUDA skip cleared in-process, and the run records the core of each case: Rust for the two non-session cases, Python for the two session cases.cargo test --features inspection,cargo fmt --checkandcargo clippy --all-targetsforrust/sglang-radix-treeat this head: 922 passed, 1 ignored (also 922/1 onmain). fmt and clippy (-D warnings) are clean, with toolchain 1.92 (therust/rust-toolchain.tomlpin).F401,F821,F823,F841,UP037) reports only sixF841s intest_unified_radix_cache_unittest.py, andmainhas the same six. The tools were run directly, not through the pre-commit hooks.Speed Tests and Profiling
None run, and no speedup is claimed. A GPU timeline check is still to do: earlier layers compute while later layers load, and allocation adds no stream dependency on unrelated copies.
A note on the shared page-envelope host pools (no backend, or
mori):Checklist
CI States
Latest PR Test (Base): ❌ Run #37544937847
Latest PR Test (Extra): ❌ Run #37544937419
Latest PR Test (AMD ROCm 10): ❌ Run #37544937603