Repository navigation
Fix unified HiCache physical transfers - #39479
Conversation
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Co-authored-by: yhzhuang <yhzhuang@fb.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Minimize both component reclaim quotas under the joint capacity predicate so a blocked compaction path cannot turn one allocation shortfall into an all-SWA eviction. Original prod_inference commit: 56c7082a92bc0c9024585387460e64b065652918 Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Co-authored-by: yhzhuang <yhzhuang@fb.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Co-authored-by: yhzhuang <yhzhuang@fb.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
# Conflicts: # python/sglang/srt/mem_cache/common.py
# Conflicts: # python/sglang/srt/mem_cache/common.py
|
This port clears our bar on gate 20✓✗6 lint✓ build✓ A✓ B✓ C 0/10⏳ other✓ · accel ✗amd,mlx,musa @merrymercy @ZYHowell — could one of you take a look and merge it if you are happy? Nothing else is waiting on our side. |
|
@ZYHowell this branch reads CONFLICTING against Suggested resolution: keep this PR's signature and host-lock wrapper, take main's body. @contextmanager
def _lock_node(self, last_node: TreeNode, *, lock_host: bool = False):
host_lock_params = (
self.tree_cache.inc_host_lock_ref(last_node).to_dec_params()
if lock_host
else None
)
try:
# Replay the acquire's receipt (SWA boundary uuid, mamba flag) so the
# release takes back exactly what this temporary lock took.
dec_lock_params = self.tree_cache.inc_lock_ref(last_node).to_dec_params()
try:
yield None
finally:
self.tree_cache.dec_lock_ref(last_node, dec_lock_params)
finally:
if host_lock_params is not None:
self.tree_cache.dec_host_lock_ref(last_node, host_lock_params)
|
Preserve host locks while replaying cache lock receipts unconditionally. Keep the logical host pool envelope and storage-format defaults.
|
This port clears our bar on gate✓ lint✓ build✓ A✓ B✓ C 0/10⏳ other✓ · accel ✗amd,npu @merrymercy @ZYHowell — could one of you take a look and merge it if you are happy? Nothing else is waiting on our side. |
|
@ZYHowell this branch reads CONFLICTING against Suggested resolution: keep both, main's method first, then this PR's block unchanged. def free_group_end(self) -> None:
pending, self.free_page_reps_group = self.free_page_reps_group, None
super().free_group_end()
if pending:
reps = torch.cat(pending)
self.free(reps, _pages=reps // self.page_size)
def _set_capacity(
self, max_slots: int, *, virtual_num_pages: Optional[int] = None
) -> None:
"""Set active ranges while retaining the captured v2p/p2v storage."""
num_pages = int(max_slots) // self.page_size
num_virtual_ids = (
num_pages if virtual_num_pages is None else int(virtual_num_pages)
)
self.num_pages = num_pages
self.num_virtual_ids = num_virtual_ids
self.max_slots = num_pages * self.page_size
self.size = self.max_slots
self.clear()
_pending_hicache_load_pages: _CapacityField[int] = _CapacityField()
def set_hicache_transfer_done_event(self, transfer_key: Hashable, event) -> None:
... # this PR's block, unchanged through free_physical
|
|
@ZYHowell gentle ping: #39479 still reads CONFLICTING against |
…ty add/add) Merges upstream main 4ab720e into yonghao/ump-hicache-physical-transfers. The only conflict was the end of MultiEndedAllocator in python/sglang/srt/mem_cache/allocator/unified_sub_pool.py, where main added _set_capacity (#41961) next to this branch's HiCache transfer block; both are kept. Pushed by the OSS sync bot on maintainer merrymercy's instruction.
|
@ZYHowell on maintainer @merrymercy's instruction I merged |
|
This port clears our bar on gate 20✓✗1 lint✓ build 9✓✗1✂3 A✓ B✓ C 0/10⏳ other✓ · accel ✗amd,npu,xpu @merrymercy @ZYHowell — could one of you take a look and merge it if you are happy? Nothing else is waiting on our side. |
|
/rerun-test test/registered/e2e/hicache/test_hicache_unified_memory.py test/registered/e2e/disaggregation/test_disaggregation_unified_memory_swa.py test/registered/e2e/disaggregation/test_disaggregation_unified_memory_tri.py |
|
Results for 🚀 |
ch-wan
left a comment
There was a problem hiding this comment.
Re-reviewed at e6487f4 against main 169e622.
Everything from the last round is addressed. Envelope selection is now opt-in (no backend or mori) both at startup and on runtime attach, the buffer-only check uses the resolved threshold, the shared-host votes are gone, and the unreachable handlers and dead code are removed. The commits added since then need another pass: SWA reload through physical reservations, transfer-done events, the write-through host refill and the load-back host lock.
Should fix
- Allocation now waits on the latest HiCache transfer-done events whenever a lazy hole exists. As far as I can tell, this makes a load-back batch's forward wait for its whole H2D. See the inline comment on
_alloc_bind_fast_or_slow. - Two new startup rejections for unified memory + HiCache:
--pp-size > 1andSGLANG_DISABLE_LAZY_COMPACTION=1. See the inline comment inkv_cache_hook.py. _swa_allocation_callbacksturns off main's SWA load binding for every unified stack. That leavesbind_swa_for_loaded_rows/free_swaand the newbinds_swa_to_fullbranch dead. See the inline comment.- The PR description no longer matches the code:
- It says the PR reuses main's virtual-owner SWA allocation and removes the allocator reservation/event bookkeeping. b55790f adds
alloc_physical/cancel_physical_reservation/free_physical,_pending_hicache_load_pagesandset_hicache_transfer_done_event, and bypassesbind_swa_for_loaded_rows. - It says buffer_only uses shared arenas for every backend except Mooncake, with rank consensus. Now only no backend or mori qualifies, and the consensus collectives were removed.
- The "known inherited limitation" for cache-mode Mooncake no longer applies.
_uses_unified_page_envelope_hostnow requiressupports_page_envelope_hostin every mode. So in cache mode, file/sim/shm/mooncake move from the envelope to separate pools, and attaching one of them at runtime to an envelope arena is rejected. Main accepts that attach. - It says the PP restrictions were removed, but
kv_cache_hook.pynow adds one. - These changes are not mentioned:
- Under write-through,
insert_hostnow refills host rows on device-only nodes. This is on the default HiCache + storage path. PrefillAddernow takes a host lock and builds a FULL load-back spec for every load-back, including non-unified HiCache.evict_to_free_tokensswitches fromevict_for_alloctoevict. This also affects unified memory without HiCache.- Host-pool retraction is now allowed for unified hybrid-SWA and for the decode radix cache.
- Attach now runs
is_write_through_compatiblebefore switching from write_back to write_through.
- Under write-through,
- It says the PR reuses main's virtual-owner SWA allocation and removes the allocator reservation/event bookkeeping. b55790f adds
- The new mechanisms have no permanent tests. That covers physical reserve/cancel with the
_pending_hicache_load_pagesmove block, the transfer-done event wait, the load-back host lock andprepare_load_back, the write-policy switch check, and the per-backend envelope selection from the last round. The changed tests mostly adapt mocks.
Cleanup / nits
- The PR's
backup_kvSWA barrier overlaps the barrier main added in #39627, which now runs first. See the inline comment inswa.rs.
| if self.lazy_compaction and self._free_phys_pages.numel() > 0: | ||
| self._wait_hicache_transfers() |
There was a problem hiding this comment.
start_loading records the load's finish_event through set_hicache_transfer_done_event (cache_controller.py:993). In _get_new_batch_prefill_raw, ready_to_load_host_cache() runs before prepare_for_extend(). So the batch's own extend allocation reaches this branch whenever either sub-pool has a lazy hole. _wait_hicache_transfers then makes the schedule stream wait on that event, and run_batch calls forward_stream.wait_stream(self.schedule_stream). As far as I can tell, the forward of every load-back batch then waits for the whole H2D, instead of waiting layer by layer through layer_done_counter. The next allocation waits on the write event from start_writing in the same way.
Scenario: --enable-unified-memory --enable-hierarchical-cache (lazy compaction is the default), with some freed pages not yet compacted. Every step with a load-back loses the H2D/compute overlap.
Could you confirm this, or share numbers showing it doesn't matter? If some path really needs the wait, for example reusing a page that is freed while a transfer still targets it, could the event be recorded only for pages freed while a transfer is in flight? _pending_reuse already does this for forward events. That would avoid waiting on every outstanding transfer.
There was a problem hiding this comment.
Confirmed, your reading is right. start_loading registers the load's finish event before the batch's own extend allocation. When either sub-pool had a lazy hole, _alloc_bind_fast_or_slow / alloc_physical then called _wait_hicache_transfers(). That made the schedule stream, and through forward_stream.wait_stream(schedule_stream) the forward too, wait for the whole H2D and for any outstanding D2H.
#42708 removes that wait from both allocation paths instead of narrowing it to pages freed during a transfer. No page that an in-flight copy uses can be on the free list:
- Write-through and buffer-only D2H sources keep a device lock until the ack has synchronized the copy's finish event. Write-back drains its D2H (
writing_check(write_back=True)) before the demote frees anything. - An H2D target stays owned until its ack:
- cache mode: by the load-back's locks, until
loading_check; - buffer-only: by the request lock on the published span; redundant SWA destinations are released by
free_physicalat the ack.
- cache mode: by the load-back's locks, until
- Page moves stay blocked:
- for an uncommitted reservation, by
_pending_hicache_load_pages; - for a queued or unacknowledged transfer, by the host-transfer move gate.
- for an uncommitted reservation, by
_flush, the eagerfreepath andfree_physicalstill wait on the registered events.
One gap in that protection is fixed too. #39479 counted reservations in _pending_hicache_load_pages only on lazy pools. The SWA pool of the FULL/SWA/Mamba tri-pool floats and is not lazy, so a Mamba allocation in the same load could move the reserved SWA page before the load was queued, and the queued copy kept the old address. The follow-up counts the floating pool's reservations as well, until the load is queued or cancelled. Such a load now fails and rolls back unless evicting a checkpoint frees a slot.
test/registered/unit/mem_cache/test_unified_hicache_transfer_lifecycle.py checks this on CPU, with a stand-in that records the stream calls:
- hole reuse and
alloc_physicaladd no wait (these tests fail on current main); - moves and frees wait first;
- the rows a copy uses are neither reused nor moved, including a floating SWA reservation in the tri-pool;
- each layer's attention read still waits on its own load event.
These are CPU ordering checks, not GPU measurements. We have not measured the overlap or any speedup on a GPU yet. One note for that measurement, about the shared page-envelope host pools (no backend, or mori):
- Each component's envelope is copied once, at that component's first mapped layer. FULL and SWA are copied separately, not all before layer 0.
- So nothing within an envelope is copied layer by layer, but a later component's envelope copy can still overlap earlier layers.
- With separate host pools, layer-by-layer overlap is possible within a component too.
| if cfg.enable_hierarchical_cache: | ||
| assert cfg.pp_size == 1, ( | ||
| "--enable-unified-memory with hierarchical cache does not support " | ||
| "pipeline parallelism (--pp-size > 1)." | ||
| ) | ||
| assert not envs.SGLANG_DISABLE_LAZY_COMPACTION.get(), ( | ||
| "--enable-unified-memory with hierarchical cache requires lazy " | ||
| "compaction so pending H2D physical reservations remain stable." | ||
| ) |
There was a problem hiding this comment.
On main, handle_unified_memory_pool requires pp_size == 1 and lazy compaction only under PD disaggregation. Here both become startup assertions for any unified-memory + HiCache run. --enable-unified-memory --enable-hierarchical-cache --pp-size 2 now fails to launch, and so does the same command with SGLANG_DISABLE_LAZY_COMPACTION=1, the A/B and rollback escape hatch. Main starts both.
The lazy-compaction assert comes from the new unbound alloc_physical reservations. Main's bind-at-allocation path has no equivalent guard. The PP assert has no stated reason, and the description says the PP restrictions were removed.
Could you either keep main's binding path for these configurations, or give the reason and list both as trade-offs in the description?
There was a problem hiding this comment.
Agreed. #42708 removes both asserts for unified memory with HiCache; the PD disaggregation asserts stay. test/registered/unit/server_args/test_unified_hicache_startup_args.py checks that --pp-size 2 and SGLANG_DISABLE_LAZY_COMPACTION=1 pass the argument layer. That is only an argument-layer check: pipeline-parallel serving with unified memory and HiCache has not been validated.
One correction on the lazy-compaction case. Main before #39479 also stopped at startup with SGLANG_DISABLE_LAZY_COMPACTION=1, just later:
- The scheduler installs the host-transfer move gate for unified memory with HiCache (
python/sglang/srt/managers/scheduler.pyL598–L609 at 6cc661f, the parent of the Fix unified HiCache physical transfers #39479 merge). install_move_gateasserts lazy compaction (python/sglang/srt/mem_cache/allocator/unified_sub_pool.pyL242–L261, same commit).SGLANG_DISABLE_LAZY_COMPACTION=1turns lazy compaction off (python/sglang/srt/mem_cache/kv_cache_configurator.pyL149–L152).
That allocator-level check predates #39479, and the follow-up leaves it unchanged.
| if isinstance(allocator, MultiEndedAllocator): | ||
| return dict( | ||
| device_alloc_fn=allocator.alloc_physical, | ||
| device_free_fn=allocator.cancel_physical_reservation, | ||
| ) |
There was a problem hiding this comment.
Every unified SWA stack builds its SWA sub-allocator as a MultiEndedAllocator or FloatMultiEndedAllocator, so this branch always wins. As a result:
bind/free_bound(bind_swa_for_loaded_rows/free_swa, passed at :504 and :1344) are never installed, and no entry getsdevice_indices_from_anchor_fn.- The anchor loop in
_resolve_device_transfersis unreachable. - The
binds_swa_to_fullbranch that 6027473 just added toBufferModePipeline(pipeline.py:1295-1330 and :1367) is unreachable too, so buffer-only staged loads quietly take theelif swa_dev is not Nonepath instead.
Could you keep one mechanism? Either drop the bind path, the swa_indices_from_anchor_fn/swa_free_from_anchor_fn plumbing and the binds_swa_to_full branch, or keep main's binding and drop the physical reservation.
There was a problem hiding this comment.
Done in #42708. There is one mechanism now: physical reservation. The follow-up removes:
bind_swa_for_loaded_rows;- the
device_indices_from_anchor_fn/swa_indices_from_anchor_fn/swa_free_from_anchor_fnplumbing; anchor_index_partsand its producers (Python core and Rust adapter);- the anchor loop in
_resolve_device_transfers; - the
binds_swa_to_fullbranch inBufferModePipeline.
free_swa stays, because free_swa_segment still uses it when page_size == 1.
New tests cover the physical path (test/registered/unit/mem_cache/test_unified_hicache_transfer_lifecycle.py):
- cache-mode load-back reserves only the missing window rows and binds them;
- a failed reservation rolls back the whole load;
- a later-pool failure cancels the reservation;
- a buffer-only staged load reserves the whole window and frees the rows of resident pages at the ack.
| if tree_core.is_write_back | ||
| && tree_core.swa_write_back_eviction_barrier_enabled | ||
| && tree_core.component_state(SWA).evict_device_last_backup != Some(x) | ||
| && !tree_core.arena.node(x).backuped() | ||
| { | ||
| // A later Full backup cannot recover SWA data after this | ||
| // internal node is tombstoned. Pause on the same cursor so | ||
| // the Controller can preserve the dirty path first. | ||
| tree_core.component_state_mut(SWA).evict_device_backup_node = Some(x); | ||
| tree_core.component_state_mut(SWA).evict_device_last_backup = Some(x); | ||
| cursor = Some(x); | ||
| break 'step None; |
There was a problem hiding this comment.
After the merge with #39627, main's SWA write-back barrier just above runs first. It already pauses on every internal node that has FULL on device and neither FULL nor SWA on host. As far as I can tell, that leaves this barrier firing only when SWA already has a host copy and FULL does not. Tombstoning device SWA loses nothing in that case, yet the controller still runs a FULL write-back for the node through backup_kv (_evict_device_next_node in unified_radix_cache.py).
Could you drop this barrier and its plumbing, or say which case it covers that main's barrier doesn't? The plumbing is swa_write_back_eviction_barrier_enabled, evict_device_backup_node/evict_device_last_backup, EvictDeviceNextNodeResult.backup_kv and the retry loop in _evict_device_next_node.
There was a problem hiding this comment.
Agreed. The barrier fired only when SWA already had a host copy, or had nothing to back up, and then ran a FULL write-back that nothing needed. #42708 drops the barrier and its plumbing:
swa_write_back_eviction_barrier_enabledand its enable call;evict_device_backup_node/evict_device_last_backup;EvictDeviceNextNodeResult.backup_kv;- the retry loop in
_evict_device_next_node.
The barrier from PR 39627 still backs up an unbacked window first. The host-pressure test keeps its four cases (Rust/session × pinned/unpinned). New tests check three cases:
- a hosted window is dropped with no D2H. This test is pinned to the Rust TreeCore, where the barrier lived, and reports a skip when that core cannot load; it fails on current main. A twin runs the same check on the Python TreeCore;
- an unbacked window is backed up first;
- under host pressure only the window is dropped.
|
Re: review, the bullets on Eviction: The FULL/SWA reclaim plan returns joint eviction quotas, which are passed together in one Write-through host refill: This is a separate cache-mode behavior change, not a requirement for unified-memory buffer-only operation. |
|
Following up on the remaining points of the review #39479 (review):
The follow-up is #42708. It is open and not merged. "The new mechanisms have no permanent tests." The follow-up adds CPU-registered tests that run the real allocator, cache, controller and page-envelope host pool, and check row contents as well as counts:
The files are:
The host lock stays. It covers a different interval from load-back's own locks: from
With the lock removed, reclaim evicts the matched SWA host copy and the load-back fails. Host-pool retraction for unified hybrid-SWA and the decode radix cache. The behavior is unchanged.
The write_back→write_through attach check. It stays. It stops a write-back host tree that write-through cannot hold from being switched to write-through: auxiliary host rows without their FULL rows, or a host suffix below a parent whose FULL rows are not on host. It does not apply to buffer-only mode. The tests check that:
Validation. CPU only: torch 2.8 with the Python TreeCore, and torch 2.11 with the Rust TreeCore built from the branch. No GPU, serving or performance runs. |
|
I've corrected this PR's description so that it matches the merged code, as the review asked. Each correction is marked inline. Struck text is what the description originally said, Corrected is how #39479 behaved as merged, and Follow-up is what #42708 changes. In short:
|
The review of #39479 asked for permanent tests of five mechanisms: physical reserve/cancel with the _pending_hicache_load_pages move block, the transfer-done event wait, the load-back host lock and prepare_load_back, the write-policy switch check, and the per-backend envelope selection. Keep a short behavior test or two for each, and the tests that replace base coverage the bind path took with it. Delete the rest, and the support code only they used. test_unified_hicache_transfer_lifecycle.py keeps 11 of its 37 tests: - reserve/cancel and the move block: test_reservation_is_counted_unbound_and_released_by_cancel, test_pending_reservation_blocks_moves_until_its_load_is_queued, and, for the floating SWA pool whose reservations did not block moves, test_state_allocation_cannot_move_a_reservation_before_its_load_is_queued; - the transfer-done event wait: test_hole_reuse_does_not_wait_for_unrelated_copies and test_moves_and_frees_wait_for_every_copy_first, which no longer fixes the order of the two waits, only that both come before the move; - the load-back host lock and prepare_load_back: test_host_match_stays_pinned_while_reclaim_evicts_host_leaves and test_shared_budget_reclaims_for_the_load_and_pending_demand; - the write-policy switch check: test_incompatible_host_trees_are_rejected_without_side_effects; - SWA load-back into physical reservations, in place of the removed TestSwaLoadAllocation cases: test_only_the_missing_window_rows_are_reserved_and_bound, test_reservation_retries_once_after_device_eviction (now only the retry itself) and test_later_pool_failure_cancels_the_swa_reservation. TestUnifiedPageEnvelopeSelection keeps test_only_no_backend_and_mori_select_the_envelope, which now reads the backend choices from the parser. The startup-args file keeps its two argument-layer tests and drops the PD one.
Keep the eviction-headroom reserve from this PR while adopting upstream sgl-project#39479, which switched the unified hybrid-SWA call site from evict_for_alloc() to evict() so cumulative reclaim quotas are fully honored. Update the unit-test assertion accordingly and refresh the mocked tree-cache guard renamed by sgl-project#42362 (is_chunk_cache -> supports_prefix_sharing).
Original PR: #37496 — previous reviews and comments.
Motivation
Unified HiCache must preserve index ownership and address lifetimes while device or shared host pages can move. Main now provides per-pool L2 translation, SWA load binding, device move gates, and Python SWA write-back demotion. This PR retains the remaining shared-host buffer and storage support while using those upstream contracts.
The preceding #39478 has merged. This PR now targets
main; the conflict resolution preserves the existing head history and incorporates mainf404db94d425b1a10dd4954e9f9a70d6ecb65402.Modifications
virtual-owner SWA allocation/rollback,independent host-transfer move gate, deferred D2H batching, and sparse layer-ID span. Remove duplicate eager physical translationand allocator reservation/event bookkeeping.alloc_physical/cancel_physical_reservation/free_physical), not main's virtual-owner SWA binding. The allocator keeps that bookkeeping:set_hicache_transfer_done_eventrecords the transfer-done events, and_pending_hicache_load_pagesblocks page moves until the load is queued. That count covered only lazy pools. In the FULL/SWA/Mamba tri-pool the SWA pool floats and is not lazy, so its reservations were not counted. A Mamba allocation could then move a reserved SWA page before its load was queued, while the queued copy kept the old address. Main's binding path stayed in the code but became unreachable._alloc_bind_fast_or_slow,alloc_physical) also waited on the registered transfer-done events whenever a lazy hole existed.and host-pool selection, removing the original extra storage-backend allowlist{None, file, sim, mori, shm}along with the superseded broad model/PP/I/O restrictions. Retain the direct external-cache-linker guard because that path bypasses the L2 translation/lifetime contract.mori, selects the shared page envelope, in every host memory mode.file,sim,shm,mooncakeand the other backends use separate host pools, in cache mode too.--pp-size 1, and lazy compaction (SGLANG_DISABLE_LAZY_COMPACTIONunset).buffer_onlyfor eligible two-ended layoutsexcept Mooncake, with byte-based staging admission, load reserve, occupancy, atomic prefetch allocation,rank consensus,and rollback.buffer_onlythat ismori. The shared-host allocation consensus collectives were removed before merge. Other consensus checks are unchanged.--hicache-storage-backend mooncakeselects compatible pools.kernel/path.Known inherited limitation: main already permits Mooncake in cache mode and selects a shared envelope for eligible FULL/SWA pools, but Mooncake's MHA metadata still expects two K/V pointers per page while that envelope supplies one. This raises even withpage_first; an uncaught backup-worker exception can prevent ACK/release completion. This PR preserves that upstream cache-mode behavior and does not claim it is supported. The buffer-only fallback prevents the newly added shared-buffer support from extending the mismatch to main's working separate-pool path. Switching an existing shared buffer to Mooncake at runtime requires restart. Other untested backend/layout combinations are not claimed compatible.Corrected: the limitation above does not apply to the merged code. In cache mode, Mooncake selects separate host pools, not the shared envelope. Other untested backend/layout combinations are still not claimed compatible.
Corrected: these merged behavior changes were not listed above:
insert_hostkeeps host slots already filled by an L3 prefetch when the matching node has device KV but no host copy. This is a cache-mode change and starts no extra D2H copy. Buffer-only completion takes the staging path beforeinsert_host. See Fix unified HiCache physical transfers #39479 (comment).PrefillAddertakes a host lock on the selected host match from beforeprepare_load_back's reclaim untilinit_load_backholds its own locks, and builds a FULL load-back spec. Follow-up: the FULL spec is built only forSharedSWAPrefillBudget, the budget that reads it. The host lock stays.evict_to_free_tokenspasses the FULL/SWA reclaim quotas together to oneevict(EvictParams(...))call instead ofevict_for_alloc. This also applies to unified memory without HiCache. See the comment linked above.is_write_through_compatible. Host trees that write-through cannot hold are rejected before any policy change. Buffer-only mode is not checked.Accuracy Tests
No model accuracy or serving runs were performed. Focused CPU contract tests passed, including the actual compiled Rust extension. Existing host-pressure/index-domain expectations were updated to the main contracts; no new permanent test cases were added.
Follow-up: #42708 adds CPU regression tests for these mechanisms: physical reservations (including the tri-pool's floating SWA pool), allocation and transfer order, the transfer lifecycles, load-back admission, the write-policy attach check, envelope selection for every storage backend choice, host-pool retraction (including decode radix retraction with a shared prefix) and SWA write-back eviction. CPU only; no GPU, serving or performance runs.
Speed Tests and Profiling
No performance runs or speedup claims. Main's move gate can defer compaction until transfers are acknowledged; main's D2H batching is retained.
Test plan
git diff --checkagainst main passed.moriis absent in the existing environment. GPU/serving, storage backend IO, multi-rank/multi-node PD, and performance/accuracy tests were not run.Checklist
Review and Merge Process
Prior discussion remains on the original PR linked above. This conflict fix preserves the current PR and branch history; review and CI must be evaluated against its new head.
CI States
Latest PR Test (Base): 🚫 Run #37174694531
Latest PR Test (Extra): ❌ Run #37174694422
Latest PR Test (AMD ROCm 10): ❌ Run #37174694602