Repository navigation
Support post-capture KV sizing for the unified hybrid-SWA pool - #41961
Conversation
|
This port clears our bar on gate 18✓✗1✂3 lint✓ build✓ A✓ B✓ C 0/10⏳ other 6✓✂5 · accel ✗amd,npu,xpu @merrymercy @ZYHowell — could one of you take a look and merge it if you are happy? Nothing else is waiting on our side. |
Post-capture KV sizing reserves the KV pool as CUDA virtual memory, captures the CUDA graphs, and only then backs the pool from measured free memory. The unified memory pool was excluded: its shared byte arena was a plain `torch.empty` that is fully backed before capture and cannot change size afterwards, so `post_capture_kv_sizing_planned` returned False for every `--enable-unified-memory` configuration and those servers kept paying the pre-capture activation reserve. Give the unified hybrid-SWA byte pool the same lifecycle: - `UnifiedKVPool` reserves its arena through `KvVmmBufferOwner` when `post_capture_active` is set, backs only the slot-0 dummy-write sink before capture, and `finalize_backing(config)` backs the final `unified_memory_pool_bytes` in place: the raw tensor the graphs captured keeps its address, the per-sub-pool slot bounds are re-derived from the final byte count, and the bs=1 feasibility floor is checked again against it. `active_allocation_bytes` is the backed prefix, so transfer engines register that rather than the reserved upper bound. - `UnifiedSWATokenToKVPoolAllocator.resize(config)` re-derives the capacities of both sub-allocators and both sub-pools after the pool is finalized, through `MultiEndedAllocator._set_capacity`, which shrinks the active page range while keeping the captured v2p/p2v tables. - The planner enables the path for unified hybrid-SWA models only. The Mamba pools and the independent draft pools keep their pre-capture sizing: they are budgeted separately and have no resize path, so they are still excluded. - `PostCaptureKVResize` carries `unified_memory_pool_bytes` and the model runner mirrors it into `memory_pool_config`. Tests: a 1-GPU unit test builds the reserved pool, finalizes it to a smaller budget, and checks every capacity the pool and allocator expose against the eagerly-built pool of that size, with the tensor address unchanged and the resize contract (finalize first, empty allocator, budget within the reservation and above the bs=1 floor) enforced; a CPU case pins `_set_capacity` bookkeeping on both grow directions. Co-authored-by: ZYHowell <ZYHowell@users.noreply.github.com>
cde5f10 to
749a6fb
Compare
|
This port clears our bar on gate✓ lint✓ build 3✓✗1✂3 A✓ B✓ C 0/10⏳ other✓ · accel ✗amd,npu @merrymercy @ZYHowell — could one of you take a look and merge it if you are happy? Nothing else is waiting on our side. |
…ty add/add) Merges upstream main 4ab720e into yonghao/ump-hicache-physical-transfers. The only conflict was the end of MultiEndedAllocator in python/sglang/srt/mem_cache/allocator/unified_sub_pool.py, where main added _set_capacity (#41961) next to this branch's HiCache transfer block; both are kept. Pushed by the OSS sync bot on maintainer merrymercy's instruction.
Motivation
Post-capture KV sizing (
SGLANG_ENABLE_POST_CAPTURE_KV_SIZING) reserves the KV pool as CUDA virtual memory, captures the CUDA graphs, and only then backs the pool from measured free memory.--enable-unified-memorywas excluded from it: the unified pool's shared byte arena was a plaintorch.emptythat is fully backed before capture and cannot change size afterwards, sopost_capture_kv_sizing_plannedreturnedFalsefor every unified configuration and those servers kept paying the pre-capture activation reserve.This gives the unified hybrid-SWA byte pool the same reserve → capture → back-in-place lifecycle the other pools already have. The final size is the shared byte budget from #36729 (
unified_memory_pool_bytes), so the post-capture solver needs no new sizing rule.Modifications
UnifiedKVPoolreserves its arena throughKvVmmBufferOwnerwhenpost_capture_activeis set and backs only the slot-0 dummy-write sink before capture (page-major slot 0 spans the largest page envelope, so the whole reserved floor is backed).finalize_backing(config)backs the finalunified_memory_pool_bytesin place: the raw tensor the graphs captured keeps its address, the per-sub-pool slot bounds are re-derived from the final byte count, and the bs=1 feasibility floor is checked again against it.active_allocation_bytesis the backed prefix, andget_contiguous_buf_infosreports it instead of the reserved upper bound, so transfer engines never register unbacked address space.UnifiedSWATokenToKVPoolAllocator.resize(config)re-derives both sub-allocators' and both sub-pools' capacities after the pool is finalized, viaMultiEndedAllocator._set_capacity, which shrinks the active page range while keeping the captured v2p/p2v tables. It refuses to run beforefinalize_backingor on a non-empty allocator.post_capture_kv_sizing_plannedplans the path for unified hybrid-SWA models only. The Mamba pools and the independent draft pools (EAGLE / standalone / DFlash) keep their pre-capture sizing: they are budgeted separately and have no resize path, so they stay excluded exactly as before.PostCaptureKVResizecarriesunified_memory_pool_bytes, andModelRunner.post_capture_resize_kv_poolmirrors it intomemory_pool_config.post_capture_activeintoinit_unified_swa_pools.Out of scope: unified Mamba / Mamba+SWA pools, draft-worker pools, and the HiSparse/DSV4 unified path, which all keep returning
Falsefrom the planner.Accuracy Tests
Not applicable: this changes how the pool's bytes are reserved and backed, not model math. The existing end-to-end guard
test/registered/mem_cache/test_post_capture_kv_sizing.pykeeps covering the non-unified path.Speed Tests and Profiling
Not run; no kernel changes. The unified pool is now backed after capture, the same as the non-unified post-capture path.
Test Plan
New tests:
test/registered/unit/mem_cache/test_unified_post_capture_sizing.py(1 GPU, base-b 1-gpu-small): builds the reserved pool, checks that only the sink is backed and thatget_contiguous_buf_infosreports the backed prefix, finalizes it to a smaller budget, and compares every capacity the pool, sub-pools and allocators expose against an eagerly-built pool of that size, with the raw tensor address and the v2p/p2v table storage unchanged, the full span writable, byte accounting clean and the whole capacity allocatable. A third case pins the contract: resize before finalize, a missing byte count, a budget above the reservation or below the bs=1 floor, resize with live allocations, and finalize/resize on an eager pool are all refused.test/registered/unit/mem_cache/test_multi_ended_allocator.py::TestMultiEndedAllocator::test_set_capacity_keeps_tables_and_shrinks_the_active_range(CPU):_set_capacitykeeps the mapping tables' storage and shrinks the active range, free-id range and allocatable span on both grow directions.Run locally on one CUDA GPU (Blackwell):
Lint:
ruff check --select=F401,F821,F823,UP037,ruff format --check(ruff 0.15.1),isort --check-only(7.0.0),codespell(2.4.1) andscripts/lint/check_registered_tests.pypass on the changed files.Not run: full-model serving with
--enable-unified-memoryandSGLANG_ENABLE_POST_CAPTURE_KV_SIZING=1, PD disaggregation, and HiCache on top of the resized unified pool.Checklist
CI States
Latest PR Test (Base): ✅ Run #36947163525⚠️ Not enabled -- add
Latest PR Test (Extra):
run-ci-extralabel to opt in.Latest PR Test (AMD ROCm 10): ❌ Run #36947163506