Skip to content

Support post-capture KV sizing for the unified hybrid-SWA pool - #41961

Merged
merrymercy merged 1 commit into
mainfrom
oss-sync-6a2f1214e8
Oct 3, 2026
Merged

merrymercy merged 1 commit into
mainfrom
oss-sync-6a2f1214e8

Conversation

@metamergebot

@metamergebot metamergebot commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

Post-capture KV sizing (SGLANG_ENABLE_POST_CAPTURE_KV_SIZING) reserves the KV pool as CUDA virtual memory, captures the CUDA graphs, and only then backs the pool from measured free memory. --enable-unified-memory was excluded from it: the unified pool's shared byte arena was a plain torch.empty that is fully backed before capture and cannot change size afterwards, so post_capture_kv_sizing_planned returned False for every unified configuration and those servers kept paying the pre-capture activation reserve.

This gives the unified hybrid-SWA byte pool the same reserve → capture → back-in-place lifecycle the other pools already have. The final size is the shared byte budget from #36729 (unified_memory_pool_bytes), so the post-capture solver needs no new sizing rule.

Modifications

  • UnifiedKVPool reserves its arena through KvVmmBufferOwner when post_capture_active is set and backs only the slot-0 dummy-write sink before capture (page-major slot 0 spans the largest page envelope, so the whole reserved floor is backed). finalize_backing(config) backs the final unified_memory_pool_bytes in place: the raw tensor the graphs captured keeps its address, the per-sub-pool slot bounds are re-derived from the final byte count, and the bs=1 feasibility floor is checked again against it. active_allocation_bytes is the backed prefix, and get_contiguous_buf_infos reports it instead of the reserved upper bound, so transfer engines never register unbacked address space.
  • UnifiedSWATokenToKVPoolAllocator.resize(config) re-derives both sub-allocators' and both sub-pools' capacities after the pool is finalized, via MultiEndedAllocator._set_capacity, which shrinks the active page range while keeping the captured v2p/p2v tables. It refuses to run before finalize_backing or on a non-empty allocator.
  • post_capture_kv_sizing_planned plans the path for unified hybrid-SWA models only. The Mamba pools and the independent draft pools (EAGLE / standalone / DFlash) keep their pre-capture sizing: they are budgeted separately and have no resize path, so they stay excluded exactly as before.
  • PostCaptureKVResize carries unified_memory_pool_bytes, and ModelRunner.post_capture_resize_kv_pool mirrors it into memory_pool_config.
  • The KV configurator passes post_capture_active into init_unified_swa_pools.

Out of scope: unified Mamba / Mamba+SWA pools, draft-worker pools, and the HiSparse/DSV4 unified path, which all keep returning False from the planner.

Accuracy Tests

Not applicable: this changes how the pool's bytes are reserved and backed, not model math. The existing end-to-end guard test/registered/mem_cache/test_post_capture_kv_sizing.py keeps covering the non-unified path.

Speed Tests and Profiling

Not run; no kernel changes. The unified pool is now backed after capture, the same as the non-unified post-capture path.

Test Plan

New tests:

  • test/registered/unit/mem_cache/test_unified_post_capture_sizing.py (1 GPU, base-b 1-gpu-small): builds the reserved pool, checks that only the sink is backed and that get_contiguous_buf_infos reports the backed prefix, finalizes it to a smaller budget, and compares every capacity the pool, sub-pools and allocators expose against an eagerly-built pool of that size, with the raw tensor address and the v2p/p2v table storage unchanged, the full span writable, byte accounting clean and the whole capacity allocatable. A third case pins the contract: resize before finalize, a missing byte count, a budget above the reservation or below the bs=1 floor, resize with live allocations, and finalize/resize on an eager pool are all refused.
  • test/registered/unit/mem_cache/test_multi_ended_allocator.py::TestMultiEndedAllocator::test_set_capacity_keeps_tables_and_shrinks_the_active_range (CPU): _set_capacity keeps the mapping tables' storage and shrinks the active range, free-id range and allocatable span on both grow directions.

Run locally on one CUDA GPU (Blackwell):

python -m pytest -q test/registered/unit/mem_cache/test_unified_post_capture_sizing.py
# 3 passed, 6 subtests passed

python -m pytest -q \
  test/registered/unit/mem_cache/test_multi_ended_allocator.py \
  test/registered/unit/mem_cache/test_unified_byte_budget_sizing.py \
  test/registered/unit/mem_cache/test_unified_swa_shared_virtual_ids.py \
  test/registered/unit/mem_cache/test_unified_tri_pool.py \
  test/registered/unit/mem_cache/test_prefill_memory_budget.py \
  test/registered/unit/model_executor/test_pool_configurator.py \
  test/registered/unit/model_executor/test_kv_canary_headroom.py \
  test/registered/unit/disaggregation/test_mooncake_efa_allocator.py
# 227 passed, 776 subtests passed

Lint: ruff check --select=F401,F821,F823,UP037, ruff format --check (ruff 0.15.1), isort --check-only (7.0.0), codespell (2.4.1) and scripts/lint/check_registered_tests.py pass on the changed files.

Not run: full-model serving with --enable-unified-memory and SGLANG_ENABLE_POST_CAPTURE_KV_SIZING=1, PD disaggregation, and HiCache on top of the resized unified pool.

Checklist


CI States

Latest PR Test (Base): ✅ Run #36947163525
Latest PR Test (Extra): ⚠️ Not enabled -- add run-ci-extra label to opt in.
Latest PR Test (AMD ROCm 10): ❌ Run #36947163506

@metamergebot

metamergebot commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator Author

This port clears our bar on cde5f1075a4acd1ea3e87574a8d2bc1f8f80f66e: every base-a and base-b job is green.
base-c and the accelerator suites are reported but not gated, because they run
on scarce hardware and upstream merges through them routinely.

gate 18✓✗1✂3 lint✓ build✓ A✓ B✓ C 0/10⏳ other 6✓✂5 · accel ✗amd,npu,xpu

@merrymercy @ZYHowell — could one of you take a look and merge it if you are happy? Nothing else is waiting on our side.

Post-capture KV sizing reserves the KV pool as CUDA virtual memory, captures
the CUDA graphs, and only then backs the pool from measured free memory. The
unified memory pool was excluded: its shared byte arena was a plain
`torch.empty` that is fully backed before capture and cannot change size
afterwards, so `post_capture_kv_sizing_planned` returned False for every
`--enable-unified-memory` configuration and those servers kept paying the
pre-capture activation reserve.

Give the unified hybrid-SWA byte pool the same lifecycle:

- `UnifiedKVPool` reserves its arena through `KvVmmBufferOwner` when
  `post_capture_active` is set, backs only the slot-0 dummy-write sink before
  capture, and `finalize_backing(config)` backs the final
  `unified_memory_pool_bytes` in place: the raw tensor the graphs captured
  keeps its address, the per-sub-pool slot bounds are re-derived from the final
  byte count, and the bs=1 feasibility floor is checked again against it.
  `active_allocation_bytes` is the backed prefix, so transfer engines register
  that rather than the reserved upper bound.
- `UnifiedSWATokenToKVPoolAllocator.resize(config)` re-derives the capacities
  of both sub-allocators and both sub-pools after the pool is finalized,
  through `MultiEndedAllocator._set_capacity`, which shrinks the active page
  range while keeping the captured v2p/p2v tables.
- The planner enables the path for unified hybrid-SWA models only. The Mamba
  pools and the independent draft pools keep their pre-capture sizing: they are
  budgeted separately and have no resize path, so they are still excluded.
- `PostCaptureKVResize` carries `unified_memory_pool_bytes` and the model
  runner mirrors it into `memory_pool_config`.

Tests: a 1-GPU unit test builds the reserved pool, finalizes it to a smaller
budget, and checks every capacity the pool and allocator expose against the
eagerly-built pool of that size, with the tensor address unchanged and the
resize contract (finalize first, empty allocator, budget within the
reservation and above the bs=1 floor) enforced; a CPU case pins
`_set_capacity` bookkeeping on both grow directions.

Co-authored-by: ZYHowell <ZYHowell@users.noreply.github.com>
@metamergebot

metamergebot commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator Author

This port clears our bar on 749a6fb07166337ddc5f02c1e9c13045b504559a: every base-a and base-b job is green.
base-c and the accelerator suites are reported but not gated, because they run
on scarce hardware and upstream merges through them routinely.

gate✓ lint✓ build 3✓✗1✂3 A✓ B✓ C 0/10⏳ other✓ · accel ✗amd,npu

@merrymercy @ZYHowell — could one of you take a look and merge it if you are happy? Nothing else is waiting on our side.

@metamergebot metamergebot added the bypass-fail-fast CI: a failing job no longer aborts its siblings (lint still gates) label Oct 3, 2026
@merrymercy
merrymercy merged commit b016ca4 into main Oct 3, 2026
423 of 498 checks passed
@merrymercy
merrymercy deleted the oss-sync-6a2f1214e8 branch October 3, 2026 16:54
@metamergebot metamergebot mentioned this pull request Oct 3, 2026
4 of 5 tasks
metamergebot added a commit that referenced this pull request Oct 4, 2026
…ty add/add)

Merges upstream main 4ab720e into yonghao/ump-hicache-physical-transfers.
The only conflict was the end of MultiEndedAllocator in
python/sglang/srt/mem_cache/allocator/unified_sub_pool.py, where main added
_set_capacity (#41961) next to this branch's HiCache transfer block; both are
kept. Pushed by the OSS sync bot on maintainer merrymercy's instruction.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bypass-fail-fast CI: a failing job no longer aborts its siblings (lint still gates) memory-pool run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants