Skip to content

Support unified memory page-envelope transfers in PD - #39477

Merged
ch-wan merged 90 commits into
mainfrom
yonghao/ump-pd-page-envelopes
Sep 19, 2026
Merged

ch-wan merged 90 commits into
mainfrom
yonghao/ump-pd-page-envelopes

Conversation

@ZYHowell

@ZYHowell ZYHowell commented Sep 14, 2026 •

Copy link
Copy Markdown
Collaborator

Original PR: #36730 — previous reviews and comments.

Motivation

Unified memory pools expose virtual token IDs while disaggregated transfer backends operate on physical buffers. PD transfers need an explicit page-envelope contract and physical index translation, including distinct target and independent-draft index vectors when their layouts differ. Decode preallocation should charge SWA only for newly allocated tail pages, including the Mamba/SWA/FULL tri-pool layout.

Targets main, with merged #36729 supplied by the existing base. This review update preserves the current base and propagates through #39478 and #39479.

Modifications

  • Expose contiguous page envelopes for unified MHA/SWA pools and translate virtual IDs before transfer. Keep attention SWA IDs kernel-facing and transfer one-region unified SWA envelopes through Mooncake's flat path.
  • Gate compaction while a transfer can still reference physical pages; provide separate target/draft index vectors and preserve the CPU-copy fallback for decode retraction.
  • Share tail allocation in UnifiedSWAAllocatorBase. Price only newly bound SWA pages, derive their IDs from the allocation, and preserve an already-bound partial prefix page. Remove the duplicate two-pool override and its page-alignment assertions.
  • Support asymmetric FULL/SWA demand in the tri-pool capacity policy, retaining the existing float-band page-grid, compaction, and relocation rules. Use one page-demand predicate for tail admission and ordinary joint capacity.
  • Use self.token_to_kv_pool_allocator consistently in the decode reservation helpers.
  • Remove the inherited merge_and_sort_free() calls from unified paged allocation: that method only updates free_pages/release_pages, while these kernels allocate from free_virtual_ids and already reject requests with insufficient virtual pages. Physical free-page sorting remains in compaction; the removed calls cannot replenish or reorder the virtual IDs used by these kernels.

Accuracy Tests

Focused allocator tests verify tail binding, physical translations, prefix preservation and release accounting. Full model accuracy and multi-node PD serving were not rerun for this review update.

Speed Tests and Profiling

No serving benchmark or profiling run; no new performance result is claimed.

Test plan

Using Python 3.12 in /data/venvs/pd-kl-tier5 with local source on PYTHONPATH:

CUDA_VISIBLE_DEVICES=0 PYTHONPATH=python \
uv --no-cache run --no-project --no-sync --python /data/venvs/pd-kl-tier5/bin/python python -m pytest -q \
  test/registered/unit/mem_cache/test_unified_tri_pool.py

CUDA_VISIBLE_DEVICES=0 PYTHONPATH=python \
uv --no-cache run --no-project --no-sync --python /data/venvs/pd-kl-tier5/bin/python python -m pytest -q \
  test/registered/kernel/mem_cache/test_unified_swa_tail_allocation.py
  • Tri-pool: 45 tests and 80 subtests passed. Capacity, binding and compaction use real CPU allocators; the focused tail tests substitute the GPU extend kernel with equivalent empty-prefix allocation.
  • GPU tail allocation: 1 test with 16 subtests passed on GB300. Includes empty/short tails, capacity beyond the symmetric budget, unaligned tails, partial prefix-page reuse, no new pages, and release accounting.
  • Both test files were rerun together at the updated Fix unified HiCache physical transfers #39479 stack tip: 46 tests and 96 subtests passed.
  • Changed-file pre-commit checks passed before committing each layer, including formatting, lint and registered-test validation. git diff --check passed.
  • Full serving, multi-node/backend PD end-to-end tests, and remote CI completion are outside this focused validation.

Checklist

  • Format code with pre-commit.
  • Extend existing allocator tests to cover the review findings.
  • Document the inherited free-list cleanup and capacity behavior.
  • Full model accuracy and serving benchmarks for this PR.

Review and Merge Process

The original PR linked at the top contains prior reviews and comments. Review and CI status for this replacement PR must be evaluated separately.


CI States

Latest PR Test (Base): ❌ Run #35379101664
Latest PR Test (Extra): 🚫 Run #35379101178
Latest PR Test (AMD ROCm 10): ❌ Run #35379101701

yhzhuang and others added 30 commits September 1, 2026 15:04
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Co-authored-by: yhzhuang <yhzhuang@fb.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Minimize both component reclaim quotas under the joint capacity predicate so a blocked compaction path cannot turn one allocation shortfall into an all-SWA eviction.

Original prod_inference commit: 56c7082a92bc0c9024585387460e64b065652918

Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Co-authored-by: yhzhuang <yhzhuang@fb.com>
# Conflicts:
#	python/sglang/srt/mem_cache/common.py
# Conflicts:
#	python/sglang/srt/mem_cache/common.py
# Conflicts:
#	python/sglang/srt/mem_cache/common.py
…pacity

# Conflicts:
#	python/sglang/srt/managers/schedule_policy.py
#	python/sglang/srt/managers/scheduler_components/invariant_checker.py
#	python/sglang/srt/mem_cache/kv_cache_configurator.py
#	python/sglang/srt/mem_cache/multi_ended_allocator.py
#	python/sglang/srt/model_executor/pool_configurator.py
# Conflicts:
#	python/sglang/srt/disaggregation/base/conn.py
#	python/sglang/srt/mem_cache/kv_cache_configurator.py
Comment thread python/sglang/srt/mem_cache/allocator/unified_hybrid_swa.py Outdated
Comment thread python/sglang/srt/mem_cache/allocator/unified_hybrid_swa.py Outdated
Comment thread python/sglang/srt/disaggregation/decode.py Outdated
Comment thread python/sglang/srt/mem_cache/allocator/unified_sub_pool.py
@ZYHowell
ZYHowell requested a review from ch-wan September 16, 2026 18:56
Comment thread python/sglang/srt/mem_cache/allocator/unified_hybrid_swa.py Outdated
Comment thread python/sglang/srt/mem_cache/allocator/unified_hybrid_swa.py Outdated
Comment thread python/sglang/srt/disaggregation/decode.py Outdated
ch-wan and others added 6 commits September 17, 2026 22:14
`DecodePreallocQueue` carried both capacity algorithms and picked between
them with `isinstance(allocator, UnifiedSWATokenToKVPoolAllocator)`: a
per-side token comparison for pools whose sides own their own buffers, and a
shared-byte reservation for the unified layout. The scheduler had to
reconstruct the allocator's own accounting to ask the second question --
adding `full_available_size() + evictable - budget` back into the demand --
and naming one class meant the answer depended on which sibling a layout
happened to subclass.

Following `check_decode_capacity`, the pool now answers and the scheduler only
calls. Three methods, each with the separate-buffer behaviour as the default
and the shared-envelope behaviour as an override:

- `prealloc_fits` prices a (full, swa) demand against the scheduler's budgets.
  The default compares per side and never reads `tree_cache`; the override
  reads what the tree could reclaim and prices the whole ask in bytes.
- `reclaim_for_prealloc` frees room and names the shortfall it could not
  meet. Defaults to the sliding-window evict on `SWATokenToKVPoolAllocator`,
  since only SWA-family allocators reach it.
- `has_shared_byte_envelope` states the layout fact behind both consequences:
  the sides are priced together, and the answer is post-reclaim, so admitting
  on it still owes the reclaim.

`swa_capacity_and_available` needed no override at all -- the base already
returned exactly the pair the static branch was recomputing.

Behaviour is unchanged for every layout, including the tri-pool: its
`can_reserve` refuses an evictable allowance, so it keeps the inherited
per-side default it already took under the old type test.

One branch remains in `_check_if_req_exceed_kv_capacity`, marked in place: its
two sides bound a different length, so folding them would change which
requests are refused.

Scheduler tests that admit requests now bind the real separate-buffer
implementations onto their allocator double, or a bare `MagicMock` returns a
truthy `Mock` and the decision under test stops being made anywhere.

Verified on cheng-wan-h200-8gpu: 266 passed / 2 skipped / 1403 subtests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…tion

Two bugs from review of the previous commit, plus its two comment-style slips.

**Tri-pool PD admission could crash the decode node.** `full_available_size()`
and `swa_available_size()` both take `schedulable_available_size()`, which
credits the peer's drainable holes, so each side reports capacity backed by
the *same* shared gap. Comparing each against its own token budget therefore
double-counts those bytes: a pair that fits neither together was admitted,
`alloc_extend_swa_tail` then priced them jointly through `_fits_page_demand`
and returned None, and `_pre_alloc` asserts on that (`KV cache is full! Bug in
memory estimation`) rather than refusing the request. The previous commit's
note that the tri-pool "keeps the per-side default it already took" described
the control flow correctly but did not make per-side admission safe, now that
tail allocation is asymmetric on one envelope.

`UnifiedMambaSWATokenToKVPoolAllocator.prealloc_fits` now prices the pair on
the float chain's own page grid before applying the budgets. The budgets still
bind -- they carry decode headroom the allocator cannot see.

**Hybrid-Mamba (non-SWA) decode retraction raised NotImplementedError.**
Unified PD is forced onto `cpu_tensor`, and `Req.offload_kv_cache` calls the
allocator's `get_cpu_copy`; `UnifiedMambaTokenToKVPoolAllocator` inherited the
raising base. It now translates virtual FULL ids to physical before the pool,
the way `UnifiedSWAKVPool` already does for the SWA layout. `UnifiedMLATokenToKVPool`
grew the physical->kernel hop its own class docstring states, since the MLA
parent indexes `kv_buffer` by kernel-facing ids -- delegating without it would
have read the wrong rows silently.

`has_shared_byte_envelope` bundled two contracts that the tri-pool answers
differently: it is one buffer, but its `prealloc_fits` prices the pool as it
stands rather than post-reclaim. Split into `prealloc_fits_assumes_reclaim`
(owed reclaim) and `prealloc_ceiling_fits` (an ask that can never fit, or None
when the pool has no ceiling of its own), which also lets
`_check_if_req_exceed_kv_capacity` drop its last layout test.

Dropped two comments that narrated the refactor rather than stating a fact the
code does not -- that rationale belongs in the PR, per the repo comment rules.

Verified on cheng-wan-h200-8gpu: 278 passed / 2 skipped / 1403 subtests. The
new tri-pool case fails without the fix (admits an infeasible pair).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The previous regression case asked for `2 * available_size()` on both sides
with artificial budgets. It did fail without the fix, but through the budget
conjunct rather than the joint grid -- it was not the shape the production bug
takes.

Scanning the whole (full, swa) space on the tri-pool fixture: at page_size 1
there is no pair that each side can host alone yet the grid refuses, in any of
five pool states. The slack only opens on a paged grid -- at page_size 4 there
are 16 such pairs, and asking for exactly each side's own `available_size` is
one of them. The case now pins that, and asserts the per-side precondition
alongside the joint refusal so the double-count is visible in the test rather
than implied. Verified red without the fix.

Also trimmed the `prealloc_fits` docstring to the two facts a reader cannot
recover here, dropping the `_pre_alloc` assert story -- same rule as the two
comments removed in the previous commit.

Verified on cheng-wan-h200-8gpu: 309 passed / 2 skipped / 1443 subtests across
18 files, including the MLA/envelope/gate tests that cover the pools the
previous commit touched and the 1-GPU SWA tail allocation test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The previous commit fixed two halves of that path by reading the code and ran
neither. Both now have a case, and both were verified red without their fix.

`UnifiedMLATokenToKVPool.get_cpu_copy` / `load_cpu_copy` round-trip through
PHYSICAL ids at page_size 1 and 4, checked against the kernel-id formula the
class docstring states. Without the rewrite the parent indexes `kv_buffer` at
a different row, so the restore returns other tokens' KV -- silently, which is
why a round-trip assert is the guard rather than a raise.

`UnifiedMambaTokenToKVPoolAllocator.get_cpu_copy` / `load_cpu_copy` hand the
pool physical ids. The case pins that the v2p table is not identity for the
allocated run first, so a delegate that passed `req_to_token`'s virtual ids
straight through -- the shape the fix replaced -- fails it rather than
coincidentally agreeing.

Both live in `test_unified_mla_views.py`, which already had the CPU unified
MLA + mamba fixture and the `_kernel_id` helper. The two call sites need a
stated `dcp_enabled` (and `attn_dcp_size` for the composite), so they take
`get_parallel().override(...)` rather than publishing a whole ServerArgs.

Verified on cheng-wan-h200-8gpu: 311 passed / 2 skipped / 1445 subtests across
18 files.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Only conflict is the import block in `test_decode_radix_lock_ref.py`, where
main added `CustomTestCase` next to the allocator-double helper this branch
introduced; both imports are kept.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Main tightened the registered-test layout: kernel tests live under
test/registered/kernels/{ops,benchmark}/<group>/. This one landed at
test/registered/kernel/mem_cache/ before that, and the sync makes it a
lint failure.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@ch-wan ch-wan added bypass-fastfail run-ci-extra CI: also run the extra suite (requires run-ci) labels Sep 18, 2026
@ch-wan

ch-wan commented Sep 18, 2026

Copy link
Copy Markdown
Collaborator

/tag-and-rerun-ci

@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label Sep 18, 2026
Bring in the unified-memory rerun group and current main updates.
@ZYHowell

Copy link
Copy Markdown
Collaborator Author

/rerun-group unified-memory

@github-actions

github-actions Bot commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-group unified-memory:

🚀 1-gpu-h100 (5 tests): ✅ View workflow run

cd test/ && python3 registered/attention/test_gemma4_unified_swa_virtual_ids.py
cd test/ && python3 registered/attention/test_unified_memory_deterministic.py
cd test/ && python3 registered/e2e/models/test_inkling_unified.py
cd test/ && python3 registered/page_major/test_page_major_gpt_oss.py
cd test/ && python3 registered/page_major/test_page_major_qwen_hybrid.py

🚀 2-gpu-h100 (4 tests): ✅ View workflow run

cd test/ && python3 registered/disaggregation/test_disaggregation_unified_memory.py
cd test/ && python3 registered/e2e/disaggregation/test_disaggregation_unified_memory_swa.py
cd test/ && python3 registered/e2e/disaggregation/test_disaggregation_unified_memory_tri.py
cd test/ && python3 registered/e2e/models/test_kimi_linear_models.py

🚀 4-gpu-h100 (1 test): ✅ View workflow run

cd test/ && python3 registered/e2e/models/test_kimi_linear_unified_memory.py

🚀 4-gpu-b200 (1 test): ✅ View workflow run

cd test/ && python3 registered/e2e/models/test_kimi_linear_unified_memory_dcp_blackwell.py

🚀 ubuntu-latest (31 tests): ❌ View workflow run

cd test/ && python3 registered/unit/disaggregation/test_unified_memory_move_gate.py
cd test/ && python3 registered/unit/layers/attention/test_kv_translate_ownership.py
cd test/ && python3 registered/unit/managers/test_prefill_adder.py
cd test/ && python3 registered/unit/managers/test_scheduler_init_req_max_new_tokens.py
cd test/ && python3 registered/unit/mem_cache/test_dsv4_unified_fp8_pool.py
cd test/ && python3 registered/unit/mem_cache/test_flashkda_strided_state_access.py
cd test/ && python3 registered/unit/mem_cache/test_full_loc_fast_path.py
cd test/ && python3 registered/unit/mem_cache/test_hisparse_allocator.py
cd test/ && python3 registered/unit/mem_cache/test_hisparse_max_token_pool_size.py
cd test/ && python3 registered/unit/mem_cache/test_kda_fused_decode_strided_state.py
cd test/ && python3 registered/unit/mem_cache/test_kv_index_translator.py
cd test/ && python3 registered/unit/mem_cache/test_multi_ended_allocator.py
cd test/ && python3 registered/unit/mem_cache/test_pd_envelope_transfer_layout.py
cd test/ && python3 registered/unit/mem_cache/test_prefill_memory_budget.py
cd test/ && python3 registered/unit/mem_cache/test_swa_locked_full_recover_unified.py
cd test/ && python3 registered/unit/mem_cache/test_unified_byte_accounting.py
cd test/ && python3 registered/unit/mem_cache/test_unified_byte_budget_sizing.py
cd test/ && python3 registered/unit/mem_cache/test_unified_capacity_memo.py
cd test/ && python3 registered/unit/mem_cache/test_unified_free_no_host_sync.py
cd test/ && python3 registered/unit/mem_cache/test_unified_mamba_views.py
cd test/ && python3 registered/unit/mem_cache/test_unified_mha_views.py
cd test/ && python3 registered/unit/mem_cache/test_unified_mla_views.py
cd test/ && python3 registered/unit/mem_cache/test_unified_npool_sweep.py
cd test/ && python3 registered/unit/mem_cache/test_unified_radix_allocation_eviction.py
cd test/ && python3 registered/unit/mem_cache/test_unified_swa_shared_virtual_ids.py
cd test/ && python3 registered/unit/mem_cache/test_unified_tri_pool.py
cd test/ && python3 registered/unit/model_executor/test_pool_configurator.py
cd test/ && python3 registered/unit/server_args/test_page_major_backend_allowlist.py
cd test/ && python3 registered/unit/server_args/test_unified_prefill_cuda_graph_gate.py
cd test/ && python3 registered/unit/server_args/test_unified_tbo_gate.py
cd test/ && python3 registered/unit/test_model_overrides.py

🚀 1-gpu-5090 (2 tests): ✅ View workflow run

cd test/ && python3 registered/unit/mem_cache/test_unified_handout_zeroing.py
cd test/ && python3 registered/unit/mem_cache/test_unified_mla_gpu_parity.py

Use the existing separate-buffer allocator double and publish the runtime configuration needed by the preallocation capacity checks. Preserve the HiSparse host-backed capacity assertions.
@ZYHowell

Copy link
Copy Markdown
Collaborator Author

/rerun-group unified-memory

@github-actions

github-actions Bot commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-group unified-memory:

🚀 1-gpu-h100 (5 tests): ✅ View workflow run

cd test/ && python3 registered/attention/test_gemma4_unified_swa_virtual_ids.py
cd test/ && python3 registered/attention/test_unified_memory_deterministic.py
cd test/ && python3 registered/e2e/models/test_inkling_unified.py
cd test/ && python3 registered/page_major/test_page_major_gpt_oss.py
cd test/ && python3 registered/page_major/test_page_major_qwen_hybrid.py

🚀 2-gpu-h100 (4 tests): ✅ View workflow run

cd test/ && python3 registered/disaggregation/test_disaggregation_unified_memory.py
cd test/ && python3 registered/e2e/disaggregation/test_disaggregation_unified_memory_swa.py
cd test/ && python3 registered/e2e/disaggregation/test_disaggregation_unified_memory_tri.py
cd test/ && python3 registered/e2e/models/test_kimi_linear_models.py

🚀 4-gpu-h100 (1 test): ✅ View workflow run

cd test/ && python3 registered/e2e/models/test_kimi_linear_unified_memory.py

🚀 4-gpu-b200 (1 test): ✅ View workflow run

cd test/ && python3 registered/e2e/models/test_kimi_linear_unified_memory_dcp_blackwell.py

🚀 ubuntu-latest (31 tests): ✅ View workflow run

cd test/ && python3 registered/unit/disaggregation/test_unified_memory_move_gate.py
cd test/ && python3 registered/unit/layers/attention/test_kv_translate_ownership.py
cd test/ && python3 registered/unit/managers/test_prefill_adder.py
cd test/ && python3 registered/unit/managers/test_scheduler_init_req_max_new_tokens.py
cd test/ && python3 registered/unit/mem_cache/test_dsv4_unified_fp8_pool.py
cd test/ && python3 registered/unit/mem_cache/test_flashkda_strided_state_access.py
cd test/ && python3 registered/unit/mem_cache/test_full_loc_fast_path.py
cd test/ && python3 registered/unit/mem_cache/test_hisparse_allocator.py
cd test/ && python3 registered/unit/mem_cache/test_hisparse_max_token_pool_size.py
cd test/ && python3 registered/unit/mem_cache/test_kda_fused_decode_strided_state.py
cd test/ && python3 registered/unit/mem_cache/test_kv_index_translator.py
cd test/ && python3 registered/unit/mem_cache/test_multi_ended_allocator.py
cd test/ && python3 registered/unit/mem_cache/test_pd_envelope_transfer_layout.py
cd test/ && python3 registered/unit/mem_cache/test_prefill_memory_budget.py
cd test/ && python3 registered/unit/mem_cache/test_swa_locked_full_recover_unified.py
cd test/ && python3 registered/unit/mem_cache/test_unified_byte_accounting.py
cd test/ && python3 registered/unit/mem_cache/test_unified_byte_budget_sizing.py
cd test/ && python3 registered/unit/mem_cache/test_unified_capacity_memo.py
cd test/ && python3 registered/unit/mem_cache/test_unified_free_no_host_sync.py
cd test/ && python3 registered/unit/mem_cache/test_unified_mamba_views.py
cd test/ && python3 registered/unit/mem_cache/test_unified_mha_views.py
cd test/ && python3 registered/unit/mem_cache/test_unified_mla_views.py
cd test/ && python3 registered/unit/mem_cache/test_unified_npool_sweep.py
cd test/ && python3 registered/unit/mem_cache/test_unified_radix_allocation_eviction.py
cd test/ && python3 registered/unit/mem_cache/test_unified_swa_shared_virtual_ids.py
cd test/ && python3 registered/unit/mem_cache/test_unified_tri_pool.py
cd test/ && python3 registered/unit/model_executor/test_pool_configurator.py
cd test/ && python3 registered/unit/server_args/test_page_major_backend_allowlist.py
cd test/ && python3 registered/unit/server_args/test_unified_prefill_cuda_graph_gate.py
cd test/ && python3 registered/unit/server_args/test_unified_tbo_gate.py
cd test/ && python3 registered/unit/test_model_overrides.py

🚀 1-gpu-5090 (2 tests): ✅ View workflow run

cd test/ && python3 registered/unit/mem_cache/test_unified_handout_zeroing.py
cd test/ && python3 registered/unit/mem_cache/test_unified_mla_gpu_parity.py

@ch-wan ch-wan removed the run-ci-extra CI: also run the extra suite (requires run-ci) label Sep 18, 2026
@ch-wan
ch-wan merged commit 5931fd6 into main Sep 19, 2026
196 of 230 checks passed
@ch-wan
ch-wan deleted the yonghao/ump-pd-page-envelopes branch September 19, 2026 00:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bypass-fastfail memory-pool run-ci CI: run the baseline test suite on this PR unified-radix-cache

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants