Skip to content

Support unified memory page-envelope transfers in PD - #36730

Closed
ZYHowell wants to merge 78 commits into
sgl-project:mainfrom
ZYHowell:yonghao/ump-pd-page-envelopes
Closed

ZYHowell wants to merge 78 commits into
sgl-project:mainfrom
ZYHowell:yonghao/ump-pd-page-envelopes

Conversation

@ZYHowell

@ZYHowell ZYHowell commented Aug 27, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

Unified memory pools expose virtual token IDs while disaggregated transfer backends operate on physical buffers. PD transfers therefore need an explicit page-envelope contract and physical index translation, including distinct target and independent-draft index vectors when their layouts differ.

Stacked after #36729; prerequisite #36403 is merged.

Modifications

  • Move the PD translation, SWA-tail allocation, and compaction gates into upstream's split allocator modules; retain the new config-field namespaces and physical/logical page distinction.
  • Expose contiguous page envelopes for unified MHA/SWA pools and translate virtual IDs before transfer.
  • Gate compaction while a transfer can still reference physical pages.
  • Add an opt-in target/draft index-vector interface for transfer backends.
  • Preserve the CPU-copy fallback for decode retraction and validate unsupported combinations at startup.
  • Keep attention SWA IDs kernel-facing while translating PD payloads to physical SWA page IDs, and transfer unified one-region SWA envelopes through Mooncake's flat path.

Accuracy Tests

Not applicable; this changes KV transport addressing rather than model math.

Speed Tests and Profiling

Not run; no transfer benchmark environment was used locally.

Test Plan

  • Updated head: e09999ea68. Merged OSS main a63efd9056 into the bottom PR and propagated normal merges; no rebase or force-push. Final merge previews also pass against main dc5f59c3a2.
  • Existing focused suite: 200 passed; 17 subtests passed.
  • Adapted the existing HiSparse fixture to config bags and Python 3.10-compatible with syntax after the first CPU CI exposed unittest.enterContext being unavailable on Python 3.10. Its 8 tests pass locally. No new tests were added for this conflict refresh.
  • Full changed-file pre-commit passes, including test registry validation.
  • Exact-head Base CI has started; results are not yet claimed green. Extra CI is not opted in (run-ci-extra is absent). The initial separate MLX run failed in the upstream, unchanged test_scheduler_mixin.py mock-ingestion contract.
  • Full-model serving, multi-node PD, and accuracy/performance benchmarks were not run locally.
Local validation commands

Python commands below were executed through an existing uv run environment (Python 3.12, Torch 2.11); environment-specific paths are omitted.

PYTHONPATH=python python -m pytest -q \
  test/registered/unit/disaggregation/test_decode_queue_cleanup.py \
  test/registered/unit/disaggregation/test_pp_hybrid_kv_transfer.py \
  test/registered/unit/disaggregation/test_unified_memory_move_gate.py \
  test/registered/unit/managers/test_priority_scheduling_disaggregation.py \
  test/registered/unit/mem_cache/test_decode_radix_lock_ref.py \
  test/registered/unit/mem_cache/test_multi_ended_allocator.py \
  test/registered/unit/mem_cache/test_unified_mha_views.py \
  test/registered/unit/mem_cache/test_unified_swa_shared_virtual_ids.py \
  test/registered/unit/mem_cache/test_swa_alloc_extend_page_estimation.py \
  test/registered/unit/mem_cache/test_swa_locked_full_recover_unified.py \
  test/registered/unit/mem_cache/test_kv_index_translator.py \
  test/registered/unit/server_args/test_unified_tbo_gate.py \
  test/registered/unit/server_args/test_unified_prefill_cuda_graph_gate.py \
  --disable-warnings --maxfail=4

PYTHONPATH=python python test/registered/unit/mem_cache/test_hisparse_max_token_pool_size.py -q

GITHUB_BASE_REF=a63efd9056b33a3d4a32dfba6262fac1d62b959a \
  python -m pre_commit run \
  --from-ref a63efd9056b33a3d4a32dfba6262fac1d62b959a --to-ref HEAD

Original commits

  • 3f245db06c33bca46e62b60c0ed0a3ed9edd3353
  • d036c1443db3c7b3ab86d4194c7e1dab5b5dbd3d

Checklist


CI States

Latest PR Test (Base): Not run yet
Latest PR Test (Extra): ⚠️ Not enabled -- add run-ci-extra label to opt in.
Latest PR Test (AMD ROCm 10): ➖ No AMD PR run found for this commit.

@ch-wan ch-wan left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Summary

This PR enables unified-memory PD for hybrid MHA/SWA by registering whole-buffer page envelopes, translating virtual ids to physical ids before Mooncake transfer, gating compaction while destinations can still be referenced, and forcing cpu-tensor retraction with startup checks for unsupported combos. The envelope addressing matches the existing MLA contract (raw_ptr + physical_page * page_bytes), SWA state still goes through translate_loc_from_full_to_swa, and spec+Mooncake is correctly refused until a backend implements separate draft indices. The remaining risk is decode-side unified SWA admission/prealloc: joint FULL/SWA budgeting is skipped unless SWA-tail mode is on, and alloc_extend_swa_tail asserts a seq/prefix/delta identity that decode HiCache does not satisfy.

Issue counts by severity

  • bugs: 2
  • suggestions: 3
  • nits: 0

Comment thread python/sglang/srt/disaggregation/decode.py Outdated
Comment thread python/sglang/srt/mem_cache/multi_ended_allocator.py Outdated
Comment thread python/sglang/srt/mem_cache/unified_memory_pool.py
Comment thread python/sglang/srt/mem_cache/unified_memory_pool.py
Comment thread python/sglang/srt/mem_cache/kv_cache_builder.py Outdated
yhzhuang and others added 9 commits September 1, 2026 15:04
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Co-authored-by: yhzhuang <yhzhuang@fb.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Minimize both component reclaim quotas under the joint capacity predicate so a blocked compaction path cannot turn one allocation shortfall into an all-SWA eviction.

Original prod_inference commit: 56c7082a92bc0c9024585387460e64b065652918

Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>

@ZYHowell ZYHowell left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the incremental change from #36729 567f4f5d0f13 to this head 606649367b4f, following the allocator if/else feedback. One remaining duplicated rejection workflow is identified below; inherited lower-PR changes are not counted again.

Validation: an isolated CPU comparison of the current method and a common rejection tail matched return values, reservation-call arguments, log/abort/output effects, and message precedence in 1,536 cases. No serving or PD transfer test was run for this review.

Comment thread python/sglang/srt/disaggregation/decode.py Outdated
@ch-wan ch-wan self-assigned this Sep 13, 2026
yhzhuang and others added 18 commits September 13, 2026 18:25
Retain shared virtual-ID lifecycle in a common base, gate eager-pool post-capture planning, and preserve capacity realization results. Keep allocator-owned admission policy and existing tri-pool behavior.
Consolidate the historical PR iterations through the last upstream merge (f83adaf), preserving its exact tree. Later review and refactor commits remain separate.
Retain shared virtual-ID lifecycle in a common base, gate eager-pool post-capture planning, and preserve capacity realization results. Keep allocator-owned admission policy and existing tri-pool behavior.

@ZYHowell ZYHowell left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the PD-specific delta from inherited capacity commit 089673a0e6 to 3c9040a7362e. The previous duplicated over-capacity rejection tail is addressed, and unified admission dispatch is consistently first. One small remaining conditional simplification is inline; it preserves the unified reservation call and optional static SWA constraint.

The isolated check matched return values and unified-call arguments in 960 boundary combinations. Validation was limited to source inspection and isolated CPU checks of the suggested simplifications; no GPU, serving, PD, or storage end-to-end validation was run.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

memory-pool run-ci CI: run the baseline test suite on this PR unified-radix-cache

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants