Skip to content

Support unified memory decode host pools - #36731

Closed
ZYHowell wants to merge 34 commits into
sgl-project:yonghao/ump-pd-page-envelopesfrom
ZYHowell:yonghao/ump-decode-host-pool
Closed

ZYHowell wants to merge 34 commits into
sgl-project:yonghao/ump-pd-page-envelopesfrom
ZYHowell:yonghao/ump-decode-host-pool

Conversation

@ZYHowell

@ZYHowell ZYHowell commented Aug 27, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

Decode retraction and KV offload need a host representation that preserves unified page envelopes. Independently sized Full and sliding-window host pools cannot safely model a device pool whose byte capacity is shared dynamically, and a fragmented shared pool must not expose stale physical addresses while pages are being moved.

Stacked after #36730 and #36729; prerequisite #36403 is merged.

Modifications

  • Retain upstream's UMBP registration/host-allocator structure while discovering unified envelope sizes without dereferencing an unallocated logical page.
  • Back Full and SWA page envelopes with one shared byte arena and allocate their aligned pages from opposite ends.
  • Keep cache-visible logical host IDs stable while compacting live physical pages when total free bytes are sufficient but the requested side has no usable extent.
  • Hold layout leases across asynchronous L2 transfers and synchronous storage access so compaction cannot move pages while a backend is using their physical addresses; nested same-thread leases are reentrant.
  • Resolve logical IDs only at transfer boundaries and keep existing non-unified and logical-anchor host pools outside the shared allocation domain.
  • Derive UMBP page bytes without probing an unallocated logical page.
  • Keep unsafe integrations fail-fast in this layer: unified hierarchical/LMCache remains rejected; unified decode-radix host-pool retraction, unified MLA decode offload, and hybrid-SWA/Mamba decode offload or retraction are rejected until their transfer paths are implemented by a follow-up.
  • Keep automatic unified decode-radix and hybrid-SWA retraction on cpu_tensor rather than selecting host_pool.

Accuracy Tests

Not applicable; this changes KV movement, compatibility validation, and capacity management rather than model math.

Speed Tests and Profiling

Not run; no host-transfer benchmark environment was used locally.

Test Plan

  • Updated head: 35519a8ea5. Merged OSS main a63efd9056 into the bottom PR and propagated normal merges; no rebase or force-push. Final merge previews also pass against main dc5f59c3a2.
  • Existing focused suite: 60 passed; 4 subtests passed.
  • Adapted the existing HiSparse fixture to config bags and Python 3.10-compatible with syntax after the first CPU CI exposed unittest.enterContext being unavailable on Python 3.10. Its 8 tests pass locally. No new tests were added for this conflict refresh.
  • Full changed-file pre-commit passes, including test registry validation.
  • Exact-head Base CI has started; results are not yet claimed green. Extra CI is not opted in (run-ci-extra is absent). The initial separate MLX run failed in the upstream, unchanged test_scheduler_mixin.py mock-ingestion contract.
  • The AMD-registered test_umbp_store.py requires optional mori.umbp, unavailable on this NVIDIA host; actual Mori backend execution remains unvalidated.
  • Full-model serving, multi-node PD, and accuracy/performance benchmarks were not run locally.
Local validation commands

Python commands below were executed through an existing uv run environment (Python 3.12, Torch 2.11); environment-specific paths are omitted.

PYTHONPATH=python python -m pytest -q \
  test/registered/unit/mem_cache/test_hybrid_pool_assembler.py \
  test/registered/unit/mem_cache/test_decode_retraction_backup.py \
  test/registered/unit/mem_cache/test_umbp_host_allocator.py \
  test/registered/unit/mem_cache/test_unified_radix_hicache_dispatch.py \
  test/registered/unit/mem_cache/test_hicache_dcp_host_pool.py \
  test/registered/unit/disaggregation/test_specv2_kvcache_offloading.py \
  test/registered/unit/disaggregation/test_unified_memory_move_gate.py \
  test/registered/unit/mem_cache/test_swa_cpu_copy_filter.py \
  --disable-warnings --maxfail=4

PYTHONPATH=python python test/registered/unit/mem_cache/test_hisparse_max_token_pool_size.py -q

GITHUB_BASE_REF=a63efd9056b33a3d4a32dfba6262fac1d62b959a \
  python -m pre_commit run \
  --from-ref a63efd9056b33a3d4a32dfba6262fac1d62b959a --to-ref HEAD

Original commits

  • 74ae793e5eff15a911205943a788d3318f08853d
  • db6489428c3099acd4871f23594067b40a19e4b8
  • d3e06da155cc99f1c78573b7de5e11790fda34a0

Checklist


CI States

Latest PR Test (Base): Not run yet
Latest PR Test (Extra): ⚠️ Not enabled -- add run-ci-extra label to opt in.
Latest PR Test (AMD ROCm 10): ➖ No AMD PR run found for this commit.

@ch-wan ch-wan left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Summary

This PR adds a shared page-envelope host arena for unified FULL/SWA KV, joint byte-budget admission (can_reserve / ensure_capacity), virtual-to-physical translation for PD, decode offload, and retraction, and a compaction gate while decode offload D2H is in flight. The envelope allocator, atomic alloc_many, and the retraction/offload translation look internally consistent. The dominant risk is that radix HiCache backup/load still feed tree virtual ids into the new physical-page host pool, and compaction is not gated on those async copies -- unlike the paths this PR did update.

Issue counts by severity

  • bugs: 2
  • suggestions: 2
  • nits: 0

Comment thread python/sglang/srt/mem_cache/unified_radix_cache.py
Comment thread python/sglang/srt/disaggregation/utils.py
Comment thread python/sglang/srt/mem_cache/kv_cache_configurator.py Outdated
Comment thread python/sglang/srt/managers/schedule_policy.py Outdated
@ZYHowell
ZYHowell force-pushed the yonghao/ump-decode-host-pool branch 2 times, most recently from c5f25f1 to d02710a Compare September 1, 2026 22:40
yhzhuang and others added 4 commits September 1, 2026 15:52
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
Co-authored-by: yhzhuang <yhzhuang@fb.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
@ZYHowell
ZYHowell force-pushed the yonghao/ump-decode-host-pool branch from d02710a to 1b75db1 Compare September 2, 2026 00:05
@ZYHowell ZYHowell mentioned this pull request Sep 2, 2026
5 tasks done
@ZYHowell

Copy link
Copy Markdown
Collaborator Author

/tag-and-rerun-ci

@ch-wan ch-wan left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Summary

Reviewed the delta from #36730 at 7a7119277c3c. The host layout lease and allocation design are coherent, but L2 index preparation runs on the transfer stream before its producer-event wait. With the direct backend, preparation copies newly produced device indices to CPU too early, so the subsequent KV transfer can use stale addresses.

Validation: instrumented both actual L2 submit methods and the unified host-pool preparation method to confirm the ordering. Exhaustive single-host-pool capacity checks also showed that fitting allocations reuse holes without compaction. No CUDA race or serving run was available. Inherited FULL-envelope transport findings remain on #36730.

Issue counts by severity

  • bugs: 1
  • suggestions: 1
  • nits: 0

Comment thread python/sglang/srt/mem_cache/pool_host/group.py Outdated
Comment thread python/sglang/srt/mem_cache/l2_transfer.py Outdated

@ch-wan ch-wan left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Summary

Reviewed 759b4282ef. The shared host arena itself is sound: allocation, freeing and compaction across both grow directions survive randomized stress without corruption, and the layout-lease protocol is the right shape for keeping transfers off relocating pages. The blocking problem is reachability. Both new flags assert the model is neither mamba-ish nor hybrid-SWA, while KVCacheConfigurator only builds a unified pool for models that are one of those two, so no configuration can enable either feature and roughly a thousand lines of new code run only from unit tests that construct the pools directly. The four remaining bugs all sit in the hybrid-SWA path that gate disables, which is the evidence that the gate is over-broad rather than deliberate -- in particular the SWA transfer reintroduces the kernel-facing/physical id confusion that 43ecd06885 fixed elsewhere in this same stack. The last one is unrelated to unified memory: a new startup rejection that lands on existing static-pool decode servers.

Validation: read both gate sites against _init_pools and kv_cache_builder's supports_host_pool; confirmed kernel_page_multiplier is 2 * layer_num (a 35-layer SWA pool logs 70 at startup), and that the write-through raise is absent from origin/main while hicache_write_policy defaults to write_through with no resolver overriding it.

Issue counts by severity

  • bugs: 5
  • suggestions: 0
  • nits: 0

"--enable-unified-memory decode KV offload does not support "
"hybrid-Mamba models."
)
assert not model_config.is_hybrid_swa, (

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[bug][P1] These gates and the pool factory admit disjoint model sets, so neither new flag can be enabled. Both new blocks assert not use_mla_backend(...), mambaish_config(model_config) is None and not model_config.is_hybrid_swa. KVCacheConfigurator._init_pools raises unless the model is hybrid-Mamba or hybrid-SWA ("--enable-unified-memory only supports hybrid Mamba and hybrid sliding-window-attention models"), so every model fails one side or the other: a hybrid model trips these asserts, a non-hybrid model trips the configurator. The resolution path agrees -- kv_cache_builder computes supports_host_pool = unified_draft_host_pool_supported and not unified_hybrid_swa and not uses_ssm_state(...), which is unconditionally False under --enable-unified-memory. The result is that pool_host/unified.py, the UnifiedSWAKVPool branch of DecodeKVCacheOffloadManager, build_hybrid_swa_pool_pair and the shared-arena reclaim machinery are reachable only from unit tests that construct the pools directly; there is no configuration in which the feature runs.

Suggestion: Drop the is_hybrid_swa assert from both blocks (that is the case the code is written for) and fix the hybrid-SWA path -- see the translate_loc_from_full_to_swa and index-filter comments, which are exactly the bugs the gate is currently hiding. If hybrid-SWA is genuinely out of scope for this PR, then the SWA branches and the shared-arena machinery should not land yet either.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The reachability analysis is correct: these UMP host-pool retraction/offload flags are still intentionally unavailable, and this PR should not be described as enabling them end to end. I am keeping the safety gates in this update. Fixing one physical-index call is not sufficient evidence to enable the whole lifecycle; ordinary UMP PD retraction continues to use cpu_tensor.

#37496 separately integrates the shared host arena with hybrid-SWA HiCache backup/load-back. That does not automatically qualify these decode flags. The scope/staging concern remains open; I am not marking it resolved or removing the gates based only on unit coverage.

Comment thread python/sglang/srt/disaggregation/decode_kvcache_offload_manager.py Outdated
return full_indices, []

swa_indices = self.kv_cache.translate_loc_from_full_to_swa(virtual_indices)
live_swa_indices = swa_indices[swa_indices > 0].to(torch.int64)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[bug] The boolean filter yields a transfer length that is not a multiple of the page size. swa_indices[swa_indices > 0] drops arbitrary positions, so live_swa_indices.numel() is unrelated to page_size. That tensor becomes PoolTransfer.device_indices, and HostPoolGroup allocates with alloc(len(transfer.device_indices)), whose _request_page_counts asserts "The requested size should be a multiple of the page size"; _to_page_indices rejects it as well. Concretely, page_size=4 with 10 of 12 window tokens mapped gives numel() == 10 and aborts. Separately, > 0 rather than >= 0 silently discards physical slot 0; that happens to be the reserved sink here, but the intent is not stated anywhere.

Suggestion: Select whole pages -- take the trailing page-aligned run of bound indices (or reshape to (-1, page_size) and keep rows that are fully bound) -- so the count is page-aligned by construction. Comment the slot-0 exclusion, or compare against the sink explicitly.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The proposed 10-of-12 mapping is not produced by this allocator/caller contract. Offload start/length and stride are page-aligned, and SWA bindings/tombstones are per virtual page, not per token. A mapped physical page is entirely above the reserved sink page; an unmapped page is entirely filtered out. Thus this filter removes whole pages on valid input. Physical slot 0 is reserved, not live KV. Aligned caller.

The new real-page regression covers 0, 4 and 12 tombstoned tokens out of 12 at page size 4, including the all-unbound case. I did not add arbitrary partial-row recovery, which would mask a broken page-binding invariant. An explicit sink/invariant comment would still be a reasonable clarity improvement.

page_hashes = self._compute_prefix_hash(incremental_tokens, prior_hash)
for transfer in pool_transfers:
page_count = transfer.device_indices.numel() // self.page_size
transfer.keys = page_hashes[-page_count:]

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[bug] page_hashes[-0:] is the whole list, not the empty one. When page_count is 0 -- reachable whenever the live index count is below one page, which the unaligned filter above makes easy -- page_hashes[-page_count:] evaluates to page_hashes[:], so the extra pool is backed up under every key in the chunk instead of none. The failure is silent: the transfer succeeds and the sidecar is populated with wrong-keyed entries that later read back as hits.

Suggestion: Guard the zero case explicitly: transfer.keys = page_hashes[-page_count:] if page_count else [].

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Correct about Python's [-0:] semantics, but the stated zero-page transfer does not reach this loop on the valid path. _resolve_offload_transfers returns no SWA transfer when all indices are unbound; otherwise the page-binding and aligned-chunk invariant makes the nonempty count at least one whole page. A malformed sub-page transfer is rejected by host allocation/index preparation before storage backup, not silently stored under every key.

The all-tombstoned case is covered by the new transfer regression and returns an empty transfer list. I have not added the defensive slice guard because the reported reachable silent-corruption path was not established. Early return.

Comment thread python/sglang/srt/arg_groups/kv_cache_hook.py

@ZYHowell ZYHowell left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the incremental change from #36730 606649367b4f to this head 1e3a2c4fd501, focusing on the earlier request to unify allocator-dependent control flow. Three opportunities remain: use the transfer hooks already added here, share direct-copy lease handling, and centralize host-transfer allocation used by write/retraction.

Validation: static call-chain and interface inspection, including the non-HostKVCache LogicalHostPool; AST checks confirm identical arguments in both D2H/H2D dispatch pairs. No GPU/HiCache E2E run was performed.

Comment thread python/sglang/srt/mem_cache/l2_transfer.py Outdated
Comment thread python/sglang/srt/mem_cache/pool_host/unified.py Outdated
Comment thread python/sglang/srt/mem_cache/unified_radix_cache.py Outdated
@ch-wan ch-wan self-assigned this Sep 13, 2026

@ZYHowell ZYHowell left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the host-pool delta 3c9040a7362e..a8e00853a927. The earlier transfer-hook, direct-copy lease, and retraction/controller allocation duplication has been addressed. Two remaining opportunities are inline: centralize shared-layout domain discovery, and give base/hybrid controllers a consistent offload call signature.

Both domain collectors produced the same ordered domains across 1,555 isolated input sequences, including repeated shared domains and logical/non-shared pools. The controller signature mismatch was confirmed from source. Validation was limited to source inspection and isolated CPU checks of the suggested simplifications; no GPU, serving, PD, or storage end-to-end validation was run.

@ZYHowell
ZYHowell changed the base branch from main to yonghao/ump-pd-page-envelopes September 14, 2026 21:37
@ZYHowell ZYHowell mentioned this pull request Sep 14, 2026
2 of 3 tasks
@ZYHowell ZYHowell closed this Sep 14, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

hicache Hierarchical Caching for SGLang memory-pool run-ci CI: run the baseline test suite on this PR unified-radix-cache

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants