Skip to content

Fix unified HiCache physical transfers - #37496

Closed
ZYHowell wants to merge 34 commits into
sgl-project:yonghao/ump-decode-host-poolfrom
ZYHowell:yonghao/ump-hicache-physical-transfers
Closed

ZYHowell wants to merge 34 commits into
sgl-project:yonghao/ump-decode-host-poolfrom
ZYHowell:yonghao/ump-hicache-physical-transfers

Conversation

@ZYHowell

@ZYHowell ZYHowell commented Sep 2, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

Unified-memory allocators expose stable logical token IDs while HiCache device transfers must address the current physical page envelopes. Passing logical IDs to H2D/D2H can copy the wrong pages after compaction, and allowing compaction while an asynchronous transfer is active can invalidate an otherwise correct translation.

This PR is stacked after #36731. OSS main was merged into #36729 and propagated through the existing stack; see the exact-head validation below.

Modifications

  • Reconcile upstream's SWA dirty-window ordering and nodes_to_load with physical translation, retain UMBP region deduplication and the TBO gate, and reject the unsupported direct external cache linker only when UMP is enabled.
  • Translate Full and SWA indices at the HiCache transfer boundary, reserve and bind SWA physical pages atomically on load-back, and roll reservations back on failure.
  • Hold shared-layout leases through asynchronous L2 transfers and synchronous L3 buffer access so compaction cannot move pages while a backend is using their addresses.
  • Support buffer_only with the same shared Full/SWA byte arena as cache mode. Staging admission, write reserve, prefetch occupancy, reclaim, and rollback account for both components by bytes rather than treating them as fixed partitions.
  • Store each unified page envelope as one UMBP object and keep ordinary MHA K/V object cardinality unchanged.
  • Keep unsupported combinations fail-fast: hierarchical unified memory currently requires PP=1, no speculative decoding, no recurrent-state or MLA model, and the kernel or direct L2 backend. L3 is limited to file, sim, mori, and shm; a pure L2 configuration remains supported.
  • Preserve UMP-off and non-unified host-pool behavior.

Added test_retraction_uses_physical_swa_transfer_indices to pin physical, rather than kernel-facing, SWA transfer IDs. Existing move-gate, buffer-mode sidecar, and decode-offload fixtures were updated for the current HiCache fields and transfer tuple.

Accuracy Tests

Not run. This change affects cache movement and capacity management, not model forward kernels.

Speed Tests and Profiling

Not run. No performance claim is made.

Test Plan

  • Updated head: 87df0bca72. Merged OSS main a63efd9056 into the bottom PR and propagated normal merges; no rebase or force-push. Final merge previews also pass against main dc5f59c3a2.
  • Existing focused suite: 1437 passed, 1363 skipped; 124 subtests passed on the reconciled implementation at b20d10f786; the final follow-up changes only the inherited HiSparse fixture's context-manager syntax.
  • Adapted the existing HiSparse fixture to config bags and Python 3.10-compatible with syntax after the first CPU CI exposed unittest.enterContext being unavailable on Python 3.10. Its 8 tests pass locally. No new tests were added for this conflict refresh.
  • Full changed-file pre-commit passes, including test registry validation, Rust clippy, and rustfmt.
  • Exact-head Base CI has started; results are not yet claimed green. Extra CI is not opted in (run-ci-extra is absent). The initial separate MLX run failed in the upstream, unchanged test_scheduler_mixin.py mock-ingestion contract.
  • The AMD-registered test_umbp_store.py requires optional mori.umbp, unavailable on this NVIDIA host; actual Mori backend execution remains unvalidated.
  • Additional 25-file inherited allocator/host regression: 392 passed, 730 subtests passed.
  • Native Rust: 846 passed, 1 ignored. Existing Rust-backed Python cache suite: 2573 run, OK with 1368 skipped.
  • Real CUDA Full/SWA D2H → relocating host compaction → H2D is byte-exact in all four combinations: page sizes 1/4 × kernel/direct (page_first/page_first_direct). UMP-on direct-linker rejection and unchanged UMP-off behavior were also checked with one-off probes.
  • Full-model serving, multi-node PD, and accuracy/performance benchmarks were not run locally; no such results are claimed.
Local validation commands

Python commands below were executed through an existing uv run environment (Python 3.12, Torch 2.11); environment-specific paths are omitted.

PYTHONPATH=python python -m pytest -q \
  test/registered/unit/mem_cache/test_buffer_mode_sidecar.py \
  test/registered/unit/mem_cache/test_hicache_staged_write_back_dispatch.py \
  test/registered/unit/mem_cache/test_unified_free_no_host_sync.py \
  test/registered/unit/mem_cache/test_unified_radix_cache_unittest.py \
  test/registered/unit/mem_cache/test_swa_locked_full_recover_unified.py \
  test/registered/unit/mem_cache/test_swa_lock_release_lifecycle.py \
  test/registered/unit/mem_cache/test_unified_radix_hicache_dispatch.py \
  test/registered/unit/mem_cache/test_decode_retraction_backup.py \
  test/registered/unit/mem_cache/test_hybrid_pool_assembler.py \
  test/registered/unit/disaggregation/test_unified_memory_move_gate.py \
  test/registered/unit/server_args/test_unified_tbo_gate.py \
  test/registered/unit/mem_cache/test_multi_ended_allocator.py \
  test/registered/unit/mem_cache/test_linker_pool_assembler.py \
  test/registered/unit/mem_cache/test_unified_cache_linker.py \
  --disable-warnings --maxfail=8

PYTHONPATH=python python test/registered/unit/mem_cache/test_hisparse_max_token_pool_size.py -q

GITHUB_BASE_REF=a63efd9056b33a3d4a32dfba6262fac1d62b959a \
  python -m pre_commit run \
  --from-ref a63efd9056b33a3d4a32dfba6262fac1d62b959a --to-ref HEAD
LIBTORCH_USE_PYTORCH=1 LIBTORCH_BYPASS_VERSION_CHECK=1 \
  CXXFLAGS="-include $PWD/rust/sglang-radix-tree/torch_2_13_compat.h" \
  cargo test --manifest-path rust/sglang-radix-tree/Cargo.toml --locked

PYTHONPATH=python python test/registered/unit/mem_cache/test_rust_unified_radix_cache_unittest.py -q

The Rust commands used a node-local Cargo/build cache and the active environment's Torch library directory in LD_LIBRARY_PATH.

Original commits

  • fb8beb6817db682da3195971b4968ae272ce5304
  • f16293b00e00fcc2470e7ee7174410977c3a5a9f
  • 136669cc5508d8e82c78da758f6977b53572ae45
  • f0af6befe3b13eb7c013fb49ee86aa94725f34bd

Checklist

  • Format code with pre-commit.
  • Run the relevant existing unit tests and the reviewer-requested physical SWA transfer regression.
  • Documentation is not needed for this bug fix.
  • Accuracy and speed benchmarks are not applicable; no claims are made.
  • Follow the SGLang code style guidance.

Review and Merge Process

  1. Ping Merge Oncalls to start the process.
  2. Get approvals from CODEOWNERS and other reviewers.
  3. Trigger CI tests with the standard PR comments.
  4. After green CI and required approvals, ask Merge Oncalls or people with Write permission to merge the PR.

CI States

Latest PR Test (Base): Not run yet
Latest PR Test (Extra): ⚠️ Not enabled -- add run-ci-extra label to opt in.
Latest PR Test (AMD ROCm 10): ➖ No AMD PR run found for this commit.

@ZYHowell

ZYHowell commented Sep 2, 2026

Copy link
Copy Markdown
Collaborator Author

/tag-and-rerun-ci

@github-actions github-actions Bot added hicache Hierarchical Caching for SGLang unified-radix-cache memory-pool run-ci CI: run the baseline test suite on this PR labels Sep 2, 2026
@ch-wan ch-wan mentioned this pull request Sep 4, 2026
2 of 5 tasks
yhzhuang and others added 5 commits September 3, 2026 22:43
Track shared Full/SWA host usage by bytes, allocate storage-hit buffers atomically across both components, and keep allocation and reclaim decisions in rank consensus. Treat unified page envelopes as one UMBP object per page and permit the supported buffer-mode configuration.

Co-authored-by: Yonghao Zhuang <yhzhuang@meta.com>
@ZYHowell

ZYHowell commented Sep 4, 2026

Copy link
Copy Markdown
Collaborator Author

/tag-and-rerun-ci

@ch-wan ch-wan left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Summary

Reviewed the delta from #36731 at 709e1df7c959. Physical translation, SWA reservation binding, transfer completion fencing and shared host accounting are present. One new correctness issue remains in the Rust adapter: an SWA-only backup asserts when FULL already has a host copy. Shared-arena prefetch also drops the existing shorter-prefix fallback under pressure.

Validation: reproduced the Rust adapter assertion with an empty FULL transfer and nonempty SWA transfer. Traced both cache-mode and buffer-mode load-back: SWA bindings are committed before start_loading; successive completion events in each direction are ordered on that direction's dedicated stream. These are CPU probes and source traces, not CUDA or full-model validation. Inherited transport/ordering findings remain on #36730 and #36731.

Issue counts by severity

  • bugs: 1
  • suggestions: 1
  • nits: 0

Comment thread python/sglang/srt/mem_cache/unified_radix_cache.py Outdated
Comment thread python/sglang/srt/mem_cache/rust_tree_core/adapter.py Outdated

@ch-wan ch-wan left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Summary

Reviewed 809821e28c. The id-space work is the good part: virtual, physical and kernel-facing are kept separate through backup, load-back, retraction and prefetch, load-back binds rather than translating so the clamp-to-sink trap is avoided, and the identity translators leave static pools untouched. Both remaining concerns are about blast radius rather than the unified transfers themselves. The new SWA write-back eviction barrier is gated on the host pool instead of on unified memory, and UnifiedRadixCache is the generic prefix cache for every hybrid-SWA model -- so it reaches static-pool write_back HiCache deployments, where, unlike the leaf path, it has no fallback once the host fills and SWA device eviction stops making progress entirely. Two catch-alls added to the shared controller change failure semantics for all HiCache users in the same way. Prior round's two items are not re-raised here.

Validation: confirmed UnifiedRadixCache is selected in registry._create_unified_radix_cache without a unified-memory gate; read the new handlers (both call logger.exception, so the degradation is visible, but a fault still becomes a cache miss); confirmed stores_page_envelope is declared on HostKVCache, memory_pool_host and UnifiedPageEnvelopeHostPool, so the getattr default is unreachable.

Issue counts by severity

  • bugs: 1
  • suggestions: 3
  • nits: 0

Comment thread python/sglang/srt/mem_cache/unified_radix_cache.py
Comment thread python/sglang/srt/managers/cache_controller.py

# State initialization
if self.buffer_pipeline is not None:
self.cache_controller.host_write_staged_tokens_fn = lambda: (

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[suggestion] Duplicate assignment, and the barrier landed under someone else's comment. self.cache_controller.host_write_staged_tokens_fn is assigned here with a body identical to the assignment six lines above, inside the buffer_pipeline construction block. One of the two is dead. Separately, the enable_swa_write_back_eviction_barrier() call a few lines below was inserted between the comment # Pre-seed the logical dropped-tokens series. and the metrics block that comment actually describes, so the comment now reads as documentation for the barrier.

Suggestion: Keep the assignment that runs on every path and drop the other; move the barrier call above the comment, with its own one-line WHY.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Partially addressed in e951bb1ab8: the barrier is now above the metrics pre-seeding comment, and is explicitly UMP-gated. The duplicate host_write_staged_tokens_fn assignment is still present and redundant; removing it (and adding the focused WHY comment) remains valid cleanup. I am not marking the whole suggestion addressed.

Comment thread python/sglang/srt/mem_cache/storage/umbp/umbp_store.py Outdated

@ZYHowell ZYHowell left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the incremental change from #36731 1e3a2c4fd501 to this head 1fcd12537490, focusing on repeated allocator/layout decisions. The UMBP layout interpretation can be shared across key generation, depth expansion, and result grouping; the SWA prefetch capability checks can reuse the existing helper.

Validation: isolated CPU probes matched current key order/count, depth expansion, and complete-page Boolean result grouping in 664 cases across envelope, MLA, ordinary MHA, and split-head layouts. No storage or HiCache E2E run was performed.

Comment thread python/sglang/srt/mem_cache/storage/umbp/umbp_store.py Outdated
Comment thread python/sglang/srt/mem_cache/unified_cache/components/swa_component.py Outdated
@ch-wan ch-wan self-assigned this Sep 13, 2026

@ZYHowell ZYHowell left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the HiCache delta a8e00853a927..8e0df01ef70a. The earlier UMBP key-layout duplication and repeated SWA shared-host predicate are addressed. Three remaining shared-host simplifications are inline: the repeated allocation/consensus/rollback block, duplicated consensus-group selection, and the minimum transfer reservation computed independently at startup and runtime.

Checks confirmed two AST-identical allocation blocks, matching consensus membership/order/fallback in 64 combinations, and matching minimum FULL-reserve arithmetic at 19 boundaries. Validation was limited to source inspection and isolated CPU checks of the suggested simplifications; no GPU, serving, PD, or storage end-to-end validation was run.

@ZYHowell
ZYHowell changed the base branch from main to yonghao/ump-decode-host-pool September 14, 2026 21:37
@ZYHowell ZYHowell mentioned this pull request Sep 14, 2026
4 of 5 tasks
@ZYHowell ZYHowell closed this Sep 14, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

hicache Hierarchical Caching for SGLang memory-pool run-ci CI: run the baseline test suite on this PR unified-radix-cache

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants