Repository navigation
[PP + HiCache] Add PP Prefetch Tickets for eager cross-stage storage prefetch - #36700
Merged
huangtingwei9988 merged 26 commits intoSep 20, 2026
Merged
huangtingwei9988 merged 26 commits into
huangtingwei9988 merged 26 commits into
Conversation
huangtingwei9988
requested review from
Ying1123,
alphabetc1,
hanming-lu,
hnyls2002,
hzh0425,
ispobock,
merrymercy,
xiezhq-hermann and
yizhang2077
as code owners
August 27, 2026 12:33
Collaborator
Author
|
/tag-and-rerun-ci |
huangtingwei9988
force-pushed
the
support_pp_hicache_ticket_prefetch
branch
from
August 31, 2026 11:22
09dd291 to
5f0f228
Compare
…t_prefetch # Conflicts: # python/sglang/srt/mem_cache/buffer_mode/pipeline.py # python/sglang/srt/mem_cache/storage/mooncake_store/mooncake_store.py
# Conflicts: # python/sglang/srt/managers/scheduler.py # python/sglang/srt/mem_cache/buffer_mode/pipeline.py # python/sglang/srt/mem_cache/unified_radix_cache.py
…cket_prefetch # Conflicts: # python/sglang/srt/mem_cache/unified_radix_cache.py
Collaborator
Author
Targeted fault-injection validation
|
|
why only support --hicache-host-memory-mode buffer_only ? |
Collaborator
Author
Because the ticket scheme prevents the host from being inserted into the radix tree, otherwise different PP ranks would hit different lengths in the radix tree, and buffer-only solutions perfectly satisfy this characteristic. @MichoChan |
xiezhq-hermann
approved these changes
Sep 16, 2026
Collaborator
Author
|
/rerun-failed-ci |
Collaborator
Author
|
/rerun-failed-ci |
Collaborator
Author
|
/rerun-failed-ci |
1 similar comment
Collaborator
Author
|
/rerun-failed-ci |
stmatengss
approved these changes
Sep 19, 2026
Collaborator
Author
|
/rerun-failed-ci |
Collaborator
Author
|
/rerun-failed-ci |
2 similar comments
Collaborator
Author
|
/rerun-failed-ci |
Collaborator
Author
|
/rerun-failed-ci |
4 of 5 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Co-author: @stepinto
Motivation
#34815 (comment)
The existing PP HiCache path couples storage prefetch to normal request propagation:
For a storage hit, TTFT therefore includes scheduler propagation delays in addition to storage I/O. The propagation cost is not just network latency: a downstream scheduler may still be finishing the previous step before it can process the request and start its local prefetch. When scheduling takes an exceptionally long time, such as during prefill's forward time, the impact on TTFT can be enormous.
This is especially visible in an alternating hit/miss workload. In the baseline, hit requests can be slower than misses because a hit pays the staged query/prefetch/readiness path while being interleaved with long miss prefills.
Existing Path
sequenceDiagram participant C as Client participant P0 as PP0 Scheduler participant I0 as PP0 Storage I/O participant P1 as PP1 Scheduler participant I1 as PP1 Storage I/O participant S as PP Completion Sync C->>P0: Request P0->>I0: Query and start local prefetch P0->>P1: Propagate request/control P1->>I1: Query and start local prefetch Note over I0,I1: PP1 I/O cannot start until the real request reaches PP1 I0-->>S: PP0 local completion I1-->>S: PP1 local completion S-->>P0: Global readiness reaches PP0 P0->>P1: Admit and run through the normal PP pathPP Prefetch Ticket Path
sequenceDiagram participant S0 as PP0 Scheduler participant F0 as PP0 Forward participant S1 as PP1 Scheduler participant F1 as PP1 Forward S0->>S0: Query PP0 and PP1 storage namespaces alt Global storage miss Note right of S0: Miss: admit immediately, no ticket or prefetch collective par Start PP0 forward S0->>F0: Run the selected miss request and Propagate scheduler state S0->>S1: Propagate the selected request S1->>F1: Prepare PP1 forward end F0->>F1: Hidden states else Global storage hit S0->>S1: Broadcast PrefetchTicket par PP0 ticket worker S0->>S0: Allocate KV + sidecars and start Mooncake GET and PP1 ticket worker S1->>S1: Allocate KV + sidecars and start Mooncake GET end Note over S0,S1: The real request remains at PP0 while both stages prefetch Note over S0,S1: Progressive KV + sidecar MIN publishes the globally safe prefix Note right of S0: Global ready: admit on the next scheduler iteration par Start PP0 forward S0->>F0: Run with staged KV + sidecars and Propagate scheduler state S0->>S1: Propagate the selected ready request S1->>F1: Bind ticket state and prepare PP1 forward end F0->>F1: Hidden states par PP0 local lifetime F0->>F0: Layer-wise H2D + forward, then local release and PP1 local lifetime F1->>F1: Layer-wise H2D + forward, then local release end endThe key change is that storage I/O no longer waits for the real request to traverse PP0 -> PP1 -> ... -> PPn.
Benchmark
Setup
DeepSeek-V4-Flash-0731TP=4,PP=4,nnodes=2dsv4hummingfp8_e4m3Mooncake and SGLang launch commands
Mooncake master on node rank 0:
Mooncake client on node rank 0:
Mooncake client on node rank 1:
SGLang on node rank 0:
nohup env \ PYTHONPATH=/tmp/pp27010-main-buffer-ticket-src/python \ SGLANG_SKIP_SGL_KERNEL_VERSION_CHECK=1 \ MOONCAKE_CLIENT=200.9.12.86:51052 \ MOONCAKE_STANDALONE_STORAGE=true \ MOONCAKE_ENABLE_SSD_OFFLOAD=false \ python3 -u -m sglang.launch_server \ --model-path /data/model-cache/DeepSeek-V4-Flash-0731 \ --tp-size 4 \ --pp-size 4 \ --nnodes 2 \ --node-rank 0 \ --dist-init-addr 10.13.2.86:32844 \ --host 0.0.0.0 \ --port 30000 \ --mem-fraction-static 0.5 \ --max-running-requests 16 \ --chunked-prefill-size 16384 \ --page-size 256 \ --kv-cache-dtype fp8_e4m3 \ --attention-backend dsv4 \ --swa-full-tokens-ratio 0.1 \ --moe-runner-backend humming \ --enable-cache-report \ --log-level info \ --enable-hierarchical-cache \ --hicache-ratio 1.1 \ --hicache-host-memory-mode buffer_only \ --hicache-write-policy write_through \ --hicache-io-backend kernel \ --hicache-storage-backend mooncake \ --hicache-storage-prefetch-policy wait_complete \ >> /tmp/pp27010-ticket-bench/sglang-rank0.log 2>&1 </dev/null &SGLang on node rank 1:
nohup env \ PYTHONPATH=/tmp/pp27010-main-buffer-ticket-src/python \ SGLANG_SKIP_SGL_KERNEL_VERSION_CHECK=1 \ MOONCAKE_CLIENT=200.9.2.82:51052 \ MOONCAKE_STANDALONE_STORAGE=true \ MOONCAKE_ENABLE_SSD_OFFLOAD=false \ python3 -u -m sglang.launch_server \ --model-path /data/model-cache/DeepSeek-V4-Flash-0731 \ --tp-size 4 \ --pp-size 4 \ --nnodes 2 \ --node-rank 1 \ --dist-init-addr 10.13.2.86:32844 \ --host 0.0.0.0 \ --port 30000 \ --mem-fraction-static 0.5 \ --max-running-requests 16 \ --chunked-prefill-size 16384 \ --page-size 256 \ --kv-cache-dtype fp8_e4m3 \ --attention-backend dsv4 \ --swa-full-tokens-ratio 0.1 \ --moe-runner-backend humming \ --enable-cache-report \ --log-level info \ --enable-hierarchical-cache \ --hicache-ratio 1.1 \ --hicache-host-memory-mode buffer_only \ --hicache-write-policy write_through \ --hicache-io-backend kernel \ --hicache-storage-backend mooncake \ --hicache-storage-prefetch-policy wait_complete \ >> /tmp/pp27010-ticket-bench/sglang-rank1.log 2>&1 </dev/null &Baseline is clean
sgl/mainafter #27010 using the existing PP HiCache cache-mode path. A validation-only Mooncake allocator guard was applied so the DSV4 logical host anchor could start; no ticket, buffer-only scheduling, or PP-key-query changes were added to the baseline.TTFT Results
Storage-hit Validation
Checklist
Review and Merge Process
/tag-and-rerun-ci,/tag-run-ci-label,/rerun-failed-ciCI States
Latest PR Test (Base): ✅ Run #35442368734
Latest PR Test (Extra): ❌ Run #35442368744
Latest PR Test (AMD ROCm 10): ⏳ Run #35442368766