Skip to content

[PP + HiCache] Add PP Prefetch Tickets for eager cross-stage storage prefetch - #36700

Merged
huangtingwei9988 merged 26 commits into
sgl-project:mainfrom
antgroup:support_pp_hicache_ticket_prefetch
Sep 20, 2026
Merged

huangtingwei9988 merged 26 commits into
sgl-project:mainfrom
antgroup:support_pp_hicache_ticket_prefetch

Conversation

@huangtingwei9988

@huangtingwei9988 huangtingwei9988 commented Aug 27, 2026 •

Copy link
Copy Markdown
Collaborator

Co-author: @stepinto

Motivation

#34815 (comment)
The existing PP HiCache path couples storage prefetch to normal request propagation:

  1. The request reaches PP0.
  2. The request propagates through the PP scheduler chain.
  3. Each downstream stage can query storage and start prefetch only after it sees the request.
  4. Prefetch readiness is propagated/synchronized before the request can safely run.

For a storage hit, TTFT therefore includes scheduler propagation delays in addition to storage I/O. The propagation cost is not just network latency: a downstream scheduler may still be finishing the previous step before it can process the request and start its local prefetch. When scheduling takes an exceptionally long time, such as during prefill's forward time, the impact on TTFT can be enormous.

This is especially visible in an alternating hit/miss workload. In the baseline, hit requests can be slower than misses because a hit pays the staged query/prefetch/readiness path while being interleaved with long miss prefills.

Existing Path

sequenceDiagram
    participant C as Client
    participant P0 as PP0 Scheduler
    participant I0 as PP0 Storage I/O
    participant P1 as PP1 Scheduler
    participant I1 as PP1 Storage I/O
    participant S as PP Completion Sync

    C->>P0: Request
    P0->>I0: Query and start local prefetch
    P0->>P1: Propagate request/control
    P1->>I1: Query and start local prefetch
    Note over I0,I1: PP1 I/O cannot start until the real request reaches PP1
    I0-->>S: PP0 local completion
    I1-->>S: PP1 local completion
    S-->>P0: Global readiness reaches PP0
    P0->>P1: Admit and run through the normal PP path
Loading

PP Prefetch Ticket Path

sequenceDiagram
    participant S0 as PP0 Scheduler
    participant F0 as PP0 Forward
    participant S1 as PP1 Scheduler
    participant F1 as PP1 Forward

    S0->>S0: Query PP0 and PP1 storage namespaces
    alt Global storage miss
        Note right of S0: Miss: admit immediately, no ticket or prefetch collective
        par Start PP0 forward
            S0->>F0: Run the selected miss request
        and Propagate scheduler state
            S0->>S1: Propagate the selected request
            S1->>F1: Prepare PP1 forward
        end
        F0->>F1: Hidden states
    else Global storage hit
        S0->>S1: Broadcast PrefetchTicket
        par PP0 ticket worker
            S0->>S0: Allocate KV + sidecars and start Mooncake GET
        and PP1 ticket worker
            S1->>S1: Allocate KV + sidecars and start Mooncake GET
        end
        Note over S0,S1: The real request remains at PP0 while both stages prefetch
        Note over S0,S1: Progressive KV + sidecar MIN publishes the globally safe prefix
        Note right of S0: Global ready: admit on the next scheduler iteration
        par Start PP0 forward
            S0->>F0: Run with staged KV + sidecars
        and Propagate scheduler state
            S0->>S1: Propagate the selected ready request
            S1->>F1: Bind ticket state and prepare PP1 forward
        end
        F0->>F1: Hidden states
        par PP0 local lifetime
            F0->>F0: Layer-wise H2D + forward, then local release
        and PP1 local lifetime
            F1->>F1: Layer-wise H2D + forward, then local release
        end
    end
Loading

The key change is that storage I/O no longer waits for the real request to traverse PP0 -> PP1 -> ... -> PPn.

Benchmark

Setup

  • Model: DeepSeek-V4-Flash-0731
  • Hardware: 2 nodes, 8 x H20 per node
  • Parallelism: TP=4, PP=4, nnodes=2
  • Attention backend: dsv4
  • MoE runner: humming
  • KV cache dtype: fp8_e4m3
  • Page size: 256
  • Input length: 16,384 tokens
  • Output length: 1 token
  • Requests per measured group: 128
  • Concurrency: 4
  • HiCache storage backend: Mooncake standalone storage over RDMA
  • Mixed workload: alternating full-prefix storage-hit and storage-miss requests
  • Every group used fresh keys; only L1/L2 were flushed between prime and measure
  • Storage-hit counts were validated from response metadata
Mooncake and SGLang launch commands

Mooncake master on node rank 0:

nohup mooncake_master \
  --rpc_port=51051 \
  --enable_http_metadata_server=true \
  --http_metadata_server_port=18082 \
  --enable_metric_reporting=false \
  --metrics_port=19003 \
  --logtostderr=true \
  >> /tmp/pp27010-ticket-bench/mooncake-master.log 2>&1 </dev/null &

Mooncake client on node rank 0:

nohup mooncake_client \
  --host=200.9.12.86:45128 \
  --port=51052 \
  --threads=16 \
  --global_segment_size=200g \
  --local_buffer_size=0 \
  --master_server_address=10.13.2.86:51051 \
  --metadata_server=http://10.13.2.86:18082/metadata \
  --protocol=rdma \
  --device_names=mlx5_bond_0 \
  --tenant_id=default \
  --enable_http_server=true \
  --http_port=18081 \
  >> /tmp/pp27010-ticket-bench/mooncake-client.log 2>&1 </dev/null &

Mooncake client on node rank 1:

nohup mooncake_client \
  --host=200.9.2.82:45128 \
  --port=51052 \
  --threads=16 \
  --global_segment_size=200g \
  --local_buffer_size=0 \
  --master_server_address=10.13.2.86:51051 \
  --metadata_server=http://10.13.2.86:18082/metadata \
  --protocol=rdma \
  --device_names=mlx5_bond_0 \
  --tenant_id=default \
  --enable_http_server=true \
  --http_port=18081 \
  >> /tmp/pp27010-ticket-bench/mooncake-client.log 2>&1 </dev/null &

SGLang on node rank 0:

nohup env \
  PYTHONPATH=/tmp/pp27010-main-buffer-ticket-src/python \
  SGLANG_SKIP_SGL_KERNEL_VERSION_CHECK=1 \
  MOONCAKE_CLIENT=200.9.12.86:51052 \
  MOONCAKE_STANDALONE_STORAGE=true \
  MOONCAKE_ENABLE_SSD_OFFLOAD=false \
  python3 -u -m sglang.launch_server \
    --model-path /data/model-cache/DeepSeek-V4-Flash-0731 \
    --tp-size 4 \
    --pp-size 4 \
    --nnodes 2 \
    --node-rank 0 \
    --dist-init-addr 10.13.2.86:32844 \
    --host 0.0.0.0 \
    --port 30000 \
    --mem-fraction-static 0.5 \
    --max-running-requests 16 \
    --chunked-prefill-size 16384 \
    --page-size 256 \
    --kv-cache-dtype fp8_e4m3 \
    --attention-backend dsv4 \
    --swa-full-tokens-ratio 0.1 \
    --moe-runner-backend humming \
    --enable-cache-report \
    --log-level info \
    --enable-hierarchical-cache \
    --hicache-ratio 1.1 \
    --hicache-host-memory-mode buffer_only \
    --hicache-write-policy write_through \
    --hicache-io-backend kernel \
    --hicache-storage-backend mooncake \
    --hicache-storage-prefetch-policy wait_complete \
  >> /tmp/pp27010-ticket-bench/sglang-rank0.log 2>&1 </dev/null &

SGLang on node rank 1:

nohup env \
  PYTHONPATH=/tmp/pp27010-main-buffer-ticket-src/python \
  SGLANG_SKIP_SGL_KERNEL_VERSION_CHECK=1 \
  MOONCAKE_CLIENT=200.9.2.82:51052 \
  MOONCAKE_STANDALONE_STORAGE=true \
  MOONCAKE_ENABLE_SSD_OFFLOAD=false \
  python3 -u -m sglang.launch_server \
    --model-path /data/model-cache/DeepSeek-V4-Flash-0731 \
    --tp-size 4 \
    --pp-size 4 \
    --nnodes 2 \
    --node-rank 1 \
    --dist-init-addr 10.13.2.86:32844 \
    --host 0.0.0.0 \
    --port 30000 \
    --mem-fraction-static 0.5 \
    --max-running-requests 16 \
    --chunked-prefill-size 16384 \
    --page-size 256 \
    --kv-cache-dtype fp8_e4m3 \
    --attention-backend dsv4 \
    --swa-full-tokens-ratio 0.1 \
    --moe-runner-backend humming \
    --enable-cache-report \
    --log-level info \
    --enable-hierarchical-cache \
    --hicache-ratio 1.1 \
    --hicache-host-memory-mode buffer_only \
    --hicache-write-policy write_through \
    --hicache-io-backend kernel \
    --hicache-storage-backend mooncake \
    --hicache-storage-prefetch-policy wait_complete \
  >> /tmp/pp27010-ticket-bench/sglang-rank1.log 2>&1 </dev/null &

Baseline is clean sgl/main after #27010 using the existing PP HiCache cache-mode path. A validation-only Mooncake allocator guard was applied so the DSV4 logical host anchor could start; no ticket, buffer-only scheduling, or PP-key-query changes were added to the baseline.

TTFT Results

Workload Baseline mean Ticket mean Mean change Baseline median Ticket median Median change
0% storage hit 2596.0 ms 1987.9 ms -23.4% 2610.8 ms 1892.4 ms -27.5%
Mixed: hit requests 2749.2 ms 1879.3 ms -31.6% 2966.3 ms 1879.9 ms -36.6%
Mixed: miss requests 2502.4 ms 1885.9 ms -24.6% 2854.3 ms 1880.2 ms -34.1%
Mixed: all requests 2625.8 ms 1882.6 ms -28.3% 2858.7 ms 1880.1 ms -34.2%

Storage-hit Validation

Workload Successful requests Storage hits
0% hit 128 / 128 0 / 128
Mixed 128 / 128 64 / 128

Checklist

Review and Merge Process

  1. Ping Merge Oncalls to start the process. See the PR Merge Process.
  2. Get approvals from CODEOWNERS and other reviewers.
  3. Trigger CI tests with comments or contact authorized users to do so.
    • Common commands include /tag-and-rerun-ci, /tag-run-ci-label, /rerun-failed-ci
  4. After green CI and required approvals, ask Merge Oncalls or people with Write permission to merge the PR.

CI States

Latest PR Test (Base): ✅ Run #35442368734
Latest PR Test (Extra): ❌ Run #35442368744
Latest PR Test (AMD ROCm 10): ⏳ Run #35442368766

@huangtingwei9988

Copy link
Copy Markdown
Collaborator Author

/tag-and-rerun-ci

@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label Aug 27, 2026
@huangtingwei9988
huangtingwei9988 force-pushed the support_pp_hicache_ticket_prefetch branch from 09dd291 to 5f0f228 Compare August 31, 2026 11:22
@huangtingwei9988 huangtingwei9988 added the run-ci-extra CI: also run the extra suite (requires run-ci) label Aug 31, 2026
…t_prefetch

# Conflicts:
#	python/sglang/srt/mem_cache/buffer_mode/pipeline.py
#	python/sglang/srt/mem_cache/storage/mooncake_store/mooncake_store.py
# Conflicts:
#	python/sglang/srt/managers/scheduler.py
#	python/sglang/srt/mem_cache/buffer_mode/pipeline.py
#	python/sglang/srt/mem_cache/unified_radix_cache.py
…cket_prefetch

# Conflicts:
#	python/sglang/srt/mem_cache/unified_radix_cache.py
@huangtingwei9988

huangtingwei9988 commented Sep 11, 2026 •

Copy link
Copy Markdown
Collaborator Author

Targeted fault-injection validation

Issue Trigger Result
Duplicate tickets on enqueue/retract/retry 16 mixed requests, with 2 hits and 2 misses forced through native retraction; another 8 requests exercised native retry All completed. No duplicate queries, tickets, or ticket consumption
Cleanup after rejection/abort 8 rejected requests and 2 explicit aborts, including rejection before downstream ticket registration Buffers were safely released and ticket state was cleared
Ticket-worker allocation failure 20 injected sidecar KeyErrors across 8 requests: first on all TP ranks of PP1, then only PP2/TP1 All 8 requests fell back to computation with zero storage hits and completed. Workers continued processing subsequent tickets

@MichoChan

Copy link
Copy Markdown

why only support --hicache-host-memory-mode buffer_only ?

@huangtingwei9988

huangtingwei9988 commented Sep 11, 2026 •

Copy link
Copy Markdown
Collaborator Author

why only support --hicache-host-memory-mode buffer_only ?

Because the ticket scheme prevents the host from being inserted into the radix tree, otherwise different PP ranks would hit different lengths in the radix tree, and buffer-only solutions perfectly satisfy this characteristic. @MichoChan

@huangtingwei9988

Copy link
Copy Markdown
Collaborator Author

/rerun-failed-ci

@huangtingwei9988

Copy link
Copy Markdown
Collaborator Author

/rerun-failed-ci

@huangtingwei9988

Copy link
Copy Markdown
Collaborator Author

/rerun-failed-ci

1 similar comment
@huangtingwei9988

Copy link
Copy Markdown
Collaborator Author

/rerun-failed-ci

@huangtingwei9988

Copy link
Copy Markdown
Collaborator Author

/rerun-failed-ci

@huangtingwei9988

Copy link
Copy Markdown
Collaborator Author

/rerun-failed-ci

2 similar comments
@huangtingwei9988

Copy link
Copy Markdown
Collaborator Author

/rerun-failed-ci

@huangtingwei9988

Copy link
Copy Markdown
Collaborator Author

/rerun-failed-ci

@huangtingwei9988
huangtingwei9988 merged commit 0207039 into sgl-project:main Sep 20, 2026
491 of 572 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

hicache Hierarchical Caching for SGLang high priority run-ci CI: run the baseline test suite on this PR run-ci-extra CI: also run the extra suite (requires run-ci) unified-radix-cache

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants