Skip to content

[HiCache] Fence full prefill CUDA-graph replay behind pending load-back - #41741

Open
rodamani wants to merge 4 commits into
sgl-project:mainfrom
modal-projects:rohan/up/hicache-prefill-graph-load-fence
Open

rodamani wants to merge 4 commits into
sgl-project:mainfrom
modal-projects:rohan/up/hicache-prefill-graph-load-fence

Conversation

@rodamani

@rodamani rodamani commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Motivation

HiCache load-back gates device-state reads with Python-level per-layer waits (layer_transfer_counter.wait_until in the KV / Mamba pool accessors). A captured CUDA graph never executes those waits on replay. The breakable backend runs attention as eager graph breaks, so the waits still fire, and tc_piecewise is already disabled with hierarchical cache. The full prefill graph backend replays the whole transformer body as one graph, so a prefill batch with a pending load-back could read KV / Mamba state before the H2D copy lands. This is the open "disable cuda graph execution if hicache loading triggered" TODO in Scheduler.get_new_batch_prefill.

Modifications

  • In prefill_cuda_graph_runner.py, under the full backend, enqueue one wait on the pending load op's final event on the forward stream before replay. The graph is kept; only that batch loses layerwise H2D/compute overlap.
  • test/registered/unit/model_executor/test_full_prefill_graph_hicache_load_fence.py.

Accuracy Tests

CPU: 3 passed. With the runner change reverted: 3 failed. CUDA-graph replay itself was not exercised on GPU for this branch.

Speed Tests and Profiling

Adds one stream wait only for full-backend prefill batches that have a pending load-back; other batches are unchanged. Not separately benchmarked.

Checklist

Review and Merge Process

  1. Ping Merge Oncalls to start the process. See the PR Merge Process.
  2. Get approvals from CODEOWNERS and other reviewers.
  3. Trigger CI tests with comments or contact authorized users to do so.
    • Common commands include /tag-and-rerun-ci, /tag-run-ci-label, /rerun-failed-ci
  4. After green CI and required approvals, ask Merge Oncalls or people with Write permission to merge the PR.

CI States

Latest PR Test (Base): ❌ Run #36767706030
Latest PR Test (Extra): ❌ Run #36767705632
Latest PR Test (AMD ROCm 10): ❌ Run #36767706183

…d-back

HiCache load-back gates device state reads with Python-level per-layer
waits (layer_transfer_counter.wait_until in the KV / mamba pool
accessors). A captured graph body never executes those waits on replay.
The breakable backend runs attention and linear attention as eager graph
breaks, so the waits still fire there, and tc_piecewise is already
disabled with hierarchical cache. The full prefill backend replays the
whole transformer body as one graph, so a prefill batch carrying a live
consumer index could read KV / mamba state before the H2D load landed
(the standing "disable cuda graph execution if hicache loading
triggered" TODO in Scheduler.get_new_batch_prefill).

Under the full backend, enqueue one wait on the pending load op's FINAL
event on the forward stream before replay. This keeps the graph and trades
the layerwise H2D/compute pipelining for that batch only.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@github-actions github-actions Bot added the hicache Hierarchical Caching for SGLang label Sep 29, 2026
@rodamani
rodamani marked this pull request as ready for review September 29, 2026 21:15
@rodamani

Copy link
Copy Markdown
Contributor Author

/rerun-test -c test_full_cuda_graph_prefill.py test_hicache_variants.py

@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test -c test_full_cuda_graph_prefill.py test_hicache_variants.py:

🚀 4-gpu-b200 (1 test): ✅ View workflow run

cd test/ && python3 registered/cuda_graph/full_prefill/test_full_cuda_graph_prefill.py

🚀 1-gpu-h100 (1 test): ✅ View workflow run

cd test/ && python3 registered/hicache/test_hicache_variants.py

⛔ test/registered/unit/model_executor/test_full_prefill_graph_hicache_load_fence.py: File not found: test/registered/unit/model_executor/test_full_prefill_graph_hicache_load_fence.py

@rodamani

Copy link
Copy Markdown
Contributor Author

/tag-and-rerun-ci

@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label Sep 30, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

hicache Hierarchical Caching for SGLang run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants