Repository navigation
Conversation
…d-back HiCache load-back gates device state reads with Python-level per-layer waits (layer_transfer_counter.wait_until in the KV / mamba pool accessors). A captured graph body never executes those waits on replay. The breakable backend runs attention and linear attention as eager graph breaks, so the waits still fire there, and tc_piecewise is already disabled with hierarchical cache. The full prefill backend replays the whole transformer body as one graph, so a prefill batch carrying a live consumer index could read KV / mamba state before the H2D load landed (the standing "disable cuda graph execution if hicache loading triggered" TODO in Scheduler.get_new_batch_prefill). Under the full backend, enqueue one wait on the pending load op's FINAL event on the forward stream before replay. This keeps the graph and trades the layerwise H2D/compute pipelining for that batch only. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
rodamani
marked this pull request as ready for review
September 29, 2026 21:15
rodamani
requested review from
Fridge003,
Ying1123,
hnyls2002,
ispobock and
merrymercy
as code owners
September 29, 2026 21:15
Contributor
Author
|
/rerun-test -c test_full_cuda_graph_prefill.py test_hicache_variants.py |
Contributor
|
Results for 🚀 🚀 ⛔ |
Contributor
Author
|
/tag-and-rerun-ci |
3 of 5 tasks
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
HiCache load-back gates device-state reads with Python-level per-layer waits (
layer_transfer_counter.wait_untilin the KV / Mamba pool accessors). A captured CUDA graph never executes those waits on replay. The breakable backend runs attention as eager graph breaks, so the waits still fire, andtc_piecewiseis already disabled with hierarchical cache. The full prefill graph backend replays the whole transformer body as one graph, so a prefill batch with a pending load-back could read KV / Mamba state before the H2D copy lands. This is the open "disable cuda graph execution if hicache loading triggered" TODO inScheduler.get_new_batch_prefill.Modifications
prefill_cuda_graph_runner.py, under the full backend, enqueue one wait on the pending load op's final event on the forward stream before replay. The graph is kept; only that batch loses layerwise H2D/compute overlap.test/registered/unit/model_executor/test_full_prefill_graph_hicache_load_fence.py.Accuracy Tests
CPU: 3 passed. With the runner change reverted: 3 failed. CUDA-graph replay itself was not exercised on GPU for this branch.
Speed Tests and Profiling
Adds one stream wait only for full-backend prefill batches that have a pending load-back; other batches are unchanged. Not separately benchmarked.
Checklist
Review and Merge Process
/tag-and-rerun-ci,/tag-run-ci-label,/rerun-failed-ciCI States
Latest PR Test (Base): ❌ Run #36767706030
Latest PR Test (Extra): ❌ Run #36767705632
Latest PR Test (AMD ROCm 10): ❌ Run #36767706183