[CUDA graph] Drain the device after the last warmup, before capture - #8
Open
jiaqiang-dot-liu wants to merge 1 commit into
Open
jiaqiang-dot-liu wants to merge 1 commit into
jiaqiang-dot-liu wants to merge 1 commit into
Conversation
Problem
Both CUDA-graph backends run two warmup iterations before capture, with a
`synchronize()` + `barrier()` at the **top** of each iteration:
```python
for warmup_step in range(2):
self._device_module.synchronize()
self._tp_group.barrier()
with self._precarve.measure():
output = forward_fn()
...
if post_warmup_hook is not None:
post_warmup_hook()
... nothing drains the last warmup's own async work ...
graph = torch.cuda.CUDAGraph()
```
Those syncs order work issued *before* each warmup. Nothing drains the **final**
warmup's own async work.
Impact
Lazy / JIT kernel compilation kicked off by a first-seen shape can still be in flight
when capture starts. A JIT module that finishes loading mid-capture issues illegal
driver calls (`cuModuleLoadData` and friends) on the capturing stream, surfacing as
`CUDA_ERROR_ILLEGAL_ADDRESS` or `cudaErrorStreamCaptureUnsupported`.
This is intermittent by nature — it depends on whether the last warmup happened to
trigger a compile and whether that compile lands before or after capture begins.
Fix
Drain the device and re-align ranks once more, immediately before the capture.
Applied to **both** backends. sgl-project#33795 fixed
`FullCudaGraphBackend`; `BreakableCudaGraphBackend` has the identical pattern and was
left unfixed.
Verification
Reasoning-only. The ordering gap is visible in the source and the fix is the same
drain sgl-project#33795 already established as correct for the sibling backend. I have not
constructed a reproducer — the failure is a race and reproducing it reliably would
mean forcing a JIT compile on the last warmup shape.
Cost is two extra synchronization points per captured graph, at capture time only.
No effect on the replay path.
---
*Provenance: originally authored by an automated kernel-optimization agent
(hyperloom session `20260806T082051Z`), rebased onto current `main` and reviewed by
hand before submission.*
jiaqiang-dot-liu
force-pushed
the
fix/cudagraph-drain-before-capture
branch
from
September 15, 2026 03:09
71c206f to
eab52c7
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Both CUDA-graph backends run two warmup iterations before capture, with a
synchronize()+barrier()at the top of each iteration:Those syncs order work issued before each warmup. Nothing drains the final
warmup's own async work.
Impact
Lazy / JIT kernel compilation kicked off by a first-seen shape can still be in flight
when capture starts. A JIT module that finishes loading mid-capture issues illegal
driver calls (
cuModuleLoadDataand friends) on the capturing stream, surfacing asCUDA_ERROR_ILLEGAL_ADDRESSorcudaErrorStreamCaptureUnsupported.This is intermittent by nature — it depends on whether the last warmup happened to
trigger a compile and whether that compile lands before or after capture begins.
Fix
Drain the device and re-align ranks once more, immediately before the capture.
Applied to both backends. sgl-project#33795 fixed
FullCudaGraphBackend;BreakableCudaGraphBackendhas the identical pattern and wasleft unfixed.
Verification
Reasoning-only. The ordering gap is visible in the source and the fix is the same
drain sgl-project#33795 already established as correct for the sibling backend. I have not
constructed a reproducer — the failure is a race and reproducing it reliably would
mean forcing a JIT compile on the last warmup shape.
Cost is two extra synchronization points per captured graph, at capture time only.
No effect on the replay path.
Provenance: originally authored by an automated kernel-optimization agent
(hyperloom session
20260806T082051Z), rebased onto currentmainand reviewed byhand before submission.
CI States
Latest PR Test (Base): ❌ Run #34923945812
Latest PR Test (Extra): ❌ Run #34923945548
Latest PR Test (AMD ROCm 10): ❌ Run #34923945923