[Profiler][GPU] Extend CUDA graph capture profiling to the V2 model runner - #54061
devalshahamd wants to merge 19 commits into
Conversation
Signed-off-by: Deval Shah <devashah@amd.com>
|
This pull request has merge conflicts that must be resolved before it can be |
Signed-off-by: Deval Shah <devashah@amd.com>
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository UI Review profile: CHILL Plan: Team Run ID: 📒 Files selected for processing (6)
🚧 Files skipped from review as they are similar to previous changes (5)
Included review availability: Your plan provides up to 10 included reviews per hour; 9 remain after this review. 📝 SummarySummary by CodeRabbit
WalkthroughThe change adds optional rank-zero torch profiler tracing for graph capture. It adds capture-step annotations, subsystem-specific trace names, lifecycle cleanup, and instrumentation for encoder, decoder, speculator, piecewise, and standard CUDA graph captures. ChangesGraph capture profiling
Estimated code review effort: 3 (Moderate) | ~20 minutes Merge Risk: ⚪ Minimal · up to This change adds optional rank-zero graph-capture profiling with separate encoder, decoder, and speculator traces while retaining disabled-mode behavior. No current merge-readiness risk is identified. Sequence Diagram(s)sequenceDiagram
participant capture_model
participant graph_capture_profiler
participant graph_capture_step
participant torch_profiler
participant capture_loop
capture_model->>graph_capture_profiler: enter subsystem profiling
graph_capture_profiler->>torch_profiler: create profiler when enabled on rank 0
capture_model->>capture_loop: run graph capture
capture_loop->>graph_capture_step: record tokens and mode
graph_capture_step->>torch_profiler: emit capture annotation
graph_capture_profiler->>torch_profiler: write compressed trace
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
@njhill @yewentao256, can you help to review this PR? |
tjtanaa
left a comment
There was a problem hiding this comment.
LGTM but let's get one more pair of eyes
|
/ci run |
|
Note GitHub couldn't provide a complete incremental comparison for this pull request, so CodeRabbit is performing a full review instead. This review may take a little longer. |
|
✅ @devalshahamd, CI is now available for this PR.
|
|
✅ Triggered Buildkite CI #87537 for commit |
Signed-off-by: devalshahamd <deval.shah@amd.com>
|
Hi @devalshahamd, the pre-commit checks have failed. Please run: uv pip install pre-commit>=4.5.1
pre-commit install
pre-commit run --all-filesThen, commit the changes and push to your branch. For future commits, |
Co-authored-by: Cursor Grok 4.6 <cursor@cursor.com> Signed-off-by: Deval Shah <devashah@amd.com>
Co-authored-by: Cursor Grok 4.6 <cursor@cursor.com> Signed-off-by: Deval Shah <devashah@amd.com>
|
This pull request has merge conflicts that must be resolved before it can be |
Signed-off-by: devalshahamd <deval.shah@amd.com>
|
Hi @devalshahamd, the pre-commit checks have failed. Please run: uv pip install pre-commit>=4.5.1
pre-commit install
pre-commit run --all-filesThen, commit the changes and push to your branch. For future commits, |
|
This pull request has merge conflicts that must be resolved before it can be |
Signed-off-by: Deval Shah <devashah@amd.com> Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor Grok 4.6 <cursor@cursor.com> Signed-off-by: Deval Shah <devashah@amd.com>
|
Hello @tjtanaa, @njhill |
|
This pull request has merge conflicts that must be resolved before it can be |
Signed-off-by: Deval Shah <devashah@amd.com> Co-authored-by: Cursor <cursoragent@cursor.com>
This PR extends the CUDA graph capture profiling added in #37524 to the V2 model runner, and to the encoder capture path in both runners. Capture tracing itself is unchanged: the
capture_torch_profilerflag, thecapture_traces/layout and thecapture_{batch_size}_{mode}annotation all behave as merged.Shared capture-profiling helper:
vllm/profiler/graph_capture.py.graph_capture_profiler(vllm_config, subsystem, label_prefix)builds one profiler for one capture subsystem and binds it to aContextVar;graph_capture_step(num_tokens, mode)enters that profiler and records the annotation for a single captured shape, and is a no-op when profiling is not active.V2 model runner:
gpu/model_runner.py:capture_model()binds a profiler per subsystem — encoder, decoder and speculator — so each writes its own trace file and annotation prefix.gpu/cudagraph_utils.py: annotates the three capture execution points inCudaGraphManager.capture()(PIECEWISE reusing attention state, PIECEWISE with fresh state, and FULL). Only the capture call sits inside the profiled region.Encoder capture (V1 and V2):
encoder_cudagraph.py: annotates the per-budget encoder graph capture.gpu_model_runner.py: V1capture_model()binds the encoder profiler. The throwaway capture used for encoder memory estimation is intentionally not included in the traces.Trace output — one profiler enter/exit per captured shape, so each shape lands in its own file:
graph_capture_rank_0.*capture_{n}_{FULL,PIECEWISE}graph_capture_rank_0_encoder.*capture_{n}_encoder_{path}graph_capture_rank_0_speculator.*capture_{n}_draft_{FULL,PIECEWISE}Unit tests:
TestGraphCaptureProfilerintests/v1/worker/test_gpu_profiler.py:graph_capture_stepis a no-op outside a binding, the bound profiler is entered once per captured shape, annotation labels are correct for all three subsystems, the binding is cleared when capture raises, the trace file name and directory are correct per subsystem, and profiling falls back tonullcontextwhen disabled or off rank 0.Purpose
#37524 instruments only the V1 decoder capture loop, so a run on the V2 model runner (
VLLM_USE_V2_MODEL_RUNNER=1) produces an emptycapture_traces/directory, and encoder capture is untraced in both runners. This closes the V2 gap reported in #53767.Note that #53767 also asks for capture profiling to go through
TorchProfilerWrapperrather than constructingtorch.profiler.profiledirectly. This PR does not change that — it keeps the merged behaviour and only extends coverage, socapture_torch_profilerstill ignores thetorch_profiler_record_shapes/_with_memory/_with_stack/_with_flops/_use_gzipfields. The main goal of graph capture tracing is to recover the CPU callstack for graph replay region. Hence, the torch profiler configs related to CPU callstack and record_shapes are set True independent of the profiler config.Test Result
Environment: 1× MI355X (gfx950), TP=1, ROCm 7.2.3, PyTorch 2.11.0, ROCm base images for v0.28.0 with this change applied. Traces collected with
--profiler-config.capture_torch_profiler True --profiler-config.detailed_trace_annotation True.1. Test summary
Tests executed:
[256,512])The gpt-oss-20b run confirms the legacy V1 path still behaves as before this change: 166 shards for
83 capture sizes × {PIECEWISE, FULL}.
Details on Qwen3-8B + EAGLE3, V2 runner:
capture_traces/directory createdVLLM_USE_V2_MODEL_RUNNER=0)2. Subsystem separation and annotation naming
Each subsystem writes its own trace-file group.
graph_capture_rank_0.*—capture_104_FULL,capture_104_PIECEWISE, … (100 distinctannotations across 100 shards)
graph_capture_rank_0_speculator.*—capture_104_draft_FULL,capture_104_draft_PIECEWISE, …(151 shards, 102 distinct annotations; the 49
*_draft_FULLshapes captured by both draftercapture invocations appear once per invocation, each in its own shard file)
graph_capture_rank_0_encoder.*—capture_256_encoder_default,capture_512_encoder_default3. Cost and opt-out (v0.28.0, Qwen3-8B + EAGLE3, 251 captured shapes)
capture_torch_profilerTrueFalseDisabling the flag returns capture time to the no-flag baseline within noise (50 s vs 52 s), so the
instrumentation is inert when off. Captured graph memory is identical in all three cases, i.e. the
profiler does not perturb what gets captured.
Essential Elements of an Effective PR Description Checklist