Skip to content

[Profiler][GPU] Extend CUDA graph capture profiling to the V2 model runner - #54061

Open
devalshahamd wants to merge 19 commits into
vllm-project:mainfrom
devalshahamd:main
Open

devalshahamd wants to merge 19 commits into
vllm-project:mainfrom
devalshahamd:main

Conversation

@devalshahamd

@devalshahamd devalshahamd commented Aug 27, 2026 •

Copy link
Copy Markdown
Contributor

This PR extends the CUDA graph capture profiling added in #37524 to the V2 model runner, and to the encoder capture path in both runners. Capture tracing itself is unchanged: the capture_torch_profiler flag, the capture_traces/ layout and the capture_{batch_size}_{mode} annotation all behave as merged.

Shared capture-profiling helper:

  • New vllm/profiler/graph_capture.py. graph_capture_profiler(vllm_config, subsystem, label_prefix) builds one profiler for one capture subsystem and binds it to a ContextVar; graph_capture_step(num_tokens, mode) enters that profiler and records the annotation for a single captured shape, and is a no-op when profiling is not active.

V2 model runner:

  • gpu/model_runner.py: capture_model() binds a profiler per subsystem — encoder, decoder and speculator — so each writes its own trace file and annotation prefix.
  • gpu/cudagraph_utils.py: annotates the three capture execution points in CudaGraphManager.capture() (PIECEWISE reusing attention state, PIECEWISE with fresh state, and FULL). Only the capture call sits inside the profiled region.

Encoder capture (V1 and V2):

  • encoder_cudagraph.py: annotates the per-budget encoder graph capture.
  • gpu_model_runner.py: V1 capture_model() binds the encoder profiler. The throwaway capture used for encoder memory estimation is intentionally not included in the traces.

Trace output — one profiler enter/exit per captured shape, so each shape lands in its own file:

Subsystem Trace file Annotation
Decoder graph_capture_rank_0.* capture_{n}_{FULL,PIECEWISE}
Encoder graph_capture_rank_0_encoder.* capture_{n}_encoder_{path}
Speculator graph_capture_rank_0_speculator.* capture_{n}_draft_{FULL,PIECEWISE}

Unit tests:

  • TestGraphCaptureProfiler in tests/v1/worker/test_gpu_profiler.py: graph_capture_step is a no-op outside a binding, the bound profiler is entered once per captured shape, annotation labels are correct for all three subsystems, the binding is cleared when capture raises, the trace file name and directory are correct per subsystem, and profiling falls back to nullcontext when disabled or off rank 0.

Purpose

#37524 instruments only the V1 decoder capture loop, so a run on the V2 model runner (VLLM_USE_V2_MODEL_RUNNER=1) produces an empty capture_traces/ directory, and encoder capture is untraced in both runners. This closes the V2 gap reported in #53767.

Note that #53767 also asks for capture profiling to go through TorchProfilerWrapper rather than constructing torch.profiler.profile directly. This PR does not change that — it keeps the merged behaviour and only extends coverage, so capture_torch_profiler still ignores the torch_profiler_record_shapes / _with_memory / _with_stack / _with_flops / _use_gzip fields. The main goal of graph capture tracing is to recover the CPU callstack for graph replay region. Hence, the torch profiler configs related to CPU callstack and record_shapes are set True independent of the profiler config.

Test Result

Environment: 1× MI355X (gfx950), TP=1, ROCm 7.2.3, PyTorch 2.11.0, ROCm base images for v0.28.0 with this change applied. Traces collected with
--profiler-config.capture_torch_profiler True --profiler-config.detailed_trace_annotation True.

1. Test summary

Tests executed:

vLLM Workload Shards (decoder + subsystem) Graph capture
v0.28.0 Qwen3-8B + EAGLE3, V2 runner 100 + 151 speculator 234 s
v0.28.0 Qwen2-VL-2B-Instruct, encoder cudagraphs (budgets [256,512]) 102 + 2 encoder 65 s
v0.28.0 gpt-oss-20b MXFP4 (MoE → legacy V1 runner) 166 222 s

The gpt-oss-20b run confirms the legacy V1 path still behaves as before this change: 166 shards for
83 capture sizes × {PIECEWISE, FULL}.

Details on Qwen3-8B + EAGLE3, V2 runner:

Build Model runner Graph-capture shards emitted
stock v0.26.0 V2 (default for this model) 0 — no capture_traces/ directory created
stock v0.26.0 V1 (VLLM_USE_V2_MODEL_RUNNER=0) 98
patched V2 (default) 251 (100 decoder + 151 speculator)

2. Subsystem separation and annotation naming

Each subsystem writes its own trace-file group.

  • graph_capture_rank_0.* — capture_104_FULL, capture_104_PIECEWISE, … (100 distinct
    annotations across 100 shards)
  • graph_capture_rank_0_speculator.* — capture_104_draft_FULL, capture_104_draft_PIECEWISE, …
    (151 shards, 102 distinct annotations; the 49 *_draft_FULL shapes captured by both drafter
    capture invocations appear once per invocation, each in its own shard file)
  • graph_capture_rank_0_encoder.* — capture_256_encoder_default, capture_512_encoder_default

3. Cost and opt-out (v0.28.0, Qwen3-8B + EAGLE3, 251 captured shapes)

capture_torch_profiler Graph capture time Captured graph memory
True 234 s 5.33 GiB
False 50 s 5.33 GiB
profiler flags absent entirely 52 s 5.33 GiB

Disabling the flag returns capture time to the no-flag baseline within noise (50 s vs 52 s), so the
instrumentation is inert when off. Captured graph memory is identical in all three cases, i.e. the
profiler does not perturb what gets captured.


Essential Elements of an Effective PR Description Checklist

  • [Y] The purpose of the PR, such as "Fix some issue (link existing issues this PR will resolve)".
  • [Y] The test plan, such as providing test command.
  • [Y] The test results, such as pasting the results comparison before and after, or e2e results
  • (Optional) The necessary documentation update.
  • (Optional) Release notes update.

Signed-off-by: Deval Shah <devashah@amd.com>
@mergify mergify Bot added nvidia mrv2 Model Runner V2 specific labels Aug 27, 2026
@devalshahamd
devalshahamd marked this pull request as ready for review August 31, 2026 21:31

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@devalshahamd devalshahamd changed the title [Profiler] Extend CUDA graph capture profiling to the V2 model runner [Profiler][GPU] Extend CUDA graph capture profiling to the V2 model runner Aug 31, 2026
@mergify

mergify Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @devalshahamd.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

Signed-off-by: Deval Shah <devashah@amd.com>
@coderabbitai

coderabbitai Bot commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Team

Run ID: 71a79504-36af-4aee-9b92-bee555e350df

📥 Commits

Reviewing files that changed from the base of the PR and between 58ad1f3 and 5268c4f.

📒 Files selected for processing (6)
  • tests/v1/worker/test_gpu_profiler.py
  • vllm/profiler/graph_capture.py
  • vllm/v1/worker/encoder_cudagraph.py
  • vllm/v1/worker/gpu/cudagraph_utils.py
  • vllm/v1/worker/gpu/model_runner.py
  • vllm/v1/worker/gpu_model_runner.py
🚧 Files skipped from review as they are similar to previous changes (5)
  • vllm/v1/worker/encoder_cudagraph.py
  • vllm/v1/worker/gpu/cudagraph_utils.py
  • vllm/profiler/graph_capture.py
  • vllm/v1/worker/gpu/model_runner.py
  • vllm/v1/worker/gpu_model_runner.py

Included review availability: Your plan provides up to 10 included reviews per hour; 9 remain after this review.


📝 Summary

Summary by CodeRabbit

  • New Features
    • Added optional profiling for CUDA/HIP graph capture operations.
    • Capture traces include token counts, capture modes, and subsystem-specific labels.
    • Traces are saved as compressed files in the configured capture-trace directory.
    • Profiling remains inactive when disabled or on non-primary ranks.
  • Bug Fixes
    • Improved cleanup when graph capture profiling encounters a failure.

Walkthrough

The change adds optional rank-zero torch profiler tracing for graph capture. It adds capture-step annotations, subsystem-specific trace names, lifecycle cleanup, and instrumentation for encoder, decoder, speculator, piecewise, and standard CUDA graph captures.

Changes

Graph capture profiling

Layer / File(s) Summary
Profiler contexts and lifecycle tests
vllm/profiler/graph_capture.py, tests/v1/worker/test_gpu_profiler.py
Adds conditional profiler creation, ContextVar binding, subsystem trace naming, capture annotations, cleanup, and tests for enabled, disabled, off-rank, and failure cases.
Capture-step instrumentation
vllm/v1/worker/encoder_cudagraph.py, vllm/v1/worker/gpu/cudagraph_utils.py
Records token counts and graph modes during encoder, piecewise, and standard CUDA graph capture.
Subsystem profiler wiring
vllm/v1/worker/gpu/model_runner.py, vllm/v1/worker/gpu_model_runner.py
Wraps encoder, decoder, and speculator capture entry points with subsystem-specific profiler contexts and labels.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Merge Risk: ⚪ Minimal · up to 5268c

This change adds optional rank-zero graph-capture profiling with separate encoder, decoder, and speculator traces while retaining disabled-mode behavior. No current merge-readiness risk is identified.

Sequence Diagram(s)

sequenceDiagram
  participant capture_model
  participant graph_capture_profiler
  participant graph_capture_step
  participant torch_profiler
  participant capture_loop
  capture_model->>graph_capture_profiler: enter subsystem profiling
  graph_capture_profiler->>torch_profiler: create profiler when enabled on rank 0
  capture_model->>capture_loop: run graph capture
  capture_loop->>graph_capture_step: record tokens and mode
  graph_capture_step->>torch_profiler: emit capture annotation
  graph_capture_profiler->>torch_profiler: write compressed trace
Loading
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 41.18% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 17 functions across 6 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly and concisely describes the main change: extending CUDA graph capture profiling to the V2 model runner. It is specific and aligned with the changeset.
Description check ✅ Passed The description directly explains the profiling helper, V2 and encoder coverage, trace separation, tests, and ROCm validation. It is fully related to the changeset.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
  • Fix all pre-merge checks with AI

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@mergify mergify Bot removed the needs-rebase label Sep 3, 2026
@tjtanaa

tjtanaa commented Sep 7, 2026

Copy link
Copy Markdown
Member

@njhill @yewentao256, can you help to review this PR?

@tjtanaa tjtanaa left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM but let's get one more pair of eyes

@github-project-automation github-project-automation Bot moved this to Ready in NVIDIA Sep 7, 2026
@tjtanaa tjtanaa added the ready ONLY add when PR is ready to merge/full CI is needed label Sep 7, 2026
@tjtanaa

tjtanaa commented Sep 7, 2026

Copy link
Copy Markdown
Member

/ci run

@coderabbitai

coderabbitai Bot commented Sep 7, 2026

Copy link
Copy Markdown
Contributor

Note

GitHub couldn't provide a complete incremental comparison for this pull request, so CodeRabbit is performing a full review instead. This review may take a little longer.

@tjtanaa tjtanaa added verified Run pre-commit for new contributors without triggering other tests and removed ready ONLY add when PR is ready to merge/full CI is needed labels Sep 7, 2026
@github-actions

github-actions Bot commented Sep 7, 2026

Copy link
Copy Markdown

✅ @devalshahamd, CI is now available for this PR.

  • /ci run starts upstream CI; /amd-ci run starts AMD CI only.
  • /ci retry retries failed jobs in the CI build for the current PR head. If the current head has no CI build, it starts a new CI build for the current head containing only jobs that failed in the latest earlier CI build for this PR.
  • /amd-ci retry retries failed jobs in AMD CI for the current PR head. Use /amd-ci run when the current head has no AMD CI build.
  • /ci cancel cancels scheduled or running CI builds for this PR branch; /amd-ci cancel does the same for AMD CI only.

@github-actions

github-actions Bot commented Sep 7, 2026

Copy link
Copy Markdown

✅ Triggered Buildkite CI #87537 for commit 5268c4f48371.

Signed-off-by: devalshahamd <deval.shah@amd.com>
@mergify

mergify Bot commented Sep 15, 2026

Copy link
Copy Markdown
Contributor

Hi @devalshahamd, the pre-commit checks have failed. Please run:

uv pip install pre-commit>=4.5.1
pre-commit install
pre-commit run --all-files

Then, commit the changes and push to your branch.

For future commits, pre-commit will run automatically on changed files before each commit.

devalshah-amd and others added 2 commits September 14, 2026 20:22
Co-authored-by: Cursor Grok 4.6 <cursor@cursor.com>
Signed-off-by: Deval Shah <devashah@amd.com>
Comment thread vllm/v1/worker/gpu/model_runner.py Outdated
Co-authored-by: Cursor Grok 4.6 <cursor@cursor.com>
Signed-off-by: Deval Shah <devashah@amd.com>
@mergify

mergify Bot commented Sep 16, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @devalshahamd.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify mergify Bot added the needs-rebase label Sep 16, 2026
Signed-off-by: devalshahamd <deval.shah@amd.com>
@mergify

mergify Bot commented Sep 16, 2026

Copy link
Copy Markdown
Contributor

Hi @devalshahamd, the pre-commit checks have failed. Please run:

uv pip install pre-commit>=4.5.1
pre-commit install
pre-commit run --all-files

Then, commit the changes and push to your branch.

For future commits, pre-commit will run automatically on changed files before each commit.

@mergify mergify Bot removed the needs-rebase label Sep 16, 2026
@mergify

mergify Bot commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @devalshahamd.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify mergify Bot added the needs-rebase label Sep 17, 2026
devalshah-amd and others added 2 commits September 17, 2026 13:37
Signed-off-by: Deval Shah <devashah@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor Grok 4.6 <cursor@cursor.com>
Signed-off-by: Deval Shah <devashah@amd.com>
@mergify mergify Bot removed the needs-rebase label Sep 17, 2026
@devalshahamd

Copy link
Copy Markdown
Contributor Author

Hello @tjtanaa, @njhill
I have addressed all comments. Now the code uses TorchProfilerWraper to construct the profiler. All graph capture profiling related functions are in one central location, which is better for maintainability in the future. Please let me know if anything else needs to be added.

@mergify

mergify Bot commented Sep 21, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @devalshahamd.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify mergify Bot added the needs-rebase label Sep 21, 2026
Signed-off-by: Deval Shah <devashah@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
@mergify mergify Bot removed the needs-rebase label Sep 24, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

dflash mrv2 Model Runner V2 specific nvidia speculative-decoding verified Run pre-commit for new contributors without triggering other tests

Projects

Status: In review
Status: Backlog

Development

Successfully merging this pull request may close these issues.

5 participants