Skip to content

[Bugfix][MRV2] Reserve encoder graph memory without decoder graphs - #58243

Open
BWAAEEEK wants to merge 1 commit into
vllm-project:mainfrom
BWAAEEEK:fix/mrv2-encoder-cudagraph-memory
Open

BWAAEEEK wants to merge 1 commit into
vllm-project:mainfrom
BWAAEEEK:fix/mrv2-encoder-cudagraph-memory

Conversation

@BWAAEEEK

Copy link
Copy Markdown
Contributor

Purpose

While following up on the MRV2 feedback on my earlier PR, I revisited CUDA graph memory accounting in the V2 runner. This led me to a missing case similar to the encoder graph memory accounting I previously worked on in #41714, but in a different runner and configuration.

MRV2 can capture encoder CUDA graphs independently of decoder graphs. With cudagraph_mm_encoder=True and cudagraph_mode=NONE, the encoder still captures graphs, but the Worker skips graph memory profiling. The V2 profiler also returns early based on decoder-only checks. Consequently, automatic KV-cache sizing does not reserve memory for those encoder graphs.

This PR:

  • Allows the Worker to profile MRV2 encoder graphs when decoder graphs are disabled, excluding enforce_eager.
  • Uses the runner's existing encoder/decoder capture predicate in the V2 profiler. The pre-bootstrap decoder configuration check remains necessary because the decoder manager has not been initialized yet.
  • Preserves MRV1 behavior: MRV1 does not perform encoder capture when decoder graph mode is NONE, so it should not reserve memory for that case.
  • Adds regression coverage for KV-budget subtraction, an uninitialized decoder manager, encoder capture without decoder descriptors, and cleanup on capture failure.

No kernels, capture sizes, batch-size limits, pool-sharing policy, or memory-estimation formulas are changed. This concerns normal multimodal generation, not the separate --mm-encoder-only deployment mode.

Related Work / Duplicate Check

I checked #49224 and its comments, open PR references to that issue, and open/merged PRs for encoder graph memory, cudagraph_mode=NONE, and MRV2 profiling.

The first two open accounting PRs overlap files with this change, but address different failures. This PR deliberately leaves their accounting-policy changes out of scope.

Test Plan

Validation used a fresh Python 3.12 environment and a standard editable installation with the official CUDA 13.0 wheel for the exact main-base commit a6c47fbbf4187e8ae5396ebdcf93234db617a4ad. Both baseline and fixed sources used the same dependencies and CUDA binaries. All 18 installed extension files matched that wheel by SHA256; loaded-library paths were also checked. There are no C++/CUDA changes in this PR.

The focused suites were run both through the environment-verification wrapper and directly with pytest. The direct command, with temporary/cache directories configured outside the repository, was:

../encoder-cg-matched-20260922/.venv/bin/python -m pytest \
  tests/v1/worker/test_gpu_worker.py \
  tests/v1/worker/test_gpu_model_runner_v2_cudagraph_profiling.py \
  tests/v1/cudagraph/test_encoder_cudagraph.py -q \
  --basetemp=../encoder-cg-matched-20260922/pytest-direct \
  --junitxml=../encoder-cg-matched-20260922/pytest-direct.xml

The regression subset was also run against an unmodified main worktree using the final tests and --import-mode=importlib.

Model checks use NVIDIA B200, BF16, TP=1, Torch 2.13.0+cu130, and the actual Qwen/Qwen2.5-VL-3B-Instruct weights at revision 66285546d2b821cf421d4f5eb2576359d3770cd3. Configuration: max_model_len=2048, max_num_seqs=2, max_num_batched_tokens=1024, gpu_memory_utilization=0.12, prefix caching disabled, FLASHINFER encoder attention, and one image/no video per prompt.

The model harness observes the real profiling/capture returns, KV budget, temporary-pool release, production recapture/replay, and generated token IDs. A separate check runs the repository's vllm bench throughput with actual multimodal requests, rather than a synthetic kernel timing loop.

Multimodal benchmark configuration

The local invocation uses a read-only wrapper that asserts one image per request before calling the unchanged benchmark CLI. Its standard-install equivalent is:

VLLM_USE_V2_MODEL_RUNNER=1 VLLM_ENABLE_V1_MULTIPROCESSING=0 \
  .venv/bin/vllm bench throughput \
  --backend vllm-chat --model Qwen/Qwen2.5-VL-3B-Instruct \
  --revision 66285546d2b821cf421d4f5eb2576359d3770cd3 \
  --dtype bfloat16 --seed 0 \
  --dataset-name random-mm --enable-multimodal-chat \
  --dataset-path synthetic-random-mm \
  --num-prompts 32 --num-warmups 4 \
  --random-input-len 32 --random-output-len 16 --random-range-ratio 0 \
  --random-mm-base-items-per-request 1 \
  --random-mm-num-mm-items-range-ratio 0 \
  --random-mm-limit-mm-per-prompt '{"image":1,"video":0}' \
  --random-mm-bucket-config '{(224,224,1):0.5,(448,448,1):0.5}' \
  --max-model-len 2048 --max-num-seqs 2 --max-num-batched-tokens 1024 \
  --gpu-memory-utilization 0.12 --no-enable-prefix-caching \
  --mm-processor-cache-gb 0 --mm-encoder-attn-backend FLASHINFER \
  --limit-mm-per-prompt '{"image":1,"video":0}' \
  --mm-processor-kwargs '{"min_pixels":3136,"max_pixels":200704}' \
  --compilation-config '{"mode":3,"cudagraph_mode":"NONE","cudagraph_mm_encoder":true,"encoder_cudagraph_token_budgets":[512,1024,2048],"encoder_cudagraph_max_vision_items_per_batch":2}'

The nonempty synthetic dataset-path marker avoids the current CLI fallback to text-only random; the random multimodal dataset does not read that path. The local run uses the pinned downloaded snapshot instead of fetching weights.

Test Result

Tested base: a6c47fbbf4187e8ae5396ebdcf93234db617a4ad.
Tested PR head: 7ac8ba13be40e48b81984a73fcfd217f2102d52e.

  • Focused CPU/GPU suites: 92 passed, including real graph allocation/capture/replay tests.
  • Unmodified submission-base negative control: 4 failed, 14 passed in the selected subset. The four failures are the Worker admission, the two V2 profiler admission cases, and the previously skipped encoder capture-error path. All pass with the fix.
  • Changed-file pre-commit hooks and the manual Python 3.12 mypy check passed.
  • uv pip check: all 228 installed packages are compatible.

Tested-base real-model checks:

Configuration Graph reservation Production capture increment Replay / outputs
Unmodified base, encoder ON / decoder NONE, budget 1024 0 MiB 262 MiB 2 graph hits
Fixed, same graph configuration 260 MiB 258 MiB 2 hits, tokens identical to base
Fixed, budgets 512/1024/2048, max vision items 2 534 MiB 530 MiB 4 hits, including a concurrent pair; tokens identical to base

Both fixed runs released their temporary encoder profiling pools completely before production recapture and had zero replay misses. In each run, the KV budget equals requested memory minus the observed forward-profile non-KV footprint minus the graph reservation. In the matched-environment single-budget comparison, the forward-profile footprint was identical: KV budget changed from 14,439,960,146 to 14,167,330,386 bytes, exactly the 260 MiB reservation.

The full nine-configuration model matrix was rerun on this submission revision and matched wheel:

Fixed configuration Applied graph reservation Result
Encoder ON, decoder NONE, compilation enabled 260 MiB Pass
Encoder ON, decoder NONE, compilation disabled 228 MiB Pass
Both graph types disabled, compilation enabled 0 MiB Pass
Both graph types disabled, compilation disabled 0 MiB Pass
Encoder and decoder graphs enabled 708 MiB Pass
Decoder graphs only 448 MiB Pass
Encoder ON, reservation opt-out 0 MiB (260 MiB measured) Pass
Eager 0 MiB Pass
Encoder budgets 512/1024/2048, max vision items 2 534 MiB Pass

All generated smoke-test token IDs matched the baseline, including the concurrent pair. Temporary encoder pools were released in every profiled configuration, and every enabled encoder graph run had zero replay misses.

The matched-environment multimodal benchmark completed four warmup and 32 measured image requests per run, with concurrency capped at two. Both runs processed 5,856 prompt tokens and generated 512 output tokens. The wrapper confirmed that every request actually contained one image.

Source Elapsed Requests/s Output tokens/s
Unmodified submission base 4.0719 s 7.8588 125.74
Fixed 3.8608 s 8.2884 132.62

Both CLI runs exited successfully and emitted the same NCCL process-group shutdown warning at exit. These are single short runs checking successful multimodal processing, not evidence of a statistically significant speedup or performance equivalence. This PR fixes a demonstrated memory-reservation omission, not a reproduced OOM or throughput bottleneck.

AI assistance was used for implementation, validation, and drafting.

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@mergify mergify Bot added nvidia mrv2 Model Runner V2 specific labels Sep 23, 2026
@mergify mergify Bot added the bug Something isn't working label Sep 23, 2026
@mergify

mergify Bot commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @BWAAEEEK.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify mergify Bot added the needs-rebase label Sep 25, 2026
@BWAAEEEK
BWAAEEEK force-pushed the fix/mrv2-encoder-cudagraph-memory branch from 7ac8ba1 to daf06cf Compare September 25, 2026 07:21
@mergify mergify Bot removed the needs-rebase label Sep 25, 2026
@mergify

mergify Bot commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @BWAAEEEK.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify mergify Bot added the needs-rebase label Sep 25, 2026
Allow the worker and V2 graph profiler to account for encoder CUDA graphs
when decoder graphs are disabled. Keep the pre-bootstrap decoder check,
MRV1 behavior, eager exclusion, and graph-reservation opt-out semantics.

Add regression coverage for profiling eligibility, KV budget subtraction,
and cleanup when encoder graph capture fails.

Assisted-by: OpenAI Codex
Signed-off-by: BWAAEEEK <jooho414@gmail.com>
@BWAAEEEK
BWAAEEEK force-pushed the fix/mrv2-encoder-cudagraph-memory branch from daf06cf to 8c75d67 Compare September 26, 2026 07:19
@mergify mergify Bot removed the needs-rebase label Sep 26, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working mrv2 Model Runner V2 specific nvidia

Projects

Status: No status

1 participant