Skip to content

[Perf] Add CUDA graphs for MiniCPM-o 4.5 input encoders - #8332

Merged
natureofnature merged 8 commits into
vllm-project:mainfrom
amy-why-3459:perf-minicpmo-encoder-cuda-graph
Oct 5, 2026
Merged

natureofnature merged 8 commits into
vllm-project:mainfrom
amy-why-3459:perf-minicpmo-encoder-cuda-graph

Conversation

@amy-why-3459

@amy-why-3459 amy-why-3459 commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose

MiniCPM-o 4.5 input encoders launch their transformer layers eagerly for every input. Add bounded, exact-shape CUDA Graph replay for the SigLIP transformer stack and stateless audio encoder/projection/pooling.

The model-local exact-shape adapter now delegates capture, static input/output buffers, replay and hit accounting to upstream EncoderCudaGraphManager through SupportsEncoderCudaGraph. One already-packed local batch is one indivisible manager item, preserving existing attention and padding behavior. By default, capture starts on the second occurrence of a shape, with at most four graphs per encoder. HF overrides configure the per-encoder graph cap, minimum call count and pre-capture free-memory floor (default 1 GiB). Admission misses are exposed in adapter statistics. Capture failures propagate and prevent reuse until worker restart. See admission and failure semantics. Upstream replay refreshes input and mask buffers; model postprocessing clones outputs so subsequent requests cannot overwrite retained embeddings. Packing, positional/mask construction, resampling and output slicing retain their existing behavior.

Keep eager execution for CPU, training/grad/autocast, nested capture, padded FlashAttention vision, FP16 Whisper, intermediate audio layers, and stateful streaming audio. --enforce-eager disables these graphs; --hf-overrides '{"encoder_cuda_graph": false}' provides an encoder-only A/B switch.

This PR contains implementation, regression tests and configuration documentation. Local benchmarks/minicpmo/ files are excluded.

Test Plan

vLLM version: 0.30.0; PyTorch 2.13.0+cu130.
vLLM-Omni base: 4af28f33bd.

pytest -q tests/model_executor/models/minicpmo_4_5/test_encoder_cuda_graph.py
pytest -q tests/model_executor/models/minicpmo_4_5/test_encoder_cuda_graph.py tests/model_executor/models/minicpmo_4_5/test_vision_batching.py tests/model_executor/models/minicpmo_4_5/test_vision_flash_attention.py tests/model_executor/models/minicpmo_4_5/test_audio_chunk_mask.py tests/model_executor/models/minicpmo_4_5/test_audio_pool_step.py tests/model_executor/models/minicpmo_4_5/test_streaming_audio_cache.py tests/model_executor/models/minicpmo_4_5/test_minicpmo_4_5_omni_llm_forward.py -m cpu
pre-commit run --files vllm_omni/model_executor/models/minicpmo_4_5/encoder_cuda_graph.py vllm_omni/model_executor/models/minicpmo_4_5/minicpmo_4_5_omni_llm.py tests/model_executor/models/minicpmo_4_5/test_encoder_cuda_graph.py tests/model_executor/models/minicpmo_4_5/test_vision_batching.py

Test Result

Latest implementation: 3732befb2d69a86f6cb80e5e59c51a54c4260b69. The model-local upstream manager integration remains in place; runner startup capture and padded token-budget batching are not introduced. Manager configuration copies remain local.

  • Scoped CPU/CUDA regression over the seven modules listed above: 258 passed, 4 skipped, 0 failed, 0 deselected, 0 collection errors. All four skips are FlashAttention2 cases because Transformers reports FA2 unavailable.
  • Environment: Python 3.12.13, vLLM 0.30.0, PyTorch 2.13.0+cu130, transformers 5.14.1, pytest 9.1.1; NVIDIA L20X, CUDA_VISIBLE_DEVICES=0, PYTHONPATH=.. The seven-module command was run without -m cpu, with -ra to report skips.
  • All applicable pre-commit checks on the four changed files passed, including Ruff, mypy, Markdown, test marks, SPDX and CUDA API checks.
  • New regressions cover configurable admission, capacity miss accounting, low-memory deferral and recovery, existing replay below the memory floor, fatal capture failure without retries/eager fallback, and zero capacity.
  • The complete repository unit-test matrix has not been run locally. These scoped results do not establish full-suite coverage or merge readiness. New-head CI and complete unit-test evidence remain required; previous-head results cannot validate this update.

Performance and accuracy

The upstream-manager A/B report measures 2d9b427bf49160c567bade1bcf9beb20cd550389. Those measurements precede the admission/failure-policy update at 3732befb2d69; A/B performance has not been remeasured for the latest head. It supersedes the earlier custom-implementation report; those earlier numbers do not describe this implementation.

Setup: one exclusively reserved NVIDIA L20X; vLLM 0.30.0, PyTorch 2.13.0+cu130 and transformers 5.14.1; the same MiniCPM-o 4.5 checkpoint and default vllm_omni/deploy/minicpmo_4_5.yaml. Only encoder_cuda_graph is toggled. Daily-Omni audio+video uses minicpm-interleave, concurrency 10, seed 0, temperature 0 and a 128-token text-output cap. Requesting 2000 prompts with --no-oversample produces 1197 requests per pass, matching the dataset size.

Protocol: exclude one full warmup pass per mode, then measure three passes per mode. Values below are mean ± sample standard deviation; percentile rows average the per-run percentiles. Enabled mode was measured first, disabled mode second. No encoder captures occurred during the six measured passes.

Metric Graph off Graph on Change
Request throughput (req/s) 1.039 ± 0.004 1.215 ± 0.022 +16.91%
Mean E2E latency (ms) 9608.64 ± 39.87 8177.76 ± 146.17 −14.89%
P99 E2E latency (ms) 16011.69 ± 427.51 13111.80 ± 2478.29 −18.11%
Mean TTFT (ms) 7796.88 ± 50.95 8009.22 ± 148.23 +2.72%
Peak device-used VRAM (GiB) 80.23 ± 0.00 87.81 ± 0.29 +9.45%
Accuracy on successful requests 78.237% ± 0.028% 78.243% ± 0.000% Similar
Accuracy including failed requests 78.084% ± 0.128% 78.112% ± 0.000% Similar
Failed requests per pass 2.33 ± 1.53 2.00 ± 0.00 See report

For paired passes, all requests successful in both modes had identical answer strings: 1195/1195, 1195/1195 and 1193/1193. This is not a zero-failure validation: disabled runs had 1/2/4 failures and enabled runs had 2/2/2. The report records stage-0 failures following multimodal cache-miss warnings; it does not establish their root cause.

Cold capture is excluded above: one startup capture took 0.166 s, and seven warmup captures took 1.302 s in total. VRAM was sampled using NVML every 100 ms and reflects device-used memory, not PyTorch allocated memory; shorter peaks may be missed. Enabled-run tail latency varied substantially. The sequential on/off experiment was not repeated in reverse order, so order effects remain a limitation.

These results show improved throughput and mean E2E latency for this workload, at the cost of higher mean TTFT and memory use. Stateful duplex audio remains eager; these end-to-end measurements do not establish a streaming-audio encoder speedup.

Relationship to #8430

This PR adds graph management around existing SigLIP and stateless-audio computations. It does not add streaming-audio KV batching, incremental Fbank, packed/fused vision kernels or candidate-space sampling.

#8430 addresses those broader Stage-0 optimizations and includes custom vision and streaming-audio graph managers. Its packed vision path calls forward_packed(), whereas this PR attaches the vision adapter to forward(). Combining the two requires explicitly integrating the graph adapter with the packed entry point; enabling both implementations does not automatically combine their benefits. The packed/fused computation could remain model-owned while upstream graph management is reused. Stateful audio additionally needs model-owned KV preparation and commit semantics.

Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
@amy-why-3459
amy-why-3459 marked this pull request as ready for review September 30, 2026 15:01
@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/model_integration.md, docs/design/module/ar_runtime.md.

Module owners: @tzhouam @fake0fan @Gaohan123

Routing: @tzhouam via module of the changed files, CODEOWNERS; @fake0fan via module of the changed files; @Gaohan123 via module of the changed files

@amy-why-3459, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit f4ff11ae6795 produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

@amy-why-3459 amy-why-3459 changed the title [Model] Add CUDA graphs for MiniCPM-o 4.5 input encoders [Perf] Add CUDA graphs for MiniCPM-o 4.5 input encoders Sep 30, 2026
@amy-why-3459

Copy link
Copy Markdown
Collaborator Author

Daily-Omni concurrency-10 A/B results for commit 4df8d0bb49375d710a44e7ac30e195b1198be1bb:

Setup: one exclusively reserved NVIDIA L20X, MiniCPM-o 4.5, vLLM 0.30.0, PyTorch 2.13.0+cu130. Both runs use vllm_omni/deploy/minicpmo_4_5.yaml; only --hf-overrides '{"encoder_cuda_graph": false/true}' changes. Encoder graphs disabled first, enabled second, on the same GPU.

Workload: Daily-Omni audio+video, minicpm-interleave packing, concurrency 10, seed 0, temperature 0, text output capped at 128 tokens. Requested 2000 prompts with --no-oversample; the dataset contains 1197 entries, so each run submitted 1197 requests.

Metric Encoder graph off Encoder graph on Change
Successful requests 1195/1197 1195/1197 —
Benchmark duration 1195.69 s 1004.49 s −15.99%
Request throughput 0.9994 req/s 1.1897 req/s +19.03%
Mean end-to-end latency 9980.12 ms 8356.58 ms −16.27%
Median end-to-end latency 9817.92 ms 8082.35 ms −17.68%
P99 end-to-end latency 16348.20 ms 12874.73 ms −21.25%
Mean TTFT 8206.26 ms 8087.80 ms −1.44%
P99 TTFT 13848.37 ms 11784.76 ms −14.90%
Total output tokens 2390 2390 —
Accuracy, successful requests 78.2427% 78.2427% unchanged
Accuracy, including failed requests 78.1119% 78.1119% unchanged

Correctness and validation:

  • Encoder regression test module on CUDA: 7 passed.
  • Scoped CPU regression suite listed in the PR description: 244 passed, 8 deselected.
  • Applicable local pre-commit hooks passed.
  • All 1194 requests successful in both runs produced identical answer strings.
  • Each run had two failures, but the failed samples were not identical: zero-based indices off [210, 833], on [210, 699]. The logs report stage-0 Stage request failed; the underlying cause has not been established. This is not a zero-failure validation.

Interpretation: This single sequential A/B pair measured a 19.03% throughput improvement and a 16.27% reduction in mean end-to-end latency. It includes cold graph capture costs and has not been repeated with reversed ordering to control for cache/order effects. Stateful streaming audio remains eager; these are end-to-end deployment results, not evidence of a stateless audio encoder speedup in the duplex path. The full repository test suite has not been run.

Client command (same for both server configurations):

vllm bench serve --omni --port 28973 --max-concurrency 10 \
  --dataset-name daily-omni --num-prompts 2000 --no-oversample \
  --trust-remote-code --temperature 0 --output-len 128 --seed 0 \
  --daily-omni-input-mode all --daily-omni-pack-mode minicpm-interleave \
  --daily-omni-video-dir /path/to/Daily-Omni/Videos \
  --daily-omni-qa-json /path/to/Daily-Omni/qa.json \
  --model /path/to/MiniCPM-o-4_5 \
  --endpoint /v1/chat/completions --backend openai-chat-omni \
  --percentile-metrics ttft,tpot,itl,e2el \
  --extra_body '{"modalities":["text"],"chat_template_kwargs":{"enable_thinking":false}}' \
  --save-result --save-detailed --result-dir /path/to/results \
  --result-filename off.json  # use on.json for the enabled run

AI assistance: OpenAI Codex inspected the saved result JSON files and logs, compared answer strings, and drafted this report.

Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
@amy-why-3459 amy-why-3459 added the ready label to trigger buildkite CI label Sep 30, 2026
@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot routing record

Assigned Strict under experiment vllm-omni-strict-5050-20260829.

@vllm-omni-review-bot vllm-omni-review-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Omni ReviewBot review

PR description

This PR adds exact-shape CUDA Graph replay for MiniCPM-o 4.5 SigLIP vision (transformer stack only) and stateless Whisper audio encode/project/pool, default-on unless --enforce-eager or hf_overrides.encoder_cuda_graph=false. A model-local EncoderCudaGraph admits up to four stream-scoped shapes (capture on the second hit) and delegates buffers/replay to upstream EncoderCudaGraphManager via a packed-batch _ExactShapeEncoder adapter. Host-side packing/masks, padded FA2 vision, FP16 Whisper, intermediate audio layers, and streaming KV audio stay eager; outputs are cloned so retained embeddings survive later replays.

Change flow

flowchart TD
  A["[EXISTING] Vision/audio encode callers<br/>vpm / get_audio_hidden_states"]:::existing --> B["[CHANGED] minicpmo_4_5_omni_llm.py<br/>gates + _encode_* extraction"]:::changed
  B --> C["[NEW] EncoderCudaGraph<br/>stream/shape admission"]:::new
  C --> D["[NEW] _ExactShapeEncoder<br/>SupportsEncoderCudaGraph adapter"]:::new
  D --> E["[EXISTING] EncoderCudaGraphManager<br/>capture / replay / clone"]:::existing
  E --> F["[EXISTING] Eager fallbacks<br/>CPU/FA2-pad/FP16/streaming"]:::existing
  classDef existing fill:#e5e7eb,stroke:#6b7280,color:#111827
  classDef changed fill:#fef3c7,stroke:#d97706,color:#451a03,stroke-width:2px
  classDef new fill:#dcfce7,stroke:#16a34a,color:#052e16,stroke-width:2px
  classDef removed fill:#fee2e2,stroke:#dc2626,color:#450a0a,stroke-width:2px
Loading

CI at 2d9b427bf491 (2026-09-30T20:58:03.064564+00:00): all required and operator-watched checks reported green. Observed Buildkite: buildkite/vllm-omni-npu-ci (failed), buildkite/vllm-omni-amd-ci (failed), buildkite/vllm-omni (passed), and 1 more.

No actionable findings.


🤖 This review was generated by InferMatrix Copilot, an open-source repo-maintenance agent for PR review, CI debugging and issue triage. Try it on your own repo, and ⭐ star it if it helped!

@amy-why-3459 amy-why-3459 added the high priority high priority issue, needs to be done asap label Oct 1, 2026

hsliuustc0106 commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

Please attach the pending performance + accuracy A/B for the upstream-manager implementation at 2d9b427bf491. The 4df8d0bb4937 report covers the previous implementation. For the same Daily-Omni workload and GPU, please include exact SHAs/configs, latency/throughput/peak VRAM over at least 3 measured runs with warmup excluded and mean ± stddev, plus answer agreement/accuracy and failed-request counts. Keep cold-capture cost separate from steady-state measurements.

@hsliuustc0106 hsliuustc0106 added the enhancement New feature or request label Oct 3, 2026
@amy-why-3459

Copy link
Copy Markdown
Collaborator Author

Daily-Omni concurrency-10 A/B for the upstream-manager implementation at 2d9b427bf49160c567bade1bcf9beb20cd550389 (base 4af28f33bd). This replaces the earlier 4df8d0bb4937 report.

Setup: one exclusively reserved NVIDIA L20X, MiniCPM-o 4.5, vLLM 0.30.0, PyTorch 2.13.0+cu130, transformers 5.14.1. Both servers use vllm_omni/deploy/minicpmo_4_5.yaml; the only change is --hf-overrides '{"encoder_cuda_graph": false/true}'. Enabled first, disabled second, on the same GPU.

Workload: Daily-Omni audio+video, minicpm-interleave, concurrency 10, seed 0, temperature 0, text output capped at 128 tokens. Requested 2000 prompts with --no-oversample; the dataset has 1197 entries, so each pass submitted 1197 requests.

Protocol: one full 1197-request warmup pass per mode is excluded. The table is three measured passes per mode, mean ± sample standard deviation (ddof=1). Percentiles are the mean of the per-run percentiles. Peak VRAM is NVML device-used memory sampled every 100 ms during the measured passes only (not PyTorch allocated memory; peaks shorter than 100 ms can be missed). Cold capture is timed separately around _capture and is not included below. No encoder capture occurred during the six measured passes.

Metric Encoder graph off Encoder graph on Change
Duration (s) 1149.84 ± 5.88 983.96 ± 17.69 −14.43%
Request throughput (req/s) 1.039 ± 0.004 1.215 ± 0.022 +16.91%
Mean E2E latency (ms) 9608.64 ± 39.87 8177.76 ± 146.17 −14.89%
Median E2E latency (ms) 9599.04 ± 62.48 8258.79 ± 74.05 −13.96%
P99 E2E latency (ms) 16011.69 ± 427.51 13111.80 ± 2478.29 −18.11%
Mean TTFT (ms) 7796.88 ± 50.95 8009.22 ± 148.23 +2.72%
Median TTFT (ms) 7573.30 ± 200.32 8095.02 ± 79.07 +6.89%
P99 TTFT (ms) 13409.80 ± 160.99 12787.43 ± 2686.36 −4.64%
Peak VRAM (GiB) 80.23 ± 0.00 87.81 ± 0.29 +9.45%
Accuracy, successful requests 78.237% ± 0.028% 78.243% ± 0.000% unchanged
Accuracy, including failed requests 78.084% ± 0.128% 78.112% ± 0.000% unchanged
Failed requests 2.33 ± 1.53 2.00 ± 0.00 —

Per-run failures (zero-based dataset indices). Every failure is a stage-0 Stage request failed immediately after a multimodal cache-miss warning:

Pass Failed Indices
off 1 1 210
off 2 2 210, 699
off 3 4 210, 699, 825, 833
on 1 2 210, 699
on 2 2 210, 699
on 3 2 210, 699

Answer agreement on requests that succeeded in both the off and on pass of the same index: 1195/1195, 1195/1195, and 1193/1193 identical. No answer string differed.

Cold capture, excluded from the table: 1 capture during on-server startup (0.166 s) and 7 captures during the on warmup pass (1.302 s total). The off server does not capture these graphs.

P99 E2E and P99 TTFT for the enabled runs have a large spread across the three passes, so those two reductions are less stable than the mean E2E and throughput results. The passes were sequential (on, then off) and were not repeated in the reverse order.

@amy-why-3459

Copy link
Copy Markdown
Collaborator Author

@hsliuustc0106 Following up on your A/B validation request: the updated performance and accuracy report now covers the upstream-manager implementation at 2d9b427bf49160c567bade1bcf9beb20cd550389, and the PR description has been updated to match.

The report includes the SHA/configuration, three measured passes per mode on the same exclusively reserved L20X and Daily-Omni workload, an excluded full warmup pass, mean ± sample standard deviation, latency/throughput/peak VRAM, accuracy, answer agreement and per-run failures. Cold capture is reported separately; no encoder captures occurred during the measured passes.

  • Throughput: +16.91%; mean E2E latency: −14.89%.
  • Tradeoffs: mean TTFT +2.72%, peak device-used VRAM +9.45%.
  • All answers for requests successful in both corresponding passes were identical (1195/1195, 1195/1195 and 1193/1193). Failed requests remain documented: off 1/2/4, on 2/2/2; their root cause has not been established.

The measurements were sequential, enabled then disabled, without a reverse-order repeat; tail latency varied across runs, and NVML memory sampling was at 100 ms intervals. Scoped CPU/CUDA tests are reported, but the full repository test suite has not been run.

Could you please take another look and let us know whether this addresses your validation request? If the implementation and evidence are satisfactory and the remaining CI/test requirements are met, please help approve and merge this PR, or identify any remaining blockers. Thank you!

@BeatSeat BeatSeat left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for this neat integration with vLLM's upstream EncoderCudaGraphManager! Left two comments regarding output tensor memory aliasing and test import formatting.

) -> None:
output = outputs["default"]
# Callers retain embeddings across subsequent replays.
dest[0] = output.clone() if clone else output

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[question / bug risk] In _ExactShapeEncoder.postprocess_encoder_output:

dest[0] = output.clone() if clone else output

The docstring states "callers get a clone so a later replay cannot overwrite retained embeddings".
However, in _capture, config.compilation_config never sets encoder_cudagraph_clone_outputs = True, and upstream EncoderCudaGraphManager.execute passes clone=self.clone_outputs which defaults to False.

When clone=False, dest[0] directly references the static CUDA graph output buffer. If an async pipeline or subsequent request triggers _encoder_graph for the same shape while a caller still retains or reads previous embeddings, the replay will overwrite the previous request's embeddings in-place in GPU memory.

Should this either unconditionally clone (dest[0] = output.clone()) or explicitly set config.compilation_config.encoder_cudagraph_clone_outputs = True in _capture to prevent memory aliasing?

@amy-why-3459 amy-why-3459 Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for checking this. I verified the upstream version used by this PR: in vLLM v0.30.0, _execute_local, the manager explicitly passes clone=True to postprocess_encoder_output. This version does not use self.clone_outputs or an encoder_cudagraph_clone_outputs configuration field. Consequently the adapter takes output.clone() on this path.

The existing retained-output tests exercise a subsequent replay with changed inputs/masks, and the graph regression module passes on CUDA. The AMD job also shows all five CUDA tests in this module passing at 2d9b427bf491, including retained outputs and cross-stream isolation. I am keeping the upstream clone contract rather than adding a configuration field that this version does not implement.

import pytest
import torch

from vllm_omni.model_executor.models.minicpmo_4_5.encoder_cuda_graph import EncoderCudaGraph

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[style / nit] ruff check reports an I001 import sorting error on this block:

I001 [*] Import block is un-sorted or un-formatted

Running ruff check --fix (removing the extra blank line before from vllm_omni...) resolves it cleanly.

@amy-why-3459 amy-why-3459 Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I checked this at 2d9b427bf491 using the repository-pinned Ruff 0.14.10 through pre-commit run ruff-check --files tests/model_executor/models/minicpmo_4_5/test_encoder_cuda_graph.py, from the repository root. It passes without modifying the file; the PR pre-commit check also passed. The blank line separates third-party imports from the local vllm_omni package in this checkout. Running Ruff outside the repository/config context can classify the local package differently, so I could not reproduce I001 under the project's configured check and have left that grouping unchanged.

@Sy0307 Sy0307 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Non-blocking feedback from validation at 2d9b427bf4: targeted tests passed (249 passed, 4 FA2 skips), and 110 sampled real-weight graph/eager encoder comparisons were bit-exact. HTTP E2E also completed on one H200 at C10 using a small Daily-Omni subset. Three follow-ups would help clarify the limits of this implementation before broader reuse:

  1. Cache admission (line 143). The first four repeated shapes retain the slots for the model's lifetime, and startup profiling consumes one vision slot. After filling the cache, a new hot shape called 100 times produced zero captures and zero hits. Could admission be configurable or support replacement with capture-cost controls, with miss accounting exposed so changing traffic does not silently lose coverage?

  2. Memory admission (line 174). Each entry retains a separate graph pool, so the entry cap alone does not define a retained-byte limit. The fixed-128-token comparison added approximately 4.73 GiB of device-used memory at the end of measured rounds; this was a snapshot measurement, not a peak or upper bound. Please consider reporting per-pool retention and adding a byte budget or available-memory check, particularly for the default-enabled path.

  3. Capture failure contract (line 153). A failed capture leaves the shape in _seen, so a subsequent call attempts capture again. Fault injection reproduced repeated propagated exceptions; no natural capture failure was observed in the model runs. Could the intended fatal/recoverable behavior be documented and covered by a focused test, with retry state handled explicitly if recovery is supported?

@hsliuustc0106

Copy link
Copy Markdown
Collaborator

@amy-why-3459 this PR is labeled ready + high priority, but CI is failing on the latest commit:

Could you please take a look and push an update to get CI green? Once the checks pass we can proceed with review/merge. Thanks!

Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
@amy-why-3459

amy-why-3459 commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator Author

Follow-up to the review of 2d9b427bf491, addressed in 3732befb2d69a86f6cb80e5e59c51a54c4260b69:

  • Cache admission: the observation is correct. --hf-overrides now exposes encoder_cuda_graph_max_graphs (default 4) and encoder_cuda_graph_min_capture_calls (default 2). The no-eviction policy and startup-profiling slot consumption are documented. Adapter statistics include capacity/warmup/ineligible/memory misses, including calls that never reach the upstream manager. This provides configurability and visibility without introducing recurring recapture costs.
  • Memory admission: the graph count is not a retained-byte bound. Added encoder_cuda_graph_min_free_bytes (default 1 GiB), checked through the Omni platform API before each new capture. Below the floor, execution stays eager; existing graphs still replay. This is a preflight check, not a pool-byte budget, allocation reservation, or guarantee against OOM. Per-pool byte accounting is not implemented; the documentation states this scope and the tuning tradeoff.
  • Capture failures: explicitly fatal for the adapter. The first exception propagates; admission state is cleared and subsequent calls fail without another capture or eager CUDA work. Added fault-injection coverage. This avoids assuming the CUDA context is safe after an arbitrary capture error.
  • Output aliasing / import sorting: checked and replied in the two inline threads. vLLM 0.30.0 passes clone=True; repository-pinned Ruff passes the original import block.
  • Performance request: correct, and the three measured runs per mode report already covers 2d9b427bf491. It does not validate this new revision's performance; the description now makes that distinction explicit.

Local validation on the new code: 258 passed, 4 skipped, 0 failed, 0 deselected, 0 collection errors across the seven modules in the PR test plan, using the same command without -m cpu, plus -ra. Four FA2 tests skipped because Transformers reports FA2 unavailable. Python 3.12.13, vLLM 0.30.0, PyTorch 2.13.0+cu130, transformers 5.14.1, pytest 9.1.1, one NVIDIA L20X. All applicable pre-commit checks passed.

I also investigated the CI failures reported on the old head:

  • NPU A5 layer job: the pod remained Pending for three hours and never started. Buildkite records Unschedulable, including insufficient NPU/CPU resources and unavailable nodes. This needs CI capacity/retry, not a change to encoder code.
  • AMD build: hard failures include Ulysses invalid device ordinal, AuK exact-equality assertions, and a diffusion shard timeout. The soft-failing engine/model job has Breeze failures; all five MiniCPM encoder graph CUDA tests passed there. These failures are outside the changed paths, but I have not established a same-environment base/head comparison, so I am not claiming they are pre-existing or waiving them.

Full-suite evidence is still outstanding. Before merge, please complete the repository unit-test matrix for 3732befb2d69a86f6cb80e5e59c51a54c4260b69: the four CPU groups in .buildkite/cuda/test-ready.yml and the applicable CUDA/AMD/NPU/Intel unit-test jobs/shards. Attach exact commands, dependency/hardware versions, passed/failed/skipped/deselected/collection-error counts, and CI/log links; explain skips and unavailable jobs. For suspected existing failures, include a same-environment base/head comparison, or obtain an explicit maintainer disposition. Scoped green tests do not establish merge readiness.

@vllm-omni-review-bot

vllm-omni-review-bot commented Oct 5, 2026 •

Copy link
Copy Markdown

Omni ReviewBot: superseded

The CI failure noted on 9de0fa6eb65e refers to an earlier head; the pull request now points at 4386db3586f7.

Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
@amy-why-3459 amy-why-3459 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Oct 5, 2026
@natureofnature
natureofnature enabled auto-merge (squash) October 5, 2026 08:04

@natureofnature natureofnature left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@natureofnature
natureofnature enabled auto-merge (squash) October 5, 2026 08:05
@natureofnature natureofnature added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Oct 5, 2026
@vllm-omni-review-bot

vllm-omni-review-bot commented Oct 5, 2026 •

Copy link
Copy Markdown

Omni ReviewBot: superseded

The CI failure noted on 781e316c5411 refers to an earlier head; the pull request now points at 880bacdf4ccd.

@amy-why-3459 amy-why-3459 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Oct 5, 2026
@natureofnature
natureofnature merged commit 9146284 into vllm-project:main Oct 5, 2026
6 of 9 checks passed
@amy-why-3459
amy-why-3459 deleted the perf-minicpmo-encoder-cuda-graph branch October 5, 2026 15:53
amy-why-3459 added a commit to amy-why-3459/vllm-omni that referenced this pull request Oct 5, 2026
Add an optional NPU encoder graph path for vision and stateless audio.
Keep the Ascend audio convolution stem eager and capture only the
transformer stack. Reuse the CUDA adapter merged in vllm-project#8332, preserving
its admission controls and shared-pool configuration.

Add opt-in exact-shape CUDA graphs for the streaming Code2Wav Conformer.
Treat CNN and attention caches as explicit inputs and outputs, clone
returned state, and key captures by stream, cache shape, last-chunk flag,
autocast state and positional-table storage. Admit repeated shapes into
bounded LRU slots with independently owned capture streams and pools.

Enable cached HiFT ISTFT and warm up waveform finalization after graph
capture to move first-use envelope checks and FFT plan creation out of
replay. Add regression coverage for replay correctness, retained outputs,
stream isolation, eviction and backend configuration forwarding.

Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request high priority high priority issue, needs to be done asap ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants