Repository navigation
[Perf] Add CUDA graphs for MiniCPM-o 4.5 input encoders - #8332
natureofnature merged 8 commits into
Conversation
Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
|
This PR appears to belong to: docs/design/module/model_integration.md, docs/design/module/ar_runtime.md. Module owners: @tzhouam @fake0fan @Gaohan123 Routing: @tzhouam via module of the changed files, CODEOWNERS; @fake0fan via module of the changed files; @Gaohan123 via module of the changed files @amy-why-3459, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Omni ReviewBot triage noteAutomated triage of commit
These are automated triage suggestions only — the final decision belongs to the maintainers. |
|
Daily-Omni concurrency-10 A/B results for commit Setup: one exclusively reserved NVIDIA L20X, MiniCPM-o 4.5, vLLM 0.30.0, PyTorch 2.13.0+cu130. Both runs use Workload: Daily-Omni audio+video,
Correctness and validation:
Interpretation: This single sequential A/B pair measured a 19.03% throughput improvement and a 16.27% reduction in mean end-to-end latency. It includes cold graph capture costs and has not been repeated with reversed ordering to control for cache/order effects. Stateful streaming audio remains eager; these are end-to-end deployment results, not evidence of a stateless audio encoder speedup in the duplex path. The full repository test suite has not been run. Client command (same for both server configurations): vllm bench serve --omni --port 28973 --max-concurrency 10 \
--dataset-name daily-omni --num-prompts 2000 --no-oversample \
--trust-remote-code --temperature 0 --output-len 128 --seed 0 \
--daily-omni-input-mode all --daily-omni-pack-mode minicpm-interleave \
--daily-omni-video-dir /path/to/Daily-Omni/Videos \
--daily-omni-qa-json /path/to/Daily-Omni/qa.json \
--model /path/to/MiniCPM-o-4_5 \
--endpoint /v1/chat/completions --backend openai-chat-omni \
--percentile-metrics ttft,tpot,itl,e2el \
--extra_body '{"modalities":["text"],"chat_template_kwargs":{"enable_thinking":false}}' \
--save-result --save-detailed --result-dir /path/to/results \
--result-filename off.json # use on.json for the enabled runAI assistance: OpenAI Codex inspected the saved result JSON files and logs, compared answer strings, and drafted this report. |
Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
Omni ReviewBot routing recordAssigned Strict under experiment |
vllm-omni-review-bot
left a comment
There was a problem hiding this comment.
Omni ReviewBot review
PR description
This PR adds exact-shape CUDA Graph replay for MiniCPM-o 4.5 SigLIP vision (transformer stack only) and stateless Whisper audio encode/project/pool, default-on unless --enforce-eager or hf_overrides.encoder_cuda_graph=false. A model-local EncoderCudaGraph admits up to four stream-scoped shapes (capture on the second hit) and delegates buffers/replay to upstream EncoderCudaGraphManager via a packed-batch _ExactShapeEncoder adapter. Host-side packing/masks, padded FA2 vision, FP16 Whisper, intermediate audio layers, and streaming KV audio stay eager; outputs are cloned so retained embeddings survive later replays.
Change flow
flowchart TD
A["[EXISTING] Vision/audio encode callers<br/>vpm / get_audio_hidden_states"]:::existing --> B["[CHANGED] minicpmo_4_5_omni_llm.py<br/>gates + _encode_* extraction"]:::changed
B --> C["[NEW] EncoderCudaGraph<br/>stream/shape admission"]:::new
C --> D["[NEW] _ExactShapeEncoder<br/>SupportsEncoderCudaGraph adapter"]:::new
D --> E["[EXISTING] EncoderCudaGraphManager<br/>capture / replay / clone"]:::existing
E --> F["[EXISTING] Eager fallbacks<br/>CPU/FA2-pad/FP16/streaming"]:::existing
classDef existing fill:#e5e7eb,stroke:#6b7280,color:#111827
classDef changed fill:#fef3c7,stroke:#d97706,color:#451a03,stroke-width:2px
classDef new fill:#dcfce7,stroke:#16a34a,color:#052e16,stroke-width:2px
classDef removed fill:#fee2e2,stroke:#dc2626,color:#450a0a,stroke-width:2px
CI at
2d9b427bf491(2026-09-30T20:58:03.064564+00:00): all required and operator-watched checks reported green. Observed Buildkite:buildkite/vllm-omni-npu-ci(failed),buildkite/vllm-omni-amd-ci(failed),buildkite/vllm-omni(passed), and 1 more.
No actionable findings.
🤖 This review was generated by InferMatrix Copilot, an open-source repo-maintenance agent for PR review, CI debugging and issue triage. Try it on your own repo, and ⭐ star it if it helped!
|
Please attach the pending performance + accuracy A/B for the upstream-manager implementation at |
|
Daily-Omni concurrency-10 A/B for the upstream-manager implementation at Setup: one exclusively reserved NVIDIA L20X, MiniCPM-o 4.5, vLLM 0.30.0, PyTorch 2.13.0+cu130, transformers 5.14.1. Both servers use Workload: Daily-Omni audio+video, Protocol: one full 1197-request warmup pass per mode is excluded. The table is three measured passes per mode, mean ± sample standard deviation (ddof=1). Percentiles are the mean of the per-run percentiles. Peak VRAM is NVML device-used memory sampled every 100 ms during the measured passes only (not PyTorch allocated memory; peaks shorter than 100 ms can be missed). Cold capture is timed separately around
Per-run failures (zero-based dataset indices). Every failure is a stage-0
Answer agreement on requests that succeeded in both the off and on pass of the same index: 1195/1195, 1195/1195, and 1193/1193 identical. No answer string differed. Cold capture, excluded from the table: 1 capture during on-server startup (0.166 s) and 7 captures during the on warmup pass (1.302 s total). The off server does not capture these graphs. P99 E2E and P99 TTFT for the enabled runs have a large spread across the three passes, so those two reductions are less stable than the mean E2E and throughput results. The passes were sequential (on, then off) and were not repeated in the reverse order. |
|
@hsliuustc0106 Following up on your A/B validation request: the updated performance and accuracy report now covers the upstream-manager implementation at The report includes the SHA/configuration, three measured passes per mode on the same exclusively reserved L20X and Daily-Omni workload, an excluded full warmup pass, mean ± sample standard deviation, latency/throughput/peak VRAM, accuracy, answer agreement and per-run failures. Cold capture is reported separately; no encoder captures occurred during the measured passes.
The measurements were sequential, enabled then disabled, without a reverse-order repeat; tail latency varied across runs, and NVML memory sampling was at 100 ms intervals. Scoped CPU/CUDA tests are reported, but the full repository test suite has not been run. Could you please take another look and let us know whether this addresses your validation request? If the implementation and evidence are satisfactory and the remaining CI/test requirements are met, please help approve and merge this PR, or identify any remaining blockers. Thank you! |
BeatSeat
left a comment
There was a problem hiding this comment.
Thanks for this neat integration with vLLM's upstream EncoderCudaGraphManager! Left two comments regarding output tensor memory aliasing and test import formatting.
| ) -> None: | ||
| output = outputs["default"] | ||
| # Callers retain embeddings across subsequent replays. | ||
| dest[0] = output.clone() if clone else output |
There was a problem hiding this comment.
[question / bug risk] In _ExactShapeEncoder.postprocess_encoder_output:
dest[0] = output.clone() if clone else outputThe docstring states "callers get a clone so a later replay cannot overwrite retained embeddings".
However, in _capture, config.compilation_config never sets encoder_cudagraph_clone_outputs = True, and upstream EncoderCudaGraphManager.execute passes clone=self.clone_outputs which defaults to False.
When clone=False, dest[0] directly references the static CUDA graph output buffer. If an async pipeline or subsequent request triggers _encoder_graph for the same shape while a caller still retains or reads previous embeddings, the replay will overwrite the previous request's embeddings in-place in GPU memory.
Should this either unconditionally clone (dest[0] = output.clone()) or explicitly set config.compilation_config.encoder_cudagraph_clone_outputs = True in _capture to prevent memory aliasing?
There was a problem hiding this comment.
Thanks for checking this. I verified the upstream version used by this PR: in vLLM v0.30.0, _execute_local, the manager explicitly passes clone=True to postprocess_encoder_output. This version does not use self.clone_outputs or an encoder_cudagraph_clone_outputs configuration field. Consequently the adapter takes output.clone() on this path.
The existing retained-output tests exercise a subsequent replay with changed inputs/masks, and the graph regression module passes on CUDA. The AMD job also shows all five CUDA tests in this module passing at 2d9b427bf491, including retained outputs and cross-stream isolation. I am keeping the upstream clone contract rather than adding a configuration field that this version does not implement.
| import pytest | ||
| import torch | ||
|
|
||
| from vllm_omni.model_executor.models.minicpmo_4_5.encoder_cuda_graph import EncoderCudaGraph |
There was a problem hiding this comment.
[style / nit] ruff check reports an I001 import sorting error on this block:
I001 [*] Import block is un-sorted or un-formatted
Running ruff check --fix (removing the extra blank line before from vllm_omni...) resolves it cleanly.
There was a problem hiding this comment.
I checked this at 2d9b427bf491 using the repository-pinned Ruff 0.14.10 through pre-commit run ruff-check --files tests/model_executor/models/minicpmo_4_5/test_encoder_cuda_graph.py, from the repository root. It passes without modifying the file; the PR pre-commit check also passed. The blank line separates third-party imports from the local vllm_omni package in this checkout. Running Ruff outside the repository/config context can classify the local package differently, so I could not reproduce I001 under the project's configured check and have left that grouping unchanged.
Sy0307
left a comment
There was a problem hiding this comment.
Non-blocking feedback from validation at 2d9b427bf4: targeted tests passed (249 passed, 4 FA2 skips), and 110 sampled real-weight graph/eager encoder comparisons were bit-exact. HTTP E2E also completed on one H200 at C10 using a small Daily-Omni subset. Three follow-ups would help clarify the limits of this implementation before broader reuse:
-
Cache admission (line 143). The first four repeated shapes retain the slots for the model's lifetime, and startup profiling consumes one vision slot. After filling the cache, a new hot shape called 100 times produced zero captures and zero hits. Could admission be configurable or support replacement with capture-cost controls, with miss accounting exposed so changing traffic does not silently lose coverage?
-
Memory admission (line 174). Each entry retains a separate graph pool, so the entry cap alone does not define a retained-byte limit. The fixed-128-token comparison added approximately 4.73 GiB of device-used memory at the end of measured rounds; this was a snapshot measurement, not a peak or upper bound. Please consider reporting per-pool retention and adding a byte budget or available-memory check, particularly for the default-enabled path.
-
Capture failure contract (line 153). A failed capture leaves the shape in
_seen, so a subsequent call attempts capture again. Fault injection reproduced repeated propagated exceptions; no natural capture failure was observed in the model runs. Could the intended fatal/recoverable behavior be documented and covered by a focused test, with retry state handled explicitly if recovery is supported?
|
@amy-why-3459 this PR is labeled Could you please take a look and push an update to get CI green? Once the checks pass we can proceed with review/merge. Thanks! |
Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
|
Follow-up to the review of
Local validation on the new code: 258 passed, 4 skipped, 0 failed, 0 deselected, 0 collection errors across the seven modules in the PR test plan, using the same command without I also investigated the CI failures reported on the old head:
Full-suite evidence is still outstanding. Before merge, please complete the repository unit-test matrix for |
Omni ReviewBot: supersededThe CI failure noted on |
Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
Omni ReviewBot: supersededThe CI failure noted on |
Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
Add an optional NPU encoder graph path for vision and stateless audio. Keep the Ascend audio convolution stem eager and capture only the transformer stack. Reuse the CUDA adapter merged in vllm-project#8332, preserving its admission controls and shared-pool configuration. Add opt-in exact-shape CUDA graphs for the streaming Code2Wav Conformer. Treat CNN and attention caches as explicit inputs and outputs, clone returned state, and key captures by stream, cache shape, last-chunk flag, autocast state and positional-table storage. Admit repeated shapes into bounded LRU slots with independently owned capture streams and pools. Enable cached HiFT ISTFT and warm up waveform finalization after graph capture to move first-use envelope checks and FFT plan creation out of replay. Add regression coverage for replay correctness, retained outputs, stream isolation, eviction and backend configuration forwarding. Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
Purpose
MiniCPM-o 4.5 input encoders launch their transformer layers eagerly for every input. Add bounded, exact-shape CUDA Graph replay for the SigLIP transformer stack and stateless audio encoder/projection/pooling.
The model-local exact-shape adapter now delegates capture, static input/output buffers, replay and hit accounting to upstream
EncoderCudaGraphManagerthroughSupportsEncoderCudaGraph. One already-packed local batch is one indivisible manager item, preserving existing attention and padding behavior. By default, capture starts on the second occurrence of a shape, with at most four graphs per encoder. HF overrides configure the per-encoder graph cap, minimum call count and pre-capture free-memory floor (default 1 GiB). Admission misses are exposed in adapter statistics. Capture failures propagate and prevent reuse until worker restart. See admission and failure semantics. Upstream replay refreshes input and mask buffers; model postprocessing clones outputs so subsequent requests cannot overwrite retained embeddings. Packing, positional/mask construction, resampling and output slicing retain their existing behavior.Keep eager execution for CPU, training/grad/autocast, nested capture, padded FlashAttention vision, FP16 Whisper, intermediate audio layers, and stateful streaming audio.
--enforce-eagerdisables these graphs;--hf-overrides '{"encoder_cuda_graph": false}'provides an encoder-only A/B switch.This PR contains implementation, regression tests and configuration documentation. Local
benchmarks/minicpmo/files are excluded.Test Plan
vLLM version: 0.30.0; PyTorch 2.13.0+cu130.
vLLM-Omni base:
4af28f33bd.Test Result
Latest implementation:
3732befb2d69a86f6cb80e5e59c51a54c4260b69. The model-local upstream manager integration remains in place; runner startup capture and padded token-budget batching are not introduced. Manager configuration copies remain local.CUDA_VISIBLE_DEVICES=0,PYTHONPATH=.. The seven-module command was run without-m cpu, with-rato report skips.Performance and accuracy
The upstream-manager A/B report measures
2d9b427bf49160c567bade1bcf9beb20cd550389. Those measurements precede the admission/failure-policy update at3732befb2d69; A/B performance has not been remeasured for the latest head. It supersedes the earlier custom-implementation report; those earlier numbers do not describe this implementation.Setup: one exclusively reserved NVIDIA L20X; vLLM 0.30.0, PyTorch 2.13.0+cu130 and transformers 5.14.1; the same MiniCPM-o 4.5 checkpoint and default
vllm_omni/deploy/minicpmo_4_5.yaml. Onlyencoder_cuda_graphis toggled. Daily-Omni audio+video usesminicpm-interleave, concurrency 10, seed 0, temperature 0 and a 128-token text-output cap. Requesting 2000 prompts with--no-oversampleproduces 1197 requests per pass, matching the dataset size.Protocol: exclude one full warmup pass per mode, then measure three passes per mode. Values below are mean ± sample standard deviation; percentile rows average the per-run percentiles. Enabled mode was measured first, disabled mode second. No encoder captures occurred during the six measured passes.
For paired passes, all requests successful in both modes had identical answer strings: 1195/1195, 1195/1195 and 1193/1193. This is not a zero-failure validation: disabled runs had 1/2/4 failures and enabled runs had 2/2/2. The report records stage-0 failures following multimodal cache-miss warnings; it does not establish their root cause.
Cold capture is excluded above: one startup capture took 0.166 s, and seven warmup captures took 1.302 s in total. VRAM was sampled using NVML every 100 ms and reflects device-used memory, not PyTorch allocated memory; shorter peaks may be missed. Enabled-run tail latency varied substantially. The sequential on/off experiment was not repeated in reverse order, so order effects remain a limitation.
These results show improved throughput and mean E2E latency for this workload, at the cost of higher mean TTFT and memory use. Stateful duplex audio remains eager; these end-to-end measurements do not establish a streaming-audio encoder speedup.
Relationship to #8430
This PR adds graph management around existing SigLIP and stateless-audio computations. It does not add streaming-audio KV batching, incremental Fbank, packed/fused vision kernels or candidate-space sampling.
#8430 addresses those broader Stage-0 optimizations and includes custom vision and streaming-audio graph managers. Its packed vision path calls
forward_packed(), whereas this PR attaches the vision adapter toforward(). Combining the two requires explicitly integrating the graph adapter with the packed entry point; enabling both implementations does not automatically combine their benefits. The packed/fused computation could remain model-owned while upstream graph management is reused. Stateful audio additionally needs model-owned KV preparation and commit semantics.