Skip to content

[Feature] Add bounded MiniCPM-o 4.5 vision encoder CUDA Graphs - #7659

Closed
0z5a wants to merge 11 commits into
vllm-project:mainfrom
0z5a:minicpmo45-encoder-cudagraph
Closed

0z5a wants to merge 11 commits into
vllm-project:mainfrom
0z5a:minicpmo45-encoder-cudagraph

Conversation

@0z5a

@0z5a 0z5a commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Adds opt-in MiniCPM-o 4.5 image/video encoder CUDA Graphs through the native vLLM capture-axis protocol. Parsed vision inputs, position IDs and masks are prepared outside replay; graph budgets remain bounded. Audio and stateful duplex retain their existing encoder behavior.

The full-checkpoint repository parity test now requires equal image/video text; output differences fail the test. Scheduler compatibility and XD-RoPE compatibility are outside this feature diff. The diff contains core implementation, regression/E2E tests, required CI wiring and relevant serving guidance.

First-wave A100 full-checkpoint E2E

Two A100 PCIe 40GB GPUs; native vLLM 0.30.0, Torch 2.13.0+cu130, Transformers 5.14.1; BF16, complete checkpoint 503e754207c94da6bb26850b4469f367c9ea3582. All three stages loaded. Stage 0 ran on GPU0; Talker/Code2Wav loaded on GPU1. These requests exercised image/video-to-text.

Each configured B ran four fresh A/P/P/A processes: eager encoder, Graph, Graph, eager encoder. Decoder configuration was identical (FULL_DECODE_ONLY); prefix caching and multimodal processor caching were off, with unique multimodal IDs. Each cell submitted 2×C requests, with at most C active. Rates below are geometric means of the two corresponding arms. Fixed 224×224 images and a two-frame video were reused across arms. Output used a 128-token cap; video descriptions reach that cap.

All 2,984 formal API requests completed and all twelve processes exited naturally. The two Graph arms recorded 1,492 hits and zero misses, with stable capture budgets. Peak reserved memory was 29.41 GiB. Native scheduled encoder/decoder batches are reported separately.

Configured B Decoder B Encoder B C Eager req/s Graph req/s Observed throughput change Exact outputs vs A0 Qualification
1 1 1 1 1.6898 1.7137 +1.42% 8/8 exact text/tokens
1 1 1 8 1.0532 1.0627 +0.90% 64/64 exact text/tokens
1 1 1 16 1.0463 1.0552 +0.86% 128/128 exact text/tokens
1 1 1 32 1.0250 1.0311 +0.59% 256/256 exact text/tokens
1 1 1 64 1.0240 1.0317 +0.76% 512/512 exact text/tokens
2 1 1 1 1.6796 1.7194 +2.37% 8/8 exact text/tokens
2 2 1 8 1.8920 1.9190 +1.43% 64/64 exact text/tokens
2 2 1 16 2.0008 2.0242 +1.17% 122/128 output differences; see controls below
2 2 1 32 1.9457 1.9709 +1.30% 256/256 exact text/tokens
2 2 1 64 1.9680 1.9923 +1.24% 512/512 exact text/tokens
4 1 1 1 1.6744 1.7247 +3.00% 8/8 exact text/tokens
4 4 3 8 3.0341 3.0976 +2.09% 54/64 output differences; see controls below
4 4 3 16 3.1844 3.4389 +7.99% 100/128 output differences; see controls below
4 4 3 32 3.4093 3.4796 +2.06% 256/256 exact text/tokens
4 4 3 64 3.5053 3.5698 +1.84% 512/512 exact text/tokens

The original B2/C16 and B4/C8,C16 comparisons contain output differences. B4 also produced 31 differences between its two eager runs. The affected speed cells remain unqualified; their original results are retained.

Separate output diagnostics

Native encoder execution was checked against an immediate eager call on the same trained model and actual vision inputs: B2 gave 38/38 bit-exact embedding comparisons; B4 gave 56/56, including actual encoder B3. Maximum absolute error was zero.

The affected cells were then rerun in four fresh A/P/P/A processes with VLLM_BATCH_INVARIANT=1 in every arm. All 376 control requests, including warmups, produced identical corresponding text and token IDs. The combination of exact encoder embeddings, eager-to-eager variation in the original runs and repeatable invariant controls supports batch-dependent native numerical variation as the explanation; it does not qualify the original mismatching speed cells.

Configured B C Invariant eager req/s Invariant Graph req/s Throughput change Exact measured outputs
2 16 0.7877 0.7936 +0.75% 128/128
4 8 1.2918 1.3075 +1.22% 64/64
4 16 1.4190 1.4373 +1.29% 128/128

Validation and limits

  • First-wave validation: 27 CPU and four native A100 CUDA core regressions passed; Ruff and diff checks passed. At that stage, production files in the strict-parity update were byte-identical to the full-model tested implementation 6b1d3bfa8549. The host-metadata follow-up is validated separately below.
  • The earlier L20/vLLM 0.29 development-build repeated-token failure remains a recorded historical failure. The current measurements use native vLLM 0.30 without private scheduler/XD-RoPE overlays.
  • A complete API response at the fixed token cap establishes the stated text parity scope. Speech generation, stateful duplex and other resolutions are outside these measurements.

Remains Draft. Raw logs, checkpoint files and private benchmark/diagnostic harnesses are excluded from the PR diff.

A100 host-metadata follow-up: complete EOS output and concurrency

Cache vision target sizes on the host once per input and prepare layout masks there before uploading the final mask. This removes repeated device-to-host metadata synchronization from encoder graph preparation. The follow-up changes only the encoder implementation and its CPU/CUDA-metadata regression coverage.

Full pinned MiniCPM-o 4.5 checkpoint on 2×A100 40GB, native three-stage initialization, image/video-to-text route, configured B1/B2/B4, encoder graphs enabled in both arms and the same decoder graph configuration. Each configuration/arm completed 240 measured requests at C8/16/32/64. B1/B2 used 6 warmup requests per arm; B4 used 8. All 1,440 measured requests reached EOS (finish_reason=stop, 512-token cap), and all 740 paired complete token/text outputs including warmups matched exactly. Audio generation was not exercised. All six processes exited naturally. Native decoder batch reached each configured limit; encoder batch maxima are listed separately.

Baseline 688a59ed20ab versus aa8c70ea577f. One unprofiled pair per B; independent CPU installers remained present. The measured changes below do not establish a meaningful E2E speedup; repeated starts remain pending. These are host-metadata comparisons with graphs enabled in both arms, separate from the first-wave eager-versus-graph results above.

Configured B Maximum decoder B Maximum encoder B C Baseline req/s Host-metadata req/s Throughput change Exact measured outputs
1 1 1 8 1.0193 1.0211 +0.18% 16/16
1 1 1 16 1.0142 1.0141 -0.02% 32/32
1 1 1 32 0.9912 0.9948 +0.37% 64/64
1 1 1 64 0.9903 0.9909 +0.06% 128/128
2 2 1 8 1.1305 1.1310 +0.05% 16/16
2 2 1 16 1.2559 1.2566 +0.06% 32/32
2 2 1 32 1.1904 1.1913 +0.07% 64/64
2 2 1 64 1.2188 1.2205 +0.14% 128/128
4 4 2 8 2.2873 2.2695 -0.78% 16/16
4 4 2 16 2.3020 2.3078 +0.25% 32/32
4 4 3 32 2.3098 2.3119 +0.09% 64/64
4 4 3 64 2.3604 2.3604 -0.00% 128/128

Separate C8 Nsight diagnostics also had 22/22 complete token/text pairs exact. Recorded cudaStreamSynchronize calls fell from 228 to 48 and device-to-host copies from 5,827 to 5,631. These diagnostic counts are not unprofiled E2E timing. Native regression results: 27 CPU + 8 CUDA passed, zero skips; CUDA coverage includes image/video, FP32/BF16, CPU/CUDA metadata, padded/multiple inputs and oversized eager fallback. Ruff, formatting and diff checks passed.

@hsliuustc0106

Copy link
Copy Markdown
Collaborator

This PR touches vllm_omni/model_executor/, tests/model_executor/, .buildkite/common/, .buildkite/cuda/, examples/online_serving/ (8 files). Based on CODEOWNERS coverage of the changed files, the most-related reviewers appear to be:

@yenuo26 @NickCao @tzhouam

Could one of you take a look when you get a chance? Thanks!

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/ar_runtime.md.

Module owners: @tzhouam @fake0fan @Gaohan123

Routing: @tzhouam via module of the changed files, CODEOWNERS; @fake0fan via module of the changed files; @Gaohan123 via module of the changed files

@0z5a, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 18, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Resolved as of acbceadf7d5a: the high-priority or low-quality signal noted on an earlier commit no longer applies.

@MrlixiangWE MrlixiangWE left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ran the CPU lanes on vLLM 0.29.0, at 2ae62ee and at the merge-base 275720b:

lane 2ae62ee 275720b
tests/core -m 'core_model and cpu' 205 passed, 2 failed 207 passed, 0 failed
tests/model_executor/models/minicpmo_4_5 -m 'core_model and cpu' 453 passed, 2 failed, 3 skipped 434 passed, 2 failed, 1 skipped
tests/worker -m 'core_model and cpu' 332 passed 332 passed

The two tests/model_executor failures are the same test_code2wav_batching ids on both sides. The two tests/core failures are new; they are the scheduler helper, commented inline.

The environment line on the top table needs correcting. It pins the tested tree to e3dddb5 / c88e76e0, but that tree reads self.uses_xdrope_dim at vllm_omni/worker/gpu_model_runner.py:693, inside _update_states, unconditionally for every newly admitted request, and never sets it; vLLM ab35354 has no such attribute. The restore arrives in 2ae62ee. So the 24-request greedy E2E, the three-stage streaming run and the 984 HTTP requests in that table did not run on the tree that line names. The "Remaining integration gates" section says those fixes lived as private overlays, which fits — it is the environment line that reads as a public-tree reproduction and is not one.

The uses_xdrope_dim guard duplicates #7635, which is mine. The two are different code at different anchors — a hasattr guard at the end of the constructor here, a getattr default right after super().__init__() there — and git merge-tree on the two heads is clean with both surviving, so whoever lands second deletes their own.

new_prompt_len_snapshot=new_prompt_len_snapshot,
)

def _accept_structured_output_tokens(self, request: Request, new_token_ids: list[int]) -> bool:

@MrlixiangWE MrlixiangWE Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This turns two existing CPU tests red. tests/core/sched/test_omni_ar_scheduler_logprobs.py calls OmniARScheduler.update_from_output unbound with a SimpleNamespace stub that binds eight mixin helpers onto itself (_MIXIN_UPDATE_HELPERS, :22-31) but not this one, so self._accept_structured_output_tokens(...) at omni_ar_scheduler.py:557 raises:

AttributeError: 'types.SimpleNamespace' object has no attribute '_accept_structured_output_tokens'

vLLM 0.29.0, pytest tests/core -m 'core_model and cpu': 207 passed at 275720b, 205 passed and 2 failed at this head — test_mid_step_stop_trims_logprob_rows_with_token_ids and test_invalid_logprobs_finish_only_the_affected_scheduler_request. That lane runs on the ready label as pytest -sv tests/ -m 'core_model and cpu' --ignore=tests/diffusion --ignore=tests/model_executor --ignore=tests/entrypoints --ignore=tests/engine.

Adding the name to that helper list fixes it. A module-level accept_structured_output_tokens(manager, request, token_ids) keeps both call sites on one helper and leaves the stub alone.

output is disabled for the request.
"""
manager = self.structured_output_manager
accept_tokens = getattr(manager, "accept_tokens", None)

@MrlixiangWE MrlixiangWE Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ten call sites build the scheduler as a bare MagicMock and then set structured_output_manager.should_advance.return_value = False (tests/core/sched/test_omni_ar_scheduler_stale_drain.py:92, test_omni_ar_scheduler_streaming.py:116,264,347, test_omni_generation_scheduler_update_session.py:133, test_omni_sched_ec_request_finish.py:126, test_omni_sched_prefill_stats_finalize.py:59, tests/distributed/omni_connectors/test_chunk_transfer_adapter.py:2394,2474,2982); five of them sit in stub factories shared by 2 to 5 tests each. With self a MagicMock, self._accept_structured_output_tokens(...) resolves to an auto-created child mock and the real helper never runs; and once it does, getattr(mock_manager, "accept_tokens", None) is never None, so the should_advance branch stays unreachable. Those lines are dead configuration, and the seven in tests/core still pass, which is why nothing reports it. Probing the class instead — hasattr(type(manager), "accept_tokens") — makes that branch reachable, and it only starts to matter once the helper stops being looked up on self.


_close_runner(omni_runner)

with OmniRunner(_MODEL, deploy_config=_EAGER_DEPLOY, trust_remote_code=True) as eager_runner:

@MrlixiangWE MrlixiangWE Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The two arms are not built the same way, so a difference between them is not isolated to the flag. The fixture builds the graph arm through iter_omni_runner, which applies stage_config_path_for_run_level (tests/helpers/runtime.py:1035) and MODEL_PREFIX (:1036); this line builds the eager arm directly with neither. At --run-level core_model, which is what the new Buildkite command uses, stage_config_path_for_run_level adds load_format: dummy to every stage (tests/helpers/stage_config.py:773-777 then 728-739), so the graph arm runs random weights while the eager arm would load the real checkpoint. On the pinned vLLM the test skips at :260 before it gets there; once the pin carries capture_axes the comparison runs with the two arms on different weights, can never be equal, and the strict xfail reports "expected failure" for a reason that has nothing to do with replay.

That also bears on the divergence in the description. As committed, this test cannot tell a replay-correctness gap from the weight asymmetry, so it does not support the xfail reason's "not an admission or capture failure" — which run level produced the recorded 'validator\n' * 32 decides that.

Both arms need the same stage_config_path_for_run_level(...) and get_model_prefix() treatment, and the parity test needs to run where that treatment loads real weights — advanced_model, which also means adding this file to the Omni · MiniCPM-o 4.5 Test command in test-merge.yml:118, since that command lists its files explicitly.

@pytest.mark.omni
@hardware_test(res={"cuda": "H100", "npu": "A3"}, num_cards=1)
@pytest.mark.parametrize("omni_runner", test_params, indirect=True)
@pytest.mark.xfail(

@MrlixiangWE MrlixiangWE Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

xfail absorbs every exception in the test, not only the text mismatch. The _stage0_uses_encoder_graph assertion at :263 — whose docstring at :159-165 says it stops the parity test from quietly comparing two eager engines — reports as the expected failure if it ever trips, and so do a failed second-engine load and an OOM inside _close_runner. Moving the guard and the engine setup out of the xfail-covered region, or marking only the final comparison, keeps them able to fail.

The component suite also sits at the opposite end of the axis from the failing case. _EncoderModel runs 6 patches padded to 256 in float32 (tests/model_executor/models/minicpmo_4_5/test_encoder_cudagraph.py:252, test_encoder_cudagraph_cuda.py:57), while the 224x224 fixture resolves to 1 slice of 1024 patches through the in-tree MiniCPMVImageProcessor at its defaults (max_slice_nums=9, scale_resolution=448, patch_size=14), which production reads from the checkpoint config — so _ceiling(1024, _PATCH_CAPS) is 1024 and the replay buffers carry no padding at all. Adding an unpadded BF16 shape to the component suite covers the configuration the E2E fails on.

graph_image = _generate(omni_runner, images=_small_image())
graph_video = _generate(omni_runner, videos=_small_video())

_close_runner(omni_runner)

@MrlixiangWE MrlixiangWE Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

omni_runner is module-scoped (tests/helpers/fixtures/runtime.py:119). Exiting it from inside a test leaves the other two tests in this module with a dead engine as soon as anything reorders them — --ff, an xdist split, or a test appended after this one. omni_runner_function exists for tests that need to own the runner.



def _ceiling(value: int, tiers: tuple[int, ...]) -> int:
# An uncaptured key is deliberately preserved for the manager's eager

@MrlixiangWE MrlixiangWE Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The manager does not fall back to eager on an uncaptured key. It takes the eager path only when the token budget does not fit; with a budget that fits and a layout key that was never captured, _run_budget_graph returns None and the next statement in _execute_local is assert graph_output is not None.

That is reachable through the ordinary image path once one slice exceeds _PATCH_CAPS[-1]. Running the in-tree MiniCPMVImageProcessor at its defaults (max_slice_nums=9, scale_resolution=448, patch_size=14), which production reads from the checkpoint config: a 5000x1 image resolves to 1 slice of 2143 patches and 64 output tokens, so it fits the 256-token budget and selects a layout that was never captured.

test_oversized_image_falls_back_without_failing does not cover this. Its 1024x1024 fixture is 7 slices at 448 output tokens, so its layout key is chosen and then never looked up: the budget misses first and the item goes eager for that reason, not the layout — the axis fallback this comment describes has no test. Either the tier set has to cover what the processor can emit, or this needs to raise with the layout in the message, so the failure names its cause instead of surfacing as an assertion inside the manager.

],
out_hidden_size=int(self.config.hidden_size),
max_frames_per_video=self.get_max_frames_per_video(),
capture_axes=(tuple(axes),),

@MrlixiangWE MrlixiangWE Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

capture_axes=(tuple(axes),) is unconditional, so when axes stays empty this returns ((),), and the manager rejects that at construction: raise ValueError("Encoder cudagraph capture axes must be non-empty."). axes is empty whenever self.vpm is None, which minicpmo_4_5_omni_llm.py:3960 does for any deploy without images. --limit-mm-per-prompt '{"image":0,"video":0,"audio":2}' plus cudagraph_mm_encoder: true therefore kills the engine during init, and the message names the capture axes, not the missing vision tower. The manager factory gates only on the flag, supports_mm_inputs and supports_encoder_cudagraph(model), and bind_minicpmo_encoder_cudagraph advertises the protocol off the class, so the instance having no vpm does not stop any of it.

Returning () instead is not enough: capture() iterates itertools.product(*self._capture_axes), which yields one empty tuple when there are no axes, and _capture_budget_graph then calls prepare_encoder_cudagraph_capture_inputs(axis_keys=()) with no modality guard, so kind, capacity, extent = axis_keys[0] raises IndexError instead. The manager has to not be built at all — an early return in bind_minicpmo_encoder_cudagraph when the thinker has no vision tower leaves the wrapper without the protocol names, so supports_encoder_cudagraph is False and the factory returns None.


_AXIS_KEY = getattr(encoder_cudagraph_defs, "ENCODER_CUDAGRAPH_AXIS_KEYS_KWARG", "encoder_cudagraph_axis_keys")
_LAYOUT_KEY = "minicpmo_encoder_layout"
_PATCH_CAPS = (256, 512, 1024, 2048)

@MrlixiangWE MrlixiangWE Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These tiers do not match what the processor emits. find_best_resize targets about 448^2 / 14^2 = 1024 patches for every slice, so the per-slice patch count clusters just over 1024 and the next tier is 2048. Running the in-tree MiniCPMVImageProcessor at its defaults (max_slice_nums=9, scale_resolution=448, patch_size=14), which production reads from the checkpoint config, at the default 256-token budget, i.e. slice caps 1/2/4:

image slices patches/slice tokens tier linear attention
224x224 1 1024 64 1 x 1024 0% 0%
448x448 1 1024 64 1 x 1024 0% 0%
512x512 3 1035 192 4 x 2048 +164% +422%
640x480 3 1036 192 4 x 2048 +164% +421%
800x600 5 1036 320 eager
1024x1024 7 1024 448 eager

Slices and the largest per-slice patch count come from the processor; tokens are slices x 64, tier is the captured capacity x extent the layout key selects, and the last two columns are arithmetic over those counts against the eager path, which pads to the batch's own max: B x L for the linear layers, B x L^2 for SiglipAttention.

Sweeping 1278 in-budget sizes on a 32-px grid from 64 to 1568 px, both ladders cost something. 620 sizes stay on the 1024 extent: the 511 whose slice count already equals a capacity tier pay +0% to +9% linear, and the 109 that are 3 slices rounded up to 4 pay +33% to +35%. The 658 that cross into 2048 pay +88% to +166% linear and +254% to +432% attention, 1 slice and 4 slices alike, with 3 slices worst because both roundings apply. Over 1024 sizes scanned from 32 to 2016 px the smallest per-slice count is 896, so the 256 and 512 entries here are never selected and 6 of the 12 captured graphs are capture time and resident buffers for nothing. Dropping those two is self-contained; a tier between 1024 and 2048 is what the distribution asks for, and changing the capacity ladder also moves test_slice_capture_tiers_follow_output_budget (:127) and test_protocol_buffers_match_encoder_entry_point's assert axes[0][1] == 4 (:256).

def encoder_eager_forward(self, mm_kwargs: dict[str, Any], path: str = "default") -> torch.Tensor:
return torch.cat(self.get_multimodal_embeddings(**mm_kwargs))

def postprocess_encoder_output(

@MrlixiangWE MrlixiangWE Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The default postprocess_encoder_output on SupportsEncoderCudaGraph already selects outputs["default"] and hands it to scatter_output_slices, whose body is this loop line for line (vllm/model_executor/models/utils.py); and this override ignores the batch_mm_kwargs it accepts. The body can go. The name has to stay in the bind list at :44, because supports_encoder_cudagraph is an isinstance check against a runtime-checkable Protocol and the wrapper only carries the names that list copies onto it.

opt-in is:

```bash
vllm serve openbmb/MiniCPM-o-4_5 --omni --trust-remote-code \

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is the only place a user learns how to turn the flag on, and it does not mention what this PR measured for it: the same 224x224 image on the same checkpoint answers 'validator\n' * 32 with the flag and correctly without it. The "development-only" paragraph lists shared metadata preparation, the dependency pin and serving/performance/memory acceptance, which reads as "not qualified yet", not as "returns wrong text today". One sentence next to the command, or hold the recipe until replay matches eager.

@0z5a

0z5a commented Sep 18, 2026

Copy link
Copy Markdown
Contributor Author

Ran the CPU lanes on vLLM 0.29.0, at 2ae62ee and at the merge-base 275720b:

lane 2ae62ee 275720b
tests/core -m 'core_model and cpu' 205 passed, 2 failed 207 passed, 0 failed
tests/model_executor/models/minicpmo_4_5 -m 'core_model and cpu' 453 passed, 2 failed, 3 skipped 434 passed, 2 failed, 1 skipped
tests/worker -m 'core_model and cpu' 332 passed 332 passed

The two tests/model_executor failures are the same test_code2wav_batching ids on both sides. The two tests/core failures are new; they are the scheduler helper, commented inline.

The environment line on the top table needs correcting. It pins the tested tree to e3dddb5 / c88e76e0, but that tree reads self.uses_xdrope_dim at vllm_omni/worker/gpu_model_runner.py:693, inside _update_states, unconditionally for every newly admitted request, and never sets it; vLLM ab35354 has no such attribute. The restore arrives in 2ae62ee. So the 24-request greedy E2E, the three-stage streaming run and the 984 HTTP requests in that table did not run on the tree that line names. The "Remaining integration gates" section says those fixes lived as private overlays, which fits — it is the environment line that reads as a public-tree reproduction and is not one.

The uses_xdrope_dim hunk overlaps #7635 line for line (gpu_model_runner.py:92-105 against 84-102); whichever lands second drops its copy.

Thanks for reviewing. I will fix them soon.

# So when the `patch_attention_mask` is full of 1s (i.e. attending to the whole sequence),
# avoiding passing the attention_mask, which is equivalent to attending to the full sequence
if not torch.any(~patch_attention_mask):
if position_ids is not None:

@MrlixiangWE MrlixiangWE Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These two keyword arguments are independent but one silently redefines the other: passing position_ids without encoder_attention_mask takes this branch, sets attention_mask to None and drops the patch_attention_mask the caller supplied, so every padded patch is attended to. Nothing asserts the pair. In production only encoder_cudagraph_forward passes position_ids, and it passes both, so this is latent.

Asserting both are present needs test_vision_capture_metadata_matches_eager to change with it: at :196 it passes encoder_attention_mask=None alongside position_ids whenever mask.all(), which is the grids1 case, so it would go red. Having it build the 4-D mask the way prepare_encoder_cudagraph_replay_buffers does — always — removes both the trap and the divergence between what the test pins and what production takes.

budget = min(4 * int(self.config.query_num), limit)
return budget, budget

def _encoder_data(self, mm_kwargs: dict[str, Any]):

@MrlixiangWE MrlixiangWE Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

_encoder_data runs the full production parse — _parse_and_validate_vision_input plus flatten_bn over the whole pixel list — and a single-batch replay reaches it four times, twice on the whole batch and twice on the selected one: _execute_local calls _get_item_specs(mm_kwargs), _select_items calls select_encoder_cudagraph_items(mm_kwargs, indices), then _run_budget_graph calls _get_item_specs on the selected kwargs and prepare_encoder_cudagraph_replay_buffers parses them once more. Three of those run inside gpu_sync_allowed(), which upstream permits because the per-item grid and patch counts are read off device tensors to size the buffers — and this model does exactly that, at int(group.sum()) in get_encoder_cudagraph_item_specs and int(torch.cat(sizes).prod(-1).max()) in select_encoder_cudagraph_items. Stashing the derivation on the dict select_encoder_cudagraph_items already returns removes the two on the selected batch; the two full-batch parses need a memo keyed on the incoming mm_kwargs, because the second of them is the call that builds that dict.

which also owns the request's token history and constraint-start
bookkeeping. Older releases gate on ``should_advance`` and accept
through the per-request ``grammar``. Dispatch on the manager surface so
both layouts work; the manager form returns ``True`` when structured

@MrlixiangWE MrlixiangWE Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This docstring is the only thing asserting that both layouts work, and neither branch has a test. The 0.29.0 pin has should_advance and no StructuredOutputManager.accept_tokens, so CI can only execute the legacy half; the L20 in the description runs ab35354, where should_advance is gone and only the manager half exists. Neither side runs both, so nothing checks the dispatch itself — including the True-when-disabled behaviour that omni_ar_scheduler.py:557 turns into RequestStatus.FINISHED_ERROR on a False. Two fake managers in a CPU test, one carrying accept_tokens on the class and one carrying only should_advance, pin all of it.

return "Describe what happens in this video in one short sentence."


def _oversized_vision_tokens() -> int:

@MrlixiangWE MrlixiangWE Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is folded from module constants only, so assert _oversized_vision_tokens() > _ENCODER_TOKEN_BUDGET at :219 passes whatever the processor does — if the slice config ever changed so the fixture fitted the budget, the test would keep reporting that it exercises the eager fallback while exercising nothing. It is also not the quantity the manager compares: output tokens are num_slices * query_num, and the processor turns this fixture into 7 slices, so 448, not 1332. Counting slices with MiniCPMVImageProcessor at the deploy's settings and multiplying by query_num keeps it in this process and makes the guard able to fail.

def get_max_frames_per_video(self) -> int:
return int(self.multimodal_config.media_io_kwargs.get("video", {}).get("num_frames", 128))

def get_encoder_cudagraph_budget_range(self, vllm_config: VllmConfig) -> tuple[int, int]:

@MrlixiangWE MrlixiangWE Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Returning the same value for both ends of the range fixes the manager's batch size at one. It derives max_batch_size = min(max_budget // min_budget, min(token_budgets)), which is min(1, 256) here, so with the shipped default the greedy packer can never put two images in one graph — the thing capture is supposed to buy. Both the README recipe and the E2E work around it by setting encoder_cudagraph_max_vision_items_per_batch by hand, without saying it is required for any batching at all.

Widening the range is not free: _generate_budgets(64, 256) returns three budgets, and with the current twelve axis combinations that is 36 captured graphs instead of 12. Either say in the README that the knob is required, or take the batch size off a range that cannot express it.

request.status = RequestStatus.FINISHED_ERROR
request.resumable = False
stopped = True
if new_token_ids and not self._accept_structured_output_tokens(request, new_token_ids):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This block, the matching one in omni_generation_scheduler.py, the new mixin helper and the uses_xdrope_dim guard in vllm_omni/worker/gpu_model_runner.py are vLLM-compatibility changes with nothing to do with the MiniCPM-o encoder. The description files them under "Compatibility fixes", and the runner guard arrives in a commit titled test: add MiniCPM-o 4.5 vision encoder CUDA Graph serving E2E. Split out, they are a [Core] bugfix that lands on the pin and stands on its own; bundled here they hold behind a feature that cannot run on the pin at all, and the scheduler half turns two green CPU tests red.

@0z5a

0z5a commented Sep 18, 2026

Copy link
Copy Markdown
Contributor Author

Split scheduler compatibility into Draft #7799; XD-RoPE remains in #7635, and both Core changes are removed from this feature PR. Addressed the constructor/platform, metadata caching, patch tiers and E2E ownership/xfail comments. SSH validation: 31 passed on vLLM main; 22 passed/9 skipped on 0.29.0. The description now identifies the historical overlays and unresolved real-checkpoint text divergence. No new serving-parity claim.

@hsliuustc0106 hsliuustc0106 added the high priority high priority issue, needs to be done asap label Sep 18, 2026

@MrlixiangWE MrlixiangWE left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The scheduler and XD-RoPE split is clean. I found three remaining issues on the current encoder-graph head: one supported request shape now fails when the flag is enabled, the default capture set contains unreachable graphs, and the graph E2E does not prove that replay occurred.

if modality == "video"
else mm_kwargs
)
if len(kwargs["pixel_values"]) == 0:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Precomputed image embeddings are still a supported MiniCPM vision input: _parse_and_validate_vision_input accepts image_embeds, and the eager path returns them directly. With the flag enabled, however, the runner still classifies this as the image modality and enters the encoder graph manager, whose first model hook reaches this unconditional pixel_values lookup and raises KeyError. Please bypass the graph manager for embedding inputs, or preserve them through selection and eager fallback, and add a flag-on image_embeds regression.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in a613806: precomputed embeddings use a separate batching key and bypass encoder capture/replay. Added a flag-on regression through the runner’s encoder execution path.

# Capture the largest intermediates first so smaller layouts can
# reuse the graph pool instead of growing fragmented segments.
axes.extend(
("vision", slices, patches)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The token-budget and slice-capacity dimensions both encode the same slice count. With the default query_num=64, the manager creates budgets 64/128/256 while this code creates slice capacities 1/2/4, so the Cartesian product captures 27 graphs after the three patch tiers. A request with s slices always has s*64 output tokens and selects ceil(s) as its slice capacity, so only (64,1), (128,2), and (256,4) can be looked up; 18 captured keys are structurally unreachable. Please derive slice capacity from the token budget or otherwise prune invalid budget/axis pairs, pin the captured key set in a test, and update the README's statement that the default captures one budget.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in a613806: slice capacity now comes from each token budget; the default capture set has exactly 9 keys. Added the default-key regression and corrected the README.

"""Check that Stage 0 received the requested encoder graph flag."""
stage0 = omni_runner.omni.engine.stage_configs[0]
compilation = stage0.engine_args.get("compilation_config")
return bool(compilation and compilation.get("cudagraph_mm_encoder", False))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This guard proves only that the flag reached Stage 0, not that an encoder graph manager was created, captured, or replayed. That distinction matters because the only real-checkpoint result is the already documented divergent output; without an execution-level assertion, the xfail cannot establish that it exercised replay. It is also misleading on NPU: _select_encoder_cudagraph_mixin binds the protocol only when current_omni_platform.is_cuda(), so these NPU cases run eager and the parity case can compare eager with eager. Please make this graph-specific coverage CUDA-only and assert a manager, capture, and hit signal in the graph arm before comparing outputs.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in a613806: CUDA-only E2E now checks the worker’s actual captured graphs and increasing image/video replay counters outside the final-text xfail. The video fixture now contains the configured 2 frames.

@0z5a

0z5a commented Sep 19, 2026

Copy link
Copy Markdown
Contributor Author

Validation for a613806 (before: acbcead). Full MiniCPM-o-4_5 weights; two L20s, with Stage 0 on one card and Stages 1/2 on the other; vLLM 0.29.1rc1.dev197+gab35354c2. Fresh A/P/P/A processes, identical media hashes, 64 output tokens per request; one warmup per case/process excluded (6 measured requests per case/variant).

Measurement Before After Speedup Time reduction
Stage 0 graph capture (logged) 6 s 3 s ≈2.0× 50%
Image 1456.95 ms 1458.64 ms 0.999× -0.12%
Two-frame video 1523.54 ms 1521.19 ms 1.002× +0.15%
Oversized fallback 1846.15 ms 1854.70 ms 0.995× -0.46%

Default captured graphs: 27 → 9. Capture logs report 6/6 s before and 3/3 s after (integer-second resolution); request medians changed by less than 0.5% on this shared host.

Validation: 26 CPU tests, 4 CUDA tests passed; complete advanced-model E2E: 2 passed, 1 xfailed. Actual capture and image/video replay assertions passed outside the existing final-text parity xfail. All 48 benchmark requests completed; oversized requests used fallback. A0/P0/P1 token sequences match; A1 baseline repeated with different video/oversized text, so exact repeatability is not claimed.

Both encoder arms use the same external structured-output compatibility helper and uses_xdrope_dim=0 prerequisite; neither is part of this PR. The host harness limits cleanup to task descendants and checks released memory against each card’s initial shared occupancy. All applicable pre-commit checks passed.

Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
Deliver the full-checkpoint end-to-end coverage the encoder-graph work was
missing, and fix the vLLM-main integration blockers that kept any serving
E2E from running against public code.

Compatibility (needed for real serving E2E on newer vLLM):

- OmniGPUModelRunner: XD-RoPE was folded into M-RoPE upstream and the runner
  attribute was removed. Every read site is a plain attribute access inside
  the inherited __init__, so admission failed with AttributeError before any
  model work. Restore the attribute with the parent\x27s value when present and
  0 otherwise.
- Schedulers: newer vLLM moved both the reasoning-end decision and token
  acceptance onto StructuredOutputManager.accept_tokens(request, tokens),
  which also owns request history and constraint bookkeeping. Dispatch on the
  manager surface, and keep the old should_advance + grammar path for the
  released pin. The generation scheduler keeps its pre-existing behaviour of
  not turning a rejection into a terminal error.

E2E (tests/e2e/offline_inference/test_minicpmo_4_5_encoder_cudagraph.py):

- Loads the real Stage 0/1/2 pipeline with compilation_config.cudagraph_mm_encoder
  and once without it, using one shared deploy profile.
- Serves in-budget image and video requests, and an image above the capture
  budget through the manager\x27s eager fallback.
- Compares greedy text between the graph engine and the flag-off engine. On an
  L20 with vLLM 0.29.1rc1.dev197 the graph engine answers the image prompt with
  a repeated token while the eager engine answers correctly. Capture itself
  succeeds (12 graphs for the 256-token budget). The test is therefore
  xfail(strict=True): it documents a real replay-correctness gap instead of
  silently passing.

CI routing: the new file joins the MiniCPM-o 4.5 offline job and its
source-file dependency group, with the job timeout raised to fit the second
engine.

Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
@0z5a
0z5a force-pushed the minicpmo45-encoder-cudagraph branch from a613806 to e4149f2 Compare September 24, 2026 06:13
0z5a and others added 2 commits September 24, 2026 15:10
Signed-off-by: 0z5a <dezhen.lu@uni-tuebingen.de>
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
@0z5a
0z5a force-pushed the minicpmo45-encoder-cudagraph branch from e4149f2 to f3f6e49 Compare September 24, 2026 07:10
@0z5a
0z5a marked this pull request as draft September 27, 2026 06:37
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
@0z5a
0z5a force-pushed the minicpmo45-encoder-cudagraph branch from 5ae7aa2 to 489f440 Compare September 27, 2026 13:55
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
0z5a and others added 5 commits October 4, 2026 01:24
Signed-off-by: 0z5a <192209249+0z5a@users.noreply.github.com>
Signed-off-by: 0z5a <192209249+0z5a@users.noreply.github.com>
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
Signed-off-by: 0z5a <192209249+0z5a@users.noreply.github.com>
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
@BeatSeat

BeatSeat commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

#8332 is a better implementation, PTAL

@0z5a 0z5a closed this Oct 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request high priority high priority issue, needs to be done asap

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants