Repository navigation
Conversation
|
This PR touches vllm_omni/model_executor/, tests/model_executor/, .buildkite/common/, .buildkite/cuda/, examples/online_serving/ (8 files). Based on CODEOWNERS coverage of the changed files, the most-related reviewers appear to be: Could one of you take a look when you get a chance? Thanks! |
|
This PR appears to belong to: docs/design/module/ar_runtime.md. Module owners: @tzhouam @fake0fan @Gaohan123 Routing: @tzhouam via module of the changed files, CODEOWNERS; @fake0fan via module of the changed files; @Gaohan123 via module of the changed files @0z5a, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Omni ReviewBot triage noteResolved as of |
There was a problem hiding this comment.
Ran the CPU lanes on vLLM 0.29.0, at 2ae62ee and at the merge-base 275720b:
| lane | 2ae62ee |
275720b |
|---|---|---|
tests/core -m 'core_model and cpu' |
205 passed, 2 failed | 207 passed, 0 failed |
tests/model_executor/models/minicpmo_4_5 -m 'core_model and cpu' |
453 passed, 2 failed, 3 skipped | 434 passed, 2 failed, 1 skipped |
tests/worker -m 'core_model and cpu' |
332 passed | 332 passed |
The two tests/model_executor failures are the same test_code2wav_batching ids on both sides. The two tests/core failures are new; they are the scheduler helper, commented inline.
The environment line on the top table needs correcting. It pins the tested tree to e3dddb5 / c88e76e0, but that tree reads self.uses_xdrope_dim at vllm_omni/worker/gpu_model_runner.py:693, inside _update_states, unconditionally for every newly admitted request, and never sets it; vLLM ab35354 has no such attribute. The restore arrives in 2ae62ee. So the 24-request greedy E2E, the three-stage streaming run and the 984 HTTP requests in that table did not run on the tree that line names. The "Remaining integration gates" section says those fixes lived as private overlays, which fits — it is the environment line that reads as a public-tree reproduction and is not one.
The uses_xdrope_dim guard duplicates #7635, which is mine. The two are different code at different anchors — a hasattr guard at the end of the constructor here, a getattr default right after super().__init__() there — and git merge-tree on the two heads is clean with both surviving, so whoever lands second deletes their own.
| new_prompt_len_snapshot=new_prompt_len_snapshot, | ||
| ) | ||
|
|
||
| def _accept_structured_output_tokens(self, request: Request, new_token_ids: list[int]) -> bool: |
There was a problem hiding this comment.
This turns two existing CPU tests red. tests/core/sched/test_omni_ar_scheduler_logprobs.py calls OmniARScheduler.update_from_output unbound with a SimpleNamespace stub that binds eight mixin helpers onto itself (_MIXIN_UPDATE_HELPERS, :22-31) but not this one, so self._accept_structured_output_tokens(...) at omni_ar_scheduler.py:557 raises:
AttributeError: 'types.SimpleNamespace' object has no attribute '_accept_structured_output_tokens'
vLLM 0.29.0, pytest tests/core -m 'core_model and cpu': 207 passed at 275720b, 205 passed and 2 failed at this head — test_mid_step_stop_trims_logprob_rows_with_token_ids and test_invalid_logprobs_finish_only_the_affected_scheduler_request. That lane runs on the ready label as pytest -sv tests/ -m 'core_model and cpu' --ignore=tests/diffusion --ignore=tests/model_executor --ignore=tests/entrypoints --ignore=tests/engine.
Adding the name to that helper list fixes it. A module-level accept_structured_output_tokens(manager, request, token_ids) keeps both call sites on one helper and leaves the stub alone.
| output is disabled for the request. | ||
| """ | ||
| manager = self.structured_output_manager | ||
| accept_tokens = getattr(manager, "accept_tokens", None) |
There was a problem hiding this comment.
Ten call sites build the scheduler as a bare MagicMock and then set structured_output_manager.should_advance.return_value = False (tests/core/sched/test_omni_ar_scheduler_stale_drain.py:92, test_omni_ar_scheduler_streaming.py:116,264,347, test_omni_generation_scheduler_update_session.py:133, test_omni_sched_ec_request_finish.py:126, test_omni_sched_prefill_stats_finalize.py:59, tests/distributed/omni_connectors/test_chunk_transfer_adapter.py:2394,2474,2982); five of them sit in stub factories shared by 2 to 5 tests each. With self a MagicMock, self._accept_structured_output_tokens(...) resolves to an auto-created child mock and the real helper never runs; and once it does, getattr(mock_manager, "accept_tokens", None) is never None, so the should_advance branch stays unreachable. Those lines are dead configuration, and the seven in tests/core still pass, which is why nothing reports it. Probing the class instead — hasattr(type(manager), "accept_tokens") — makes that branch reachable, and it only starts to matter once the helper stops being looked up on self.
|
|
||
| _close_runner(omni_runner) | ||
|
|
||
| with OmniRunner(_MODEL, deploy_config=_EAGER_DEPLOY, trust_remote_code=True) as eager_runner: |
There was a problem hiding this comment.
The two arms are not built the same way, so a difference between them is not isolated to the flag. The fixture builds the graph arm through iter_omni_runner, which applies stage_config_path_for_run_level (tests/helpers/runtime.py:1035) and MODEL_PREFIX (:1036); this line builds the eager arm directly with neither. At --run-level core_model, which is what the new Buildkite command uses, stage_config_path_for_run_level adds load_format: dummy to every stage (tests/helpers/stage_config.py:773-777 then 728-739), so the graph arm runs random weights while the eager arm would load the real checkpoint. On the pinned vLLM the test skips at :260 before it gets there; once the pin carries capture_axes the comparison runs with the two arms on different weights, can never be equal, and the strict xfail reports "expected failure" for a reason that has nothing to do with replay.
That also bears on the divergence in the description. As committed, this test cannot tell a replay-correctness gap from the weight asymmetry, so it does not support the xfail reason's "not an admission or capture failure" — which run level produced the recorded 'validator\n' * 32 decides that.
Both arms need the same stage_config_path_for_run_level(...) and get_model_prefix() treatment, and the parity test needs to run where that treatment loads real weights — advanced_model, which also means adding this file to the Omni · MiniCPM-o 4.5 Test command in test-merge.yml:118, since that command lists its files explicitly.
| @pytest.mark.omni | ||
| @hardware_test(res={"cuda": "H100", "npu": "A3"}, num_cards=1) | ||
| @pytest.mark.parametrize("omni_runner", test_params, indirect=True) | ||
| @pytest.mark.xfail( |
There was a problem hiding this comment.
xfail absorbs every exception in the test, not only the text mismatch. The _stage0_uses_encoder_graph assertion at :263 — whose docstring at :159-165 says it stops the parity test from quietly comparing two eager engines — reports as the expected failure if it ever trips, and so do a failed second-engine load and an OOM inside _close_runner. Moving the guard and the engine setup out of the xfail-covered region, or marking only the final comparison, keeps them able to fail.
The component suite also sits at the opposite end of the axis from the failing case. _EncoderModel runs 6 patches padded to 256 in float32 (tests/model_executor/models/minicpmo_4_5/test_encoder_cudagraph.py:252, test_encoder_cudagraph_cuda.py:57), while the 224x224 fixture resolves to 1 slice of 1024 patches through the in-tree MiniCPMVImageProcessor at its defaults (max_slice_nums=9, scale_resolution=448, patch_size=14), which production reads from the checkpoint config — so _ceiling(1024, _PATCH_CAPS) is 1024 and the replay buffers carry no padding at all. Adding an unpadded BF16 shape to the component suite covers the configuration the E2E fails on.
| graph_image = _generate(omni_runner, images=_small_image()) | ||
| graph_video = _generate(omni_runner, videos=_small_video()) | ||
|
|
||
| _close_runner(omni_runner) |
There was a problem hiding this comment.
omni_runner is module-scoped (tests/helpers/fixtures/runtime.py:119). Exiting it from inside a test leaves the other two tests in this module with a dead engine as soon as anything reorders them — --ff, an xdist split, or a test appended after this one. omni_runner_function exists for tests that need to own the runner.
|
|
||
|
|
||
| def _ceiling(value: int, tiers: tuple[int, ...]) -> int: | ||
| # An uncaptured key is deliberately preserved for the manager's eager |
There was a problem hiding this comment.
The manager does not fall back to eager on an uncaptured key. It takes the eager path only when the token budget does not fit; with a budget that fits and a layout key that was never captured, _run_budget_graph returns None and the next statement in _execute_local is assert graph_output is not None.
That is reachable through the ordinary image path once one slice exceeds _PATCH_CAPS[-1]. Running the in-tree MiniCPMVImageProcessor at its defaults (max_slice_nums=9, scale_resolution=448, patch_size=14), which production reads from the checkpoint config: a 5000x1 image resolves to 1 slice of 2143 patches and 64 output tokens, so it fits the 256-token budget and selects a layout that was never captured.
test_oversized_image_falls_back_without_failing does not cover this. Its 1024x1024 fixture is 7 slices at 448 output tokens, so its layout key is chosen and then never looked up: the budget misses first and the item goes eager for that reason, not the layout — the axis fallback this comment describes has no test. Either the tier set has to cover what the processor can emit, or this needs to raise with the layout in the message, so the failure names its cause instead of surfacing as an assertion inside the manager.
| ], | ||
| out_hidden_size=int(self.config.hidden_size), | ||
| max_frames_per_video=self.get_max_frames_per_video(), | ||
| capture_axes=(tuple(axes),), |
There was a problem hiding this comment.
capture_axes=(tuple(axes),) is unconditional, so when axes stays empty this returns ((),), and the manager rejects that at construction: raise ValueError("Encoder cudagraph capture axes must be non-empty."). axes is empty whenever self.vpm is None, which minicpmo_4_5_omni_llm.py:3960 does for any deploy without images. --limit-mm-per-prompt '{"image":0,"video":0,"audio":2}' plus cudagraph_mm_encoder: true therefore kills the engine during init, and the message names the capture axes, not the missing vision tower. The manager factory gates only on the flag, supports_mm_inputs and supports_encoder_cudagraph(model), and bind_minicpmo_encoder_cudagraph advertises the protocol off the class, so the instance having no vpm does not stop any of it.
Returning () instead is not enough: capture() iterates itertools.product(*self._capture_axes), which yields one empty tuple when there are no axes, and _capture_budget_graph then calls prepare_encoder_cudagraph_capture_inputs(axis_keys=()) with no modality guard, so kind, capacity, extent = axis_keys[0] raises IndexError instead. The manager has to not be built at all — an early return in bind_minicpmo_encoder_cudagraph when the thinker has no vision tower leaves the wrapper without the protocol names, so supports_encoder_cudagraph is False and the factory returns None.
|
|
||
| _AXIS_KEY = getattr(encoder_cudagraph_defs, "ENCODER_CUDAGRAPH_AXIS_KEYS_KWARG", "encoder_cudagraph_axis_keys") | ||
| _LAYOUT_KEY = "minicpmo_encoder_layout" | ||
| _PATCH_CAPS = (256, 512, 1024, 2048) |
There was a problem hiding this comment.
These tiers do not match what the processor emits. find_best_resize targets about 448^2 / 14^2 = 1024 patches for every slice, so the per-slice patch count clusters just over 1024 and the next tier is 2048. Running the in-tree MiniCPMVImageProcessor at its defaults (max_slice_nums=9, scale_resolution=448, patch_size=14), which production reads from the checkpoint config, at the default 256-token budget, i.e. slice caps 1/2/4:
| image | slices | patches/slice | tokens | tier | linear | attention |
|---|---|---|---|---|---|---|
| 224x224 | 1 | 1024 | 64 | 1 x 1024 | 0% | 0% |
| 448x448 | 1 | 1024 | 64 | 1 x 1024 | 0% | 0% |
| 512x512 | 3 | 1035 | 192 | 4 x 2048 | +164% | +422% |
| 640x480 | 3 | 1036 | 192 | 4 x 2048 | +164% | +421% |
| 800x600 | 5 | 1036 | 320 | eager | ||
| 1024x1024 | 7 | 1024 | 448 | eager |
Slices and the largest per-slice patch count come from the processor; tokens are slices x 64, tier is the captured capacity x extent the layout key selects, and the last two columns are arithmetic over those counts against the eager path, which pads to the batch's own max: B x L for the linear layers, B x L^2 for SiglipAttention.
Sweeping 1278 in-budget sizes on a 32-px grid from 64 to 1568 px, both ladders cost something. 620 sizes stay on the 1024 extent: the 511 whose slice count already equals a capacity tier pay +0% to +9% linear, and the 109 that are 3 slices rounded up to 4 pay +33% to +35%. The 658 that cross into 2048 pay +88% to +166% linear and +254% to +432% attention, 1 slice and 4 slices alike, with 3 slices worst because both roundings apply. Over 1024 sizes scanned from 32 to 2016 px the smallest per-slice count is 896, so the 256 and 512 entries here are never selected and 6 of the 12 captured graphs are capture time and resident buffers for nothing. Dropping those two is self-contained; a tier between 1024 and 2048 is what the distribution asks for, and changing the capacity ladder also moves test_slice_capture_tiers_follow_output_budget (:127) and test_protocol_buffers_match_encoder_entry_point's assert axes[0][1] == 4 (:256).
| def encoder_eager_forward(self, mm_kwargs: dict[str, Any], path: str = "default") -> torch.Tensor: | ||
| return torch.cat(self.get_multimodal_embeddings(**mm_kwargs)) | ||
|
|
||
| def postprocess_encoder_output( |
There was a problem hiding this comment.
The default postprocess_encoder_output on SupportsEncoderCudaGraph already selects outputs["default"] and hands it to scatter_output_slices, whose body is this loop line for line (vllm/model_executor/models/utils.py); and this override ignores the batch_mm_kwargs it accepts. The body can go. The name has to stay in the bind list at :44, because supports_encoder_cudagraph is an isinstance check against a runtime-checkable Protocol and the wrapper only carries the names that list copies onto it.
| opt-in is: | ||
|
|
||
| ```bash | ||
| vllm serve openbmb/MiniCPM-o-4_5 --omni --trust-remote-code \ |
There was a problem hiding this comment.
This is the only place a user learns how to turn the flag on, and it does not mention what this PR measured for it: the same 224x224 image on the same checkpoint answers 'validator\n' * 32 with the flag and correctly without it. The "development-only" paragraph lists shared metadata preparation, the dependency pin and serving/performance/memory acceptance, which reads as "not qualified yet", not as "returns wrong text today". One sentence next to the command, or hold the recipe until replay matches eager.
Thanks for reviewing. I will fix them soon. |
| # So when the `patch_attention_mask` is full of 1s (i.e. attending to the whole sequence), | ||
| # avoiding passing the attention_mask, which is equivalent to attending to the full sequence | ||
| if not torch.any(~patch_attention_mask): | ||
| if position_ids is not None: |
There was a problem hiding this comment.
These two keyword arguments are independent but one silently redefines the other: passing position_ids without encoder_attention_mask takes this branch, sets attention_mask to None and drops the patch_attention_mask the caller supplied, so every padded patch is attended to. Nothing asserts the pair. In production only encoder_cudagraph_forward passes position_ids, and it passes both, so this is latent.
Asserting both are present needs test_vision_capture_metadata_matches_eager to change with it: at :196 it passes encoder_attention_mask=None alongside position_ids whenever mask.all(), which is the grids1 case, so it would go red. Having it build the 4-D mask the way prepare_encoder_cudagraph_replay_buffers does — always — removes both the trap and the divergence between what the test pins and what production takes.
| budget = min(4 * int(self.config.query_num), limit) | ||
| return budget, budget | ||
|
|
||
| def _encoder_data(self, mm_kwargs: dict[str, Any]): |
There was a problem hiding this comment.
_encoder_data runs the full production parse — _parse_and_validate_vision_input plus flatten_bn over the whole pixel list — and a single-batch replay reaches it four times, twice on the whole batch and twice on the selected one: _execute_local calls _get_item_specs(mm_kwargs), _select_items calls select_encoder_cudagraph_items(mm_kwargs, indices), then _run_budget_graph calls _get_item_specs on the selected kwargs and prepare_encoder_cudagraph_replay_buffers parses them once more. Three of those run inside gpu_sync_allowed(), which upstream permits because the per-item grid and patch counts are read off device tensors to size the buffers — and this model does exactly that, at int(group.sum()) in get_encoder_cudagraph_item_specs and int(torch.cat(sizes).prod(-1).max()) in select_encoder_cudagraph_items. Stashing the derivation on the dict select_encoder_cudagraph_items already returns removes the two on the selected batch; the two full-batch parses need a memo keyed on the incoming mm_kwargs, because the second of them is the call that builds that dict.
| which also owns the request's token history and constraint-start | ||
| bookkeeping. Older releases gate on ``should_advance`` and accept | ||
| through the per-request ``grammar``. Dispatch on the manager surface so | ||
| both layouts work; the manager form returns ``True`` when structured |
There was a problem hiding this comment.
This docstring is the only thing asserting that both layouts work, and neither branch has a test. The 0.29.0 pin has should_advance and no StructuredOutputManager.accept_tokens, so CI can only execute the legacy half; the L20 in the description runs ab35354, where should_advance is gone and only the manager half exists. Neither side runs both, so nothing checks the dispatch itself — including the True-when-disabled behaviour that omni_ar_scheduler.py:557 turns into RequestStatus.FINISHED_ERROR on a False. Two fake managers in a CPU test, one carrying accept_tokens on the class and one carrying only should_advance, pin all of it.
| return "Describe what happens in this video in one short sentence." | ||
|
|
||
|
|
||
| def _oversized_vision_tokens() -> int: |
There was a problem hiding this comment.
This is folded from module constants only, so assert _oversized_vision_tokens() > _ENCODER_TOKEN_BUDGET at :219 passes whatever the processor does — if the slice config ever changed so the fixture fitted the budget, the test would keep reporting that it exercises the eager fallback while exercising nothing. It is also not the quantity the manager compares: output tokens are num_slices * query_num, and the processor turns this fixture into 7 slices, so 448, not 1332. Counting slices with MiniCPMVImageProcessor at the deploy's settings and multiplying by query_num keeps it in this process and makes the guard able to fail.
| def get_max_frames_per_video(self) -> int: | ||
| return int(self.multimodal_config.media_io_kwargs.get("video", {}).get("num_frames", 128)) | ||
|
|
||
| def get_encoder_cudagraph_budget_range(self, vllm_config: VllmConfig) -> tuple[int, int]: |
There was a problem hiding this comment.
Returning the same value for both ends of the range fixes the manager's batch size at one. It derives max_batch_size = min(max_budget // min_budget, min(token_budgets)), which is min(1, 256) here, so with the shipped default the greedy packer can never put two images in one graph — the thing capture is supposed to buy. Both the README recipe and the E2E work around it by setting encoder_cudagraph_max_vision_items_per_batch by hand, without saying it is required for any batching at all.
Widening the range is not free: _generate_budgets(64, 256) returns three budgets, and with the current twelve axis combinations that is 36 captured graphs instead of 12. Either say in the README that the knob is required, or take the batch size off a range that cannot express it.
| request.status = RequestStatus.FINISHED_ERROR | ||
| request.resumable = False | ||
| stopped = True | ||
| if new_token_ids and not self._accept_structured_output_tokens(request, new_token_ids): |
There was a problem hiding this comment.
This block, the matching one in omni_generation_scheduler.py, the new mixin helper and the uses_xdrope_dim guard in vllm_omni/worker/gpu_model_runner.py are vLLM-compatibility changes with nothing to do with the MiniCPM-o encoder. The description files them under "Compatibility fixes", and the runner guard arrives in a commit titled test: add MiniCPM-o 4.5 vision encoder CUDA Graph serving E2E. Split out, they are a [Core] bugfix that lands on the pin and stands on its own; bundled here they hold behind a feature that cannot run on the pin at all, and the scheduler half turns two green CPU tests red.
|
Split scheduler compatibility into Draft #7799; XD-RoPE remains in #7635, and both Core changes are removed from this feature PR. Addressed the constructor/platform, metadata caching, patch tiers and E2E ownership/xfail comments. SSH validation: 31 passed on vLLM main; 22 passed/9 skipped on 0.29.0. The description now identifies the historical overlays and unresolved real-checkpoint text divergence. No new serving-parity claim. |
MrlixiangWE
left a comment
There was a problem hiding this comment.
The scheduler and XD-RoPE split is clean. I found three remaining issues on the current encoder-graph head: one supported request shape now fails when the flag is enabled, the default capture set contains unreachable graphs, and the graph E2E does not prove that replay occurred.
| if modality == "video" | ||
| else mm_kwargs | ||
| ) | ||
| if len(kwargs["pixel_values"]) == 0: |
There was a problem hiding this comment.
Precomputed image embeddings are still a supported MiniCPM vision input: _parse_and_validate_vision_input accepts image_embeds, and the eager path returns them directly. With the flag enabled, however, the runner still classifies this as the image modality and enters the encoder graph manager, whose first model hook reaches this unconditional pixel_values lookup and raises KeyError. Please bypass the graph manager for embedding inputs, or preserve them through selection and eager fallback, and add a flag-on image_embeds regression.
There was a problem hiding this comment.
Fixed in a613806: precomputed embeddings use a separate batching key and bypass encoder capture/replay. Added a flag-on regression through the runner’s encoder execution path.
| # Capture the largest intermediates first so smaller layouts can | ||
| # reuse the graph pool instead of growing fragmented segments. | ||
| axes.extend( | ||
| ("vision", slices, patches) |
There was a problem hiding this comment.
The token-budget and slice-capacity dimensions both encode the same slice count. With the default query_num=64, the manager creates budgets 64/128/256 while this code creates slice capacities 1/2/4, so the Cartesian product captures 27 graphs after the three patch tiers. A request with s slices always has s*64 output tokens and selects ceil(s) as its slice capacity, so only (64,1), (128,2), and (256,4) can be looked up; 18 captured keys are structurally unreachable. Please derive slice capacity from the token budget or otherwise prune invalid budget/axis pairs, pin the captured key set in a test, and update the README's statement that the default captures one budget.
There was a problem hiding this comment.
Fixed in a613806: slice capacity now comes from each token budget; the default capture set has exactly 9 keys. Added the default-key regression and corrected the README.
| """Check that Stage 0 received the requested encoder graph flag.""" | ||
| stage0 = omni_runner.omni.engine.stage_configs[0] | ||
| compilation = stage0.engine_args.get("compilation_config") | ||
| return bool(compilation and compilation.get("cudagraph_mm_encoder", False)) |
There was a problem hiding this comment.
This guard proves only that the flag reached Stage 0, not that an encoder graph manager was created, captured, or replayed. That distinction matters because the only real-checkpoint result is the already documented divergent output; without an execution-level assertion, the xfail cannot establish that it exercised replay. It is also misleading on NPU: _select_encoder_cudagraph_mixin binds the protocol only when current_omni_platform.is_cuda(), so these NPU cases run eager and the parity case can compare eager with eager. Please make this graph-specific coverage CUDA-only and assert a manager, capture, and hit signal in the graph arm before comparing outputs.
There was a problem hiding this comment.
Fixed in a613806: CUDA-only E2E now checks the worker’s actual captured graphs and increasing image/video replay counters outside the final-text xfail. The video fixture now contains the configured 2 frames.
|
Validation for a613806 (before: acbcead). Full MiniCPM-o-4_5 weights; two L20s, with Stage 0 on one card and Stages 1/2 on the other; vLLM 0.29.1rc1.dev197+gab35354c2. Fresh A/P/P/A processes, identical media hashes, 64 output tokens per request; one warmup per case/process excluded (6 measured requests per case/variant).
Default captured graphs: 27 → 9. Capture logs report 6/6 s before and 3/3 s after (integer-second resolution); request medians changed by less than 0.5% on this shared host. Validation: 26 CPU tests, 4 CUDA tests passed; complete advanced-model E2E: 2 passed, 1 xfailed. Actual capture and image/video replay assertions passed outside the existing final-text parity xfail. All 48 benchmark requests completed; oversized requests used fallback. A0/P0/P1 token sequences match; A1 baseline repeated with different video/oversized text, so exact repeatability is not claimed. Both encoder arms use the same external structured-output compatibility helper and |
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
Deliver the full-checkpoint end-to-end coverage the encoder-graph work was missing, and fix the vLLM-main integration blockers that kept any serving E2E from running against public code. Compatibility (needed for real serving E2E on newer vLLM): - OmniGPUModelRunner: XD-RoPE was folded into M-RoPE upstream and the runner attribute was removed. Every read site is a plain attribute access inside the inherited __init__, so admission failed with AttributeError before any model work. Restore the attribute with the parent\x27s value when present and 0 otherwise. - Schedulers: newer vLLM moved both the reasoning-end decision and token acceptance onto StructuredOutputManager.accept_tokens(request, tokens), which also owns request history and constraint bookkeeping. Dispatch on the manager surface, and keep the old should_advance + grammar path for the released pin. The generation scheduler keeps its pre-existing behaviour of not turning a rejection into a terminal error. E2E (tests/e2e/offline_inference/test_minicpmo_4_5_encoder_cudagraph.py): - Loads the real Stage 0/1/2 pipeline with compilation_config.cudagraph_mm_encoder and once without it, using one shared deploy profile. - Serves in-budget image and video requests, and an image above the capture budget through the manager\x27s eager fallback. - Compares greedy text between the graph engine and the flag-off engine. On an L20 with vLLM 0.29.1rc1.dev197 the graph engine answers the image prompt with a repeated token while the eager engine answers correctly. Capture itself succeeds (12 graphs for the 256-token budget). The test is therefore xfail(strict=True): it documents a real replay-correctness gap instead of silently passing. CI routing: the new file joins the MiniCPM-o 4.5 offline job and its source-file dependency group, with the job timeout raised to fit the second engine. Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
a613806 to
e4149f2
Compare
Signed-off-by: 0z5a <dezhen.lu@uni-tuebingen.de> Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
e4149f2 to
f3f6e49
Compare
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
5ae7aa2 to
489f440
Compare
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
Signed-off-by: 0z5a <192209249+0z5a@users.noreply.github.com>
Signed-off-by: 0z5a <192209249+0z5a@users.noreply.github.com>
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
Signed-off-by: 0z5a <192209249+0z5a@users.noreply.github.com>
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
|
#8332 is a better implementation, PTAL |
Adds opt-in MiniCPM-o 4.5 image/video encoder CUDA Graphs through the native vLLM capture-axis protocol. Parsed vision inputs, position IDs and masks are prepared outside replay; graph budgets remain bounded. Audio and stateful duplex retain their existing encoder behavior.
The full-checkpoint repository parity test now requires equal image/video text; output differences fail the test. Scheduler compatibility and XD-RoPE compatibility are outside this feature diff. The diff contains core implementation, regression/E2E tests, required CI wiring and relevant serving guidance.
First-wave A100 full-checkpoint E2E
Two A100 PCIe 40GB GPUs; native vLLM 0.30.0, Torch 2.13.0+cu130, Transformers 5.14.1; BF16, complete checkpoint
503e754207c94da6bb26850b4469f367c9ea3582. All three stages loaded. Stage 0 ran on GPU0; Talker/Code2Wav loaded on GPU1. These requests exercised image/video-to-text.Each configured B ran four fresh A/P/P/A processes: eager encoder, Graph, Graph, eager encoder. Decoder configuration was identical (
FULL_DECODE_ONLY); prefix caching and multimodal processor caching were off, with unique multimodal IDs. Each cell submitted 2×C requests, with at most C active. Rates below are geometric means of the two corresponding arms. Fixed 224×224 images and a two-frame video were reused across arms. Output used a 128-token cap; video descriptions reach that cap.All 2,984 formal API requests completed and all twelve processes exited naturally. The two Graph arms recorded 1,492 hits and zero misses, with stable capture budgets. Peak reserved memory was 29.41 GiB. Native scheduled encoder/decoder batches are reported separately.
The original B2/C16 and B4/C8,C16 comparisons contain output differences. B4 also produced 31 differences between its two eager runs. The affected speed cells remain unqualified; their original results are retained.
Separate output diagnostics
Native encoder execution was checked against an immediate eager call on the same trained model and actual vision inputs: B2 gave 38/38 bit-exact embedding comparisons; B4 gave 56/56, including actual encoder B3. Maximum absolute error was zero.
The affected cells were then rerun in four fresh A/P/P/A processes with
VLLM_BATCH_INVARIANT=1in every arm. All 376 control requests, including warmups, produced identical corresponding text and token IDs. The combination of exact encoder embeddings, eager-to-eager variation in the original runs and repeatable invariant controls supports batch-dependent native numerical variation as the explanation; it does not qualify the original mismatching speed cells.Validation and limits
6b1d3bfa8549. The host-metadata follow-up is validated separately below.Remains Draft. Raw logs, checkpoint files and private benchmark/diagnostic harnesses are excluded from the PR diff.
A100 host-metadata follow-up: complete EOS output and concurrency
Cache vision target sizes on the host once per input and prepare layout masks there before uploading the final mask. This removes repeated device-to-host metadata synchronization from encoder graph preparation. The follow-up changes only the encoder implementation and its CPU/CUDA-metadata regression coverage.
Full pinned MiniCPM-o 4.5 checkpoint on 2×A100 40GB, native three-stage initialization, image/video-to-text route, configured B1/B2/B4, encoder graphs enabled in both arms and the same decoder graph configuration. Each configuration/arm completed 240 measured requests at C8/16/32/64. B1/B2 used 6 warmup requests per arm; B4 used 8. All 1,440 measured requests reached EOS (
finish_reason=stop, 512-token cap), and all 740 paired complete token/text outputs including warmups matched exactly. Audio generation was not exercised. All six processes exited naturally. Native decoder batch reached each configured limit; encoder batch maxima are listed separately.Baseline
688a59ed20abversusaa8c70ea577f. One unprofiled pair per B; independent CPU installers remained present. The measured changes below do not establish a meaningful E2E speedup; repeated starts remain pending. These are host-metadata comparisons with graphs enabled in both arms, separate from the first-wave eager-versus-graph results above.Separate C8 Nsight diagnostics also had 22/22 complete token/text pairs exact. Recorded
cudaStreamSynchronizecalls fell from 228 to 48 and device-to-host copies from 5,827 to 5,631. These diagnostic counts are not unprofiled E2E timing. Native regression results: 27 CPU + 8 CUDA passed, zero skips; CUDA coverage includes image/video, FP32/BF16, CPU/CUDA metadata, padded/multiple inputs and oversized eager fallback. Ruff, formatting and diff checks passed.