Repository navigation
[Bugfix] Add opt-in MiniCPM-o 4.5 Stage-0 sliding window - #7631
Conversation
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
0af1690 to
b9b5769
Compare
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
|
@natureofnature PTAL |
|
This PR appears to belong to: docs/design/module/ar_runtime.md. Module owners: @tzhouam @fake0fan @Gaohan123 Routing: @tzhouam via module of the changed files, CODEOWNERS; @fake0fan via module of the changed files; @Gaohan123 via module of the changed files @0z5a, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Omni ReviewBot triage noteAutomated triage of commit
These are automated triage suggestions only — the final decision belongs to the maintainers. |
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
amy-why-3459
left a comment
There was a problem hiding this comment.
I found two P2 issues in the window-history lifecycle and boundary accounting, detailed inline. Both were reproduced on CPU using the actual scheduler/window methods with the existing fake Stage-0 model; these were not live GPU reproductions.
Validation: 137 existing CPU tests passed across test_omni_ar_scheduler_streaming.py, test_native_duplex_hooks.py, test_minicpmo_window_plugin.py, and test_minicpmo_4_5_omni.py. The environment reported a vLLM/vLLM-Omni version mismatch warning. No GPU end-to-end test was run for this review.
Non-blocking maintainability suggestion: the generic scheduler now contains roughly 150 lines of MiniCPM-specific window policy and repeats the 8000/6000/24/500 defaults from MiniCPMO45DuplexWindowConfig. Consider extracting model-specific window planning into a dedicated module and sharing those defaults, while retaining KV release and request replacement under scheduler ownership.
| embeds = list(pending.embeds) | ||
| embeds.extend(self._embed_token(token_id) for token_id in generated) | ||
| embeds.extend(self._embed_token(token_id) for token_id in closure_token_ids) | ||
| state.window_units.append(_MiniCPMO45WindowUnit(embeds=embeds, token_ids=token_ids)) |
There was a problem hiding this comment.
[P2] Do not retain window history when sliding_window_mode is off
_finalize_window_unit() unconditionally appends the completed unit's embeddings to state.window_units. History is only removed by _window_replacement_parts() when a replacement plan is present, but the scheduler returns early for the default off mode and never supplies such a plan. Consequently, ordinary sessions now retain an additional, growing history of audio/vision embeddings even though windowing is disabled. These tensor references remain live for the session lifetime, adding memory pressure, especially with camera input.
CPU reproduction using the existing fake Stage-0 runtime: run 100 one-unit appends without any window configuration or stage0_window plan; len(state.window_units) is 99.
Please gate history collection on windowing being enabled, passing that setting to the worker from the first append so enabled windows can still preserve their initial history. Add a regression asserting that off-mode appends do not accumulate completed window units.
There was a problem hiding this comment.
Fixed in f918052. History collection is now configured before the first append: default/off sessions retain no completed or pending window history, while basic/context preserve the initial append. Added 100-append regressions for all four configurations. The four CPU regression files pass (145 tests), and the current L20 run passes all six selected session-reopen/burst-flush E2E cases.
| if isinstance(token_id, int) | ||
| } | ||
| units = list(getattr(session, "_minicpmo45_window_units", ())) | ||
| units.append( |
There was a problem hiding this comment.
[P2] Use the same unit boundaries in the scheduler and worker
This records one history unit per streaming update, spanning the previous append's entire confirmed prompt/output interval. In contrast, _stage_prefill_embeddings_only() finalizes a history unit for each audio chunk in its processing loop. An append can contain multiple chunks (for example, flushing buffered audio), so drop_units no longer identifies the same token span in both histories. A subsequent window replacement can then fail the worker's replacement_prompt_len check.
CPU reproduction using the existing fake Stage-0 model and the actual scheduler planner: the first append contains two processor chunks and produces [1, 11, 2, 1, 11]. After a confirmed output token and terminator, the next append triggers basic windowing. The scheduler plans drop_units=1, dropped_tokens=8, replacement_prompt_len=2, but the worker drops only its first chunk unit and raises:
MiniCPM-o Stage-0 window rebuild length mismatch: worker=7, scheduler=2
Please define the history boundary consistently on both sides, either grouping worker history by append or explicitly communicating/accounting for individual chunk boundaries. Add coverage for a multi-chunk append followed by a basic/context window replacement, including buffered flush input.
There was a problem hiding this comment.
Fixed in f918052 by grouping every processor chunk from one append into a single worker history unit, including internal closure tokens. Added actual-scheduler/worker regressions for basic/context replacement after multi-chunk appends and buffered carry/flush. CPU: 145 passed. Real-checkpoint L20: six selected E2E cases passed, including off/basic/context burst input with final flush; no window-length mismatch in the successful server log. Exact environment and scope are documented in the PR body.
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
|
@codex review |
linyueqian
left a comment
There was a problem hiding this comment.
Reviewed at f918052d (static read of the diff against main; no fork code executed here). Scope check for the default path first: the scheduler now computes confirmed_end, segment_output_ids and completed_terminator on every streaming update for every AR model, but they are only consumed behind the MiniCPM gate (duplex.data_plane, sliding_window_mode in basic/context, seq > 1), so block allocation, _free_request_blocks and num_computed_tokens on the ungated path are unchanged, and the worker keeps no unit history unless window_enabled is set from the first append's session config. Both of amy's P2s are addressed at this head: history collection is keyed off the session config in _prepare_session_context, and the worker folds every processor chunk of one append into a single unit including the internal closure tokens, matching the scheduler's one-unit-per-accepted-append boundary; the new lifecycle and multi-chunk tests cover both.
Unit accounting reads consistent to me: the scheduler's span [open_start, base_len + 2) deliberately includes the two closure tokens the worker re-injects at the head of the next append, replacement_prompt_len = projected - dropped equals the worker's rebuild of prefix (plus marker and previous in context mode), suffix, retained units and the pending unit, and the worker raises on a length mismatch rather than silently shifting positions, which is the right failure mode for a KV rebuild. The one correctness gap I found is inline: the replacement branch returns before _streaming_update_overflows, so a rebuilt prompt that no longer fits max_model_len reaches the worker instead of finishing the session with context_length_exceeded.
The merge risk is the e2e coverage. The nine new cases in test_minicpmo_4_5_window.py are advanced_model and run only in the merge-only Omni · MiniCPM-o 4.5 Duplex Test step (about 8 minutes on main today against a 50 minute timeout, so wall time is fine), which means their first CI execution happens on main after this merges, and the validation in the description is one L20 with custom memory limits, not the H100 CI environment. Before merging, please attach an H100 run of pytest tests/e2e/online_serving/test_minicpmo_4_5_window.py -m 'advanced_model and cuda' --run-level advanced_model at this head, or have a maintainer run that merge step on the branch; a merge-only test broke main earlier this week (#7704) and we want to avoid a repeat.
Verdict: requesting changes for the overflow bypass and the red pre-commit (both inline); the rest are suggestions and one question about the off reservation change. The general lane is still pending at f918052d. Happy to approve once those two are fixed, the lane passes, and the H100 e2e evidence is on the PR.
| segment_output_ids=segment_output_ids, | ||
| completed_terminator=completed_terminator, | ||
| ): | ||
| self._release_replaced_streaming_prompt_cache(session) |
There was a problem hiding this comment.
[important] This branch returns before _streaming_update_overflows (line 890), so a rebuilt prompt is never checked against max_model_len. The replacement length is bounded only by the client's window settings, not by the model: in context mode it is prefix + marker + previous + up to context_max_units retained units (a unit with camera frames is hundreds of tokens) + the new append, and in basic mode it is basic_window_low_tokens + the new append with no upper bound on those values beyond low < high. When that exceeds max_model_len - sample_room, _replace_streaming_session installs the prompt with num_computed_tokens = 0 and the next admission hits exactly the worker failure the overflow guard's docstring describes (input batch copy fails, the EngineCore dies with every session on it) instead of finishing this one session with context_length_exceeded. Please check len(update.prompt_token_ids) + sample_room against max_model_len here before replacing (the existing helper computes projected from the old prompt, so it needs a replacement-aware variant), and consider clamping basic_window_high_tokens to the model length at session creation. The new scheduler tests build the scheduler without max_model_len and mock _free_request_blocks, so they cannot catch this.
There was a problem hiding this comment.
Fixed in 4dd174f.
The replacement branch now runs _streaming_update_overflows with projected_len = stage0_window["replacement_prompt_len"]: a rebuild replaces the whole prompt, so the projection is the new length, not the session prompt plus an extension. Over the line, the session finishes with the same context_length_exceeded error output as a plain extension instead of handing the prompt to the worker.
Regressions at 06d71be:
test_stage0_window_rebuild_that_overflows_max_model_len_finishes_the_sessiondrives a 21-token plan withnum_sampled_tokens_per_step = 1againstmax_model_len = 21through the real finish path and asserts the terminal ERROR output, the emptied admission queues and the recorded reason.test_stage0_window_rebuild_that_leaves_room_to_sample_replaces_the_promptasserts that 25 (plan + one sampled token) still replaces through the normal path.
| previous_len = int(getattr(session, "_minicpmo45_window_previous_len", 0) or 0) | ||
| new_preserve_len = prefix_len + previous_len + suffix_len if mode == "context" else preserve_len | ||
| session._minicpmo45_window_open_start = new_preserve_len + sum(int(unit["length"]) for unit in retained_units) | ||
| return True |
There was a problem hiding this comment.
[important] pre-commit is red on this file: ruff-format wants a blank line after this return True before the STREAMING_CONTEXT_OVERFLOW_STOP_REASON comment block (the job reports 1 file reformatted). A pre-commit run --all-files locally reproduces it.
| } | ||
| if completed_terminator is not None: | ||
| duplex["stage0_window"]["completed_terminator_token_id"] = int(completed_terminator) | ||
| prefix_len = int(runtime_config.get("duplex_window_prefix_tokens", preserve_len) or preserve_len) |
There was a problem hiding this comment.
[suggestion] or preserve_len treats a legitimate zero prefix count as missing and substitutes the whole first-append context length (prefix, reference audio and suffix), so new_preserve_len would count the suffix twice and the worker's length check at stage0.py would raise on the first rebuild. duplex_window_prefix_tokens is always set by _apply_first_append_context_tokens, so today this only bites with an empty prefix, but an explicit is None check is cheaper than debugging that mismatch later; the same pattern is at line 972 for _minicpmo45_window_open_start.
There was a problem hiding this comment.
Fixed in 4dd174f, both sites, with explicit is None checks.
The regressions show the difference rather than just the branch taken: with the recorded open_start = 0 the plan reports a 12-token dropped span and a 19-token replacement; with it unset the preserve_len fallback reports 9 and 22. Treating 0 as missing therefore shifts the previous region by the context reserve exactly as you described.
| previous = stage0_window.get("previous_token_ids") | ||
| previous_ids = [int(token_id) for token_id in previous] if isinstance(previous, list) else [] | ||
| if previous_ids: | ||
| marker_ids = self._encode_text("\n\nprevious: ") |
There was a problem hiding this comment.
[suggestion] The worker re-tokenizes "\n\nprevious: " here while the scheduler sizes the same block from duplex_window_previous_marker_token_ids, which the plugin produced with its own tokenizer; _encode_text returns [] when the runtime tokenizer has no encode, and any divergence only surfaces as the RuntimeError below. Carrying the marker ids in the stage0_window plan (or reading them from runtime_config) keeps one source of truth for both sides.
There was a problem hiding this comment.
Fixed in 6d39337.
The plan now carries previous_marker_token_ids, taken from duplex_window_previous_marker_token_ids (the ids the plugin produced and the scheduler sized the previous region with), and _window_replacement_parts embeds and appends those ids. _encode_text stays for the session context prefix and suffix. If a plan ever arrives without the key, the worker rebuilds without the marker and the existing replacement_prompt_len check raises instead of silently shifting positions.
| token_budget += 1 | ||
| if final and duplex_payload_is_exact_chunks(payload): | ||
| token_budget += 12 | ||
| # Serving already pads the final residual audio. Stage0 does not |
There was a problem hiding this comment.
[question] Dropping the final reservation is unconditional, so a default off session's one-unit exact final append now reserves 13 slots instead of 25. I follow the reasoning (Stage0 never emitted the silent unit, so the 12 slots were never filled), but this is the one change in the PR that reaches every MiniCPM-o duplex session regardless of window settings. Can you confirm what happened to those 12 unfilled slots on main (trimmed via num_input_tokens, or prefilled as scheduler-token positions?) and that test_window_buffered_flush[off] passed at this head? If they were prefilled, this is a silent behaviour fix for off worth its own line in the description.
There was a problem hiding this comment.
On main those 12 slots were prefilled, not trimmed. MiniCPMO45OmniForConditionalGeneration pads the rebuilt embeddings up to duplex_prompt_len and prepends the pad block (full_req_embeds = torch.cat([pad_embeds, full_req_embeds])), so the runner scheduled [num_computed_tokens, prompt_len) and the 12 pad tokens became real KV positions immediately before the audio. The audio therefore sat at a later RoPE position and the pads were inside the model's attention context - none of it removed afterwards.
That matches your reasoning that Stage0 never emitted the silent unit, so the removal is exact: the scheduler's num_prompt_tokens now equals the worker's embedding count and no padding or truncation happens on an exact-chunk final append.
The off case of test_window_buffered_flush is unaffected by the window code path (see the PR-body note: the L20 host here cannot boot the three-stage server with its prebuilt vLLM wheel, so I could not re-run the module end to end). test_final_exact_chunk_append_reserves_one_unit_not_two now asserts a final exact-chunk append reserves exactly what a non-final append of the same audio does for seq 1/2/3.
Will fix them. H100 e2e run may take some time to settle, since I have no access to h100. Maybe I will rent some h100 to finish this when it's okay. Volunteers for H100 e2e test are welcome. |
The Stage-0 window replacement branch returned before ``_streaming_update_overflows``, so the rebuilt prompt was bounded only by the client's window settings: in ``context`` mode it is prefix + marker + ``previous`` + up to ``context_max_units`` retained units (a camera unit is hundreds of tokens) + the new append, and in ``basic`` mode it is ``basic_window_low_tokens`` + the append, with no upper bound beyond ``low < high``. A prompt over ``max_model_len - sample_room`` reached the worker instead of finishing the session with ``context_length_exceeded``. The plan already carries ``replacement_prompt_len``, so the guard now takes that as an explicit projection: a replacement is not the session's current prompt plus an extension. Two scheduler regressions cover the overflow and the one-slot-larger replacement that still applies. Also stop treating a legitimate zero ``_minicpmo45_window_open_start`` (an empty context prefix) as missing: ``or preserve_len`` substituted the whole first-append context length, which double-counts the suffix and makes the worker's rebuild-length check raise on the first replacement. Same for the recorded ``duplex_window_prefix_tokens``. Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
The worker re-tokenized "\n\nprevious: " with its own tokenizer while the scheduler sized the same block from ``duplex_window_previous_marker_token_ids``, which the plugin produced with the runtime tokenizer. Any divergence only surfaced as the rebuild-length RuntimeError. The window plan now carries the marker ids the scheduler counted, so both sides use one source of truth; ``_encode_text`` stays for the session context prefix and suffix. Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
The unconditional extra 12 slots for a ``final`` exact-chunk append were removed because Stage0 never fed the silent second unit. Serving pads the final residual audio itself, so the reservation must equal one unit per chunk (``<unit>`` + its audio) plus the closure pair for every unit after the first. Add the off-mode regression that asserts a final append reserves exactly what a non-final append of the same audio does. Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
The overflow regression needs a plan whose replacement length is at the sample-room boundary, so pin num_sampled_tokens_per_step and the append length instead of relying on the window defaults. The zero-open_start regressions now drive the planner with the confirmed output span it reads, and assert the dropped span and projection that the two fallbacks produce. Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
linyueqian
left a comment
There was a problem hiding this comment.
Re-reviewed at 06d71be7 (static read of the four follow-up commits; nothing executed). Everything from the first pass is addressed. The replacement branch now runs _streaming_update_overflows with the plan's replacement_prompt_len as the projection before releasing the cache, so a rebuilt window that no longer leaves room to sample finishes the session with context_length_exceeded instead of reaching the worker, and the two new scheduler tests pin both sides of that boundary. The open_start and duplex_window_prefix_tokens fallbacks are explicit is None checks, with a test for a legitimate zero. The previous: marker ids travel in the stage0_window plan and the worker embeds those instead of re-tokenizing, so the length the scheduler computed and the sequence the worker builds come from one source. ruff format is clean on every touched file, and the final-append reservation change now has a test stating the one-unit intent for off as well as the window modes.
What this approval does not cover is the merge-only e2e: test_minicpmo_4_5_window.py still runs for the first time on main after the merge, and the author has no H100 access. I would like a maintainer with H100 access to run the Omni · MiniCPM-o 4.5 Duplex Test merge step on this branch (or an equivalent run of the new module) before this is merged, and I will hold the merge until that run or an explicit maintainer decision to take the risk. Approving the content.
|
H100 run of the merge-only step, at
This is the H100 evidence @linyueqian asked for; the CPU set is |
|
After the first 24 units, each time we will do a re-prefilling. How much impact does that have on the latency? |
Yes, for context mode that's correct. For audio-only input, the retained 24 units are relatively small (roughly 12 Stage-0 positions per unit, plus the preserved prefix/suffix and up to 500 tokens of compacted previous: text), but camera units can make the rebuilt prompt substantially larger. So the actual cost should be measured against replacement_prompt_len, especially for audio+video. Also volunteers are welcome for this h100 test, self-funded H100 is not very affordable to me. Maybe I will have another benchmark when it's ok. |
Have you already solved the CI issues? The CI still fails, due to some out-of-the-scope issue, you may rebase to the latest main first |
will check it soon |
Signed-off-by: 0z5a <dezhen.lu@uni-tuebingen.de>
|
Merged main and reran the core/MiniCPM-o CPU selection on the SSH host: 853 passed, 1 skipped (Ruff check/format passed). This includes main's switch from realtime to chat TTS benchmarking and the CUDA-only MAGI2 test guard, matching the old NVIDIA/AMD failures. The old Intel run failed during checkout with |
|
Ran the whole merge-only "Omni · MiniCPM-o 4.5 Duplex Test" step at |
|
@0z5a heads-up on |
…t#7631) Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de> Signed-off-by: 0z5a <dezhen.lu@uni-tuebingen.de> Co-authored-by: 0z5a <dezhen.lu@uni-tuebingen.de> Signed-off-by: Matthieu Laneuville <matthieu.laneuville@surf.nl>
…t#7631) Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de> Signed-off-by: 0z5a <dezhen.lu@uni-tuebingen.de> Co-authored-by: 0z5a <dezhen.lu@uni-tuebingen.de>
Purpose
Add opt-in MiniCPM-o 4.5 Thinker KV windows for #7251.
offremains the default.basicdrops old units at token watermarks;contextalso carries compacted speak text in aprevious:region. Both preserve the system/reference-audio prefix and rebuild retained positions through scheduler-owned re-prefill.The scheduler derives window boundaries from confirmed canonical tokens and passes accepted terminators to the worker. This keeps queued async updates consistent after resumable output lists are cleared. Native
tts_boshandoffs use segment-end hidden-state alignment, matching the existing speak path: an earlier listen decision already folded into the rebuilt prompt must not be counted twice. Otherwise a basic window with reference audio can forward two text tokens with only one hidden row and terminate Talker.This changes Stage-0 LLM KV, separate from the PCM/mel buffer, streaming audio-encoder reset, and Talker window. It uses the engine-owned duplex session/plugin architecture.
Test Plan
Published commit:
f8e646960dea18d1d83785a9c7acf06a7daff9c1. Its tree995eef5abcafe1d7b08fa9c8023e6e9fc23f3139is identical to tested local commitaeaf7e09bd1f7a06501b91a641d3c6614f5ea239.Real checkpoint, BF16, one NVIDIA L20, vLLM 0.29.0, PyTorch 2.13/CUDA 13. Three-stage
minicpmo_4_5.yaml, decoder and native Code2Wav graphs enabled. L20 Stage-0 overrides: KV 2 GiB, model length 8192, batched tokens 4096, image/audio/video limits 2, video fps 1 and frames 2. These are explicit memory limits, not default-deployment acceptance. The existing package-version mismatch warning remains.tts_bosalignment casesThe real-service run used the committed test bodies with a private fixture pointing at the already-started service and local checkpoint reference audio. Normal CI uses
omni_server; the new suite is included in the CUDA advanced-model merge job. Fixtures repeat the repository WAV and use a synthetic camera clip held at its last frame, so these tests establish window execution/completion and session lifecycle, not generation quality.With project/test dependencies installed in
.venv, local regression:CI-equivalent E2E (one H100 with the normal model-cache/checkpoint setup; not run on H100 in this report):
uv run --python .venv/bin/python --no-project python -m pytest \ tests/e2e/online_serving/test_minicpmo_4_5_window.py \ -m 'advanced_model and cuda' --run-level advanced_model -vEarlier canonical-boundary validation at
6817745: scheduler/worker 103 passed and session/plugin/serving 95 passed. The new commit adds the Talker boundary fix and committed real-model E2E coverage. Raw failing/fixed server logs and JUnit are retained. Draft for review.Review follow-up validation
Commit:
f918052dbcdc9c4e9af5e961bfde8154f83c2cae.Current-commit results:
test_minicpmo_4_5_window.py -k 'not continuous_input': 6 passed, 3 deselected, 200.50 s. This runs the three session-reopen cases and all three new burst/flush cases. The existing continuous-input cases were not rerun at this commit; their earlier results above remain attributed to the earlier commit.git diff --checkpassed.The current L20 run used the same explicit Stage-0 memory overrides documented above, the real local BF16 checkpoint, and a private fixture selecting the already-started service/local reference audio. Public test bodies were unchanged. Dependencies were pinned to vLLM 0.29.0 and matching FlashInfer 0.6.16.post3 without changing the shared environment. This is L20 validation, not H100/default-deployment validation. CPU and E2E logs, JUnit and the successful server log are retained on jk01 under
/home/kxqandccx/minicpmo45-window-e2e/evidence/review7631-*.The non-blocking scheduler-policy extraction suggestion is left for a separate refactor; this follow-up addresses the two reproduced correctness issues.
Review follow-up (f918052 -> 06d71be)
4dd174f2,6d393371,6b384947,06d71be7.max_model_len. The replacement branch returned before_streaming_update_overflows, so a rebuilt prompt was bounded only by the client's window settings. It now runs the same guard withprojected_len = stage0_window["replacement_prompt_len"], because a replacement is the new prompt length, not the session prompt plus an extension. One regression drives a 21-token plan againstmax_model_len=21and asserts the terminalcontext_length_exceededoutput through the real finish path; a second asserts that 25 (plan + the token one step samples) still replaces.open_start._minicpmo45_window_open_start or preserve_lenandruntime_config.get("duplex_window_prefix_tokens", preserve_len) or preserve_lentreated a legitimate 0 as missing. Both are explicitis Nonechecks now. The regressions show the difference: an empty prefix reports a 12-token dropped span and a 19-token replacement, the unset fallback 9 and 22.previous_marker_token_idsfromduplex_window_previous_marker_token_ids, and the worker embeds those ids instead of re-tokenizing"\n\nprevious: ", so the scheduler'spreviouslength and the worker's rebuild have one source of truth.MiniCPMO45OmniForConditionalGenerationpads the rebuilt embeddings up toduplex_prompt_lenand prepends the pad block, so onmainthose slots were 12 real pad-token KV positions immediately ahead of the audio (audio shifted to a later RoPE position, pads inside the model's attention context). Theoffcase oftest_window_buffered_flushpasses at this head, andtest_final_exact_chunk_append_reserves_one_unit_not_twoasserts a final append reserves exactly what a non-final one of the same audio does.ruff@0.14.10 format --checkandruff checkare clean on all five changed files (1 file reformattedgone).Validation at
06d71be7(jk01, 8x L20; a purpose-built vLLM 0.29.1rc1 + torch environment, not the shared one):pytest -m "core_model and cpu" tests/core/sched/test_omni_ar_scheduler_streaming.py tests/worker/test_native_duplex_hooks.py tests/engine/duplex/test_minicpmo_window_plugin.py tests/model_executor/stage_input_processors/test_minicpmo_4_5_omni.py: 152 passed, 3.04 s.pytest tests/core/sched/test_omni_ar_scheduler_streaming.py tests/engine/duplex/test_minicpmo_window_plugin.py tests/worker/test_native_duplex_hooks.py tests/model_executor/stage_input_processors/test_minicpmo_4_5_omni.py tests/engine/duplex/: 496 passed, 1 skipped, 14.55 s.test_minicpmo_4_5_window.py -m 'advanced_model and cuda' --run-level advanced_model: 9 passed, 981.35 s (16:21). One H100 80GB HBM3, vLLM 0.29.0 / vLLM-Omni06d71be7, torch 2.13.0+cu130, real BF16 checkpoint, the exact merge-step command with no-kfilter and no custom Stage-0 overrides. All three modules passed: session-reopen (basic / context / context+reference audio), continuous input (basic / context / context+camera) and buffered flush (off / basic / context). See the PR comment for details.