Skip to content

[Bugfix] Add opt-in MiniCPM-o 4.5 Stage-0 sliding window - #7631

Merged
linyueqian merged 10 commits into
vllm-project:mainfrom
0z5a:stage0-sliding-window
Sep 22, 2026
Merged

linyueqian merged 10 commits into
vllm-project:mainfrom
0z5a:stage0-sliding-window

Conversation

@0z5a

@0z5a 0z5a commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Add opt-in MiniCPM-o 4.5 Thinker KV windows for #7251. off remains the default. basic drops old units at token watermarks; context also carries compacted speak text in a previous: region. Both preserve the system/reference-audio prefix and rebuild retained positions through scheduler-owned re-prefill.

The scheduler derives window boundaries from confirmed canonical tokens and passes accepted terminators to the worker. This keeps queued async updates consistent after resumable output lists are cleared. Native tts_bos handoffs use segment-end hidden-state alignment, matching the existing speak path: an earlier listen decision already folded into the rebuilt prompt must not be counted twice. Otherwise a basic window with reference audio can forward two text tokens with only one hidden row and terminate Talker.

This changes Stage-0 LLM KV, separate from the PCM/mel buffer, streaming audio-encoder reset, and Talker window. It uses the engine-owned duplex session/plugin architecture.

Test Plan

Published commit: f8e646960dea18d1d83785a9c7acf06a7daff9c1. Its tree 995eef5abcafe1d7b08fa9c8023e6e9fc23f3139 is identical to tested local commit aeaf7e09bd1f7a06501b91a641d3c6614f5ea239.

Real checkpoint, BF16, one NVIDIA L20, vLLM 0.29.0, PyTorch 2.13/CUDA 13. Three-stage minicpmo_4_5.yaml, decoder and native Code2Wav graphs enabled. L20 Stage-0 overrides: KV 2 GiB, model length 8192, batched tokens 4096, image/audio/video limits 2, video fps 1 and frames 2. These are explicit memory limits, not default-deployment acceptance. The existing package-version mismatch warning remains.

Validation Result
Thinker/Talker handoff, pipeline and serving regressions (5 files) 69 passed, 1 skipped; 3.92 s
New tts_bos alignment cases 6 passed within the regression suite
Basic short input, then a new session Passed; two 10.94 s inputs
Context short input, then a new session Passed; two 10.94 s inputs
Context + reference audio, then a new session Passed; two 10.94 s inputs, audio returned
Basic continuous input + reference audio Passed; 131.26 s input, audio returned
Context continuous input + reference audio Passed; 131.26 s input, audio returned
Context continuous audio/video + reference audio Passed; 131.26 s input, at least 120 frames, audio returned
Complete public E2E test bodies against the live L20 service 6 passed; 613.47 s
End-of-suite health Healthy; no client/server errors in the fixed run
Ruff, formatting, test marks, CI YAML, diff checks Passed

The real-service run used the committed test bodies with a private fixture pointing at the already-started service and local checkpoint reference audio. Normal CI uses omni_server; the new suite is included in the CUDA advanced-model merge job. Fixtures repeat the repository WAV and use a synthetic camera clip held at its last frame, so these tests establish window execution/completion and session lifecycle, not generation quality.

With project/test dependencies installed in .venv, local regression:

uv run --python .venv/bin/python --no-project python -m pytest \
  tests/model_executor/stage_input_processors/test_minicpmo_4_5_omni.py \
  tests/model_executor/models/minicpmo_4_5/test_llm2tts.py \
  tests/model_executor/models/minicpmo_4_5/test_pipeline.py \
  tests/entrypoints/openai_api/test_serving_chat_minicpmo45_stream_output_kinds.py \
  tests/entrypoints/openai_api/test_serving_chat_minicpmo45_reference_audio.py \
  --run-level core_model -q

CI-equivalent E2E (one H100 with the normal model-cache/checkpoint setup; not run on H100 in this report):

uv run --python .venv/bin/python --no-project python -m pytest \
  tests/e2e/online_serving/test_minicpmo_4_5_window.py \
  -m 'advanced_model and cuda' --run-level advanced_model -v

Earlier canonical-boundary validation at 6817745: scheduler/worker 103 passed and session/plugin/serving 95 passed. The new commit adds the Talker boundary fix and committed real-model E2E coverage. Raw failing/fixed server logs and JUnit are retained. Draft for review.

Review follow-up validation

Commit: f918052dbcdc9c4e9af5e961bfde8154f83c2cae.

  • Window history collection is initialized from the first append's session configuration. Default/off sessions retain neither completed window units nor a pending history unit; basic/context preserve history from the initial append.
  • Worker history now groups all processor chunks of one append into one unit, matching the scheduler's accepted-append boundaries, including internal closure tokens and buffered flush input.
  • Added eight CPU regressions: absent/off/basic/context history lifecycle over 100 appends, and basic/context replacement after multi-chunk input with and without buffered carry/flush.
  • Added three real-model E2E cases for burst input plus final flush in off/basic/context modes.

Current-commit results:

  • The four scheduler, worker, window-plugin and Stage-0 input-processor regression files: 145 passed, 2.94 s.
  • L20 real-model test_minicpmo_4_5_window.py -k 'not continuous_input': 6 passed, 3 deselected, 200.50 s. This runs the three session-reopen cases and all three new burst/flush cases. The existing continuous-input cases were not rerun at this commit; their earlier results above remain attributed to the earlier commit.
  • Final service health: HTTP 200. Successful-run server log: no ERROR, traceback or window-length mismatch; 75 basic and 45 context KV replacements observed.
  • All changed-file pre-commit hooks and git diff --check passed.

The current L20 run used the same explicit Stage-0 memory overrides documented above, the real local BF16 checkpoint, and a private fixture selecting the already-started service/local reference audio. Public test bodies were unchanged. Dependencies were pinned to vLLM 0.29.0 and matching FlashInfer 0.6.16.post3 without changing the shared environment. This is L20 validation, not H100/default-deployment validation. CPU and E2E logs, JUnit and the successful server log are retained on jk01 under /home/kxqandccx/minicpmo45-window-e2e/evidence/review7631-*.

The non-blocking scheduler-policy extraction suggestion is left for a separate refactor; this follow-up addresses the two reproduced correctness issues.

Review follow-up (f918052 -> 06d71be)

4dd174f2, 6d393371, 6b384947, 06d71be7.

  • Window rebuild vs max_model_len. The replacement branch returned before _streaming_update_overflows, so a rebuilt prompt was bounded only by the client's window settings. It now runs the same guard with projected_len = stage0_window["replacement_prompt_len"], because a replacement is the new prompt length, not the session prompt plus an extension. One regression drives a 21-token plan against max_model_len=21 and asserts the terminal context_length_exceeded output through the real finish path; a second asserts that 25 (plan + the token one step samples) still replaces.
  • Zero open_start. _minicpmo45_window_open_start or preserve_len and runtime_config.get("duplex_window_prefix_tokens", preserve_len) or preserve_len treated a legitimate 0 as missing. Both are explicit is None checks now. The regressions show the difference: an empty prefix reports a 12-token dropped span and a 19-token replacement, the unset fallback 9 and 22.
  • Previous-marker ids. The plan now carries previous_marker_token_ids from duplex_window_previous_marker_token_ids, and the worker embeds those ids instead of re-tokenizing "\n\nprevious: ", so the scheduler's previous length and the worker's rebuild have one source of truth.
  • Final exact-chunk append. Confirmed and covered: the 12 removed slots were never trimmed. MiniCPMO45OmniForConditionalGeneration pads the rebuilt embeddings up to duplex_prompt_len and prepends the pad block, so on main those slots were 12 real pad-token KV positions immediately ahead of the audio (audio shifted to a later RoPE position, pads inside the model's attention context). The off case of test_window_buffered_flush passes at this head, and test_final_exact_chunk_append_reserves_one_unit_not_two asserts a final append reserves exactly what a non-final one of the same audio does.
  • ruff-format. ruff@0.14.10 format --check and ruff check are clean on all five changed files (1 file reformatted gone).

Validation at 06d71be7 (jk01, 8x L20; a purpose-built vLLM 0.29.1rc1 + torch environment, not the shared one):

  • pytest -m "core_model and cpu" tests/core/sched/test_omni_ar_scheduler_streaming.py tests/worker/test_native_duplex_hooks.py tests/engine/duplex/test_minicpmo_window_plugin.py tests/model_executor/stage_input_processors/test_minicpmo_4_5_omni.py: 152 passed, 3.04 s.
  • pytest tests/core/sched/test_omni_ar_scheduler_streaming.py tests/engine/duplex/test_minicpmo_window_plugin.py tests/worker/test_native_duplex_hooks.py tests/model_executor/stage_input_processors/test_minicpmo_4_5_omni.py tests/engine/duplex/: 496 passed, 1 skipped, 14.55 s.
  • H100 real-model test_minicpmo_4_5_window.py -m 'advanced_model and cuda' --run-level advanced_model: 9 passed, 981.35 s (16:21). One H100 80GB HBM3, vLLM 0.29.0 / vLLM-Omni 06d71be7, torch 2.13.0+cu130, real BF16 checkpoint, the exact merge-step command with no -k filter and no custom Stage-0 overrides. All three modules passed: session-reopen (basic / context / context+reference audio), continuous input (basic / context / context+camera) and buffered flush (off / basic / context). See the PR comment for details.

@0z5a 0z5a changed the title [MiniCPM-o 4.5] Add opt-in Stage-0 sliding window [Bugfix] Add opt-in MiniCPM-o 4.5 Stage-0 sliding window Sep 16, 2026
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
@0z5a
0z5a force-pushed the stage0-sliding-window branch from 0af1690 to b9b5769 Compare September 16, 2026 10:52
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
@amy-why-3459

Copy link
Copy Markdown
Collaborator

@natureofnature PTAL

@0z5a
0z5a marked this pull request as ready for review September 16, 2026 14:40
@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/ar_runtime.md.

Module owners: @tzhouam @fake0fan @Gaohan123

Routing: @tzhouam via module of the changed files, CODEOWNERS; @fake0fan via module of the changed files; @Gaohan123 via module of the changed files

@0z5a, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 16, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit a792ef078285 produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
@hsliuustc0106 hsliuustc0106 added bug Something isn't working core related to core module: cache, scheduler, engine, worker, modelrunner labels Sep 16, 2026

@amy-why-3459 amy-why-3459 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I found two P2 issues in the window-history lifecycle and boundary accounting, detailed inline. Both were reproduced on CPU using the actual scheduler/window methods with the existing fake Stage-0 model; these were not live GPU reproductions.

Validation: 137 existing CPU tests passed across test_omni_ar_scheduler_streaming.py, test_native_duplex_hooks.py, test_minicpmo_window_plugin.py, and test_minicpmo_4_5_omni.py. The environment reported a vLLM/vLLM-Omni version mismatch warning. No GPU end-to-end test was run for this review.

Non-blocking maintainability suggestion: the generic scheduler now contains roughly 150 lines of MiniCPM-specific window policy and repeats the 8000/6000/24/500 defaults from MiniCPMO45DuplexWindowConfig. Consider extracting model-specific window planning into a dedicated module and sharing those defaults, while retaining KV release and request replacement under scheduler ownership.

embeds = list(pending.embeds)
embeds.extend(self._embed_token(token_id) for token_id in generated)
embeds.extend(self._embed_token(token_id) for token_id in closure_token_ids)
state.window_units.append(_MiniCPMO45WindowUnit(embeds=embeds, token_ids=token_ids))

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Do not retain window history when sliding_window_mode is off

_finalize_window_unit() unconditionally appends the completed unit's embeddings to state.window_units. History is only removed by _window_replacement_parts() when a replacement plan is present, but the scheduler returns early for the default off mode and never supplies such a plan. Consequently, ordinary sessions now retain an additional, growing history of audio/vision embeddings even though windowing is disabled. These tensor references remain live for the session lifetime, adding memory pressure, especially with camera input.

CPU reproduction using the existing fake Stage-0 runtime: run 100 one-unit appends without any window configuration or stage0_window plan; len(state.window_units) is 99.

Please gate history collection on windowing being enabled, passing that setting to the worker from the first append so enabled windows can still preserve their initial history. Add a regression asserting that off-mode appends do not accumulate completed window units.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in f918052. History collection is now configured before the first append: default/off sessions retain no completed or pending window history, while basic/context preserve the initial append. Added 100-append regressions for all four configurations. The four CPU regression files pass (145 tests), and the current L20 run passes all six selected session-reopen/burst-flush E2E cases.

if isinstance(token_id, int)
}
units = list(getattr(session, "_minicpmo45_window_units", ()))
units.append(

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Use the same unit boundaries in the scheduler and worker

This records one history unit per streaming update, spanning the previous append's entire confirmed prompt/output interval. In contrast, _stage_prefill_embeddings_only() finalizes a history unit for each audio chunk in its processing loop. An append can contain multiple chunks (for example, flushing buffered audio), so drop_units no longer identifies the same token span in both histories. A subsequent window replacement can then fail the worker's replacement_prompt_len check.

CPU reproduction using the existing fake Stage-0 model and the actual scheduler planner: the first append contains two processor chunks and produces [1, 11, 2, 1, 11]. After a confirmed output token and terminator, the next append triggers basic windowing. The scheduler plans drop_units=1, dropped_tokens=8, replacement_prompt_len=2, but the worker drops only its first chunk unit and raises:

MiniCPM-o Stage-0 window rebuild length mismatch: worker=7, scheduler=2

Please define the history boundary consistently on both sides, either grouping worker history by append or explicitly communicating/accounting for individual chunk boundaries. Add coverage for a multi-chunk append followed by a basic/context window replacement, including buffered flush input.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in f918052 by grouping every processor chunk from one append into a single worker history unit, including internal closure tokens. Added actual-scheduler/worker regressions for basic/context replacement after multi-chunk appends and buffered carry/flush. CPU: 145 passed. Real-checkpoint L20: six selected E2E cases passed, including off/basic/context burst input with final flush; no window-length mismatch in the successful server log. Exact environment and scope are documented in the PR body.

Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
@0z5a
0z5a requested a review from amy-why-3459 September 17, 2026 04:50
@natureofnature

Copy link
Copy Markdown
Collaborator

@codex review

@linyueqian linyueqian added the ready label to trigger buildkite CI label Sep 17, 2026

@linyueqian linyueqian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed at f918052d (static read of the diff against main; no fork code executed here). Scope check for the default path first: the scheduler now computes confirmed_end, segment_output_ids and completed_terminator on every streaming update for every AR model, but they are only consumed behind the MiniCPM gate (duplex.data_plane, sliding_window_mode in basic/context, seq > 1), so block allocation, _free_request_blocks and num_computed_tokens on the ungated path are unchanged, and the worker keeps no unit history unless window_enabled is set from the first append's session config. Both of amy's P2s are addressed at this head: history collection is keyed off the session config in _prepare_session_context, and the worker folds every processor chunk of one append into a single unit including the internal closure tokens, matching the scheduler's one-unit-per-accepted-append boundary; the new lifecycle and multi-chunk tests cover both.

Unit accounting reads consistent to me: the scheduler's span [open_start, base_len + 2) deliberately includes the two closure tokens the worker re-injects at the head of the next append, replacement_prompt_len = projected - dropped equals the worker's rebuild of prefix (plus marker and previous in context mode), suffix, retained units and the pending unit, and the worker raises on a length mismatch rather than silently shifting positions, which is the right failure mode for a KV rebuild. The one correctness gap I found is inline: the replacement branch returns before _streaming_update_overflows, so a rebuilt prompt that no longer fits max_model_len reaches the worker instead of finishing the session with context_length_exceeded.

The merge risk is the e2e coverage. The nine new cases in test_minicpmo_4_5_window.py are advanced_model and run only in the merge-only Omni · MiniCPM-o 4.5 Duplex Test step (about 8 minutes on main today against a 50 minute timeout, so wall time is fine), which means their first CI execution happens on main after this merges, and the validation in the description is one L20 with custom memory limits, not the H100 CI environment. Before merging, please attach an H100 run of pytest tests/e2e/online_serving/test_minicpmo_4_5_window.py -m 'advanced_model and cuda' --run-level advanced_model at this head, or have a maintainer run that merge step on the branch; a merge-only test broke main earlier this week (#7704) and we want to avoid a repeat.

Verdict: requesting changes for the overflow bypass and the red pre-commit (both inline); the rest are suggestions and one question about the off reservation change. The general lane is still pending at f918052d. Happy to approve once those two are fixed, the lane passes, and the H100 e2e evidence is on the PR.

segment_output_ids=segment_output_ids,
completed_terminator=completed_terminator,
):
self._release_replaced_streaming_prompt_cache(session)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[important] This branch returns before _streaming_update_overflows (line 890), so a rebuilt prompt is never checked against max_model_len. The replacement length is bounded only by the client's window settings, not by the model: in context mode it is prefix + marker + previous + up to context_max_units retained units (a unit with camera frames is hundreds of tokens) + the new append, and in basic mode it is basic_window_low_tokens + the new append with no upper bound on those values beyond low < high. When that exceeds max_model_len - sample_room, _replace_streaming_session installs the prompt with num_computed_tokens = 0 and the next admission hits exactly the worker failure the overflow guard's docstring describes (input batch copy fails, the EngineCore dies with every session on it) instead of finishing this one session with context_length_exceeded. Please check len(update.prompt_token_ids) + sample_room against max_model_len here before replacing (the existing helper computes projected from the old prompt, so it needs a replacement-aware variant), and consider clamping basic_window_high_tokens to the model length at session creation. The new scheduler tests build the scheduler without max_model_len and mock _free_request_blocks, so they cannot catch this.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 4dd174f.

The replacement branch now runs _streaming_update_overflows with projected_len = stage0_window["replacement_prompt_len"]: a rebuild replaces the whole prompt, so the projection is the new length, not the session prompt plus an extension. Over the line, the session finishes with the same context_length_exceeded error output as a plain extension instead of handing the prompt to the worker.

Regressions at 06d71be:

  • test_stage0_window_rebuild_that_overflows_max_model_len_finishes_the_session drives a 21-token plan with num_sampled_tokens_per_step = 1 against max_model_len = 21 through the real finish path and asserts the terminal ERROR output, the emptied admission queues and the recorded reason.
  • test_stage0_window_rebuild_that_leaves_room_to_sample_replaces_the_prompt asserts that 25 (plan + one sampled token) still replaces through the normal path.

previous_len = int(getattr(session, "_minicpmo45_window_previous_len", 0) or 0)
new_preserve_len = prefix_len + previous_len + suffix_len if mode == "context" else preserve_len
session._minicpmo45_window_open_start = new_preserve_len + sum(int(unit["length"]) for unit in retained_units)
return True

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[important] pre-commit is red on this file: ruff-format wants a blank line after this return True before the STREAMING_CONTEXT_OVERFLOW_STOP_REASON comment block (the job reports 1 file reformatted). A pre-commit run --all-files locally reproduces it.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 4dd174f (the blank line after return True).

ruff@0.14.10 format --check and ruff check are clean on all five changed files at 06d71be.

}
if completed_terminator is not None:
duplex["stage0_window"]["completed_terminator_token_id"] = int(completed_terminator)
prefix_len = int(runtime_config.get("duplex_window_prefix_tokens", preserve_len) or preserve_len)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[suggestion] or preserve_len treats a legitimate zero prefix count as missing and substitutes the whole first-append context length (prefix, reference audio and suffix), so new_preserve_len would count the suffix twice and the worker's length check at stage0.py would raise on the first rebuild. duplex_window_prefix_tokens is always set by _apply_first_append_context_tokens, so today this only bites with an empty prefix, but an explicit is None check is cheaper than debugging that mismatch later; the same pattern is at line 972 for _minicpmo45_window_open_start.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 4dd174f, both sites, with explicit is None checks.

The regressions show the difference rather than just the branch taken: with the recorded open_start = 0 the plan reports a 12-token dropped span and a 19-token replacement; with it unset the preserve_len fallback reports 9 and 22. Treating 0 as missing therefore shifts the previous region by the context reserve exactly as you described.

previous = stage0_window.get("previous_token_ids")
previous_ids = [int(token_id) for token_id in previous] if isinstance(previous, list) else []
if previous_ids:
marker_ids = self._encode_text("\n\nprevious: ")

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[suggestion] The worker re-tokenizes "\n\nprevious: " here while the scheduler sizes the same block from duplex_window_previous_marker_token_ids, which the plugin produced with its own tokenizer; _encode_text returns [] when the runtime tokenizer has no encode, and any divergence only surfaces as the RuntimeError below. Carrying the marker ids in the stage0_window plan (or reading them from runtime_config) keeps one source of truth for both sides.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 6d39337.

The plan now carries previous_marker_token_ids, taken from duplex_window_previous_marker_token_ids (the ids the plugin produced and the scheduler sized the previous region with), and _window_replacement_parts embeds and appends those ids. _encode_text stays for the session context prefix and suffix. If a plan ever arrives without the key, the worker rebuilds without the marker and the existing replacement_prompt_len check raises instead of silently shifting positions.

token_budget += 1
if final and duplex_payload_is_exact_chunks(payload):
token_budget += 12
# Serving already pads the final residual audio. Stage0 does not

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[question] Dropping the final reservation is unconditional, so a default off session's one-unit exact final append now reserves 13 slots instead of 25. I follow the reasoning (Stage0 never emitted the silent unit, so the 12 slots were never filled), but this is the one change in the PR that reaches every MiniCPM-o duplex session regardless of window settings. Can you confirm what happened to those 12 unfilled slots on main (trimmed via num_input_tokens, or prefilled as scheduler-token positions?) and that test_window_buffered_flush[off] passed at this head? If they were prefilled, this is a silent behaviour fix for off worth its own line in the description.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On main those 12 slots were prefilled, not trimmed. MiniCPMO45OmniForConditionalGeneration pads the rebuilt embeddings up to duplex_prompt_len and prepends the pad block (full_req_embeds = torch.cat([pad_embeds, full_req_embeds])), so the runner scheduled [num_computed_tokens, prompt_len) and the 12 pad tokens became real KV positions immediately before the audio. The audio therefore sat at a later RoPE position and the pads were inside the model's attention context - none of it removed afterwards.

That matches your reasoning that Stage0 never emitted the silent unit, so the removal is exact: the scheduler's num_prompt_tokens now equals the worker's embedding count and no padding or truncation happens on an exact-chunk final append.

The off case of test_window_buffered_flush is unaffected by the window code path (see the PR-body note: the L20 host here cannot boot the three-stage server with its prebuilt vLLM wheel, so I could not re-run the module end to end). test_final_exact_chunk_append_reserves_one_unit_not_two now asserts a final exact-chunk append reserves exactly what a non-final append of the same audio does for seq 1/2/3.

@0z5a

0z5a commented Sep 17, 2026

Copy link
Copy Markdown
Contributor Author

Reviewed at f918052d (static read of the diff against main; no fork code executed here). Scope check for the default path first: the scheduler now computes confirmed_end, segment_output_ids and completed_terminator on every streaming update for every AR model, but they are only consumed behind the MiniCPM gate (duplex.data_plane, sliding_window_mode in basic/context, seq > 1), so block allocation, _free_request_blocks and num_computed_tokens on the ungated path are unchanged, and the worker keeps no unit history unless window_enabled is set from the first append's session config. Both of amy's P2s are addressed at this head: history collection is keyed off the session config in _prepare_session_context, and the worker folds every processor chunk of one append into a single unit including the internal closure tokens, matching the scheduler's one-unit-per-accepted-append boundary; the new lifecycle and multi-chunk tests cover both.

Unit accounting reads consistent to me: the scheduler's span [open_start, base_len + 2) deliberately includes the two closure tokens the worker re-injects at the head of the next append, replacement_prompt_len = projected - dropped equals the worker's rebuild of prefix (plus marker and previous in context mode), suffix, retained units and the pending unit, and the worker raises on a length mismatch rather than silently shifting positions, which is the right failure mode for a KV rebuild. The one correctness gap I found is inline: the replacement branch returns before _streaming_update_overflows, so a rebuilt prompt that no longer fits max_model_len reaches the worker instead of finishing the session with context_length_exceeded.

The merge risk is the e2e coverage. The nine new cases in test_minicpmo_4_5_window.py are advanced_model and run only in the merge-only Omni · MiniCPM-o 4.5 Duplex Test step (about 8 minutes on main today against a 50 minute timeout, so wall time is fine), which means their first CI execution happens on main after this merges, and the validation in the description is one L20 with custom memory limits, not the H100 CI environment. Before merging, please attach an H100 run of pytest tests/e2e/online_serving/test_minicpmo_4_5_window.py -m 'advanced_model and cuda' --run-level advanced_model at this head, or have a maintainer run that merge step on the branch; a merge-only test broke main earlier this week (#7704) and we want to avoid a repeat.

Verdict: requesting changes for the overflow bypass and the red pre-commit (both inline); the rest are suggestions and one question about the off reservation change. The general lane is still pending at f918052d. Happy to approve once those two are fixed, the lane passes, and the H100 e2e evidence is on the PR.

Will fix them. H100 e2e run may take some time to settle, since I have no access to h100. Maybe I will rent some h100 to finish this when it's okay.

Volunteers for H100 e2e test are welcome.

The Stage-0 window replacement branch returned before
``_streaming_update_overflows``, so the rebuilt prompt was bounded only by
the client's window settings: in ``context`` mode it is prefix + marker +
``previous`` + up to ``context_max_units`` retained units (a camera unit is
hundreds of tokens) + the new append, and in ``basic`` mode it is
``basic_window_low_tokens`` + the append, with no upper bound beyond
``low < high``. A prompt over ``max_model_len - sample_room`` reached the
worker instead of finishing the session with ``context_length_exceeded``.

The plan already carries ``replacement_prompt_len``, so the guard now takes
that as an explicit projection: a replacement is not the session's current
prompt plus an extension. Two scheduler regressions cover the overflow and
the one-slot-larger replacement that still applies.

Also stop treating a legitimate zero ``_minicpmo45_window_open_start`` (an
empty context prefix) as missing: ``or preserve_len`` substituted the whole
first-append context length, which double-counts the suffix and makes the
worker's rebuild-length check raise on the first replacement. Same for the
recorded ``duplex_window_prefix_tokens``.

Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
The worker re-tokenized "\n\nprevious: " with its own tokenizer while the
scheduler sized the same block from
``duplex_window_previous_marker_token_ids``, which the plugin produced with
the runtime tokenizer. Any divergence only surfaced as the rebuild-length
RuntimeError. The window plan now carries the marker ids the scheduler
counted, so both sides use one source of truth; ``_encode_text`` stays for
the session context prefix and suffix.

Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
The unconditional extra 12 slots for a ``final`` exact-chunk append were
removed because Stage0 never fed the silent second unit. Serving pads the
final residual audio itself, so the reservation must equal one unit per
chunk (``<unit>`` + its audio) plus the closure pair for every unit after
the first. Add the off-mode regression that asserts a final append reserves
exactly what a non-final append of the same audio does.

Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
The overflow regression needs a plan whose replacement length is at the
sample-room boundary, so pin num_sampled_tokens_per_step and the append
length instead of relying on the window defaults. The zero-open_start
regressions now drive the planner with the confirmed output span it reads,
and assert the dropped span and projection that the two fallbacks produce.

Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
@linyueqian linyueqian added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 17, 2026

@linyueqian linyueqian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-reviewed at 06d71be7 (static read of the four follow-up commits; nothing executed). Everything from the first pass is addressed. The replacement branch now runs _streaming_update_overflows with the plan's replacement_prompt_len as the projection before releasing the cache, so a rebuilt window that no longer leaves room to sample finishes the session with context_length_exceeded instead of reaching the worker, and the two new scheduler tests pin both sides of that boundary. The open_start and duplex_window_prefix_tokens fallbacks are explicit is None checks, with a test for a legitimate zero. The previous: marker ids travel in the stage0_window plan and the worker embeds those instead of re-tokenizing, so the length the scheduler computed and the sequence the worker builds come from one source. ruff format is clean on every touched file, and the final-append reservation change now has a test stating the one-unit intent for off as well as the window modes.

What this approval does not cover is the merge-only e2e: test_minicpmo_4_5_window.py still runs for the first time on main after the merge, and the author has no H100 access. I would like a maintainer with H100 access to run the Omni · MiniCPM-o 4.5 Duplex Test merge step on this branch (or an equivalent run of the new module) before this is merged, and I will hold the merge until that run or an explicit maintainer decision to take the risk. Approving the content.

@0z5a

0z5a commented Sep 18, 2026

Copy link
Copy Markdown
Contributor Author

H100 run of the merge-only step, at 06d71be7:

pytest tests/e2e/online_serving/test_minicpmo_4_5_window.py \
  -m 'advanced_model and cuda' --run-level advanced_model
================== 9 passed, 15 warnings in 981.35s (0:16:21) ==================
  • H100 80GB HBM3 (driver 610.57.04), one stage-0 replica, BF16 real checkpoint openbmb/MiniCPM-o-4_5 (snapshot 503e7542), local checkpoint via MODEL_PREFIX.
  • vLLM 0.29.0 / vLLM-Omni 0.1.dev1+g06d71be73 (commit 06d71be7 on this branch), torch 2.13.0+cu130. Exactly the test command from .buildkite/cuda/test-merge.yml, no -k filter and no custom Stage-0 memory overrides.
  • All nine cases passed: test_window_rebuild_and_next_session (basic, context, context+reference audio), test_window_continuous_input (basic, context, context+camera), test_window_buffered_flush (off, basic, context). The off case matches the reservation change in plugin.py.
  • Environment only: step_audio2 (provides the cosyvoice2 classes Stage 2 registers), onnxruntime, s3tokenizer, hf_xet were installed beyond requirements/common.txt; no source changes were needed to boot or run.

This is the H100 evidence @linyueqian asked for; the CPU set is 152 passed on the same commit.

@BeatSeat

Copy link
Copy Markdown
Contributor

After the first 24 units, each time we will do a re-prefilling. How much impact does that have on the latency?

@0z5a

0z5a commented Sep 18, 2026

Copy link
Copy Markdown
Contributor Author

After the first 24 units, each time we will do a re-prefilling. How much impact does that have on the latency?

Yes, for context mode that's correct.

For audio-only input, the retained 24 units are relatively small (roughly 12 Stage-0 positions per unit, plus the preserved prefix/suffix and up to 500 tokens of compacted previous: text), but camera units can make the rebuilt prompt substantially larger. So the actual cost should be measured against replacement_prompt_len, especially for audio+video.

Also volunteers are welcome for this h100 test, self-funded H100 is not very affordable to me. Maybe I will have another benchmark when it's ok.

@BeatSeat

Copy link
Copy Markdown
Contributor

After the first 24 units, each time we will do a re-prefilling. How much impact does that have on the latency?

Yes, for context mode that's correct.

For audio-only input, the retained 24 units are relatively small (roughly 12 Stage-0 positions per unit, plus the preserved prefix/suffix and up to 500 tokens of compacted previous: text), but camera units can make the rebuilt prompt substantially larger. So the actual cost should be measured against replacement_prompt_len, especially for audio+video.

Also volunteers are welcome for this h100 test, self-funded H100 is not very affordable to me. Maybe I will have another benchmark when it's ok.

Have you already solved the CI issues? The CI still fails, due to some out-of-the-scope issue, you may rebase to the latest main first

@0z5a

0z5a commented Sep 18, 2026

Copy link
Copy Markdown
Contributor Author

After the first 24 units, each time we will do a re-prefilling. How much impact does that have on the latency?

Yes, for context mode that's correct.
For audio-only input, the retained 24 units are relatively small (roughly 12 Stage-0 positions per unit, plus the preserved prefix/suffix and up to 500 tokens of compacted previous: text), but camera units can make the rebuilt prompt substantially larger. So the actual cost should be measured against replacement_prompt_len, especially for audio+video.
Also volunteers are welcome for this h100 test, self-funded H100 is not very affordable to me. Maybe I will have another benchmark when it's ok.

Have you already solved the CI issues? The CI still fails, due to some out-of-the-scope issue, you may rebase to the latest main first

will check it soon

Signed-off-by: 0z5a <dezhen.lu@uni-tuebingen.de>
@0z5a

0z5a commented Sep 18, 2026

Copy link
Copy Markdown
Contributor Author

Merged main and reran the core/MiniCPM-o CPU selection on the SSH host: 853 passed, 1 skipped (Ruff check/format passed). This includes main's switch from realtime to chat TTS benchmarking and the CUDA-only MAGI2 test guard, matching the old NVIDIA/AMD failures. The old Intel run failed during checkout with No space left on device; the other AMD job timed out during collection. Fresh CI is pending; those infrastructure failures are not reported as fixed locally.

@linyueqian

Copy link
Copy Markdown
Collaborator

Ran the whole merge-only "Omni · MiniCPM-o 4.5 Duplex Test" step at a792ef07 on H20 cards (96 GB, the same three-stage single-GPU deploy YAML CI uses), with a fresh vLLM 0.29.0 environment built the way docker/Dockerfile.ci builds it and the exact commands from .buildkite/cuda/test-merge.yml. The new tests/e2e/online_serving/test_minicpmo_4_5_window.py gives 9 passed, 0 failed in 16 min 18 s: all three parametrizations each of rebuild-and-next-session, continuous input (24 repeats, over two minutes of audio, with and without camera frames) and buffered flush (off, basic, context). The existing tests/e2e/online_serving/test_minicpmo_4_5_duplex.py on the same environment gives 11 passed, 1 deselected, 0 failed in 5 min 48 s, so the scheduler changes do not regress the current duplex behaviour. Together with the author's H100 run at 06d71be7 this is the second independent pass of the new step; from my side the only remaining blocker is a green general lane at this head.

@linyueqian
linyueqian merged commit 8bc0c38 into vllm-project:main Sep 22, 2026
6 of 9 checks passed
@linyueqian

Copy link
Copy Markdown
Collaborator

@0z5a heads-up on main after the merge: the merge-only Omni · MiniCPM-o 4.5 Duplex Test step now fails 13 of 20 cases on every build that carries this PR (15824, 15827, 15828). Four of those failures predate us and come from #7633 (see the note there), but the other nine are the window suite, and the first of them kills stage 2: during test_window_rebuild_and_next_session[basic-False] the Code2Wav engine core dies with RuntimeError: MiniCPMO45Code2WavBatchError {"cache_epoch": 1, "chunk_seq": 1, "reason": "new_epoch_requires_first_chunk"} (minicpmo_4_5_code2wav.py, the guard that a new cache epoch must start at chunk_seq 0), and every later window case then fails with stage 2 has no live replica. The same nine cases passed 9/9 on an H20 at a792ef07 against the 09-18 main this branch was based on, so this is an interaction with what landed in between: on the MiniCPM-o path that is #7633 (duplex PCM reservation and input changes), #7781 and #7166, on top of your llm2tts alignment for the native duplex handoff. It looks like the window rebuild advances the Talker's cache epoch while the first chunk of the new epoch never reaches Code2Wav (or arrives as chunk_seq 1). Could you reproduce against current main today? If a fix is not quick we should revert this on main and re-land it once the step is green, since a stage-2 crash takes the whole duplex step down for everyone.

@0z5a
0z5a deleted the stage0-sliding-window branch September 22, 2026 09:41
mlaneuville pushed a commit to mlaneuville/vllm-omni that referenced this pull request Sep 22, 2026
…t#7631)

Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
Signed-off-by: 0z5a <dezhen.lu@uni-tuebingen.de>
Co-authored-by: 0z5a <dezhen.lu@uni-tuebingen.de>
Signed-off-by: Matthieu Laneuville <matthieu.laneuville@surf.nl>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
…t#7631)

Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
Signed-off-by: 0z5a <dezhen.lu@uni-tuebingen.de>
Co-authored-by: 0z5a <dezhen.lu@uni-tuebingen.de>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working core related to core module: cache, scheduler, engine, worker, modelrunner ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: MiniCPM-o 4.5 native full-duplex: long video input drifts, drops APM cache, and breaks the Realtime session

7 participants