Skip to content

[Bugfix][MiniCPM-o] Fix async-chunk snapshot replacement and prompt cleanup - #6406

Merged
amy-why-3459 merged 9 commits into
vllm-project:mainfrom
natureofnature:fix/minicpmo-async-chunk-rollover-20260820
Aug 21, 2026
Merged

amy-why-3459 merged 9 commits into
vllm-project:mainfrom
natureofnature:fix/minicpmo-async-chunk-rollover-20260820

Conversation

@natureofnature

@natureofnature natureofnature commented Aug 20, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose

This is the second correctness-only split from #5102, following the resumable cleanup fix in #6360. It fixes MiniCPM-o 4.5 async-chunk snapshot replacement and explicit prompt-replacement cleanup without including the OmniInteract benchmark client, Nightly workflow, production YAML changes, final-input lifecycle work, force-listen policy, or automatic context-capacity rollover.

What broke

  1. Code2Wav runtime metadata was merged across chunks. A previous terminal audio payload could remain in the model buffer and be processed again by a later control-only boundary.
  2. When a producer explicitly replaced a streaming prompt at a new assistant turn, the scheduler did not release the old KV/encoder state or clear the connector watermark. The replacement could therefore inherit cache ownership and chunk progress from the previous prompt.

Reproduction

The targeted tests exercise both failures directly:

  • send a terminal Code2Wav snapshot followed by an empty/control-only replacement and verify that the old audio is not replayed;
  • explicitly replace a running or computed Talker prompt and verify re-admission, one-time KV/encoder release, stale-output fencing, and connector-watermark reset.

Root cause

The Code2Wav path had only incremental merge semantics for runtime additional information. Separately, an explicit async prompt replacement updated the request payload without first going through the scheduler's cache-release and connector-reset path.

Fix

  • Add an explicit producer marker and a MiniCPM Code2Wav model opt-in for snapshot replacement. Unmarked generation/diffusion payloads retain their existing merge behavior.
  • Emit a scheduler-ready empty replacement snapshot for duplicate terminal boundaries so old terminal audio cannot be replayed, while retaining sibling fields such as codes.ref when 1-D audio is represented by prompt placeholders.
  • Requeue an explicitly replaced running prompt through normal scheduler admission, release its old KV/encoder state exactly once, mark in-flight output stale once, and reset the connector watermark.
  • Pin the NPU runner's inherited snapshot-replacement contract with an Ascend test; OmniNPUModelRunner already derives from OmniGPUModelRunner and explicitly delegates Omni state updates to it.

Scope decision

This revision deliberately removes the automatic Talker context-capacity rollover that was previously in this PR. We do not yet have a targeted long single-Talker-turn E2E showing that automatic replacement is required, and replacing an accumulated prompt is a model-semantics decision rather than a safe transport-only fix. If the long-video E2E demonstrates a real context-limit failure, that policy can return in a separate PR with model-quality and continuity evidence. Explicit turn-boundary replacement and its cache/watermark cleanup remain in this PR.

Test Plan

vLLM Version: target pin 0.27.0; compatibility unit-test container reports vLLM-Omni 0.25.0

vLLM-Omni Commit: d5e206cd6b4cf5744fff19ef6a5a77d07f1d3d9e

python3 -m pytest -q \
  tests/core/sched/test_omni_ar_scheduler_streaming.py::test_explicit_streaming_payload_replaces_placeholder_prompt \
  tests/core/sched/test_omni_ar_scheduler_streaming.py::test_explicit_model_intermediate_prompt_replacement_releases_cache_and_watermark \
  tests/core/sched/test_omni_ar_scheduler_streaming.py::test_ready_async_chunk_prompt_replacement_releases_stale_kv_once \
  tests/distributed/omni_connectors/test_chunk_transfer_adapter.py::test_load_poll_ar_requeues_explicitly_replaced_running_prompt \
  tests/distributed/omni_connectors/test_chunk_transfer_adapter.py::test_load_poll_generation_segment_marker_replaces_previous_chunk \
  tests/distributed/omni_connectors/test_chunk_transfer_adapter.py::test_load_poll_generation_empty_replacement_snapshot_is_ready \
  tests/distributed/omni_connectors/test_chunk_transfer_adapter.py::test_load_poll_generation_without_snapshot_marker_keeps_incremental_state \
  tests/model_executor/models/minicpmo_4_5/test_code2wav_batching.py::test_empty_final_ignores_generation_scheduler_placeholder_token \
  tests/model_executor/stage_input_processors/test_minicpmo_4_5_async_chunk.py::test_duplex_turn_end_closes_epoch_and_next_turn_restarts_sequence \
  tests/model_executor/stage_input_processors/test_minicpmo_4_5_omni.py::test_native_duplex_continuation_appends_only_new_talker_condition \
  tests/worker/test_omni_gpu_model_runner.py::test_update_additional_information_deserializes_new_request_payload \
  tests/worker/test_omni_gpu_model_runner.py::test_streaming_new_request_marker_replaces_terminal_chunk_snapshot \
  tests/worker/test_omni_gpu_model_runner.py::test_cached_empty_marker_replaces_terminal_chunk_snapshot

python3 -m pytest -q \
  tests/core/sched/test_omni_ar_scheduler_streaming.py \
  tests/distributed/omni_connectors/test_chunk_transfer_adapter.py \
  tests/model_executor/models/minicpmo_4_5/test_code2wav_batching.py \
  tests/model_executor/stage_input_processors/test_minicpmo_4_5_async_chunk.py \
  tests/model_executor/stage_input_processors/test_minicpmo_4_5_omni.py \
  tests/worker/test_omni_gpu_model_runner.py

python3 -m pytest -q tests/model_executor/models/minicpmo_4_5/test_talker_batching.py

git diff --name-only upstream/main...HEAD -- '*.py' | xargs python3 -m ruff check
git diff --name-only upstream/main...HEAD -- '*.py' | xargs python3 -m ruff format --check
git diff --check upstream/main...HEAD

Test Result

  • Targeted regression: 13 passed, 17 warnings in 1.79s
  • Reviewer stale-seed regression: 2 passed, 17 warnings in 0.11s
  • Affected related suite: 181 passed, 17 warnings in 2.92s
  • Talker chunk-budget regression: 24 passed, 17 warnings in 0.33s
  • NPU snapshot inheritance contract: added under tests/platforms/npu/ for Ascend CI
  • Ruff: All checks passed; all changed Python files are formatted
  • compileall and git diff --check: passed

The snapshot replacement path requires both an explicit payload marker and the MiniCPM Code2Wav model opt-in. Explicit prompt replacement is producer-marked; there is no automatic context-limit policy in this PR.

Signed-off-by: natureofnature <wzliu@connect.hku.hk>
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@natureofnature

Copy link
Copy Markdown
Collaborator Author

@R2-Y PTAL

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to be related to model: minicpm.

Model owners: @y-null

@natureofnature, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@amy-why-3459

Copy link
Copy Markdown
Collaborator

Thanks — this is the right second split from #5102. Talker capacity
rollover (min of model/TTS len, generation reserve 26, requeue through
admission, one-shot KV/encoder free + watermark reset) looks correct
and the AR/adapter tests pin it. Snapshot replace as an explicit
producer marker plus the Code2Wav opt-in is the right default so
other generation/diffusion paths keep merging.

Request changes on the Code2Wav replay path.

The duplicate turn_end payload is only

OmniPayloadStruct(
    meta=_MiniCPMO45MetaStruct(replace_runtime_additional_information=True),
    request_id=request_id,
)

finished / is_segment_finished are unset. Generation poll then
increments get_req_chunk and hits

if not has_new_ids and not payload_finished:
    return False

so the empty snapshot is written onto request.additional_information
but never marked ready. The runner buffer is only replaced when that
request is actually scheduled into _update_additional_information.
If Code2Wav is still running and this step is scheduled without a
ready update, forward() still sees the previous terminal
codes.audio — the replay this PR claims to fix.

The producer test only checks the marker exists; the runner test
injects a scheduled_new_reqs snapshot. Please mark the duplicate
boundary ready (is_segment_finished=True, or treat the replace
flag as ready) and add a generation-poll test that an empty
duplicate snapshot becomes ready and leaves no leftover audio in
the runner buffer. Keeping codes.ref under the 1-D placeholder
is fine.

Should-fix:

  • NPU generation runner does not inherit
    _replace_intermediate_buffer; snapshot replace is CUDA-only
    unless you lift it to a shared mixin.
  • Confirm _free_request_blocks exists on the vLLM pin (the PR
    says 0.25.0; recent schedulers have it, older ones only have
    _free_blocks).
  • Prefix-cache / async-scheduling remain follow-up, as noted.

Happy to re-review once the empty snapshot is actually consumed.

Signed-off-by: natureofnature <wzliu@connect.hku.hk>
@natureofnature

natureofnature commented Aug 20, 2026 •

Copy link
Copy Markdown
Collaborator Author

Self-review and reviewer update for c04fc9fc:

  • Fixed the blocking replay path: a duplicate native-duplex turn_end now emits an empty replacement snapshot with is_segment_finished=True, so generation poll marks it ready and the runner consumes it instead of retaining the previous terminal codes.audio.
  • Added regressions for the producer marker, generation-poll readiness, and both scheduled-new and scheduled-cached runner replacement. The cached-request test specifically verifies that the old terminal audio is removed.
  • Confirmed the NPU path shares this implementation: OmniNPUModelRunner derives from OmniGPUModelRunner, and its _update_states() explicitly delegates to OmniGPUModelRunner._update_states(). Added an Ascend contract test that exercises the inherited snapshot replacement.
  • Corrected the documented target vLLM version to 0.27.0 and verified that its scheduler provides _free_request_blocks() with deferred in-flight block release. The 0.25.0 value referred only to the compatibility unit-test container.

Validation after the fix:

  • targeted regression: 13 passed, 16 warnings in 1.92s
  • affected related suite: 149 passed, 21 deselected, 16 warnings in 2.65s

I also rechecked the scope: snapshot replacement remains gated by the explicit producer marker plus the MiniCPM Code2Wav model opt-in; unmarked generation/diffusion models retain merge semantics. @amy-why-3459 PTAL when convenient.

@natureofnature

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: c04fc9fcad

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread vllm_omni/core/sched/omni_ar_scheduler.py
Signed-off-by: natureofnature <wzliu@connect.hku.hk>
@natureofnature

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 4ec3817f3d

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread vllm_omni/core/sched/omni_scheduler_mixin.py Outdated
@hsliuustc0106 hsliuustc0106 added bug Something isn't working omni code related to omni models labels Aug 21, 2026
Signed-off-by: natureofnature <wzliu@connect.hku.hk>
Signed-off-by: natureofnature <wzliu@connect.hku.hk>
@natureofnature natureofnature changed the title [Bugfix][MiniCPM-o] Fix async-chunk prompt rollover [Bugfix][MiniCPM-o] Fix async-chunk snapshot replacement and prompt cleanup Aug 21, 2026
Signed-off-by: natureofnature <wzliu@connect.hku.hk>
@natureofnature

Copy link
Copy Markdown
Collaborator Author

@codex review

@natureofnature

Copy link
Copy Markdown
Collaborator Author

@tzhouam PTAL

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Swish!

Reviewed commit: d5e206cd6b

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@natureofnature

Copy link
Copy Markdown
Collaborator Author

@Gaohan123 PTAL

@amy-why-3459

Copy link
Copy Markdown
Collaborator

Thanks for the follow-up. Rechecked d5e206c.

The previous blocking replay path is fixed: a duplicate native-duplex
turn_end now emits an empty replacement with is_segment_finished=True,
generation poll marks it ready, and both scheduled-new and
scheduled-cached runner updates drop the previous terminal codes.audio.
The generation scheduler already forwards request.additional_information
onto cached reqs, so a still-running Code2Wav step can consume the
snapshot. NPU was my miss — OmniNPUModelRunner inherits
OmniGPUModelRunner and NPUGenerationModelRunner goes through
_preprocess() → _update_additional_information. The 0.27.0
_free_request_blocks pin is also fine.

Approve with comments.

Should-fix / follow-up:

  • _clear_chunk_ready still drops replaced_streaming_prompt_ids only on
    scheduled_new_reqs. Mixin now consumes the marker, so this is leftover
    backup; please also clear it on cached admission or document that mixin
    is the only consumer.
  • _update_request_as_session still does num_stale += in_flight and then
    _replace_streaming_session assigns over it. Harmless after d5e206c,
    but the += is dead on the replace path and will double-count again if
    someone reverts the assign.
  • Automatic Talker capacity rollover is correctly out of this PR. Please
    keep that as a separate [Core][Benchmark]Omniinteract for Minicpm-o4.5 #5102 follow-up if the long-turn E2E actually
    hits TTS max_position_embeddings; this split should not be read as
    having fixed that overflow.

Snapshot replace remaining opt-in (producer marker + Code2Wav flag) is
the right default. Happy to approve.

@amy-why-3459 amy-why-3459 added the ready label to trigger buildkite CI label Aug 21, 2026
@amy-why-3459 amy-why-3459 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Aug 21, 2026
@amy-why-3459
amy-why-3459 merged commit 4f7c08c into vllm-project:main Aug 21, 2026
8 of 9 checks passed
fan2956 pushed a commit to fan2956/vllm-omni that referenced this pull request Aug 23, 2026
…leanup (vllm-project#6406)

Signed-off-by: natureofnature <wzliu@connect.hku.hk>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>
fan2956 pushed a commit to fan2956/vllm-omni that referenced this pull request Aug 23, 2026
…leanup (vllm-project#6406)

Signed-off-by: natureofnature <wzliu@connect.hku.hk>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>
AndyZhou952 pushed a commit to AndyZhou952/vllm-omni that referenced this pull request Aug 26, 2026
…leanup (vllm-project#6406)

Signed-off-by: natureofnature <wzliu@connect.hku.hk>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>
Signed-off-by: AndyZhou952 <jzhoubc@connect.ust.hk>
JoseCarlosGarcia95 pushed a commit to valendra-tech/vllm-omni that referenced this pull request Sep 5, 2026
…leanup (vllm-project#6406)

Signed-off-by: natureofnature <wzliu@connect.hku.hk>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>
chickeyton added a commit to chickeyton/vllm-omni that referenced this pull request Sep 22, 2026
…odec chunk

`OmniGPUModelRunner._update_additional_information` skipped a newly
scheduled request's `additional_information` whenever the request also
carried a non-empty `model_intermediate_buffer` and the model opted into
replace semantics. That guard exists so a fresh boundary marker is not
clobbered by the stale chunk it supersedes (vllm-project#6406), but it also fires when
the buffer is only the request's own setup state.

MiniCPM-o 4.5 Code2Wav is the one model with
`replace_runtime_additional_information = True`. Async-chunk pre-warm gives
its Stage-2 requests a setup buffer holding `duplex` and
`global_request_id`, and every codec chunk that admits such a request
travels in `additional_information`. So the first chunk of every Talker
generation reached the vocoder with all producer metadata stripped and
`_parse_item` substituted defaults, which killed the Stage-2 engine two
different ways:

* A real codec window lost `cache_epoch` and `chunk_seq`, so the vocoder
  recorded epoch 0 / chunk 0 for what was really epoch 1 / chunk 0. The next
  chunk arrived correctly labelled and raised
  `new_epoch_requires_first_chunk` (vllm-project#7979).
* The control-only segment boundary is one placeholder token plus
  `code_flat_numel == 0`. Losing that zero made the placeholder look like
  codec data, so a one-frame non-final window raised
  `chunk_below_lookahead_window` (vllm-project#7978).

Decide by what the buffer contains: a buffer carrying an `OmniPayload`
section is a producer snapshot and still wins, while a buffer holding only
setup state no longer shadows the chunk.

Two Code2Wav changes keep a stage-wide failure off the table when a chunk
does go missing anyway. A stream-position gap now resets or skips the
affected request instead of raising, since the transport can legitimately
lose a payload and killing the engine drops every session on the replica; a
backward position is still refused, because replaying it would advance the
codec cache twice. A non-final window narrower than the encoder's
pre-lookahead kernel is withheld and rides on the next window rather than
being vocoded alone.

Signed-off-by: chickeyton <ngton2014@gmail.com>

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
chickeyton added a commit to chickeyton/vllm-omni that referenced this pull request Sep 22, 2026
…odec chunk

`OmniGPUModelRunner._update_additional_information` skipped a newly
scheduled request's `additional_information` whenever the request also
carried a non-empty `model_intermediate_buffer` and the model opted into
replace semantics. That guard exists so a fresh boundary marker is not
clobbered by the stale chunk it supersedes (vllm-project#6406), but it also fires when
the buffer is only the request's own setup state.

MiniCPM-o 4.5 Code2Wav is the one model with
`replace_runtime_additional_information = True`. Async-chunk pre-warm gives
its Stage-2 requests a setup buffer holding `duplex` and
`global_request_id`, and every codec chunk that admits such a request
travels in `additional_information`. So the first chunk of every Talker
generation reached the vocoder with all producer metadata stripped and
`_parse_item` substituted defaults, which killed the Stage-2 engine two
different ways:

* A real codec window lost `cache_epoch` and `chunk_seq`, so the vocoder
  recorded epoch 0 / chunk 0 for what was really epoch 1 / chunk 0. The next
  chunk arrived correctly labelled and raised
  `new_epoch_requires_first_chunk` (vllm-project#7979).
* The control-only segment boundary is one placeholder token plus
  `code_flat_numel == 0`. Losing that zero made the placeholder look like
  codec data, so a one-frame non-final window raised
  `chunk_below_lookahead_window` (vllm-project#7978).

Decide by what the buffer contains: a buffer carrying an `OmniPayload`
section is a producer snapshot and still wins, while a buffer holding only
setup state no longer shadows the chunk.

Two Code2Wav changes keep a stage-wide failure off the table when a chunk
does go missing anyway. A stream-position gap now resets or skips the
affected request instead of raising, since the transport can legitimately
lose a payload and killing the engine drops every session on the replica; a
backward position is still refused, because replaying it would advance the
codec cache twice. A non-final window narrower than the encoder's
pre-lookahead kernel is withheld and rides on the next window rather than
being vocoded alone.

Signed-off-by: chickeyton <ngton2014@gmail.com>

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
…leanup (vllm-project#6406)

Signed-off-by: natureofnature <wzliu@connect.hku.hk>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working omni code related to omni models ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants