Skip to content

[WIP][Bugfix][MiniCPM-o] Stop pre-warm state from dropping the admitting codec chunk - #7995

Closed
chickeyton wants to merge 1 commit into
vllm-project:mainfrom
chickeyton:fix_7978_7979
Closed

chickeyton wants to merge 1 commit into
vllm-project:mainfrom
chickeyton:fix_7978_7979

Conversation

@chickeyton

@chickeyton chickeyton commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Fixes #7979 and #7978.

#7978 and #7979 share the same root cause. They were filed separately because they raise different MiniCPMO45Code2WavBatchError reasons, but both are the same lost payload: the codec chunk that admits a Stage-2 request is dropped before the vocoder ever sees its producer metadata. new_epoch_requires_first_chunk is what that loss looks like when the dropped chunk is a real codec window; chunk_below_lookahead_window is what it looks like when the dropped chunk is the control-only segment boundary. One change to the model runner removes both. That is why they are fixed together here rather than in two PRs.

Both are merge-CI failures of Omni · MiniCPM-o 4.5 Duplex Test that kill the Stage-2 engine core and take every duplex session on that replica with it (stage 2 has no live replica).

Root cause (one bug, two signatures)

OmniGPUModelRunner._update_additional_information skips a newly scheduled request's additional_information whenever that request also carries a non-empty model_intermediate_buffer and the model opts into replace semantics:

if isinstance(model_buffer, dict) and model_buffer:
    update_buffer(new_req.req_id, model_buffer)
    if replace:
        continue          # <- the chunk that admitted this request is dropped

The guard was added in #6406 so a fresh boundary marker is not clobbered by the stale chunk it supersedes. It also fires, however, when the buffer holds nothing but the request's own setup state.

MiniCPM-o 4.5 Code2Wav is the only model with replace_runtime_additional_information = True. Async-chunk pre-warm gives its Stage-2 requests a setup buffer holding duplex and global_request_id, and the codec chunk that admits such a request arrives in additional_information. So the first chunk of every Talker generation reaches the vocoder with all producer metadata stripped, and _parse_item substitutes defaults.

Instrumented run of the failing CI test at c7cd37b6, probing both sources on every new request:

NEWREQDIAG buffer_keys=['duplex', 'global_request_id', 'request_id']
           payload_keys=['meta', 'request_id']
           payload_meta=['cache_epoch', 'chunk_seq', 'code_flat_numel', 'codec_chunk_frames',
                         'codec_left_context_frames', 'duplex_epoch', 'duplex_turn_id',
                         'is_segment_finished', 'last_chunk', 'left_context_size',
                         'llm_output_text_utf8', 'replace_runtime_additional_information',
                         'req_id', 'request_id', 'stream_finished', 'tts_is_last_chunk', 'turn_end']

The chunk is present and is thrown away. That single loss produces both reported failures, depending only on which chunk happens to admit the request:

The second signature reproduces directly against unpatched main: feed Code2Wav one boundary placeholder whose producer meta was stripped and it raises that exact JSON.

Changes

vllm_omni/worker/gpu_model_runner.py — decide by what the buffer contains, not by which source it is. A buffer carrying an OmniPayload section (meta, codes, ids, embed, hidden_states, latent) is a producer snapshot and still wins, which keeps #6406's boundary-marker behaviour. A buffer holding only setup state no longer shadows the chunk that admitted the request. OmniNPUModelRunner inherits this.

vllm_omni/model_executor/models/minicpmo_4_5/minicpmo_4_5_code2wav.py — two containment changes, so a missing chunk can never again take the stage down:

  • A stream-position gap resets or skips the affected request instead of raising. The Talker-to-Code2Wav transport has documented silent-drop paths, so a forward gap is reachable, and the response to it must not be an exception that kills every session on the replica. A backward position is still refused, because replaying it would advance the codec cache twice; the request's step is dropped instead.
  • A non-final window narrower than the encoder's pre-lookahead kernel is withheld and rides on the next window rather than being vocoded alone. The frames are re-chunked, not lost, and the final window is zero-padded by the encoder.

batched_token2wav.py only renames _pre_lookahead_len to pre_lookahead_len so the model layer can read it.

Test Plan

Hardware: L20X, vLLM 0.29.0, openbmb/MiniCPM-o-4_5, vllm_omni/deploy/minicpmo_4_5.yaml. Baseline and patched runs use the merge-CI command:

pytest -s -v \
  tests/e2e/online_serving/test_minicpmo_4_5_duplex.py \
  tests/e2e/online_serving/test_minicpmo_4_5_window.py \
  -m 'advanced_model and cuda' --run-level 'advanced_model'

Unit: tests/worker/, tests/model_executor/models/minicpmo_4_5/, tests/model_executor/stage_input_processors/, tests/engine/duplex/. Lint: ruff check and ruff format --check.

Test Result

End-to-end, same command, same machine, same deploy config:

Tree Result Code2Wav fatals stage 2 has no live replica
main (c7cd37b6) 13 failed, 7 passed, 25:33 new_epoch_requires_first_chunk kills Stage 2 33
this PR 20 passed, 0 failed, 28:24 none none

The baseline reproduces merge build #15824 exactly: same error JSON, same 13/7 split, same cascade into DuplexSessionClosedError: expired: request_cleanup and stage 2 has no live replica for the following cases.

With the patch, Code2Wav logs no stream-position recovery warning at all: the runner change removes the desync rather than papering over it, and the two Code2Wav changes stay unexercised as the safety net they are.

The four test_minicpmo_4_5_duplex.py cases tracked separately in #7962 (TimeoutError: Timed out waiting for duplex CI speech response.created, plus the resume assertion) also pass here. That follows from the same cause: the stripped chunk also loses duplex_turn_id, turn_end and tts_is_last_chunk, so the first chunk of each turn could not be attributed to its response.

Unit tests, CUDA_VISIBLE_DEVICES="":

pytest -q tests/worker/ tests/model_executor/models/minicpmo_4_5/           tests/model_executor/stage_input_processors/ tests/engine/duplex/
6 failed, 1825 passed, 16 skipped, 3 subtests passed in 110.12s

The same 6 fail identically on main at c7cd37b6 (4 in test_batched_omni_output.py, 2 in test_cuda_graph_wrapper.py); they need a CUDA device and this run had none. Nothing else regressed.

Every new test fails on unpatched production code and passes with the fix. Running the 7 new cases with vllm_omni/ reverted to c7cd37b6 and only tests/ from this branch gives 7 failed; with the fix applied, 7 passed.

pre-commit run --from-ref c7cd37b6 --to-ref HEAD passes end to end, including ruff check, ruff format, mypy-3.10, SPDX, forbidden imports, the torch.cuda ratchet, and the test-marks hook.

New regression coverage:

  • test_new_request_setup_buffer_does_not_shadow_its_admitting_chunk and ..._a_boundary_placeholder pin the runner contract for both signatures.
  • test_streaming_new_request_marker_replaces_terminal_chunk_snapshot is unchanged and still passes, so [Bugfix][MiniCPM-o] Fix async-chunk snapshot replacement and prompt cleanup #6406's behaviour is preserved.
  • Code2Wav gains cases for a new epoch mid-stream, a chunk with no cached state, a skipped chunk_seq, a replayed chunk, a retired epoch, short-window withholding and its flush, and a boundary placeholder whose metadata was stripped.

🤖 Generated with Claude Code

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/ar_runtime.md.

Module owners: @tzhouam @fake0fan @Gaohan123

Routing: @tzhouam via module of the changed files, CODEOWNERS; @fake0fan via module of the changed files; @Gaohan123 via module of the changed files

@chickeyton, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 22, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit 6755cd547438 produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

@chickeyton

Copy link
Copy Markdown
Contributor Author

Pre-check report

Self-review of fix_7978_7979 against c7cd37b6 (current upstream/main tip). Full mode, bug-fix type. No blocking issues remain.

Dimension Result
PR title format ✓ [Bugfix] prefix, no WIP
Code quality ⚠ 2 warnings, 0 blocking
Examples policy ✓ no new Python files under examples/
Simplification ⚠ 1 warning
PR description integrity ✓ after correction, see below
Bug-fix checklist ✓
Local pre-commit gates ✓ full run passes
Dead code / import hygiene ✓
Branch rebased ✓ 0 commits behind upstream/main

Verdict: 0 blocking, 3 warnings.

Blocking issues found and fixed

Three ✗ were found on the first pass and are fixed in the current head:

  1. mypy-3.10 failure. A test spy used list.append's return value inside a tuple expression (func-returns-value at test_code2wav_batching.py:1447). Replaced with a named function.
  2. Mixed line endings. Five files carried CRLF from a Windows working tree; the mixed-line-ending hook rewrote them and they are now committed normalized.
  3. PR description integrity. The unit-test numbers in the description came from a narrower selection than the command shown next to them. The description now carries the exact command, the exact summary line, and an explicit statement that the 6 failures are pre-existing on main.

Remaining warnings, with rationale

Code quality — one new .clone(). minicpmo_4_5_code2wav.py copies withheld codec frames once per withheld window. The copy is required for correctness, not convenience: those frames outlive the step, and the ids they are sliced from belong to the runner's input buffer, which it reuses. A comment at the call site says so. This is a handful of int64 tokens, not a latent, and the path is off the steady-state route.

Code quality — five new SimpleNamespace lines in tests. They fake vLLM's SchedulerOutput and NewRequestData, matching how every surrounding case in test_omni_gpu_model_runner.py already constructs them. Diverging in two new cases would have been the more surprising choice.

Simplification — the two Code2Wav changes are defence in depth. Once the model-runner change lands, neither is needed to make the reported failures go away: the final end-to-end run logs no stream-position recovery warning at all, because the desync no longer happens. They are here because both issues ask for the stage to survive a bad chunk rather than take every session on the replica down with it, and because the transport has documented silent-drop paths that can produce the same gap from a different direction. Both paths are reachable and covered by tests. A reviewer who would rather land the one-line ownership fix alone can say so and I will split them out.

What was verified

Every new test fails without the fix. Reverting vllm_omni/ to c7cd37b6 while keeping tests/ from this branch turns all 7 new cases red; with the fix they pass. That is the property that matters for a regression test, and it is checked rather than assumed.

Pre-commit passes end to end, not skipped: ruff check, ruff format, mypy-3.10, typos, SPDX headers, forbidden imports, the torch.cuda ratchet, and the test CI-marks hook. The diff adds no **kwargs string plumbing, no broad except, no Any annotations, no event-loop blocking, no new dependency, and no new file.

The end-to-end evidence is in the PR description: 13 failed / 7 passed on main, 20 passed / 0 failed on this branch, same command and deploy config on the same machine, with the baseline matching merge build 15824 exactly.

@twu3202

twu3202 commented Sep 22, 2026

Copy link
Copy Markdown
Contributor

#7992 just merged and closed #7979 along with #7962, so here's how I think the two fit. The setup buffer in your NEWREQDIAG (duplex, global_request_id, request_id) is the model_intermediate_buffer from Stage0's append prompt, which the pre-warm had been copying into the Stage-2 placeholder since #7633. #7992 stops that copy, so the placeholder has no buffer and the if replace: continue guard never fires. That's how it got to 20/20 without touching the runner.

The runner check and the Code2Wav containment here still look worth having, so a stray setup buffer or a dropped chunk can't take the stage down again. I haven't run this on top of #7992. Your trace also fills in the part I could only guess at on #7978, thanks.

…odec chunk

`OmniGPUModelRunner._update_additional_information` skipped a newly
scheduled request's `additional_information` whenever the request also
carried a non-empty `model_intermediate_buffer` and the model opted into
replace semantics. That guard exists so a fresh boundary marker is not
clobbered by the stale chunk it supersedes (vllm-project#6406), but it also fires when
the buffer is only the request's own setup state.

MiniCPM-o 4.5 Code2Wav is the one model with
`replace_runtime_additional_information = True`. Async-chunk pre-warm gives
its Stage-2 requests a setup buffer holding `duplex` and
`global_request_id`, and every codec chunk that admits such a request
travels in `additional_information`. So the first chunk of every Talker
generation reached the vocoder with all producer metadata stripped and
`_parse_item` substituted defaults, which killed the Stage-2 engine two
different ways:

* A real codec window lost `cache_epoch` and `chunk_seq`, so the vocoder
  recorded epoch 0 / chunk 0 for what was really epoch 1 / chunk 0. The next
  chunk arrived correctly labelled and raised
  `new_epoch_requires_first_chunk` (vllm-project#7979).
* The control-only segment boundary is one placeholder token plus
  `code_flat_numel == 0`. Losing that zero made the placeholder look like
  codec data, so a one-frame non-final window raised
  `chunk_below_lookahead_window` (vllm-project#7978).

Decide by what the buffer contains: a buffer carrying an `OmniPayload`
section is a producer snapshot and still wins, while a buffer holding only
setup state no longer shadows the chunk.

Two Code2Wav changes keep a stage-wide failure off the table when a chunk
does go missing anyway. A stream-position gap now resets or skips the
affected request instead of raising, since the transport can legitimately
lose a payload and killing the engine drops every session on the replica; a
backward position is still refused, because replaying it would advance the
codec cache twice. A non-final window narrower than the encoder's
pre-lookahead kernel is withheld and rides on the next window rather than
being vocoded alone.

Signed-off-by: chickeyton <ngton2014@gmail.com>

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
prompt_wav=item.prompt_wav,
token2wav=item.previous.token2wav,
)
for item in held_items

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Track the stream position even when the first window is withheld

When the first window of a request or a new cache epoch is shorter than the lookahead minimum, item.previous is None, so this condition skips recording its epoch and chunk sequence even though _held_tokens retains its frames. This breaks the recovery logic: replaying the initial chunk appends the same frames again; switching epochs before any decode mixes the old held frames into the new epoch; and, when an older epoch already has cached state, the next chunk of the new epoch triggers the reset again and discards that new epoch's held frames.

Please retain the epoch and sequence for withheld windows independently of whether a decoder cache has been initialized, and use that position for replay rejection and epoch transitions. Add regression coverage where the first window of a request or epoch is short; the existing withholding tests establish a decoder cache before withholding.

@Gaohan123 Gaohan123 added this to the v0.30.0 milestone Sep 22, 2026
_OMNI_PAYLOAD_SECTIONS = frozenset({"meta", "codes", "ids", "embed", "hidden_states", "latent"})


def _is_replacement_snapshot(buffer: dict) -> bool:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please move it to utils

@chickeyton chickeyton changed the title [Bugfix][MiniCPM-o] Stop pre-warm state from dropping the admitting codec chunk [WIP][Bugfix][MiniCPM-o] Stop pre-warm state from dropping the admitting codec chunk Sep 22, 2026
@chickeyton

Copy link
Copy Markdown
Contributor Author

#7992 just merged and closed #7979 along with #7962, so here's how I think the two fit. The setup buffer in your NEWREQDIAG (duplex, global_request_id, request_id) is the model_intermediate_buffer from Stage0's append prompt, which the pre-warm had been copying into the Stage-2 placeholder since #7633. #7992 stops that copy, so the placeholder has no buffer and the if replace: continue guard never fires. That's how it got to 20/20 without touching the runner.

The runner check and the Code2Wav containment here still look worth having, so a stray setup buffer or a dropped chunk can't take the stage down again. I haven't run this on top of #7992. Your trace also fills in the part I could only guess at on #7978, thanks.

I am testing it again, maybe some of the changes is not needed anymore

@chickeyton

chickeyton commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor Author

Update after rebasing onto f7e28348: this PR is no longer needed for #7978 / #7979

#7992 ("Keep resumable duplex prompts out of async-chunk prewarm") landed on main while this PR was open and fixes both issues on its own. I re-verified rather than assumed, and I am posting the result here rather than quietly dropping the branch.

Evidence

Two end-to-end runs, launched in parallel on separate GPUs of the same host, same command and same deploy config as merge CI:

pytest -s -v \
  tests/e2e/online_serving/test_minicpmo_4_5_duplex.py \
  tests/e2e/online_serving/test_minicpmo_4_5_window.py \
  -m 'advanced_model and cuda' --run-level 'advanced_model'
Tree Result Code2Wav fatals stage 2 has no live replica
main @ c7cd37b6 (before #7992) 13 failed, 7 passed, 25:33 new_epoch_requires_first_chunk kills Stage 2 33
main @ f7e28348, without this PR 20 passed, 0 failed, 35:28 none none
main @ f7e28348, with this PR 20 passed, 0 failed, 34:11 none none

test_window_rebuild_and_next_session, the case that produced the #7979 crash, passes on plain f7e28348. This PR changes nothing on the current main.

Why #7992 fixes it

It is the only production change among the four commits added since c7cd37b6 that touches this path. It stops Stage-0's prompt being stored on the request state for resumable submissions, and that prompt is exactly what _prewarm_async_chunk_stages copied into the Stage-2 request via copy_request_snapshot(original_prompt). So the pre-warm buffer holding duplex and global_request_id no longer exists.

That is the same defect this PR describes, removed from the producer side instead of the consumer side. With an empty buffer, _update_additional_information falls through to the chunk's own additional_information and no producer metadata is stripped, so neither new_epoch_requires_first_chunk nor chunk_below_lookahead_window can be reached that way.

One thing that did not change

The ownership rule in OmniGPUModelRunner._update_additional_information is untouched by #7992. Running this PR's two runner regression tests against f7e28348's production code still gives:

FAILED tests/worker/test_omni_gpu_model_runner.py::test_new_request_setup_buffer_does_not_shadow_its_admitting_chunk
FAILED tests/worker/test_omni_gpu_model_runner.py::test_new_request_setup_buffer_does_not_shadow_a_boundary_placeholder
2 failed

A newly scheduled request's payload is still discarded whenever the request also carries a setup-only model_intermediate_buffer and the model uses replace semantics. #7992 removes the one input shape that reached it in this pipeline; it does not change the rule. The bug is latent rather than fixed, and any future path that puts a non-payload buffer on a new async-chunk request reintroduces the same silent metadata loss.

@chickeyton

chickeyton commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor Author

@Gaohan123 after testing it again, I found that issue #7978 and #7979 are fixed by PR #7992 already, you may consider closing issue #7978 and this PR

@Gaohan123

Copy link
Copy Markdown
Collaborator

Thanks for your contribution

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

4 participants