Skip to content

[Bugfix] Fix Qwen3-TTS prefill probe on MRv2 runner - #8065

Merged
Gaohan123 merged 2 commits into
vllm-project:mainfrom
Sy0307:fix/8049-prefill-probe
Sep 23, 2026
Merged

Gaohan123 merged 2 commits into
vllm-project:mainfrom
Sy0307:fix/8049-prefill-probe

Conversation

@Sy0307

@Sy0307 Sy0307 commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose

Fixes #8049.

The Qwen3-TTS chunked-prefill phase-contract probe still accessed the legacy runner's _talker_mtp_forward. The default MRv2 OmniARModelRunner runs MTP through model_state._run_batched_mtp, so the probe failed during initialization before checking the audio result.

Update the test-only worker extension to observe the active runner's preprocess and MTP paths. Preserve the legacy path, and verify on MRv2 that prefill rows do not enter MTP and that the one-token prefill tail embedding reaches the next stage unchanged.

The first two Read the Docs builds for this PR failed because three external Read the Docs inventories returned HTTP 429. Disable those optional inventories in mkdocs.yml so strict documentation builds no longer depend on rate-limited hosts.

Test Plan

vLLM Version: 0.30.0

vLLM-Omni Commit: 6a14af3be (based on 57d3e9fdf82c06094ee0c0b39c79d24b247fe8c3)

  • H200 physical GPU 2: tests/e2e/features/prefill_phase/test_chunked_prefill_tail.py::test_one_token_prefill_tail_matches_unchunked_audio with -m 'advanced_model and cuda' --run-level advanced_model.
  • pre-commit run --files mkdocs.yml tests/e2e/features/prefill_phase/worker_extension.py
  • git diff --check origin/main...HEAD

Test Result

  • Targeted H200 regression: 1 passed. Full and 20+1 chunked prefill each produced 46 codec frames and 88,320 audio samples; maximum audio difference was 0.
  • Changed-file pre-commit and whitespace checks: passed.
  • Wider local Qwen3-TTS nightly functional run: 13 passed, 2 audio-similarity checks failed (Base and CustomVoice no_async_chunk). Both failures generated audio and persisted on isolated rerun; these tests do not use the changed probe. Two-GPU TTS performance cases: 8 passed.
  • A local all-files pre-commit scan hit unrelated existing Markdown and mypy errors; it did not leave changes in this branch. CI will provide the repository check for this PR.

Signed-off-by: Sy03 <1370724210@qq.com>
@Sy0307 Sy0307 added the merge-test label to trigger buildkite merge test CI label Sep 23, 2026
@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to be related to model: qwen-tts.

Model owners: @FayeSpica @NickCao @yenuo26

Routing: @FayeSpica via semantic router; @NickCao via CODEOWNERS; @yenuo26 via CODEOWNERS

@Sy0307, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 23, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit 6a14af3be632 produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: CI is red on this head

@Sy0307 required checks failed on 943d8393aed0:

Please fix the failure and push again; this note is updated in place when the head goes green or moves.

@Sy0307 Sy0307 closed this Sep 23, 2026
@Sy0307 Sy0307 reopened this Sep 23, 2026
Signed-off-by: Sy03 <1370724210@qq.com>

@Gaohan123 Gaohan123 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Thanks

@Gaohan123 Gaohan123 added ready label to trigger buildkite CI and removed merge-test label to trigger buildkite merge test CI labels Sep 23, 2026
@Gaohan123
Gaohan123 enabled auto-merge (squash) September 23, 2026 09:43
@Gaohan123
Gaohan123 merged commit 8953d08 into vllm-project:main Sep 23, 2026
7 of 9 checks passed
RyanYun09 added a commit to RyanYun09/vllm-omni that referenced this pull request Sep 23, 2026
Read the Docs builds (--strict via fail_on_warning: true) intermittently
fail while mkdocstrings downloads external objects.inv from typing-extensions,
pillow, and psutil: the RTD IP range gets rate-limited / the connection resets,
which surfaces as a non-zero exit from 'python -m mkdocs build --strict'.

Mirror the upstream fix (vllm-project/vllm-omni PR vllm-project#8065) that comments these
three inventories out, keeping only docs.python.org and numpy, which are
reliable. Local 'mkdocs build --clean --site-dir ... --strict' passes with
zero warnings.

Signed-off-by: RyanYun09 <318555231+RyanYun09@users.noreply.github.com>
LinzeShi added a commit to LinzeShi/vllm-omni that referenced this pull request Sep 23, 2026
Apply the mkdocs.yml change from upstream commit 8953d08 (vllm-project#8065). Disable the three optional inventories returning HTTP 429 while retaining strict documentation checks.

Signed-off-by: Linze-Shi <linzeshi0@gmail.com>
abtonmoy added a commit to abtonmoy/vllm-omni that referenced this pull request Oct 4, 2026
Syncs the branch with main so Read the Docs can build this head. The RTD
check has been red on an 09-13 build that failed at `git checkout` with
"reference is not a tree", because upstream `refs/pull/4831/head` is still
at the pre-09-13 head and the commit was absent from what RTD fetched.

This also picks up vllm-project#8065, which commented out the typing-extensions,
pillow and psutil mkdocstrings inventories because Read the Docs
rate-limits them in strict builds. This branch still had all three
enabled under `fail_on_warning: true`.

No change to the encode-parallel delta itself.

Signed-off-by: Abdul Basit Tonmoy <abdulbasittonmoy@gmail.com>
psv666 added a commit to psv666/vllm-omni that referenced this pull request Oct 8, 2026
Explicitly disable derived TPOT for Realtime chat, which only observes
client response timings. Add regression coverage for explicit and VAD
turns, including tokenizer fallback and TPOT goodput rejection.

List the benchmark wrapper, dataset, and duplex client in the nightly
job dependencies. Align the external inventory configuration with the
ReadTheDocs HTTP 429 fix already merged in upstream PR vllm-project#8065.

Update the timing comments and benchmark documentation to match the
actual metric calculation.

Signed-off-by: psv666 <2693925048@qq.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Merge CI, Qwen3-TTS CustomVoice - chunked-prefill probe fails because OmniARModelRunner lacks _talker_mtp_forward

3 participants