Repository navigation
[Bugfix] Fix Qwen3-TTS prefill probe on MRv2 runner - #8065
Conversation
Signed-off-by: Sy03 <1370724210@qq.com>
|
This PR appears to be related to model: qwen-tts. Model owners: @FayeSpica @NickCao @yenuo26 Routing: @FayeSpica via semantic router; @NickCao via CODEOWNERS; @yenuo26 via CODEOWNERS @Sy0307, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Omni ReviewBot triage noteAutomated triage of commit
These are automated triage suggestions only — the final decision belongs to the maintainers. |
Omni ReviewBot: CI is red on this head@Sy0307 required checks failed on Please fix the failure and push again; this note is updated in place when the head goes green or moves. |
Signed-off-by: Sy03 <1370724210@qq.com>
Read the Docs builds (--strict via fail_on_warning: true) intermittently fail while mkdocstrings downloads external objects.inv from typing-extensions, pillow, and psutil: the RTD IP range gets rate-limited / the connection resets, which surfaces as a non-zero exit from 'python -m mkdocs build --strict'. Mirror the upstream fix (vllm-project/vllm-omni PR vllm-project#8065) that comments these three inventories out, keeping only docs.python.org and numpy, which are reliable. Local 'mkdocs build --clean --site-dir ... --strict' passes with zero warnings. Signed-off-by: RyanYun09 <318555231+RyanYun09@users.noreply.github.com>
Apply the mkdocs.yml change from upstream commit 8953d08 (vllm-project#8065). Disable the three optional inventories returning HTTP 429 while retaining strict documentation checks. Signed-off-by: Linze-Shi <linzeshi0@gmail.com>
Syncs the branch with main so Read the Docs can build this head. The RTD check has been red on an 09-13 build that failed at `git checkout` with "reference is not a tree", because upstream `refs/pull/4831/head` is still at the pre-09-13 head and the commit was absent from what RTD fetched. This also picks up vllm-project#8065, which commented out the typing-extensions, pillow and psutil mkdocstrings inventories because Read the Docs rate-limits them in strict builds. This branch still had all three enabled under `fail_on_warning: true`. No change to the encode-parallel delta itself. Signed-off-by: Abdul Basit Tonmoy <abdulbasittonmoy@gmail.com>
Explicitly disable derived TPOT for Realtime chat, which only observes client response timings. Add regression coverage for explicit and VAD turns, including tokenizer fallback and TPOT goodput rejection. List the benchmark wrapper, dataset, and duplex client in the nightly job dependencies. Align the external inventory configuration with the ReadTheDocs HTTP 429 fix already merged in upstream PR vllm-project#8065. Update the timing comments and benchmark documentation to match the actual metric calculation. Signed-off-by: psv666 <2693925048@qq.com>
Purpose
Fixes #8049.
The Qwen3-TTS chunked-prefill phase-contract probe still accessed the legacy runner's
_talker_mtp_forward. The default MRv2OmniARModelRunnerruns MTP throughmodel_state._run_batched_mtp, so the probe failed during initialization before checking the audio result.Update the test-only worker extension to observe the active runner's preprocess and MTP paths. Preserve the legacy path, and verify on MRv2 that prefill rows do not enter MTP and that the one-token prefill tail embedding reaches the next stage unchanged.
The first two Read the Docs builds for this PR failed because three external Read the Docs inventories returned HTTP 429. Disable those optional inventories in
mkdocs.ymlso strict documentation builds no longer depend on rate-limited hosts.Test Plan
vLLM Version: 0.30.0
vLLM-Omni Commit:
6a14af3be(based on57d3e9fdf82c06094ee0c0b39c79d24b247fe8c3)tests/e2e/features/prefill_phase/test_chunked_prefill_tail.py::test_one_token_prefill_tail_matches_unchunked_audiowith-m 'advanced_model and cuda' --run-level advanced_model.pre-commit run --files mkdocs.yml tests/e2e/features/prefill_phase/worker_extension.pygit diff --check origin/main...HEADTest Result
BaseandCustomVoiceno_async_chunk). Both failures generated audio and persisted on isolated rerun; these tests do not use the changed probe. Two-GPU TTS performance cases: 8 passed.