Skip to content

[Model] Default Qwen3-TTS to experimental Model Runner V2 - #7930

Merged
Sy0307 merged 2 commits into
vllm-project:mainfrom
Sy0307:sy03/qwen3-tts-mrv2-default
Sep 21, 2026
Merged

Sy0307 merged 2 commits into
vllm-project:mainfrom
Sy0307:sy03/qwen3-tts-mrv2-default

Conversation

@Sy0307

@Sy0307 Sy0307 commented Sep 21, 2026

Copy link
Copy Markdown
Collaborator

Summary

Make the bundled Qwen3-TTS deploy profile select Model Runner V2 on CUDA, and document MRV2 as experimental with an explicit opt-out. Follow-up to #7781, which enabled MRV2 as opt-in only.

  • vllm_omni/deploy/qwen3_tts.yaml: model_runner: v2, stage-0 talker prefill bound to 512 tokens (the validated basic MRV2 profile), and platforms: sections that pin V1 on NPU, XPU, ROCm and MUSA.
  • qwen3_tts_mrv2.yaml: kept as an explicit MRV2 profile; header updated because the base file now selects V2 as well. Removing the now-redundant profile can be a separate cleanup.
  • Config tests: the "default profiles are explicit opt-in" test is split — qwen3_tts_high_concurrency.yaml stays V1 with its opt-in MRV2 sibling, and a new test pins the Qwen3-TTS default plus per-platform fallbacks.
  • Docs: docs/configuration/stage_configs.md and recipes/Qwen/Qwen3-TTS.md note the experimental default and how to opt out.

Scope

Qwen3-TTS only. Qwen3-Omni and MOSS keep V1; their opt-in MRV2 profiles are untouched, and no other model family is affected.

Testing

Static checks passed locally: YAML validation, a standalone simulation of config resolution (CUDA → v2; NPU/XPU/ROCm/MUSA → v1 with no NotImplementedError), py_compile, and the repository pre-commit hooks on the changed files. The config test suite was not executed here (no vLLM 0.29.0 environment on this machine); CI should exercise the config tests and the CUDA Qwen3-TTS e2e lanes that use the default profile.

Risks and known limits

  • MRV2 is still experimental. The open limits from [Core][Model] Enable Qwen3-TTS MRV2 and optimize the TTS pipeline #7781 apply to default deployments after this PR: codec budget exhaustion, streaming continuity, and single-GPU underruns with C128 RTF above 1. Reverting the default is a one-line change.
  • CUDA e2e lanes that pass qwen3_tts.yaml now exercise MRV2, so V1 coverage for the default profile moves to explicit model_runner: v1 configurations.
  • Non-CUDA backends are unchanged (pinned to V1 in the same file).

AI assistance: drafted with opencode; human review requested.

The bundled Qwen3-TTS deploy profile now selects `model_runner: v2` on
CUDA, matching the validated basic MRV2 profile: stage 0 (talker) bounds
per-step prefill work to 512 tokens, and each stage selects the native
Talker AR / Code2Wav generation runner.

MRV2 stays experimental for this model. NPU, XPU, ROCm and MUSA keep V1
through the `platforms:` sections, so non-CUDA backends see no behavior
change. V1 can still be requested with `model_runner: v1`, and
`qwen3_tts_mrv2.yaml` is kept as an explicit profile that resolves to the
same runner selection as the default.

The config test that asserted "default profiles are explicit opt-in" is
rewritten: `qwen3_tts_high_concurrency.yaml` still defaults to V1 with an
opt-in MRV2 sibling, and a new test pins the Qwen3-TTS default plus its
per-platform fallbacks.

Signed-off-by: Sy03 <1370724210@qq.com>
@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/vllm_omni_config.md, docs/design/module/entrypoints.md.

Module owners: @lishunyang12 @alex-jw-brooks @linyueqian

Routing: @lishunyang12 via module of the changed files, CODEOWNERS; @alex-jw-brooks via module of the changed files; @linyueqian via module of the changed files

@Sy0307, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@Sy0307

Sy0307 commented Sep 21, 2026 •

Copy link
Copy Markdown
Collaborator Author

Seed-TTS validation on H200 (seed-tts-eval/unique-666, en, --seed 42 --disable-shuffle), PR head @ 5cb97a75, vLLM 0.29.0 + torch 2.13.0.

Protocol

  • Server: /root/pr7930-head (PR head tree, vllm_omni.deploy.qwen3_tts.yaml with the shipped model_runner: v2 + stage-0 512 keep). V1 arm uses the same YAML with model_runner: v1.
  • Bench client: vllm bench serve --omni --dataset-name seed-tts --dataset-path /root/.cache/huggingface/qwen3-tts-cold-reference-20260917/unique-seedtts --seed-tts-locale en --seed-tts-file-ref-audio --seed 42 --disable-shuffle.
  • VLLM_OMNI_EVENT_DRIVEN_ORCH=0 serves --allowed-local-media-path /root.
  • Steps on GPU 3: warmup C64/666 (discarded) → 2× measured C64/666 → 1× measured C128/1088 → 1× V1 C64/666 (all 666/666 completed).

Results

Arm Concurrency / requests Throughput (audio-s/s) TTFP p50 (ms) TTFP p99 (ms) Mean RTF Total audio duration (s)
V2 warmup (discarded) C64 / 666 39.0 3433 4554 1.670 2775
V2 measured r1 C64 / 666 80.9 1434 2014 0.823 2782
V2 measured r2 C64 / 666 81.1 1450 1714 0.820 2772
V2 measured mean C64 / 666 81.0 1442 1864 0.821 2777
V1 control C64 / 666 39.9 3473 6290 1.682 2789
V2 C128 / 1088 82.8 2866 4724 1.607 4575

Interpretation

  • Default V2 on this run roughly 2× V1 at C64 (+103% audio-s/s, mean RTF 0.82 vs 1.68, TTFP p50/p99 –58%/–70%). Warmup round matches V1 (≤2.3% lower), consistent with AMx graph amortization.
  • At C128 / 666 more requests the V2 curve stalls (RTF mean 1.61, >1). Matches the “B4 throughput tradeoff warning” in [Core][Model] Enable Qwen3-TTS MRV2 and optimize the TTS pipeline #7781 on single-GPU.
  • Only concurrency C64 and C128 were measured at Base. No MPS, no 1.7B-CustomVoice/VoiceDesign requalification, no WER / speaker-sim or listening check. For lab use on H100/H200 the default profile is consistent with the earlier MRV2 numbers reported on [Core][Model] Enable Qwen3-TTS MRV2 and optimize the TTS pipeline #7781.

Artifacts: /root/pr7930-{warm-c64,c64-r1,c64-r2,c128,v1-c64}/{server,bench-cli,r}.json|log inside vllm-minghui.

@Sy0307

Sy0307 commented Sep 21, 2026 •

Copy link
Copy Markdown
Collaborator Author

Rebased onto current upstream/main 856d19f42, force-pushed the branch, and re-ran the H200 Seed-TTS Pi A/B against the new head (b24a17810). Same protocol as the previous comment: /root/pr7930-head tree, vLLM 0.29.0 / torch 2.13.0, seed-tts-eval/unique-666 en, --seed 42 --disable-shuffle, VLLM_OMNI_EVENT_DRIVEN_ORCH=0, --allowed-local-media-path /root.

Arm Concurrency / requests Throughput (audio-s/s) TTFP p50 (ms) TTFP p99 (ms) Mean RTF Total audio (s)
V2 warmup (discarded) C64 / 666 38.99 3433 4554 1.670 2775
V2 measured r1 C64 / 666 80.86 1434 2014 0.823 2782
V2 measured r2 C64 / 666 81.11 1450 1714 0.820 2772
V2 mean (C64) C64 / 666 80.98 1442 1864 0.821 2777
V1 control C64 / 666 39.91 3473 6290 1.682 2789
V2 C128 / 1088 82.67 2894 4647 1.614 4586

All 666/666 and 1088/1088 requests completed with non-empty audio throughout (no codec/budget failure observed in this round).

Summary deltas (V2 vs V1 on C64):

  • Throughput: +102.9%
  • TTFP p50: –58.5%; TTFP p99: –70.4%
  • Mean RTF ratio: 0.488

C128 V2 mean RTF >1 (1.614) and audio-s/s only 1.02× C64, so C128 observation continues to match the B4 limits stated in #7781.

@Sy0307
Sy0307 enabled auto-merge (squash) September 21, 2026 17:20
Signed-off-by: Sy03 <1370724210@qq.com>
@Sy0307 Sy0307 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 21, 2026
@Sy0307
Sy0307 merged commit 60a933b into vllm-project:main Sep 21, 2026
7 of 9 checks passed
y-null pushed a commit to y-null/vllm-omni that referenced this pull request Sep 22, 2026
…ct#7930)

Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: y-null <y-null@users.noreply.github.com>
y-null pushed a commit to y-null/vllm-omni that referenced this pull request Sep 22, 2026
@linyueqian

Copy link
Copy Markdown
Collaborator

@Sy0307 heads-up: the merge-only TTS · Qwen3-TTS CustomVoice Test step has been red on main since this landed. Build 15793 (d8d162d8, the previous merge that ran the step) passed it; build 15796 (60a933b9, this PR's own merge build) and 15797 both hit the job timeout (exit 124). The online cases pass; the run stalls in tests/e2e/offline_inference/test_qwen3_tts_customvoice.py::test_text_to_audio_001[no_cuda_graph], where the stage 1 coordinator logs Request ... timed out waiting for input (waited > 600s) and marks the request FINISHED_ERROR, so under the new Model Runner V2 default the offline no-CUDA-graph path never delivers stage 0 output to stage 1. Since this PR only changes deploy YAML defaults, the quickest way to get main green is to restore the legacy runner as the default in qwen3_tts.yaml (keeping qwen3_tts_mrv2.yaml for the opt-in) while the offline path is fixed; happy to review that or a forward fix today.

@linyueqian

Copy link
Copy Markdown
Collaborator

Addendum: the TTS · Qwen3-TTS Base Test step is broken by the same change. It passed on 15792 and 15793 and has ended errored on every main build that ran it since 15796 (15797, 15824, 15829, 15833): both job attempts hit the job timeout and are terminated by signal while the stage engine cores are still up, the same stall shape as the CustomVoice offline case, so with Model Runner V2 as the default both Qwen3-TTS merge-only steps are red. That strengthens the case for restoring the legacy runner default in qwen3_tts.yaml today while the offline path is fixed.

mlaneuville pushed a commit to mlaneuville/vllm-omni that referenced this pull request Sep 22, 2026
…ct#7930)

Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Matthieu Laneuville <matthieu.laneuville@surf.nl>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants