Repository navigation
[Model] Default Qwen3-TTS to experimental Model Runner V2 - #7930
Conversation
The bundled Qwen3-TTS deploy profile now selects `model_runner: v2` on CUDA, matching the validated basic MRV2 profile: stage 0 (talker) bounds per-step prefill work to 512 tokens, and each stage selects the native Talker AR / Code2Wav generation runner. MRV2 stays experimental for this model. NPU, XPU, ROCm and MUSA keep V1 through the `platforms:` sections, so non-CUDA backends see no behavior change. V1 can still be requested with `model_runner: v1`, and `qwen3_tts_mrv2.yaml` is kept as an explicit profile that resolves to the same runner selection as the default. The config test that asserted "default profiles are explicit opt-in" is rewritten: `qwen3_tts_high_concurrency.yaml` still defaults to V1 with an opt-in MRV2 sibling, and a new test pins the Qwen3-TTS default plus its per-platform fallbacks. Signed-off-by: Sy03 <1370724210@qq.com>
|
This PR appears to belong to: docs/design/module/vllm_omni_config.md, docs/design/module/entrypoints.md. Module owners: @lishunyang12 @alex-jw-brooks @linyueqian Routing: @lishunyang12 via module of the changed files, CODEOWNERS; @alex-jw-brooks via module of the changed files; @linyueqian via module of the changed files @Sy0307, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
|
Seed-TTS validation on H200 ( Protocol
Results
Interpretation
Artifacts: |
|
Rebased onto current upstream/main
All 666/666 and 1088/1088 requests completed with non-empty audio throughout (no codec/budget failure observed in this round). Summary deltas (V2 vs V1 on C64):
C128 V2 mean RTF >1 (1.614) and audio-s/s only 1.02× C64, so C128 observation continues to match the B4 limits stated in #7781. |
Signed-off-by: Sy03 <1370724210@qq.com>
…ct#7930) Signed-off-by: Sy03 <1370724210@qq.com> Signed-off-by: y-null <y-null@users.noreply.github.com>
…ct#7930) Signed-off-by: Sy03 <1370724210@qq.com>
|
@Sy0307 heads-up: the merge-only |
|
Addendum: the |
…ct#7930) Signed-off-by: Sy03 <1370724210@qq.com> Signed-off-by: Matthieu Laneuville <matthieu.laneuville@surf.nl>
…ct#7930) Signed-off-by: Sy03 <1370724210@qq.com>
Summary
Make the bundled Qwen3-TTS deploy profile select Model Runner V2 on CUDA, and document MRV2 as experimental with an explicit opt-out. Follow-up to #7781, which enabled MRV2 as opt-in only.
vllm_omni/deploy/qwen3_tts.yaml:model_runner: v2, stage-0 talker prefill bound to 512 tokens (the validated basic MRV2 profile), andplatforms:sections that pin V1 on NPU, XPU, ROCm and MUSA.qwen3_tts_mrv2.yaml: kept as an explicit MRV2 profile; header updated because the base file now selects V2 as well. Removing the now-redundant profile can be a separate cleanup.qwen3_tts_high_concurrency.yamlstays V1 with its opt-in MRV2 sibling, and a new test pins the Qwen3-TTS default plus per-platform fallbacks.docs/configuration/stage_configs.mdandrecipes/Qwen/Qwen3-TTS.mdnote the experimental default and how to opt out.Scope
Qwen3-TTS only. Qwen3-Omni and MOSS keep V1; their opt-in MRV2 profiles are untouched, and no other model family is affected.
Testing
Static checks passed locally: YAML validation, a standalone simulation of config resolution (CUDA → v2; NPU/XPU/ROCm/MUSA → v1 with no
NotImplementedError),py_compile, and the repository pre-commit hooks on the changed files. The config test suite was not executed here (no vLLM 0.29.0 environment on this machine); CI should exercise the config tests and the CUDA Qwen3-TTS e2e lanes that use the default profile.Risks and known limits
qwen3_tts.yamlnow exercise MRV2, so V1 coverage for the default profile moves to explicitmodel_runner: v1configurations.AI assistance: drafted with opencode; human review requested.