Repository navigation
[Core][Model] Enable Qwen3-TTS MRV2 and optimize the TTS pipeline - #7781
Conversation
0356d05 to
84ae3e4
Compare
Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Sy03 <1370724210@qq.com>
96e24d3 to
d24cc56
Compare
linyueqian
left a comment
There was a problem hiding this comment.
Ran E2E streaming A/B on 1× H100 80 GB with Qwen3-TTS-12Hz-1.7B-Base and vLLM 0.29.0. All 3,328 timed requests across four configurations completed with zero request failures or stream aborts. Please address the inline correctness findings and broken regression tests before merging.
The matched comparison uses decoder batch=2 for both V1 and MRV2:
| Concurrency | V1 req/s | MRV2 req/s | Change | Mean first audio, V1 → MRV2 (ms) | Mean request RTF, V1 → MRV2 |
|---|---|---|---|---|---|
| 1 | 2.480 | 2.569 | +3.6% | 56.8 → 59.5 | 0.0915 → 0.0899 |
| 16 | 15.934 | 16.836 | +5.7% | 173.9 → 164.0 | 0.2201 → 0.2096 |
| 64 | 24.060 | 24.595 | +2.2% | 509.6 → 774.7 | 0.5757 → 0.5657 |
At C64, mean first-audio latency increases 52.0%; the average per-run p95 rises from 1,218 to 1,505 ms (+23.6%). With decoder batch=1 on both sides, MRV2 throughput regresses 10.8% (20.588 → 18.366 req/s). The 19.5% gain from comparing V1/B1 against MRV2/B2 includes the batching change; it should not be presented as an MRV2-only gain. RTF is request wall time divided by generated audio duration, lower is better.
Protocol: two repeats per concurrency, respectively 32/128/256 timed requests and 8/16/64 warmups; fixed arm order; 30 public English smoke texts and one warmed clone_2.wav reference. Same full deploy settings: AR max sequences 64 and batched tokens 512, decoder max sequences 10, prefix caching/code-predictor prefix graphs/TF32 disabled, reference context 72 frames. PCM streaming at 24 kHz with a 2,048-token cap. Server and dataset seed 0; timing requests did not set per-request seeds. These are descriptive warm-reference results, not a sustained-load or broad-workload qualification.
Separate quality checks completed for 30 paired prompts per arm (120 WAVs total, C16, API seed 42). For matched B2, ASR WER is 0.265% → 0.531% (1 → 2 word errors / 377 words), and UTMOSv2 predicted MOS is 3.434 → 3.527. The paired MOS difference has a 95% utterance-bootstrap interval of [−0.001, +0.192], so this does not establish a quality improvement. ASR/MOS scoring ran outside performance timing.
Measured base 6df62d51573ca5be7826321ba2de1cd071d997f8 against head 96e24d3b180c3985f2598c83dc4838328dc0b136, using the same pinned benchmark client. Relevant Qwen3-TTS paths and all four finding files are unchanged at current head d24cc563b2342579eb80284e31348fd29c61460c, checked statically. This comparison includes the PR's model/pipeline changes as well as MRV2. The two inline test issues were also reproduced: four missing-fixture errors and two payload-shape failures.
Keep V1 as default and MRV2 experimental for Qwen3-TTS in 0.29.0; Qwen3-Omni and other flagship models remain for 0.30.0. The high-concurrency first-audio regression needs attention before promoting this path.
Declare scheduler-owned fields, narrow optional request IDs, and make test doubles explicit. Preserve external-only request ID resolution and reject unidentified requests. Retain the reviewed lifecycle tests and remove unsupported MPS benefit wording. Validation: 381 targeted CPU tests passed; all non-mypy hooks passed. Same-environment mypy comparison against d4ffde1 has no new diagnostics (base 159, branch 112); inherited diagnostics remain. Signed-off-by: Sy03 <1370724210@qq.com>
Preserve request-local reference tensors after batch list selection and clean receiver state for requests cancelled before model admission. Remove unreachable PP branches and unused span symbols; consolidate regression coverage below 3000 added test lines. Validation: 321 CPU tests and 10 CUDA tests passed. Non-mypy hooks passed. Same-environment mypy comparison: base 159, branch 112, zero new diagnostics; inherited failures remain. Signed-off-by: Sy03 <1370724210@qq.com>
|
This PR appears to belong to: docs/design/module/ar_runtime.md. Module owners: @fake0fan @Gaohan123 @tzhouam Routing: @fake0fan via module of the changed files, CODEOWNERS; @Gaohan123 via module of the changed files, CODEOWNERS; @tzhouam via module of the changed files, CODEOWNERS @Sy0307, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Omni ReviewBot triage noteAutomated triage of commit
These are automated triage suggestions only — the final decision belongs to the maintainers. |
Omni ReviewBot: three questions on the performance claim@Sy0307 this PR reads as a performance or value claim:
Before the full evidence checklist, three short questions:
When you answer, the evidence that settles it is: base and head SHA, hardware, model, workload, warm-up and repeat count, mean or percentiles with their spread, and a correctness/quality-equivalence signal; an end-to-end claim also needs stage attribution. |
Performance update: single- and two-H200 deploymentsAdding the operator's cache-capacity-aligned results below. These historical measurements use main Both sides use Qwen3-TTS-12Hz-1.7B-Base, vLLM 0.29.0 and SeedTTS English ×1,088, matching GPU allocation, request seed/order, and cache entry limits within each comparison. Full warmup rounds are excluded. Single-card: both stages on H200 GPU4. Dual-card: stages split across GPUs4/5. No MPS in the table. The single-card protocol reports three rounds with warmup discarded; dual-card uses a full excluded C128 dataset warmup before C64/C128.
Audio-s/s means generated audio duration divided by benchmark wall time, not requests/s. Throughput increase is Additional single-card results:
The improvement has limits: single-card underrun rises by 0.32 s / +123.1% at C64 and 0.90 s / +360% at C128; there is one unsuccessful request in each listed single-card run, and C128 RTF remains >1. Earlier first audio therefore does not establish smooth real-time playback or production readiness. Failure causes are not assigned from these aggregate numbers. Matching cache capacity removes that specific confounder, but this remains a comparison of complete profiles, including shared model optimizations and decoder settings—not proof of an MRV2-only speedup. The independent H100 matched-B2 result remains useful for its different single-reference workload; these H200 results do not invalidate it. A same-commit V1/B2 versus MRV2/B2 ablation would separate runner-related effects. No consistent MPS gain is established: the previous dual-card PR measurements changed from 241.1 to 230.5 audio-s/s at C64 and from 222.1 to 218.1 at C128 with MPS. No numerical single-card MPS comparison was supplied here. The PR description now includes these detailed results and qualifications. |
Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Sy03 <1370724210@qq.com>
…lm-project#7781) Signed-off-by: Sy03 <1370724210@qq.com> Signed-off-by: Matthieu Laneuville <matthieu.laneuville@surf.nl>
…lm-project#7781) Signed-off-by: Sy03 <1370724210@qq.com>
Summary
Enable opt-in Qwen3-TTS on native Model Runner V2 with its pipeline optimizations. V1 remains the default; vLLM stays at 0.29.0, with no dependency changes. Extracted from #4582; #7088 remains independent.
Qwen3-Omni/MOSS activation and experimental wakeup/transport switches are outside this PR.
Latest correctness and scope fixes
Head:
b9648399f(base used for validation and diff accounting:d4ffde1a6).d4ffde1a6, including fixtures, mocks and comments; below the requested 3,000 limit.Validation
Existing H200 environment, vLLM 0.29.0:
Reproduce CPU validation from repository root in an environment with vLLM 0.29.0, torch, pytest and pytest-mock:
python -m pytest tests/worker_v2 tests/core/sched \ tests/model_executor/models/qwen3_tts/test_code_predictor_dtype.py \ tests/model_executor/models/qwen3_tts/test_qwen3_tts_talker_preprocess.py \ -m 'core_model and cpu' -qPerformance evidence and limits
The following historical H200 results were supplied by the benchmark operator. They compare main + aligned reference-cache capacity (V1) against the PR implementation with MRV2/B2, on the same GPU(s) and workload within each deployment comparison. Full-dataset warmup rounds are excluded.
Provenance: main
d4ffde1a6with API cache capacity patched to 1024 entries / 512 MiB and model artifact cache to 1024 entries; PR source/root/mrv2-cleanup-20260918, reported as500f5626plus the then-uncommitted cleanup worktree, using B2. These are not benchmarks of current headb9648399f. The working-tree snapshot is part of the measured revision; the commit alone does not reproduce it. The figures below are the operator's supplied summaries; raw logs and repetition aggregation have not been independently re-audited for this update.Protocol:
Qwen/Qwen3-TTS-12Hz-1.7B-Base, SeedTTS English ×1,088, vLLM 0.29.0, same dataset seed/order and requests per comparison, no MPS for the main tables. Single-GPU runs colocate both stages on H200 GPU4 (reported three rounds, warmup discarded). Two-GPU runs use H200 GPUs4/5 with separate stages; each side first completes an excluded 1,088-request C128 warmup, then measures C64/C128. The supplied summaries do not specify whether the single-GPU retained values are means or a selected repeat.Throughput and first-audio latency
Throughput is generated audio seconds per wall-clock second (audio-s/s), not requests/s. Latencies are milliseconds. Gain =
(PR / main - 1) × 100%.Additional single-GPU measurements
RTF, underrun and GPU-utilization figures retain the operator's reported aggregation; they are not per-request guarantees. GPU utilization is sampled device utilization, not SM occupancy.
Tradeoffs: single-GPU throughput improves by approximately 73–76%, but reported underrun worsens by 0.32 s at C64 (+123.1%) and 0.90 s at C128 (+360%). Each listed single-GPU run has one unsuccessful request; its cause is not established by this summary. C128 PR RTF remains above 1. These observations do not support a blanket production-readiness or uninterrupted-playback claim.
The earlier two-GPU MPS comparison was 241.1 → 230.5 audio-s/s at C64 and 222.1 → 218.1 at C128: no demonstrated throughput benefit. The operator also reports no single-GPU MPS gain, but supplied no numerical single-GPU MPS rows here. B4 has no established consistent advantage.
These are whole-profile improvements, not an isolated MRV2 runner ablation: matching reference-cache capacity does not align cache representation, all model optimizations, or decoder graph/batch settings. They also do not contradict the independent single-H100, single-reference matched-B2 review; hardware, topology, workload and metrics differ. A same-commit PR V1/B2 versus PR MRV2/B2 comparison is needed to isolate the runner-related contribution.
Two-GPU raw artifact locations recorded by the operator, inside
vllm-minghui:/root/fair-main-{c64,c128,warmup}/,/root/fair-pr-{c64,c128,warmup}/, and/root/fair-mps2-pr-{c64,c128,warm}/. A single-GPU raw artifact path was not supplied with these results. These cache-aligned figures supersede the earlier capacity-unaligned table in this description; the underlying historical runs remain distinct.Known limits
AI assistance: Codex assisted extraction, implementation, tests, documentation and acceptance. Human review requested.