Skip to content

[Core][Model] Enable Qwen3-TTS MRV2 and optimize the TTS pipeline - #7781

Merged
Sy0307 merged 25 commits into
vllm-project:mainfrom
Sy0307:sy03/qwen3-tts-mrv2
Sep 21, 2026
Merged

Sy0307 merged 25 commits into
vllm-project:mainfrom
Sy0307:sy03/qwen3-tts-mrv2

Conversation

@Sy0307

@Sy0307 Sy0307 commented Sep 18, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Enable opt-in Qwen3-TTS on native Model Runner V2 with its pipeline optimizations. V1 remains the default; vLLM stays at 0.29.0, with no dependency changes. Extracted from #4582; #7088 remains independent.

  • Add shared MRV2 AR/generation runners, request GPU state, chunk readiness, and ordered asynchronous output delivery with cancellation/error handling.
  • Adapt Qwen3-TTS Talker/MTP and Code2Wav, retaining per-request sampling state and GPU-resident intermediate data.
  • Optimize request snapshots, compact reference-audio storage, fixed-token projections and token uploads; align API/model cache entry limits to 1024 (API payload limit 512 MiB).
  • Provide opt-in basic/B2 profiles and an experimental B4 profile. B2 sets both decode batch capacity and captured graph shapes. MPS is optional deployment documentation, with no throughput-benefit claim.

Qwen3-Omni/MOSS activation and experimental wakeup/transport switches are outside this PR.

Latest correctness and scope fixes

Head: b9648399f (base used for validation and diff accounting: d4ffde1a6).

  • Preserve request-scoped reference-code tensors after selecting a batch list element. A mixed prefill/decode batch must not re-slice reference frames when their length happens to match the token/padded-token count.
  • Clean up receiver state when a request is cancelled before its first chunk/model admission; terminal emission remains limited to live model requests.
  • Remove unreachable pipeline-parallel branches (PP>1 is explicitly rejected) and unused span symbols. Retain output ownership, backpressure and terminal ordering.
  • Consolidate duplicate/implementation-detail tests. 2,993 added test lines relative to d4ffde1a6, including fixtures, mocks and comments; below the requested 3,000 limit.

Validation

Existing H200 environment, vLLM 0.29.0:

  • 321 CPU tests passed: MRV2 workers, schedulers, code predictor and Talker preprocessing; 2 CUDA cases deselected in this lane.
  • 10 CUDA tests passed: intermediate-buffer residency, actual decoder graph capture/replay, and asynchronous output ownership.
  • Non-mypy pre-commit checks and whitespace checks passed. Commit identity/DCO verified.
  • Same-environment mypy comparison on the selected changed files: baseline 159 diagnostics, PR 112, zero new diagnostics. This is not a full mypy pass; the local mypy hook was explicitly skipped for commit after separate baseline attribution.

Reproduce CPU validation from repository root in an environment with vLLM 0.29.0, torch, pytest and pytest-mock:

python -m pytest tests/worker_v2 tests/core/sched \
  tests/model_executor/models/qwen3_tts/test_code_predictor_dtype.py \
  tests/model_executor/models/qwen3_tts/test_qwen3_tts_talker_preprocess.py \
  -m 'core_model and cpu' -q

Performance evidence and limits

The following historical H200 results were supplied by the benchmark operator. They compare main + aligned reference-cache capacity (V1) against the PR implementation with MRV2/B2, on the same GPU(s) and workload within each deployment comparison. Full-dataset warmup rounds are excluded.

Provenance: main d4ffde1a6 with API cache capacity patched to 1024 entries / 512 MiB and model artifact cache to 1024 entries; PR source /root/mrv2-cleanup-20260918, reported as 500f5626 plus the then-uncommitted cleanup worktree, using B2. These are not benchmarks of current head b9648399f. The working-tree snapshot is part of the measured revision; the commit alone does not reproduce it. The figures below are the operator's supplied summaries; raw logs and repetition aggregation have not been independently re-audited for this update.

Protocol: Qwen/Qwen3-TTS-12Hz-1.7B-Base, SeedTTS English ×1,088, vLLM 0.29.0, same dataset seed/order and requests per comparison, no MPS for the main tables. Single-GPU runs colocate both stages on H200 GPU4 (reported three rounds, warmup discarded). Two-GPU runs use H200 GPUs4/5 with separate stages; each side first completes an excluded 1,088-request C128 warmup, then measures C64/C128. The supplied summaries do not specify whether the single-GPU retained values are means or a selected repeat.

Throughput and first-audio latency

Throughput is generated audio seconds per wall-clock second (audio-s/s), not requests/s. Latencies are milliseconds. Gain = (PR / main - 1) × 100%.

Deployment Concurrency Completed / attempted, main → PR Audio-s/s, main → PR Ratio Throughput gain TTFP p50 ms, main → PR p50 reduction TTFP p99 ms, main → PR p99 reduction
1× H200 C64 1087/1088 → 1087/1088 55.4 → 97.3 1.76× +75.6% 3946 → 724 81.7% 4859 → 1050 78.4%
1× H200 C128 1087/1088 → 1087/1088 56.3 → 97.4 1.73× +73.0% 8782 → 1470 83.3% 9730 → 3665 62.3%
2× H200 C64 1088/1088 → 1088/1088 66.8 → 241.1 3.61× +260.9% 3141 → 239 92.4% 5100 → 658 87.1%
2× H200 C128 1088/1088 → 1088/1088 67.3 → 222.1 3.30× +230.0% 7244 → 1208 83.3% 12971 → 2178 83.2%

Additional single-GPU measurements

RTF, underrun and GPU-utilization figures retain the operator's reported aggregation; they are not per-request guarantees. GPU utilization is sampled device utilization, not SM occupancy.

Concurrency E2EL p50 ms, main → PR Reduction RTF, main → PR Reported underrun s, main → PR Sampled GPU utilization, main → PR
C64 4762 → 2433 48.9% 1.22 → 0.58 0.26 → 0.58 82% → 76%
C128 9581 → 4570 52.3% 2.35 → 1.19 0.25 → 1.15 80% → 76%

Tradeoffs: single-GPU throughput improves by approximately 73–76%, but reported underrun worsens by 0.32 s at C64 (+123.1%) and 0.90 s at C128 (+360%). Each listed single-GPU run has one unsuccessful request; its cause is not established by this summary. C128 PR RTF remains above 1. These observations do not support a blanket production-readiness or uninterrupted-playback claim.

The earlier two-GPU MPS comparison was 241.1 → 230.5 audio-s/s at C64 and 222.1 → 218.1 at C128: no demonstrated throughput benefit. The operator also reports no single-GPU MPS gain, but supplied no numerical single-GPU MPS rows here. B4 has no established consistent advantage.

These are whole-profile improvements, not an isolated MRV2 runner ablation: matching reference-cache capacity does not align cache representation, all model optimizations, or decoder graph/batch settings. They also do not contradict the independent single-H100, single-reference matched-B2 review; hardware, topology, workload and metrics differ. A same-commit PR V1/B2 versus PR MRV2/B2 comparison is needed to isolate the runner-related contribution.

Two-GPU raw artifact locations recorded by the operator, inside vllm-minghui: /root/fair-main-{c64,c128,warmup}/, /root/fair-pr-{c64,c128,warmup}/, and /root/fair-mps2-pr-{c64,c128,warm}/. A single-GPU raw artifact path was not supplied with these results. These cache-aligned figures supersede the earlier capacity-unaligned table in this description; the underlying historical runs remain distinct.

Known limits

  • Codec budget-exhaustion/repetitive-generation failures remain under investigation; that investigation is currently paused. This PR does not claim to fix or eliminate them.
  • Audio quality and streaming continuity still require broader qualification; early first audio alone does not establish uninterrupted playback.
  • The basic profile uses a 512-token prefill budget. A historical large-budget startup failure was mitigated by that configuration, not established as generally fixed.
  • Ready for review does not mean these deployment/quality limitations are resolved or that the PR should auto-merge.

AI assistance: Codex assisted extraction, implementation, tests, documentation and acceptance. Human review requested.

@Sy0307
Sy0307 force-pushed the sy03/qwen3-tts-mrv2 branch from 0356d05 to 84ae3e4 Compare September 18, 2026 09:02
Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Sy03 <1370724210@qq.com>
@Sy0307
Sy0307 force-pushed the sy03/qwen3-tts-mrv2 branch from 96e24d3 to d24cc56 Compare September 18, 2026 18:25

@linyueqian linyueqian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ran E2E streaming A/B on 1× H100 80 GB with Qwen3-TTS-12Hz-1.7B-Base and vLLM 0.29.0. All 3,328 timed requests across four configurations completed with zero request failures or stream aborts. Please address the inline correctness findings and broken regression tests before merging.

The matched comparison uses decoder batch=2 for both V1 and MRV2:

Concurrency V1 req/s MRV2 req/s Change Mean first audio, V1 → MRV2 (ms) Mean request RTF, V1 → MRV2
1 2.480 2.569 +3.6% 56.8 → 59.5 0.0915 → 0.0899
16 15.934 16.836 +5.7% 173.9 → 164.0 0.2201 → 0.2096
64 24.060 24.595 +2.2% 509.6 → 774.7 0.5757 → 0.5657

At C64, mean first-audio latency increases 52.0%; the average per-run p95 rises from 1,218 to 1,505 ms (+23.6%). With decoder batch=1 on both sides, MRV2 throughput regresses 10.8% (20.588 → 18.366 req/s). The 19.5% gain from comparing V1/B1 against MRV2/B2 includes the batching change; it should not be presented as an MRV2-only gain. RTF is request wall time divided by generated audio duration, lower is better.

Protocol: two repeats per concurrency, respectively 32/128/256 timed requests and 8/16/64 warmups; fixed arm order; 30 public English smoke texts and one warmed clone_2.wav reference. Same full deploy settings: AR max sequences 64 and batched tokens 512, decoder max sequences 10, prefix caching/code-predictor prefix graphs/TF32 disabled, reference context 72 frames. PCM streaming at 24 kHz with a 2,048-token cap. Server and dataset seed 0; timing requests did not set per-request seeds. These are descriptive warm-reference results, not a sustained-load or broad-workload qualification.

Separate quality checks completed for 30 paired prompts per arm (120 WAVs total, C16, API seed 42). For matched B2, ASR WER is 0.265% → 0.531% (1 → 2 word errors / 377 words), and UTMOSv2 predicted MOS is 3.434 → 3.527. The paired MOS difference has a 95% utterance-bootstrap interval of [−0.001, +0.192], so this does not establish a quality improvement. ASR/MOS scoring ran outside performance timing.

Measured base 6df62d51573ca5be7826321ba2de1cd071d997f8 against head 96e24d3b180c3985f2598c83dc4838328dc0b136, using the same pinned benchmark client. Relevant Qwen3-TTS paths and all four finding files are unchanged at current head d24cc563b2342579eb80284e31348fd29c61460c, checked statically. This comparison includes the PR's model/pipeline changes as well as MRV2. The two inline test issues were also reproduced: four missing-fixture errors and two payload-shape failures.

Keep V1 as default and MRV2 experimental for Qwen3-TTS in 0.29.0; Qwen3-Omni and other flagship models remain for 0.30.0. The high-concurrency first-audio regression needs attention before promoting this path.

Comment thread vllm_omni/worker_v2/omni_data_plane.py
Comment thread vllm_omni/worker_v2/omni_ar_model_runner.py Outdated
Comment thread tests/model_executor/models/qwen3_tts/test_qwen3_tts_stateless_capture.py Outdated
Comment thread tests/worker_v2/test_omni_data_plane.py
Declare scheduler-owned fields, narrow optional request IDs, and make test doubles explicit. Preserve external-only request ID resolution and reject unidentified requests. Retain the reviewed lifecycle tests and remove unsupported MPS benefit wording.

Validation: 381 targeted CPU tests passed; all non-mypy hooks passed. Same-environment mypy comparison against d4ffde1 has no new diagnostics (base 159, branch 112); inherited diagnostics remain.
Signed-off-by: Sy03 <1370724210@qq.com>
Preserve request-local reference tensors after batch list selection and clean receiver state for requests cancelled before model admission. Remove unreachable PP branches and unused span symbols; consolidate regression coverage below 3000 added test lines.

Validation: 321 CPU tests and 10 CUDA tests passed. Non-mypy hooks passed. Same-environment mypy comparison: base 159, branch 112, zero new diagnostics; inherited failures remain.
Signed-off-by: Sy03 <1370724210@qq.com>
@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/ar_runtime.md.

Module owners: @fake0fan @Gaohan123 @tzhouam

Routing: @fake0fan via module of the changed files, CODEOWNERS; @Gaohan123 via module of the changed files, CODEOWNERS; @tzhouam via module of the changed files, CODEOWNERS

@Sy0307, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 18, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit 2e3c7fe2c171 produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 18, 2026 •

Copy link
Copy Markdown

Omni ReviewBot: three questions on the performance claim

@Sy0307 this PR reads as a performance or value claim:

  • title: [Core][Model] Enable Qwen3-TTS MRV2 and optimize the TTS pipeline
  • claim: | Deployment | Concurrency | Completed / attempted, main → PR | Audio-s/s, main → PR | Ratio | Throughput gain | TTFP p50 ms, main → PR | p50 reduction | TTFP p99 ms, main → PR | p99 reduction…
  • claim: Tradeoffs: single-GPU throughput improves by approximately 73–76%, but reported underrun worsens by 0.32 s at C64 (+123.1%) and 0.90 s at C128 (+360%).

Before the full evidence checklist, three short questions:

  1. Bottleneck — what is the current bottleneck, and which profile, trace or per-stage measurement shows it?
  2. Value — what does the change buy the user or the system (latency, throughput, memory, cost), and at which workload?
  3. A/B or ablation — is there a same-workload, same-head/config comparison that isolates each main claim on its own? For stacked optimizations, one number per item rather than a blended delta.

When you answer, the evidence that settles it is: base and head SHA, hardware, model, workload, warm-up and repeat count, mean or percentiles with their spread, and a correctness/quality-equivalence signal; an end-to-end claim also needs stage attribution.

@Sy0307

Sy0307 commented Sep 18, 2026

Copy link
Copy Markdown
Collaborator Author

Performance update: single- and two-H200 deployments

Adding the operator's cache-capacity-aligned results below. These historical measurements use main d4ffde1a6 + cache patches and PR 500f5626 + the cleanup worktree/B2; they are not a fresh benchmark of current head b9648399f. Raw-log verification and the exact retained-round aggregation remain to be supplied for this update.

Both sides use Qwen3-TTS-12Hz-1.7B-Base, vLLM 0.29.0 and SeedTTS English ×1,088, matching GPU allocation, request seed/order, and cache entry limits within each comparison. Full warmup rounds are excluded. Single-card: both stages on H200 GPU4. Dual-card: stages split across GPUs4/5. No MPS in the table. The single-card protocol reports three rounds with warmup discarded; dual-card uses a full excluded C128 dataset warmup before C64/C128.

Deployment Concurrency Audio-s/s main → PR Ratio / increase TTFP p50 ms main → PR TTFP p99 ms main → PR Completed/attempted main → PR
1× H200 C64 55.4 → 97.3 1.76× / +75.6% 3946 → 724 (−81.7%) 4859 → 1050 (−78.4%) 1087/1088 → 1087/1088
1× H200 C128 56.3 → 97.4 1.73× / +73.0% 8782 → 1470 (−83.3%) 9730 → 3665 (−62.3%) 1087/1088 → 1087/1088
2× H200 C64 66.8 → 241.1 3.61× / +260.9% 3141 → 239 (−92.4%) 5100 → 658 (−87.1%) 1088/1088 → 1088/1088
2× H200 C128 67.3 → 222.1 3.30× / +230.0% 7244 → 1208 (−83.3%) 12971 → 2178 (−83.2%) 1088/1088 → 1088/1088

Audio-s/s means generated audio duration divided by benchmark wall time, not requests/s. Throughput increase is (PR/main − 1) × 100%.

Additional single-card results:

Concurrency E2EL p50 ms main → PR RTF main → PR Reported underrun s main → PR Sampled GPU utilization main → PR
C64 4762 → 2433 (−48.9%) 1.22 → 0.58 0.26 → 0.58 82% → 76%
C128 9581 → 4570 (−52.3%) 2.35 → 1.19 0.25 → 1.15 80% → 76%

The improvement has limits: single-card underrun rises by 0.32 s / +123.1% at C64 and 0.90 s / +360% at C128; there is one unsuccessful request in each listed single-card run, and C128 RTF remains >1. Earlier first audio therefore does not establish smooth real-time playback or production readiness. Failure causes are not assigned from these aggregate numbers.

Matching cache capacity removes that specific confounder, but this remains a comparison of complete profiles, including shared model optimizations and decoder settings—not proof of an MRV2-only speedup. The independent H100 matched-B2 result remains useful for its different single-reference workload; these H200 results do not invalidate it. A same-commit V1/B2 versus MRV2/B2 ablation would separate runner-related effects.

No consistent MPS gain is established: the previous dual-card PR measurements changed from 241.1 to 230.5 audio-s/s at C64 and from 222.1 to 218.1 at C128 with MPS. No numerical single-card MPS comparison was supplied here. The PR description now includes these detailed results and qualifications.

@Sy0307 Sy0307 added amd-test Used to trigger AMD CI separately. cuda-test Used to trigger vllm-omni cuda CI separately. ready label to trigger buildkite CI and removed cuda-test Used to trigger vllm-omni cuda CI separately. amd-test Used to trigger AMD CI separately. ready label to trigger buildkite CI labels Sep 20, 2026
Signed-off-by: Sy03 <1370724210@qq.com>
@Sy0307 Sy0307 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 20, 2026
@hsliuustc0106 hsliuustc0106 removed the new model add new model label Sep 20, 2026
Signed-off-by: Sy03 <1370724210@qq.com>
@Sy0307 Sy0307 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 21, 2026
@Sy0307
Sy0307 merged commit a9e2117 into vllm-project:main Sep 21, 2026
9 checks passed
mlaneuville pushed a commit to mlaneuville/vllm-omni that referenced this pull request Sep 22, 2026
…lm-project#7781)

Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Matthieu Laneuville <matthieu.laneuville@surf.nl>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

core related to core module: cache, scheduler, engine, worker, modelrunner high priority high priority issue, needs to be done asap ready label to trigger buildkite CI tts code related to tts models

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants