Add non-async chunk support for Qwen3-TTS by linyueqian · Pull Request #1678 · vllm-project/vllm-omni

linyueqian · 2026-03-05T06:30:29Z

Summary

Add talker2code2wav non-async stage input processor for Qwen3-TTS
Add qwen3_tts_no_async_chunk.yaml stage config
Add benchmark scripts to compare async_chunk on vs off (TTFP, E2E, RTF)

How to test

# Serve with non-async chunk config
vllm-omni serve <model> \
    --stage-configs-path vllm_omni/model_executor/stage_configs/qwen3_tts_no_async_chunk.yaml \
    --trust-remote-code --omni

# Test with curl
curl http://localhost:8000/v1/audio/speech \
    -H "Content-Type: application/json" \
    -d '{"input": "Hello world", "voice": "Vivian", "response_format": "wav"}' \
    -o test.wav

# Run full benchmark (async on vs off)
GPU_DEVICE=0 MODEL=<model> bash benchmarks/qwen3-tts/vllm_omni/run_async_chunk_benchmark.sh

cc @Sy0307 can you test this on your machine?

Signed-off-by: linyueqian <linyueqian@outlook.com>

chatgpt-codex-connector

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: bc935bd85c

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

Open a pull request for review
Mark a draft as ready
Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

chatgpt-codex-connector · 2026-03-05T06:37:34Z

+        audio_codes = output.multimodal_output["audio_codes"].to(torch.long)
+        # Code2Wav expects codebook-major flat: [Q*num_frames]
+        codec_codes = audio_codes.transpose(0, 1).cpu().reshape(-1).tolist()


Filter zero-padded frames before flattening audio codes

The new non-async talker2code2wav path forwards every row from multimodal_output["audio_codes"] directly to code2wav, but Qwen3-TTS emits all-zero rows for invalid/EOS layer-0 steps (the async path explicitly drops these via _extract_last_frame). Including those rows adds padded codec frames to the final decode, which can produce trailing silence and systematically inflate latency/RTF measurements for completed requests.

Useful? React with 👍 / 👎.

Signed-off-by: linyueqian <linyueqian@outlook.com>

hsliuustc0106 · 2026-03-05T07:02:24Z

@@ -0,0 +1,67 @@
+async_chunk: false


is this called default args @lishunyang12

Sy0307 · 2026-03-05T07:10:20Z

cc @Sy0307 can you test this on your machine?

I will test it later. Thanks for your contributing.

Sy0307 · 2026-03-05T08:45:18Z

Test results on 5090:

Metric	async_chunk ON	async_chunk OFF
TTFP mean	1062.4 ms	1065.9 ms
TTFP median	1062.5 ms	1068.5 ms
E2E mean	3001.2 ms	2945.2 ms
E2E median	2992.4 ms	2991.3 ms
E2E p90	3777.4 ms	3500.1 ms
RTF mean	0.523	0.524
Audio throughput	1.91 s/s	1.91 s/s

hsliuustc0106 · 2026-03-05T11:26:00Z

I think we need to add test for this model in different modes

Signed-off-by: linyueqian <linyueqian@outlook.com>

linyueqian · 2026-03-05T17:58:53Z

Added nightly e2e tests for the non-async-chunk mode (English + Chinese) in tests/e2e/online_serving/test_qwen3_tts.py, with a corresponding Buildkite step in test-nightly.yml.

Signed-off-by: linyueqian <linyueqian@outlook.com>

hsliuustc0106

Review

Rating: 8/10 | Verdict: ✅ Approved

Summary

Comprehensive feature addition enabling non-async chunk mode for Qwen3-TTS. Includes stage config, input processor, benchmarking tools, and E2E tests. Code quality is solid, but missing documentation and benchmark results.

Highlights

✅ Complete implementation: config, processor, benchmarks, tests
✅ Excellent benchmark tooling (717 lines across 3 files)
✅ E2E tests for both English and Chinese TTS
✅ Added to nightly CI pipeline
✅ Clear code structure with good comments

Minor Issues (non-blocking)

Missing benchmark results: PR adds benchmark scripts but provides no actual performance data. Include at least one comparison table or plot showing async_chunk off vs on.
No documentation: Missing guidance on when to use async_chunk vs non-async_chunk mode. Add a note in README or a docs page explaining the trade-offs.
Processor lacks unit tests: 31-line talker2code2wav function has no unit tests. Consider adding tests for edge cases (empty codes, invalid shapes).
Magic numbers: Line 36-37 uses hardcoded values (valid_mask logic). Could benefit from constants or comments explaining the codec structure.

Recommendation

Ready to merge. Code quality is high and functionality is complete. Documentation and benchmark results can be added in follow-up.

Reviewed by OpenClaw with vllm-omni-skills 🦐

hsliuustc0106 · 2026-03-06T01:48:51Z

+            OmniTokensPrompt(
+                prompt_token_ids=codec_codes,
+                multi_modal_data=None,
+                mm_processor_kwargs=None,


The zero-padding filter logic is critical for correctness but has no unit test. Add a test case verifying that: (1) all-zero frames are filtered, (2) valid frames are preserved, (3) edge cases like all-zero or no-zero input work correctly.

hsliuustc0106 · 2026-03-06T01:48:51Z

+echo " Plots:"
+echo "   - ${RESULT_DIR}/qwen3_tts_async_chunk_ttfp.png"
+echo "   - ${RESULT_DIR}/qwen3_tts_async_chunk_all.png"
+echo "============================================================"


The benchmark script is excellent, but the PR description should include actual benchmark results. Run the script once and paste the comparison table or link to a plot showing TTFP/E2E/RTF improvements.

hsliuustc0106 · 2026-03-06T01:48:51Z

@@ -0,0 +1,67 @@
+async_chunk: false


Add a comment at the top explaining when to use this config vs the default async_chunk config. Example: "Use this config when streaming is not required and you want simpler processing. Disables chunk-level streaming for lower complexity but potentially higher latency."

hsliuustc0106 · 2026-03-06T01:48:51Z

+            "--enforce-eager",
+            "--disable-log-stats",
+        ],
+    ) as server:


The new test class duplicates setup code from the async_chunk tests. Consider extracting common fixtures or using pytest.mark.parametrize to reduce duplication.

### vllm-omni-api - Source: [PR #1724](vllm-project/vllm-omni#1724) - Revert "[Profile] Adding metrics for Diffusion/DiT Single diffusion Pipeline (#668)" - Changes: - New feature: Revert "[Profile] Adding metrics for Diffusion/DiT Single diffusion Pipeline (#668)" ### vllm-omni-contrib - Source: [PR #1724](vllm-project/vllm-omni#1724) - Revert "[Profile] Adding metrics for Diffusion/DiT Single diffusion Pipeline (#668)" - Changes: - New feature: Revert "[Profile] Adding metrics for Diffusion/DiT Single diffusion Pipeline (#668)" ### vllm-omni-api - Source: [PR #1716](vllm-project/vllm-omni#1716) - [Feature]: Add vae-patch-parallel CLI argument in online serving - Changes: - New feature: [Feature]: Add vae-patch-parallel CLI argument in online serving ### vllm-omni-contrib - Source: [PR #1716](vllm-project/vllm-omni#1716) - [Feature]: Add vae-patch-parallel CLI argument in online serving - Changes: - New feature: [Feature]: Add vae-patch-parallel CLI argument in online serving ### vllm-omni-contrib - Source: [PR #1693](vllm-project/vllm-omni#1693) - [skip CI][Docs] Add TTS model developer guide - Changes: - New feature: [skip CI][Docs] Add TTS model developer guide ### vllm-omni-audio-tts - Source: [PR #1688](vllm-project/vllm-omni#1688) - [MiMo-Audio] Bugfix tp lg than 1 - Changes: - Bug fix: [MiMo-Audio] Bugfix tp lg than 1 ### vllm-omni-distributed - Source: [PR #1688](vllm-project/vllm-omni#1688) - [MiMo-Audio] Bugfix tp lg than 1 - Changes: - Bug fix: [MiMo-Audio] Bugfix tp lg than 1 ### vllm-omni-perf - Source: [PR #1688](vllm-project/vllm-omni#1688) - [MiMo-Audio] Bugfix tp lg than 1 - Changes: - Bug fix: [MiMo-Audio] Bugfix tp lg than 1 ### vllm-omni-perf - Source: [PR #1687](vllm-project/vllm-omni#1687) - [BugFix] Return proper HTTP status for ErrorResponse in create_speech - Changes: - Bug fix: [BugFix] Return proper HTTP status for ErrorResponse in create_speech ### vllm-omni-distributed - Source: [PR #1687](vllm-project/vllm-omni#1687) - [BugFix] Return proper HTTP status for ErrorResponse in create_speech - Changes: - Bug fix: [BugFix] Return proper HTTP status for ErrorResponse in create_speech ### vllm-omni-api - Source: [PR #1687](vllm-project/vllm-omni#1687) - [BugFix] Return proper HTTP status for ErrorResponse in create_speech - Changes: - Bug fix: [BugFix] Return proper HTTP status for ErrorResponse in create_speech - Additions: - `/v1/audio/speech` ### vllm-omni-quantization - Source: [PR #1687](vllm-project/vllm-omni#1687) - [BugFix] Return proper HTTP status for ErrorResponse in create_speech - Changes: - Bug fix: [BugFix] Return proper HTTP status for ErrorResponse in create_speech ### vllm-omni-cicd - Source: [PR #1683](vllm-project/vllm-omni#1683) - [CI] Remove high concurrency tests before issue #1374 fixed. - Changes: - Bug fix: [CI] Remove high concurrency tests before issue #1374 fixed. ### vllm-omni-audio-tts - Source: [PR #1678](vllm-project/vllm-omni#1678) - Add non-async chunk support for Qwen3-TTS - Changes: - New feature: Add non-async chunk support for Qwen3-TTS ### vllm-omni-cicd - Source: [PR #1678](vllm-project/vllm-omni#1678) - Add non-async chunk support for Qwen3-TTS - Changes: - New feature: Add non-async chunk support for Qwen3-TTS ### vllm-omni-cicd - Source: [PR #1677](vllm-project/vllm-omni#1677) - Replace hard-coded cuda generator with current_omni_platform.device_type ### vllm-omni-perf - Source: [PR #1677](vllm-project/vllm-omni#1677) - Replace hard-coded cuda generator with current_omni_platform.device_type ### vllm-omni-serving - Source: [PR #1675](vllm-project/vllm-omni#1675) - [Misc] remove logits_processor_pattern this field, because vllm have … ### vllm-omni-cicd - Source: [PR #1666](vllm-project/vllm-omni#1666) - [Cleanup] Move cosyvoice3 tests to model subdirectory ### vllm-omni-audio-tts - Source: [PR #1664](vllm-project/vllm-omni#1664) - [Bugfix] Fix all-silence TTS output: use float32 for speech tokenizer decoder - Changes: - Bug fix: [Bugfix] Fix all-silence TTS output: use float32 for speech tokenizer decoder ### vllm-omni-cicd - Source: [PR #1664](vllm-project/vllm-omni#1664) - [Bugfix] Fix all-silence TTS output: use float32 for speech tokenizer decoder - Changes: - Bug fix: [Bugfix] Fix all-silence TTS output: use float32 for speech tokenizer decoder ### vllm-omni-distributed - Source: [PR #1656](vllm-project/vllm-omni#1656) - [Optimize][Qwen3-Omni] Reduce inter-packet latency in async chunk ### vllm-omni-contrib - Source: [PR #1656](vllm-project/vllm-omni#1656) - [Optimize][Qwen3-Omni] Reduce inter-packet latency in async chunk ### vllm-omni-quantization - Source: [PR #1652](vllm-project/vllm-omni#1652) - [UX] Add progress bar for diffusion models - Changes: - New feature: [UX] Add progress bar for diffusion models ### vllm-omni-perf - Source: [PR #1652](vllm-project/vllm-omni#1652) - [UX] Add progress bar for diffusion models - Changes: - New feature: [UX] Add progress bar for diffusion models ### vllm-omni-distributed - Source: [PR #1651](vllm-project/vllm-omni#1651) - docs: Announce vllm-omni-skills community project ### vllm-omni-quantization - Source: [PR #1651](vllm-project/vllm-omni#1651) - docs: Announce vllm-omni-skills community project ### vllm-omni-perf - Source: [PR #1651](vllm-project/vllm-omni#1651) - docs: Announce vllm-omni-skills community project ### vllm-omni-contrib - Source: [PR #1649](vllm-project/vllm-omni#1649) - [Misc] update wechat ### vllm-omni-perf - Source: [PR #1642](vllm-project/vllm-omni#1642) - [chore] add _repeated_blocks for regional compilation support - Changes: - New feature: [chore] add _repeated_blocks for regional compilation support ### vllm-omni-api - Source: [PR #1641](vllm-project/vllm-omni#1641) - [Bugfix] Add TTS request validation to prevent engine crashes - Changes: - New feature: [Bugfix] Add TTS request validation to prevent engine crashes ### vllm-omni-cicd - Source: [PR #1641](vllm-project/vllm-omni#1641) - [Bugfix] Add TTS request validation to prevent engine crashes - Changes: - New feature: [Bugfix] Add TTS request validation to prevent engine crashes ### vllm-omni-image-gen - Source: [PR #1640](vllm-project/vllm-omni#1640) - [FP8 Quantization] Add FP8 quantization support for Flux transformer - Changes: - New feature: [FP8 Quantization] Add FP8 quantization support for Flux transformer - Additions: - text-to-image - Text-to-Image - Flux ### vllm-omni-quantization - Source: [PR #1640](vllm-project/vllm-omni#1640) - [FP8 Quantization] Add FP8 quantization support for Flux transformer - Changes: - New feature: [FP8 Quantization] Add FP8 quantization support for Flux transformer - Additions: - FP8 support or improvements ### vllm-omni-contrib - Source: [PR #1640](vllm-project/vllm-omni#1640) - [FP8 Quantization] Add FP8 quantization support for Flux transformer - Changes: - New feature: [FP8 Quantization] Add FP8 quantization support for Flux transformer ### vllm-omni-perf - Source: [PR #1640](vllm-project/vllm-omni#1640) - [FP8 Quantization] Add FP8 quantization support for Flux transformer - Changes: - New feature: [FP8 Quantization] Add FP8 quantization support for Flux transformer ### vllm-omni-contrib - Source: [PR #1631](vllm-project/vllm-omni#1631) - [BugFix] Fix LongCat Sequence Parallelism / Small Cleanup - Changes: - Bug fix: [BugFix] Fix LongCat Sequence Parallelism / Small Cleanup ### vllm-omni-cicd - Source: [PR #1628](vllm-project/vllm-omni#1628) - [Test][Qwen3-Omni]Modify Qwen3-Omni benchmark test cases ### vllm-omni-perf - Source: [PR #1628](vllm-project/vllm-omni#1628) - [Test][Qwen3-Omni]Modify Qwen3-Omni benchmark test cases ### vllm-omni-perf - Source: [PR #1619](vllm-project/vllm-omni#1619) - [Bugfix] Fix Qwen3-TTS code predictor crash due to missing vLLM config context - Changes: - Bug fix: [Bugfix] Fix Qwen3-TTS code predictor crash due to missing vLLM config context ### vllm-omni-perf - Source: [PR #1617](vllm-project/vllm-omni#1617) - [Refactor][Perf] Qwen3-TTS: re-prefill Code Predictor with torch.compile + enable Code2Wav decoder CUDA Graph - Changes: - Performance improvement: [Refactor][Perf] Qwen3-TTS: re-prefill Code Predictor with torch.compile + enable Code2Wav decoder CUDA Graph ### vllm-omni-contrib - Source: [PR #1615](vllm-project/vllm-omni#1615) - [Doc] Fix links in the configuration doc - Changes: - Bug fix: [Doc] Fix links in the configuration doc ### vllm-omni-audio-tts - Source: [PR #1614](vllm-project/vllm-omni#1614) - perf: replace per-element .item() GPU syncs with batch .tolist() in TTS code predictor - Changes: - Performance improvement: perf: replace per-element .item() GPU syncs with batch .tolist() in TTS code predictor ### vllm-omni-perf - Source: [PR #1614](vllm-project/vllm-omni#1614) - perf: replace per-element .item() GPU syncs with batch .tolist() in TTS code predictor - Changes: - Performance improvement: perf: replace per-element .item() GPU syncs with batch .tolist() in TTS code predictor ### vllm-omni-image-gen - Source: [PR #1609](vllm-project/vllm-omni#1609) - [Bugfix] Fix filepath resolution for model with subdir and GLM-Image generation - Changes: - Bug fix: [Bugfix] Fix filepath resolution for model with subdir and GLM-Image generation - Additions: - GLM-Image - GLM-Image - GLM-Image - GLM-Image - GLM-Image - GLM-Image - GLM-Image - GLM-Image ### vllm-omni-api - Source: [PR #1609](vllm-project/vllm-omni#1609) - [Bugfix] Fix filepath resolution for model with subdir and GLM-Image generation - Changes: - Bug fix: [Bugfix] Fix filepath resolution for model with subdir and GLM-Image generation ### vllm-omni-perf - Source: [PR #1609](vllm-project/vllm-omni#1609) - [Bugfix] Fix filepath resolution for model with subdir and GLM-Image generation - Changes: - Bug fix: [Bugfix] Fix filepath resolution for model with subdir and GLM-Image generation ### vllm-omni-contrib - Source: [PR #1604](vllm-project/vllm-omni#1604) - [Model]: support Helios from ByteDance ### vllm-omni-perf - Source: [PR #1604](vllm-project/vllm-omni#1604) - [Model]: support Helios from ByteDance ### vllm-omni-serving - Source: [PR #1602](vllm-project/vllm-omni#1602) - [Bugfix] fix kernel error for qwen3-omni - Changes: - Bug fix: [Bugfix] fix kernel error for qwen3-omni ### vllm-omni-distributed - Source: [PR #1598](vllm-project/vllm-omni#1598) - [BugFix] Fix load_weights error when loading HunyuanImage3.0 - Changes: - Bug fix: [BugFix] Fix load_weights error when loading HunyuanImage3.0 ### vllm-omni-image-gen - Source: [PR #1598](vllm-project/vllm-omni#1598) - [BugFix] Fix load_weights error when loading HunyuanImage3.0 - Changes: - Bug fix: [BugFix] Fix load_weights error when loading HunyuanImage3.0 - Additions: - HunyuanImage3 - HunyuanImage3Pipeline - HunyuanImage3 - HunyuanImage-3 - HunyuanImage-3 - HunyuanImage-3 - HunyuanImage3Pipeline - HunyuanImage3Pipeline - HunyuanImage3Pipeline - HunyuanImage3Pipeline - HunyuanImage3Pipeline - HunyuanImage3Pipeline - HunyuanImage3Pipeline - HunyuanImage3Pipeline - HunyuanImage-3 ### vllm-omni-quantization - Source: [PR #1598](vllm-project/vllm-omni#1598) - [BugFix] Fix load_weights error when loading HunyuanImage3.0 - Changes: - Bug fix: [BugFix] Fix load_weights error when loading HunyuanImage3.0 ### vllm-omni-perf - Source: [PR #1598](vllm-project/vllm-omni#1598) - [BugFix] Fix load_weights error when loading HunyuanImage3.0 - Changes: - Bug fix: [BugFix] Fix load_weights error when loading HunyuanImage3.0 ### vllm-omni-audio-tts - Source: [PR #1583](vllm-project/vllm-omni#1583) - [Feat][Qwen3TTS] reduce TTFA with flexible initial phase - Changes: - New feature: [Feat][Qwen3TTS] reduce TTFA with flexible initial phase ### vllm-omni-api - Source: [PR #1583](vllm-project/vllm-omni#1583) - [Feat][Qwen3TTS] reduce TTFA with flexible initial phase - Changes: - New feature: [Feat][Qwen3TTS] reduce TTFA with flexible initial phase ### vllm-omni-cicd - Source: [PR #1583](vllm-project/vllm-omni#1583) - [Feat][Qwen3TTS] reduce TTFA with flexible initial phase - Changes: - New feature: [Feat][Qwen3TTS] reduce TTFA with flexible initial phase ### vllm-omni-contrib - Source: [PR #1583](vllm-project/vllm-omni#1583) - [Feat][Qwen3TTS] reduce TTFA with flexible initial phase - Changes: - New feature: [Feat][Qwen3TTS] reduce TTFA with flexible initial phase ### vllm-omni-api - Source: [PR #1579](vllm-project/vllm-omni#1579) - [1/N][Refactor] Clean up dead code in output processor ### vllm-omni-serving - Source: [PR #1579](vllm-project/vllm-omni#1579) - [1/N][Refactor] Clean up dead code in output processor ### vllm-omni-distributed - Source: [PR #1578](vllm-project/vllm-omni#1578) - [Feature][Bagel] Add CFG parallel mode - Changes: - New feature: [Feature][Bagel] Add CFG parallel mode ### vllm-omni-cicd - Source: [PR #1578](vllm-project/vllm-omni#1578) - [Feature][Bagel] Add CFG parallel mode - Changes: - New feature: [Feature][Bagel] Add CFG parallel mode ### vllm-omni-perf - Source: [PR #1578](vllm-project/vllm-omni#1578) - [Feature][Bagel] Add CFG parallel mode - Changes: - New feature: [Feature][Bagel] Add CFG parallel mode ### vllm-omni-contrib - Source: [PR #1576](vllm-project/vllm-omni#1576) - 0.16.0 release ### vllm-omni-audio-tts - Source: [PR #1570](vllm-project/vllm-omni#1570) - [bugfix] Fix unexpected argument 'is_finished' in function llm2code2wav_async_chunk of mimo-audio - Changes: - Bug fix: [bugfix] Fix unexpected argument 'is_finished' in function llm2code2wav_async_chunk of mimo-audio ### vllm-omni-api - Source: [PR #1566](vllm-project/vllm-omni#1566) - [Bugfix] Import InputPreprocessor into Renderer - Changes: - Bug fix: [Bugfix] Import InputPreprocessor into Renderer ### vllm-omni-distributed - Source: [PR #1539](vllm-project/vllm-omni#1539) - [Debug] Enable curl retry aligned with openai ### vllm-omni-quantization - Source: [PR #1539](vllm-project/vllm-omni#1539) - [Debug] Enable curl retry aligned with openai ### vllm-omni-perf - Source: [PR #1539](vllm-project/vllm-omni#1539) - [Debug] Enable curl retry aligned with openai ### vllm-omni-image-gen - Source: [PR #1537](vllm-project/vllm-omni#1537) - [NPU] [Features] [Bugfix] Support mindiesd adaln - Changes: - New feature: [NPU] [Features] [Bugfix] Support mindiesd adaln - Additions: - mindiesd - mindiesd - Qwen-Image-Edit-2509 - mindiesd - mindiesd - mindiesd - mindiesd ### vllm-omni-perf - Source: [PR #1537](vllm-project/vllm-omni#1537) - [NPU] [Features] [Bugfix] Support mindiesd adaln - Changes: - New feature: [NPU] [Features] [Bugfix] Support mindiesd adaln ### vllm-omni-serving - Source: [PR #1536](vllm-project/vllm-omni#1536) - [Bugfix] Fix transformers 5.x compat issues in online TTS serving - Changes: - Bug fix: [Bugfix] Fix transformers 5.x compat issues in online TTS serving ### vllm-omni-perf - Source: [PR #1536](vllm-project/vllm-omni#1536) - [Bugfix] Fix transformers 5.x compat issues in online TTS serving - Changes: - Bug fix: [Bugfix] Fix transformers 5.x compat issues in online TTS serving

Signed-off-by: linyueqian <linyueqian@outlook.com> Signed-off-by: lishunyang <lishunyang12@163.com>

Add non-async chunk support and benchmark scripts for Qwen3-TTS

bc935bd

Signed-off-by: linyueqian <linyueqian@outlook.com>

linyueqian requested a review from hsliuustc0106 as a code owner March 5, 2026 06:30

chatgpt-codex-connector Bot reviewed Mar 5, 2026

View reviewed changes

Filter zero-padded frames in non-async talker2code2wav

71607a5

Signed-off-by: linyueqian <linyueqian@outlook.com>

hsliuustc0106 reviewed Mar 5, 2026

View reviewed changes

Add nightly e2e tests for Qwen3-TTS non-async-chunk mode

f066eae

Signed-off-by: linyueqian <linyueqian@outlook.com>

linyueqian added 2 commits March 5, 2026 13:13

Fix E402 lint: add future annotations import

a45c8bc

Signed-off-by: linyueqian <linyueqian@outlook.com>

Fix E402: move OmniTokensPrompt import inside function

fc6aec9

Signed-off-by: linyueqian <linyueqian@outlook.com>

linyueqian added ready label to trigger buildkite CI labels Mar 5, 2026

hsliuustc0106 approved these changes Mar 6, 2026

View reviewed changes

hsliuustc0106 merged commit 3bc89de into vllm-project:main Mar 6, 2026
7 checks passed

linyueqian mentioned this pull request Mar 10, 2026

[RFC]: TTS Development Roadmap - March 2026 #1795

Open

lishunyang12 pushed a commit to lishunyang12/vllm-omni that referenced this pull request Mar 11, 2026

Add non-async chunk support for Qwen3-TTS (vllm-project#1678)

178c117

Signed-off-by: linyueqian <linyueqian@outlook.com> Signed-off-by: lishunyang <lishunyang12@163.com>

linyueqian mentioned this pull request Mar 11, 2026

[Bug]: Qwen3-TTS generations hang indefinitely after the first few requests #1525

Closed

1 task

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Add non-async chunk support for Qwen3-TTS#1678

Add non-async chunk support for Qwen3-TTS#1678
hsliuustc0106 merged 5 commits intovllm-project:mainfrom
linyueqian:feat/qwen3-tts-non-async-chunk

linyueqian commented Mar 5, 2026 •

edited

Loading

Uh oh!

chatgpt-codex-connector Bot left a comment

Uh oh!

chatgpt-codex-connector Bot Mar 5, 2026

Uh oh!

hsliuustc0106 Mar 5, 2026

Uh oh!

Sy0307 commented Mar 5, 2026

Uh oh!

Sy0307 commented Mar 5, 2026

Uh oh!

hsliuustc0106 commented Mar 5, 2026

Uh oh!

linyueqian commented Mar 5, 2026

Uh oh!

hsliuustc0106 left a comment

Uh oh!

hsliuustc0106 Mar 6, 2026

Uh oh!

hsliuustc0106 Mar 6, 2026

Uh oh!

hsliuustc0106 Mar 6, 2026

Uh oh!

hsliuustc0106 Mar 6, 2026

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

3 participants

Conversation

linyueqian commented Mar 5, 2026 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Summary

How to test

Uh oh!

chatgpt-codex-connector Bot left a comment

Choose a reason for hiding this comment

💡 Codex Review

Uh oh!

chatgpt-codex-connector Bot Mar 5, 2026

Choose a reason for hiding this comment

Uh oh!

hsliuustc0106 Mar 5, 2026

Choose a reason for hiding this comment

Uh oh!

Sy0307 commented Mar 5, 2026

Uh oh!

Sy0307 commented Mar 5, 2026

Uh oh!

hsliuustc0106 commented Mar 5, 2026

Uh oh!

linyueqian commented Mar 5, 2026

Uh oh!

hsliuustc0106 left a comment

Choose a reason for hiding this comment

Review

Summary

Highlights

Minor Issues (non-blocking)

Recommendation

Uh oh!

hsliuustc0106 Mar 6, 2026

Choose a reason for hiding this comment

Uh oh!

hsliuustc0106 Mar 6, 2026

Choose a reason for hiding this comment

Uh oh!

hsliuustc0106 Mar 6, 2026

Choose a reason for hiding this comment

Uh oh!

hsliuustc0106 Mar 6, 2026

Choose a reason for hiding this comment

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

3 participants

linyueqian commented Mar 5, 2026 •

edited

Loading