Repository navigation
[Frontend] Opt-in WebSocket TTS split_granularity and session seed - #7046
Conversation
|
This PR appears to belong to: docs/design/module/entrypoints.md. Module owners: @alex-jw-brooks @linyueqian @NickCao @rk9595, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Omni ReviewBot triage noteAutomated triage of commit
These are automated triage suggestions only — the final decision belongs to the maintainers. |
hsliuustc0106
left a comment
There was a problem hiding this comment.
Review findings:
- The splitter mishandles punctuation runs/closing quotes, abbreviations, and long incremental buffers.
- Session configuration can change after a sentence has already been emitted but before input.done, producing mixed settings within one utterance.
GitHub currently reports this PR as non-mergeable; a local three-way merge also conflicts in the docs and stream test file. py_compile and whitespace checks pass, but targeted pytest collection was blocked by the environment's incompatible Torch/NumPy installation.
|
|
||
| if msg_type == "session.config": | ||
| if text_parts: | ||
| if splitter.has_buffered_text(): |
There was a problem hiding this comment.
[P1] Reject reconfiguration after a split unit has already been emitted. In sentence/clause mode, feed() can return a unit while leaving _buf empty, so this check allows session.config before input.done. The next unit then uses different voice/settings but keeps the same utterance_index and total_sentences; track an active flush separately and reset it only after input.done.
|
|
||
| latin = ch in _LATIN_TERMINATORS | ||
| if latin: | ||
| complete = j > i + 1 or flush |
There was a problem hiding this comment.
[P1] Preserve punctuation runs and closing delimiters when flushing. Because flush makes every terminator complete independently, Wait... becomes three units and He said "Hello." leaves the quote as a separate unit. That causes extra TTS requests and malformed text for common input.done payloads; consume the punctuation/closing-delimiter run before emitting one unit.
|
|
||
| from __future__ import annotations | ||
|
|
||
| # ASCII `.!?` are abbreviation-prone; they need following whitespace (or a |
There was a problem hiding this comment.
[P2] Add actual abbreviation handling. Requiring whitespace after . does not distinguish sentence endings from Dr. Smith, e.g. this, or initials, so these are emitted as separate TTS requests despite this comment identifying abbreviations as a concern.
| terminators = self._terminators() | ||
| if terminators is None: | ||
| return [] | ||
| units, self._buf = extract_complete_units(self._buf, terminators, flush=False) |
There was a problem hiding this comment.
[P2] Avoid rescanning the entire unfinished buffer on every input.text message. A long punctuation-free sentence sent one token/word at a time makes extract_complete_units O(n^2) and can consume substantial CPU; retain a scan cursor/state for incremental parsing.
Restore per-sentence/clause synthesis as an explicit session option so STT/LLM clients can start audio before input.done, without changing the default one-request-per-flush path from vllm-project#4731. Honor seed on the WS speech session and document Indic/CJK/Arabic boundaries. Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu> Co-authored-by: Cursor <cursoragent@cursor.com>
Splitter: - Absorb punctuation runs and closing delimiters into the unit they close, so `Wait...` and `He said "Hello."` stay one TTS request instead of three. - Skip abbreviations (`Dr.`, `e.g.`), initials (`J. R.`) and thousands separators (`1,000`) as boundaries, alongside the existing decimal check. - Require a known next character before closing a unit, and keep a scan cursor across feeds so a long punctuation-free sentence arriving token by token is scanned once instead of rescanned on every input.text (O(n^2)). Handler: - Track the open utterance explicitly instead of inferring it from buffered text: in sentence/clause mode feed() can empty the buffer after emitting a unit, which let a session.config land mid-utterance and synthesize the rest of the same utterance_index in a different voice. Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu> Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QZN6qW9eif4sM11xHCqSMa
feba351 to
4f9cad2
Compare
Self-reviewRebased onto [P1] [P1] Punctuation runs and closing delimiters. The splitter now absorbs the run of terminators and closing delimiters that follows a boundary, so [P2] Abbreviations. Added a real check: the token before a [P2] O(n²) rescanning. What I checked
Not addressed: nothing from the review. Remaining known limitation is that abbreviation detection is a fixed list plus a single-letter rule, not locale-aware — over-suppression at worst merges two sentences into one request, and |
linyueqian
left a comment
There was a problem hiding this comment.
The default buffering behavior and session-seed forwarding look correct. In sentence and clause mode, abbreviation detection can repeatedly rescan the full prefix of one accepted text frame, making the synchronous splitter quadratic and able to stall the shared WebSocket event loop. That needs a bounded scan before merge; the Arabic semicolon mismatch is only a small follow-up.
| if buffer[index] != ".": | ||
| return False | ||
| start = index | ||
| while start > 0 and (buffer[start - 1].isalnum() or buffer[start - 1] == "."): |
There was a problem hiding this comment.
[important] Bound this backward scan or track the current token start. For input such as a.a.a.a... without whitespace, every period walks the entire preceding prefix, making the splitter quadratic. One accepted WebSocket text frame can approach 128 KiB, so sentence or clause mode can perform billions of synchronous character checks and stall every session sharing this event loop.
| "؟", # Arabic question mark | ||
| } | ||
| ) | ||
| _SCRIPT_CLAUSE_TERMINATORS = _SCRIPT_SENTENCE_TERMINATORS | { |
There was a problem hiding this comment.
[suggestion] Add U+061B ARABIC SEMICOLON (؛) here or correct the module contract. The docstring says Arabic semicolons are covered, but clause mode includes the Arabic comma and question mark only, so text ending at ؛ stays buffered until another boundary or input.done.
There was a problem hiding this comment.
Done in e4c89b3: ؛ U+061B is now in _SCRIPT_CLAUSE_TERMINATORS, so it also flows into the run-absorption set. The split_granularity row in docs/serving/speech_api.md omitted it too and is corrected; covered by test_clause_mode_splits_on_arabic_semicolon.
The abbreviation check added in 4f9cad2 walked back over the whole alnum/dot prefix at every '.', so one accepted WebSocket frame of 'a.a.a...' was quadratic and could stall the shared event loop. The walk is now capped at the longest known abbreviation: hitting the cap on a token character can only rule a token out, never produce a false positive, so splitting behaviour is unchanged for real text. Clause mode now also splits on U+061B ARABIC SEMICOLON, which the module docstring already claimed; the docs table is corrected to match. The regression test counts character checks rather than wall-clock time so it cannot flake on a loaded CI runner. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RT6EQ5o56Rc4B3igw2GQ98 Signed-off-by: Rakesh Kariya <rakesh.kariya@somaiya.edu>
Round 2Pushed [important] Unbounded abbreviation lookback. Correct, and a regression from the round-1 abbreviation fix. [suggestion] Arabic semicolon. Added Checked: 18 splitter tests pass standalone ( @hsliuustc0106 your round-1 review is still marked changes-requested; all four findings were addressed in |
linyueqian
left a comment
There was a problem hiding this comment.
The abbreviation lookback is now bounded, and the Arabic semicolon handling matches the docs. Both review findings are fixed, with regression tests for the long-token case, the scan bound, and Arabic clause splitting.
CI at this head. The lanes had never run when this was first reviewed, because the PR carried no ready label; with it added, general build 14829, Intel 7953, and NPU 6570 are green. AMD build 11553 hard-failed exactly one step, mi300_1: Diffusion · Model Test, with its other 13 steps passing.
That AMD failure is inherited rather than caused here, and today gives it unusually strong corroboration: five unrelated PRs hit the identical single step in builds 11556, 11558, 11561, 11562, and 11563, across frontend, TTS, and refactor changes with disjoint file sets. It is the known SenseNova paged-decode failure on a ROCm image lacking the CUDA flash-attention extensions. This PR adds opt-in WebSocket TTS split_granularity and touches no diffusion code.
Validation: static review of the delta at e4c89b34, the abbreviation scan bound, the clause-splitting paths, and the new regression tests, plus public Buildkite step metadata for all four lanes. Codex and Claude agree both findings are resolved. I did not execute the fork's code.
…llm-project#7046) Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu> Signed-off-by: Rakesh Kariya <rakesh.kariya@somaiya.edu> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…llm-project#7046) Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu> Signed-off-by: Rakesh Kariya <rakesh.kariya@somaiya.edu> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: wenjie.yan <wenjyan@outlook.com>
* [Bugfix][Examples] Use --profiler-config flag in offline TTS examples (vllm-project#6763) Signed-off-by: Asthenia <asthenia0412@gmail.com> Co-authored-by: Asthenia <asthenia0412@gmail.com> * [Bugfix] Skip HWR store-size scans when no limit is configured (vllm-project#7131) Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [CI][ROCm] Route LTX2 Ulysses parity to two-GPU lane (vllm-project#7234) Signed-off-by: andyluo7 <andy.luo@amd.com> * [Bugfix][Model] GR00T-N1.7: honor the per-request seed for flow-matching noise (vllm-project#7253) Signed-off-by: liangmengh <liangmengh@nvidia.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> * Add vLLM-Omni library info to Hugging Face Hub requests (vllm-project#5381) Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * [Bugfix][NPU] Limit MiniMax H3 modulation grid size (vllm-project#6794) Signed-off-by: KrystalRay <keeleiray@gmail.com> Co-authored-by: KrystalRay <keeleiray@gmail.com> * [Bugfix] Build the forced-aligner prompt without a chat template (word timestamps one bin late) (vllm-project#7240) Signed-off-by: Tianyao Wu <rayroy31@gmail.com> * [Refactor][Diffusion] Resolve offload topology through one plan resolver (vllm-project#7209) Signed-off-by: specture724 <specture724@gmail.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * [Doc] Add AI usage policy for contributions (vllm-project#7305) Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com> * [Bugfix][MiMo-Audio] Align code2wav decode with tokenizer device (vllm-project#6539) Signed-off-by: chaosansui <zzc15560846421@163.com> Signed-off-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com> * [Bugfix][MiniCPM-o] Fix the audio_embeds input path (vllm-project#5730) Signed-off-by: eval-dev <0xe5bca0@gmail.com> Signed-off-by: eval <74645252+eval-dev@users.noreply.github.com> * [Feat][OmniVoice]Support Varlen Attn, Request-Batch and Step-Execution (vllm-project#6408) Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com> * [Model] Add Audio8 TTS Preview 0.6B (DualAR, 44.1 kHz codec) (vllm-project#6157) Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com> Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com> * [Bugfix][Frontend] Accept the msgpack-numpy package's numpy markers on the OpenPI endpoint (vllm-project#6051) Signed-off-by: zjli2013 <leezhengjiang@126.com> Co-authored-by: Cursor <cursoragent@cursor.com> * [Frontend] Opt-in WebSocket TTS split_granularity and session seed (vllm-project#7046) Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu> Signed-off-by: Rakesh Kariya <rakesh.kariya@somaiya.edu> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * [Bugfix][Frontend] Clear the P0 multimodal cache through the renderer (vllm-project#7003) Signed-off-by: ZenAlexa <zimingwang945@gmail.com> * [Bugfix][Frontend] Enforce image pixel limits for video input references (vllm-project#6963) Signed-off-by: BANANASJIM <bananasjim1@gmail.com> * [Bugfix][TTS] Isolate shared Higgs v3 reference encode from request cancellation (vllm-project#7076) Signed-off-by: Allen Wu <allenwu2795@gmail.com> Co-authored-by: TRAE CLI <traecli@bytedance.com> * [Bugfix][CosyVoice3] Resolve hash snapshot pipeline (vllm-project#6896) Signed-off-by: xutianle <xutianle@fudan.edu.cn> * [CI] Skip Qwen3-Omni Server VAD multi-turn realtime test (vllm-project#7279) (vllm-project#7314) Signed-off-by: wangyu <410167048@qq.com> * [Bugfix][Magi2] Allow import without an active Triton driver (vllm-project#7239) Signed-off-by: andyluo7 <andy.luo@amd.com> * [Core] Split Omni connector model runner mixin (vllm-project#6903) Signed-off-by: natureofnature <wzliu@connect.hku.hk> * [Bugfix] Make LTX vocoder decoding deterministic (vllm-project#7231) Signed-off-by: mglyn <1203789601@qq.com> * [Doc] [Recipe] Add FLUX.1-schnell recipe for RTX 5090 32GB (vllm-project#7299) Signed-off-by: Sparks-M <41097544+Sparks-M@users.noreply.github.com> * [Doc] Qwen3-TTS: add 0.6B on 1x A100 40GB (vllm-project#7289) Signed-off-by: chi030303 <106855944+chi030303@users.noreply.github.com> * [Perf][Model] Add optimized LTX-2.5 DiffVAE operators (vllm-project#7308) Signed-off-by: mglyn <1203789601@qq.com> * [2/N] Add a minimal temporal chunk callback for MiniMax-H3 (vllm-project#7017) Signed-off-by: specture724 <specture724@gmail.com> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> * [Feature][Diffusion] Expose detailed pipeline timings (vllm-project#6822) Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com> * [Bugfix] Resolve vllm-project#6931 hub FA3 on torch 2.13 via kernels 0.16.1 (vllm-project#7185) Signed-off-by: NumberWan <wantszkin2003@gmail.com> * [Bugfix][Ascend] fix npu 310/a5 bugs (vllm-project#6685) Signed-off-by: zouyizhou <zouyizhou@huawei.com> * [Bugfix][Engine] Group overlapping device stages into one sequential init component (vllm-project#7328) Signed-off-by: ZhengWG <zwg0606@gmail.com> Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com> * fix: reserve Qwen3-Omni NVFP4 backend fix (vllm-project#7200) Signed-off-by: kunkunblueberry <1833921874@qq.com> * [BugFix] Add field validators for /v1/audio/generate request (vllm-project#4741) Signed-off-by: Shaun Walsh <shaunwalsh24@gmail.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> Co-authored-by: Nick Cao <ncao@redhat.com> * [CI][ROCm] Match CUDA/NPU L2/L3 label routing (vllm-project#6966) Signed-off-by: andyluo7 <andy.luo@amd.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [CI/Build] Avoid duplicate stage CLI deploy config (vllm-project#7007) Signed-off-by: mershi <mershi@tencent.com> Co-authored-by: mershi <mershi@tencent.com> * [CI/Build][ROCm] Normalize SenseNova paged-decode hardware markers (vllm-project#6935) Signed-off-by: andyluo7 <andy.luo@amd.com> * [Model] Skip unused frame packing in Wan2.2 S2V (vllm-project#7155) Signed-off-by: hyw <yuweih205@gmail.com> * [Doc] Add dual DGX Spark MiniMax-H3 results (vllm-project#7343) Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com> * [Model] Optimize MOSS-TTS Local batched execution and streaming codec (vllm-project#7202) Signed-off-by: Sy03 <1370724210@qq.com> * [Bugfix][XPU] Restore N-D output shape for W8A16 FP8 linear (vllm-project#7301) Signed-off-by: Joshna Medisetty <joshna.medisetty@intel.com> Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com> * [Doc] Document num_outputs_per_prompt for /v1/videos (vllm-project#7341) Signed-off-by: Guangjian <hiro20833@gmail.com> * [Skills] Add perf-evidence isolation, stage-attribution, and realtime-contract requirements (vllm-project#6820) Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com> Co-authored-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com> * [Bugfix] Allow LLM replicas on different GPUs to initialize concurrently (vllm-project#7292) Signed-off-by: Gao Han <hgaoaf@connect.ust.hk> Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com> * [CI/Build] Stabilize LTX2 vocoder autocast test on ROCm (vllm-project#7336) Signed-off-by: andyluo7 <andy.luo@amd.com> * [NPU][CI] Add A5 and 310P CI support (vllm-project#6875) Signed-off-by: Weiming Liao <liaowm5@gmail.com> Co-authored-by: wangyu <53896905+yenuo26@users.noreply.github.com> * [Kernel] Enable LTX DiffVAE fusions on SM100 and SM103 (vllm-project#7350) Signed-off-by: mglyn <1203789601@qq.com> * [Bugfix][MiniCPM-o] Align structured chat content with native omni rendering (vllm-project#7344) Signed-off-by: Sy03 <1370724210@qq.com> * [Rebase] Rebase to vLLM 0.29.0 (vllm-project#7230) Signed-off-by: tzhouam <tzhouam@connect.ust.hk> Signed-off-by: Zhou Taichang <tzhouam@connect.ust.hk> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * [Refactor] P0.2: Migrate API server helpers out of api_server (vllm-project#5453) Signed-off-by: herotai214 <herotai214@gmail.com> * [CI] Stabilize Qwen3-Omni Server VAD E2E (vllm-project#7356) Signed-off-by: LHXuuu <xulianhao.xlh@antgroup.com> Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com> * [CI/Build] Diff-aware source_file_dependencies for CUDA/NPU pipelines (vllm-project#6597) Signed-off-by: wangyu <410167048@qq.com> Co-authored-by: Cursor <cursoragent@cursor.com> * [Core][Diffusion] Add a typed pre-D2H video media contract (vllm-project#6615) Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com> Signed-off-by: Samit <285365963@qq.com> Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com> Co-authored-by: Samit <285365963@qq.com> * [Bugfix] Bound HWR domain initialization lock waits (vllm-project#7128) Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Bugfix] Escalate diffusion worker shutdown and retain survivors (vllm-project#7126) Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Misc] Add standalone safetensors retention diagnostic (vllm-project#7145) Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [CI] Isolate layerwise offload memory measurements (vllm-project#6938) Signed-off-by: andyluo7 <andy.luo@amd.com> * [Model] Add Cosmos3 mixed W8A8/W8A16 and W4A4/W4A16 denoising (vllm-project#6560) Signed-off-by: Rahul Steiger <rsteiger@aws-cmh-slurm-1-vscode-04.cm.cluster> Signed-off-by: Wojciech Kutak <wkutak@nvidia.com> Co-authored-by: Rahul Steiger <rsteiger@nvidia.com> * [Test] Use public render_jinja_template in MiniCPM-o native template test (vllm-project#7362) Signed-off-by: tly <2200895168@qq.com> * [Bugfix] Fix video prewarm cache retention and cancel-restart delay (vllm-project#7363) Signed-off-by: psv666 <2693925048@qq.com> * Cosmos3 action policy improvements (vllm-project#6460) Signed-off-by: Maciej Bala <mbala@nvidia.com> Signed-off-by: MaciejBalaNV <mbala@nvidia.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [BugFix][CI] Restore diff-aware source filtering for post-merge L3 (vllm-project#7371) Signed-off-by: wangyu <410167048@qq.com> * [Bugfix] Fail when a diffusion LoRA adapter binds no layer (vllm-project#7349) Signed-off-by: Guangjian <hiro20833@gmail.com> * [Bugfix] Fix host-memory leak on aborted /v1/images/generations (vllm-project#6462) (vllm-project#6561) Signed-off-by: summer <128961079+zhang-keliang@users.noreply.github.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Refactor] Declare model-local KV held outside the paged manager (vllm-project#6171) Signed-off-by: Yueqian Lin <linyueqian@outlook.com> * [Realtime] Emit current (non-beta) OpenAI audio/transcript event names (vllm-project#7339) Signed-off-by: Nick Cao <ncao@redhat.com> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> * [Bugfix][Core] Clean up failed HWR atomic metadata writes (vllm-project#6956) Signed-off-by: BANANASJIM <bananasjim1@gmail.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Bugfix] Keep MiniMax-H3 reference audio budgets separate (vllm-project#7281) Signed-off-by: david6666666 <530634352@qq.com> * [Bugfix] Fix Helios USP: per-component split for correct sequence parallelism (vllm-project#6930) Signed-off-by: yancaocn <yancaochn@163.com> Co-authored-by: yancaocn <yancaochn@163.com> * [Perf][Diffusion] Optimize HSDP startup via Rank-0 shared weight loading and accelerated LoRA delta computation (vllm-project#7005) Signed-off-by: samithuang <285365963@qq.com> * [Example] Migrate HunyuanImage-3.0 to model_extras + shared task examples (vllm-project#5559) Signed-off-by: suyanli220 <suyanli220@gmail.com> Signed-off-by: suyan.li <suyan.li@bytedance.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: suyan.li <suyan.li@bytedance.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Model] Avoid scalar synchronizations in GLM-Image preparation (vllm-project#7172) Signed-off-by: hyw <yuweih205@gmail.com> * [Model][ERNIE-Image] Delay AdaLN modulation broadcast (vllm-project#7171) Signed-off-by: hyw <yuweih205@gmail.com> * [Kernel][MiniMax-H3] Run Q/K RMSNorm-RoPE in one launch (vllm-project#7167) Signed-off-by: hyw <yuweih205@gmail.com> * [CI][ROCm] Align AMD image with vLLM 0.29 (vllm-project#7395) Signed-off-by: andyluo7 <andy.luo@amd.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Bugfix] Add embed_multimodal to MiniCPM-o 4.5 omni LLM class (vllm-project#7384) Signed-off-by: Guangjian <hiro20833@gmail.com> * [Model] Add LingBot World Ulysses sequence parallelism (vllm-project#6841) Signed-off-by: wtz2333 <2955110911@qq.com> Co-authored-by: Zhou Taichang <tzhouam@connect.ust.hk> * [Feature][TTS] Add Speech API streaming metrics (vllm-project#6853) Signed-off-by: XIN GAO <1037396230@qq.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Bugfix][Model] Fix FLUX.2 Klein multi-image edit metadata (vllm-project#7430) Signed-off-by: QI JIA <qi.jia@shengshu.ai> Co-authored-by: QI JIA <qi.jia@shengshu.ai> Co-authored-by: Cursor <cursoragent@cursor.com> * [BugFix] Fix leftovers of the legacy OpenAI realtime API event names (vllm-project#7426) Signed-off-by: Nick Cao <ncao@redhat.com> Co-authored-by: Codex <noreply@openai.com> * [Model] Add Tencent AuK speech generation and editing (encoder + diffusion pipeline) (vllm-project#7385) Signed-off-by: Yueqian Lin <linyueqian@outlook.com> Co-authored-by: Sy03 <1370724210@qq.com> * [XPU][Docker] Align XPU image and CI with vLLM v0.29.0 (vllm-project#7441) Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com> * [Bugfix] Add explicit error when using CFGP with distilled Cosmos3 models (vllm-project#7427) Signed-off-by: Maciej Bala <mbala@nvidia.com> * [Perf][Diffusion] Run MammothModa2 DiT attention through the shared attention layer (vllm-project#7094) Signed-off-by: MrlixiangWE <mrdanaer@gmail.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Bugfix] Give model CLI flags typed owners in the Omni config (vllm-project#7390) Signed-off-by: Guangjian <hiro20833@gmail.com> Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com> * [Bugfix] Require a model for `vllm serve --omni` (fixes vllm-project#4158) (vllm-project#4167) Signed-off-by: abinggo <107740309+abinggo@users.noreply.github.com> * [Bugfix] Send a downstream terminal chunk when a parked stage ends (vllm-project#6889) Signed-off-by: psv666 <2693925048@qq.com> * [NPU] upgrade to v0.29.0 (vllm-project#7433) Signed-off-by: Weiming Liao <liaowm5@gmail.com> * [Bugfix][Model][Lance] Support decoded video frames in video editing (vllm-project#5128) Signed-off-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com> Co-authored-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com> * [Refactor][Diffusion] Remove model-specific names from LoRA and ModelOpt loader defaults (vllm-project#5907) Signed-off-by: Alicia <115451386+congw729@users.noreply.github.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * Optimize CosyVoice3 Stage1 flow batching (vllm-project#4876) Signed-off-by: gerayking <399geray@gmail.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [3/N] Encode streamed video on the worker with bounded batching (vllm-project#7018) Signed-off-by: specture724 <specture724@gmail.com> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> * [Kernel][Boogu-Image] Fuse Q/K RMSNorm + interleaved RoPE via fused_qk_norm_rope (vllm-project#6982) Signed-off-by: Qihan Kang <rollykanggg@gmail.com> * [Bugfix][Frontend] Honor output_compression on the image generations route (vllm-project#7447) Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com> Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com> Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com> --------- Signed-off-by: Asthenia <asthenia0412@gmail.com> Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com> Signed-off-by: andyluo7 <andy.luo@amd.com> Signed-off-by: liangmengh <liangmengh@nvidia.com> Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> Signed-off-by: KrystalRay <keeleiray@gmail.com> Signed-off-by: Tianyao Wu <rayroy31@gmail.com> Signed-off-by: specture724 <specture724@gmail.com> Signed-off-by: hsliuustc0106 <liuhongsheng4@huawei.com> Signed-off-by: chaosansui <zzc15560846421@163.com> Signed-off-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com> Signed-off-by: eval-dev <0xe5bca0@gmail.com> Signed-off-by: eval <74645252+eval-dev@users.noreply.github.com> Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com> Signed-off-by: NancyFyong <NancyFyong@users.noreply.github.com> Signed-off-by: zjli2013 <leezhengjiang@126.com> Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu> Signed-off-by: Rakesh Kariya <rakesh.kariya@somaiya.edu> Signed-off-by: ZenAlexa <zimingwang945@gmail.com> Signed-off-by: BANANASJIM <bananasjim1@gmail.com> Signed-off-by: Allen Wu <allenwu2795@gmail.com> Signed-off-by: xutianle <xutianle@fudan.edu.cn> Signed-off-by: wangyu <410167048@qq.com> Signed-off-by: natureofnature <wzliu@connect.hku.hk> Signed-off-by: mglyn <1203789601@qq.com> Signed-off-by: Sparks-M <41097544+Sparks-M@users.noreply.github.com> Signed-off-by: chi030303 <106855944+chi030303@users.noreply.github.com> Signed-off-by: Bo Li <22713281+bobboli@users.noreply.github.com> Signed-off-by: NumberWan <wantszkin2003@gmail.com> Signed-off-by: zouyizhou <zouyizhou@huawei.com> Signed-off-by: ZhengWG <zwg0606@gmail.com> Signed-off-by: kunkunblueberry <1833921874@qq.com> Signed-off-by: Shaun Walsh <shaunwalsh24@gmail.com> Signed-off-by: mershi <mershi@tencent.com> Signed-off-by: hyw <yuweih205@gmail.com> Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com> Signed-off-by: Sy03 <1370724210@qq.com> Signed-off-by: Joshna Medisetty <joshna.medisetty@intel.com> Signed-off-by: Joshna-Medisetty <joshna.medisetty@intel.com> Signed-off-by: Guangjian <hiro20833@gmail.com> Signed-off-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com> Signed-off-by: Gao Han <hgaoaf@connect.ust.hk> Signed-off-by: Weiming Liao <liaowm5@gmail.com> Signed-off-by: tzhouam <tzhouam@connect.ust.hk> Signed-off-by: Zhou Taichang <tzhouam@connect.ust.hk> Signed-off-by: herotai214 <herotai214@gmail.com> Signed-off-by: LHXuuu <xulianhao.xlh@antgroup.com> Signed-off-by: Samit <285365963@qq.com> Signed-off-by: Rahul Steiger <rsteiger@aws-cmh-slurm-1-vscode-04.cm.cluster> Signed-off-by: Wojciech Kutak <wkutak@nvidia.com> Signed-off-by: tly <2200895168@qq.com> Signed-off-by: psv666 <2693925048@qq.com> Signed-off-by: Maciej Bala <mbala@nvidia.com> Signed-off-by: MaciejBalaNV <mbala@nvidia.com> Signed-off-by: summer <128961079+zhang-keliang@users.noreply.github.com> Signed-off-by: Yueqian Lin <linyueqian@outlook.com> Signed-off-by: Nick Cao <ncao@redhat.com> Signed-off-by: david6666666 <530634352@qq.com> Signed-off-by: yancaocn <yancaochn@163.com> Signed-off-by: samithuang <285365963@qq.com> Signed-off-by: suyanli220 <suyanli220@gmail.com> Signed-off-by: suyan.li <suyan.li@bytedance.com> Signed-off-by: wtz2333 <2955110911@qq.com> Signed-off-by: XIN GAO <1037396230@qq.com> Signed-off-by: QI JIA <qi.jia@shengshu.ai> Signed-off-by: MrlixiangWE <mrdanaer@gmail.com> Signed-off-by: abinggo <107740309+abinggo@users.noreply.github.com> Signed-off-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com> Signed-off-by: Alicia <115451386+congw729@users.noreply.github.com> Signed-off-by: gerayking <399geray@gmail.com> Signed-off-by: Qihan Kang <rollykanggg@gmail.com> Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com> Signed-off-by: José Carlos <jose@valendra.tech> Co-authored-by: Yancy <138764723+Asthenia0412@users.noreply.github.com> Co-authored-by: Asthenia <asthenia0412@gmail.com> Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com> Co-authored-by: andyluo7 <43718156+andyluo7@users.noreply.github.com> Co-authored-by: liangmenghuang <liangmengh@nvidia.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com> Co-authored-by: Lei Ke <1141466880@qq.com> Co-authored-by: KrystalRay <keeleiray@gmail.com> Co-authored-by: Tianyao Wu <54675599+twu3202@users.noreply.github.com> Co-authored-by: Anjie Hou <149605198+specture724@users.noreply.github.com> Co-authored-by: Zhichao Zhang <60429419+smartDream-chao@users.noreply.github.com> Co-authored-by: eval <74645252+eval-dev@users.noreply.github.com> Co-authored-by: boatman <1930807094@qq.com> Co-authored-by: NancyFyong <88076188+NancyFyong@users.noreply.github.com> Co-authored-by: NancyFyong <NancyFyong@users.noreply.github.com> Co-authored-by: zhengjia <ZJLi2013@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Rakesh Kariya <83279947+rk9595@users.noreply.github.com> Co-authored-by: Ziming Wang <125807850+ZenAlexa@users.noreply.github.com> Co-authored-by: Jim Ban <77719403+BANANASJIM@users.noreply.github.com> Co-authored-by: Allen Wu <85376543+EchoHayate@users.noreply.github.com> Co-authored-by: TRAE CLI <traecli@bytedance.com> Co-authored-by: xutianle <24210290017@m.fudan.edu.cn> Co-authored-by: wangyu <53896905+yenuo26@users.noreply.github.com> Co-authored-by: NATURE <wzliu@connect.hku.hk> Co-authored-by: Mu GuanLin <1203789601@qq.com> Co-authored-by: Sparks <41097544+Sparks-M@users.noreply.github.com> Co-authored-by: chi030303 <106855944+chi030303@users.noreply.github.com> Co-authored-by: Bo Li <22713281+bobboli@users.noreply.github.com> Co-authored-by: NumberWan <wantszkin2003@gmail.com> Co-authored-by: zyz111222 <zouyizhou@huawei.com> Co-authored-by: Zheng Wengang <zwg0606@gmail.com> Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com> Co-authored-by: kunkun <72174834+kunkunblueberry@users.noreply.github.com> Co-authored-by: Shaun Walsh <153730091+Shaun-Walsh@users.noreply.github.com> Co-authored-by: Nick Cao <ncao@redhat.com> Co-authored-by: shiyichuan <93317314+CarrotSwordsman@users.noreply.github.com> Co-authored-by: mershi <mershi@tencent.com> Co-authored-by: hyw <109567717+yuweih205@users.noreply.github.com> Co-authored-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com> Co-authored-by: Sy03 <1370724210@qq.com> Co-authored-by: Joshna-Medisetty <joshna.medisetty@intel.com> Co-authored-by: Guangjian Dong <163994576+Hiro208@users.noreply.github.com> Co-authored-by: hsliu_ustc <hsliu_ustc@noreply.gitcode.com> Co-authored-by: Gao Han <hgaoaf@connect.ust.hk> Co-authored-by: Weiming Liao <liaowm5@gmail.com> Co-authored-by: Zhou Taichang <tzhouam@connect.ust.hk> Co-authored-by: herotai214 <68222888+herotai214@users.noreply.github.com> Co-authored-by: LHXuuu <xulianhao.xlh@antgroup.com> Co-authored-by: Samit <285365963@qq.com> Co-authored-by: wkutak <wkutak@nvidia.com> Co-authored-by: Rahul Steiger <rsteiger@nvidia.com> Co-authored-by: tlysanhuo <166924864+tlysanhuo@users.noreply.github.com> Co-authored-by: psv666 <150513104+psv666@users.noreply.github.com> Co-authored-by: MaciejBalaNV <mbala@nvidia.com> Co-authored-by: summer <128961079+zhang-keliang@users.noreply.github.com> Co-authored-by: Yueqian Lin <70319226+linyueqian@users.noreply.github.com> Co-authored-by: WeiQing Chen <40507679+david6666666@users.noreply.github.com> Co-authored-by: Yan Cao <31481315+yancaocn@users.noreply.github.com> Co-authored-by: yancaocn <yancaochn@163.com> Co-authored-by: SuyanLi <126558907+suyanli220@users.noreply.github.com> Co-authored-by: suyan.li <suyan.li@bytedance.com> Co-authored-by: wtz2333 <2955110911@qq.com> Co-authored-by: GXIN <37653830+gxxx-hum@users.noreply.github.com> Co-authored-by: Qi Jia <kuafou@gmail.com> Co-authored-by: QI JIA <qi.jia@shengshu.ai> Co-authored-by: Codex <noreply@openai.com> Co-authored-by: DanaerLee <mrdanaer@gmail.com> Co-authored-by: longguo <107740309+abinggo@users.noreply.github.com> Co-authored-by: junpengw67-max <junpengw67@gmail.com> Co-authored-by: 吴俊鹏 <248679769+junpengw67-max@users.noreply.github.com> Co-authored-by: Alicia <115451386+congw729@users.noreply.github.com> Co-authored-by: geray <48796550+gerayking@users.noreply.github.com> Co-authored-by: KANG Qihan <3149604185@qq.com>
…llm-project#7046) Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu> Signed-off-by: Rakesh Kariya <rakesh.kariya@somaiya.edu> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Summary
/v1/audio/speech/streamsynthesizes one request perinput.done. Docs still described per-sentence splitting from [Feature][TTS] Streaming Text Input for Qwen3-TTS via WebSocket #1230.split_granularity(nonedefault /sentence/clause) so STT/LLM clients can start audio beforeinput.done, without changing the default path..!?, CJK, Indic danda (।॥), and Arabic marks.seedfrom WebSocketsession.config(HTTP speech already had it).This is a new opt-in feature plus a docs/API alignment fix, not a crash bug. No dedicated issue.
AI assistance: implementation drafted with Cursor; human submitter is rk9595. Please review splitter edge cases (decimals, Latin abbreviations).
Test plan
pytest tests/entrypoints/openai_api/test_speech_text_splitter.py tests/entrypoints/openai_api/test_serving_speech_stream.py -m "core_model and cpu"streaming_speech_client.py --split-granularity sentencewith Indic text such asनमस्ते। कैसे हो?none) still one request per flush