Repository navigation
[Bugfix] Segment long IndexTTS-2.5 speech requests - #7567
voidreaming wants to merge 1 commit into
Conversation
Signed-off-by: voidreaming <2707786374wsj@gmail.com>
|
This PR appears to belong to: docs/design/module/profiling.md, docs/design/module/model_integration.md, docs/design/module/entrypoints.md. Module owners: @gcanlin @tzhouam @linyueqian Routing: @gcanlin via module of the changed files, semantic router; @tzhouam via module of the changed files, semantic router; @linyueqian via module of the changed files, CODEOWNERS @voidreaming, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Omni ReviewBot triage noteAutomated triage of commit
These are automated triage suggestions only — the final decision belongs to the maintainers. |
linyueqian
left a comment
There was a problem hiding this comment.
Reviewed at d3366e12 (static read against main; nothing executed). The design is the right one for a bug fix: normalize once, split under the 120-token budget with the paired pronunciation markers kept atomic, run the segments sequentially with per-segment clone()d sampling params (both IndexTTS-2.5 stages use vLLM SamplingParams, so the clone contract holds), abort the in-flight segment on any failure before aclose(), and hand serving_speech.py one aggregate OmniRequestOutput whose _multimodal_output is a 1-D waveform plus sr and whose metrics carry the summed num_tokens_out, which is exactly what the non-streaming encoder and the usage counter consume. Single-segment IndexTTS-2.5 requests and every other adapter keep the ordinary generate() call because additional_prompts defaults to empty.
One correctness item and one test gap are inline: the splitter can emit a whitespace-only segment at the budget edge, which then fails prompt preparation and turns a valid long request into a 400. Process item: the branch conflicts with current main in serving_speech.py, where #7544 landed the streaming word-timestamp changes; when you rebase, keep main's OutputPolicy handling and _speech_output_policies, the streaming cumulative_audio plumbing and the RequestOutputKind.CUMULATIVE coercion for streaming timestamps, and insert the additional_prompts branch around the generate() call without dropping those. I will re-review on the rebased head.
| def append_unit(unit: str) -> None: | ||
| nonlocal current | ||
| if current and not fits(current + unit): | ||
| segments.append(current) |
There was a problem hiding this comment.
[important] append_unit only checks truthiness, so a whitespace-only unit becomes its own segment. The punctuation split keeps the delimiter on the left part, so text ending in punctuation plus a space or newline ("甲。\n", or "甲. " when normalization is off and trailing spaces survive) yields ["甲。", "\n"]; when the filled segment cannot take that trailing unit it is committed alone, and _prepare_segments then hits prepare_indextts25_text's empty-text check and the whole request fails with a 400 before any audio is generated. Skip units whose strip() is empty when starting a segment, and drop trailing whitespace that cannot attach to current instead of committing it.
|
|
||
| @pytest.mark.asyncio | ||
| @pytest.mark.parametrize("failure", ["exception", "empty", "length", "unfinished", "nan", "rate"]) | ||
| async def test_segment_failure_does_not_return_partial_audio(failure): |
There was a problem hiding this comment.
[suggestion] The failure test always fails segment 0, and the sample-rate test fails the final segment without asserting the abort, so the path this module is really for (a middle segment failing after earlier ones succeeded, with abort(segment_id) called on the failing one and the error surfaced without partial audio) is not pinned. One parametrized case that fails segment 1 of 3 and asserts the abort call and the raised error would cover it.
Omni ReviewBot: no human activity for 7 days@voidreaming this pull request has had no human commit, comment or review since 2026-09-17. Please confirm the current plan and next step. The author or a maintainer decides whether to change the PR state. To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline. |
Omni ReviewBot: no human activity for 14 days@voidreaming this pull request has had no human commit, comment or review since 2026-09-17. Please consider marking this PR as draft until work can resume. The author or a maintainer decides whether to change the PR state. To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline. |
|
@voidreaming this PR is labeled Could you please take a look and push an update to get CI green? Once the checks pass we can proceed with review/merge. Thanks! |
Omni ReviewBot routing recordAssigned Strict on cursor (cursor-grok-4.6-high) under experiment |
Purpose
Addresses #7378: long Chinese input can produce incorrect speech with
IndexTTS-2.5 through
/v1/audio/speech.Normalize text once and split it under a 120-token budget while preserving
pronunciation annotations. Generate segments sequentially and concatenate
their waveforms with 200 ms pauses into one audio response.
Multi-segment requests support non-streaming audio without word timestamps.
Generation limits apply per segment; failed or incomplete segments fail
the request.
Test Plan
vLLM Version: 0.29.0
vLLM-Omni Commit:
d3366e12, based on5761ad55.audio assembly, usage accounting, and failure/cancellation handling.
with IndexTTS-2.5.
Test Result
with correct speech. Short-text synthesis also passed.
AI assistance: Codex assisted with implementation and validation.
The contributor reviewed the changes and listened to the generated audio.
Reproduction commands
With the native IndexTTS-2.5 model bundle and optional dependencies installed,
start the server from the repository root:
CUDA_VISIBLE_DEVICES=3 vllm serve "$MODEL" \ --omni --trust-remote-code --host 127.0.0.1 --port 8093 \ --served-model-name IndexTTS-2.5 \ --deploy-config vllm_omni/deploy/indextts2_5.yamlSave the reporter's original Shenzhen passage as
input.txt, and obtain thesame reference recording:
curl --fail --location \ https://github.com/user-attachments/files/32059486/ref_audio.wav \ --output ref_audio.wav python examples/online_serving/text_to_speech/indextts2/speech_client.py \ --api-base http://127.0.0.1:8093 --model-version 2.5 --model IndexTTS-2.5 \ --lang zh --text "$(cat input.txt)" --ref-audio ref_audio.wav --output full.wav