Repository navigation
[Perf][CosyVoice3] Bound the flow left context with a sliding window - #6588
BruceLoveDecimal wants to merge 7 commits into
Conversation
|
Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits. |
|
This PR appears to belong to: docs/design/module/model_integration.md. Module owners: @tzhouam @gcanlin @BruceLoveDecimal, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Omni ReviewBot: no human activity for 10 days@BruceLoveDecimal this pull request has had no human commit, comment or review since 2026-08-29. Per repository policy it may be closed if it stays inactive. To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline. |
5a6fe21 to
b500e54
Compare
b500e54 to
e5e94f9
Compare
13f6252 to
3787a7c
Compare
The streaming code2wav path resent the whole emitted-token prefix to the flow on every chunk, so per-chunk flow cost grew linearly with elapsed audio and the per-request cost was O(T^2); long streams also outgrew the TensorRT estimator's 3000-frame profile. vllm-project#7521 bounded the HiFT vocoder side of the same pipeline and left this side untouched. Resend at most ``codec_left_context_frames`` emitted tokens (25 by default, one second of audio; ``<= 0`` restores the unbounded behaviour) as left context, and shrink ``left_context_size`` to match so only the new frames are emitted. The deploy config already carried the key; it now takes effect. This is a quality trade-off rather than a free win: the DiT attends over the full sequence, so dropping earlier prefix changes the mel of the new chunk. WER/SIM were neutral in the PR's A/B; widen the window if seams are audible. Measured on an RTX PRO 6000 with the default TensorRT estimator against a main that already includes vllm-project#7521 (30 requests, three utterance lengths, reference audio served from a local file): at concurrency 1 the per-chunk latency no longer climbs from 142 ms to 1182 ms over a long utterance but stays near 230 ms, and RTF drops from 0.171 to 0.113; at concurrency 8 the median per-chunk latency drops from 2008 ms to 700 ms (p95 4025 ms to 963 ms), RTF from 1.25 to 0.38, and audio throughput rises from 6.0x to 21x realtime. Signed-off-by: liuqihao <liuqihao970610@gmail.com>
3787a7c to
a70ef32
Compare
…tion-indexed buffer Upstream CosyVoice seeds one noise buffer (``torch.randn([1, 80, 50 * 300])`` under seed 0) and slices it by mel position, so the flow starts from the same noise at a given position on every call. The port replaced that with a fresh ``torch.randn`` per call, which made every regenerated left context differ from the frames it was emitted from, and made the same seed produce different audio across requests. Restore the buffer (built with a local generator so the global RNG is not touched, and kept out of the checkpoint) and index it by absolute position: the prompt always maps to the start of the buffer, and the frames after it map to ``prompt_len + offset`` where the offset is the absolute mel index of the first resent token. The streaming code2wav tracks how many tokens each request has emitted, so a bounded left context reuses exactly the noise its frames were first generated with, per row when flow batching is on. The offline path keeps offset 0, which is upstream's behaviour bit for bit. Signed-off-by: liuqihao <liuqihao970610@gmail.com>
With a fixed seed a CosyVoice3 request is byte-reproducible on its own, but under concurrency the same seed yields different audio: the talker's logits depend on which other requests share its batch, and a one-ulp difference is enough to flip a sampled token. Audio durations of the same text spread from 37 s to 65 s at concurrency 8 while being identical at concurrency 1, which places the variance in the talker, not in code2wav. vLLM's batch-invariant mode is the fix, but the talker could not run in it: ``llm_decoder`` has ``speech_token_size + 200 = 6761`` outputs, cuBLASLt rejects that odd leading dimension for fp16/bf16 GEMMs, and batch-invariant mode forbids the cuBLAS fallback. Pad the head to a multiple of 8 and slice the logits back, but only in that mode: a different N can change cuBLAS's algorithm choice, and the default path stays bit-identical (same c=1 audio before and after). The deploy config documents the opt-in as a commented per-stage ``env`` on the talker stage. Measured on an RTX PRO 6000 (30 requests, three utterance lengths, seed 42): byte-identity violations at concurrency 8 go from 27/30 to 0/30 and the audio is identical across concurrency 1 and 8. The cost is real, which is why it stays opt-in: about 2x talker latency at concurrency 1 (RTF 0.114 -> 0.225) and about 35% less throughput at concurrency 8 (21.0x -> 13.7x realtime). Signed-off-by: liuqihao <liuqihao970610@gmail.com>
a70ef32 to
c54a72c
Compare
Omni ReviewBot triage noteAutomated triage of commit
These are automated triage suggestions only — the final decision belongs to the maintainers. |
Omni ReviewBot routing recordAssigned Strict on zcode (GLM-5.3-Flash) under experiment |
|
@BruceLoveDecimal Hi, is this PR still active and waiting for review? Please have a look at the latest main, resolve the conflict, and fix the DCO CI. |
Thanks for keeping tracking of this PR. I'll rebase my PR and then re-run the benchmark since the main branch had new updates about cozyvoice3. After finishing it, i'll ping you and please help review it. Thanks! |
Omni ReviewBot attempt recordReview attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: 00f174b4-0f70-45c6-9115-1d04e3ada2df) — check |
Omni ReviewBot attempt recordReview attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: 50508f51-d8ac-45e1-b64f-14f1dbc2c3fc) — check |
Omni ReviewBot attempt recordReview attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: 7e189a7e-281e-4c23-b520-ad08ea0c39d4) — check |
vllm-omni-review-bot
left a comment
There was a problem hiding this comment.
Omni ReviewBot review
0 actionable finding(s).
CI at
74004ce108d8(2026-10-10T03:00:04.529796+00:00): required check(s) blocking:buildkite/vllm-omni(missing).
Note: The assigned review arm
strict/zcode/GLM-5.3-Flashcould not complete this review, so it was produced by the fallback armdirect/cursor/auto. It is excluded from the routing experiment.
Full review analysis
PR description
CosyVoice3 streaming code2wav no longer feeds the flow the entire emitted token prefix on every chunk. The async chunk processor resends at most codec_left_context_frames tokens (default 25; a nonpositive value keeps the old full prefix) and records that shorter left-context length on the payload. The flow draws its initial noise from one CPU seed-0 buffer, indexed by absolute mel position and by a per-request flow_emitted_tokens counter, so a slid window reuses the noise those frames were first generated with; the offline path still uses offset 0. Separately, when VLLM_BATCH_INVARIANT is set at weight load, the talker llm_decoder matmul is copied out to a multiple of 8 and the extra logits are sliced off before the existing EOS merge. The default head path is unchanged.
Change flow
flowchart LR
hidden["[EXISTING] Talker hidden states"]:::existing
head["[NEW] Optional 8-aligned<br/>llm_decoder logits"]:::new
window["[CHANGED] Chunk processor<br/>bounded left context"]:::changed
noise["[NEW] Seed-0 position-indexed<br/>flow noise"]:::new
code2wav["[CHANGED] Code2wav stream state<br/>flow_emitted_tokens"]:::changed
audio["[EXISTING] PCM chunk"]:::existing
hidden --> head --> window --> code2wav --> audio
noise --> code2wav
classDef existing fill:#e5e7eb,stroke:#6b7280,color:#111827
classDef changed fill:#fef3c7,stroke:#d97706,color:#451a03,stroke-width:2px
classDef new fill:#dcfce7,stroke:#16a34a,color:#052e16,stroke-width:2px
classDef removed fill:#fee2e2,stroke:#dc2626,color:#450a0a,stroke-width:2px
No actionable findings.
🤖 This review was generated by InferMatrix Copilot, an open-source repo-maintenance agent for PR review, CI debugging and issue triage. Try it on your own repo, and ⭐ star it if it helped!
Purpose
Part of #6870 C6 task
CosyVoice3's streaming code2wav resends the whole emitted-token prefix to the flow on every chunk. #7521 bounded the HiFT side, but the flow side still grows. That causes three problems:
torch.randnon every call, while upstream uses one fixed noise buffer. So the regenerated left context no longer matches the frames already emitted, and the same seed gives different audio even at c=1.VLLM_BATCH_INVARIANTwould fix that, but cuBLASLt rejectsllm_decoder's 6761-wide output for fp16/bf16, and that mode forbids the cuBLAS fallback.Modifications
Sliding-window flow left context (
stage_input_processors/cosyvoice3.py)codec_left_context_framesemitted tokens as left context (default 25, i.e. 1 s) and shrinkleft_context_sizeto match.<= 0restores the unbounded behaviour.Fixed, position-indexed initial noise (
code2wav_core/cfm.py,cosyvoice3_code2wav.py)torch.randn([1, 80, 50 * 300])buffer under seed 0. It uses a private CPU generator and is not a checkpoint key (persistent=False).p + noise_offset.flow_emitted_tokensper request (per row when flow batching is on), so a windowed left context reuses the exact noise its frames were first generated with.Batch-invariant talker head (
cosyvoice3.py,deploy/cosyvoice3.yaml)VLLM_BATCH_INVARIANTis set: padllm_decoderto a multiple of 8 withpad_vocab_size, then slice the logits back. The default path is unchanged and bit-identical.envon the talker stage. It costs ~2x talker latency at c=1 and ~35% throughput at c=8.Test Plan
Base:
main@ 1b87115. Each arm runs once onmainand once on this PR.Deploy config: default
vllm_omni/deploy/cosyvoice3.yamlmax_num_seqs 8on both stagescodec_chunk_frames 15,codec_left_context_frames 25COSYVOICE3_BATCH_FLOWoff,VLLM_BATCH_INVARIANTunsetHW: 1x H800 PCIe 80 GB, both stages on the same GPU · vLLM 0.29.0 · TensorRT 11.3
Workload
zhaochenyang20/seed-tts-eval, Englishmeta.lst(1,088 rows).ref_audiodata URL +ref_text, requestseed=42./v1/audio/speech.SIM uses the official seed-tts-eval
cal_simprotocol: UniSpeech WavLM-large speaker verification withwavlm_large_finetune.pth, prompt WAV vs synthesized WAV.Metrics
Caveat for c=16/32: these exceed
max_num_seqs 8and trip a pre-existingmainissue. For those cells, both arms carry the same one-line fix,MetaStruct.resumable: bool | None = None(see Test Result).Test Result
Unit: 120 passed (CosyVoice3 suite incl. the new fixed-noise and talker-head tests), 0 failed.
H800: Seed-TTS EN with per-request reference voices
mainvs this PR on a realistic streaming workload:zhaochenyang20/seed-tts-evalEnglish rows, each request cloning its own reference voice (inlineref_audio+ref_text), requestseed=42, default deploy config (TensorRT estimator,max_num_seqs 8). Driven byvllm bench serve --omni --backend openai-audio-speech --dataset-name seed-tts --seed 42 --num-warmups 2, one server per arm. Seed-TTS utterances are short (mean ~5 s of audio), so this measures the everyday regime rather than the long-utterance O(T²) case.HW: 1x H800 PCIe 80 GB, both stages on the same GPU · vLLM 0.29.0 · TensorRT 11.3 · base:
main@ 1b87115Full English set (1,088 rows), this PR
torch 2.13.0+cu129 on driver 570 with the CUDA 13.4 forward-compat runtime. Each server gets a discarded 100-request c=8 warmup before the cells. All 8 cells: 1088/1088 completed, no engine errors.
Accuracy: WER / SIM (H800, Seed-TTS EN full set, 1,088 rows, c=4)
Same H800 setup as the accuracy row in the Test Plan. Protocol:
--seed-tts-wer-eval).cal_simprotocol, i.e. UniSpeech WavLM-large speaker verification withwavlm_large_finetune.pth, prompt WAV vs synthesized WAV.H800: long-form Seed-TTS EN (100 rows, ~30 s each)
The full-set run above uses short utterances (mean ~5 s), so the window rarely engages. This run uses long-form rows on the same H800 box, same torch/driver,
main@ 1b87115 vs this PR @ 3787a7c.Workload
zhaochenyang20/seed-tts-evalEN, in the samemeta.lstformat.ref_audio+ref_text.mainfinishes every request and the two arms stay comparable. This run does not exercise the >60 s crash.seed=42.All 6 cells: 100/100 completed, no engine errors, no TensorRT profile errors.
Perf
main's c=1 cell was run twice (the first run had no WER dependency installed). E2E mean was 5462 ms and 5441 ms.Accuracy: WER / SIM
Same protocol as above: Whisper-large-v3 with seed-tts-eval normalization, which transcribes audio > 30 s in sequential long-form mode, and the official
cal_simWavLM-large SV for SIM.long009(Whisper long-form repetition loop on this PR)