Skip to content

[Perf][CosyVoice3] Bound the flow left context with a sliding window - #6588

Open
BruceLoveDecimal wants to merge 7 commits into
vllm-project:mainfrom
BruceLoveDecimal:feat/cozyvoice3_cudagraph
Open

BruceLoveDecimal wants to merge 7 commits into
vllm-project:mainfrom
BruceLoveDecimal:feat/cozyvoice3_cudagraph

Conversation

@BruceLoveDecimal

@BruceLoveDecimal BruceLoveDecimal commented Aug 24, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Part of #6870 C6 task
CosyVoice3's streaming code2wav resends the whole emitted-token prefix to the flow on every chunk. #7521 bounded the HiFT side, but the flow side still grows. That causes three problems:

  • Cost and crashes: per-chunk flow cost grows with elapsed audio, so each request costs O(T²). Once a stream passes ~60 s it outgrows the TensorRT estimator's 3000-frame profile, and the stage-1 engine core dies with every in-flight request.
  • Non-deterministic flow: the port draws fresh torch.randn on every call, while upstream uses one fixed noise buffer. So the regenerated left context no longer matches the frames already emitted, and the same seed gives different audio even at c=1.
  • Talker can't run batch-invariant: under concurrency, the sampled tokens depend on which other requests share the batch. VLLM_BATCH_INVARIANT would fix that, but cuBLASLt rejects llm_decoder's 6761-wide output for fp16/bf16, and that mode forbids the cuBLAS fallback.

Modifications

  1. Sliding-window flow left context (stage_input_processors/cosyvoice3.py)

    • Resend at most codec_left_context_frames emitted tokens as left context (default 25, i.e. 1 s) and shrink left_context_size to match. <= 0 restores the unbounded behaviour.
    • The deploy config already had this key; it now takes effect.
    • Trade-off: the DiT is non-causal, so a shorter window also changes the mel of the new frames. WER/SIM stay neutral in the tests; widen the window if seams are audible.
  2. Fixed, position-indexed initial noise (code2wav_core/cfm.py, cosyvoice3_code2wav.py)

    • Restore upstream's torch.randn([1, 80, 50 * 300]) buffer under seed 0. It uses a private CPU generator and is not a checkpoint key (persistent=False).
    • Index the buffer by absolute mel position: prompt frames map to the start of the buffer, later frames to p + noise_offset.
    • Streaming state records flow_emitted_tokens per request (per row when flow batching is on), so a windowed left context reuses the exact noise its frames were first generated with.
    • The offline path keeps offset 0 and matches upstream bit for bit.
  3. Batch-invariant talker head (cosyvoice3.py, deploy/cosyvoice3.yaml)

    • Only when VLLM_BATCH_INVARIANT is set: pad llm_decoder to a multiple of 8 with pad_vocab_size, then slice the logits back. The default path is unchanged and bit-identical.
    • The deploy config adds a commented opt-in env on the talker stage. It costs ~2x talker latency at c=1 and ~35% throughput at c=8.

Test Plan

Base: main @ 1b87115. Each arm runs once on main and once on this PR.

Deploy config: default vllm_omni/deploy/cosyvoice3.yaml

  • max_num_seqs 8 on both stages
  • codec_chunk_frames 15, codec_left_context_frames 25
  • TensorRT fp16 estimator
  • COSYVOICE3_BATCH_FLOW off, VLLM_BATCH_INVARIANT unset

HW: 1x H800 PCIe 80 GB, both stages on the same GPU · vLLM 0.29.0 · TensorRT 11.3

run torch / driver this PR
Perf, full set, c=1/8/16/32 torch 2.13.0+cu129 on driver 570 with the CUDA 13.4 forward-compat runtime 3787a7c
Accuracy (WER/SIM) torch 2.13.0+cu130 on driver 595 (CUDA 13.2) 3787a7c

Workload

  • zhaochenyang20/seed-tts-eval, English meta.lst (1,088 rows).
  • Each request clones its own reference voice: inline ref_audio data URL + ref_text, request seed=42.
  • Audio streams as PCM over /v1/audio/speech.
  • One server per arm. For the full-set run, each server gets a discarded 100-request c=8 warmup first.
vllm serve $MODEL --omni --deploy-config vllm_omni/deploy/cosyvoice3.yaml --port 8000

# Perf: full set at c=1/8/16/32
vllm bench serve --omni --port 8000 --model $MODEL --tokenizer $MODEL/CosyVoice-BlankEN \
  --backend openai-audio-speech --endpoint /v1/audio/speech \
  --dataset-name seed-tts --dataset-path /path/to/seed-tts-eval --seed-tts-locale en \
  --num-prompts 1088 --num-warmups 2 --max-concurrency $C --request-rate inf --seed 42 \
  --extra-body '{"seed": 42}' \
  --percentile-metrics ttft,e2el,audio_rtf,audio_ttfp,audio_duration,audio_underrun --metric-percentiles 50,95,99

# Accuracy: full set at c=4
# WER uses Whisper-large-v3 with seed-tts-eval normalization; synthesized WAVs are saved for SIM
SEED_TTS_WER_SAVE_AUDIO_DIR=/path/to/wavs vllm bench serve ... \
  --num-prompts 1088 --max-concurrency 4 --seed-tts-wer-eval --seed-tts-wer-save-items

SIM uses the official seed-tts-eval cal_sim protocol: UniSpeech WavLM-large speaker verification with wavlm_large_finetune.pth, prompt WAV vs synthesized WAV.

Metrics

  • TTFA: time to first PCM byte.
  • E2E: request start to stream close.
  • RTF: E2E / audio duration.
  • req/s: completed requests / wall time.
  • Throughput: total audio seconds / wall time, including queueing.

Caveat for c=16/32: these exceed max_num_seqs 8 and trip a pre-existing main issue. For those cells, both arms carry the same one-line fix, MetaStruct.resumable: bool | None = None (see Test Result).

Test Result

Unit: 120 passed (CosyVoice3 suite incl. the new fixed-noise and talker-head tests), 0 failed.

H800: Seed-TTS EN with per-request reference voices

main vs this PR on a realistic streaming workload: zhaochenyang20/seed-tts-eval English rows, each request cloning its own reference voice (inline ref_audio + ref_text), request seed=42, default deploy config (TensorRT estimator, max_num_seqs 8). Driven by vllm bench serve --omni --backend openai-audio-speech --dataset-name seed-tts --seed 42 --num-warmups 2, one server per arm. Seed-TTS utterances are short (mean ~5 s of audio), so this measures the everyday regime rather than the long-utterance O(T²) case.

HW: 1x H800 PCIe 80 GB, both stages on the same GPU · vLLM 0.29.0 · TensorRT 11.3 · base: main @ 1b87115

Full English set (1,088 rows), this PR

torch 2.13.0+cu129 on driver 570 with the CUDA 13.4 forward-compat runtime. Each server gets a discarded 100-request c=8 warmup before the cells. All 8 cells: 1088/1088 completed, no engine errors.

image
conc metric main this PR delta
1 req/s 1.255 1.316 +4.8%
1 throughput (audio s/s) 6.22x 6.52x +4.8%
1 TTFA mean / p50 / p95 (ms) 254 / 254 / 290 248 / 249 / 284 −2%
1 E2E mean / p50 / p95 (ms) 797 / 794 / 1174 760 / 759 / 1096 −4.6% / −4.4% / −6.7%
1 RTF mean / p50 / p95 0.163 / 0.160 / 0.189 0.157 / 0.154 / 0.183 −4%
8 req/s 2.300 2.493 +8.4%
8 throughput (audio s/s) 11.52x 12.35x +7.2%
8 TTFA mean / p50 / p95 (ms) 1241 / 1255 / 1624 1105 / 1120 / 1461 −11% / −11% / −10%
8 E2E mean / p50 / p95 (ms) 3457 / 3436 / 4915 3205 / 3184 / 4563 −7.3%
8 RTF mean / p50 / p95 0.714 / 0.694 / 0.934 0.665 / 0.651 / 0.840 −7% / −6% / −10%
16 req/s 2.050 2.143 +4.5%
16 TTFA mean / p50 / p95 (ms) 3365 / 3274 / 4648 3493 / 3370 / 4879 +4% / +3% / +5%
16 E2E mean / p50 / p95 (ms) 7779 / 7521 / 10986 7447 / 7323 / 10021 −4.3% / −2.6% / −8.8%
16 RTF mean / p50 / p95 1.415 / 1.357 / 2.020 1.232 / 1.158 / 1.767 −13% / −15% / −13%
32 req/s 1.915 2.070 +8.1%
32 TTFA mean / p50 / p95 (ms) 11353 / 11351 / 14462 10519 / 10523 / 12315 −7% / −7% / −15%
32 E2E mean / p50 / p95 (ms) 16617 / 16638 / 21238 15339 / 15378 / 18403 −7.7% / −7.6% / −13.3%
32 RTF mean / p50 / p95 2.892 / 2.725 / 4.478 2.466 / 2.311 / 3.756 −15% / −15% / −16%

Accuracy: WER / SIM (H800, Seed-TTS EN full set, 1,088 rows, c=4)

Same H800 setup as the accuracy row in the Test Plan. Protocol:

  • WER: Whisper-large-v3 with seed-tts-eval normalization (--seed-tts-wer-eval).
  • SIM: the official cal_sim protocol, i.e. UniSpeech WavLM-large speaker verification with wavlm_large_finetune.pth, prompt WAV vs synthesized WAV.
metric main this PR
completed 1088/1088 1088/1088
WER mean 0.0219 0.0217
WER median 0.0000 0.0000
rows with WER > 0 / > 0.1 191 / 70 181 / 70
SIM mean 0.7011 0.7036
SIM median 0.7082 0.7134
mean audio duration 4.92 s 4.94 s

H800: long-form Seed-TTS EN (100 rows, ~30 s each)

The full-set run above uses short utterances (mean ~5 s), so the window rarely engages. This run uses long-form rows on the same H800 box, same torch/driver, main @ 1b87115 vs this PR @ 3787a7c.

Workload

  • 100 rows built from zhaochenyang20/seed-tts-eval EN, in the same meta.lst format.
  • Each row uses a different reference voice: a distinct prompt WAV (≤ 5.5 s) and its transcript, sent inline as ref_audio + ref_text.
  • Target text: distinct seed-tts EN target sentences joined to 65–104 words, no sentence reused. Synthesized audio is 30.0 s mean, ~37 s p95.
  • Prompt + target stays under the TensorRT estimator's 3000-frame (60 s) profile, so main finishes every request and the two arms stay comparable. This run does not exercise the >60 s crash.
  • One server per arm with a discarded 8-request c=8 warmup, then c=1/4/8. Each cell runs WER and saves WAVs for SIM. Default deploy config, request seed=42.
SEED_TTS_WER_SAVE_ITEMS=1 SEED_TTS_WER_SAVE_AUDIO_DIR=/path/to/wavs_c$C \
vllm bench serve --omni --port 8000 --model $MODEL --tokenizer $MODEL/CosyVoice-BlankEN \
  --backend openai-audio-speech --endpoint /v1/audio/speech \
  --dataset-name seed-tts --dataset-path /path/to/seedtts_long --seed-tts-locale en \
  --num-prompts 100 --num-warmups 2 --max-concurrency $C --request-rate inf --seed 42 \
  --extra-body '{"seed": 42}' --seed-tts-wer-eval \
  --percentile-metrics ttft,e2el,audio_rtf,audio_ttfp,audio_duration,audio_underrun --metric-percentiles 50,95,99

All 6 cells: 100/100 completed, no engine errors, no TensorRT profile errors.

Perf

conc metric main this PR delta
1 req/s 0.184 0.270 +46.8%
1 throughput (audio s/s) 5.51x 8.10x +46.8%
1 TTFA mean / p50 / p95 (ms) 254 / 254 / 303 256 / 255 / 300 ±0
1 E2E mean / p50 / p95 (ms) 5441 / 5315 / 7502 3705 / 3693 / 4529 −31.9% / −30.5% / −39.6%
1 RTF mean / p50 / p95 0.180 / 0.178 / 0.217 0.124 / 0.123 / 0.129 −31.4% / −30.9% / −40.5%
1 streaming continuity OK 100% 100%
4 req/s 0.264 0.509 +92.8%
4 throughput (audio s/s) 7.82x 15.26x +95.0%
4 TTFA mean / p50 / p95 (ms) 1188 / 1000 / 2275 496 / 511 / 602 −58% / −49% / −74%
4 E2E mean / p50 / p95 (ms) 15075 / 14726 / 19510 7807 / 7734 / 9641 −48.2% / −47.5% / −50.6%
4 RTF mean / p50 / p95 0.508 / 0.514 / 0.568 0.261 / 0.262 / 0.271 −48.7% / −49.1% / −52.3%
4 underrun mean / p95 (s) 0.174 / 0.814 0.001 / 0.000
4 streaming continuity OK 67% 100%
8 req/s 0.271 0.588 +117.2%
8 throughput (audio s/s) 8.11x 17.79x +119.3%
8 TTFA mean / p50 / p95 (ms) 3358 / 3459 / 5235 869 / 904 / 1146 −74% / −74% / −78%
8 E2E mean / p50 / p95 (ms) 29136 / 28987 / 36944 13376 / 13540 / 16709 −54.1% / −53.3% / −54.8%
8 RTF mean / p50 / p95 0.973 / 0.973 / 1.100 0.442 / 0.446 / 0.461 −54.6% / −54.2% / −58.0%
8 underrun mean / p95 (s) 1.586 / 3.413 0.071 / 0.428
8 streaming continuity OK 6% 75%
  • The gain grows with concurrency and is far larger than on the short full set (+5–8%). That fits per-chunk flow cost no longer growing with elapsed audio.
  • At c=1 every row has the same audio duration on both arms (mean 30.02 s), so the talker emitted the same tokens and the difference is code2wav alone. At c=4/8 the talker is batch-variant, and mean durations differ by ~1% (29.62 vs 29.95 s at c=4, 29.97 vs 30.26 s at c=8).
  • main's c=1 cell was run twice (the first run had no WER dependency installed). E2E mean was 5462 ms and 5441 ms.

Accuracy: WER / SIM

Same protocol as above: Whisper-large-v3 with seed-tts-eval normalization, which transcribes audio > 30 s in sequential long-form mode, and the official cal_sim WavLM-large SV for SIM.

conc metric main this PR
1 WER mean 0.0627 0.0619
1 rows with WER > 0 / > 0.1 79 / 28 79 / 28
1 SIM mean 0.7468 0.7694
4 WER mean 0.0634 0.0570
4 rows with WER > 0 / > 0.1 82 / 22 77 / 20
4 SIM mean 0.7455 0.7689
8 WER mean 0.0573 0.0726
8 WER mean, excluding long009 (Whisper long-form repetition loop on this PR) 0.0574 0.0494
8 rows with WER > 0 / > 0.1 78 / 24 79 / 19
8 SIM mean 0.7422 0.7700

@BruceLoveDecimal BruceLoveDecimal changed the title [Perf][TTS] Bound the flow left context with a sliding window [Perf][CosyVoice3] Add CUDA graphs for the flow DiT estimator Aug 24, 2026
@hsliuustc0106 hsliuustc0106 added the enhancement New feature or request label Aug 25, 2026
@Sy0307

Sy0307 commented Aug 25, 2026

Copy link
Copy Markdown
Collaborator

Will review it with #4876 and #5673. You can also check your PR if it overlaps with these PRs.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/model_integration.md.

Module owners: @tzhouam @gcanlin

@BruceLoveDecimal, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: no human activity for 10 days

@BruceLoveDecimal this pull request has had no human commit, comment or review since 2026-08-29. Per repository policy it may be closed if it stays inactive.

To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline.

@BruceLoveDecimal
BruceLoveDecimal marked this pull request as draft September 14, 2026 18:17
@BruceLoveDecimal
BruceLoveDecimal force-pushed the feat/cozyvoice3_cudagraph branch from 5a6fe21 to b500e54 Compare September 21, 2026 06:39
@BruceLoveDecimal BruceLoveDecimal changed the title [Perf][CosyVoice3] Add CUDA graphs for the flow DiT estimator [Perf][CosyVoice3] Bound the flow left context with a sliding window Sep 21, 2026
@BruceLoveDecimal
BruceLoveDecimal force-pushed the feat/cozyvoice3_cudagraph branch from b500e54 to e5e94f9 Compare September 21, 2026 07:37
@BruceLoveDecimal
BruceLoveDecimal force-pushed the feat/cozyvoice3_cudagraph branch 2 times, most recently from 13f6252 to 3787a7c Compare September 22, 2026 19:06
The streaming code2wav path resent the whole emitted-token prefix to the
flow on every chunk, so per-chunk flow cost grew linearly with elapsed
audio and the per-request cost was O(T^2); long streams also outgrew the
TensorRT estimator's 3000-frame profile. vllm-project#7521 bounded the HiFT vocoder
side of the same pipeline and left this side untouched.

Resend at most ``codec_left_context_frames`` emitted tokens (25 by default,
one second of audio; ``<= 0`` restores the unbounded behaviour) as left
context, and shrink ``left_context_size`` to match so only the new frames
are emitted. The deploy config already carried the key; it now takes effect.
This is a quality trade-off rather than a free win: the DiT attends over
the full sequence, so dropping earlier prefix changes the mel of the new
chunk. WER/SIM were neutral in the PR's A/B; widen the window if seams are
audible.

Measured on an RTX PRO 6000 with the default TensorRT estimator against a
main that already includes vllm-project#7521 (30 requests, three utterance lengths,
reference audio served from a local file):
at concurrency 1 the per-chunk latency no longer climbs from 142 ms to
1182 ms over a long utterance but stays near 230 ms, and RTF drops from
0.171 to 0.113; at concurrency 8 the median per-chunk latency drops from
2008 ms to 700 ms (p95 4025 ms to 963 ms), RTF from 1.25 to 0.38, and
audio throughput rises from 6.0x to 21x realtime.

Signed-off-by: liuqihao <liuqihao970610@gmail.com>
@BruceLoveDecimal
BruceLoveDecimal force-pushed the feat/cozyvoice3_cudagraph branch from 3787a7c to a70ef32 Compare September 23, 2026 18:43
…tion-indexed buffer

Upstream CosyVoice seeds one noise buffer (``torch.randn([1, 80, 50 * 300])``
under seed 0) and slices it by mel position, so the flow starts from the
same noise at a given position on every call. The port replaced that with
a fresh ``torch.randn`` per call, which made every regenerated left
context differ from the frames it was emitted from, and made the same seed
produce different audio across requests.

Restore the buffer (built with a local generator so the global RNG is not
touched, and kept out of the checkpoint) and index it by absolute
position: the prompt always maps to the start of the buffer, and the
frames after it map to ``prompt_len + offset`` where the offset is the
absolute mel index of the first resent token. The streaming code2wav
tracks how many tokens each request has emitted, so a bounded left
context reuses exactly the noise its frames were first generated with,
per row when flow batching is on. The offline path keeps offset 0, which
is upstream's behaviour bit for bit.

Signed-off-by: liuqihao <liuqihao970610@gmail.com>
With a fixed seed a CosyVoice3 request is byte-reproducible on its own,
but under concurrency the same seed yields different audio: the talker's
logits depend on which other requests share its batch, and a one-ulp
difference is enough to flip a sampled token. Audio durations of the same
text spread from 37 s to 65 s at concurrency 8 while being identical at
concurrency 1, which places the variance in the talker, not in code2wav.

vLLM's batch-invariant mode is the fix, but the talker could not run in it:
``llm_decoder`` has ``speech_token_size + 200 = 6761`` outputs, cuBLASLt
rejects that odd leading dimension for fp16/bf16 GEMMs, and batch-invariant
mode forbids the cuBLAS fallback. Pad the head to a multiple of 8 and
slice the logits back, but only in that mode: a different N can change
cuBLAS's algorithm choice, and the default path stays bit-identical (same
c=1 audio before and after).

The deploy config documents the opt-in as a commented per-stage ``env``
on the talker stage. Measured on an RTX PRO 6000 (30 requests, three
utterance lengths, seed 42): byte-identity violations at concurrency 8 go
from 27/30 to 0/30 and the audio is identical across concurrency 1 and 8.
The cost is real, which is why it stays opt-in: about 2x talker latency
at concurrency 1 (RTF 0.114 -> 0.225) and about 35% less throughput at
concurrency 8 (21.0x -> 13.7x realtime).

Signed-off-by: liuqihao <liuqihao970610@gmail.com>
@BruceLoveDecimal
BruceLoveDecimal force-pushed the feat/cozyvoice3_cudagraph branch from a70ef32 to c54a72c Compare September 23, 2026 19:20
@BruceLoveDecimal
BruceLoveDecimal marked this pull request as ready for review September 23, 2026 19:57
@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 24, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit 74004ce108d8 produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot routing record

Assigned Strict on zcode (GLM-5.3-Flash) under experiment fleet-strict-cursor-grok46-zcode-glm53flash-5050-c5-z10-20261002.

@timzsu

timzsu commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

@BruceLoveDecimal Hi, is this PR still active and waiting for review? Please have a look at the latest main, resolve the conflict, and fix the DCO CI.

@BruceLoveDecimal

Copy link
Copy Markdown
Contributor Author

@BruceLoveDecimal Hi, is this PR still active and waiting for review? Please have a look at the latest main, resolve the conflict, and fix the DCO CI.

Thanks for keeping tracking of this PR. I'll rebase my PR and then re-run the benchmark since the main branch had new updates about cozyvoice3. After finishing it, i'll ping you and please help review it. Thanks!

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: 00f174b4-0f70-45c6-9115-1d04e3ada2df) — check zcode login and the CLI version; retrying strict/zcode/GLM-5.3-Flash in 120s (try ).

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: 50508f51-d8ac-45e1-b64f-14f1dbc2c3fc) — check zcode login and the CLI version; retrying strict/zcode/GLM-5.3-Flash in 600s (try ).

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: 7e189a7e-281e-4c23-b520-ad08ea0c39d4) — check zcode login and the CLI version; falling back to direct/cursor/auto).

@vllm-omni-review-bot vllm-omni-review-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Omni ReviewBot review

0 actionable finding(s).

CI at 74004ce108d8 (2026-10-10T03:00:04.529796+00:00): required check(s) blocking: buildkite/vllm-omni (missing).

Note: The assigned review arm strict/zcode/GLM-5.3-Flash could not complete this review, so it was produced by the fallback arm direct/cursor/auto. It is excluded from the routing experiment.

Full review analysis

PR description

CosyVoice3 streaming code2wav no longer feeds the flow the entire emitted token prefix on every chunk. The async chunk processor resends at most codec_left_context_frames tokens (default 25; a nonpositive value keeps the old full prefix) and records that shorter left-context length on the payload. The flow draws its initial noise from one CPU seed-0 buffer, indexed by absolute mel position and by a per-request flow_emitted_tokens counter, so a slid window reuses the noise those frames were first generated with; the offline path still uses offset 0. Separately, when VLLM_BATCH_INVARIANT is set at weight load, the talker llm_decoder matmul is copied out to a multiple of 8 and the extra logits are sliced off before the existing EOS merge. The default head path is unchanged.

Change flow

flowchart LR
  hidden["[EXISTING] Talker hidden states"]:::existing
  head["[NEW] Optional 8-aligned<br/>llm_decoder logits"]:::new
  window["[CHANGED] Chunk processor<br/>bounded left context"]:::changed
  noise["[NEW] Seed-0 position-indexed<br/>flow noise"]:::new
  code2wav["[CHANGED] Code2wav stream state<br/>flow_emitted_tokens"]:::changed
  audio["[EXISTING] PCM chunk"]:::existing
  hidden --> head --> window --> code2wav --> audio
  noise --> code2wav
  classDef existing fill:#e5e7eb,stroke:#6b7280,color:#111827
  classDef changed fill:#fef3c7,stroke:#d97706,color:#451a03,stroke-width:2px
  classDef new fill:#dcfce7,stroke:#16a34a,color:#052e16,stroke-width:2px
  classDef removed fill:#fee2e2,stroke:#dc2626,color:#450a0a,stroke-width:2px
Loading

No actionable findings.


🤖 This review was generated by InferMatrix Copilot, an open-source repo-maintenance agent for PR review, CI debugging and issue triage. Try it on your own repo, and ⭐ star it if it helped!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants