Skip to content

[Model] Add opt-in packed Flow and batched HiFT inference for CosyVoice3 - #8224

Merged
Sy0307 merged 38 commits into
vllm-project:mainfrom
Sy0307:perf/cosyvoice3-packed-inference
Sep 29, 2026
Merged

Sy0307 merged 38 commits into
vllm-project:mainfrom
Sy0307:perf/cosyvoice3-packed-inference

Conversation

@Sy0307

@Sy0307 Sy0307 commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

CosyVoice3 currently pays per-request Flow/HiFT overhead and repeatedly preprocesses unnamed reference audio. This change adds opt-in Hopper packed inference profiles, batches full-response HiFT, reuses content-addressed reference conditioning, and preserves live request conditioning outside AR graph replay. Existing RAS remains the default; an explicit standard-sampling mode supports controlled cross-engine comparisons.

Scope

  • Pack valid Flow frames and run compiled DiT through vLLM FA3; retain all ten Euler steps.
  • Mainline GPU FP32 F0 and cross-request HiFT batching; a separate chunk-causal streaming profile with request-owned GPU HiFT state.
  • Correct conditioning row alignment for mixed prefill/decode, and avoid unnecessary hidden-state payloads in packed profiles.
  • Vectorized random RAS with request-owned RNG; standard sampling preserves all control logits and applies repetition penalties to generated tokens only.
  • Bounded content-addressed reference cache, deployment YAMLs and focused CPU/GPU/API coverage.

Packed DiT is adapted from Apache-2.0 SGLang-Omni commit 127f34b57446a3cb5588ec987da0771a3be675be; this is not a claim of algorithmic originality. GPU F0 now directly follows the mainline implementation, including its CPU opt-out; it is not a new contribution here. Existing Flow batching is not claimed as new.

Stack and attribution

Shared prerequisite #8184 is now merged as 001f8a80b; this branch includes main f145f66c8. Original #7871 weight-normalization commits are reused prerequisites, not independent contributions. Existing #4876 Flow batching, #7521 HiFT caching and #6424 correctness fixes remain baseline functionality.

Current model head: 537287103. Model-only review excludes the shared prerequisite and original #7871 commits.

Historical validation (before the mainline merge)

  • Changed-file repository-pinned pre-commit: passed, including Ruff, mypy, Markdown, inventory/CI checks.
  • H200 physical GPU 0, vLLM 0.30.0, frozen pr-source-v5: 52 model/cache/config tests, 8 GPU tests, 13 API tests passed.
  • C64 Seed-TTS EN1088 aligned performance and quality validation: completed: full response cold/hot/pooled 137.93/176.06/154.41 audio-s/s versus 98.52/100.73/99.60; streaming first nonempty PCM P50/P99 4.111/7.616 s versus 4.996/8.520 s. Only vLLM is rerun against the existing SGLang-Omni results. Stage-0 effective settings are T=0.7, p=0.8, k=20, repetition penalty=1.21, seed=42, ordinary sampling.

Limits

Packed kernels/GEMM shapes are not bitwise equivalent to the default path. These profiles are opt-in and Hopper-specific, with separate streaming and full-response configurations. Identical sampling settings do not imply identical RNG trajectories; SGLang silence filtering and payload size also differ. Report WER, full-audio speaker similarity, cold-reference and hot-reference throughput separately. Cross-framework performance is not a measured net gain versus this PR's dependency base. No claim of zero quality cost or 200 audio-s/s is made.

Aligned quality (vLLM/SGL): full WER 2.051%/1.854%, SIM 70.339/70.164; stream WER 2.079%/1.874%, SIM 70.249/70.465. Paired bootstrap difference intervals cross zero; this is not proof of equivalence. Each mode completed 4352 requests without synthesis failures. Long ASR requests rejected with 503 were retried serially from unchanged saved audio. Cold-reference rounds follow 64 warmups and are not all cache misses.

Current review fixes and validation (2026-09-29)

Use the mainline GPU FP32 F0 path, removing our FP64 variant. Fold F0 weight normalization on its actual inference device (GPU by default, CPU for COSYVOICE3_F0_ON_CPU=1), preserving the mainline forward semantics. Regression tests caught the prior CPU-fold/GPU-forward mismatch; after the fix all 200 Cosy model tests passed.

Request-finished cache removal is synchronized and tolerates missing/repeated IDs. The final-chunk batch shortcut is limited to packed mode. Packed parity tests now compare to the original Flow implementation, and streaming compile coverage invokes real compilation. Full-response and streaming profiles explicitly select their intended Flow backend.

Real checkpoint streaming-profile E2E passed six requests (cold stream, four concurrent streams and one complete response) from Seed-TTS EN on physical H200 GPU 2. After merging concurrent MRv2 Talker/RAS work, both full-response and streaming profiles passed six requests each on physical GPUs 7 and 0. Cancellation followed by a successful request passed on both profiles. This is functional validation, not an updated performance/quality run. The final output-contract fix passes changed-file pre-commit. The newly inherited patch.py also exposes 12 existing mypy errors in previously present monkey patches; these are not CUDA failures and remain outside the CUDA-only follow-up. The previous head passed CUDA CI; current-head validation is listed below.

The historical performance and quality figures above were not remeasured after replacing F0 with the mainline implementation. Cold packed compilation remains a startup/first-request limitation; the cross-framework numbers do not replace a same-repository base/head ablation.

AI assistance: implementation, review fixes and validation were assisted by Codex.

Concurrent MRv2 integration follow-up

Preserved remote commits 2c00cf246 and eb7f2b373 through a merge rather than overwriting their MRv2 Talker, packed Flow graph and fused RAS work. Fresh real-model E2E found a startup failure: the MRv2 Talker returned OmniOutput when packed Flow was enabled only in the codec process. The MRv2 sampler setup now explicitly selects tensor-only forward output, leaving V1 behavior intact. A two-case regression covers both paths; 215 merged model tests plus these two new cases and 48 API/model-state tests passed.

These tests validate the merged implementation functionally, not the old performance claims. After consolidating tiny-model/input fixtures and dropping lower-value duplicate/config checks, the independent added regression increment is 1,000 lines, excluding 547 reused #7871 lines. All 49 tests in the four cleaned files passed on physical H200 GPU 7. No production code changed in this cleanup; the real-model E2E results remain applicable.

Final CUDA CI inventory repair and API protocol correction

Build 16336 passed 26 steps; its remaining Other Test failures were stale environment inventory counts and missing Flow/reference-prefetch entries from the concurrent MRv2 work. Registered all six controls (including the dynamically accessed graph capacity limit) and reviewed their dispositions; all 13 inventory tests passed on the H200 host. Runtime inference behavior is unchanged. Replacement CUDA CI 16337 passed all 27 steps.

The first E2E driver's sixth TTS request used a collecting client but still specified stream_format=audio, which selects server streaming. Corrected the driver and separately tested real nonstreaming API calls with this field omitted: both streaming and full-response deployment profiles passed, with server logs explicitly recording stream=false. Each profile also passed stream cancellation followed by a successful true nonstreaming request. The cold nonstreaming calls included approximately 50 seconds of compilation; these are not steady-state latency measurements.

Main synchronization and DCO (2026-09-29)

Head af478e1c8f878c84fedbe16b900aa8d15a81cbca merges main after the shared framework landed. The environment inventory keeps both main's SeedVR2 variables and the CosyVoice3 controls; expected model-specific/promote counts are 74/45. No CosyVoice3 inference implementation changed in this synchronization.

Added missing Sy03 sign-offs to two owned commits, preserving all source trees. In the subsequent metadata-only update, two Claude coauthor trailers were removed at the author’s request; all other message text and Sy03 DCO sign-offs remain. GitHub DCO now passes. Changed-file checks pass except the 12 existing patch.py mypy diagnostics: clean main and this head produce byte-identical diagnostic output under the same environment; no additional diagnostics.

Fresh H200 validation: 226 model/environment regression tests passed on physical GPU 0. The streaming and full-response deployments each passed six real-checkpoint Seed-TTS EN reference-audio requests on physical GPUs 1/2: one cold stream, four concurrent streams, one true nonstream request. This is functional E2E, not a new throughput or audio-quality qualification.

Current-head CUDA validation

CUDA Buildkite 16352 passed all 27 steps for 53728710390af51412880f053c2b9ebb17f6d5af: https://buildkite.com/vllm/vllm-omni/builds/16352. This metadata-only head has the same source tree as the H200-tested af478e1c8. GitHub DCO passes. Squash auto-merge is enabled; immediate merging was blocked by the base-branch policy while other checks remain queued/pending. This is not a claim that the PR has merged.

Sy0307 and others added 10 commits September 27, 2026 03:32
…kpoint load

CosyVoice3's HiFT vocoder keeps 77 weight-norm parametrizations on frozen
inference weights, so every decode recomputes g*v/||v|| per layer: profiled
as 77 aten::_weight_norm + 77 CUDA weight_norm_fwd_first_dim_kernel events
per chunk (RFC 6870, workstream C4).

The shipped remove_weight_norm() was unusable: it calls the legacy
torch.nn.utils spelling on modules built with the parametrizations API
(ValueError on any current torch), invokes a non-existent
SourceModuleHnNSF.remove_weight_norm(), and walks source_downs layers that
were never weight-normed. Nothing called it.

Fix and wire it into the load path:

- per-layer _fold_weight_norm() supports both APIs (parametrizations API;
  legacy WeightNorm forward-pre-hook via isinstance, same idea as the
  merged indextts2 _strip_weight_norm), returns 0/1, is idempotent, and
  normalizes the materialized weight to a frozen nn.Parameter regardless
  of the caller's grad mode (no_grad/inference_mode callers would
  otherwise get a buffer or an inference tensor);
- ResBlock/HiFTGenerator.remove_weight_norm() return the folded count so a
  silent no-op shows up in the log (same convention as minimax_music3);
- CosyVoice3Code2Wav.load_weights() folds after the strict checkpoint load
  and device move;
- f0_predictor's 5 weight norms are deliberately left parametrized
  (pinned-CPU precision path; RFC C5, to be done with PR 7518).
  step_audio2 and glm_tts import HiFTGenerator but never fold, so they
  are unaffected.

One-time post-load transform: folding rewrites the state_dict schema
(parametrizations.weight.original0/1 -> weight). vllm-omni has no
reload_weights path and sleep/wake restores physical pages via
CuMemAllocator without re-invoking the loader, so nothing in-tree loads
twice; a second load of the original checkpoint would fail the strict
load_state_dict loudly.

Validation (Runpod, vllm 0.29.0, torch 2.13.0+cu130, base d4ffde1):
- tests/model_executor/models/cosyvoice3/: 114 passed (RTX 2000 Ada host);
  shared-consumer checks: minicpmo cuda-graph test 13 passed, step_audio2
  and glm_tts imports OK
- pre-commit run --files <3 changed files>: all gates passed (ruff,
  typos, mypy-3.10, CI marks, SPDX, forbidden imports, torch.cuda guard)
- new L1 regressions (16 cases): fold exactness atol=rtol=0 across
  parametrized and legacy APIs, 22050 and 24000 Hz, finalize and
  streaming; fold counts, idempotency, Parameter-ness (including folding
  under inference_mode); shipped-config generator folds exactly 77 with
  only the 5 f0 layers left; load_weights folds loaded (not initial)
  weights for modern and legacy checkpoint key schemas. All 16 fail on
  unpatched d4ffde1.
- real checkpoint (FunAudioLLM/Fun-CosyVoice3-0.5B-2512, GPU): production
  load_weights logs 'Folded 77 weight-norm layers'; inference() on the
  real hift weights is bit-identical before vs after fold; generator 0 /
  f0 5 parametrizations left; folded weight is a Parameter on CUDA
- CUDA profiler (real-sized generator, decode path): weight-norm events
  231 -> 0; inference output bit-identical (torch.equal=True)
- wall-clock A/B, paired interleaved rounds, alternating arm order,
  bootstrap 95% CI: RTX A4500 (200 rounds, initial fold revision):
  decode() 3.83 ms [3.74, 3.89] chunk 41 (200/200 faster) and 4.51 ms
  [4.33, 4.62] chunk 191 (197/200); RTX 2000 Ada (30 rounds, final head):
  decode() 2.82 and 2.35 ms (30/30 each). Full inference() path is
  dominated by unchanged f0_predictor jitter (C2/7518 scope); no e2e RTF
  claim made (7521's harness cannot resolve this effect on shared GPUs).

Refs: RFC 6870 (C4).
Related: PR 6927 (generic fold helper for MiniCPM-o),
PR 7518 (C2/C5 f0 predictor).

Signed-off-by: zack <huixindaddy@yahoo.com>
Apply the shared normalizer suggested by @0z5a to both weight-norm
APIs so folding under inference_mode cannot leave inference Parameters.
Preserve existing ordinary Parameter flags.

Cover both APIs across grad contexts, cold loading into a legacy-created
vocoder, the loader's transition to eval, and multichunk finalization.
Keep the generator-only 77-layer scope and five F0 norms unchanged.

Related: vllm-project#7871
Signed-off-by: zack <huixindaddy@yahoo.com>
Materialize F0 weights on CPU in FP32 using the shared folding helpers,
while keeping generator and F0 folding as separate entry points.

Add joint C4/C5 loading and streaming regressions for modern and legacy
weight normalization, including strict loading and idempotence.

Signed-off-by: zack <huixindaddy@yahoo.com>
Signed-off-by: Sy03 <1370724210@qq.com>
@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/ar_runtime.md, docs/design/module/model_integration.md.

Module owners: @tzhouam @fake0fan @Gaohan123

Routing: @tzhouam via module of the changed files, semantic router, CODEOWNERS; @fake0fan via module of the changed files, CODEOWNERS; @Gaohan123 via module of the changed files

@Sy0307, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit 53728710390a produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

@Sy0307 Sy0307 added the ready label to trigger buildkite CI label Sep 28, 2026
@Sy0307 Sy0307 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 29, 2026
…coverage

Signed-off-by: Sy03 <1370724210@qq.com>
@Sy0307 Sy0307 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 29, 2026
…nt controls

Signed-off-by: Sy03 <1370724210@qq.com>
@Sy0307 Sy0307 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 29, 2026
…entory

Signed-off-by: Sy03 <1370724210@qq.com>
@Sy0307
Sy0307 force-pushed the perf/cosyvoice3-packed-inference branch 2 times, most recently from aed26af to af478e1 Compare September 29, 2026 14:56
@Sy0307 Sy0307 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 29, 2026
@Sy0307
Sy0307 force-pushed the perf/cosyvoice3-packed-inference branch from af478e1 to 5372871 Compare September 29, 2026 15:11
@Sy0307 Sy0307 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 29, 2026
@Sy0307
Sy0307 enabled auto-merge (squash) September 29, 2026 15:39
@Sy0307
Sy0307 merged commit 1ea0ee3 into vllm-project:main Sep 29, 2026
7 of 9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

high priority high priority issue, needs to be done asap ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants