Repository navigation
[Model] Add opt-in packed Flow and batched HiFT inference for CosyVoice3 - #8224
Conversation
Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Sy03 <1370724210@qq.com>
…kpoint load CosyVoice3's HiFT vocoder keeps 77 weight-norm parametrizations on frozen inference weights, so every decode recomputes g*v/||v|| per layer: profiled as 77 aten::_weight_norm + 77 CUDA weight_norm_fwd_first_dim_kernel events per chunk (RFC 6870, workstream C4). The shipped remove_weight_norm() was unusable: it calls the legacy torch.nn.utils spelling on modules built with the parametrizations API (ValueError on any current torch), invokes a non-existent SourceModuleHnNSF.remove_weight_norm(), and walks source_downs layers that were never weight-normed. Nothing called it. Fix and wire it into the load path: - per-layer _fold_weight_norm() supports both APIs (parametrizations API; legacy WeightNorm forward-pre-hook via isinstance, same idea as the merged indextts2 _strip_weight_norm), returns 0/1, is idempotent, and normalizes the materialized weight to a frozen nn.Parameter regardless of the caller's grad mode (no_grad/inference_mode callers would otherwise get a buffer or an inference tensor); - ResBlock/HiFTGenerator.remove_weight_norm() return the folded count so a silent no-op shows up in the log (same convention as minimax_music3); - CosyVoice3Code2Wav.load_weights() folds after the strict checkpoint load and device move; - f0_predictor's 5 weight norms are deliberately left parametrized (pinned-CPU precision path; RFC C5, to be done with PR 7518). step_audio2 and glm_tts import HiFTGenerator but never fold, so they are unaffected. One-time post-load transform: folding rewrites the state_dict schema (parametrizations.weight.original0/1 -> weight). vllm-omni has no reload_weights path and sleep/wake restores physical pages via CuMemAllocator without re-invoking the loader, so nothing in-tree loads twice; a second load of the original checkpoint would fail the strict load_state_dict loudly. Validation (Runpod, vllm 0.29.0, torch 2.13.0+cu130, base d4ffde1): - tests/model_executor/models/cosyvoice3/: 114 passed (RTX 2000 Ada host); shared-consumer checks: minicpmo cuda-graph test 13 passed, step_audio2 and glm_tts imports OK - pre-commit run --files <3 changed files>: all gates passed (ruff, typos, mypy-3.10, CI marks, SPDX, forbidden imports, torch.cuda guard) - new L1 regressions (16 cases): fold exactness atol=rtol=0 across parametrized and legacy APIs, 22050 and 24000 Hz, finalize and streaming; fold counts, idempotency, Parameter-ness (including folding under inference_mode); shipped-config generator folds exactly 77 with only the 5 f0 layers left; load_weights folds loaded (not initial) weights for modern and legacy checkpoint key schemas. All 16 fail on unpatched d4ffde1. - real checkpoint (FunAudioLLM/Fun-CosyVoice3-0.5B-2512, GPU): production load_weights logs 'Folded 77 weight-norm layers'; inference() on the real hift weights is bit-identical before vs after fold; generator 0 / f0 5 parametrizations left; folded weight is a Parameter on CUDA - CUDA profiler (real-sized generator, decode path): weight-norm events 231 -> 0; inference output bit-identical (torch.equal=True) - wall-clock A/B, paired interleaved rounds, alternating arm order, bootstrap 95% CI: RTX A4500 (200 rounds, initial fold revision): decode() 3.83 ms [3.74, 3.89] chunk 41 (200/200 faster) and 4.51 ms [4.33, 4.62] chunk 191 (197/200); RTX 2000 Ada (30 rounds, final head): decode() 2.82 and 2.35 ms (30/30 each). Full inference() path is dominated by unchanged f0_predictor jitter (C2/7518 scope); no e2e RTF claim made (7521's harness cannot resolve this effect on shared GPUs). Refs: RFC 6870 (C4). Related: PR 6927 (generic fold helper for MiniCPM-o), PR 7518 (C2/C5 f0 predictor). Signed-off-by: zack <huixindaddy@yahoo.com>
Apply the shared normalizer suggested by @0z5a to both weight-norm APIs so folding under inference_mode cannot leave inference Parameters. Preserve existing ordinary Parameter flags. Cover both APIs across grad contexts, cold loading into a legacy-created vocoder, the loader's transition to eval, and multichunk finalization. Keep the generator-only 77-layer scope and five F0 norms unchanged. Related: vllm-project#7871 Signed-off-by: zack <huixindaddy@yahoo.com>
Materialize F0 weights on CPU in FP32 using the shared folding helpers, while keeping generator and F0 folding as separate entry points. Add joint C4/C5 loading and streaming regressions for modern and legacy weight normalization, including strict loading and idempotence. Signed-off-by: zack <huixindaddy@yahoo.com>
Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Sy03 <1370724210@qq.com>
Signed-off-by: Sy03 <1370724210@qq.com>
|
This PR appears to belong to: docs/design/module/ar_runtime.md, docs/design/module/model_integration.md. Module owners: @tzhouam @fake0fan @Gaohan123 Routing: @tzhouam via module of the changed files, semantic router, CODEOWNERS; @fake0fan via module of the changed files, CODEOWNERS; @Gaohan123 via module of the changed files @Sy0307, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Omni ReviewBot triage noteAutomated triage of commit
These are automated triage suggestions only — the final decision belongs to the maintainers. |
Signed-off-by: Sy03 <1370724210@qq.com>
…coverage Signed-off-by: Sy03 <1370724210@qq.com>
…nt controls Signed-off-by: Sy03 <1370724210@qq.com>
…entory Signed-off-by: Sy03 <1370724210@qq.com>
aed26af to
af478e1
Compare
af478e1 to
5372871
Compare
CosyVoice3 currently pays per-request Flow/HiFT overhead and repeatedly preprocesses unnamed reference audio. This change adds opt-in Hopper packed inference profiles, batches full-response HiFT, reuses content-addressed reference conditioning, and preserves live request conditioning outside AR graph replay. Existing RAS remains the default; an explicit standard-sampling mode supports controlled cross-engine comparisons.
Scope
Packed DiT is adapted from Apache-2.0 SGLang-Omni commit
127f34b57446a3cb5588ec987da0771a3be675be; this is not a claim of algorithmic originality. GPU F0 now directly follows the mainline implementation, including its CPU opt-out; it is not a new contribution here. Existing Flow batching is not claimed as new.Stack and attribution
Shared prerequisite #8184 is now merged as
001f8a80b; this branch includes mainf145f66c8. Original #7871 weight-normalization commits are reused prerequisites, not independent contributions. Existing #4876 Flow batching, #7521 HiFT caching and #6424 correctness fixes remain baseline functionality.Current model head:
537287103. Model-only review excludes the shared prerequisite and original #7871 commits.Historical validation (before the mainline merge)
pr-source-v5: 52 model/cache/config tests, 8 GPU tests, 13 API tests passed.Limits
Packed kernels/GEMM shapes are not bitwise equivalent to the default path. These profiles are opt-in and Hopper-specific, with separate streaming and full-response configurations. Identical sampling settings do not imply identical RNG trajectories; SGLang silence filtering and payload size also differ. Report WER, full-audio speaker similarity, cold-reference and hot-reference throughput separately. Cross-framework performance is not a measured net gain versus this PR's dependency base. No claim of zero quality cost or 200 audio-s/s is made.
Aligned quality (vLLM/SGL): full WER 2.051%/1.854%, SIM 70.339/70.164; stream WER 2.079%/1.874%, SIM 70.249/70.465. Paired bootstrap difference intervals cross zero; this is not proof of equivalence. Each mode completed 4352 requests without synthesis failures. Long ASR requests rejected with 503 were retried serially from unchanged saved audio. Cold-reference rounds follow 64 warmups and are not all cache misses.
Current review fixes and validation (2026-09-29)
Use the mainline GPU FP32 F0 path, removing our FP64 variant. Fold F0 weight normalization on its actual inference device (GPU by default, CPU for
COSYVOICE3_F0_ON_CPU=1), preserving the mainline forward semantics. Regression tests caught the prior CPU-fold/GPU-forward mismatch; after the fix all 200 Cosy model tests passed.Request-finished cache removal is synchronized and tolerates missing/repeated IDs. The final-chunk batch shortcut is limited to packed mode. Packed parity tests now compare to the original Flow implementation, and streaming compile coverage invokes real compilation. Full-response and streaming profiles explicitly select their intended Flow backend.
Real checkpoint streaming-profile E2E passed six requests (cold stream, four concurrent streams and one complete response) from Seed-TTS EN on physical H200 GPU 2. After merging concurrent MRv2 Talker/RAS work, both full-response and streaming profiles passed six requests each on physical GPUs 7 and 0. Cancellation followed by a successful request passed on both profiles. This is functional validation, not an updated performance/quality run. The final output-contract fix passes changed-file pre-commit. The newly inherited
patch.pyalso exposes 12 existing mypy errors in previously present monkey patches; these are not CUDA failures and remain outside the CUDA-only follow-up. The previous head passed CUDA CI; current-head validation is listed below.The historical performance and quality figures above were not remeasured after replacing F0 with the mainline implementation. Cold packed compilation remains a startup/first-request limitation; the cross-framework numbers do not replace a same-repository base/head ablation.
AI assistance: implementation, review fixes and validation were assisted by Codex.
Concurrent MRv2 integration follow-up
Preserved remote commits
2c00cf246andeb7f2b373through a merge rather than overwriting their MRv2 Talker, packed Flow graph and fused RAS work. Fresh real-model E2E found a startup failure: the MRv2 Talker returnedOmniOutputwhen packed Flow was enabled only in the codec process. The MRv2 sampler setup now explicitly selects tensor-only forward output, leaving V1 behavior intact. A two-case regression covers both paths; 215 merged model tests plus these two new cases and 48 API/model-state tests passed.These tests validate the merged implementation functionally, not the old performance claims. After consolidating tiny-model/input fixtures and dropping lower-value duplicate/config checks, the independent added regression increment is 1,000 lines, excluding 547 reused #7871 lines. All 49 tests in the four cleaned files passed on physical H200 GPU 7. No production code changed in this cleanup; the real-model E2E results remain applicable.
Final CUDA CI inventory repair and API protocol correction
Build 16336 passed 26 steps; its remaining Other Test failures were stale environment inventory counts and missing Flow/reference-prefetch entries from the concurrent MRv2 work. Registered all six controls (including the dynamically accessed graph capacity limit) and reviewed their dispositions; all 13 inventory tests passed on the H200 host. Runtime inference behavior is unchanged. Replacement CUDA CI 16337 passed all 27 steps.
The first E2E driver's sixth TTS request used a collecting client but still specified
stream_format=audio, which selects server streaming. Corrected the driver and separately tested real nonstreaming API calls with this field omitted: both streaming and full-response deployment profiles passed, with server logs explicitly recordingstream=false. Each profile also passed stream cancellation followed by a successful true nonstreaming request. The cold nonstreaming calls included approximately 50 seconds of compilation; these are not steady-state latency measurements.Main synchronization and DCO (2026-09-29)
Head
af478e1c8f878c84fedbe16b900aa8d15a81cbcamerges main after the shared framework landed. The environment inventory keeps both main's SeedVR2 variables and the CosyVoice3 controls; expected model-specific/promote counts are 74/45. No CosyVoice3 inference implementation changed in this synchronization.Added missing Sy03 sign-offs to two owned commits, preserving all source trees. In the subsequent metadata-only update, two Claude coauthor trailers were removed at the author’s request; all other message text and Sy03 DCO sign-offs remain. GitHub DCO now passes. Changed-file checks pass except the 12 existing patch.py mypy diagnostics: clean main and this head produce byte-identical diagnostic output under the same environment; no additional diagnostics.
Fresh H200 validation: 226 model/environment regression tests passed on physical GPU 0. The streaming and full-response deployments each passed six real-checkpoint Seed-TTS EN reference-audio requests on physical GPUs 1/2: one cold stream, four concurrent streams, one true nonstream request. This is functional E2E, not a new throughput or audio-quality qualification.
Current-head CUDA validation
CUDA Buildkite 16352 passed all 27 steps for
53728710390af51412880f053c2b9ebb17f6d5af: https://buildkite.com/vllm/vllm-omni/builds/16352. This metadata-only head has the same source tree as the H200-testedaf478e1c8. GitHub DCO passes. Squash auto-merge is enabled; immediate merging was blocked by the base-branch policy while other checks remain queued/pending. This is not a claim that the PR has merged.