Skip to content

[Model] Optimize CosyVoice3 TensorRT stream handoff - #5673

Merged
linyueqian merged 3 commits into
vllm-project:mainfrom
EchoHayate:perf/cosyvoice-trt-stream-sync
Aug 26, 2026
Merged

linyueqian merged 3 commits into
vllm-project:mainfrom
EchoHayate:perf/cosyvoice-trt-stream-sync

Conversation

@EchoHayate

@EchoHayate EchoHayate commented Aug 1, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

CosyVoice3's TensorRT CFM estimator performed two host-side CUDA synchronizations for every diffusion step:

  1. synchronize the caller stream before switching to the estimator stream;
  2. synchronize the estimator stream immediately after execute_async_v3.

This blocked the CPU on every step and prevented normal CUDA stream overlap. The context pool also stored the result of torch.cuda.stream(...), which is a StreamContext rather than a raw torch.cuda.Stream; StreamContext does not expose wait_stream, synchronize, or cuda_stream.

This PR:

  • stores a real torch.cuda.Stream with each TensorRT execution context;
  • replaces both host synchronizations with caller-to-estimator and estimator-to-caller stream dependencies;
  • records all raw-pointer input/output tensors on the estimator stream so the caching allocator cannot reuse them before TensorRT finishes;
  • records the estimator output on the caller stream before the dtype conversion/consumer path;
  • adds regression coverage for zero host synchronizations, bidirectional dependencies, the CUDA stream handle passed to TensorRT, context/stream pair reuse, and pool release.

Test Plan

vLLM Version: 0.26.0

vLLM-Omni Commits:

  • base: 185e3d6df60276166349e69fdc1171fd0e50c79f
  • PR head: 738b5cb326bc15fa73a4dfd39c55fd5fa794dee3

Focused tests:

python -m pytest -q -rs \
  tests/model_executor/models/cosyvoice3/test_cosyvoice3_components.py \
  -k TestCFM

Controlled stream-handoff benchmark:

python benchmarks/tts/benchmark_cosyvoice3_trt_streams.py \
  --device cuda:0 \
  --steps 100 \
  --warmup-steps 10 \
  --repeats 7 \
  --producer-cycles 1000000 \
  --estimator-cycles 2000000

Real TensorRT engine validation used the pinned fp16 estimator ONNX, built/loaded a serialized plan, and executed 10 warmups plus 40 fixed-input repetitions with shape [2, 80, 500].

Full-model validation used production Omni with the two-stage CosyVoice3 deploy config, reference-audio voice cloning, fixed sampling seed 42, and three synchronous generations for both base and head.

Static validation:

python -m py_compile \
  benchmarks/tts/benchmark_cosyvoice3_trt_streams.py \
  tests/model_executor/models/cosyvoice3/test_cosyvoice3_components.py \
  vllm_omni/model_executor/models/cosyvoice3/code2wav_core/cfm.py \
  vllm_omni/model_executor/models/cosyvoice3/flow_estimator_trt.py

uvx --offline ruff check \
  benchmarks/tts/benchmark_cosyvoice3_trt_streams.py \
  tests/model_executor/models/cosyvoice3/test_cosyvoice3_components.py \
  vllm_omni/model_executor/models/cosyvoice3/code2wav_core/cfm.py \
  vllm_omni/model_executor/models/cosyvoice3/flow_estimator_trt.py

uvx --offline ruff format --check \
  benchmarks/tts/benchmark_cosyvoice3_trt_streams.py \
  tests/model_executor/models/cosyvoice3/test_cosyvoice3_components.py \
  vllm_omni/model_executor/models/cosyvoice3/code2wav_core/cfm.py \
  vllm_omni/model_executor/models/cosyvoice3/flow_estimator_trt.py

Test Results

Environment and pinned artifacts

  • NVIDIA A100 80GB PCIe
  • NVIDIA driver 535.261.03
  • PyTorch 2.11.0
  • CUDA 13.1
  • TensorRT Python/runtime 10.15.1.29
  • CosyVoice3 checkpoint: FunAudioLLM/Fun-CosyVoice3-0.5B-2512
    • revision 29e01c4e8d000f4bcd70751be16fa94bf3d85a18
    • 20 files, 9,747,516,745 bytes
  • fp16 estimator ONNX: yuekai/Fun-CosyVoice3-0.5B-2512-FP16-ONNX
    • revision 9564f2b28071b2b9c97d3e39d813ff969b485576
    • SHA-256 3b2052fb9be6d857afe1e1fa63a24dd79d3ed5ba7f1fa557bd13200b0f1f8fe8

Generated TensorRT plans:

  • flow_estimator_23b57a1ddd37630f.plan: 666,513,060 bytes, SHA-256 a385b7ed47be5ee77d408545012adb66b7df71547f7f8893d088a11de5306bfd
  • campplus_02c8d5975de1eae7.plan: 31,098,780 bytes, SHA-256 f7452b52c95a57408da22dac035171510155ceeb7410b8db890c74195e63790b

Focused tests and controlled microbenchmark

Focused pytest:

3 passed, 12 deselected, 19 warnings in 0.11s

The controlled benchmark isolates stream handoff and intentionally exaggerates host synchronization overhead:

Metric Host synchronize() Stream dependencies Change
Median host submission latency per step 2.1531 ms 0.0439 ms 49.04x faster
Median GPU elapsed time per 100 steps 219.9491 ms 213.6760 ms -2.85%
Peak allocation delta during measured loop 0 bytes 0 bytes unchanged
Final tensor value 100 100 exact parity in all 7 rounds

This table is a mechanism-level microbenchmark, not an end-to-end throughput result.

Real TensorRT engine A/B

Both commits built/loaded and executed the real estimator engine. The fixed-input output arrays were exactly equal:

  • output shape [2, 80, 500], dtype float32
  • output SHA-256 for both base and head: 261f9662041d8f554f20938e48f13af3cf25d768d43be4314c665a5ed47565a0
  • MAE 0.0
Metric Base Head Change
Median host submission, 40 repeats 5.6082 ms 5.4075 ms -3.58%
GPU elapsed 224.5079 ms 216.8074 ms -3.43%
Peak GPU memory 16,453 MiB 16,493 MiB +40 MiB

The real-engine result confirms a small directional reduction in host and GPU time, not the 49x end-to-end effect suggested by the controlled microbenchmark.

Full two-stage CosyVoice3 E2E

Both base and head logs contain:

  • CosyVoice3: using TensorRT flow-decoder estimator (code2wav)
  • Loaded flow-estimator TensorRT engine

Each commit generated three finite, non-silent 24 kHz WAVs. Every run had 192,960 samples (8.04 seconds), and each corresponding base/head decoded waveform was exactly equal (MAE 0.0, max absolute difference 0.0).

Steady metric (runs 2-3 only) Base Head Change
Median request latency 1.2746 s 1.3520 s +6.07%
Median RTF 0.1585 0.1682 +6.07%
Peak GPU memory for full run 45,129 MiB 46,121 MiB +992 MiB

This is only two steady-state samples per commit, so it is a correctness and directional E2E check, not a stable throughput claim. It does not show an E2E latency improvement in this run; full-model timing noise and work outside the estimator handoff dominate the small engine-level difference.

Compatibility overlay used for the E2E comparison

The exact base and head commits both hit the known unrelated CosyVoice3 inter-stage bf16 serialization failure tracked by #5739. The production-code patch from #5755 commit 6cd7a8ed2fd0fa617ea4c58da219997b1d8f0e0d was applied identically to both sides before the final E2E comparison.

  • original data_entry_keys.py SHA-256 on both sides: 1a2195d87717fcfbf0760c36bf7b3dad6009b0307e74bf26cf4ac256d3ad8215
  • overlaid file SHA-256 on both sides: 4678ccdde9162f1df85c3ca6c23868a13cfc8a10818e8413f800240345f11a2a
  • bf16 and fp16 dtype/shape/value round-trip checks passed on both sides

This overlay is held constant and is not part of this PR's measured delta.

Validation Boundary

  • The focused benchmark proves the stream-dependency mechanism removes host synchronizations.
  • The real-engine A/B proves fixed-input numerical parity and a small directional engine-level timing improvement.
  • The full E2E proves both commits run the pinned real CosyVoice3 checkpoint through the TensorRT code2wav path and produce identical non-silent audio samples.
  • The three-run E2E sample is too small to claim stable throughput or a model-level speedup.

Replace per-step host synchronization with CUDA stream dependencies and keep TensorRT pointer buffers alive across asynchronous execution. Store the raw CUDA stream in the context pool so wait_stream and cuda_stream are available at runtime.

Signed-off-by: Allen Wu <allenwu2795@gmail.com>
Co-authored-by: TRAE CLI <noreply@bytedance.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@hsliuustc0106 hsliuustc0106 added tts code related to tts models Kernel optimization Codes related to optimize kernel execution to improve hardware utilization labels Aug 2, 2026
@EchoHayate

Copy link
Copy Markdown
Contributor Author

Real TensorRT/checkpoint validation update (2026-08-11):

I reran base 185e3d6df60276166349e69fdc1171fd0e50c79f and head 738b5cb326bc15fa73a4dfd39c55fd5fa794dee3 on an A100 80GB PCIe with TensorRT 10.15.1.29, PyTorch 2.11.0, and CUDA 13.1.

Pinned artifacts:

  • FunAudioLLM/Fun-CosyVoice3-0.5B-2512 revision 29e01c4e8d000f4bcd70751be16fa94bf3d85a18 (20 files, 9,747,516,745 bytes)
  • fp16 estimator ONNX revision 9564f2b28071b2b9c97d3e39d813ff969b485576, SHA-256 3b2052fb9be6d857afe1e1fa63a24dd79d3ed5ba7f1fa557bd13200b0f1f8fe8
  • generated flow plan: 666,513,060 bytes, SHA-256 a385b7ed47be5ee77d408545012adb66b7df71547f7f8893d088a11de5306bfd

Real engine A/B (10 warmups, 40 repeats, fixed [2, 80, 500] input):

Metric Base Head
Median host submission 5.6082 ms 5.4075 ms
GPU elapsed 224.5079 ms 216.8074 ms
Peak GPU memory 16,453 MiB 16,493 MiB

The output arrays are exactly equal (same SHA-256 261f9662041d8f554f20938e48f13af3cf25d768d43be4314c665a5ed47565a0, MAE 0.0). This is about a 3.58% host-submission and 3.43% GPU-time reduction in the real engine path; the earlier 49x result remains only a controlled stream-handoff microbenchmark.

Full production Omni two-stage CosyVoice3 voice-clone E2E:

  • base and head both activated and loaded the TensorRT flow-estimator path
  • three finite, non-silent 24 kHz WAVs per commit
  • every output was 192,960 samples / 8.04 seconds
  • corresponding decoded base/head waveforms were exactly equal (MAE and max absolute difference both 0.0)
  • steady median latency (runs 2-3): base 1.2746 s, head 1.3520 s
  • steady median RTF: base 0.1585, head 0.1682
  • peak GPU memory: base 45,129 MiB, head 46,121 MiB

The E2E sample is intentionally reported as a correctness/directional check, not a throughput claim: with only two steady samples it does not show a model-level speedup.

Both exact commits initially hit the unrelated known bf16 transport bug #5739. I applied the production patch from #5755 commit 6cd7a8ed2fd0fa617ea4c58da219997b1d8f0e0d identically to base and head; bf16/fp16 dtype, shape, and value round trips passed on both. The overlaid production file has the same SHA-256 on both sides, so it is held constant and excluded from the PR delta.

Validation summary SHA-256: ba45ad846e4b06e062206cdde6e277ade7abcaec3fb9f416eefa33d2ca3f0d38.

Signed-off-by: Allen Wu <allenwu2795@gmail.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
@EchoHayate

Copy link
Copy Markdown
Contributor Author

Updated the branch after the main merge in 68641ae1:

  • migrated the benchmark memory/device calls to the device-agnostic APIs required by current main;
  • pre-commit, DCO, and Python 3.11/3.12 builds are green;
  • the production TensorRT stream-handoff code is unchanged.

I also checked the overlap with #4876 and #6588:

  • [Perf][CosyVoice3] Bound the flow left context with a sliding window #6588 enables CUDA graph replay only when the estimator is a torch.nn.Module; the TensorRT estimator in this PR remains on the separate non-module branch. The two changes share cfm.py textually but address distinct runtime paths.
  • Optimize CosyVoice3 Stage1 flow batching #4876 changes the same TensorRT branch to support dynamic 2 * batch_size shapes. When it is rebased, that branch should retain this PRs caller-to-estimator / estimator-to-caller wait_stream dependencies, raw CUDA stream handle, and record_stream lifetime protection.

A low-conflict landing order would be #5673 first, followed by rebasing the batching and CUDA-graph work while preserving those boundaries. I can adapt this patch if a different order is preferred.

@linyueqian linyueqian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the stream handoff and ran the changed path end to end on an NVIDIA H20 with TensorRT 11.2.1.2 (one major newer than the 10.15 in the test plan), vLLM 0.27.0, torch 2.13.0+cu132, at the current head 68641ae.

On the code: the bidirectional wait_stream ordering is sufficient for the solver's in-place reused CFG buffers (the x_in[:] = x write of the next step lands on the caller stream, which has already waited on the estimator stream), and the record_stream calls are load-bearing rather than defensive, since .to(io_dtype).contiguous() returns the caller-allocated buffer unchanged when the dtype already matches. Releasing the context back to the pool with work still in flight is safe because each context is paired with its own dedicated stream and enqueue snapshots the tensor addresses.

Results at 68641ae:

  • test_cosyvoice3_components.py -k TestCFM: 3 passed
  • benchmark_cosyvoice3_trt_streams.py (defaults): median host submission 1.536 ms to 0.021 ms per step (71.6x), GPU elapsed -1.8%, exact parity, zero peak-allocation delta
  • tests/e2e/offline_inference/test_cosyvoice3_expansion.py --run-level advanced_model: 2 passed (sync and async_chunk), with CosyVoice3: using TensorRT flow-decoder estimator in the log, so the TRT branch is what actually ran
  • the same e2e at core_model fails with duration=1.840s on both this head and its main-side merge parent 51b7565 (identical sample count), so that failure is inherited from the dummy-weight run level, not from this PR

One heads-up unrelated to this PR: passing a resolved HF snapshot path to Omni currently crashes at stage init on main with "Diffusers pipeline index not found", because name-based model detection became basename-only in #5036 and a snapshot dir basename is a bare revision SHA. I worked around it via VLLM_OMNI_COSYVOICE3_MODEL_DIR for these runs; fix coming separately.

Adding the ready label so Buildkite picks up the new cpu/core_model-marked tests.

@linyueqian linyueqian added the ready label to trigger buildkite CI label Aug 25, 2026
@linyueqian
linyueqian merged commit f90e992 into vllm-project:main Aug 26, 2026
8 of 9 checks passed
AndyZhou952 pushed a commit to AndyZhou952/vllm-omni that referenced this pull request Aug 26, 2026
Signed-off-by: Allen Wu <allenwu2795@gmail.com>
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: Sy03 <1370724210@qq.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Signed-off-by: AndyZhou952 <jzhoubc@connect.ust.hk>
JoseCarlosGarcia95 pushed a commit to valendra-tech/vllm-omni that referenced this pull request Sep 5, 2026
Signed-off-by: Allen Wu <allenwu2795@gmail.com>
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: Sy03 <1370724210@qq.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
Signed-off-by: Allen Wu <allenwu2795@gmail.com>
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: Sy03 <1370724210@qq.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Kernel optimization Codes related to optimize kernel execution to improve hardware utilization ready label to trigger buildkite CI tts code related to tts models

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants