Repository navigation
[Model] Optimize CosyVoice3 TensorRT stream handoff - #5673
Conversation
Replace per-step host synchronization with CUDA stream dependencies and keep TensorRT pointer buffers alive across asynchronous execution. Store the raw CUDA stream in the context pool so wait_stream and cuda_stream are available at runtime. Signed-off-by: Allen Wu <allenwu2795@gmail.com> Co-authored-by: TRAE CLI <noreply@bytedance.com>
|
Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits. |
|
Real TensorRT/checkpoint validation update (2026-08-11): I reran base Pinned artifacts:
Real engine A/B (10 warmups, 40 repeats, fixed
The output arrays are exactly equal (same SHA-256 Full production
The E2E sample is intentionally reported as a correctness/directional check, not a throughput claim: with only two steady samples it does not show a model-level speedup. Both exact commits initially hit the unrelated known bf16 transport bug #5739. I applied the production patch from #5755 commit Validation summary SHA-256: |
Signed-off-by: Allen Wu <allenwu2795@gmail.com> Co-authored-by: TRAE CLI <traecli@bytedance.com>
|
Updated the branch after the main merge in
I also checked the overlap with #4876 and #6588:
A low-conflict landing order would be #5673 first, followed by rebasing the batching and CUDA-graph work while preserving those boundaries. I can adapt this patch if a different order is preferred. |
linyueqian
left a comment
There was a problem hiding this comment.
Reviewed the stream handoff and ran the changed path end to end on an NVIDIA H20 with TensorRT 11.2.1.2 (one major newer than the 10.15 in the test plan), vLLM 0.27.0, torch 2.13.0+cu132, at the current head 68641ae.
On the code: the bidirectional wait_stream ordering is sufficient for the solver's in-place reused CFG buffers (the x_in[:] = x write of the next step lands on the caller stream, which has already waited on the estimator stream), and the record_stream calls are load-bearing rather than defensive, since .to(io_dtype).contiguous() returns the caller-allocated buffer unchanged when the dtype already matches. Releasing the context back to the pool with work still in flight is safe because each context is paired with its own dedicated stream and enqueue snapshots the tensor addresses.
Results at 68641ae:
test_cosyvoice3_components.py -k TestCFM: 3 passedbenchmark_cosyvoice3_trt_streams.py(defaults): median host submission 1.536 ms to 0.021 ms per step (71.6x), GPU elapsed -1.8%, exact parity, zero peak-allocation deltatests/e2e/offline_inference/test_cosyvoice3_expansion.py --run-level advanced_model: 2 passed (sync and async_chunk), withCosyVoice3: using TensorRT flow-decoder estimatorin the log, so the TRT branch is what actually ran- the same e2e at core_model fails with duration=1.840s on both this head and its main-side merge parent 51b7565 (identical sample count), so that failure is inherited from the dummy-weight run level, not from this PR
One heads-up unrelated to this PR: passing a resolved HF snapshot path to Omni currently crashes at stage init on main with "Diffusers pipeline index not found", because name-based model detection became basename-only in #5036 and a snapshot dir basename is a bare revision SHA. I worked around it via VLLM_OMNI_COSYVOICE3_MODEL_DIR for these runs; fix coming separately.
Adding the ready label so Buildkite picks up the new cpu/core_model-marked tests.
Signed-off-by: Allen Wu <allenwu2795@gmail.com> Co-authored-by: TRAE CLI <noreply@bytedance.com> Co-authored-by: Sy03 <1370724210@qq.com> Co-authored-by: TRAE CLI <traecli@bytedance.com> Signed-off-by: AndyZhou952 <jzhoubc@connect.ust.hk>
Signed-off-by: Allen Wu <allenwu2795@gmail.com> Co-authored-by: TRAE CLI <noreply@bytedance.com> Co-authored-by: Sy03 <1370724210@qq.com> Co-authored-by: TRAE CLI <traecli@bytedance.com>
Signed-off-by: Allen Wu <allenwu2795@gmail.com> Co-authored-by: TRAE CLI <noreply@bytedance.com> Co-authored-by: Sy03 <1370724210@qq.com> Co-authored-by: TRAE CLI <traecli@bytedance.com>
Purpose
CosyVoice3's TensorRT CFM estimator performed two host-side CUDA synchronizations for every diffusion step:
execute_async_v3.This blocked the CPU on every step and prevented normal CUDA stream overlap. The context pool also stored the result of
torch.cuda.stream(...), which is aStreamContextrather than a rawtorch.cuda.Stream;StreamContextdoes not exposewait_stream,synchronize, orcuda_stream.This PR:
torch.cuda.Streamwith each TensorRT execution context;Test Plan
vLLM Version:
0.26.0vLLM-Omni Commits:
185e3d6df60276166349e69fdc1171fd0e50c79f738b5cb326bc15fa73a4dfd39c55fd5fa794dee3Focused tests:
Controlled stream-handoff benchmark:
Real TensorRT engine validation used the pinned fp16 estimator ONNX, built/loaded a serialized plan, and executed 10 warmups plus 40 fixed-input repetitions with shape
[2, 80, 500].Full-model validation used production
Omniwith the two-stage CosyVoice3 deploy config, reference-audio voice cloning, fixed sampling seed42, and three synchronous generations for both base and head.Static validation:
Test Results
Environment and pinned artifacts
535.261.032.11.013.110.15.1.29FunAudioLLM/Fun-CosyVoice3-0.5B-251229e01c4e8d000f4bcd70751be16fa94bf3d85a189,747,516,745bytesyuekai/Fun-CosyVoice3-0.5B-2512-FP16-ONNX9564f2b28071b2b9c97d3e39d813ff969b4855763b2052fb9be6d857afe1e1fa63a24dd79d3ed5ba7f1fa557bd13200b0f1f8fe8Generated TensorRT plans:
flow_estimator_23b57a1ddd37630f.plan:666,513,060bytes, SHA-256a385b7ed47be5ee77d408545012adb66b7df71547f7f8893d088a11de5306bfdcampplus_02c8d5975de1eae7.plan:31,098,780bytes, SHA-256f7452b52c95a57408da22dac035171510155ceeb7410b8db890c74195e63790bFocused tests and controlled microbenchmark
Focused pytest:
The controlled benchmark isolates stream handoff and intentionally exaggerates host synchronization overhead:
synchronize()This table is a mechanism-level microbenchmark, not an end-to-end throughput result.
Real TensorRT engine A/B
Both commits built/loaded and executed the real estimator engine. The fixed-input output arrays were exactly equal:
[2, 80, 500], dtypefloat32261f9662041d8f554f20938e48f13af3cf25d768d43be4314c665a5ed47565a00.0The real-engine result confirms a small directional reduction in host and GPU time, not the 49x end-to-end effect suggested by the controlled microbenchmark.
Full two-stage CosyVoice3 E2E
Both base and head logs contain:
CosyVoice3: using TensorRT flow-decoder estimator (code2wav)Loaded flow-estimator TensorRT engineEach commit generated three finite, non-silent 24 kHz WAVs. Every run had
192,960samples (8.04seconds), and each corresponding base/head decoded waveform was exactly equal (MAE 0.0, max absolute difference0.0).This is only two steady-state samples per commit, so it is a correctness and directional E2E check, not a stable throughput claim. It does not show an E2E latency improvement in this run; full-model timing noise and work outside the estimator handoff dominate the small engine-level difference.
Compatibility overlay used for the E2E comparison
The exact base and head commits both hit the known unrelated CosyVoice3 inter-stage bf16 serialization failure tracked by #5739. The production-code patch from #5755 commit
6cd7a8ed2fd0fa617ea4c58da219997b1d8f0e0dwas applied identically to both sides before the final E2E comparison.data_entry_keys.pySHA-256 on both sides:1a2195d87717fcfbf0760c36bf7b3dad6009b0307e74bf26cf4ac256d3ad82154678ccdde9162f1df85c3ca6c23868a13cfc8a10818e8413f800240345f11a2aThis overlay is held constant and is not part of this PR's measured delta.
Validation Boundary