Repository navigation
Conversation
…S path vllm-project#7413 moved the Seed-TTS perf entry from /v1/chat/completions (with the TTS chat template) to /v1/realtime. A model-native duplex session only speaks once it has heard enough audio units, and the benchmark client feeds silence in real time, so every request pays a listening phase before the first audio packet, which the shared metric then counts in TTFT/e2el. Restore the entry to the chat-completions TTS path (revert the config to its pre-vllm-project#7413 state). Measured on L20X, same harness: 1x32 mean_audio_rtf 1.35 -> 0.31, mean_ttft_ms 2461 -> 122, mean_e2el_ms 5472 -> 1260; 4x64 rtf 1.37 -> 0.62, ttft 2668 -> 181.
|
This PR matches CODEOWNERS paths: /tests/. Code owners: @NickCao @yenuo26 Routing: @NickCao via CODEOWNERS; @yenuo26 via CODEOWNERS @y-null, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
The pre-vllm-project#7413 baseline values (H100/A3) do not match what the restored chat-completions path measures on current hardware, so leave the field unset instead of restoring stale numbers. Only the request path and its extra_body are restored.
Omni ReviewBot triage noteAutomated triage of commit
These are automated triage suggestions only — the final decision belongs to the maintainers. |
|
@y-null Please resolve conflicts |
|
+1 on restoring the chat-completions entry without the pre-#7413 baseline block. From the nightly data in #7865 (details there), those numbers are the Aug 13 run: before the benchmark's Two suggestions for the follow-up re-baseline, building on @tlysanhuo's point in #7865:
Happy to provide the |
|
done |
|
done |
Problem
The MiniCPM-o 4.5 Seed-TTS perf series regressed sharply on the performance board on 9/16:
mean_audio_rtfwent from 0.29-0.45 to 1.32-1.35, andmean_ttft_msfrom ~150ms to ~2400ms. No other model shows a comparable change.Cause
#7413 moved the Seed-TTS perf entry off the chat-completions TTS path and onto the duplex realtime path:
backendopenai-chat-omniopenai-realtime-ttsendpoint/v1/chat/completions/v1/realtimeextra_bodymodalities=[text,audio],chat_template_kwargs.use_tts_template=trueA model-native duplex session only speaks after it has heard enough audio units, and the benchmark client feeds those units (silence) in real time. Every request therefore pays an extra listening phase before the first audio packet, and the shared metric counts that phase in TTFT/e2el — a ~1.4-1.9s per-request overhead that accounts for the whole regression. The chat-completions TTS path drives the same model from text directly and has no such phase.
Fix
Restore the Seed-TTS perf entry to the chat-completions TTS path. No source change.
The pre-#7413
baselineblock is deliberately not restored: those values (H100/A3, measured months ago) no longer describe what this path produces on current hardware, so the field stays unset rather than re-introducing stale reference numbers.Results
L20X, 1 card, same harness and same seed-tts dataset:
mean_audio_rtfmean_ttft_msmean_e2el_msdurationTTFT returns to the pre-#7413 level (~122ms) and rtf is better than the old baseline (0.45), because the restored path also benefits from the later code2wav/CUDA-graph work.
Note
This restores the TTS-facing perf series. If the duplex realtime path should keep its own perf coverage, it is better covered by a separate entry rather than by sharing one series with a different request protocol — the two are not comparable, and mixing them is what made the board look like a framework regression.