Skip to content

[Perf] Restore MiniCPM-o 4.5 Seed-TTS perf to the chat-completions TTS path - #7806

Closed
y-null wants to merge 2 commits into
vllm-project:mainfrom
y-null:fix/minicpmo-seed-tts-perf
Closed

y-null wants to merge 2 commits into
vllm-project:mainfrom
y-null:fix/minicpmo-seed-tts-perf

Conversation

@y-null

@y-null y-null commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Problem

The MiniCPM-o 4.5 Seed-TTS perf series regressed sharply on the performance board on 9/16: mean_audio_rtf went from 0.29-0.45 to 1.32-1.35, and mean_ttft_ms from ~150ms to ~2400ms. No other model shows a comparable change.

Cause

#7413 moved the Seed-TTS perf entry off the chat-completions TTS path and onto the duplex realtime path:

field before after
backend openai-chat-omni openai-realtime-tts
endpoint /v1/chat/completions /v1/realtime
extra_body modalities=[text,audio], chat_template_kwargs.use_tts_template=true dropped

A model-native duplex session only speaks after it has heard enough audio units, and the benchmark client feeds those units (silence) in real time. Every request therefore pays an extra listening phase before the first audio packet, and the shared metric counts that phase in TTFT/e2el — a ~1.4-1.9s per-request overhead that accounts for the whole regression. The chat-completions TTS path drives the same model from text directly and has no such phase.

Fix

Restore the Seed-TTS perf entry to the chat-completions TTS path. No source change.

The pre-#7413 baseline block is deliberately not restored: those values (H100/A3, measured months ago) no longer describe what this path produces on current hardware, so the field stays unset rather than re-introducing stale reference numbers.

Results

L20X, 1 card, same harness and same seed-tts dataset:

config mean_audio_rtf mean_ttft_ms mean_e2el_ms duration
before, 1x32 1.35 2461 5472 175s
after, 1x32 0.31 122 1260 40s
before, 4x64 1.37 2668 5690 —
after, 4x64 0.62 181 2620 43s

TTFT returns to the pre-#7413 level (~122ms) and rtf is better than the old baseline (0.45), because the restored path also benefits from the later code2wav/CUDA-graph work.

Note

This restores the TTS-facing perf series. If the duplex realtime path should keep its own perf coverage, it is better covered by a separate entry rather than by sharing one series with a different request protocol — the two are not comparable, and mixing them is what made the board look like a framework regression.

…S path

vllm-project#7413 moved the Seed-TTS perf entry from /v1/chat/completions (with the TTS chat template) to /v1/realtime. A model-native duplex session only speaks once it has heard enough audio units, and the benchmark client feeds silence in real time, so every request pays a listening phase before the first audio packet, which the shared metric then counts in TTFT/e2el.

Restore the entry to the chat-completions TTS path (revert the config to its pre-vllm-project#7413 state). Measured on L20X, same harness: 1x32 mean_audio_rtf 1.35 -> 0.31, mean_ttft_ms 2461 -> 122, mean_e2el_ms 5472 -> 1260; 4x64 rtf 1.37 -> 0.62, ttft 2668 -> 181.
@vllm-omni-review-bot

Copy link
Copy Markdown

This PR matches CODEOWNERS paths: /tests/.

Code owners: @NickCao @yenuo26

Routing: @NickCao via CODEOWNERS; @yenuo26 via CODEOWNERS

@y-null, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

The pre-vllm-project#7413 baseline values (H100/A3) do not match what the restored chat-completions path measures on current hardware, so leave the field unset instead of restoring stale numbers. Only the request path and its extra_body are restored.
@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit 5189873dcbf0 produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

@hsliuustc0106 hsliuustc0106 added tts code related to tts models benchmark/profiler/metrics/logger Codes related to benchmarks, profiler, metrics and logger system labels Sep 18, 2026
@Gaohan123 Gaohan123 added this to the v0.30.0 milestone Sep 19, 2026
@Gaohan123

Copy link
Copy Markdown
Collaborator

@y-null Please resolve conflicts

@MinhaoLi0318

Copy link
Copy Markdown
Contributor

+1 on restoring the chat-completions entry without the pre-#7413 baseline block. From the nightly data in #7865 (details there), those numbers are the Aug 13 run: before the benchmark's ref_audio took effect (#6628), before the duplex YAML (#6619) and before the Talker sampler changes (#6346/#6458), so they cannot be a reference for what this entry measures today.

Two suggestions for the follow-up re-baseline, building on @tlysanhuo's point in #7865:

  1. Keep two benchmark entries rather than one: seed-tts (per-request reference audio, the voice-cloning path) and seed-tts-text (same prompts, default voice). The cloning path currently carries ~+160 ms of stage-2 first-packet work per request that the non-cloning path does not, so a single cloning-tied baseline would hide regressions on the other path.
  2. Take the new baseline from a 7-day window only after the stage-2 reference-conditioning fix lands (tracked in [Bug][Perf]: minicpm, performance metrics regressed by more than 10% compared to the baseline in some scenarios #7865); re-baselining on today's numbers would lock in that cost.

Happy to provide the seed-tts-text numbers from the same code once the A/B in #7865 runs, so the second entry does not have to wait for a nightly window.

@y-null y-null closed this Sep 23, 2026
@y-null

y-null commented Sep 23, 2026

Copy link
Copy Markdown
Contributor Author

done

@y-null y-null reopened this Sep 23, 2026
@y-null

y-null commented Sep 23, 2026

Copy link
Copy Markdown
Contributor Author

done

@y-null y-null closed this Sep 23, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

benchmark/profiler/metrics/logger Codes related to benchmarks, profiler, metrics and logger system enhancement New feature or request tts code related to tts models

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants