Skip to content

[Refactor] Add Ming Flash Omni TTS adapter - #5746

Merged
linyueqian merged 3 commits into
vllm-project:mainfrom
sphinxkkkbc:refactor/migrate_ming_flash_omni_tts
Aug 10, 2026
Merged

linyueqian merged 3 commits into
vllm-project:mainfrom
sphinxkkkbc:refactor/migrate_ming_flash_omni_tts

Conversation

@sphinxkkkbc

@sphinxkkkbc sphinxkkkbc commented Aug 4, 2026 •

Copy link
Copy Markdown
Contributor

PLEASE FILL IN THE PR DESCRIPTION HERE.

Purpose

Register a serving adapter for ming_flash_omni_tts and remove its legacy validation and request-building branches from serving_speech.py.Preserves the existing request validation, prompt construction, and generation behavior without changing the adapter contract.

This PR moves request validation and routing into the adapter while continuing to reuse the existing _build_ming_flash_omni_prompt() helper. Consistent with the other migrated adapters, relocating the model-specific prompt builder is left to a follow-up PR, but it can be included here if preferred.

Removed test_ming_flash_omni_not_migrated() as it is no longer applicable

Test Plan

pytest -v tests/entrypoints/openai_api/test_serving_speech.py
pytest -v tests/entrypoints/openai_api/test_tts_adapter.py
Covers basic request validation and prompt-building delegation through the adapter.

vLLM Version:
0.26.0
vLLM-Omni Commit:
9bd1189

Test Result

207 passed
9 passed

cc @linyueqian for review

BEFORE SUBMITTING: read CONTRIBUTING.md and run the precheck-pr skill with the code agent for a self-check against project conventions.
(anything written below this line will be removed by GitHub Actions)

Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@sphinxkkkbc sphinxkkkbc changed the title [Refactor]Add ming_flash_omni_tts Adapter [Refactor] Add Ming Flash Omni TTS adapter Aug 4, 2026
@hsliuustc0106 hsliuustc0106 added refactor refactoring for better code scalability and quality tts code related to tts models labels Aug 5, 2026
@linyueqian linyueqian added the ready label to trigger buildkite CI label Aug 9, 2026
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
@sphinxkkkbc
sphinxkkkbc requested a review from NickCao as a code owner August 9, 2026 02:44
@sphinxkkkbc

Copy link
Copy Markdown
Contributor Author

Update ce17357 Fix CI Errors: Remove stale test test_ming_flash_omni_not_migrated() from test_tts_adapter as it is no longer applicable.

@linyueqian linyueqian added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Aug 10, 2026

@linyueqian linyueqian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm

@linyueqian
linyueqian enabled auto-merge (squash) August 10, 2026 01:32
@linyueqian
linyueqian merged commit 11f633a into vllm-project:main Aug 10, 2026
6 of 9 checks passed
yuekaizhang added a commit to yuekaizhang/vllm-omni that referenced this pull request Aug 10, 2026
The pre-commit ratchet 'check-tts-adapter-migration' fails on every PR:
vllm-project#5746 reduced serving_speech.py to 27 _tts_model_type comparisons but left
MAX_MODEL_TYPE_BRANCHES at 29, and the hook demands the budget be lowered
to hold the new ground. One-line ratchet update as instructed by the hook
output; unrelated to this PR's model work but required for a green
pipeline (this branch also merges current main to pick the hook up).

Verified post-merge on 1x H100: sync path bit-identical (196/196 tokens,
same WAV bytes), streaming path 196/196 with 99.98% sample-equal audio;
38 CPU tests pass; the ratchet script exits 0.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Yuekai Zhang <zhangyuekai@foxmail.com>
linyueqian added a commit to linyueqian/vllm-omni that referenced this pull request Aug 10, 2026
The ratchet added in vllm-project#5682 errored in both directions: over budget and
under it. The under-budget case demanded a hand-edited constant, so
removing a branch -- the whole point of the tool -- broke the build.

That is what happened when vllm-project#5746 migrated the last legacy detector. It
was green on its own base, which predated the ratchet, and turned main
red on merge. main stayed red for 5h19m across 5 commits, and an
unrelated diffusion PR (vllm-project#5843) had to carry the fix to get its own CI
green.

Now only an increase fails. A decrease prints a reminder to lower the
budget and exits 0.

Two smaller fixes found while testing this:

- pre-commit hides output from hooks that pass, so the reminder would
  have been invisible. The hook is now verbose.
- the hook's files: pattern did not include the checker itself, so a
  commit that only edited a budget was never checked by it. Added.

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
linyueqian added a commit to linyueqian/vllm-omni that referenced this pull request Aug 15, 2026
MiniCPM-o's streaming audio cache was permanently dead. The vendored Whisper
layer read the cache back out of the attention return value:

    past_key_values = attn_out[2] if len(attn_out) > 2 else None

transformers 4.x returned the cache as a third element; v5 returns only
(attn_output, attn_weights), so attn_out[2] was unreachable and this rebound
the cache to None on every layer -- the encoder re-encoded its whole prefix on
every chunk. v5 Cache objects are mutated in place, so the object passed in is
already current and the read-back is both broken and unnecessary. Measured
against the real transformers 5.14.1 WhisperAttention with two 5-position
chunks: before, cache length None -> None; after, 5 -> 10.

The TTS ratchet test asserted count == budget, re-imposing through pytest what
vllm-project#6008 removed from the checker: with 27 branches and a budget of 27, removing
one branch gives 26 == 27 -> FAIL unless the constant is hand-edited too. That
is what broke main for 5h19m in vllm-project#5746. Budgets are a ceiling.

Adds model_local_kv.py: a post-load declaration protocol. Geometry is not
always static -- ming_flash_omni builds its cache from a checkpoint-side
Qwen2Config -- and no cache is bounded by max_model_len, so only the model
knows its own bound. scope x max_live_instances rather than a single
preallocated/grows flag, because MiMo's 704 KiB cache costs 179 MiB once
replicated per graph bucket. Describes only; allocation stays with
RFC vllm-project#5244 / PR vllm-project#6094.

Part of vllm-project#4855.

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ready label to trigger buildkite CI refactor refactoring for better code scalability and quality tts code related to tts models

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants