Repository navigation
[Refactor] Add Ming Flash Omni TTS adapter - #5746
Merged
linyueqian merged 3 commits intoAug 10, 2026
Merged
linyueqian merged 3 commits into
linyueqian merged 3 commits into
Conversation
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
sphinxkkkbc
requested review from
alex-jw-brooks,
tzhouam and
yenuo26
as code owners
August 4, 2026 02:40
|
Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits. |
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Contributor
Author
|
Update ce17357 Fix CI Errors: Remove stale test |
linyueqian
enabled auto-merge (squash)
August 10, 2026 01:32
yuekaizhang
added a commit
to yuekaizhang/vllm-omni
that referenced
this pull request
Aug 10, 2026
The pre-commit ratchet 'check-tts-adapter-migration' fails on every PR: vllm-project#5746 reduced serving_speech.py to 27 _tts_model_type comparisons but left MAX_MODEL_TYPE_BRANCHES at 29, and the hook demands the budget be lowered to hold the new ground. One-line ratchet update as instructed by the hook output; unrelated to this PR's model work but required for a green pipeline (this branch also merges current main to pick the hook up). Verified post-merge on 1x H100: sync path bit-identical (196/196 tokens, same WAV bytes), streaming path 196/196 with 99.98% sample-equal audio; 38 CPU tests pass; the ratchet script exits 0. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Yuekai Zhang <zhangyuekai@foxmail.com>
This was referenced Aug 10, 2026
linyueqian
added a commit
to linyueqian/vllm-omni
that referenced
this pull request
Aug 10, 2026
The ratchet added in vllm-project#5682 errored in both directions: over budget and under it. The under-budget case demanded a hand-edited constant, so removing a branch -- the whole point of the tool -- broke the build. That is what happened when vllm-project#5746 migrated the last legacy detector. It was green on its own base, which predated the ratchet, and turned main red on merge. main stayed red for 5h19m across 5 commits, and an unrelated diffusion PR (vllm-project#5843) had to carry the fix to get its own CI green. Now only an increase fails. A decrease prints a reminder to lower the budget and exits 0. Two smaller fixes found while testing this: - pre-commit hides output from hooks that pass, so the reminder would have been invisible. The hook is now verbose. - the hook's files: pattern did not include the checker itself, so a commit that only edited a budget was never checked by it. Added. Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
This was referenced Aug 11, 2026
Merged
linyueqian
added a commit
to linyueqian/vllm-omni
that referenced
this pull request
Aug 15, 2026
MiniCPM-o's streaming audio cache was permanently dead. The vendored Whisper
layer read the cache back out of the attention return value:
past_key_values = attn_out[2] if len(attn_out) > 2 else None
transformers 4.x returned the cache as a third element; v5 returns only
(attn_output, attn_weights), so attn_out[2] was unreachable and this rebound
the cache to None on every layer -- the encoder re-encoded its whole prefix on
every chunk. v5 Cache objects are mutated in place, so the object passed in is
already current and the read-back is both broken and unnecessary. Measured
against the real transformers 5.14.1 WhisperAttention with two 5-position
chunks: before, cache length None -> None; after, 5 -> 10.
The TTS ratchet test asserted count == budget, re-imposing through pytest what
vllm-project#6008 removed from the checker: with 27 branches and a budget of 27, removing
one branch gives 26 == 27 -> FAIL unless the constant is hand-edited too. That
is what broke main for 5h19m in vllm-project#5746. Budgets are a ceiling.
Adds model_local_kv.py: a post-load declaration protocol. Geometry is not
always static -- ming_flash_omni builds its cache from a checkpoint-side
Qwen2Config -- and no cache is bounded by max_model_len, so only the model
knows its own bound. scope x max_live_instances rather than a single
preallocated/grows flag, because MiMo's 704 KiB cache costs 179 MiB once
replicated per graph bucket. Describes only; allocation stays with
RFC vllm-project#5244 / PR vllm-project#6094.
Part of vllm-project#4855.
Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
khairulkabir1661
pushed a commit
to khairulkabir1661/vllm-omni
that referenced
this pull request
Sep 25, 2026
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
PLEASE FILL IN THE PR DESCRIPTION HERE.
Purpose
Register a serving adapter for
ming_flash_omni_ttsand remove its legacy validation and request-building branches fromserving_speech.py.Preserves the existing request validation, prompt construction, and generation behavior without changing the adapter contract.This PR moves request validation and routing into the adapter while continuing to reuse the existing
_build_ming_flash_omni_prompt()helper. Consistent with the other migrated adapters, relocating the model-specific prompt builder is left to a follow-up PR, but it can be included here if preferred.Removed
test_ming_flash_omni_not_migrated()as it is no longer applicableTest Plan
pytest -v tests/entrypoints/openai_api/test_serving_speech.pypytest -v tests/entrypoints/openai_api/test_tts_adapter.pyCovers basic request validation and prompt-building delegation through the adapter.
vLLM Version:
0.26.0
vLLM-Omni Commit:
9bd1189
Test Result
207 passed
9 passed
cc @linyueqian for review
BEFORE SUBMITTING: read CONTRIBUTING.md and run the precheck-pr skill with the code agent for a self-check against project conventions.
(anything written below this line will be removed by GitHub Actions)