Skip to content

[Migrate] Move sampling parameter overrides from legacy dispatches to TTS adapters - #5272

Merged
linyueqian merged 24 commits into
vllm-project:mainfrom
sphinxkkkbc:feat/serving_speech_clean_up
Aug 17, 2026
Merged

linyueqian merged 24 commits into
vllm-project:mainfrom
sphinxkkkbc:feat/serving_speech_clean_up

Conversation

@sphinxkkkbc

@sphinxkkkbc sphinxkkkbc commented Jul 21, 2026 •

Copy link
Copy Markdown
Contributor

PLEASE FILL IN THE PR DESCRIPTION HERE.

Purpose

M3 of #4855. This PR moves model-specific sampling_params_list mutations from serving_speech into the corresponding TTS adapters throughapply_sampling_overrides. Removed model-specific dispatch logic in serving_speech while preserving the existing behavior for legacy dispatch paths.

Compatibility Note

The legacy glm_tts and cosyvoice3 paths previously applied their sampling_params_list overrides before merging request.extra_params. With the unified adapter flow, the order is now: stream coercion -> extra_params -> apply_sampling_overrides -> seed, matches the contract documented in the adapter base class.

These adapters do not read or modify sampling_params_list[0].extra_args, so changing the order of these two independent operations does not change their behavior.

For cosyvoice3, the legacy flow applied the dynamic token limits in two steps:

sampling_params_list[0].min_tokens = max(1, int(text_token_len * min_ratio))
sampling_params_list[0].max_tokens = min(2048, int(text_token_len * max_ratio))

When request.max_new_tokens was provided, it then applied the request-level cap:

sampling_params_list[0].min_tokens = min(
    getattr(sampling_params_list[0], "min_tokens", 0),
    request.max_new_tokens,
)
sampling_params_list[0].max_tokens = request.max_new_tokens

The migrated CosyVoice3Adapter.apply_sampling_overrides() consolidates these operations and preserves the previous behavior.

Test Plan

pytest -v tests/entrypoints/openai_api/test_serving_speech.py
pytest -v tests/worker/test_gpu_ar_model_runner.py

vLLM Version: 0.26.0

vLLM-Omni Commit: 4fc8892

Test Result

204 passed, 17 warnings in 1.44s
23 passed, 16 warnings in 1.14s

BEFORE SUBMITTING: read CONTRIBUTING.md and run the precheck-pr skill with the code agent for a self-check against project conventions.

(anything written below this line will be removed by GitHub Actions)

Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@sphinxkkkbc

Copy link
Copy Markdown
Contributor Author

@linyueqian please feel free to take a look, thanks

@hsliuustc0106 hsliuustc0106 added tts code related to tts models refactor refactoring for better code scalability and quality labels Jul 22, 2026
TEXT_EOS_TOKEN_ID,
)

hf_config = self.engine_client.model_config.hf_config

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Any Ming-TTS request with the normal non-empty sampling list will crash here: TTSModelAdapter.__init__ only sets self.ctx, so this raises AttributeError when _prepare_speech_generation() invokes the override. Please use self.ctx.server.engine_client (or self.ctx.engine_client) and add a Ming override regression test; the current 199-test suite does not exercise this path.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

fixed.

before ddc2296:
1 failed, 199 passed, 17 warnings in 1.34s

after:
200 passed, 17 warnings in 1.44s

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verified at head — self.ctx.server is in place and the regression test now asserts stop_token_ids == [TEXT_EOS_TOKEN_ID] and max_tokens == 8 for max_new_tokens=7. Thanks.

Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>

@linyueqian linyueqian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the full diff against the pre-PR dispatch code. The migration looks faithful: the adapter registry names cover exactly the deleted _SAMPLING_MAX_TOKENS_TTS_MODEL_TYPES set plus the ming_tts/glm_tts branches, models that never had overrides keep the identity default, and since the extra_params merge only writes extra_args, moving the cosyvoice3/GLM blocks after it is behavior-neutral as the description says. Nice to see 144 lines of dispatch leave serving_speech.py.

A few non-blocking comments, mostly cleanup:

  • Three adapters still carry "stays in the orchestrator tail" NOTEs that this PR makes false; see the inline comment on cosyvoice3.py.
  • Two more stale spots in serving_speech.py are outside the diff so I could not anchor comments there: the comment at line 452 still references the deleted _apply_cosyvoice3_dynamic_tokens, and the "Sampling overrides ... remain in the orchestrator tail" wording in the adapter-resolution comment (~line 505) and the pre-build comment (~line 2999) no longer holds after this PR.
  • Seven adapters now carry an identical max_new_tokens-only override (fish, qwen3, voxtral, higgs v2/v3, indextts2, voxcpm2). A small shared helper on the base class (keeping the base default as identity so unmigrated adapters are unaffected) would remove the copy-paste. Fine to defer if you prefer strict move-only migration PRs.
  • Tiny style nit: in cosyvoice3.py::apply_sampling_overrides the two explanatory comment lines sit above the docstring; folding them into the docstring reads better.

LGTM once the stale comments are tidied up.

@@ -50,3 +55,60 @@ async def build(
# NOTE: CosyVoice3 dynamic-token sampling stays in the orchestrator tail

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This NOTE is now stale: the dynamic-token sampling moves into this adapter's apply_sampling_overrides in this very PR. Same for the equivalent NOTEs at glm_tts.py:50 and ming_tts.py:62. Worth deleting all three here so the next reader is not sent looking for logic in the orchestrator tail that no longer exists.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

removed

)
asyncio.run(ming_tts_server._prepare_speech_generation(request))

assert ming_tts_server._tts_model_type == "ming_tts"

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggestion: also assert the override outcome so a logic regression fails, not just a crash. The real apply_sampling_overrides runs in this test (only build is mocked), so you can grab ming_tts_server.engine_client.generate.call_args.kwargs["sampling_params_list"] and check stop_token_ids == [TEXT_EOS_TOKEN_ID], plus a variant with max_new_tokens=N asserting max_tokens == N + 1. Now that the cosyvoice3/GLM logic lives in adapters it is unit-testable in isolation for the first time, so direct tests for those two would also be a natural follow-up.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

added. And yes, unit tests for the cosyvoice3 and glm-tts adapters should be added as a follow up.

Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
@sphinxkkkbc

Copy link
Copy Markdown
Contributor Author

Updated to the latest main after Audex merge and resolved the resulting conflicts, reran the tests locally with 204 passed. @linyueqian, could you please take a look?

Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
@sphinxkkkbc

Copy link
Copy Markdown
Contributor Author

Update:

While reviewing serving_speech.py, found Audex sampling-parameter helpers that had not been fully migrated in the previous commit. To complete the migration, the existing apply_sampling_overrides interface now accepts a request_id parameter.

To keep this PR focused, the update only covers the sampling-parameter logic already within its scope and does not change any capability-related behavior or state.

Tests:

  • pytest -v tests/entrypoints/openai_api/test_audex_serving_guards.py — 49 passed
  • pytest -v tests/entrypoints/openai_api/test_serving_speech.py — 204 passed

}
)
cosyvoice3_server._apply_cosyvoice3_dynamic_tokens = mocker.MagicMock(side_effect=lambda spl, req: spl)
cosyvoice3_server._adapter.apply_sampling_overrides = mocker.MagicMock(

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Can we keep the real CosyVoice3Adapter.apply_sampling_overrides here and assert the resulting min_tokens/max_tokens? Replacing it with an identity mock means this test only proves dispatch reaches a stub, so a regression in the newly migrated dynamic-token logic (including max_new_tokens capping) would still pass. A focused adapter test or an assertion on generate(...).sampling_params_list would close this.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added the real-path sampling_params tests for GLM-TTS and CosyVoice3.This was previously suggested as a follow-up but didn't land.

pytest -v tests/entrypoints/openai_api/test_serving_speech.py — 205 passed

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Confirmed — the CosyVoice3/GLM tests now run the real apply_sampling_overrides (only prompt building and tokenizer internals are stubbed) and assert the resulting min_tokens/max_tokens. Closes this for me.

Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
@linyueqian linyueqian added the ready label to trigger buildkite CI label Aug 9, 2026
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
@sphinxkkkbc
sphinxkkkbc requested a review from NickCao as a code owner August 9, 2026 02:53
@linyueqian linyueqian added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Aug 10, 2026
…pter models, remove related check/test/pre-commit hooks

Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
@sphinxkkkbc

Copy link
Copy Markdown
Contributor Author

@linyueqian, Update e0ee6a5:

Removed LEGACY_TTS_DETECTORS, LegacyDetector, and the related tests and pre-commit checks. TTS detection now relies exclusively on registered adapters, so this path no longer supports models without an adapter. If transitional compatibility is still desired, happy to restore the legacy mechanism and only update the affected tests.

Test:

pytest -v tests/entrypoints/openai_api/test_serving_speech.py \
       tests/worker/test_gpu_ar_model_runner.py \
       tests/entrypoints/openai_api/test_tts_detection.py

353 passed

Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
@linyueqian linyueqian added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Aug 10, 2026
@linyueqian

Copy link
Copy Markdown
Collaborator

Sorry — the conflict here is mine. #6008 landed while this was open and we both edited tools/pre_commit/check_tts_adapter.py. It is the only conflicting file.

You deleted _legacy_detectors (correct — Ming is migrated, so it has no callers left), and I changed _check's signature in the same region.

Resolution: keep both. Take main's _check (it now takes a notices list and no longer errors when the count drops), and keep your deletion of _legacy_detectors.

One thing that may save you work: your tight max_model_type_branches to 20 commits are no longer needed for CI to pass. #6008 changed the ratchet so a count below budget prints a reminder and exits 0, instead of failing. Tightening is still welcome, just not required to go green — and it can no longer break anyone else's build when it lands.

Happy to push the resolution myself if you'd rather not deal with someone else's mess. CI on your current head was clean (build 13412, 22 jobs, 0 failures), so this is the only thing standing between it and a review.

@sphinxkkkbc

Copy link
Copy Markdown
Contributor Author

Conflicts resolved, @linyueqian. We’d better lower the threshold whenever the branch count drops during refactoring.

@linyueqian

Copy link
Copy Markdown
Collaborator

Agreed, and thanks for resolving it.

Lowering it as you go is still the right habit — it is just no longer something CI forces on you mid-refactor. #6008 changed the ratchet so a count below budget prints a reminder and exits 0; only an increase fails. That was the bug behind the two outages: the old check treated a dropped count as an error, so removing branches broke main until someone hand-edited the constant.

So: tighten it in the PR that does the removing when it is convenient, and if you forget, the next person sees the reminder instead of a red build.

@sphinxkkkbc

Copy link
Copy Markdown
Contributor Author

@linyueqian Conflicts resolved after IndexTTS 2.5 merge.

pytest -v tests/entrypoints/openai_api/test_serving_speech.py tests/worker/test_gpu_ar_model_runner.py tests/entrypoints/openai_api/test_tts_detection.py

All 359 tests passed. Can this PR move forward?

@linyueqian linyueqian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-reviewed the full diff at 245d82ae against the pre-PR orchestrator tail. The migration is faithful. Three nits below, no blockers.

What I checked, since the diff grew a lot since my July pass (Audex helpers migrated, LegacyDetector removed, budget lowered to 20):

  • Registry coverage is exact. The deleted _SAMPLING_MAX_TOKENS_TTS_MODEL_TYPES had 11 names; all 11 now resolve to adapters calling apply_max_new_tokens, plus the two special cases (ming_tts, glm_tts). Nothing gained or lost an override.
  • No inheritance hazard. IndexTTS25Adapter is the only adapter that subclasses another registered adapter. MingFlashOmniTTSAdapter, MossTTS*, StepAudio2Adapter, CovoAudioAdapter and OmniVoiceAdapter all extend ARTTSAdapter directly, so none of them silently inherits Ming's stop_token_ids or a max_tokens override it did not have before.
  • The reordering is neutral. The extra_params merge writes only temperature/top_p/top_k and extra_args; the cosyvoice3 and GLM blocks touch only min_tokens/max_tokens/seed. Disjoint, so moving them after it cannot interact.
  • The Audex in-place sampling_params_list[0] = ... is safe. serving_speech.py:2982 builds a fresh list(...) and coerce_param_message_types mutates that copy in place, so the engine's shared default_sampling_params_list is never reindexed. Elements are still shared, which is why the injectors' deepcopy is load-bearing; that is preserved.
  • The CosyVoice3 consolidation is algebraically identical. Old: min = min(max(1, L*min_ratio), max_new_tokens), max = max_new_tokens. New: the same, because the min_tokens read inside the new max_new_tokens branch is the value just assigned from the ratio.
  • The ratchet is honest. python3.12 tools/pre_commit/check_tts_adapter.py exits 0 at this head and the branch count is exactly 20, matching MAX_MODEL_TYPE_BRANCHES, so test_budgets_are_the_real_counts holds.

One more stale-doc item that is not anchorable to a diff hunk: tts_adapters/base.py:10-12 still says "the remaining models stay on the legacy path until individually migrated", which this PR is the one to make false.

Validation gap worth clearing before merge: there is no Buildkite build on 245d82ae. The only checks on this SHA are build (3.11), build (3.12), pre-commit, DCO and readthedocs, none of which run the five test files this PR changes. The ready label predates the latest push, so the label event that normally kicks Buildkite has not fired for this head. The 359-passed number is a local run only. Re-triggering CI would make the equivalence argument above test-backed rather than review-backed.

params["duration_factor"] = [1.0 / speed]
return params

def apply_sampling_overrides(

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P3] This is byte-identical to the method added at IndexTTS2Adapter line 244, and IndexTTS25Adapter(IndexTTS2Adapter) already inherits it. Safe to delete the subclass copy.

sampling_params_list[0].min_tokens = max(1, int(text_token_len * min_ratio))
sampling_params_list[0].max_tokens = min(2048, int(text_token_len * max_ratio))

if request.max_new_tokens is not None:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] This branch is the only genuinely consolidated logic in the PR (everything else is a verbatim move), and it is the only one with no test.

test_prepare_speech_generation_cosyvoice3 exercises the ratio path only (min_tokens == 10, max_tokens == 2000, no max_new_tokens), and the fish_speech tests cover the shared apply_max_new_tokens helper, but nothing exercises CosyVoice3's own min_tokens = min(min_tokens, max_new_tokens) clamp.

A max_new_tokens=5 variant of the existing test closes it: with min_token_text_ratio=1 and text_token_len=10 the uncapped min_tokens is 10, so the assertion becomes min_tokens == 5 and max_tokens == 5.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1, the max_new_tokens=5 variant asserting min_tokens == 5 / max_tokens == 5 closes this for me.

if not path.is_file():
print(f"check_tts_adapter: expected file not found: {path}", file=sys.stderr)
return 1
if not serving_path.is_file():

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P3] Two leftovers from dropping the second counter:

  • line 136 above still says "we always audit the same two files", now one.
  • .pre-commit-config.yaml:69 still lists vllm_omni/entrypoints/openai/tts_adapters/__init__.py in the hook's files: pattern, so editing that file triggers a gate that no longer reads it.

Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
@sphinxkkkbc

Copy link
Copy Markdown
Contributor Author

@linyueqian PTAL, comments resolved, thanks!

@linyueqian linyueqian added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Aug 16, 2026
@linyueqian
linyueqian enabled auto-merge (squash) August 16, 2026 03:25

@linyueqian linyueqian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 3 on 740a5f61. All three items from my last pass are resolved, and I verified each rather than taking the diff at face value:

  • The files: regex dropping tts_adapters/__init__.py is right. tools/pre_commit/check_tts_adapter.py:28 only ever reads SERVING_SPEECH, so __init__.py was never part of the audit surface.
  • Deleting IndexTTS25Adapter.apply_sampling_overrides is behavior-preserving. The inherited IndexTTS2Adapter.apply_sampling_overrides (indextts2.py:244-251) is identical, return apply_max_new_tokens(sampling_params_list, request).
  • The parametrized test_prepare_speech_generation_cosyvoice3 now covers both branches (None -> (10, 2000), 5 -> (5, 5)), which was the gap.

No new findings. The migration still reads faithful to the legacy dispatch behavior.

One blocker left, and it is not the code: Buildkite has never run on this head. The status rollup carries GitHub Actions only (build 3.11/3.12, DCO, pre-commit, readthedocs), none of which execute tests. The ready label is on, but its last label event was 2026-08-10T20:14Z while the head commit landed 2026-08-12T17:06Z, and a fork PR needs a label event after the push to trigger a build. So the changed test in tests/entrypoints/openai_api/test_serving_speech.py has not actually executed anywhere.

Markers on that file are correct (core_model, cpu), so it will run once a build fires. I have toggled ready off and back on to kick one off. Happy to approve once it is green.

@linyueqian linyueqian added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Aug 17, 2026
@sphinxkkkbc

Copy link
Copy Markdown
Contributor Author

The latest buildkite/vllm-omni run at [2026-08-16T03:41:38Z] is green. amd-ci and intel-ci failures appear unrelated to this PR. @linyueqian Thanks!

@linyueqian linyueqian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm

@linyueqian
linyueqian merged commit 65b3d41 into vllm-project:main Aug 17, 2026
7 of 9 checks passed
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
… TTS adapters (vllm-project#5272)

Signed-off-by: boatman <109857087+sphinxkkkbc@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ready label to trigger buildkite CI refactor refactoring for better code scalability and quality tts code related to tts models

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants