Skip to content

[Bugfix][MiniCPM-o] Cap offline Talker generation at remaining context - #6458

Merged
amy-why-3459 merged 1 commit into
vllm-project:mainfrom
amy-why-3459:bugfix_ci
Aug 22, 2026
Merged

amy-why-3459 merged 1 commit into
vllm-project:mainfrom
amy-why-3459:bugfix_ci

Conversation

@amy-why-3459

@amy-why-3459 amy-why-3459 commented Aug 21, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

After #6346 made the Talker a real single-vocab codec LM, three leftover mismatches with upstream MiniCPMTTS.generate() still produce silent or stuttering tails on long answers. This PR scores the codec repetition penalty the way upstream does, restores forced EOS after vLLM's min_tokens processor blanks it, and caps offline generation at upstream's max_new_token.

Motivation

Fix: #6428
#6346 wired Talker logits into vLLM's Sampler. That was necessary, but the Sampler is not a drop-in for MiniCPMTTS.generate() / generate_chunk():

  1. Whole-stream presence penalty vs 16-frame frequency penalty. vLLM taxes a code once it has ever been sampled, for the rest of the stream. A codec request is thousands of frames long, so every seen code ends up scaled by the same flat factor while unseen codes stay untouched; the tail then drifts off the speech manifold into near-silence. Upstream's CustomRepetitionPenaltyLogitsProcessorRepeat(penalty, num_code, 16) (built by gen_logits()) taxes penalty ** freq over the last 16 frames only.

  2. Forced EOS vs MinTokensLogitsProcessor. Codec EOS is a stage stop_token_ids entry. When the model forces EOS (empty speech, or the recorded budget), that processor then masks EOS for the first min_tokens steps, the row becomes all -inf, and the request keeps decoding until the length cap.

  3. Offline length. Without MiniCPMTTS.generate's max_new_token=2048, a request that never samples codec EOS keeps emitting for the remaining Talker context — often twice as long as upstream — which is audible as a long silent tail.

Changes

Talker (minicpmo_4_5_omni_tts.py)

  • Keep a per-request recent_codes window of 16 (upstream past_window) and score _apply_batched_repetition_penalty in sample() before the Sampler, so the penalty lands ahead of top-k / top-p.
  • Accept a per-row repetition_penalty (mirrors upstream's per-request sampling_params.repetition_penalty) and neutralize the Sampler's own pass (repetition_penalties = 1) so the penalty is scored once.
  • Stash forced-EOS rows across MinTokensLogitsProcessor and overwrite the sampled ids afterwards, so a finished request actually emits codec EOS and releases.
  • Offline budget is min(2048, remaining Talker context). Native duplex is unchanged: one generate_chunk is still 25 frames + the terminating sample.

Deploy YAML

Stage-1 default_sampling_params now track utils.TTSSamplingParams:

  • temperature: 0.8 (unchanged)
  • top_p: 0.8 → 0.85
  • top_k: 100 → 25
  • repetition_penalty: 1.02 → 1.05
  • min_tokens: 50 (upstream min_new_token)
  • max_tokens: 4096 (Sampler ceiling; Talker still clamps to 2048 / remaining context)

repetition_penalty: 1.05 is required: the field's semantics changed from "whole-stream presence" to "16-frame frequency", so the old 1.02 (chosen to keep the old penalty from being too harsh) is no longer a meaningful number.

top_k / top_p are not required for the code change to be correct. They were also #6346 placeholders, not a tuned baseline, and they are aligned here so the default request matches upstream. They do change acoustics and can move seed-tts / accuracy numbers; minicpmo_4_5_duplex.yaml does not need a matching edit — it only overlays min_tokens: 0 / max_tokens: 4096 and inherits the rest.

min_p / win_size / tau_r stay out: they are declared on TTSSamplingParams but never read by gen_logits().

Warper order is still not identical (upstream is top-p then top-k with min_tokens_to_keep=3; vLLM is top-k then top-p). On 0.85 / 25 that difference is small; this PR only aligns the numbers.

Tests

  • Windowed penalty matches the request-local reference, forgets codes older than 16, honours no_penalties, and is consumed once so the next step cannot rescore a stale row.
  • Forced EOS that min_tokens blanked is restored on the sampled ids; unforced rows are left to the Sampler.
  • Offline prefill records max_tokens = 2048 when context is not the binding constraint, and still clamps to remaining context when it is.
  • Config factory asserts top_k=25, top_p=0.85, repetition_penalty=1.05 on every MiniCPM-o 4.5 deploy YAML (duplex overlay is allowed to drop min_tokens to 0).

Tested

  • tests/model_executor/models/minicpmo_4_5/test_talker_batching.py
  • tests/config/test_config_factory.py
  • offline /v1/chat/completions long-form TTS (listen for silent / stuttering tail)
  • native duplex still chunks at 26 and does not hold the Talker request open
<frozen importlib._bootstrap>:488
  <frozen importlib._bootstrap>:488: DeprecationWarning: builtin type SwigPyObject has no __module__ attribute

../.venv/lib/python3.12/site-packages/torch/jit/_script.py:365: 14 warnings
  /data/why/.venv/lib/python3.12/site-packages/torch/jit/_script.py:365: DeprecationWarning: `torch.jit.script_method` is deprecated. Please switch to `torch.compile` or `torch.export`.
    warnings.warn(

-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
--- Running Summary
================= 10 passed, 16 warnings in 4456.09s (1:14:16) =================
sys:1: DeprecationWarning: builtin type swigvarlink has no __module__ attribute

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to be related to model: minicpm.

Model owners: @y-null

@amy-why-3459, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
@amy-why-3459

Copy link
Copy Markdown
Collaborator Author

@Gaohan123 @natureofnature @y-null PTAL

@amy-why-3459 amy-why-3459 added ready label to trigger buildkite CI omni-test npu-test labels Aug 21, 2026

_REPETITION_PENALTY_CHUNK_SIZE = 16
# ``past_window`` of MiniCPMTTS's codec repetition penalty: both generate() and
# generate_chunk() build it through gen_logits(), which hardcodes

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does this file need to be added to the source_file_dependencies of ready & merge?

@yenuo26 yenuo26 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@amy-why-3459
amy-why-3459 merged commit 42059f1 into vllm-project:main Aug 22, 2026
8 of 9 checks passed
fan2956 pushed a commit to fan2956/vllm-omni that referenced this pull request Aug 23, 2026
fan2956 pushed a commit to fan2956/vllm-omni that referenced this pull request Aug 23, 2026
AndyZhou952 pushed a commit to AndyZhou952/vllm-omni that referenced this pull request Aug 26, 2026
vllm-project#6458)

Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
Signed-off-by: AndyZhou952 <jzhoubc@connect.ust.hk>
JoseCarlosGarcia95 pushed a commit to valendra-tech/vllm-omni that referenced this pull request Sep 5, 2026
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

npu-test ready label to trigger buildkite CI

Projects

None yet

3 participants