Repository navigation
[Bugfix][MiniCPM-o] Cap offline Talker generation at remaining context - #6458
Merged
Merged
Conversation
amy-why-3459
requested review from
NickCao,
Sy0307,
alex-jw-brooks,
gcanlin,
linyueqian,
lishunyang12,
tzhouam and
yenuo26
as code owners
August 21, 2026 12:50
|
Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits. |
|
This PR appears to be related to model: minicpm. Model owners: @y-null @amy-why-3459, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
amy-why-3459
force-pushed
the
bugfix_ci
branch
from
August 21, 2026 13:26
39b65b9 to
83fc166
Compare
Collaborator
Author
yenuo26
reviewed
Aug 22, 2026
|
|
||
| _REPETITION_PENALTY_CHUNK_SIZE = 16 | ||
| # ``past_window`` of MiniCPMTTS's codec repetition penalty: both generate() and | ||
| # generate_chunk() build it through gen_logits(), which hardcodes |
Collaborator
There was a problem hiding this comment.
Does this file need to be added to the source_file_dependencies of ready & merge?
fan2956
pushed a commit
to fan2956/vllm-omni
that referenced
this pull request
Aug 23, 2026
vllm-project#6458) Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
fan2956
pushed a commit
to fan2956/vllm-omni
that referenced
this pull request
Aug 23, 2026
vllm-project#6458) Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
1 task done
AndyZhou952
pushed a commit
to AndyZhou952/vllm-omni
that referenced
this pull request
Aug 26, 2026
vllm-project#6458) Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com> Signed-off-by: AndyZhou952 <jzhoubc@connect.ust.hk>
JoseCarlosGarcia95
pushed a commit
to valendra-tech/vllm-omni
that referenced
this pull request
Sep 5, 2026
vllm-project#6458) Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
khairulkabir1661
pushed a commit
to khairulkabir1661/vllm-omni
that referenced
this pull request
Sep 25, 2026
vllm-project#6458) Signed-off-by: amy-why-3459 <wuhaiyan17@huawei.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
After #6346 made the Talker a real single-vocab codec LM, three leftover mismatches with upstream
MiniCPMTTS.generate()still produce silent or stuttering tails on long answers. This PR scores the codec repetition penalty the way upstream does, restores forced EOS after vLLM'smin_tokensprocessor blanks it, and caps offline generation at upstream'smax_new_token.Motivation
Fix: #6428
#6346wired Talker logits into vLLM's Sampler. That was necessary, but the Sampler is not a drop-in forMiniCPMTTS.generate()/generate_chunk():Whole-stream presence penalty vs 16-frame frequency penalty. vLLM taxes a code once it has ever been sampled, for the rest of the stream. A codec request is thousands of frames long, so every seen code ends up scaled by the same flat factor while unseen codes stay untouched; the tail then drifts off the speech manifold into near-silence. Upstream's
CustomRepetitionPenaltyLogitsProcessorRepeat(penalty, num_code, 16)(built bygen_logits()) taxespenalty ** freqover the last 16 frames only.Forced EOS vs
MinTokensLogitsProcessor. Codec EOS is a stagestop_token_idsentry. When the model forces EOS (empty speech, or the recorded budget), that processor then masks EOS for the firstmin_tokenssteps, the row becomes all-inf, and the request keeps decoding until the length cap.Offline length. Without
MiniCPMTTS.generate'smax_new_token=2048, a request that never samples codec EOS keeps emitting for the remaining Talker context — often twice as long as upstream — which is audible as a long silent tail.Changes
Talker (
minicpmo_4_5_omni_tts.py)recent_codeswindow of 16 (upstreampast_window) and score_apply_batched_repetition_penaltyinsample()before the Sampler, so the penalty lands ahead of top-k / top-p.repetition_penalty(mirrors upstream's per-requestsampling_params.repetition_penalty) and neutralize the Sampler's own pass (repetition_penalties = 1) so the penalty is scored once.MinTokensLogitsProcessorand overwrite the sampled ids afterwards, so a finished request actually emits codec EOS and releases.min(2048, remaining Talker context). Native duplex is unchanged: onegenerate_chunkis still 25 frames + the terminating sample.Deploy YAML
Stage-1
default_sampling_paramsnow trackutils.TTSSamplingParams:temperature: 0.8 (unchanged)top_p: 0.8 → 0.85top_k: 100 → 25repetition_penalty: 1.02 → 1.05min_tokens: 50 (upstreammin_new_token)max_tokens: 4096 (Sampler ceiling; Talker still clamps to 2048 / remaining context)repetition_penalty: 1.05is required: the field's semantics changed from "whole-stream presence" to "16-frame frequency", so the old 1.02 (chosen to keep the old penalty from being too harsh) is no longer a meaningful number.top_k/top_pare not required for the code change to be correct. They were also #6346 placeholders, not a tuned baseline, and they are aligned here so the default request matches upstream. They do change acoustics and can move seed-tts / accuracy numbers;minicpmo_4_5_duplex.yamldoes not need a matching edit — it only overlaysmin_tokens: 0/max_tokens: 4096and inherits the rest.min_p/win_size/tau_rstay out: they are declared onTTSSamplingParamsbut never read bygen_logits().Warper order is still not identical (upstream is top-p then top-k with
min_tokens_to_keep=3; vLLM is top-k then top-p). On 0.85 / 25 that difference is small; this PR only aligns the numbers.Tests
no_penalties, and is consumed once so the next step cannot rescore a stale row.min_tokensblanked is restored on the sampled ids; unforced rows are left to the Sampler.max_tokens = 2048when context is not the binding constraint, and still clamps to remaining context when it is.top_k=25,top_p=0.85,repetition_penalty=1.05on every MiniCPM-o 4.5 deploy YAML (duplex overlay is allowed to dropmin_tokensto 0).Tested
tests/model_executor/models/minicpmo_4_5/test_talker_batching.pytests/config/test_config_factory.py/v1/chat/completionslong-form TTS (listen for silent / stuttering tail)