Repository navigation
[Bugfix] Restore MiniCPM native duplex Seed-TTS startup - #3
Draft
NolenLiang wants to merge 1 commit into
Draft
NolenLiang wants to merge 1 commit into
NolenLiang wants to merge 1 commit into
Conversation
Mark nonempty initial text pending when preparing the native duplex session, and cover silent input, forced listening, and post-turn state with regression tests. Signed-off-by: bcsdhjew <nliang@nvidia.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Proposed fix for vllm-project#8138. This draft is hosted in the contributor fork for review; its base is the validated upstream snapshot.
MiniCPM native duplex Seed-TTS can receive its initial text but remain in listen mode for every silent audio unit, so the benchmark never starts its first response. When Stage 0 embeds nonempty initial text, this patch also marks it as pending user content. The existing generation lifecycle consumes that content and restores listening after the response ends.
The change is confined to session initialization and regression tests. It preserves explicit forced-listen prefixes and vllm-project#7974's protection against state changes from discarded asynchronous lookahead. It uses the existing native duplex session; it does not restore the removed turn-based Realtime mode.
AI assistance: Codex assisted with diagnosis, implementation, regression tests, validation orchestration, and this description. Claude performed an independent review.
Test Plan
vLLM Version: 0.30.0, PyTorch 2.13.0+cu130, Transformers 5.14.1.
vLLM-Omni Commit: Issue revision
6fb2b36before/after the patch; candidate also tested on frozen mainb308e19. Historical controls use pristined486bf0and its child7a8d956(vllm-project#7974).GPU: one GB200. Model:
openbmb/MiniCPM-o-4_5, revision503e754. Dataset:zhaochenyang20/seed-tts-eval, revision8f5e1aa.The original benchmark uses
vllm_omni/deploy/minicpmo_4_5.yaml, the first four English targets without shuffling, concurrency 1, two warmups, and one measured request. Only model and dataset paths are mapped to local copies inlocal_config.json; assertions, timeouts, input data, and deployment settings are retained. The E2E run maps the model throughMODEL_PREFIXand retains the original test source. Commands, run from each source checkout:The isolated no-reference text-only comparison runs this original test once per source, with a fresh server and the same model mapping:
MODEL_PREFIX=/model-prefix python3 -m pytest -s -vv \ tests/e2e/online_serving/test_minicpmo_4_5_duplex.py::test_duplex_seeded_text_to_text_needs_no_reference_voice \ -m "H100 and cards_1" --run-level advanced_modelThe hardware expressions select existing test marks; the local GPU is GB200. CPU regressions exercise context preparation, silent input, sampling policy, forced listening, and consumption of seeded content. A separate diagnostic replays actual policy/sampler methods with synthetic embeddings and logits across four source snapshots, with seeded/unseeded input and with/without a discarded lookahead callback.
Test Result
d486bf07a8d9566fb2b366fb2b36+ patchThe adjacent pristine GPU comparison used the same GPU, package versions, model, reference audio, and selected inputs, confirming the Seed-TTS regression at vllm-project#7974. Source inspection and the synthetic-logits replay explain a possible mechanism: before vllm-project#7974, a discarded lookahead token could clear
current_turn_ended, allowing the next silent unit to generate despite the missing pending-input flag. vllm-project#7974 freezes that discarded frame's decisions. No historical GPU token/state trace was captured, so the replay is not direct evidence of the exact token sequence used by the older successful GPU run.The five passing E2E cases cover English/Chinese seeded speech, both long-output cases, and sequential-session independence. The remaining no-reference text-only case was then compared in isolation:
d486bf0response.doneand nonempty text/transcript7a8d956(vllm-project#7974)response.doneb308e19+ patchresponse.doneEach run collected exactly one original test, with zero errors or skips. Environment/input comparability checks matched, including the GPU, vLLM/package versions, model revision/metadata, and test/helper hashes. The test, helper, deployment configuration, assertions, and timeouts were unchanged. This controlled comparison confirms the no-reference text-only regression at vllm-project#7974 under the same runtime and shows that this patch does not repair it. The additional underlying cause has not been isolated.
That unchanged text-only test previously passed in build 15495 (
f5e4f5f, vllm-project#7654) and build 15824 (8bc0c38, vllm-project#7631), then failed in build 16048 (7a8d956). The earlier passes used vLLM 0.29.0 and followed other tests in the full module; the failing build used 0.30.0. Those historical CI observations alone were confounded by runtime and test order; the isolated comparison above reproduces the source boundary with vLLM 0.30.0 fixed for all three runs.This patch's demonstrated fix is Seed-TTS startup. It does not establish a fix for the entire Duplex Test step or vllm-project#8123. Original CI uses H100; local GPU validation uses GB200. Generated audio is model output, and completion checks do not establish transcription accuracy or perceptual quality.
Publication commit:
426159d; its complete source tree matches the validated two-file candidate. Available local lint/type-check tools passed, but their versions differ from the pinned pre-commit environments. The full pinned suite has not been run locally.