Skip to content

[Bugfix] Restore MiniCPM native duplex Seed-TTS startup - #3

Draft
NolenLiang wants to merge 1 commit into
review-base/minicpm-seed-8138-b308e19from
fix/minicpm-native-text-seed-8138-draft
Draft

NolenLiang wants to merge 1 commit into
review-base/minicpm-seed-8138-b308e19from
fix/minicpm-native-text-seed-8138-draft

Conversation

@NolenLiang

@NolenLiang NolenLiang commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

Current upstream PR: vllm-project/vllm-omni#8152. This fork draft retains the earlier validation record.

Purpose

Proposed fix for vllm-project#8138. This draft is hosted in the contributor fork for review; its base is the validated upstream snapshot.

MiniCPM native duplex Seed-TTS can receive its initial text but remain in listen mode for every silent audio unit, so the benchmark never starts its first response. When Stage 0 embeds nonempty initial text, this patch also marks it as pending user content. The existing generation lifecycle consumes that content and restores listening after the response ends.

The change is confined to session initialization and regression tests. It preserves explicit forced-listen prefixes and vllm-project#7974's protection against state changes from discarded asynchronous lookahead. It uses the existing native duplex session; it does not restore the removed turn-based Realtime mode.

AI assistance: Codex assisted with diagnosis, implementation, regression tests, validation orchestration, and this description. Claude performed an independent review.

Test Plan

vLLM Version: 0.30.0, PyTorch 2.13.0+cu130, Transformers 5.14.1.

vLLM-Omni Commit: Issue revision 6fb2b36 before/after the patch; candidate also tested on frozen main b308e19. Historical controls use pristine d486bf0 and its child 7a8d956 (vllm-project#7974).

GPU: one GB200. Model: openbmb/MiniCPM-o-4_5, revision 503e754. Dataset: zhaochenyang20/seed-tts-eval, revision 8f5e1aa.

The original benchmark uses vllm_omni/deploy/minicpmo_4_5.yaml, the first four English targets without shuffling, concurrency 1, two warmups, and one measured request. Only model and dataset paths are mapped to local copies in local_config.json; assertions, timeouts, input data, and deployment settings are retained. The E2E run maps the model through MODEL_PREFIX and retains the original test source. Commands, run from each source checkout:

python -m pytest -s -vv tests/dfx/perf/scripts/run_benchmark.py \
  -m "H100 and B200 and cards_1" --run-level full_model \
  --test-config-file local_config.json
python -m pytest tests/worker/test_native_duplex_hooks.py
python -m pytest -s -vv tests/e2e/online_serving/test_minicpmo_4_5_duplex.py \
  -k "seeded or sequential_sessions_are_independent" \
  -m "H100 and cards_1" --run-level advanced_model

The isolated no-reference text-only comparison runs this original test once per source, with a fresh server and the same model mapping:

MODEL_PREFIX=/model-prefix python3 -m pytest -s -vv \
  tests/e2e/online_serving/test_minicpmo_4_5_duplex.py::test_duplex_seeded_text_to_text_needs_no_reference_voice \
  -m "H100 and cards_1" --run-level advanced_model

The hardware expressions select existing test marks; the local GPU is GB200. CPU regressions exercise context preparation, silent input, sampling policy, forced listening, and consumption of seeded content. A separate diagnostic replays actual policy/sampler methods with synthetic embeddings and logits across four source snapshots, with seeded/unseeded input and with/without a discarded lookahead callback.

Test Result

Check Result
Original benchmark, pristine d486bf0 Passed: 1/1 measured request, four audio targets
Original benchmark, pristine 7a8d956 Failed: no response starts in either warmup or the measured request; 0/1 measured
Original benchmark, issue revision 6fb2b36 Same no-start failure in both warmups and measured request; 0/1 measured
Original benchmark, 6fb2b36 + patch Passed: 1/1 measured request, four audio targets, no error records
Eight new regression cases on issue baseline 3 failed, 5 passed; the failures exercise missing seeded-input state
Patched native duplex hooks test file, including frozen main 91 passed
Four-source CPU replay All 16 controls passed
Existing seeded E2E cases, frozen main + patch 5 passed, 1 failed
Ruff, formatting, and patch checks Passed

The adjacent pristine GPU comparison used the same GPU, package versions, model, reference audio, and selected inputs, confirming the Seed-TTS regression at vllm-project#7974. Source inspection and the synthetic-logits replay explain a possible mechanism: before vllm-project#7974, a discarded lookahead token could clear current_turn_ended, allowing the next silent unit to generate despite the missing pending-input flag. vllm-project#7974 freezes that discarded frame's decisions. No historical GPU token/state trace was captured, so the replay is not direct evidence of the exact token sequence used by the older successful GPU run.

The five passing E2E cases cover English/Chinese seeded speech, both long-output cases, and sequential-session independence. The remaining no-reference text-only case was then compared in isolation:

Source, fixed vLLM 0.30.0 Original isolated test result
Pristine d486bf0 Passed: response.done and nonempty text/transcript
Pristine 7a8d956 (vllm-project#7974) Failed: 12 listen events, missing response.done
b308e19 + patch Failed: 12 listen events, missing response.done

Each run collected exactly one original test, with zero errors or skips. Environment/input comparability checks matched, including the GPU, vLLM/package versions, model revision/metadata, and test/helper hashes. The test, helper, deployment configuration, assertions, and timeouts were unchanged. This controlled comparison confirms the no-reference text-only regression at vllm-project#7974 under the same runtime and shows that this patch does not repair it. The additional underlying cause has not been isolated.

That unchanged text-only test previously passed in build 15495 (f5e4f5f, vllm-project#7654) and build 15824 (8bc0c38, vllm-project#7631), then failed in build 16048 (7a8d956). The earlier passes used vLLM 0.29.0 and followed other tests in the full module; the failing build used 0.30.0. Those historical CI observations alone were confounded by runtime and test order; the isolated comparison above reproduces the source boundary with vLLM 0.30.0 fixed for all three runs.

This patch's demonstrated fix is Seed-TTS startup. It does not establish a fix for the entire Duplex Test step or vllm-project#8123. Original CI uses H100; local GPU validation uses GB200. Generated audio is model output, and completion checks do not establish transcription accuracy or perceptual quality.

Publication commit: 426159d; its complete source tree matches the validated two-file candidate. Available local lint/type-check tools passed, but their versions differ from the pinned pre-commit environments. The full pinned suite has not been run locally.

Mark nonempty initial text pending when preparing the native duplex session, and cover silent input, forced listening, and post-turn state with regression tests.

Signed-off-by: bcsdhjew <nliang@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant