Skip to content

[Core][Model] Fit PersonaPlex into the Unified Full-duplex Framework (RFC #7181 PR 2) - #7695

Merged
linyueqian merged 4 commits into
vllm-project:mainfrom
chickeyton:fit_peronaplex
Sep 27, 2026
Merged

linyueqian merged 4 commits into
vllm-project:mainfrom
chickeyton:fit_peronaplex

Conversation

@chickeyton

@chickeyton chickeyton commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

RFC #7181, PR 2. Ports PersonaPlex onto the DuplexModelPlugin contract so a PersonaPlex deployment is served by DuplexOmni over /v1/realtime?duplex=1 again, re-enabling the PersonaPlex duplex suites that #7413 had to skip (#7636 Issue 16). The other #7636 follow-up issues are handled in a separate PR.

Purpose

Framework extensions (model-neutral; MiniCPM-o behaviour unchanged)

  • DuplexModelPlugin.silence_unit_payload(): the runner's turn-continuation unit and the startup warmup use the plugin's own format/rate/length instead of a fixed 16 kHz payload (PersonaPlex units are 1920 samples @ 24 kHz).
  • SessionEmitter.auto_responds() is implied when a model takes no client commits (supports_client_commit=False), so a stock Realtime client works on a lockstep model without extra_body.auto_response.
  • DefaultDuplexModelSessionState, DuplexDataPlaneContext and a default validate_client_extra_body remove the per-model copies of shared session code.

Shared toolbox at vllm_omni/model_executor/common/: request_outputs.py (RequestOutput / multimodal_output readers, cumulative-delta helpers), audio/pcm.py (pcm_f32le decode/count/materialise) and duplex/ (append-payload validation, FixedFramePcmAppendBuffer, CumulativeAudioTextDataPlane). MiniCPM-o adopts the shared helpers; PersonaPlex subclasses the duplex building blocks.

PersonaPlex port: PersonaPlexDuplexPlugin replaces the pre-framework runtime extension + serving adapter pair; the Stage 0 lockstep runtime is keyed by (session_id, epoch) so response.cancel / output_audio_buffer.clear restart the model context on a fresh Stage 0 request (the old epoch's Mimi encoder is recycled, the voice/persona prefill replayed); the pipeline declares duplex_plugin. PersonaPlex advertises no client commits, no barge-in, no resume and no chat route, so /v1/chat/completions is not mounted for it.

Tests: PersonaPlex duplex suites rewritten for the plugin (test_plugin.py, rekeyed test_stage0_runtime.py), toolbox suites under tests/model_executor/common/, a runner scenario driven by the real plugin, a warmup test, plugin-default tests, and a GPU pytest wrapper; the E2E driver uses server-allocated session ids.

Docs: full_duplex_api.md, realtime_duplex_api.md, fullduplex.md, a rewritten fullduplex-personaplex.md, the PersonaPlex example README, recipe and supported-models row.

Test Plan

CPU (no GPU needed):

pytest tests/engine/duplex tests/entrypoints/duplex tests/entrypoints/openai_api/test_duplex_api_server.py        tests/entrypoints/test_duplex_omni.py tests/engine/test_duplex_omni_engine.py tests/engine/test_duplex_orchestrator.py        tests/engine/test_duplex_import_boundary.py tests/clients tests/model_executor/common        tests/model_executor/models/personaplex tests/model_executor/models/minicpmo_4_5/duplex        tests/model_executor/stage_input_processors/test_personaplex.py -n 0

GPU (one card, gated nvidia/personaplex-7b-v1 checkout; default vllm_omni/deploy/personaplex.yaml):

vllm serve /path/to/personaplex-7b-v1 --omni --deploy-config vllm_omni/deploy/personaplex.yaml --port 8099
python tests/e2e/online_serving/personaplex_realtime_duplex.py --model /path/to/personaplex-7b-v1   --input-wav tests/assets/minicpmo_4_5/response_required_16k.wav --output-dir /tmp/pplex-e2e
PERSONAPLEX_MODEL_PATH=/path/to/personaplex-7b-v1 pytest tests/e2e/online_serving/test_personaplex_duplex.py --run-level advanced_model

vLLM Version: 0.30.0

vLLM-Omni Commit: c0e6088 (this branch, rebased on main 5634f03e)

Test Result

L20X (1 GPU), Python 3.12, vLLM 0.30.0, torch 2.13.0+cu130, VLLM_USE_FLASHINFER_SAMPLER=0 (the fresh venv had no FlashInfer JIT cache); re-run in full after the rebase onto main 5634f03e:

Run Result
CPU duplex suite above (-n 0) 878 passed, 4 skipped
E2E driver on a live server: two paced sessions (94 / 188 frames) + replacement exit 0; all three sessions audible, whole 80 ms frames, frame coverage 0.97-0.99; third concurrent open refused with resource_exhausted; ids server-allocated and distinct
test_personaplex_duplex.py --run-level advanced_model (own server via the omni_server fixture) 1 passed
response.cancel mid-stream new response under epoch 1, audio deltas every 0.4 s for the rest of the input, no errors
DuplexClient + vllm_omni.clients.personaplex.create_duplex_session_config on the live server correct session payload and capabilities, 5.28 s of PCM16 audio back, no errors
pre-commit run --from-ref upstream/main --to-ref HEAD all hooks pass (incl. mypy-3.10, SPDX, forbidden imports, test marks)

Known behaviour, documented: on PersonaPlex a cancel restarts the conversation context (a fresh Stage 0 request replays the voice/persona prefill).

BEFORE SUBMITTING: read CONTRIBUTING.md and run the precheck-pr skill with the code agent for a self-check against project conventions.
(anything written below this line will be removed by GitHub Actions)

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/model_integration.md, docs/design/module/ar_runtime.md.

Module owners: @tzhouam @gcanlin @fake0fan

Routing: @tzhouam via module of the changed files, semantic router, CODEOWNERS; @gcanlin via module of the changed files, semantic router; @fake0fan via module of the changed files

@chickeyton, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 17, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit 768f6b81a587 produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

@hsliuustc0106 hsliuustc0106 added high priority high priority issue, needs to be done asap core related to core module: cache, scheduler, engine, worker, modelrunner omni code related to omni models labels Sep 17, 2026
@chickeyton
chickeyton force-pushed the fit_peronaplex branch 3 times, most recently from 7518637 to 6831782 Compare September 21, 2026 03:19
@chickeyton chickeyton changed the title [WIP][Core][Model] Fit PersonaPlex into the Unified Full-duplex Framework … [Core][Model] Fit PersonaPlex into the Unified Full-duplex Framework (RFC #7181 PR 2) Sep 21, 2026
@chickeyton

chickeyton commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor Author

Pre-check report (precheck-pr, full mode)

Branch fit_peronaplex @ 68317826, diffed against upstream/main @ 1b87115c (51 files). Type: General (framework + model port + tests + docs; no new model registry entries, no diffusion, no perf claims).

Dimension Result
PR title format ✓ [Core][Model] … with the model named; WIP tag removed
Code quality ⚠ see notes: 1 test-only **kwargs passthrough; 1 except Exception as exc that re-raises as DuplexRuntimeConfigError (personaplex/duplex/plugin.py); Any/SimpleNamespace only in tests, no production Any; the two copies added (SamplingParams.clone() at session open, deepcopy of runtime config on session.update) are per-session, not per-step; no blocking calls on the serving path
Examples policy ✓ no new Python example paths (only examples/online_serving/personaplex/README.md edited)
Simplification ✓ two evidence-backed cleanups already applied in this head: an unused PersonaPlexPcmAppendReservation alias and an unused EncodeAudio re-export were removed; six best-effort except Exception swallows in the shared request_outputs.py helpers narrowed to the tensor/array errors they handle. Remaining candidate (warning): MiniCPMO45DataPlaneContext is now a field-less subclass of the framework DuplexDataPlaneContext, kept so MiniCPM-o's data-plane code and tests need not move
PR desc integrity ✓ every path and claim in the description is in the diff; validation table matches the runs below
Registry/config ✓ pipeline declares duplex_plugin; the three legacy pipeline fields stay on PipelineConfig for Nemotron (comment updated); deploy YAML unchanged
Dead code ✓ ruff F401/F841/F811 clean on all changed files; pre-framework runtime_extension.py / serving_adapter.py and the collect_ignore conftest deleted
Test coverage ✓ core changes (engine/duplex/plugin.py, session/emitter.py, session/model_channel.py, entrypoints/duplex/warmup.py) each have new tests; PersonaPlex suites re-enabled (#7636 Issue 16)
Local pre-commit gates ✓ pre-commit run --from-ref upstream/main --to-ref HEAD: ruff check/format, typos, markdownlint, mypy-3.10, test CI marks, SPDX, forbidden imports, torch.cuda all pass (mypy-3.10 initially reported 16 typing errors in the new test files; fixed in this head)
CI gates ✓ DCO passes on this head (sign-off added); build / pre-commit workflows pending at the time of writing
Unrelated changes ✓ diff limited to the duplex framework, the shared toolbox, PersonaPlex, MiniCPM-o adoption of shared helpers, their tests and docs
Accuracy / benchmark — not applicable (existing model; no perf claims). Audible-output and whole-frame checks come from the E2E driver

Verdict: 0 blocking | 2 warnings (code-quality notes above; the MiniCPM-o context subclass).

Validation on L20X (vLLM 0.29.0, torch 2.11.0+cu129, VLLM_USE_FLASHINFER_SAMPLER=0):

Run Result
CPU duplex suite (rebased tree, -n 0) 790 passed, 4 skipped
Affected suites after the typing fixes (PersonaPlex, plugin defaults, toolbox, runner scenario, wrapper) 170 passed, 5 skipped; the final head only narrows three exception catches and drops two unused re-exports, its full re-run is in progress on the same server
E2E driver, live server (two paced sessions + replacement) exit 0, audible whole-frame output, overflow refused with resource_exhausted
test_personaplex_duplex.py --run-level advanced_model 1 passed
response.cancel mid-stream new epoch, audio deltas every 0.4 s for the remaining input, no errors
DuplexClient + create_duplex_session_config on the live server 5.28 s of PCM16 audio back, no errors

Host note for reviewers reproducing this: the machine runs canhazgpu guard --enforce, so the server must be launched under canhazgpu run; the omni_server fixture loads dummy weights unless --run-level advanced_model|full_model is passed.

@chickeyton

chickeyton commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor Author

Follow-up to the pre-check report: the re-run on the final head 68317826 finished. pre-commit run --from-ref upstream/main --to-ref HEAD: all hooks pass. Affected suites (tests/engine/duplex, tests/entrypoints/duplex, tests/model_executor/common, tests/model_executor/models/personaplex, tests/model_executor/models/minicpmo_4_5/duplex, tests/engine/test_duplex_import_boundary.py, tests/e2e/online_serving/test_personaplex_duplex.py collected on CPU): 649 passed, 5 skipped. Verdict unchanged: 0 blocking, 2 warnings.

…(RFC vllm-project#7181 PR 2)

Port PersonaPlex onto the DuplexModelPlugin contract so a PersonaPlex
deployment is served by DuplexOmni over /v1/realtime?duplex=1 again.

Framework extensions (model-neutral, MiniCPM-o unchanged):
- DuplexModelPlugin.silence_unit_payload(): the runner's turn-continuation
  unit and the startup warmup use the plugin's own format/rate/length
  instead of a fixed 16 kHz payload.
- SessionEmitter.auto_responds() is implied when a model takes no client
  commits (supports_client_commit=False), so stock Realtime clients work
  on a lockstep model.
- DefaultDuplexModelSessionState, DuplexDataPlaneContext and a default
  validate_client_extra_body remove the per-model copies of shared code.

Shared toolbox at vllm_omni/model_executor/common/: request-output readers
(request_outputs.py), pcm_f32le helpers (audio/pcm.py), and the duplex
building blocks (payload validation, FixedFramePcmAppendBuffer,
CumulativeAudioTextDataPlane). MiniCPM-o adopts the shared helpers.

PersonaPlex port: PersonaPlexDuplexPlugin replaces the pre-framework
runtime extension + serving adapter pair; the Stage 0 lockstep runtime is
keyed by (session_id, epoch) so a response.cancel restarts the context on
a fresh Stage 0 request; the pipeline declares duplex_plugin.

Tests: the PersonaPlex duplex suites are re-enabled (vllm-project#7636 Issue 16), with
toolbox suites, a runner scenario driven by the real plugin, a warmup test
and a GPU pytest wrapper; the E2E driver uses server-allocated session ids.

Validated on L20X: 657 CPU tests pass; the E2E driver (two paced sessions,
admission overflow, slot recycling, audible 24 kHz output), the pytest
wrapper with real weights, and a mid-stream response.cancel all pass.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: chickeyton <ngton2014@gmail.com>
@chickeyton

Copy link
Copy Markdown
Contributor Author

Rebased onto main 5634f03e (62 commits, including the vLLM 0.30.0 rebase in #7820) as c0e6088f; the only conflicts were the lru_cache import in engine/duplex/plugin.py and entrypoints/duplex/warmup.py, where the silent-frames path now takes the plugin's own silence unit (warmup_silence_unit) inside the new video_turn / silent_frames structure.

Re-verified on the same L20X host with a venv upgraded to vLLM 0.30.0 (torch 2.13.0+cu130):

Run Result
pre-commit run --from-ref upstream/main --to-ref HEAD all hooks pass
CPU duplex suite (-n 0) 878 passed, 4 skipped
E2E driver, live server (two paced sessions + replacement) exit 0; audible whole-frame output (coverage 0.97-0.99), third open refused with resource_exhausted
test_personaplex_duplex.py --run-level advanced_model 1 passed
response.cancel mid-stream new response under epoch 1, further audio deltas, no errors
DuplexClient + create_duplex_session_config 5.28 s of PCM16 audio back, no errors

🤖 Generated with Claude Code

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
@linyueqian

Copy link
Copy Markdown
Collaborator

@chickeyton I pushed dd629d0df onto this branch: a merge of current main (6715c57b2) that resolves the conflicts. Resolution:

  • vllm_omni/engine/duplex/plugin.py: took main's DefaultDuplexModelSessionState (now slots=True) and kept this PR's DuplexDataPlaneContext, which the shared PersonaPlex data plane still uses.
  • minicpmo_4_5/duplex/session.py: took main's version.
  • docs/design/fullduplex.md, docs/serving/full_duplex_api.md: took main's side; please re-apply any PersonaPlex wording you want there.

Verified on the merged tree with vLLM 0.30.0 on one Hopper-class GPU: the duplex CPU suites (tests/model_executor/models/personaplex tests/engine/duplex tests/entrypoints/duplex tests/model_executor/common tests/model_executor/models/minicpmo_4_5/duplex tests/engine/test_duplex_import_boundary.py) pass, and the realtime lifecycle driver passes on personaplex.yaml (two sessions, third refused with resource_exhausted, slot reuse, audible output).

PersonaPlex duplex has been refused on main since #7413, so it would be good to land this. #8192 builds on it for multi-session.

@linyueqian linyueqian added the ready label to trigger buildkite CI label Sep 27, 2026
@linyueqian
linyueqian enabled auto-merge (squash) September 27, 2026 01:45

@linyueqian linyueqian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving to unblock PersonaPlex duplex, which has been refused on main since #7413.

Verified on dd629d0df (this branch merged with current main), vLLM 0.30.0, one Hopper-class GPU:

  • Duplex CPU suites pass (PersonaPlex, engine/duplex, entrypoints/duplex, model_executor/common, MiniCPM-o duplex, import boundary).
  • Realtime lifecycle driver passes on personaplex.yaml: two paced sessions, third refused with resource_exhausted, slot reuse, audible output.
  • Paced load: N=1 realtime. Concurrency beyond that is addressed in #8192.

… a package-name clash

tests/model_executor and tests/model_executor/models have no __init__.py, so
tests/model_executor/common and tests/model_executor/models/common were both
imported as the top-level package 'common' and the second one failed to collect.

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
@linyueqian

Copy link
Copy Markdown
Collaborator

Also pushed 5bd4875cf: renamed tests/model_executor/common/ to tests/model_executor/executor_common/. tests/model_executor/ and tests/model_executor/models/ have no __init__.py, so this package and the existing tests/model_executor/models/common/ were both imported as top-level common, and the Model Executor jobs on CUDA and ROCm failed to collect tests/model_executor/models/common/test_*.py. Both directories now collect and pass together (148 passed).

@linyueqian linyueqian added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 27, 2026
…able_cpu_offload

vllm-project#7579 added a read of the legacy enable_cpu_offload flag and vllm-project#7327 then banned
such reads outside the compatibility layer, so test_legacy_flag_readers fails on
main whenever the diffusion suite runs. enable_cpu_offload maps to
OffloadStrategy.MODEL_LEVEL; use the resolved policy as flux2 and wan2.2 do.

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
@linyueqian
linyueqian requested a review from wtomin as a code owner September 27, 2026 02:37
@linyueqian linyueqian added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 27, 2026
@linyueqian

Copy link
Copy Markdown
Collaborator

Pushed 768f6b81a to clear the CUDA "Simple · Diffusion Test" failure, which is not from this PR: test_legacy_flag_readers fails on main too, because #7579 reads enable_cpu_offload in pipeline_qwen_image.py and #7327 later banned such reads outside the compatibility layer. The Qwen-Image pipeline now uses resolve_offload_strategy(od_config) is OffloadStrategy.MODEL_LEVEL, the same pattern as flux2 and wan2.2. test_legacy_flag_readers.py and test_qwen_image_pipeline_device.py pass locally.

@linyueqian
linyueqian merged commit 798d52c into vllm-project:main Sep 27, 2026
7 of 9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

core related to core module: cache, scheduler, engine, worker, modelrunner high priority high priority issue, needs to be done asap omni code related to omni models ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants