Repository navigation
[Core][Model] Fit PersonaPlex into the Unified Full-duplex Framework (RFC #7181 PR 2) - #7695
Conversation
|
This PR appears to belong to: docs/design/module/model_integration.md, docs/design/module/ar_runtime.md. Module owners: @tzhouam @gcanlin @fake0fan Routing: @tzhouam via module of the changed files, semantic router, CODEOWNERS; @gcanlin via module of the changed files, semantic router; @fake0fan via module of the changed files @chickeyton, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Omni ReviewBot triage noteAutomated triage of commit
These are automated triage suggestions only — the final decision belongs to the maintainers. |
7518637 to
6831782
Compare
Pre-check report (precheck-pr, full mode)Branch
Verdict: 0 blocking | 2 warnings (code-quality notes above; the MiniCPM-o context subclass). Validation on L20X (vLLM 0.29.0, torch 2.11.0+cu129,
Host note for reviewers reproducing this: the machine runs |
|
Follow-up to the pre-check report: the re-run on the final head |
…(RFC vllm-project#7181 PR 2) Port PersonaPlex onto the DuplexModelPlugin contract so a PersonaPlex deployment is served by DuplexOmni over /v1/realtime?duplex=1 again. Framework extensions (model-neutral, MiniCPM-o unchanged): - DuplexModelPlugin.silence_unit_payload(): the runner's turn-continuation unit and the startup warmup use the plugin's own format/rate/length instead of a fixed 16 kHz payload. - SessionEmitter.auto_responds() is implied when a model takes no client commits (supports_client_commit=False), so stock Realtime clients work on a lockstep model. - DefaultDuplexModelSessionState, DuplexDataPlaneContext and a default validate_client_extra_body remove the per-model copies of shared code. Shared toolbox at vllm_omni/model_executor/common/: request-output readers (request_outputs.py), pcm_f32le helpers (audio/pcm.py), and the duplex building blocks (payload validation, FixedFramePcmAppendBuffer, CumulativeAudioTextDataPlane). MiniCPM-o adopts the shared helpers. PersonaPlex port: PersonaPlexDuplexPlugin replaces the pre-framework runtime extension + serving adapter pair; the Stage 0 lockstep runtime is keyed by (session_id, epoch) so a response.cancel restarts the context on a fresh Stage 0 request; the pipeline declares duplex_plugin. Tests: the PersonaPlex duplex suites are re-enabled (vllm-project#7636 Issue 16), with toolbox suites, a runner scenario driven by the real plugin, a warmup test and a GPU pytest wrapper; the E2E driver uses server-allocated session ids. Validated on L20X: 657 CPU tests pass; the E2E driver (two paced sessions, admission overflow, slot recycling, audible 24 kHz output), the pytest wrapper with real weights, and a mid-stream response.cancel all pass. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: chickeyton <ngton2014@gmail.com>
6831782 to
c0e6088
Compare
|
Rebased onto main Re-verified on the same L20X host with a venv upgraded to vLLM 0.30.0 (torch 2.13.0+cu130):
🤖 Generated with Claude Code |
Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
|
@chickeyton I pushed
Verified on the merged tree with vLLM 0.30.0 on one Hopper-class GPU: the duplex CPU suites ( PersonaPlex duplex has been refused on |
linyueqian
left a comment
There was a problem hiding this comment.
Approving to unblock PersonaPlex duplex, which has been refused on main since #7413.
Verified on dd629d0df (this branch merged with current main), vLLM 0.30.0, one Hopper-class GPU:
- Duplex CPU suites pass (PersonaPlex,
engine/duplex,entrypoints/duplex,model_executor/common, MiniCPM-o duplex, import boundary). - Realtime lifecycle driver passes on
personaplex.yaml: two paced sessions, third refused withresource_exhausted, slot reuse, audible output. - Paced load: N=1 realtime. Concurrency beyond that is addressed in #8192.
… a package-name clash tests/model_executor and tests/model_executor/models have no __init__.py, so tests/model_executor/common and tests/model_executor/models/common were both imported as the top-level package 'common' and the second one failed to collect. Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
|
Also pushed |
…able_cpu_offload vllm-project#7579 added a read of the legacy enable_cpu_offload flag and vllm-project#7327 then banned such reads outside the compatibility layer, so test_legacy_flag_readers fails on main whenever the diffusion suite runs. enable_cpu_offload maps to OffloadStrategy.MODEL_LEVEL; use the resolved policy as flux2 and wan2.2 do. Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
|
Pushed |
RFC #7181, PR 2. Ports PersonaPlex onto the
DuplexModelPlugincontract so a PersonaPlex deployment is served byDuplexOmniover/v1/realtime?duplex=1again, re-enabling the PersonaPlex duplex suites that #7413 had to skip (#7636 Issue 16). The other #7636 follow-up issues are handled in a separate PR.Purpose
Framework extensions (model-neutral; MiniCPM-o behaviour unchanged)
DuplexModelPlugin.silence_unit_payload(): the runner's turn-continuation unit and the startup warmup use the plugin's own format/rate/length instead of a fixed 16 kHz payload (PersonaPlex units are 1920 samples @ 24 kHz).SessionEmitter.auto_responds()is implied when a model takes no client commits (supports_client_commit=False), so a stock Realtime client works on a lockstep model withoutextra_body.auto_response.DefaultDuplexModelSessionState,DuplexDataPlaneContextand a defaultvalidate_client_extra_bodyremove the per-model copies of shared session code.Shared toolbox at
vllm_omni/model_executor/common/:request_outputs.py(RequestOutput / multimodal_output readers, cumulative-delta helpers),audio/pcm.py(pcm_f32le decode/count/materialise) andduplex/(append-payload validation,FixedFramePcmAppendBuffer,CumulativeAudioTextDataPlane). MiniCPM-o adopts the shared helpers; PersonaPlex subclasses the duplex building blocks.PersonaPlex port:
PersonaPlexDuplexPluginreplaces the pre-framework runtime extension + serving adapter pair; the Stage 0 lockstep runtime is keyed by(session_id, epoch)soresponse.cancel/output_audio_buffer.clearrestart the model context on a fresh Stage 0 request (the old epoch's Mimi encoder is recycled, the voice/persona prefill replayed); the pipeline declaresduplex_plugin. PersonaPlex advertises no client commits, no barge-in, no resume and no chat route, so/v1/chat/completionsis not mounted for it.Tests: PersonaPlex duplex suites rewritten for the plugin (
test_plugin.py, rekeyedtest_stage0_runtime.py), toolbox suites undertests/model_executor/common/, a runner scenario driven by the real plugin, a warmup test, plugin-default tests, and a GPU pytest wrapper; the E2E driver uses server-allocated session ids.Docs:
full_duplex_api.md,realtime_duplex_api.md,fullduplex.md, a rewrittenfullduplex-personaplex.md, the PersonaPlex example README, recipe and supported-models row.Test Plan
CPU (no GPU needed):
GPU (one card, gated
nvidia/personaplex-7b-v1checkout; defaultvllm_omni/deploy/personaplex.yaml):vLLM Version: 0.30.0
vLLM-Omni Commit: c0e6088 (this branch, rebased on main
5634f03e)Test Result
L20X (1 GPU), Python 3.12, vLLM 0.30.0, torch 2.13.0+cu130,
VLLM_USE_FLASHINFER_SAMPLER=0(the fresh venv had no FlashInfer JIT cache); re-run in full after the rebase onto main5634f03e:-n 0)resource_exhausted; ids server-allocated and distincttest_personaplex_duplex.py --run-level advanced_model(own server via theomni_serverfixture)response.cancelmid-streamDuplexClient+vllm_omni.clients.personaplex.create_duplex_session_configon the live serverpre-commit run --from-ref upstream/main --to-ref HEADKnown behaviour, documented: on PersonaPlex a cancel restarts the conversation context (a fresh Stage 0 request replays the voice/persona prefill).
BEFORE SUBMITTING: read CONTRIBUTING.md and run the precheck-pr skill with the code agent for a self-check against project conventions.
(anything written below this line will be removed by GitHub Actions)