Repository navigation
[Core] Graduate MiniCPM-o 4.5 and PersonaPlex full-duplex serving out of experimental - #6196
Conversation
|
Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits. |
|
@Sy0307 Please review, see if that compactible with nemo-voicechat-11B |
|
This PR appears to be related to model: minicpm. Model owners: @y-null @chickeyton, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
fd52837 to
52cacaf
Compare
linyueqian
left a comment
There was a problem hiding this comment.
Ran the relocated suites and the duplex e2e on an H200-class GPU server (vLLM 0.27.0, torch 2.13.0+cu130, Python 3.12) against 52cacaf07.
| Batch | Result |
|---|---|
| Relocated duplex unit batch (18 suites) | 622 passed, 0 failed, on both vLLM 0.26.0 and 0.27.0 |
tests/e2e/online_serving/test_minicpmo_4_5_duplex.py |
3 passed, 308s, -m cuda --run-level advanced_model |
tests/e2e/online_serving/personaplex_realtime_duplex.py |
passed, errors: [] |
The MiniCPM run covers all three tests rather than only the core_model one CI selects. The PersonaPlex driver returned two concurrent sessions with identical capability payloads, resource_exhausted on the third, a recycled slot, and non-silent 24 kHz mono output on all three sessions (7.28 / 14.88 / 7.28 s, RMS around 0.03, frame coverage 0.968 to 0.989).
One environment note that is not caused by this PR: on vLLM 0.26.0 every staged server dies at startup with TypeError: GPUModelRunner.maybe_randomize_inputs() got an unexpected keyword argument 'randomize_inputs'. I reproduced that byte for byte on a clean main worktree at the merge base, so it is a version floor rather than a regression here. Stage 2 also needs s3tokenizer and step-audio2 / hyperpyyaml, which live in the dev extra rather than requirements/cuda.txt.
The relocation itself checks out. All 5,693 vllm_omni.* import statements at this head resolve statically, no experimental.fullduplex production references remain (including in the tree merged against current main), all four plugin dotted strings resolve, and the lazy to eager import change introduces no cycles.
Findings below, three of them inline.
[Medium] Module design metadata not updated by the move.
docs/design/module/engine_orchestration.md:46 still lists tests/e2e/features/fullduplex/engine/** in validation_paths, which this PR empties. Neither that page nor docs/design/module/entrypoints.md claims the new locations: vllm_omni/engine/duplex/**, vllm_omni/entrypoints/duplex/**, vllm_omni/entrypoints/duplex_request_client.py. Worth noting entrypoints/openai/duplex_capability.py was covered by the vllm_omni/entrypoints/openai/** glob before the move and now sits outside every declared path. Everything else lands inside an existing glob, so this is the full list.
[Low] Two public symbols removed.
DuplexCapabilities.minicpmo45_native() becomes the module function minicpmo45_native_capabilities(), and entrypoints/duplex/serving.py drops _minicpmo_session_state / _minicpmo_sessions / _minicpmo_data_plane, which carried a comment describing them as accessors for downstream tests and extensions. Both are good changes and no in-repo caller is left. Flagging only because a caller that follows the stated relocation contract and updates just the import path gets an AttributeError.
[Low] Title. [WIP] is not one of the prefixes listed in docs/contributing/README.md, and the PR title becomes the squash headline on main. Both commits already carry [Core] subjects.
On the scope claim.
The description says a pure relocation with no behavior change, but the PersonaPlex demo tier is deleted rather than moved: the standalone server, the elastic batching from #4771, the offline example, the headless client, and run_server.sh. The README and recipe do say so, so this is a description accuracy point rather than a silent removal. Two things would help reviewers who only read the body: state the deletion there, and say what replaces the --batch-size concurrency, since vllm_omni/deploy/personaplex.yaml ships max_sessions: 2 and the unified path admits one scheduler request per session rather than advancing B slots on a shared 80 ms tick.
Separately, both PersonaPlex docs this PR rewrites still tell users to run --stage-configs-path, which the CLI now rejects (tests/entrypoints/test_serve.py:60 asserts the rejection). --deploy-config is the working flag and is what I used. That is repo-wide debt across several recipes rather than something introduced here, but this PR removes run_server.sh, so the non-working command becomes the only documented way to serve PersonaPlex.
| from vllm_omni.experimental.fullduplex.client import ( | ||
| from vllm_omni.metrics import definitions as defs | ||
| from vllm_omni.metrics.utils import coerce_positive_int_scalar | ||
| from vllm_omni.model_executor.models.minicpmo_4_5.duplex.client import ( |
There was a problem hiding this comment.
[High] This is generic benchmark plumbing importing a model-private module.
client.py is a generic Realtime WebSocket client with three non-MiniCPM consumers: this file, tests/e2e/online_serving/personaplex_realtime_duplex.py:21, and the MiniCPM demos. Filing it under models/minicpmo_4_5/duplex/ means importing it executes vllm_omni/model_executor/models/minicpmo_4_5/duplex/__init__.py, which eagerly loads adapter, input, policy, and stage0. So vllm-omni bench serve (vllm_omni/benchmarks/serve.py:11 -> vllm_omni/entrypoints/cli/benchmark/serve.py:13) now pulls in the MiniCPM duplex serving adapter and Stage-0 runtime, and a PersonaPlex-only validation path depends on MiniCPM's package.
vllm_omni/entrypoints/duplex/client.py would sit next to the existing consumers and keep the plugin boundary this PR is establishing.
|
|
||
| forbidden_prefixes = ( | ||
| "vllm_omni.model_executor.models.minicpmo_4_5.duplex", | ||
| "vllm_omni.model_executor.models.minicpmo_4_5.duplex.client", |
There was a problem hiding this comment.
[High] The probe above imports only vllm_omni.engine.async_omni_engine, vllm_omni.engine.orchestrator, and vllm_omni.entrypoints.async_omni, so this forbidden list does not cover vllm_omni/benchmarks/patch/patch.py:61, which imports vllm_omni.model_executor.models.minicpmo_4_5.duplex.client unconditionally. The invariant reads as enforced, but the one shipped module that violates it is outside the scope of the check.
This entry is also redundant: the match is name == prefix or name.startswith(prefix + "."), so line 51 already covers ...duplex.client.
| from vllm_omni.outputs import OmniRequestOutput | ||
|
|
||
| DUPLEX_OUTPUT_DECISION_KEY = "_vllm_omni.experimental.fullduplex.duplex_output_decision" | ||
| DUPLEX_OUTPUT_DECISION_KEY = "_vllm_omni.duplex.duplex_output_decision" |
There was a problem hiding this comment.
[Low] The value changed here, from _vllm_omni.experimental.fullduplex.duplex_output_decision.
The writer (vllm_omni/engine/orchestrator.py:1622) and the readers (vllm_omni/entrypoints/duplex_request_client.py:359, vllm_omni/model_executor/models/minicpmo_4_5/duplex/data_plane.py:151) all move together, so nothing breaks in-tree. Calling it out because this key rides in OmniRequestOutput._custom_output across the engine to client boundary, which makes it a value change under a "no behavior change" description.
There was a problem hiding this comment.
the rename is intentional, and you're right that it deserved a callout under the PR description rather than riding silently. I've added a behavior note to the description.
Rationale for keeping the new value rather than preserving the old string: the key rides in the underscore-private OmniRequestOutput._custom_output, which isn't a public wire contract, and the writer (engine/orchestrator.py) and all readers (entrypoints/duplex_request_client.py, models/minicpmo_4_5/duplex/data_plane.py) move atomically in this PR — there is no supported mixed-version engine/frontend deployment that could observe the old key. Keeping _vllm_omni.experimental.fullduplex.duplex_output_decision would embed a permanent experimental.fullduplex reference in the stable tree, which is exactly what this migration removes.
441d725 to
ba2c1af
Compare
| handshake_queue = self._add_subscriber() | ||
| try: | ||
| self._reader_task = asyncio.create_task(self._read_loop(), name="duplex-client-reader") | ||
| await self.send( |
There was a problem hiding this comment.
[High] session_id alone cannot resume or take over an existing session.
Entering the client always sends session.update, while the server only enters its resume path when the first event is session.resume, which additionally requires incarnation, resume_token, and last_received_server_event_seq. Therefore, constructing a new DuplexClient with a known session_id will attempt to create that session (and may collide with the existing attachment) rather than resume it, contrary to the public API description and realtime_duplex_api.md.
Please either add explicit resume credentials and select the session.resume handshake on initial connection, or narrow the documentation to say that only automatic reconnect within the same DuplexClient instance is supported and that session_id merely chooses the ID for a new session.
There was a problem hiding this comment.
You're right — the entry handshake is always session.update, so a fresh
DuplexClient given a known session_id would try to create that session, not
resume it. Fixed in 458a342 by narrowing the claim on both surfaces:
- realtime_duplex_api.md now says session_id only names the session the
client creates, that resume is supported solely as automatic reconnect
within the same DuplexClient instance (via ReconnectPolicy), and that
re-attaching from a new client or process requires the wire-level
session.resume handshake (resume_token, incarnation,
last_received_server_event_seq), which this client does not expose. - The DuplexClient class docstring states the same contract.
and, I will propose to drop incarnation and client side chose session id in the follow up refactor RFC, so the resume logics will be refactored later
| for queue in list(self._subscribers): | ||
| await queue.put(event) | ||
|
|
||
| if isinstance(event, SessionClosed): |
| if request_started_at_s is not None: | ||
| # Lazily import so this module stays importable without the | ||
| # vllm_omni metrics stack (the client is otherwise standalone). | ||
| from vllm_omni.metrics.definitions import compute_audio_rtf |
There was a problem hiding this comment.
[Medium] This import breaks the lightweight dependency boundary advertised for the public client.
Importing vllm_omni.metrics.definitions first executes vllm_omni.metrics.__init__, which pulls in the Prometheus/statistics stack and dependencies such as prometheus_client and prettytable. In an environment containing only the documented client dependencies (pybase64 and websockets), the client lifecycle tests pass, but the three tests that call EventCollector.timing_summary() fail on these server-side dependencies.
Since compute_audio_rtf() is only a guarded division of generation time by audio duration, please calculate it locally or move the calculation to a genuinely dependency-free helper so that vllm_omni.clients.duplex remains usable with its documented dependency set.
There was a problem hiding this comment.
vllm_omni.metrics import was removed, now the function outputs audio_generation_ms and audio_duration_ms which are sufficient for computing the metric Audio RTF by the caller with compute_audio_rtf()
…sume Review fixes for the DuplexClient library (PR vllm-project#6196, round 24): - EventCollector.timing_summary no longer imports vllm_omni.metrics.definitions for the RTF value: that import executes vllm_omni.metrics.__init__, which pulls the Prometheus/statistics stack (prometheus_client, prettytable) into an environment that only carries the documented client dependencies (pybase64 + websockets). The RTF is the same guarded division, now computed locally with a comment pinning it to the server-side metric definition. A new import-boundary probe exercises timing_summary in a subprocess and fails if importing the client package pulls any vllm_omni module outside vllm_omni.clients. - The constructor's session_id only names the session to create; it cannot resume or take over an existing one (the entry handshake is always session.update, while the server-side resume path additionally needs resume_token, incarnation and last_received_server_event_seq). The class docstring and realtime_duplex_api.md now say exactly that: resume is automatic reconnect within the same client instance, and cross-client takeover requires the wire-level session.resume handshake, which this client does not expose. Signed-off-by: chickeyton <ngton2014@gmail.com> Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Mmoj3yYiiBZAdLFDiE4fKp
…dary probe timing_summary() takes after_s/response_id as keywords; the new client import-boundary probe passed response_id positionally and failed on TypeError instead of exercising the RTF path. Signed-off-by: chickeyton <ngton2014@gmail.com> Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Mmoj3yYiiBZAdLFDiE4fKp
… in callers Follow-up to the round-24 review fix for the client dependency boundary: instead of duplicating the RTF formula inside the client, EventCollector.timing_summary now reports raw measurements only (ttft_ms, ttfp_ms, audio_generation_ms, audio_duration_ms) and leaves derived metrics to the caller. The benchmarks (omniinteract and the Seed-TTS realtime path in the benchmark patch) derive the RTF from those raw fields through a shared helper that calls the canonical vllm_omni.metrics.definitions.compute_audio_rtf, so the metric has one definition in the tree and the client keeps its documented dependency set (pybase64 + websockets). summarize_session_request_metrics is unchanged: it averages whatever caller-assembled dicts provide, and its docstring now says rtf is caller-computed. Signed-off-by: chickeyton <ngton2014@gmail.com> Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Mmoj3yYiiBZAdLFDiE4fKp
Signed-off-by: Gao Han <hgaoaf@connect.ust.hk>
Keep turn-metrics on the graduated duplex_request_client path after vllm-project#6196. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: guozhihao-224 <guozhihaoemail@gmail.com>
… of experimental (vllm-project#6196) Signed-off-by: chickeyton <ngton2014@gmail.com> Signed-off-by: Gao Han <hgaoaf@connect.ust.hk> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Gao Han <hgaoaf@connect.ust.hk>
… of experimental (vllm-project#6196) Signed-off-by: chickeyton <ngton2014@gmail.com> Signed-off-by: Gao Han <hgaoaf@connect.ust.hk> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Gao Han <hgaoaf@connect.ust.hk> Signed-off-by: ZhengWG <zwg0606@gmail.com>
The /v1/realtime duplex output projector emitted the deprecated beta Realtime API names (response.audio.delta, response.audio_transcript.delta, response.audio.done, response.audio_transcript.done). The official openai Python SDK's non-beta openai.types.realtime module, and clients built against it such as livekit-plugins-openai's non-Azure code path, only recognize the current names (response.output_audio.delta, etc.) -- so an unmodified modern OpenAI Realtime client connected to any duplex model (PersonaPlex, MiniCPM-o 4.5, Nemotron VoiceChat) never saw any audio or transcript output at all. docs/serving/realtime_duplex_api.md claimed this dialect "deliberately emits and accepts both the beta and the GA spellings at once" for these events, alongside the genuinely-dual conversation.item.added/.created. Checked history: at 5d7fdf9 (vllm-project#6196), which first wrote that claim, the code already only ever emitted the beta name -- the doc was wrong from the start, not a later regression. Switch the four event names to their current spelling and update every in-repo consumer (the first-party DuplexClient, benchmark/example clients, and tests) to match, and correct the doc claim. Session-field and conversation.item.* dual beta/GA spelling support is unaffected -- only these four output event types were beta-only. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The /v1/realtime duplex output projector emitted the deprecated beta Realtime API names (response.audio.delta, response.audio_transcript.delta, response.audio.done, response.audio_transcript.done). The official openai Python SDK's non-beta openai.types.realtime module, and clients built against it such as livekit-plugins-openai's non-Azure code path, only recognize the current names (response.output_audio.delta, etc.) -- so an unmodified modern OpenAI Realtime client connected to any duplex model (PersonaPlex, MiniCPM-o 4.5, Nemotron VoiceChat) never saw any audio or transcript output at all. docs/serving/realtime_duplex_api.md claimed this dialect "deliberately emits and accepts both the beta and the GA spellings at once" for these events, alongside the genuinely-dual conversation.item.added/.created. Checked history: at 5d7fdf9 (vllm-project#6196), which first wrote that claim, the code already only ever emitted the beta name -- the doc was wrong from the start, not a later regression. Switch the four event names to their current spelling and update every in-repo consumer: the first-party DuplexClient and its example (barge_in_client.py), benchmark clients, the MiniCPM demo script and its browser JS, the server's own duplex warmup probe (api_server.py, which would otherwise time out on every startup waiting for an event that no longer arrives), design/serving docs, and tests across the duplex, PersonaPlex, and MiniCPM suites. Session-field and conversation.item.* dual beta/GA spelling support is unaffected -- only these four output event types were beta-only. Left the separate legacy (pre-duplex) Qwen3-Omni realtime fallback and the unrelated video-stream protocol untouched; they don't share this code or its OpenAI SDK type contract. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The /v1/realtime duplex output projector emitted the deprecated beta Realtime API names (response.audio.delta, response.audio_transcript.delta, response.audio.done, response.audio_transcript.done). The official openai Python SDK's non-beta openai.types.realtime module, and clients built against it such as livekit-plugins-openai's non-Azure code path, only recognize the current names (response.output_audio.delta, etc.) -- so an unmodified modern OpenAI Realtime client connected to any duplex model (PersonaPlex, MiniCPM-o 4.5, Nemotron VoiceChat) never saw any audio or transcript output at all. docs/serving/realtime_duplex_api.md claimed this dialect "deliberately emits and accepts both the beta and the GA spellings at once" for these events, alongside the genuinely-dual conversation.item.added/.created. Checked history: at 5d7fdf9 (vllm-project#6196), which first wrote that claim, the code already only ever emitted the beta name -- the doc was wrong from the start, not a later regression. Switch the four event names to their current spelling and update every in-repo consumer: the first-party DuplexClient and its example (barge_in_client.py), benchmark clients, the MiniCPM demo script and its browser JS, the server's own duplex warmup probe (api_server.py, which would otherwise time out on every startup waiting for an event that no longer arrives), design/serving docs, and tests across the duplex, PersonaPlex, and MiniCPM suites. Session-field and conversation.item.* dual beta/GA spelling support is unaffected -- only these four output event types were beta-only. Left the separate legacy (pre-duplex) Qwen3-Omni realtime fallback and the unrelated video-stream protocol untouched; they don't share this code or its OpenAI SDK type contract. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Signed-off-by: Nick Cao <ncao@redhat.com>
The /v1/realtime duplex output projector emitted the deprecated beta Realtime API names (response.audio.delta, response.audio_transcript.delta, response.audio.done, response.audio_transcript.done). The official openai Python SDK's non-beta openai.types.realtime module, and clients built against it such as livekit-plugins-openai's non-Azure code path, only recognize the current names (response.output_audio.delta, etc.) -- so an unmodified modern OpenAI Realtime client connected to any duplex model (PersonaPlex, MiniCPM-o 4.5, Nemotron VoiceChat) never saw any audio or transcript output at all. docs/serving/realtime_duplex_api.md claimed this dialect "deliberately emits and accepts both the beta and the GA spellings at once" for these events, alongside the genuinely-dual conversation.item.added/.created. Checked history: at 5d7fdf9 (vllm-project#6196), which first wrote that claim, the code already only ever emitted the beta name -- the doc was wrong from the start, not a later regression. Switch the four event names to their current spelling and update every in-repo consumer: the first-party DuplexClient and its example (barge_in_client.py), benchmark clients, the MiniCPM demo script and its browser JS, the server's own duplex warmup probe (api_server.py, which would otherwise time out on every startup waiting for an event that no longer arrives), design/serving docs, and tests across the duplex, PersonaPlex, and MiniCPM suites. Session-field and conversation.item.* dual beta/GA spelling support is unaffected -- only these four output event types were beta-only. Left the separate legacy (pre-duplex) Qwen3-Omni realtime fallback and the unrelated video-stream protocol untouched; they don't share this code or its OpenAI SDK type contract. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Signed-off-by: Nick Cao <ncao@redhat.com>
The /v1/realtime duplex output projector emitted the deprecated beta Realtime API names (response.audio.delta, response.audio_transcript.delta, response.audio.done, response.audio_transcript.done). The official openai Python SDK's non-beta openai.types.realtime module, and clients built against it such as livekit-plugins-openai's non-Azure code path, only recognize the current names (response.output_audio.delta, etc.) -- so an unmodified modern OpenAI Realtime client connected to any duplex model (PersonaPlex, MiniCPM-o 4.5, Nemotron VoiceChat) never saw any audio or transcript output at all. docs/serving/realtime_duplex_api.md claimed this dialect "deliberately emits and accepts both the beta and the GA spellings at once" for these events, alongside the genuinely-dual conversation.item.added/.created. Checked history: at 5d7fdf9 (vllm-project#6196), which first wrote that claim, the code already only ever emitted the beta name -- the doc was wrong from the start, not a later regression. Switch the four event names to their current spelling and update every in-repo consumer: the first-party DuplexClient and its example (barge_in_client.py), benchmark clients, the MiniCPM demo script and its browser JS, the server's own duplex warmup probe (api_server.py, which would otherwise time out on every startup waiting for an event that no longer arrives), design/serving docs, and tests across the duplex, PersonaPlex, and MiniCPM suites. Session-field and conversation.item.* dual beta/GA spelling support is unaffected -- only these four output event types were beta-only. Left the separate legacy (pre-duplex) Qwen3-Omni realtime fallback and the unrelated video-stream protocol untouched; they don't share this code or its OpenAI SDK type contract. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Signed-off-by: Nick Cao <ncao@redhat.com>
The /v1/realtime duplex output projector emitted the deprecated beta Realtime API names (response.audio.delta, response.audio_transcript.delta, response.audio.done, response.audio_transcript.done). The official openai Python SDK's non-beta openai.types.realtime module, and clients built against it such as livekit-plugins-openai's non-Azure code path, only recognize the current names (response.output_audio.delta, etc.) -- so an unmodified modern OpenAI Realtime client connected to any duplex model (PersonaPlex, MiniCPM-o 4.5, Nemotron VoiceChat) never saw any audio or transcript output at all. docs/serving/realtime_duplex_api.md claimed this dialect "deliberately emits and accepts both the beta and the GA spellings at once" for these events, alongside the genuinely-dual conversation.item.added/.created. Checked history: at 5d7fdf9 (vllm-project#6196), which first wrote that claim, the code already only ever emitted the beta name -- the doc was wrong from the start, not a later regression. Switch the four event names to their current spelling and update every in-repo consumer: the first-party DuplexClient and its example (barge_in_client.py), benchmark clients, the MiniCPM demo script and its browser JS, the server's own duplex warmup probe (api_server.py, which would otherwise time out on every startup waiting for an event that no longer arrives), design/serving docs, and tests across the duplex, PersonaPlex, and MiniCPM suites. Session-field and conversation.item.* dual beta/GA spelling support is unaffected -- only these four output event types were beta-only. Left the separate legacy (pre-duplex) Qwen3-Omni realtime fallback and the unrelated video-stream protocol untouched; they don't share this code or its OpenAI SDK type contract. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Signed-off-by: Nick Cao <ncao@redhat.com>
… of experimental (vllm-project#6196) Signed-off-by: chickeyton <ngton2014@gmail.com> Signed-off-by: Gao Han <hgaoaf@connect.ust.hk> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Gao Han <hgaoaf@connect.ust.hk>
As the PR 1 of [RFC]: Moving MiniCPM-o 4.5 and PersonaPlex out of Experimental
The full-duplex stack currently lives under
vllm_omni/experimental/fullduplex/. It was placed there when MiniCPM-o 4.5 duplex serving landed (#3907), but the situation has changed since then:/v1/realtime?duplex=1endpoint, has unit and E2E coverage wired into CI, and has been validated end to end on GPUs.experimental/fullduplex/is not model code at all — it is a model-neutral engine control plane and a model-neutral serving stack. Keeping generic infrastructure under an experimental namespace sends the wrong signal to users and makes the import paths awkward for the stable modules that already depend on it.The
joyvl/demo and the smallcore/scaffold it uses are genuinely demo-only and are not part of this proposal; they stay inexperimental/for now.Understanding of the current code structure: slides
Purpose
Pure Relocation
No behavior change
experimental/fullduplex/engine/vllm_omni/engine/duplex/openai/vllm_omni/entrypoints/duplex/minicpmo45/vllm_omni/model_executor/models/minicpmo_4_5/duplex/personaplex/(production glue only)vllm_omni/model_executor/models/personaplex/duplex/output.py,model_executor.py,request_client.pyvllm_omni/outputs/duplex.py,vllm_omni/model_executor/duplex_sampling.py,vllm_omni/entrypoints/duplex_request_client.pyExisting stable modules that need touching — in every case the change is an import-path or config-string rewrite, not a logic change:
vllm_omni/engine/orchestrator.pyengine/duplex/vllm_omni/engine/async_omni_engine.pyvllm_omni/entrypoints/async_omni.pyvllm_omni/entrypoints/openai/api_server.pyvllm_omni/worker/gpu_ar_model_runner.pyvllm_omni/model_executor/stage_input_processors/minicpmo_4_5_omni.pyvllm_omni/model_executor/models/minicpmo_4_5/(pipeline + model files)duplex_runtime_extension/duplex_serving_adapterdotted strings retargeted; model-file imports updatedvllm_omni/model_executor/models/personaplex/(pipeline + talker).buildkite/cuda/test-ready.yml,test-merge.ymlexamples/,tests/e2e/online_serving/)Points worth calling out:
duplex_runtime_extension/duplex_serving_adapterdotted-string paths in each model's pipeline config. The generic serving and engine code never imports model code — this is enforced by import-boundary tests that move along with everything else./v1/duplexstack already does in production. Only the 8 production glue files graduate.core_model and cpu, so they keep running in the existing CPU blocks without CI command changes.docs/design/with paths updated.A working branch with the full move exists and I am happy to open the PR(s) once there is agreement on the target layout.
New
DuplexClientAPI — public interface descriptionAdd
DuplexClientpython API for the/v1/realtime?duplex=1WebSocket endpoint. The package is client-side only and is never imported by
runtime/serving code (enforced by
tests/engine/test_duplex_import_boundary.pyand the AST guard in
tests/entrypoints/openai_api/test_duplex_capability.py).vllm_omni.clients.duplex.DuplexClientAnd add 2 utility functions for creating session config:
Benchmarking Update
The benchmark's playback reporting did not behave like a real listener: it
acknowledged each response's audio once, after the fact, instead of
reporting progress continuously while the audio played. In multi-turn
duplex sessions the conversation moves on while earlier audio is still
playing, so an after-the-fact acknowledgement can arrive after a later
user turn has been committed — which the server correctly refuses,
because accepting it could reorder conversation history. A benchmark must
exercise the protocol the way a real client does, so playback
acknowledgement was changed to report cumulative progress live,
checkpointing each response as soon as its audio starts. This makes
multi-turn cases complete under the server's history-ordering contract
instead of tripping over it, and measures the serving stack under
realistic client behavior.
Test Plan
vLLM Version: 0.28.0
1. Unittests
2. MiniCPM-o 4.5 barge in test
sequenceDiagram autonumber participant U as user (client script) participant A as assistant (server + model) Note over U,A: GREET U->>A: start a voice session (with the reference voice) A-->>U: ready Note over U,A: ASK U->>A: ask a question (speak the question WAV) Note right of A: decides on its own to answer<br/>(no VAD — the model chooses) A-->>U: starts answering aloud Note over U,A: INTERRUPT Note over U: listens for ~2 seconds par both talk at once U->>A: interrupts with a follow-up and A-->>U: still speaking the first answer end alt the interruption is substantial Note right of A: stops mid-sentence (barge-in),<br/>first answer is cut off else it was just a short remark Note right of A: finishes the first answer,<br/>then takes the follow-up end Note over U,A: ANSWER AGAIN A-->>U: answers the follow-up aloud Note over U,A: HANG UP U->>A: goodbye (close the session) Note over U: saves one recording per answer<br/>plus a summary of what happenedstart server
client
python examples/online_serving/barge_in_client.py \ --url ws://127.0.0.1:8098/v1/realtime \ --model openbmb/MiniCPM-o-4_5 \ --ref-audio "$HF_HOME/hub/models--openbmb--MiniCPM-o-4_5/snapshots/<rev>/assets/HT_ref_audio.wav" \ --question-wav tests/assets/minicpmo_4_5/response_required_16k.wav \ --interrupt-wav tests/assets/minicpmo_4_5/soft_interrupt_16k.wav \ --output-dir ./barge_in_outTest Result
1. Unittests
At tip
dae63b7a(after merging upstream main376b28ac), on the vLLM0.28.0 venv (torch 2.13.0+cu130), 4× L20X host:
demos, config factory): 518 passed, 0 failed
ruff check/ruff format --check, forbidden-imports, markdownlint,and typos hooks clean on every changed file
Live e2e (validated at
03f8d8a5, 1× L20X per run):test_duplex_websocket_protocol_smoke)response_required,video_input,resume_and_takeover)test_duplex_client_live.py(handshake, speakresponse via
ResponseHandle, playback acks, forced transport drop withautomatic
session.resume, 5 s heartbeat, clean close): 1 passed2. MiniCPM-o 4.5 barge in test
barge_in_test_files.zip
PASS
barge-in flow:
Inputs (what the client streamed):
Outputs (what the server produced):
Duplex E2E benchmark tables (PR #6522 format)
OmniInteract realtime benchmark, one deterministic case per subset,
--max-concurrency 1 --num-warmups 0, exactly the setup of PR #6522's"E2E Benchmark Result" section. Both runs: MiniCPM-o 4.5
(
vllm_omni/deploy/minicpmo_4_5.yaml), one NVIDIA L20X on test server 1,.venv-fd3-v028(vLLM 0.28.0), datasetlucky-lance/OmniInteract(pre-extracted), ref audio
MiniCPM-o-4_5/assets/HT_ref_audio.wav,2026-09-03. The per-subset deterministic selection drew the same cases in
both runs (
1q1a/0028.mp4,1q1a_math/0013.mp4,1qna/Blender_Banana_Pancakes-21_44.mp4), so the tables are directlycomparable.
fullduplex3@dae63b7a1q1a1q1a_math1qnaAll three runs completed with
Successful requests: 1,Failed requests: 0, streaming continuity OK rate 100 %, and a per-case.doneartifact; every case is included inofficial_eval_manifest.jsonl.The benchmark runner on this branch drives the public
vllm_omni.clients.duplex.DuplexClient.upstream
main@d1a5a7c11q1a1q1a_math1qnaAll three runs completed with
Successful requests: 1,Failed requests: 0, streaming continuity OK rate 100 %, and a per-case.doneartifact; every case is included inofficial_eval_manifest.jsonl.This run used the upstream benchmark runner on its legacy probe client
(the venv's editable install was re-pointed at the upstream checkout for
the whole pass and verified via the import path, then restored).
Why the ~8× difference in TTFT and Audio TTFP ?: Due to implementation difference in measurement, timestamping location. Main's legacy probe client stamps each frame inside its WebSocket recv loop — approximating true wire inter-arrival spacing (~0.1–0.3 ms). Fullduplex3's benchmark stamps in EventCollector.add when the consumer task drains the subscriber queue — when both frames are already queued, they get stamped back-to-back in one wakeup (~20–50 µs), so the queue hop compresses the gap. Fullduplex3 under-measures an interval main measures more faithfully; neither interval carries model-latency information.
Conclusion: There is no performance difference by this PR
Pre-check report for
fullduplex3@dae63b7a(PR #6196) — re-runRun 2026-09-03 (second pass, after the title/description fixes were applied
on GitHub) with the repo's
precheck-prskill, full mode, diffedagainst the upstream-main merge-base
376b28ac(178 files). The branch tipis unchanged since the first pass, so the code dimensions were re-swept and
match; the GitHub-side blockers were re-verified against the live PR.
Nothing was posted to GitHub; this report is local only.
claims present (benchmark tables in the PR body)
Cleared since the first pass
[Prefix][Core] Graduate MiniCPM-o 4.5 and PersonaPlex full-duplex serving out of experimentalappend_audio(final=…), no "data-URL JPEGs";stacked_video_framesand theresponses()error contract present), the Behavior notes section exists with both theduplex_output_decisionkey rename and thenative_duplexwire-name generalization, the Test Plan/Result read vLLM 0.28.0 / 768 passed /minicpmo_4_5.yaml(no 0.27.0 / 671 / deleted-yaml references), and the embedded 2026-08-17 pre-check block is goneDimension results
[Core]prefix, no WIP/Draft; commit subjects all correctly prefixedvllm_omni/clients/duplex.pytransport/cleanup boundaries: 3 re-raise as typed errors (send,_default_connect,__aenter__), 1 is the read-loop's resume boundary, 2 are best-effort cleanup swallows annotated# noqa: BLE001. None on fail-fast paths — deliberate shapes, discussed in review. 0 kwargs string-lookup plumbing in production, 0Anyhints, 0 hot-pathclone/deepcopy, 0 event-loop blocking, 0 newtorch.cudacall sitesexamples/online_serving/barge_in_client.py— model-neutral entrypoint (model via--model, behavior viavllm_omni.clientspresets), permitted by the policy; advisory: the per-preset glue in-script (_PRESET_DEFAULT_CHUNK_MS,_convert_to_session_format) would sit better on the preset modulesomniinteract._RealtimeSessionandpatch._RealtimeTTSProbeare two thin DuplexClient+EventCollector wrappers with overlapping shape — a shared probe-session helper invllm_omni.clientswould remove the duplication (advisory follow-up)check-spdx-header,check-mark,check-torch-cuda-call,check-forbidden-imports,check-buildkiteall Passed; ruff check/format,markdownlint-cli2,typosclean; DCO sign-off on all commits; noallowed_files/MAX_MODEL_TYPE_BRANCHESgrowth. ⚠mypy-3.10(a manual-stage hook, not a default gate) remains red across the relocated duplex serving files — the same untyped legacy patterns that moved; mypy's exclude list only covers model trees, so these files were equally red at their old experimental paths376b28ac(upstream main has since moved 6 commits, none touching the duplex path); GitHub reportsmergeable: true. ⚠mergeable_state: unstable— status checks need a fresh CI run on headdae63b7a(matches the reviewer's "re-run CI" note; not something the contributor fixes in the diff)Verdict
0 blocking | 4 warnings — ready for review.
All four warnings are advisory, none require pre-merge code changes:
boundaries already discussed in review.
a reasonable follow-up.
(
_RealtimeSession/_RealtimeTTSProbe) is a reasonable follow-up.ratchet posture, inherited by relocation.
The only remaining action item is external: trigger a fresh CI run on head
dae63b7ato clearmergeable_state: unstable.BEFORE SUBMITTING: read CONTRIBUTING.md and run the precheck-pr skill with the code agent for a self-check against project conventions.
(anything written below this line will be removed by GitHub Actions)