Repository navigation
[Core][Duplex] Ephemeral turn-commit framework primitives - #7682
mjZhaoElaine wants to merge 3 commits into
Conversation
The unified full-duplex framework (vllm-project#7413) assumes a resident, resumable stage0 request per session. Turn-commit models (AURA vllm-project#7633, Qwen3-Omni issue vllm-project#7636 follow-up 5) instead need one ordinary, non-resumable request per committed turn. Land the model-agnostic primitives so those plugins can port without forking session internals: - DuplexStageSubmission.resumable + duplex_ephemeral_stage_request_id: turn-scoped ...r.stage0_t{N} ids when the plugin capability supports_core_resumable_request is off; helpers.stage0_request_id and manager.stage_request_id/ensure_stage_request branch on it, and the orchestrator threads the flag into the engine-core request. - model_channel: re-bind a stale ephemeral id (a listen-only turn that never completed) by mechanically completing the turn, minting a fresh id, and aborting the orphaned stage request; bind the submitted id to the session; no silence continuation for a non-resumable stage0. - append_task: only the stable ...r.stage0 placeholder is cleared after a listen-only append; turn-scoped ids stay bound for the re-bind. - DuplexModelPlugin.observe_stage_output: opt-in hook projecting an intermediate stage to the client without short-circuiting the pipeline (the runner still forwards it downstream); stream close and continuation stay with the final stage or a direct decide_output decision. L1/L2 coverage in tests/engine/duplex/test_session_runner_ephemeral_bind.py: id shape and capability branching, two committed turns getting fresh stage0 ids through the real session runner, re-bind with stale-request abort, and observe projection that still forwards. Signed-off-by: Mengjie Zhao <zmj0129@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com>
The manager harness is a MiniCPM-shaped resident Stage0, so FakePlugin must opt into supports_core_resumable_request. Pin observe-finished, silence-continuation refusal, listen-only keep-bound, and orchestrator threading of resumable=False. Signed-off-by: Mengjie Zhao <zmj0129@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com>
Abort only the stale Stage0 request on re-bind so leftover downstream bindings stay, and compare the resident placeholder by exact id instead of a suffix. Cover a real listen-only submit followed by the next commit. Signed-off-by: Mengjie Zhao <zmj0129@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com>
|
This PR appears to belong to: docs/design/module/engine_orchestration.md, docs/design/module/observability.md. Module owners: @fake0fan @tzhouam @lishunyang12 Routing: @fake0fan via module of the changed files, semantic router, CODEOWNERS; @tzhouam via module of the changed files, semantic router, CODEOWNERS; @lishunyang12 via module of the changed files @mjZhaoElaine, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Omni ReviewBot triage noteAutomated triage of commit
These are automated triage suggestions only — the final decision belongs to the maintainers. |
Omni ReviewBot: no human activity for 7 days@mjZhaoElaine this pull request has had no human commit, comment or review since 2026-09-16. Please confirm the current plan and next step. The author or a maintainer decides whether to change the PR state. To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline. |
Omni ReviewBot: no human activity for 14 days@mjZhaoElaine this pull request has had no human commit, comment or review since 2026-09-16. Please consider marking this PR as draft until work can resume. The author or a maintainer decides whether to change the PR state. To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline. |
Omni ReviewBot routing recordAssigned Strict on cursor (cursor-grok-4.6-high) under experiment |
vllm-omni-review-bot
left a comment
There was a problem hiding this comment.
Omni ReviewBot review
6 actionable finding(s).
CI at
d3c8979b3bc3(2026-10-10T02:07:28.098712+00:00): required check(s) blocking:buildkite/vllm-omni(missing).
Full review analysis
Scan:
| Category | Result |
|---|---|
| Tests / verification | 4 finding(s) below |
| Security | no finding reported |
| Docs / comments | 2 finding(s) below |
| Behavior / compatibility | no finding reported |
| Correctness | no finding reported |
Validated:
- [resolved] test_session_manager FakePlugin now sets supports_core_resumable_request=True (tests/engine/duplex/test_session_manager.py:200) so the MiniCPM-shaped harness does not silently switch to ...r.stage0_tN. Residual: tests/engine/duplex/test_duplex_plugin.py:85 still returns DuplexCapabilities() (False) but only load/validate, never mints stage ids.
- [resolved] Second commit during an active response: DEFER_ACTIVE_RESPONSE needs overlap_speech_ms>0, then runner.py:1727 skips flush while response_in_progress. Residual: barge-in/overlap still out of scope as stated.
- [claim-verified] MiniCPM-o 4.5 / PersonaPlex / Nemotron set supports_core_resumable_request=True (minicpmo_4_5/duplex/capabilities.py:29, personaplex serving_adapter.py:129, nemotron serving_adapter.py:215).
- [claim-verified] vllm_omni/engine/duplex/ is in omni_minicpmo_4_5_duplex source_file_dependencies — MiniCPM GPU duplex CI will run even though the author skipped it
- [claim-verified] PersonaPlex adapter sets True: personaplex/duplex/serving_adapter.py:129. Nemotron adapter sets True: nemotron_voicechat/duplex/serving_adapter.py:215.
- [claim-verified][22c] Off-fixture census: only minicpmo_4_5/pipeline.py:27 declares duplex_plugin=; PersonaPlex/Nemotron do not reach DuplexSessionManager, so they do not inherit the new False id-minting until a port.
6 actionable finding(s).
Verdict: REQUEST CHANGES
Findings
- **[P1] This hunk pops only
(0, stale_ephemeral_id)thencleanup([stale_id], abort=T…** —vllm_omni/engine/duplex/session/model_channel.py`
Evidence for This hunk pops only `(0, stale_ephemeral_id)` then `cleanup([stale_id], abort=T…
This hunk pops only (0, stale_ephemeral_id) then cleanup([stale_id], abort=True), and the new comment claims downstream bindings are left alone. Unchanged production DuplexOrchestrator._on_stage_submitted binds Stage1 as (stage_id, request_id) with the same Stage0 id (test_forwarded_stage_requests_are_bound_and_barge_in_aborts_them already pins both pools aborting that id). cleanup(abort=True) calls _abort_request_ids, which fans the id to every stage pool, so in-flight TTS for the turn is killed while (1, stale_id) can remain in session.request_resources. The new tests bind leftover Stage1 as a different ...r.stage1 id that _on_stage_submitted never uses. The pop also happens before cleanup succeeds; a failed abort is swallowed and the next submit proceeds with the engine request still running; _close_runner snapshots resource_request_ids() after the pop so a Stage0-only leak is not retried. Drop every resource key whose request_id equals the stale id, restore those keys and raise if cleanup fails, and stop claiming Stage1 is left alone.
Evidence: Trigger: non-resumable ephemeral session whose Stage0 id is already submitted (listen-only append or incomplete turn); next commit hits the new rebind. Adverse: abort of that id kills in-flight Stage1 TTS (same request_id), leftover (1, stale_id) stays in session.request_resources, and a failed abort is swallowed so the next submit proceeds while close will not retry a Stage0-only leak.
vllm_omni/engine/duplex/session/model_channel.py:217-219 # Downstream stage bindings are left alone — they are created by / # orchestrator forward, not by this Stage0 append path.
vllm_omni/engine/duplex/session/model_channel.py:224 session.request_resources.pop((stage_id, stale_ephemeral_id), None)
vllm_omni/engine/duplex/session/model_channel.py:226-233 await self._ctx.stage_port.cleanup([stale_ephemeral_id], abort=True) then except Exception: logs and does not re-raise.
Unchanged by this diff, present in the PR-time tree: vllm_omni/engine/duplex_orchestrator.py:125 runner.session.bind_stage_request(stage_id, request_id, fence=fence) — Stage1 key is (1, same Stage0 id).
Unchanged: vllm_omni/engine/orchestrator.py:2199-2202 _on_stage_submitted(next_logical, req_id, replica_id, req_state) — forward reuses req_id.
Unchanged: vllm_omni/engine/duplex_orchestrator.py:344-345 cleanup → _cleanup_request_ids; vllm_omni/engine/orchestrator.py:1465-1466 if abort: abort_outputs = await self._abort_request_ids(cleanup_ids); vllm_omni/engine/orchestrator.py:626-627 for pool in self.stage_pools: stage_outputs = await pool.abort_requests(request_ids).
Unchanged: vllm_omni/engine/duplex/session/manager.py:598-599 _close_runner snapshots session.resource_request_ids() after the pop.
Unchanged proof of same-id fan-out: tests/engine/test_duplex_orchestrator.py:375-382 _on_stage_submitted(1, request_id, ...) then barge-in clients[0].abort_calls == [[request_id]] and clients[1].abort_calls == [[request_id]].
In this diff: tests/engine/duplex/test_session_runner_ephemeral_bind.py:697-712 binds leftover Stage1 as duplex_resource_request_id(..., "stage1") (a different id) and asserts it is not cleaned — not the production (1, first_id) binding.
Suggestion: stale_records = {
key: session.request_resources.pop(key)
for key in list(session.request_resources)
if key[1] == stale_ephemeral_id
}
try:
await self._ctx.stage_port.cleanup([stale_ephemeral_id], abort=True)
except Exception:
session.request_resources.update(stale_records)
logger.warning(
"duplex abort of stale ephemeral request failed session=%s id=%s",
session.session_id,
stale_ephemeral_id,
exc_info=True,
)
raise
Evidence for The leftover abort-scope bind (`leftover_stage1 = duplex_resource_request_id(..…
The leftover abort-scope bind (leftover_stage1 = duplex_resource_request_id(..., "stage1") at :697, twin :746) mints a resident ...r.stage1 id that DuplexOrchestrator._on_stage_submitted never creates — forward binds Stage1 as (1, first_id) (test_forwarded_stage_requests_are_bound_and_barge_in_aborts_them). leftover_stage1 not in cleanups cannot fail if abort wiped the real downstream binding: rebind only cleanup([first_id], abort=True), and aborting that shared id already kills every stage pool. Bind (1, first_id) to match production; after popping every request_resources key for first_id, assert (1, first_id) is gone rather than that a synthetic ...r.stage1 was spared.
Evidence: Trigger: after a finished Stage1 delivery with no model_turn_id (or a listen-only Stage0 submit), the next commit rebinds a stale ephemeral Stage0 id. tests/engine/duplex/test_session_runner_ephemeral_bind.py:697-698 leftover_stage1 = duplex_resource_request_id(h.session.fence, "stage1") / h.session.bind_stage_request(1, leftover_stage1, fence=h.session.fence) (twin :746-747) then :711-712 assert leftover_stage1 not in {rid for ids, _abort in h.port.cleanups for rid in ids} / assert (1, leftover_stage1) in h.session.request_resources. That leftover id is not production: unchanged by this diff, present in the PR-time tree: vllm_omni/engine/duplex_orchestrator.py:125 runner.session.bind_stage_request(stage_id, request_id, fence=fence) and tests/engine/test_duplex_orchestrator.py:375 orchestrator._on_stage_submitted(1, request_id, replica_id, request_state) with :377 assert session.stage_request_submitted(1, request_id) — Stage1 is (1, first_id). Rebind only drops Stage0: vllm_omni/engine/duplex/session/model_channel.py:224 session.request_resources.pop((stage_id, stale_ephemeral_id), None) and :226 await self._ctx.stage_port.cleanup([stale_ephemeral_id], abort=True). Abort of that shared id already tears down every pool — unchanged by this diff: vllm_omni/engine/orchestrator.py:617 """Forward abort requests to all stage pools. / :626-627 for pool in self.stage_pools: / stage_outputs = await pool.abort_requests(request_ids) or []; tests/engine/test_duplex_orchestrator.py:381-382 assert clients[0].abort_calls == [[request_id]] / assert clients[1].abort_calls == [[request_id]]. Adverse effect: the leftover assertion cannot fail if abort wiped the real (1, first_id) binding, and it pins a false "downstream stays" contract while abort(first_id) already kills Stage1 in every pool and leaves (1, first_id) as stale session bookkeeping.
Suggestion: h.session.bind_stage_request(1, first_id, fence=h.session.fence)
# The next commit cannot reuse the finished t0 id: the turn is
# completed mechanically, a fresh t1 id is minted, and abort of the
# shared Stage0 request_id tears down every stage bound under it.
await h.run(append_audio())
await h.run(commands.Commit(create_response=True))
assert len(h.port.submissions) == 2
second = h.port.submissions[1]
assert second.context.request_id.endswith(".r.stage0_t1")
assert second.already_submitted is False
assert h.session.turn_id == 1
assert h.port.cleanups == [([first_id], True)]
assert (1, first_id) not in h.session.request_resources
Evidence for The rebind gate at model_channel.py:213 only aborts when the *current* turn-sco…
The rebind gate at model_channel.py:213 only aborts when the current turn-scoped Stage0 id is still submitted. After a completed turn, complete_model_turn advances turn_id and does not drop request_resources, so the next append mints ...r.stage0_t{N+1} and skips this block (pinned by assert h.port.cleanups == []). Orchestrator cleanup is if request_finished and not req_state.session_owned (unchanged), and duplex ensure_request always sets session_owned=True, so each submit_initial-per-turn leaves a finished id in request_states / request_resources / _request_index and keeps running_counter incremented until session close. MiniCPM reused one resident id; this path accumulates one finished Stage0 id per turn. Release previously submitted ephemeral Stage0 ids (pop those (stage, id) keys, stage_port.cleanup, unregister_request) when other submitted Stage0 resources exist — not only when the newly minted id is already submitted — and drop the happy-path empty-cleanups assert. The draft snippet still sits behind the current-id submitted check and would not run on the completed-turn path.
Evidence: Trigger: two committed turns with supports_core_resumable_request=False (test_ephemeral_two_committed_turns_get_fresh_stage0_ids). After turn 0 finishes with model_turn_id, turn_id advances and the next commit mints a new id, so the rebind/abort block never runs and the finished t0 id stays registered until session close.
vllm_omni/engine/duplex/session/model_channel.py:213 if not resumable and session.stage_request_submitted(stage_id, request_id): — gate keyed on the id just minted from the current fence.turn_id.
unchanged by this diff, present in the PR-time tree: vllm_omni/engine/duplex/session/engine_session.py:639-644 def complete_model_turn(self, turn_id: int) -> None: / if completed_turn_id >= self.turn_id: / self.turn_id = completed_turn_id + 1 / self.sync_fence() — advances turn_id only; does not pop request_resources.
unchanged by this diff, present in the PR-time tree: vllm_omni/engine/duplex/session/engine_session.py:540-544 def clear_request(...) / self._response.active_request_id = None — end-of-turn clear does not unregister stage resources.
unchanged by this diff, present in the PR-time tree: vllm_omni/engine/orchestrator.py:1727 if request_finished and not req_state.session_owned: — finished duplex ids are not passed to _cleanup_request_ids (the only pop of request_states and running_counter.decrement).
unchanged by this diff, present in the PR-time tree: vllm_omni/engine/duplex_orchestrator.py:286 session_owned=True, and :337 self._register_running_request(request_state) — every new ephemeral submit_initial creates a session-owned state and increments the running counter.
tests/engine/duplex/test_session_runner_ephemeral_bind.py:673 assert h.port.cleanups == [] — happy path pins that the finished t0 Stage0 id is not aborted/released before t1 submit.
Suggestion: stale_ephemeral_ids = [
resource.request_id
for (sid, rid), resource in list(session.request_resources.items())
if sid == stage_id and resource.submitted and rid != request_id
]
if not resumable and session.stage_request_submitted(stage_id, request_id):
Evidence for `supports_core_resumable_request` now selects Stage0 id shape (`stage0_request_…
supports_core_resumable_request now selects Stage0 id shape (stage0_request_id / ensure_stage_request / ModelChannel.append: resident ...r.stage0 vs ephemeral ...r.stage0_t{N}), but DuplexCapabilities' class docstring still only says the scheduler can resume the same request id. Document the minting contract there; default False is a silent switch to turn-scoped ids (why FakePlugin had to pin True).
Evidence: Unchanged by this diff, present in the PR-time tree: vllm_omni/engine/duplex/config.py:73-75 supports_core_resumable_request means the scheduler can resume the same request id across streaming updates, but it is not a KV lease by itself. — no id-shape/minting language. Same file:96 supports_core_resumable_request: bool = False. This PR makes that flag the minting switch: vllm_omni/engine/duplex/session/helpers.py:43-47 if session.capabilities.supports_core_resumable_request: return duplex_resource_request_id(fence, "stage0") else return duplex_ephemeral_stage_request_id(fence, stage_id=0); manager.py:428-429 resumable = bool(session.capabilities.supports_core_resumable_request) then stage_request_id(..., resumable=resumable); model_channel.py:211-212 same. Trigger: a plugin leaves the dataclass default (or reads only the scheduler-resume docstring). Adverse/unmet: Stage0 is minted as ...r.stage0_t{N} instead of resident ...r.stage0 without the public capability contract saying so. contracts.py:83-84 (in-diff, on DuplexStageSubmission.resumable, not the plugin-facing field) already admits the capability selects the request-id shape.
Evidence for This head branches Stage0 on `supports_core_resumable_request` (dataclass defau…
This head branches Stage0 on supports_core_resumable_request (dataclass default False): False mints ...r.stage0_t{N} and forces already_submitted=False, so the port submit_initials a new ordinary request each committed turn. The owner design page (docs/design/fullduplex.md:297-300) still states the old universal contract (one preregistered resumable Stage0, submit_initial then submit_update), and duplex_orchestrator.py:8 still says the stage port submits resumable Stage0 requests. Update both; a plugin author following the owner page will implement the stale resident path.
Evidence: Trigger: a plugin that leaves DuplexCapabilities.supports_core_resumable_request at its default (or sets False). Unmet requirement: the owner duplex design page still documents only the old resident Stage0 contract after this PR made the default path ephemeral. Unchanged by this diff, present in the PR-time tree: docs/design/fullduplex.md:297-300 the resumable Stage0 request is preregistered at open with \DuplexOrchestratorRequestState(session_owned=True, session_id, fence)`, submitted with `submit_initial` on the first unit and `submit_update` afterwards. Unchanged by this diff: vllm_omni/engine/duplex_orchestrator.py:8 to submit resumable Stage0 requests. Unchanged by this diff: vllm_omni/engine/duplex/config.py:96 supports_core_resumable_request: bool = False. This PR: vllm_omni/engine/duplex/session/helpers.py:43-47 if session.capabilities.supports_core_resumable_request: return duplex_resource_request_id(fence, "stage0")elsereturn duplex_ephemeral_stage_request_id(fence, stage_id=0); vllm_omni/engine/duplex/session/model_channel.py:238 already_submitted = False if not resumable else session.stage_request_submitted(stage_id, request_id)` — False therefore never takes the documented submit_update-on-resident-...r.stage0 path.
Suggestion: cleanup, abort_requests): Stage0 is preregistered at
open with DuplexOrchestratorRequestState(session_owned=True, session_id, fence). When supports_core_resumable_request is True the resident ...r.stage0
id is submit_initial on the first unit and submit_update afterwards; when it is False each committed turn is a new ordinary ...r.stage0_t{N} submit_initial.
Evidence for test_listen_only_submit_rebinds_on_the_next_commit (and the missing-model_turn_…
test_listen_only_submit_rebinds_on_the_next_commit (and the missing-model_turn_id twin at line 709) never assert h.session.active_request_id == second.context.request_id after the rebind commit. Commit(create_response=True) hits _start_append with final=True, which binds helpers.stage0_request_id while turn_id is still 0 (...stage0_t0). The new session.bind_request(request_id) at model_channel.py:282 is what retargets onto ...stage0_t1; drop it and active_request_id stays on the aborted t0 id while these tests still pass.
Evidence: Trigger: after a still-bound turn-0 ephemeral Stage0 submit, Commit(create_response=True) rebinds. Unmet requirement: neither rebind test locks post-rebind active_request_id, so they miss a regression of the new bind. tests/engine/duplex/test_session_runner_ephemeral_bind.py:740 assert h.session.active_request_id == first_id (only the pre-rebind t0 id). tests/engine/duplex/test_session_runner_ephemeral_bind.py:757 assert h.session.turn_id == 1 then :758 assert h.port.cleanups == [([first_id], True)] — no active_request_id == second.context.request_id. Twin at :709 assert h.session.turn_id == 1 with the same omission. Unchanged by this diff, present in the PR-time tree: vllm_omni/engine/duplex/session/runner.py:937 request_id = helpers.stage0_request_id(self.session, append_epoch) then :938-939 if final or precreate_response: / session.bind_request(request_id) — Commit(create_response=True) uses this with turn_id still 0. helpers.py:42 fence = DuplexFence(session.session_id, epoch=epoch, turn_id=session.turn_id) so that pre-bind is ...stage0_t0. vllm_omni/engine/duplex/session/model_channel.py:282 session.bind_request(request_id) is what retargets onto ...stage0_t1 after aborting t0; without it active_request_id stays on the aborted request and both tests still pass.
Suggestion: assert h.session.turn_id == 1
assert h.session.active_request_id == second.context.request_id
🤖 This review was generated by InferMatrix Copilot, an open-source repo-maintenance agent for PR review, CI debugging and issue triage. Try it on your own repo, and ⭐ star it if it helped!
Omni ReviewBot: finding feedback**[p1] The leftover abort-scope bind ( See the review for details. If you are the PR author and disagree, react 👎 here; the maintainer will see your disagreement. |
Omni ReviewBot: finding feedback**[p1] This hunk pops only See the review for details. If you are the PR author and disagree, react 👎 here; the maintainer will see your disagreement. |
Omni ReviewBot: finding feedback[p1] The rebind gate at model_channel.py:213 only aborts when the current turn-sco… — See the review for details. If you are the PR author and disagree, react 👎 here; the maintainer will see your disagreement. |
Summary
Toward #7636 (Qwen3-Omni Server VAD restore). The unified full-duplex framework (#7413) assumes one resident, resumable Stage0 request per session. Turn-commit models (Qwen3-Omni, and the same shape as AURA #7633) need one ordinary, non-resumable request per committed turn instead.
This PR lands the model-agnostic primitives, gated on the existing
supports_core_resumable_requestcapability so MiniCPM-o 4.5 / PersonaPlex / Nemotron stay on today's resident Stage0 path:...r.stage0_t{N}) plusDuplexStageSubmission.resumablewhen the capability is off.model_turn_id): complete the turn, mint a fresh id, abort only that Stage0 request. Downstream bindings created by orchestrator forward are left alone.is_stable_stage0_placeholder), not a.r.stage0suffix.DuplexModelPlugin.observe_stage_output: default-off hook that projects an intermediate stage to the client and still forwards it downstream (unlikedecide_output)....r.stage0placeholder is auto-cleared after a listen-only append.supports_core_resumable_requestdefaults toFalseonDuplexCapabilities. That default is now load-bearing: a test or plugin that leaves it unset gets ephemeral ids. MiniCPM-o 4.5 / PersonaPlex / Nemotron already set itTrue. The MiniCPM-shapedFakePluginintests/engine/duplex/test_session_manager.pydoes the same so the resident-path harness does not silently switch.Overlapped-input / barge-in machinery is intentionally out of scope. A second commit while a response is already in progress is deferred by the existing
DEFER_ACTIVE_RESPONSE/response_in_progresspath; it does not abort the in-flight request.Follow-up
This PR is the model-agnostic slice only. A follow-up will add the Qwen3-Omni
DuplexModelPluginand aqwen3_omni_moe_duplexpipeline variant (opt-in via the existingpipeline:field plussession_mode: duplex; no Issue 20 change). Re-enabling theserver_vadRealtime E2E is a later PR (Issue 18 of #7636). Overlapped-input / barge-in stays out of this series.Testing Done
CPU-only L1/L2 on RunPod (vLLM 0.28.0, vLLM-Omni
d3c8979b3, Python 3.12.3, RTX PRO 4500 Blackwell / CUDA 13.0). Recipe: profilevo-s01-ephemeral-bind --cpu-only --bootstrap.Step 10 now includes the listen-only submit → next commit case and leftover-Stage1 abort-scope asserts (11 → 13). Step 20 is the existing duplex suite plus orchestrator / omni-engine /
tests/entrypoints/duplex.Commands equivalent to CI Simple · Engine&Entrypoints:
vLLM Version: 0.28.0
Testing Gap: GPU / real stage-pool submit of
resumable=Falseis not in this PR (needs the Qwen3-Omni plugin follow-up and the later realtime E2E). MiniCPM-o 4.5 GPU duplex e2e is unchanged and not re-run here.Made with Cursor