Repository navigation
[Core][Model] Unified Full-Duplex MRv2 Support and MiniCPM-o 4.5 Adaptation #8536
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Open
BeatSeat
wants to merge
35
commits into
vllm-project:main
Choose a base branch
from
BeatSeat:feat/duplex-mrv2-minicpmo45
base: main
Could not load branches
Branch not found: {{ refName }}
Loading
Could not load tags
Nothing to show
Loading
Are you sure you want to change the base?
Some commits from the old base branch may be removed from the timeline,
and old review comments may become outdated.
+2,076
−164
Open
Changes from all commits
Commits
Show all changes
35 commits
Select commit
Hold shift + click to select a range
3642cc6
[Core][Model] Support MiniCPM-o 4.5 full-duplex serving on Model Runn…
BeatSeat 9ac7d31
fix(minicpmo_4_5): wait for terminal codec completion before duplex t…
BeatSeat dd9cc2e
fix(minicpmo): preserve duplex boundary metadata and keep streaming T…
BeatSeat af23ffe
fix(minicpmo): complete three-stage MRv2 duplex streaming
BeatSeat c41754a
fix(mrv2): preserve streaming request metadata in native transport
BeatSeat 15fa874
Merge main and preserve MRv2 duplex codec history
BeatSeat 2a63626
ci(minicpmo): run CUDA duplex merge and nightly coverage on MRv2
BeatSeat 3ef0ca2
Merge main after the MRv2 codec penalty fix landed
BeatSeat ea4e3c8
chore(minicpmo): narrow duplex MRv2 integration scope
BeatSeat 35115b4
test(minicpmo): remove stale duplex graph-bucket assertion
BeatSeat a7c1174
fix(minicpmo): isolate MRv2 duplex sessions and codec completion
BeatSeat dd7aa53
ci(minicpmo): add MRv2 ready E2E and retain V1 coverage
BeatSeat a3a5a9d
test(minicpmo): cover duplex concurrency reorder and slot reuse
BeatSeat d062622
perf(minicpmo): defer duplex history reads and retain async scheduling
BeatSeat cfb1107
refactor(minicpmo): keep MRv2 output channels and reuse shared error …
BeatSeat 9e38fc1
refactor(minicpmo): publish duplex Thinker payload through the MRv2 o…
BeatSeat 281ee6d
refactor(minicpmo): keep duplex sampling adaptation model-local
BeatSeat 63869f1
fix(worker_v2): suppress queued outputs of a failed preprocess request
BeatSeat c908fb3
perf(minicpmo45): keep MRv2 duplex meta tokens on the host
BeatSeat 5c436d8
fix(minicpmo45): keep the MRv2 Code2Wav stream open across duplex turns
BeatSeat b8d50ee
Merge main into the MRv2 duplex branch
BeatSeat 0632929
Merge main into the MRv2 duplex branch
BeatSeat 91a9f8f
docs(minicpmo45): point duplex MRv2 users at the duplex overlay
BeatSeat 24c6a88
revert(worker_v2): drop request-scoped preprocess error isolation fro…
BeatSeat 04c9aaa
perf(minicpmo45): cut MRv2 duplex host work that grows with sessions
BeatSeat 2e6e4c3
ci: drop the removed request_error module from the MRv2 duplex job deps
BeatSeat 0eade7d
Merge main into the MRv2 duplex branch
BeatSeat e9f7816
Free a MiniCPM-o duplex session cancelled before its first MRv2 prefi…
BeatSeat 8973921
Merge branch 'main' into feat/duplex-mrv2-minicpmo45
amy-why-3459 f775f78
Read each KV group's block table once per MiniCPM-o reanchor
BeatSeat 1e9aee1
Share the MiniCPM-o Talker native-duplex row metadata between V1 and …
BeatSeat 2731d8b
Parameterize the MiniCPM-o duplex final-segment chunk tests by delta …
BeatSeat db14d6f
Keep MiniCPM-o Stage-0 special token ids for rows without a unit on MRv2
BeatSeat 1258035
Drop pending MiniCPM-o codec frames when an MRv2 Talker request is ab…
BeatSeat 56f9f24
Keep the MRv2 multimodal encoder for the MiniCPM-o duplex Thinker
BeatSeat File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,49 @@ | ||
| # SPDX-License-Identifier: Apache-2.0 | ||
| # SPDX-FileCopyrightText: Copyright contributors to the vLLM-Omni project | ||
| """Resolve the MiniCPM-o three-stage duplex MRv2 deployment contract.""" | ||
|
|
||
| import pytest | ||
|
|
||
| from tests.helpers.stage_config import get_deploy_config_path | ||
| from vllm_omni.config.stage_config import ( | ||
| _apply_platform_overrides, | ||
| load_deploy_config, | ||
| merge_pipeline_deploy, | ||
| ) | ||
| from vllm_omni.model_executor.models.minicpmo_4_5.pipeline import MINICPMO_4_5_PIPELINE | ||
| from vllm_omni.platforms import current_omni_platform | ||
|
|
||
| pytestmark = [pytest.mark.core_model, pytest.mark.cpu] | ||
|
|
||
| _DEPLOY = "minicpmo_4_5_duplex_mrv2.yaml" | ||
|
|
||
|
|
||
| def _resolve_cuda_stages(monkeypatch): | ||
| # Resolve CUDA overrides even when this test runs on a non-CUDA host. | ||
| monkeypatch.setattr(current_omni_platform, "device_name", "cuda") | ||
| config = _apply_platform_overrides(load_deploy_config(get_deploy_config_path(_DEPLOY)), platform="cuda") | ||
| return config, merge_pipeline_deploy(MINICPMO_4_5_PIPELINE, config) | ||
|
|
||
|
|
||
| def test_duplex_mrv2_cuda_profile(monkeypatch) -> None: | ||
| config, stages = _resolve_cuda_stages(monkeypatch) | ||
| assert config.session_mode == "duplex" | ||
| # All three stages use MRv2; there is no hidden Talker fallback. | ||
| assert [s.yaml_engine_args["use_v2_model_runner"] for s in stages] == [True, True, True] | ||
| # Stage 0 duplex preprocessing stays on the synchronous path while the | ||
| # downstream stages keep streaming chunk transfer. | ||
| assert [s.yaml_engine_args["async_chunk"] for s in stages] == [False, True, True] | ||
| # Preserve asynchronous AR scheduling, including the Talker. | ||
| assert all(s.yaml_engine_args["async_scheduling"] for s in stages[:2]) | ||
| assert stages[2].yaml_engine_args.get("async_scheduling") is not False | ||
| # Retain the base profile capacity and Talker KV budget. | ||
| assert [s.yaml_engine_args["max_num_seqs"] for s in stages] == [16, 16, 16] | ||
| assert stages[1].yaml_engine_args["kv_cache_memory_bytes"] == 4 * 1024**3 | ||
|
|
||
|
|
||
| @pytest.mark.parametrize("platform", ["npu", "xpu", "rocm", "musa"]) | ||
| def test_duplex_mrv2_profile_keeps_non_cuda_on_v1(monkeypatch, platform) -> None: | ||
| monkeypatch.setattr(current_omni_platform, "device_name", platform) | ||
| config = _apply_platform_overrides(load_deploy_config(get_deploy_config_path(_DEPLOY)), platform=platform) | ||
| stages = merge_pipeline_deploy(MINICPMO_4_5_PIPELINE, config) | ||
| assert [s.yaml_engine_args["use_v2_model_runner"] for s in stages] == [False, False, False] |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
121 changes: 121 additions & 0 deletions
121
tests/e2e/online_serving/test_minicpmo_4_5_duplex_mrv2.py
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,121 @@ | ||
| # SPDX-License-Identifier: Apache-2.0 | ||
| # SPDX-FileCopyrightText: Copyright contributors to the vLLM-Omni project | ||
| """Real-weight MRv2 duplex concurrency guard, additional to the V1 CI jobs.""" | ||
|
|
||
| import asyncio | ||
| from pathlib import Path | ||
|
|
||
| import pytest | ||
|
|
||
| from tests.e2e.online_serving.helpers.minicpmo_4_5_duplex import ( | ||
| MODEL, | ||
| multi_session_args, | ||
| realtime_url, | ||
| resolve_ref_audio, | ||
| validated_input_wav, | ||
| ) | ||
| from tests.e2e.online_serving.run_minicpmo_realtime_duplex_multi_session import run_multi_session | ||
| from tests.e2e.online_serving.test_minicpmo_4_5_duplex import _run_seeded_text_to_audio | ||
| from tests.helpers.mark import hardware_test | ||
| from tests.helpers.runtime import OmniServerParams | ||
| from tests.helpers.stage_config import get_deploy_config_path | ||
|
|
||
| pytestmark = [pytest.mark.omni, pytest.mark.advanced_model] | ||
|
|
||
| _SERVER = OmniServerParams( | ||
| model=MODEL, | ||
| stage_config_path=get_deploy_config_path("minicpmo_4_5_duplex_mrv2.yaml"), | ||
| use_stage_cli=False, | ||
| server_args=["--trust-remote-code"], | ||
| ) | ||
|
|
||
|
|
||
| @hardware_test(res={"cuda": "H100"}, num_cards=1) | ||
| @pytest.mark.parametrize("omni_server", [pytest.param(_SERVER, id="mrv2-real-weights")], indirect=True) | ||
| @pytest.mark.parametrize("sessions", [1, 2, 4]) | ||
| def test_mrv2_duplex_overlapping_turns(omni_server, tmp_path: Path, sessions: int): | ||
| args = multi_session_args( | ||
| omni_server=omni_server, | ||
| input_wav=validated_input_wav(), | ||
| ref_audio=resolve_ref_audio(), | ||
| output_dir=tmp_path / f"concurrency_{sessions}", | ||
| response_required=True, | ||
| ) | ||
| args.sessions = sessions | ||
| args.turns = 2 | ||
| args.turn_duration_ms = [args.first_turn_ms] * args.turns | ||
| args.synchronized_start = True | ||
| args.disconnect_session_index = None | ||
| args.takeover_session_index = None | ||
| result = asyncio.run(run_multi_session(args)) | ||
| assert result["ok"] is True, result | ||
| assert result["session_count"] == sessions | ||
| assert result["identity_isolation_ok"] is True | ||
| assert not result["failures"] | ||
| for session in result["sessions"]: | ||
| assert session["audio_delta_count"] > 0, session | ||
| assert session["done_count"] == args.turns, session | ||
| assert session["error_count"] == 0, session | ||
|
|
||
|
|
||
| @hardware_test(res={"cuda": "H100"}, num_cards=1) | ||
| @pytest.mark.parametrize("omni_server", [pytest.param(_SERVER, id="mrv2-real-weights")], indirect=True) | ||
| def test_mrv2_duplex_new_session_during_long_response(omni_server): | ||
| # Unequal prompts and a delayed second session prevent identical clients | ||
| # from advancing in lockstep and hiding shared condition-state bugs. | ||
| ref_audio = resolve_ref_audio() | ||
|
|
||
| async def speak(text, delay): | ||
| await asyncio.sleep(delay) | ||
| return await _run_seeded_text_to_audio( | ||
| url=realtime_url(omni_server), | ||
| model=omni_server.model, | ||
| ref_audio=ref_audio, | ||
| text=text, | ||
| silence_seconds=30.0, | ||
| ) | ||
|
|
||
| async def overlap(): | ||
| return await asyncio.gather( | ||
| speak("What is the capital of China? Answer in about 40 words.", 0), | ||
| speak("What is the capital of France? Answer in one short sentence.", 2), | ||
| ) | ||
|
|
||
| first, second = asyncio.run(overlap()) | ||
| for result in (first, second): | ||
| assert "response.done" in result["event_types"], result | ||
| assert result["audio_bytes"] > 0, result | ||
| assert str(result["transcript"]).strip(), result | ||
| assert first["audio_bytes"] > 96_000, first | ||
| assert str(first["transcript"]).strip() != str(second["transcript"]).strip() | ||
|
|
||
|
|
||
| @hardware_test(res={"cuda": "H100"}, num_cards=1) | ||
| @pytest.mark.parametrize("omni_server", [pytest.param(_SERVER, id="mrv2-real-weights")], indirect=True) | ||
| def test_mrv2_duplex_server_answers_image_chat_from_the_image(omni_server, openai_client): | ||
| """The duplex Thinker also serves /v1/chat/completions: its image features must reach the prompt.""" | ||
| import base64 | ||
| import io | ||
|
|
||
| from PIL import Image | ||
|
|
||
| buffer = io.BytesIO() | ||
| Image.new("RGB", (224, 224), (255, 0, 0)).save(buffer, format="JPEG") | ||
| image_url = "data:image/jpeg;base64," + base64.b64encode(buffer.getvalue()).decode("ascii") | ||
| request_config = { | ||
| "model": omni_server.model, | ||
| "messages": [ | ||
| { | ||
| "role": "user", | ||
| "content": [ | ||
| {"type": "image_url", "image_url": {"url": image_url}}, | ||
| {"type": "text", "text": "What color is this image? Answer with one word."}, | ||
| ], | ||
| } | ||
| ], | ||
| "stream": True, | ||
| "modalities": ["text"], | ||
| "key_words": {"text": ["red"]}, | ||
| "extra_body": {"chat_template_kwargs": {"enable_thinking": False}}, | ||
| } | ||
| openai_client.send_omni_request(request_config) |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
[P2] test_mrv2_talker_reuses_confirmed_prompt_window only admits streaming_condition…
Evidence and suggested fix
test_mrv2_talker_reuses_confirmed_prompt_window only admits streaming_condition_seq 0→1. On MRv2 the stub sets chunk_transfer_adapter=None, so a skipped seq (0→3) takes the new adapter-less except at omni_ar_scheduler.py:1065-1068 (
_streaming_context_overflow[req_id]=...thenfinish_requests(..., FINISHED_ERROR)), the EngineCore-crash guard this PR added._make_schedulerusesOmniARScheduler.__new__and never allocates that dict (production__init__does). V1 already coversadapter.record_receive_failurevia test_talker_invalid_condition_sequence_does_not_advance_tracking; this else branch is unrun. Add a_native_data_plane=Trueskip-seq test that initializes_streaming_context_overflow = {}and assertsfinish_requestsplus the overflow record.Evidence: Trigger: MRv2 Talker
_update_request_as_sessionwith_native_data_plane=True,chunk_transfer_adapter is None, and streaming_condition_seq skip 0→3. Unmet requirement: the new adapter-less EngineCore-crash guard is never executed; only monotonic 0→1 is covered.tests/core/sched/test_omni_ar_scheduler_streaming.py:38
sched = OmniARScheduler.__new__(OmniARScheduler)— stub never sets_streaming_context_overflow.tests/core/sched/test_omni_ar_scheduler_streaming.py:45
sched.chunk_transfer_adapter = None— native path cannot takerecord_receive_failure.tests/core/sched/test_omni_ar_scheduler_streaming.py:1628
update = _make_talker_update(20, reserve=10, condition_seq=1)and :1625"streaming_condition_seq": 0— only 0→1.tests/core/sched/test_omni_ar_scheduler_streaming.py:1634
sched.finish_requests.assert_not_called()— success-only assertion; no skip-seq counterpart exists (grep found notest_mrv2_native_prompt/ seq-skip native test).vllm_omni/core/sched/omni_ar_scheduler.py:1063-1068
if chunk_transfer_adapter is not None: chunk_transfer_adapter.record_receive_failure(...)/else: self._streaming_context_overflow[req_id] = (session.client_index, str(exc)); if not session.is_finished(): self.finish_requests((req_id,), RequestStatus.FINISHED_ERROR)— this else is new in the PR.vllm_omni/model_executor/stage_input_processors/minicpmo_4_5_omni.py:1204-1206
or seq != previous_seq + 1: raise ValueError("native Talker streaming_condition_seq must advance by one").Unchanged by this diff, present in the PR-time tree: vllm_omni/core/sched/omni_ar_scheduler.py:162
self._streaming_context_overflow: dict[str, tuple[int, str]] = {}— production init allocates the dict, so this is a missing-test finding, not a live KeyError.Suggestion: sched.finish_requests.assert_not_called()
def test_mrv2_native_prompt_seq_skip_finishes_request() -> None:
sched = _make_scheduler(stage_id=1, session_mode="duplex")
sched._native_data_plane = True
sched._streaming_context_overflow = {}
sched.vllm_config.model_config.custom_process_next_stage_input_func = (
"vllm_omni.model_executor.stage_input_processors.minicpmo_4_5_omni.tts2code2wav_async_chunk"
)
sched.vllm_config.model_config.hf_config = SimpleNamespace(
model_type="minicpmtts", max_position_embeddings=100, attention_type="full_attention"
)
sched.max_model_len = 100
sched.finish_requests = MagicMock()
session = _make_request()
session.status = RequestStatus.WAITING_FOR_STREAMING_REQ
sched.num_waiting_for_streaming_input = 1
session.model_intermediate_buffer = {
"native_duplex": True,
"meta": {"next_stage_prompt_len": 20, "streaming_condition_seq": 0},
}
sched._update_request_as_session(session, _make_talker_update(20, reserve=10, condition_seq=3))
sched.finish_requests.assert_called_once()
assert session.request_id in sched._streaming_context_overflow