Repository navigation
[Bugfix][MiniCPM-o] Restore the duplex tests removed in #7413 - #7872
chickeyton wants to merge 4 commits into
Conversation
…st in vllm-project#7413 vllm-project#7413 rebuilt the duplex serving stack around engine-resident sessions and dropped the per-response request timing that vllm-project#7242 had added for the MiniCPM-o Omni-DuplexEval / OmniInteract benchmarks: the engine no longer emitted ``response_request_metrics``, so ``DuplexClient`` and the benchmark summaries silently fell back to the client's receive clock, and the TPOT weighting regressed to counting a one-token segment as one interval. Port the timing into the engine session: the model-turn request start is recorded when the model channel submits the append, the first text / audio of the response fix TTFT / TTFP, and the numbers ride response.created (under response.metadata.duplex_event) and every speak / audio delta (under metadata.vllm_omni), exactly where the client already reads them. Restore the interval-weighted TPOT mean. Re-home the MiniCPM-o plugin policy unit tests deleted with tests/engine/duplex/test_duplex_runtime.py, make the duplex e2e gate assert the server-side anchor, fix the pre-existing mypy errors in the two touched test files, and drop three doc passages that still describe the session-per-request chat adapter removed before vllm-project#7413 merged. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: chickeyton <ngton2014@gmail.com>
Pre-check report (
|
|
This PR appears to belong to: docs/design/module/engine_orchestration.md, docs/design/module/observability.md. Module owners: @tzhouam @fake0fan @lishunyang12 Routing: @tzhouam via module of the changed files, CODEOWNERS; @fake0fan via module of the changed files; @lishunyang12 via module of the changed files @chickeyton, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Omni ReviewBot triage noteAutomated triage of commit
These are automated triage suggestions only — the final decision belongs to the maintainers. |
|
@yenuo26 PTAL |
Omni ReviewBot: no human activity for 7 days@chickeyton this pull request has had no human commit, comment or review since 2026-09-20. Please confirm the current plan and next step. The author or a maintainer decides whether to change the PR state. To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline. |
Omni ReviewBot: no human activity for 14 days@chickeyton this pull request has had no human commit, comment or review since 2026-09-20. Please consider marking this PR as draft until work can resume. The author or a maintainer decides whether to change the PR state. To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline. |
Omni ReviewBot routing recordAssigned Strict on cursor (cursor-grok-4.6-high) under experiment |
What broke
Since #7413 a duplex session never reports the server-side per-response TTFT/TTFP that #7242 added for the MiniCPM-o Omni-DuplexEval / OmniInteract benchmarks.
response.createdand the audio deltas carry noresponse_request_metrics, sovllm_omni.clients.duplex.EventCollectorandbenchmarks/duplex_session_metrics.pysilently fall back to the client's receive clock ("source": "client_monotonic_receive"), whiledocs/cli/bench/serve.mdstill promises "Per-response TTFT and TTFP start when the server begins executing the native model-turn request". The same commit's TPOT weighting fix was lost too: a one-token segment is weighted as one interval again and skews the response mean.Repro
The strengthened e2e gate fails on
mainthe same way: every response'srequest_metrics.sourceisclient_monotonic_receiveinstead ofserver_request_start_and_client_receive.CUDA_VISIBLE_DEVICES=0 pytest -s -v tests/e2e/online_serving/test_minicpmo_4_5_duplex.py \ tests/e2e/online_serving/test_duplex_client_live.py -m "core_model and cuda" --run-level core_modelRoot cause
#7413 replaced
entrypoints/duplex/{protocol,runtime_bridge,session_runner}.pywith the engine-residentDuplexEngineSession/ModelChannel, and the port did not carry overmark_model_turn_request_started,mark_response_first_outputs, theresponse_request_metricsattachment, or themax(tokens - 1, 0)TPOT weight. The client side of #7242 survived, so the feature degraded silently instead of failing.Fix
engine_session.py: per-turn request-start ledger onResponseState,mark_model_turn_request_started/mark_response_first_outputs, binding atbegin_response, pruning atcomplete_model_turn, resets on end / barge-in / close, and the interval-weighted TPOT mean.model_channel.py: anchors the clock when the append is submitted and attaches the metrics toresponse.created(lands underresponse.metadata.duplex_event) and to every speak / audio delta (metadata.vllm_omni), the two places the client already reads.tests/engine/duplex/test_duplex_runtime.py, and the duplex e2e gate now asserts the server-side anchor. The 11 pre-existing mypy-3.10 errors in the two touched test files are fixed so pre-commit passes.DuplexChatCompletionsAdapterthat was removed before [Core][Frontend] Unified Full-duplex Framework #7413 merged; they now match the chat route that exists.Tests restored
Thirteen tests across four files: ten restore tests #7413 deleted, three are new.
From
tests/entrypoints/openai_api/test_duplex_handler.py(deleted by #7413), now intests/engine/duplex/test_engine_session.pytest_response_timing_binds_latest_request_start_for_model_turnDuplexEngineSessiontest_response_tpot_fallback_ignores_single_token_segmenttest_response_tpot_keeps_token_weighted_value_when_itls_existFrom the same file, now in
tests/engine/duplex/test_session_runner.pytest_continuous_response_metrics_accumulate_only_owned_model_units(theresponse_request_metricshalf)test_response_request_metrics_are_measured_from_the_model_turn_request_start, driven through the runner harness with the real MiniCPM-o pluginFrom
tests/engine/duplex/test_duplex_runtime.py(deleted by #7413), now in the newtests/model_executor/models/minicpmo_4_5/duplex/test_plugin_policy.pytest_minicpmo_extension_owns_stage_sampling_overridestest_plugin_owns_stage_sampling_overrides_without_mutating_defaultstest_minicpmo_output_decision_uses_raw_streaming_token_snapshottest_listen_decision_uses_the_raw_streaming_token_snapshottest_minicpmo_output_decision_ignores_output_level_token_history(x2 params)test_listen_decision_ignores_output_level_token_historytest_minicpmo_output_decision_uses_completion_token_ids(x2 params)test_listen_decision_uses_completion_token_idstest_minicpmo_output_decision_uses_completion_stop_reasontest_listen_decision_uses_completion_stop_reasontest_duplex_scheduler_token_budget_estimates_pcm_slotstest_scheduler_token_budget_estimates_pcm_slotstest_duplex_scheduler_token_budget_ignores_client_budget_fieldstest_scheduler_token_budget_ignores_client_budget_fieldsNew, no deleted counterpart
test_response_timing_is_empty_without_a_request_start_and_is_scoped_to_one_response(test_engine_session.py): lifecycle of the timing state across end, turn completion, and barge-in.test_a_speak_segment_is_not_a_listen_decision(test_plugin_policy.py): negative case for the listen decision.tests/e2e/online_serving/test_minicpmo_4_5_duplex.py:_assert_request_metricsnow requiressource == "server_request_start_and_client_receive"and the native request-start measurement origin, so the CI e2e gate catches this regression.Deliberately not restored
test_duplex_scheduler_token_budget_one_frame_is_one_block,..._matches_stage0_hd_slicing,..._unreadable_base_frame_keeps_hd_fallbackandtest_duplex_hd_slice_count_matches_official_gridfromtest_duplex_runtime.py: [Bugfix][MiniCPM-o] Size the duplex HD slice reservation from the frame #7654 re-covers them intests/model_executor/models/minicpmo_4_5/duplex/test_scheduler_token_budget.py.test_minicpmo_*handler tests, which targeted the serving architecture [Core][Frontend] Unified Full-duplex Framework #7413 removed.Validation (1x NVIDIA L20X, vllm 0.29.0, torch 2.13.0+cu129)
pre-commit run --files <changed>(all hooks incl. mypy-3.10)tests/engine/duplex tests/clients tests/model_executor/models/minicpmo_4_5 tests/benchmarks/{test_omniinteract.py,patch,metrics} tests/entrypoints/duplex tests/engine/test_duplex_orchestrator.py tests/engine/test_duplex_omni_engine.py tests/entrypoints/test_duplex_omni.py tests/entrypoints/openai_api/test_duplex_api_server.py tests/entrypoints/openai/test_duplex_session_attachment.py tests/worker/test_native_duplex_hooks.py tests/protocol tests/engine/test_duplex_import_boundary.py tests/engine/test_output_processor.py-m "core_model and cuda" --run-level core_modeltest_engine_session.py,test_session_runner.py,test_plugin_policy.py)Not restored on purpose: the frame-size-aware vision slot budget from #7271 was also dropped by #7413 but has since been re-fixed by #7654; the duplex Seed-TTS perf baselines removed in #7413 are left out because the native per-unit path changed.
🤖 Generated with Claude Code