Skip to content

[Frontend] Emit per-turn --log-stats tables for native duplex - #6892

Open
guozhihao-224 wants to merge 8 commits into
vllm-project:mainfrom
guozhihao-224:frontend/duplex-turn-log-metrics
Open

guozhihao-224 wants to merge 8 commits into
vllm-project:mainfrom
guozhihao-224:frontend/duplex-turn-log-metrics

Conversation

@guozhihao-224

@guozhihao-224 guozhihao-224 commented Sep 1, 2026 •

Copy link
Copy Markdown

Purpose

Implements the Phase-1 slice of #6614. Does not close that issue (Related to #6614); Prometheus / audio TTFP / session rollup remain PR2+.

Chat /v1/chat/completions already prints Overall Summary / RequestE2EStats / [OmniTiming] / StageRequestStats when --log-stats is on. Native full-duplex (/v1/realtime?duplex=1) did not, because work rides a long-lived stage-0 request through append_duplex_input_async instead of generate().

This PR treats one assistant turn (response_id) as one chat completion:

  • DuplexSession.begin_response / end_response drive begin_turn / finalize_turn on the request client
  • ingest stage-0 + TTS into a per-turn aggregator keyed by response_id
  • auto_response snapshots that arrive before begin_response are buffered and flushed at begin; arrival_ts is commit/first-append (or first buffered stage_submit_ts)
  • _cancel_active_response maps native sources onto log reasons (session_close/disconnect* → close, timeout/new_response/input.cancel → cancel, real barge-in stays barge_in)
  • collapsed vllm_tpot_ms is token-weighted (weight = max(num_tokens_out - 1, 1)); one [OmniTiming] row per turn includes reason
  • connector TX that landed on the long-lived session aggregator is copied into the turn table
  • finalize on done, barge-in, close, and _abort (before pop), idempotent
  • table request_id column is the turn's response_id
  • chat generate() / _log_summary_and_cleanup are unchanged
  • chat fallback ids (chatcmpl-*) and existing fake duplex test ids do not duplex-finalize (is_duplex_resource_request_id requires duplex-s. + 8-part resource format)
  • hooks are native-runtime only; hook / summary failures are swallowed so a close path cannot drop request_states

Out of scope (RFC PR2+): Prometheus, audio TTFP, session rollup, protocol label, listen counter.

Test Plan

vLLM Version: 0.27.1 (aligned with Omni 0.27.0rc2; torch 2.13.0+cu130)
vLLM-Omni Commit: (fill after push)

# L1: new tests + related duplex regression
pytest -n 0 -s -v tests/metrics/test_duplex_turn_metrics.py \
  tests/metrics/test_stats.py \
  tests/entrypoints/test_async_omni_duplex.py \
  tests/entrypoints/openai_api/test_duplex_handler.py::test_cancel_active_response_maps_session_close_finished_reason \
  tests/entrypoints/openai_api/test_duplex_handler.py::test_cancel_active_response_maps_timeout_to_cancel

# L3: MiniCPM-o 4.5 native duplex, including auto_response if available
CUDA_VISIBLE_DEVICES=1 pytest -s -v tests/e2e/online_serving/test_minicpmo_4_5_duplex.py \
  -k 'test_duplex_single_session' \
  -m 'advanced_model and cuda' --run-level 'advanced_model'

Test Result

  • L1 tests/metrics/test_duplex_turn_metrics.py + tests/metrics/test_stats.py + cancel mapping tests + tests/entrypoints/test_async_omni_duplex.py: passed locally after the review follow-up
  • L3 MiniCPM-o 4.5 native duplex on H200 GPU 1 (earlier response_required run): 1 passed; auto_response path not re-run in this follow-up

L1 coverage: resource-id filter, turn accumulate/collapse with token-weighted tpot, two turns → two tables, log_stats off, auto_response metrics flushed at begin, session_close cancel → reason=close then later close() silent, barge-in / abort finalize before pop, one [OmniTiming] identity row, session TX copied into the turn aggregator, chat _log_summary_and_cleanup does not use duplex_turn, listen snapshots are not double-counted, hook / summary exceptions do not break protocol state.

Known pre-existing stats quirk, not in this PR: serving_time_to_first_output_ms looks like epoch-ms. Prometheus / audio TTFP are RFC PR2.

Not run: L2 MiniCPM-o protocol smoke (handshake only; it does not assert turn tables).

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/observability.md.

Module owners: @lishunyang12 @vraiti

@guozhihao-224, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 1, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit 5f33fbfb2069 produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

@guozhihao-224
guozhihao-224 force-pushed the frontend/duplex-turn-log-metrics branch from f40ac21 to db00e28 Compare September 1, 2026 03:36
@guozhihao-224

Copy link
Copy Markdown
Author

Self-review

Checked against the Phase-1 slice of #6614 (log tables only; Prometheus / audio TTFP are follow-ups).

  • One assistant turn (response_id) owns one OrchestratorAggregator. Chat generate() / _log_summary_and_cleanup and stats.py are unchanged.
  • Finalize is idempotent and runs before pop on response.done, barge-in, close, and _abort. Close / hook / summary failures must not drop request_states.
  • is_duplex_resource_request_id requires duplex-s. + the 8-part engine id, so chatcmpl-* and existing fake test ids (duplex-sid-e0-stage0) do not duplex-finalize.
  • Native-runtime hooks only. Listen OutputMessages that already have a duplex decision are not double-counted; TTS finished outputs still ingest.
  • Local: L1 53 passed on the duplex/metrics tests; L3 MiniCPM-o 4.5 single-session emitted two turn tables with reason=stop and table request_id == response_id.

@vllm-omni-review-bot

@shiy1022

shiy1022 commented Sep 1, 2026

Copy link
Copy Markdown

Reviewed at db00e28.

_normalize_finished_reason (vllm_omni/metrics/duplex_turn.py:34) collapses every non-stop termination to "abort". Six paths reach it: barge-in / input.cancel / response.cancel (request_client.py:220-222), session close (request_client.py:241), _abort (async_omni.py), runtime lifecycle teardown and the precreated-response cleanup (both finished_reason="abort"), and a superseded turn in begin_turn_metrics. Barge-in and session close are normal operation, so a healthy deployment reports mostly abort and the field can't separate interruption from failure — which is the distinction #6614's motivation asks for ("abort/barge-in turns are invisible to /metrics"). review-pr/references/modules/observability.md says to treat metric lifecycle as a public contract, and PR2 puts counters on this field, so widening later is a breaking change. signal() already has the triggering event in hand at :220; emitting barge_in / cancel / close / error from there costs little now.

accumulate_turn_stage_metrics (duplex_turn.py:76) takes no stage_submit_ts and stamps time.time() at receipt (:94), while the chat path at omni_base.py:483 prefers msg.stage_submit_ts and only falls back to now. The duplex tables therefore include orchestrator queue delay that the chat tables exclude, so the two aren't comparable — the stated goal of #6614. Same at the second call site, omni_base.py:551. Observability contract check: "align emitted values with their actual timing and lifecycle owner." Threading submit_ts into the helper looks sufficient.

Question rather than a finding: _collapse_stage_events (duplex_turn.py:119) skips any merged key that isn't an attribute of StageRequestStats, with one hand-written alias at :114. A field added to _merge_stage_metric_event later would be absent from duplex tables while still present in chat tables, and both would still look complete. Worth a warning on unmapped keys given the fail-closed gate in #6494?

Non-goals in #6614 look respected: keyed on response_id not the stage-0 resource id, segments collapsed rather than per-chunk tables, no wire-format change, Prometheus deferred. L3 evidence is pinned to hardware and versions.

@amy-why-3459

Copy link
Copy Markdown
Collaborator

@ZacheryAU PTAL

@guozhihao-224

Copy link
Copy Markdown
Author

Reviewed at db00e28.

_normalize_finished_reason (vllm_omni/metrics/duplex_turn.py:34) collapses every non-stop termination to "abort". Six paths reach it: barge-in / input.cancel / response.cancel (request_client.py:220-222), session close (request_client.py:241), _abort (async_omni.py), runtime lifecycle teardown and the precreated-response cleanup (both finished_reason="abort"), and a superseded turn in begin_turn_metrics. Barge-in and session close are normal operation, so a healthy deployment reports mostly abort and the field can't separate interruption from failure — which is the distinction #6614's motivation asks for ("abort/barge-in turns are invisible to /metrics"). review-pr/references/modules/observability.md says to treat metric lifecycle as a public contract, and PR2 puts counters on this field, so widening later is a breaking change. signal() already has the triggering event in hand at :220; emitting barge_in / cancel / close / error from there costs little now.

accumulate_turn_stage_metrics (duplex_turn.py:76) takes no stage_submit_ts and stamps time.time() at receipt (:94), while the chat path at omni_base.py:483 prefers msg.stage_submit_ts and only falls back to now. The duplex tables therefore include orchestrator queue delay that the chat tables exclude, so the two aren't comparable — the stated goal of #6614. Same at the second call site, omni_base.py:551. Observability contract check: "align emitted values with their actual timing and lifecycle owner." Threading submit_ts into the helper looks sufficient.

Question rather than a finding: _collapse_stage_events (duplex_turn.py:119) skips any merged key that isn't an attribute of StageRequestStats, with one hand-written alias at :114. A field added to _merge_stage_metric_event later would be absent from duplex tables while still present in chat tables, and both would still look complete. Worth a warning on unmapped keys given the fail-closed gate in #6494?

Non-goals in #6614 look respected: keyed on response_id not the stage-0 resource id, segments collapsed rather than per-chunk tables, no wire-format change, Prometheus deferred. L3 evidence is pinned to hardware and versions.

Thanks for the review — addressed in ce11c40.

  1. finished_reason: logs now pass through stop / barge_in / cancel / close / abort / error instead of collapsing everything non-stop to abort. signal() maps barge_in and *.cancel; session close uses close; native data-plane / append failures use error; _abort stays abort. Unknown values still map to abort. Prometheus buckets stay in PR2, but these are the values PR2 should reuse so we don't have to widen the contract later.

  2. stage_submit_ts: accumulate_turn_stage_metrics now takes it and uses the same first-ts rule as chat (submit_ts if set, else receipt time.time()). Both OmniBase ingest sites forward msg.stage_submit_ts.

  3. Unmapped collapse keys: left the merge contract unchanged and log dropped keys at debug. A warning on every turn felt too noisy given the existing audio_frames alias; happy to promote it if you'd rather fail closed in this PR.

L1: 57 passed (test_duplex_turn_metrics.py plus the duplex protocol / async omni tests).

@shiy1022

shiy1022 commented Sep 1, 2026

Copy link
Copy Markdown

Verified at ce11c40; both look right. _KNOWN_FINISHED_REASONS with signal() mapping barge_in / cancel at :222, close at :242, and the superseded turn staying abort. And first_value = stage_submit_ts if ... else now now matches omni_base.py:493, with both ingest sites forwarding.

Debug is fine for the unmapped keys — your noise point is right, a per-turn warning would be worse. If you want it visible without the spam, logger.warning_once from vllm.logger fires once per process, though nothing in vllm_omni uses it yet so it would be a new pattern here. Either way, not blocking.

@amy-why-3459

Copy link
Copy Markdown
Collaborator

Reviewed ce11c40 against the Phase-1 slice of #6614.

The skeleton is right: one assistant turn (response_id) owns one aggregator, chat generate() / _log_summary_and_cleanup / stats.py are untouched, finalize is idempotent and runs before pop, and listen OutputMessages that already have a duplex decision are not double-counted. ce11c40 fixed the helper — _KNOWN_FINISHED_REASONS plus stage_submit_ts matching chat — and the unit tests pin those two. Request changes on the landing path.

  1. Do not Fixes #6614.

    db00e28 says Fixes #6614 (PR1: log tables only). GitHub closes from the commit message; the parenthetical does not save it. Goals 2/3 (Prometheus request families, audio TTFP) are still PR2, and the body already says so. Use Related #6614 / “implements the Phase-1 slice”.

  2. Native cancel still lands as barge_in.

    The first finalize on the serving path is not signal() / close(). _cancel_active_response (the SessionRunner mixin every native WS cancel goes through) calls end_response(finished_reason="barge_in"). The hook finalizes immediately and clears duplex_turn. _abort (abort), signal() (barge_in/cancel), and close() (close) are then no-ops. Callers already pass reason: timeout, session_close, disconnect, disconnect_grace_expired, new_response. A healthy WS close or idle timeout therefore prints barge_in. That is the distinction [RFC]: Duplex Request Metrics #6614 asked for and the contract you told PR2 to reuse.

    Thread _cancel_active_response's reason into end_response and map it (session_close/disconnect* → close, timeout/new_response → cancel or abort). session_runner also folds input.cancel / response.cancel into "barge_in" before this function; fixing only DuplexRequestClient.signal does not change the table. Add a serving test: open turn → cancel with session_close → one table with close, later close() is silent.

Please also take these:

  • auto_response / START_AUTO_RESPONSE uses precreate_response=False. begin_response runs on the first speak packet in runtime_bridge, which is after the output handler has already dropped the triggering stage-0 StageMetricsMessage (test_metrics_before_begin_are_dropped makes that the contract). L3 only ran response_required, where precreate opens the turn before append. Open the aggregator at commit / first append, or buffer stage-0 until begin. Pass arrival_ts from commit; right now e2e is time.time() at begin_response.
  • _collapse_stage_events last-wins vllm_tpot_ms via _merge_stage_metric_event. The WS sidecar and RFC §4 are token-weighted. test_accumulate_appends_segments_collapse_at_finalize locks 20.0 instead of 16.0. Document that duplex stage_gen_time_ms is turn-accumulated, not one chat decode step — the metrics.md change is almost all table reflow and missed that.

@amy-why-3459

Copy link
Copy Markdown
Collaborator

@natureofnature @Sy0307 PTAL

@ZacheryAU ZacheryAU left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed at ce11c40.

Comment thread vllm_omni/metrics/duplex_turn.py Outdated
@shiy1022

shiy1022 commented Sep 1, 2026

Copy link
Copy Markdown

@amy-why-3459 is right, and my earlier suggestion pointed at the wrong site. Confirmed at ce11c40: _cancel_active_response (serving.py:1716) hardcodes finished_reason="barge_in" at :1736, while its own reason parameter — one of barge_in / disconnect / disconnect_grace_expired / new_response / session_close / timeout across six call sites — is only used in the emitted event dict. And session_runner.py:879-881 already folds input.cancel / response.cancel into "barge_in" before that. So the reason is discarded twice upstream of signal(), and fixing signal() alone can't change the table. I traced the helper's call sites, not the execution order on the native path.

@hsliuustc0106 hsliuustc0106 added frontend code related to entrypoint benchmark/profiler/metrics/logger Codes related to benchmarks, profiler, metrics and logger system labels Sep 2, 2026

@Sy0307 Sy0307 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Three targeted findings below. The duplicate [OmniTiming] issue is already covered by the existing thread.

Comment thread vllm_omni/experimental/fullduplex/request_client.py Outdated
Comment thread vllm_omni/experimental/fullduplex/openai/serving.py Outdated
Comment thread vllm_omni/metrics/duplex_turn.py
@guozhihao-224
guozhihao-224 force-pushed the frontend/duplex-turn-log-metrics branch from ce11c40 to 381f9c3 Compare September 2, 2026 04:21
@guozhihao-224

Copy link
Copy Markdown
Author

Thanks @amy-why-3459 @Sy0307 @ZacheryAU @shiy1022 — this update is the Phase-1 review follow-up. It does not close #6614 (Related to #6614).

  • Native cancel no longer finalizes as barge_in. _cancel_active_response maps session_close / disconnect* → close, timeout / new_response / input.cancel / response.cancel → cancel; real barge-in stays barge_in. A later close() after that finalize is a no-op.
  • auto_response: stage snapshots that arrive before begin_response are buffered and flushed at begin. arrival_ts is commit / first append (or the first buffered stage_submit_ts), not time.time() at first speak.
  • Collapsed vllm_tpot_ms is token-weighted like the WS sidecar (weight = max(num_tokens_out - 1, 1)).
  • One [OmniTiming] row per turn, including response / turn / reason.
  • Connector TX that landed on the long-lived session aggregator is copied into the turn table (delta vs begin baseline).
  • Docs: duplex stage_gen_time_ms is turn-accumulated, not one chat decode step.

Prometheus / audio TTFP / session rollup stay out of this PR.

Native duplex keeps one long-lived stage-0 request, so chat generate()
never finalizes a turn. Key the aggregator by response_id and log the
same tables as /v1/chat/completions on done, barge-in, close, and abort.

Related to vllm-project#6614 (PR1: log tables only)

Signed-off-by: guozhihao-224 <guozhihaoemail@gmail.com>
Keep barge-in/cancel/close distinct from abort in log tables, and use
stage_submit_ts for wall-clock first-ts so duplex matches chat.

Signed-off-by: guozhihao-224 <guozhihaoemail@gmail.com>
@amy-why-3459 amy-why-3459 added the ready label to trigger buildkite CI label Sep 4, 2026
@amy-why-3459

Copy link
Copy Markdown
Collaborator

Reviewed 166bfe9 against the Phase-1 slice of #6614 and the earlier request-changes thread.

The landing path caught up: _cancel_active_response maps session_close/disconnect* → close and timeout/new_response/input.cancel → cancel; auto_response snapshots are buffered until begin_response; collapsed vllm_tpot_ms is token-weighted; one [OmniTiming] row; session TX is copied in. Do not Fixes #6614. Requesting changes on CI and on the dump operators will actually read.

  1. CUDA L1 is red. buildkite/vllm-omni #14610 failed 2/26 jobs in 16 minutes — that is the core_model and cpu surface, not a long GPU suite. test_abort_emits_table_once_then_close_is_noop stubs abort_async=mocker.AsyncMock() with no return_value. Current _abort does abort_outputs = await self.engine.abort_async(...) or [] and iterates it; a truthy AsyncMock never becomes []. Stub return_value=[]. Re-run the CI command, not only the five files in the body:

    pytest -n 0 -s -v -m "core_model and cpu" --ignore=tests/diffusion --ignore=tests/model_executor

    AMD #11371 is a different pipeline (Simple · Other Test there is a copy of the diffusion job and does not run tests/metrics / tests/entrypoints). Paste the two CUDA node ids after they are green. Do not treat AMD as this PR until then.

  2. The printed tables are not the RFC §7 shape, and they are not the chat tables with an extra identity line.

    RequestE2E / Stage / Transfer titles are [request_id=resp-...]. [OmniTiming] is req=duplex-s....r.stage0 response=resp-... turn=N reason=stop. grep request_id=duplex-s misses the tables. RFC puts the engine id + response + turn + reason on the StageRequestStats title as well. Either merge timing_identity into all three titles, or keep table request_id as the resource id and add response=.

    serving_time_to_first_output_ms is still in the Stage table. You already flagged it as epoch-ms. It is not in STAGE_EXCLUDE and is non-zero, so the zero-filter prints 1,757,xxx,xxx.000 on the first duplex --log-stats dump. Phase 1 does not have to fix the metric; it does have to stop printing it on duplex turns.

    When there is no preprocess_ms (typical duplex), OmniTiming also prints total=Xs engine=Xs. The RFC example does not.

Also: e2e_stage_i_wall_time_ms is first-submit → last-receipt (idle between segments included); stage_gen_time_ms is the summed compute. Say that in metrics.md. The auto_response buffer flushes listen units into the first speak table — RFC Q1 was no e2e for listen-only. on_begin still does not pass arrival_ts, so response_required e2e is begin_response now, not commit. _merge_stage_metric_event is now globally weighted; test_stats.py does not pin it, and the body still says stats.py is unchanged. Gate queue_turn_stage_metrics on is_duplex_resource_request_id so chat completions do not fill duplex_turn_pending.

Signed-off-by: guozhihao-224 <guozhihaoemail@gmail.com>
@guozhihao-224
guozhihao-224 force-pushed the frontend/duplex-turn-log-metrics branch from 166bfe9 to 226b43d Compare September 4, 2026 04:14
@guozhihao-224

Copy link
Copy Markdown
Author

Thanks @amy-why-3459 — addressed on this push.

CI. test_abort_emits_table_once_then_close_is_noop now stubs abort_async(return_value=[]). _abort still leaves request_states so generate() can consume the synthetic abort; the test asserts the turn table is emitted once and a second finalize is a no-op.

Dump titles. Stage / E2E / Transfer titles now include the engine resource request_id plus response= / turn= / reason=, so grep request_id=duplex-s hits the tables. Row keys stay response_id.

TTFO / OmniTiming. Duplex dumps omit serving_time_to_first_output_ms (still an epoch-ms clock; Prometheus audio TTFP stays PR2+). [OmniTiming] skips total= / engine= when there is no preprocess_ms; chat without timing_identity is unchanged.

Also in this pass. Native append stamps arrival_ts before begin_response and on_begin forwards it. Chat completions no longer fill duplex_turn_pending. Token-weighted vllm_tpot_ms is pinned in test_stats.py. Docs record wall vs compute, listen-before-speak attribution to the first speak table, and that listen-only still emits no table.

Local CUDA L1 Other Test: abort / dump / tpot cases passed. Remaining 8 failures (test_serve_cli, transcript helpers, torch profiler excel) are unrelated to this PR.

Keep turn-metrics on the graduated duplex_request_client path after vllm-project#6196.

Co-authored-by: Cursor <cursoragent@cursor.com>

Signed-off-by: guozhihao-224 <guozhihaoemail@gmail.com>
@Sy0307

Sy0307 commented Sep 5, 2026

Copy link
Copy Markdown
Collaborator

RFC #6494 tracks this as Phase 1 under #6895/#6614. The current scope is per-turn log tables with stable reason/timestamp/aggregation semantics; Prometheus, audio TTFP, and session rollups remain Phase 2. Please keep #6614 open until those follow-ups have an owner and a separate implementation path.

@amy-why-3459 amy-why-3459 added ready label to trigger buildkite CI merge-test label to trigger buildkite merge test CI and removed ready label to trigger buildkite CI labels Sep 5, 2026
@amy-why-3459

Copy link
Copy Markdown
Collaborator

@guozhihao-224 CUDA CI is still red on c2d664d. buildkite/vllm-omni #14682 failed 2/35 jobs in 29 minutes: https://buildkite.com/vllm/vllm-omni/builds/14682

GHA / DCO / Intel / NPU / RTD are green. This is the remaining merge gate. Your 226b43d note said the abort stub is fixed, but also that 8 Other Test failures (test_serve_cli, transcript helpers, torch profiler excel) were left as unrelated. Simple · Other Test is one job (-m 'core_model and cpu' --ignore=tests/diffusion --ignore=tests/model_executor). Those failures keep the job red. Please make that surface green, or show the two failed node ids plus the first traceback if they are infra/flake.

Please re-run the CI command, not only the five files in the body:

pytest -n 0 -s -v -m "core_model and cpu" --ignore=tests/diffusion --ignore=tests/model_executor

AMD #11429 (3/19) is a different pipeline. Do not treat it as this PR, and do not use it to explain away the CUDA red. Paste the two CUDA node ids after they are green.

@amy-why-3459

Copy link
Copy Markdown
Collaborator

@guozhihao-224 This PR currently has merge conflicts with main. Could you please update your branch with the latest main, resolve the conflicts, and rerun the relevant checks? Thank you!

…log-metrics

# Conflicts:
#	tests/entrypoints/openai_api/test_duplex_handler.py
#	vllm_omni/entrypoints/client_request_state.py
#	vllm_omni/entrypoints/duplex/protocol.py
#	vllm_omni/entrypoints/duplex/serving.py

Signed-off-by: guozhihao-224 <guozhihaoemail@gmail.com>
finalize_turn_metrics is now safe to call for any aborted request id.
Cleanup of the per-turn aggregator is restricted to native duplex stage
resource ids (mirroring begin_turn_metrics), so ordinary AR aborts no
longer dereference missing turn state on their request states.

Fixes the persistent Buildkite "Simple · Engine&Entrypoints" failures
(test_abort_* failing on SimpleNamespace request states) and adds a
regression test.

Signed-off-by: guozhihao-224 <guozhihaoemail@gmail.com>
@guozhihao-224

Copy link
Copy Markdown
Author

Root-caused the persistent Buildkite failures by reproducing the CI locally (vllm 0.29.0 + core_model and cpu markers):

  • Simple · Engine&Entrypoints failed on tests/entrypoints/test_async_omni.py test_abort_*: _abort calls finalize_duplex_turn_metrics for every aborted request, and DuplexRequestClient.finalize_turn_metrics dereferenced req_state.duplex_turn, which does not exist on plain AR request states.
  • Fixed in f78b259 by guarding finalize_turn_metrics with is_duplex_resource_request_id (mirroring begin_turn_metrics), plus a regression test.
  • Re-ran the group locally: 2721 passed, 0 failed.

Remaining local-only failures (tests/profile/*, tests/helpers/* missing opencc, tests/benchmarks/test_serve_cli.py) are environment gaps that the CI image's [dev] install covers.

@linyueqian linyueqian added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 14, 2026
request_id="",
final_output_type=final_output_type,
)
req_state.duplex_turn_pending.append(

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Collapse pending listen metrics per stage instead of retaining every snapshot.

wall_start = float(req_state.duplex_turn_arrival_ts)
if wall_start <= 0:
wall_start = time.time()
turn = DuplexTurnMetrics(

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Skip per-turn metric collection when log_stats is disabled.

@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: CI is red on this head

@guozhihao-224 required checks failed on 5f33fbfb2069:

Please fix the failure and push again; this note is updated in place when the head goes green or moves.

@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: no human activity for 14 days

@guozhihao-224 this pull request has had no human commit, comment or review since 2026-09-15. Please consider marking this PR as draft until work can resume. The author or a maintainer decides whether to change the PR state.

To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline.

@hsliuustc0106

Copy link
Copy Markdown
Collaborator

@guozhihao-224 this PR is labeled ready + high priority, but CI is failing on the latest commit:

Could you please take a look and push an update to get CI green? Once the checks pass we can proceed with review/merge. Thanks!

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot routing record

Assigned Strict on cursor (cursor-grok-4.6-high) under experiment fleet-strict-cursor-grok46-zcode-glm53flash-5050-c5-z10-20261002.

@hsliuustc0106

Copy link
Copy Markdown
Collaborator

@guozhihao-224 gentle ping — this high-priority PR has seen no updates for 24 days and review feedback is waiting. Could you share a status or ETA when you get a chance? Thanks!

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as timeout (Strict attempt outlived its budget; falling back to direct/cursor/auto).

@vllm-omni-review-bot vllm-omni-review-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Omni ReviewBot review

Changes since the previous review

  • 3 new inline finding(s); 2 finding(s) below.

CI at 5f33fbfb2069 (2026-10-09T21:56:54.462422+00:00): verification incomplete; required-check status is unknown. Observed Buildkite: buildkite/vllm-omni-npu-ci (passed), buildkite/vllm-omni-amd-ci (failed), buildkite/vllm-omni (failed), and 1 more.

Note: The assigned review arm strict/cursor/cursor-grok-4.6-high could not complete this review, so it was produced by the fallback arm direct/cursor/auto. It is excluded from the routing experiment.

Full review analysis

PR description

With --log-stats, native full-duplex now prints the same end-to-end, stage, and transfer tables once per assistant turn. A long-lived stage-0 session request stays open; begin_response / end_response open and finalize a separate aggregator keyed by response_id, including buffered stage snapshots that arrive before the response starts. Chat generate() cleanup is unchanged. Cancel, close, barge-in, and abort each print one table with a distinct reason, and the stage title carries the engine request id plus response, turn, and reason.

Change flow

flowchart LR
  A["[EXISTING] Native duplex append / response"]:::existing
  B["[CHANGED] Session begin and end hooks"]:::changed
  C["[CHANGED] OmniBase stage ingest"]:::changed
  D["[NEW] Per-turn aggregator"]:::new
  E["[CHANGED] log-stats table dump"]:::changed
  A --> B
  A --> C
  C --> D
  B --> D
  D --> E
  classDef existing fill:#e5e7eb,stroke:#6b7280,color:#111827
  classDef changed fill:#fef3c7,stroke:#d97706,color:#451a03,stroke-width:2px
  classDef new fill:#dcfce7,stroke:#16a34a,color:#052e16,stroke-width:2px
  classDef removed fill:#fee2e2,stroke:#dc2626,color:#450a0a,stroke-width:2px
Loading

Findings

  • [P2] Duplex turn metrics are collected when log-stats is off — vllm_omni/entrypoints/omni_base.py:460
    Existing thread: #6892 (comment)
Evidence for Duplex turn metrics are collected when log-stats is off

OmniBase defaults log_stats to false, but _accumulate_duplex_turn_metrics only checks is_duplex_resource_request_id. Every stage-0/TTS snapshot for a duplex resource id is copied into duplex_turn_pending or the open turn aggregator, and begin_turn_metrics still builds that aggregator. --log-stats off therefore pays per-chunk copy and list growth on the native output path and only skips build_and_log_summary. Return before queue_turn_stage_metrics when log_stats is false, and skip begin_turn_metrics the same way.

- **[P2] Pre-response snapshots are retained one-for-one** — `vllm_omni/metrics/duplex_turn.py:160` Existing thread:
Evidence for Pre-response snapshots are retained one-for-one

Before begin_response, queue_turn_stage_metrics appends every StageRequestStats copy to duplex_turn_pending. A listen-only auto-response never opens a turn, so this list grows for the whole session until close() drops request_states. Finalize does not clear it when duplex_turn is still none. Collapse into one merged row per stage with _merge_stage_metric_event as each snapshot arrives so a long listen cannot retain every chunk.


🤖 This review was generated by InferMatrix Copilot, an open-source repo-maintenance agent for PR review, CI debugging and issue triage. Try it on your own repo, and ⭐ star it if it helped!

emitted = finalize_duplex_turn_metrics(turn, reason=reason)
req_state.duplex_turn = None
req_state.duplex_turn_pending = []
req_state.duplex_turn_arrival_ts = None

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Turn finalize wipes the next turn's arrival stamp

Evidence and suggested fix

On a normal auto-response handoff, start_native_append stamps duplex_turn_arrival_ts before the append, then _end_active_response_before_future_model_turn ends the open response and the speak path calls begin_response. finalize_turn_metrics sets duplex_turn_arrival_ts = None after logging the old turn. on_begin then calls mark_turn_arrival, sees an empty stamp, and stores time.time() at speak. The commit/append t0 collected for the next turn is discarded in that same call, so e2e_total_ms starts at begin_response instead of the first append. begin_turn_metrics already clears the stamp it consumes; finalize should leave a stamp that was set after the current turn opened.

elif event_type == "response.cancel":
cancel_reason = "client_cancelled"
else:
cancel_reason = "cancel"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] input.cancel now changes the audio.cancelled reason

Evidence and suggested fix

input.cancel used to pass reason="barge_in" into _cancel_active_response. The else branch now passes "cancel". That string is both mapped to the log finished_reason and copied into audio.cancelled.reason (serving.py sends "reason": reason). _signal_runtime_session for this same branch still signals the engine with "barge_in". Clients that treat audio.cancelled.reason == "barge_in" as user cancel now see cancel, which this PR does not describe as a wire change. Keep the existing event reason and map it only through finished_reason_for_cancel.

from vllm_omni.outputs import OmniRequestOutput
from vllm_omni.outputs.duplex import attach_duplex_output_decision

pytestmark = [pytest.mark.core_model, pytest.mark.cpu]

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Reported L1 commands do not match the CPU guard selectors

Evidence and suggested fix

Changed entrypoint and metrics paths are guarded by .buildkite/cuda/test-ready.yml (the merge pipeline uses the same commands). Simple · Engine&Entrypoints Test runs pytest -sv tests/entrypoints tests/engine -m 'core_model and cpu'. Simple · Other Test runs pytest -sv tests/ -m 'core_model and cpu' --ignore=tests/diffusion --ignore=tests/model_executor --ignore=tests/entrypoints --ignore=tests/engine, which includes tests/metrics. The PR Test Result reports only tests/metrics/test_duplex_turn_metrics.py, tests/metrics/test_stats.py, tests/entrypoints/test_async_omni_duplex.py, and two test_duplex_handler.py cases. It states that L2 protocol smoke and the auto_response L3 rerun were skipped, not that these CPU marker jobs were skipped. Re-run those two selectors or record that gap in the test result.

@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: finding feedback

[p1] Reported L1 commands do not match the CPU guard selectors — tests/metrics/test_duplex_turn_metrics.py:30

See the review for details. If you are the PR author and disagree, react 👎 here; the maintainer will see your disagreement.

@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: finding feedback

[p1] Turn finalize wipes the next turn's arrival stamp — vllm_omni/entrypoints/duplex_request_client.py:356

See the review for details. If you are the PR author and disagree, react 👎 here; the maintainer will see your disagreement.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

benchmark/profiler/metrics/logger Codes related to benchmarks, profiler, metrics and logger system frontend code related to entrypoint high priority high priority issue, needs to be done asap merge-test label to trigger buildkite merge test CI ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants