Skip to content

[Test][Core] Pin the server-VAD endpoint rules across resets and append sizes - #7653

Closed
twu3202 wants to merge 3 commits into
vllm-project:mainfrom
twu3202:duplex-vad-clear-speech-resets-silence
Closed

twu3202 wants to merge 3 commits into
vllm-project:mainfrom
twu3202:duplex-vad-clear-speech-resets-silence

Conversation

@twu3202

@twu3202 twu3202 commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Issue 4 of #7636.

#7585 landed the same rule this PR was opened for: a frame at or above the
activation threshold cancels a pending endpoint. So vad.py here is now main's,
and what is left is coverage for the rules around it.

Three gaps on main:

  • reset() mid-turn. It keeps the session clock, so the next turn's
    speech_start_ms still refers to the session timeline, but it must not
    inherit a silence candidate from the turn it replaced.
  • Whether an endpoint moves when the same audio arrives in different append
    sizes. tests/entrypoints/openai_api/test_server_vad.py checked that for the
    serving-side pipeline; [Core][Frontend] Unified Full-duplex Framework #7413 removed that file along with the policy it
    tested. The engine scores whole appends, so the endpoints have to come out the
    same frame by frame, in 200 ms appends, in odd sizes, and at 24 kHz.
  • The endpoint rules driven through the session runner. The server-VAD cases
    there hand the runner a fixed TurnDetectionResult, so nothing composes the
    real rules with what the session does on an endpoint.

The parity table also gets a row for clear speech inside the silence timer, so
that rule is pinned against the policy this detector replaced, and its
provenance comment now says 99ff4f307~1; it pointed at e2d2617f, which does
not exist in this repository.

Test Plan

vLLM Version: 0.29.0

vLLM-Omni Commit: 242c3cd

  1. pytest -q tests/engine/duplex/test_vad_backend.py tests/engine/duplex/test_session_runner.py -m 'core_model and cpu'
  2. The same two files with vllm_omni/engine/duplex/vad.py restored to
    fa639b889, the version before [Frontend] Add a shared realtime web UI for MiniCPM-o and Qwen3-Omni (#7222) #7585, to check the new cases can fail.
  3. pytest -q tests/entrypoints tests/engine -m 'core_model and cpu', the
    Simple · Engine&Entrypoints Test selection.

Test Result

  1. 67 passed.
  2. 9 failed, 58 passed. Six are added here: the new parity row, the reset case
    that keeps the endpoint after the pause, the cut-up utterance's absolute
    endpoint, and the three runner cases. The other three are the cases [Frontend] Add a shared realtime web UI for MiniCPM-o and Qwen3-Omni (#7222) #7585
    added. test_endpoints_do_not_depend_on_where_the_audio_is_cut passes on
    both, by design: it compares cuttings against each other, not against an
    expected endpoint.
  3. 3016 passed, 4 skipped, 1 xfailed, 0 failed. On 242c3cdde without this
    branch: 3005 passed, with the same skips. The 11 are the cases added here.

SileroStreamingVAD adds to the silence clock on every frame once a
candidate exists, frames at or above the activation threshold included. A
short pause mid-utterance therefore keeps counting through the speech that
follows it, and the next quiet frame commits the turn.

tests/assets/minicpmo_4_5/response_required_16k.wav shows what that costs.
Scored with the pinned Silero v6.2 ONNX graph, no stretch of frames below
the activation threshold after speech starts is longer than 224 ms, so at
silence_duration_ms=500 the clip contains no endpoint at all. Driven
through a live duplex server in 200 ms appends, main still reports
speech_stopped at 5344 ms; after this change it reports none. Under
server_vad that report is should_commit, so main ends the user's turn
mid-word.

The serving-side ThresholdEndpointPolicy this detector was ported from
reset _silence_samples whenever a frame reached the activation threshold,
and so does Silero's own VADIterator, which clears temp_end. Restore that
rule. Frames between the two thresholds are unchanged: they still keep the
clock running and still cannot close a turn on their own. That band is thin
in practice -- 3 frames across 25.6 s of speech in soft_interrupt_16k.wav,
against 52 clear frames landing inside running silence candidates -- so
merging the top two bands was acting almost entirely on clear speech.

Checked differentially against the replaced policy, extracted at
99ff4f3~1: over continuous thresholds, paddings, silence and min-speech
durations, main disagrees on 16% of 25000 random sequences and this change
on 0 of 50000.

The parity case and test that pinned the old behaviour used 0.45, a
hysteresis-band frame rather than clear speech; both keep their
expectations and are renamed to say which band they cover. The golden
table's provenance is corrected to 99ff4f3~1, the commit before the
serving module was deleted; the ref it named is not an object in this
repository.

Part of vllm-project#7636 (issue 4).

Co-authored-by: Claude <noreply@anthropic.com>
Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/engine_orchestration.md, docs/design/module/observability.md.

Module owners: @fake0fan @tzhouam @lishunyang12

Routing: @fake0fan via module of the changed files, semantic router, CODEOWNERS; @tzhouam via module of the changed files, semantic router, CODEOWNERS; @lishunyang12 via module of the changed files

@twu3202, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 16, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Resolved as of 9aab70fd3040: the high-priority or low-quality signal noted on an earlier commit no longer applies.

@twu3202

twu3202 commented Sep 16, 2026

Copy link
Copy Markdown
Contributor Author

Self-review done.

  • The early continue only skips the silence accumulator. A frame that is already _speech_active has nothing else to do in that loop, and candidate state is zeroed at activation.
  • Frames in the hysteresis band behave exactly as before. The renamed test still pins that.
  • The offline traces come from the pinned Silero v6.2 ONNX graph, and the live-server numbers were taken on this commit.
  • tests/engine and tests/entrypoints at core_model and cpu pass on this branch.

@hsliuustc0106 hsliuustc0106 added bug Something isn't working core related to core module: cache, scheduler, engine, worker, modelrunner labels Sep 16, 2026
twu3202 and others added 2 commits September 17, 2026 13:43
…s append sizes

Two gaps in the coverage for the clear-speech reset.

The detector tests score one frame per call, and the runner's server-VAD tests
feed it a scripted `TurnDetectionResult`, so nothing ran the endpoint rules the
way a session does. `open_harness` now takes a `vad_backend_provider`, and two
cases drive a scripted Silero backend through the real detector and runner: with
`auto_response` off a 192 ms or 352 ms pause mid-utterance must not commit, and
with it on the utterance must stay one speech segment. Both fail on main.

vllm-project#7413 also dropped the chunk-boundary check in
`tests/entrypoints/openai_api/test_server_vad.py`. It is back, against
`SileroStreamingVAD`: 200 ms appends, odd sizes and 24 kHz PCM16, with and
without a mid-frame reset, must give the same endpoints as frame-by-frame input.
That holds on main too; the case next to it checks that the sequence does
exercise the reset.

Co-authored-by: Claude <noreply@anthropic.com>
Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
main already carries the endpoint rule through vllm-project#7585, so keep main's vad.py and
the tests it added, and drop the local duplicate of its cancel case.

Co-Authored-By: Claude <noreply@anthropic.com>
Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
@twu3202 twu3202 changed the title [Bugfix][Core] Let clear speech cancel a pending server-VAD endpoint [Test][Core] Pin the server-VAD endpoint rules across resets and append sizes Sep 18, 2026
@twu3202

twu3202 commented Sep 18, 2026

Copy link
Copy Markdown
Contributor Author

#7585 landed the same rule while this was open, so I merged main in and dropped the fix: vad.py here is now main's, and I removed the case that duplicates the one #7585 added.

What is left is coverage: the reset() interaction, endpoints not moving with the append size (what #7413 dropped with test_server_vad.py), and the rules driven through the session runner. Retitled and rewrote the body for that.

@twu3202

twu3202 commented Sep 18, 2026

Copy link
Copy Markdown
Contributor Author

Closing as suggested — the rule itself is in main via #7585. The coverage here (the reset() interaction, endpoints not moving with the append size, and the rules through the session runner) is not, so if anyone wants it later it can come back as a standalone test PR.

@twu3202 twu3202 closed this Sep 18, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working core related to core module: cache, scheduler, engine, worker, modelrunner

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants