Repository navigation
Conversation
SileroStreamingVAD adds to the silence clock on every frame once a candidate exists, frames at or above the activation threshold included. A short pause mid-utterance therefore keeps counting through the speech that follows it, and the next quiet frame commits the turn. tests/assets/minicpmo_4_5/response_required_16k.wav shows what that costs. Scored with the pinned Silero v6.2 ONNX graph, no stretch of frames below the activation threshold after speech starts is longer than 224 ms, so at silence_duration_ms=500 the clip contains no endpoint at all. Driven through a live duplex server in 200 ms appends, main still reports speech_stopped at 5344 ms; after this change it reports none. Under server_vad that report is should_commit, so main ends the user's turn mid-word. The serving-side ThresholdEndpointPolicy this detector was ported from reset _silence_samples whenever a frame reached the activation threshold, and so does Silero's own VADIterator, which clears temp_end. Restore that rule. Frames between the two thresholds are unchanged: they still keep the clock running and still cannot close a turn on their own. That band is thin in practice -- 3 frames across 25.6 s of speech in soft_interrupt_16k.wav, against 52 clear frames landing inside running silence candidates -- so merging the top two bands was acting almost entirely on clear speech. Checked differentially against the replaced policy, extracted at 99ff4f3~1: over continuous thresholds, paddings, silence and min-speech durations, main disagrees on 16% of 25000 random sequences and this change on 0 of 50000. The parity case and test that pinned the old behaviour used 0.45, a hysteresis-band frame rather than clear speech; both keep their expectations and are renamed to say which band they cover. The golden table's provenance is corrected to 99ff4f3~1, the commit before the serving module was deleted; the ref it named is not an object in this repository. Part of vllm-project#7636 (issue 4). Co-authored-by: Claude <noreply@anthropic.com> Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
|
This PR appears to belong to: docs/design/module/engine_orchestration.md, docs/design/module/observability.md. Module owners: @fake0fan @tzhouam @lishunyang12 Routing: @fake0fan via module of the changed files, semantic router, CODEOWNERS; @tzhouam via module of the changed files, semantic router, CODEOWNERS; @lishunyang12 via module of the changed files @twu3202, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Omni ReviewBot triage noteResolved as of |
|
Self-review done.
|
…s append sizes Two gaps in the coverage for the clear-speech reset. The detector tests score one frame per call, and the runner's server-VAD tests feed it a scripted `TurnDetectionResult`, so nothing ran the endpoint rules the way a session does. `open_harness` now takes a `vad_backend_provider`, and two cases drive a scripted Silero backend through the real detector and runner: with `auto_response` off a 192 ms or 352 ms pause mid-utterance must not commit, and with it on the utterance must stay one speech segment. Both fail on main. vllm-project#7413 also dropped the chunk-boundary check in `tests/entrypoints/openai_api/test_server_vad.py`. It is back, against `SileroStreamingVAD`: 200 ms appends, odd sizes and 24 kHz PCM16, with and without a mid-frame reset, must give the same endpoints as frame-by-frame input. That holds on main too; the case next to it checks that the sequence does exercise the reset. Co-authored-by: Claude <noreply@anthropic.com> Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
main already carries the endpoint rule through vllm-project#7585, so keep main's vad.py and the tests it added, and drop the local duplicate of its cancel case. Co-Authored-By: Claude <noreply@anthropic.com> Signed-off-by: Tianyao Wu <rayroy31@gmail.com>
|
#7585 landed the same rule while this was open, so I merged main in and dropped the fix: What is left is coverage: the |
|
Closing as suggested — the rule itself is in main via #7585. The coverage here (the |
Purpose
Issue 4 of #7636.
#7585 landed the same rule this PR was opened for: a frame at or above the
activation threshold cancels a pending endpoint. So
vad.pyhere is now main's,and what is left is coverage for the rules around it.
Three gaps on main:
reset()mid-turn. It keeps the session clock, so the next turn'sspeech_start_msstill refers to the session timeline, but it must notinherit a silence candidate from the turn it replaced.
sizes.
tests/entrypoints/openai_api/test_server_vad.pychecked that for theserving-side pipeline; [Core][Frontend] Unified Full-duplex Framework #7413 removed that file along with the policy it
tested. The engine scores whole appends, so the endpoints have to come out the
same frame by frame, in 200 ms appends, in odd sizes, and at 24 kHz.
there hand the runner a fixed
TurnDetectionResult, so nothing composes thereal rules with what the session does on an endpoint.
The parity table also gets a row for clear speech inside the silence timer, so
that rule is pinned against the policy this detector replaced, and its
provenance comment now says
99ff4f307~1; it pointed ate2d2617f, which doesnot exist in this repository.
Test Plan
vLLM Version: 0.29.0
vLLM-Omni Commit: 242c3cd
pytest -q tests/engine/duplex/test_vad_backend.py tests/engine/duplex/test_session_runner.py -m 'core_model and cpu'vllm_omni/engine/duplex/vad.pyrestored tofa639b889, the version before [Frontend] Add a shared realtime web UI for MiniCPM-o and Qwen3-Omni (#7222) #7585, to check the new cases can fail.pytest -q tests/entrypoints tests/engine -m 'core_model and cpu', theSimple · Engine&Entrypoints Testselection.Test Result
that keeps the endpoint after the pause, the cut-up utterance's absolute
endpoint, and the three runner cases. The other three are the cases [Frontend] Add a shared realtime web UI for MiniCPM-o and Qwen3-Omni (#7222) #7585
added.
test_endpoints_do_not_depend_on_where_the_audio_is_cutpasses onboth, by design: it compares cuttings against each other, not against an
expected endpoint.
242c3cddewithout thisbranch: 3005 passed, with the same skips. The 11 are the cases added here.