fix(stt): thread confidence thresholds into faster-whisper's own gate (#74178) - #74193
PRATHAMESH75 wants to merge 1 commit into
Conversation
…NousResearch#74178) build_local_transcribe_kwargs read stt.local.no_speech_prob_threshold / stt.local.logprob_threshold only for Hermes' post-filter (_is_hallucinated_segment). faster-whisper's model.transcribe() never received them, so its internal defaults (no_speech_threshold=0.6, log_prob_threshold=-1.0) always applied and silently dropped low-confidence segments before they reached the post-filter — making those config knobs dead for the first gate. Non-English speech decodes at a lower avg_logprob, so the English-tuned defaults discard whole utterances (empty transcript despite correct capture and language detection). Map the same config values through to model.transcribe() so both gates stay in sync and the knobs work. Defaults are unchanged, so behavior is identical unless a user tunes them. Fixes NousResearch#74178
|
Thanks for the focused, well-covered threshold propagation fix. The current-head premise is real: Problems
Suggested changes
The two added kwargs otherwise fit the pinned faster-whisper API and the shared kwargs helper reaches both local transcribe calls ( Automated hermes-sweeper review. |
|
Good catch on the closing keyword. This PR only threads the confidence thresholds into faster-whisper's own gate; it intentionally leaves the default-VAD path ( I've switched the PR body to a non-closing |
|
Merged via #77516 — thank you @PRATHAMESH75. Your commit was cherry-picked, so you remain the author in git history. Verified before merging: both kwargs checked against faster-whisper's own source (they gate segment retention exactly as you described), the premise re-confirmed on current main ( |
What & why
build_local_transcribe_kwargs()readsstt.local.no_speech_prob_thresholdandstt.local.logprob_thresholdfrom config, but only uses them in Hermes' ownpost-filter (
_is_hallucinated_segmentvia_join_confident_segments). It neverpassed them to faster-whisper's
model.transcribe(), so the library's internaldefaults (
no_speech_threshold=0.6,log_prob_threshold=-1.0) always applied andsilently dropped low-confidence segments before they ever reached our
post-filter. That made those two config knobs dead for the first (and decisive) gate.
Non-English speech decodes at a lower
avg_logprob, so the English-tuned defaultsdiscard whole utterances — the caller gets an empty transcript even though the mic
captured audio and language detection succeeded (
lang=zh, prob=1.000).Fix
Map the same config values through to
model.transcribe()so faster-whisper'sinternal gate and Hermes' post-filter use identical thresholds. Reuses the existing
_confidence_thresholds()helper (same parsing + fallback the post-filter alreadyuses), so there's a single source of truth. Defaults are unchanged (0.6 / -1.0),
so behavior is byte-for-byte identical unless a user tunes them — this only makes
the previously-dead knobs actually reach the model.
Users hitting this on non-English speech can now relax the gate, e.g.:
Tests
tests/tools/test_stt_silence_hallucinations.py:test_confidence_thresholds_default_to_faster_whisper_values— defaults preserved.test_confidence_thresholds_configurable_reach_model_gate— tuned values reach the kwargs.test_confidence_thresholds_garbage_falls_back— bad config falls back to defaults.test_hardened_kwargs_reach_modelto assert both thresholds reachmodel.transcribe().Note: 3 pre-existing failures in
tests/tools/test_transcription_tools.py(
test_config_device_and_compute_type_passed_to_whisper,test_config_defaults_to_auto_when_not_set,test_cublas_status_not_supported_retries_on_cpu) reproduce on a pristineupstream/maincheckout with none of this change — they assert the device stringpassed to the
WhisperModelconstructor (auto/auto) and are environment-specific(local CUDA-less resolution), unrelated to this diff, which only adds
transcribe()kwargs.Scope
This narrows the confidence-threshold gate only; it deliberately does not touch the default-VAD path (
tools/transcription_tools.py:1533), which is the other half of #74178 (VAD can strip all audio before decoding). Referencing #74178 with a non-closing keyword so that remaining VAD work stays tracked after this merges.Refs #74178