Skip to content

fix(stt): thread confidence thresholds into faster-whisper's own gate (#74178) - #74193

Closed
PRATHAMESH75 wants to merge 1 commit into
NousResearch:mainfrom
PRATHAMESH75:fix/local-whisper-pass-thresholds
Closed

PRATHAMESH75 wants to merge 1 commit into
NousResearch:mainfrom
PRATHAMESH75:fix/local-whisper-pass-thresholds

Conversation

@PRATHAMESH75

@PRATHAMESH75 PRATHAMESH75 commented Jul 29, 2026 •

Copy link
Copy Markdown

What & why

build_local_transcribe_kwargs() reads stt.local.no_speech_prob_threshold and
stt.local.logprob_threshold from config, but only uses them in Hermes' own
post-filter (_is_hallucinated_segment via _join_confident_segments). It never
passed them to faster-whisper's model.transcribe(), so the library's internal
defaults (no_speech_threshold=0.6, log_prob_threshold=-1.0) always applied and
silently dropped low-confidence segments before they ever reached our
post-filter. That made those two config knobs dead for the first (and decisive) gate.

Non-English speech decodes at a lower avg_logprob, so the English-tuned defaults
discard whole utterances — the caller gets an empty transcript even though the mic
captured audio and language detection succeeded (lang=zh, prob=1.000).

Fix

Map the same config values through to model.transcribe() so faster-whisper's
internal gate and Hermes' post-filter use identical thresholds. Reuses the existing
_confidence_thresholds() helper (same parsing + fallback the post-filter already
uses), so there's a single source of truth. Defaults are unchanged (0.6 / -1.0),
so behavior is byte-for-byte identical unless a user tunes them
— this only makes
the previously-dead knobs actually reach the model.

Users hitting this on non-English speech can now relax the gate, e.g.:

hermes config set stt.local.no_speech_prob_threshold 0.9
hermes config set stt.local.logprob_threshold -2.0

Tests

tests/tools/test_stt_silence_hallucinations.py:

  • test_confidence_thresholds_default_to_faster_whisper_values — defaults preserved.
  • test_confidence_thresholds_configurable_reach_model_gate — tuned values reach the kwargs.
  • test_confidence_thresholds_garbage_falls_back — bad config falls back to defaults.
  • Extended test_hardened_kwargs_reach_model to assert both thresholds reach model.transcribe().
tests/tools/test_stt_silence_hallucinations.py  21 passed

Note: 3 pre-existing failures in tests/tools/test_transcription_tools.py
(test_config_device_and_compute_type_passed_to_whisper,
test_config_defaults_to_auto_when_not_set,
test_cublas_status_not_supported_retries_on_cpu) reproduce on a pristine
upstream/main checkout with none of this change — they assert the device string
passed to the WhisperModel constructor (auto/auto) and are environment-specific
(local CUDA-less resolution), unrelated to this diff, which only adds
transcribe() kwargs.

Scope

This narrows the confidence-threshold gate only; it deliberately does not touch the default-VAD path (tools/transcription_tools.py:1533), which is the other half of #74178 (VAD can strip all audio before decoding). Referencing #74178 with a non-closing keyword so that remaining VAD work stays tracked after this merges.

Refs #74178

…NousResearch#74178)

build_local_transcribe_kwargs read stt.local.no_speech_prob_threshold /
stt.local.logprob_threshold only for Hermes' post-filter
(_is_hallucinated_segment). faster-whisper's model.transcribe() never
received them, so its internal defaults (no_speech_threshold=0.6,
log_prob_threshold=-1.0) always applied and silently dropped
low-confidence segments before they reached the post-filter — making
those config knobs dead for the first gate.

Non-English speech decodes at a lower avg_logprob, so the English-tuned
defaults discard whole utterances (empty transcript despite correct
capture and language detection). Map the same config values through to
model.transcribe() so both gates stay in sync and the knobs work.
Defaults are unchanged, so behavior is identical unless a user tunes them.

Fixes NousResearch#74178
@alt-glitch alt-glitch added type/bug Something isn't working P2 Medium — degraded but workaround exists tool/tts Text-to-speech and transcription comp/cli CLI entry point, hermes_cli/, setup wizard labels Jul 29, 2026
@teknium1

Copy link
Copy Markdown
Collaborator

Thanks for the focused, well-covered threshold propagation fix. The current-head premise is real: build_local_transcribe_kwargs() returns without these keys at tools/transcription_tools.py:1513-1552, while _confidence_thresholds() is currently used only by the post-filter at tools/transcription_tools.py:1555-1605.

Problems

  • Fixes #74178 is too broad. That issue reports a primary case where the default VAD removes all audio before decoding; current main still enables VAD at tools/transcription_tools.py:1533, and this PR intentionally does not change that path. Merging with the closing keyword would close remaining work.

Suggested changes

The two added kwargs otherwise fit the pinned faster-whisper API and the shared kwargs helper reaches both local transcribe calls (tools/transcription_tools.py:1642-1665).

Automated hermes-sweeper review.

@teknium1 teknium1 added sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades sweeper:blast-moderate Sweeper blast radius: moderate — a subsystem or single platform type/perf Performance improvement or optimization labels Jul 30, 2026
@PRATHAMESH75

Copy link
Copy Markdown
Author

Good catch on the closing keyword. This PR only threads the confidence thresholds into faster-whisper's own gate; it intentionally leaves the default-VAD path (tools/transcription_tools.py:1533) untouched, which is the other, primary half of #74178 (VAD stripping all audio before decode). Merging with Fixes would have closed that remaining work.

I've switched the PR body to a non-closing Refs #74178 and added a Scope section spelling out exactly which half this addresses, so the unresolved VAD behavior stays tracked. No code change — the threshold propagation and its tests are unchanged.

@kshitijk4poor

Copy link
Copy Markdown

Merged via #77516 — thank you @PRATHAMESH75. Your commit was cherry-picked, so you remain the author in git history.

Verified before merging: both kwargs checked against faster-whisper's own source (they gate segment retention exactly as you described), the premise re-confirmed on current main (build_local_transcribe_kwargs still omitted both), and a mutation check proved your 4 wiring tests fail without the production change. The fix makes stt.local.no_speech_prob_threshold / stt.local.logprob_threshold work end-to-end for the first time — non-English speech was the real casualty of the hardcoded internal gate. No changes to your code; the salvage exists only because your fork's base was too old for current required CI checks to run.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/cli CLI entry point, hermes_cli/, setup wizard P2 Medium — degraded but workaround exists sweeper:blast-moderate Sweeper blast radius: moderate — a subsystem or single platform sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades tool/tts Text-to-speech and transcription type/bug Something isn't working type/perf Performance improvement or optimization

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants