Align the silence scale range between Validate() and ScaleSilence - #3747
Conversation
After k2-fsa#3745, ScaleSilence accepts [0.01, 2] but OfflineTtsConfig::Validate() still only rejects silence_scale < 0.001. Values in [0.001, 0.01) and above 2 therefore pass validation and are then silently skipped at run time, so the user asks for scaling and gets unscaled audio. Use one pair of constants for both, document the range in the --tts-silence-scale help text, and raise the upper bound from 2 to 10. Rejecting in Validate() also makes bad input fail before synthesis rather than after it. The check in ScaleSilence stays, because callers of the Generate API pass GeneratedAudioConfig::silence_scale directly and bypass Validate().
📝 WalkthroughWalkthroughThe TTS silence scale option now accepts values from 0.01 through 10.0. Shared bounds are used by audio scaling and configuration validation, and the command-line help and error messages describe the updated range. ChangesTTS silence scale range
Estimated code review effort: 2 (Simple) | ~10 minutes Possibly related PRs
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Code Review
This pull request increases the maximum allowed silence scale for offline text-to-speech from 2.0 to 10.0, introducing kMinSilenceScale and kMaxSilenceScale constants to enforce this range during validation and scaling. The command-line option description is also updated to reflect this new range. The reviewer suggests moving these constants to the header file offline-tts.h to improve maintainability and allow other parts of the codebase to access them.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
| static constexpr float kMinSilenceScale = 0.01f; | ||
| static constexpr float kMaxSilenceScale = 10.0f; |
There was a problem hiding this comment.
To improve maintainability and allow other parts of the codebase (or external language bindings) to access these validation limits, consider moving kMinSilenceScale and kMaxSilenceScale to the header file offline-tts.h (for example, as public static constexpr members of OfflineTtsConfig or within the sherpa_onnx namespace).
Currently, defining them as static constexpr in the .cc file hides them from other translation units, making it harder to reuse them for validation or documentation elsewhere.
There was a problem hiding this comment.
🧹 Nitpick comments (1)
sherpa-onnx/csrc/offline-tts.cc (1)
41-41: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low valueConsider using
1.0ffor explicit float comparison.
scale == 1compares afloatagainst anintliteral. While1is exactly representable and the implicit conversion is harmless, using1.0fis more idiomatic and avoids relying on implicit promotion.♻️ Optional style tweak
- if (scale == 1) { + if (scale == 1.0f) {🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@sherpa-onnx/csrc/offline-tts.cc` at line 41, In the scale comparison within the relevant TTS function, replace the integer literal 1 in `if (scale == 1)` with the explicit float literal `1.0f`.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Nitpick comments:
In `@sherpa-onnx/csrc/offline-tts.cc`:
- Line 41: In the scale comparison within the relevant TTS function, replace the
integer literal 1 in `if (scale == 1)` with the explicit float literal `1.0f`.
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro
Run ID: b15503ac-3697-44a6-9064-9db2ae84c274
📒 Files selected for processing (1)
sherpa-onnx/csrc/offline-tts.cc
csukuangfj
left a comment
There was a problem hiding this comment.
Thanks for your contribution!
Follow-up to #3745, from a review comment on it.
After #3745,
ScaleSilenceaccepts[0.01, 2], butOfflineTtsConfig::Validate()still only rejects
silence_scale < 0.001. So--tts-silence-scale=0.005or=5passes validation and is then silently skipped at run time: the user asks for
scaling and gets unscaled audio back, with only a log line.
This puts both checks on one pair of constants, documents the range in the
--tts-silence-scalehelp text, and makes bad input fail before synthesis ratherthan after it. The check inside
ScaleSilencestays, because callers of theGenerateAPI passGeneratedAudioConfig::silence_scaledirectly and bypassValidate().One question: should the upper bound be 2, or higher?
I raised it to 10 here. Happy to change it back to 2 — it is a one-line
revert, and it is your call.
The reason to raise it: nothing breaks until far higher. Overflow of the
int32_tconversion needspause_length * scale > 2^31, which for a 1 s pause at24 kHz means a scale around 89000, and for a 5 s pause around 17900. At
scale = 10a 1 s pause becomes 10 s, which is 1 MB of float32. Only NaN,infinity and negative values are actually dangerous, and those are rejected
regardless of where the ceiling sits.
2also rules out reasonable uses of scaling a pause up — long dramatic pausesin an audiobook, say. For comparison,
speedhas no upper bound at all, onlyspeed <= 0is rejected.If you prefer to keep the tighter limit, say so and I will set both constants
back to 2.
Verified
Built and run against
kokoro-multi-lang-v1_0.--tts-silence-scale0.0099Validate(), then silently skippedValidate()0.012.02.001Validate(), then silently skipped10Validate(), then silently skipped10.001Validate(), then silently skippedValidate()0,-1,nan,inf,1e9Validate()catches0/-1;nan/inf/1e9skipped inScaleSilenceValidate()scale = 10produces 19.25 s from a 4.17 s clip, and the count of samples abovethe 0.01 silence threshold is unchanged (
-1), so the added audio is silence.Checked with
clang-format --dry-run --Werror; no changes.Summary by CodeRabbit