Skip to content

Fix ScaleSilence copying past the pause interval when scale > 1 - #3744

Merged
csukuangfj merged 1 commit into
k2-fsa:masterfrom
ekenberg:pr/fix-scale-silence
Jul 10, 2026
Merged

csukuangfj merged 1 commit into
k2-fsa:masterfrom
ekenberg:pr/fix-scale-silence

Conversation

@ekenberg

@ekenberg ekenberg commented Jul 10, 2026 •

Copy link
Copy Markdown
Contributor

GeneratedAudio::ScaleSilence() lengthens each detected pause by copying
n = (interval.end - interval.start) * scale samples starting at interval.start:

int32_t n = static_cast<int32_t>((interval.end - interval.start) * scale);

ans.samples.insert(ans.samples.end(), samples.begin() + interval.start,
                   samples.begin() + interval.start + n);

When scale > 1, n exceeds the length of the interval, so the copy runs past
the end of the pause. Three consequences:

  1. The following word's onset is duplicated into the pause. Audible as a
    stutter after every pause.
  2. Out-of-bounds read on a trailing pause. If the last interval runs to the
    end of the buffer then interval.end == num_samples, and
    samples.begin() + interval.start + n is past samples.end().
  3. The pause is not lengthened at all in that trailing case — the requested
    scaling silently does not happen.

Copying the pause once and padding with zeros gives the intended behaviour. For
scale <= 1 nothing changes, and scale == 1 still returns early before
reaching this code.

Measured

Built at 97293f0, Kokoro kokoro-multi-lang-v1_0, --sid=10, text
"First sentence here. Second sentence follows. Third one ends it."

Same binary before and after the patch, --tts-silence-scale=2.0, compared
against the --tts-silence-scale=1.0 output as reference (scale 1 returns early,
so it is unscaled ground truth). Counting samples above the same 0.01 threshold
ScaleSilence itself uses to detect silence:

Output samples > 0.01 vs reference
reference, scale=1.0 58996 —
unpatched, scale=2.0 87375 +28379
patched, scale=2.0 59098 +102

The unpatched build invents 28379 speech-like samples — 1.18 s of audio that
the reference does not contain. That is duplicated speech pasted into the pauses.
The patched build adds silence only; the +102 residue is boundary rounding.

Applying the same pause detector to the outputs is another way to see it: the
reference has 4 pauses and the patched output has 4, but the unpatched output
reports 7 — the duplicated words split each pause into two.

The out-of-bounds read (2) follows by construction from
interval.end == num_samples; it was not separately exercised under a sanitizer.

Reproducing

sherpa-onnx-offline-tts \
  --kokoro-model=model.onnx --kokoro-voices=voices.bin \
  --kokoro-tokens=tokens.txt --kokoro-data-dir=espeak-ng-data \
  --kokoro-lexicon=lexicon-us-en.txt \
  --sid=10 --tts-silence-scale=2.0 \
  --output-filename=out.wav \
  "First sentence here. Second sentence follows. Third one ends it."

The onset of each sentence is audibly repeated inside the pause preceding it.


Checked with clang-format --dry-run --Werror using the repo's own
.clang-format; it reports no changes.

Summary by CodeRabbit

  • Bug Fixes
    • Fixed text-to-speech pause scaling when increasing silence duration.
    • Prevented speech audio from being incorrectly included in extended pauses.
    • Ensured longer pauses are filled with silence for cleaner, more accurate audio output.

GeneratedAudio::ScaleSilence() lengthens each detected pause by copying
n = (interval.end - interval.start) * scale samples starting at
interval.start.

When scale > 1 that is n > interval.end - interval.start, so the copy
runs past the end of the pause and into the audio that follows it. Two
consequences:

- The beginning of the next word is duplicated, which is audible as a
  stutter after every pause.
- On the last interval, interval.end == num_samples, so
  samples.begin() + interval.start + n is past samples.end() and the
  insert reads out of bounds.

Copy the pause once and pad with zeros to the requested length instead.
For scale <= 1 the behaviour is unchanged, and scale == 1 still returns
early before reaching this code.
@dosubot dosubot Bot added the size:S This PR changes 10-29 lines, ignoring generated files. label Jul 10, 2026
@coderabbitai

coderabbitai Bot commented Jul 10, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 6d19f2af-93fc-4565-a843-44955a54c941

📥 Commits

Reviewing files that changed from the base of the PR and between 97293f0 and 94d2e62.

📒 Files selected for processing (1)
  • sherpa-onnx/csrc/offline-tts.cc

📝 Walkthrough

Walkthrough

GeneratedAudio::ScaleSilence now separates original and scaled interval lengths, preventing out-of-bounds reads when enlarging silence intervals by appending zero-valued samples.

Changes

Silence scaling

Layer / File(s) Summary
Safe silence interval extension
sherpa-onnx/csrc/offline-tts.cc
ScaleSilence inserts only samples within the original interval and appends zeros when the scaled silence is longer than the source interval.

Estimated code review effort: 2 (Simple) | ~10 minutes

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly summarizes the main fix in ScaleSilence for pause scaling above 1.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request updates the silence scaling logic in GeneratedAudio::ScaleSilence to prevent out-of-bounds reads and audio artifacts when the scale factor is greater than 1 by copying the pause once and appending silence. The reviewer identified a potential issue where a negative or NaN scale factor could result in a negative sample count, leading to undefined behavior, and suggested clamping the sample count to a minimum of zero.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment on lines +82 to +85
int32_t len = interval.end - interval.start;
int32_t n = static_cast<int32_t>(len * scale);

if (n <= len) {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

If scale is negative or NaN, n can be negative (or a large negative value like INT_MIN). This would make n <= len true, and the iterator arithmetic samples.begin() + interval.start + n would point to a memory location before samples.begin() + interval.start, leading to undefined behavior or a crash during std::vector::insert.

To prevent this, we should clamp n to be at least 0.

    int32_t len = interval.end - interval.start;
    int32_t n = static_cast<int32_t>(len * scale);
    if (n < 0) {
      n = 0;
    }

    if (n <= len) {

@csukuangfj csukuangfj left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for your contribution!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:S This PR changes 10-29 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants