Skip to content

Add ten-vad Rust API example to remove silences from a file - #3778

Merged
csukuangfj merged 2 commits into
k2-fsa:masterfrom
sportiz91:rust-ten-vad-example
Jul 23, 2026
Merged

csukuangfj merged 2 commits into
k2-fsa:masterfrom
sportiz91:rust-ten-vad-example

Conversation

@sportiz91

@sportiz91 sportiz91 commented Jul 21, 2026 •

Copy link
Copy Markdown
Contributor

Part of #3210 ("Add ten-vad example to remove silences from a file").

This ports rust-api-examples/examples/silero_vad_remove_silence.rs to ten-vad, following the existing Java (TenVadRemoveSilence.java) and Pascal (remove_silence_ten_vad.pas) examples:

  • examples/ten_vad_remove_silence.rs — uses TenVadModelConfig with an explicit window_size = 256 (matching the other language examples) and feeds the VAD in 256-sample chunks.
  • run-ten-vad-remove-silence.sh — downloads ten-vad.onnx and the test wave file if needed, then runs the example.
  • README.md — added example 49 to the table and the "Run it" section.
  • .github/scripts/test-rust.sh — runs the new script right after the silero VAD one.

Tested on Linux x86_64 (Rust 1.94, default static linking):

Input WAV: sample rate: 16000, num samples: 4358370, duration: 272.40s
Saved speech-only audio to ./no-silence-ten-vad.wav

=== Summary ===
Input:  sample rate = 16000, samples = 4358370, duration = 272.40s
Output: sample rate = 16000, samples = 2513376, duration = 157.09s
Removed non-speech: 42.33% of input removed

Summary by CodeRabbit

  • New Features
    • Added a Rust CLI example to remove non-speech segments from WAV files using ten-vad, writing a speech-only output WAV and printing a processing summary.
    • Added a convenience script to fetch the required model and input audio (when missing) and run the example.
  • Documentation
    • Updated the Rust API examples README with “Example 49” and the command to execute the new silence-removal workflow.
  • Tests
    • Included the new silence-removal run script in the Rust test/run sequence.

@dosubot dosubot Bot added the size:L This PR changes 100-499 lines, ignoring generated files. label Jul 21, 2026
@coderabbitai

coderabbitai Bot commented Jul 21, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 67d74fdd-b23c-45ec-8af4-f9b6e34f2ea1

📥 Commits

Reviewing files that changed from the base of the PR and between 4a28f0c and beffc47.

📒 Files selected for processing (2)
  • rust-api-examples/examples/ten_vad_remove_silence.rs
  • rust-api-examples/run-ten-vad-remove-silence.sh
🚧 Files skipped from review as they are similar to previous changes (2)
  • rust-api-examples/run-ten-vad-remove-silence.sh
  • rust-api-examples/examples/ten_vad_remove_silence.rs

📝 Walkthrough

Walkthrough

Adds a Rust ten-vad example that removes non-speech audio from WAV files, a script that downloads required assets and runs it, README documentation, and integration into the Rust test sequence.

Changes

Ten-VAD silence removal

Layer / File(s) Summary
Rust VAD processing pipeline
rust-api-examples/examples/ten_vad_remove_silence.rs
Adds CLI arguments, ten-vad configuration, chunked waveform processing, speech-segment collection, filtered WAV output, and duration reporting.
Example runner and asset setup
rust-api-examples/run-ten-vad-remove-silence.sh
Downloads ten-vad.onnx and lei-jun-test.wav when missing, then runs the Rust example with the corresponding paths.
Documentation and test registration
rust-api-examples/README.md, .github/scripts/test-rust.sh
Documents Example 49 and adds the runner to the Rust test sequence.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Sequence Diagram(s)

sequenceDiagram
  participant Runner
  participant RustExample
  participant VoiceActivityDetector
  participant OutputWAV
  Runner->>RustExample: provide input, output, and model paths
  RustExample->>VoiceActivityDetector: accept_waveform audio chunks
  VoiceActivityDetector-->>RustExample: return detected speech segments
  RustExample->>VoiceActivityDetector: flush remaining audio
  VoiceActivityDetector-->>RustExample: return remaining speech segments
  RustExample->>OutputWAV: write speech-only samples
Loading

Possibly related PRs

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly matches the main change: a new ten-vad Rust API example for removing silence from audio files.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a new Rust API example and script for removing non-speech segments from audio files using ten-vad. Feedback on the changes includes validating that the input WAV file has a sample rate of 16000 Hz to prevent abrupt process termination, pre-allocating capacity for the speech samples vector to optimize performance, and adding the -f flag to curl commands in the shell script to ensure failures are handled gracefully.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment on lines +38 to +41
println!(
"Input WAV: sample rate: {}, num samples: {}, duration: {:.2}s",
sample_rate, input_num_samples, input_duration
);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Since ten-vad strictly requires a sample rate of 16000 Hz (and the underlying C++ implementation will abruptly terminate the process with exit(-1) if it is not), it is highly recommended to validate the sample rate in Rust and return a graceful error.

Suggested change
println!(
"Input WAV: sample rate: {}, num samples: {}, duration: {:.2}s",
sample_rate, input_num_samples, input_duration
);
println!(
"Input WAV: sample rate: {}, num samples: {}, duration: {:.2}s",
sample_rate, input_num_samples, input_duration
);
if sample_rate != 16000 {
anyhow::bail!(
"ten-vad expects a sample rate of 16000 Hz, but the input file has {} Hz",
sample_rate
);
}

let vad = VoiceActivityDetector::create(&vad_config, 30.0)
.ok_or_else(|| anyhow::anyhow!("Failed to create VoiceActivityDetector"))?;

let mut speech_samples = Vec::new();

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Reserving capacity for speech_samples based on input_num_samples avoids multiple reallocations and data copies as the vector grows, improving performance.

Suggested change
let mut speech_samples = Vec::new();
let mut speech_samples = Vec::with_capacity(input_num_samples as usize);

Comment on lines +5 to +11
if [ ! -f "./ten-vad.onnx" ]; then
curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/ten-vad.onnx
fi

if [ ! -f ./lei-jun-test.wav ]; then
curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/lei-jun-test.wav
fi

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

By default, curl does not return a non-zero exit code on HTTP errors (like 404 or 500) unless -f or --fail is specified. Adding -f ensures the script fails immediately if the download fails, preventing subsequent commands from trying to use invalid/empty files.

Suggested change
if [ ! -f "./ten-vad.onnx" ]; then
curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/ten-vad.onnx
fi
if [ ! -f ./lei-jun-test.wav ]; then
curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/lei-jun-test.wav
fi
if [ ! -f "./ten-vad.onnx" ]; then
curl -SLf -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/ten-vad.onnx
fi
if [ ! -f ./lei-jun-test.wav ]; then
curl -SLf -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/lei-jun-test.wav
fi

@csukuangfj csukuangfj left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for your contribution!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L This PR changes 100-499 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants