Add ten-vad Rust API example to remove silences from a file - #3778
Conversation
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (2)
🚧 Files skipped from review as they are similar to previous changes (2)
📝 WalkthroughWalkthroughAdds a Rust ChangesTen-VAD silence removal
Estimated code review effort: 3 (Moderate) | ~20 minutes Sequence Diagram(s)sequenceDiagram
participant Runner
participant RustExample
participant VoiceActivityDetector
participant OutputWAV
Runner->>RustExample: provide input, output, and model paths
RustExample->>VoiceActivityDetector: accept_waveform audio chunks
VoiceActivityDetector-->>RustExample: return detected speech segments
RustExample->>VoiceActivityDetector: flush remaining audio
VoiceActivityDetector-->>RustExample: return remaining speech segments
RustExample->>OutputWAV: write speech-only samples
Possibly related PRs
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Code Review
This pull request introduces a new Rust API example and script for removing non-speech segments from audio files using ten-vad. Feedback on the changes includes validating that the input WAV file has a sample rate of 16000 Hz to prevent abrupt process termination, pre-allocating capacity for the speech samples vector to optimize performance, and adding the -f flag to curl commands in the shell script to ensure failures are handled gracefully.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
| println!( | ||
| "Input WAV: sample rate: {}, num samples: {}, duration: {:.2}s", | ||
| sample_rate, input_num_samples, input_duration | ||
| ); |
There was a problem hiding this comment.
Since ten-vad strictly requires a sample rate of 16000 Hz (and the underlying C++ implementation will abruptly terminate the process with exit(-1) if it is not), it is highly recommended to validate the sample rate in Rust and return a graceful error.
| println!( | |
| "Input WAV: sample rate: {}, num samples: {}, duration: {:.2}s", | |
| sample_rate, input_num_samples, input_duration | |
| ); | |
| println!( | |
| "Input WAV: sample rate: {}, num samples: {}, duration: {:.2}s", | |
| sample_rate, input_num_samples, input_duration | |
| ); | |
| if sample_rate != 16000 { | |
| anyhow::bail!( | |
| "ten-vad expects a sample rate of 16000 Hz, but the input file has {} Hz", | |
| sample_rate | |
| ); | |
| } |
| let vad = VoiceActivityDetector::create(&vad_config, 30.0) | ||
| .ok_or_else(|| anyhow::anyhow!("Failed to create VoiceActivityDetector"))?; | ||
|
|
||
| let mut speech_samples = Vec::new(); |
There was a problem hiding this comment.
| if [ ! -f "./ten-vad.onnx" ]; then | ||
| curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/ten-vad.onnx | ||
| fi | ||
|
|
||
| if [ ! -f ./lei-jun-test.wav ]; then | ||
| curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/lei-jun-test.wav | ||
| fi |
There was a problem hiding this comment.
By default, curl does not return a non-zero exit code on HTTP errors (like 404 or 500) unless -f or --fail is specified. Adding -f ensures the script fails immediately if the download fails, preventing subsequent commands from trying to use invalid/empty files.
| if [ ! -f "./ten-vad.onnx" ]; then | |
| curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/ten-vad.onnx | |
| fi | |
| if [ ! -f ./lei-jun-test.wav ]; then | |
| curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/lei-jun-test.wav | |
| fi | |
| if [ ! -f "./ten-vad.onnx" ]; then | |
| curl -SLf -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/ten-vad.onnx | |
| fi | |
| if [ ! -f ./lei-jun-test.wav ]; then | |
| curl -SLf -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/lei-jun-test.wav | |
| fi |
csukuangfj
left a comment
There was a problem hiding this comment.
Thank you for your contribution!
Part of #3210 ("Add ten-vad example to remove silences from a file").
This ports
rust-api-examples/examples/silero_vad_remove_silence.rsto ten-vad, following the existing Java (TenVadRemoveSilence.java) and Pascal (remove_silence_ten_vad.pas) examples:examples/ten_vad_remove_silence.rs— usesTenVadModelConfigwith an explicitwindow_size = 256(matching the other language examples) and feeds the VAD in 256-sample chunks.run-ten-vad-remove-silence.sh— downloadsten-vad.onnxand the test wave file if needed, then runs the example.README.md— added example 49 to the table and the "Run it" section..github/scripts/test-rust.sh— runs the new script right after the silero VAD one.Tested on Linux x86_64 (Rust 1.94, default static linking):
Summary by CodeRabbit
ten-vad, writing a speech-only output WAV and printing a processing summary.