Add a Whisper example for the Rust API - #3661
Conversation
Add a Whisper offline recognition example under rust-api-examples, including the run-whisper.sh helper script, README entry, and CI registration in test-rust.sh.
There was a problem hiding this comment.
Code Review
This pull request adds a new Rust example for offline speech recognition using Whisper, including a runner script, integration into the test suite, and documentation. The review feedback suggests improving the robustness of the example by handling empty language arguments as None to ensure correct auto-detection, and guarding against division by zero when calculating the Real-Time Factor (RTF) if the audio duration is zero.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
| recognizer_config.model_config.whisper = OfflineWhisperModelConfig { | ||
| encoder: Some(args.encoder.clone()), | ||
| decoder: Some(args.decoder.clone()), | ||
| language: Some(args.language.clone()), | ||
| task: Some(args.task.clone()), | ||
| tail_paddings: 0, | ||
| enable_token_timestamps: false, | ||
| enable_segment_timestamps: false, | ||
| }; |
There was a problem hiding this comment.
If args.language is empty, passing Some("") to the configuration will result in a non-null pointer to an empty string being passed to the underlying C API. To ensure auto-detection works correctly and robustly across all platforms, it is more idiomatic to pass None when the language string is empty.
| recognizer_config.model_config.whisper = OfflineWhisperModelConfig { | |
| encoder: Some(args.encoder.clone()), | |
| decoder: Some(args.decoder.clone()), | |
| language: Some(args.language.clone()), | |
| task: Some(args.task.clone()), | |
| tail_paddings: 0, | |
| enable_token_timestamps: false, | |
| enable_segment_timestamps: false, | |
| }; | |
| let language = if args.language.is_empty() { | |
| None | |
| } else { | |
| Some(args.language.clone()) | |
| }; | |
| recognizer_config.model_config.whisper = OfflineWhisperModelConfig { | |
| encoder: Some(args.encoder.clone()), | |
| decoder: Some(args.decoder.clone()), | |
| language, | |
| task: Some(args.task.clone()), | |
| tail_paddings: 0, | |
| enable_token_timestamps: false, | |
| enable_segment_timestamps: false, | |
| }; |
| println!("Decoded text: {}", result.text); | ||
|
|
||
| let total_elapsed = creation_elapsed + recognition_elapsed; | ||
| let rtf = recognition_elapsed / audio_duration; |
There was a problem hiding this comment.
If audio_duration is 0.0 (e.g., due to an empty or invalid WAV file), dividing by it will result in NaN or Infinity. Adding a guard ensures that the Real-Time Factor (RTF) is safely calculated as 0.0 in such cases.
let rtf = if audio_duration > 0.0 {
recognition_elapsed / audio_duration
} else {
0.0
};|
Need the big picture first? Review this PR in Change Stack to see what changed before going file by file. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (1)
🚧 Files skipped from review as they are similar to previous changes (1)
📝 WalkthroughWalkthroughAdds a non-streaming Whisper offline ASR example: a Rust CLI example that decodes WAV files with sherpa-onnx Whisper, a bootstrap script to fetch models and run the example, README additions, and a CI step to execute the script during Rust tests. ChangesRust Whisper Offline ASR Example
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~20 minutes Possibly related PRs
Suggested reviewers
Poem
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 3
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@rust-api-examples/examples/whisper.rs`:
- Around line 49-50: Validate that the CLI numeric options are positive before
building the runtime/config: check the num_threads argument (struct field
num_threads from the CLI parser) and return an error or exit early if <= 0, and
apply the same positive check to the other numeric CLI fields referenced nearby
(the numeric args around lines 53-54 and 76) so invalid non-positive values are
rejected before constructing the backend config or initializing threads. Ensure
the validation happens right after parsing CLI args and before any call that
builds the config or initializes the backend.
- Around line 95-118: The None branch after calling stream.get_result()
currently only prints an error and exits with status 0; change it to fail fast
by terminating the process with a non-zero exit code (e.g., call
std::process::exit(1) or return Err from main) so CI detects failures; update
the else block that contains eprintln!("Failed to get recognition result") (the
branch after stream.get_result()) to exit with a non-zero status instead of
allowing normal success.
In `@rust-api-examples/run-whisper.sh`:
- Around line 6-9: The script run-whisper.sh currently gates download on only
tiny-encoder.onnx; update it to check for the full set of extracted artifacts
(e.g., tiny-encoder.onnx plus tiny-decoder.onnx, tokens.txt or tokens.bin, and
any wav/model files referenced later) and trigger the download if any are
missing, and replace the curl invocation with a resilient form (use curl --fail
and retry logic, e.g., --retry and backoff flags) and verify the tar extraction
succeeded before removing the .tar.bz2; reference the artifact names
(tiny-encoder.onnx, tiny-decoder.onnx, tokens.*) and the download/extract block
in run-whisper.sh when making these changes.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro
Run ID: e922aa26-879b-405c-a812-b2e3bdd1798e
📒 Files selected for processing (4)
.github/scripts/test-rust.shrust-api-examples/README.mdrust-api-examples/examples/whisper.rsrust-api-examples/run-whisper.sh
| #[arg(long, default_value_t = 2)] | ||
| num_threads: i32, |
There was a problem hiding this comment.
Validate --num-threads is positive before building config.
Non-positive values are currently accepted and forwarded directly, which can trigger backend init failures.
Suggested fix
fn main() {
let args = Args::parse();
+ if args.num_threads <= 0 {
+ eprintln!("--num-threads must be > 0");
+ std::process::exit(2);
+ }
let wave = Wave::read(&args.wav).expect("Failed to read WAV file");Also applies to: 53-54, 76-76
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@rust-api-examples/examples/whisper.rs` around lines 49 - 50, Validate that
the CLI numeric options are positive before building the runtime/config: check
the num_threads argument (struct field num_threads from the CLI parser) and
return an error or exit early if <= 0, and apply the same positive check to the
other numeric CLI fields referenced nearby (the numeric args around lines 53-54
and 76) so invalid non-positive values are rejected before constructing the
backend config or initializing threads. Ensure the validation happens right
after parsing CLI args and before any call that builds the config or initializes
the backend.
| if let Some(result) = stream.get_result() { | ||
| println!("Decoded text: {}", result.text); | ||
|
|
||
| let total_elapsed = creation_elapsed + recognition_elapsed; | ||
| let rtf = recognition_elapsed / audio_duration; | ||
| println!("\n=== Performance Summary ==="); | ||
| println!("Audio duration : {:.3} seconds", audio_duration); | ||
| println!("Recognizer creation time: {:.3} seconds", creation_elapsed); | ||
| println!( | ||
| "Recognition time : {:.3} seconds", | ||
| recognition_elapsed | ||
| ); | ||
| println!("Total elapsed time : {:.3} seconds", total_elapsed); | ||
| println!( | ||
| "Real-Time Factor (RTF) : {:.3} (recognition_elapsed / audio_duration = {:.3} / {:.3})", | ||
| rtf, recognition_elapsed, audio_duration | ||
| ); | ||
| println!( | ||
| "Number of threads : {}", | ||
| recognizer_config.model_config.num_threads | ||
| ); | ||
| } else { | ||
| eprintln!("Failed to get recognition result"); | ||
| } |
There was a problem hiding this comment.
Fail fast when decoding returns no result.
The None branch only logs and exits with status 0, so CI can pass even when recognition fails. Return a non-zero exit code there.
Suggested fix
} else {
eprintln!("Failed to get recognition result");
+ std::process::exit(1);
}📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| if let Some(result) = stream.get_result() { | |
| println!("Decoded text: {}", result.text); | |
| let total_elapsed = creation_elapsed + recognition_elapsed; | |
| let rtf = recognition_elapsed / audio_duration; | |
| println!("\n=== Performance Summary ==="); | |
| println!("Audio duration : {:.3} seconds", audio_duration); | |
| println!("Recognizer creation time: {:.3} seconds", creation_elapsed); | |
| println!( | |
| "Recognition time : {:.3} seconds", | |
| recognition_elapsed | |
| ); | |
| println!("Total elapsed time : {:.3} seconds", total_elapsed); | |
| println!( | |
| "Real-Time Factor (RTF) : {:.3} (recognition_elapsed / audio_duration = {:.3} / {:.3})", | |
| rtf, recognition_elapsed, audio_duration | |
| ); | |
| println!( | |
| "Number of threads : {}", | |
| recognizer_config.model_config.num_threads | |
| ); | |
| } else { | |
| eprintln!("Failed to get recognition result"); | |
| } | |
| if let Some(result) = stream.get_result() { | |
| println!("Decoded text: {}", result.text); | |
| let total_elapsed = creation_elapsed + recognition_elapsed; | |
| let rtf = recognition_elapsed / audio_duration; | |
| println!("\n=== Performance Summary ==="); | |
| println!("Audio duration : {:.3} seconds", audio_duration); | |
| println!("Recognizer creation time: {:.3} seconds", creation_elapsed); | |
| println!( | |
| "Recognition time : {:.3} seconds", | |
| recognition_elapsed | |
| ); | |
| println!("Total elapsed time : {:.3} seconds", total_elapsed); | |
| println!( | |
| "Real-Time Factor (RTF) : {:.3} (recognition_elapsed / audio_duration = {:.3} / {:.3})", | |
| rtf, recognition_elapsed, audio_duration | |
| ); | |
| println!( | |
| "Number of threads : {}", | |
| recognizer_config.model_config.num_threads | |
| ); | |
| } else { | |
| eprintln!("Failed to get recognition result"); | |
| std::process::exit(1); | |
| } |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@rust-api-examples/examples/whisper.rs` around lines 95 - 118, The None branch
after calling stream.get_result() currently only prints an error and exits with
status 0; change it to fail fast by terminating the process with a non-zero exit
code (e.g., call std::process::exit(1) or return Err from main) so CI detects
failures; update the else block that contains eprintln!("Failed to get
recognition result") (the branch after stream.get_result()) to exit with a
non-zero status instead of allowing normal success.
| if [ ! -f ./sherpa-onnx-whisper-tiny/tiny-encoder.onnx ]; then | ||
| curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/sherpa-onnx-whisper-tiny.tar.bz2 | ||
| tar xvf sherpa-onnx-whisper-tiny.tar.bz2 | ||
| rm sherpa-onnx-whisper-tiny.tar.bz2 |
There was a problem hiding this comment.
Harden artifact existence checks and download reliability.
The gate only checks tiny-encoder.onnx; if decoder/tokens/wav are missing, download is skipped and the run fails later. Also prefer curl --fail + retries for transient CI/network failures.
Suggested fix
-if [ ! -f ./sherpa-onnx-whisper-tiny/tiny-encoder.onnx ]; then
- curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/sherpa-onnx-whisper-tiny.tar.bz2
+if [ ! -f ./sherpa-onnx-whisper-tiny/tiny-encoder.onnx ] || \
+ [ ! -f ./sherpa-onnx-whisper-tiny/tiny-decoder.onnx ] || \
+ [ ! -f ./sherpa-onnx-whisper-tiny/tiny-tokens.txt ] || \
+ [ ! -f ./sherpa-onnx-whisper-tiny/test_wavs/0.wav ]; then
+ curl --fail --location --retry 3 --retry-delay 2 -O \
+ https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/sherpa-onnx-whisper-tiny.tar.bz2
tar xvf sherpa-onnx-whisper-tiny.tar.bz2
rm sherpa-onnx-whisper-tiny.tar.bz2
fi📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| if [ ! -f ./sherpa-onnx-whisper-tiny/tiny-encoder.onnx ]; then | |
| curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/sherpa-onnx-whisper-tiny.tar.bz2 | |
| tar xvf sherpa-onnx-whisper-tiny.tar.bz2 | |
| rm sherpa-onnx-whisper-tiny.tar.bz2 | |
| if [ ! -f ./sherpa-onnx-whisper-tiny/tiny-encoder.onnx ] || \ | |
| [ ! -f ./sherpa-onnx-whisper-tiny/tiny-decoder.onnx ] || \ | |
| [ ! -f ./sherpa-onnx-whisper-tiny/tiny-tokens.txt ] || \ | |
| [ ! -f ./sherpa-onnx-whisper-tiny/test_wavs/0.wav ]; then | |
| curl --fail --location --retry 3 --retry-delay 2 -O \ | |
| https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/sherpa-onnx-whisper-tiny.tar.bz2 | |
| tar xvf sherpa-onnx-whisper-tiny.tar.bz2 | |
| rm sherpa-onnx-whisper-tiny.tar.bz2 | |
| fi |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@rust-api-examples/run-whisper.sh` around lines 6 - 9, The script
run-whisper.sh currently gates download on only tiny-encoder.onnx; update it to
check for the full set of extracted artifacts (e.g., tiny-encoder.onnx plus
tiny-decoder.onnx, tokens.txt or tokens.bin, and any wav/model files referenced
later) and trigger the download if any are missing, and replace the curl
invocation with a resilient form (use curl --fail and retry logic, e.g., --retry
and backoff flags) and verify the tar extraction succeeded before removing the
.tar.bz2; reference the artifact names (tiny-encoder.onnx, tiny-decoder.onnx,
tokens.*) and the download/extract block in run-whisper.sh when making these
changes.
| @@ -0,0 +1,119 @@ | |||
| // Copyright (c) 2026 Xiaomi Corporation | |||
There was a problem hiding this comment.
Can you update it to use your own information?
There was a problem hiding this comment.
Done, I’ve already updated it with my own information. Thanks for the reminder!
csukuangfj
left a comment
There was a problem hiding this comment.
Thank you for your contribution!
Part of #3210
Add a Whisper offline recognition example under rust-api-examples, including the run-whisper.sh helper script, README entry, and CI registration in test-rust.sh.
Summary by CodeRabbit
New Features
Documentation
Tests