Skip to content

Add a Whisper example for the Rust API - #3661

Merged
csukuangfj merged 2 commits into
k2-fsa:masterfrom
WuYujie666:rust-whisper-example
Jun 6, 2026
Merged

csukuangfj merged 2 commits into
k2-fsa:masterfrom
WuYujie666:rust-whisper-example

Conversation

@WuYujie666

@WuYujie666 WuYujie666 commented Jun 5, 2026 •

Copy link
Copy Markdown
Contributor

Part of #3210
Add a Whisper offline recognition example under rust-api-examples, including the run-whisper.sh helper script, README entry, and CI registration in test-rust.sh.

$ ./run-whisper.sh
+ '[' '!' -f ./sherpa-onnx-whisper-tiny/tiny-encoder.onnx ']'
+ cargo run --example whisper -- --wav ./sherpa-onnx-whisper-tiny/test_wavs/0.wav --encoder ./sherpa-onnx-whisper-tiny/tiny-encoder.onnx --decoder ./sherpa-onnx-whisper-tiny/tiny-decoder.onnx --tokens ./sherpa-onnx-whisper-tiny/tiny-tokens.txt --language en --num-threads 2
   Compiling rust-api-examples v1.13.2 (C:\code\sherpa-onnx\rust-api-examples)
    Finished `dev` profile [unoptimized + debuginfo] target(s) in 3.97s                          
     Running `target\debug\examples\whisper.exe --wav ./sherpa-onnx-whisper-tiny/test_wavs/0.wav --encoder ./sherpa-onnx-whisper-tiny/tiny-encoder.onnx --decoder ./sherpa-onnx-whisper-tiny/tiny-decoder.onnx --tokens ./sherpa-onnx-whisper-tiny/tiny-tokens.txt --language en --num-threads 2`
Creating recognizer ...
Recognizer created in 0.458 seconds.
Decoded text:  After early nightfall, the yellow lamps would light up here and there the squalid quarter of the brothels.

=== Performance Summary ===
Audio duration          : 6.625 seconds
Recognizer creation time: 0.458 seconds
Recognition time        : 0.406 seconds
Total elapsed time      : 0.864 seconds
Real-Time Factor (RTF)  : 0.061 (recognition_elapsed / audio_duration = 0.406 / 6.625)
Number of threads       : 2

Summary by CodeRabbit

  • New Features

    • Added a new CLI example showcasing offline Whisper-based speech recognition with performance output (elapsed time and real-time factor).
  • Documentation

    • Updated examples index and run instructions with a new entry and a "Run it" command for the Whisper example.
  • Tests

    • Added Whisper ASR execution to the test pipeline so Whisper functionality runs as part of automated tests.

Add a Whisper offline recognition example under rust-api-examples, including the run-whisper.sh helper script, README entry, and CI registration in test-rust.sh.
@dosubot dosubot Bot added the size:L This PR changes 100-499 lines, ignoring generated files. label Jun 5, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request adds a new Rust example for offline speech recognition using Whisper, including a runner script, integration into the test suite, and documentation. The review feedback suggests improving the robustness of the example by handling empty language arguments as None to ensure correct auto-detection, and guarding against division by zero when calculating the Real-Time Factor (RTF) if the audio duration is zero.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment on lines +63 to +71
recognizer_config.model_config.whisper = OfflineWhisperModelConfig {
encoder: Some(args.encoder.clone()),
decoder: Some(args.decoder.clone()),
language: Some(args.language.clone()),
task: Some(args.task.clone()),
tail_paddings: 0,
enable_token_timestamps: false,
enable_segment_timestamps: false,
};

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

If args.language is empty, passing Some("") to the configuration will result in a non-null pointer to an empty string being passed to the underlying C API. To ensure auto-detection works correctly and robustly across all platforms, it is more idiomatic to pass None when the language string is empty.

Suggested change
recognizer_config.model_config.whisper = OfflineWhisperModelConfig {
encoder: Some(args.encoder.clone()),
decoder: Some(args.decoder.clone()),
language: Some(args.language.clone()),
task: Some(args.task.clone()),
tail_paddings: 0,
enable_token_timestamps: false,
enable_segment_timestamps: false,
};
let language = if args.language.is_empty() {
None
} else {
Some(args.language.clone())
};
recognizer_config.model_config.whisper = OfflineWhisperModelConfig {
encoder: Some(args.encoder.clone()),
decoder: Some(args.decoder.clone()),
language,
task: Some(args.task.clone()),
tail_paddings: 0,
enable_token_timestamps: false,
enable_segment_timestamps: false,
};

println!("Decoded text: {}", result.text);

let total_elapsed = creation_elapsed + recognition_elapsed;
let rtf = recognition_elapsed / audio_duration;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

If audio_duration is 0.0 (e.g., due to an empty or invalid WAV file), dividing by it will result in NaN or Infinity. Adding a guard ensures that the Real-Time Factor (RTF) is safely calculated as 0.0 in such cases.

        let rtf = if audio_duration > 0.0 {
            recognition_elapsed / audio_duration
        } else {
            0.0
        };

@coderabbitai

coderabbitai Bot commented Jun 5, 2026 •

Copy link
Copy Markdown

Need the big picture first? Review this PR in Change Stack to see what changed before going file by file.

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: de69e668-664a-490d-a3ac-dba5435f88d6

📥 Commits

Reviewing files that changed from the base of the PR and between b5597ef and 87510d9.

📒 Files selected for processing (1)
  • rust-api-examples/examples/whisper.rs
🚧 Files skipped from review as they are similar to previous changes (1)
  • rust-api-examples/examples/whisper.rs

📝 Walkthrough

Walkthrough

Adds a non-streaming Whisper offline ASR example: a Rust CLI example that decodes WAV files with sherpa-onnx Whisper, a bootstrap script to fetch models and run the example, README additions, and a CI step to execute the script during Rust tests.

Changes

Rust Whisper Offline ASR Example

Layer / File(s) Summary
Whisper Example Implementation
rust-api-examples/examples/whisper.rs
Args Clap parser for model/token/wav/language/task/provider/debug/threads; main loads WAV, builds OfflineRecognizerConfig with OfflineWhisperModelConfig, creates OfflineRecognizer, decodes a stream, and prints decoded text plus timing and RTF metrics.
Example Bootstrap Script & Documentation
rust-api-examples/run-whisper.sh, rust-api-examples/README.md
run-whisper.sh conditionally downloads/extracts Whisper tiny ONNX assets and invokes cargo run --example whisper with model/token paths, --language en, and --num-threads 2; README adds Example 48 to the index and Run-it section with the ./run-whisper.sh command.
CI Test Pipeline Integration
.github/scripts/test-rust.sh
Inserts ./run-whisper.sh execution into the Rust test runner immediately after the spoken-language-identification step.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

  • k2-fsa/sherpa-onnx#3352: Also adds a new example script invocation to .github/scripts/test-rust.sh though for a different example.
  • k2-fsa/sherpa-onnx#3209: Similarly extends the Rust test pipeline to run additional offline ASR example scripts.

Suggested reviewers

  • csukuangfj

Poem

🐰 I hopped in with models snug and light,
Tiny Whisper ready to transcribe the night,
Rust streams hummed, tokens danced in line,
Audio to text — a rabbit's little sign.
CI clicks, results appear — oh what a sight!

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'Add a Whisper example for the Rust API' clearly and accurately summarizes the main change—adding a new Whisper offline recognition example to the Rust API examples.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@rust-api-examples/examples/whisper.rs`:
- Around line 49-50: Validate that the CLI numeric options are positive before
building the runtime/config: check the num_threads argument (struct field
num_threads from the CLI parser) and return an error or exit early if <= 0, and
apply the same positive check to the other numeric CLI fields referenced nearby
(the numeric args around lines 53-54 and 76) so invalid non-positive values are
rejected before constructing the backend config or initializing threads. Ensure
the validation happens right after parsing CLI args and before any call that
builds the config or initializes the backend.
- Around line 95-118: The None branch after calling stream.get_result()
currently only prints an error and exits with status 0; change it to fail fast
by terminating the process with a non-zero exit code (e.g., call
std::process::exit(1) or return Err from main) so CI detects failures; update
the else block that contains eprintln!("Failed to get recognition result") (the
branch after stream.get_result()) to exit with a non-zero status instead of
allowing normal success.

In `@rust-api-examples/run-whisper.sh`:
- Around line 6-9: The script run-whisper.sh currently gates download on only
tiny-encoder.onnx; update it to check for the full set of extracted artifacts
(e.g., tiny-encoder.onnx plus tiny-decoder.onnx, tokens.txt or tokens.bin, and
any wav/model files referenced later) and trigger the download if any are
missing, and replace the curl invocation with a resilient form (use curl --fail
and retry logic, e.g., --retry and backoff flags) and verify the tar extraction
succeeded before removing the .tar.bz2; reference the artifact names
(tiny-encoder.onnx, tiny-decoder.onnx, tokens.*) and the download/extract block
in run-whisper.sh when making these changes.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: e922aa26-879b-405c-a812-b2e3bdd1798e

📥 Commits

Reviewing files that changed from the base of the PR and between 8799533 and b5597ef.

📒 Files selected for processing (4)
  • .github/scripts/test-rust.sh
  • rust-api-examples/README.md
  • rust-api-examples/examples/whisper.rs
  • rust-api-examples/run-whisper.sh

Comment on lines +49 to +50
#[arg(long, default_value_t = 2)]
num_threads: i32,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

Validate --num-threads is positive before building config.

Non-positive values are currently accepted and forwarded directly, which can trigger backend init failures.

Suggested fix
 fn main() {
     let args = Args::parse();
+    if args.num_threads <= 0 {
+        eprintln!("--num-threads must be > 0");
+        std::process::exit(2);
+    }

     let wave = Wave::read(&args.wav).expect("Failed to read WAV file");

Also applies to: 53-54, 76-76

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@rust-api-examples/examples/whisper.rs` around lines 49 - 50, Validate that
the CLI numeric options are positive before building the runtime/config: check
the num_threads argument (struct field num_threads from the CLI parser) and
return an error or exit early if <= 0, and apply the same positive check to the
other numeric CLI fields referenced nearby (the numeric args around lines 53-54
and 76) so invalid non-positive values are rejected before constructing the
backend config or initializing threads. Ensure the validation happens right
after parsing CLI args and before any call that builds the config or initializes
the backend.

Comment on lines +95 to +118
if let Some(result) = stream.get_result() {
println!("Decoded text: {}", result.text);

let total_elapsed = creation_elapsed + recognition_elapsed;
let rtf = recognition_elapsed / audio_duration;
println!("\n=== Performance Summary ===");
println!("Audio duration : {:.3} seconds", audio_duration);
println!("Recognizer creation time: {:.3} seconds", creation_elapsed);
println!(
"Recognition time : {:.3} seconds",
recognition_elapsed
);
println!("Total elapsed time : {:.3} seconds", total_elapsed);
println!(
"Real-Time Factor (RTF) : {:.3} (recognition_elapsed / audio_duration = {:.3} / {:.3})",
rtf, recognition_elapsed, audio_duration
);
println!(
"Number of threads : {}",
recognizer_config.model_config.num_threads
);
} else {
eprintln!("Failed to get recognition result");
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Fail fast when decoding returns no result.

The None branch only logs and exits with status 0, so CI can pass even when recognition fails. Return a non-zero exit code there.

Suggested fix
     } else {
         eprintln!("Failed to get recognition result");
+        std::process::exit(1);
     }
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
if let Some(result) = stream.get_result() {
println!("Decoded text: {}", result.text);
let total_elapsed = creation_elapsed + recognition_elapsed;
let rtf = recognition_elapsed / audio_duration;
println!("\n=== Performance Summary ===");
println!("Audio duration : {:.3} seconds", audio_duration);
println!("Recognizer creation time: {:.3} seconds", creation_elapsed);
println!(
"Recognition time : {:.3} seconds",
recognition_elapsed
);
println!("Total elapsed time : {:.3} seconds", total_elapsed);
println!(
"Real-Time Factor (RTF) : {:.3} (recognition_elapsed / audio_duration = {:.3} / {:.3})",
rtf, recognition_elapsed, audio_duration
);
println!(
"Number of threads : {}",
recognizer_config.model_config.num_threads
);
} else {
eprintln!("Failed to get recognition result");
}
if let Some(result) = stream.get_result() {
println!("Decoded text: {}", result.text);
let total_elapsed = creation_elapsed + recognition_elapsed;
let rtf = recognition_elapsed / audio_duration;
println!("\n=== Performance Summary ===");
println!("Audio duration : {:.3} seconds", audio_duration);
println!("Recognizer creation time: {:.3} seconds", creation_elapsed);
println!(
"Recognition time : {:.3} seconds",
recognition_elapsed
);
println!("Total elapsed time : {:.3} seconds", total_elapsed);
println!(
"Real-Time Factor (RTF) : {:.3} (recognition_elapsed / audio_duration = {:.3} / {:.3})",
rtf, recognition_elapsed, audio_duration
);
println!(
"Number of threads : {}",
recognizer_config.model_config.num_threads
);
} else {
eprintln!("Failed to get recognition result");
std::process::exit(1);
}
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@rust-api-examples/examples/whisper.rs` around lines 95 - 118, The None branch
after calling stream.get_result() currently only prints an error and exits with
status 0; change it to fail fast by terminating the process with a non-zero exit
code (e.g., call std::process::exit(1) or return Err from main) so CI detects
failures; update the else block that contains eprintln!("Failed to get
recognition result") (the branch after stream.get_result()) to exit with a
non-zero status instead of allowing normal success.

Comment on lines +6 to +9
if [ ! -f ./sherpa-onnx-whisper-tiny/tiny-encoder.onnx ]; then
curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/sherpa-onnx-whisper-tiny.tar.bz2
tar xvf sherpa-onnx-whisper-tiny.tar.bz2
rm sherpa-onnx-whisper-tiny.tar.bz2

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Harden artifact existence checks and download reliability.

The gate only checks tiny-encoder.onnx; if decoder/tokens/wav are missing, download is skipped and the run fails later. Also prefer curl --fail + retries for transient CI/network failures.

Suggested fix
-if [ ! -f ./sherpa-onnx-whisper-tiny/tiny-encoder.onnx ]; then
-  curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/sherpa-onnx-whisper-tiny.tar.bz2
+if [ ! -f ./sherpa-onnx-whisper-tiny/tiny-encoder.onnx ] || \
+   [ ! -f ./sherpa-onnx-whisper-tiny/tiny-decoder.onnx ] || \
+   [ ! -f ./sherpa-onnx-whisper-tiny/tiny-tokens.txt ] || \
+   [ ! -f ./sherpa-onnx-whisper-tiny/test_wavs/0.wav ]; then
+  curl --fail --location --retry 3 --retry-delay 2 -O \
+    https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/sherpa-onnx-whisper-tiny.tar.bz2
   tar xvf sherpa-onnx-whisper-tiny.tar.bz2
   rm sherpa-onnx-whisper-tiny.tar.bz2
 fi
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
if [ ! -f ./sherpa-onnx-whisper-tiny/tiny-encoder.onnx ]; then
curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/sherpa-onnx-whisper-tiny.tar.bz2
tar xvf sherpa-onnx-whisper-tiny.tar.bz2
rm sherpa-onnx-whisper-tiny.tar.bz2
if [ ! -f ./sherpa-onnx-whisper-tiny/tiny-encoder.onnx ] || \
[ ! -f ./sherpa-onnx-whisper-tiny/tiny-decoder.onnx ] || \
[ ! -f ./sherpa-onnx-whisper-tiny/tiny-tokens.txt ] || \
[ ! -f ./sherpa-onnx-whisper-tiny/test_wavs/0.wav ]; then
curl --fail --location --retry 3 --retry-delay 2 -O \
https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/sherpa-onnx-whisper-tiny.tar.bz2
tar xvf sherpa-onnx-whisper-tiny.tar.bz2
rm sherpa-onnx-whisper-tiny.tar.bz2
fi
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@rust-api-examples/run-whisper.sh` around lines 6 - 9, The script
run-whisper.sh currently gates download on only tiny-encoder.onnx; update it to
check for the full set of extracted artifacts (e.g., tiny-encoder.onnx plus
tiny-decoder.onnx, tokens.txt or tokens.bin, and any wav/model files referenced
later) and trigger the download if any are missing, and replace the curl
invocation with a resilient form (use curl --fail and retry logic, e.g., --retry
and backoff flags) and verify the tar extraction succeeded before removing the
.tar.bz2; reference the artifact names (tiny-encoder.onnx, tiny-decoder.onnx,
tokens.*) and the download/extract block in run-whisper.sh when making these
changes.

Comment thread rust-api-examples/examples/whisper.rs Outdated
@@ -0,0 +1,119 @@
// Copyright (c) 2026 Xiaomi Corporation

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you update it to use your own information?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done, I’ve already updated it with my own information. Thanks for the reminder!

@csukuangfj csukuangfj left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for your contribution!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L This PR changes 100-499 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants