Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .github/scripts/test-rust.sh
Original file line number Diff line number Diff line change
Expand Up @@ -67,6 +67,7 @@ rm -rf sherpa-onnx-online-punct-*
rm -rf sherpa-onnx-kws-zipformer-wenetspeech-3.3M-2024-01-01-mobile

./run-spoken-language-identification.sh
./run-whisper.sh
rm -rf sherpa-onnx-whisper-tiny spoken-language-identification-test-wavs

./run-offline-punctuation.sh
Expand Down
7 changes: 7 additions & 0 deletions rust-api-examples/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -90,6 +90,7 @@ in your own Cargo project, see
| 45 | [zipformer_transducer_simulate_streaming_microphone](#example-45-simulated-streaming-asr-with-zipformer-transducer-and-vad-from-microphone) | Simulated streaming ASR with Zipformer transducer and VAD from microphone |
| 46 | [zipformer_transducer_simulate_streaming_microphone](#example-46-simulated-streaming-asr-with-zipformer-transducer-japanese-and-vad-from-microphone) | Simulated streaming ASR with Zipformer transducer (Japanese) and VAD from microphone |
| 47 | [qwen3_asr_simulate_streaming_microphone](#example-47-simulated-streaming-asr-with-qwen3-asr-and-vad-from-microphone) | Simulated streaming ASR with Qwen3 ASR and VAD from microphone |
| 48 | [whisper](#example-48-asr-with-non-streaming-whisper) | Non-streaming ASR with Whisper (multilingual) |

## Run it

Expand Down Expand Up @@ -415,3 +416,9 @@ Qwen3 ASR recognizer on each detected segment.
```bash
./run-qwen3-asr-simulate-streaming-microphone.sh
```

### Example 48: ASR with non-streaming Whisper

```bash
./run-whisper.sh
```
119 changes: 119 additions & 0 deletions rust-api-examples/examples/whisper.rs
Original file line number Diff line number Diff line change
@@ -0,0 +1,119 @@
// Copyright (c) 2026 Yujie Wu
//
// This file demonstrates how to use Whisper with sherpa-onnx's Rust API
// for offline speech recognition.
//
// See ../README.md for how to run it.

use clap::Parser;
use sherpa_onnx::{OfflineRecognizer, OfflineRecognizerConfig, OfflineWhisperModelConfig, Wave};
use std::time::Instant;

/// Whisper offline ASR example
#[derive(Parser, Debug)]
#[command(author, version, about, long_about = None)]
struct Args {
/// Path to encoder ONNX model
#[arg(long)]
encoder: String,

/// Path to decoder ONNX model
#[arg(long)]
decoder: String,

/// Path to tokens file
#[arg(long)]
tokens: String,

/// Path to input WAV file
#[arg(long)]
wav: String,

/// Recognition language, e.g. "en", "zh". Leave empty for auto-detection
#[arg(long, default_value = "")]
language: String,

/// Task type: "transcribe" or "translate" (translate to English)
#[arg(long, default_value = "transcribe")]
task: String,

/// Provider (default: cpu)
#[arg(long, default_value = "cpu")]
provider: String,

/// Enable debug logs
#[arg(long, default_value_t = false)]
debug: bool,

/// Number of threads
#[arg(long, default_value_t = 2)]
num_threads: i32,
Comment on lines +49 to +50

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

Validate --num-threads is positive before building config.

Non-positive values are currently accepted and forwarded directly, which can trigger backend init failures.

Suggested fix
 fn main() {
     let args = Args::parse();
+    if args.num_threads <= 0 {
+        eprintln!("--num-threads must be > 0");
+        std::process::exit(2);
+    }

     let wave = Wave::read(&args.wav).expect("Failed to read WAV file");

Also applies to: 53-54, 76-76

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@rust-api-examples/examples/whisper.rs` around lines 49 - 50, Validate that
the CLI numeric options are positive before building the runtime/config: check
the num_threads argument (struct field num_threads from the CLI parser) and
return an error or exit early if <= 0, and apply the same positive check to the
other numeric CLI fields referenced nearby (the numeric args around lines 53-54
and 76) so invalid non-positive values are rejected before constructing the
backend config or initializing threads. Ensure the validation happens right
after parsing CLI args and before any call that builds the config or initializes
the backend.

}

fn main() {
let args = Args::parse();

let wave = Wave::read(&args.wav).expect("Failed to read WAV file");
let audio_duration = wave.samples().len() as f64 / wave.sample_rate() as f64;

// Create default recognizer config
let mut recognizer_config = OfflineRecognizerConfig::default();

// Set the Whisper model
recognizer_config.model_config.whisper = OfflineWhisperModelConfig {
encoder: Some(args.encoder.clone()),
decoder: Some(args.decoder.clone()),
language: Some(args.language.clone()),
task: Some(args.task.clone()),
tail_paddings: 0,
enable_token_timestamps: false,
enable_segment_timestamps: false,
};
Comment on lines +63 to +71

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

If args.language is empty, passing Some("") to the configuration will result in a non-null pointer to an empty string being passed to the underlying C API. To ensure auto-detection works correctly and robustly across all platforms, it is more idiomatic to pass None when the language string is empty.

Suggested change
recognizer_config.model_config.whisper = OfflineWhisperModelConfig {
encoder: Some(args.encoder.clone()),
decoder: Some(args.decoder.clone()),
language: Some(args.language.clone()),
task: Some(args.task.clone()),
tail_paddings: 0,
enable_token_timestamps: false,
enable_segment_timestamps: false,
};
let language = if args.language.is_empty() {
None
} else {
Some(args.language.clone())
};
recognizer_config.model_config.whisper = OfflineWhisperModelConfig {
encoder: Some(args.encoder.clone()),
decoder: Some(args.decoder.clone()),
language,
task: Some(args.task.clone()),
tail_paddings: 0,
enable_token_timestamps: false,
enable_segment_timestamps: false,
};


recognizer_config.model_config.tokens = Some(args.tokens.clone());
recognizer_config.model_config.provider = Some(args.provider.clone());
recognizer_config.model_config.debug = args.debug;
recognizer_config.model_config.num_threads = args.num_threads;

// Measure recognizer creation time
println!("Creating recognizer ...");
let start_creation = Instant::now();
let recognizer =
OfflineRecognizer::create(&recognizer_config).expect("Failed to create OfflineRecognizer");
let creation_elapsed = start_creation.elapsed().as_secs_f64();
println!("Recognizer created in {:.3} seconds.", creation_elapsed);

let stream = recognizer.create_stream();

// Measure recognition time
let start_recognition = Instant::now();
stream.accept_waveform(wave.sample_rate(), wave.samples());
recognizer.decode(&stream);
let recognition_elapsed = start_recognition.elapsed().as_secs_f64();

// Get recognition result
if let Some(result) = stream.get_result() {
println!("Decoded text: {}", result.text);

let total_elapsed = creation_elapsed + recognition_elapsed;
let rtf = recognition_elapsed / audio_duration;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

If audio_duration is 0.0 (e.g., due to an empty or invalid WAV file), dividing by it will result in NaN or Infinity. Adding a guard ensures that the Real-Time Factor (RTF) is safely calculated as 0.0 in such cases.

        let rtf = if audio_duration > 0.0 {
            recognition_elapsed / audio_duration
        } else {
            0.0
        };

println!("\n=== Performance Summary ===");
println!("Audio duration : {:.3} seconds", audio_duration);
println!("Recognizer creation time: {:.3} seconds", creation_elapsed);
println!(
"Recognition time : {:.3} seconds",
recognition_elapsed
);
println!("Total elapsed time : {:.3} seconds", total_elapsed);
println!(
"Real-Time Factor (RTF) : {:.3} (recognition_elapsed / audio_duration = {:.3} / {:.3})",
rtf, recognition_elapsed, audio_duration
);
println!(
"Number of threads : {}",
recognizer_config.model_config.num_threads
);
} else {
eprintln!("Failed to get recognition result");
}
Comment on lines +95 to +118

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Fail fast when decoding returns no result.

The None branch only logs and exits with status 0, so CI can pass even when recognition fails. Return a non-zero exit code there.

Suggested fix
     } else {
         eprintln!("Failed to get recognition result");
+        std::process::exit(1);
     }
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
if let Some(result) = stream.get_result() {
println!("Decoded text: {}", result.text);
let total_elapsed = creation_elapsed + recognition_elapsed;
let rtf = recognition_elapsed / audio_duration;
println!("\n=== Performance Summary ===");
println!("Audio duration : {:.3} seconds", audio_duration);
println!("Recognizer creation time: {:.3} seconds", creation_elapsed);
println!(
"Recognition time : {:.3} seconds",
recognition_elapsed
);
println!("Total elapsed time : {:.3} seconds", total_elapsed);
println!(
"Real-Time Factor (RTF) : {:.3} (recognition_elapsed / audio_duration = {:.3} / {:.3})",
rtf, recognition_elapsed, audio_duration
);
println!(
"Number of threads : {}",
recognizer_config.model_config.num_threads
);
} else {
eprintln!("Failed to get recognition result");
}
if let Some(result) = stream.get_result() {
println!("Decoded text: {}", result.text);
let total_elapsed = creation_elapsed + recognition_elapsed;
let rtf = recognition_elapsed / audio_duration;
println!("\n=== Performance Summary ===");
println!("Audio duration : {:.3} seconds", audio_duration);
println!("Recognizer creation time: {:.3} seconds", creation_elapsed);
println!(
"Recognition time : {:.3} seconds",
recognition_elapsed
);
println!("Total elapsed time : {:.3} seconds", total_elapsed);
println!(
"Real-Time Factor (RTF) : {:.3} (recognition_elapsed / audio_duration = {:.3} / {:.3})",
rtf, recognition_elapsed, audio_duration
);
println!(
"Number of threads : {}",
recognizer_config.model_config.num_threads
);
} else {
eprintln!("Failed to get recognition result");
std::process::exit(1);
}
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@rust-api-examples/examples/whisper.rs` around lines 95 - 118, The None branch
after calling stream.get_result() currently only prints an error and exits with
status 0; change it to fail fast by terminating the process with a non-zero exit
code (e.g., call std::process::exit(1) or return Err from main) so CI detects
failures; update the else block that contains eprintln!("Failed to get
recognition result") (the branch after stream.get_result()) to exit with a
non-zero status instead of allowing normal success.

}
18 changes: 18 additions & 0 deletions rust-api-examples/run-whisper.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
#!/usr/bin/env bash
set -ex

# see
# https://github.com/k2-fsa/sherpa-onnx/releases/tag/asr-models
if [ ! -f ./sherpa-onnx-whisper-tiny/tiny-encoder.onnx ]; then
curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/sherpa-onnx-whisper-tiny.tar.bz2
tar xvf sherpa-onnx-whisper-tiny.tar.bz2
rm sherpa-onnx-whisper-tiny.tar.bz2
Comment on lines +6 to +9

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Harden artifact existence checks and download reliability.

The gate only checks tiny-encoder.onnx; if decoder/tokens/wav are missing, download is skipped and the run fails later. Also prefer curl --fail + retries for transient CI/network failures.

Suggested fix
-if [ ! -f ./sherpa-onnx-whisper-tiny/tiny-encoder.onnx ]; then
-  curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/sherpa-onnx-whisper-tiny.tar.bz2
+if [ ! -f ./sherpa-onnx-whisper-tiny/tiny-encoder.onnx ] || \
+   [ ! -f ./sherpa-onnx-whisper-tiny/tiny-decoder.onnx ] || \
+   [ ! -f ./sherpa-onnx-whisper-tiny/tiny-tokens.txt ] || \
+   [ ! -f ./sherpa-onnx-whisper-tiny/test_wavs/0.wav ]; then
+  curl --fail --location --retry 3 --retry-delay 2 -O \
+    https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/sherpa-onnx-whisper-tiny.tar.bz2
   tar xvf sherpa-onnx-whisper-tiny.tar.bz2
   rm sherpa-onnx-whisper-tiny.tar.bz2
 fi
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
if [ ! -f ./sherpa-onnx-whisper-tiny/tiny-encoder.onnx ]; then
curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/sherpa-onnx-whisper-tiny.tar.bz2
tar xvf sherpa-onnx-whisper-tiny.tar.bz2
rm sherpa-onnx-whisper-tiny.tar.bz2
if [ ! -f ./sherpa-onnx-whisper-tiny/tiny-encoder.onnx ] || \
[ ! -f ./sherpa-onnx-whisper-tiny/tiny-decoder.onnx ] || \
[ ! -f ./sherpa-onnx-whisper-tiny/tiny-tokens.txt ] || \
[ ! -f ./sherpa-onnx-whisper-tiny/test_wavs/0.wav ]; then
curl --fail --location --retry 3 --retry-delay 2 -O \
https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/sherpa-onnx-whisper-tiny.tar.bz2
tar xvf sherpa-onnx-whisper-tiny.tar.bz2
rm sherpa-onnx-whisper-tiny.tar.bz2
fi
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@rust-api-examples/run-whisper.sh` around lines 6 - 9, The script
run-whisper.sh currently gates download on only tiny-encoder.onnx; update it to
check for the full set of extracted artifacts (e.g., tiny-encoder.onnx plus
tiny-decoder.onnx, tokens.txt or tokens.bin, and any wav/model files referenced
later) and trigger the download if any are missing, and replace the curl
invocation with a resilient form (use curl --fail and retry logic, e.g., --retry
and backoff flags) and verify the tar extraction succeeded before removing the
.tar.bz2; reference the artifact names (tiny-encoder.onnx, tiny-decoder.onnx,
tokens.*) and the download/extract block in run-whisper.sh when making these
changes.

fi

cargo run --example whisper -- \
--wav ./sherpa-onnx-whisper-tiny/test_wavs/0.wav \
--encoder ./sherpa-onnx-whisper-tiny/tiny-encoder.onnx \
--decoder ./sherpa-onnx-whisper-tiny/tiny-decoder.onnx \
--tokens ./sherpa-onnx-whisper-tiny/tiny-tokens.txt \
--language en \
--num-threads 2