Add Rust API for speaker diarization - #3370
Conversation
|
Caution Review failedThe pull request is closed. ℹ️ Recent review info⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (9)
📝 WalkthroughWalkthroughThis PR introduces offline speaker diarization functionality to the Rust API. It adds FFI bindings, safe Rust wrappers with resource management, a complete example with automated model downloading, documentation, and CI integration for testing the new feature. Changes
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~25 minutes Possibly related PRs
Suggested labels
Poem
✨ Finishing Touches🧪 Generate unit tests (beta)
📝 Coding Plan
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request significantly expands the Rust API capabilities by introducing comprehensive support for offline speaker diarization. This new feature allows users to identify and segment speech by different speakers in an audio recording, providing a robust tool for multi-speaker analysis. The changes include new FFI bindings, a high-level Rust API wrapper, and a runnable example, making it accessible for Rust developers. Highlights
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
There was a problem hiding this comment.
Code Review
This pull request introduces a Rust API for speaker diarization, complete with FFI bindings, a high-level safe wrapper, and a usage example. The implementation is well-structured. I have a couple of suggestions for improvement: one to correct a typo in a download URL within a shell script, and another to add a safeguard against potential integer overflow when handling audio sample lengths in the Rust wrapper.
| fi | ||
|
|
||
| if [ ! -f ./3dspeaker_speech_eres2net_base_sv_zh-cn_3dspeaker_16k.onnx ]; then | ||
| curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/speaker-recongition-models/3dspeaker_speech_eres2net_base_sv_zh-cn_3dspeaker_16k.onnx |
There was a problem hiding this comment.
There's a typo in the URL: recongition should be recognition. While the current URL with the typo works because a release with the typo exists, it's better to use the corrected URL for future-proofing and clarity, as a corrected release tag also exists.
| curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/speaker-recongition-models/3dspeaker_speech_eres2net_base_sv_zh-cn_3dspeaker_16k.onnx | |
| curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/speaker-recognition-models/3dspeaker_speech_eres2net_base_sv_zh-cn_3dspeaker_16k.onnx |
| let ptr = unsafe { | ||
| sys::SherpaOnnxOfflineSpeakerDiarizationProcess( | ||
| self.ptr, | ||
| samples.as_ptr(), | ||
| samples.len() as i32, | ||
| ) | ||
| }; |
There was a problem hiding this comment.
The cast samples.len() as i32 can truncate if the number of samples is larger than i32::MAX. This could lead to incorrect data being processed or a panic. While it's unlikely to happen with typical audio files (it would require an audio file of over 37 hours at 16kHz), it's safer to handle this potential overflow. Using try_into() provides a more robust way to perform the conversion and handle the error gracefully.
let n = match samples.len().try_into() {
Ok(v) => v,
Err(_) => {
// The number of samples is too large to fit in an i32.
return None;
}
};
let ptr = unsafe {
sys::SherpaOnnxOfflineSpeakerDiarizationProcess(
self.ptr,
samples.as_ptr(),
n,
)
};There was a problem hiding this comment.
Pull request overview
Adds a Rust API surface (plus FFI bindings and examples) for running offline speaker diarization using pyannote-based segmentation and speaker embeddings.
Changes:
- Introduces
OfflineSpeakerDiarizationRust wrapper + config/segment/result types. - Adds
sherpa-onnx-sysFFI bindings for offline speaker diarization C APIs. - Adds a runnable Rust example + CI script hook to exercise the new API.
Reviewed changes
Copilot reviewed 9 out of 9 changed files in this pull request and generated 3 comments.
Show a summary per file
| File | Description |
|---|---|
| sherpa-onnx/rust/sherpa-onnx/src/speaker_embedding.rs | Exposes to_sys for crate-internal reuse by diarization config serialization |
| sherpa-onnx/rust/sherpa-onnx/src/offline_speaker_diarization.rs | New safe-ish Rust wrapper around the diarization FFI and its configuration/result types |
| sherpa-onnx/rust/sherpa-onnx/src/lib.rs | Exposes the new diarization module from the Rust crate |
| sherpa-onnx/rust/sherpa-onnx-sys/src/offline_speaker_diarization.rs | New raw FFI declarations/structs for diarization APIs |
| sherpa-onnx/rust/sherpa-onnx-sys/src/lib.rs | Re-exports the new diarization sys module |
| rust-api-examples/run-offline-speaker-diarization.sh | Script to download models/sample and run the new example |
| rust-api-examples/examples/offline_speaker_diarization.rs | New example demonstrating offline speaker diarization |
| rust-api-examples/README.md | Documents the new example entry |
| .github/scripts/test-rust.sh | Adds the new example script to CI runs |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| unsafe impl Send for OfflineSpeakerDiarization {} | ||
|
|
There was a problem hiding this comment.
unsafe impl Send is a strong thread-safety claim. Unless the underlying sys::OfflineSpeakerDiarization is explicitly documented as safe to move across threads (and its destructor / set_config / process are thread-safe with respect to the handle), this can introduce undefined behavior in multithreaded use. Consider removing Send, or (if the C API guarantees it) add a comment pointing to that guarantee and consider whether Sync is also appropriate or should remain intentionally absent.
| unsafe impl Send for OfflineSpeakerDiarization {} |
| let ptr = unsafe { | ||
| sys::SherpaOnnxOfflineSpeakerDiarizationProcess( | ||
| self.ptr, | ||
| samples.as_ptr(), | ||
| samples.len() as i32, |
There was a problem hiding this comment.
Casting samples.len() from usize to i32 can truncate on large inputs, potentially passing a negative/incorrect n to the C API. Use a checked conversion (e.g., try_into()) and return None (or an error) when the buffer length doesn't fit into i32.
| let ptr = unsafe { | |
| sys::SherpaOnnxOfflineSpeakerDiarizationProcess( | |
| self.ptr, | |
| samples.as_ptr(), | |
| samples.len() as i32, | |
| // Ensure the buffer length fits into i32 before passing it to the C API. | |
| let n: i32 = match samples.len().try_into() { | |
| Ok(v) => v, | |
| Err(_) => { | |
| // Length does not fit into i32; do not call the C API. | |
| return None; | |
| } | |
| }; | |
| let ptr = unsafe { | |
| sys::SherpaOnnxOfflineSpeakerDiarizationProcess( | |
| self.ptr, | |
| samples.as_ptr(), | |
| n, |
| fi | ||
|
|
||
| if [ ! -f ./3dspeaker_speech_eres2net_base_sv_zh-cn_3dspeaker_16k.onnx ]; then | ||
| curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/speaker-recongition-models/3dspeaker_speech_eres2net_base_sv_zh-cn_3dspeaker_16k.onnx |
There was a problem hiding this comment.
The path segment speaker-recongition-models appears misspelled (likely speaker-recognition-models). If the release tag uses the correct spelling, this will 404 and break CI. Please verify the release URL and correct the typo if needed.
| curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/speaker-recongition-models/3dspeaker_speech_eres2net_base_sv_zh-cn_3dspeaker_16k.onnx | |
| curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/speaker-recognition-models/3dspeaker_speech_eres2net_base_sv_zh-cn_3dspeaker_16k.onnx |
Summary by CodeRabbit
New Features
Documentation
Tests