Skip to content

Add Rust API for speaker diarization - #3370

Merged
csukuangfj merged 1 commit into
k2-fsa:masterfrom
csukuangfj:rust-speaker-diarization
Mar 20, 2026
Merged

csukuangfj merged 1 commit into
k2-fsa:masterfrom
csukuangfj:rust-speaker-diarization

Conversation

@csukuangfj

@csukuangfj csukuangfj commented Mar 20, 2026 •

Copy link
Copy Markdown
Collaborator

Summary by CodeRabbit

  • New Features

    • Added offline speaker diarization capability to identify and separate individual speakers in audio files using advanced speaker segmentation and embedding models.
  • Documentation

    • Added comprehensive example demonstrating speaker diarization functionality with accompanying script for automated model download and example execution.
  • Tests

    • Updated CI test suite to validate offline speaker diarization functionality.

Copilot AI review requested due to automatic review settings March 20, 2026 03:40
@dosubot dosubot Bot added the size:L This PR changes 100-499 lines, ignoring generated files. label Mar 20, 2026
@coderabbitai

coderabbitai Bot commented Mar 20, 2026 •

Copy link
Copy Markdown

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 1b7a65a4-f4c6-4461-9226-6d2ed25a5482

📥 Commits

Reviewing files that changed from the base of the PR and between 9629e18 and 7421551.

📒 Files selected for processing (9)
  • .github/scripts/test-rust.sh
  • rust-api-examples/README.md
  • rust-api-examples/examples/offline_speaker_diarization.rs
  • rust-api-examples/run-offline-speaker-diarization.sh
  • sherpa-onnx/rust/sherpa-onnx-sys/src/lib.rs
  • sherpa-onnx/rust/sherpa-onnx-sys/src/offline_speaker_diarization.rs
  • sherpa-onnx/rust/sherpa-onnx/src/lib.rs
  • sherpa-onnx/rust/sherpa-onnx/src/offline_speaker_diarization.rs
  • sherpa-onnx/rust/sherpa-onnx/src/speaker_embedding.rs

📝 Walkthrough

Walkthrough

This PR introduces offline speaker diarization functionality to the Rust API. It adds FFI bindings, safe Rust wrappers with resource management, a complete example with automated model downloading, documentation, and CI integration for testing the new feature.

Changes

Cohort / File(s) Summary
CI and Documentation
.github/scripts/test-rust.sh, rust-api-examples/README.md
Added CI execution of offline speaker diarization example script with cleanup. Documented new Example 33 in README with description and usage instructions.
FFI Bindings
sherpa-onnx/rust/sherpa-onnx-sys/src/lib.rs, sherpa-onnx/rust/sherpa-onnx-sys/src/offline_speaker_diarization.rs
Declared C-compatible configuration structs for pyannote segmentation, embedding extraction, and fast clustering. Defined opaque handles and full extern "C" interface for creating/destroying diarization objects, processing audio, and querying results.
High-level Rust Wrappers
sherpa-onnx/rust/sherpa-onnx/src/lib.rs, sherpa-onnx/rust/sherpa-onnx/src/offline_speaker_diarization.rs, sherpa-onnx/rust/sherpa-onnx/src/speaker_embedding.rs
Implemented safe Rust wrappers with configuration structs, Default and to_sys conversions, resource management via Drop, and APIs for creating diarization objects, processing samples, and retrieving sorted segments. Made SpeakerEmbeddingExtractorConfig::to_sys public to crate scope for configuration composition.
Example Code and Scripts
rust-api-examples/examples/offline_speaker_diarization.rs, rust-api-examples/run-offline-speaker-diarization.sh
Added runnable Rust example demonstrating configuration, initialization, audio processing, and result iteration. Included Bash script to verify/download required ONNX models and sample audio, then execute the example.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related PRs

  • k2-fsa/sherpa-onnx#3209: Modifies .github/scripts/test-rust.sh in the same manner to add execution of new example-run scripts for additional offline ASR examples.

Suggested labels

size:XL

Poem

🐰 A diarization dream comes into view,
With FFI bridges and wrappers so true,
Safe Rust embraces the speaker's refrain,
While four voices harmonize—speakers maintain!
Models download, samples sing clear,
Another great feature for Rust devs to cheer! 🎤

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
📝 Coding Plan
  • Generate coding plan for human review comments

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist

Copy link
Copy Markdown

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly expands the Rust API capabilities by introducing comprehensive support for offline speaker diarization. This new feature allows users to identify and segment speech by different speakers in an audio recording, providing a robust tool for multi-speaker analysis. The changes include new FFI bindings, a high-level Rust API wrapper, and a runnable example, making it accessible for Rust developers.

Highlights

  • New Rust API for Offline Speaker Diarization: Introduced a comprehensive Rust API for performing offline speaker diarization, enabling the identification and segmentation of speech by different speakers in an audio recording.
  • New Example and Documentation: Added a new Rust example (offline_speaker_diarization.rs) and a corresponding run-offline-speaker-diarization.sh script to demonstrate the usage of the new API, along with updates to the README.md.
  • CI/CD Integration: Updated the CI/CD test script (.github/scripts/test-rust.sh) to include testing and cleanup for the new speaker diarization functionality.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@csukuangfj
csukuangfj merged commit cdf8e60 into k2-fsa:master Mar 20, 2026
1 of 2 checks passed
@csukuangfj
csukuangfj deleted the rust-speaker-diarization branch March 20, 2026 03:41

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a Rust API for speaker diarization, complete with FFI bindings, a high-level safe wrapper, and a usage example. The implementation is well-structured. I have a couple of suggestions for improvement: one to correct a typo in a download URL within a shell script, and another to add a safeguard against potential integer overflow when handling audio sample lengths in the Rust wrapper.

fi

if [ ! -f ./3dspeaker_speech_eres2net_base_sv_zh-cn_3dspeaker_16k.onnx ]; then
curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/speaker-recongition-models/3dspeaker_speech_eres2net_base_sv_zh-cn_3dspeaker_16k.onnx

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

There's a typo in the URL: recongition should be recognition. While the current URL with the typo works because a release with the typo exists, it's better to use the corrected URL for future-proofing and clarity, as a corrected release tag also exists.

Suggested change
curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/speaker-recongition-models/3dspeaker_speech_eres2net_base_sv_zh-cn_3dspeaker_16k.onnx
curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/speaker-recognition-models/3dspeaker_speech_eres2net_base_sv_zh-cn_3dspeaker_16k.onnx

Comment on lines +153 to +159
let ptr = unsafe {
sys::SherpaOnnxOfflineSpeakerDiarizationProcess(
self.ptr,
samples.as_ptr(),
samples.len() as i32,
)
};

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The cast samples.len() as i32 can truncate if the number of samples is larger than i32::MAX. This could lead to incorrect data being processed or a panic. While it's unlikely to happen with typical audio files (it would require an audio file of over 37 hours at 16kHz), it's safer to handle this potential overflow. Using try_into() provides a more robust way to perform the conversion and handle the error gracefully.

        let n = match samples.len().try_into() {
            Ok(v) => v,
            Err(_) => {
                // The number of samples is too large to fit in an i32.
                return None;
            }
        };
        let ptr = unsafe {
            sys::SherpaOnnxOfflineSpeakerDiarizationProcess(
                self.ptr,
                samples.as_ptr(),
                n,
            )
        };

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds a Rust API surface (plus FFI bindings and examples) for running offline speaker diarization using pyannote-based segmentation and speaker embeddings.

Changes:

  • Introduces OfflineSpeakerDiarization Rust wrapper + config/segment/result types.
  • Adds sherpa-onnx-sys FFI bindings for offline speaker diarization C APIs.
  • Adds a runnable Rust example + CI script hook to exercise the new API.

Reviewed changes

Copilot reviewed 9 out of 9 changed files in this pull request and generated 3 comments.

Show a summary per file
File Description
sherpa-onnx/rust/sherpa-onnx/src/speaker_embedding.rs Exposes to_sys for crate-internal reuse by diarization config serialization
sherpa-onnx/rust/sherpa-onnx/src/offline_speaker_diarization.rs New safe-ish Rust wrapper around the diarization FFI and its configuration/result types
sherpa-onnx/rust/sherpa-onnx/src/lib.rs Exposes the new diarization module from the Rust crate
sherpa-onnx/rust/sherpa-onnx-sys/src/offline_speaker_diarization.rs New raw FFI declarations/structs for diarization APIs
sherpa-onnx/rust/sherpa-onnx-sys/src/lib.rs Re-exports the new diarization sys module
rust-api-examples/run-offline-speaker-diarization.sh Script to download models/sample and run the new example
rust-api-examples/examples/offline_speaker_diarization.rs New example demonstrating offline speaker diarization
rust-api-examples/README.md Documents the new example entry
.github/scripts/test-rust.sh Adds the new example script to CI runs

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +128 to +129
unsafe impl Send for OfflineSpeakerDiarization {}

Copilot AI Mar 20, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

unsafe impl Send is a strong thread-safety claim. Unless the underlying sys::OfflineSpeakerDiarization is explicitly documented as safe to move across threads (and its destructor / set_config / process are thread-safe with respect to the handle), this can introduce undefined behavior in multithreaded use. Consider removing Send, or (if the C API guarantees it) add a comment pointing to that guarantee and consider whether Sync is also appropriate or should remain intentionally absent.

Suggested change
unsafe impl Send for OfflineSpeakerDiarization {}

Copilot uses AI. Check for mistakes.
Comment on lines +153 to +157
let ptr = unsafe {
sys::SherpaOnnxOfflineSpeakerDiarizationProcess(
self.ptr,
samples.as_ptr(),
samples.len() as i32,

Copilot AI Mar 20, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Casting samples.len() from usize to i32 can truncate on large inputs, potentially passing a negative/incorrect n to the C API. Use a checked conversion (e.g., try_into()) and return None (or an error) when the buffer length doesn't fit into i32.

Suggested change
let ptr = unsafe {
sys::SherpaOnnxOfflineSpeakerDiarizationProcess(
self.ptr,
samples.as_ptr(),
samples.len() as i32,
// Ensure the buffer length fits into i32 before passing it to the C API.
let n: i32 = match samples.len().try_into() {
Ok(v) => v,
Err(_) => {
// Length does not fit into i32; do not call the C API.
return None;
}
};
let ptr = unsafe {
sys::SherpaOnnxOfflineSpeakerDiarizationProcess(
self.ptr,
samples.as_ptr(),
n,

Copilot uses AI. Check for mistakes.
fi

if [ ! -f ./3dspeaker_speech_eres2net_base_sv_zh-cn_3dspeaker_16k.onnx ]; then
curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/speaker-recongition-models/3dspeaker_speech_eres2net_base_sv_zh-cn_3dspeaker_16k.onnx

Copilot AI Mar 20, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The path segment speaker-recongition-models appears misspelled (likely speaker-recognition-models). If the release tag uses the correct spelling, this will 404 and break CI. Please verify the release URL and correct the typo if needed.

Suggested change
curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/speaker-recongition-models/3dspeaker_speech_eres2net_base_sv_zh-cn_3dspeaker_16k.onnx
curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/speaker-recognition-models/3dspeaker_speech_eres2net_base_sv_zh-cn_3dspeaker_16k.onnx

Copilot uses AI. Check for mistakes.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L This PR changes 100-499 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants