Skip to content

Add C API for source separation - #3404

Merged
csukuangfj merged 1 commit into
k2-fsa:masterfrom
csukuangfj:c-api-source-separation
Mar 25, 2026
Merged

csukuangfj merged 1 commit into
k2-fsa:masterfrom
csukuangfj:c-api-source-separation

Conversation

@csukuangfj

@csukuangfj csukuangfj commented Mar 25, 2026 •

Copy link
Copy Markdown
Collaborator

Summary by CodeRabbit

  • New Features

    • Added offline source separation functionality supporting Spleeter and UVR models
    • Added multi-channel audio file reading and writing support
    • Expanded C API with source separation operations and multi-channel audio utilities
  • Tests

    • Added CI/CD pipeline tests for source separation example programs

@dosubot dosubot Bot added the size:XL This PR changes 500-999 lines, ignoring generated files. label Mar 25, 2026
@gemini-code-assist

Copy link
Copy Markdown

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly expands the sherpa-onnx C API by integrating offline source separation capabilities. It enables users to perform audio source separation using popular models like Spleeter and UVR directly through a C interface, complete with multi-channel audio input/output and example usage. This addition enhances the library's utility for audio processing applications requiring vocal/instrumental separation.

Highlights

  • New C API for Source Separation: Introduced a comprehensive C API for offline source separation, allowing developers to integrate source separation capabilities into C/C++ applications.
  • Spleeter and UVR Model Support: The new C API supports both Spleeter (2-stems) and UVR (MDX-Net) models, providing flexibility for different source separation tasks.
  • Multi-Channel Wave Handling: Added new C API functions for reading and writing multi-channel WAVE files, essential for processing and outputting separated audio stems.
  • C API Examples: Provided new C API examples (source-separation-spleeter-c-api.c and source-separation-uvr-c-api.c) demonstrating how to use the source separation functionality with Spleeter and UVR models.
Ignored Files
  • Ignored by pattern: .github/workflows/** (1)
    • .github/workflows/c-api.yaml
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@coderabbitai

coderabbitai Bot commented Mar 25, 2026 •

Copy link
Copy Markdown

Caution

Review failed

Pull request was closed or merged during review

📝 Walkthrough

Walkthrough

This PR introduces offline source separation capabilities to Sherpa-ONNX's C API, supporting both Spleeter and UVR models. New C API functions handle multi-channel audio I/O, source separation configuration, engine creation/destruction, and processing. Two example C programs demonstrate end-to-end usage with each model. Build system, CI workflows, and symbol exports are updated accordingly.

Changes

Cohort / File(s) Summary
Workflow & Build
.github/workflows/c-api.yaml, c-api-examples/CMakeLists.txt
Added CI steps to compile and test source separation examples (Spleeter and UVR), with artifact uploads. Extended CMake to build two new executables linked against the C API library.
C API Examples
c-api-examples/source-separation-spleeter-c-api.c, c-api-examples/source-separation-uvr-c-api.c
Two new example programs demonstrating offline source separation: Spleeter separates audio into vocals and accompaniment; UVR separates into vocals and non-vocals. Both read input WAV, process via C API, and write output stems.
C API Surface
sherpa-onnx/c-api/c-api.h, sherpa-onnx/c-api/c-api.cc
Added comprehensive offline source separation C API: configuration structs, engine lifecycle (create/destroy), processing functions, metadata queries, and multi-channel waveform I/O (read/write/free). Includes OHOS variant for resource management.
Symbol Exports
sherpa-onnx/c-api/sherpa-onnx-symbols-c.exp
Exported nine new symbols for source separation functions and multi-channel wave operations (create, destroy, process, get stems, get sample rate, read/write/free wave).
Multi-Channel Wave I/O
sherpa-onnx/csrc/wave-writer.h, sherpa-onnx/csrc/wave-writer.cc
Added WriteWaveMultiChannel overloads for file and buffer output supporting interleaved multi-channel float samples. Fixed sample conversion to use symmetric clamp instead of one-sided min, preventing negative overflow.

Sequence Diagram(s)

sequenceDiagram
    participant Example as Example Program
    participant CAPI as C API Layer
    participant Engine as Source Separation<br/>Engine
    participant WaveIO as Wave I/O

    Example->>WaveIO: SherpaOnnxReadWaveMultiChannel(input.wav)
    WaveIO-->>Example: Multi-channel wave data
    
    Example->>CAPI: SherpaOnnxCreateOfflineSourceSeparation(config)
    CAPI->>Engine: Initialize with model path
    Engine-->>CAPI: Engine instance
    CAPI-->>Example: Engine handle

    Example->>CAPI: SherpaOnnxOfflineSourceSeparationProcess(engine, samples)
    CAPI->>Engine: Process multi-channel input
    Engine-->>CAPI: Separation output (stems)
    CAPI-->>Example: Output structure with stems

    Example->>WaveIO: SherpaOnnxWriteWaveMultiChannel(stem[0], output.wav)
    WaveIO-->>Example: File written
    
    Example->>WaveIO: SherpaOnnxWriteWaveMultiChannel(stem[1], output.wav)
    WaveIO-->>Example: File written

    Example->>CAPI: SherpaOnnxDestroySourceSeparationOutput(output)
    CAPI-->>Example: Resources freed
    
    Example->>CAPI: SherpaOnnxDestroyOfflineSourceSeparation(engine)
    CAPI->>Engine: Cleanup
    Engine-->>CAPI: Engine destroyed
    CAPI-->>Example: Cleanup complete
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Poem

🐰 With whiskers held high, we separate the sound,
Vocals and beats now neatly unbound,
Through Spleeter's fine mesh and UVR's keen eye,
Multi-channel melodies dance in the sky! 🎵

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 9.09% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'Add C API for source separation' directly and clearly summarizes the main change—the PR adds a new C API surface for offline source separation with support for multiple models (Spleeter and UVR) and corresponding example programs.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces C API examples for source separation using Spleeter and UVR models. It includes new C API functions for reading and writing multi-channel wave files, managing the source separation engine, and processing audio. The changes also include updates to CMakeLists.txt to build these examples and improvements to wave file writing with std::clamp. The review comments suggest improving maintainability in the example files by replacing magic numbers with sizeof calculations for array sizes.


// Write each stem to a separate multi-channel wave file.
const char *stem_names[] = {"vocals", "accompaniment"};
for (int32_t s = 0; s < output->num_stems && s < 2; ++s) {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

To improve maintainability and avoid using a magic number (2), it's better to calculate the number of stem names directly from the stem_names array using sizeof.

Suggested change
for (int32_t s = 0; s < output->num_stems && s < 2; ++s) {
for (int32_t s = 0; s < output->num_stems && s < sizeof(stem_names) / sizeof(stem_names[0]); ++s) {


// Write each stem to a separate multi-channel wave file.
const char *stem_names[] = {"uvr-vocals", "uvr-non-vocals"};
for (int32_t s = 0; s < output->num_stems && s < 2; ++s) {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

To improve maintainability and avoid using a magic number (2), it's better to calculate the number of stem names directly from the stem_names array using sizeof.

Suggested change
for (int32_t s = 0; s < output->num_stems && s < 2; ++s) {
for (int32_t s = 0; s < output->num_stems && s < sizeof(stem_names) / sizeof(stem_names[0]); ++s) {

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XL This PR changes 500-999 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant