Skip to content

Improve C/C++ API documentation, examples, and CI - #3626

Merged
csukuangfj merged 8 commits into
k2-fsa:masterfrom
csukuangfj:c-api-doc
May 20, 2026
Merged

csukuangfj merged 8 commits into
k2-fsa:masterfrom
csukuangfj:c-api-doc

Conversation

@csukuangfj

@csukuangfj csukuangfj commented May 19, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Comprehensive C API documentation and example improvements for sherpa-onnx.

Documentation

  • Add 13 Doxygen .dox files in sherpa-onnx/c-api/docs/ covering all model families: offline ASR (20 models), online ASR (5 models), TTS (7 models), VAD (2 models), audio tagging
    (2 models), punctuation, speech enhancement (GTCRN + DPDFNet), source separation (Spleeter + UVR), speaker diarization, speaker embedding, spoken language identification,
    keyword spotting, and linear resampler
  • Add @see cross-references to all 66 key structs/functions in c-api.h and to all 13 dox files
  • Add multi-model doc examples to c-api.h creator APIs: TTS (7 families), offline ASR (16 families), online ASR (4 families), VAD (Silero + Ten), audio tagging (Zipformer +
    CED), speech denoiser (GTCRN + DPDFNet), source separation (Spleeter + UVR)
  • Fix doc errors: wrong function names (SherpaOnnxDestroyWave → SherpaOnnxFreeWave), wrong function signatures, typos (offline-sepaker → offline-speaker), wrong struct field
    paths, bad formatting
  • Update punctuation model to int8 version (sherpa-onnx-punct-ct-transformer-zh-en-vocab272727-2024-04-12-int8)
  • Update mainpage.md with comprehensive example links organized by feature

New C++ API wrappers (cxx-api.h / cxx-api.cc)

  • SpokenLanguageIdentification — Create, CreateStream, Compute
  • SpeakerEmbeddingExtractor — Create, Dim, CreateStream, IsReady, ComputeEmbedding
  • SpeakerEmbeddingManager — Create, Add, AddList, Remove, Search, GetBestMatches, Verify, Contains, NumSpeakers, GetAllSpeakers
  • OfflineSpeakerDiarization — Create, GetSampleRate, SetConfig, Process (with progress callback)

New examples

C API:

  • nemo-ctc-c-api.c — NeMo CTC model
  • nemo-giga-am-v2-c-api.c — GigaAM v2 Russian model
  • streaming-nemotron-c-api.c — Nemotron streaming model

C++ API:

  • paraformer-cxx-api.cc — offline Paraformer
  • streaming-paraformer-cxx-api.cc — streaming Paraformer
  • nemo-ctc-cxx-api.cc — NeMo CTC
  • nemo-giga-am-v2-cxx-api.cc — GigaAM v2 Russian
  • streaming-nemotron-cxx-api.cc — Nemotron streaming
  • offline-tts-piper-cxx-api.cc — Piper VITS TTS
  • offline-speaker-diarization-cxx-api.cc — speaker diarization
  • speaker-identification-cxx-api.cc — speaker embedding/identification
  • spoken-language-identification-cxx-api.cc — spoken language ID

CI improvements

  • Refactor c-api.yaml from 1 monolithic job into 6 parallel jobs: build, test-asr-offline, test-asr-streaming, test-tts, test-vad-punct, test-other
  • Refactor cxx-api.yaml from 1 monolithic job into 6 parallel jobs: same structure
  • Add 15 new CI tests (10 C API + 5 C++ API) covering: offline punctuation, audio tagging, speaker identification, spoken language ID, speaker diarization, NeMo Parakeet,
    streaming CTC buffered tokens, streaming paraformer buffered tokens, streaming zipformer buffered tokens hotwords, keywords spotter buffered tokens, NeMo CTC, Paraformer CXX,
    Offline TTS Piper, streaming paraformer CXX
  • Fix CI issues: correct test wav filename for GigaAM v2 (example.wav), fix sr-data archive directory name (sr-data-main), create hotwords file for buffered tokens test, copy
    keywords file for KWS test

Code cleanup

  • Refactor 20 C API examples to remove redundant memset on sub-structs (memset only on top-level config)

Summary by CodeRabbit

  • New Features

    • Added many new model families and features: expanded ASR (offline/streaming), TTS (including Piper), VAD, punctuation, source separation, audio tagging, keyword spotting, speech enhancement, speaker embedding/identification, offline speaker diarization, and spoken-language identification.
  • Tests

    • CI test matrix reorganized and expanded to cover offline/streaming ASR, TTS, VAD, punctuation, speaker features, and other audio capabilities with centralized build artifacts.
  • Documentation

    • Large set of new and updated docs and example guides for models, APIs, and workflows.

Review Change Stack

@dosubot dosubot Bot added the size:XXL This PR changes 1000+ lines, ignoring generated files. label May 19, 2026
@coderabbitai

coderabbitai Bot commented May 19, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

Refactors CI workflows to centralize build artifacts and library paths; adds many C/C++ example programs and CMake targets; consolidates example config patterns; introduces C++ audio-analysis RAII wrappers (language ID, speaker embedding/manager, diarization); and expands Doxygen docs with model-specific pages and header cross-links.

Changes

CI Workflows and Build Infrastructure

Layer / File(s) Summary
Build job consolidation and artifact handling
.github/workflows/c-api.yaml, .github/workflows/cxx-api.yaml
Renamed primary build jobs to build, added artifact uploads, and centralized library path setup via $GITHUB_ENV for downstream test jobs.
Offline ASR and streaming test updates
.github/workflows/c-api.yaml, .github/workflows/cxx-api.yaml
Updated many test steps to download/extract newer model artifacts, reorganized coverage across test-asr-offline, test-asr-streaming, and added test-vad-punct/test-other; removed repeated per-step dependency diagnostics.
Test scripts
.github/scripts/*, .github/scripts/test-python.sh
Updated SenseVoice and punctuation model download URLs/cleanup in CI helper scripts to use int8-quantized and revised tarball names; adjusted subtitle-generation calls to use model.int8.onnx.

C API Examples

Layer / File(s) Summary
C API example build targets
c-api-examples/CMakeLists.txt
Added new C API example executables (nemo-giga-am-v2-c-api, nemo-ctc-c-api, streaming-nemotron-c-api) linked to sherpa-onnx-c-api.
Configuration pattern consolidation in existing C API examples
c-api-examples/*.c
Refactored many examples to populate recognizer_config.model_config fields directly instead of building intermediate ModelConfig structs, simplifying example initialization across ASR families.
New C API example programs
c-api-examples/nemo-ctc-c-api.c, c-api-examples/nemo-giga-am-v2-c-api.c, c-api-examples/streaming-nemotron-c-api.c
Added end-to-end C examples demonstrating NeMo CTC, NeMo GigaAM v2, and streaming Nemotron workflows.
Punctuation model variant update
c-api-examples/add-punctuation-c-api.c
Switched example to int8 CT-Transformer model and updated model.int8.onnx path.

C++ API Examples

Layer / File(s) Summary
C++ API example build targets
cxx-api-examples/CMakeLists.txt
Added multiple new executable targets for Nemo/Paraformer variants, streaming Nemotron, offline TTS Piper, and audio-analysis examples.
New C++ example programs
cxx-api-examples/*.cc
Added offline/streaming ASR and TTS examples and new audio-analysis examples (spoken language ID, speaker identification, offline diarization).

C++ API New Audio Analysis Features

Layer / File(s) Summary
C++ API headers for audio analysis features
sherpa-onnx/c-api/cxx-api.h
Introduced types and RAII wrappers for SpokenLanguageIdentification, SpeakerEmbeddingExtractor, SpeakerEmbeddingManager, and OfflineSpeakerDiarization, plus supporting config/result types and a progress callback alias.
C++ API implementation for audio analysis
sherpa-onnx/c-api/cxx-api.cc
Implemented wrappers that convert C++ configs to C structs, manage C object lifecycles, translate C results to C++ types, handle callback bridging, and free C-allocated auxiliary data.

API Documentation Infrastructure

Layer / File(s) Summary
Model-specific Doxygen documentation pages
sherpa-onnx/c-api/docs/*.dox
Added 13 new Doxygen pages covering offline/online ASR, TTS, VAD, audio tagging, punctuation, speech enhancement, source separation, speaker diarization/embedding, spoken-language ID, keyword spotting, and the resampler.
Doxygen configuration and mainpage updates
sherpa-onnx/c-api/Doxyfile, sherpa-onnx/c-api/mainpage.md
Extended Doxyfile INPUT and reorganized mainpage with a “Model-specific documentation” section and grouped example program references.
C API header documentation cross-linking
sherpa-onnx/c-api/c-api.h
Inserted Doxygen @see references across many API groups to link related create/destroy helpers and result/stream helpers.
Exported C API symbol ordering
sherpa-onnx/c-api/sherpa-onnx-symbols-c.exp
Reordered exported symbol entries to group related create/destroy pairs and related stream/option functions.

🎯 4 (Complex) | ⏱️ ~60 minutes

Possibly related PRs:

"A rabbit typed at midnight, neat and spry,
CI carrots stacked in orderly supply,
Examples hopped to C and C++, ready to play,
Language, speakers, diarize — new functions on display,
Docs and tests aligned — a burrow bright with cheer!"

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request adds several new C and C++ API examples for various models, including NeMo CTC, GigaAM v2, streaming Nemotron, and others, along with corresponding documentation updates. My review identified a compilation error in the Piper TTS example where a non-existent field was accessed, as well as some minor code quality improvements regarding namespace usage and error handling in the speaker identification example.

Comment thread cxx-api-examples/offline-tts-piper-cxx-api.cc Outdated

#include "sherpa-onnx/c-api/cxx-api.h"

using namespace sherpa_onnx::cxx; // NOLINT

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Using using namespace at the global scope in a source file is generally discouraged as it can lead to name collisions. It is better to move it inside main() to be consistent with other examples in this pull request.

Wave wave = ReadWave(wav_filename);
if (wave.samples.empty()) {
std::cerr << "Failed to read " << wav_filename << "\n";
exit(-1);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Using exit(-1) is non-standard. For failure, it is recommended to use exit(EXIT_FAILURE) from or return a non-zero value from main(). This also applies to line 42.

Suggested change
exit(-1);
exit(EXIT_FAILURE);

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 5

🧹 Nitpick comments (6)
.github/workflows/c-api.yaml (1)

1650-1665: 💤 Low value

Remove disabled step with constant if: false condition.

The static analysis tool (actionlint) flags the if: false as a constant expression. If the ffmpeg test is not intended to run, it's cleaner to remove or comment out the entire step rather than leaving it with a hardcoded false condition.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In @.github/workflows/c-api.yaml around lines 1650 - 1665, Remove the disabled
GitHub Actions step named "Test ffmpeg" that uses the constant condition if:
false (the whole YAML block starting with the step name "Test ffmpeg"); either
delete the entire step or comment it out so actionlint no longer sees a constant
expression, and keep any needed test logic elsewhere or behind a proper
conditional matrix entry or input-driven if condition.
cxx-api-examples/streaming-paraformer-cxx-api.cc (1)

17-22: ⚡ Quick win

Include <algorithm> explicitly for std::min.

Line 71 uses std::min, which requires <algorithm>. Avoid relying on transitive includes.

Proposed fix
 `#include` <chrono>  // NOLINT
+#include <algorithm>
 `#include` <cstdio>
 `#include` <iostream>
 `#include` <string>
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@cxx-api-examples/streaming-paraformer-cxx-api.cc` around lines 17 - 22, The
code uses std::min (referenced at the call to std::min in
streaming-paraformer-cxx-api.cc) but doesn't include <algorithm>, so add an
explicit `#include` <algorithm> to the top include block (alongside <chrono>,
<cstdio>, <iostream>, <string>) to avoid relying on transitive includes and
ensure std::min is declared.
cxx-api-examples/offline-tts-piper-cxx-api.cc (2)

62-62: 💤 Low value

Clarify silence_scale initialization.

gen_config.silence_scale is assigned from config.silence_scale, but config.silence_scale was never explicitly set. This relies on default initialization, which may be intentional but is unclear to readers.

♻️ Consider explicit initialization

Either set config.silence_scale explicitly before this line, or set gen_config.silence_scale directly:

-gen_config.silence_scale = config.silence_scale;
+gen_config.silence_scale = 1.0;  // or appropriate default
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@cxx-api-examples/offline-tts-piper-cxx-api.cc` at line 62, The assignment
gen_config.silence_scale = config.silence_scale uses config.silence_scale
without it being explicitly initialized; either initialize config.silence_scale
to the intended default value earlier (e.g., set config.silence_scale =
<desired_value> before the assignment) or avoid relying on config and set
gen_config.silence_scale directly to the intended default/constant when
constructing gen_config so the value is explicit and readable.

58-58: ⚡ Quick win

Add error checking after TTS creation.

The code doesn't verify that OfflineTts::Create succeeded. If the model files are missing or invalid, subsequent operations on tts may fail or crash.

🛡️ Proposed validation check
 auto tts = OfflineTts::Create(config);
+if (!tts.Get()) {
+  std::cerr << "Failed to create TTS model\n";
+  return -1;
+}
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@cxx-api-examples/offline-tts-piper-cxx-api.cc` at line 58, The code calls
OfflineTts::Create(config) but doesn't check the return value; verify that the
returned pointer or optional (variable tts) is valid before using it and handle
failure by logging an error and exiting or returning an error status. Locate the
OfflineTts::Create call and after it check if tts is null/empty or indicates
failure (depending on its return type) and then call process logger/print an
error like "Failed to create OfflineTts" along with any error info and
abort/return to avoid subsequent dereference of tts.
cxx-api-examples/speaker-identification-cxx-api.cc (1)

62-62: ⚡ Quick win

Add error checking after manager creation.

For consistency with the extractor creation (lines 54-57), verify that SpeakerEmbeddingManager::Create succeeded before proceeding with enrollment operations.

🛡️ Proposed validation check
 SpeakerEmbeddingManager manager = SpeakerEmbeddingManager::Create(dim);
+if (!manager.Get()) {
+  std::cerr << "Failed to create speaker embedding manager\n";
+  return -1;
+}
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@cxx-api-examples/speaker-identification-cxx-api.cc` at line 62, Verify the
result of SpeakerEmbeddingManager::Create(dim) just like the extractor creation:
check that the returned manager is valid (non-null or has_value depending on its
return type) before using it for enrollment, and if creation failed log an error
via the same logger and exit/return early; update references to manager in
enrollment code to assume a valid manager only after this validation.
sherpa-onnx/c-api/cxx-api.h (1)

1867-1888: ⚡ Quick win

Add buffer-size and termination documentation to the C++ wrapper header.

The C API documentation (in c-api.h) clearly specifies that AddList() expects a NULL-terminated pointer array, AddListFlattened()'s n parameter is the embedding count (not float count), and Add()/Search()/Verify() operate on vectors of exactly dim elements. However, the C++ wrapper header comments at lines 1869–1887 omit these details, leaving C++ callers to either guess the contract or dig into the C API docs. Add explicit dimension and termination requirements to the C++ method comments to match the clarity of the C API layer.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@sherpa-onnx/c-api/cxx-api.h` around lines 1867 - 1888, Update the C++ wrapper
comments for the SpeakerManager methods to explicitly document buffer sizes and
termination: for Add(const std::string &name, const float *v) state that v must
point to exactly dim floats (embedding length); for AddList(const std::string
&name, const float **v) state that v is a NULL-terminated array of pointers,
each pointing to an embedding of exactly dim floats; for AddListFlattened(const
std::string &name, const float *v, int32_t n) state that v is a flat array
containing n embeddings (so length = n * dim) and that n is the number of
embeddings (not number of floats); and for Search(const float *v, float
threshold) and Verify(const std::string &name, const float *v, float threshold)
state that v must point to exactly dim floats; update the brief comments above
those method declarations (Add, AddList, AddListFlattened, Search, Verify) to
include these requirements.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@cxx-api-examples/offline-speaker-diarization-cxx-api.cc`:
- Line 15: The wget command in offline-speaker-diarization-cxx-api.cc contains a
misspelling in the release path: replace "speaker-recongition-models" with
"speaker-recognition-models" in the commented download URL; update the comment
line that starts with "wget
https://github.com/k2-fsa/sherpa-onnx/releases/download/..." so the path is
corrected to the proper release directory name.

In `@sherpa-onnx/c-api/cxx-api.cc`:
- Around line 1448-1454: The bridge DiarizationProgressCallback currently
dereferences and invokes OfflineSpeakerDiarizationProgressCallback
unconditionally and lets exceptions cross the C boundary; update the bridge
(DiarizationProgressCallback) to first check that the pointer is non-null and
that the std::function is non-empty before calling, and wrap the invocation in a
try/catch that swallows or logs exceptions so they don't unwind past the C
callback boundary. Additionally, when registering the callback before calling
SherpaOnnxOfflineSpeakerDiarizationProcessWithCallback, only pass the pointer if
the supplied OfflineSpeakerDiarizationProgressCallback is non-empty (otherwise
pass nullptr) so the bridge can skip invocation safely. Ensure you reference the
symbols DiarizationProgressCallback, OfflineSpeakerDiarizationProgressCallback,
and SherpaOnnxOfflineSpeakerDiarizationProcessWithCallback when making the
changes.

In `@sherpa-onnx/c-api/docs/punctuation.dox`:
- Around line 26-30: Update the doc snippet to use the consistent
SherpaOnnxOffline* API names: replace SherpaOfflinePunctuationAddPunct with
SherpaOnnxOfflinePunctuationAddPunct and replace
SherpaOfflinePunctuationFreeText with SherpaOnnxOfflinePunctuationFreeText so
all offline punctuation calls (e.g., the call that produces result and the free
call) match the SherpaOnnxOffline naming used elsewhere (leave
SherpaOnnxDestroyOfflinePunctuation as-is).

In `@sherpa-onnx/c-api/docs/source-separation.dox`:
- Around line 28-30: The examples create a SherpaOnnxOfflineSourceSeparation
instance via SherpaOnnxCreateOfflineSourceSeparation(&config) but never destroy
it; add a call to SherpaOnnxDestroyOfflineSourceSeparation(ss) after the usage
of the ss variable in both snippets (the snippet around the creation at lines
28-30 and the other example at 47-49) to properly free resources and avoid
leaking the SherpaOnnxOfflineSourceSeparation object.

In `@sherpa-onnx/c-api/docs/speech-enhancement.dox`:
- Around line 45-47: Add explicit destroy calls for every denoiser/audio
allocation shown in the snippets: after
SherpaOnnxCreateOfflineSpeechDenoiser(&config) call invoke
SherpaOnnxDestroyOfflineSpeechDenoiser(sd); do the same for the online denoiser
example by pairing SherpaOnnxCreateOnlineSpeechDenoiser(...) with
SherpaOnnxDestroyOnlineSpeechDenoiser(...) and for any denoised audio buffers
pair their creation with SherpaOnnxDestroyDenoisedAudio(...); place the destroy
calls immediately after the last use of each object to make ownership and
lifetimes explicit.

---

Nitpick comments:
In @.github/workflows/c-api.yaml:
- Around line 1650-1665: Remove the disabled GitHub Actions step named "Test
ffmpeg" that uses the constant condition if: false (the whole YAML block
starting with the step name "Test ffmpeg"); either delete the entire step or
comment it out so actionlint no longer sees a constant expression, and keep any
needed test logic elsewhere or behind a proper conditional matrix entry or
input-driven if condition.

In `@cxx-api-examples/offline-tts-piper-cxx-api.cc`:
- Line 62: The assignment gen_config.silence_scale = config.silence_scale uses
config.silence_scale without it being explicitly initialized; either initialize
config.silence_scale to the intended default value earlier (e.g., set
config.silence_scale = <desired_value> before the assignment) or avoid relying
on config and set gen_config.silence_scale directly to the intended
default/constant when constructing gen_config so the value is explicit and
readable.
- Line 58: The code calls OfflineTts::Create(config) but doesn't check the
return value; verify that the returned pointer or optional (variable tts) is
valid before using it and handle failure by logging an error and exiting or
returning an error status. Locate the OfflineTts::Create call and after it check
if tts is null/empty or indicates failure (depending on its return type) and
then call process logger/print an error like "Failed to create OfflineTts" along
with any error info and abort/return to avoid subsequent dereference of tts.

In `@cxx-api-examples/speaker-identification-cxx-api.cc`:
- Line 62: Verify the result of SpeakerEmbeddingManager::Create(dim) just like
the extractor creation: check that the returned manager is valid (non-null or
has_value depending on its return type) before using it for enrollment, and if
creation failed log an error via the same logger and exit/return early; update
references to manager in enrollment code to assume a valid manager only after
this validation.

In `@cxx-api-examples/streaming-paraformer-cxx-api.cc`:
- Around line 17-22: The code uses std::min (referenced at the call to std::min
in streaming-paraformer-cxx-api.cc) but doesn't include <algorithm>, so add an
explicit `#include` <algorithm> to the top include block (alongside <chrono>,
<cstdio>, <iostream>, <string>) to avoid relying on transitive includes and
ensure std::min is declared.

In `@sherpa-onnx/c-api/cxx-api.h`:
- Around line 1867-1888: Update the C++ wrapper comments for the SpeakerManager
methods to explicitly document buffer sizes and termination: for Add(const
std::string &name, const float *v) state that v must point to exactly dim floats
(embedding length); for AddList(const std::string &name, const float **v) state
that v is a NULL-terminated array of pointers, each pointing to an embedding of
exactly dim floats; for AddListFlattened(const std::string &name, const float
*v, int32_t n) state that v is a flat array containing n embeddings (so length =
n * dim) and that n is the number of embeddings (not number of floats); and for
Search(const float *v, float threshold) and Verify(const std::string &name,
const float *v, float threshold) state that v must point to exactly dim floats;
update the brief comments above those method declarations (Add, AddList,
AddListFlattened, Search, Verify) to include these requirements.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 3d7c18b5-ded6-4686-84d0-a832bfe448a0

📥 Commits

Reviewing files that changed from the base of the PR and between 6acf367 and 2cc5bc5.

📒 Files selected for processing (55)
  • .github/workflows/c-api.yaml
  • .github/workflows/cxx-api.yaml
  • c-api-examples/CMakeLists.txt
  • c-api-examples/add-punctuation-c-api.c
  • c-api-examples/fire-red-asr-ctc-c-api.c
  • c-api-examples/funasr-nano-c-api.c
  • c-api-examples/keywords-spotter-buffered-tokens-keywords-c-api.c
  • c-api-examples/medasr-ctc-c-api.c
  • c-api-examples/nemo-ctc-c-api.c
  • c-api-examples/nemo-giga-am-v2-c-api.c
  • c-api-examples/omnilingual-asr-ctc-c-api.c
  • c-api-examples/paraformer-c-api.c
  • c-api-examples/qwen3-asr-c-api.c
  • c-api-examples/sense-voice-c-api.c
  • c-api-examples/sense-voice-with-hr-c-api.c
  • c-api-examples/streaming-ctc-buffered-tokens-c-api.c
  • c-api-examples/streaming-nemotron-c-api.c
  • c-api-examples/streaming-paraformer-buffered-tokens-c-api.c
  • c-api-examples/streaming-paraformer-c-api.c
  • c-api-examples/streaming-t-one-ctc-c-api.c
  • c-api-examples/streaming-zipformer-buffered-tokens-hotwords-c-api.c
  • c-api-examples/streaming-zipformer-c-api.c
  • c-api-examples/vad-sense-voice-c-api.c
  • c-api-examples/wenet-ctc-c-api.c
  • c-api-examples/whisper-c-api.c
  • c-api-examples/zipformer-c-api.c
  • cxx-api-examples/CMakeLists.txt
  • cxx-api-examples/nemo-ctc-cxx-api.cc
  • cxx-api-examples/nemo-giga-am-v2-cxx-api.cc
  • cxx-api-examples/offline-speaker-diarization-cxx-api.cc
  • cxx-api-examples/offline-tts-piper-cxx-api.cc
  • cxx-api-examples/paraformer-cxx-api.cc
  • cxx-api-examples/speaker-identification-cxx-api.cc
  • cxx-api-examples/spoken-language-identification-cxx-api.cc
  • cxx-api-examples/streaming-nemotron-cxx-api.cc
  • cxx-api-examples/streaming-paraformer-cxx-api.cc
  • sherpa-onnx/c-api/Doxyfile
  • sherpa-onnx/c-api/c-api.h
  • sherpa-onnx/c-api/cxx-api.cc
  • sherpa-onnx/c-api/cxx-api.h
  • sherpa-onnx/c-api/docs/audio-tagging.dox
  • sherpa-onnx/c-api/docs/keyword-spotting.dox
  • sherpa-onnx/c-api/docs/offline-asr.dox
  • sherpa-onnx/c-api/docs/online-asr.dox
  • sherpa-onnx/c-api/docs/punctuation.dox
  • sherpa-onnx/c-api/docs/resampler.dox
  • sherpa-onnx/c-api/docs/source-separation.dox
  • sherpa-onnx/c-api/docs/speaker-diarization.dox
  • sherpa-onnx/c-api/docs/speaker-embedding.dox
  • sherpa-onnx/c-api/docs/speech-enhancement.dox
  • sherpa-onnx/c-api/docs/spoken-language-id.dox
  • sherpa-onnx/c-api/docs/tts.dox
  • sherpa-onnx/c-api/docs/vad.dox
  • sherpa-onnx/c-api/mainpage.md
  • sherpa-onnx/c-api/sherpa-onnx-symbols-c.exp

// tar xvf sherpa-onnx-pyannote-segmentation-3-0.tar.bz2
// rm sherpa-onnx-pyannote-segmentation-3-0.tar.bz2
//
// wget https://github.com/k2-fsa/sherpa-onnx/releases/download/speaker-recongition-models/3dspeaker_speech_eres2net_base_sv_zh-cn_3dspeaker_16k.onnx

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

Fix typo in the release URL path.

Line 15 uses speaker-recongition-models; this typo is likely to break the example download command.

Proposed fix
-// wget https://github.com/k2-fsa/sherpa-onnx/releases/download/speaker-recongition-models/3dspeaker_speech_eres2net_base_sv_zh-cn_3dspeaker_16k.onnx
+// wget https://github.com/k2-fsa/sherpa-onnx/releases/download/speaker-recognition-models/3dspeaker_speech_eres2net_base_sv_zh-cn_3dspeaker_16k.onnx
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
// wget https://github.com/k2-fsa/sherpa-onnx/releases/download/speaker-recongition-models/3dspeaker_speech_eres2net_base_sv_zh-cn_3dspeaker_16k.onnx
// wget https://github.com/k2-fsa/sherpa-onnx/releases/download/speaker-recognition-models/3dspeaker_speech_eres2net_base_sv_zh-cn_3dspeaker_16k.onnx
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@cxx-api-examples/offline-speaker-diarization-cxx-api.cc` at line 15, The wget
command in offline-speaker-diarization-cxx-api.cc contains a misspelling in the
release path: replace "speaker-recongition-models" with
"speaker-recognition-models" in the commented download URL; update the comment
line that starts with "wget
https://github.com/k2-fsa/sherpa-onnx/releases/download/..." so the path is
corrected to the proper release directory name.

Comment thread sherpa-onnx/c-api/cxx-api.cc
Comment on lines +26 to +30
const char *result = SherpaOfflinePunctuationAddPunct(
punct, "how are you i am fine thank you");
printf("%s\n", result);
SherpaOfflinePunctuationFreeText(result);
SherpaOnnxDestroyOfflinePunctuation(punct);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Fix offline punctuation function names in the snippet.

The offline sample uses inconsistent symbols and likely won’t compile as written (SherpaOffline... vs SherpaOnnx... used elsewhere). Please use the SherpaOnnxOffline... function names.

Proposed doc fix
-const char *result = SherpaOfflinePunctuationAddPunct(
+const char *result = SherpaOnnxOfflinePunctuationAddPunct(
     punct, "how are you i am fine thank you");
 printf("%s\n", result);
-SherpaOfflinePunctuationFreeText(result);
+SherpaOnnxOfflinePunctuationFreeText(result);
 SherpaOnnxDestroyOfflinePunctuation(punct);
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
const char *result = SherpaOfflinePunctuationAddPunct(
punct, "how are you i am fine thank you");
printf("%s\n", result);
SherpaOfflinePunctuationFreeText(result);
SherpaOnnxDestroyOfflinePunctuation(punct);
const char *result = SherpaOnnxOfflinePunctuationAddPunct(
punct, "how are you i am fine thank you");
printf("%s\n", result);
SherpaOnnxOfflinePunctuationFreeText(result);
SherpaOnnxDestroyOfflinePunctuation(punct);
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@sherpa-onnx/c-api/docs/punctuation.dox` around lines 26 - 30, Update the doc
snippet to use the consistent SherpaOnnxOffline* API names: replace
SherpaOfflinePunctuationAddPunct with SherpaOnnxOfflinePunctuationAddPunct and
replace SherpaOfflinePunctuationFreeText with
SherpaOnnxOfflinePunctuationFreeText so all offline punctuation calls (e.g., the
call that produces result and the free call) match the SherpaOnnxOffline naming
used elsewhere (leave SherpaOnnxDestroyOfflinePunctuation as-is).

Comment thread sherpa-onnx/c-api/docs/source-separation.dox
Comment thread sherpa-onnx/c-api/docs/speech-enhancement.dox
@dosubot dosubot Bot added size:XL This PR changes 500-999 lines, ignoring generated files. and removed size:XXL This PR changes 1000+ lines, ignoring generated files. labels May 20, 2026

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In @.github/scripts/test-python.sh:
- Line 252: The script downloads the SenseVoice int8 archive but later calls the
CLI with --sense-voice=$repo/model.onnx which doesn't exist in that archive;
update the references so the subtitle-generation steps use $repo/model.int8.onnx
(or add simple detection logic to prefer model.int8.onnx when present) instead
of $repo/model.onnx; change both occurrences of --sense-voice=$repo/model.onnx
to --sense-voice=$repo/model.int8.onnx (or implement a short check that picks
model.int8.onnx if tokens.txt plus model.int8.onnx are present).
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: fa6ce343-68d9-46fd-abcf-9a30bf9c0f6f

📥 Commits

Reviewing files that changed from the base of the PR and between 173ae50 and e6a5242.

📒 Files selected for processing (19)
  • .github/scripts/test-c-api.sh
  • .github/scripts/test-cxx-api.sh
  • .github/scripts/test-python.sh
  • .github/workflows/c-api.yaml
  • .github/workflows/cxx-api.yaml
  • c-api-examples/sense-voice-c-api.c
  • c-api-examples/sense-voice-with-hr-c-api.c
  • c-api-examples/vad-sense-voice-c-api.c
  • cxx-api-examples/sense-voice-cxx-api.cc
  • cxx-api-examples/sense-voice-simulate-streaming-alsa-cxx-api.cc
  • cxx-api-examples/sense-voice-simulate-streaming-microphone-cxx-api.cc
  • cxx-api-examples/sense-voice-with-hr-cxx-api.cc
  • python-api-examples/generate-subtitles.py
  • python-api-examples/non_streaming_server.py
  • python-api-examples/offline-sense-voice-ctc-decode-files-with-hr.py
  • python-api-examples/offline-sense-voice-ctc-decode-files.py
  • python-api-examples/simulate-streaming-sense-voice-microphone.py
  • python-api-examples/two-pass-wss.py
  • python-api-examples/vad-with-non-streaming-asr.py
✅ Files skipped from review due to trivial changes (5)
  • python-api-examples/simulate-streaming-sense-voice-microphone.py
  • python-api-examples/two-pass-wss.py
  • python-api-examples/vad-with-non-streaming-asr.py
  • python-api-examples/generate-subtitles.py
  • python-api-examples/non_streaming_server.py

Comment thread .github/scripts/test-python.sh

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
.github/scripts/test-python.sh (1)

282-282: 💤 Low value

Consider quoting variable expansions for defensive coding.

Shellcheck recommends quoting $repo/model.int8.onnx to prevent potential word splitting or globbing issues. While unlikely to cause problems in this controlled CI environment, it's a defensive best practice.

🛡️ Proposed fix
   python3 ./python-api-examples/generate-subtitles.py \
     --silero-vad-model=./silero_vad.onnx \
-    --sense-voice=$repo/model.int8.onnx \
+    --sense-voice="$repo/model.int8.onnx" \
     --tokens=$repo/tokens.txt \
@@
   python3 ./python-api-examples/generate-subtitles.py \
     --silero-vad-model=./silero_vad.onnx \
-    --sense-voice=$repo/model.int8.onnx \
+    --sense-voice="$repo/model.int8.onnx" \
     --tokens=$repo/tokens.txt \

Note: Many other variable expansions throughout this script also lack quotes. A comprehensive fix would quote them all, but that's beyond the scope of this PR.

Also applies to: 296-296

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In @.github/scripts/test-python.sh at line 282, The command flag
--sense-voice=$repo/model.int8.onnx should use a quoted variable expansion to
avoid word-splitting or globbing; update the invocation that sets --sense-voice
to use "--sense-voice=\"$repo/model.int8.onnx\"" (and likewise quote the
analogous expansion at the other occurrence mentioned) so the shell treats the
path as a single argument.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In @.github/scripts/test-python.sh:
- Line 282: The command flag --sense-voice=$repo/model.int8.onnx should use a
quoted variable expansion to avoid word-splitting or globbing; update the
invocation that sets --sense-voice to use
"--sense-voice=\"$repo/model.int8.onnx\"" (and likewise quote the analogous
expansion at the other occurrence mentioned) so the shell treats the path as a
single argument.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: c3aa1141-b1e9-4d5f-9da2-8e8338629bf1

📥 Commits

Reviewing files that changed from the base of the PR and between e6a5242 and 9d51fd3.

📒 Files selected for processing (1)
  • .github/scripts/test-python.sh

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XL This PR changes 500-999 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant