Skip to content

Add CXX API for PocketTTS - #3128

Merged
csukuangfj merged 2 commits into
k2-fsa:masterfrom
csukuangfj:cxx-api-pocket-tts
Feb 4, 2026
Merged

csukuangfj merged 2 commits into
k2-fsa:masterfrom
csukuangfj:cxx-api-pocket-tts

Conversation

@csukuangfj

@csukuangfj csukuangfj commented Feb 4, 2026 •

Copy link
Copy Markdown
Collaborator

Summary by CodeRabbit

  • New Features

    • PocketTTS English text-to-speech support and example usage added.
    • TTS generation now accepts richer configuration including reference audio, sample rate, and advanced options for finer control.
  • API

    • Updated TTS generation interfaces to accept the new unified generation configuration and return improved result objects.
  • Testing

    • CI now runs PocketTTS build and end-to-end generation, capturing generated WAV artifacts.

@dosubot dosubot Bot added the size:L This PR changes 100-499 lines, ignoring generated files. label Feb 4, 2026
@csukuangfj
csukuangfj requested a review from Copilot February 4, 2026 03:41
@coderabbitai

coderabbitai Bot commented Feb 4, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

Adds PocketTTS support: introduces a unified GenerationConfig with reference-audio fields, updates C and C++ TTS APIs and overloads, adds a C++ PocketTTS example and C example updates, wires Pocket model parameters in the CXX layer, and adds a CI step to build/run the pocket-tts C++ example and upload WAV artifacts.

Changes

Cohort / File(s) Summary
CI Testing
.github/workflows/cxx-api.yaml
Added "Test PocketTTS" step: builds pocket-tts-en-cxx-api, sets LD_LIBRARY_PATH/DYLD_LIBRARY_PATH, downloads/extracts pocket-tts model tarball, runs the binary, captures WAV output, and uploads artifacts per OS.
C API
sherpa-onnx/c-api/c-api.h, sherpa-onnx/c-api/c-api.cc
Renamed GenerationConfig → SherpaOnnxGenerationConfig; updated SherpaOnnxOfflineTtsGenerateWithConfig signatures to use new type; added reference audio fields to the C config struct.
CXX API
sherpa-onnx/c-api/cxx-api.h, sherpa-onnx/c-api/cxx-api.cc
Added C++ GenerationConfig (incl. reference_audio, reference_sample_rate, extra map); added Generate/Generate2 overloads accepting config; serialize extra to JSON; wire Pocket-specific model file fields into OfflineTts::Create.
Build / Examples
cxx-api-examples/CMakeLists.txt, cxx-api-examples/pocket-tts-en-cxx-api.cc, c-api-examples/pocket-tts-en-c-api.c
Added pocket-tts-en-cxx-api target; new C++ example demonstrating config-driven generation and progress callback; updated C example to use SherpaOnnxGenerationConfig and removed unused sid logging.

Sequence Diagram(s)

sequenceDiagram
    participant Example as C++ Example
    participant CXX as sherpa-onnx CXX API
    participant CApi as sherpa-onnx C API
    participant Model as PocketTTS Model / Filesystem

    Note over Example,CXX: Generate(text, GenerationConfig)
    Example->>CXX: OfflineTts::Generate(text, config, callback)
    CXX->>CApi: SherpaOnnxOfflineTtsGenerateWithConfig(tts, text, sherpaConfig, callback)
    CApi->>Model: load model artifacts (lm, encoder, decoder, vocab, token_scores)
    CApi->>Model: run generation (uses reference_audio if provided)
    Model-->>CApi: generated audio buffer
    CApi-->>CXX: SherpaOnnxGeneratedAudio result
    CXX-->>Example: GeneratedAudio -> write WAV
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related issues

  • Issue #3084: Directly related — same GenerationConfig-based Generate API change and added OfflineTts::Generate overloads.

Possibly related PRs

  • PR #3087: Adds/uses GenerationConfig with reference_audio and PocketTTS wiring; strongly related to example and config fields here.
  • PR #3088: Extends GenerationConfig/PocketTTS API surface (Python bindings and examples); complements C/C++ API changes.
  • PR #3115: Implements config-driven TTS generation and updates SherpaOnnxOfflineTtsGenerateWithConfig signatures; overlaps C API edits.

Suggested reviewers

  • Copilot

Poem

🐰 Pocket voice in tiny hop and trill,
Configs in order, callbacks fit the bill.
Reference waves guide each clever tone,
From C++ burrow, a warm new song is grown. 🥕🎶

🚥 Pre-merge checks | ✅ 2 | ❌ 1
❌ Failed checks (1 warning)
Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The PR title 'Add CXX API for PocketTTS' accurately summarizes the main changes: adding C++ API support for PocketTTS, including new example files, CMake configuration, and API extensions.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing touches
  • 📝 Generate docstrings
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist

Copy link
Copy Markdown

Summary of Changes

Hello @csukuangfj, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly expands the sherpa-onnx library by adding a comprehensive C++ API for PocketTTS, a text-to-speech model. The changes involve introducing new C++ configuration structures for generation, updating the underlying C-API to support these new configurations, and providing a practical C++ example to guide users. This enhancement aims to offer greater flexibility and ease of use for C++ developers working with PocketTTS.

Highlights

  • New C++ API for PocketTTS: Introduced a dedicated C++ API for PocketTTS, allowing developers to integrate English Text-to-Speech capabilities more seamlessly into C++ applications.
  • Enhanced Generation Configuration: The C++ API now supports a more flexible GenerationConfig struct, enabling the use of reference audio, custom sample rates, and model-specific extra parameters via an unordered_map.
  • C-API Struct Renaming for Clarity: The GenerationConfig struct in the C-API has been renamed to SherpaOnnxGenerationConfig to prevent naming conflicts and improve clarity when interacting with the C++ API.
  • New C++ Example: A new C++ example (pocket-tts-en-cxx-api.cc) has been added, demonstrating how to use the PocketTTS C++ API, including model loading, text generation, and progress callbacks.
Changelog
  • c-api-examples/pocket-tts-en-c-api.c
    • Updated GenerationConfig to SherpaOnnxGenerationConfig to align with the renamed C-API struct.
    • Removed an unused sid variable.
  • cxx-api-examples/CMakeLists.txt
    • Added a new executable target pocket-tts-en-cxx-api to build the new PocketTTS C++ example.
  • cxx-api-examples/pocket-tts-en-cxx-api.cc
    • New file: Implements a C++ example demonstrating English Text-to-Speech using PocketTTS, including model configuration, text input, reference audio usage, and progress callbacks.
  • sherpa-onnx/c-api/c-api.cc
    • Updated function signatures for SherpaOnnxOfflineTtsGenerateInternal and SherpaOnnxOfflineTtsGenerateWithConfig to use const SherpaOnnxGenerationConfig *.
  • sherpa-onnx/c-api/c-api.h
    • Renamed the C-API struct GenerationConfig to struct SherpaOnnxGenerationConfig.
  • sherpa-onnx/c-api/cxx-api.cc
    • Included nlohmann/json.hpp to facilitate handling of extra parameters in the C++ GenerationConfig.
    • Added logic within OfflineTts::Create to correctly map PocketTTS model paths from the C++ configuration to the underlying C-API configuration.
    • Implemented new overloads for OfflineTts::Generate and OfflineTts::Generate2 methods to accept the newly defined C++ GenerationConfig struct.
    • Applied minor formatting adjustments to funasr_nano configuration assignments.
  • sherpa-onnx/c-api/cxx-api.h
    • Included <unordered_map> for the new C++ GenerationConfig struct.
    • Defined a new C++ struct GenerationConfig which includes fields for reference_audio (as std::vector<float>), reference_sample_rate, and an unordered_map<std::string, std::string> extra for model-specific parameters.
    • Added new Generate and Generate2 method overloads to the OfflineTts class that accept the new C++ GenerationConfig.
Ignored Files
  • Ignored by pattern: .github/workflows/** (1)
    • .github/workflows/cxx-api.yaml
Activity
  • No specific activity (comments, reviews, or progress updates) was provided in the context for this pull request.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a C++ API for PocketTTS, which is a great addition. The changes include a new C++ example, updates to CMake, and modifications to the C/C++ APIs to support a more flexible generation configuration.

My review focuses on improving the newly introduced GenerationConfig in the C++ API for better flexibility and correctness. I've identified a potential issue with how extra parameters are handled, which could lead to parsing errors for non-string values. I've also found a critical issue with struct initialization that could lead to undefined behavior. I've provided specific suggestions to address these points in the relevant files.

Overall, the changes are well-structured, and with a few adjustments, the new API will be more robust and easier to use.

const GenerationConfig &config,
OfflineTtsCallback callback /*= nullptr*/,
void *arg /*= nullptr*/) const {
SherpaOnnxGenerationConfig c;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

critical

The SherpaOnnxGenerationConfig struct c is not initialized. This can lead to undefined behavior as its members will contain garbage values, especially pointers which might not be nullptr. It's important to zero-initialize C-style structs.

Suggested change
SherpaOnnxGenerationConfig c;
SherpaOnnxGenerationConfig c{};


#include <memory>
#include <string>
#include <unordered_map>

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

To support a more flexible GenerationConfig::extra field that can handle various data types (not just strings), it would be beneficial to use nlohmann::json. Please include its header here. I've added another comment with more details on the GenerationConfig struct itself.

#include "nlohmann/json.hpp"

int32_t num_steps = 5; // number of steps in flow matching

// extra attrs , model specific
std::unordered_map<std::string, std::string> extra;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The current design of GenerationConfig::extra as std::unordered_map<std::string, std::string> is restrictive as it only allows string values. This can lead to issues when numerical or boolean parameters are needed, as they are passed as strings and may be parsed incorrectly downstream. For instance, a value like "10" becomes "\"10\"" after JSON serialization, which fails to parse as a number.

To improve flexibility and prevent such parsing errors, I recommend changing the type of extra to nlohmann::json and initializing it as an empty object. This allows for storing various data types (numbers, booleans, strings) correctly.

  nlohmann::json extra = nlohmann::json::object();

Comment on lines +522 to +524
nlohmann::json j = config.extra;
std::string s = j.dump();
c.extra = s.c_str();

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Following the suggested change in cxx-api.h to use nlohmann::json for GenerationConfig::extra, this logic for JSON conversion should be updated. You can now directly dump config.extra if it's not empty, which simplifies the code and correctly handles various data types.

Suggested change
nlohmann::json j = config.extra;
std::string s = j.dump();
c.extra = s.c_str();
std::string s;
if (!config.extra.empty()) {
s = config.extra.dump();
c.extra = s.c_str();
}

Wave wave = ReadWave(reference_audio_file);
cfg.reference_audio = std::move(wave.samples);
cfg.reference_sample_rate = wave.sample_rate;
cfg.extra["max_reference_audio_len"] = "10";

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

With the proposed change to use nlohmann::json for GenerationConfig::extra, you should use a numeric type for this parameter instead of a string to avoid parsing issues and leverage the flexibility of the new API design.

Suggested change
cfg.extra["max_reference_audio_len"] = "10";
cfg.extra["max_reference_audio_len"] = 10.0;

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Fix all issues with AI agents
In `@cxx-api-examples/pocket-tts-en-cxx-api.cc`:
- Around line 67-86: ReadWave can fail and leave Wave.sample_rate uninitialized,
so after calling ReadWave(reference_audio_file) validate the returned Wave
(e.g., ensure wave.sample_rate > 0 and wave.samples is non-empty) before
assigning cfg.reference_audio and cfg.reference_sample_rate; if invalid, handle
the error path (log/throw/return or set sane defaults) so
GenerationConfig::reference_sample_rate is never assigned an uninitialized value
— update the code around the ReadWave call and the assignments to
cfg.reference_audio and cfg.reference_sample_rate in the block that constructs
GenerationConfig (where ReadWave, Wave, and
GenerationConfig::reference_sample_rate are used).

In `@sherpa-onnx/c-api/cxx-api.cc`:
- Around line 512-535: The code dereferences the result of
SherpaOnnxOfflineTtsGenerateWithConfig without checking for nullptr; modify the
call site (where SherpaOnnxOfflineTtsGenerateWithConfig is invoked and its
result stored in audio) to check if audio == nullptr and return an
empty/appropriate GeneratedAudio (e.g., samples empty and sample_rate 0) instead
of dereferencing; also guard the reference_audio assignment (c.reference_audio =
config.reference_audio.data()) so you only set c.reference_audio and
c.reference_audio_len when config.reference_audio is non-empty to avoid
triggering the C API validation path that returns nullptr.

In `@sherpa-onnx/c-api/cxx-api.h`:
- Around line 487-498: The GenerationConfig struct has reference_sample_rate
uninitialized which can lead to indeterminate values; update the struct
definition (GenerationConfig) to give reference_sample_rate a default value
(e.g., = 0) to match the Java API and ensure deterministic behavior, leaving
other fields unchanged; modify the declaration of reference_sample_rate in
GenerationConfig to include the default initializer.

Comment on lines +67 to +86
auto tts = OfflineTts::Create(config);
GenerationConfig cfg;
cfg.speed = 1.0;

std::string reference_audio_file =
"./sherpa-onnx-pocket-tts-int8-2026-01-26/test_wavs/bria.wav";

Wave wave = ReadWave(reference_audio_file);
cfg.reference_audio = std::move(wave.samples);
cfg.reference_sample_rate = wave.sample_rate;
cfg.extra["max_reference_audio_len"] = "10";

#if 0
// If you don't want to use a callback, then please enable this branch
GeneratedAudio audio = tts.Generate(text, cfg);
#else
GeneratedAudio audio = tts.Generate(text, cfg, ProgressCallback);
#endif

WriteWave(filename, {audio.samples, audio.sample_rate});

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Search for ReadWave function and GenerationConfig struct definitions
echo "=== Searching for ReadWave ==="
rg -n "Wave ReadWave" --type cc --type h -C 3

echo "=== Searching for GenerationConfig struct ==="
rg -n "struct GenerationConfig|class GenerationConfig" --type cc --type h -C 3

Repository: k2-fsa/sherpa-onnx

Length of output: 209


🏁 Script executed:

#!/bin/bash
# Search for ReadWave function and GenerationConfig struct definitions
echo "=== Searching for ReadWave ==="
rg "Wave ReadWave" -C 3

echo "=== Searching for GenerationConfig struct ==="
rg "struct GenerationConfig|class GenerationConfig" -C 3

echo "=== Check the file structure ==="
find . -name "cxx-api.h" -o -name "cxx-api.cc" | head -10

Repository: k2-fsa/sherpa-onnx

Length of output: 2788


🏁 Script executed:

#!/bin/bash
# Get full ReadWave implementation
echo "=== ReadWave implementation ==="
rg -A 15 "^Wave ReadWave" sherpa-onnx/c-api/cxx-api.cc

echo "=== Wave struct definition ==="
rg -B 2 -A 5 "^struct Wave" sherpa-onnx/c-api/cxx-api.h

echo "=== GenerationConfig in offline-tts.h ==="
rg -A 20 "^struct GenerationConfig" sherpa-onnx/csrc/offline-tts.h

echo "=== GenerationConfig in cxx-api.h ==="
rg -A 20 "^struct GenerationConfig" sherpa-onnx/c-api/cxx-api.h

Repository: k2-fsa/sherpa-onnx

Length of output: 2209


🏁 Script executed:

#!/bin/bash
# Check the actual file being reviewed
echo "=== File context around lines 67-86 ==="
sed -n '60,90p' cxx-api-examples/pocket-tts-en-cxx-api.cc

echo "=== Check includes to see which GenerationConfig is used ==="
head -20 cxx-api-examples/pocket-tts-en-cxx-api.cc

Repository: k2-fsa/sherpa-onnx

Length of output: 1807


🏁 Script executed:

#!/bin/bash
# Get full includes section
echo "=== Full includes in the file ==="
head -50 cxx-api-examples/pocket-tts-en-cxx-api.cc | grep -A 50 "include"

echo "=== Check if ReadWave and Wave are from cxx-api ==="
rg "^#include.*cxx-api" cxx-api-examples/pocket-tts-en-cxx-api.cc

Repository: k2-fsa/sherpa-onnx

Length of output: 1282


Add validation after ReadWave to prevent undefined behavior from uninitialized sample_rate.

ReadWave() performs default initialization on failure, leaving the Wave struct's sample_rate field uninitialized. Assigning this garbage value to GenerationConfig::reference_sample_rate causes undefined behavior. Check the result before use and ensure proper initialization.

Suggested fix
-  GenerationConfig cfg;
+  GenerationConfig cfg{};
   cfg.speed = 1.0;
   
   std::string reference_audio_file =
       "./sherpa-onnx-pocket-tts-int8-2026-01-26/test_wavs/bria.wav";
   
   Wave wave = ReadWave(reference_audio_file);
+  if (wave.samples.empty() || wave.sample_rate <= 0) {
+    fprintf(stderr, "Failed to read %s\n", reference_audio_file.c_str());
+    return -1;
+  }
+  cfg.reference_sample_rate = wave.sample_rate;
   cfg.reference_audio = std::move(wave.samples);
-  cfg.reference_sample_rate = wave.sample_rate;
   cfg.extra["max_reference_audio_len"] = "10";
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
auto tts = OfflineTts::Create(config);
GenerationConfig cfg;
cfg.speed = 1.0;
std::string reference_audio_file =
"./sherpa-onnx-pocket-tts-int8-2026-01-26/test_wavs/bria.wav";
Wave wave = ReadWave(reference_audio_file);
cfg.reference_audio = std::move(wave.samples);
cfg.reference_sample_rate = wave.sample_rate;
cfg.extra["max_reference_audio_len"] = "10";
#if 0
// If you don't want to use a callback, then please enable this branch
GeneratedAudio audio = tts.Generate(text, cfg);
#else
GeneratedAudio audio = tts.Generate(text, cfg, ProgressCallback);
#endif
WriteWave(filename, {audio.samples, audio.sample_rate});
auto tts = OfflineTts::Create(config);
GenerationConfig cfg{};
cfg.speed = 1.0;
std::string reference_audio_file =
"./sherpa-onnx-pocket-tts-int8-2026-01-26/test_wavs/bria.wav";
Wave wave = ReadWave(reference_audio_file);
if (wave.samples.empty() || wave.sample_rate <= 0) {
fprintf(stderr, "Failed to read %s\n", reference_audio_file.c_str());
return -1;
}
cfg.reference_sample_rate = wave.sample_rate;
cfg.reference_audio = std::move(wave.samples);
cfg.extra["max_reference_audio_len"] = "10";
`#if` 0
// If you don't want to use a callback, then please enable this branch
GeneratedAudio audio = tts.Generate(text, cfg);
`#else`
GeneratedAudio audio = tts.Generate(text, cfg, ProgressCallback);
`#endif`
WriteWave(filename, {audio.samples, audio.sample_rate});
🤖 Prompt for AI Agents
In `@cxx-api-examples/pocket-tts-en-cxx-api.cc` around lines 67 - 86, ReadWave can
fail and leave Wave.sample_rate uninitialized, so after calling
ReadWave(reference_audio_file) validate the returned Wave (e.g., ensure
wave.sample_rate > 0 and wave.samples is non-empty) before assigning
cfg.reference_audio and cfg.reference_sample_rate; if invalid, handle the error
path (log/throw/return or set sane defaults) so
GenerationConfig::reference_sample_rate is never assigned an uninitialized value
— update the code around the ReadWave call and the assignments to
cfg.reference_audio and cfg.reference_sample_rate in the block that constructs
GenerationConfig (where ReadWave, Wave, and
GenerationConfig::reference_sample_rate are used).

Comment thread sherpa-onnx/c-api/cxx-api.cc
Comment thread sherpa-onnx/c-api/cxx-api.h

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Fix all issues with AI agents
In `@sherpa-onnx/c-api/cxx-api.cc`:
- Around line 500-504: The early-return path creates a GeneratedAudio ans and
returns it without initializing its sample_rate (and potentially other fields);
update both early-return sites (the null-audio checks in the function that
declares GeneratedAudio ans) to set ans.sample_rate to a safe default (e.g., 0)
before returning, and ensure any other relevant members are set to known
defaults so callers do not read indeterminate values from GeneratedAudio.
🧹 Nitpick comments (1)
sherpa-onnx/c-api/cxx-api.cc (1)

560-569: Prefer std::make_shared for the new Generate2 overload.

It avoids a separate allocation and is the usual C++ idiom.

♻️ Proposed refactor
-  GeneratedAudio *ans = new GeneratedAudio;
-  ans->samples = std::move(audio.samples);
-  ans->sample_rate = audio.sample_rate;
-
-  return std::shared_ptr<GeneratedAudio>(ans);
+  auto ans = std::make_shared<GeneratedAudio>();
+  ans->samples = std::move(audio.samples);
+  ans->sample_rate = audio.sample_rate;
+  return ans;

Comment on lines 500 to +504
GeneratedAudio ans;

if (!audio) {
return ans;
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Find the GeneratedAudio struct definition
find . -type f -name "*.h" -o -name "*.cc" | head -20

Repository: k2-fsa/sherpa-onnx

Length of output: 968


🏁 Script executed:

# Search for GeneratedAudio struct definition
rg "struct GeneratedAudio|class GeneratedAudio" --type cpp --type cc -A 10

Repository: k2-fsa/sherpa-onnx

Length of output: 88


🏁 Script executed:

# Read the specific lines in cxx-api.cc
grep -n "GeneratedAudio" sherpa-onnx/c-api/cxx-api.cc | head -20

Repository: k2-fsa/sherpa-onnx

Length of output: 762


🏁 Script executed:

# Read the specific lines around 500-504
sed -n '495,515p' sherpa-onnx/c-api/cxx-api.cc

Repository: k2-fsa/sherpa-onnx

Length of output: 657


🏁 Script executed:

# Read the specific lines around 536-540
sed -n '530,545p' sherpa-onnx/c-api/cxx-api.cc

Repository: k2-fsa/sherpa-onnx

Length of output: 528


🏁 Script executed:

# Find the GeneratedAudio struct definition
rg "struct GeneratedAudio" --type cpp -A 15

Repository: k2-fsa/sherpa-onnx

Length of output: 1991


Initialize GeneratedAudio.sample_rate on early return paths.

The GeneratedAudio struct contains an int32_t sample_rate member without a default constructor. When returning early on the null-audio path with GeneratedAudio ans; return ans;, the sample_rate field is left uninitialized (indeterminate value), leading to undefined behavior if callers read it. Initialize it to a safe default before returning.

Proposed fix
  GeneratedAudio ans;

  if (!audio) {
+   ans.sample_rate = 0;
    return ans;
  }

This applies to both locations (lines 500-504 and 536-540).

📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
GeneratedAudio ans;
if (!audio) {
return ans;
}
GeneratedAudio ans;
if (!audio) {
ans.sample_rate = 0;
return ans;
}
🤖 Prompt for AI Agents
In `@sherpa-onnx/c-api/cxx-api.cc` around lines 500 - 504, The early-return path
creates a GeneratedAudio ans and returns it without initializing its sample_rate
(and potentially other fields); update both early-return sites (the null-audio
checks in the function that declares GeneratedAudio ans) to set ans.sample_rate
to a safe default (e.g., 0) before returning, and ensure any other relevant
members are set to known defaults so callers do not read indeterminate values
from GeneratedAudio.

@csukuangfj
csukuangfj merged commit e80f1b0 into k2-fsa:master Feb 4, 2026
9 of 21 checks passed
@csukuangfj
csukuangfj deleted the cxx-api-pocket-tts branch February 4, 2026 04:03
This was referenced Mar 10, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L This PR changes 100-499 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants