Skip to content

Add async js API for tts generate. - #3133

Merged
csukuangfj merged 2 commits into
k2-fsa:masterfrom
csukuangfj:nodejs-async-tts-generate
Feb 4, 2026
Merged

csukuangfj merged 2 commits into
k2-fsa:masterfrom
csukuangfj:nodejs-async-tts-generate

Conversation

@csukuangfj

@csukuangfj csukuangfj commented Feb 4, 2026 •

Copy link
Copy Markdown
Collaborator

Summary by CodeRabbit

  • New Features

    • Asynchronous text‑to‑speech generation with an onProgress callback that can cancel mid‑generation.
    • Updated example shows non‑blocking usage and per‑chunk progress/logging.
  • Bug Fixes / Improvements

    • More reliable cancellation handling and pre‑generation cleanup to avoid stale tasks.

@dosubot dosubot Bot added the size:L This PR changes 100-499 lines, ignoring generated files. label Feb 4, 2026
@gemini-code-assist

Copy link
Copy Markdown

Summary of Changes

Hello @csukuangfj, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly enhances the Node.js binding for Sherpa-Onnx's Text-to-Speech functionality by introducing an asynchronous generation API. This allows applications to perform TTS synthesis in a non-blocking manner, providing real-time progress updates and the ability to cancel ongoing operations. The changes ensure a more responsive and interactive user experience by integrating asynchronous patterns across the C++ backend, the Node.js wrapper, and an updated example.

Highlights

  • Asynchronous TTS Generation API: Introduced a new generateAsync method in the Node.js OfflineTts class, enabling non-blocking Text-to-Speech synthesis with progress reporting.
  • Progress Reporting and Cancellation: The generateAsync API now supports an onProgress callback, allowing JavaScript consumers to receive partial audio chunks, monitor generation progress, and cancel the operation by returning false or 0.
  • C++ Backend Enhancements: The C++ implementation was updated to correctly interpret cancellation signals from JavaScript callbacks (handling both boolean false and numeric 0) and to perform proper memory cleanup for processed data within the TtsGenerateWorker.
  • Updated Example Usage: The test_tts_non_streaming_kitten_en.js example script was refactored to demonstrate the usage of the new generateAsync API, showcasing its progress reporting and cancellation capabilities.
Changelog
  • harmony-os/SherpaOnnxHar/sherpa_onnx/src/main/cpp/non-streaming-tts.cc
    • Modified InvokeJsCallback to correctly interpret boolean false and number 0 return values from JavaScript callbacks as cancellation signals.
    • Reordered data->processed = true; assignment to ensure cancellation logic is evaluated first.
    • Added a cleanup loop in TtsGenerateWorker::Execute to delete and erase processed data items from data_list_, preventing memory leaks.
  • nodejs-addon-examples/test_tts_non_streaming_kitten_en.js
    • Refactored the synchronous TTS generation example into an async function main() to demonstrate the new asynchronous API.
    • Implemented tts.generateAsync with an onProgress callback to show progress reporting and cancellation.
    • Updated console output to reflect asynchronous generation metrics.
    • Added JSDoc comments for the createOfflineTts function.
    • Included error handling for the main asynchronous function.
  • scripts/node-addon-api/lib/non-streaming-tts.js
    • Added a new generateAsync method to the OfflineTts class, exposing the asynchronous TTS generation functionality.
    • Implemented logic within generateAsync to wrap the onProgress callback, translating JavaScript return values (0 or false) into the C++ expected values for cancellation.
    • Added comprehensive JSDoc comments for the new generateAsync method, detailing its parameters and return type.
Activity
  • No specific activity (comments, reviews, or progress updates) has been recorded for this pull request yet.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@coderabbitai

coderabbitai Bot commented Feb 4, 2026 •

Copy link
Copy Markdown

Caution

Review failed

The pull request is closed.

📝 Walkthrough

Walkthrough

Adds asynchronous non-streaming TTS: Node wrapper gets OfflineTts.generateAsync with onProgress and cancellation semantics; C++ addon switches callback flags to atomic, refines cancellation handling and cleans processed callback entries; example updated to use the new async API. (≤50 words)

Changes

Cohort / File(s) Summary
C++ Native Binding
harmony-os/SherpaOnnxHar/sherpa_onnx/src/main/cpp/non-streaming-tts.cc
Added #include <atomic>; changed processed/cancelled to std::atomic<bool>; adjusted InvokeJsCallback to set cancelled from callback return and set processed afterwards; added pre-loop cleanup to remove/delete already-processed entries from data_list_.
Node.js API Wrapper
scripts/node-addon-api/lib/non-streaming-tts.js
Added OfflineTts.generateAsync(obj) returning a Promise<GeneratedAudio>; accepts optional onProgress callback which is wrapped to canonicalize return values (0/false → cancel).
Test / Example
nodejs-addon-examples/test_tts_non_streaming_kitten_en.js
Rewrote example to use generateAsync with onProgress, moved timing/RTF logging to async flow, added top-level async main and error handling, and prints initial speaker/sample-rate info.

Sequence Diagram

sequenceDiagram
    participant App as JavaScript App
    participant Wrapper as Node.js Wrapper
    participant Native as C++ Addon
    participant Worker as Worker Thread

    App->>Wrapper: generateAsync({ onProgress })
    Wrapper->>Native: addon.offlineTtsGenerateAsync(wrappedCallback)
    Native->>Worker: enqueue generation task

    loop per audio chunk
        Worker->>Native: emit progress event
        Native->>Wrapper: call wrappedCallback(progress)
        Wrapper->>App: invoke user onProgress(progress)
        App-->>Wrapper: return value (continue or 0/false)
        Wrapper-->>Native: return canonicalized boolean
        Native->>Native: set cancelled if return is 0/false
        alt cancelled
            Native-->>Worker: stop processing
        else
            Worker->>Worker: continue processing
        end
    end

    Worker->>Native: generation finished
    Native-->>Wrapper: return GeneratedAudio
    Wrapper-->>App: resolve Promise with GeneratedAudio
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~22 minutes

Suggested labels

size:M

Poem

🐰 I hopped through code with careful paws,
Async blossoms, progress applause,
Callbacks wrapped, cancellations neat,
Processed crumbs swept off the seat,
The TTS hums — a tidy little cause.

🚥 Pre-merge checks | ✅ 2 | ❌ 1
❌ Failed checks (1 warning)
Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 50.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the main change: adding an asynchronous JavaScript API for TTS generation, which is reflected in all modified files.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing touches
  • 📝 Generate docstrings
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces an asynchronous generateAsync API for TTS in the Node.js addon, which is a great feature for non-blocking operations. The implementation includes a progress callback mechanism.

My review has identified a critical thread safety issue in the C++ addon that could lead to data races and undefined behavior. I've also suggested an improvement to the JavaScript example file to make its behavior less confusing for users.

Comment on lines +474 to +482
for (auto it = _this->data_list_.begin();
it != _this->data_list_.end();) {
if ((*it)->processed) {
delete *it;
it = _this->data_list_.erase(it);
} else {
++it;
}
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

critical

There is a data race on the data_list_ member. The TtsGenerateWorker::Execute method runs on a worker thread, while InvokeJsCallback is executed on the main Node.js event loop thread.

  • The worker thread reads (*it)->processed here and d->cancelled in the next loop.
  • The main thread writes to data->processed and data->cancelled in InvokeJsCallback.

This concurrent read/write from different threads without synchronization is a data race and can lead to undefined behavior.

To fix this, the processed and cancelled members of the TtsCallbackData struct should be declared as std::atomic<bool>.

// In non-streaming-tts.cc
#include <atomic>

struct TtsCallbackData {
  std::vector<float> samples;
  float progress;
  std::atomic<bool> processed = {false};
  std::atomic<bool> cancelled = {false};
};

With this change, the concurrent access will be thread-safe.

Comment on lines +55 to +56
// Return 0 or false to cancel, any other value to continue
return false;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The example currently cancels the TTS generation immediately after the first chunk by returning false from the onProgress callback. This can be confusing for users who run the example expecting to get a fully generated audio file.

To make the example more intuitive, I suggest changing it to continue generation by default, and improving the comment to clarify how cancellation works.

      // Return true to continue generation.
      // You can return false or 0 to cancel the generation.
      return true;

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Fix all issues with AI agents
In `@harmony-os/SherpaOnnxHar/sherpa_onnx/src/main/cpp/non-streaming-tts.cc`:
- Around line 430-438: The TtsCallbackData fields processed and cancelled are
written on the JS thread inside the TSFN callback and read on the worker thread
in Execute, causing a data race; fix it by changing TtsCallbackData::processed
and TtsCallbackData::cancelled to std::atomic<bool> (or wrap accesses with a
mutex) and update all reads/writes in the TSFN callback and the Execute method
to use atomic operations (or lock/unlock the mutex) so synchronization is
enforced for TtsCallbackData, ensuring thread-safe visibility between the TSFN
callback and Execute.

In `@nodejs-addon-examples/test_tts_non_streaming_kitten_en.js`:
- Around line 50-57: The onProgress callback currently returns false (in
onProgress), which per the API cancels generation and yields incomplete audio;
change the callback to return a truthy value (for example return true) or remove
the explicit return so generation continues; update the onProgress
implementation where samples and progress are handled to return true instead of
false.
🧹 Nitpick comments (1)
scripts/node-addon-api/lib/non-streaming-tts.js (1)

27-43: JSDoc nested comment workaround is fragile.

The escaped *\/ inside the JSDoc block is a workaround for nested comments, but some documentation generators may not handle this correctly. Consider extracting the callback type to a separate @typedef for cleaner documentation.

Comment thread nodejs-addon-examples/test_tts_non_streaming_kitten_en.js
@csukuangfj
csukuangfj merged commit 58a4680 into k2-fsa:master Feb 4, 2026
1 check was pending
@csukuangfj
csukuangfj deleted the nodejs-async-tts-generate branch February 4, 2026 10:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L This PR changes 100-499 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant