Skip to content

Fix OHOS APIs for TTS and ASR - #3356

Merged
csukuangfj merged 1 commit into
k2-fsa:masterfrom
csukuangfj:fix-ohos
Mar 19, 2026
Merged

csukuangfj merged 1 commit into
k2-fsa:masterfrom
csukuangfj:fix-ohos

Conversation

@csukuangfj

@csukuangfj csukuangfj commented Mar 19, 2026 •

Copy link
Copy Markdown
Collaborator

Summary by CodeRabbit

Release Notes

  • New Features
    • Added support for new speech recognition models: FireRed ASR, FunASR-Nano, Omnilingual ASR, and MedASR.
    • Introduced new text-to-speech voices: Zipvoice, Pocket, and Supertonic.
    • Enhanced TTS generation with configurable parameters including speed and voice selection.
    • Extended voice activity detection with provider configuration support.
    • Added support for in-memory buffer input for streaming recognition.

@csukuangfj
csukuangfj requested a review from Copilot March 19, 2026 03:36
@dosubot dosubot Bot added the size:L This PR changes 100-499 lines, ignoring generated files. label Mar 19, 2026
@gemini-code-assist

Copy link
Copy Markdown

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly expands the capabilities of the SherpaOnnx HarmonyOS module by integrating support for a wider range of ASR and TTS models. It introduces new configuration options for both streaming and non-streaming speech processing, providing developers with more flexibility and control over model selection and speech generation parameters. These updates aim to enhance the robustness and versatility of speech-related functionalities within HarmonyOS applications.

Highlights

  • Expanded ASR Model Support: Added configurations for several new offline Automatic Speech Recognition (ASR) models, including FireRedAsr, FunASRNano, MedAsrCtc, and OmnilingualAsrCtc, enhancing the range of supported ASR capabilities.
  • New TTS Model Integrations: Introduced new Text-to-Speech (TTS) model configurations for Zipvoice, Pocket, and Supertonic models, broadening the available TTS options.
  • Advanced TTS Generation Control: Implemented new TtsGenerationConfig and TtsInputWithConfig types, along with corresponding generateWithConfig and generateAsyncWithConfig methods in the OfflineTts class, providing more granular control over speech generation parameters.
  • Streaming ASR Configuration Enhancements: Updated streaming ASR configurations to include tokensBuf, tokensBufSize, hotwordsBuf, and hotwordsBufSize properties, offering finer control over online model and recognizer behavior.
  • VAD Provider Flexibility: Extended the VadConfig to include a provider property, allowing for specification of the Voice Activity Detection (VAD) processing backend.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@csukuangfj
csukuangfj merged commit 8f24b52 into k2-fsa:master Mar 19, 2026
1 of 2 checks passed
@csukuangfj
csukuangfj deleted the fix-ohos branch March 19, 2026 03:37
@coderabbitai

coderabbitai Bot commented Mar 19, 2026 •

Copy link
Copy Markdown

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 1b9b8e84-514d-45c5-bd75-3147a7cea1c7

📥 Commits

Reviewing files that changed from the base of the PR and between c6425d5 and 0879d6e.

📒 Files selected for processing (6)
  • harmony-os/SherpaOnnxHar/sherpa_onnx/Index.ets
  • harmony-os/SherpaOnnxHar/sherpa_onnx/src/main/cpp/types/libsherpa_onnx/Index.d.ts
  • harmony-os/SherpaOnnxHar/sherpa_onnx/src/main/ets/components/NonStreamingAsr.ets
  • harmony-os/SherpaOnnxHar/sherpa_onnx/src/main/ets/components/NonStreamingTts.ets
  • harmony-os/SherpaOnnxHar/sherpa_onnx/src/main/ets/components/StreamingAsr.ets
  • harmony-os/SherpaOnnxHar/sherpa_onnx/src/main/ets/components/Vad.ets

📝 Walkthrough

Walkthrough

The PR extends the public API surface of the sherpa-onnx HarmonyOS binding by adding new ASR model configuration classes (FireRed, FunASR Nano, Omnilingual, MedASR), TTS model configuration classes (Pocket, Supertonic, Zipvoice), generation config types, and new TTS generation methods with configuration support. Additional fields are added to existing config classes for in-memory buffers and provider selection.

Changes

Cohort / File(s) Summary
Module Exports
harmony-os/SherpaOnnxHar/sherpa_onnx/Index.ets
Extended public API surface by exporting 5 new ASR model config classes and 5 new TTS-related types (model configs and generation config classes).
Native Bindings
harmony-os/SherpaOnnxHar/sherpa_onnx/src/main/cpp/types/libsherpa_onnx/Index.d.ts
Added two new synchronous and asynchronous native function declarations for TTS generation with configuration parameters.
Non-Streaming ASR Configs
harmony-os/SherpaOnnxHar/sherpa_onnx/src/main/ets/components/NonStreamingAsr.ets
Added 5 new model configuration classes (FireRed variants, FunASR Nano, Omnilingual, MedASR) with fields for model paths, generation parameters, and prompts. Updated existing Whisper and Moonshine configs with timestamp and decoder options, and extended OfflineModelConfig with 5 new properties.
Non-Streaming TTS Configs & Methods
harmony-os/SherpaOnnxHar/sherpa_onnx/src/main/ets/components/NonStreamingTts.ets
Added 3 new TTS model configuration classes (Zipvoice, Pocket, Supertonic), a TtsGenerationConfig class with generation parameters, TtsInputWithConfig class pairing text with config, and two new OfflineTts methods (sync/async) for generation with config.
Streaming ASR Buffer Config
harmony-os/SherpaOnnxHar/sherpa_onnx/src/main/ets/components/StreamingAsr.ets
Added in-memory buffer fields (tokensBuf/size and hotwordsBuf/size) to OnlineModelConfig and OnlineRecognizerConfig for supplying pre-loaded data.
VAD Provider Config
harmony-os/SherpaOnnxHar/sherpa_onnx/src/main/ets/components/Vad.ets
Added provider field to VadConfig with default value 'cpu' and updated constructor signature.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Suggested labels

size:M

Poem

🐰 New models hop into the sherpa nest,
ZipVoice, Pocket, Supertonic—all the best!
Configs dance with FireRed and FunASR Nano grace,
The API expands to fill every space! 🎉

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
📝 Coding Plan
  • Generate coding plan for human review comments

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

The pull request introduces new ASR and TTS model configurations and corresponding API exports across several files. New model configurations for OfflineFireRedAsrCtcModelConfig, OfflineFireRedAsrModelConfig, OfflineFunASRNanoModelConfig, OfflineMedAsrCtcModelConfig, OfflineOmnilingualAsrCtcModelConfig, OfflineTtsPocketModelConfig, OfflineTtsSupertonicModelConfig, and OfflineTtsZipvoiceModelConfig have been added. Additionally, TtsGenerationConfig and TtsInputWithConfig are introduced for more flexible TTS generation. The VadConfig and OnlineModelConfig classes also received minor updates to include new properties. Overall, the changes are well-structured and extend the functionality of the SherpaOnnxHar library.

Comment on lines +75 to +76
public temperature: number = 0;
public topP: number = 0;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The default values for temperature and topP are set to 0. For language models, temperature typically controls the randomness of predictions (a value of 0 usually means deterministic output), and topP is a nucleus sampling parameter (a value of 0 might disable it or lead to unexpected behavior). It's generally recommended to use small positive values for temperature (e.g., 0.7-1.0) and values between 0 and 1 for topP (e.g., 0.8-0.95) to allow for more natural and diverse text generation. Please consider adjusting these defaults to more appropriate values for typical LLM usage.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Updates HarmonyOS ETS bindings/configs to match newer Sherpa-ONNX OHOS APIs for TTS and ASR, adding missing configuration fields and exposing new generation functions.

Changes:

  • Add provider support to VAD config.
  • Extend streaming ASR configs with token/hotword buffer fields.
  • Expand offline TTS/ASR model configs and expose “generate with config” APIs + type/export updates.

Reviewed changes

Copilot reviewed 5 out of 6 changed files in this pull request and generated 6 comments.

Show a summary per file
File Description
harmony-os/SherpaOnnxHar/sherpa_onnx/src/main/ets/components/Vad.ets Adds provider to VadConfig and threads it through constructor.
harmony-os/SherpaOnnxHar/sherpa_onnx/src/main/ets/components/StreamingAsr.ets Adds buffer + size fields to model/recognizer configs.
harmony-os/SherpaOnnxHar/sherpa_onnx/src/main/ets/components/NonStreamingTts.ets Adds new TTS model configs and new “with config” generation inputs/APIs.
harmony-os/SherpaOnnxHar/sherpa_onnx/src/main/ets/components/NonStreamingAsr.ets Adds new offline ASR model config types/fields.
harmony-os/SherpaOnnxHar/sherpa_onnx/src/main/cpp/types/libsherpa_onnx/Index.d.ts Exposes new native TTS functions in TS typings.
harmony-os/SherpaOnnxHar/sherpa_onnx/Index.ets Re-exports newly added config types and inputs.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines 61 to 74
public sampleRate: number;
public debug: boolean;
public numThreads: number;
public provider: string = 'cpu';

public constructor(sileroVad: SileroVadConfig, tenVad: TenVadConfig, sampleRate: number, debug: boolean, numThreads: number) {
public constructor(sileroVad: SileroVadConfig, tenVad: TenVadConfig, sampleRate: number, debug: boolean,
numThreads: number, provider: string = 'cpu') {
this.sileroVad = sileroVad;
this.tenVad = tenVad;
this.sampleRate = sampleRate;
this.debug = debug;
this.numThreads = numThreads;
this.provider = provider;
}
@@ -67,6 +67,8 @@ export class OnlineModelConfig {
public modelType: string = '';
public modelingUnit: string = "cjkchar";
public bpeVocab: string = '';
Comment on lines +94 to 96
public hotwordsBuf: string = '';
public hotwordsBufSize: number = 0;
public hr: HomophoneReplacerConfig = new HomophoneReplacerConfig();
public referenceSampleRate: number = 0;
public referenceText: string = '';
public numSteps: number = 5;
public extra: object = {};
Comment on lines +132 to +139
public enableExternalBuffer: boolean = true;
public callback?: (data: TtsCallbackData) => number;
}

export class TtsInputWithConfig {
public text: string = '';
public generationConfig: TtsGenerationConfig = new TtsGenerationConfig();
public enableExternalBuffer: boolean = true;
}

generateWithConfig(input: TtsInputWithConfig): TtsOutput {
return offlineTtsGenerateWithConfig(this.handle, input) as TtsOutput;
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L This PR changes 100-499 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants