Skip to content

Feature/ort iobinding Fire Red ASR - #3010

Closed
Wasser1462 wants to merge 10 commits into
k2-fsa:masterfrom
Wasser1462:feature/ort-iobinding-fire-red
Closed

Wasser1462 wants to merge 10 commits into
k2-fsa:masterfrom
Wasser1462:feature/ort-iobinding-fire-red

Conversation

@Wasser1462

@Wasser1462 Wasser1462 commented Jan 9, 2026 •

Copy link
Copy Markdown
Collaborator

FireRedASR Performance Optimization with ONNX Runtime I/O Binding

Summary

This PR addresses sherpa-onnx issue #2943 by adopting ONNX Runtime I/O Binding (Ort::IoBinding) for both the encoder and decoder in FireRedASR offline ASR model to minimize GPU↔CPU transfers and improve end-to-end GPU latency.

Benchmark

Test Audio: sherpa-onnx-fire-red-asr-large-zh_en-2025-02-16/test_wavs/0.wav
GPU: RTX 4090
Provider: CUDA

Version RTF (Real-Time Factor)
Before 0.152
After 0.083

Summary by CodeRabbit

  • New Features

    • Added FunASR-Nano model support for offline speech recognition with configurable LLM components and generation parameters (temperature, top_p, max_new_tokens).
    • Added Python API method to construct recognizers with FunASR-Nano configuration.
    • Added C++ and Python example scripts demonstrating FunASR-Nano usage.
  • Documentation

    • Added CLI usage examples and help text for FunASR-Nano model setup and configuration.

✏️ Tip: You can customize this high-level summary in your review settings.

@dosubot dosubot Bot added the size:XXL This PR changes 1000+ lines, ignoring generated files. label Jan 9, 2026
@coderabbitai

coderabbitai Bot commented Jan 9, 2026 •

Copy link
Copy Markdown

Caution

Review failed

The pull request is closed.

📝 Walkthrough

Walkthrough

This PR introduces comprehensive FunASR-nano offline speech recognition support across the sherpa-onnx library, including new model configuration structures, tokenizer implementation, offline recognizer backend, C++ and Python API bindings, example programs, and platform-specific constructors for Android and OHOS environments.

Changes

Cohort / File(s) Summary
FunASR-nano Model Configuration (C API)
sherpa-onnx/c-api/c-api.h, sherpa-onnx/c-api/c-api.cc
Introduces new struct SherpaOnnxOfflineFunASRNanoModelConfig with fields for encoder adaptor, LLM prefill/decode, embedding, tokenizer, prompts, and generation parameters; initializes config fields with sensible defaults.
FunASR-nano Model Configuration (C++ API)
sherpa-onnx/c-api/cxx-api.h, sherpa-onnx/c-api/cxx-api.cc
Adds public struct OfflineFunASRNanoModelConfig with default values for generation parameters; extends C++ recognizer config population with FunASR-nano field mappings.
Model Configuration Core
sherpa-onnx/csrc/offline-funasr-nano-model-config.h, sherpa-onnx/csrc/offline-funasr-nano-model-config.cc
Implements configuration validation (file existence checks, parameter range validation), command-line registration, and string serialization for FunASR-nano setup.
Offline Model Config Integration
sherpa-onnx/csrc/offline-model-config.h, sherpa-onnx/csrc/offline-model-config.cc
Extends OfflineModelConfig with new funasr_nano member; makes tokens validation conditional; adds FunASR-nano validation and ToString delegation.
Core FunASR-nano Model
sherpa-onnx/csrc/offline-funasr-nano-model.h, sherpa-onnx/csrc/offline-funasr-nano-model.cc
Implements OfflineFunASRNanoModel managing ONNX Runtime sessions for encoder adaptor, LLM prefill/decode, and embedding; includes CUDA IO binding support, metadata extraction, and type-aware forward passes (942 lines of complex logic).
FunASR-nano Tokenizer
sherpa-onnx/csrc/funasr-nano-tokenizer.h, sherpa-onnx/csrc/funasr-nano-tokenizer.cc
Implements lightweight ByteLevel-BPE tokenizer with platform-specific constructors (Android/OHOS), JSON parsing, Trie-based added-token matching, BPE caching, and comprehensive UTF-8 handling (1368 lines).
FunASR-nano Offline Recognizer
sherpa-onnx/csrc/offline-recognizer-funasr-nano-impl.h, sherpa-onnx/csrc/offline-recognizer-funasr-nano-impl.cc
Implements offline recognizer backend with feature config, LFR processing, source ID construction, autoregressive text generation using encoder/prefill/decode models, token sampling, and result post-processing (436 lines).
Recognizer Factory Integration
sherpa-onnx/csrc/offline-recognizer-impl.cc
Routes recognizer creation to OfflineRecognizerFunASRNanoImpl when FunASR-nano config encoder adaptor is non-empty.
Offline Fire-Red ASR Enhancement
sherpa-onnx/csrc/offline-fire-red-asr-model.cc
Adds CUDA IO binding support with conditional memory binding for encoder/decoder outputs to optimize GPU memory usage.
Float Conversion Utilities
sherpa-onnx/csrc/onnx-utils.h, sherpa-onnx/csrc/onnx-utils.cc
Adds bidirectional conversion functions between IEEE 754 half-precision (FP16) and single-precision (FP32) floats.
Python Bindings (Config)
sherpa-onnx/python/csrc/offline-funasr-nano-model-config.h, sherpa-onnx/python/csrc/offline-funasr-nano-model-config.cc
Exposes OfflineFunASRNanoModelConfig to Python via pybind11 with read/write field bindings and string representation.
Python Bindings (Integration)
sherpa-onnx/python/csrc/offline-model-config.cc, sherpa-onnx/python/csrc/CMakeLists.txt
Registers FunASR-nano config bindings; extends OfflineModelConfig binding with funasr_nano member and constructor parameter.
Python API Exposure
sherpa-onnx/python/sherpa_onnx/__init__.py, sherpa-onnx/python/sherpa_onnx/offline_recognizer.py
Exports OfflineFunASRNanoModelConfig; adds from_funasr_nano() classmethod to OfflineRecognizer for convenient Python API access with all required parameters.
C++ Example
cxx-api-examples/funasr-nano-cxx-api.cc, cxx-api-examples/CMakeLists.txt
New C++ example program demonstrating FunASR-nano usage via sherpa-onnx C++ API with command-line interface, audio file processing, and timing measurement.
Python Example
python-api-examples/offline-funasr-nano-decode-files.py
New Python example demonstrating offline FunASR-nano decoding via sherpa-onnx Python API with batch file processing and result printing.
Build System
sherpa-onnx/csrc/CMakeLists.txt
Adds four new source files (offline-funasr-nano-model-config.cc, offline-funasr-nano-model.cc, offline-recognizer-funasr-nano-impl.cc, funasr-nano-tokenizer.cc) to unconditional core library compilation.
CLI Usage Documentation
sherpa-onnx/csrc/sherpa-onnx-offline.cc, sherpa-onnx/csrc/sherpa-onnx-vad-with-offline-asr.cc
Extends usage/help text with FunASR-nano model documentation and command-line examples.

Sequence Diagram(s)

sequenceDiagram
    actor User
    participant OfflineRecognizer
    participant OfflineFunASRNanoImpl
    participant OfflineFunASRNanoModel
    participant FunASRNanoTokenizer
    participant ORT as ONNX Runtime<br/>(Encoder/LLM/Embedding)
    
    User->>OfflineRecognizer: Create(config)
    OfflineRecognizer->>OfflineFunASRNanoImpl: new (config)
    OfflineFunASRNanoImpl->>OfflineFunASRNanoModel: new (config)
    OfflineFunASRNanoModel->>ORT: Load encoder, prefill, decode, embedding sessions
    OfflineFunASRNanoImpl->>FunASRNanoTokenizer: new (tokenizer_dir)
    FunASRNanoTokenizer->>FunASRNanoTokenizer: Load vocab, merges, tokenizer.json
    
    User->>OfflineRecognizer: CreateStream()
    OfflineRecognizer->>OfflineFunASRNanoImpl: CreateStream()
    OfflineFunASRNanoImpl-->>OfflineRecognizer: OfflineStream
    
    User->>OfflineRecognizer: Decode(stream with audio)
    OfflineRecognizer->>OfflineFunASRNanoImpl: DecodeStreams(streams)
    OfflineFunASRNanoImpl->>OfflineFunASRNanoModel: ForwardEncoderAdaptor(features)
    ORT-->>OfflineFunASRNanoModel: encoder_output
    
    OfflineFunASRNanoImpl->>FunASRNanoTokenizer: Encode(system/user prompts)
    FunASRNanoTokenizer-->>OfflineFunASRNanoImpl: token_ids
    
    loop Autoregressive Generation (up to max_new_tokens)
        OfflineFunASRNanoImpl->>OfflineFunASRNanoModel: ForwardLLMPrefill(inputs_embeds, attention_mask)
        ORT-->>OfflineFunASRNanoModel: logits, past_key_values
        OfflineFunASRNanoImpl->>OfflineFunASRNanoModel: ForwardLLMDecode(inputs_embeds, attention_mask, cache)
        ORT-->>OfflineFunASRNanoModel: next_logits, updated_cache
        OfflineFunASRNanoImpl->>OfflineFunASRNanoImpl: SampleToken(logits)
    end
    
    OfflineFunASRNanoImpl->>FunASRNanoTokenizer: Decode(generated_token_ids)
    FunASRNanoTokenizer-->>OfflineFunASRNanoImpl: transcription_text
    
    OfflineFunASRNanoImpl->>OfflineRecognizer: Set result on stream
    OfflineRecognizer->>User: GetResult() with transcription
Loading

Estimated code review effort

🎯 5 (Critical) | ⏱️ ~120 minutes

Possibly related PRs

  • Add funASR-Nano support #2936 — Both PRs add identical FunASR-nano support infrastructure (model, tokenizer, recognizer impl, config, bindings, and examples) with the same design and components.
  • FunASR-nano: switch to unified KV-cache LLM #2995 — Both PRs implement FunASR-nano integration but with conflicting field structures (main PR uses separate llm_prefill/llm_decode while retrieved PR consolidates to single llm model).
  • Fix building for HarmonyOS #2972 — Both PRs modify the same FunASR-nano offline recognizer implementation files and platform-specific Android/OHOS constructors.

Suggested reviewers

  • csukuangfj

Poem

🐰 Hop, hop—the tokenizer's here with BPE so fine,
FunASR-nano flows through LLM and time,
From prompt to speech, with prefill and decode divine,
CUDA bindings gleam and ONNX models align! ✨

✨ Finishing touches
  • 📝 Generate docstrings

📜 Recent review details

Configuration used: defaults

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between e061d42 and e9cee7f.

📒 Files selected for processing (30)
  • cxx-api-examples/CMakeLists.txt
  • cxx-api-examples/funasr-nano-cxx-api.cc
  • python-api-examples/offline-funasr-nano-decode-files.py
  • sherpa-onnx/c-api/c-api.cc
  • sherpa-onnx/c-api/c-api.h
  • sherpa-onnx/c-api/cxx-api.cc
  • sherpa-onnx/c-api/cxx-api.h
  • sherpa-onnx/csrc/CMakeLists.txt
  • sherpa-onnx/csrc/funasr-nano-tokenizer.cc
  • sherpa-onnx/csrc/funasr-nano-tokenizer.h
  • sherpa-onnx/csrc/offline-fire-red-asr-model.cc
  • sherpa-onnx/csrc/offline-funasr-nano-model-config.cc
  • sherpa-onnx/csrc/offline-funasr-nano-model-config.h
  • sherpa-onnx/csrc/offline-funasr-nano-model.cc
  • sherpa-onnx/csrc/offline-funasr-nano-model.h
  • sherpa-onnx/csrc/offline-model-config.cc
  • sherpa-onnx/csrc/offline-model-config.h
  • sherpa-onnx/csrc/offline-recognizer-funasr-nano-impl.cc
  • sherpa-onnx/csrc/offline-recognizer-funasr-nano-impl.h
  • sherpa-onnx/csrc/offline-recognizer-impl.cc
  • sherpa-onnx/csrc/onnx-utils.cc
  • sherpa-onnx/csrc/onnx-utils.h
  • sherpa-onnx/csrc/sherpa-onnx-offline.cc
  • sherpa-onnx/csrc/sherpa-onnx-vad-with-offline-asr.cc
  • sherpa-onnx/python/csrc/CMakeLists.txt
  • sherpa-onnx/python/csrc/offline-funasr-nano-model-config.cc
  • sherpa-onnx/python/csrc/offline-funasr-nano-model-config.h
  • sherpa-onnx/python/csrc/offline-model-config.cc
  • sherpa-onnx/python/sherpa_onnx/__init__.py
  • sherpa-onnx/python/sherpa_onnx/offline_recognizer.py

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@Wasser1462 Wasser1462 changed the title Feature/ort iobinding fire red Feature/ort iobinding Fire Red ASR Jan 9, 2026
@gemini-code-assist

Copy link
Copy Markdown

Summary of Changes

Hello @Wasser1462, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly enhances the project's capabilities by introducing a new, efficient FunASR-nano model, complete with its own tokenizer and API examples. Concurrently, it delivers a substantial performance boost to the existing FireRedASR model through optimized GPU-CPU data handling, demonstrating a marked improvement in processing speed.

Highlights

  • FireRedASR Performance Optimization: Implemented ONNX Runtime I/O Binding for both the encoder and decoder of the FireRedASR model. This significantly reduces GPU↔CPU data transfers, leading to a substantial improvement in the Real-Time Factor (RTF) from 0.152 to 0.083 on an RTX 4090 with CUDA.
  • FunASR-nano Model Integration: Introduced comprehensive support for the FunASR-nano offline ASR model. This includes its encoder adaptor, LLM prefill/decode, and embedding components, expanding the range of supported ASR architectures.
  • Custom Qwen3 Tokenizer: Developed a self-contained ByteLevel-BPE tokenizer specifically for Qwen3, which is utilized by the FunASR-nano model. This new tokenizer handles added tokens via a Trie longest-match approach and includes byte-level encoding/decoding.
  • API Examples and Utilities: Added new C++ and Python API examples to demonstrate the usage of the FunASR-nano model. Additionally, new utility functions for converting between FP16 and FP32 floating-point formats were introduced, crucial for mixed-precision model inference.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@Wasser1462 Wasser1462 closed this Jan 9, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces two significant and largely independent changes. The first is an excellent performance optimization for the FireRedASR model using ONNX Runtime I/O Binding, which shows impressive benchmark improvements. The second is the addition of a new model, FunASR-nano, including its tokenizer, model implementation, recognizer, and various examples.

While both features are valuable, combining them into a single pull request makes it difficult to review and understand the scope of changes. I strongly recommend splitting this PR into two separate ones:

  1. One PR for the FireRedASR I/O binding optimization.
  2. A second PR for adding the FunASR-nano model support.

This separation will allow for a more focused review of each feature and will make the git history cleaner and more understandable. My detailed comments below cover aspects of both features, but I urge you to consider splitting the PR before merging.

./test.wav
)usage";

if (argc < 6) {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The argument count check argc < 6 seems incorrect. The usage message indicates 5 required flag arguments and 1 positional audio file argument, for a total of 6 arguments. This means argc should be at least 7 (including the program name). An argc of 6 would mean one argument is missing. Please consider changing the check to argc < 7 for a more accurate pre-condition check.

Suggested change
if (argc < 6) {
if (argc < 7) {

Comment on lines +660 to +667
while (j < text.size()) {
size_t t = j;
uint32_t cx = 0;
size_t nx = 0;
if (!Utf8Next(text, &t, &cx, &nx)) break;
if (!is_punct_like(cx)) break;
j += nx;
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The indentation in this while loop is misleading and could lead to confusion. The if statements without braces and inconsistent indentation make the control flow hard to follow. For better readability and maintainability, please consider using braces and consistent indentation.

        while (j < text.size()) {
          size_t t = j;
          uint32_t cx = 0;
          size_t nx = 0;
          if (!Utf8Next(text, &t, &cx, &nx)) {
            break;
          }
          if (!is_punct_like(cx)) {
            break;
          }
          j += nx;
        }

system_prompt, user_prompt, audio_token_len, fbank_beg_idx,
fake_token_len);
int32_t context_len = static_cast<int32_t>(source_ids.size());
const int32_t max_seq_len = 2048;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The value 2048 for max_seq_len is a magic number. It would be better to define it as a named constant (e.g., kMaxSeqLen) to improve readability and make it easier to change if needed. Consider defining it as a static constexpr at a suitable scope (like class or file level).

@Wasser1462
Wasser1462 deleted the feature/ort-iobinding-fire-red branch January 9, 2026 02:24
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL This PR changes 1000+ lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant