Feature/ort iobinding Fire Red ASR - #3010
Wasser1462 wants to merge 10 commits into
Conversation
|
Caution Review failedThe pull request is closed. 📝 WalkthroughWalkthroughThis PR introduces comprehensive FunASR-nano offline speech recognition support across the sherpa-onnx library, including new model configuration structures, tokenizer implementation, offline recognizer backend, C++ and Python API bindings, example programs, and platform-specific constructors for Android and OHOS environments. Changes
Sequence Diagram(s)sequenceDiagram
actor User
participant OfflineRecognizer
participant OfflineFunASRNanoImpl
participant OfflineFunASRNanoModel
participant FunASRNanoTokenizer
participant ORT as ONNX Runtime<br/>(Encoder/LLM/Embedding)
User->>OfflineRecognizer: Create(config)
OfflineRecognizer->>OfflineFunASRNanoImpl: new (config)
OfflineFunASRNanoImpl->>OfflineFunASRNanoModel: new (config)
OfflineFunASRNanoModel->>ORT: Load encoder, prefill, decode, embedding sessions
OfflineFunASRNanoImpl->>FunASRNanoTokenizer: new (tokenizer_dir)
FunASRNanoTokenizer->>FunASRNanoTokenizer: Load vocab, merges, tokenizer.json
User->>OfflineRecognizer: CreateStream()
OfflineRecognizer->>OfflineFunASRNanoImpl: CreateStream()
OfflineFunASRNanoImpl-->>OfflineRecognizer: OfflineStream
User->>OfflineRecognizer: Decode(stream with audio)
OfflineRecognizer->>OfflineFunASRNanoImpl: DecodeStreams(streams)
OfflineFunASRNanoImpl->>OfflineFunASRNanoModel: ForwardEncoderAdaptor(features)
ORT-->>OfflineFunASRNanoModel: encoder_output
OfflineFunASRNanoImpl->>FunASRNanoTokenizer: Encode(system/user prompts)
FunASRNanoTokenizer-->>OfflineFunASRNanoImpl: token_ids
loop Autoregressive Generation (up to max_new_tokens)
OfflineFunASRNanoImpl->>OfflineFunASRNanoModel: ForwardLLMPrefill(inputs_embeds, attention_mask)
ORT-->>OfflineFunASRNanoModel: logits, past_key_values
OfflineFunASRNanoImpl->>OfflineFunASRNanoModel: ForwardLLMDecode(inputs_embeds, attention_mask, cache)
ORT-->>OfflineFunASRNanoModel: next_logits, updated_cache
OfflineFunASRNanoImpl->>OfflineFunASRNanoImpl: SampleToken(logits)
end
OfflineFunASRNanoImpl->>FunASRNanoTokenizer: Decode(generated_token_ids)
FunASRNanoTokenizer-->>OfflineFunASRNanoImpl: transcription_text
OfflineFunASRNanoImpl->>OfflineRecognizer: Set result on stream
OfflineRecognizer->>User: GetResult() with transcription
Estimated code review effort🎯 5 (Critical) | ⏱️ ~120 minutes Possibly related PRs
Suggested reviewers
Poem
✨ Finishing touches
📜 Recent review detailsConfiguration used: defaults Review profile: CHILL Plan: Pro 📒 Files selected for processing (30)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Summary of ChangesHello @Wasser1462, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request significantly enhances the project's capabilities by introducing a new, efficient FunASR-nano model, complete with its own tokenizer and API examples. Concurrently, it delivers a substantial performance boost to the existing FireRedASR model through optimized GPU-CPU data handling, demonstrating a marked improvement in processing speed. Highlights
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here. You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension. Footnotes
|
There was a problem hiding this comment.
Code Review
This pull request introduces two significant and largely independent changes. The first is an excellent performance optimization for the FireRedASR model using ONNX Runtime I/O Binding, which shows impressive benchmark improvements. The second is the addition of a new model, FunASR-nano, including its tokenizer, model implementation, recognizer, and various examples.
While both features are valuable, combining them into a single pull request makes it difficult to review and understand the scope of changes. I strongly recommend splitting this PR into two separate ones:
- One PR for the FireRedASR I/O binding optimization.
- A second PR for adding the FunASR-nano model support.
This separation will allow for a more focused review of each feature and will make the git history cleaner and more understandable. My detailed comments below cover aspects of both features, but I urge you to consider splitting the PR before merging.
| ./test.wav | ||
| )usage"; | ||
|
|
||
| if (argc < 6) { |
There was a problem hiding this comment.
The argument count check argc < 6 seems incorrect. The usage message indicates 5 required flag arguments and 1 positional audio file argument, for a total of 6 arguments. This means argc should be at least 7 (including the program name). An argc of 6 would mean one argument is missing. Please consider changing the check to argc < 7 for a more accurate pre-condition check.
| if (argc < 6) { | |
| if (argc < 7) { |
| while (j < text.size()) { | ||
| size_t t = j; | ||
| uint32_t cx = 0; | ||
| size_t nx = 0; | ||
| if (!Utf8Next(text, &t, &cx, &nx)) break; | ||
| if (!is_punct_like(cx)) break; | ||
| j += nx; | ||
| } |
There was a problem hiding this comment.
The indentation in this while loop is misleading and could lead to confusion. The if statements without braces and inconsistent indentation make the control flow hard to follow. For better readability and maintainability, please consider using braces and consistent indentation.
while (j < text.size()) {
size_t t = j;
uint32_t cx = 0;
size_t nx = 0;
if (!Utf8Next(text, &t, &cx, &nx)) {
break;
}
if (!is_punct_like(cx)) {
break;
}
j += nx;
}| system_prompt, user_prompt, audio_token_len, fbank_beg_idx, | ||
| fake_token_len); | ||
| int32_t context_len = static_cast<int32_t>(source_ids.size()); | ||
| const int32_t max_seq_len = 2048; |
There was a problem hiding this comment.
FireRedASR Performance Optimization with ONNX Runtime I/O Binding
Summary
This PR addresses sherpa-onnx issue #2943 by adopting ONNX Runtime I/O Binding (
Ort::IoBinding) for both the encoder and decoder in FireRedASR offline ASR model to minimize GPU↔CPU transfers and improve end-to-end GPU latency.Benchmark
Test Audio:
sherpa-onnx-fire-red-asr-large-zh_en-2025-02-16/test_wavs/0.wavGPU: RTX 4090
Provider: CUDA
Summary by CodeRabbit
New Features
Documentation
✏️ Tip: You can customize this high-level summary in your review settings.