Add C++ runtime for Paraformer on Ascend NPU. - #2741
Conversation
|
Caution Review failedThe pull request is closed. Note Other AI code review bot(s) detectedCodeRabbit has detected other AI code review bot(s) in this pull request and will avoid duplicating their findings in the review comments. This may lead to a less comprehensive review. WalkthroughAdds Ascend NPU support for Paraformer: new Ascend model implementation, recognizer integration, CMake wiring, and updated Paraformer model config validation for three-file .om format. Also minor include/comment cleanups and error-message updates. Changes
Sequence Diagram(s)sequenceDiagram
actor User
participant Factory as OfflineRecognizer (factory)
participant Impl as OfflineRecognizerParaformerAscendImpl
participant Model as OfflineParaformerModelAscend
participant Device as AscendDevice
User->>Factory: Create(...)
Factory->>Impl: instantiate (Ascend path)
User->>Factory: CreateStream()/DecodeStreams(streams)
Factory->>Impl: DecodeStreams(streams)
rect rgb(235,245,255)
Impl->>Impl: Gather frames
Impl->>Model: Run(features)
rect rgb(220,235,255)
Note over Model,Device: Ascend multi-stage inference
Model->>Device: Upload features
Device->>Model: Encoder (encoder.om)
Device->>Model: Predictor (predictor.om)
Device->>Model: Decoder (decoder.om)
Model->>Device: Compute acoustic embedding / collect logits
Device-->>Model: logits (host)
end
Model-->>Impl: logits (vector<float>)
Impl->>Impl: Argmax → token IDs, post-process
Impl-->>Factory: Set stream result
end
Factory-->>User: GetResult(stream)
Estimated code review effort🎯 4 (Complex) | ⏱️ ~45 minutes Areas needing extra attention:
Possibly related PRs
Suggested labels
Poem
Pre-merge checks and finishing touches✅ Passed checks (2 passed)
📜 Recent review detailsConfiguration used: CodeRabbit UI Review profile: CHILL Plan: Pro 📒 Files selected for processing (4)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 2
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
sherpa-onnx/csrc/offline-recognizer-paraformer-impl.h (1)
25-81: Restore internal linkage (or inline) forConvertDropping
staticleft this function with external linkage while still being defined in the header. Any translation unit that includes this header now emits a global definition, so linking succeeds only if exactly one TU includes it. Once another TU (e.g., the new Ascend recognizer) includes this header, the build breaks with multiple-definition errors. Please keep it header-only but mark itinline(or move it to a.ccfile with a declaration).Apply this diff:
-OfflineRecognitionResult Convert(const OfflineParaformerDecoderResult &src, - const SymbolTable &sym_table) { +inline OfflineRecognitionResult Convert( + const OfflineParaformerDecoderResult &src, + const SymbolTable &sym_table) {
🧹 Nitpick comments (1)
sherpa-onnx/csrc/ascend/offline-recognizer-paraformer-ascend-impl.h (1)
17-118: Remove the unused decoder member.
decoder_(and its include) are never used—the greedy path is implemented inline withstd::max_element. Dropping the dead member will keep the class lean and avoid future confusion.Apply this diff:
-#include "sherpa-onnx/csrc/rknn/offline-ctc-greedy-search-decoder-rknn.h" ... - std::unique_ptr<OfflineCtcGreedySearchDecoderRknn> decoder_;
📜 Review details
Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro
📒 Files selected for processing (11)
sherpa-onnx/csrc/CMakeLists.txt(1 hunks)sherpa-onnx/csrc/ascend/offline-paraformer-model-ascend.cc(1 hunks)sherpa-onnx/csrc/ascend/offline-paraformer-model-ascend.h(1 hunks)sherpa-onnx/csrc/ascend/offline-recognizer-paraformer-ascend-impl.h(1 hunks)sherpa-onnx/csrc/ascend/offline-recognizer-sense-voice-ascend-impl.h(1 hunks)sherpa-onnx/csrc/ascend/offline-sense-voice-model-ascend.cc(1 hunks)sherpa-onnx/csrc/ascend/utils.cc(1 hunks)sherpa-onnx/csrc/offline-paraformer-model-config.cc(1 hunks)sherpa-onnx/csrc/offline-paraformer-model-config.h(1 hunks)sherpa-onnx/csrc/offline-recognizer-impl.cc(3 hunks)sherpa-onnx/csrc/offline-recognizer-paraformer-impl.h(3 hunks)
🧰 Additional context used
🧠 Learnings (1)
📚 Learning: 2025-08-06T04:23:50.237Z
Learnt from: litongjava
Repo: k2-fsa/sherpa-onnx PR: 2440
File: sherpa-onnx/java-api/src/main/java/com/k2fsa/sherpa/onnx/core/Core.java:4-6
Timestamp: 2025-08-06T04:23:50.237Z
Learning: The sherpa-onnx JNI library files are stored in Hugging Face repository at https://huggingface.co/csukuangfj/sherpa-onnx-libs under versioned directories like jni/1.12.7/, and the actual Windows JNI library filename is "sherpa-onnx-jni.dll" as defined in Core.java constants.
Applied to files:
sherpa-onnx/csrc/offline-recognizer-impl.cc
🧬 Code graph analysis (5)
sherpa-onnx/csrc/ascend/offline-paraformer-model-ascend.h (2)
sherpa-onnx/csrc/ascend/offline-recognizer-paraformer-ascend-impl.h (1)
sherpa_onnx(20-42)sherpa-onnx/csrc/ascend/offline-paraformer-model-ascend.cc (15)
OfflineParaformerModelAscend(424-426)OfflineParaformerModelAscend(428-428)OfflineParaformerModelAscend(440-441)OfflineParaformerModelAscend(445-446)Run(430-433)Run(430-431)features(130-176)features(130-130)features(181-210)features(181-181)VocabSize(435-437)VocabSize(435-435)Impl(82-98)Impl(82-82)Impl(101-128)
sherpa-onnx/csrc/ascend/offline-recognizer-paraformer-ascend-impl.h (4)
sherpa-onnx/csrc/ascend/offline-paraformer-model-ascend.h (2)
sherpa_onnx(12-37)OfflineParaformerModelAscend(14-35)sherpa-onnx/csrc/offline-recognizer-paraformer-impl.h (4)
sherpa_onnx(23-203)Convert(25-81)- `` (123-125)
DecodeStreams(127-183)sherpa-onnx/csrc/offline-recognizer-impl.cc (8)
OfflineRecognizerImpl(508-559)OfflineRecognizerImpl(562-612)OfflineRecognizerImpl(641-642)OfflineRecognizerImpl(649-650)ApplyInverseTextNormalization(614-625)ApplyInverseTextNormalization(614-615)ApplyHomophoneReplacer(627-634)ApplyHomophoneReplacer(627-628)sherpa-onnx/csrc/rknn/offline-ctc-greedy-search-decoder-rknn.h (1)
OfflineCtcGreedySearchDecoderRknn(14-24)
sherpa-onnx/csrc/offline-recognizer-paraformer-impl.h (3)
sherpa-onnx/csrc/offline-recognizer-transducer-impl.h (1)
Convert(33-115)sherpa-onnx/csrc/offline-recognizer-ctc-impl.h (1)
Convert(25-289)sherpa-onnx/csrc/online-recognizer-paraformer-impl.h (1)
Convert(25-80)
sherpa-onnx/csrc/offline-paraformer-model-config.cc (3)
sherpa-onnx/csrc/offline-tts.cc (2)
Validate(121-150)Validate(121-121)sherpa-onnx/csrc/online-model-config.cc (2)
Validate(55-162)Validate(55-55)sherpa-onnx/csrc/file-utils.cc (2)
FileExists(16-18)FileExists(16-16)
sherpa-onnx/csrc/ascend/offline-paraformer-model-ascend.cc (2)
sherpa-onnx/csrc/ascend/offline-sense-voice-model-ascend.cc (10)
out(180-180)features(53-108)features(53-54)data(119-125)data(119-119)max_num_frames_(145-155)in(157-193)in(157-157)Run(220-223)Run(220-221)sherpa-onnx/csrc/ascend/offline-paraformer-model-ascend.h (1)
OfflineParaformerModelAscend(14-35)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (19)
- GitHub Check: ubuntu-latest Debug static tts-ON
- GitHub Check: ubuntu-latest Debug static tts-OFF
- GitHub Check: ubuntu-latest Debug shared tts-OFF
- GitHub Check: ubuntu-latest Release static tts-ON
- GitHub Check: ubuntu-latest Release shared tts-ON
- GitHub Check: Release shared-OFF tts-ON
- GitHub Check: Release shared-OFF tts-OFF
- GitHub Check: swift (macos-latest)
- GitHub Check: Debug shared-ON tts-ON
- GitHub Check: Debug shared-OFF tts-ON
- GitHub Check: Debug shared-OFF tts-OFF
- GitHub Check: Debug shared-ON tts-OFF
- GitHub Check: swift (macos-13)
- GitHub Check: Release shared-ON tts-ON
- GitHub Check: Release shared-ON tts-OFF
- GitHub Check: Debug shared tts-ON
- GitHub Check: Release static tts-ON
- GitHub Check: rknn shared OFF
- GitHub Check: rknn shared ON
🔇 Additional comments (1)
sherpa-onnx/csrc/ascend/offline-paraformer-model-ascend.h (1)
14-35: Interface mirrors existing backends.The PIMPL layout and minimal public surface match the other offline Paraformer model headers. Looks good.
| std::vector<float> ApplyLFR(std::vector<float> in) const { | ||
| int32_t lfr_window_size = 7; | ||
| int32_t lfr_window_shift = 6; | ||
| int32_t in_feat_dim = 80; | ||
|
|
||
| int32_t in_num_frames = in.size() / in_feat_dim; | ||
| int32_t out_num_frames = | ||
| (in_num_frames - lfr_window_size) / lfr_window_shift + 1; | ||
|
|
||
| if (out_num_frames > max_num_frames_) { | ||
| SHERPA_ONNX_LOGE( | ||
| "Number of input frames %d is too large. Truncate it to %d frames.", | ||
| out_num_frames, max_num_frames_); | ||
|
|
||
| SHERPA_ONNX_LOGE( | ||
| "Recognition result may be truncated/incomplete. Please select a " | ||
| "model accepting longer audios."); | ||
|
|
||
| out_num_frames = max_num_frames_; | ||
| } | ||
|
|
||
| int32_t out_feat_dim = in_feat_dim * lfr_window_size; | ||
|
|
||
| std::vector<float> out(out_num_frames * out_feat_dim); | ||
|
|
||
| const float *p_in = in.data(); | ||
| float *p_out = out.data(); | ||
|
|
||
| for (int32_t i = 0; i != out_num_frames; ++i) { | ||
| std::copy(p_in, p_in + out_feat_dim, p_out); | ||
|
|
||
| p_out += out_feat_dim; | ||
| p_in += lfr_window_shift * in_feat_dim; | ||
| } |
There was a problem hiding this comment.
Guard LFR for short utterances to avoid OOB reads.
When the input has fewer than lfr_window_size frames, (in_num_frames - lfr_window_size) / lfr_window_shift + 1 still yields 1 because of integer truncation. The loop then copies 560 floats from a buffer that may contain fewer than 560 entries, causing an out-of-bounds read.
Apply this fix:
int32_t in_num_frames = in.size() / in_feat_dim;
+ if (in_num_frames < lfr_window_size) {
+ return {};
+ }
int32_t out_num_frames =
(in_num_frames - lfr_window_size) / lfr_window_shift + 1;📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| std::vector<float> ApplyLFR(std::vector<float> in) const { | |
| int32_t lfr_window_size = 7; | |
| int32_t lfr_window_shift = 6; | |
| int32_t in_feat_dim = 80; | |
| int32_t in_num_frames = in.size() / in_feat_dim; | |
| int32_t out_num_frames = | |
| (in_num_frames - lfr_window_size) / lfr_window_shift + 1; | |
| if (out_num_frames > max_num_frames_) { | |
| SHERPA_ONNX_LOGE( | |
| "Number of input frames %d is too large. Truncate it to %d frames.", | |
| out_num_frames, max_num_frames_); | |
| SHERPA_ONNX_LOGE( | |
| "Recognition result may be truncated/incomplete. Please select a " | |
| "model accepting longer audios."); | |
| out_num_frames = max_num_frames_; | |
| } | |
| int32_t out_feat_dim = in_feat_dim * lfr_window_size; | |
| std::vector<float> out(out_num_frames * out_feat_dim); | |
| const float *p_in = in.data(); | |
| float *p_out = out.data(); | |
| for (int32_t i = 0; i != out_num_frames; ++i) { | |
| std::copy(p_in, p_in + out_feat_dim, p_out); | |
| p_out += out_feat_dim; | |
| p_in += lfr_window_shift * in_feat_dim; | |
| } | |
| std::vector<float> ApplyLFR(std::vector<float> in) const { | |
| int32_t lfr_window_size = 7; | |
| int32_t lfr_window_shift = 6; | |
| int32_t in_feat_dim = 80; | |
| int32_t in_num_frames = in.size() / in_feat_dim; | |
| if (in_num_frames < lfr_window_size) { | |
| return {}; | |
| } | |
| int32_t out_num_frames = | |
| (in_num_frames - lfr_window_size) / lfr_window_shift + 1; | |
| if (out_num_frames > max_num_frames_) { | |
| SHERPA_ONNX_LOGE( | |
| "Number of input frames %d is too large. Truncate it to %d frames.", | |
| out_num_frames, max_num_frames_); | |
| SHERPA_ONNX_LOGE( | |
| "Recognition result may be truncated/incomplete. Please select a " | |
| "model accepting longer audios."); | |
| out_num_frames = max_num_frames_; | |
| } | |
| int32_t out_feat_dim = in_feat_dim * lfr_window_size; | |
| std::vector<float> out(out_num_frames * out_feat_dim); | |
| const float *p_in = in.data(); | |
| float *p_out = out.data(); | |
| for (int32_t i = 0; i != out_num_frames; ++i) { | |
| std::copy(p_in, p_in + out_feat_dim, p_out); | |
| p_out += out_feat_dim; | |
| p_in += lfr_window_shift * in_feat_dim; | |
| } |
🤖 Prompt for AI Agents
In sherpa-onnx/csrc/ascend/offline-paraformer-model-ascend.cc around lines 362
to 395, the LFR logic can read out-of-bounds when in_num_frames <
lfr_window_size because the out_num_frames formula yields 1; change the
computation to guard short utterances by checking if in_num_frames <
lfr_window_size and if so set out_num_frames = 0 (or return an empty vector)
before allocating/copying, otherwise compute out_num_frames = (in_num_frames -
lfr_window_size) / lfr_window_shift + 1 and proceed; ensure you also clamp
out_num_frames to max_num_frames_ and skip the copy-loop when out_num_frames ==
0 to avoid OOB reads.
There was a problem hiding this comment.
Pull Request Overview
This pull request adds support for Paraformer models on Ascend NPU for offline speech recognition. The changes extend the existing Ascend NPU support (which previously only supported SenseVoice models) to also handle Paraformer models.
Key changes:
- Added Ascend NPU implementation for Paraformer model support
- Made the
Convertfunction non-static to allow external usage - Replaced
exit(-1)calls withSHERPA_ONNX_EXIT(-1)for better error handling - Enhanced error messages with documentation links
Reviewed Changes
Copilot reviewed 11 out of 11 changed files in this pull request and generated 3 comments.
Show a summary per file
| File | Description |
|---|---|
| sherpa-onnx/csrc/offline-recognizer-paraformer-impl.h | Removed static keyword from Convert function to enable external usage |
| sherpa-onnx/csrc/offline-recognizer-impl.cc | Added Paraformer model support for Ascend NPU, improved error messages with documentation links |
| sherpa-onnx/csrc/offline-paraformer-model-config.h | Added documentation comment for Ascend NPU model path format |
| sherpa-onnx/csrc/offline-paraformer-model-config.cc | Added validation logic for Ascend NPU model files (.om format) |
| sherpa-onnx/csrc/ascend/utils.cc | Added missing <utility> include |
| sherpa-onnx/csrc/ascend/offline-sense-voice-model-ascend.cc | Reordered includes alphabetically |
| sherpa-onnx/csrc/ascend/offline-recognizer-sense-voice-ascend-impl.h | Fixed comment reference from "online" to "offline" |
| sherpa-onnx/csrc/ascend/offline-recognizer-paraformer-ascend-impl.h | New implementation file for Paraformer Ascend NPU recognizer |
| sherpa-onnx/csrc/ascend/offline-paraformer-model-ascend.h | New header for Paraformer Ascend NPU model |
| sherpa-onnx/csrc/ascend/offline-paraformer-model-ascend.cc | New implementation of Paraformer model for Ascend NPU |
| sherpa-onnx/csrc/CMakeLists.txt | Added new Paraformer Ascend source file to build |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| if (EndsWith(model, ".onnx") && !FileExists(model)) { | ||
| SHERPA_ONNX_LOGE("Paraformer model '%s' does not exist", model.c_str()); | ||
| return false; | ||
| } else if (EndsWith(model, ".om")) { |
There was a problem hiding this comment.
The pattern \"*.om\" is treated as a literal string, not a wildcard pattern. The asterisk will be matched as a literal character. This should likely be \".om\" to check if the model path ends with the .om extension.
|
|
||
| namespace sherpa_onnx { | ||
|
|
||
| // defined in ../online-recognizer-paraformer-impl.h |
There was a problem hiding this comment.
The comment incorrectly references 'online-recognizer-paraformer-impl.h' but should reference 'offline-recognizer-paraformer-impl.h' since this is for offline recognition, not online. This was correctly fixed for SenseVoice in line 22 of offline-recognizer-sense-voice-ascend-impl.h.
| // defined in ../online-recognizer-paraformer-impl.h | |
| // defined in ../offline-recognizer-paraformer-impl.h |
| OfflineRecognizerConfig config_; | ||
| SymbolTable symbol_table_; | ||
| std::unique_ptr<OfflineParaformerModelAscend> model_; | ||
| std::unique_ptr<OfflineCtcGreedySearchDecoderRknn> decoder_; |
There was a problem hiding this comment.
The member variable decoder_ is declared but never initialized or used in the class implementation. This appears to be dead code that should be removed.
| std::unique_ptr<OfflineCtcGreedySearchDecoderRknn> decoder_; |
See also #2697
Usage
Build sherpa-onnx
Please follow
https://k2-fsa.github.io/sherpa/onnx/ascend/install.html
Download a test model
Run it
./bin/sherpa-onnx-offline \ --provider=ascend \ --paraformer="sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/encoder.om,sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/predictor.om,sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/decoder.om" \ --tokens=sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/tokens.txt \ sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/test_wavs/0.wav \ sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/test_wavs/1.wav \ sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/test_wavs/2.wav \ sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/test_wavs/3-sichuan.wavNote that we have split the model into 3 parts. To make the interface compatible with onnx models, we pass the three models at once; the files are separated with a comma.
The output is given below:
Summary by CodeRabbit
New Features
Bug Fixes / Improvements
Documentation