Skip to content

Add C++ runtime for Paraformer on Ascend NPU. - #2741

Merged
csukuangfj merged 8 commits into
k2-fsa:masterfrom
csukuangfj:ascend-npu-paraformer-cpp
Nov 3, 2025
Merged

csukuangfj merged 8 commits into
k2-fsa:masterfrom
csukuangfj:ascend-npu-paraformer-cpp

Conversation

@csukuangfj

@csukuangfj csukuangfj commented Nov 3, 2025 •

Copy link
Copy Markdown
Collaborator

See also #2697

Usage

Build sherpa-onnx

Please follow
https://k2-fsa.github.io/sherpa/onnx/ascend/install.html

Download a test model

wget https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/sherpa-onnx-ascend-910B2-paraformer-zh-2023-03-28.tar.bz2

tar xvf sherpa-onnx-ascend-910B2-paraformer-zh-2023-03-28.tar.bz2

Run it

./bin/sherpa-onnx-offline \
  --provider=ascend \
  --paraformer="sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/encoder.om,sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/predictor.om,sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/decoder.om" \
  --tokens=sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/tokens.txt \
  sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/test_wavs/0.wav \
  sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/test_wavs/1.wav \
  sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/test_wavs/2.wav \
  sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/test_wavs/3-sichuan.wav

Note that we have split the model into 3 parts. To make the interface compatible with onnx models, we pass the three models at once; the files are separated with a comma.

The output is given below:

/root/open-source/sherpa-onnx/sherpa-onnx/csrc/parse-options.cc:Read:372 ./bin/sherpa-onnx-offline --provider=ascend --paraformer=sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/encoder.om,sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/predictor.om,sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/decoder.om --tokens=sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/tokens.txt sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/test_wavs/0.wav sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/test_wavs/1.wav sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/test_wavs/2.wav sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/test_wavs/3-sichuan.wav

OfflineRecognizerConfig(feat_config=FeatureExtractorConfig(sampling_rate=16000, feature_dim=80, low_freq=20, high_freq=-400, dither=0, normalize_samples=True, snip_edges=False), model_config=OfflineModelConfig(transducer=OfflineTransducerModelConfig(encoder_filename="", decoder_filename="", joiner_filename=""), paraformer=OfflineParaformerModelConfig(model="sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/encoder.om,sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/predictor.om,sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/decoder.om"), nemo_ctc=OfflineNemoEncDecCtcModelConfig(model=""), whisper=OfflineWhisperModelConfig(encoder="", decoder="", language="", task="transcribe", tail_paddings=-1), fire_red_asr=OfflineFireRedAsrModelConfig(encoder="", decoder=""), tdnn=OfflineTdnnModelConfig(model=""), zipformer_ctc=OfflineZipformerCtcModelConfig(model=""), wenet_ctc=OfflineWenetCtcModelConfig(model=""), sense_voice=OfflineSenseVoiceModelConfig(model="", language="auto", use_itn=False), moonshine=OfflineMoonshineModelConfig(preprocessor="", encoder="", uncached_decoder="", cached_decoder=""), dolphin=OfflineDolphinModelConfig(model=""), canary=OfflineCanaryModelConfig(encoder="", decoder="", src_lang="", tgt_lang="", use_pnc=True), telespeech_ctc="", tokens="sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/tokens.txt", num_threads=2, debug=False, provider="ascend", model_type="", modeling_unit="cjkchar", bpe_vocab=""), lm_config=OfflineLMConfig(model="", scale=0.5, lodr_scale=0.01, lodr_fst="", lodr_backoff_id=-1), ctc_fst_decoder_config=OfflineCtcFstDecoderConfig(graph="", max_active=3000), decoding_method="greedy_search", max_active_paths=4, hotwords_file="", hotwords_score=1.5, blank_penalty=0, rule_fsts="", rule_fars="", hr=HomophoneReplacerConfig(lexicon="", rule_fsts=""))
Creating recognizer ...
Started
Done!

sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/test_wavs/0.wav
{"lang": "", "emotion": "", "event": "", "text": "对我做了介绍啊那么我想说的是呢大家如果对我的研究感兴趣呢嗯", "timestamps": [], "durations": [], "tokens":["对", "我", "做", "
了", "介", "绍", "啊", "那", "么", "我", "想", "说", "的", "是", "呢", "大", "家", "如", "果", "对", "我", "的", "研", "究", "感", "兴", "趣", "呢", "嗯"], "words": []}
----
sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/test_wavs/1.wav
{"lang": "", "emotion": "", "event": "", "text": "重点呢想谈三个问题首先呢就是这一轮全球金融动荡的表现", "timestamps": [], "durations": [], "tokens":["重", "点", "呢", "想", "
谈", "三", "个", "问", "题", "首", "先", "呢", "就", "是", "这", "一", "轮", "全", "球", "金", "融", "动", "荡", "的", "表", "现"], "words": []}
----
sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/test_wavs/2.wav
{"lang": "", "emotion": "", "event": "", "text": "深入的分析这一次全球金融动荡背后的根源", "timestamps": [], "durations": [], "tokens":["深", "入", "的", "分", "析", "这", "一", "次", "全", "球", "金", "融", "动", "荡", "背", "后", "的", "根", "源"], "words": []}
----
sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28/test_wavs/3-sichuan.wav
{"lang": "", "emotion": "", "event": "", "text": "自己就是在那个在那个就是在情节里面就是感觉是演的特别好就是好像很真实一样你知道吧", "timestamps": [], "durations": [], "tokens":["自", "己", "就", "是", "在", "那", "个", "在", "那", "个", "就", "是", "在", "情", "节", "里", "面", "就", "是", "感", "觉", "是", "演", "的", "特", "别", "好", "就", "是", "好", "像", "很", "真", "实", "一", "样", "你", "知", "道", "吧"], "words": []}
----
num threads: 2
decoding method: greedy_search
Elapsed seconds: 0.264 s
Real time factor (RTF): 0.264 / 23.133 = 0.011

Summary by CodeRabbit

  • New Features

    • Added support for running Paraformer offline speech recognition on Ascend NPU accelerators.
    • Extended model configuration to accept Ascend multi-file model format (encoder/predictor/decoder).
  • Bug Fixes / Improvements

    • Improved provider selection and clearer error messages with installation guidance for Ascend-enabled paths.
    • Enhanced model validation with detailed checks and helpful error reporting.
  • Documentation

    • Added usage comments describing Ascend model format expectations.

@dosubot dosubot Bot added the size:XL This PR changes 500-999 lines, ignoring generated files. label Nov 3, 2025
@coderabbitai

coderabbitai Bot commented Nov 3, 2025 •

Copy link
Copy Markdown

Caution

Review failed

The pull request is closed.

Note

Other AI code review bot(s) detected

CodeRabbit has detected other AI code review bot(s) in this pull request and will avoid duplicating their findings in the review comments. This may lead to a less comprehensive review.

Walkthrough

Adds Ascend NPU support for Paraformer: new Ascend model implementation, recognizer integration, CMake wiring, and updated Paraformer model config validation for three-file .om format. Also minor include/comment cleanups and error-message updates.

Changes

Cohort / File(s) Change Summary
Ascend Paraformer Model Core
sherpa-onnx/csrc/ascend/offline-paraformer-model-ascend.h, sherpa-onnx/csrc/ascend/offline-paraformer-model-ascend.cc
New OfflineParaformerModelAscend class (PIMPL). Implements Ascend-based encoder/predictor/decoder pipeline, device memory management, host-device transfers, LFR feature handling, thread-safety, and platform-specific constructor instantiations.
Ascend Recognizer Integration
sherpa-onnx/csrc/ascend/offline-recognizer-paraformer-ascend-impl.h
New OfflineRecognizerParaformerAscendImpl class extending recognizer impl; enforces greedy decoding, creates streams, runs model, converts logits→tokens, and applies post-processing (ITN, homophone replacement).
Build Configuration
sherpa-onnx/csrc/CMakeLists.txt
Adds ./ascend/offline-paraformer-model-ascend.cc to Ascend NPU build sources when SHERPA_ONNX_ENABLE_ASCEND_NPU is enabled.
Paraformer Configuration
sherpa-onnx/csrc/offline-paraformer-model-config.h, sherpa-onnx/csrc/offline-paraformer-model-config.cc
Documentation update and enhanced validation: supports comma-separated three-file .om format (encoder.om,predictor.om,decoder.om) and validates ordering/existence; retains ONNX handling.
Recognizer Creation Logic
sherpa-onnx/csrc/offline-recognizer-impl.cc
Adjusts Ascend provider selection logic to consider non-empty Paraformer config, enables Paraformer Ascend instantiation, and improves error messages with installation guidance URLs.
Paraformer Utilities
sherpa-onnx/csrc/offline-recognizer-paraformer-impl.h
Convert made non-static; replaced exit(-1) with SHERPA_ONNX_EXIT(-1) in unsupported decoding-paths.
Maintenance / Minor Fixes
sherpa-onnx/csrc/ascend/offline-sense-voice-model-ascend.cc, sherpa-onnx/csrc/ascend/utils.cc, sherpa-onnx/csrc/ascend/offline-recognizer-sense-voice-ascend-impl.h
Header include additions/ordering and a corrected inline comment; no behavioral changes.

Sequence Diagram(s)

sequenceDiagram
    actor User
    participant Factory as OfflineRecognizer (factory)
    participant Impl as OfflineRecognizerParaformerAscendImpl
    participant Model as OfflineParaformerModelAscend
    participant Device as AscendDevice

    User->>Factory: Create(...) 
    Factory->>Impl: instantiate (Ascend path)
    User->>Factory: CreateStream()/DecodeStreams(streams)
    Factory->>Impl: DecodeStreams(streams)

    rect rgb(235,245,255)
      Impl->>Impl: Gather frames
      Impl->>Model: Run(features)
      rect rgb(220,235,255)
        Note over Model,Device: Ascend multi-stage inference
        Model->>Device: Upload features
        Device->>Model: Encoder (encoder.om)
        Device->>Model: Predictor (predictor.om)
        Device->>Model: Decoder (decoder.om)
        Model->>Device: Compute acoustic embedding / collect logits
        Device-->>Model: logits (host)
      end
      Model-->>Impl: logits (vector<float>)
      Impl->>Impl: Argmax → token IDs, post-process
      Impl-->>Factory: Set stream result
    end

    Factory-->>User: GetResult(stream)
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Areas needing extra attention:

  • offline-paraformer-model-ascend.cc — ACL calls, buffer lifecycle, memory copies, thread-safety.
  • offline-recognizer-paraformer-ascend-impl.h — decoding loop, token conversion, post-processing correctness.
  • offline-paraformer-model-config.cc — parsing/validation of comma-separated .om paths and error messages.
  • offline-recognizer-impl.cc — provider-selection logic and error-message accuracy.

Possibly related PRs

Suggested labels

size:XXL

Poem

🐇 I hopped into code at dawn,
Ascend files stitched, a brand-new lawn.
Encoder hums, predictors play,
Decoders dance and logits sway.
Tiny rabbit cheers — models drawn! 🥕

Pre-merge checks and finishing touches

✅ Passed checks (2 passed)
Check name Status Explanation
Title Check ✅ Passed The pull request title "Add C++ runtime for Paraformer on Ascend NPU" is clear, specific, and directly aligned with the changeset. The modifications comprehensively implement this objective: new C++ classes for the OfflineParaformerModelAscend and OfflineRecognizerParaformerAscendImpl are introduced, configuration handling is extended to support Ascend's multi-file model format, the build system is updated to include the new source files, and integration with the offline recognizer is completed. The title accurately captures the primary and most important change from the developer's perspective without being vague or overly broad.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.

📜 Recent review details

Configuration used: CodeRabbit UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 321c92e and 7c4e61c.

📒 Files selected for processing (4)
  • sherpa-onnx/csrc/ascend/offline-paraformer-model-ascend.cc (1 hunks)
  • sherpa-onnx/csrc/ascend/offline-recognizer-paraformer-ascend-impl.h (1 hunks)
  • sherpa-onnx/csrc/ascend/offline-sense-voice-model-ascend.cc (2 hunks)
  • sherpa-onnx/csrc/offline-paraformer-model-config.cc (1 hunks)

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
sherpa-onnx/csrc/offline-recognizer-paraformer-impl.h (1)

25-81: Restore internal linkage (or inline) for Convert

Dropping static left this function with external linkage while still being defined in the header. Any translation unit that includes this header now emits a global definition, so linking succeeds only if exactly one TU includes it. Once another TU (e.g., the new Ascend recognizer) includes this header, the build breaks with multiple-definition errors. Please keep it header-only but mark it inline (or move it to a .cc file with a declaration).

Apply this diff:

-OfflineRecognitionResult Convert(const OfflineParaformerDecoderResult &src,
-                                 const SymbolTable &sym_table) {
+inline OfflineRecognitionResult Convert(
+    const OfflineParaformerDecoderResult &src,
+    const SymbolTable &sym_table) {
🧹 Nitpick comments (1)
sherpa-onnx/csrc/ascend/offline-recognizer-paraformer-ascend-impl.h (1)

17-118: Remove the unused decoder member.

decoder_ (and its include) are never used—the greedy path is implemented inline with std::max_element. Dropping the dead member will keep the class lean and avoid future confusion.

Apply this diff:

-#include "sherpa-onnx/csrc/rknn/offline-ctc-greedy-search-decoder-rknn.h"
...
-  std::unique_ptr<OfflineCtcGreedySearchDecoderRknn> decoder_;
📜 Review details

Configuration used: CodeRabbit UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between d5e5ffb and 321c92e.

📒 Files selected for processing (11)
  • sherpa-onnx/csrc/CMakeLists.txt (1 hunks)
  • sherpa-onnx/csrc/ascend/offline-paraformer-model-ascend.cc (1 hunks)
  • sherpa-onnx/csrc/ascend/offline-paraformer-model-ascend.h (1 hunks)
  • sherpa-onnx/csrc/ascend/offline-recognizer-paraformer-ascend-impl.h (1 hunks)
  • sherpa-onnx/csrc/ascend/offline-recognizer-sense-voice-ascend-impl.h (1 hunks)
  • sherpa-onnx/csrc/ascend/offline-sense-voice-model-ascend.cc (1 hunks)
  • sherpa-onnx/csrc/ascend/utils.cc (1 hunks)
  • sherpa-onnx/csrc/offline-paraformer-model-config.cc (1 hunks)
  • sherpa-onnx/csrc/offline-paraformer-model-config.h (1 hunks)
  • sherpa-onnx/csrc/offline-recognizer-impl.cc (3 hunks)
  • sherpa-onnx/csrc/offline-recognizer-paraformer-impl.h (3 hunks)
🧰 Additional context used
🧠 Learnings (1)
📚 Learning: 2025-08-06T04:23:50.237Z
Learnt from: litongjava
Repo: k2-fsa/sherpa-onnx PR: 2440
File: sherpa-onnx/java-api/src/main/java/com/k2fsa/sherpa/onnx/core/Core.java:4-6
Timestamp: 2025-08-06T04:23:50.237Z
Learning: The sherpa-onnx JNI library files are stored in Hugging Face repository at https://huggingface.co/csukuangfj/sherpa-onnx-libs under versioned directories like jni/1.12.7/, and the actual Windows JNI library filename is "sherpa-onnx-jni.dll" as defined in Core.java constants.

Applied to files:

  • sherpa-onnx/csrc/offline-recognizer-impl.cc
🧬 Code graph analysis (5)
sherpa-onnx/csrc/ascend/offline-paraformer-model-ascend.h (2)
sherpa-onnx/csrc/ascend/offline-recognizer-paraformer-ascend-impl.h (1)
  • sherpa_onnx (20-42)
sherpa-onnx/csrc/ascend/offline-paraformer-model-ascend.cc (15)
  • OfflineParaformerModelAscend (424-426)
  • OfflineParaformerModelAscend (428-428)
  • OfflineParaformerModelAscend (440-441)
  • OfflineParaformerModelAscend (445-446)
  • Run (430-433)
  • Run (430-431)
  • features (130-176)
  • features (130-130)
  • features (181-210)
  • features (181-181)
  • VocabSize (435-437)
  • VocabSize (435-435)
  • Impl (82-98)
  • Impl (82-82)
  • Impl (101-128)
sherpa-onnx/csrc/ascend/offline-recognizer-paraformer-ascend-impl.h (4)
sherpa-onnx/csrc/ascend/offline-paraformer-model-ascend.h (2)
  • sherpa_onnx (12-37)
  • OfflineParaformerModelAscend (14-35)
sherpa-onnx/csrc/offline-recognizer-paraformer-impl.h (4)
  • sherpa_onnx (23-203)
  • Convert (25-81)
  • `` (123-125)
  • DecodeStreams (127-183)
sherpa-onnx/csrc/offline-recognizer-impl.cc (8)
  • OfflineRecognizerImpl (508-559)
  • OfflineRecognizerImpl (562-612)
  • OfflineRecognizerImpl (641-642)
  • OfflineRecognizerImpl (649-650)
  • ApplyInverseTextNormalization (614-625)
  • ApplyInverseTextNormalization (614-615)
  • ApplyHomophoneReplacer (627-634)
  • ApplyHomophoneReplacer (627-628)
sherpa-onnx/csrc/rknn/offline-ctc-greedy-search-decoder-rknn.h (1)
  • OfflineCtcGreedySearchDecoderRknn (14-24)
sherpa-onnx/csrc/offline-recognizer-paraformer-impl.h (3)
sherpa-onnx/csrc/offline-recognizer-transducer-impl.h (1)
  • Convert (33-115)
sherpa-onnx/csrc/offline-recognizer-ctc-impl.h (1)
  • Convert (25-289)
sherpa-onnx/csrc/online-recognizer-paraformer-impl.h (1)
  • Convert (25-80)
sherpa-onnx/csrc/offline-paraformer-model-config.cc (3)
sherpa-onnx/csrc/offline-tts.cc (2)
  • Validate (121-150)
  • Validate (121-121)
sherpa-onnx/csrc/online-model-config.cc (2)
  • Validate (55-162)
  • Validate (55-55)
sherpa-onnx/csrc/file-utils.cc (2)
  • FileExists (16-18)
  • FileExists (16-16)
sherpa-onnx/csrc/ascend/offline-paraformer-model-ascend.cc (2)
sherpa-onnx/csrc/ascend/offline-sense-voice-model-ascend.cc (10)
  • out (180-180)
  • features (53-108)
  • features (53-54)
  • data (119-125)
  • data (119-119)
  • max_num_frames_ (145-155)
  • in (157-193)
  • in (157-157)
  • Run (220-223)
  • Run (220-221)
sherpa-onnx/csrc/ascend/offline-paraformer-model-ascend.h (1)
  • OfflineParaformerModelAscend (14-35)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (19)
  • GitHub Check: ubuntu-latest Debug static tts-ON
  • GitHub Check: ubuntu-latest Debug static tts-OFF
  • GitHub Check: ubuntu-latest Debug shared tts-OFF
  • GitHub Check: ubuntu-latest Release static tts-ON
  • GitHub Check: ubuntu-latest Release shared tts-ON
  • GitHub Check: Release shared-OFF tts-ON
  • GitHub Check: Release shared-OFF tts-OFF
  • GitHub Check: swift (macos-latest)
  • GitHub Check: Debug shared-ON tts-ON
  • GitHub Check: Debug shared-OFF tts-ON
  • GitHub Check: Debug shared-OFF tts-OFF
  • GitHub Check: Debug shared-ON tts-OFF
  • GitHub Check: swift (macos-13)
  • GitHub Check: Release shared-ON tts-ON
  • GitHub Check: Release shared-ON tts-OFF
  • GitHub Check: Debug shared tts-ON
  • GitHub Check: Release static tts-ON
  • GitHub Check: rknn shared OFF
  • GitHub Check: rknn shared ON
🔇 Additional comments (1)
sherpa-onnx/csrc/ascend/offline-paraformer-model-ascend.h (1)

14-35: Interface mirrors existing backends.

The PIMPL layout and minimal public surface match the other offline Paraformer model headers. Looks good.

Comment on lines +362 to +395
std::vector<float> ApplyLFR(std::vector<float> in) const {
int32_t lfr_window_size = 7;
int32_t lfr_window_shift = 6;
int32_t in_feat_dim = 80;

int32_t in_num_frames = in.size() / in_feat_dim;
int32_t out_num_frames =
(in_num_frames - lfr_window_size) / lfr_window_shift + 1;

if (out_num_frames > max_num_frames_) {
SHERPA_ONNX_LOGE(
"Number of input frames %d is too large. Truncate it to %d frames.",
out_num_frames, max_num_frames_);

SHERPA_ONNX_LOGE(
"Recognition result may be truncated/incomplete. Please select a "
"model accepting longer audios.");

out_num_frames = max_num_frames_;
}

int32_t out_feat_dim = in_feat_dim * lfr_window_size;

std::vector<float> out(out_num_frames * out_feat_dim);

const float *p_in = in.data();
float *p_out = out.data();

for (int32_t i = 0; i != out_num_frames; ++i) {
std::copy(p_in, p_in + out_feat_dim, p_out);

p_out += out_feat_dim;
p_in += lfr_window_shift * in_feat_dim;
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🔴 Critical

Guard LFR for short utterances to avoid OOB reads.

When the input has fewer than lfr_window_size frames, (in_num_frames - lfr_window_size) / lfr_window_shift + 1 still yields 1 because of integer truncation. The loop then copies 560 floats from a buffer that may contain fewer than 560 entries, causing an out-of-bounds read.

Apply this fix:

     int32_t in_num_frames = in.size() / in_feat_dim;
+    if (in_num_frames < lfr_window_size) {
+      return {};
+    }
     int32_t out_num_frames =
         (in_num_frames - lfr_window_size) / lfr_window_shift + 1;
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
std::vector<float> ApplyLFR(std::vector<float> in) const {
int32_t lfr_window_size = 7;
int32_t lfr_window_shift = 6;
int32_t in_feat_dim = 80;
int32_t in_num_frames = in.size() / in_feat_dim;
int32_t out_num_frames =
(in_num_frames - lfr_window_size) / lfr_window_shift + 1;
if (out_num_frames > max_num_frames_) {
SHERPA_ONNX_LOGE(
"Number of input frames %d is too large. Truncate it to %d frames.",
out_num_frames, max_num_frames_);
SHERPA_ONNX_LOGE(
"Recognition result may be truncated/incomplete. Please select a "
"model accepting longer audios.");
out_num_frames = max_num_frames_;
}
int32_t out_feat_dim = in_feat_dim * lfr_window_size;
std::vector<float> out(out_num_frames * out_feat_dim);
const float *p_in = in.data();
float *p_out = out.data();
for (int32_t i = 0; i != out_num_frames; ++i) {
std::copy(p_in, p_in + out_feat_dim, p_out);
p_out += out_feat_dim;
p_in += lfr_window_shift * in_feat_dim;
}
std::vector<float> ApplyLFR(std::vector<float> in) const {
int32_t lfr_window_size = 7;
int32_t lfr_window_shift = 6;
int32_t in_feat_dim = 80;
int32_t in_num_frames = in.size() / in_feat_dim;
if (in_num_frames < lfr_window_size) {
return {};
}
int32_t out_num_frames =
(in_num_frames - lfr_window_size) / lfr_window_shift + 1;
if (out_num_frames > max_num_frames_) {
SHERPA_ONNX_LOGE(
"Number of input frames %d is too large. Truncate it to %d frames.",
out_num_frames, max_num_frames_);
SHERPA_ONNX_LOGE(
"Recognition result may be truncated/incomplete. Please select a "
"model accepting longer audios.");
out_num_frames = max_num_frames_;
}
int32_t out_feat_dim = in_feat_dim * lfr_window_size;
std::vector<float> out(out_num_frames * out_feat_dim);
const float *p_in = in.data();
float *p_out = out.data();
for (int32_t i = 0; i != out_num_frames; ++i) {
std::copy(p_in, p_in + out_feat_dim, p_out);
p_out += out_feat_dim;
p_in += lfr_window_shift * in_feat_dim;
}
🤖 Prompt for AI Agents
In sherpa-onnx/csrc/ascend/offline-paraformer-model-ascend.cc around lines 362
to 395, the LFR logic can read out-of-bounds when in_num_frames <
lfr_window_size because the out_num_frames formula yields 1; change the
computation to guard short utterances by checking if in_num_frames <
lfr_window_size and if so set out_num_frames = 0 (or return an empty vector)
before allocating/copying, otherwise compute out_num_frames = (in_num_frames -
lfr_window_size) / lfr_window_shift + 1 and proceed; ensure you also clamp
out_num_frames to max_num_frames_ and skip the copy-loop when out_num_frames ==
0 to avoid OOB reads.

Comment thread sherpa-onnx/csrc/offline-paraformer-model-config.cc Outdated

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This pull request adds support for Paraformer models on Ascend NPU for offline speech recognition. The changes extend the existing Ascend NPU support (which previously only supported SenseVoice models) to also handle Paraformer models.

Key changes:

  • Added Ascend NPU implementation for Paraformer model support
  • Made the Convert function non-static to allow external usage
  • Replaced exit(-1) calls with SHERPA_ONNX_EXIT(-1) for better error handling
  • Enhanced error messages with documentation links

Reviewed Changes

Copilot reviewed 11 out of 11 changed files in this pull request and generated 3 comments.

Show a summary per file
File Description
sherpa-onnx/csrc/offline-recognizer-paraformer-impl.h Removed static keyword from Convert function to enable external usage
sherpa-onnx/csrc/offline-recognizer-impl.cc Added Paraformer model support for Ascend NPU, improved error messages with documentation links
sherpa-onnx/csrc/offline-paraformer-model-config.h Added documentation comment for Ascend NPU model path format
sherpa-onnx/csrc/offline-paraformer-model-config.cc Added validation logic for Ascend NPU model files (.om format)
sherpa-onnx/csrc/ascend/utils.cc Added missing <utility> include
sherpa-onnx/csrc/ascend/offline-sense-voice-model-ascend.cc Reordered includes alphabetically
sherpa-onnx/csrc/ascend/offline-recognizer-sense-voice-ascend-impl.h Fixed comment reference from "online" to "offline"
sherpa-onnx/csrc/ascend/offline-recognizer-paraformer-ascend-impl.h New implementation file for Paraformer Ascend NPU recognizer
sherpa-onnx/csrc/ascend/offline-paraformer-model-ascend.h New header for Paraformer Ascend NPU model
sherpa-onnx/csrc/ascend/offline-paraformer-model-ascend.cc New implementation of Paraformer model for Ascend NPU
sherpa-onnx/csrc/CMakeLists.txt Added new Paraformer Ascend source file to build

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

if (EndsWith(model, ".onnx") && !FileExists(model)) {
SHERPA_ONNX_LOGE("Paraformer model '%s' does not exist", model.c_str());
return false;
} else if (EndsWith(model, ".om")) {

Copilot AI Nov 3, 2025

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The pattern \"*.om\" is treated as a literal string, not a wildcard pattern. The asterisk will be matched as a literal character. This should likely be \".om\" to check if the model path ends with the .om extension.

Copilot uses AI. Check for mistakes.

namespace sherpa_onnx {

// defined in ../online-recognizer-paraformer-impl.h

Copilot AI Nov 3, 2025

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The comment incorrectly references 'online-recognizer-paraformer-impl.h' but should reference 'offline-recognizer-paraformer-impl.h' since this is for offline recognition, not online. This was correctly fixed for SenseVoice in line 22 of offline-recognizer-sense-voice-ascend-impl.h.

Suggested change
// defined in ../online-recognizer-paraformer-impl.h
// defined in ../offline-recognizer-paraformer-impl.h

Copilot uses AI. Check for mistakes.
OfflineRecognizerConfig config_;
SymbolTable symbol_table_;
std::unique_ptr<OfflineParaformerModelAscend> model_;
std::unique_ptr<OfflineCtcGreedySearchDecoderRknn> decoder_;

Copilot AI Nov 3, 2025

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The member variable decoder_ is declared but never initialized or used in the class implementation. This appears to be dead code that should be removed.

Suggested change
std::unique_ptr<OfflineCtcGreedySearchDecoderRknn> decoder_;

Copilot uses AI. Check for mistakes.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XL This PR changes 500-999 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants