Skip to content

Add C++ runtime for Parakeet CTC with QNN. - #3688

Merged
csukuangfj merged 2 commits into
k2-fsa:masterfrom
csukuangfj:cpp-qnn-parakeet-ctc
Jun 18, 2026
Merged

csukuangfj merged 2 commits into
k2-fsa:masterfrom
csukuangfj:cpp-qnn-parakeet-ctc

Conversation

@csukuangfj

@csukuangfj csukuangfj commented Jun 18, 2026 •

Copy link
Copy Markdown
Collaborator

See https://k2-fsa.github.io/sherpa/onnx/qnn/run-executables-on-your-phone-binary.html
for how to run it.

Usage:

./sherpa-onnx-offline \
  --provider=qnn \
  --nemo-ctc-model=./libmodel.so \
  --nemo-ctc.qnn-backend-lib=./libQnnHtp.so \
  --nemo-ctc.qnn-context-binary=./model.bin \
  --nemo-ctc.qnn-system-lib=./libQnnSystem.so \
  --debug=1 \
  --tokens=./tokens.txt \
  ./0.wav

Logs:

/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/parse-options.cc:Read:373 ./sherpa-onnx-offline --provider=qnn --nemo-ctc-model=./libmodel.so --nemo-ctc.qnn-backend-lib=.
/libQnnHtp.so --nemo-ctc.qnn-context-binary=./model.bin --nemo-ctc.qnn-system-lib=./libQnnSystem.so --debug=1 --tokens=./tokens.txt ./0.wav

OfflineRecognizerConfig(feat_config=FeatureExtractorConfig(sampling_rate=16000, feature_dim=80, low_freq=20, high_freq=-400, dither=0, normalize_samples=True, snip_edges=False),
model_config=OfflineModelConfig(transducer=OfflineTransducerModelConfig(encoder_filename="", decoder_filename="", joiner_filename=""), paraformer=OfflineParaformerModelConfig(mod
el=""), nemo_ctc=OfflineNemoEncDecCtcModelConfig(model="./libmodel.so", qnn_config=QnnConfig(backend_lib="./libQnnHtp.so", context_binary="./model.bin", system_lib="./libQnnSyste
m.so"), ), whisper=OfflineWhisperModelConfig(encoder="", decoder="", language="", task="transcribe", tail_paddings=-1, enable_token_timestamps=False, enable_segment_timestamps=Fa
lse), fire_red_asr=OfflineFireRedAsrModelConfig(encoder="", decoder=""), tdnn=OfflineTdnnModelConfig(model=""), zipformer_ctc=OfflineZipformerCtcModelConfig(model=""), wenet_ctc=
OfflineWenetCtcModelConfig(model=""), sense_voice=OfflineSenseVoiceModelConfig(model="", language="auto", use_itn=False), moonshine=OfflineMoonshineModelConfig(preprocessor="", e
ncoder="", uncached_decoder="", cached_decoder="", merged_decoder=""), dolphin=OfflineDolphinModelConfig(model=""), canary=OfflineCanaryModelConfig(encoder="", decoder="", src_la
ng="", tgt_lang="", use_pnc=True), cohere_transcribe=OfflineCohereTranscribeModelConfig(encoder="", decoder="", language="", use_punct=True, use_itn=True), omnilingual=OfflineOmn
ilingualAsrCtcModelConfig(model=""), funasr_nano=OfflineFunASRNanoModelConfig(encoder_adaptor="", llm="", embedding="", tokenizer="", system_prompt="You are a helpful assistant."
, user_prompt="语音转写:", max_new_tokens=512, temperature=1e-06, top_p=0.8, seed=42, language="", itn=True, hotwords=""), medasr=OfflineMedAsrCtcModelConfig(model=""), fire_red
_asr_ctc=OfflineFireRedAsrCtcModelConfig(model=""), qwen3_asr=OfflineQwen3ASRModelConfig(conv_frontend="", encoder="", decoder="", tokenizer="", hotwords="", max_total_len=512, m
ax_new_tokens=128, temperature=1e-06, top_p=0.8, seed=42), telespeech_ctc="", tokens="./tokens.txt", num_threads=2, debug=True, provider="qnn", model_type="", modeling_unit="cjkc
har", bpe_vocab=""), lm_config=OfflineLMConfig(model="", scale=0.5, lodr_scale=0.01, lodr_fst="", lodr_backoff_id=-1), ctc_fst_decoder_config=OfflineCtcFstDecoderConfig(graph="",
 max_active=3000), decoding_method="greedy_search", max_active_paths=4, hotwords_file="", hotwords_score=1.5, blank_penalty=0, rule_fsts="", rule_fars="", hr=HomophoneReplacerCon
fig(lexicon="", rule_fsts=""))
Creating recognizer ...
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/qnn-backend.cc:InitQnnInterface:100 loaded ./libQnnHtp.so
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/qnn-backend.cc:InitQnnInterface:114 Got QnnInterface_getProviders
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/qnn-backend.cc:InitQnnInterface:136 QNN_API_VERSION_MAJOR: 2
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/qnn-backend.cc:InitQnnInterface:137 QNN_API_VERSION_MINOR: 29
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/qnn-backend.cc:InitQnnInterface:138 QNN_API_VERSION_PATCH: 0
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/qnn-backend.cc:InitQnnInterface:161 ---0----
backendId: 6
coreApiVersion.major: 2
coreApiVersion.minor: 30
coreApiVersion.patch: 0
backendApiVersion.major: 5
backendApiVersion.minor: 40
backendApiVersion.patch: 0

/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/qnn-backend.cc:InitQnnInterface:179 backend build ID: v2.40.0.251030114326_189385
     0.0ms [WARN   ] QnnDsp <W> Initializing HtpProvider
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/offline-parakeet-ctc-model-qnn.cc:Impl:66 Init from context binary './model.bin'
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/qnn-model.cc:LoadSystemLib:65 loaded ./libQnnSystem.so
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/qnn-model.cc:LoadSystemLib:101 QNN_SYSTEM_API_VERSION_MAJOR: 1
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/qnn-model.cc:LoadSystemLib:103 QNN_SYSTEM_API_VERSION_MINOR: 5
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/qnn-model.cc:LoadSystemLib:107 systemApiVersion.major: 1
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/qnn-model.cc:LoadSystemLib:111 systemApiVersion.minor: 6
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/utils.cc:CopyGraphsInfo:481 version: 3
     0.0ms [WARN   ] QnnDsp <W> m_CFBCallbackInfoObj is not initialized, return emptyList
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/qnn-model.cc:InitInputTensors:570 input 0
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/utils.cc:PrintTensor:426   id: 1
  name: x
  type: QNN_TENSOR_TYPE_APP_WRITE
  data format: 0
  data type: QNN_DATATYPE_UFIXED_POINT_16
  quantize info:
    encodingDefinition: 0x1
    quantizationEncoding: QNN_QUANTIZATION_ENCODING_SCALE_OFFSET
     scale: 0.000124647
     offset: -31906
  rank: 3
  dimensions: 1, 1000, 80,
  memType: QNN_TENSORMEMTYPE_RAW
 memType raw data size: 160000
  isDynamicDimensions: False
  isProduced: 0

/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/qnn-model.cc:InitOutputTensors:591 output 2106
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/utils.cc:PrintTensor:426   id: 3757
  name: transpose_292
  type: QNN_TENSOR_TYPE_APP_READ
  data format: 0
  data type: QNN_DATATYPE_UFIXED_POINT_16
  quantize info:
    encodingDefinition: 0x1
    quantizationEncoding: QNN_QUANTIZATION_ENCODING_SCALE_OFFSET
     scale: 0.00219824
     offset: -3019
  rank: 3
  dimensions: 1, 125, 1025,
  memType: QNN_TENSORMEMTYPE_RAW
 memType raw data size: 256250
  isDynamicDimensions: False
  isProduced: 0

/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/qnn-model.cc:InitOutputTensors:591 output 2107
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/utils.cc:PrintTensor:426   id: 3758
  name: log_probs
  type: QNN_TENSOR_TYPE_APP_READ
  data format: 0
  data type: QNN_DATATYPE_UFIXED_POINT_16
  quantize info:
    encodingDefinition: 0x1
    quantizationEncoding: QNN_QUANTIZATION_ENCODING_SCALE_OFFSET
     scale: 0.00219739
     offset: -65535
  rank: 3
  dimensions: 1, 125, 1025,
  memType: QNN_TENSORMEMTYPE_RAW
 memType raw data size: 256250
  isDynamicDimensions: False
  isProduced: 0


/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/qnn-model.cc:AllocateBuffer:612 Allocate 817009250 bytes, or 779.161 MB
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/qnn-model.cc:SetupPointers:633 Setup pointers successfully.
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/offline-parakeet-ctc-model-qnn.cc:CheckModel:203 max_num_frames: 1000
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/offline-parakeet-ctc-model-qnn.cc:CheckModel:204 feat_dim: 80
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/offline-parakeet-ctc-model-qnn.cc:CheckModel:205 vocab_size: 1025
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/offline-parakeet-ctc-model-qnn.cc:CheckModel:206 subsampling_factor: 8
recognizer created in 4.852 s
Started
Done!

./0.wav
{"lang": "", "emotion": "", "event": "", "text": "well i don't wish to see it any more observed phoebe turning away her eyes it is certainly very like the old portrait it", "time
stamps": [0.48, 0.72, 0.88, 0.96, 1.04, 1.12, 1.20, 1.36, 1.44, 1.60, 1.76, 1.92, 2.16, 2.32, 2.40, 2.48, 2.64, 2.72, 2.88, 2.96, 3.36, 3.44, 3.60, 3.76, 4.00, 4.16, 4.24, 4.32,
5.04, 5.20, 5.44, 5.68, 5.84, 6.08, 6.32, 6.40, 6.64, 6.72, 6.80, 6.96, 7.60], "durations": [], "tokens":[" well", " i", " don", "'", "t", " w", "ish", " to", " see", " it", " an
y", " more", " ob", "s", "er", "ved", " ph", "o", "e", "be", " t", "ur", "ning", " away", " her", " e", "y", "es", " it", " is", " certain", "ly", " very", " like", " the", " old
", " p", "ort", "ra", "it", " it"], "ys_log_probs": [], "words": []}
----
num threads: 2
decoding method: greedy_search
Elapsed seconds: 0.518 s
Real time factor (RTF): 0.518 / 7.435 = 0.070
     0.0ms [WARN   ] QnnDsp <W> m_CFBCallbackInfoObj is not initialized, return emptyList

Summary by CodeRabbit

Release Notes

  • New Features

    • Added QNN-accelerated support for NeMo CTC (Parakeet) models in offline recognition.
  • Improvements

    • Enhanced NeMo CTC model configuration to better handle QNN context binaries, including clearer validation and more informative config output.
    • Improved decoding robustness for per-frame selection during greedy CTC decoding.

@dosubot dosubot Bot added the size:XL This PR changes 500-999 lines, ignoring generated files. label Jun 18, 2026
@coderabbitai

coderabbitai Bot commented Jun 18, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

Adds QNN backend support for NeMo CTC (Parakeet) offline speech recognition. A new OfflineParakeetCtcModelQnn class handles model loading and inference via QNN, OfflineRecognizerParakeetCtcQnnImpl wires feature extraction and greedy CTC decoding, OfflineNemoEncDecCtcModelConfig gains a QnnConfig field, and the recognizer factory is updated to route to the new implementation. The RKNN CTC greedy decoder's argmax is also refactored to use MaxElementIndex.

Changes

QNN Parakeet CTC Offline Recognizer

Layer / File(s) Summary
NeMo CTC config extended with QnnConfig
sherpa-onnx/csrc/offline-nemo-enc-dec-ctc-model-config.h, sherpa-onnx/csrc/offline-nemo-enc-dec-ctc-model-config.cc
Adds QnnConfig qnn_config member to OfflineNemoEncDecCtcModelConfig; wires qnn_config.Register() under a prefixed ParseOptions; reworks Validate() to branch on model/context_binary/.so/.bin combinations; conditionally serializes qnn_config in ToString().
OfflineParakeetCtcModelQnn class (Pimpl)
sherpa-onnx/csrc/qnn/offline-parakeet-ctc-model-qnn.h, sherpa-onnx/csrc/qnn/offline-parakeet-ctc-model-qnn.cc
Declares OfflineParakeetCtcModelQnn with Pimpl pattern; implements Impl construction via QnnBackend with model-lib vs context-binary init paths, PostInit/CheckModel for tensor validation, Run with truncation/padding, dimension accessors, and Android/OHOS template instantiations.
RKNN CTC greedy decoder argmax refactor
sherpa-onnx/csrc/rknn/offline-ctc-greedy-search-decoder-rknn.cc
Adds math.h include and replaces std::max_element + std::distance argmax with MaxElementIndex() from math.h; explicitly types the result as int64_t.
OfflineRecognizerParakeetCtcQnnImpl
sherpa-onnx/csrc/qnn/offline-recognizer-parakeet-ctc-qnn-impl.h
Declares Convert() helper and defines OfflineRecognizerParakeetCtcQnnImpl: Init() configures librosa-compatible fbank and resolves blank token ID; DecodeStreams normalizes features, runs QNN inference, applies greedy CTC decoding, ITN, and homophone replacement, and writes OfflineRecognitionResult into each stream.
Factory wiring and CMake registration
sherpa-onnx/csrc/CMakeLists.txt, sherpa-onnx/csrc/offline-recognizer-impl.cc
Appends offline-parakeet-ctc-model-qnn.cc to QNN CMake sources; extends both OfflineRecognizerImpl::Create overloads to detect nemo_ctc + non-empty context_binary and construct OfflineRecognizerParakeetCtcQnnImpl, with updated error messages.

Sequence Diagram(s)

sequenceDiagram
    participant App
    participant OfflineRecognizerImpl
    participant OfflineRecognizerParakeetCtcQnnImpl
    participant OfflineParakeetCtcModelQnn
    participant OfflineCtcGreedySearchDecoderRknn

    App->>OfflineRecognizerImpl: Create(config, provider="qnn")
    OfflineRecognizerImpl->>OfflineRecognizerParakeetCtcQnnImpl: construct(config)
    OfflineRecognizerParakeetCtcQnnImpl->>OfflineParakeetCtcModelQnn: construct(model_config)
    OfflineParakeetCtcModelQnn-->>OfflineRecognizerParakeetCtcQnnImpl: model ready

    App->>OfflineRecognizerParakeetCtcQnnImpl: DecodeStreams(streams)
    OfflineRecognizerParakeetCtcQnnImpl->>OfflineRecognizerParakeetCtcQnnImpl: normalize features
    OfflineRecognizerParakeetCtcQnnImpl->>OfflineParakeetCtcModelQnn: Run(features) → log_probs
    OfflineParakeetCtcModelQnn-->>OfflineRecognizerParakeetCtcQnnImpl: log_probs
    OfflineRecognizerParakeetCtcQnnImpl->>OfflineCtcGreedySearchDecoderRknn: Decode(log_probs) → CtcResult
    OfflineCtcGreedySearchDecoderRknn-->>OfflineRecognizerParakeetCtcQnnImpl: CtcResult
    OfflineRecognizerParakeetCtcQnnImpl->>OfflineRecognizerParakeetCtcQnnImpl: Convert + ITN + HomophoneReplacer
    OfflineRecognizerParakeetCtcQnnImpl-->>App: OfflineRecognitionResult written to stream
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~50 minutes

Possibly related PRs

  • k2-fsa/sherpa-onnx#2766: Both PRs update sherpa-onnx/csrc/CMakeLists.txt under SHERPA_ONNX_ENABLE_QNN to append new QNN source files to the core library build.
  • k2-fsa/sherpa-onnx#3655: Both PRs extend the QNN provider branch in offline-recognizer-impl.cc to add new model-type construction paths in the same factory selection code.
  • k2-fsa/sherpa-onnx#2931: Both PRs modify offline-recognizer-impl.cc to add new QNN recognizer branches and update the unsupported-model error messages.

🐇 A new Parakeet took flight today,
Through QNN's gate it found its way.
Blank tokens checked, log probs returned,
From context binary, features learned.
The bunny cheers with twitching nose—
One more CTC decoder goes! 🎉

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title directly and clearly summarizes the main change: adding C++ runtime support for Parakeet CTC with QNN backend, which is the primary objective across all modified files.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces support for Qualcomm Neural Network (QNN) offline Parakeet CTC models, including the implementation of the QNN model wrapper, the offline recognizer implementation, and integration into the existing offline recognizer factory. Feedback focuses on resolving a member variable shadowing issue in OfflineRecognizerParakeetCtcQnnImpl that could lead to out-of-sync configurations, and adding input/output tensor shape validation in OfflineParakeetCtcModelQnn to prevent potential division-by-zero errors.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment on lines +34 to +53
explicit OfflineRecognizerParakeetCtcQnnImpl(
const OfflineRecognizerConfig &config)
: OfflineRecognizerImpl(config),
config_(config),
symbol_table_(config_.model_config.tokens),
model_(std::make_unique<OfflineParakeetCtcModelQnn>(
config.model_config)) {
Init();
}

template <typename Manager>
OfflineRecognizerParakeetCtcQnnImpl(Manager *mgr,
const OfflineRecognizerConfig &config)
: OfflineRecognizerImpl(mgr, config),
config_(config),
symbol_table_(mgr, config_.model_config.tokens),
model_(std::make_unique<OfflineParakeetCtcModelQnn>(
mgr, config.model_config)) {
Init();
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The member variable config_ shadows the base class member OfflineRecognizerImpl::config_. This shadowing causes OfflineRecognizerImpl::SetConfig to only update the base class's config_, leaving the derived class's config_ out of sync and breaking functionality when SetConfig is called.

Please remove config_(config) from the initializer list and use the passed-in config parameter directly to initialize other members.

  explicit OfflineRecognizerParakeetCtcQnnImpl(
      const OfflineRecognizerConfig &config)
      : OfflineRecognizerImpl(config),
        symbol_table_(config.model_config.tokens),
        model_(std::make_unique<OfflineParakeetCtcModelQnn>(
            config.model_config)) {
    Init();
  }

  template <typename Manager>
  OfflineRecognizerParakeetCtcQnnImpl(Manager *mgr,
                                      const OfflineRecognizerConfig &config)
      : OfflineRecognizerImpl(mgr, config),
        symbol_table_(mgr, config.model_config.tokens),
        model_(std::make_unique<OfflineParakeetCtcModelQnn>(
            mgr, config.model_config)) {
    Init();
  }

Comment on lines +134 to +135
OfflineRecognizerConfig config_;
SymbolTable symbol_table_;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Remove the shadowed config_ member variable to ensure that the base class's config_ is used consistently throughout the class.

Suggested change
OfflineRecognizerConfig config_;
SymbolTable symbol_table_;
SymbolTable symbol_table_;

Comment on lines +190 to +191
max_num_frames_ = x_shape[1];
feat_dim_ = x_shape[2];

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

We should validate that max_num_frames_ and feat_dim_ are positive integers to prevent potential division-by-zero or other undefined behaviors in Run().

    max_num_frames_ = x_shape[1];
    feat_dim_ = x_shape[2];

    if (max_num_frames_ <= 0 || feat_dim_ <= 0) {
      SHERPA_ONNX_LOGE("Invalid model input shape: max_num_frames=%d, feat_dim=%d",
                       max_num_frames_, feat_dim_);
      SHERPA_ONNX_EXIT(-1);
    }

Comment on lines +198 to +201
auto out_shape = model_->TensorShape("log_probs");
vocab_size_ = out_shape[2];

subsampling_factor_ = max_num_frames_ / out_shape[1];

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

We should validate the output tensor shape of log_probs to ensure it is 3-dimensional and has positive dimensions. This prevents potential division-by-zero when calculating subsampling_factor_ and ensures vocab_size_ is valid.

    auto out_shape = model_->TensorShape("log_probs");
    if (out_shape.size() != 3 || out_shape[1] <= 0 || out_shape[2] <= 0) {
      SHERPA_ONNX_LOGE("Invalid output shape for log_probs");
      SHERPA_ONNX_EXIT(-1);
    }
    vocab_size_ = out_shape[2];

    subsampling_factor_ = max_num_frames_ / out_shape[1];

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@sherpa-onnx/csrc/offline-nemo-enc-dec-ctc-model-config.cc`:
- Around line 16-17: Update the help text in the po->Register call for
"nemo-ctc-model" in offline-nemo-enc-dec-ctc-model-config.cc to reflect that the
option now accepts compiled artifacts (such as libmodel.so) for QNN usage, not
just model.onnx files. Replace the current description that specifically
mentions "model.onnx" with a more generic description that accurately describes
the supported formats and clarifies that it works with both ONNX and compiled
artifacts in the QNN code path.

In `@sherpa-onnx/csrc/qnn/offline-parakeet-ctc-model-qnn.cc`:
- Around line 198-201: The code accesses out_shape[2] and out_shape[1] without
validating that the tensor has sufficient rank or that the values are positive,
which can cause out-of-bounds access or division by zero errors. Add validation
checks after calling model_->TensorShape("log_probs") to ensure out_shape has at
least 3 elements and that out_shape[1] is greater than zero before using it in
the subsampling_factor_ calculation. If validation fails, log an appropriate
error and return early.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: b644fc7a-9a21-4b87-97f6-faa4c16886fb

📥 Commits

Reviewing files that changed from the base of the PR and between 9961aee and c85b37e.

📒 Files selected for processing (8)
  • sherpa-onnx/csrc/CMakeLists.txt
  • sherpa-onnx/csrc/offline-nemo-enc-dec-ctc-model-config.cc
  • sherpa-onnx/csrc/offline-nemo-enc-dec-ctc-model-config.h
  • sherpa-onnx/csrc/offline-recognizer-impl.cc
  • sherpa-onnx/csrc/qnn/offline-parakeet-ctc-model-qnn.cc
  • sherpa-onnx/csrc/qnn/offline-parakeet-ctc-model-qnn.h
  • sherpa-onnx/csrc/qnn/offline-recognizer-parakeet-ctc-qnn-impl.h
  • sherpa-onnx/csrc/rknn/offline-ctc-greedy-search-decoder-rknn.cc

Comment thread sherpa-onnx/csrc/offline-nemo-enc-dec-ctc-model-config.cc Outdated
Comment thread sherpa-onnx/csrc/qnn/offline-parakeet-ctc-model-qnn.cc

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
sherpa-onnx/csrc/qnn/offline-parakeet-ctc-model-qnn.cc (1)

82-99: 💤 Low value

Consider validating that feature vector size is an exact multiple of feat_dim_.

Integer division at line 83 silently truncates if features.size() % feat_dim_ != 0. The subsequent resize() preserves orphan floats, causing frame misalignment. While callers should pass properly aligned data, a defensive check would catch upstream bugs early.

💡 Optional defensive check
 std::vector<float> Run(std::vector<float> features) {
+  if (features.size() % feat_dim_ != 0) {
+    SHERPA_ONNX_LOGE(
+        "Feature vector size %d is not a multiple of feat_dim %d",
+        static_cast<int32_t>(features.size()), feat_dim_);
+  }
   int32_t num_frames = features.size() / feat_dim_;
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@sherpa-onnx/csrc/qnn/offline-parakeet-ctc-model-qnn.cc` around lines 82 - 99,
The Run method performs integer division of features.size() by feat_dim_ without
validating that the feature vector size is an exact multiple of feat_dim_, which
can cause silent truncation and frame misalignment. Add a defensive check
immediately after entering the Run method to validate that features.size()
modulo feat_dim_ equals zero, and log an error or throw an exception if this
validation fails to catch upstream bugs early.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In `@sherpa-onnx/csrc/qnn/offline-parakeet-ctc-model-qnn.cc`:
- Around line 82-99: The Run method performs integer division of features.size()
by feat_dim_ without validating that the feature vector size is an exact
multiple of feat_dim_, which can cause silent truncation and frame misalignment.
Add a defensive check immediately after entering the Run method to validate that
features.size() modulo feat_dim_ equals zero, and log an error or throw an
exception if this validation fails to catch upstream bugs early.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: c7bd9f7d-4a8d-4fdf-8718-9c9dd5c0bace

📥 Commits

Reviewing files that changed from the base of the PR and between c85b37e and 07f49e7.

📒 Files selected for processing (3)
  • sherpa-onnx/csrc/offline-nemo-enc-dec-ctc-model-config.cc
  • sherpa-onnx/csrc/qnn/offline-parakeet-ctc-model-qnn.cc
  • sherpa-onnx/csrc/qnn/offline-recognizer-parakeet-ctc-qnn-impl.h
🚧 Files skipped from review as they are similar to previous changes (2)
  • sherpa-onnx/csrc/qnn/offline-recognizer-parakeet-ctc-qnn-impl.h
  • sherpa-onnx/csrc/offline-nemo-enc-dec-ctc-model-config.cc

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XL This PR changes 500-999 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant