Skip to content

Add C++ runtime for Whisper with Qualcomm NPU using QNN. - #3699

Merged
csukuangfj merged 2 commits into
k2-fsa:masterfrom
csukuangfj:cpp-qnn-whisper
Jun 24, 2026
Merged

csukuangfj merged 2 commits into
k2-fsa:masterfrom
csukuangfj:cpp-qnn-whisper

Conversation

@csukuangfj

@csukuangfj csukuangfj commented Jun 24, 2026 •

Copy link
Copy Markdown
Collaborator

Usage

See https://k2-fsa.github.io/sherpa/onnx/qnn/run-executables-on-your-phone-binary.html

wget https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models-qnn-binary-2/sherpa-onnx-qnn-SM8850-binary-whisper-tiny.tar.bz2
tar xvf sherpa-onnx-qnn-SM8850-binary-whisper-tiny.tar.bz2
rm sherpa-onnx-qnn-SM8850-binary-whisper-tiny.tar.bz2
ls -lh sherpa-onnx-qnn-SM8850-binary-whisper-tiny
total 239680
-rw-r--r--@ 1 fangjun  staff    93M 24 Jun 12:31 decoder.bin
-rw-r--r--@ 1 fangjun  staff    20M 24 Jun 12:31 encoder.bin
drwxr-xr-x@ 5 fangjun  staff   160B 24 Jun 12:31 test_wavs
-rw-r--r--@ 1 fangjun  staff   798K 24 Jun 12:31 tokens.txt

now run it on your phone:

./sherpa-onnx-offline \
  --debug=1 \
  --whisper.qnn-backend-lib=./libQnnHtp.so \
  --whisper.qnn-context-binary=./sherpa-onnx-qnn-SM8850-binary-whisper-tiny/encoder.bin,./sherpa-onnx-qnn-SM8850-binary-whisper-tiny/decoder.bin \
  --whisper.qnn-system-lib=./libQnnSystem.so \
  --tokens=./sherpa-onnx-qnn-SM8850-binary-whisper-tiny/tokens.txt \
  --provider=qnn \
  ./sherpa-onnx-qnn-SM8850-binary-whisper-tiny/test_wavs/1.wav

Output logs:

/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/qnn-model.cc:AllocateBuffer:612 Allocate 24158572 bytes, or 23.039 MB
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/qnn-model.cc:SetupPointers:633 Setup pointers successfully.
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/offline-whisper-model-qnn.cc:PostInitEncoder:396 model type: tiny, prefix: tiny
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/offline-whisper-model-qnn.cc:PostInitEncoder:437 feat_dim_: 80
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/offline-whisper-model-qnn.cc:PostInitEncoder:438 num_frames_: 3000
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/offline-whisper-model-qnn.cc:PostInitEncoder:439 num_out_frames_: 1500
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/offline-whisper-model-qnn.cc:PostInitEncoder:440 n_text_layer_: 4
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/offline-whisper-model-qnn.cc:PostInitEncoder:441 n_text_state_: 384
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/offline-whisper-model-qnn.cc:PostInitDecoder:487 mask_size_: 448
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/offline-whisper-model-qnn.cc:PostInitDecoder:488 n_text_state_: 384
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/offline-whisper-model-qnn.cc:PostInitDecoder:489 vocab_size_: 51865
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/offline-whisper-model-qnn.cc:PostInitDecoder:490 self_kv_size_: 172032
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/offline-whisper-model-qnn.cc:PostInitDecoder:491 self_kv_stride_: 448
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/offline-whisper-model-qnn.cc:InitSotSequence:563 sot_sequence: 50258 50259 50359 50363
eot: 50257

/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/offline-recognizer-whisper-tpl-impl.h:Init:69 use greedy_search
recognizer created in 0.587 s
Started
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/offline-whisper-model-qnn.cc:Run:152 Detecting language.
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/qnn/offline-whisper-model-qnn.cc:Run:159 Detected Language: en
Done!

./sherpa-onnx-qnn-SM8850-binary-whisper-tiny/test_wavs/1.wav
{"lang": "en", "emotion": "", "event": "", "text": " God, as a direct consequence of the sin which man thus punished, had given her a lovely child whose place was on that same dishonored bosom to connect her parrot forever with the race and descent of mortals, and to be finally a blessed soul in heaven.", "timestamps": [], "durations": [], "tokens":[" God", ",", " as", " a", " direct", " consequence", " of", " the", " sin", " which", " man", " thus", " punished", ",", " had", " given", " her", " a", " lovely", " child", " whose", " place", " was", " on", " that", " same", " dishon", "ored", " bos", "om", " to", " connect", " her", " parrot", " forever", " with", " the", " race", " and", " descent", " of", " mort", "als", ",", " and", " to", " be", " finally", " a", " blessed", " soul", " in", " heaven", "."], "ys_log_probs": [], "words": []}
----
num threads: 2
decoding method: greedy_search
Elapsed seconds: 1.535 s
Real time factor (RTF): 1.535 / 16.715 = 0.092
     0.0ms [WARN   ] QnnDsp <W> m_CFBCallbackInfoObj is not initialized, return emptyList
     0.0ms [WARN   ] QnnDsp <W> m_CFBCallbackInfoObj is not initialized, return emptyList

We use SM8850 in the above example. You can select more models for different SOC from
https://github.com/k2-fsa/sherpa-onnx/releases/tag/asr-models-qnn-binary-2

Summary by CodeRabbit

  • New Features

    • Added QNN-backed offline decoding support for Whisper models.
    • Expanded QNN Whisper model detection so Whisper QNN artifacts are selected and run automatically.
    • Improved Whisper QNN artifact naming and per-platform package directory prefixes for clearer outputs.
  • Bug Fixes

    • Strengthened validation for QNN Whisper configurations, including context-binary format checks and required encoder/decoder .so presence.
    • Updated supported-model error messaging to include Whisper QNN options.

@coderabbitai

coderabbitai Bot commented Jun 24, 2026 •

Copy link
Copy Markdown

Review Change Stack

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 10f525f2-33cf-4adf-b5c1-8dd52787516f

📥 Commits

Reviewing files that changed from the base of the PR and between 3e5368e and b2e9219.

📒 Files selected for processing (1)
  • sherpa-onnx/csrc/qnn/offline-whisper-model-qnn.cc

📝 Walkthrough

Walkthrough

Adds QNN Whisper config support, a new offline QNN Whisper decoder implementation, factory and build wiring, and export workflow renames for Whisper QNN artifacts.

Changes

Whisper QNN Integration

Layer / File(s) Summary
OfflineWhisperModelConfig QNN contract
sherpa-onnx/csrc/offline-whisper-model-config.h, sherpa-onnx/csrc/offline-whisper-model-config.cc
Adds QnnConfig qnn_config, extends validation for QNN Whisper artifacts, and updates ToString() to include QNN config when present.
OfflineWhisperModelQnn API and initialization
sherpa-onnx/csrc/qnn/offline-whisper-model-qnn.h, sherpa-onnx/csrc/qnn/offline-whisper-model-qnn.cc
Declares the QNN Whisper model API and implements loading, tensor setup, dimension extraction, token initialization, and wrapper state.
Whisper QNN inference loop
sherpa-onnx/csrc/qnn/offline-whisper-model-qnn.cc
Implements Run() with mel padding, encoder execution, KV cache updates, optional language detection, and decoder token generation.
Factory dispatch and build registration
sherpa-onnx/csrc/offline-recognizer-impl.cc, sherpa-onnx/csrc/CMakeLists.txt
Wires Whisper QNN into both recognizer factory paths and adds the QNN source file to the build.
CI workflow artifact naming
.github/workflows/export-whisper-qnn.yaml
Updates the trigger branch and Whisper QNN artifact directory naming.

Sequence Diagram(s)

sequenceDiagram
    participant App
    participant OfflineRecognizerImpl
    participant OfflineWhisperModelQnn
    participant QnnEncoderModel
    participant QnnDecoderModel

    App->>OfflineRecognizerImpl: Create(config, provider="qnn")
    OfflineRecognizerImpl->>OfflineWhisperModelQnn: construct(config)
    OfflineWhisperModelQnn->>QnnEncoderModel: Init encoder
    OfflineWhisperModelQnn->>QnnDecoderModel: Init decoder
    OfflineWhisperModelQnn-->>OfflineRecognizerImpl: ready

    App->>OfflineRecognizerImpl: Decode(audio)
    OfflineRecognizerImpl->>OfflineWhisperModelQnn: Run(features)
    OfflineWhisperModelQnn->>QnnEncoderModel: RunEncoder(mel)
    QnnEncoderModel-->>OfflineWhisperModelQnn: cross-KV tensors
    loop token generation
        OfflineWhisperModelQnn->>QnnDecoderModel: RunDecoder(token, KV cache, mask)
        QnnDecoderModel-->>OfflineWhisperModelQnn: logits
    end
    OfflineWhisperModelQnn-->>OfflineRecognizerImpl: transcript result
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes

Possibly related PRs

Suggested labels

size:XXL

🐇 I hopped through QNN with a whispering grin,
Mel frames and token trails spun neatly within.
A cache full of carrots, a decoder that sings,
Now Whisper runs faster on springy QNN springs.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 6.90% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the main change: adding a C++ Whisper runtime backed by QNN for Qualcomm NPU.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@dosubot dosubot Bot added the size:XL This PR changes 500-999 lines, ignoring generated files. label Jun 24, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request adds support for the Whisper model using the Qualcomm Neural Network (QNN) backend in sherpa-onnx, including configuration validation and the core model implementation. The review feedback highlights several critical safety improvements to prevent potential crashes and undefined behavior. Specifically, it is recommended to add bounds checks on the logits vector during language detection, verify tensor sizes before updating the KV cache or transposing encoder outputs, ensure shape vectors are not empty during decoder initialization, validate that cross_k_shape has at least three dimensions, and confirm that feat_dim_ is greater than zero to avoid a division-by-zero error.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment on lines +644 to +656
const auto &all_lang_ids = GetAllWhisperLanguageTokenIds();
int32_t lang_id = all_lang_ids[0];
float this_logit = logits[lang_id];

for (int32_t i = 1; i != static_cast<int32_t>(all_lang_ids.size()); ++i) {
int32_t id = all_lang_ids[i];
float p = logits[id];

if (p > this_logit) {
this_logit = p;
lang_id = id;
}
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

In DetectLanguage, the code directly accesses logits[lang_id] and logits[id] using language token IDs from GetAllWhisperLanguageTokenIds(). However, if the loaded model has a smaller vocabulary size (e.g., a pruned or custom model) such that logits.size() is less than the language token IDs (which can be up to 50357), this will cause an out-of-bounds vector access and crash the application.

Please add a safety check to ensure that we only query token IDs that are within the bounds of the logits vector.

    const auto &all_lang_ids = GetAllWhisperLanguageTokenIds();
    int32_t lang_id = -1;
    float this_logit = 0.0f;

    for (int32_t id : all_lang_ids) {
      if (id >= static_cast<int32_t>(logits.size())) {
        continue;
      }
      float p = logits[id];
      if (lang_id == -1 || p > this_logit) {
        this_logit = p;
        lang_id = id;
      }
    }

    if (lang_id == -1) {
      SHERPA_ONNX_LOGE("No valid language token found in logits (logits size: %d)",
                       static_cast<int32_t>(logits.size()));
      SHERPA_ONNX_EXIT(-1);
    }

Comment on lines +619 to +632
auto delta_k =
decoder_model_->GetOutputTensorData(decoder_this_self_k_[i]);
float *self_k = self_kv_data + (i * 2) * self_kv_size_;
for (size_t r = 0; r != delta_k.size(); ++r) {
self_k[r * self_kv_stride_ + offset] += delta_k[r];
}

// Update self_v
auto delta_v =
decoder_model_->GetOutputTensorData(decoder_this_self_v_[i]);
float *self_v = self_kv_data + (i * 2 + 1) * self_kv_size_;
for (size_t r = 0; r != delta_v.size(); ++r) {
self_v[r * self_kv_stride_ + offset] += delta_v[r];
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

In UpdateSelfKvCache, the code retrieves delta_k and delta_v from the decoder model and copies their elements into self_k and self_v assuming their size matches n_text_state_. If a mismatched or invalid model is loaded where the output tensor size of decoder_this_self_k_ or decoder_this_self_v_ is larger than n_text_state_, this loop will write out of bounds of the allocated self_kv buffer, leading to memory corruption or crashes.

Please add a defensive check to verify that the sizes of delta_k and delta_v match n_text_state_ before copying.

      auto delta_k = 
          decoder_model_->GetOutputTensorData(decoder_this_self_k_[i]);
      if (delta_k.size() != static_cast<size_t>(n_text_state_)) {
        SHERPA_ONNX_LOGE("Mismatched self_k size: expected %d, got %d",
                         n_text_state_, static_cast<int32_t>(delta_k.size()));
        SHERPA_ONNX_EXIT(-1);
      }
      float *self_k = self_kv_data + (i * 2) * self_kv_size_;
      for (size_t r = 0; r != delta_k.size(); ++r) {
        self_k[r * self_kv_stride_ + offset] += delta_k[r];
      }

      // Update self_v
      auto delta_v = 
          decoder_model_->GetOutputTensorData(decoder_this_self_v_[i]);
      if (delta_v.size() != static_cast<size_t>(n_text_state_)) {
        SHERPA_ONNX_LOGE("Mismatched self_v size: expected %d, got %d",
                         n_text_state_, static_cast<int32_t>(delta_v.size()));
        SHERPA_ONNX_EXIT(-1);
      }
      float *self_v = self_kv_data + (i * 2 + 1) * self_kv_size_;
      for (size_t r = 0; r != delta_v.size(); ++r) {
        self_v[r * self_kv_stride_ + offset] += delta_v[r];
      }

Comment on lines +111 to +117
auto cross_k = encoder_model_->GetOutputTensorData(encoder_cross_k_[i]);
auto cross_v = encoder_model_->GetOutputTensorData(encoder_cross_v_[i]);

cross_kv[i * 2] =
Transpose(cross_k.data(), num_out_frames_, n_text_state_);
cross_kv[i * 2 + 1] =
Transpose(cross_v.data(), num_out_frames_, n_text_state_);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

In the encoder output processing loop, Transpose is called on cross_k.data() and cross_v.data() with dimensions num_out_frames_ and n_text_state_. If the model output tensor size does not match num_out_frames_ * n_text_state_, Transpose will perform an out-of-bounds read, leading to undefined behavior or crashes.

Please add a defensive check to verify that the sizes of cross_k and cross_v match the expected dimensions before transposing.

      auto cross_k = encoder_model_->GetOutputTensorData(encoder_cross_k_[i]);
      auto cross_v = encoder_model_->GetOutputTensorData(encoder_cross_v_[i]);

      if (cross_k.size() != static_cast<size_t>(num_out_frames_ * n_text_state_)) {
        SHERPA_ONNX_LOGE("Mismatched cross_k size for layer %d: expected %d, got %d",
                         i, num_out_frames_ * n_text_state_, static_cast<int32_t>(cross_k.size()));
        SHERPA_ONNX_EXIT(-1);
      }
      if (cross_v.size() != static_cast<size_t>(num_out_frames_ * n_text_state_)) {
        SHERPA_ONNX_LOGE("Mismatched cross_v size for layer %d: expected %d, got %d",
                         i, num_out_frames_ * n_text_state_, static_cast<int32_t>(cross_v.size()));
        SHERPA_ONNX_EXIT(-1);
      }

      cross_kv[i * 2] = 
          Transpose(cross_k.data(), num_out_frames_, n_text_state_);
      cross_kv[i * 2 + 1] = 
          Transpose(cross_v.data(), num_out_frames_, n_text_state_);

Comment on lines +464 to +476
std::vector<int32_t> mask_shape = decoder_model_->TensorShape(mask_name);
mask_size_ = mask_shape[0];

std::string logits_name = prefix_ + "_logits";
if (!decoder_model_->HasTensor(logits_name)) {
SHERPA_ONNX_LOGE("Decoder does not have output tensor '%s'",
logits_name.c_str());
SHERPA_ONNX_EXIT(-1);
}

std::vector<int32_t> logits_shape =
decoder_model_->TensorShape(logits_name);
vocab_size_ = logits_shape.back();

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

In PostInitDecoder, the code retrieves mask_shape and logits_shape from the decoder model and accesses their elements without checking if the shape vectors are empty. If a malformed or incompatible model is loaded, TensorShape might return an empty vector, and accessing mask_shape[0] or logits_shape.back() will cause undefined behavior or crashes.

Please add safety checks to ensure these shape vectors are not empty before accessing their elements.

    std::vector<int32_t> mask_shape = decoder_model_->TensorShape(mask_name);
    if (mask_shape.empty()) {
      SHERPA_ONNX_LOGE("Decoder mask shape is empty");
      SHERPA_ONNX_EXIT(-1);
    }
    mask_size_ = mask_shape[0];

    std::string logits_name = prefix_ + "_logits";
    if (!decoder_model_->HasTensor(logits_name)) {
      SHERPA_ONNX_LOGE("Decoder does not have output tensor '%s'",
                       logits_name.c_str());
      SHERPA_ONNX_EXIT(-1);
    }

    std::vector<int32_t> logits_shape = 
        decoder_model_->TensorShape(logits_name);
    if (logits_shape.empty()) {
      SHERPA_ONNX_LOGE("Decoder logits shape is empty");
      SHERPA_ONNX_EXIT(-1);
    }
    vocab_size_ = logits_shape.back();

Comment on lines +430 to +434
std::vector<int32_t> cross_k_shape =
encoder_model_->TensorShape("cross_k_0");

num_out_frames_ = cross_k_shape[1];
n_text_state_ = cross_k_shape[2];

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

In PostInitEncoder, the code retrieves cross_k_shape and accesses its elements at indices 1 and 2 without verifying that the shape vector has at least 3 dimensions. If the shape vector is smaller, this will cause an out-of-bounds access and crash.

Please add a check to ensure cross_k_shape has at least 3 dimensions.

    std::vector<int32_t> cross_k_shape = 
        encoder_model_->TensorShape("cross_k_0");
    if (cross_k_shape.size() < 3) {
      SHERPA_ONNX_LOGE("Expected cross_k_0 to have at least 3 dimensions");
      SHERPA_ONNX_EXIT(-1);
    }

    num_out_frames_ = cross_k_shape[1];
    n_text_state_ = cross_k_shape[2];

Comment on lines +417 to +418
num_frames_ = mel_shape[1];
feat_dim_ = mel_shape[2];

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

In PostInitEncoder, after extracting num_frames_ and feat_dim_ from mel_shape, there is no check to ensure they are positive. If feat_dim_ is 0, a division-by-zero crash (SIGFPE) will occur in Run when computing num_frames = features.size() / feat_dim_.

Please add a check to ensure num_frames_ and feat_dim_ are greater than 0.

    num_frames_ = mel_shape[1];
    feat_dim_ = mel_shape[2];
    if (num_frames_ <= 0 || feat_dim_ <= 0) {
      SHERPA_ONNX_LOGE("Invalid encoder mel input shape: [%d, %d, %d]",
                       mel_shape[0], num_frames_, feat_dim_);
      SHERPA_ONNX_EXIT(-1);
    }

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 4

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In @.github/workflows/export-whisper-qnn.yaml:
- Line 413: The workflow shell step that builds and uses the d variable from
matrix values needs hardening because unquoted `${{ matrix.* }}` content is
later expanded through `$d`, which can trigger shell metacharacter
interpretation. Update the affected shell script blocks that assign and consume
`d` in the export-whisper-qnn workflow so the matrix-derived values are safely
quoted/escaped before any shell use, and ensure all subsequent references to `d`
preserve that quoting behavior.

In `@sherpa-onnx/csrc/offline-recognizer-impl.cc`:
- Around line 210-214: The Whisper dispatch in OfflineRecognizerImpl is too
broad and sends any non-empty whisper.encoder config into
OfflineWhisperModelQnn, even when the model is not actually a QNN artifact.
Update the Create() branching in OfflineRecognizerImpl to use the same
QNN-artifact check as OfflineWhisperModelConfig::Validate()—only dispatch to
OfflineRecognizerWhisperTplImpl<OfflineWhisperModelQnn> when the encoder/decoder
are QNN .so files or whisper.qnn_config.context_binary is set—so normal ONNX
Whisper configs do not enter the QNN path.

In `@sherpa-onnx/csrc/qnn/offline-whisper-model-qnn.cc`:
- Around line 65-71: The Manager-based ctor in OfflineWhisperModelQnn::Impl is
currently an always-failing stub, but OfflineRecognizerImpl::Create(Manager *,
...) now routes QNN Whisper through it. Either implement real Manager-backed
model loading in this constructor or remove/gate the Manager overload so the
factory rejects Android/OHOS asset-manager usage explicitly instead of aborting
after dispatch; use the OfflineRecognizerImpl::Create and Impl(Manager *, const
OfflineModelConfig &) symbols to locate the path.
- Around line 430-434: Validate the tensor ranks before indexing shapes in the
QNN Whisper model init path, since `cross_k_shape[1]`, `cross_k_shape[2]`,
`mask_shape[0]`, and `logits_shape.back()` can otherwise access out of bounds on
mismatched artifacts. Add explicit rank checks in the same style as the existing
`mel_shape` validation inside `OfflineWhisperModelQnn` before reading these
dimensions, and fail with a clear validation error if the encoder/decoder tensor
shapes do not match expectations. Apply the same safeguard around the related
shape reads in the later init block that uses `mask_shape` and `logits_shape`.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: c22c89c0-b439-41e6-b4d7-cc9b76848a0c

📥 Commits

Reviewing files that changed from the base of the PR and between 8c0586e and 3e5368e.

📒 Files selected for processing (7)
  • .github/workflows/export-whisper-qnn.yaml
  • sherpa-onnx/csrc/CMakeLists.txt
  • sherpa-onnx/csrc/offline-recognizer-impl.cc
  • sherpa-onnx/csrc/offline-whisper-model-config.cc
  • sherpa-onnx/csrc/offline-whisper-model-config.h
  • sherpa-onnx/csrc/qnn/offline-whisper-model-qnn.cc
  • sherpa-onnx/csrc/qnn/offline-whisper-model-qnn.h

echo "collect results"

d=sherpa-onnx-qnn-${{ matrix.soc }}-binary-${{ matrix.model_name }}
d=sherpa-onnx-qnn-${{ matrix.soc }}-binary-whisper-${{ matrix.model_name }}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔒 Security & Privacy | 🟠 Major | ⚡ Quick win

Quote templated matrix values before shell expansion.

Line 413 and Lines 430-432 interpolate ${{ matrix.* }} into bash variables that are later used unquoted ($d). This can allow shell metacharacter expansion and break or execute unintended commands if matrix values change upstream.

Suggested hardening diff
-          d=sherpa-onnx-qnn-${{ matrix.soc }}-binary-whisper-${{ matrix.model_name }}
-          mkdir -p $d
-          mkdir -p $d/test_wavs
-          cp -v non_quant/binary/*.bin $d/
-          cp -v $model_dir/tokens.txt $d
-          cp -v $model_dir/*.wav $d/test_wavs
+          d="sherpa-onnx-qnn-${{ matrix.soc }}-binary-whisper-${{ matrix.model_name }}"
+          mkdir -p "$d"
+          mkdir -p "$d/test_wavs"
+          cp -v non_quant/binary/*.bin "$d"/
+          cp -v "$model_dir"/tokens.txt "$d"
+          cp -v "$model_dir"/*.wav "$d"/test_wavs
@@
-              d=sherpa-onnx-qnn-whisper-${{ matrix.model_name }}-linux-x64
+              d="sherpa-onnx-qnn-whisper-${{ matrix.model_name }}-linux-x64"
             elif [[ $p == aarch64-android ]]; then
-              d=sherpa-onnx-qnn-whisper-${{ matrix.model_name }}-android-aarch64
+              d="sherpa-onnx-qnn-whisper-${{ matrix.model_name }}-android-aarch64"

Also applies to: 430-432

🧰 Tools
🪛 zizmor (1.26.1)

[warning] 413-413: code injection via template expansion (template-injection): may expand into attacker-controllable code

(template-injection)


[warning] 413-413: code injection via template expansion (template-injection): may expand into attacker-controllable code

(template-injection)


[warning] 413-413: code injection via template expansion (template-injection): may expand into attacker-controllable code

(template-injection)


[warning] 413-413: code injection via template expansion (template-injection): may expand into attacker-controllable code

(template-injection)

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In @.github/workflows/export-whisper-qnn.yaml at line 413, The workflow shell
step that builds and uses the d variable from matrix values needs hardening
because unquoted `${{ matrix.* }}` content is later expanded through `$d`, which
can trigger shell metacharacter interpretation. Update the affected shell script
blocks that assign and consume `d` in the export-whisper-qnn workflow so the
matrix-derived values are safely quoted/escaped before any shell use, and ensure
all subsequent references to `d` preserve that quoting behavior.

Source: Linters/SAST tools

Comment on lines +210 to +214
} else if (!config.model_config.whisper.encoder.empty() ||
!config.model_config.whisper.qnn_config.context_binary
.empty()) {
return std::make_unique<
OfflineRecognizerWhisperTplImpl<OfflineWhisperModelQnn>>(config);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Gate QNN Whisper dispatch on actual QNN artifacts.

These branches instantiate OfflineWhisperModelQnn for any non-empty whisper.encoder, but OfflineWhisperModelConfig::Validate() only recognizes Whisper-as-QNN when the encoder/decoder are .so or whisper.qnn_config.context_binary is set. With provider == "qnn" and a normal ONNX encoder path, Create() routes into the QNN codepath and OfflineWhisperModelQnn then tries to load that ONNX file as a QNN model library. Reuse the same QNN-artifact predicate here so unsupported configs fail at dispatch instead of during init.

Also applies to: 576-580

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@sherpa-onnx/csrc/offline-recognizer-impl.cc` around lines 210 - 214, The
Whisper dispatch in OfflineRecognizerImpl is too broad and sends any non-empty
whisper.encoder config into OfflineWhisperModelQnn, even when the model is not
actually a QNN artifact. Update the Create() branching in OfflineRecognizerImpl
to use the same QNN-artifact check as OfflineWhisperModelConfig::Validate()—only
dispatch to OfflineRecognizerWhisperTplImpl<OfflineWhisperModelQnn> when the
encoder/decoder are QNN .so files or whisper.qnn_config.context_binary is set—so
normal ONNX Whisper configs do not enter the QNN path.

Source: Learnings

Comment on lines +65 to +71
template <typename Manager>
Impl(Manager *mgr, const OfflineModelConfig &config) : config_(config) {
SHERPA_ONNX_LOGE(
"Please copy all files from assets to SD card and set assetManager to "
"null");
SHERPA_ONNX_EXIT(-1);
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🩺 Stability & Availability | 🟠 Major | ⚡ Quick win

Remove or gate the unsupported manager overload.

This ctor always exits, but the new OfflineRecognizerImpl::Create(Manager *, ...) branch now selects it for QNN Whisper. Any Android/OHOS asset-manager path will abort immediately after successful factory dispatch. Either implement manager-backed loading here or reject this case in the factory with an explicit unsupported-path message.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@sherpa-onnx/csrc/qnn/offline-whisper-model-qnn.cc` around lines 65 - 71, The
Manager-based ctor in OfflineWhisperModelQnn::Impl is currently an
always-failing stub, but OfflineRecognizerImpl::Create(Manager *, ...) now
routes QNN Whisper through it. Either implement real Manager-backed model
loading in this constructor or remove/gate the Manager overload so the factory
rejects Android/OHOS asset-manager usage explicitly instead of aborting after
dispatch; use the OfflineRecognizerImpl::Create and Impl(Manager *, const
OfflineModelConfig &) symbols to locate the path.

Comment on lines +430 to +434
std::vector<int32_t> cross_k_shape =
encoder_model_->TensorShape("cross_k_0");

num_out_frames_ = cross_k_shape[1];
n_text_state_ = cross_k_shape[2];

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🩺 Stability & Availability | 🟠 Major | ⚡ Quick win

Validate tensor ranks before indexing these shapes.

cross_k_shape[1], cross_k_shape[2], mask_shape[0], and logits_shape.back() assume the encoder/decoder artifacts are structurally correct. A mismatched model/context pair will turn that into out-of-bounds access during init instead of a clean validation error. Mirror the existing mel_shape rank check before reading these indices.

Also applies to: 464-476

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@sherpa-onnx/csrc/qnn/offline-whisper-model-qnn.cc` around lines 430 - 434,
Validate the tensor ranks before indexing shapes in the QNN Whisper model init
path, since `cross_k_shape[1]`, `cross_k_shape[2]`, `mask_shape[0]`, and
`logits_shape.back()` can otherwise access out of bounds on mismatched
artifacts. Add explicit rank checks in the same style as the existing
`mel_shape` validation inside `OfflineWhisperModelQnn` before reading these
dimensions, and fail with a clear validation error if the encoder/decoder tensor
shapes do not match expectations. Apply the same safeguard around the related
shape reads in the later init block that uses `mask_shape` and `logits_shape`.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XL This PR changes 500-999 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant