Skip to content

Load QNN context binary for faster startup - #2877

Merged
csukuangfj merged 5 commits into
k2-fsa:masterfrom
csukuangfj:qnn-binary
Dec 9, 2025
Merged

csukuangfj merged 5 commits into
k2-fsa:masterfrom
csukuangfj:qnn-binary

Conversation

@csukuangfj

@csukuangfj csukuangfj commented Dec 9, 2025 •

Copy link
Copy Markdown
Collaborator

See the doc at
https://k2-fsa.github.io/sherpa/onnx/qnn/run-executables-on-your-phone-binary.html

Screenshot 2025-12-09 at 17 30 55

Summary by CodeRabbit

Release Notes

  • New Features

    • Added support for QNN context binaries in Sense Voice configurations
    • Introduced SM8850 device support for Android Sense Voice models
  • Improvements

    • Upgraded QNN toolkit to version 2.40.0.251030
    • Implemented dynamic build matrix in model export workflows
    • Enhanced model packaging and artifact organization

✏️ Tip: You can customize this high-level summary in your review settings.

@dosubot dosubot Bot added the size:XL This PR changes 500-999 lines, ignoring generated files. label Dec 9, 2025
@gemini-code-assist

Copy link
Copy Markdown

Summary of Changes

Hello @csukuangfj, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request enhances the sherpa-onnx library's integration with Qualcomm Neural Network (QNN) by enabling the direct loading of pre-compiled QNN context binaries for SenseVoice models. This change is designed to reduce model loading and compilation times, resulting in faster application startup on compatible Qualcomm devices. The PR also includes new scripts to facilitate the generation of device-specific QNN configurations, streamlining the development workflow.

Highlights

  • QNN Context Binary Loading: Introduced the capability to load pre-compiled QNN context binaries directly, which significantly speeds up model initialization and application startup on Qualcomm devices by bypassing full model compilation.
  • Flexible Model Configuration: Updated validation logic across C++ and Kotlin APIs to support configurations where only a QNN context binary is provided, making the SenseVoice model setup more versatile.
  • QNN Configuration Generation Tools: Added new Python scripts (device_info.py, generate_config.py, generate_sense_voice.py) to manage Qualcomm SOC information and automate the generation of necessary QNN backend and HTP configuration files.
Ignored Files
  • Ignored by pattern: .github/workflows/** (1)
    • .github/workflows/export-sense-voice-to-qnn.yaml
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@coderabbitai

coderabbitai Bot commented Dec 9, 2025 •

Copy link
Copy Markdown

Warning

Rate limit exceeded

@csukuangfj has exceeded the limit for the number of commits or files that can be reviewed per hour. Please wait 19 minutes and 39 seconds before requesting another review.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

📥 Commits

Reviewing files that changed from the base of the PR and between 1dc6586 and 0809172.

📒 Files selected for processing (2)
  • .github/workflows/export-sense-voice-to-qnn.yaml (12 hunks)
  • scripts/qnn/device_info.py (1 hunks)

Walkthrough

This PR introduces QNN (Qualcomm Neural Network) support for SenseVoice with new device information and configuration generation scripts. It significantly overhauls the export workflow to use dynamic matrix-driven SOC-aware builds, modernizes binary versions, and extends model validation to support QNN context binaries as alternatives to full models across C++, Kotlin, and Android codebases.

Changes

Cohort / File(s) Summary
QNN Device & Config Scripts
scripts/qnn/device_info.py, scripts/qnn/generate_config.py, scripts/qnn/generate_sense_voice.py
Introduces new Python infrastructure: device_info.py models Qualcomm QNN device metadata (Chipset enum, HtpArch, HtpInfo, SocInfo dataclasses, and device registry); generate_config.py generates HTP backend and device configuration JSON files; generate_sense_voice.py enumerates all (SOC, duration, framework) combinations to build a dynamic matrix for CI/CD.
GitHub Actions Workflow Overhaul
.github/workflows/export-sense-voice-to-qnn.yaml
Replaces static matrix with dynamic generation via new generate_build_matrix job; upgrades toolkit and binary versions (2.33.0.250327 → 2.40.0.251030); adds SOC dimension to job naming and matrix expansion; introduces context-binary generation, structured artifact packaging (directories, tar.bz2), README/LICENSE/info.txt handling, and conditional multi-target release steps for csukuangfj and k2-fsa repos.
Gitignore Update
.gitignore
Adds ignore pattern sherpa-onnx-qnn-* to exclude QNN-related build artifacts.
Android Asset Handling
android/.../SimulateStreamingAsr.kt
Updates conditional to copy contextBinary asset when either senseVoice.model is non-empty or contextBinary asset exists, allowing QNN support without a full model file.
C++ Model Validation & Factory Logic
sherpa-onnx/csrc/offline-model-config.cc, sherpa-onnx/csrc/offline-recognizer-impl.cc, sherpa-onnx/csrc/offline-sense-voice-model-config.cc
Extends validation and factory conditions to accept context_binary in QNN config as an alternative to a full model; allows QNN SenseVoice instantiation when only context binary is supplied.
Kotlin API Model Selection
sherpa-onnx/kotlin-api/OfflineRecognizer.kt
Adds new case 9022 for QNN-based SenseVoice configuration targeting SM8850 devices with context binary support.

Sequence Diagram

sequenceDiagram
    participant CI as GitHub Actions
    participant MatrixGen as generate_build_matrix Job
    participant Build as export-sense-voice-to-qnn Job
    participant Pkg as Packaging Step
    participant Release as Release Step
    
    CI->>MatrixGen: Trigger
    MatrixGen->>MatrixGen: Run generate_sense_voice.py<br/>(enumerate SOC × duration × framework)
    MatrixGen-->>CI: Output dynamic matrix
    
    CI->>Build: Spawn per-SOC build instances<br/>(needs: generate_build_matrix)
    loop For each SOC in matrix
        Build->>Build: Setup Python 3.10
        Build->>Build: Download toolkit v2.40.0.251030
        Build->>Build: Generate HTP config<br/>(generate_config.py)
        Build->>Build: Build model binary
        Build->>Build: Generate context binary<br/>(qnn-context-binary-generator)
        Build->>Pkg: Model + context binary ready
    end
    
    Pkg->>Pkg: Create directory structure<br/>(d/{SOC}/{language})
    Pkg->>Pkg: Copy README, LICENSE, model, binary
    Pkg->>Pkg: Generate info.txt
    Pkg->>Pkg: Tar/Bz2 packaging
    Pkg->>Pkg: Organize into binary/ and so/
    
    Pkg->>Release: Artifacts staged
    Release->>Release: Publish to csukuangfj & k2-fsa<br/>with tags (asr-models-qnn,<br/>asr-models-qnn-binary)
    Release-->>CI: Release complete
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes

  • .github/workflows/export-sense-voice-to-qnn.yaml: Major workflow restructuring with dynamic matrix generation, toolkit/binary version upgrades, new artifact handling pipeline, and multi-target release logic. Requires verification of matrix flow and packaging correctness.
  • scripts/qnn/generate_sense_voice.py and scripts/qnn/generate_config.py: New infrastructure scripts that need validation of SOC enumeration logic, config schema correctness, and integration with CI matrix generation.
  • C++ validation logic (offline-model-config.cc, offline-recognizer-impl.cc, offline-sense-voice-model-config.cc): Multiple conditional updates around context_binary support—verify OR logic correctness and absence of regressions in error handling.
  • Kotlin/Android changes: Simple but distributed updates; verify asset-copy conditions align with runtime expectations.

Possibly related PRs

Poem

🐇 Hopping through configs, SOCs aligned,
Dynamic matrices, QNN designs refined,
Context binaries dance, no models in sight,
Sense Voice on Qualcomm—oh, what a delight!
From device_info to releases so true,
A rabbit's applause—this build's quite askew! 🎉

Pre-merge checks and finishing touches

❌ Failed checks (1 warning)
Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 8.33% which is insufficient. The required threshold is 80.00%. You can run @coderabbitai generate docstrings to improve docstring coverage.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'Load QNN context binary for faster startup' accurately describes the main change across the PR, which adds QNN context binary support throughout the codebase for improved performance.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 7

🧹 Nitpick comments (2)
.github/scripts/export-qnn/generate_sense_voice.py (1)

20-39: Consider the potential size of the generated configuration matrix.

The script generates all combinations of SOCs, durations (11 values), and frameworks (2 values), resulting in 22 configurations per SOC. Depending on the number of SOCs in soc_info_dict, this could create a large CI matrix that may impact build times and resource usage.

If CI execution time or cost becomes a concern, consider:

  • Prioritizing a subset of durations for automated testing (e.g., 10, 20, 30 seconds)
  • Testing all combinations only for specific SOCs or on-demand
  • Using a staged approach where critical combinations run first
scripts/qnn/generate_config.py (1)

47-97: Consider validating qnn_sdk_root path exists.

The function assumes qnn_sdk_root is valid. If an invalid path is provided, the JSON files will be written with broken paths that will only fail at runtime when QNN tools try to load the shared library.

 def generate_config(
     soc_name: str,
     graph_name: str,
     output_dir: str,
     qnn_sdk_root: str,
 ):
     if soc_name not in soc_info_dict:
         raise ValueError(
             f"Unsupported SOC {soc_name}. Supported: - {sorted(list(soc_info_dict.keys()))}"
         )
     soc = soc_info_dict[soc_name]

+    qnn_sdk_path = Path(qnn_sdk_root)
+    if not qnn_sdk_path.exists():
+        raise ValueError(f"QNN SDK root does not exist: {qnn_sdk_root}")
+
     output_dir = Path(output_dir).absolute()
📜 Review details

Configuration used: CodeRabbit UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between b229e03 and 1dc6586.

📒 Files selected for processing (12)
  • .github/scripts/export-qnn/device_info.py (1 hunks)
  • .github/scripts/export-qnn/generate_config.py (1 hunks)
  • .github/scripts/export-qnn/generate_sense_voice.py (1 hunks)
  • .github/workflows/export-sense-voice-to-qnn.yaml (12 hunks)
  • .gitignore (1 hunks)
  • android/SherpaOnnxSimulateStreamingAsr/app/src/main/java/com/k2fsa/sherpa/onnx/simulate/streaming/asr/SimulateStreamingAsr.kt (1 hunks)
  • scripts/qnn/device_info.py (1 hunks)
  • scripts/qnn/generate_config.py (1 hunks)
  • sherpa-onnx/csrc/offline-model-config.cc (1 hunks)
  • sherpa-onnx/csrc/offline-recognizer-impl.cc (2 hunks)
  • sherpa-onnx/csrc/offline-sense-voice-model-config.cc (2 hunks)
  • sherpa-onnx/kotlin-api/OfflineRecognizer.kt (1 hunks)
🧰 Additional context used
🧠 Learnings (2)
📚 Learning: 2025-08-06T04:23:50.237Z
Learnt from: litongjava
Repo: k2-fsa/sherpa-onnx PR: 2440
File: sherpa-onnx/java-api/src/main/java/com/k2fsa/sherpa/onnx/core/Core.java:4-6
Timestamp: 2025-08-06T04:23:50.237Z
Learning: The sherpa-onnx JNI library files are stored in Hugging Face repository at https://huggingface.co/csukuangfj/sherpa-onnx-libs under versioned directories like jni/1.12.7/, and the actual Windows JNI library filename is "sherpa-onnx-jni.dll" as defined in Core.java constants.

Applied to files:

  • .gitignore
📚 Learning: 2025-08-06T04:18:47.981Z
Learnt from: litongjava
Repo: k2-fsa/sherpa-onnx PR: 2440
File: sherpa-onnx/java-api/src/main/java/com/k2fsa/sherpa/onnx/core/Core.java:4-6
Timestamp: 2025-08-06T04:18:47.981Z
Learning: In sherpa-onnx Java API, the native library names in Core.java (WIN_NATIVE_LIBRARY_NAME = "sherpa-onnx-jni.dll", UNIX_NATIVE_LIBRARY_NAME = "libsherpa-onnx-jni.so", MACOS_NATIVE_LIBRARY_NAME = "libsherpa-onnx-jni.dylib") are copied directly from the compiled binary filenames and should not be changed to match other libraries' naming conventions.

Applied to files:

  • .gitignore
🧬 Code graph analysis (2)
sherpa-onnx/csrc/offline-sense-voice-model-config.cc (1)
sherpa-onnx/kotlin-api/OfflineRecognizer.kt (7)
  • model (23-25)
  • model (27-29)
  • model (31-33)
  • model (35-38)
  • model (40-42)
  • model (44-46)
  • model (76-81)
scripts/qnn/generate_config.py (3)
.github/scripts/export-qnn/generate_config.py (3)
  • get_args (14-44)
  • generate_config (47-97)
  • _test (100-107)
scripts/qnn/device_info.py (1)
  • _test (98-108)
.github/scripts/export-qnn/device_info.py (1)
  • _test (98-108)
🪛 actionlint (1.7.9)
.github/workflows/export-sense-voice-to-qnn.yaml

27-27: workflow command "set-output" was deprecated. use echo "{name}={value}" >> $GITHUB_OUTPUT instead: https://docs.github.com/en/actions/using-workflows/workflow-commands-for-github-actions

(deprecated-commands)

🪛 Cppcheck (2.18.0)
sherpa-onnx/csrc/offline-recognizer-impl.cc

[information] Limiting analysis of branches. Use --check-level=exhaustive to analyze all branches.

(normalCheckLevelMaxBranches)


[information] Limiting analysis of branches. Use --check-level=exhaustive to analyze all branches.

(normalCheckLevelMaxBranches)

🪛 Flake8 (7.3.0)
scripts/qnn/device_info.py

[error] 3-3: 'enum.unique' imported but unused

(F401)

🪛 Ruff (0.14.8)
scripts/qnn/device_info.py

29-29: An enum class should not be decorated with @dataclass

(RUF049)


52-52: An enum class should not be decorated with @dataclass

(RUF049)

scripts/qnn/generate_config.py

54-56: Avoid specifying long messages outside the exception class

(TRY003)

⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (21)
  • GitHub Check: Release static tts-OFF
  • GitHub Check: Release static tts-ON
  • GitHub Check: Release shared tts-ON
  • GitHub Check: Debug shared-ON tts-ON
  • GitHub Check: Release shared-ON tts-OFF
  • GitHub Check: Debug shared-OFF tts-OFF
  • GitHub Check: Release shared-OFF tts-OFF
  • GitHub Check: Release shared-OFF tts-ON
  • GitHub Check: rknn shared OFF
  • GitHub Check: Debug shared-ON tts-OFF
  • GitHub Check: Release shared-ON tts-ON
  • GitHub Check: Debug shared-OFF tts-ON
  • GitHub Check: rknn shared ON
  • GitHub Check: ubuntu-24.04 3.10
  • GitHub Check: ubuntu-24.04 3.13
  • GitHub Check: ubuntu-24.04 3.11
  • GitHub Check: ubuntu-24.04 3.12
  • GitHub Check: swift (macos-latest)
  • GitHub Check: ubuntu-24.04 3.9
  • GitHub Check: ubuntu-24.04 3.8
  • GitHub Check: swift (macos-13)
🔇 Additional comments (10)
.gitignore (1)

166-166: Pattern addition aligns with PR objectives and existing conventions.

The new sherpa-onnx-qnn-* pattern appropriately ignores QNN-related artifacts and binaries, consistent with similar model artifact patterns already in the file (e.g., sherpa-onnx-sense-voice-*, sherpa-onnx-vits-*). This complements the PR's goal to support loading QNN context binaries for faster startup.

sherpa-onnx/csrc/offline-model-config.cc (1)

138-139: LGTM! Validation now supports QNN context binary path.

The updated condition correctly triggers sense_voice validation when either a model or a QNN context binary is provided, enabling the faster startup path described in the PR objectives.

sherpa-onnx/csrc/offline-sense-voice-model-config.cc (1)

32-67: LGTM! Comprehensive validation for QNN context binary support.

The validation logic correctly handles both traditional model-based and new context-binary-based paths:

  • When context_binary is empty, requires a valid model file
  • When model is empty but context_binary is provided, verifies the context binary exists
  • Properly delegates to qnn_config.Validate() for QNN-specific validation
sherpa-onnx/csrc/offline-recognizer-impl.cc (1)

174-175: LGTM! Factory logic updated consistently across both overloads.

The QNN branch condition now correctly instantiates the OfflineSenseVoiceModelQnn implementation when either sense_voice.model or sense_voice.qnn_config.context_binary is provided. This change is applied consistently in both Create() overloads and aligns with the validation logic updates.

Also applies to: 495-496

android/SherpaOnnxSimulateStreamingAsr/app/src/main/java/com/k2fsa/sherpa/onnx/simulate/streaming/asr/SimulateStreamingAsr.kt (1)

135-149: LGTM! Asset copying logic correctly handles QNN context binary scenarios.

The updated logic properly handles both model-based and context-binary-only paths:

  • Checks if either the model or context binary asset exists before attempting to copy
  • Only copies the model if it's non-empty
  • Always prepares the context binary path (which may not exist initially but will be created on first run)

This aligns with the QNN context binary workflow where the binary can be generated and cached after the first run.

.github/workflows/export-sense-voice-to-qnn.yaml (2)

14-34: LGTM - Dynamic matrix generation looks good.

The matrix generation job correctly outputs the matrix via $GITHUB_OUTPUT and the deprecated set-output is appropriately commented out as a reference.


567-605: Release conditions use different SOC filters - verify intent.

The release steps have asymmetric SOC filtering:

  • Lines 568 and 590: csukuangfj and k2-fsa only release .so files when matrix.soc == 'SM8850'
  • Lines 579 and 599: Binary releases happen for all SOCs

Is this intentional that .so releases are limited to SM8850 only while binaries are released for all SOCs?

scripts/qnn/generate_config.py (2)

1-11: LGTM - Clean script structure with good documentation references.

The script is well-organized with clear CLI arguments and follows the same pattern as .github/scripts/export-qnn/generate_config.py.


100-107: Hardcoded test path is fine for local testing.

The _test() function with hardcoded path is acceptable for development/debugging purposes and is properly commented out in __main__.

scripts/qnn/device_info.py (1)

63-95: LGTM - Well-structured device info data.

The HtpInfo, SocInfo dataclasses and the soc_info_list/soc_info_dict structures are well-organized and provide a clean mapping of SOC metadata.

Comment on lines 319 to +338
echo "collect results"

d=sherpa-onnx-qnn-${{ matrix.soc}}-binary-$t-seconds-sense-voice-zh-en-ja-ko-yue-2024-07-17-int8
mkdir -p $d
mkdir -p $d/test_wavs

cp -v README.md $d
cp -v LICENSE $d
cp -v binary/model.bin $d/
cp -v tokens.txt $d
cp -v *.wav $d/test_wavs

echo "num_frames=$num_frames" > $d/info.txt
echo "target=$p" >> $d/info.txt

ls -lh $d
tar cjfv $d.tar.bz2 $d
ls -lh *.tar.bz2
rm -rf $d
mv *.tar.bz2 ../../../binary/

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Undefined variable $p will cause incorrect output.

At line 332, $p is referenced but it's not defined until the for loop at line 341. This will result in an empty value being written to info.txt.

Apply this diff to fix the issue:

           echo "num_frames=$num_frames" > $d/info.txt
-          echo "target=$p" >> $d/info.txt
+          echo "target=binary" >> $d/info.txt

Or if you want to indicate the SOC:

           echo "num_frames=$num_frames" > $d/info.txt
-          echo "target=$p" >> $d/info.txt
+          echo "soc=${{ matrix.soc }}" >> $d/info.txt
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
echo "collect results"
d=sherpa-onnx-qnn-${{ matrix.soc}}-binary-$t-seconds-sense-voice-zh-en-ja-ko-yue-2024-07-17-int8
mkdir -p $d
mkdir -p $d/test_wavs
cp -v README.md $d
cp -v LICENSE $d
cp -v binary/model.bin $d/
cp -v tokens.txt $d
cp -v *.wav $d/test_wavs
echo "num_frames=$num_frames" > $d/info.txt
echo "target=$p" >> $d/info.txt
ls -lh $d
tar cjfv $d.tar.bz2 $d
ls -lh *.tar.bz2
rm -rf $d
mv *.tar.bz2 ../../../binary/
echo "collect results"
d=sherpa-onnx-qnn-${{ matrix.soc}}-binary-$t-seconds-sense-voice-zh-en-ja-ko-yue-2024-07-17-int8
mkdir -p $d
mkdir -p $d/test_wavs
cp -v README.md $d
cp -v LICENSE $d
cp -v binary/model.bin $d/
cp -v tokens.txt $d
cp -v *.wav $d/test_wavs
echo "num_frames=$num_frames" > $d/info.txt
echo "target=binary" >> $d/info.txt
ls -lh $d
tar cjfv $d.tar.bz2 $d
ls -lh *.tar.bz2
rm -rf $d
mv *.tar.bz2 ../../../binary/
🤖 Prompt for AI Agents
.github/workflows/export-sense-voice-to-qnn.yaml around lines 319 to 338: the
script writes "target=$p" to info.txt but $p is not defined yet (defined later
in the for loop), so replace that usage with a defined value — either write the
SOC from the matrix (e.g., matrix.soc) or ensure p is set before this block;
update the echo to use a defined variable (e.g., target=${p:-${{ matrix.soc }}})
or move the info.txt write into the loop after p is assigned so the correct
target is recorded.

Comment thread .github/workflows/export-sense-voice-to-qnn.yaml Outdated
Comment on lines +509 to +527
echo "collect results"

d=sherpa-onnx-qnn-${{ matrix.soc }}-binary-$t-seconds-sense-voice-zh-en-ja-ko-yue-2025-09-09-int8
mkdir -p $d
mkdir -p $d/test_wavs

cp -v README.md $d
cp -v binary/model.bin $d/
cp -v tokens.txt $d
cp -v *.wav $d/test_wavs

echo "num_frames=$num_frames" > $d/info.txt
echo "target=$p" >> $d/info.txt

ls -lh $d
tar cjfv $d.tar.bz2 $d
ls -lh *.tar.bz2
rm -rf $d
mv *.tar.bz2 ../../../binary/

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Undefined variable $p - same issue as FunASR section.

At line 521, $p is referenced but it's not defined until the for loop at line 530. This mirrors the same bug in the FunASR section.

Apply this diff to fix the issue:

           echo "num_frames=$num_frames" > $d/info.txt
-          echo "target=$p" >> $d/info.txt
+          echo "target=binary" >> $d/info.txt
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
echo "collect results"
d=sherpa-onnx-qnn-${{ matrix.soc }}-binary-$t-seconds-sense-voice-zh-en-ja-ko-yue-2025-09-09-int8
mkdir -p $d
mkdir -p $d/test_wavs
cp -v README.md $d
cp -v binary/model.bin $d/
cp -v tokens.txt $d
cp -v *.wav $d/test_wavs
echo "num_frames=$num_frames" > $d/info.txt
echo "target=$p" >> $d/info.txt
ls -lh $d
tar cjfv $d.tar.bz2 $d
ls -lh *.tar.bz2
rm -rf $d
mv *.tar.bz2 ../../../binary/
echo "collect results"
d=sherpa-onnx-qnn-${{ matrix.soc }}-binary-$t-seconds-sense-voice-zh-en-ja-ko-yue-2025-09-09-int8
mkdir -p $d
mkdir -p $d/test_wavs
cp -v README.md $d
cp -v binary/model.bin $d/
cp -v tokens.txt $d
cp -v *.wav $d/test_wavs
echo "num_frames=$num_frames" > $d/info.txt
echo "target=binary" >> $d/info.txt
ls -lh $d
tar cjfv $d.tar.bz2 $d
ls -lh *.tar.bz2
rm -rf $d
mv *.tar.bz2 ../../../binary/
🤖 Prompt for AI Agents
.github/workflows/export-sense-voice-to-qnn.yaml around lines 509 to 527: the
script writes "target=$p" but $p is undefined at this point (it’s set later in
the for-loop), causing the same bug as in the FunASR section; fix by using the
already-defined variable $t (used earlier to build the directory name) or
otherwise ensure $p is defined before use — specifically replace the "target=$p"
line with "target=$t" so the target value is correct and defined.

Comment on lines +1 to +4
#!/usr/bin/env python3
from dataclasses import dataclass
from enum import IntEnum, unique

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Remove unused import unique.

The unique decorator is imported but never used.

 #!/usr/bin/env python3
 from dataclasses import dataclass
-from enum import IntEnum, unique
+from enum import IntEnum
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
#!/usr/bin/env python3
from dataclasses import dataclass
from enum import IntEnum, unique
#!/usr/bin/env python3
from dataclasses import dataclass
from enum import IntEnum
🧰 Tools
🪛 Flake8 (7.3.0)

[error] 3-3: 'enum.unique' imported but unused

(F401)

🤖 Prompt for AI Agents
In scripts/qnn/device_info.py around lines 1 to 4, the imported symbol `unique`
from the enum module is unused; remove the `unique` import from the imports list
(keep the other imports intact) so the file no longer imports an unused name and
update any linter if necessary.

Comment thread scripts/qnn/device_info.py Outdated
Comment thread scripts/qnn/device_info.py Outdated
Comment on lines +1117 to +1133
9022 -> {
// for my Xiaomi 17 Pro
val modelDir = "sherpa-onnx-qnn-SM8850-binary-10-seconds-sense-voice-zh-en-ja-ko-yue-2024-07-17-int8"
return OfflineModelConfig(
provider = "qnn",
senseVoice = OfflineSenseVoiceModelConfig(
qnnConfig = QnnConfig(
// Please copy libQnnHtp.so and libQnnSystem.so to jniLibs/arm64-v8a by yourself
backendLib = "libQnnHtp.so",
systemLib = "libQnnSystem.so",
contextBinary = "$modelDir/model.bin",
),
),
tokens = "$modelDir/tokens.txt",
debug = true,
)
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Consider generalizing this device-specific test configuration.

The comment "for my Xiaomi 17 Pro" indicates this is a personal test case for a specific device. While it demonstrates the new QNN context binary feature correctly, including developer-specific test configurations in production code can clutter the codebase.

Consider either:

  • Removing the personal comment and making the configuration more generic
  • Moving this to an examples directory or test suite
  • Using it as a template that users can customize
🤖 Prompt for AI Agents
In sherpa-onnx/kotlin-api/OfflineRecognizer.kt around lines 1117 to 1133, this
block contains a device-specific comment and test configuration ("for my Xiaomi
17 Pro") that should not remain in production code; remove the personal comment,
generalize or rename the modelDir to a neutral example name (e.g.,
"sherpa-onnx-qnn-model-..."), or move the entire case into an examples/tests
file and replace it here with a concise template comment describing how to
customize QNN provider fields; ensure the example still documents the required
backendLib/systemLib/contextBinary fields and keep debug flag as appropriate.

@csukuangfj
csukuangfj merged commit 8fac37f into k2-fsa:master Dec 9, 2025
26 of 27 checks passed

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

The pull request introduces support for loading QNN context binaries to accelerate startup times for SenseVoice models on Android. This involves modifications in C++ and Kotlin code to handle the new context_binary configuration field, ensuring proper validation and asset management. Additionally, new Python scripts are added to generate configuration files related to QNN device information and model settings. The changes appear to be well-integrated and logically sound, enhancing the flexibility and performance of QNN model loading.

@@ -0,0 +1,112 @@
#!/usr/bin/env python3
from dataclasses import dataclass
from enum import IntEnum, unique

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The unique decorator is imported from the enum module but is not used anywhere in this file. Removing unused imports helps maintain code cleanliness and avoids potential confusion.

Suggested change
from enum import IntEnum, unique
from enum import IntEnum

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XL This PR changes 500-999 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant