Skip to content

Add Android demo for Paraformer ASR with Qualcomm NPU. - #2932

Merged
csukuangfj merged 1 commit into
k2-fsa:masterfrom
csukuangfj:qnn-paraformer-android
Dec 24, 2025
Merged

csukuangfj merged 1 commit into
k2-fsa:masterfrom
csukuangfj:qnn-paraformer-android

Conversation

@csukuangfj

@csukuangfj csukuangfj commented Dec 24, 2025 •

Copy link
Copy Markdown
Collaborator

Summary by CodeRabbit

  • New Features

    • Added QNN-accelerated paraformer model variants for improved performance
    • Introduced support for int8 quantized model variants (sense-voice, paraformer)
    • Added new paraformer model options with optimized configurations
  • Bug Fixes

    • Enhanced asset file validation and multi-asset copy handling
    • Improved error messaging for QNN model configuration setup
  • Documentation

    • Updated offset calculation comments in audio processing

✏️ Tip: You can customize this high-level summary in your review settings.

@dosubot dosubot Bot added the size:L This PR changes 100-499 lines, ignoring generated files. label Dec 24, 2025
@gemini-code-assist

Copy link
Copy Markdown

Summary of Changes

Hello @csukuangfj, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly advances the Android demo by integrating Paraformer ASR models with Qualcomm NPU acceleration. The core changes involve extending the model configuration to support QNN-specific parameters for Paraformer, introducing robust asset management utilities for handling multi-file QNN models, and updating the Android application's initialization logic to prepare these models for NPU inference. Additionally, the build scripts have been adjusted to reflect the new model additions and naming conventions, ensuring a comprehensive update for NPU-powered ASR on Android.

Highlights

  • Paraformer ASR with Qualcomm NPU Support: Integrated Paraformer Automatic Speech Recognition (ASR) models with Qualcomm Neural Processing Unit (NPU) support into the Android demo.
  • Enhanced Asset Management: Introduced new utility functions (assetListExists, copyAssetListToInternalStorage) in Kotlin to efficiently handle and copy multiple model assets, which is crucial for complex QNN models.
  • QNN Configuration for Paraformer: Extended the OfflineParaformerModelConfig in both Java and Kotlin APIs to include QnnConfig, enabling detailed configuration for Paraformer models running on QNN.
  • Automated QNN Model Deployment: Implemented logic within the Android demo to automatically copy Paraformer QNN model components (e.g., encoder, predictor, decoder, context binaries) from assets to internal storage, streamlining deployment.
  • Updated Model Build Scripts: Modified Python scripts responsible for generating APKs to incorporate new Paraformer QNN models and ensure correct naming conventions for existing sense-voice models.
  • Improved Error Messaging: Enhanced the error message in the C++ QNN Paraformer model implementation to guide users on manual asset copying when assetManager is null.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@csukuangfj
csukuangfj requested a review from Copilot December 24, 2025 09:20
@coderabbitai

coderabbitai Bot commented Dec 24, 2025 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

This PR adds Qualcomm Neural Network (QNN) backend support for Paraformer ASR models across Android, Java, Kotlin, and C++ layers, including asset validation utilities, model configuration extensions, JNI config parsing updates, and new int8 model variants for APK generation scripts.

Changes

Cohort / File(s) Summary
QNN Configuration Structure
sherpa-onnx/java-api/src/main/java/com/k2fsa/sherpa/onnx/OfflineParaformerModelConfig.java, sherpa-onnx/kotlin-api/OfflineRecognizer.kt
Added QnnConfig field with getter/setter to OfflineParaformerModelConfig; extended Kotlin model config with qnnConfig field; added new model cases (9023, 9024, 9025) supporting QNN-based paraformer with encoder/predictor/decoder binaries.
Android Asset & QNN Handling
android/SherpaOnnxSimulateStreamingAsr/app/src/main/java/com/k2fsa/sherpa/onnx/simulate/streaming/asr/SimulateStreamingAsr.kt
Introduced assetListExists() and copyAssetListToInternalStorage() utilities for multi-asset validation and copying; extended QNN flow to handle paraformer context assets and tokens based on model or asset existence.
C++ QNN Runtime
sherpa-onnx/csrc/qnn/offline-paraformer-model-qnn.cc, sherpa-onnx/jni/offline-recognizer.cc
Updated error message in paraformer QNN constructor for asset SD card handling; added QNN config parsing in JNI for paraformer and sense_voice sections (backend_lib, context_binary, system_lib).
APK Model Configuration
scripts/apk/generate-asr-2pass-apk-script.py, scripts/apk/generate-qnn-vad-asr-apk-script.py, scripts/apk/generate-vad-asr-apk-script.py
Updated sense-voice model to int8 variant; added two new QNN paraformer int8 models (idx 9023, 9024) with standard cleanup scripts; adjusted cleanup logic to preserve model.onnx in second-model contexts.
Minor UI Update
android/SherpaOnnxSimulateStreamingAsr/app/src/main/java/com/k2fsa/sherpa/onnx/simulate/streaming/asr/screens/Home.kt
Updated comment from "offset 0.25s" to "offset 0.4s" with no functional behavior change.

Sequence Diagram

sequenceDiagram
    participant App as Android App
    participant Asset as Asset Manager
    participant Storage as Internal Storage
    participant JNI as JNI Layer
    participant CPP as C++ Runtime
    
    rect rgb(230, 245, 230)
    Note over App,Storage: Asset Validation & Setup Phase
    App->>Asset: assetListExists(paths)
    Asset-->>App: ✓ All assets present
    App->>Asset: copyAssetListToInternalStorage(paths)
    Asset->>Storage: Copy tokens, encoder, predictor
    Asset->>Storage: Copy context_binary files
    Storage-->>App: ✓ Assets copied
    end
    
    rect rgb(230, 240, 255)
    Note over App,JNI: Configuration Parsing Phase
    App->>JNI: getOfflineConfig() with QnnConfig
    JNI->>CPP: Parse paraformer.qnnConfig
    CPP->>CPP: Read backend_lib, context_binary, system_lib
    CPP-->>JNI: ✓ Config parsed
    JNI-->>App: ✓ Ready for inference
    end
    
    rect rgb(255, 245, 230)
    Note over CPP: QNN Model Execution
    App->>CPP: Create recognizer with QNN paraformer
    CPP->>CPP: Initialize QNN backend with loaded configs
    CPP-->>App: ✓ Ready for ASR
    end
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related PRs

Poem

🐰 Whiskers twitch with quantum cheer,
QNN paths now crystal clear!
Paraformer hops with int8 grace,
Assets bundled, configs in place—
Neural networks bound at last! ✨

Pre-merge checks and finishing touches

❌ Failed checks (1 warning)
Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 5.00% which is insufficient. The required threshold is 80.00%. You can run @coderabbitai generate docstrings to improve docstring coverage.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately reflects the main focus of the pull request by highlighting the addition of Android demo support for Paraformer ASR with Qualcomm NPU, which aligns with the core changes across multiple files.
✨ Finishing touches
  • 📝 Generate docstrings
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment

📜 Recent review details

Configuration used: defaults

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between cdbf0a8 and c8f68bf.

📒 Files selected for processing (9)
  • android/SherpaOnnxSimulateStreamingAsr/app/src/main/java/com/k2fsa/sherpa/onnx/simulate/streaming/asr/SimulateStreamingAsr.kt
  • android/SherpaOnnxSimulateStreamingAsr/app/src/main/java/com/k2fsa/sherpa/onnx/simulate/streaming/asr/screens/Home.kt
  • scripts/apk/generate-asr-2pass-apk-script.py
  • scripts/apk/generate-qnn-vad-asr-apk-script.py
  • scripts/apk/generate-vad-asr-apk-script.py
  • sherpa-onnx/csrc/qnn/offline-paraformer-model-qnn.cc
  • sherpa-onnx/java-api/src/main/java/com/k2fsa/sherpa/onnx/OfflineParaformerModelConfig.java
  • sherpa-onnx/jni/offline-recognizer.cc
  • sherpa-onnx/kotlin-api/OfflineRecognizer.kt
🧰 Additional context used
🧠 Learnings (2)
📚 Learning: 2025-08-06T04:23:50.237Z
Learnt from: litongjava
Repo: k2-fsa/sherpa-onnx PR: 2440
File: sherpa-onnx/java-api/src/main/java/com/k2fsa/sherpa/onnx/core/Core.java:4-6
Timestamp: 2025-08-06T04:23:50.237Z
Learning: The sherpa-onnx JNI library files are stored in Hugging Face repository at https://huggingface.co/csukuangfj/sherpa-onnx-libs under versioned directories like jni/1.12.7/, and the actual Windows JNI library filename is "sherpa-onnx-jni.dll" as defined in Core.java constants.

Applied to files:

  • sherpa-onnx/kotlin-api/OfflineRecognizer.kt
  • sherpa-onnx/jni/offline-recognizer.cc
📚 Learning: 2025-08-06T04:18:47.981Z
Learnt from: litongjava
Repo: k2-fsa/sherpa-onnx PR: 2440
File: sherpa-onnx/java-api/src/main/java/com/k2fsa/sherpa/onnx/core/Core.java:4-6
Timestamp: 2025-08-06T04:18:47.981Z
Learning: In sherpa-onnx Java API, the native library names in Core.java (WIN_NATIVE_LIBRARY_NAME = "sherpa-onnx-jni.dll", UNIX_NATIVE_LIBRARY_NAME = "libsherpa-onnx-jni.so", MACOS_NATIVE_LIBRARY_NAME = "libsherpa-onnx-jni.dylib") are copied directly from the compiled binary filenames and should not be changed to match other libraries' naming conventions.

Applied to files:

  • sherpa-onnx/jni/offline-recognizer.cc
🔇 Additional comments (19)
android/SherpaOnnxSimulateStreamingAsr/app/src/main/java/com/k2fsa/sherpa/onnx/simulate/streaming/asr/screens/Home.kt (1)

166-167: Accurate documentation of the speech offset.

The comment correctly documents the actual 6400-sample offset, which equals 0.4 seconds at the 16 kHz sample rate. This aligns the documentation with the implementation.

sherpa-onnx/csrc/qnn/offline-paraformer-model-qnn.cc (1)

98-104: Error message improvement provides clear guidance.

The updated error message is much better than a TODO placeholder—it gives users concrete steps to work around the limitation. Since the QNN backend appears to require file paths rather than asset streams, this workaround is reasonable.

However, verify that this limitation and workaround are documented in user-facing materials (README, examples, or API docs). Users encountering this constructor might expect it to work since it's publicly exposed via template instantiations.

scripts/apk/generate-qnn-vad-asr-apk-script.py (1)

402-404: LGTM!

The completion of the cmd block for the previous model entry ensures consistency with the structure used by all other models in the list.

android/SherpaOnnxSimulateStreamingAsr/app/src/main/java/com/k2fsa/sherpa/onnx/simulate/streaming/asr/SimulateStreamingAsr.kt (4)

26-41: LGTM! Asset list validation logic is sound.

The implementation correctly validates all comma-separated asset paths, handles edge cases (blank input, empty strings), and follows idiomatic Kotlin patterns.


76-91: LGTM! Asset list copying implementation is correct.

The function properly handles comma-separated asset paths, maintains consistent parsing logic with assetListExists, and correctly reconstructs the comma-separated output.


162-165: LGTM! QNN backend validation extended correctly for paraformer.

The condition properly extends the existing validation pattern to include paraformer models, ensuring at least one QNN backend library is configured.


206-225: LGTM! Paraformer asset copying logic is well-structured.

The implementation correctly handles paraformer's multi-file architecture (encoder, predictor, decoder) by leveraging the new list-based utility functions, and maintains consistency with the existing senseVoice and zipformerCtc patterns.

sherpa-onnx/jni/offline-recognizer.cc (2)

102-116: LGTM! QNN config parsing for paraformer implemented correctly.

The JNI field access pattern is consistent with the existing zipformerCtc QNN config parsing, and properly reads all required QNN configuration fields.


187-188: LGTM! Variable reuse is consistent with existing pattern.

The reassignment of qnn_config and qnn_config_cls for sense_voice follows the same pattern used for zipformerCtc, appropriately reusing variables across model-specific config sections.

sherpa-onnx/kotlin-api/OfflineRecognizer.kt (4)

25-25: LGTM! QNN config field added consistently.

The addition of qnnConfig field to OfflineParaformerModelConfig is consistent with the existing pattern in OfflineSenseVoiceModelConfig and OfflineZipformerCtcModelConfig.


409-409: LGTM! Formatting improvements for readability.

Line splits for long modelDir strings improve code readability without changing functionality.

Also applies to: 726-727, 758-759, 780-781, 798-799, 816-817, 833-834, 850-851, 867-868, 884-885, 901-902, 918-919, 935-936, 952-953, 970-971, 988-989, 1006-1007, 1023-1024, 1040-1041, 1057-1058, 1074-1075, 1091-1092, 1108-1109, 1125-1126, 1142-1144


1160-1177: LGTM! QNN paraformer configuration is well-structured.

The configuration correctly specifies comma-separated model files (encoder, predictor, decoder) and context binaries for QNN-accelerated paraformer, with helpful comments explaining the binary caching mechanism.


1179-1213: LGTM! Additional QNN paraformer configurations are correct.

Case 9024 follows the same pattern as 9023 with a newer model version. Case 9025 appropriately handles device-specific precompiled binaries for Xiaomi 17 Pro by specifying only context binaries without model files.

sherpa-onnx/java-api/src/main/java/com/k2fsa/sherpa/onnx/OfflineParaformerModelConfig.java (1)

7-7: LGTM! QNN config field integration follows Java best practices.

The implementation properly adds the qnnConfig field with correct encapsulation (private final), appropriate default value in the Builder, and follows the existing builder pattern for method chaining.

Also applies to: 11-11, 22-24, 28-28, 39-42

scripts/apk/generate-asr-2pass-apk-script.py (4)

409-416: LGTM: Consistent model pairing for Chinese second-pass ASR.

The int8 sense-voice model is correctly added to the second_zh list, enabling it to be paired with first-pass Chinese streaming models.


421-446: LGTM: Valid cross-language model pairing.

The int8 sense-voice model is correctly paired with the English streaming zipformer. This is valid since sense-voice supports multiple languages including English.


75-87: Inconsistency between AI summary and code.

The AI summary states "Removed the removal of model.onnx from the second-model cleanup script," but line 82 still contains rm -fv model.onnx for the paraformer model. This line was not modified in this PR (not marked with ~). The AI summary may be inaccurate or referring to a different context.


117-131: Verify that the int8 model variant exists at the download URL.

The model name references an int8 quantized variant which aligns with QNN/NPU optimization goals. The model is available and documented at:
https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/sherpa-onnx-sense-voice-zh-en-ja-ko-yue-int8-2024-07-17.tar.bz2

scripts/apk/generate-vad-asr-apk-script.py (1)

97-114: Verify Android app compatibility with the int8 model variant at idx=15.

The model archive sherpa-onnx-sense-voice-zh-en-ja-ko-yue-int8-2024-07-17.tar.bz2 exists and is available at the expected release URL. It contains model.int8.onnx (not model.onnx), which correctly aligns with the cleanup script's behavior of not removing model.onnx. Confirm that the Android app code properly handles this int8 model variant with the assigned index of 15.

Warning

Review ran into problems

🔥 Problems

Git: Failed to clone repository. Please run the @coderabbitai full review command to re-trigger a full review. If the issue persists, set path_filters to include or exclude specific files.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request adds support for Paraformer ASR models with Qualcomm NPU acceleration to the Android demo. The changes include updating Kotlin and Java data classes to include QNN configuration for Paraformer models, adding JNI bindings to read the new configuration, and implementing logic in the Android demo to handle copying of multiple model files required by Paraformer QNN models. The code is well-structured, and I have one suggestion to refactor some duplicated code in the Android demo to improve maintainability.

Comment on lines +76 to +91
fun copyAssetListToInternalStorage(
paths: String,
context: Context
): String {
if (paths.isBlank()) return paths

val pathList = paths.split(",")
.map { it.trim() }
.filter { it.isNotEmpty() }

val copiedPaths = pathList.map { path ->
copyAssetToInternalStorage(path, context)
}

return copiedPaths.joinToString(",")
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The logic for splitting and cleaning the comma-separated path string is duplicated in both assetListExists and copyAssetListToInternalStorage. To improve maintainability and reduce redundancy, you could extract this logic into a private helper function that can be called from both places.

For example:

private fun splitPaths(paths: String): List<String> {
    return paths.split(",")
        .map { it.trim() }
        .filter { it.isNotEmpty() }
}

@csukuangfj
csukuangfj merged commit 6cb5005 into k2-fsa:master Dec 24, 2025
3 checks passed
@csukuangfj
csukuangfj deleted the qnn-paraformer-android branch December 24, 2025 09:24

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR adds Android demo support for Paraformer ASR models running on Qualcomm NPU (Neural Processing Unit). The changes enable Paraformer models to utilize Qualcomm's QNN (Qualcomm Neural Network) SDK for hardware acceleration on compatible Android devices.

Key changes include:

  • Added QNN configuration support to OfflineParaformerModelConfig across Kotlin, Java, and JNI layers
  • Implemented three new Paraformer model configurations (indices 9023, 9024, 9025) with QNN backend
  • Extended asset handling to support comma-separated model paths required by Paraformer's multi-component architecture (encoder, predictor, decoder)

Reviewed changes

Copilot reviewed 9 out of 9 changed files in this pull request and generated no comments.

Show a summary per file
File Description
sherpa-onnx/kotlin-api/OfflineRecognizer.kt Added QnnConfig field to OfflineParaformerModelConfig; implemented three new Paraformer+QNN model configurations; updated model name for SenseVoice to int8 variant; reformatted long lines for readability
sherpa-onnx/jni/offline-recognizer.cc Added JNI code to read QNN configuration for Paraformer models; reused qnn_config variables following existing pattern for other models
sherpa-onnx/java-api/src/main/java/com/k2fsa/sherpa/onnx/OfflineParaformerModelConfig.java Added QnnConfig field with getter/setter to support QNN configuration in Java API
sherpa-onnx/csrc/qnn/offline-paraformer-model-qnn.cc Updated error message to provide clearer instructions for AssetManager usage
scripts/apk/generate-vad-asr-apk-script.py Updated SenseVoice model name to int8 variant; removed deletion of non-existent model.onnx file
scripts/apk/generate-qnn-vad-asr-apk-script.py Added two new Paraformer model configurations for QNN-enabled APK generation
scripts/apk/generate-asr-2pass-apk-script.py Updated SenseVoice model name references to int8 variant for consistency
android/SherpaOnnxSimulateStreamingAsr/.../Home.kt Fixed comment to correctly reflect 0.4s offset (was incorrectly labeled as 0.25s)
android/SherpaOnnxSimulateStreamingAsr/.../SimulateStreamingAsr.kt Added helper functions to handle comma-separated asset paths; extended QNN setup logic to support Paraformer models with multi-component structure

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L This PR changes 100-499 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants