Skip to content

Export Whisper models to QNN - #3697

Merged
csukuangfj merged 1 commit into
k2-fsa:masterfrom
csukuangfj:export-whisper-qnn
Jun 24, 2026
Merged

csukuangfj merged 1 commit into
k2-fsa:masterfrom
csukuangfj:export-whisper-qnn

Conversation

@csukuangfj

@csukuangfj csukuangfj commented Jun 23, 2026 •

Copy link
Copy Markdown
Collaborator

C++ runtime will be added in a separate PR.

Summary by CodeRabbit

  • New Features

    • Added automated Whisper model export workflow for generating hardware-optimized, deployable binaries across multiple device architectures.
  • Improvements

    • Enhanced artifact management and cleanup in model export pipelines.
    • Improved code organization through export and validation utility consolidation.

@dosubot dosubot Bot added the size:M This PR changes 30-99 lines, ignoring generated files. label Jun 23, 2026
@coderabbitai

coderabbitai Bot commented Jun 23, 2026 •

Copy link
Copy Markdown

Review Change Stack

Caution

Review failed

An error occurred during the review process. Please try again later.

📝 Walkthrough

Walkthrough

Adds a new GitHub Actions workflow (export-whisper-qnn.yaml) and supporting scripts to export Whisper models to ONNX on macOS and then convert them to QNN binaries for multiple SoCs on Ubuntu. Refactors shared rknn scripts to expose get_parser(), adds --wav support to test_onnx.py, and wires QNN scripts as symlinks to the shared rknn implementations. Minor fixes to parakeet and whisper-to-onnx workflows are also included.

Changes

Whisper QNN Export Pipeline

Layer / File(s) Summary
Shared rknn script refactoring for reuse
scripts/whisper/rknn/export_onnx.py, scripts/whisper/rknn/test_onnx.py, scripts/whisper/rknn/.gitignore
export_onnx.py extracts get_parser() from get_args() to make the parser importable. test_onnx.py switches to get_parser(), adds required --wav, expands model loading to per-model local-file branches with missing-file errors, uses compute_feat(args.wav), and reorders the decoder loop so idx is computed before the self-cache update. .gitignore excludes test_onnx_raw.py.
QNN symlinks and build-matrix generator
scripts/whisper/qnn/export_onnx.py, scripts/whisper/qnn/test_onnx.py, scripts/whisper/qnn/test_torch.py, .github/scripts/export-qnn/generate_whisper.py
The three qnn/ scripts are added as symlinks to their ../rknn/ counterparts. generate_whisper.py defines a Config dataclass and iterates soc_info_dict (excluding SM8350) paired with a fixed Whisper model list, printing a JSON {"include": [...]} matrix.
Workflow: triggers, ONNX job, and matrix generation
.github/workflows/export-whisper-qnn.yaml (lines 1–123)
Defines push/manual-dispatch triggers with per-ref concurrency cancellation. The onnx job runs on macOS across a model-name matrix, installs deps, conditionally downloads pretrained sources, exports ONNX, validates via test_onnx.py, and uploads compressed encoder/decoder/token/WAV artifacts. The generate_build_matrix job runs generate_whisper.py and exposes its JSON output.
Workflow: QNN job setup, conversion, packaging, and release
.github/workflows/export-whisper-qnn.yaml (lines 124–517)
The qnn job consumes the SoC/model matrix, downloads ONNX artifacts, sets up Python/NDK/QNN SDK, installs dependencies, runs qnn-onnx-converter → qnn-model-lib-generator → qnn-context-binary-generator for encoder and decoder, packages outputs into tarballs, uploads JSON configs, and conditionally publishes release assets to k2-fsa/sherpa-onnx based on repository_owner, matrix.soc, and matrix.model_name.

Minor Workflow Fixes

Layer / File(s) Summary
Parakeet archive cleanup and whisper-to-onnx diagnostic ls
.github/workflows/export-parakeet-ctc-qnn.yaml, .github/workflows/export-whisper-to-onnx.yaml
Adds rm $src.tar.bz2 after tar xvf in the parakeet workflow to remove the archive post-extraction. Inserts an ls -lh diagnostic step in the whisper-to-onnx workflow after the model download block.

Sequence Diagram(s)

sequenceDiagram
  participant branch as export-whisper-qnn branch
  participant onnx_job as onnx job (macOS)
  participant matrix_job as generate_build_matrix (Ubuntu)
  participant qnn_job as qnn job (Ubuntu)
  participant qnn_sdk as QNN SDK tools
  participant release as sherpa-onnx releases

  branch->>onnx_job: trigger per model_name matrix
  onnx_job->>onnx_job: install PyTorch/Whisper, download pretrained sources
  onnx_job->>onnx_job: export_onnx.py → test_onnx.py
  onnx_job->>onnx_job: upload encoder/decoder/tokens/WAV artifact

  branch->>matrix_job: trigger
  matrix_job->>matrix_job: generate_whisper.py → JSON matrix
  matrix_job-->>qnn_job: matrix (soc × model_name)

  qnn_job->>onnx_job: download ONNX artifact
  qnn_job->>qnn_sdk: qnn-onnx-converter (encoder + decoder)
  qnn_sdk-->>qnn_job: C/graph configs
  qnn_job->>qnn_sdk: qnn-model-lib-generator
  qnn_sdk-->>qnn_job: .so model libraries
  qnn_job->>qnn_sdk: qnn-context-binary-generator
  qnn_sdk-->>qnn_job: context binaries
  qnn_job->>qnn_job: package tarballs, upload JSON configs
  qnn_job->>release: publish tar.bz2 (conditional)
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes

Possibly related PRs

  • k2-fsa/sherpa-onnx#2986: Adds scripts/whisper/ascend-npu/export_onnx.py as a delegation to the same rknn/export_onnx.py that this PR refactors with get_parser().

Suggested labels

size:L

🐇 A new path through Qualcomm's gate,
The Whisper models now migrate!
With symlinks neat and parsers split,
Each SoC gets its perfect fit.
JSON matrices, tarballs galore —
This bunny hops to export more! 🎉

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 20.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'Export Whisper models to QNN' directly and clearly describes the main objective of the pull request, which is to add functionality for exporting Whisper models to QNN format across multiple new workflows, scripts, and configurations.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces Whisper model export and testing scripts for QNN and RKNN, including a configuration generator and updates to ONNX export and testing utilities. Key changes include exposing the argument parser in export_onnx.py, adding support for custom WAV files and distilled/special Whisper models in test_onnx.py, and fixing the KV cache update order. Feedback focuses on resolving a potential ModuleNotFoundError in generate_whisper.py by correctly setting up the system path for importing device_info, and refactoring the repetitive model-loading logic in test_onnx.py using a dictionary-based lookup to improve maintainability.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment on lines +4 to +8
import json

from device_info import soc_info_dict
from dataclasses import asdict, dataclass
import itertools

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The script imports device_info directly, but device_info.py is located in scripts/qnn/device_info.py. When running this script, it will fail with a ModuleNotFoundError because scripts/qnn is not in the Python search path. Additionally, itertools is imported but never used.

We can resolve this by dynamically adding the scripts/qnn directory to sys.path before importing device_info, and removing the unused itertools import.

Suggested change
import json
from device_info import soc_info_dict
from dataclasses import asdict, dataclass
import itertools
import json
import sys
from dataclasses import asdict, dataclass
from pathlib import Path
# Add scripts/qnn to sys.path to import device_info
sys.path.insert(0, str(Path(__file__).resolve().parents[3] / "scripts" / "qnn"))
from device_info import soc_info_dict

Comment on lines +127 to +207
name = args.model
if name == "distil-medium.en":
filename = "./distil-medium-en-original-model.bin"
if not Path(filename).is_file():
raise ValueError(
"""
Please go to https://huggingface.co/distil-whisper/distil-medium.en
to download original-model.bin
You can use the following command to do that:

wget -O distil-medium-en-original-model.bin https://huggingface.co/distil-whisper/distil-medium.en/resolve/main/original-model.bin
"""
)
torch_model = whisper.load_model(filename)
elif name == "distil-large-v2":
filename = "./distil-large-v2-original-model.bin"
if not Path(filename).is_file():
raise ValueError(
"""
Please go to https://huggingface.co/distil-whisper/distil-large-v2
to download original-model.bin
You can use the following command to do that:

wget -O distil-large-v2-original-model.bin https://huggingface.co/distil-whisper/distil-large-v2/resolve/main/original-model.bin
"""
)
torch_model = whisper.load_model(filename)
elif name == "distil-large-v3":
filename = "./distil-large-v3-original-model.bin"
if not Path(filename).is_file():
raise ValueError(
"""
Please go to https://huggingface.co/distil-whisper/distil-large-v3-openai
to download model.bin
You can use the following command to do that:

wget -O distil-large-v3-original-model.bin https://huggingface.co/distil-whisper/distil-large-v3-openai/resolve/main/model.bin
"""
)
torch_model = whisper.load_model(filename)
elif name == "distil-large-v3.5":
filename = "./distil-large-v3.5-original-model.bin"
if not Path(filename).is_file():
raise ValueError(
"""
Please go to https://huggingface.co/distil-whisper/distil-large-v3.5-openai/
to download model.bin
You can use the following command to do that:

wget -O distil-large-v3.5-original-model.bin https://huggingface.co/distil-whisper/distil-large-v3.5-openai/resolve/main/model.bin
"""
)
torch_model = whisper.load_model(filename)
elif name == "distil-small.en":
filename = "./distil-small-en-original-model.bin"
if not Path(filename).is_file():
raise ValueError(
"""
Please go to https://huggingface.co/distil-whisper/distil-small.en
to download original-model.bin
You can use the following command to do that:

wget -O distil-small-en-original-model.bin https://huggingface.co/distil-whisper/distil-small.en/resolve/main/original-model.bin
"""
)
torch_model = whisper.load_model(filename)
elif name == "medium-aishell":
filename = "./medium-aishell.pt"
if not Path(filename).is_file():
raise ValueError(
"""
Please go to https://huggingface.co/yuekai/icefall_asr_aishell_whisper/tree/main/exp_medium
to download whisper-medium-aishell1-epoch-10-avg-4.pt
You can use the following command to do that:

wget -O medium-aishell.pt https://huggingface.co/yuekai/icefall_asr_aishell_whisper/resolve/main/exp_medium/whisper-medium-aishell1-epoch-10-avg-4.pt
"""
)
torch_model = whisper.load_model(filename)
else:
torch_model = whisper.load_model(name)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The if/elif/else block for loading special/distilled models is highly repetitive and contains duplicated multi-line error messages. This can be simplified significantly by using a dictionary mapping model names to their metadata (local filename, download URL, repository, and download name). This improves readability, reduces boilerplate, and makes it much easier to add new models in the future.

    name = args.model
    special_models = {
        "distil-medium.en": {
            "filename": "./distil-medium-en-original-model.bin",
            "url": "https://huggingface.co/distil-whisper/distil-medium.en/resolve/main/original-model.bin",
            "repo": "https://huggingface.co/distil-whisper/distil-medium.en",
            "download_name": "original-model.bin",
        },
        "distil-large-v2": {
            "filename": "./distil-large-v2-original-model.bin",
            "url": "https://huggingface.co/distil-whisper/distil-large-v2/resolve/main/original-model.bin",
            "repo": "https://huggingface.co/distil-whisper/distil-large-v2",
            "download_name": "original-model.bin",
        },
        "distil-large-v3": {
            "filename": "./distil-large-v3-original-model.bin",
            "url": "https://huggingface.co/distil-whisper/distil-large-v3-openai/resolve/main/model.bin",
            "repo": "https://huggingface.co/distil-whisper/distil-large-v3-openai",
            "download_name": "model.bin",
        },
        "distil-large-v3.5": {
            "filename": "./distil-large-v3.5-original-model.bin",
            "url": "https://huggingface.co/distil-whisper/distil-large-v3.5-openai/resolve/main/model.bin",
            "repo": "https://huggingface.co/distil-whisper/distil-large-v3.5-openai/",
            "download_name": "model.bin",
        },
        "distil-small.en": {
            "filename": "./distil-small-en-original-model.bin",
            "url": "https://huggingface.co/distil-whisper/distil-small.en/resolve/main/original-model.bin",
            "repo": "https://huggingface.co/distil-whisper/distil-small.en",
            "download_name": "original-model.bin",
        },
        "medium-aishell": {
            "filename": "./medium-aishell.pt",
            "url": "https://huggingface.co/yuekai/icefall_asr_aishell_whisper/resolve/main/exp_medium/whisper-medium-aishell1-epoch-10-avg-4.pt",
            "repo": "https://huggingface.co/yuekai/icefall_asr_aishell_whisper/tree/main/exp_medium",
            "download_name": "whisper-medium-aishell1-epoch-10-avg-4.pt",
        },
    }

    if name in special_models:
        info = special_models[name]
        filename = info["filename"]
        if not Path(filename).is_file():
            raise ValueError(
                f"Please go to {info['repo']} to download {info['download_name']}\n"
                f"You can use the following command to do that:\n\n"
                f"wget -O {filename} {info['url']}"
            )
        torch_model = whisper.load_model(filename)
    else:
        torch_model = whisper.load_model(name)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M This PR changes 30-99 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant