Skip to content

Add CI to export Whisper models to Ascend NPU - #3008

Merged
csukuangfj merged 1 commit into
k2-fsa:masterfrom
csukuangfj:ci-export-whisper-ascend-npu
Jan 8, 2026
Merged

csukuangfj merged 1 commit into
k2-fsa:masterfrom
csukuangfj:ci-export-whisper-ascend-npu

Conversation

@csukuangfj

@csukuangfj csukuangfj commented Jan 8, 2026 •

Copy link
Copy Markdown
Collaborator

You can find them at
https://huggingface.co/k2-fsa/sherpa-onnx-models/tree/main/asr-models/ascend-npu/whisper
or at
https://github.com/k2-fsa/sherpa-onnx/releases/tag/asr-models-ascend

Summary by CodeRabbit

Release Notes

  • New Features

    • Added automated publishing pipelines to distribute exported models to HuggingFace and ModelScope repositories
  • Chores

    • Enhanced CI/CD workflows for Ascend NPU model export with improved intermediate artifact cleanup
    • Expanded environment setup with Git and Git LFS support for export workflows
    • Added diagnostic logging to model conversion processes for troubleshooting

✏️ Tip: You can customize this high-level summary in your review settings.

@dosubot dosubot Bot added the size:XL This PR changes 500-999 lines, ignoring generated files. label Jan 8, 2026
@coderabbitai

coderabbitai Bot commented Jan 8, 2026 •

Copy link
Copy Markdown

Caution

Review failed

The pull request is closed.

📝 Walkthrough

Walkthrough

This pull request introduces a new Whisper model export workflow for Ascend NPU alongside a new Python generator script that creates model configuration matrices. It updates existing export workflows (Paraformer, SenseVoice, Zipformer) with artifact cleanup and publishing stages to HuggingFace and ModelScope repositories. Minor debug logging is added to test and export scripts.

Changes

Cohort / File(s) Summary
New Whisper Export Infrastructure
.github/scripts/export-ascend/generate_whisper.py
Introduces Config dataclass with cann, soc_version, model, and image fields. Implements main() to enumerate CANN/SOC versions, generate model combinations via itertools.product(), compute images via get_image(), and output JSON matrix for CI consumption.
New Whisper Export Workflow
.github/workflows/export-whisper-to-ascend-npu.yaml
New comprehensive CI/CD workflow that generates build matrix, conditionally runs export jobs per model/CANN/SOC combination in containerized environment, downloads model weights from HuggingFace, exports to ONNX format, converts to Ascend OM via atc tool, packages artifacts as tar.bz2, and publishes to HuggingFace and ModelScope.
Existing Export Workflows (3 files)
.github/workflows/export-paraformer-to-ascend-npu.yaml, .github/workflows/export-sense-voice-to-ascend-npu.yaml, .github/workflows/export-zipformer-ctc-to-ascend-20250703.yaml
Update trigger branch references, install git and git-lfs, add cleanup steps to remove intermediate artifacts (\.pt, \.onnx, \*.om) at conversion and packaging stages, and introduce HuggingFace and ModelScope publishing stages with LFS tracking and repository commits.
Debug Logging
scripts/whisper/ascend-npu/test_om.py, scripts/whisper/rknn/export_onnx.py
Add console output statements: detailed token/tensor diagnostics in decoding loop (test_om.py) and confirmation message after token file export (export_onnx.py).

Sequence Diagram

sequenceDiagram
    participant GH as GitHub Actions
    participant PyGen as Python Generator
    participant Matrix as Build Matrix
    participant Exporter as Export Container
    participant HF as HuggingFace
    participant MS as ModelScope

    GH->>PyGen: Invoke generate_whisper.py
    PyGen->>PyGen: Enumerate CANN/SOC versions
    PyGen->>PyGen: Generate model combinations
    PyGen->>GH: Output JSON matrix
    GH->>Matrix: Create matrix job
    Matrix->>Exporter: Trigger export job per config
    Exporter->>Exporter: Download model weights
    Exporter->>Exporter: Export to ONNX
    Exporter->>Exporter: Convert ONNX to OM (atc)
    Exporter->>Exporter: Package tar.bz2 artifact
    Exporter->>HF: Publish with git-lfs
    Exporter->>MS: Publish with git-lfs
    HF-->>Exporter: Confirm push
    MS-->>Exporter: Confirm push
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~50 minutes

Possibly related PRs

  • Export models for Ascend 910B4 #2878: Adds similar export-matrix generator script under .github/scripts/export-ascend/ that defines Config dataclass and reuses same helpers (get_cann_version, get_soc_version, get_image).
  • Export models for CANN 8.2 #2745: Modifies Ascend export CI workflows to incorporate CANN version matrix handling and conditional execution based on repository ownership.
  • Export models to Ascend 910B3 #2761: Updates multiple .github/workflows/export-*-to-ascend-npu.yaml files with publishing and cleanup patterns similar to those introduced here.

Suggested labels

size:XXL, ci/cd, ascend-npu

Poem

🐰 Whisper models hop to Ascend with glee,
CANN versions spinning in matrix spree,
ONNX converts, OM files take flight,
GitHub Actions runs the workflow just right—
To HuggingFace and ModelScope they'll go,
Let the rabbit-powered exports flow! 🚀

✨ Finishing touches
  • 📝 Generate docstrings

📜 Recent review details

Configuration used: defaults

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 42aa0b8 and 482ddc6.

📒 Files selected for processing (7)
  • .github/scripts/export-ascend/generate_whisper.py
  • .github/workflows/export-paraformer-to-ascend-npu.yaml
  • .github/workflows/export-sense-voice-to-ascend-npu.yaml
  • .github/workflows/export-whisper-to-ascend-npu.yaml
  • .github/workflows/export-zipformer-ctc-to-ascend-20250703.yaml
  • scripts/whisper/ascend-npu/test_om.py
  • scripts/whisper/rknn/export_onnx.py

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist

Copy link
Copy Markdown

Summary of Changes

Hello @csukuangfj, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly expands the continuous integration capabilities by adding support for exporting Whisper speech recognition models to Huawei Ascend NPUs. It introduces a new Python script to systematically generate export configurations across different hardware and software environments, ensuring broader compatibility and streamlined model deployment. Minor adjustments were also made to existing test and export utilities to improve usability and debugging.

Highlights

  • New CI Script for Ascend NPU Export Configurations: A new Python script has been introduced to automate the generation of configurations for exporting Whisper models to Ascend NPUs, supporting various CANN and SoC versions.
  • Enhanced Ascend NPU Model Testing Script: The test_om.py script for Ascend NPU Whisper models now includes a detailed usage example in its docstring and a debug print statement within the decoder loop for better visibility during execution.
  • Improved ONNX Export Script Feedback: The export_onnx.py script for Whisper models now provides a confirmation message after successfully saving the token files, improving user feedback.
Ignored Files
  • Ignored by pattern: .github/workflows/** (4)
    • .github/workflows/export-paraformer-to-ascend-npu.yaml
    • .github/workflows/export-sense-voice-to-ascend-npu.yaml
    • .github/workflows/export-whisper-to-ascend-npu.yaml
    • .github/workflows/export-zipformer-ctc-to-ascend-20250703.yaml
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@csukuangfj
csukuangfj merged commit e061d42 into k2-fsa:master Jan 8, 2026
1 check was pending
@csukuangfj
csukuangfj deleted the ci-export-whisper-ascend-npu branch January 8, 2026 08:22

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request adds a CI workflow to export Whisper models to Ascend NPU. The changes include a new Python script for generating CI matrix configurations and modifications to related test and export scripts. The overall implementation is good, but I've identified a minor issue with a leftover debugging print statement that should be removed to ensure clean CI logs.

for t in model.sot_sequence:
token = np.array([[t]], dtype=np.int32) # sot
mask = causal_mask_1d(offset.item(), model.n_text_ctx)
print(t, model.sot_sequence, token, mask.shape, len(cross_kv), len(self_kv))

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

This print statement appears to be for debugging. It should be removed to keep the script's output clean, especially in a CI environment where excessive logging can make it difficult to find important information.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XL This PR changes 500-999 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant