Skip to content

Export Paraformer ASR models from FunASR to Ascend NPU 910B - #2697

Merged
csukuangfj merged 21 commits into
k2-fsa:masterfrom
csukuangfj:ascend-npu
Oct 17, 2025
Merged

csukuangfj merged 21 commits into
k2-fsa:masterfrom
csukuangfj:ascend-npu

Conversation

@csukuangfj

@csukuangfj csukuangfj commented Oct 17, 2025 •

Copy link
Copy Markdown
Collaborator

Where to find the exported models

You can download them from
https://github.com/k2-fsa/sherpa-onnx/releases/tag/asr-models

Screenshot 2025-10-18 at 00 32 50

Note: The model is exported using CANN 8.0 for Linux aarch64 and 910B NPU. If you have a different setting, please use the export script from us to re-export the model.

Usage

We will add C++ implementation and APIs for other programming languages in separate PRs.

At present, you can use Python APIs from Ascend to work with the provided models. We have included a script, test_om.py inside the model files. The following is an example.

Screenshot 2025-10-18 at 00 22 01

Summary by CodeRabbit

  • New Features
    • Export Paraformer ASR models to Ascend NPU (ONNX → compiled bundles) and provide end-to-end test tooling to run compiled components.
  • Bug Fixes
    • Fixed workflow cleanup commands to correctly remove intended temporary directories.
  • Chores
    • Added GitHub Actions workflows to automate model export, compilation, packaging, and release uploads for Ascend and RKNN targets.

@dosubot dosubot Bot added the size:XL This PR changes 500-999 lines, ignoring generated files. label Oct 17, 2025
@coderabbitai

coderabbitai Bot commented Oct 17, 2025 •

Copy link
Copy Markdown

Caution

Review failed

The pull request is closed.

Walkthrough

Adds a GitHub Actions workflow to export Paraformer models to Ascend NPU and several scripts to export encoder/decoder/predictor to ONNX, compile to OM, package artifacts, and run an end-to-end OM-based inference test.

Changes

Cohort / File(s) Summary
GitHub Actions workflow
\.github/workflows/export-paraformer-to-ascend-npu.yaml
New workflow triggered on pushes to ascend-npu and workflow_dispatch; matrix over frameworks; checks out, sets up Python, verifies Ascend environment, installs deps, exports ONNX, runs ATC to produce .om, packages tarballs, and creates Release uploads for two repo owners.
ONNX export — encoder
scripts/paraformer/ascend-npu/export_encoder_onnx.py
New script: parses CMVN (load_cmvn), loads model and state_dict (load_model), exports model.encoder to encoder.onnx with opset 14 and dynamic time axis.
ONNX export — decoder
scripts/paraformer/ascend-npu/export_decoder_onnx.py
New script: loads model (via encoder loader), constructs dummy encoder_out and acoustic_embedding, exports model.decoder to decoder.onnx with opset 14 and dynamic axes for sequence lengths.
ONNX export — predictor (CIF)
scripts/paraformer/ascend-npu/export_predictor_onnx.py
New script: monkey-patches CifPredictorV2.forward to a modified implementation, exports model.predictor to predictor.onnx with opset 14 and dynamic time axis.
OM inference & utilities
scripts/paraformer/ascend-npu/test_om.py
New end-to-end inference script using InferSession for encoder/predictor/decoder OM models: feature computation, streaming window via as_strided, token loading, acoustic embedding accumulation, and decoding to text.
Torch model shim
scripts/paraformer/ascend-npu/torch_model.py
New file referencing ../rknn/torch_model.py (single-line import/redirect).
Workflow fixes in RKNN exports
.github/workflows/export-paraformer-to-rknn.yaml, .github/workflows/export-sense-voice-to-rknn.yaml
Replaced rm -rf d with rm -rf $d in cleanup steps to remove directory referenced by variable d instead of a literal d.

Sequence Diagram

sequenceDiagram
    autonumber
    participant User as Trigger (push / manual)
    participant GH as GitHub Actions
    participant Container as Ascend Container
    participant Scripts as Export Scripts
    participant ATC as ATC Compiler
    participant Pack as Packaging
    participant Release as GitHub Release

    User->>GH: trigger workflow
    GH->>Container: run job (ubuntu, py3.8) for each framework
    Container->>Scripts: run export_encoder_onnx.py / export_predictor_onnx.py / export_decoder_onnx.py
    Scripts->>Scripts: produce encoder.onnx / predictor.onnx / decoder.onnx
    Scripts->>ATC: invoke atc to compile .onnx -> .om (shape params)
    ATC->>Pack: compiled .om files returned
    Pack->>Pack: create tarball (framework-specific)
    Pack->>Release: upload artifacts as Release (tag: asr-models)
    Note over ATC,Pack: Success and artifacts moved to repo root for release
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Multiple new scripts with nontrivial logic (CMVN parsing, state_dict handling, monkey-patching, audio feature streaming, InferSession orchestration) plus a complex CI workflow and packaging/upload steps.

Possibly related PRs

Suggested labels

size:L

Poem

🐰 I hopped through ONNX paths tonight,
Exported encoders in the moonlight,
Predictor patched, decoder spun,
ATC baked .om and tarballs done,
A tiny rabbit cheers: export — well done! 🎩✨

Pre-merge checks and finishing touches

❌ Failed checks (1 warning)
Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 7.14% which is insufficient. The required threshold is 80.00%. You can run @coderabbitai generate docstrings to improve docstring coverage.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title Check ✅ Passed The PR title "Export Paraformer ASR models to Ascend NPU 910B" directly and accurately summarizes the main objective of the changeset. The changes include a GitHub Actions workflow to export Paraformer models to ONNX format, multiple Python scripts for exporting encoder/decoder/predictor components, and a test script for validating the exported models on Ascend NPU 910B hardware. The title is concise, specific, and uses clear language without vague terms or noise. A teammate reviewing the git history would immediately understand that this PR adds export functionality for Paraformer ASR models targeting Ascend NPU 910B hardware.

📜 Recent review details

Configuration used: CodeRabbit UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 1e3f4ff and 59804f8.

📒 Files selected for processing (3)
  • .github/workflows/export-paraformer-to-ascend-npu.yaml (1 hunks)
  • .github/workflows/export-paraformer-to-rknn.yaml (2 hunks)
  • .github/workflows/export-sense-voice-to-rknn.yaml (2 hunks)

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 7

🧹 Nitpick comments (8)
.github/workflows/export-paraformer-to-ascend-npu.yaml (4)

112-121: OM output filenames likely don’t match the later cp commands.

You set --output=predictor|decoder|encoder (ATC typically emits predictor.om, …), but later copy *_linux_aarch64.om. This will fail if those suffixed names aren’t produced.

Two options:

  • Align outputs to the expected names:
- --output=predictor
+ --output=predictor_linux_aarch64
...
- --output=decoder
+ --output=decoder_linux_aarch64
...
- --output=encoder
+ --output=encoder_linux_aarch64
  • Or copy the unsuffixed artifact:
- cp -v encoder_linux_aarch64.om $d/encoder.om
+ cp -v encoder.om $d/encoder.om

Please pick one and keep consistent across both branches.

Also applies to: 123-131, 134-142, 202-210, 213-221, 224-232, 153-155, 243-245


9-11: Typo in concurrency group name.

export-paraformer-to-ascend-nput-… → …-ascend-npu-…. Harmless but noisy in UI.


14-14: Job id is misleading.

export-paraformer-to-rknn should read …to-ascend-npu to match purpose.


55-57: Env path note.

You export an x86_64 devlib path; that’s fine for running ATC in this x86 container, but please ensure it matches the container image layout/version to avoid loader issues when tags change.

Also applies to: 109-111, 199-201

scripts/paraformer/ascend-npu/export_decoder_onnx.py (1)

14-31: LGTM; dynamic axes and I/O names align with test_om.

No blockers. Consider adding do_constant_folding=True to help ATC by simplifying constants.

scripts/paraformer/ascend-npu/export_predictor_onnx.py (2)

9-29: Monkey‑patching the class is brittle; prefer a wrapper module.

Patching CifPredictorV2.forward has global side‑effects and couples ordering to __main__. Wrap instead to expose only alphas without altering the original class.

Example minimal change:

- if __name__ == "__main__":
- 
-     def modified_predictor_forward(self: CifPredictorV2, hidden: torch.Tensor):
-         ...
-     CifPredictorV2.forward = modified_predictor_forward
+ class PredictorAlphas(torch.nn.Module):
+     def __init__(self, predictor: CifPredictorV2):
+         super().__init__()
+         self.p = predictor
+     def forward(self, hidden: torch.Tensor):
+         context = hidden.transpose(1, 2)
+         queries = self.p.pad(context)
+         output = torch.relu(self.p.cif_conv1d(queries)).transpose(1, 2)
+         alphas = torch.sigmoid(self.p.cif_output(output))
+         alphas = torch.nn.functional.relu(alphas * self.p.smooth_factor - self.p.noise_threshold)
+         return alphas.squeeze(-1)

Then export PredictorAlphas(model.predictor) in main(). This avoids global mutation and preserves the original predictor for other uses.


31-52: I/O contract matches test_om; seed for reproducibility is good.

Once the wrapper above is applied, keep the same dynamic axes and names.

scripts/paraformer/ascend-npu/test_om.py (1)

103-135: CIF integration edge cases.

If alpha sums are slightly < N due to thresholding, last residual may be dropped. Consider optional tail handling or a small epsilon to flush residual at the end.

📜 Review details

Configuration used: CodeRabbit UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between f20ebb4 and 1e3f4ff.

📒 Files selected for processing (6)
  • .github/workflows/export-paraformer-to-ascend-npu.yaml (1 hunks)
  • scripts/paraformer/ascend-npu/export_decoder_onnx.py (1 hunks)
  • scripts/paraformer/ascend-npu/export_encoder_onnx.py (1 hunks)
  • scripts/paraformer/ascend-npu/export_predictor_onnx.py (1 hunks)
  • scripts/paraformer/ascend-npu/test_om.py (1 hunks)
  • scripts/paraformer/ascend-npu/torch_model.py (1 hunks)
🧰 Additional context used
🧬 Code graph analysis (4)
scripts/paraformer/ascend-npu/export_encoder_onnx.py (1)
scripts/paraformer/ascend-npu/torch_model.py (1)
  • Paraformer (1225-1281)
scripts/paraformer/ascend-npu/test_om.py (4)
sherpa-onnx/csrc/ten-vad-model.cc (1)
  • frame_opts (237-255)
scripts/paraformer/ascend-npu/export_encoder_onnx.py (1)
  • main (64-85)
scripts/paraformer/ascend-npu/export_decoder_onnx.py (1)
  • main (10-32)
scripts/paraformer/ascend-npu/export_predictor_onnx.py (1)
  • main (32-52)
scripts/paraformer/ascend-npu/export_predictor_onnx.py (2)
scripts/paraformer/ascend-npu/export_encoder_onnx.py (1)
  • load_model (30-60)
scripts/paraformer/ascend-npu/torch_model.py (14)
  • CifPredictorV2 (1111-1222)
  • forward (41-98)
  • forward (114-118)
  • forward (272-291)
  • forward (324-329)
  • forward (350-352)
  • forward (480-508)
  • forward (583-623)
  • forward (654-702)
  • forward (810-827)
  • forward (853-855)
  • forward (986-1032)
  • forward (1140-1194)
  • forward (1258-1281)
scripts/paraformer/ascend-npu/export_decoder_onnx.py (3)
scripts/paraformer/ascend-npu/export_encoder_onnx.py (2)
  • load_model (30-60)
  • main (64-85)
scripts/paraformer/ascend-npu/test_om.py (1)
  • main (137-172)
scripts/paraformer/ascend-npu/export_predictor_onnx.py (1)
  • main (32-52)
🪛 Ruff (0.14.0)
scripts/paraformer/ascend-npu/test_om.py

131-131: Avoid specifying long messages outside the exception class

(TRY003)

🔇 Additional comments (1)
scripts/paraformer/ascend-npu/export_encoder_onnx.py (1)

63-85: Encoder export looks correct.

Dynamic time axis on input/output is consistent with ATC shapes.

Comment thread .github/workflows/export-paraformer-to-ascend-npu.yaml
ls -lh $d
tar cjfv $d.tar.bz2 $d
ls -lh *.tar.bz2
rm -rf d

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

rm targets a literal “d” directory, not the variable.

Currently rm -rf d leaves the packaged folder behind and may bloat the workspace.

- rm -rf d
+ rm -rf "$d"

Also applies to: 253-253

🤖 Prompt for AI Agents
.github/workflows/export-paraformer-to-ascend-npu.yaml lines 163 and 253: the
script currently runs rm -rf d which deletes a literal "d" directory instead of
the variable; replace each occurrence with rm -rf "$d" (or rm -rf "${d}") to
expand the shell variable safely (quote it to handle spaces) so the intended
packaged folder is removed.

Comment on lines +12 to +27
def load_cmvn(filename) -> Tuple[List[float], List[float]]:
neg_mean = None
inv_stddev = None

with open(filename) as f:
for line in f:
if not line.startswith("<LearnRateCoef>"):
continue
t = line.split()[3:-1]

if neg_mean is None:
neg_mean = list(map(lambda x: float(x), t))
else:
inv_stddev = list(map(lambda x: float(x), t))

return neg_mean, inv_stddev

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Fragile CMVN parsing; add validation and clearer parsing.

If <LearnRateCoef> lines are missing or format changes, neg_mean/inv_stddev become None and crash later.

-def load_cmvn(filename) -> Tuple[List[float], List[float]]:
-    neg_mean = None
-    inv_stddev = None
-    with open(filename) as f:
+def load_cmvn(filename) -> Tuple[List[float], List[float]]:
+    neg_mean = None
+    inv_stddev = None
+    with open(filename, encoding="utf-8") as f:
         for line in f:
             if not line.startswith("<LearnRateCoef>"):
                 continue
             t = line.split()[3:-1]
             if neg_mean is None:
                 neg_mean = list(map(lambda x: float(x), t))
             else:
                 inv_stddev = list(map(lambda x: float(x), t))
-    return neg_mean, inv_stddev
+    if neg_mean is None or inv_stddev is None:
+        raise ValueError(f"Failed to parse CMVN stats from {filename}")
+    return neg_mean, inv_stddev
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
def load_cmvn(filename) -> Tuple[List[float], List[float]]:
neg_mean = None
inv_stddev = None
with open(filename) as f:
for line in f:
if not line.startswith("<LearnRateCoef>"):
continue
t = line.split()[3:-1]
if neg_mean is None:
neg_mean = list(map(lambda x: float(x), t))
else:
inv_stddev = list(map(lambda x: float(x), t))
return neg_mean, inv_stddev
def load_cmvn(filename) -> Tuple[List[float], List[float]]:
neg_mean = None
inv_stddev = None
with open(filename, encoding="utf-8") as f:
for line in f:
if not line.startswith("<LearnRateCoef>"):
continue
t = line.split()[3:-1]
if neg_mean is None:
neg_mean = list(map(lambda x: float(x), t))
else:
inv_stddev = list(map(lambda x: float(x), t))
if neg_mean is None or inv_stddev is None:
raise ValueError(f"Failed to parse CMVN stats from {filename}")
return neg_mean, inv_stddev

Comment on lines +23 to +39
features = np.stack(
[online_fbank.get_frame(i) for i in range(online_fbank.num_frames_ready)]
)
assert features.data.contiguous is True
assert features.dtype == np.float32, features.dtype
print("features sum", features.sum(), features.shape)

window_size = 7 # lfr_m
window_shift = 6 # lfr_n

T = (features.shape[0] - window_size) // window_shift + 1
features = np.lib.stride_tricks.as_strided(
features,
shape=(T, features.shape[1] * window_size),
strides=((window_shift * features.shape[1]) * 4, 4),
)
return np.copy(features)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

as_strided math hardcodes 4‑byte stride; derive from dtype and guard short utterances.

Hardcoding 4 assumes float32 and C‑contiguity. Also, very short inputs can produce negative/zero T.

-    assert features.data.contiguous is True
-    assert features.dtype == np.float32, features.dtype
+    assert features.flags["C_CONTIGUOUS"]
+    assert features.dtype == np.float32, features.dtype
...
-    T = (features.shape[0] - window_size) // window_shift + 1
+    total = features.shape[0]
+    if total < window_size:
+        # pad with zeros to at least one window
+        pad = window_size - total
+        features = np.pad(features, ((0, pad), (0, 0)), mode="constant")
+        total = features.shape[0]
+    T = (total - window_size) // window_shift + 1
-    features = np.lib.stride_tricks.as_strided(
+    item = features.dtype.itemsize
+    features = np.lib.stride_tricks.as_strided(
         features,
         shape=(T, features.shape[1] * window_size),
-        strides=((window_shift * features.shape[1]) * 4, 4),
+        strides=((window_shift * features.shape[1]) * item, item),
     )
🤖 Prompt for AI Agents
In scripts/paraformer/ascend-npu/test_om.py around lines 23-39, the as_strided
call hardcodes 4-byte strides and assumes C-contiguous float32 inputs and also
can compute a non-positive T for very short utterances; change to compute byte
strides from features.dtype.itemsize (or use features.strides to derive
per-column byte stride), ensure features is C-contiguous (or call
np.ascontiguousarray) and validate T = (features.shape[0] - window_size) //
window_shift + 1 is >= 1 before calling as_strided (handle short utterances by
returning an appropriate empty array or padding), and then use the computed byte
stride values instead of the hardcoded 4.

Comment on lines +42 to +49
def load_tokens():
ans = dict()
i = 0
with open("tokens.txt", encoding="utf-8") as f:
for line in f:
ans[i] = line.strip().split()[0]
i += 1
return ans

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Token mapping ignores IDs in tokens.txt; may decode wrong text.

Most tokens.txt are “token id”. Incrementing i assumes sorted by id. Parse the id and build by index.

-def load_tokens():
-    ans = dict()
-    i = 0
-    with open("tokens.txt", encoding="utf-8") as f:
-        for line in f:
-            ans[i] = line.strip().split()[0]
-            i += 1
-    return ans
+def load_tokens():
+    # returns dict[id] = token
+    ans = {}
+    with open("tokens.txt", encoding="utf-8") as f:
+        for line in f:
+            parts = line.strip().split()
+            if not parts:
+                continue
+            token = parts[0]
+            idx = int(parts[1]) if len(parts) > 1 else len(ans)
+            ans[idx] = token
+    return ans

Usage below then remains correct (tokens[i] by predicted id).

Also applies to: 168-172

🤖 Prompt for AI Agents
In scripts/paraformer/ascend-npu/test_om.py around lines 42 to 49 (and similarly
lines 168 to 172), the load_tokens function ignores token IDs in tokens.txt and
instead assigns tokens sequentially by i, which breaks mapping when file lines
aren’t ordered by id; update the parser to split each line into token and id (or
id and token depending on file), parse the id as an integer, and store the token
in ans[id] so the dictionary keys are true token IDs; ensure you handle empty
lines and strip whitespace, and keep usage unchanged (tokens[id] by predicted
id).

Comment on lines +53 to +59
def __init__(self):
print("init encoder")
self.encoder = InferSession(device_id=0, model_path="./encoder.om", debug=False)
self.decoder = InferSession(device_id=1, model_path="./decoder.om", debug=False)
self.predictor = InferSession(
device_id=0, model_path="./predictor.om", debug=False
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Assumes two devices; decoder on device_id=1 will fail on single‑card setups.

Default all to device 0, or read from env.

-        self.decoder = InferSession(device_id=1, model_path="./decoder.om", debug=False)
+        self.decoder = InferSession(device_id=0, model_path="./decoder.om", debug=False)

Or make device ids configurable via env vars.

🤖 Prompt for AI Agents
In scripts/paraformer/ascend-npu/test_om.py around lines 53 to 59, the
constructor hardcodes the decoder to device_id=1 which will fail on single‑card
setups; change this to default to device 0 or make all device IDs configurable
via environment variables. Replace the literal device_id values with either a
single default (0) for all three sessions or read per-component env vars (e.g.,
ENCODER_DEVICE_ID, DECODER_DEVICE_ID, PREDICTOR_DEVICE_ID) with a fallback to
"0" and cast to int before passing to InferSession. Ensure the code uses the
resolved device_id variables for each session so the script works on single‑card
systems and can be overridden in multi‑card environments.

@@ -0,0 +1 @@
../rknn/torch_model.py No newline at end of file

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🔴 Critical

Broken module: this is not valid Python (won’t import).

The file contains only a path string. Importers (from torch_model import …) will fail at runtime.

Replace with a thin re-export shim that points to ../rknn/torch_model.py:

-../rknn/torch_model.py
+import os
+import sys
+from pathlib import Path
+
+_rknn_dir = Path(__file__).parent.parent / "rknn"
+sys.path.insert(0, str(_rknn_dir.resolve()))
+
+# Re-export everything from the RKNN version
+from torch_model import *  # noqa: F401,F403

This keeps existing relative imports working when scripts cd into this directory, without requiring package-relative imports.

📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
../rknn/torch_model.py
import os
import sys
from pathlib import Path
_rknn_dir = Path(__file__).parent.parent / "rknn"
sys.path.insert(0, str(_rknn_dir.resolve()))
# Re-export everything from the RKNN version
from torch_model import * # noqa: F401,F403
🤖 Prompt for AI Agents
In scripts/paraformer/ascend-npu/torch_model.py (line 1), the file currently
contains only a path string and is not valid Python; replace it with a thin
re-export shim that loads the real module at ../rknn/torch_model.py and
re-exports its public names so existing imports ("from torch_model import ...")
continue to work when scripts cd into this directory; implement the shim by
resolving the absolute path to ../rknn/torch_model.py relative to __file__, use
importlib.util.spec_from_file_location and module_from_spec to load that file as
a module, execute it, then update the current module's globals() with the loaded
module's public attributes (or use setattr for __all__ if present) so callers
can import the same symbols.

@csukuangfj
csukuangfj merged commit 027871d into k2-fsa:master Oct 17, 2025
1 check was pending
@csukuangfj
csukuangfj deleted the ascend-npu branch October 17, 2025 16:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XL This PR changes 500-999 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant