Export Paraformer ASR models from FunASR to Ascend NPU 910B - #2697
Conversation
|
Caution Review failedThe pull request is closed. WalkthroughAdds a GitHub Actions workflow to export Paraformer models to Ascend NPU and several scripts to export encoder/decoder/predictor to ONNX, compile to OM, package artifacts, and run an end-to-end OM-based inference test. Changes
Sequence DiagramsequenceDiagram
autonumber
participant User as Trigger (push / manual)
participant GH as GitHub Actions
participant Container as Ascend Container
participant Scripts as Export Scripts
participant ATC as ATC Compiler
participant Pack as Packaging
participant Release as GitHub Release
User->>GH: trigger workflow
GH->>Container: run job (ubuntu, py3.8) for each framework
Container->>Scripts: run export_encoder_onnx.py / export_predictor_onnx.py / export_decoder_onnx.py
Scripts->>Scripts: produce encoder.onnx / predictor.onnx / decoder.onnx
Scripts->>ATC: invoke atc to compile .onnx -> .om (shape params)
ATC->>Pack: compiled .om files returned
Pack->>Pack: create tarball (framework-specific)
Pack->>Release: upload artifacts as Release (tag: asr-models)
Note over ATC,Pack: Success and artifacts moved to repo root for release
Estimated code review effort🎯 4 (Complex) | ⏱️ ~45 minutes Multiple new scripts with nontrivial logic (CMVN parsing, state_dict handling, monkey-patching, audio feature streaming, InferSession orchestration) plus a complex CI workflow and packaging/upload steps. Possibly related PRs
Suggested labels
Poem
Pre-merge checks and finishing touches❌ Failed checks (1 warning)
✅ Passed checks (2 passed)
📜 Recent review detailsConfiguration used: CodeRabbit UI Review profile: CHILL Plan: Pro 📒 Files selected for processing (3)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 7
🧹 Nitpick comments (8)
.github/workflows/export-paraformer-to-ascend-npu.yaml (4)
112-121: OM output filenames likely don’t match the later cp commands.You set
--output=predictor|decoder|encoder(ATC typically emitspredictor.om, …), but later copy*_linux_aarch64.om. This will fail if those suffixed names aren’t produced.Two options:
- Align outputs to the expected names:
- --output=predictor + --output=predictor_linux_aarch64 ... - --output=decoder + --output=decoder_linux_aarch64 ... - --output=encoder + --output=encoder_linux_aarch64
- Or copy the unsuffixed artifact:
- cp -v encoder_linux_aarch64.om $d/encoder.om + cp -v encoder.om $d/encoder.omPlease pick one and keep consistent across both branches.
Also applies to: 123-131, 134-142, 202-210, 213-221, 224-232, 153-155, 243-245
9-11: Typo in concurrency group name.
export-paraformer-to-ascend-nput-…→…-ascend-npu-…. Harmless but noisy in UI.
14-14: Job id is misleading.
export-paraformer-to-rknnshould read…to-ascend-nputo match purpose.
55-57: Env path note.You export an x86_64 devlib path; that’s fine for running ATC in this x86 container, but please ensure it matches the container image layout/version to avoid loader issues when tags change.
Also applies to: 109-111, 199-201
scripts/paraformer/ascend-npu/export_decoder_onnx.py (1)
14-31: LGTM; dynamic axes and I/O names align with test_om.No blockers. Consider adding
do_constant_folding=Trueto help ATC by simplifying constants.scripts/paraformer/ascend-npu/export_predictor_onnx.py (2)
9-29: Monkey‑patching the class is brittle; prefer a wrapper module.Patching
CifPredictorV2.forwardhas global side‑effects and couples ordering to__main__. Wrap instead to expose only alphas without altering the original class.Example minimal change:
- if __name__ == "__main__": - - def modified_predictor_forward(self: CifPredictorV2, hidden: torch.Tensor): - ... - CifPredictorV2.forward = modified_predictor_forward + class PredictorAlphas(torch.nn.Module): + def __init__(self, predictor: CifPredictorV2): + super().__init__() + self.p = predictor + def forward(self, hidden: torch.Tensor): + context = hidden.transpose(1, 2) + queries = self.p.pad(context) + output = torch.relu(self.p.cif_conv1d(queries)).transpose(1, 2) + alphas = torch.sigmoid(self.p.cif_output(output)) + alphas = torch.nn.functional.relu(alphas * self.p.smooth_factor - self.p.noise_threshold) + return alphas.squeeze(-1)Then export
PredictorAlphas(model.predictor)inmain(). This avoids global mutation and preserves the original predictor for other uses.
31-52: I/O contract matches test_om; seed for reproducibility is good.Once the wrapper above is applied, keep the same dynamic axes and names.
scripts/paraformer/ascend-npu/test_om.py (1)
103-135: CIF integration edge cases.If alpha sums are slightly < N due to thresholding, last residual may be dropped. Consider optional tail handling or a small epsilon to flush residual at the end.
📜 Review details
Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro
📒 Files selected for processing (6)
.github/workflows/export-paraformer-to-ascend-npu.yaml(1 hunks)scripts/paraformer/ascend-npu/export_decoder_onnx.py(1 hunks)scripts/paraformer/ascend-npu/export_encoder_onnx.py(1 hunks)scripts/paraformer/ascend-npu/export_predictor_onnx.py(1 hunks)scripts/paraformer/ascend-npu/test_om.py(1 hunks)scripts/paraformer/ascend-npu/torch_model.py(1 hunks)
🧰 Additional context used
🧬 Code graph analysis (4)
scripts/paraformer/ascend-npu/export_encoder_onnx.py (1)
scripts/paraformer/ascend-npu/torch_model.py (1)
Paraformer(1225-1281)
scripts/paraformer/ascend-npu/test_om.py (4)
sherpa-onnx/csrc/ten-vad-model.cc (1)
frame_opts(237-255)scripts/paraformer/ascend-npu/export_encoder_onnx.py (1)
main(64-85)scripts/paraformer/ascend-npu/export_decoder_onnx.py (1)
main(10-32)scripts/paraformer/ascend-npu/export_predictor_onnx.py (1)
main(32-52)
scripts/paraformer/ascend-npu/export_predictor_onnx.py (2)
scripts/paraformer/ascend-npu/export_encoder_onnx.py (1)
load_model(30-60)scripts/paraformer/ascend-npu/torch_model.py (14)
CifPredictorV2(1111-1222)forward(41-98)forward(114-118)forward(272-291)forward(324-329)forward(350-352)forward(480-508)forward(583-623)forward(654-702)forward(810-827)forward(853-855)forward(986-1032)forward(1140-1194)forward(1258-1281)
scripts/paraformer/ascend-npu/export_decoder_onnx.py (3)
scripts/paraformer/ascend-npu/export_encoder_onnx.py (2)
load_model(30-60)main(64-85)scripts/paraformer/ascend-npu/test_om.py (1)
main(137-172)scripts/paraformer/ascend-npu/export_predictor_onnx.py (1)
main(32-52)
🪛 Ruff (0.14.0)
scripts/paraformer/ascend-npu/test_om.py
131-131: Avoid specifying long messages outside the exception class
(TRY003)
🔇 Additional comments (1)
scripts/paraformer/ascend-npu/export_encoder_onnx.py (1)
63-85: Encoder export looks correct.Dynamic time axis on input/output is consistent with ATC shapes.
| ls -lh $d | ||
| tar cjfv $d.tar.bz2 $d | ||
| ls -lh *.tar.bz2 | ||
| rm -rf d |
There was a problem hiding this comment.
rm targets a literal “d” directory, not the variable.
Currently rm -rf d leaves the packaged folder behind and may bloat the workspace.
- rm -rf d
+ rm -rf "$d"Also applies to: 253-253
🤖 Prompt for AI Agents
.github/workflows/export-paraformer-to-ascend-npu.yaml lines 163 and 253: the
script currently runs rm -rf d which deletes a literal "d" directory instead of
the variable; replace each occurrence with rm -rf "$d" (or rm -rf "${d}") to
expand the shell variable safely (quote it to handle spaces) so the intended
packaged folder is removed.
| def load_cmvn(filename) -> Tuple[List[float], List[float]]: | ||
| neg_mean = None | ||
| inv_stddev = None | ||
|
|
||
| with open(filename) as f: | ||
| for line in f: | ||
| if not line.startswith("<LearnRateCoef>"): | ||
| continue | ||
| t = line.split()[3:-1] | ||
|
|
||
| if neg_mean is None: | ||
| neg_mean = list(map(lambda x: float(x), t)) | ||
| else: | ||
| inv_stddev = list(map(lambda x: float(x), t)) | ||
|
|
||
| return neg_mean, inv_stddev |
There was a problem hiding this comment.
Fragile CMVN parsing; add validation and clearer parsing.
If <LearnRateCoef> lines are missing or format changes, neg_mean/inv_stddev become None and crash later.
-def load_cmvn(filename) -> Tuple[List[float], List[float]]:
- neg_mean = None
- inv_stddev = None
- with open(filename) as f:
+def load_cmvn(filename) -> Tuple[List[float], List[float]]:
+ neg_mean = None
+ inv_stddev = None
+ with open(filename, encoding="utf-8") as f:
for line in f:
if not line.startswith("<LearnRateCoef>"):
continue
t = line.split()[3:-1]
if neg_mean is None:
neg_mean = list(map(lambda x: float(x), t))
else:
inv_stddev = list(map(lambda x: float(x), t))
- return neg_mean, inv_stddev
+ if neg_mean is None or inv_stddev is None:
+ raise ValueError(f"Failed to parse CMVN stats from {filename}")
+ return neg_mean, inv_stddev📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| def load_cmvn(filename) -> Tuple[List[float], List[float]]: | |
| neg_mean = None | |
| inv_stddev = None | |
| with open(filename) as f: | |
| for line in f: | |
| if not line.startswith("<LearnRateCoef>"): | |
| continue | |
| t = line.split()[3:-1] | |
| if neg_mean is None: | |
| neg_mean = list(map(lambda x: float(x), t)) | |
| else: | |
| inv_stddev = list(map(lambda x: float(x), t)) | |
| return neg_mean, inv_stddev | |
| def load_cmvn(filename) -> Tuple[List[float], List[float]]: | |
| neg_mean = None | |
| inv_stddev = None | |
| with open(filename, encoding="utf-8") as f: | |
| for line in f: | |
| if not line.startswith("<LearnRateCoef>"): | |
| continue | |
| t = line.split()[3:-1] | |
| if neg_mean is None: | |
| neg_mean = list(map(lambda x: float(x), t)) | |
| else: | |
| inv_stddev = list(map(lambda x: float(x), t)) | |
| if neg_mean is None or inv_stddev is None: | |
| raise ValueError(f"Failed to parse CMVN stats from {filename}") | |
| return neg_mean, inv_stddev |
| features = np.stack( | ||
| [online_fbank.get_frame(i) for i in range(online_fbank.num_frames_ready)] | ||
| ) | ||
| assert features.data.contiguous is True | ||
| assert features.dtype == np.float32, features.dtype | ||
| print("features sum", features.sum(), features.shape) | ||
|
|
||
| window_size = 7 # lfr_m | ||
| window_shift = 6 # lfr_n | ||
|
|
||
| T = (features.shape[0] - window_size) // window_shift + 1 | ||
| features = np.lib.stride_tricks.as_strided( | ||
| features, | ||
| shape=(T, features.shape[1] * window_size), | ||
| strides=((window_shift * features.shape[1]) * 4, 4), | ||
| ) | ||
| return np.copy(features) |
There was a problem hiding this comment.
as_strided math hardcodes 4‑byte stride; derive from dtype and guard short utterances.
Hardcoding 4 assumes float32 and C‑contiguity. Also, very short inputs can produce negative/zero T.
- assert features.data.contiguous is True
- assert features.dtype == np.float32, features.dtype
+ assert features.flags["C_CONTIGUOUS"]
+ assert features.dtype == np.float32, features.dtype
...
- T = (features.shape[0] - window_size) // window_shift + 1
+ total = features.shape[0]
+ if total < window_size:
+ # pad with zeros to at least one window
+ pad = window_size - total
+ features = np.pad(features, ((0, pad), (0, 0)), mode="constant")
+ total = features.shape[0]
+ T = (total - window_size) // window_shift + 1
- features = np.lib.stride_tricks.as_strided(
+ item = features.dtype.itemsize
+ features = np.lib.stride_tricks.as_strided(
features,
shape=(T, features.shape[1] * window_size),
- strides=((window_shift * features.shape[1]) * 4, 4),
+ strides=((window_shift * features.shape[1]) * item, item),
)🤖 Prompt for AI Agents
In scripts/paraformer/ascend-npu/test_om.py around lines 23-39, the as_strided
call hardcodes 4-byte strides and assumes C-contiguous float32 inputs and also
can compute a non-positive T for very short utterances; change to compute byte
strides from features.dtype.itemsize (or use features.strides to derive
per-column byte stride), ensure features is C-contiguous (or call
np.ascontiguousarray) and validate T = (features.shape[0] - window_size) //
window_shift + 1 is >= 1 before calling as_strided (handle short utterances by
returning an appropriate empty array or padding), and then use the computed byte
stride values instead of the hardcoded 4.
| def load_tokens(): | ||
| ans = dict() | ||
| i = 0 | ||
| with open("tokens.txt", encoding="utf-8") as f: | ||
| for line in f: | ||
| ans[i] = line.strip().split()[0] | ||
| i += 1 | ||
| return ans |
There was a problem hiding this comment.
Token mapping ignores IDs in tokens.txt; may decode wrong text.
Most tokens.txt are “token id”. Incrementing i assumes sorted by id. Parse the id and build by index.
-def load_tokens():
- ans = dict()
- i = 0
- with open("tokens.txt", encoding="utf-8") as f:
- for line in f:
- ans[i] = line.strip().split()[0]
- i += 1
- return ans
+def load_tokens():
+ # returns dict[id] = token
+ ans = {}
+ with open("tokens.txt", encoding="utf-8") as f:
+ for line in f:
+ parts = line.strip().split()
+ if not parts:
+ continue
+ token = parts[0]
+ idx = int(parts[1]) if len(parts) > 1 else len(ans)
+ ans[idx] = token
+ return ansUsage below then remains correct (tokens[i] by predicted id).
Also applies to: 168-172
🤖 Prompt for AI Agents
In scripts/paraformer/ascend-npu/test_om.py around lines 42 to 49 (and similarly
lines 168 to 172), the load_tokens function ignores token IDs in tokens.txt and
instead assigns tokens sequentially by i, which breaks mapping when file lines
aren’t ordered by id; update the parser to split each line into token and id (or
id and token depending on file), parse the id as an integer, and store the token
in ans[id] so the dictionary keys are true token IDs; ensure you handle empty
lines and strip whitespace, and keep usage unchanged (tokens[id] by predicted
id).
| def __init__(self): | ||
| print("init encoder") | ||
| self.encoder = InferSession(device_id=0, model_path="./encoder.om", debug=False) | ||
| self.decoder = InferSession(device_id=1, model_path="./decoder.om", debug=False) | ||
| self.predictor = InferSession( | ||
| device_id=0, model_path="./predictor.om", debug=False | ||
| ) |
There was a problem hiding this comment.
Assumes two devices; decoder on device_id=1 will fail on single‑card setups.
Default all to device 0, or read from env.
- self.decoder = InferSession(device_id=1, model_path="./decoder.om", debug=False)
+ self.decoder = InferSession(device_id=0, model_path="./decoder.om", debug=False)Or make device ids configurable via env vars.
🤖 Prompt for AI Agents
In scripts/paraformer/ascend-npu/test_om.py around lines 53 to 59, the
constructor hardcodes the decoder to device_id=1 which will fail on single‑card
setups; change this to default to device 0 or make all device IDs configurable
via environment variables. Replace the literal device_id values with either a
single default (0) for all three sessions or read per-component env vars (e.g.,
ENCODER_DEVICE_ID, DECODER_DEVICE_ID, PREDICTOR_DEVICE_ID) with a fallback to
"0" and cast to int before passing to InferSession. Ensure the code uses the
resolved device_id variables for each session so the script works on single‑card
systems and can be overridden in multi‑card environments.
| @@ -0,0 +1 @@ | |||
| ../rknn/torch_model.py No newline at end of file | |||
There was a problem hiding this comment.
Broken module: this is not valid Python (won’t import).
The file contains only a path string. Importers (from torch_model import …) will fail at runtime.
Replace with a thin re-export shim that points to ../rknn/torch_model.py:
-../rknn/torch_model.py
+import os
+import sys
+from pathlib import Path
+
+_rknn_dir = Path(__file__).parent.parent / "rknn"
+sys.path.insert(0, str(_rknn_dir.resolve()))
+
+# Re-export everything from the RKNN version
+from torch_model import * # noqa: F401,F403This keeps existing relative imports working when scripts cd into this directory, without requiring package-relative imports.
📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| ../rknn/torch_model.py | |
| import os | |
| import sys | |
| from pathlib import Path | |
| _rknn_dir = Path(__file__).parent.parent / "rknn" | |
| sys.path.insert(0, str(_rknn_dir.resolve())) | |
| # Re-export everything from the RKNN version | |
| from torch_model import * # noqa: F401,F403 |
🤖 Prompt for AI Agents
In scripts/paraformer/ascend-npu/torch_model.py (line 1), the file currently
contains only a path string and is not valid Python; replace it with a thin
re-export shim that loads the real module at ../rknn/torch_model.py and
re-exports its public names so existing imports ("from torch_model import ...")
continue to work when scripts cd into this directory; implement the shim by
resolving the absolute path to ../rknn/torch_model.py relative to __file__, use
importlib.util.spec_from_file_location and module_from_spec to load that file as
a module, execute it, then update the current module's globals() with the loaded
module's public attributes (or use setattr for __all__ if present) so callers
can import the same symbols.
Where to find the exported models
You can download them from
https://github.com/k2-fsa/sherpa-onnx/releases/tag/asr-models
sherpa-onnx-ascend-910B-paraformer-zh-2025-10-07.tar.bz2 is converted from https://www.modelscope.cn/models/iic/speech_paraformer-large_asr_nat-zh-cn-16k-common-vocab8404-pytorch
sherpa-onnx-ascend-910B-paraformer-zh-2023-03-28.tar.bz2 is converted from https://huggingface.co/ASLP-lab/WSChuan-ASR/tree/main/Paraformer-large-Chuan
Note: The model is exported using CANN 8.0 for Linux aarch64 and 910B NPU. If you have a different setting, please use the export script from us to re-export the model.
Usage
We will add C++ implementation and APIs for other programming languages in separate PRs.
At present, you can use Python APIs from Ascend to work with the provided models. We have included a script,
test_om.pyinside the model files. The following is an example.Summary by CodeRabbit