Skip to content

fix: 修复了 Ascend NPU 上下文导致的崩溃问题 - #3512

Merged
csukuangfj merged 1 commit into
k2-fsa:masterfrom
cn-tao:fix-ascend-crash
Apr 15, 2026
Merged

csukuangfj merged 1 commit into
k2-fsa:masterfrom
cn-tao:fix-ascend-crash

Conversation

@cn-tao

@cn-tao cn-tao commented Apr 14, 2026 •

Copy link
Copy Markdown
Contributor

鲲鹏aarch64 CPU,Ascend 910A NPU卡,CANN 8.2.RC1华为官方容器。使用sherpa-onnx-offline工作正常。sherpa-onnx-offline-websocket-server启动服务正常,使用python客户程序连接时服务崩溃。

root@localhost:/opt/sherpa-onnx/build# ./bin/sherpa-onnx-offline-websocket-server --port=6006 --num-work-threads=5 --provider=ascend --paraformer="/data/paraformer/sherpa-onnx-ascend-910B-cann-8.2-paraformer-zh-2025-10-07/encoder.om,/data/paraformer/sherpa-onnx-ascend-910B-cann-8.2-paraformer-zh-2025-10-07/predictor.om,/data/paraformer/sherpa-onnx-ascend-910B-cann-8.2-paraformer-zh-2025-10-07/decoder.om" --tokens=/data/paraformer/sherpa-onnx-ascend-910B-cann-8.2-paraformer-zh-2025-10-07/tokens.txt --log-file=./log.txt --max-batch-size=5
onnxruntime cpuid_info warning: Unknown CPU vendor. cpuinfo_vendor value: 15
/opt/sherpa-onnx/sherpa-onnx/csrc/parse-options.cc:Read:373 ./bin/sherpa-onnx-offline-websocket-server --port=6006 --num-work-threads=5 --provider=ascend --paraformer=/data/paraformer/sherpa-onnx-ascend-910B-cann-8.2-paraformer-zh-2025-10-07/encoder.om,/data/paraformer/sherpa-onnx-ascend-910B-cann-8.2-paraformer-zh-2025-10-07/predictor.om,/data/paraformer/sherpa-onnx-ascend-910B-cann-8.2-paraformer-zh-2025-10-07/decoder.om --tokens=/data/paraformer/sherpa-onnx-ascend-910B-cann-8.2-paraformer-zh-2025-10-07/tokens.txt --log-file=./log.txt --max-batch-size=5

/opt/sherpa-onnx/sherpa-onnx/csrc/offline-websocket-server.cc:main:93 Started!
/opt/sherpa-onnx/sherpa-onnx/csrc/offline-websocket-server.cc:main:94 Listening on: 6006
/opt/sherpa-onnx/sherpa-onnx/csrc/offline-websocket-server.cc:main:95 Number of work threads: 5
[2026-04-13 11:32:21] [connect] WebSocket Connection 192.168.1.41:51694 v13 "Python/3.9 websockets/15.0.1" / 101
/opt/sherpa-onnx/sherpa-onnx/csrc/offline-websocket-server-impl.cc:OnOpen:170 Number of active connections: 1
/opt/sherpa-onnx/sherpa-onnx/csrc/offline-websocket-server-impl.cc:Decode:67 size: 1
/opt/sherpa-onnx/sherpa-onnx/csrc/ascend/offline-paraformer-model-ascend.cc:RunEncoder:149 Return code is: 107002
/opt/sherpa-onnx/sherpa-onnx/csrc/ascend/offline-paraformer-model-ascend.cc:RunEncoder:149 Error message: EE1001: [PID: 118] 2026-04-13-11:32:21.618.180 The argument is invalid.Reason: rtMemcpy execute failed, reason=[context pointer null]
Solution: 1.Check the input parameter range of the function. 2.Check the function invocation relationship.
TraceBack (most recent call last):
ctx is NULL![FUNC:MemCopySync][FILE:api_impl.cc][LINE:2555]
The argument is invalid.Reason: rtMemcpy execute failed, reason=[context pointer null]
synchronized memcpy failed, kind = 1, runtime result = 107002[FUNC:ReportCallError][FILE:log_inner.cpp][LINE:161]
ctx is NULL![FUNC:GetDevErrMsg][FILE:api_impl.cc][LINE:6147]
The argument is invalid.Reason: rtGetDevMsg execute failed, reason=[context pointer null]

/opt/sherpa-onnx/sherpa-onnx/csrc/ascend/offline-paraformer-model-ascend.cc:RunEncoder:149 Failed to call aclrtMemcpy

错误原因:Ascend NPU 的上下文(aclrtContext)与线程绑定。当多工作线程并发调用模型时,aclrtMemcpy 和 aclmdlExecute调用时当前线程没有有效的上下文。
解决方案:在sherpa-onnx/sherpa-onnx/csrc/ascend/offline-*-model-ascend.cc 4个文件的 std::lock_guardstd::mutex lock(mutex_); 行之后加上:
aclError ret_set_ctx = aclrtSetCurrentContext(*context_);
SHERPA_ONNX_ASCEND_CHECK(ret_set_ctx, "Failed to call aclrtSetCurrentContext");
经测试在昇腾910A NPU 上正常工作。

Summary by CodeRabbit

  • Bug Fixes
    • Enhanced reliability of Ascend hardware-accelerated inference across Paraformer, SenseVoice, Whisper, and Zipformer CTC models by strengthening execution context management during runtime operations.

@dosubot dosubot Bot added the size:S This PR changes 10-29 lines, ignoring generated files. label Apr 14, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request updates several Ascend-based offline model implementations, including Paraformer, SenseVoice, Whisper, and Zipformer CTC, to explicitly set the current ACL context using aclrtSetCurrentContext before processing. These additions are placed within existing mutex-protected blocks and include error handling via the SHERPA_ONNX_ASCEND_CHECK macro. I have no feedback to provide as there were no review comments to evaluate.

@coderabbitai

coderabbitai Bot commented Apr 14, 2026 •

Copy link
Copy Markdown

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 6e1dc838-f04f-4fd3-8697-617930a95918

📥 Commits

Reviewing files that changed from the base of the PR and between 2cceb3b and f11f15e.

📒 Files selected for processing (4)
  • sherpa-onnx/csrc/ascend/offline-paraformer-model-ascend.cc
  • sherpa-onnx/csrc/ascend/offline-sense-voice-model-ascend.cc
  • sherpa-onnx/csrc/ascend/offline-whisper-model-ascend.cc
  • sherpa-onnx/csrc/ascend/offline-zipformer-ctc-model-ascend.cc

📝 Walkthrough

Walkthrough

Added explicit Ascend context-setting calls to four offline model implementations. Each Run() method now invokes aclrtSetCurrentContext(*context_) within the existing mutex lock and checks for errors using SHERPA_ONNX_ASCEND_CHECK before proceeding with model execution.

Changes

Cohort / File(s) Summary
Ascend Context Initialization in Run Methods
sherpa-onnx/csrc/ascend/offline-paraformer-model-ascend.cc, sherpa-onnx/csrc/ascend/offline-sense-voice-model-ascend.cc, sherpa-onnx/csrc/ascend/offline-whisper-model-ascend.cc, sherpa-onnx/csrc/ascend/offline-zipformer-ctc-model-ascend.cc
Added per-call aclrtSetCurrentContext() invocations at the start of each Run() method with corresponding error checking via SHERPA_ONNX_ASCEND_CHECK. Ensures the correct Ascend context is active before device operations.

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~10 minutes

Possibly related PRs

Suggested labels

size:M

Poem

🐰 Context switching, swift and sure,
Each run now guards with Ascend's cure,
Four models strengthened, threads aligned,
Device operations, safely confined!

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately describes the main change: fixing Ascend NPU context crashes by explicitly setting the current context in the Run methods across four model implementation files.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@csukuangfj csukuangfj left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for your contribution!

@csukuangfj
csukuangfj merged commit dadac36 into k2-fsa:master Apr 15, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:S This PR changes 10-29 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants