Conversation
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. Tip 💡 Consider Linking a Related Issue or RFCYour PR title contains the [BugFix] tag, indicating a bug fix or new feature. Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:
🙏 Thanks for helping us keep the project well-organized! |
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
|
Warning Gemini encountered an error creating the summary. You can try again by commenting |
triton ascen extension 中部分 op (如 insert_slice) 在当前环境中 不可用,_resolve_triton_ascend_op 抛异常会导致整个模块加载失败。 改为返回 None,使非 triton 路径可正常工作。 Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> Signed-off-by: Jinge-Ma <majinge1762401198@foxmail.com>
- 21 个 A 类测试: patch_routed_experts_capturer, apply capture 逻辑, multistream capture 逻辑, ModelRunner 集成, Dense 模型兼容性 - 3 个 B 类 NPU 测试: RoutedExpertsCapturer 核心逻辑 (capture 写入 buffer, clear_buffer 清零, save_captured_experts device->host) - 全部 24 个测试通过 Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> Signed-off-by: Jinge-Ma <majinge1762401198@foxmail.com>
- Remove unused variable topk_ids_before in test_capture_post_allgather. - Rename ambiguous variable 'l' to 'layer' in test_has_fused_moe_no_layers. - Apply ruff format. Signed-off-by: Jinge-Ma <majinge1762401198@foxmail.com>
The module is part of vllm-ascend but mypy with --follow-imports skip treats it as third-party. Add type: ignore[import-untyped] to the three in-test imports. Signed-off-by: Jinge-Ma <majinge1762401198@foxmail.com>
Signed-off-by: Jinge-Ma <majinge1762401198@foxmail.com>
|
Thanks for the review patience. After rebasing onto current What I found during conflict resolution:
So both pieces of this PR are already covered upstream in a more comprehensive way, and the test file targets a structure that no longer exists. Closing as superseded. Will follow up if any gap in |
What this PR does / why we need it
--enable-return-routed-expertsis partially supported on vllm-ascend, butthe
multistream_overlap_gate=Truebranch ofAscendFusedMoE.forward_impldoes not call
capturer.capture(), causing routed_experts data to besilently lost for MoE models with shared experts (which commonly enable
multistream overlap to overlap shared experts with routed experts).
This PR:
capturer.capture(layer_id, topk_ids)call in themultistream path of
AscendFusedMoE.forward_impl, mirroring thestandard path's parameters.
triton_utilsreturnNoneon parse failure instead of raisingan exception — graceful degradation when triton is installed but has no
active driver on Ascend (common in container environments without
/dev/davinci*).Does this PR introduce any user-facing change?
No.
--enable-return-routed-expertsnow correctly returns routed expertsdata for both standard and multistream paths.
How was this patch tested?
Unit tests (
tests/ut/ops/test_routed_experts_capture.py, A+B classes):RoutedExpertsCapturer, run both paths, assertcapture()called withcorrect
layer_idandtopk_ids).All A-class tests pass on CPU-only environments.
End-to-end on Ascend 910B2C + Qwen1.5-MoE-A2.7B-Chat:
routed_expertsreturnsnp.ndarray(num_tokens, 24, 4)[0, 60)Performance (warmup 2 + bench 5):
RE adds -1.7% to -6.8% overhead from per-layer
topk_idswrites to devicebuffer, which is amortized at high concurrency. Notably, RE stabilizes
throughput at high concurrency (std drops from 94.58 to 6.75 at 16 req)
because the capture synchronization reduces non-determinism.
Use case
Required for RLHF scenarios that need routing replay (e.g.,
GRPO/RLHF training that re-executes forward with the same expert routing
indices sampled during generation).
Files changed
Checklist
Signed-off-by: Jinge-Ma majinge1762401198@foxmail.com