[Bugfix][Revert][Attention] Revert AscendC MsaIndexScore MiniMax-M3 path - #17020
Conversation
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request reverts the recent addition of the AscendC-based index score decoding path for MiniMax-M3. The revert was necessary because the custom operator failed to register correctly in the A5 nightly validation environment, causing all precision test cases to fail. The system now falls back to the previous implementation, ensuring stability while the underlying build and registration issues are addressed. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. |
There was a problem hiding this comment.
Code Review
Suggested PR Title:\n\nmarkdown\n[Attention][Misc] Simplify MsaIndexScore tiling and remove replicated decode logic\n\n\nSuggested PR Summary:\n\nmarkdown\n### What this PR does / why we need it?\nThis PR simplifies the `MsaIndexScore` operator by removing the complex `kvChunks` and `launchAic` estimation logic from the host tiling, and removing the replicated decode logic (`minimax_m3_index_decode_replicated`) for A5. It cleans up the `MsaTask` structure and simplifies the epilogue logic for both arch22 and arch35.\n\nHowever, a critical bug was introduced in the arch22 epilogue (`msa_index_score_epilogue.h`) where the removal of `OpenStageWindow` and `AdvanceStageWindow` leaves `stageBlkEnd_` uninitialized, causing `FlushStage()` to always return early and break attention score outputs on Atlas A2/A3 during decode.\n\n### Does this PR introduce _any_ user-facing change?\nNo user-facing API changes are introduced, but it simplifies internal scheduling and removes unused replicated decode paths.\n\n### How was this patch tested?\nThe PR adds a comprehensive end-to-end accuracy self-check test suite in `test_aclnn_msa_index_score.cpp` covering 40 test cases on Ascend 950 and 36 cases on A2/A3.\n
e8787ad to
cdad1ba
Compare
|
/rerun Rerun (failed jobs only):
|
Revert the AscendC MsaIndexScore MiniMax-M3 path until the custom operator build and registration issue is resolved. Signed-off-by: yingyingzizi <1084697284@qq.com>
cdad1ba to
57c08a1
Compare
…ath (vllm-project#17020) ### What this PR does / why we need it? Revert PR vllm-project#16426, `[Performance] Enable AscendC index score decode for MiniMax-M3`. The A5 nightly validation cannot currently register or execute `npu_msa_index_score`. The latest A5 run reached pytest and all 24 precision cases failed because `torch.ops._C_ascend.npu_msa_index_score` was unavailable. Reverting the change removes the failing AscendC MiniMax-M3 index-score path until the custom-operator build and registration issue is resolved. This PR reverts merge commit `fefa33767507514577b80affaa90b5b77ce8ac8a` and preserves unrelated changes already present on `main`, including subsequent nightly matrix maintenance. ### Does this PR introduce _any_ user-facing change? Yes. MiniMax-M3 index-score execution returns to the pre-vllm-project#16426 implementation. The AscendC MsaIndexScore optimization and its associated source, documentation, integration changes, and precision test are removed. ### How was this patch tested? - Reverted merge commit `fefa33767507514577b80affaa90b5b77ce8ac8a` with parent `-m 1`. - Resolved conflicts by preserving current `main` workflow changes and removing only the vllm-project#16426 changes. - `git diff --check` passed. - The revert commit includes a `Signed-off-by` line. - A5 nightly failure: https://github.com/vllm-project/vllm-ascend/actions/runs/35500139842/job/106055635575 Related: vllm-project#16426, vllm-project#16955, vllm-project#17013. - vLLM main: vllm-project/vllm@84030bb Signed-off-by: yingyingzizi <1084697284@qq.com>
…ath (vllm-project#17020) ### What this PR does / why we need it? Revert PR vllm-project#16426, `[Performance] Enable AscendC index score decode for MiniMax-M3`. The A5 nightly validation cannot currently register or execute `npu_msa_index_score`. The latest A5 run reached pytest and all 24 precision cases failed because `torch.ops._C_ascend.npu_msa_index_score` was unavailable. Reverting the change removes the failing AscendC MiniMax-M3 index-score path until the custom-operator build and registration issue is resolved. This PR reverts merge commit `fefa33767507514577b80affaa90b5b77ce8ac8a` and preserves unrelated changes already present on `main`, including subsequent nightly matrix maintenance. ### Does this PR introduce _any_ user-facing change? Yes. MiniMax-M3 index-score execution returns to the pre-vllm-project#16426 implementation. The AscendC MsaIndexScore optimization and its associated source, documentation, integration changes, and precision test are removed. ### How was this patch tested? - Reverted merge commit `fefa33767507514577b80affaa90b5b77ce8ac8a` with parent `-m 1`. - Resolved conflicts by preserving current `main` workflow changes and removing only the vllm-project#16426 changes. - `git diff --check` passed. - The revert commit includes a `Signed-off-by` line. - A5 nightly failure: https://github.com/vllm-project/vllm-ascend/actions/runs/35500139842/job/106055635575 Related: vllm-project#16426, vllm-project#16955, vllm-project#17013. - vLLM main: vllm-project/vllm@84030bb Signed-off-by: yingyingzizi <1084697284@qq.com>
…ath (vllm-project#17020) ### What this PR does / why we need it? Revert PR vllm-project#16426, `[Performance] Enable AscendC index score decode for MiniMax-M3`. The A5 nightly validation cannot currently register or execute `npu_msa_index_score`. The latest A5 run reached pytest and all 24 precision cases failed because `torch.ops._C_ascend.npu_msa_index_score` was unavailable. Reverting the change removes the failing AscendC MiniMax-M3 index-score path until the custom-operator build and registration issue is resolved. This PR reverts merge commit `fefa33767507514577b80affaa90b5b77ce8ac8a` and preserves unrelated changes already present on `main`, including subsequent nightly matrix maintenance. ### Does this PR introduce _any_ user-facing change? Yes. MiniMax-M3 index-score execution returns to the pre-vllm-project#16426 implementation. The AscendC MsaIndexScore optimization and its associated source, documentation, integration changes, and precision test are removed. ### How was this patch tested? - Reverted merge commit `fefa33767507514577b80affaa90b5b77ce8ac8a` with parent `-m 1`. - Resolved conflicts by preserving current `main` workflow changes and removing only the vllm-project#16426 changes. - `git diff --check` passed. - The revert commit includes a `Signed-off-by` line. - A5 nightly failure: https://github.com/vllm-project/vllm-ascend/actions/runs/35500139842/job/106055635575 Related: vllm-project#16426, vllm-project#16955, vllm-project#17013. - vLLM main: vllm-project/vllm@84030bb Signed-off-by: yingyingzizi <1084697284@qq.com>
What this PR does / why we need it?
Revert PR #16426,
[Performance] Enable AscendC index score decode for MiniMax-M3.The A5 nightly validation cannot currently register or execute
npu_msa_index_score. The latest A5 run reached pytest and all 24 precision cases failed becausetorch.ops._C_ascend.npu_msa_index_scorewas unavailable. Reverting the change removes the failing AscendC MiniMax-M3 index-score path until the custom-operator build and registration issue is resolved.This PR reverts merge commit
fefa33767507514577b80affaa90b5b77ce8ac8aand preserves unrelated changes already present onmain, including subsequent nightly matrix maintenance.Does this PR introduce any user-facing change?
Yes. MiniMax-M3 index-score execution returns to the pre-#16426 implementation. The AscendC MsaIndexScore optimization and its associated source, documentation, integration changes, and precision test are removed.
How was this patch tested?
fefa33767507514577b80affaa90b5b77ce8ac8awith parent-m 1.mainworkflow changes and removing only the [Performance] Enable AscendC index score decode for MiniMax-M3 #16426 changes.git diff --checkpassed.Signed-off-byline.Related: #16426, #16955, #17013.