Repository navigation
[Test] Add GLM-5.2 SFA DCP nightly guard - #14683
Conversation
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request updates the nightly testing infrastructure by migrating from a deprecated DeepSeek configuration to a new GLM-5.2 W4A8C8 guard. The changes align the test environment with the required 1M layout and enable advanced features to ensure comprehensive coverage for the replicated-indexer path within the A3 nightly matrix. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Ignored Files
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. |
There was a problem hiding this comment.
Code Review
This pull request replaces the DeepSeek-V3.2 configuration with a GLM-5.2-W4A8C8 nightly test configuration, adjusting environment variables, context parallel sizes, block sizes, and compilation settings. Feedback highlights a critical issue where the speculative decoding method is still configured as deepseek_mtp, which is incompatible with the GLM-5.2 model and will cause a runtime error. Additionally, a suggested PR title and summary have been provided in accordance with the repository's style guide.
| - "--speculative-config" | ||
| - '{"num_speculative_tokens": 3, "method": "deepseek_mtp"}' | ||
| - '{"num_speculative_tokens": 3, "method": "deepseek_mtp", "enforce_eager": true}' |
There was a problem hiding this comment.
The speculative decoding method is set to deepseek_mtp, but the model has been changed to Eco-Tech/GLM-5.2-w4a8c8 (which is a glm_moe_dsa model). The deepseek_mtp speculative method is highly specific to DeepSeek-V3/V4 models that feature Multi-Token Prediction (MTP) modules. Since GLM-5.2 does not have DeepSeek's MTP architecture, this configuration will cause a runtime error during initialization. Please remove the speculative config if speculative decoding is not supported/needed for this test, or correct it to a compatible speculative method.
| @@ -3,14 +3,17 @@ | |||
| # ========================================== | |||
|
|
|||
| test_cases: | |||
There was a problem hiding this comment.
Suggested PR Title:
[Test][Feature] Add GLM-5.2 SFA DCP nightly guardSuggested PR Summary:
### What this PR does / why we need it?
This PR replaces the obsolete DeepSeek-V3.2 SFA+DCP configuration with a new GLM-5.2 W4A8C8 nightly guard. It configures a single-node 1M context layout (DP1/PP1/TP16/PCP1/DCP16) with block and interleave sizes of 128. It also enables SFA C8, LI C8, DSA-CP, and speculative decoding to cover the replicated-indexer path, and registers the case in the A3 nightly matrix.
Fixes #13747
### Does this PR introduce _any_ user-facing change?
No, this is a test-only change adding a nightly guard configuration.
### How was this patch tested?
- Validated the YAML and embedded JSON configurations.
- Verified the 1M context, TP/DCP/PCP, block/interleave, and SFA settings with static assertions.
- Ran `git diff --check`.References
- Follow the Pull Request Summary Style Guide to provide a suggested PR Title and Summary in the specified markdown format. (link)
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
7703bc4 to
8428448
Compare
|
/nightly glm-5.2-w4a8c8-sfa-dcp
|
8428448 to
fd97c2a
Compare
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
241d977 to
dd72b34
Compare
|
/nightly glm-5.2-w4a8c8-sfa-dcp |
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com>
Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com>
dd72b34 to
e7f22e0
Compare
|
/nightly glm-5.2-w4a8c8-sfa-dcp
|
Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com>
|
/nightly glm-5.2-w4a8c8-sfa-dcp
|
|
/nightly glm-5.2-w4a8c8-sfa-dcp
|
While this branch was open, main grew its own implementations of both features it was proposing: - vllm-project#15913 GLM-Next KV cache management, which keeps the incomplete pool in an absolute-position FP32 Glm5NextStateCache rather than in this branch's paged Glm5NextTailCache, and stores completed pools as unquantized BF16 - vllm-project#15669 decoupled the indexer from SFA, restructuring the very plumbing this branch's kpool backend was wired into - AscendDflash2Proposer, selected by is_dflash2_draft() under method dflash Those designs are mutually exclusive with this branch's, not textually conflicting with it, so every conflict is resolved in favour of main and the branch's own kpool backend, tail cache and DFlash2 patch are dropped. The follow-up commit re-adds only what main still lacks. Also drops CI-config churn this branch had picked up from a stale tree: it reverted its own parent vllm-project#14683 by deleting GLM-5.2-W4A8C8-SFA-DCP.yaml and moving nightly entries into weekly, and added an unreferenced DeepSeek-V3.2-W8A8-DCP.yaml. Signed-off-by: yiminghub2024 <482890@qq.com> Co-authored-by: Cursor <cursoragent@cursor.com>
## What this PR does - replaces the DeepSeek-V3.2 SFA+DCP nightly config with a GLM-5.2 W4A8C8 guard - follows the GLM-5.2 single-node 1M layout: DP1/PP1/TP16/PCP1/DCP16 with block/interleave size 128 - uses the reference 1M SFA setting (`enable_sparse_sfa_c8: false`) with DSA-CP, LI C8, MTP5, and graph capture sizes `[6, 24, 192]` - registers the case in the A3 nightly matrix ## How this patch was tested - parsed the YAML and embedded JSON configurations - validated the 1M, TP/DCP/PCP, block/interleave, MTP5, and graph-capture settings with static assertions - ran `git diff --check` - installed the branch against vLLM `ba07e4a48fc951300d97eb506217dd530583dea3` on A3 and confirmed TP8/DCP8 startup resolves MTP5 with graph capture sizes `[24, 192]` Related: vllm-project#13747 - vLLM main: vllm-project/vllm@ba07e4a --------- Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com>
## What this PR does - replaces the DeepSeek-V3.2 SFA+DCP nightly config with a GLM-5.2 W4A8C8 guard - follows the GLM-5.2 single-node 1M layout: DP1/PP1/TP16/PCP1/DCP16 with block/interleave size 128 - uses the reference 1M SFA setting (`enable_sparse_sfa_c8: false`) with DSA-CP, LI C8, MTP5, and graph capture sizes `[6, 24, 192]` - registers the case in the A3 nightly matrix ## How this patch was tested - parsed the YAML and embedded JSON configurations - validated the 1M, TP/DCP/PCP, block/interleave, MTP5, and graph-capture settings with static assertions - ran `git diff --check` - installed the branch against vLLM `ba07e4a48fc951300d97eb506217dd530583dea3` on A3 and confirmed TP8/DCP8 startup resolves MTP5 with graph capture sizes `[24, 192]` Related: vllm-project#13747 - vLLM main: vllm-project/vllm@ba07e4a --------- Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com> Signed-off-by: like-0517 <ithwlike@126.com>
What this PR does
enable_sparse_sfa_c8: false) with DSA-CP, LI C8, MTP5, and graph capture sizes[6, 24, 192]How this patch was tested
git diff --checkba07e4a48fc951300d97eb506217dd530583dea3on A3 and confirmed TP8/DCP8 startup resolves MTP5 with graph capture sizes[24, 192]Related: #13747