Skip to content

[Test] Add GLM-5.2 SFA DCP nightly guard - #14683

Merged
wenjun91 merged 3 commits into
vllm-project:mainfrom
pisceskkk:codex/glm52-sfa-dcp-nightly
Sep 7, 2026
Merged

wenjun91 merged 3 commits into
vllm-project:mainfrom
pisceskkk:codex/glm52-sfa-dcp-nightly

Conversation

@pisceskkk

@pisceskkk pisceskkk commented Aug 20, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does

  • replaces the DeepSeek-V3.2 SFA+DCP nightly config with a GLM-5.2 W4A8C8 guard
  • follows the GLM-5.2 single-node 1M layout: DP1/PP1/TP16/PCP1/DCP16 with block/interleave size 128
  • uses the reference 1M SFA setting (enable_sparse_sfa_c8: false) with DSA-CP, LI C8, MTP5, and graph capture sizes [6, 24, 192]
  • registers the case in the A3 nightly matrix

How this patch was tested

  • parsed the YAML and embedded JSON configurations
  • validated the 1M, TP/DCP/PCP, block/interleave, MTP5, and graph-capture settings with static assertions
  • ran git diff --check
  • installed the branch against vLLM ba07e4a48fc951300d97eb506217dd530583dea3 on A3 and confirmed TP8/DCP8 startup resolves MTP5 with graph capture sizes [24, 192]

Related: #13747

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request updates the nightly testing infrastructure by migrating from a deprecated DeepSeek configuration to a new GLM-5.2 W4A8C8 guard. The changes align the test environment with the required 1M layout and enable advanced features to ensure comprehensive coverage for the replicated-indexer path within the A3 nightly matrix.

Highlights

  • Configuration Update: Replaced the obsolete DeepSeek-V3.2 SFA+DCP configuration with a new GLM-5.2 W4A8C8 nightly guard.
  • Architecture Alignment: Updated the configuration to follow the GLM-5.2 single-node 1M layout, including specific TP/DCP/PCP settings and a block/interleave size of 128.
  • Feature Enablement: Enabled SFA C8, LI C8, DSA-CP, and MTP3 to support the replicated-indexer path in the A3 nightly matrix.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Ignored Files
  • Ignored by pattern: .github/workflows/** (1)
    • .github/workflows/configs/nightly_config.yaml
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request replaces the DeepSeek-V3.2 configuration with a GLM-5.2-W4A8C8 nightly test configuration, adjusting environment variables, context parallel sizes, block sizes, and compilation settings. Feedback highlights a critical issue where the speculative decoding method is still configured as deepseek_mtp, which is incompatible with the GLM-5.2 model and will cause a runtime error. Additionally, a suggested PR title and summary have been provided in accordance with the repository's style guide.

Comment on lines +73 to +74
- "--speculative-config"
- '{"num_speculative_tokens": 3, "method": "deepseek_mtp"}'
- '{"num_speculative_tokens": 3, "method": "deepseek_mtp", "enforce_eager": true}'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

critical

The speculative decoding method is set to deepseek_mtp, but the model has been changed to Eco-Tech/GLM-5.2-w4a8c8 (which is a glm_moe_dsa model). The deepseek_mtp speculative method is highly specific to DeepSeek-V3/V4 models that feature Multi-Token Prediction (MTP) modules. Since GLM-5.2 does not have DeepSeek's MTP architecture, this configuration will cause a runtime error during initialization. Please remove the speculative config if speculative decoding is not supported/needed for this test, or correct it to a compatible speculative method.

@@ -3,14 +3,17 @@
# ==========================================

test_cases:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Suggested PR Title:

[Test][Feature] Add GLM-5.2 SFA DCP nightly guard

Suggested PR Summary:

### What this PR does / why we need it?

This PR replaces the obsolete DeepSeek-V3.2 SFA+DCP configuration with a new GLM-5.2 W4A8C8 nightly guard. It configures a single-node 1M context layout (DP1/PP1/TP16/PCP1/DCP16) with block and interleave sizes of 128. It also enables SFA C8, LI C8, DSA-CP, and speculative decoding to cover the replicated-indexer path, and registers the case in the A3 nightly matrix.

Fixes #13747

### Does this PR introduce _any_ user-facing change?

No, this is a test-only change adding a nightly guard configuration.

### How was this patch tested?

- Validated the YAML and embedded JSON configurations.
- Verified the 1M context, TP/DCP/PCP, block/interleave, and SFA settings with static assertions.
- Ran `git diff --check`.
References
  1. Follow the Pull Request Summary Style Guide to provide a suggested PR Title and Summary in the specified markdown format. (link)

@pisceskkk
pisceskkk marked this pull request as ready for review August 20, 2026 12:53
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@weiguihua2

weiguihua2 commented Aug 25, 2026 •

Copy link
Copy Markdown
Collaborator

/nightly glm-5.2-w4a8c8-sfa-dcp
nightly command triggered.

@pisceskkk
pisceskkk force-pushed the codex/glm52-sfa-dcp-nightly branch from 8428448 to fd97c2a Compare August 25, 2026 01:20
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@pisceskkk

pisceskkk commented Aug 26, 2026 •

Copy link
Copy Markdown
Contributor Author

/nightly glm-5.2-w4a8c8-sfa-dcp
[Bot]: nightly command failed: you do not have permission. Only users with triage+ permission can trigger /nightly.

@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com>
Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com>
@pisceskkk
pisceskkk force-pushed the codex/glm52-sfa-dcp-nightly branch from dd72b34 to e7f22e0 Compare August 31, 2026 01:32
@weiguihua2

weiguihua2 commented Aug 31, 2026 •

Copy link
Copy Markdown
Collaborator

/nightly glm-5.2-w4a8c8-sfa-dcp
nightly command triggered.

Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com>
@weiguihua2

weiguihua2 commented Aug 31, 2026 •

Copy link
Copy Markdown
Collaborator

/nightly glm-5.2-w4a8c8-sfa-dcp
nightly command triggered.

@wenjun91

wenjun91 commented Sep 7, 2026 •

Copy link
Copy Markdown
Collaborator

/nightly glm-5.2-w4a8c8-sfa-dcp
nightly command triggered.

@wenjun91
wenjun91 merged commit 6b0dd25 into vllm-project:main Sep 7, 2026
11 checks passed
yiminghub2024 added a commit to yiminghub2024/vllm-ascend that referenced this pull request Sep 10, 2026
While this branch was open, main grew its own implementations of both features
it was proposing:

  - vllm-project#15913 GLM-Next KV cache management, which keeps the incomplete pool in an
    absolute-position FP32 Glm5NextStateCache rather than in this branch's
    paged Glm5NextTailCache, and stores completed pools as unquantized BF16
  - vllm-project#15669 decoupled the indexer from SFA, restructuring the very plumbing this
    branch's kpool backend was wired into
  - AscendDflash2Proposer, selected by is_dflash2_draft() under method dflash

Those designs are mutually exclusive with this branch's, not textually
conflicting with it, so every conflict is resolved in favour of main and the
branch's own kpool backend, tail cache and DFlash2 patch are dropped. The
follow-up commit re-adds only what main still lacks.

Also drops CI-config churn this branch had picked up from a stale tree: it
reverted its own parent vllm-project#14683 by deleting GLM-5.2-W4A8C8-SFA-DCP.yaml and
moving nightly entries into weekly, and added an unreferenced
DeepSeek-V3.2-W8A8-DCP.yaml.

Signed-off-by: yiminghub2024 <482890@qq.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
sunny-rain-63 pushed a commit to sunny-rain-63/vllm-ascend that referenced this pull request Sep 12, 2026
## What this PR does

- replaces the DeepSeek-V3.2 SFA+DCP nightly config with a GLM-5.2
W4A8C8 guard
- follows the GLM-5.2 single-node 1M layout: DP1/PP1/TP16/PCP1/DCP16
with block/interleave size 128
- uses the reference 1M SFA setting (`enable_sparse_sfa_c8: false`) with
DSA-CP, LI C8, MTP5, and graph capture sizes `[6, 24, 192]`
- registers the case in the A3 nightly matrix

## How this patch was tested

- parsed the YAML and embedded JSON configurations
- validated the 1M, TP/DCP/PCP, block/interleave, MTP5, and
graph-capture settings with static assertions
- ran `git diff --check`
- installed the branch against vLLM
`ba07e4a48fc951300d97eb506217dd530583dea3` on A3 and confirmed TP8/DCP8
startup resolves MTP5 with graph capture sizes `[24, 192]`

Related: vllm-project#13747

- vLLM main:
vllm-project/vllm@ba07e4a

---------

Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com>
like-0517 pushed a commit to like-0517/vllm-ascend that referenced this pull request Sep 15, 2026
## What this PR does

- replaces the DeepSeek-V3.2 SFA+DCP nightly config with a GLM-5.2
W4A8C8 guard
- follows the GLM-5.2 single-node 1M layout: DP1/PP1/TP16/PCP1/DCP16
with block/interleave size 128
- uses the reference 1M SFA setting (`enable_sparse_sfa_c8: false`) with
DSA-CP, LI C8, MTP5, and graph capture sizes `[6, 24, 192]`
- registers the case in the A3 nightly matrix

## How this patch was tested

- parsed the YAML and embedded JSON configurations
- validated the 1M, TP/DCP/PCP, block/interleave, MTP5, and
graph-capture settings with static assertions
- ran `git diff --check`
- installed the branch against vLLM
`ba07e4a48fc951300d97eb506217dd530583dea3` on A3 and confirmed TP8/DCP8
startup resolves MTP5 with graph capture sizes `[24, 192]`

Related: vllm-project#13747

- vLLM main:
vllm-project/vllm@ba07e4a

---------

Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com>
Signed-off-by: like-0517 <ithwlike@126.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants