Skip to content

[CI][Kimi K3] Reduce pull request test to four NPUs - #15403

Closed
maoxx241 wants to merge 1 commit into
vllm-project:mainfrom
maoxx241:codex/k3-four-card-ci-20260831
Closed

maoxx241 wants to merge 1 commit into
vllm-project:mainfrom
maoxx241:codex/k3-four-card-ci-20260831

Conversation

@maoxx241

@maoxx241 maoxx241 commented Aug 31, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?

Follow-up to #14454. The Kimi K3 pull-request functional test currently requires a dedicated 16-NPU runner while its four scenarios can be preserved on one four-NPU A3 runner.

This PR:

  • moves the test from sixteen_card to four_card and reuses the existing A3 four-card partition;
  • keeps production tensor widths and 16 routed experts, while reducing the dummy target to two layers (one KDA and one MLA), the draft to one layer, the vision tower to one layer, and routed top-k to four;
  • retains Block5 TP4, quantized GQA DP2/TP2, MTP TP4 with an image, and MLA P2/D2 coverage;
  • keeps the discriminating request boundaries for chunked prefill, aligned hybrid-cache page creation, both DP engines, speculative draft metrics, remote KV transfer, and decode-side local fallback;
  • adjusts the reduced TP2 GQA draft head layout so its cache page size matches the target MLA cache page;
  • keeps the CI estimate at 360 seconds.

The test uses dummy weights and is a functional smoke test. It does not validate checkpoint loading, QuaRot, model accuracy, or speculative acceptance rate.

Does this PR introduce any user-facing change?

No. This only reduces the hardware requirement and runtime of an existing pull-request test.

How was this patch tested?

  • bash format.sh ci
  • A3, four logical NPUs, current PR head 7e0871722f4d897cf4ec3aa75bdf7b1b74153cd6:
pytest -sv tests/e2e/pull_request/four_card/test_kimi_k3.py --durations=10
4 passed, 14 warnings in 274.41s
end-to-end elapsed: 292s

Observed evidence includes four-rank startup, real completion requests on both DP engines, speculative draft counters, MTP image generation, Mooncake KV transfer, and a decode-side request without remote KV.

Keep the four Kimi K3 functional scenarios while reducing the dummy target, draft, and request matrix so they fit an A3 four-card runner in roughly six minutes.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request optimizes the CI testing infrastructure for the Kimi K3 model by reducing the hardware footprint required for functional smoke tests. By downsizing the test model configuration and adjusting the test scenarios, the PR enables the execution of these tests on smaller, more readily available hardware partitions without sacrificing the validation of key architectural features and functional boundaries.

Highlights

  • Hardware Requirement Reduction: The Kimi K3 functional test has been migrated from a 16-NPU requirement to a 4-NPU requirement, allowing it to run on standard A3 four-card partitions.
  • Test Configuration Optimization: The test model architecture was scaled down to two layers while maintaining production tensor widths and key functional coverage, including Block5, GQA, MLA, and MTP.
  • Functional Coverage Preservation: The reduced test continues to validate critical features such as chunked prefill, hybrid-cache page creation, remote KV transfer, and speculative decoding metrics.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Ignored Files
  • Ignored by pattern: .github/workflows/** (1)
    • .github/workflows/scripts/test_config.yaml
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Test][Misc] Scale down Kimi K3 functional tests from 16 NPUs to 4 NPUs

Suggested PR Summary:

### What this PR does / why we need it?
This PR scales down the single-node Kimi K3 functional tests from 16 logical NPUs (TP16) to 4 logical NPUs (TP4/TP2). It reduces the model size (layers from 6 to 2, max model length from 2048 to 1024) and draft model configurations to allow these tests to run efficiently on smaller hardware configurations (such as 4-card setups) while still exercising key features like Block5, quantized GQA, legacy MLA, and MTP.

### Does this PR introduce _any_ user-facing change?
No, this PR only modifies end-to-end functional tests.

### How was this patch tested?
The changes modify the existing E2E tests in `tests/e2e/pull_request/four_card/test_kimi_k3.py`.

I have no feedback to provide on the review comments as none were present.

@wangxiyuan wangxiyuan added the ready-precise run selected e2e test for pr label Aug 31, 2026

@wangxiyuan wangxiyuan left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@github-actions

github-actions Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants