Repository navigation
Conversation
Keep the four Kimi K3 functional scenarios while reducing the dummy target, draft, and request matrix so they fit an A3 four-card runner in roughly six minutes. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request optimizes the CI testing infrastructure for the Kimi K3 model by reducing the hardware footprint required for functional smoke tests. By downsizing the test model configuration and adjusting the test scenarios, the PR enables the execution of these tests on smaller, more readily available hardware partitions without sacrificing the validation of key architectural features and functional boundaries. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Ignored Files
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. |
There was a problem hiding this comment.
Code Review
Suggested PR Title:
[Test][Misc] Scale down Kimi K3 functional tests from 16 NPUs to 4 NPUsSuggested PR Summary:
### What this PR does / why we need it?
This PR scales down the single-node Kimi K3 functional tests from 16 logical NPUs (TP16) to 4 logical NPUs (TP4/TP2). It reduces the model size (layers from 6 to 2, max model length from 2048 to 1024) and draft model configurations to allow these tests to run efficiently on smaller hardware configurations (such as 4-card setups) while still exercising key features like Block5, quantized GQA, legacy MLA, and MTP.
### Does this PR introduce _any_ user-facing change?
No, this PR only modifies end-to-end functional tests.
### How was this patch tested?
The changes modify the existing E2E tests in `tests/e2e/pull_request/four_card/test_kimi_k3.py`.I have no feedback to provide on the review comments as none were present.
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
What this PR does / why we need it?
Follow-up to #14454. The Kimi K3 pull-request functional test currently requires a dedicated 16-NPU runner while its four scenarios can be preserved on one four-NPU A3 runner.
This PR:
sixteen_cardtofour_cardand reuses the existing A3 four-card partition;The test uses dummy weights and is a functional smoke test. It does not validate checkpoint loading, QuaRot, model accuracy, or speculative acceptance rate.
Does this PR introduce any user-facing change?
No. This only reduces the hardware requirement and runtime of an existing pull-request test.
How was this patch tested?
bash format.sh ci7e0871722f4d897cf4ec3aa75bdf7b1b74153cd6:Observed evidence includes four-rank startup, real completion requests on both DP engines, speculative draft counters, MTP image generation, Mooncake KV transfer, and a decode-side request without remote KV.