Repository navigation
Conversation
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request introduces support for the IQuest-Coder-V1-40B-Instruct model on Ascend NPU platforms. By aliasing the model's custom architecture to the existing LlamaForCausalLM implementation, it allows for seamless deployment. The changes include necessary registry patches, updated documentation for users, and new test configurations to ensure performance and accuracy on Ascend hardware. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
There was a problem hiding this comment.
Code Review
This pull request adds support for the IQuest-Coder-V1-40B-Instruct model on Ascend NPUs. It introduces a platform patch to register IQuestCoderForCausalLM as an alias of LlamaForCausalLM so that the model can be served natively without --hf-overrides. Additionally, it adds end-to-end test configurations, updates the test groups, and provides a comprehensive tutorial document. The reviewer provided a suggested PR title and summary formatted according to the repository's style guide.
| :::{toctree} | ||
| :maxdepth: 1 | ||
|
|
||
| IQuest-Coder-V1-40B-Instruct | ||
| ::: |
There was a problem hiding this comment.
Suggested PR Title:
[Doc][Feature] Adapt IQuest-Coder-V1-40B-Instruct to Ascend NPUSuggested PR Summary:
### What this PR does / why we need it?
This PR registers `IQuest-Coder-V1-40B-Instruct` (a LLaMA-family 40B code model) as `LlamaForCausalLM` via a `vllm-ascend` registry alias, allowing it to be served natively on Ascend NPUs without requiring `--hf-overrides`. It also adds the end-to-end test configuration and a tutorial document.
Fixes #10662
Fixes #9079
### Does this PR introduce _any_ user-facing change?
Yes, `IQuestLab/IQuest-Coder-V1-40B-Instruct` can now be served on Ascend NPU out of the box (with `--trust-remote-code`). Note that it requires `tensor_parallel_size >= 2` as a single 64GB card cannot hold the 40B weights.
### How was this patch tested?
Verified both Eager and ACLGraph modes on Atlas A2 / 910B3 with TP=2 (NPU devices 4 and 6) generating correct code (fibonacci, quicksort, linked-list, binary-search complexity).References
- The repository style guide requires generating a PR Title and PR Summary in markdown code blocks during review. (link)
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. Tip 💡 Consider Linking a Related Issue or RFCYour PR title contains the [Feature] tag, indicating a bug fix or new feature. Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:
🙏 Thanks for helping us keep the project well-organized! |
cb84c08 to
0c2db08
Compare
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
0c2db08 to
c11a515
Compare
c11a515 to
4d1fdf7
Compare
|
Hi @jyoung6652, could you please review this PR when you have a moment? This PR adapts IQuest-Coder-V1-40B-Instruct to Ascend NPU via a single-file, zero-kernel registry alias ( |
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
4d1fdf7 to
ae2eb49
Compare
|
Rebased onto the latest |
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
ae2eb49 to
8d1f7a0
Compare
|
Rebased onto the latest Pinging the assigned reviewers @Yikun @wangxiyuan @LCAIZJ for a look when you have time. Thanks! |
|
Hi @Yikun @wangxiyuan @LCAIZJ, gentle ping on this PR. It has been rebased onto the latest main and is now conflict-free (mergeable=true, CI passing). The change is small and self-contained: it registers IQuest-Coder-V1-40B-Instruct (a LLaMA-family 40B code model) via a vllm-ascend registry alias, plus its e2e test config and tutorial doc. Would appreciate your review when you get a chance. Thanks! |
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
8d1f7a0 to
3fbc162
Compare
|
Rebased onto the latest |
|
Hi @Yikun @wangxiyuan @LCAIZJ, gentle ping for review/approval of this IQuest-Coder-V1-40B-Instruct adaptation (task #15, tracked in #10662 / #9079). Status:
This PR is up to date and ready for review. Thanks! |
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
3fbc162 to
06bc713
Compare
|
Rebased onto the latest Upstream added a block of documentation comments for the GLM-Next patches right after the The |
Register IQuest-Coder-V1-40B-Instruct (a LLaMA-family 40B code model) as LlamaForCausalLM via a vllm-ascend registry alias, and add the e2e test config plus tutorial doc. - IQuestLab/IQuest-Coder-V1-40B-Instruct can now be served on Ascend NPU out of the box (no --hf-overrides needed), only --trust-remote-code. - Requires tensor_parallel_size >= 2; a single 910B3 (64GB) card cannot hold the 40B weights. - Verified both Eager and ACLGraph modes on Atlas A2 / 910B3 with TP=2 (NPU devices 4 and 6) generate correct code (fibonacci, quicksort, linked-list, binary-search complexity). Fixes vllm-project#10662 Fixes vllm-project#9079 Signed-off-by: WSY435 <WSY435@users.noreply.github.com>
06bc713 to
b7a525b
Compare
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
What
Register IQuest-Coder-V1-40B-Instruct (a LLaMA-family 40B code model) as
LlamaForCausalLM via a vllm-ascend registry alias, and add the e2e test
config plus tutorial doc.
User-facing change
out of the box (no --hf-overrides needed), only --trust-remote-code.
the 40B weights.
How tested
(NPU devices 4 and 6) generate correct code (fibonacci, quicksort,
linked-list, binary-search complexity).
Fixes #10662
Fixes #9079