Skip to content

[Doc][Feature] Adapt IQuest-Coder-V1-40B-Instruct to Ascend NPU - #12855

Open
WSY435 wants to merge 1 commit into
vllm-project:mainfrom
WSY435:support-iquest-coder-v1-40b-instruct
Open

WSY435 wants to merge 1 commit into
vllm-project:mainfrom
WSY435:support-iquest-coder-v1-40b-instruct

Conversation

@WSY435

@WSY435 WSY435 commented Jul 25, 2026 •

Copy link
Copy Markdown

What

Register IQuest-Coder-V1-40B-Instruct (a LLaMA-family 40B code model) as
LlamaForCausalLM via a vllm-ascend registry alias, and add the e2e test
config plus tutorial doc.

User-facing change

  • IQuestLab/IQuest-Coder-V1-40B-Instruct can now be served on Ascend NPU
    out of the box (no --hf-overrides needed), only --trust-remote-code.
  • Requires tensor_parallel_size >= 2; a single 910B3 (64GB) card cannot hold
    the 40B weights.

How tested

  • Verified both Eager and ACLGraph modes on Atlas A2 / 910B3 with TP=2
    (NPU devices 4 and 6) generate correct code (fibonacci, quicksort,
    linked-list, binary-search complexity).

Fixes #10662
Fixes #9079

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request introduces support for the IQuest-Coder-V1-40B-Instruct model on Ascend NPU platforms. By aliasing the model's custom architecture to the existing LlamaForCausalLM implementation, it allows for seamless deployment. The changes include necessary registry patches, updated documentation for users, and new test configurations to ensure performance and accuracy on Ascend hardware.

Highlights

  • Model Support: Registered IQuest-Coder-V1-40B-Instruct as a LlamaForCausalLM alias in the vllm-ascend registry, enabling native support without requiring --hf-overrides.
  • Documentation: Added a comprehensive tutorial for deploying the 40B model on Ascend NPUs, including environment setup, hardware requirements, and verification steps.
  • Testing and Evaluation: Integrated end-to-end accuracy testing configurations for the model, enabling automated evaluation on Ascend hardware.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request adds support for the IQuest-Coder-V1-40B-Instruct model on Ascend NPUs. It introduces a platform patch to register IQuestCoderForCausalLM as an alias of LlamaForCausalLM so that the model can be served natively without --hf-overrides. Additionally, it adds end-to-end test configurations, updates the test groups, and provides a comprehensive tutorial document. The reviewer provided a suggested PR title and summary formatted according to the repository's style guide.

Comment thread docs/source/tutorials/models/index.md Outdated
Comment on lines +5 to +9
:::{toctree}
:maxdepth: 1

IQuest-Coder-V1-40B-Instruct
:::

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Suggested PR Title:

[Doc][Feature] Adapt IQuest-Coder-V1-40B-Instruct to Ascend NPU

Suggested PR Summary:

### What this PR does / why we need it?

This PR registers `IQuest-Coder-V1-40B-Instruct` (a LLaMA-family 40B code model) as `LlamaForCausalLM` via a `vllm-ascend` registry alias, allowing it to be served natively on Ascend NPUs without requiring `--hf-overrides`. It also adds the end-to-end test configuration and a tutorial document.

Fixes #10662
Fixes #9079

### Does this PR introduce _any_ user-facing change?

Yes, `IQuestLab/IQuest-Coder-V1-40B-Instruct` can now be served on Ascend NPU out of the box (with `--trust-remote-code`). Note that it requires `tensor_parallel_size >= 2` as a single 64GB card cannot hold the 40B weights.

### How was this patch tested?

Verified both Eager and ACLGraph modes on Atlas A2 / 910B3 with TP=2 (NPU devices 4 and 6) generating correct code (fibonacci, quicksort, linked-list, binary-search complexity).
References
  1. The repository style guide requires generating a PR Title and PR Summary in markdown code blocks during review. (link)

@github-actions github-actions Bot added documentation Improvements or additions to documentation module:tests labels Jul 25, 2026
@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.


Tip

💡 Consider Linking a Related Issue or RFC

Your PR title contains the [Feature] tag, indicating a bug fix or new feature.

Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:

  • Fixes #<issue_number>
  • Closes #<issue_number>
  • Resolves #<issue_number>
  • Refs #<rfc_or_issue_number> (for RFCs)

🙏 Thanks for helping us keep the project well-organized!

@WSY435
WSY435 force-pushed the support-iquest-coder-v1-40b-instruct branch from cb84c08 to 0c2db08 Compare July 25, 2026 08:41
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@WSY435
WSY435 force-pushed the support-iquest-coder-v1-40b-instruct branch from 0c2db08 to c11a515 Compare August 12, 2026 01:33
@WSY435
WSY435 force-pushed the support-iquest-coder-v1-40b-instruct branch from c11a515 to 4d1fdf7 Compare August 12, 2026 02:12
@WSY435

WSY435 commented Aug 12, 2026

Copy link
Copy Markdown
Author

Hi @jyoung6652, could you please review this PR when you have a moment?

This PR adapts IQuest-Coder-V1-40B-Instruct to Ascend NPU via a single-file, zero-kernel registry alias (IQuestCoderForCausalLM -> LlamaForCausalLM), plus the e2e accuracy config and a tutorial doc. The branch is now rebased onto the latest main and all CI checks pass.

It fixes #10662 and #9079. Thanks a lot!

@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@WSY435
WSY435 force-pushed the support-iquest-coder-v1-40b-instruct branch from 4d1fdf7 to ae2eb49 Compare August 28, 2026 03:58
@WSY435

WSY435 commented Aug 28, 2026

Copy link
Copy Markdown
Author

Rebased onto the latest main and resolved the conflict in docs/source/tutorials/models/index.md (the toctree was migrated to the markdown-link format upstream). The merge-conflicts label should clear automatically once GitHub re-scans. CI is re-running on the new HEAD. Thanks for the review - pinging @jyoung6652 to take another look.

@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@WSY435
WSY435 force-pushed the support-iquest-coder-v1-40b-instruct branch from ae2eb49 to 8d1f7a0 Compare August 29, 2026 00:55
@WSY435

WSY435 commented Aug 29, 2026

Copy link
Copy Markdown
Author

Rebased onto the latest main (6953f266) and resolved the conflict in docs/source/tutorials/models/index.md. Upstream has since removed the Kimi K3 entry from the toctree, so this PR now adds only the IQuest-Coder-V1-40B-Instruct entry to match the current docs structure. The merge-conflicts label should clear automatically. CI is re-running on the new HEAD.

Pinging the assigned reviewers @Yikun @wangxiyuan @LCAIZJ for a look when you have time. Thanks!

@WSY435

WSY435 commented Aug 29, 2026

Copy link
Copy Markdown
Author

Hi @Yikun @wangxiyuan @LCAIZJ, gentle ping on this PR. It has been rebased onto the latest main and is now conflict-free (mergeable=true, CI passing). The change is small and self-contained: it registers IQuest-Coder-V1-40B-Instruct (a LLaMA-family 40B code model) via a vllm-ascend registry alias, plus its e2e test config and tutorial doc. Would appreciate your review when you get a chance. Thanks!

@github-actions

github-actions Bot commented Sep 3, 2026

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@WSY435
WSY435 force-pushed the support-iquest-coder-v1-40b-instruct branch from 8d1f7a0 to 3fbc162 Compare September 3, 2026 05:04
@WSY435

WSY435 commented Sep 3, 2026

Copy link
Copy Markdown
Author

Rebased onto the latest main (78f3a90f, GLM-5.3-Flash on Ascend 950, #15127) and resolved the conflict in vllm_ascend/patch/platform/__init__.py. Upstream added import ...patch_glm5next_config while this PR adds import ...patch_iquestcoder_config; both imports are now kept. The merge-conflicts label should clear after the bot re-checks. The PR is conflict-free again and ready for review. Thanks!

@WSY435

WSY435 commented Sep 3, 2026 •

Copy link
Copy Markdown
Author

Hi @Yikun @wangxiyuan @LCAIZJ, gentle ping for review/approval of this IQuest-Coder-V1-40B-Instruct adaptation (task #15, tracked in #10662 / #9079).

Status:

  • Rebased onto latest main (78f3a90f, GLM-5.3-Flash [Feature][Model] Support GLM-5.3-Flash on Ascend 950 #15127) and conflict-free.
  • pre-commit and select-tests both pass.
  • The ci-gate check is currently red only because the PR touches vllm_ascend/** source and needs the ready-precise (or ready-all) label to trigger the precision test run — could a maintainer please add that label?

This PR is up to date and ready for review. Thanks!

@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@WSY435
WSY435 force-pushed the support-iquest-coder-v1-40b-instruct branch from 3fbc162 to 06bc713 Compare September 10, 2026 07:23
@WSY435

WSY435 commented Sep 10, 2026

Copy link
Copy Markdown
Author

Rebased onto the latest main (49808096, #15737 - Complete SP MoE guide and temporary FlashComm switch) and resolved the conflict in vllm_ascend/patch/platform/__init__.py.

Upstream added a block of documentation comments for the GLM-Next patches right after the patch_glm5next_config import, while this PR adds import ...patch_iquestcoder_config at the same location. The resolution keeps both: the upstream comment block is preserved and the IQuest import is placed alongside the other imports. ruff check and ruff format --check both pass.

The merge-conflicts label should clear after the bot re-checks. The PR is conflict-free again and ready for review. Thanks!

Register IQuest-Coder-V1-40B-Instruct (a LLaMA-family 40B code model) as
LlamaForCausalLM via a vllm-ascend registry alias, and add the e2e test
config plus tutorial doc.

- IQuestLab/IQuest-Coder-V1-40B-Instruct can now be served on Ascend NPU
  out of the box (no --hf-overrides needed), only --trust-remote-code.
- Requires tensor_parallel_size >= 2; a single 910B3 (64GB) card cannot hold
  the 40B weights.

- Verified both Eager and ACLGraph modes on Atlas A2 / 910B3 with TP=2
  (NPU devices 4 and 6) generate correct code (fibonacci, quicksort,
  linked-list, binary-search complexity).

Fixes vllm-project#10662
Fixes vllm-project#9079

Signed-off-by: WSY435 <WSY435@users.noreply.github.com>
@WSY435
WSY435 force-pushed the support-iquest-coder-v1-40b-instruct branch from 06bc713 to b7a525b Compare September 10, 2026 08:29
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation merge-conflicts module:tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Contribution] 任务 #15:IQuest-Coder-V1-40B 代码大模型适配 [Contribution] vLLM-Ascend 外部开发者任务池

1 participant