Conversation
… layers Signed-off-by: gcw_OTyZJ2Ko <wanwanpeng377@gmail.com>
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request addresses a model loading failure encountered when using DeepSeek-V4-Flash-Vision-Exp checkpoints on vllm-ascend. The issue stemmed from an attempt to load a text router correction bias into hash-routed MoE layers that do not contain this parameter. The changes introduce logic to identify these specific layers during the weight loading process and safely skip the bias parameters, ensuring successful model initialization without impacting regular layers. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. Tip 💡 Consider Linking a Related Issue or RFCYour PR title contains the [BugFix] tag, indicating a bug fix or new feature. Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:
🙏 Thanks for helping us keep the project well-organized! |
There was a problem hiding this comment.
Code Review
Suggested PR Title:
[Models][BugFix] Skip text router correction bias for hash-routed MoE layersSuggested PR Summary:
### What this PR does / why we need it?
This PR skips loading the text router correction bias (`.gate.e_score_correction_bias` and legacy `.gate.bias`) for hash-routed MoE layers in DeepSeek-V4. Hash-routed MoE layers route text tokens through `tid2eid` and image tokens through `bias_vl`, meaning they do not keep a text correction bias parameter. Attempting to load it would result in a `KeyError`.
Additionally, a unit test has been added to verify that loading weights correctly skips these parameters for hash-routed layers while still loading them for regular layers.
Feedback: A runtime `AttributeError` was identified because `logger.info_once` is not a valid method on the vLLM logger. It should be replaced with `logger.info`.
### Does this PR introduce _any_ user-facing change?
No.
### How was this patch tested?
Added a new unit test `test_deepseek_v4_load_weights_skips_hash_layer_gate_bias` in `tests/ut/models/test_deepseek_v4_moe.py`.| logger.info_once( | ||
| "Skipping text router correction bias for hash-routed MoE layer %d", | ||
| layer_idx, | ||
| ) |
There was a problem hiding this comment.
The logger object (imported from vllm.logger) is a standard Python logging.Logger instance. While vLLM adds a custom warning_once method to the logger class, it does not provide an info_once method. Calling logger.info_once will raise an AttributeError at runtime during weight loading.\n\nPlease use logger.info instead.
logger.info(\n \"Skipping text router correction bias for hash-routed MoE layer %d\",\n layer_idx,\n )… layers Signed-off-by: gcw_OTyZJ2Ko <wanwanpeng377@gmail.com>
… layers Signed-off-by: gcw_OTyZJ2Ko <wanwanpeng377@gmail.com>
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
Signed-off-by: ivyilike <102584127+ivyilike@users.noreply.github.com>
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
What this PR does / why we need it?
While validating the official deepseek-ai/DeepSeek-V4-Flash-Vision-Exp release (FP8 dense + MXFP4 experts), weight loading crashes with KeyError: 'model.layers.1.mlp.gate.e_score_correction_bias'
: hash-routed MoE layers keep no such parameter, butload_weightsunconditionally remaps.gate.biasand indexesparams_dict` with it. The ModelSlim W8A8 QuaRot conversion used to validate #15457 does not carry this bias, which is why the path was never exercised.Skip the weight for hash-routed layers under both checkpoint namings; regular layers keep loading it unchanged, and checkpoints without the bias are unaffected.
Related roadmap: #15462.
Does this PR introduce any user-facing change?
Yes — the official DeepSeek-V4-Flash-Vision-Exp release can now be loaded on
vllm-ascend (on FP4-capable hardware). No API change.
How was this patch tested?
UT:
test_deepseek_v4_load_weights_skips_hash_layer_gate_bias(new).vLLM main: vllm-project/vllm@b2f6858