Skip to content

[BugFix][Model] Skip text router correction bias for DeepSeek V4 hash layers - #16029

Open
ivyilike wants to merge 5 commits into
vllm-project:mainfrom
ivyilike:dsv4_vision
Open

ivyilike wants to merge 5 commits into
vllm-project:mainfrom
ivyilike:dsv4_vision

Conversation

@ivyilike

@ivyilike ivyilike commented Sep 8, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?

While validating the official deepseek-ai/DeepSeek-V4-Flash-Vision-Exp release (FP8 dense + MXFP4 experts), weight loading crashes with KeyError: 'model.layers.1.mlp.gate.e_score_correction_bias': hash-routed MoE layers keep no such parameter, but load_weightsunconditionally remaps.gate.biasand indexesparams_dict` with it. The ModelSlim W8A8 QuaRot conversion used to validate #15457 does not carry this bias, which is why the path was never exercised.
Skip the weight for hash-routed layers under both checkpoint namings; regular layers keep loading it unchanged, and checkpoints without the bias are unaffected.
Related roadmap: #15462.

Does this PR introduce any user-facing change?

Yes — the official DeepSeek-V4-Flash-Vision-Exp release can now be loaded on
vllm-ascend (on FP4-capable hardware). No API change.

How was this patch tested?

… layers

Signed-off-by: gcw_OTyZJ2Ko <wanwanpeng377@gmail.com>
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request addresses a model loading failure encountered when using DeepSeek-V4-Flash-Vision-Exp checkpoints on vllm-ascend. The issue stemmed from an attempt to load a text router correction bias into hash-routed MoE layers that do not contain this parameter. The changes introduce logic to identify these specific layers during the weight loading process and safely skip the bias parameters, ensuring successful model initialization without impacting regular layers.

Highlights

  • Weight Loading Fix: Implemented a skip mechanism for the 'e_score_correction_bias' parameter in hash-routed MoE layers to prevent KeyError during model loading.
  • Compatibility: Ensured compatibility with official DeepSeek-V4-Flash-Vision-Exp checkpoints by handling both legacy and current naming conventions for gate biases.
  • Testing: Added a new unit test, 'test_deepseek_v4_load_weights_skips_hash_layer_gate_bias', to verify that hash-routed layers correctly bypass the bias loading process.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@github-actions

github-actions Bot commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.


Tip

💡 Consider Linking a Related Issue or RFC

Your PR title contains the [BugFix] tag, indicating a bug fix or new feature.

Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:

  • Fixes #<issue_number>
  • Closes #<issue_number>
  • Resolves #<issue_number>
  • Refs #<rfc_or_issue_number> (for RFCs)

🙏 Thanks for helping us keep the project well-organized!

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Models][BugFix] Skip text router correction bias for hash-routed MoE layers

Suggested PR Summary:

### What this PR does / why we need it?

This PR skips loading the text router correction bias (`.gate.e_score_correction_bias` and legacy `.gate.bias`) for hash-routed MoE layers in DeepSeek-V4. Hash-routed MoE layers route text tokens through `tid2eid` and image tokens through `bias_vl`, meaning they do not keep a text correction bias parameter. Attempting to load it would result in a `KeyError`.

Additionally, a unit test has been added to verify that loading weights correctly skips these parameters for hash-routed layers while still loading them for regular layers.

Feedback: A runtime `AttributeError` was identified because `logger.info_once` is not a valid method on the vLLM logger. It should be replaced with `logger.info`.

### Does this PR introduce _any_ user-facing change?

No.

### How was this patch tested?

Added a new unit test `test_deepseek_v4_load_weights_skips_hash_layer_gate_bias` in `tests/ut/models/test_deepseek_v4_moe.py`.

Comment on lines +1183 to +1186
logger.info_once(
"Skipping text router correction bias for hash-routed MoE layer %d",
layer_idx,
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The logger object (imported from vllm.logger) is a standard Python logging.Logger instance. While vLLM adds a custom warning_once method to the logger class, it does not provide an info_once method. Calling logger.info_once will raise an AttributeError at runtime during weight loading.\n\nPlease use logger.info instead.

                    logger.info(\n                        \"Skipping text router correction bias for hash-routed MoE layer %d\",\n                        layer_idx,\n                    )

… layers

Signed-off-by: gcw_OTyZJ2Ko <wanwanpeng377@gmail.com>
… layers

Signed-off-by: gcw_OTyZJ2Ko <wanwanpeng377@gmail.com>
@ivyilike
ivyilike marked this pull request as ready for review September 8, 2026 08:41
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

Signed-off-by: ivyilike <102584127+ivyilike@users.noreply.github.com>
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant