Skip to content

[DeepSeek][ROCm] V4.1-Flash: enable INT4 QuickReduce on MI355X - #1061

Merged
Fangzhou-Ai merged 1 commit into
vllm-project:mainfrom
ahmed-bsod:ahmed/int4-qr
Oct 3, 2026
Merged

Fangzhou-Ai merged 1 commit into
vllm-project:mainfrom
ahmed-bsod:ahmed/int4-qr

Conversation

@ahmed-bsod

Copy link
Copy Markdown
Contributor

No description provided.

@vercel

vercel Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
vllm-recipes Ready Ready Preview Oct 2, 2026 8:17pm UTC

Request Review

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request updates the metadata and adds a ROCm strategy override for the DeepSeek-V4.1-Flash model. A review comment correctly identifies that the environment variable name used for QuickReduce quantization is incorrect and should be changed from VLLM_ROCM_QUICK_REDUCE_QUANTIZATION to VLLM_ROCM_QUICK_REDUCE_QUANT to prevent the setting from being ignored.

Comment on lines +281 to +282
# INT4 QuickReduce
VLLM_ROCM_QUICK_REDUCE_QUANTIZATION: "INT4"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The environment variable used by vLLM to configure QuickReduce quantization is VLLM_ROCM_QUICK_REDUCE_QUANT, not VLLM_ROCM_QUICK_REDUCE_QUANTIZATION. Using the incorrect variable name will result in the setting being ignored.

          # INT4 QuickReduce
          VLLM_ROCM_QUICK_REDUCE_QUANT: "INT4"

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hallucinating

Signed-off-by: Muhammad Ahmed <Muhammad.Ahmed@amd.com>
@Fangzhou-Ai

Copy link
Copy Markdown
Collaborator

Thanks @ahmed-bsod Let me verify it first before merging

@Fangzhou-Ai

Copy link
Copy Markdown
Collaborator

This branch was successfully deployed

1 active deployment
Preview — c310fc80 Deployed Oct 2, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants