Skip to content

docs: add optional MiniMax-M3 GQA FlashInfer config - #1047

Merged
esmeetu merged 1 commit into
vllm-project:mainfrom
xinli-sw:docs/minimax-m3-blackwell-agentic-options
Sep 30, 2026
Merged

esmeetu merged 1 commit into
vllm-project:mainfrom
xinli-sw:docs/minimax-m3-blackwell-agentic-options

Conversation

@xinli-sw

@xinli-sw xinli-sw commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Adds an optional GQA example: FlashInfer draft attention when it benefits the workload, plus local argmax reduction to reduce communication traffic.

Validation: YAML and example config parse successfully.

@vercel

vercel Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
vllm-recipes Ready Ready Preview Sep 29, 2026 9:45pm UTC

Request Review

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request updates the MiniMax-M3 model metadata and adds configuration instructions for B200/B300 agentic serving with the NVFP4 variant, specifically detailing a GQA option using FlashInfer draft attention. The review feedback correctly identifies that the suggested cudagraph_mode value of "FULL_AND_PIECEWISE" is invalid in vLLM and will cause a server crash, recommending a valid mode such as "PIECEWISE" instead.

Comment thread models/MiniMaxAI/MiniMax-M3.yaml Outdated

The NVFP4 variant already enables FP8 KV and FlashInfer target attention.
For variable batch sizes, also consider `--max-num-seqs` around 1.5× expected
concurrency and `--compilation-config '{"cudagraph_mode":"FULL_AND_PIECEWISE"}'`.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The cudagraph_mode value "FULL_AND_PIECEWISE" is not a valid mode in vLLM's CompilationConfig. The supported modes are typically "FULL", "PIECEWISE", "NONE", or "FULL_DECODE_ONLY". Using an invalid value will cause a validation error and crash the server on startup. Please correct this to a valid mode, such as "PIECEWISE" or "FULL".

  For variable batch sizes, also consider `--max-num-seqs` around 1.5× expected
  concurrency and `--compilation-config '{"cudagraph_mode":"PIECEWISE"}'`.

@xinli-sw
xinli-sw force-pushed the docs/minimax-m3-blackwell-agentic-options branch from 8ed0426 to 980e24e Compare September 29, 2026 21:26
@xinli-sw xinli-sw changed the title docs: add optional MiniMax-M3 Blackwell agentic tuning docs: add optional MiniMax-M3 GQA FlashInfer config Sep 29, 2026
Signed-off-by: Xin Li <xinli@nvidia.com>
@esmeetu

esmeetu commented Sep 30, 2026

Copy link
Copy Markdown
Member

lgtm

@esmeetu
esmeetu merged commit 5bcc249 into vllm-project:main Sep 30, 2026
4 checks passed

This branch was successfully deployed

1 active deployment
Preview — 0e29a2a7 Deployed Sep 29, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants