Conversation
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
There was a problem hiding this comment.
Code Review
This pull request updates the MiniMax-M3 model metadata and adds configuration instructions for B200/B300 agentic serving with the NVFP4 variant, specifically detailing a GQA option using FlashInfer draft attention. The review feedback correctly identifies that the suggested cudagraph_mode value of "FULL_AND_PIECEWISE" is invalid in vLLM and will cause a server crash, recommending a valid mode such as "PIECEWISE" instead.
|
|
||
| The NVFP4 variant already enables FP8 KV and FlashInfer target attention. | ||
| For variable batch sizes, also consider `--max-num-seqs` around 1.5× expected | ||
| concurrency and `--compilation-config '{"cudagraph_mode":"FULL_AND_PIECEWISE"}'`. |
There was a problem hiding this comment.
The cudagraph_mode value "FULL_AND_PIECEWISE" is not a valid mode in vLLM's CompilationConfig. The supported modes are typically "FULL", "PIECEWISE", "NONE", or "FULL_DECODE_ONLY". Using an invalid value will cause a validation error and crash the server on startup. Please correct this to a valid mode, such as "PIECEWISE" or "FULL".
For variable batch sizes, also consider `--max-num-seqs` around 1.5× expected
concurrency and `--compilation-config '{"cudagraph_mode":"PIECEWISE"}'`.8ed0426 to
980e24e
Compare
Signed-off-by: Xin Li <xinli@nvidia.com>
980e24e to
0e29a2a
Compare
|
lgtm |
Adds an optional GQA example: FlashInfer draft attention when it benefits the workload, plus local argmax reduction to reduce communication traffic.
Validation: YAML and example config parse successfully.