-
Notifications
You must be signed in to change notification settings - Fork 315
[Klaud Cold] minimaxm3-fp8-gb300-dynamo-vllm-mtp: day-zero GB300 MXFP8 EAGLE3 MTP + FULL_DECODE_ONLY CG / 新增 GB300 MXFP8 EAGLE3 MTP 配方,解码启用 FULL_DECODE_ONLY 图模式 #2486
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Merged
functionstackx
merged 4 commits into
main
from
feat/minimaxm3-fp8-gb300-dynamo-vllm-mtp-dayzero
Aug 4, 2026
Merged
Changes from 1 commit
Commits
Show all changes
4 commits
Select commit
Hold shift + click to select a range
4d74fd8
minimaxm3-fp8-gb300-dynamo-vllm-mtp: day-zero GB300 EAGLE3 MTP, FULL_…
xinli-sw d7f6ae6
fix: use compilation-config JSON for FULL_DECODE_ONLY cudagraph mode
xinli-sw 1f1b743
remove cutlass MSA and FULL_DECODE_ONLY from c1/c4/c8 configs
xinli-sw 0ae1f94
Merge remote-tracking branch 'origin/main' into pr-2486-merge
functionstackx File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
116 changes: 116 additions & 0 deletions
116
.../srt-slurm-recipes/vllm/minimax-m3-gb300-fp8/8k1k/mtp/1p1d-dep2-dep8-eagle3-c64-8k1k.yaml
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,116 @@ | ||
| name: "minimax-m3-vllm-disagg-gb300-1p1d-dep2-dep8-mxfp8-8k1k-eagle3-c64" | ||
|
|
||
| model: | ||
| path: "minimax-m3-mxfp8" | ||
| container: "vllm/vllm-openai:nightly-5e35a6f4f9bbc217c599692157ca985c894373f7" | ||
| precision: "fp8" | ||
|
|
||
| identity: | ||
| model: | ||
| repo: "MiniMaxAI/MiniMax-M3-MXFP8" | ||
| revision: "c5454eb03678d8710e54a4e0fc681b9f3b4a3dba" | ||
| container: | ||
| image: "vllm/vllm-openai:nightly-5e35a6f4f9bbc217c599692157ca985c894373f7" | ||
| frameworks: | ||
| dynamo: "1.4.0.dev20260730" | ||
| vllm: "0.26.1rc1.dev255+g5e35a6f4f" | ||
|
|
||
| dynamo: | ||
| install: true | ||
| version: "1.4.0.dev20260730" | ||
| request_plane: "nats" | ||
|
|
||
| health_check: | ||
| max_attempts: 720 | ||
| interval_seconds: 10 | ||
|
|
||
| sbatch_directives: | ||
| mem: "0" | ||
| cpus-per-task: "72" | ||
|
|
||
| srun_options: | ||
| mem: "0" | ||
|
|
||
| resources: | ||
| gpu_type: "gb300" | ||
| gpus_per_node: 4 | ||
| prefill_nodes: 1 | ||
| decode_nodes: 2 | ||
| prefill_workers: 1 | ||
| decode_workers: 1 | ||
| gpus_per_prefill: 2 | ||
| gpus_per_decode: 8 | ||
|
|
||
| frontend: | ||
| type: "dynamo" | ||
| enable_multiple_frontends: false | ||
|
|
||
| backend: | ||
| type: "vllm" | ||
| connector: null | ||
|
|
||
| prefill_environment: &worker-environment | ||
| VLLM_ENGINE_READY_TIMEOUT_S: "3600" | ||
| VLLM_FLOAT32_MATMUL_PRECISION: "high" | ||
| VLLM_FLASHINFER_ALLREDUCE_BACKEND: "mnnvl" | ||
| UCX_CUDA_IPC_ENABLE_MNNVL: "y" | ||
| UCX_MODULE_DIR: "/usr/local/lib/python3.12/dist-packages/nixl_cu13.libs/ucx" | ||
| UCX_RNDV_PIPELINE_ERROR_HANDLING: "y" | ||
| NCCL_CUMEM_ENABLE: "1" | ||
| NCCL_MNNVL_ENABLE: "1" | ||
| NCCL_NVLS_ENABLE: "1" | ||
|
|
||
| decode_environment: *worker-environment | ||
|
|
||
| vllm_config: | ||
| prefill: | ||
| no-enable-flashinfer-autotune: true | ||
| kv-transfer-config: '{"kv_connector": "NixlConnector", "kv_role": "kv_both"}' | ||
| attention-config: '{"backend": "FLASHINFER", "use_trtllm_attention": true, "indexer_kv_dtype": "fp8"}' | ||
| speculative-config: '{"method":"eagle3","model":"Inferact/MiniMax-M3-EAGLE3-GQA","num_speculative_tokens":3,"attention_backend":"FLASH_ATTN"}' | ||
| tensor-parallel-size: 1 | ||
| data-parallel-size: 2 | ||
| data-parallel-rpc-port: 13345 | ||
| enable-expert-parallel: true | ||
| trust-remote-code: true | ||
| no-enable-prefix-caching: true | ||
| block-size: 128 | ||
| gpu-memory-utilization: 0.90 | ||
| max-model-len: 9472 | ||
| language-model-only: true | ||
| kv-cache-dtype: "fp8" | ||
| stream-interval: 32 | ||
| max-cudagraph-capture-size: 2048 | ||
| max-num-batched-tokens: 16384 | ||
|
|
||
| decode: | ||
| no-enable-flashinfer-autotune: true | ||
| kv-transfer-config: '{"kv_connector": "NixlConnector", "kv_role": "kv_both"}' | ||
| attention-config: '{"backend": "FLASHINFER", "use_trtllm_attention": true, "indexer_kv_dtype": "fp8", "minimax_m3_msa_decode_backend": "cutlass"}' | ||
| speculative-config: '{"method":"eagle3","model":"Inferact/MiniMax-M3-EAGLE3-GQA","num_speculative_tokens":3,"attention_backend":"FLASH_ATTN"}' | ||
| tensor-parallel-size: 1 | ||
| data-parallel-size: 8 | ||
| data-parallel-rpc-port: 13345 | ||
| enable-expert-parallel: true | ||
| trust-remote-code: true | ||
| no-enable-prefix-caching: true | ||
| block-size: 128 | ||
| gpu-memory-utilization: 0.90 | ||
| max-model-len: 9472 | ||
| language-model-only: true | ||
| kv-cache-dtype: "fp8" | ||
| stream-interval: 32 | ||
| max-num-seqs: 1024 | ||
| max-num-batched-tokens: 16384 | ||
| max-cudagraph-capture-size: 2048 | ||
| cudagraph_mode: FULL_DECODE_ONLY | ||
|
|
||
| benchmark: | ||
| type: "sa-bench" | ||
| isl: 8192 | ||
| osl: 1024 | ||
| concurrencies: "64" | ||
| req_rate: "inf" | ||
| num_warmup_mult: 2 | ||
| random_range_ratio: 0.8 | ||
| use_chat_template: true |
114 changes: 114 additions & 0 deletions
114
...de/srt-slurm-recipes/vllm/minimax-m3-gb300-fp8/8k1k/mtp/1p1d-dep2-tp4-eagle3-c1-8k1k.yaml
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,114 @@ | ||
| name: "minimax-m3-vllm-disagg-gb300-1p1d-dep2-tp4-mxfp8-8k1k-eagle3-c1" | ||
|
|
||
| model: | ||
| path: "minimax-m3-mxfp8" | ||
| container: "vllm/vllm-openai:nightly-5e35a6f4f9bbc217c599692157ca985c894373f7" | ||
| precision: "fp8" | ||
|
|
||
| identity: | ||
| model: | ||
| repo: "MiniMaxAI/MiniMax-M3-MXFP8" | ||
| revision: "c5454eb03678d8710e54a4e0fc681b9f3b4a3dba" | ||
| container: | ||
| image: "vllm/vllm-openai:nightly-5e35a6f4f9bbc217c599692157ca985c894373f7" | ||
| frameworks: | ||
| dynamo: "1.4.0.dev20260730" | ||
| vllm: "0.26.1rc1.dev255+g5e35a6f4f" | ||
|
|
||
| dynamo: | ||
| install: true | ||
| version: "1.4.0.dev20260730" | ||
| request_plane: "nats" | ||
|
|
||
| health_check: | ||
| max_attempts: 720 | ||
| interval_seconds: 10 | ||
|
|
||
| sbatch_directives: | ||
| mem: "0" | ||
| cpus-per-task: "72" | ||
|
|
||
| srun_options: | ||
| mem: "0" | ||
|
|
||
| resources: | ||
| gpu_type: "gb300" | ||
| gpus_per_node: 4 | ||
| prefill_nodes: 1 | ||
| decode_nodes: 1 | ||
| prefill_workers: 1 | ||
| decode_workers: 1 | ||
| gpus_per_prefill: 2 | ||
| gpus_per_decode: 4 | ||
|
|
||
| frontend: | ||
| type: "dynamo" | ||
| enable_multiple_frontends: false | ||
|
|
||
| backend: | ||
| type: "vllm" | ||
| connector: null | ||
|
|
||
| prefill_environment: &worker-environment | ||
| VLLM_ENGINE_READY_TIMEOUT_S: "3600" | ||
| VLLM_FLOAT32_MATMUL_PRECISION: "high" | ||
| VLLM_FLASHINFER_ALLREDUCE_BACKEND: "mnnvl" | ||
| UCX_CUDA_IPC_ENABLE_MNNVL: "y" | ||
| UCX_MODULE_DIR: "/usr/local/lib/python3.12/dist-packages/nixl_cu13.libs/ucx" | ||
| UCX_RNDV_PIPELINE_ERROR_HANDLING: "y" | ||
| NCCL_CUMEM_ENABLE: "1" | ||
| NCCL_MNNVL_ENABLE: "1" | ||
| NCCL_NVLS_ENABLE: "1" | ||
|
|
||
| decode_environment: *worker-environment | ||
|
|
||
| vllm_config: | ||
| prefill: | ||
| no-enable-flashinfer-autotune: true | ||
| kv-transfer-config: '{"kv_connector": "NixlConnector", "kv_role": "kv_both"}' | ||
| attention-config: '{"backend": "FLASHINFER", "use_trtllm_attention": true, "indexer_kv_dtype": "fp8"}' | ||
| speculative-config: '{"method":"eagle3","model":"Inferact/MiniMax-M3-EAGLE3-GQA","num_speculative_tokens":3,"attention_backend":"FLASH_ATTN"}' | ||
| tensor-parallel-size: 1 | ||
| data-parallel-size: 2 | ||
| data-parallel-rpc-port: 13345 | ||
| enable-expert-parallel: true | ||
| trust-remote-code: true | ||
| no-enable-prefix-caching: true | ||
| block-size: 128 | ||
| gpu-memory-utilization: 0.90 | ||
| max-model-len: 9472 | ||
| language-model-only: true | ||
| kv-cache-dtype: "fp8" | ||
| stream-interval: 32 | ||
| max-cudagraph-capture-size: 2048 | ||
| max-num-batched-tokens: 16384 | ||
|
|
||
| decode: | ||
| no-enable-flashinfer-autotune: true | ||
| kv-transfer-config: '{"kv_connector": "NixlConnector", "kv_role": "kv_both"}' | ||
| attention-config: '{"backend": "FLASHINFER", "use_trtllm_attention": true, "indexer_kv_dtype": "fp8", "minimax_m3_msa_decode_backend": "cutlass"}' | ||
| speculative-config: '{"method":"eagle3","model":"Inferact/MiniMax-M3-EAGLE3-GQA","num_speculative_tokens":3,"attention_backend":"FLASH_ATTN"}' | ||
| tensor-parallel-size: 4 | ||
| enable-expert-parallel: false | ||
| trust-remote-code: true | ||
| no-enable-prefix-caching: true | ||
| block-size: 128 | ||
| gpu-memory-utilization: 0.90 | ||
| max-model-len: 9472 | ||
| language-model-only: true | ||
| kv-cache-dtype: "fp8" | ||
| stream-interval: 32 | ||
| max-num-seqs: 1024 | ||
| max-num-batched-tokens: 16384 | ||
| max-cudagraph-capture-size: 2048 | ||
| cudagraph_mode: FULL_DECODE_ONLY | ||
|
Check failure on line 104 in benchmarks/multi_node/srt-slurm-recipes/vllm/minimax-m3-gb300-fp8/8k1k/mtp/1p1d-dep2-tp4-eagle3-c1-8k1k.yaml
|
||
|
|
||
| benchmark: | ||
| type: "sa-bench" | ||
| isl: 8192 | ||
| osl: 1024 | ||
| concurrencies: "1" | ||
| req_rate: "inf" | ||
| num_warmup_mult: 2 | ||
| random_range_ratio: 0.8 | ||
| use_chat_template: true | ||
114 changes: 114 additions & 0 deletions
114
...de/srt-slurm-recipes/vllm/minimax-m3-gb300-fp8/8k1k/mtp/1p1d-dep2-tp4-eagle3-c8-8k1k.yaml
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,114 @@ | ||
| name: "minimax-m3-vllm-disagg-gb300-1p1d-dep2-tp4-mxfp8-8k1k-eagle3-c8" | ||
|
|
||
| model: | ||
| path: "minimax-m3-mxfp8" | ||
| container: "vllm/vllm-openai:nightly-5e35a6f4f9bbc217c599692157ca985c894373f7" | ||
| precision: "fp8" | ||
|
|
||
| identity: | ||
| model: | ||
| repo: "MiniMaxAI/MiniMax-M3-MXFP8" | ||
| revision: "c5454eb03678d8710e54a4e0fc681b9f3b4a3dba" | ||
| container: | ||
| image: "vllm/vllm-openai:nightly-5e35a6f4f9bbc217c599692157ca985c894373f7" | ||
| frameworks: | ||
| dynamo: "1.4.0.dev20260730" | ||
| vllm: "0.26.1rc1.dev255+g5e35a6f4f" | ||
|
|
||
| dynamo: | ||
| install: true | ||
| version: "1.4.0.dev20260730" | ||
| request_plane: "nats" | ||
|
|
||
| health_check: | ||
| max_attempts: 720 | ||
| interval_seconds: 10 | ||
|
|
||
| sbatch_directives: | ||
| mem: "0" | ||
| cpus-per-task: "72" | ||
|
|
||
| srun_options: | ||
| mem: "0" | ||
|
|
||
| resources: | ||
| gpu_type: "gb300" | ||
| gpus_per_node: 4 | ||
| prefill_nodes: 1 | ||
| decode_nodes: 1 | ||
| prefill_workers: 1 | ||
| decode_workers: 1 | ||
| gpus_per_prefill: 2 | ||
| gpus_per_decode: 4 | ||
|
|
||
| frontend: | ||
| type: "dynamo" | ||
| enable_multiple_frontends: false | ||
|
|
||
| backend: | ||
| type: "vllm" | ||
| connector: null | ||
|
|
||
| prefill_environment: &worker-environment | ||
| VLLM_ENGINE_READY_TIMEOUT_S: "3600" | ||
| VLLM_FLOAT32_MATMUL_PRECISION: "high" | ||
| VLLM_FLASHINFER_ALLREDUCE_BACKEND: "mnnvl" | ||
| UCX_CUDA_IPC_ENABLE_MNNVL: "y" | ||
| UCX_MODULE_DIR: "/usr/local/lib/python3.12/dist-packages/nixl_cu13.libs/ucx" | ||
| UCX_RNDV_PIPELINE_ERROR_HANDLING: "y" | ||
| NCCL_CUMEM_ENABLE: "1" | ||
| NCCL_MNNVL_ENABLE: "1" | ||
| NCCL_NVLS_ENABLE: "1" | ||
|
|
||
| decode_environment: *worker-environment | ||
|
|
||
| vllm_config: | ||
| prefill: | ||
| no-enable-flashinfer-autotune: true | ||
| kv-transfer-config: '{"kv_connector": "NixlConnector", "kv_role": "kv_both"}' | ||
| attention-config: '{"backend": "FLASHINFER", "use_trtllm_attention": true, "indexer_kv_dtype": "fp8"}' | ||
| speculative-config: '{"method":"eagle3","model":"Inferact/MiniMax-M3-EAGLE3-GQA","num_speculative_tokens":3,"attention_backend":"FLASH_ATTN"}' | ||
| tensor-parallel-size: 1 | ||
| data-parallel-size: 2 | ||
| data-parallel-rpc-port: 13345 | ||
| enable-expert-parallel: true | ||
| trust-remote-code: true | ||
| no-enable-prefix-caching: true | ||
| block-size: 128 | ||
| gpu-memory-utilization: 0.90 | ||
| max-model-len: 9472 | ||
| language-model-only: true | ||
| kv-cache-dtype: "fp8" | ||
| stream-interval: 32 | ||
| max-cudagraph-capture-size: 2048 | ||
| max-num-batched-tokens: 16384 | ||
|
|
||
| decode: | ||
| no-enable-flashinfer-autotune: true | ||
| kv-transfer-config: '{"kv_connector": "NixlConnector", "kv_role": "kv_both"}' | ||
| attention-config: '{"backend": "FLASHINFER", "use_trtllm_attention": true, "indexer_kv_dtype": "fp8", "minimax_m3_msa_decode_backend": "cutlass"}' | ||
| speculative-config: '{"method":"eagle3","model":"Inferact/MiniMax-M3-EAGLE3-GQA","num_speculative_tokens":3,"attention_backend":"FLASH_ATTN"}' | ||
| tensor-parallel-size: 4 | ||
| enable-expert-parallel: false | ||
| trust-remote-code: true | ||
| no-enable-prefix-caching: true | ||
| block-size: 128 | ||
| gpu-memory-utilization: 0.90 | ||
| max-model-len: 9472 | ||
| language-model-only: true | ||
| kv-cache-dtype: "fp8" | ||
| stream-interval: 32 | ||
| max-num-seqs: 1024 | ||
| max-num-batched-tokens: 16384 | ||
| max-cudagraph-capture-size: 2048 | ||
| cudagraph_mode: FULL_DECODE_ONLY | ||
|
|
||
| benchmark: | ||
| type: "sa-bench" | ||
| isl: 8192 | ||
| osl: 1024 | ||
| concurrencies: "8" | ||
| req_rate: "inf" | ||
| num_warmup_mult: 2 | ||
| random_range_ratio: 0.8 | ||
| use_chat_template: true |
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
🔴 All 11 new decode
vllm_configblocks in this PR setcudagraph_mode: FULL_DECODE_ONLYas a bare top-level key (e.g. line 104 in1p1d-dep2-tp4-eagle3-c1-8k1k.yaml, and the same line in the other 10 sibling files), but every other recipe in this repo enables that mode by embedding it inside thecompilation-configJSON blob (e.g.compilation-config: '{"cudagraph_mode":"FULL_DECODE_ONLY",...}'). Since srt-slurm passes top-levelvllm_configkeys through as--<key>CLI flags, this likely emits an unrecognized--cudagraph_modeflag (crashing decode server startup) or is silently dropped — either way defeating the PR's sole stated purpose of enabling FULL_DECODE_ONLY CUDA graphs on decode.Extended reasoning...
The bug: every one of the 11 new decode
vllm_configblocks added by this PR setscudagraph_mode: FULL_DECODE_ONLYas a bare, top-level recipe key sitting alongsidemax-cudagraph-capture-size,kv-cache-dtype, etc. For example, in1p1d-dep2-tp4-eagle3-c1-8k1k.yaml:Why this deviates from the established pattern: grepping the repo shows
cudagraph_modeis enabled in 100+ places across dozens of recipe files, and in every single one of them it is embedded inside acompilation-configJSON string, never set as a bare key. For instance, the siblingminimax-m3/b200-fp4/8k1k/*.yamlrecipes (same model family) do:The same pattern holds for the deepseek-v4 and kimi-k2.5-fp4 recipe families. This is because
cudagraph_modeis a field of vLLM'sCompilationConfig, not a standalone engine-args flag — there is no--cudagraph-mode/--cudagraph_modeCLI argument in vLLM's argparse; it can only be set via--compilation-config(JSON) or the-Odot-notation shorthand. That's precisely why every other recipe author routed it through the JSON blob instead of a bare key.How this breaks at runtime: the other 20+ keys in these same
vllm_configblocks (no-enable-flashinfer-autotune,kv-transfer-config,max-cudagraph-capture-size,tensor-parallel-size, etc.) are all kebab-case, matching srt-slurm's convention of passing top-levelvllm_configkeys through verbatim as--<key>CLI flags tovllm serve.cudagraph_modeis the only snake_case key in these files — it looks like the author copy-pasted the JSON field name but forgot to wrap it insidecompilation-config. Under that passthrough convention, this key gets emitted as--cudagraph_mode FULL_DECODE_ONLY. Since vLLM has no such CLI flag, this either (a) is rejected by argparse as an unrecognized argument, crashing decode-server startup, or (b) is silently dropped by the recipe-to-CLI translation, in which case FULL_DECODE_ONLY is simply never applied.Step-by-step proof:
compilation-config: '{"cudagraph_mode":"FULL_DECODE_ONLY",...}'— confirmed by grep across 70+ files, including the same-modelminimax-m3/b200-fp4sibling recipes.cudagraph_mode: FULL_DECODE_ONLYkey at decode-block scope (e.g. line 104 of1p1d-dep2-tp4-eagle3-c1-8k1k.yaml), with none of the 11 decode blocks touchingcompilation-configat all.vllm_configkeys to--<key>vLLM CLI flags 1:1 (evidenced by every other kebab-case key in the same block mapping directly to a real vLLM flag).--cudagraph-mode/--cudagraph_modeengine argument — that field only exists insideCompilationConfig, settable via--compilation-configJSON.The fix: merge the key into the decode
compilation-configJSON blob, consistent with every other recipe, e.g. addcompilation-config: '{"cudagraph_mode":"FULL_DECODE_ONLY"}'to each of the 11 decode blocks (or fold it into an existingcompilation-configentry if one is later added) instead of the current barecudagraph_mode: FULL_DECODE_ONLYkey.