Skip to content
Merged
Show file tree
Hide file tree
Changes from 3 commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
28 changes: 15 additions & 13 deletions .github/configs/nvidia-master.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -12919,7 +12919,7 @@ minimaxm3-fp8-b200-vllm:
# (MSA sparse/index cache); weights are pre-staged at /scratch/fsw/models/MiniMax-M3-NVFP4
# (launch_b200-dgxc.sh resolves MODEL_PATH for minimaxm3-fp4).
minimaxm3-fp4-b200-vllm:
image: vllm/vllm-openai:vllm-minimax-m3-perf-x86_64-13.0.1-8b00f41
image: vllm/vllm-openai:nightly-93d8f834dd8acf33eb0e2a75b2711b628cb6e226
model: nvidia/MiniMax-M3-NVFP4
model-prefix: minimaxm3
runner: b200-dgxc
Expand All @@ -12931,21 +12931,23 @@ minimaxm3-fp4-b200-vllm:
- isl: 1024
osl: 1024
search-space:
- { tp: 8, conc-start: 1, conc-end: 64 }
- { tp: 8, ep: 8, conc-start: 1, conc-end: 512 }
- { tp: 4, conc-start: 1, conc-end: 64 }
- { tp: 4, ep: 4, conc-start: 64, conc-end: 512 }
- { tp: 4, ep: 4, dp-attn: true, conc-start: 128, conc-end: 512 }
- { tp: 8, ep: 8, dp-attn: true, conc-start: 256, conc-end: 1024 }
- { tp: 8, conc-start: 1, conc-end: 8 }
- { tp: 4, conc-start: 1, conc-end: 256 }
- { tp: 2, conc-start: 1, conc-end: 256 }
- { tp: 2, ep: 2, conc-start: 1, conc-end: 256 }
- { tp: 4, ep: 4, conc-start: 64, conc-end: 1024 }
- { tp: 4, ep: 4, dp-attn: true, conc-start: 64, conc-end: 1024 }
- { tp: 2, ep: 2, dp-attn: true, conc-start: 64, conc-end: 256 }
- isl: 8192
osl: 1024
search-space:
- { tp: 8, conc-start: 1, conc-end: 64 }
- { tp: 8, ep: 8, conc-start: 1, conc-end: 256 }
- { tp: 4, conc-start: 1, conc-end: 64 }
- { tp: 4, ep: 4, conc-start: 64, conc-end: 256 }
- { tp: 4, ep: 4, dp-attn: true, conc-start: 64, conc-end: 128 }
- { tp: 8, ep: 8, dp-attn: true, conc-start: 128, conc-end: 256 }
- { tp: 8, conc-start: 1, conc-end: 8 }
- { tp: 4, conc-start: 1, conc-end: 256 }
- { tp: 2, conc-start: 1, conc-end: 256 }
- { tp: 2, ep: 2, conc-start: 32, conc-end: 512 }
- { tp: 4, ep: 4, conc-start: 128, conc-end: 512 }
- { tp: 4, ep: 4, dp-attn: true, conc-start: 256, conc-end: 512 }
- { tp: 2, ep: 4, dp-attn: true, conc-start: 64, conc-end: 256 }

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 The new 8k1k entry { tp: 2, ep: 4, dp-attn: true, conc-start: 64, conc-end: 256 } at line 12950 mislabels EP: with dp-attn: true, minimaxm3_fp4_b200.sh launches vLLM with --tensor-parallel-size=1 --data-parallel-size=$TP --enable-expert-parallel and never passes EP_SIZE, so world_size = TP×1 = 2 and vLLM runs with EP=2 — but process_result.py reads EP_SIZE from env and records the result as ep=4, poisoning cross-config comparisons on the leaderboard. It is also the only entry in the file where ep > tp and ep != 1; every other dp-attn row has ep == tp. Likely a typo — either tp: 4, ep: 4, dp-attn: true (matching the sibling row on line 12949) or tp: 2, ep: 2, dp-attn: true.

Extended reasoning...

What's wrong

Line 12950 adds:

- { tp: 2, ep: 4, dp-attn: true, conc-start: 64, conc-end: 256 }

The ep: 4 label on this row is a lie: the actual runtime will use EP=2. Here is the code path.

Trigger path

1. Shell launcher — benchmarks/single_node/fixed_seq_len/minimaxm3_fp4_b200.sh:37-43:

if [ "${DP_ATTENTION}" = "true" ]; then
  PARALLEL_ARGS="--tensor-parallel-size=1 --data-parallel-size=$TP --enable-expert-parallel"
elif [ "$EP_SIZE" -gt 1 ]; then
  PARALLEL_ARGS="--tensor-parallel-size=$TP --enable-expert-parallel"
else
  PARALLEL_ARGS="--tensor-parallel-size=$TP"
fi

When DP_ATTENTION=true, EP_SIZE is used only for the elif's boolean check — it is never forwarded to vLLM via --expert-parallel-size. vLLM derives EP from world_size = data-parallel-size × tensor-parallel-size = TP × 1.

2. Slurm allocation — runners/launch_b200-dgxc.sh:435 requests --gres=gpu:$TP = 2 GPUs. Only 2 physical GPUs are allocated, so EP=4 is physically impossible on this node.

3. Result labeling — utils/process_result.py:110-118 reads EP_SIZE directly from env and records it as the ep field. utils/matrix_logic/generate_sweep_configs.py:483-484 propagates the YAML ep into that env var. Nothing corrects for the dp-attn case, so the label survives verbatim.

Step-by-step proof for this entry

  1. Sweep generator emits an env with TP=2, EP_SIZE=4, DP_ATTENTION=true.
  2. Slurm allocates 2 GPUs (gpu:$TP).
  3. Shell hits the first branch → launches vLLM with --tensor-parallel-size=1 --data-parallel-size=2 --enable-expert-parallel. EP_SIZE=4 is never referenced.
  4. vLLM's world_size = 2×1 = 2, so it uses EP=2.
  5. Benchmark completes. process_result.py reads EP_SIZE from env → writes ep: 4 into the result JSON.
  6. Leaderboard ingests an ep=4 row that was actually an EP=2 run.

Why validation doesn't catch this

utils/matrix_logic/validation.py SingleNodeSearchSpaceEntry permits ep independent of tp; there is no ep <= tp or ep == tp for dp-attn constraint. So this passes generate-sweep unnoticed.

Uniqueness / typo signal

Across the entire nvidia-master.yaml, this is the only entry with ep > tp and ep != 1. Every other dp-attn: true row has ep == tp (2/2, 4/4, 8/8). Combined with the row immediately above (tp: 4, ep: 4, dp-attn: true), the natural read is a typo where either tp should be 4 or ep should be 2.

Fix

One-character YAML edit — pick whichever config the author actually intended:

- { tp: 4, ep: 4, dp-attn: true, conc-start: 64, conc-end: 256 }   # matches row 12949 pattern
# or
- { tp: 2, ep: 2, dp-attn: true, conc-start: 64, conc-end: 256 }   # matches every other dp-attn row


# EAGLE3 speculative-decoding (spec-decoding: mtp) variant of
# minimaxm3-fp8-b200-vllm, pairing MiniMaxAI/MiniMax-M3-MXFP8 with the
Expand Down
7 changes: 7 additions & 0 deletions perf-changelog.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -4400,3 +4400,10 @@
description:
- "Bump SGLang image from lmsysorg/sglang:deepseek-v4-blackwell (digest sha256:df18bfc4...) to mainline nightly lmsysorg/sglang:nightly-dev-cu13-20260628-da802ddc."
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/1923

- config-keys:
- minimaxm3-fp4-b200-vllm
description:
- "Update Minimax M3 b200 vllm image tag"
- "Update search space to cover more configs"
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/1978
Loading