Tune MiniMax-M3 AgentX on B200 with updated vLLM / 更新 B200 上的 MiniMax-M3 AgentX 配置 - #3435
Conversation
|
Thanks for the contribution!
中文感谢你的贡献!
|
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=36470674070 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=36470674070 |
There was a problem hiding this comment.
I reviewed this PR and didn't find any bugs. Because it removes an engine-source patch on the strength of a new pinned vLLM nightly, a human maintainer's look would still be worthwhile.
What was reviewed: confirmed PORT/EVAL_ONLY are genuinely supplied by the b200-nscale and b300-dsxe launch scripts (not caller-invented defaults) before being added to check_env_vars; verified the CPU_OFFLOAD_BYTES switch from 1024^3 to 1000^3 matches infx/matrix/generate.py's BYTES_PER_GB=1_000_000_000 decimal convention; checked the derived MAX_NUM_SEQS/MAX_CUDAGRAPH_CAPTURE_SIZE formula against the same pattern already used in the sibling minimaxm3_fp8_h100/h200 scripts; confirmed the perf-changelog.yaml entry only appends (prior bytes untouched) and its "XXX" pr-link placeholder is filled in later by infx/workflows/prepare_perf_changelog_merge.py, matching existing convention.
Extended reasoning...
The diff tunes two single-node vLLM benchmark scripts and their master-config entries (GPU/host KV utilization, batching/concurrency, image pin) and deletes an obsolete monkey-patch of vLLM's installed SimpleCPUOffloadConnector source, with a matching perf-changelog entry. No auth/crypto/permission surface is touched; it's config and shell-script arithmetic. I independently verified the env-var validation, the binary-to-decimal GB unit fix, and the derived cudagraph/max-num-seqs formula against an existing sibling script's precedent, all of which check out. The one thing I could not verify in this environment (no network access) is the PR's core factual claim that vLLM nightly-29468dde fixes the heterogeneous-KV-layer-region bug the deleted patch worked around rather than silently reintroducing it — that needs a maintainer with vLLM upstream context or a live run to confirm.
This review covers commit 89d5ec0, which is no longer the latest commit on this pull request; later commits are not covered by it.
* feat(srt): run single-node AgentX on native srt-slurm Single-node AgentX points with an srt-recipe now take the native single-node path fixed-sequence points already use. The adapter accepts AgentX points and DSpark speculation. Single- and multi-node AgentX share one client, benchmarks/srt_agentic.sh (moved from benchmarks/multi_node/agentic_srt.sh), which replays one CONC point or a CONC_LIST batch against the srt-slurm frontend and reads engine metrics from the workers behind a router frontend. The H200 DeepSeek-V4.1-Flash SGLang config is the first one ported. * fix(amd): drop a duplicate DSV4 prefill key and restore the MI355X Qwen3.5 fixed-sequence recipe srt-slurm parses recipes strictly and rejected the repeated disable-cuda-graph key. Removing the unported MI355X Qwen3.5 AgentX recipe had also deleted the fixed-sequence recipe beside it. * feat(srt): accept vLLM points in the single-node adapter vLLM points validate their topology as tensor x data parallel GPUs, with DP attention as data-parallel ranks and expert parallelism as enable-expert-parallel, and eval-only runs set max-model-len. * feat(agentx): port the DSV4.1 Flash B300 SGLang AgentX config to srt-slurm * feat(agentx): let srt-slurm AgentX recipes apply the chat template client-side Several legacy AMD AgentX scripts appended --apply-chat-template to the replay command. AIPERF_APPLY_CHAT_TEMPLATE=true in a recipe's benchmark env keeps that behavior on the shared srt-slurm client. * feat(agentx): run GB200/GB300 single-node AgentX natively and port DSV4.1 Flash SGLang there GB launchers submit recipe points through launch_srt_single_node with an aarch64 srt-slurm setup; the squash is imported on a compute tray first. * feat(agentx): let srt-slurm AgentX recipes apply the chat template client-side Several legacy AMD AgentX scripts appended --apply-chat-template to the replay command. AIPERF_APPLY_CHAT_TEMPLATE=true in a recipe's benchmark env keeps that behavior on the shared srt-slurm client. * feat(agentx): let srt-slurm AgentX recipes apply the chat template client-side Several legacy AMD AgentX scripts appended --apply-chat-template to the replay command. AIPERF_APPLY_CHAT_TEMPLATE=true in a recipe's benchmark env keeps that behavior on the shared srt-slurm client. * feat(agentx): let srt-slurm AgentX recipes set the AIPerf benchmark grace period The GLM-5.2 MI325X legacy script bounded the post-window drain with --benchmark-grace-period 1800. AIPERF_BENCHMARK_GRACE_PERIOD in a recipe's benchmark env passes it through the shared client. * feat(agentx): port the DSV4.1 Flash B200 vLLM AgentX config to srt-slurm * feat(srt): force ATOM AgentX golden acceptance with its server flag ATOM pins acceptance with --spec-decode-acceptance-length rather than an environment variable. The adapter now sets it from the golden curve for AgentX throughput and removes it for evals, reading the draft model for the MiniMax GQA curve and the probabilistic sampler for Kimi DSpark. * feat(srt): accept draft_model speculation in the single-node adapter DeepSeek-V4 configs mark their bundled DSpark draft as draft_model; the recipe still speculates natively, so the point binds as speculative. * fix(mi355x): mount the shared HF cache for native AgentX checkpoints The legacy MI355X AgentX path read DeepSeek-V4-Pro (vLLM/ATOM), DeepSeek-V4.1 Flash, MiniMax-M3 and GLM-5.2-FP8 from /it-share rather than node-local NVMe. The native srt-slurm path keeps that mount for those AgentX checkpoints. * feat(agentx): port the Qwen3.5 MI300X SGLang AgentX config to srt-slurm * feat(agentx): port the Qwen3.5 MI325X SGLang AgentX config to srt-slurm * feat(agentx): port the GLM-5.2 MI325X SGLang AgentX config to srt-slurm * feat(agentx): port the DSV4.1 Flash H100 SGLang AgentX configs to srt-slurm * feat(agentx): port the Qwen3.5 H100 SGLang AgentX configs to srt-slurm * feat(agentx): port the Qwen3.5 H200 SGLang HiCache EP1 AgentX config to srt-slurm * feat(srt): accept the draft_model label for single-node DSpark points * feat(agentx): port the DSV4.1 Flash MI300X vLLM AgentX config to srt-slurm * feat(agentx): port the DSV4.1 Flash MI325X vLLM AgentX config to srt-slurm * fix(srt): bind single-node draft_model and EAGLE3 AgentX points The DSV4 B300 DSpark config labels its points spec-decoding: draft_model and the MiniMax-M3 TRT-LLM config drafts with EAGLE3; accept both as speculative recipes. Golden acceptance already resolves both curves. * feat(srt): accept vLLM EAGLE3 speculation in the single-node adapter * feat(srt): accept EAGLE3 speculation in the single-node adapter MiniMax-M3 vLLM AgentX recipes speculate with the EAGLE3 GQA draft, which the golden acceptance lookup already maps to minimaxm3_eagle3_gqa. * feat(srt): accept EAGLE3 speculation in the single-node adapter MiniMax-M3 vLLM AgentX recipes speculate with the EAGLE3 GQA draft, which the golden acceptance lookup already maps to minimaxm3_eagle3_gqa. * feat(agentx): port the MiniMax-M3 MI325X vLLM AgentX config to srt-slurm * feat(agentx): port the Qwen3.5 FP4 B300 SGLang AgentX config to srt-slurm * feat(agentx): port the Qwen3.5 FP8 B300 SGLang AgentX config to srt-slurm * feat(agentx): port the GLM-5.2 FP4 B300 SGLang AgentX config to srt-slurm * feat(agentx): port the GLM-5.2 FP8 B300 SGLang AgentX config to srt-slurm * feat(agentx): port the DSV4 Pro B300 SGLang AgentX config to srt-slurm * feat(agentx): port the MiniMax-M3 B300 TRT-LLM AgentX config to srt-slurm * fix(b300): serve Qwen3.8-Flash-Next NVFP4 natively from the writable model root The checkpoint is not staged on node-local NVMe; the legacy AgentX path downloads it to the shared writable model directory. * feat(agentx): port the Qwen3.8-Flash-Next FP4 B300 SGLang AgentX config to srt-slurm * feat(srt): accept draft_model speculation in the single-node adapter DeepSeek-V4 configs mark their bundled DSpark draft as draft_model; the recipe still speculates natively, so the point binds as speculative. * feat(srt): force ATOM AgentX golden acceptance with its server flag ATOM pins acceptance with --spec-decode-acceptance-length rather than an environment variable. The adapter now sets it from the golden curve for AgentX throughput and removes it for evals, reading the draft model for the MiniMax GQA curve and the probabilistic sampler for Kimi DSpark. * fix(mi355x): mount the shared HF cache for native AgentX checkpoints The legacy MI355X AgentX path read DeepSeek-V4-Pro (vLLM/ATOM), DeepSeek-V4.1 Flash, MiniMax-M3 and GLM-5.2-FP8 from /it-share rather than node-local NVMe. The native srt-slurm path keeps that mount for those AgentX checkpoints. * feat(srt): accept draft-model labeled speculation in the single-node adapter The DSV4 B200 vLLM AgentX matrix labels its native DSpark drafter draft_model rather than mtp. * feat(agentx): port the DSV4.1 Flash H100 and H200 vLLM AgentX configs to srt-slurm * feat(agentx): port the MiniMax-M3 H100 and H200 vLLM AgentX configs to srt-slurm * feat(agentx): port the Qwen3.5 MI355X SGLang AgentX config to srt-slurm * feat(agentx): port the DSV4.1 Flash MI355X SGLang AgentX config to srt-slurm * feat(agentx): port the GLM-5.2 FP4 MI355X SGLang AgentX config to srt-slurm * feat(agentx): port the GLM-5.2 FP8 MI355X SGLang AgentX config to srt-slurm * feat(agentx): port the DSV4 MI355X SGLang AgentX config to srt-slurm * feat(agentx): port the DSV4 MI355X ATOM AgentX config to srt-slurm * feat(agentx): port the GLM-5.2 MI355X ATOM AgentX config to srt-slurm The DCP4 LMCache entry stays on the legacy script: srtctl reserves ATOM's kv-transfer-config for disaggregated workers. * feat(agentx): port the Kimi-K3 MI355X ATOM AgentX config to srt-slurm The DCP8 LMCache entries stay on the legacy script: srtctl reserves ATOM's kv-transfer-config for disaggregated workers. * feat(agentx): port the MiniMax-M3 MI355X ATOM AgentX config to srt-slurm The LMCache entries stay on the legacy script: srtctl reserves ATOM's kv-transfer-config for disaggregated workers. * feat(srt): accept EAGLE3 speculation in the single-node adapter * feat(agentx): port the DSV4.1 Flash B300 vLLM AgentX config to srt-slurm * feat(agentx): port the MiniMax-M3 B200 and B300 vLLM AgentX configs to srt-slurm * feat(agentx): port the DSV4 B200 and B300 vLLM AgentX configs to srt-slurm * feat(agentx): port the Qwen3.5 FP4 B200 SGLang AgentX config to srt-slurm * feat(agentx): port the Qwen3.5 FP8 B200 SGLang AgentX config to srt-slurm * feat(agentx): port the Qwen3.8 Next FP4 B200 SGLang AgentX config to srt-slurm * fix(h100): point native single-node uv caches at shared NFS * feat(agentx): port the GLM-5.2 FP4 B200 SGLang AgentX config to srt-slurm * feat(agentx): port the GLM-5.2 FP8 B200 SGLang AgentX config to srt-slurm * feat(agentx): port the DSV4 FP4 B200 SGLang AgentX config to srt-slurm * feat(agentx): port the MiniMax-M3 FP4 B200 TRT-LLM AgentX config to srt-slurm * feat(agentx): let srt-slurm AgentX recipes set the AIPerf benchmark grace period The GLM-5.2 MI325X legacy script bounded the post-window drain with --benchmark-grace-period 1800. AIPERF_BENCHMARK_GRACE_PERIOD in a recipe's benchmark env passes it through the shared client. * feat(agentx): port the DSV4.1 Flash GB200 and GB300 vLLM AgentX configs to srt-slurm * feat(agentx): port the MiniMax-M3 MI300X vLLM AgentX config to srt-slurm The LMCache DRAM point starts one MP server per TP rank from a setup script inside the worker container, as the legacy script did. * feat(agentx): port the DSV4.1 Flash MI355X vLLM AgentX config to srt-slurm * feat(agentx): port the MiniMax-M3 MI355X vLLM AgentX config to srt-slurm * feat(srt): accept vLLM decode context parallelism and non-drafting points in the single-node adapter Kimi-K3 B300 serves TP8 with decode-context-parallel-size 8, and stops drafting above conc 16 while its matrix labels every point mtp; such a variant declares SPEC_DECODING in its benchmark env. * feat(agentx): port the Kimi-K3 B300 vLLM AgentX config to srt-slurm * feat(agentx): port the Kimi-K3 MI355X vLLM AgentX DCP1 points to srt-slurm The DCP8 no-draft points keep the legacy script: the single-node adapter requires DCP_SIZE=1 and speculation matching the matrix label. * fix(matrix): read node counts from named srt-slurm override variants recipe_node_count returned None for any CONFIG_FILE carrying a selector, so rows pointing at `file.yaml:override_<name>` fell back to the master topology estimate. Resolve `base` and `override_<name>` the way srtctl does (deep merge, null deletes, top-level schema carried into the variant) so the recipe allocation stays authoritative. Zip groups and non-schema-2 variant files keep the estimate. * refactor(agentx): consolidate DSV4 multi-node AgentX recipes into override variants Each per-configuration recipe becomes an `override_<name>` block over a shared `base` in one `*-variants.yaml` per master-config entry, and the master entries select it with `CONFIG_FILE=...:override_<name>`. Every selected variant resolves, through the pinned srtctl, to exactly the recipe it replaces, including its original `name`. Power recipes with top-level telemetry stay standalone because launchers detect them as text. * refactor(agentx): consolidate GLM-5.2 multi-node AgentX recipes into override variants Each per-configuration recipe becomes an `override_<name>` block over a shared `base` in one `*-variants.yaml` per master-config entry, and the master entries select it with `CONFIG_FILE=...:override_<name>`. Every selected variant resolves, through the pinned srtctl, to exactly the recipe it replaces, including its original `name`. Power recipes with top-level telemetry stay standalone because launchers detect them as text. * refactor(agentx): consolidate Kimi-K3 GB200 AgentX recipes into override variants Each per-configuration recipe becomes an `override_<name>` block over a shared `base` in one `*-variants.yaml` per master-config entry, and the master entries select it with `CONFIG_FILE=...:override_<name>`. Every selected variant resolves, through the pinned srtctl, to exactly the recipe it replaces, including its original `name`. Power recipes with top-level telemetry stay standalone because launchers detect them as text. * refactor(agentx): consolidate MiniMax-M3 multi-node AgentX recipes into override variants Each per-configuration recipe becomes an `override_<name>` block over a shared `base` in one `*-variants.yaml` per master-config entry, and the master entries select it with `CONFIG_FILE=...:override_<name>`. Every selected variant resolves, through the pinned srtctl, to exactly the recipe it replaces, including its original `name`. Power recipes with top-level telemetry stay standalone because launchers detect them as text. * refactor(agentx): consolidate Qwen3.5 multi-node AgentX recipes into override variants Each per-configuration recipe becomes an `override_<name>` block over a shared `base` in one `*-variants.yaml` per master-config entry, and the master entries select it with `CONFIG_FILE=...:override_<name>`. Every selected variant resolves, through the pinned srtctl, to exactly the recipe it replaces, including its original `name`. Power recipes with top-level telemetry stay standalone because launchers detect them as text. * docs(recipes): describe AgentX override-variant bundles * feat(agentx): port the DSV4 MI355X vLLM AgentX config to srt-slurm The DEP8 points run behind srt-slurm's vLLM Router frontend with consistent-hash session routing, as the legacy script's router did. * fix(agentx): give the MiniMax-M3 Hopper Mooncake master time to install * fix(agentx): give the AMD SGLang AgentX recipes an hour to become healthy The MI300X Qwen3.5 server needed ~27 minutes to load and finish first-request kernel tuning, past srt-slurm's 1800 s default. * fix(agentx): extend the DSV4.1 Flash H100 SGLang health window for cold NFS loads * fix(agentx): extend the MiniMax-M3 H100 vLLM load window for cold NFS loads * fix(h100): keep one uv cache per runner for native single-node jobs * fix(b300): give srt-slurm jobs the workflow time limit instead of srtctl's one hour The B300 profile set no default_time_limit, so srtctl submitted native single-node AgentX jobs with --time=01:00:00 and a c32 point timed out. * fix(agentx): report RDMA port states when the Kimi-K3 B300 Mooncake rail probe fails * fix(mi355x): let native single-node points use a squash staged on /it-share Node-local /var/lib/squash is not visible from the runner host, so native points always pulled the image, which fails once a nightly tag is pruned. * fix(agentx): probe DSXE rdmap RDMA rails for the Kimi-K3 B300 Mooncake store * fix(srt): let single-node AgentX evals run without a fixed-sequence context * fix(srt): evaluate single-node AgentX points with the workflow's framework at native context The single-node post-eval required MAX_MODEL_LEN and forced lm-eval, but benchmark_lib clears MAX_MODEL_LEN for AgentX, so every AgentX eval-only point failed. AgentX now runs run_eval as the multi-node post-eval does. * fix(b300): download Qwen3.8-Flash-Next NVFP4 into the shared HF cache The checkpoint is staged neither on node-local NVMe nor in the writable model root, so native jobs resolve it by HF id like the other launchers. * fix(srt): stage single-node AgentX eval artifacts once * fix(srt): stage single-node AgentX eval artifacts once * fix(b300): read models without node-local staging from the shared model root DeepSeek-V4.1-Flash is not in STAGED_MODELS and exists on /scratch only on some nodes, so native points failed wherever it was missing. * fix(agentx): give GB200/GB300 DSV4.1 Flash servers a two-hour health budget Cold weight loading from the shared HF cache plus graph capture took just under 30 minutes on GB200 TP2, past srt-slurm's 1800 s default. * chore(agentx): keep the pruned-image MI355X vLLM configs on their legacy scripts Their nightly images are gone from Docker Hub; the legacy path still runs from node-local squashes. The recipes stay in place for when a squash is staged on /it-share. * fix(agentx): give the DSV4.1 Flash GB200 and GB300 vLLM recipes a two-hour readiness window * fix(agentx): let Mooncake pick the GID on DSXE InfiniBand rails for Kimi-K3 B300 * fix(agentx): give the MiniMax-M3 B300 TRT-LLM server a two-hour health budget * style(agentx): write srt-slurm recipe overrides as block YAML Content is unchanged; flow mappings in override variants become block mappings. vLLM compilation/kv-transfer/speculative configs stay quoted JSON strings, which srtctl passes to vLLM verbatim. * fix(agentx): give the Qwen3.8-Flash-Next B300 SGLang server a two-hour health budget * fix(srt): keep GLM-5.2's 150-step SWE-bench budget for single-node AgentX evals The legacy GLM-5.2 scripts raised SWEBENCH_AGENT_STEP_LIMIT to 150 for eval-only runs. srt-slurm's post-eval only sees the runner's environment, so the recipe cannot set it; srt_eval.sh does. * fix(agentx): give the MiniMax-M3 B200 TRT-LLM server a two-hour health budget * fix(b200): default the Qwen3.8-Flash-Next NVFP4 path when no pool setting names it * chore(minimaxm3): drop the vLLM SimpleCPUOffload patch The MiniMax-M3 B200/B300 DRAM points no longer patch vLLM's SimpleCPUOffload worker. The pinned nightly is pruned, so these points cannot run until the image moves to a build that no longer needs the patch (see #3131, #3132, #3435). * style(srt): format single-node adapter 按 Ruff 格式化单节点适配器。 --------- Co-authored-by: adibarra <93070681+adibarra@users.noreply.github.com>
|
InferenceX has switched away from unmaintainable bash scripts to YAML files that don't repeat the same stuff over and over again. Please merge the latest |
|
Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding |
baa8b19 to
da2ed2b
Compare
更新 B200 MiniMax-M3 AgentX 的 vLLM 镜像,并为 EAGLE3 草稿启用本地 argmax 归约。
da2ed2b to
7cbda30
Compare
|
Hi @Ankur-singh , this submission is ready |
|
/use 36470674070 |
kedarpotdar-nv
left a comment
There was a problem hiding this comment.
As a PR reviewer and CODEOWNER, I have reviewed this and have:
- Verified that as of the moment of typing this, this is the latest version of PR_REVIEW_CHECKLIST.md
- Verified that the general code quality meets the InferenceX standard and does not make the code quality any worse.
- Verified that this PR has passed PR validation. Please link to GitHub Action workflow that shows this.
- Verified that this PR passes evals. Please link to GitHub Action workflow that shows this.
- Verified that speculative decoding PRs uses chat templates to align the AL distribution to real world
- Verified that every draft model and draft head is served as it ships: the draft that ships with the served checkpoint, at its stored precision, through the pinned upstream image's default handling, with the shipped and effective draft precision recorded in the additional detail section. No submission-side quantization, dtype override, checkpoint substitution, or patch may lower draft precision below that default, regardless of eval results or AL. Explicitly verified that
SGLANG_NVFP4_CKPT_FP8_NEXTN_MOEis not enabled in the effective recipe, including inherited settings; enabling it is prohibited going forward, and historical runs do not grant an exception. See Draft-model precision for what counts as the default and the MLPerf comparison. - For agentic workloads: verified that speculative-decoding configs (EAGLE / MTP / draft models) run with simulated synthetic acceptance, with the acceptance-length value taken from the committed golden AL curve in infx/golden_al_distribution/ for that model, thinking mode, and draft length. A submission may choose any supported draft length, but it may not substitute a different acceptance target.
- Verified against the current MODELS.md that this PR does not submit a deprecated model, scenario, or model-scenario combination.
- Verified that the model architecture isn't changed with benchmark hacks like using --hf-overrides to skipping indexer for every x layers on models that don't natively support this. As a general rule, we won't accept optimizations that reduces the number of model architecture FLOPs. Anything that makes that same computation run faster is fair game; target/verifier FLOPs at lower precisions is fine, given that the config passes private evals, but this does not permit lowering draft-model or draft-head precision below what ships. As an general north star princple, we should only use optimizations which is used in production by customers that care about accuracy
- If an company claims that they support vLLM/SGLang as first class LLM inference engines on their hardware, I have verified that the respective vLLM submission made using upstream https://hub.docker.com/u/vllm docker repo, upstream SGLang https://hub.docker.com/u/lmsysorg docker repo. The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet as supported by vLLM/SGLang community maintainers
- If an company claims that they support vLLM/SGLang as first class upstream in-tree LLM inference engines on their hardware, I have have verified that the respective vLLM/SGLang submission has been made before additional frameworks (TRT-LLM, ATOM, etc.). The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL8, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet.
- Verified that every single-node vLLM/SGLang recipe in this PR is documented in the official vLLM recipes and/or the SGLang cookbook:
- I linked the corresponding upstream PR in the vLLM recipe repo or SGLang repo and verified that it is MERGED before this InferenceX PR merges. An opened, draft, or closed-without-merge upstream PR does not satisfy this requirement. If the matching recipe was already published, I linked the published recipe/cookbook page in the additional detail section below.
- Verified that this PR does not patch the inference engine or serving stack — the pinned image must run as shipped. This covers .patch files / git apply / patch, inline patches embedded in benchmark scripts (e.g. a python3/sed heredoc that rewrites installed engine sources before serving), in-place edits of site-packages, monkey-patching, overwriting container files, and installing forked/rebuilt engine wheels on top of the pinned image. The only exception is a patch covered by a filled-out waiver at docs/waiver/
<PR_NUMBER>.md— named after the PR that introduces the patch and filed in that same PR, stating what is patched, why the unmodified upstream image cannot run this benchmark, the upstream PR/issue link, and the removal plan — which I have linked below in the additional detail section. - If this PR uses
append-only: true, verified that it only adds generated points or recipe variants inside a selected existing config/scenario and existing same-image visual curve: every previously generated point remains present with the same recipe, no prior point is removed or rerun, and every benchmark-affecting change in the complete diff can affect only the corresponding newly appended points (never an existing point), regardless of which file contains it. - If any of the above criteria cannot reasonably be satisfied, I have provided additional reasoning below.
- Reported measured throughput/E2EL Pareto counts and evidence per affected curve (≥5 points strongly recommended). Below 5 or unverifiable: tag a core maintainer for review; recorded admin bypass required before merge. N/A if no curves are affected. Details.
Additional detail section:
- Validation and evals: Run Sweep 36470674070 completed successfully on current head
7cbda30a90e59d6a941a52fd8f4305699578dabb; all 15 agentic eval jobs passed with scores 0.97–0.98. Authorized reuse command:/use 36470674070. - Upstream recipe: vLLM recipes PR #1047 is merged, and the published MiniMax-M3 recipe documents the NVFP4 target and the exact GQA EAGLE3 combination with
FLASHINFERdraft attention anduse_local_argmax_reduction: trueused here. The target FlashInfer/TRT-LLM attention path and FP8 target/indexer KV cache are also documented. - Chat/AL: the AgentX client uses chat completions with
thinking_mode: enabled. Throughput runs use vLLM synthetic rejection with acceptance length2.78, matching the committed MiniMax-M3 EAGLE3-GQA thinking-on K=3 golden curve; evals use real verification. - Draft precision:
Inferact/MiniMax-M3-EAGLE3-GQArevision96692486b5fd38ebf8fd2a5f6bb53427d30819a8ships BF16 tensors and is loaded at stored BF16 by the pinned upstreamvllm/vllm-openai:nightly-af7f9488c2210d67e1033ecdc845b087ee7fe92bdefault path. There is no draft quantization, dtype override, checkpoint substitution, or engine patch. FP8 KV cache is applied consistently to target and draft and does not lower draft weights/activations.SGLANG_NVFP4_CKPT_FP8_NEXTN_MOEis not present or enabled. - No architecture-FLOP hacks or serving-stack patches are introduced.
append-only: trueis not used. - Pareto coverage: the B200 MiniMax-M3 AgentX P90 curve from run 36470674070 has 15 valid measured points and 9 throughput/E2EL frontier points (9/5 recommended minimum).
Signed: kedarpotdar-nv
Description
vllm/vllm-openai:nightly-af7f9488c2210d67e1033ecdc845b087ee7fe92b.use_local_argmax_reductionfor the EAGLE3 draft. Keep FlashInfer draft attention and FP8 KV, size GPU and host KV for the B200 sweep, and enable chunked prefill.Validation
git diff --checkpassed.AI model disclosure
Related Issue
None.
Type of Change
Checklist
perf-changelog.yamland have not edited historical entriesOWNER/MEMBER/COLLABORATOR) has commented/use <run_id>(or the legacy/reuse-sweep-run) on this PR. Do this only once there is a final full sweep that is all green with evals passing, since after this comment the sweep label will no longer automatically kick off new sweeps. Remove and re-add the label to force one.中文
将 MiniMax-M3 B200 AgentX 配方迁移到当前单节点 SRT 结构,固定 vLLM nightly-af7f9488c2210d67e1033ecdc845b087ee7fe92b 镜像。为 EAGLE3 草稿启用
use_local_argmax_reduction,保留 FlashInfer 注意力和 FP8 KV,调整 B200 的 GPU 与主机 KV 预算,并启用分块预填充。本 PR 仅包含 B200;B300 在独立 PR 中。未修改文档或 CI 流程文件。已生成全部 15 个 B200 矩阵点,性能变更日志校验和
git diff --check均通过。GPU 吞吐与评测等待 sweep。AI 模型:GPT-6;运行环境未提供可核实的完整版本标识。用于配置修改、源码审查、本地验证及准备 PR。无关联 issue。改动类型:错误修复和配置更新。