Skip to content

Tune MiniMax-M3 AgentX on B200 with updated vLLM / 更新 B200 上的 MiniMax-M3 AgentX 配置 - #3435

Merged
Oseltamivir merged 2 commits into
mainfrom
codex/minimaxm3-b200-b300-vllm-agentx
Sep 30, 2026
Merged

Oseltamivir merged 2 commits into
mainfrom
codex/minimaxm3-b200-b300-vllm-agentx

Conversation

@xinli-sw

@xinli-sw xinli-sw commented Sep 25, 2026 •

Copy link
Copy Markdown
Collaborator

Description

  • Move the MiniMax-M3 B200 AgentX recipe to the current single-node SRT layout and pin vllm/vllm-openai:nightly-af7f9488c2210d67e1033ecdc845b087ee7fe92b.
  • Enable use_local_argmax_reduction for the EAGLE3 draft. Keep FlashInfer draft attention and FP8 KV, size GPU and host KV for the B200 sweep, and enable chunked prefill.
  • This PR contains B200 only. B300 is in a separate PR. No documentation or CI procedure files changed.

Validation

  • Generated all 15 B200 matrix points.
  • Perf changelog validator and git diff --check passed.
  • GPU throughput and eval results are pending the sweep.

AI model disclosure

  • Model/version: GPT-6; the exact runtime version identifier was not exposed.
  • Role: Configuration changes, source review, local validation, and PR preparation.

Related Issue

None.

Type of Change

  • Bug fix
  • New feature
  • Configuration change
  • Documentation update
  • Other

Checklist

  • I have completed the AI model disclosure and kept it current
  • I have tested my changes locally
  • I have updated documentation if necessary (none needed for this configuration change)
  • For every change that can affect benchmark performance and every recipe addition or modification, I have appended a new entry to the physical end of perf-changelog.yaml and have not edited historical entries
  • Before merging via reuse, an authorized maintainer (OWNER/MEMBER/COLLABORATOR) has commented /use <run_id> (or the legacy /reuse-sweep-run) on this PR. Do this only once there is a final full sweep that is all green with evals passing, since after this comment the sweep label will no longer automatically kick off new sweeps. Remove and re-add the label to force one.
中文

将 MiniMax-M3 B200 AgentX 配方迁移到当前单节点 SRT 结构,固定 vLLM nightly-af7f9488c2210d67e1033ecdc845b087ee7fe92b 镜像。为 EAGLE3 草稿启用 use_local_argmax_reduction,保留 FlashInfer 注意力和 FP8 KV,调整 B200 的 GPU 与主机 KV 预算,并启用分块预填充。本 PR 仅包含 B200;B300 在独立 PR 中。未修改文档或 CI 流程文件。

已生成全部 15 个 B200 矩阵点,性能变更日志校验和 git diff --check 均通过。GPU 吞吐与评测等待 sweep。

AI 模型:GPT-6;运行环境未提供可核实的完整版本标识。用于配置修改、源码审查、本地验证及准备 PR。无关联 issue。改动类型:错误修复和配置更新。

@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

@github-actions

github-actions Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this PR and didn't find any bugs. Because it removes an engine-source patch on the strength of a new pinned vLLM nightly, a human maintainer's look would still be worthwhile.

What was reviewed: confirmed PORT/EVAL_ONLY are genuinely supplied by the b200-nscale and b300-dsxe launch scripts (not caller-invented defaults) before being added to check_env_vars; verified the CPU_OFFLOAD_BYTES switch from 1024^3 to 1000^3 matches infx/matrix/generate.py's BYTES_PER_GB=1_000_000_000 decimal convention; checked the derived MAX_NUM_SEQS/MAX_CUDAGRAPH_CAPTURE_SIZE formula against the same pattern already used in the sibling minimaxm3_fp8_h100/h200 scripts; confirmed the perf-changelog.yaml entry only appends (prior bytes untouched) and its "XXX" pr-link placeholder is filled in later by infx/workflows/prepare_perf_changelog_merge.py, matching existing convention.

Extended reasoning...

The diff tunes two single-node vLLM benchmark scripts and their master-config entries (GPU/host KV utilization, batching/concurrency, image pin) and deletes an obsolete monkey-patch of vLLM's installed SimpleCPUOffloadConnector source, with a matching perf-changelog entry. No auth/crypto/permission surface is touched; it's config and shell-script arithmetic. I independently verified the env-var validation, the binary-to-decimal GB unit fix, and the derived cudagraph/max-num-seqs formula against an existing sibling script's precedent, all of which check out. The one thing I could not verify in this environment (no network access) is the PR's core factual claim that vLLM nightly-29468dde fixes the heterogeneous-KV-layer-region bug the deleted patch worked around rather than silently reintroducing it — that needs a maintainer with vLLM upstream context or a live run to confirm.

This review covers commit 89d5ec0, which is no longer the latest commit on this pull request; later commits are not covered by it.

cquil11 added a commit that referenced this pull request Sep 25, 2026
The MiniMax-M3 B200/B300 DRAM points no longer patch vLLM's SimpleCPUOffload
worker. The pinned nightly is pruned, so these points cannot run until the
image moves to a build that no longer needs the patch (see #3131, #3132, #3435).
adibarra added a commit that referenced this pull request Sep 26, 2026
* feat(srt): run single-node AgentX on native srt-slurm

Single-node AgentX points with an srt-recipe now take the native
single-node path fixed-sequence points already use. The adapter
accepts AgentX points and DSpark speculation. Single- and multi-node
AgentX share one client, benchmarks/srt_agentic.sh (moved from
benchmarks/multi_node/agentic_srt.sh), which replays one CONC point or
a CONC_LIST batch against the srt-slurm frontend and reads engine
metrics from the workers behind a router frontend. The H200
DeepSeek-V4.1-Flash SGLang config is the first one ported.

* fix(amd): drop a duplicate DSV4 prefill key and restore the MI355X Qwen3.5 fixed-sequence recipe

srt-slurm parses recipes strictly and rejected the repeated
disable-cuda-graph key. Removing the unported MI355X Qwen3.5 AgentX
recipe had also deleted the fixed-sequence recipe beside it.

* feat(srt): accept vLLM points in the single-node adapter

vLLM points validate their topology as tensor x data parallel GPUs,
with DP attention as data-parallel ranks and expert parallelism as
enable-expert-parallel, and eval-only runs set max-model-len.

* feat(agentx): port the DSV4.1 Flash B300 SGLang AgentX config to srt-slurm

* feat(agentx): let srt-slurm AgentX recipes apply the chat template client-side

Several legacy AMD AgentX scripts appended --apply-chat-template to the
replay command. AIPERF_APPLY_CHAT_TEMPLATE=true in a recipe's benchmark
env keeps that behavior on the shared srt-slurm client.

* feat(agentx): run GB200/GB300 single-node AgentX natively and port DSV4.1 Flash SGLang there

GB launchers submit recipe points through launch_srt_single_node with an
aarch64 srt-slurm setup; the squash is imported on a compute tray first.

* feat(agentx): let srt-slurm AgentX recipes apply the chat template client-side

Several legacy AMD AgentX scripts appended --apply-chat-template to the
replay command. AIPERF_APPLY_CHAT_TEMPLATE=true in a recipe's benchmark
env keeps that behavior on the shared srt-slurm client.

* feat(agentx): let srt-slurm AgentX recipes apply the chat template client-side

Several legacy AMD AgentX scripts appended --apply-chat-template to the
replay command. AIPERF_APPLY_CHAT_TEMPLATE=true in a recipe's benchmark
env keeps that behavior on the shared srt-slurm client.

* feat(agentx): let srt-slurm AgentX recipes set the AIPerf benchmark grace period

The GLM-5.2 MI325X legacy script bounded the post-window drain with
--benchmark-grace-period 1800. AIPERF_BENCHMARK_GRACE_PERIOD in a
recipe's benchmark env passes it through the shared client.

* feat(agentx): port the DSV4.1 Flash B200 vLLM AgentX config to srt-slurm

* feat(srt): force ATOM AgentX golden acceptance with its server flag

ATOM pins acceptance with --spec-decode-acceptance-length rather than an
environment variable. The adapter now sets it from the golden curve for
AgentX throughput and removes it for evals, reading the draft model for the
MiniMax GQA curve and the probabilistic sampler for Kimi DSpark.

* feat(srt): accept draft_model speculation in the single-node adapter

DeepSeek-V4 configs mark their bundled DSpark draft as draft_model; the
recipe still speculates natively, so the point binds as speculative.

* fix(mi355x): mount the shared HF cache for native AgentX checkpoints

The legacy MI355X AgentX path read DeepSeek-V4-Pro (vLLM/ATOM), DeepSeek-V4.1
Flash, MiniMax-M3 and GLM-5.2-FP8 from /it-share rather than node-local NVMe.
The native srt-slurm path keeps that mount for those AgentX checkpoints.

* feat(agentx): port the Qwen3.5 MI300X SGLang AgentX config to srt-slurm

* feat(agentx): port the Qwen3.5 MI325X SGLang AgentX config to srt-slurm

* feat(agentx): port the GLM-5.2 MI325X SGLang AgentX config to srt-slurm

* feat(agentx): port the DSV4.1 Flash H100 SGLang AgentX configs to srt-slurm

* feat(agentx): port the Qwen3.5 H100 SGLang AgentX configs to srt-slurm

* feat(agentx): port the Qwen3.5 H200 SGLang HiCache EP1 AgentX config to srt-slurm

* feat(srt): accept the draft_model label for single-node DSpark points

* feat(agentx): port the DSV4.1 Flash MI300X vLLM AgentX config to srt-slurm

* feat(agentx): port the DSV4.1 Flash MI325X vLLM AgentX config to srt-slurm

* fix(srt): bind single-node draft_model and EAGLE3 AgentX points

The DSV4 B300 DSpark config labels its points spec-decoding: draft_model and
the MiniMax-M3 TRT-LLM config drafts with EAGLE3; accept both as speculative
recipes. Golden acceptance already resolves both curves.

* feat(srt): accept vLLM EAGLE3 speculation in the single-node adapter

* feat(srt): accept EAGLE3 speculation in the single-node adapter

MiniMax-M3 vLLM AgentX recipes speculate with the EAGLE3 GQA draft,
which the golden acceptance lookup already maps to minimaxm3_eagle3_gqa.

* feat(srt): accept EAGLE3 speculation in the single-node adapter

MiniMax-M3 vLLM AgentX recipes speculate with the EAGLE3 GQA draft,
which the golden acceptance lookup already maps to minimaxm3_eagle3_gqa.

* feat(agentx): port the MiniMax-M3 MI325X vLLM AgentX config to srt-slurm

* feat(agentx): port the Qwen3.5 FP4 B300 SGLang AgentX config to srt-slurm

* feat(agentx): port the Qwen3.5 FP8 B300 SGLang AgentX config to srt-slurm

* feat(agentx): port the GLM-5.2 FP4 B300 SGLang AgentX config to srt-slurm

* feat(agentx): port the GLM-5.2 FP8 B300 SGLang AgentX config to srt-slurm

* feat(agentx): port the DSV4 Pro B300 SGLang AgentX config to srt-slurm

* feat(agentx): port the MiniMax-M3 B300 TRT-LLM AgentX config to srt-slurm

* fix(b300): serve Qwen3.8-Flash-Next NVFP4 natively from the writable model root

The checkpoint is not staged on node-local NVMe; the legacy AgentX path
downloads it to the shared writable model directory.

* feat(agentx): port the Qwen3.8-Flash-Next FP4 B300 SGLang AgentX config to srt-slurm

* feat(srt): accept draft_model speculation in the single-node adapter

DeepSeek-V4 configs mark their bundled DSpark draft as draft_model; the
recipe still speculates natively, so the point binds as speculative.

* feat(srt): force ATOM AgentX golden acceptance with its server flag

ATOM pins acceptance with --spec-decode-acceptance-length rather than an
environment variable. The adapter now sets it from the golden curve for
AgentX throughput and removes it for evals, reading the draft model for the
MiniMax GQA curve and the probabilistic sampler for Kimi DSpark.

* fix(mi355x): mount the shared HF cache for native AgentX checkpoints

The legacy MI355X AgentX path read DeepSeek-V4-Pro (vLLM/ATOM), DeepSeek-V4.1
Flash, MiniMax-M3 and GLM-5.2-FP8 from /it-share rather than node-local NVMe.
The native srt-slurm path keeps that mount for those AgentX checkpoints.

* feat(srt): accept draft-model labeled speculation in the single-node adapter

The DSV4 B200 vLLM AgentX matrix labels its native DSpark drafter
draft_model rather than mtp.

* feat(agentx): port the DSV4.1 Flash H100 and H200 vLLM AgentX configs to srt-slurm

* feat(agentx): port the MiniMax-M3 H100 and H200 vLLM AgentX configs to srt-slurm

* feat(agentx): port the Qwen3.5 MI355X SGLang AgentX config to srt-slurm

* feat(agentx): port the DSV4.1 Flash MI355X SGLang AgentX config to srt-slurm

* feat(agentx): port the GLM-5.2 FP4 MI355X SGLang AgentX config to srt-slurm

* feat(agentx): port the GLM-5.2 FP8 MI355X SGLang AgentX config to srt-slurm

* feat(agentx): port the DSV4 MI355X SGLang AgentX config to srt-slurm

* feat(agentx): port the DSV4 MI355X ATOM AgentX config to srt-slurm

* feat(agentx): port the GLM-5.2 MI355X ATOM AgentX config to srt-slurm

The DCP4 LMCache entry stays on the legacy script: srtctl reserves ATOM's
kv-transfer-config for disaggregated workers.

* feat(agentx): port the Kimi-K3 MI355X ATOM AgentX config to srt-slurm

The DCP8 LMCache entries stay on the legacy script: srtctl reserves ATOM's
kv-transfer-config for disaggregated workers.

* feat(agentx): port the MiniMax-M3 MI355X ATOM AgentX config to srt-slurm

The LMCache entries stay on the legacy script: srtctl reserves ATOM's
kv-transfer-config for disaggregated workers.

* feat(srt): accept EAGLE3 speculation in the single-node adapter

* feat(agentx): port the DSV4.1 Flash B300 vLLM AgentX config to srt-slurm

* feat(agentx): port the MiniMax-M3 B200 and B300 vLLM AgentX configs to srt-slurm

* feat(agentx): port the DSV4 B200 and B300 vLLM AgentX configs to srt-slurm

* feat(agentx): port the Qwen3.5 FP4 B200 SGLang AgentX config to srt-slurm

* feat(agentx): port the Qwen3.5 FP8 B200 SGLang AgentX config to srt-slurm

* feat(agentx): port the Qwen3.8 Next FP4 B200 SGLang AgentX config to srt-slurm

* fix(h100): point native single-node uv caches at shared NFS

* feat(agentx): port the GLM-5.2 FP4 B200 SGLang AgentX config to srt-slurm

* feat(agentx): port the GLM-5.2 FP8 B200 SGLang AgentX config to srt-slurm

* feat(agentx): port the DSV4 FP4 B200 SGLang AgentX config to srt-slurm

* feat(agentx): port the MiniMax-M3 FP4 B200 TRT-LLM AgentX config to srt-slurm

* feat(agentx): let srt-slurm AgentX recipes set the AIPerf benchmark grace period

The GLM-5.2 MI325X legacy script bounded the post-window drain with
--benchmark-grace-period 1800. AIPERF_BENCHMARK_GRACE_PERIOD in a
recipe's benchmark env passes it through the shared client.

* feat(agentx): port the DSV4.1 Flash GB200 and GB300 vLLM AgentX configs to srt-slurm

* feat(agentx): port the MiniMax-M3 MI300X vLLM AgentX config to srt-slurm

The LMCache DRAM point starts one MP server per TP rank from a setup
script inside the worker container, as the legacy script did.

* feat(agentx): port the DSV4.1 Flash MI355X vLLM AgentX config to srt-slurm

* feat(agentx): port the MiniMax-M3 MI355X vLLM AgentX config to srt-slurm

* feat(srt): accept vLLM decode context parallelism and non-drafting points in the single-node adapter

Kimi-K3 B300 serves TP8 with decode-context-parallel-size 8, and stops
drafting above conc 16 while its matrix labels every point mtp; such a
variant declares SPEC_DECODING in its benchmark env.

* feat(agentx): port the Kimi-K3 B300 vLLM AgentX config to srt-slurm

* feat(agentx): port the Kimi-K3 MI355X vLLM AgentX DCP1 points to srt-slurm

The DCP8 no-draft points keep the legacy script: the single-node
adapter requires DCP_SIZE=1 and speculation matching the matrix label.

* fix(matrix): read node counts from named srt-slurm override variants

recipe_node_count returned None for any CONFIG_FILE carrying a selector,
so rows pointing at `file.yaml:override_<name>` fell back to the master
topology estimate. Resolve `base` and `override_<name>` the way srtctl
does (deep merge, null deletes, top-level schema carried into the
variant) so the recipe allocation stays authoritative. Zip groups and
non-schema-2 variant files keep the estimate.

* refactor(agentx): consolidate DSV4 multi-node AgentX recipes into override variants

Each per-configuration recipe becomes an `override_<name>` block over a
shared `base` in one `*-variants.yaml` per master-config entry, and the
master entries select it with `CONFIG_FILE=...:override_<name>`. Every
selected variant resolves, through the pinned srtctl, to exactly the
recipe it replaces, including its original `name`. Power recipes with
top-level telemetry stay standalone because launchers detect them as text.

* refactor(agentx): consolidate GLM-5.2 multi-node AgentX recipes into override variants

Each per-configuration recipe becomes an `override_<name>` block over a
shared `base` in one `*-variants.yaml` per master-config entry, and the
master entries select it with `CONFIG_FILE=...:override_<name>`. Every
selected variant resolves, through the pinned srtctl, to exactly the
recipe it replaces, including its original `name`. Power recipes with
top-level telemetry stay standalone because launchers detect them as text.

* refactor(agentx): consolidate Kimi-K3 GB200 AgentX recipes into override variants

Each per-configuration recipe becomes an `override_<name>` block over a
shared `base` in one `*-variants.yaml` per master-config entry, and the
master entries select it with `CONFIG_FILE=...:override_<name>`. Every
selected variant resolves, through the pinned srtctl, to exactly the
recipe it replaces, including its original `name`. Power recipes with
top-level telemetry stay standalone because launchers detect them as text.

* refactor(agentx): consolidate MiniMax-M3 multi-node AgentX recipes into override variants

Each per-configuration recipe becomes an `override_<name>` block over a
shared `base` in one `*-variants.yaml` per master-config entry, and the
master entries select it with `CONFIG_FILE=...:override_<name>`. Every
selected variant resolves, through the pinned srtctl, to exactly the
recipe it replaces, including its original `name`. Power recipes with
top-level telemetry stay standalone because launchers detect them as text.

* refactor(agentx): consolidate Qwen3.5 multi-node AgentX recipes into override variants

Each per-configuration recipe becomes an `override_<name>` block over a
shared `base` in one `*-variants.yaml` per master-config entry, and the
master entries select it with `CONFIG_FILE=...:override_<name>`. Every
selected variant resolves, through the pinned srtctl, to exactly the
recipe it replaces, including its original `name`. Power recipes with
top-level telemetry stay standalone because launchers detect them as text.

* docs(recipes): describe AgentX override-variant bundles

* feat(agentx): port the DSV4 MI355X vLLM AgentX config to srt-slurm

The DEP8 points run behind srt-slurm's vLLM Router frontend with
consistent-hash session routing, as the legacy script's router did.

* fix(agentx): give the MiniMax-M3 Hopper Mooncake master time to install

* fix(agentx): give the AMD SGLang AgentX recipes an hour to become healthy

The MI300X Qwen3.5 server needed ~27 minutes to load and finish
first-request kernel tuning, past srt-slurm's 1800 s default.

* fix(agentx): extend the DSV4.1 Flash H100 SGLang health window for cold NFS loads

* fix(agentx): extend the MiniMax-M3 H100 vLLM load window for cold NFS loads

* fix(h100): keep one uv cache per runner for native single-node jobs

* fix(b300): give srt-slurm jobs the workflow time limit instead of srtctl's one hour

The B300 profile set no default_time_limit, so srtctl submitted native
single-node AgentX jobs with --time=01:00:00 and a c32 point timed out.

* fix(agentx): report RDMA port states when the Kimi-K3 B300 Mooncake rail probe fails

* fix(mi355x): let native single-node points use a squash staged on /it-share

Node-local /var/lib/squash is not visible from the runner host, so native
points always pulled the image, which fails once a nightly tag is pruned.

* fix(agentx): probe DSXE rdmap RDMA rails for the Kimi-K3 B300 Mooncake store

* fix(srt): let single-node AgentX evals run without a fixed-sequence context

* fix(srt): evaluate single-node AgentX points with the workflow's framework at native context

The single-node post-eval required MAX_MODEL_LEN and forced lm-eval, but
benchmark_lib clears MAX_MODEL_LEN for AgentX, so every AgentX eval-only
point failed. AgentX now runs run_eval as the multi-node post-eval does.

* fix(b300): download Qwen3.8-Flash-Next NVFP4 into the shared HF cache

The checkpoint is staged neither on node-local NVMe nor in the writable
model root, so native jobs resolve it by HF id like the other launchers.

* fix(srt): stage single-node AgentX eval artifacts once

* fix(srt): stage single-node AgentX eval artifacts once

* fix(b300): read models without node-local staging from the shared model root

DeepSeek-V4.1-Flash is not in STAGED_MODELS and exists on /scratch only on
some nodes, so native points failed wherever it was missing.

* fix(agentx): give GB200/GB300 DSV4.1 Flash servers a two-hour health budget

Cold weight loading from the shared HF cache plus graph capture took just
under 30 minutes on GB200 TP2, past srt-slurm's 1800 s default.

* chore(agentx): keep the pruned-image MI355X vLLM configs on their legacy scripts

Their nightly images are gone from Docker Hub; the legacy path still runs
from node-local squashes. The recipes stay in place for when a squash is
staged on /it-share.

* fix(agentx): give the DSV4.1 Flash GB200 and GB300 vLLM recipes a two-hour readiness window

* fix(agentx): let Mooncake pick the GID on DSXE InfiniBand rails for Kimi-K3 B300

* fix(agentx): give the MiniMax-M3 B300 TRT-LLM server a two-hour health budget

* style(agentx): write srt-slurm recipe overrides as block YAML

Content is unchanged; flow mappings in override variants become block
mappings. vLLM compilation/kv-transfer/speculative configs stay quoted JSON
strings, which srtctl passes to vLLM verbatim.

* fix(agentx): give the Qwen3.8-Flash-Next B300 SGLang server a two-hour health budget

* fix(srt): keep GLM-5.2's 150-step SWE-bench budget for single-node AgentX evals

The legacy GLM-5.2 scripts raised SWEBENCH_AGENT_STEP_LIMIT to 150 for
eval-only runs. srt-slurm's post-eval only sees the runner's environment,
so the recipe cannot set it; srt_eval.sh does.

* fix(agentx): give the MiniMax-M3 B200 TRT-LLM server a two-hour health budget

* fix(b200): default the Qwen3.8-Flash-Next NVFP4 path when no pool setting names it

* chore(minimaxm3): drop the vLLM SimpleCPUOffload patch

The MiniMax-M3 B200/B300 DRAM points no longer patch vLLM's SimpleCPUOffload
worker. The pinned nightly is pruned, so these points cannot run until the
image moves to a build that no longer needs the patch (see #3131, #3132, #3435).

* style(srt): format single-node adapter

按 Ruff 格式化单节点适配器。

---------

Co-authored-by: adibarra <93070681+adibarra@users.noreply.github.com>
@functionstackx

Copy link
Copy Markdown
Collaborator

InferenceX has switched away from unmaintainable bash scripts to YAML files that don't repeat the same stuff over and over again. Please merge the latest main into this PR: we have migrated single-node AgentX onto native srt-slurm (#3428), so AgentX configs are now declarative YAML recipes, not per-config 1000+ line bash slop scripts. Please also delete the old benchmarks/single_node/** scripts (see this recipe for the new format).

@functionstackx

Copy link
Copy Markdown
Collaborator

Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding

@xinli-sw
xinli-sw force-pushed the codex/minimaxm3-b200-b300-vllm-agentx branch from baa8b19 to da2ed2b Compare September 28, 2026 12:43
更新 B200 MiniMax-M3 AgentX 的 vLLM 镜像,并为 EAGLE3 草稿启用本地 argmax 归约。
@xinli-sw
xinli-sw force-pushed the codex/minimaxm3-b200-b300-vllm-agentx branch from da2ed2b to 7cbda30 Compare September 28, 2026 19:12
@xinli-sw xinli-sw changed the title Tune MiniMax-M3 AgentX on B200 and B300 with updated vLLM / 更新 B200 与 B300 上的 MiniMax-M3 AgentX 配置 Tune MiniMax-M3 AgentX on B200 with updated vLLM / 更新 B200 上的 MiniMax-M3 AgentX 配置 Sep 28, 2026
@xinli-sw xinli-sw mentioned this pull request Sep 28, 2026
6 of 10 tasks
@xinli-sw

Copy link
Copy Markdown
Collaborator Author

Hi @Ankur-singh , this submission is ready

@xinli-sw

Copy link
Copy Markdown
Collaborator Author

/use 36470674070

@kedarpotdar-nv kedarpotdar-nv left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

As a PR reviewer and CODEOWNER, I have reviewed this and have:

  • Verified that as of the moment of typing this, this is the latest version of PR_REVIEW_CHECKLIST.md
  • Verified that the general code quality meets the InferenceX standard and does not make the code quality any worse.
  • Verified that this PR has passed PR validation. Please link to GitHub Action workflow that shows this.
  • Verified that this PR passes evals. Please link to GitHub Action workflow that shows this.
  • Verified that speculative decoding PRs uses chat templates to align the AL distribution to real world
  • Verified that every draft model and draft head is served as it ships: the draft that ships with the served checkpoint, at its stored precision, through the pinned upstream image's default handling, with the shipped and effective draft precision recorded in the additional detail section. No submission-side quantization, dtype override, checkpoint substitution, or patch may lower draft precision below that default, regardless of eval results or AL. Explicitly verified that SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE is not enabled in the effective recipe, including inherited settings; enabling it is prohibited going forward, and historical runs do not grant an exception. See Draft-model precision for what counts as the default and the MLPerf comparison.
  • For agentic workloads: verified that speculative-decoding configs (EAGLE / MTP / draft models) run with simulated synthetic acceptance, with the acceptance-length value taken from the committed golden AL curve in infx/golden_al_distribution/ for that model, thinking mode, and draft length. A submission may choose any supported draft length, but it may not substitute a different acceptance target.
  • Verified against the current MODELS.md that this PR does not submit a deprecated model, scenario, or model-scenario combination.
  • Verified that the model architecture isn't changed with benchmark hacks like using --hf-overrides to skipping indexer for every x layers on models that don't natively support this. As a general rule, we won't accept optimizations that reduces the number of model architecture FLOPs. Anything that makes that same computation run faster is fair game; target/verifier FLOPs at lower precisions is fine, given that the config passes private evals, but this does not permit lowering draft-model or draft-head precision below what ships. As an general north star princple, we should only use optimizations which is used in production by customers that care about accuracy
  • If an company claims that they support vLLM/SGLang as first class LLM inference engines on their hardware, I have verified that the respective vLLM submission made using upstream https://hub.docker.com/u/vllm docker repo, upstream SGLang https://hub.docker.com/u/lmsysorg docker repo. The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet as supported by vLLM/SGLang community maintainers
  • If an company claims that they support vLLM/SGLang as first class upstream in-tree LLM inference engines on their hardware, I have have verified that the respective vLLM/SGLang submission has been made before additional frameworks (TRT-LLM, ATOM, etc.). The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL8, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet.
  • Verified that every single-node vLLM/SGLang recipe in this PR is documented in the official vLLM recipes and/or the SGLang cookbook:
    • I linked the corresponding upstream PR in the vLLM recipe repo or SGLang repo and verified that it is MERGED before this InferenceX PR merges. An opened, draft, or closed-without-merge upstream PR does not satisfy this requirement. If the matching recipe was already published, I linked the published recipe/cookbook page in the additional detail section below.
  • Verified that this PR does not patch the inference engine or serving stack — the pinned image must run as shipped. This covers .patch files / git apply / patch, inline patches embedded in benchmark scripts (e.g. a python3/sed heredoc that rewrites installed engine sources before serving), in-place edits of site-packages, monkey-patching, overwriting container files, and installing forked/rebuilt engine wheels on top of the pinned image. The only exception is a patch covered by a filled-out waiver at docs/waiver/<PR_NUMBER>.md — named after the PR that introduces the patch and filed in that same PR, stating what is patched, why the unmodified upstream image cannot run this benchmark, the upstream PR/issue link, and the removal plan — which I have linked below in the additional detail section.
  • If this PR uses append-only: true, verified that it only adds generated points or recipe variants inside a selected existing config/scenario and existing same-image visual curve: every previously generated point remains present with the same recipe, no prior point is removed or rerun, and every benchmark-affecting change in the complete diff can affect only the corresponding newly appended points (never an existing point), regardless of which file contains it.
  • If any of the above criteria cannot reasonably be satisfied, I have provided additional reasoning below.
  • Reported measured throughput/E2EL Pareto counts and evidence per affected curve (≥5 points strongly recommended). Below 5 or unverifiable: tag a core maintainer for review; recorded admin bypass required before merge. N/A if no curves are affected. Details.

Additional detail section:

  • Validation and evals: Run Sweep 36470674070 completed successfully on current head 7cbda30a90e59d6a941a52fd8f4305699578dabb; all 15 agentic eval jobs passed with scores 0.97–0.98. Authorized reuse command: /use 36470674070.
  • Upstream recipe: vLLM recipes PR #1047 is merged, and the published MiniMax-M3 recipe documents the NVFP4 target and the exact GQA EAGLE3 combination with FLASHINFER draft attention and use_local_argmax_reduction: true used here. The target FlashInfer/TRT-LLM attention path and FP8 target/indexer KV cache are also documented.
  • Chat/AL: the AgentX client uses chat completions with thinking_mode: enabled. Throughput runs use vLLM synthetic rejection with acceptance length 2.78, matching the committed MiniMax-M3 EAGLE3-GQA thinking-on K=3 golden curve; evals use real verification.
  • Draft precision: Inferact/MiniMax-M3-EAGLE3-GQA revision 96692486b5fd38ebf8fd2a5f6bb53427d30819a8 ships BF16 tensors and is loaded at stored BF16 by the pinned upstream vllm/vllm-openai:nightly-af7f9488c2210d67e1033ecdc845b087ee7fe92b default path. There is no draft quantization, dtype override, checkpoint substitution, or engine patch. FP8 KV cache is applied consistently to target and draft and does not lower draft weights/activations. SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE is not present or enabled.
  • No architecture-FLOP hacks or serving-stack patches are introduced. append-only: true is not used.
  • Pareto coverage: the B200 MiniMax-M3 AgentX P90 curve from run 36470674070 has 15 valid measured points and 9 throughput/E2EL frontier points (9/5 recommended minimum).

Signed: kedarpotdar-nv

@Oseltamivir Oseltamivir left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm

@Oseltamivir
Oseltamivir merged commit 3e25510 into main Sep 30, 2026
25 checks passed
@Oseltamivir
Oseltamivir deleted the codex/minimaxm3-b200-b300-vllm-agentx branch September 30, 2026 01:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

4 participants