Skip to content

[Klaud Cold] Update minimaxm3-fp8-mi300x-vllm-agentic-mtp vLLM ROCm image to v0.29.0 / 将 minimaxm3-fp8-mi300x-vllm-agentic-mtp 的 vLLM ROCm 镜像更新至 v0.29.0 - #3063

Open
Klaud-Cold wants to merge 3 commits into
mainfrom
klaud/auto-75240093e18ec8ca-4f7436dbcc1df172
Open

[Klaud Cold] Update minimaxm3-fp8-mi300x-vllm-agentic-mtp vLLM ROCm image to v0.29.0 / 将 minimaxm3-fp8-mi300x-vllm-agentic-mtp 的 vLLM ROCm 镜像更新至 v0.29.0#3063
Klaud-Cold wants to merge 3 commits into
mainfrom
klaud/auto-75240093e18ec8ca-4f7436dbcc1df172

Conversation

@Klaud-Cold

@Klaud-Cold Klaud-Cold commented Sep 13, 2026

Copy link
Copy Markdown
Collaborator

Goal: Update vLLM image from vllm/vllm-openai-rocm:nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36 to vllm/vllm-openai-rocm:v0.29.0.
Baseline: 2026-09-08 · vllm/vllm-openai-rocm:nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36
AgentX · TP8/EP1 · Mean latency · Sources: API 1, API 2, API 3, API 4, API 5

Concurrency Total tok/s/GPU ↑ Output tok/s/GPU ↑ TTFT ms ↓ TPOT ms ↓
2 916.81 7.21 1,388.27 21.56
4 1,207.64 9.51 1,122.22 29.97
6 1,616.44 13.04 1,021.15 37.91
8 2,493.5 17.49 1,038.12 44.28
10 2,698.1 20.43 2,989.5 54.09
16 2,416.61 19.11 37,641.34 69.79
Eval Score ↑ Samples
gsm8k/em_strict · c2 97.19% 1,319
gsm8k/em_strict · c4 97.27% 1,319
gsm8k/em_strict · c6 97.27% 1,319
gsm8k/em_strict · c8 97.19% 1,319
gsm8k/em_strict · c10 97.35% 1,319
gsm8k/em_strict · c16 97.19% 1,319
中文

**目标:**将 vLLM 镜像从 vllm/vllm-openai-rocm:nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36 更新为 vllm/vllm-openai-rocm:v0.29.0
**基线:**2026-09-08 · vllm/vllm-openai-rocm:nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36
AgentX · TP8/EP1 · 平均延迟 · 来源: API 1, API 2, API 3, API 4, API 5;数值及异常说明见上表。

… v0.29.0

Move the MI300X MiniMax-M3 MXFP8 AgentX EAGLE3 recipe from the unstable
commit nightly vllm/vllm-openai-rocm:nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36
to the v0.29.0 release image vllm/vllm-openai-rocm:v0.29.0
(digest sha256:e5e47f6aaab675c252c381f0dac237b31b10d87bb74d092b07fb4065efd7f5a1).
Recipe, search space and evals are unchanged.

将 MI300X MiniMax-M3 MXFP8 AgentX EAGLE3 配方的镜像从不稳定的提交 nightly
vllm/vllm-openai-rocm:nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36
更新为 v0.29.0 发布镜像 vllm/vllm-openai-rocm:v0.29.0
(digest sha256:e5e47f6aaab675c252c381f0dac237b31b10d87bb74d092b07fb4065efd7f5a1)。
配方、搜索空间与评测保持不变。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Baseline: 2026-09-08 · vllm/vllm-openai-rocm:nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36

AgentX · TP8/EP1 · Mean latency

Concurrency Total tok/s/GPU ↑ Output tok/s/GPU ↑ TTFT ms ↓ TPOT ms ↓
2 916.81 7.21 1,388.27 21.56
4 1,207.64 9.51 1,122.22 29.97
6 1,616.44 13.04 1,021.15 37.91
8 2,493.5 17.49 1,038.12 44.28
10 2,698.1 20.43 2,989.5 54.09
16 2,416.61 19.11 37,641.34 69.79
Eval Score ↑ Samples
gsm8k/em_strict · c2 97.19% 1,319
gsm8k/em_strict · c4 97.27% 1,319
gsm8k/em_strict · c6 97.27% 1,319
gsm8k/em_strict · c8 97.19% 1,319
gsm8k/em_strict · c10 97.35% 1,319
gsm8k/em_strict · c16 97.19% 1,319
中文

**基线:**2026-09-08 · vllm/vllm-openai-rocm:nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36;数值及异常说明见表格。

@Klaud-Cold

Klaud-Cold commented Sep 13, 2026

Copy link
Copy Markdown
Collaborator Author

Initial attempt · Passed · Run 34729027628 / attempt 1 · 2026-09-13 02:35 UTC
vllm/vllm-openai-rocm:v0.29.0 · fa902f1d8543 · AgentX · TP8/EP1 · Mean latency
Change: Move the master image from the commit nightly nightly-d9105ea8… (vLLM main at d9105ea8, 2026-09-07) to the v0.29.0 release image (digest sha256:e5e47f6a…, pushed 2026-09-09), measured at PR head 4c0198b0 via ref; the release branch left main on 2026-08-31, so it lacks the ROCm MiniMax-M3 decode indexer/top-k optimization (#54682) and the AITER v0.1.21.post1 base bump (#52826), while every recipe flag, env var and minimax_m3 parser exists at v0.29.0 and the recipe script is unchanged.

Concurrency Output tok/s/GPU ↑ TTFT ms ↓ TPOT ms ↓
2 7.25 (+0.6%) 1,378.3 (-0.7%) 22.32 (+3.5%)
16 15.16 (-20.7%) 59,393.45 (+57.8%) 78.79 (+12.9%)
Eval Score ↑ Samples
minimax_m3_smoke/em_strict · c2 100% (N/A) N/A/1 (old/new)
minimax_m3_smoke/em_strict · c16 100% (N/A) N/A/1 (old/new)

Note: All rows: Δ N/A: no matched eval baseline.

Next: Append the perf-changelog entry, validate the full matrix and start the final full sweep with full-sweep-enabled.

中文

初次尝试 · 已通过 · Run 34729027628 / attempt 1 · 2026-09-13 02:35 UTC
vllm/vllm-openai-rocm:v0.29.0 · fa902f1d8543 · AgentX · TP8/EP1 · 平均延迟
**变更:**将主镜像从提交 nightly nightly-d9105ea8…(vLLM main 于 2026-09-07 的 d9105ea8)更新为 v0.29.0 发布镜像(digest sha256:e5e47f6a…,2026-09-09 推送),通过 ref 在 PR 头提交 4c0198b0 上测量;该发布分支于 2026-08-31 从 main 分出,因此不含 ROCm MiniMax-M3 解码索引/top-k 优化(#54682)与 AITER v0.1.21.post1 基础镜像升级(#52826),但配方所用全部参数、环境变量及 minimax_m3 解析器均存在于 v0.29.0,配方脚本未改动。实测数值及异常说明见上表。
**下一步:**追加 perf-changelog 条目,校验完整矩阵,并以 full-sweep-enabled 启动最终完整 sweep。

github-actions Bot and others added 2 commits September 13, 2026 02:36
…tic-mtp v0.29.0 image bump

Record the vLLM ROCm image move from nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36
to the v0.29.0 release image for PR #3063 after the smoke run passed.

为 PR #3063 记录 vLLM ROCm 镜像从 nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36
更新至 v0.29.0 发布镜像的 perf-changelog 条目(冒烟运行已通过)。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@Klaud-Cold

Klaud-Cold commented Sep 13, 2026

Copy link
Copy Markdown
Collaborator Author

Final full sweep · Passed · Run 34733463164 / attempt 1 · 2026-09-13 05:14 UTC
vllm/vllm-openai-rocm:v0.29.0 · 25797d57b054 · AgentX · TP8/EP1 · Mean latency
Change: Validate the complete updated-image family at this head.

Concurrency Output tok/s/GPU ↑ TTFT ms ↓ TPOT ms ↓
2 7.88 (+9.3%) 1,035.56 (-25.4%) 19.73 (-8.5%)
4 9.25 (-2.8%) 1,165.75 (+3.9%) 30.5 (+1.8%)
6 12.88 (-1.2%) 1,022.99 (+0.2%) 38.2 (+0.8%)
8 16.81 (-3.9%) 1,047.72 (+0.9%) 45.06 (+1.8%)
10 20.13 (-1.5%) 3,499.99 (+17.1%) 55.51 (+2.6%)
16 15.23 (-20.3%) 55,727.99 (+48%) 78.89 (+13%)

Note: c16: 2 request errors.

Eval Score ↑ Samples
minimax_m3_smoke/em_strict · c2 100% (N/A) N/A/1 (old/new)
minimax_m3_smoke/em_strict · c4 100% (N/A) N/A/1 (old/new)
minimax_m3_smoke/em_strict · c6 100% (N/A) N/A/1 (old/new)
minimax_m3_smoke/em_strict · c8 100% (N/A) N/A/1 (old/new)
minimax_m3_smoke/em_strict · c10 100% (N/A) N/A/1 (old/new)
minimax_m3_smoke/em_strict · c16 100% (N/A) N/A/1 (old/new)

Note: All rows: Δ N/A: no matched eval baseline.

Next: Mark ready for maintainer review.

中文

最终完整 sweep · 已通过 · Run 34733463164 / attempt 1 · 2026-09-13 05:14 UTC
vllm/vllm-openai-rocm:v0.29.0 · 25797d57b054 · AgentX · TP8/EP1 · 平均延迟
**变更:**验证此提交更新镜像后的完整配置族。实测数值及异常说明见上表。
**下一步:**标记为就绪,等待维护者审查。

@github-actions

Copy link
Copy Markdown
Contributor

@Klaud-Cold
Klaud-Cold marked this pull request as ready for review September 13, 2026 05:14
@Klaud-Cold
Klaud-Cold requested a review from a team September 13, 2026 05:14
@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

validated · Repairs: 0 · Runs: 34729027628, 34733460912, 34733463164
All owned runs ended. Full sweep verified; ready for review.

中文

validated · 修复次数:0 · 运行:34729027628, 34733460912, 34733463164
所有自有运行均已结束。完整 sweep 已验证;已就绪,等待审查。

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good, straightforward image-bump config change.

What was reviewed: the image field bump in configs/amd-master.yaml for minimaxm3-fp8-mi300x-vllm-agentic-mtp (nightly tag to v0.29.0 release tag); confirmed the matching perf-changelog.yaml entry is appended at the true tail without disturbing prior bytes, and that it follows the existing config-keys/description/pr-link schema with paired English/Chinese descriptions like neighboring entries. This recipe is single-node (multinode: false), so the multi-node model.container == image consistency rule doesn't apply here.

Extended reasoning...

Overview

The diff touches exactly two files: configs/amd-master.yaml (a one-line image: bump from a ROCm nightly tag to the v0.29.0 release tag for the minimaxm3-fp8-mi300x-vllm-agentic-mtp recipe) and perf-changelog.yaml (a new append-only entry documenting that same change, with English and Chinese description bullets, config-keys, and a pr-link). No benchmark scripts, search-space definitions, evals, or other recipes were modified.

Security risks

None. This is a config data change pointing to an upstream, tagged Docker Hub image (vllm/vllm-openai-rocm:v0.29.0), not a build script, credential, or executable code path. No injection, auth, or data-exposure surface is touched.

Level of scrutiny

This warrants light scrutiny consistent with a mechanical version-bump PR. I verified the changelog entry lands strictly after the prior entry at line 7443 (no reordering or deletion of existing content) and that its shape (config-keys / description list with EN+ZH bullets / pr-link) mirrors the immediately preceding entries in the file. The multinode: false setting on this recipe means the "model.container must equal image" multi-node rule cited in repo conventions doesn't apply, so there's no cross-field consistency issue to flag.

Other factors

No bug-hunting findings were reported, and my own reading of the diff didn't surface anything beyond what's already documented in the changelog entry (release-branch exclusions of two upstream vLLM PRs, noted as informational rather than a functional gap in this change). The change is small, self-contained, and follows an established repository pattern for image-bump PRs.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

1 participant