Skip to content

[Klaud Cold] Update dsv4-fp4-gb200-dynamo-vllm-agentic-mtp2-agg vLLM image to v0.29.0 / 将 dsv4-fp4-gb200-dynamo-vllm-agentic-mtp2-agg 的 vLLM 镜像更新至 v0.29.0 - #3033

Closed
Klaud-Cold wants to merge 2 commits into
mainfrom
klaud/auto-dfbdb56b9f982230-dc3e0d44d25b0bb7
Closed

Klaud-Cold wants to merge 2 commits into
mainfrom
klaud/auto-dfbdb56b9f982230-dc3e0d44d25b0bb7

Conversation

@Klaud-Cold

@Klaud-Cold Klaud-Cold commented Sep 12, 2026

Copy link
Copy Markdown
Collaborator

Goal: Update vLLM image from vllm/vllm-openai:nightly-dev-arm64-cu13.0.1-426e59f to vllm/vllm-openai:v0.29.0.
Baseline: 2026-08-18 · vllm/vllm-openai:nightly-dev-arm64-cu13.0.1-426e59f
AgentX · PTP8x1/DTP8x0 · Mean latency · Sources: API 1, API 2

Concurrency Total tok/s/GPU ↑ Output tok/s/GPU ↑ TTFT ms ↓ TPOT ms ↓
1 N/A N/A N/A N/A
4 N/A N/A N/A N/A
8 N/A N/A N/A N/A

Note: All rows: unavailable.

Eval: N/A

中文

**目标:**将 vLLM 镜像从 vllm/vllm-openai:nightly-dev-arm64-cu13.0.1-426e59f 更新为 vllm/vllm-openai:v0.29.0
**基线:**2026-08-18 · vllm/vllm-openai:nightly-dev-arm64-cu13.0.1-426e59f
AgentX · PTP8x1/DTP8x0 · 平均延迟 · 来源: API 1, API 2;数值及异常说明见上表。

…o v0.29.0

Replace the locally built vllm/vllm-openai:nightly-dev-arm64-cu13.0.1-426e59f
image with the vLLM v0.29.0 release for the GB200 DeepSeek-V4-Pro FP4
Dynamo-vLLM AgentX MTP aggregate family. The master image and both dedicated
recipes (model.container and identity.container.image) move together; every
point, topology, speculation and workload setting is unchanged.

将 GB200 DeepSeek-V4-Pro FP4 Dynamo-vLLM AgentX MTP 聚合配置族的镜像从本地构建的
vllm/vllm-openai:nightly-dev-arm64-cu13.0.1-426e59f 更新为 vLLM v0.29.0 正式版。
主配置镜像与两个专用配方(model.container 与 identity.container.image)同步更新,
所有测点、拓扑、投机解码与负载设置保持不变。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Baseline: 2026-08-18 · vllm/vllm-openai:nightly-dev-arm64-cu13.0.1-426e59f

AgentX · PTP8x1/DTP8x0 · Mean latency

Concurrency Total tok/s/GPU ↑ Output tok/s/GPU ↑ TTFT ms ↓ TPOT ms ↓
1 N/A N/A N/A N/A
4 N/A N/A N/A N/A
8 N/A N/A N/A N/A

Note: All rows: unavailable.

中文

**基线:**2026-08-18 · vllm/vllm-openai:nightly-dev-arm64-cu13.0.1-426e59f;数值及异常说明见表格。

@Klaud-Cold

Klaud-Cold commented Sep 12, 2026

Copy link
Copy Markdown
Collaborator Author

Initial attempt · Failed · Run 34662875814 / attempt 1 · 2026-09-12 01:46 UTC
vllm/vllm-openai:v0.29.0 · 127c84be90f1 · AgentX · PTP8x1/DTP8x0 · Mean latency
Change: Replace the locally built dev image (build commit unknown, 426e59f absent from upstream vLLM, built 2026-07-26, CUDA 13.0.1, FlashInfer 0.6.14) with the vLLM v0.29.0 release (commit 98dff2a8, CUDA 13.0.2, FlashInfer 0.6.18, torch 2.13.0) in the master image and both recipes with every flag kept: v0.29.0 source still accepts FLASHINFER_MLA_SPARSE_DSV4, the deprecated use_fp4_indexer_cache alias (vllm-project/vllm#52550) and the VLLM_PREFIX_CACHE_RETENTION_INTERVAL env (vllm-project/vllm#52216), MTP on the forced V2 runner, and its DP hash already ignores numa_bind so the srt-slurm v1.0.45 setup patch is skipped.

Concurrency Output tok/s/GPU ↑ TTFT ms ↓ TPOT ms ↓
1 N/A N/A N/A
8 N/A N/A N/A

Note: c1: failed.
Note: c1: 2 request errors.
Note: All rows: Δ N/A: point unavailable.
Note: c8: cancelled.
Note: c8: request errors unavailable.

Next: Repair 1: lower gpu-memory-utilization from 0.94 to 0.90 in both recipes to restore runtime headroom for the sparse indexer workspace, then rerun the smoke.

中文

初次尝试 · 失败 · Run 34662875814 / attempt 1 · 2026-09-12 01:46 UTC
vllm/vllm-openai:v0.29.0 · 127c84be90f1 · AgentX · PTP8x1/DTP8x0 · 平均延迟
**变更:**将主配置镜像与两个配方从本地构建的开发镜像(构建提交未知,426e59f 不在上游 vLLM 中,构建于 2026-07-26,CUDA 13.0.1,FlashInfer 0.6.14)更新为 vLLM v0.29.0 正式版(提交 98dff2a8,CUDA 13.0.2,FlashInfer 0.6.18,torch 2.13.0),所有参数保持不变:v0.29.0 源码仍接受 FLASHINFER_MLA_SPARSE_DSV4、已弃用的 use_fp4_indexer_cache 别名(https://github.com/vllm-project/vllm/pull/52550)与 VLLM_PREFIX_CACHE_RETENTION_INTERVAL 环境变量(https://github.com/vllm-project/vllm/pull/52216),MTP 在强制启用的 V2 runner 上受支持,且其 DP 哈希已忽略 numa_bind,因此 srt-slurm v1.0.45 的 setup 补丁会被跳过。实测数值及异常说明见上表。
**下一步:**修复 1:将两个配方的 gpu-memory-utilization 从 0.94 降至 0.90,为稀疏索引器工作区恢复运行时显存余量,然后重新运行冒烟测试。

…or vLLM v0.29.0

With vLLM v0.29.0 the memory profiler measures only 3.8 GiB of peak activation
and sizes the KV cache to 62.5 GiB at 0.94, so the first long AgentX prefill
OOMed on TP rank 7 while allocating the 1 GiB sparse indexer logits workspace.
Reserve headroom by lowering gpu-memory-utilization to 0.90 in both dedicated
recipes; the KV cache stays far above the old image's 31.4 GiB.

vLLM v0.29.0 的显存分析仅测得 3.8 GiB 峰值激活,在 0.94 下将 KV 缓存扩大到 62.5 GiB,
导致首个长上下文 AgentX prefill 在 TP rank 7 分配 1 GiB 稀疏索引器 logits 工作区时 OOM。
将两个专用配方的 gpu-memory-utilization 降至 0.90 以保留余量;KV 缓存仍远高于旧镜像的 31.4 GiB。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Capacity-deferred · 2026-09-12 01:47 UTC
Initial attempt (Run 34662875814 / attempt 1) failed: c1 hit a CUDA OOM on TP rank 7 while allocating the sparse-indexer logits workspace after vLLM v0.29.0 sized the KV cache to 62.5 GiB at gpu-memory-utilization 0.94; c8 was cancelled by fail-fast.
Repair 1 (2b1f6386ae1b, gpu-memory-utilization 0.94 → 0.90 in both recipes) is pushed but was not dispatched: the gb200 capacity gate failed on two consecutive checks. Owned runs are terminal; the PR is closed and the branch released so a later autosweep can retry.

中文

容量受限延期 · 2026-09-12 01:47 UTC
初次尝试(Run 34662875814 / attempt 1)失败:vLLM v0.29.0 在 gpu-memory-utilization 0.94 下将 KV 缓存扩大到 62.5 GiB,c1 在 TP rank 7 分配稀疏索引器 logits 工作区时发生 CUDA OOM;c8 被 fail-fast 取消。
修复 1(2b1f6386ae1b,两个配方的 gpu-memory-utilization 由 0.94 降至 0.90)已推送但未派发:gb200 容量检查连续两次未通过。自有运行均已结束;PR 关闭并释放分支,以便后续自动巡检重试。

@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

capacity-deferred · Cleanup pending. Stop and confirm owned runs before closing.

中文

capacity-deferred · 清理待完成。先停止并确认自有运行结束,再关闭 PR。

@Klaud-Cold Klaud-Cold closed this Sep 12, 2026
@Klaud-Cold
Klaud-Cold deleted the klaud/auto-dfbdb56b9f982230-dc3e0d44d25b0bb7 branch September 12, 2026 01:48
@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

capacity-deferred · Repairs: 1 · Runs: 34662875814
All owned runs ended. PR closed; branch deleted for retry.

中文

capacity-deferred · 修复次数:1 · 运行:34662875814
所有自有运行均已结束。PR 已关闭;分支已删除,可重新尝试。

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

1 participant