[Klaud Cold] Update dsv4-fp4-gb200-dynamo-vllm-agentic-mtp2-agg vLLM image to v0.29.0 / 将 dsv4-fp4-gb200-dynamo-vllm-agentic-mtp2-agg 的 vLLM 镜像更新至 v0.29.0 - #3033
Conversation
…o v0.29.0 Replace the locally built vllm/vllm-openai:nightly-dev-arm64-cu13.0.1-426e59f image with the vLLM v0.29.0 release for the GB200 DeepSeek-V4-Pro FP4 Dynamo-vLLM AgentX MTP aggregate family. The master image and both dedicated recipes (model.container and identity.container.image) move together; every point, topology, speculation and workload setting is unchanged. 将 GB200 DeepSeek-V4-Pro FP4 Dynamo-vLLM AgentX MTP 聚合配置族的镜像从本地构建的 vllm/vllm-openai:nightly-dev-arm64-cu13.0.1-426e59f 更新为 vLLM v0.29.0 正式版。 主配置镜像与两个专用配方(model.container 与 identity.container.image)同步更新, 所有测点、拓扑、投机解码与负载设置保持不变。 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
|
Baseline: 2026-08-18 · AgentX · PTP8x1/DTP8x0 · Mean latency
Note: All rows: unavailable. 中文**基线:**2026-08-18 · |
|
Initial attempt · Failed · Run 34662875814 / attempt 1 · 2026-09-12 01:46 UTC
Note: c1: failed. Next: Repair 1: lower gpu-memory-utilization from 0.94 to 0.90 in both recipes to restore runtime headroom for the sparse indexer workspace, then rerun the smoke. 中文初次尝试 · 失败 · Run 34662875814 / attempt 1 · 2026-09-12 01:46 UTC |
…or vLLM v0.29.0 With vLLM v0.29.0 the memory profiler measures only 3.8 GiB of peak activation and sizes the KV cache to 62.5 GiB at 0.94, so the first long AgentX prefill OOMed on TP rank 7 while allocating the 1 GiB sparse indexer logits workspace. Reserve headroom by lowering gpu-memory-utilization to 0.90 in both dedicated recipes; the KV cache stays far above the old image's 31.4 GiB. vLLM v0.29.0 的显存分析仅测得 3.8 GiB 峰值激活,在 0.94 下将 KV 缓存扩大到 62.5 GiB, 导致首个长上下文 AgentX prefill 在 TP rank 7 分配 1 GiB 稀疏索引器 logits 工作区时 OOM。 将两个专用配方的 gpu-memory-utilization 降至 0.90 以保留余量;KV 缓存仍远高于旧镜像的 31.4 GiB。 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
Capacity-deferred · 2026-09-12 01:47 UTC 中文容量受限延期 · 2026-09-12 01:47 UTC |
|
capacity-deferred · Cleanup pending. Stop and confirm owned runs before closing. 中文capacity-deferred · 清理待完成。先停止并确认自有运行结束,再关闭 PR。 |
|
capacity-deferred · Repairs: 1 · Runs: 34662875814 中文capacity-deferred · 修复次数:1 · 运行:34662875814 |
Goal: Update vLLM image from
vllm/vllm-openai:nightly-dev-arm64-cu13.0.1-426e59ftovllm/vllm-openai:v0.29.0.Baseline: 2026-08-18 ·
vllm/vllm-openai:nightly-dev-arm64-cu13.0.1-426e59fAgentX · PTP8x1/DTP8x0 · Mean latency · Sources: API 1, API 2
Note: All rows: unavailable.
Eval: N/A
中文
**目标:**将 vLLM 镜像从
vllm/vllm-openai:nightly-dev-arm64-cu13.0.1-426e59f更新为vllm/vllm-openai:v0.29.0。**基线:**2026-08-18 ·
vllm/vllm-openai:nightly-dev-arm64-cu13.0.1-426e59fAgentX · PTP8x1/DTP8x0 · 平均延迟 · 来源: API 1, API 2;数值及异常说明见上表。