[Klaud Cold] Update dsv4-fp4-b200-dynamo-sglang SGLang image to v0.5.19-cu130 / 将 dsv4-fp4-b200-dynamo-sglang 的 SGLang 镜像更新至 v0.5.19-cu130 - #2908
Conversation
…9-cu130 Replace the unstable nightly lmsysorg/sglang:nightly-dev-cu13-20260821-f825d729 with the digest-pinned release lmsysorg/sglang:v0.5.19-cu130 in the master entry and its ten referenced B200 8k/1k disaggregated srt-slurm recipes. Model, precision, topology, workloads, Dynamo revision and resources are unchanged; the generated matrix differs only in the image. 将 dsv4-fp4-b200-dynamo-sglang 的 SGLang 镜像从不稳定的 nightly lmsysorg/sglang:nightly-dev-cu13-20260821-f825d729 更新为按摘要固定的正式发布版 lmsysorg/sglang:v0.5.19-cu130,同时更新主配置条目及其引用的十个 B200 8k/1k 分离式 srt-slurm 配方。模型、精度、拓扑、负载、Dynamo 版本与资源均保持不变, 生成的矩阵仅镜像不同。 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
|
Status: Capacity deferral, session terminating. No benchmark or eval run was dispatched for this PR, so there is nothing to cancel and no GPU time was used. Confirmed finding: The required Action / next step: Per the Klaud Cold capacity rule this PR stays in draft with no sweep labels, is closed, and its remote branch Links: commit d0a00ef · baseline producer run https://github.com/SemiAnalysisAI/InferenceX/actions/runs/33447526958 状态: 因容量原因推迟,会话终止。本 PR 未派发任何基准或评测运行,因此无需取消,也未使用 GPU 时间。 已确认的发现: 在首次定向派发 操作/下一步: 按照 Klaud Cold 容量规则,本 PR 保持草稿状态且不带任何 sweep 标签,随后关闭并删除其远程分支 链接: 提交 d0a00ef · 基线生产者运行 https://github.com/SemiAnalysisAI/InferenceX/actions/runs/33447526958 |
Status / 状态
Current status: Deferred for capacity. The image change was committed and this draft PR opened, but the mandatory
check-capacity --cluster b200-nscalerecheck run immediately before the first targeted dispatch failed twice (2026-09-09T02:30:01Z and 02:30:18Z), after passing at branch-creation time. No benchmark or eval run was dispatched, so no GPU time was used and there are no owned runs to cancel.Next step: None in this session. Per the Klaud Cold capacity rule the PR is returned to draft (it never left draft), closed, and its remote branch deleted so a later auto-sweep can reselect this candidate when
b200-nscaleis eligible again. No recovery is awaited or promised.Green targeted benchmarks and default evals do not prove that global repository checks pass.
Change / 变更
configs/nvidia-master.yaml:dsv4-fp4-b200-dynamo-sglang(DeepSeek-V4-Pro FP4, B200, Dynamo + SGLang, disaggregated, 8k/1k, no speculative decoding, runnercluster:b200-nscale).lmsysorg/sglang:nightly-dev-cu13-20260821-f825d729→lmsysorg/sglang:v0.5.19-cu130@sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9in the master entry and its ten referencedbenchmarks/multi_node/srt-slurm-recipes/sglang/deepseek-v4/8k1k/disagg-b200-*.yamlrecipes (model.containerequals the masterimage).86f84b94, NIXL KV transfer, workloads, resources, environment flags and recipe references. The generated matrix differs frommainonly in the image (10 recipes, 12 concurrency points, node counts 2/2/2/5/3/5/5/6/7/8).Supporting evidence / 支持证据
latest-imagesreportsdsv4/b200/dynamo-sglang/fp4/spec_method=none/ disagg / 8192-1024 onlmsysorg/sglang:nightly-dev-cu13-20260821-f825d729, dated 2026-09-01;framework-releasesreports SGLangv0.5.19.sgl-project/sglangtagv0.5.19(commit59f20bff, released 2026-09-05) is 631 commits ahead of and contains the nightly commitf825d729, so the release is a superset of the code currently pinned. Docker Hub tagv0.5.19-cu130(CUDA 13.0, linux/amd64 + arm64) was pushed 2026-09-04 and is pinned by its manifest-list digest. Digest-pinnedlmsysorg/sglang:<tag>@sha256:...images already have published Dynamo-SGLang results on GB200/GB300 (v0.5.14-cu130@sha256:5027e95b…), so the reference form is proven on srt-slurm.configs/nvidia-master.yaml(rechecked immediately before the branch was claimed; PR [NV] Refresh GB300 DeepSeek-V4-Pro AgentX with SGLang DSpark6 / 使用 SGLang DSpark6 更新 GB300 DeepSeek-V4-Pro AgentX #2623's removal of the same tag is in the GB300 agentic families only).generate_sweep_configs.py test-config --config-keys dsv4-fp4-b200-dynamo-sglangoutput is identical tomainexcept the image;utils/matrix_logictest_validation.py+test_generate_sweep_configs.pypass (291 tests).Baseline (published 2026-09-01) / 基线(发布于 2026-09-01)
GET /api/v1/benchmarks?model=DeepSeek-V4-Pro&date=2026-09-01&exact=true(noview) andGET /api/v1/workflow-info?date=2026-09-01;GET /api/v1/evaluations?model=DeepSeek-V4-Profiltered to the same identity.dsv4/b200/dynamo-sglang/fp4/spec_method=none/ disagg / multinode / 8192-1024 /single_turn/ imagelmsysorg/sglang:nightly-dev-cu13-20260821-f825d729/ routerdynamo-router@86f84b94/ KVnixl.a430c17bab40fe1bf9c6e94d8b69c3d032dd0143, created 2026-08-31T22:43Z).curve_workflow_run_id=2391is a logical curve snapshot, not the producer.tput_per_gputotal tok/s/GPU; output =output_tput_per_gpu; latencies in seconds):Initial attempt / 初始尝试
lmsysorg/sglang:v0.5.19-cu130@sha256:d6e72886…at commitd0a00ef89457fea20fbf9f94da0ff6f06a9d3d38.check-capacity --cluster b200-nscalerecheck required immediately before dispatch exited non-zero twice at 2026-09-09T02:30Z; the same check had passed before edits and branch creation a few minutes earlier.Repairs / 修复
Final full sweep / 最终完整 sweep
perf-changelog.yamlentry was appended, no sweep label was applied, and the PR never left draft.状态
当前状态: 因容量原因推迟。镜像变更已提交并创建了本草稿 PR,但在首次定向派发前必须执行的
check-capacity --cluster b200-nscale复查连续两次失败(2026-09-09T02:30:01Z 与 02:30:18Z),而该检查在创建分支时曾通过。未派发任何基准或评测运行,因此未使用 GPU 时间,也没有需要取消的运行。下一步: 本会话内无后续操作。按照 Klaud Cold 容量规则,PR 保持草稿状态(从未转为 ready)、关闭,并删除其远程分支,以便
b200-nscale恢复可用后由后续 auto-sweep 重新选择该候选。不等待也不承诺恢复。定向基准测试与默认评测通过并不证明仓库的全局检查通过。
变更
configs/nvidia-master.yaml:dsv4-fp4-b200-dynamo-sglang(DeepSeek-V4-Pro FP4、B200、Dynamo + SGLang、分离式、8k/1k、无投机解码,runnercluster:b200-nscale)。lmsysorg/sglang:nightly-dev-cu13-20260821-f825d729→lmsysorg/sglang:v0.5.19-cu130@sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,同时更新主配置条目及其引用的十个benchmarks/multi_node/srt-slurm-recipes/sglang/deepseek-v4/8k1k/disagg-b200-*.yaml配方(model.container与主配置image一致)。86f84b94、NIXL KV 传输、负载、资源、环境变量与配方引用。生成的矩阵与main仅镜像不同(10 个配方、12 个并发点,节点数 2/2/2/5/3/5/5/6/7/8)。支持证据
latest-images报告dsv4/b200/dynamo-sglang/fp4/spec_method=none/ 分离式 / 8192-1024 使用lmsysorg/sglang:nightly-dev-cu13-20260821-f825d729,日期 2026-09-01;framework-releases报告 SGLangv0.5.19。sgl-project/sglang标签v0.5.19(提交59f20bff,2026-09-05 发布)领先 nightly 提交f825d729631 个提交且包含该提交,因此该发布版是当前所固定代码的超集。Docker Hub 标签v0.5.19-cu130(CUDA 13.0,linux/amd64 + arm64)于 2026-09-04 推送,并按 manifest-list 摘要固定。按摘要固定的lmsysorg/sglang:<tag>@sha256:...镜像已在 GB200/GB300 上有公开发布的 Dynamo-SGLang 结果(v0.5.14-cu130@sha256:5027e95b…),该引用形式在 srt-slurm 上已被验证。configs/nvidia-master.yaml中的该系列键(在占用分支前已再次核查;PR [NV] Refresh GB300 DeepSeek-V4-Pro AgentX with SGLang DSpark6 / 使用 SGLang DSpark6 更新 GB300 DeepSeek-V4-Pro AgentX #2623 删除同一标签的位置仅在 GB300 agentic 系列中)。generate_sweep_configs.py test-config --config-keys dsv4-fp4-b200-dynamo-sglang的输出与main仅镜像不同;utils/matrix_logic的test_validation.py+test_generate_sweep_configs.py通过(291 个测试)。基线(发布于 2026-09-01)
GET /api/v1/benchmarks?model=DeepSeek-V4-Pro&date=2026-09-01&exact=true(无view)与GET /api/v1/workflow-info?date=2026-09-01;GET /api/v1/evaluations?model=DeepSeek-V4-Pro按同一身份过滤。dsv4/b200/dynamo-sglang/fp4/spec_method=none/ 分离式 / 多节点 / 8192-1024 /single_turn/ 镜像lmsysorg/sglang:nightly-dev-cu13-20260821-f825d729/ 路由dynamo-router@86f84b94/ KVnixl。a430c17bab40fe1bf9c6e94d8b69c3d032dd0143,创建于 2026-08-31T22:43Z)。curve_workflow_run_id=2391是逻辑曲线快照,不是生产者。tput_per_gpu,输出吞吐 =output_tput_per_gpu,延迟单位为秒)。初始尝试
lmsysorg/sglang:v0.5.19-cu130@sha256:d6e72886…,提交d0a00ef89457fea20fbf9f94da0ff6f06a9d3d38。check-capacity --cluster b200-nscale复查于 2026-09-09T02:30Z 连续两次返回非零;几分钟前在编辑与创建分支之前该检查曾通过。修复
最终完整 sweep
perf-changelog.yaml条目,未应用任何 sweep 标签,PR 从未离开草稿状态。🤖 Generated with Claude Code