[Klaud Cold] Update qwen3.5-fp8-h200-sglang SGLang image to v0.5.19-cu130 / 将 qwen3.5-fp8-h200-sglang 的 SGLang 镜像升级至 v0.5.19-cu130 - #2965
Conversation
…-cu130 Bump the master image for the qwen3.5-fp8-h200-sglang family from lmsysorg/sglang:v0.5.14-cu130 to lmsysorg/sglang:v0.5.19-cu130 (Docker Hub digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, tag commit sgl-project/sglang@0bcd822). The recipe script, model, TP8/EP8 topology and the 8k1k concurrency range are unchanged. 将 qwen3.5-fp8-h200-sglang 家族的主配置镜像从 lmsysorg/sglang:v0.5.14-cu130 升级到 lmsysorg/sglang:v0.5.19-cu130(Docker Hub 摘要 sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, 标签提交 sgl-project/sglang@0bcd822)。配方脚本、模型、TP8/EP8 拓扑和 8k1k 并发范围均保持不变。 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
Initial attempt
Server log:
初始尝试
服务日志:所有 rank 均出现
|
Capacity deferral before the final full sweep
最终完整扫描前的容量推迟
|
|
Klaud Cold: capacity-deferred. Finishing cleanup; owned child runs will be stopped and checked before closure. Klaud Cold:capacity-deferred。正在完成清理;将先停止并确认自有子运行的状态,再关闭 PR。 |
|
Klaud Cold: capacity-deferred. All owned runs are terminal. Repairs: 0. Runs: 34477386674, 34487565175. PR closed; branch deleted for retry. Klaud Cold:capacity-deferred。所有自有运行均已结束。修复次数:0。运行:34477386674, 34487565175。 PR 已关闭;分支已删除,后续可以重试。 |
Update the
qwen3.5-fp8-h200-sglangmaster image fromlmsysorg/sglang:v0.5.14-cu130to the current SGLang releaselmsysorg/sglang:v0.5.19-cu130(Docker Hub digestsha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, tag commit sgl-project/sglang@0bcd822). The recipe script, model, TP8/EP8 topology and the 8k1k concurrency range are unchanged.Baseline
benchmarks?model=Qwen-3.5-397B-A17B&date=2026-07-04&exact=true,workflow-info?date=2026-07-04,evaluations?model=Qwen-3.5-397B-A17B&date=2026-07-04&exact=true)lmsysorg/sglang:v0.5.14-cu130(digestsha256:5027e95bf6ec536856b1b52a91d1f35ff5c564ab83e8a94758a169ff09bb8df3, tag commit sgl-project/sglang@49e384c)Qwen/Qwen3.5-397B-A17B-FP8, SGLang FP8, TP8 EP8, no speculative decoding, fixed-seq-len 8k1k (ISL 8192 / OSL 1024), concurrency 4 / 8 / 16 / 32 / 64, random datasete016f0c14bf2aab899de08b538d9b343069b07a5, changelog PR #2061); benchmark result IDs 432218, 432222, 432231, 432227, 432220em_strict0.9651 /em_flexible0.9583 (evaluation ID 8294) and at concurrency 64,em_strict0.9689 /em_flexible0.9583 (evaluation ID 8293), both from the same producer run. The public evaluations feed labels these rowsdisagg: truewith 64 prefill/decode GPUs, which does not match the single-node TP8 recipe; the benchmark rows above carry the correctdisagg: falseidentity.将
qwen3.5-fp8-h200-sglang的主配置镜像从lmsysorg/sglang:v0.5.14-cu130升级到当前 SGLang 发布版lmsysorg/sglang:v0.5.19-cu130(Docker Hub 摘要sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,标签提交 sgl-project/sglang@0bcd822)。配方脚本、模型、TP8/EP8 拓扑以及 8k1k 并发范围均保持不变。基线
benchmarks?model=Qwen-3.5-397B-A17B&date=2026-07-04&exact=true,workflow-info?date=2026-07-04,evaluations?model=Qwen-3.5-397B-A17B&date=2026-07-04&exact=true)lmsysorg/sglang:v0.5.14-cu130(摘要sha256:5027e95bf6ec536856b1b52a91d1f35ff5c564ab83e8a94758a169ff09bb8df3,标签提交 sgl-project/sglang@49e384c)Qwen/Qwen3.5-397B-A17B-FP8,SGLang FP8,TP8 EP8,无投机解码,固定序列长度 8k1k(ISL 8192 / OSL 1024),并发 4 / 8 / 16 / 32 / 64,随机数据集e016f0c14bf2aab899de08b538d9b343069b07a5,changelog PR #2061);基准结果 ID 432218、432222、432231、432227、432220em_strict0.9651 /em_flexible0.9583(评测 ID 8294);并发 64 的 gsm8k,em_strict0.9689 /em_flexible0.9583(评测 ID 8293),均来自同一数据来源运行。公共评测接口将这些行标记为disagg: true且 64 个 prefill/decode GPU,与单节点 TP8 配方不符;上表基准行的disagg: false身份是正确的。🤖 Generated with Claude Code