[Klaud Cold] Update dsr1-fp8-h100-dynamo-sglang SGLang image to v0.5.19-cu130 / 将 dsr1-fp8-h100-dynamo-sglang 的 SGLang 镜像升级至 v0.5.19-cu130 - #2967
Conversation
Bump the master image for the H100 DeepSeek-R1 FP8 Dynamo-SGLang disaggregated MTP family from lmsysorg/sglang:v0.5.8-cu130 to lmsysorg/sglang:v0.5.19-cu130 (Docker Hub digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, sglang tag commit 0bcd822377da7b5718e674eaf9c870d349424dd1). Model, topology, speculative decoding, workloads and recipe references are unchanged. 将 H100 DeepSeek-R1 FP8 Dynamo-SGLang 分离式 MTP 配方的主配置镜像从 lmsysorg/sglang:v0.5.8-cu130 升级到 lmsysorg/sglang:v0.5.19-cu130 (Docker Hub 摘要 sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, sglang 标签提交 0bcd822377da7b5718e674eaf9c870d349424dd1)。模型、拓扑、 投机解码、工作负载与配方引用均保持不变。 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
Initial attempt
初始尝试
|
|
Klaud Cold: failed. Finishing cleanup; owned child runs will be stopped and checked before closure. Klaud Cold:failed。正在完成清理;将先停止并确认自有子运行的状态,再关闭 PR。 |
|
Klaud Cold: failed. All owned runs are terminal. Repairs: 0. Runs: —. PR closed; the exact-candidate branch is retained for manual review. The interruption does not prove image incompatibility. Klaud Cold:failed。所有自有运行均已结束。修复次数:0。运行:—。 PR 已关闭;保留该候选的分支,等待人工审查。运行中断不能证明镜像不兼容。 |
Update the
dsr1-fp8-h100-dynamo-sglangmaster image fromlmsysorg/sglang:v0.5.8-cu130to the current SGLang releaselmsysorg/sglang:v0.5.19-cu130(Docker Hub digestsha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, tag commit sgl-project/sglang@0bcd822). Model, disaggregated 1P/1D TP16 topology, EAGLE/MTP settings, the two upstream srt-slurm recipe references and the 8k1k concurrency lists are unchanged.Baseline
benchmarks?model=DeepSeek-R1-0528&date=2026-02-13&exact=true,workflow-info?date=2026-02-13,evaluations?model=DeepSeek-R1-0528&date=2026-02-13&exact=true)lmsysorg/sglang:v0.5.8-cu130(Docker Hub digestsha256:ef0d14df76c2c90ce651c274bc607600d09426128231722a96d52bb5472e1ebf, tag commit sgl-project/sglang@0189f41). Baseline mismatch: the published rows carry this label, butrunners/launch_h100-dgxc-slurm.shmaps every H100 dynamo-sglang job to the staged squash filelmsysorg_sglang_v0.5.8.post1-cu130.sqsh, so the producing container waslmsysorg/sglang:v0.5.8.post1-cu130(digestsha256:b6f9f50829ec45428db4451616978c45eb3d8d91488a702e78bccd06154d6bef).cluster:h100-dgxc,deepseek-ai/DeepSeek-R1-0528FP8, Dynamo + SGLang disaggregated, NIXL KV transfer, EAGLE MTP (2 steps, top-k 1, 3 draft tokens), fixed-seq-len 8k1k (ISL 8192 / OSL 1024), two deployment shapes of 4 nodes each: 1 prefill worker TP16 EP1 plus 1 decode worker TP16 EP1 (upstream reciperecipes/h100/8k1k/mtp/h100-fp8-1p1d-max-tp-mtp.yaml, concurrency 1-128) and 1 prefill worker TP16 EP1 plus 1 decode worker TP16 EP16 with DP attention (recipes/h100/8k1k/mtp/h100-fp8-1p1d-max-dep-mtp.yaml, concurrency 1-64), both from NVIDIA/srt-slurmsa-submission-q2-202678d48618f771603c6a06e61b0083362814d30919, changelog PRs #643 and #644); 15 benchmark points, result IDs below1P TP16 EP1 / 1D TP16 EP1 (TEP)
1P TP16 EP1 / 1D TP16 EP16 DP-attention (DEP)
em_strict0.9545 /em_flexible0.9553, DEP concurrency 32em_strict0.9575 /em_flexible0.9583) and are not part of the frozen baseline.将
dsr1-fp8-h100-dynamo-sglang的主配置镜像从lmsysorg/sglang:v0.5.8-cu130升级到当前 SGLang 发布版lmsysorg/sglang:v0.5.19-cu130(Docker Hub 摘要sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,标签提交 sgl-project/sglang@0bcd822)。模型、分离式 1P/1D TP16 拓扑、EAGLE/MTP 设置、两个上游 srt-slurm 配方引用以及 8k1k 并发列表均保持不变。基线
benchmarks?model=DeepSeek-R1-0528&date=2026-02-13&exact=true,workflow-info?date=2026-02-13,evaluations?model=DeepSeek-R1-0528&date=2026-02-13&exact=true)lmsysorg/sglang:v0.5.8-cu130(Docker Hub 摘要sha256:ef0d14df76c2c90ce651c274bc607600d09426128231722a96d52bb5472e1ebf,标签提交 sgl-project/sglang@0189f41)。基线不一致:已发布数据行标注的是该镜像,但runners/launch_h100-dgxc-slurm.sh将所有 H100 dynamo-sglang 任务映射到预置的 squash 文件lmsysorg_sglang_v0.5.8.post1-cu130.sqsh,因此实际产出数据的容器是lmsysorg/sglang:v0.5.8.post1-cu130(摘要sha256:b6f9f50829ec45428db4451616978c45eb3d8d91488a702e78bccd06154d6bef)。cluster:h100-dgxc,deepseek-ai/DeepSeek-R1-0528FP8,Dynamo + SGLang 分离式部署,NIXL KV 传输,EAGLE MTP(2 步、top-k 1、3 个草稿 token),固定序列长度 8k1k(ISL 8192 / OSL 1024),两种各占 4 节点的部署形态:1 个 prefill worker TP16 EP1 加 1 个 decode worker TP16 EP1(上游配方recipes/h100/8k1k/mtp/h100-fp8-1p1d-max-tp-mtp.yaml,并发 1-128),以及 1 个 prefill worker TP16 EP1 加 1 个 decode worker TP16 EP16 并启用 DP attention(recipes/h100/8k1k/mtp/h100-fp8-1p1d-max-dep-mtp.yaml,并发 1-64),均来自 NVIDIA/srt-slurmsa-submission-q2-202678d48618f771603c6a06e61b0083362814d30919,changelog PR #643 与 #644);共 15 个基准点,结果 ID 见下表1P TP16 EP1 / 1D TP16 EP1(TEP)
1P TP16 EP1 / 1D TP16 EP16 DP-attention(DEP)
em_strict0.9545 /em_flexible0.9553,DEP 并发 32em_strict0.9575 /em_flexible0.9583),不属于冻结的基线。🤖 Generated with Claude Code