[Klaud Cold] Update dsr1-fp4-b300-sglang SGLang image to v0.5.19-cu130 / 将 dsr1-fp4-b300-sglang 的 SGLang 镜像更新至 v0.5.19-cu130 - #2951
Conversation
Update the dsr1-fp4-b300-sglang master image from lmsysorg/sglang:v0.5.12-cu130 to lmsysorg/sglang:v0.5.19-cu130, pinned by manifest digest. Model, precision, topology, workloads and the recipe script are unchanged. 将 dsr1-fp4-b300-sglang 的 SGLang 镜像从 lmsysorg/sglang:v0.5.12-cu130 更新至 lmsysorg/sglang:v0.5.19-cu130(按 manifest 摘要固定)。模型、精度、拓扑、 工作负载和配方脚本均保持不变。 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
Initial attempt — failed (infrastructure: model not staged on DSXE)
初次尝试 — 失败(基础设施:模型未在 DSXE 上预置)
|
|
Stop: readiness-blocked. The targeted smoke (run 34422578516) failed because 终止:readiness-blocked。 定向冒烟(run 34422578516)失败,原因是 |
|
Klaud Cold: readiness-blocked. Finishing cleanup; owned child runs will be stopped and checked before closure. Klaud Cold:readiness-blocked。正在完成清理;将先停止并确认自有子运行的状态,再关闭 PR。 |
|
Klaud Cold: readiness-blocked. All owned runs are terminal. Repairs: 0. Runs: 34422578516. PR closed; branch deleted for retry. Klaud Cold:readiness-blocked。所有自有运行均已结束。修复次数:0。运行:34422578516。 PR 已关闭;分支已删除,后续可以重试。 |
Refresh the
dsr1-fp4-b300-sglangsingle-node SGLang image fromlmsysorg/sglang:v0.5.12-cu130tolmsysorg/sglang:v0.5.19-cu130(digestsha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, SGLang commit0bcd822= tagv0.5.19). Model, precision, topology, workloads andbenchmarks/single_node/fixed_seq_len/dsr1_fp4_b300.share unchanged.Baseline
workflow-info?date=2026-05-22andbenchmarks?model=DeepSeek-R1-0528&date=2026-05-22&exact=true, filtered to hardwareb300, frameworksglang, precisionfp4, specnone, non-disagg, ISL/OSL 8192/1024).lmsysorg/sglang:v0.5.12-cu130(SGLang commit127b9e3= tagv0.5.12, CUDA 13.0.1, FlashInfer 0.6.11.post1).nvidia/DeepSeek-R1-0528-FP4-V2, fixed-seq-len 8k/1k single turn, TP4/EP4 conc 1–128 and TP8/EP8 conc 1–16 (13 points), runnerb300(clusterb300-dsxe), evals gsm8k at TP4 conc 64/128.6a4fe2a1272bcd6c9df8ea6539f593088ba90437(#1534). Note: that sweep's overall conclusion is recorded asfailure(it covered six DSR1 SGLang configs); the 13 dsr1-fp4-b300-sglang points were ingested and are the currently published curve. Logical curve snapshot id 1943 is not the producer.Published evals (gsm8k,
evaluations?date=2026-05-22, same producer run; the API labels these rowsdisagg: truealthough the recipe is aggregated single-node):将
dsr1-fp4-b300-sglang单节点 SGLang 镜像从lmsysorg/sglang:v0.5.12-cu130更新至lmsysorg/sglang:v0.5.19-cu130(摘要sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,SGLang 提交0bcd822即标签v0.5.19)。模型、精度、拓扑、工作负载以及benchmarks/single_node/fixed_seq_len/dsr1_fp4_b300.sh均保持不变。基线
workflow-info?date=2026-05-22与benchmarks?model=DeepSeek-R1-0528&date=2026-05-22&exact=true,按硬件b300、框架sglang、精度fp4、无投机解码、非分离式、ISL/OSL 8192/1024 过滤)。lmsysorg/sglang:v0.5.12-cu130(SGLang 提交127b9e3即标签v0.5.12,CUDA 13.0.1,FlashInfer 0.6.11.post1)。nvidia/DeepSeek-R1-0528-FP4-V2,fixed-seq-len 8k/1k 单轮,TP4/EP4 并发 1–128 与 TP8/EP8 并发 1–16(共 13 个点),runnerb300(集群b300-dsxe),评测为 TP4 并发 64/128 的 gsm8k。6a4fe2a1272bcd6c9df8ea6539f593088ba90437(#1534)。说明:该 sweep 的整体结论记录为failure(涵盖六个 DSR1 SGLang 配置),但这 13 个 dsr1-fp4-b300-sglang 点已被摄入并构成当前发布曲线。逻辑曲线快照 id 1943 并非生产运行。基准数值见上方英文表格(吞吐 tok/s/GPU、输出吞吐、中位 TTFT/TPOT/E2EL)。
已发布评测(gsm8k,
evaluations?date=2026-05-22,同一生产运行;API 将这些行标记为disagg: true,但该配方实际为聚合式单节点):TP4/EP4 并发 64 em_strict 0.9575 / em_flexible 0.9575;并发 128 em_strict 0.9545 / em_flexible 0.9560。