[Klaud Cold] Update dsr1-fp8-h200-sglang-mtp SGLang image to v0.5.19-cu130 / 将 dsr1-fp8-h200-sglang-mtp 的 SGLang 镜像升级至 v0.5.19-cu130 - #2955
Conversation
Update the dsr1-fp8-h200-sglang-mtp master image from lmsysorg/sglang:v0.5.12-cu130 to lmsysorg/sglang:v0.5.19-cu130 (digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, sglang commit 0bcd822377da7b5718e674eaf9c870d349424dd1). Model, TP8/EP1 topology, EAGLE/MTP settings, workloads and the recipe script are unchanged. 将 dsr1-fp8-h200-sglang-mtp 的主配置镜像从 lmsysorg/sglang:v0.5.12-cu130 更新至 lmsysorg/sglang:v0.5.19-cu130(摘要 sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, sglang 提交 0bcd822377da7b5718e674eaf9c870d349424dd1)。模型、TP8/EP1 拓扑、EAGLE/MTP 设置、 工作负载和配方脚本均保持不变。 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
Initial attempt
Mean speculative accept length over the smoke's 4838 decode batches was 2.49 of a maximum 3 (baseline accept length N/A: not published).
初始尝试
冒烟运行 4838 个解码批次的平均投机接受长度为 2.49(上限 3)(基线接受长度 N/A:未发布)。聚合结果中
|
Append the perf-changelog entry for updating dsr1-fp8-h200-sglang-mtp from lmsysorg/sglang:v0.5.12-cu130 to lmsysorg/sglang:v0.5.19-cu130 (PR #2955). 为 dsr1-fp8-h200-sglang-mtp 从 lmsysorg/sglang:v0.5.12-cu130 更新至 lmsysorg/sglang:v0.5.19-cu130(PR #2955)追加 perf-changelog 条目。 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Final full sweep
最终完整 sweep
|
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=34450728410 |
|
Klaud Cold: validated. All owned runs are terminal. Repairs: 0. Runs: 34446190942, 34450701759, 34450728410. The full sweep is verified; this PR remains ready for review. Klaud Cold:validated。所有自有运行均已结束。修复次数:0。运行:34446190942, 34450701759, 34450728410。 完整 sweep 已通过验证;PR 保持就绪,等待审查。 |
|
Claude finished @Klaud-Cold's task in 1m 14s —— View job Review of PR #2955
LGTM - no blocking issues found. Image-only bump of |
There was a problem hiding this comment.
I reviewed this PR and didn't find any bugs. Because configs/nvidia-master.yaml is CODEOWNER-owned (per .github/CODEOWNERS: @ ankur-singh, @ kedarpotdar-nv, @ InferenceX/core) and image bumps require verifying the tag is an unmodified upstream SGLang build, a human/CODEOWNER look would still be worthwhile.
What was reviewed:
- The image bump itself (
v0.5.12-cu130->v0.5.19-cu130) for thedsr1-fp8-h200-sglang-mtprecipe; all other fields (model, model-prefix, runner, TP8/EP1 topology, MTP settings) are unchanged. - The
perf-changelog.yamlentry: correctly appended at the physical tail, no historical entries edited, schema (config-keys/description/pr-link) matches surrounding entries. - No
model.containerfield exists for this single-node recipe, so the container/image parity rule for multi-node srt-slurm recipes does not apply here.
Extended reasoning...
Overview
This PR is a minimal, mechanical two-file change: a single Docker image tag bump for the dsr1-fp8-h200-sglang-mtp recipe in configs/nvidia-master.yaml (v0.5.12-cu130 to v0.5.19-cu130), and a matching append-only entry in perf-changelog.yaml. I confirmed via git diff that no other fields in the recipe changed (model, model-prefix, runner, TP8/EP1 topology, and EAGLE/MTP settings are untouched), and that the changelog entry was appended strictly at the physical tail without modifying any prior entries, using the same config-keys/description/pr-link schema as neighboring entries.
Security risks
None identified. This is a config value change (image tag) plus a documentation-style changelog append; there is no code execution path, credential handling, or auth/permissions logic involved in the diff itself.
Level of scrutiny
This would normally warrant a light-touch review as a version-bump-only change with a satisfied changelog requirement. However, configs/nvidia-master.yaml is explicitly CODEOWNER-owned (@ ankur-singh, @ kedarpotdar-nv, @ InferenceX/core per .github/CODEOWNERS), and the repository's own review checklist calls for verifying that a bumped SGLang image is unmodified upstream and that any corresponding cookbook PR is merged — something I cannot verify from the diff alone. That combination (CODEOWNER path + external verification requirement) is why I'm deferring rather than approving, per the guideline to not auto-approve CODEOWNER-owned paths.
Other factors
The bug-hunting pass reported no findings, and there's no unresolved third-party objection visible in the timeline. The jump spans several SGLang minor versions (0.5.12 to 0.5.19) on an EAGLE/MTP recipe, which is a reasonable magnitude for a human with domain context to sanity-check even though nothing in the repo's known regression notes (KLAUD_DEBUG.md) points at H200-specific issues in that range.
|
/reuse-sweep-run 34450728410 |
将 main 合并到 PR #2955,保留已验证的配方并复用完整扫描结果。
Update the
dsr1-fp8-h200-sglang-mtpmaster image fromlmsysorg/sglang:v0.5.12-cu130to the current SGLang releaselmsysorg/sglang:v0.5.19-cu130(Docker Hub digestsha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, tag commit sgl-project/sglang@0bcd822). The recipe script, model, TP8/EP1 topology, EAGLE/MTP settings and the 8k1k concurrency range are unchanged.Baseline
benchmarks?model=DeepSeek-R1-0528&date=2026-05-20&exact=true,workflow-info?date=2026-05-20,evaluations?model=DeepSeek-R1-0528&date=2026-05-20&exact=true)lmsysorg/sglang:v0.5.12-cu130(digestsha256:42194170546745092e74cd5f81ad32a7c6e944c7111fe7bf13588152277ff356, tag commit sgl-project/sglang@127b9e3)deepseek-ai/DeepSeek-R1-0528, SGLang FP8, TP8 EP1, EAGLE MTP (2 steps, top-k 1, 3 draft tokens), fixed-seq-len 8k1k (ISL 8192 / OSL 1024), concurrency 4 / 8 / 16 / 32 / 64, random dataset with chat template7ec590983c089c6262b1958c8325e9d06ef7f62c, changelog PR #1523); benchmark result IDs 413852, 413854, 413845, 413847, 413850em_strict0.9545 /em_flexible0.9560 (evaluation IDs 6795 and 6794, same producer run). The public evaluations feed labels these rowsdisagg: true, which does not match the single-node TP8 recipe; the benchmark rows above carry the correctdisagg: falseidentity.将
dsr1-fp8-h200-sglang-mtp的主配置镜像从lmsysorg/sglang:v0.5.12-cu130升级到当前 SGLang 发布版lmsysorg/sglang:v0.5.19-cu130(Docker Hub 摘要sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,标签提交 sgl-project/sglang@0bcd822)。配方脚本、模型、TP8/EP1 拓扑、EAGLE/MTP 设置以及 8k1k 并发范围均保持不变。基线
benchmarks?model=DeepSeek-R1-0528&date=2026-05-20&exact=true,workflow-info?date=2026-05-20,evaluations?model=DeepSeek-R1-0528&date=2026-05-20&exact=true)lmsysorg/sglang:v0.5.12-cu130(摘要sha256:42194170546745092e74cd5f81ad32a7c6e944c7111fe7bf13588152277ff356,标签提交 sgl-project/sglang@127b9e3)deepseek-ai/DeepSeek-R1-0528,SGLang FP8,TP8 EP1,EAGLE MTP(2 步、top-k 1、3 个草稿 token),固定序列长度 8k1k(ISL 8192 / OSL 1024),并发 4 / 8 / 16 / 32 / 64,随机数据集并使用聊天模板7ec590983c089c6262b1958c8325e9d06ef7f62c,changelog PR #1523);基准结果 ID 413852、413854、413845、413847、413850em_strict0.9545 /em_flexible0.9560(评测 ID 6795 与 6794,同一数据来源运行)。公共评测接口将这些行标记为disagg: true,与单节点 TP8 配方不符;上表基准行的disagg: false身份是正确的。🤖 Generated with Claude Code
Note
Low Risk
Benchmark config and changelog only; no application runtime or security-sensitive code paths.
Overview
Bumps the SGLang container image for the
dsr1-fp8-h200-sglang-mtpbenchmark recipe fromlmsysorg/sglang:v0.5.12-cu130tov0.5.19-cu130inconfigs/nvidia-master.yaml.Adds a matching
perf-changelog.yamlentry (PR #2955) noting the new digest and that the launch script, TP8/EP1 topology, EAGLE/MTP settings, and 8k/1k workload are unchanged.Reviewed by Cursor Bugbot for commit 735b758. Bugbot is set up for automated code reviews on this repo. Configure here.