[Klaud Cold] Update qwen3.5-fp8-h100-sglang-mtp SGLang image to v0.5.19-cu130 / 将 qwen3.5-fp8-h100-sglang-mtp 的 SGLang 镜像升级至 v0.5.19-cu130 - #2952
Conversation
…0.5.19-cu130 Bump the qwen3.5-fp8-h100-sglang-mtp master image from lmsysorg/sglang:v0.5.14-cu130 to lmsysorg/sglang:v0.5.19-cu130 (Docker Hub digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, tag commit sgl-project/sglang@0bcd822). Recipe script, model, TP8/EP8 topology, EAGLE/MTP settings and the 8k1k concurrency range are unchanged. 将 qwen3.5-fp8-h100-sglang-mtp 的主配置镜像从 lmsysorg/sglang:v0.5.14-cu130 升级到 lmsysorg/sglang:v0.5.19-cu130(Docker Hub 摘要 sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, 标签提交 sgl-project/sglang@0bcd822)。 配方脚本、模型、TP8/EP8 拓扑、EAGLE/MTP 设置以及 8k1k 并发范围保持不变。 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
Initial attempt
初始尝试
|
Repair 1/5
The Triton compile-cache race did not recur; together with the attempt-1 benchmark point this completes the targeted smoke (startup/compatibility evidence, not a full curve).
修复 1/5
Triton 编译缓存竞争未再出现;结合第 1 次尝试的基准点,定向冒烟测试完成(属于启动/兼容性证据,不是完整曲线)。
|
…v0.5.19-cu130 Append the perf-changelog.yaml entry for the qwen3.5-fp8-h100-sglang-mtp SGLang image update to lmsysorg/sglang:v0.5.19-cu130 (PR #2952). 为 qwen3.5-fp8-h100-sglang-mtp 的 SGLang 镜像升级至 lmsysorg/sglang:v0.5.19-cu130(PR #2952)追加 perf-changelog.yaml 条目。 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Final full sweep
最终完整扫描
|
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=34429032159 |
|
Klaud Cold: validated. All owned runs are terminal. Repairs: 1. Runs: 34422693420, 34429017350, 34429032159. The full sweep is verified; this PR remains ready for review. Klaud Cold:validated。所有自有运行均已结束。修复次数:1。运行:34422693420, 34429017350, 34429032159。 完整 sweep 已通过验证;PR 保持就绪,等待审查。 |
|
Claude finished @Klaud-Cold's task in 1m 13s —— View job Review: PR #2952
LGTM - no blocking issues found The diff is exactly two hunks: the |
There was a problem hiding this comment.
Looks good, straightforward config change.
What was reviewed: the configs/nvidia-master.yaml hunk (confirmed only the image tag for qwen3.5-fp8-h100-sglang-mtp changed, all other recipe fields untouched) and the perf-changelog.yaml addition (confirmed it is a single append at the file tail with no reordering or byte changes to prior entries). Also checked for a model.container field that would need to mirror the image bump — this recipe schema has no container field anywhere in the file, so that concern doesn't apply.
Extended reasoning...
Overview
The PR changes exactly two files: configs/nvidia-master.yaml, where the qwen3.5-fp8-h100-sglang-mtp recipe's image field is bumped from lmsysorg/sglang:v0.5.14-cu130 to lmsysorg/sglang:v0.5.19-cu130, and perf-changelog.yaml, which gets one new entry appended at the end of the file documenting the bump, referencing the target image digest, and noting the recipe script/topology/EAGLE-MTP settings are unchanged. I ran the diff directly and confirmed both hunks match this description precisely — no other lines in the recipe block changed, and the changelog addition is a clean append with no modification to existing bytes.
Security risks
None. This is a data-only configuration change (a container image tag bump and a changelog entry) with no code paths, credentials, or access-control logic involved.
Level of scrutiny
Low. This is exactly the kind of small, mechanical, self-contained change the approval guidelines call out as approvable — a version bump with no ambiguity and a matching, correctly-formed changelog entry. The repo's own review conventions (append-only changelog, image-bump-must-be-a-real-image-change) are both satisfied per my direct read of the diff.
Other factors
The repository convention note raised the question of whether a model.container field also needs updating alongside image. I checked and this recipe/config schema has no container field anywhere in the file, so that check does not apply here. No CODEOWNERS restriction is evident for these two files, and the bug-hunting pass reported zero findings.
|
/reuse-sweep-run 34429032159 |
将 main 合并到 PR #2952,保留已验证的配方并复用完整扫描结果。
Update the
qwen3.5-fp8-h100-sglang-mtpmaster image fromlmsysorg/sglang:v0.5.14-cu130to the current SGLang releaselmsysorg/sglang:v0.5.19-cu130(Docker Hub digestsha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, tag commit sgl-project/sglang@0bcd822). The recipe script, model, TP8/EP8 topology, EAGLE/MTP settings and the 8k1k concurrency range are unchanged.Baseline
benchmarks?model=Qwen-3.5-397B-A17B&date=2026-07-05&exact=true,workflow-info?date=2026-07-05)lmsysorg/sglang:v0.5.14-cu130(digestsha256:5027e95bf6ec536856b1b52a91d1f35ff5c564ab83e8a94758a169ff09bb8df3, tag commit sgl-project/sglang@49e384c)Qwen/Qwen3.5-397B-A17B-FP8, SGLang FP8, TP8 EP8, EAGLE MTP (3 steps, top-k 1, 4 draft tokens), fixed-seq-len 8k1k (ISL 8192 / OSL 1024), concurrency 4 / 8 / 16 / 32, random dataset with chat template2eaf828b84ebe5cac3c68fabd8a18be99fc6d924, changelog PR #2060); benchmark result IDs 432614, 432616, 432593, 432601em_strict0.9697 /em_flexible0.9598 (evaluation ID 8335, same producer run). The public evaluations feed labels this rowdisagg: truewith 64 prefill/decode GPUs, which does not match the single-node TP8 recipe; the benchmark rows above carry the correctdisagg: falseidentity.将
qwen3.5-fp8-h100-sglang-mtp的主配置镜像从lmsysorg/sglang:v0.5.14-cu130升级到当前 SGLang 发布版lmsysorg/sglang:v0.5.19-cu130(Docker Hub 摘要sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,标签提交 sgl-project/sglang@0bcd822)。配方脚本、模型、TP8/EP8 拓扑、EAGLE/MTP 设置以及 8k1k 并发范围均保持不变。基线
benchmarks?model=Qwen-3.5-397B-A17B&date=2026-07-05&exact=true,workflow-info?date=2026-07-05)lmsysorg/sglang:v0.5.14-cu130(摘要sha256:5027e95bf6ec536856b1b52a91d1f35ff5c564ab83e8a94758a169ff09bb8df3,标签提交 sgl-project/sglang@49e384c)Qwen/Qwen3.5-397B-A17B-FP8,SGLang FP8,TP8 EP8,EAGLE MTP(3 步、top-k 1、4 个草稿 token),固定序列长度 8k1k(ISL 8192 / OSL 1024),并发 4 / 8 / 16 / 32,随机数据集并使用聊天模板2eaf828b84ebe5cac3c68fabd8a18be99fc6d924,changelog PR #2060);基准结果 ID 432614、432616、432593、432601em_strict0.9697 /em_flexible0.9598(评测 ID 8335,同一数据来源运行)。公共评测接口将该行标记为disagg: true且 64 个 prefill/decode GPU,与单节点 TP8 配方不符;上表基准行的disagg: false身份是正确的。Note
Low Risk
Config-only SGLang container image version bump for a single benchmark key; no application logic or security-sensitive code changes.
Overview
Bumps the
qwen3.5-fp8-h100-sglang-mtpbenchmark recipe innvidia-master.yamlfromlmsysorg/sglang:v0.5.14-cu130tolmsysorg/sglang:v0.5.19-cu130.Adds a matching
perf-changelog.yamlentry noting the digest and that the recipe script, TP8/EP8 topology, and EAGLE/MTP settings are unchanged.Reviewed by Cursor Bugbot for commit 9cda017. Bugbot is set up for automated code reviews on this repo. Configure here.