Skip to content

[Klaud Cold] Update dsr1-fp8-h200-sglang-mtp SGLang image to v0.5.19-cu130 / 将 dsr1-fp8-h200-sglang-mtp 的 SGLang 镜像升级至 v0.5.19-cu130 - #2955

Merged
adibarra merged 3 commits into
mainfrom
klaud/auto-9a4c66cbfae8eddb-a32566a003a59219
Sep 10, 2026
Merged

[Klaud Cold] Update dsr1-fp8-h200-sglang-mtp SGLang image to v0.5.19-cu130 / 将 dsr1-fp8-h200-sglang-mtp 的 SGLang 镜像升级至 v0.5.19-cu130#2955
adibarra merged 3 commits into
mainfrom
klaud/auto-9a4c66cbfae8eddb-a32566a003a59219

Conversation

@Klaud-Cold

@Klaud-Cold Klaud-Cold commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator

Update the dsr1-fp8-h200-sglang-mtp master image from lmsysorg/sglang:v0.5.12-cu130 to the current SGLang release lmsysorg/sglang:v0.5.19-cu130 (Docker Hub digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, tag commit sgl-project/sglang@0bcd822). The recipe script, model, TP8/EP1 topology, EAGLE/MTP settings and the 8k1k concurrency range are unchanged.

Baseline

  • Published date: 2026-05-20 (benchmarks?model=DeepSeek-R1-0528&date=2026-05-20&exact=true, workflow-info?date=2026-05-20, evaluations?model=DeepSeek-R1-0528&date=2026-05-20&exact=true)
  • Old image: lmsysorg/sglang:v0.5.12-cu130 (digest sha256:42194170546745092e74cd5f81ad32a7c6e944c7111fe7bf13588152277ff356, tag commit sgl-project/sglang@127b9e3)
  • Workload / topology: single-node H200, deepseek-ai/DeepSeek-R1-0528, SGLang FP8, TP8 EP1, EAGLE MTP (2 steps, top-k 1, 3 draft tokens), fixed-seq-len 8k1k (ISL 8192 / OSL 1024), concurrency 4 / 8 / 16 / 32 / 64, random dataset with chat template
  • Producer: run 26135241056 (head 7ec590983c089c6262b1958c8325e9d06ef7f62c, changelog PR #1523); benchmark result IDs 413852, 413854, 413845, 413847, 413850
Conc Total tok/s/GPU Output tok/s/GPU Median TTFT (s) Median TPOT (ms) Median E2E (s)
4 431.2 48.0 0.325 10.09 9.34
8 650.4 73.1 0.343 13.39 12.39
16 932.9 103.4 0.372 18.74 17.70
32 1237.8 138.4 0.372 28.21 26.25
64 1622.4 180.0 0.605 44.23 40.98
  • Published eval: gsm8k at concurrency 32 and 64, em_strict 0.9545 / em_flexible 0.9560 (evaluation IDs 6795 and 6794, same producer run). The public evaluations feed labels these rows disagg: true, which does not match the single-node TP8 recipe; the benchmark rows above carry the correct disagg: false identity.

dsr1-fp8-h200-sglang-mtp 的主配置镜像从 lmsysorg/sglang:v0.5.12-cu130 升级到当前 SGLang 发布版 lmsysorg/sglang:v0.5.19-cu130(Docker Hub 摘要 sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,标签提交 sgl-project/sglang@0bcd822)。配方脚本、模型、TP8/EP1 拓扑、EAGLE/MTP 设置以及 8k1k 并发范围均保持不变。

基线

  • 发布日期: 2026-05-20(benchmarks?model=DeepSeek-R1-0528&date=2026-05-20&exact=trueworkflow-info?date=2026-05-20evaluations?model=DeepSeek-R1-0528&date=2026-05-20&exact=true
  • 旧镜像: lmsysorg/sglang:v0.5.12-cu130(摘要 sha256:42194170546745092e74cd5f81ad32a7c6e944c7111fe7bf13588152277ff356,标签提交 sgl-project/sglang@127b9e3
  • 工作负载 / 拓扑: 单节点 H200,deepseek-ai/DeepSeek-R1-0528,SGLang FP8,TP8 EP1,EAGLE MTP(2 步、top-k 1、3 个草稿 token),固定序列长度 8k1k(ISL 8192 / OSL 1024),并发 4 / 8 / 16 / 32 / 64,随机数据集并使用聊天模板
  • 数据来源: 运行 26135241056(head 7ec590983c089c6262b1958c8325e9d06ef7f62c,changelog PR #1523);基准结果 ID 413852、413854、413845、413847、413850
并发 总吞吐 tok/s/GPU 输出吞吐 tok/s/GPU TTFT 中位数 (s) TPOT 中位数 (ms) 端到端中位数 (s)
4 431.2 48.0 0.325 10.09 9.34
8 650.4 73.1 0.343 13.39 12.39
16 932.9 103.4 0.372 18.74 17.70
32 1237.8 138.4 0.372 28.21 26.25
64 1622.4 180.0 0.605 44.23 40.98
  • 已发布评测: 并发 32 与 64 的 gsm8k,em_strict 0.9545 / em_flexible 0.9560(评测 ID 6795 与 6794,同一数据来源运行)。公共评测接口将这些行标记为 disagg: true,与单节点 TP8 配方不符;上表基准行的 disagg: false 身份是正确的。

🤖 Generated with Claude Code


Note

Low Risk
Benchmark config and changelog only; no application runtime or security-sensitive code paths.

Overview
Bumps the SGLang container image for the dsr1-fp8-h200-sglang-mtp benchmark recipe from lmsysorg/sglang:v0.5.12-cu130 to v0.5.19-cu130 in configs/nvidia-master.yaml.

Adds a matching perf-changelog.yaml entry (PR #2955) noting the new digest and that the launch script, TP8/EP1 topology, EAGLE/MTP settings, and 8k/1k workload are unchanged.

Reviewed by Cursor Bugbot for commit 735b758. Bugbot is set up for automated code reviews on this repo. Configure here.

Update the dsr1-fp8-h200-sglang-mtp master image from lmsysorg/sglang:v0.5.12-cu130
to lmsysorg/sglang:v0.5.19-cu130 (digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,
sglang commit 0bcd822377da7b5718e674eaf9c870d349424dd1). Model, TP8/EP1 topology,
EAGLE/MTP settings, workloads and the recipe script are unchanged.

将 dsr1-fp8-h200-sglang-mtp 的主配置镜像从 lmsysorg/sglang:v0.5.12-cu130 更新至
lmsysorg/sglang:v0.5.19-cu130(摘要 sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,
sglang 提交 0bcd822377da7b5718e674eaf9c870d349424dd1)。模型、TP8/EP1 拓扑、EAGLE/MTP 设置、
工作负载和配方脚本均保持不变。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

@Klaud-Cold

Klaud-Cold commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator Author

Initial attempt

  • Image / SHA: lmsysorg/sglang:v0.5.19-cu130 (digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9) at PR head 63c66bccd36b7443c2a27dba9e66d44778688257
  • Change: master image only (configs/nvidia-master.yaml, key dsr1-fp8-h200-sglang-mtp); benchmarks/single_node/fixed_seq_len/dsr1_fp8_h200_mtp.sh unchanged, no runtime patching
  • Targeted run: e2e run 34446190942 (e2e-tests.yml from main, ref=63c66bcc, test-config --trim-conc: TP8 EP1 MTP at concurrency 4 plus the smoke gsm8k eval; dispatched 2026-09-10T06:39Z) — completed, success: benchmark job on h200-cw_00, eval-only job on h200-cw_01, collectors and success-rate check all green. Server log shows The server is fired up and ready to roll! on both nodes with the resolved settings attention_backend=flashinfer, speculative_algorithm=EAGLE, draft path auto-set to the target model, steps 2 / top-k 1 / draft tokens 3, mem_fraction_static=0.82, disable_radix_cache=True; the only new log line is the expected SGLANG_ENABLE_SPEC_V2 has been removed warning.
  • Benchmark result (concurrency 4, 8k1k, vs published 2026-05-20 baseline, result ID 413852):
Metric Baseline v0.5.12 v0.5.19 Delta
Total tok/s/GPU 431.2 586.7 +36.1%
Output tok/s/GPU 48.0 65.3 +36.1%
Median TTFT (s) 0.325 0.311 -4.3%
Median TPOT (ms) 10.09 7.23 -28.3%
Median E2E (s) 9.34 6.65 -28.8%

Mean speculative accept length over the smoke's 4838 decode batches was 2.49 of a maximum 3 (baseline accept length N/A: not published). power_valid=1 in the aggregate.

  • Eval result: gsm8k at concurrency 4 (smoke selection), em_strict 0.9606 / em_flexible 0.9621, n = 1319, infrastructure_success=true. Published baseline eval was 0.9545 / 0.9560 at concurrency 32 and 64, so the numbers are indicative, not a like-for-like point.
  • Next step: append the perf-changelog.yaml entry, recheck capacity and start the final full sweep (all five concurrency points plus the default evals) with full-sweep-enabled on the draft PR.
  • Upstream source comparison: v0.5.12 @ 127b9e3 (2026-05-16) → v0.5.19 @ 0bcd822 (2026-09-04). Provenance: both images carry ai.sglang.build.commit / org.opencontainers.image.revision labels equal to those tag commits, and v0.5.19 and v0.5.19-cu130 share one Docker Hub digest pushed 2026-09-04 (docker/Dockerfile clones --branch v${SGL_VERSION}).
  • Coupled dependencies: torch 2.11.0 → 2.13.0, flashinfer 0.6.11.post1 → 0.6.18, CUDA base 13.0.1 → 13.0.3 (same CUDA 13.0 / driver ≥ 535 requirement as the old cu130 image).
  • Flag audit: every launch flag used by the script (--speculative-algorithm EAGLE, --speculative-num-steps, --speculative-num-draft-tokens, --speculative-eagle-topk, --attention-backend flashinfer, --mem-fraction-static, --disable-radix-cache, --chunked-prefill-size, --max-prefill-tokens, --cuda-graph-max-bs, --max-running-requests, --stream-interval, --decode-log-interval, --ep-size, --data-parallel-size, --trust-remote-code, --context-length) is still defined in v0.5.19 (server_args.py plus the new arg_groups/ hooks). For DeepseekV3ForCausalLM, speculative_hook.py still auto-sets the MTP draft path and honors the explicit steps/top-k/draft-token values. SGLANG_ENABLE_SPEC_V2=1 (set by the script) is now a removal warning only: spec V2 is always on since v0.5.13 (#25464), whereas v0.5.12 gated V2 behind that env var, so the executed speculative path is unchanged.
  • Behavior deltas noted: CUTLASS FP8 blockwise GEMM was deleted for SM90 in v0.5.16 (#30438) and v0.5.19 adds an SM90 FP8 decode routing fix (#37018), so H200 FP8 GEMM kernels differ from the baseline image; unified radix tree is the default in v0.5.19 (#35081; moot under --disable-radix-cache); compiled-kernel caches moved under SGLANG_CACHE_DIR in v0.5.18 (#32434). The server log also reports no tuned Triton MoE config for E=257,N=256 fp8 on H200 under Triton 3.7.1 and falls back to the previous Triton version's config; this is informational, not an error. The sibling qwen3.5-fp8-h100-sglang-mtp ([Klaud Cold] Update qwen3.5-fp8-h100-sglang-mtp SGLang image to v0.5.19-cu130 / 将 qwen3.5-fp8-h100-sglang-mtp 的 SGLang 镜像升级至 v0.5.19-cu130 #2952) already ran its concurrency-4 EAGLE MTP smoke on v0.5.19-cu130 on Hopper (SM90).
  • Decision: image-only bump; no in-scope flag changes required.

初始尝试

  • 镜像 / SHA: lmsysorg/sglang:v0.5.19-cu130(摘要 sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9),PR head 63c66bccd36b7443c2a27dba9e66d44778688257
  • 改动: 仅主配置镜像(configs/nvidia-master.yaml 中的 dsr1-fp8-h200-sglang-mtp);benchmarks/single_node/fixed_seq_len/dsr1_fp8_h200_mtp.sh 未改动,无任何运行时补丁
  • 定向运行: e2e 运行 34446190942(从 main 触发 e2e-tests.ymlref=63c66bcc,test-config --trim-conc:TP8 EP1 MTP 并发 4 加冒烟 gsm8k 评测;于 2026-09-10T06:39Z 触发)—— 已完成,成功:基准任务在 h200-cw_00,仅评测任务在 h200-cw_01,收集器与成功率检查全部通过。两台节点的服务日志均出现 The server is fired up and ready to roll!,解析后的设置为 attention_backend=flashinferspeculative_algorithm=EAGLE、草稿模型路径自动设为目标模型、步数 2 / top-k 1 / 草稿 token 3、mem_fraction_static=0.82disable_radix_cache=True;唯一新增日志为预期中的 SGLANG_ENABLE_SPEC_V2 has been removed 提示。
  • 基准结果(并发 4,8k1k,对比 2026-05-20 已发布基线,结果 ID 413852):
指标 基线 v0.5.12 v0.5.19 变化
总吞吐 tok/s/GPU 431.2 586.7 +36.1%
输出吞吐 tok/s/GPU 48.0 65.3 +36.1%
TTFT 中位数 (s) 0.325 0.311 -4.3%
TPOT 中位数 (ms) 10.09 7.23 -28.3%
端到端中位数 (s) 9.34 6.65 -28.8%

冒烟运行 4838 个解码批次的平均投机接受长度为 2.49(上限 3)(基线接受长度 N/A:未发布)。聚合结果中 power_valid=1

  • 评测结果: 并发 4 的 gsm8k(冒烟选择),em_strict 0.9606 / em_flexible 0.9621,n = 1319,infrastructure_success=true。已发布基线评测为并发 32 与 64 的 0.9545 / 0.9560,因此该数值仅供参考,不是同一并发点的对比。
  • 下一步: 追加 perf-changelog.yaml 条目,重新检查容量,并在草稿 PR 上以 full-sweep-enabled 启动最终完整 sweep(全部五个并发点及默认评测)。
  • 上游源码对比: v0.5.12 @ 127b9e3(2026-05-16)→ v0.5.19 @ 0bcd822(2026-09-04)。来源确认:两个镜像的 ai.sglang.build.commit / org.opencontainers.image.revision 标签均等于对应标签提交,且 v0.5.19v0.5.19-cu130 在 Docker Hub 上共享同一摘要(2026-09-04 推送);docker/Dockerfile--branch v${SGL_VERSION} 克隆。
  • 耦合依赖: torch 2.11.0 → 2.13.0,flashinfer 0.6.11.post1 → 0.6.18,CUDA 基础镜像 13.0.1 → 13.0.3(与旧 cu130 镜像相同的 CUDA 13.0 / 驱动 ≥ 535 要求)。
  • 参数审计: 脚本使用的全部启动参数(--speculative-algorithm EAGLE--speculative-num-steps--speculative-num-draft-tokens--speculative-eagle-topk--attention-backend flashinfer--mem-fraction-static--disable-radix-cache--chunked-prefill-size--max-prefill-tokens--cuda-graph-max-bs--max-running-requests--stream-interval--decode-log-interval--ep-size--data-parallel-size--trust-remote-code--context-length)在 v0.5.19 中仍有定义(server_args.py 及新的 arg_groups/ 钩子)。对于 DeepseekV3ForCausalLMspeculative_hook.py 仍会自动设置 MTP 草稿模型路径,并沿用显式的步数 / top-k / 草稿 token 数。脚本设置的 SGLANG_ENABLE_SPEC_V2=1 现仅触发移除提示:自 v0.5.13 起 spec V2 始终开启(#25464),而 v0.5.12 由该环境变量控制 V2,因此实际执行的投机解码路径不变。
  • 已注意到的行为差异: v0.5.16 删除了 SM90 的 CUTLASS FP8 blockwise GEMM(#30438),v0.5.19 加入 SM90 FP8 解码路由修复(#37018),因此 H200 FP8 GEMM 内核与基线镜像不同;v0.5.19 默认使用统一 radix 树(#35081,在 --disable-radix-cache 下无影响);v0.5.18 起编译内核缓存迁移至 SGLANG_CACHE_DIR#32434)。服务日志还提示 H200 上 Triton 3.7.1 没有 E=257,N=256 fp8 的调优 MoE 配置并回退到上一 Triton 版本的配置;这只是信息提示,不是错误。同类配方 qwen3.5-fp8-h100-sglang-mtp[Klaud Cold] Update qwen3.5-fp8-h100-sglang-mtp SGLang image to v0.5.19-cu130 / 将 qwen3.5-fp8-h100-sglang-mtp 的 SGLang 镜像升级至 v0.5.19-cu130 #2952)已在 Hopper(SM90)上用 v0.5.19-cu130 完成并发 4 的 EAGLE MTP 冒烟。
  • 决定: 仅升级镜像;无需在范围内调整参数。

Append the perf-changelog entry for updating dsr1-fp8-h200-sglang-mtp from
lmsysorg/sglang:v0.5.12-cu130 to lmsysorg/sglang:v0.5.19-cu130 (PR #2955).

为 dsr1-fp8-h200-sglang-mtp 从 lmsysorg/sglang:v0.5.12-cu130 更新至
lmsysorg/sglang:v0.5.19-cu130(PR #2955)追加 perf-changelog 条目。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@Klaud-Cold

Klaud-Cold commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator Author

Final full sweep

  • Image / SHA: lmsysorg/sglang:v0.5.19-cu130 at PR head 7db0a243f3dec67134c66fd7362fcf11152b6e5e (adds the perf-changelog.yaml entry for dsr1-fp8-h200-sglang-mtp on top of the image bump; validated locally with utils/validate_perf_changelog.py against base 01cea4db)
  • Trigger: full-sweep-enabled applied to the draft PR at 2026-09-10T07:35Z after a passing capacity check. Run 34450728410 (run-sweep.yml, labeled event, attempt 1) is the final sweep; the push-triggered run 34450701759 was cancelled by the workflow's concurrency group before running any benchmark.
  • Scope / outcome: full matrix, TP8 EP1 MTP 8k1k at concurrency 4 / 8 / 16 / 32 / 64 plus the default gsm8k evals at concurrency 32 and 64 — completed, success. All five benchmark jobs ran on h200-dgxc-slurm_00_04, both eval jobs on h200-cw_00 / _01; check-changelog, setup, collect-results, collect-evals, compare-results, upload-changelog-metadata and calc-success-rate are green. results_bmk/agg_bmk.json holds all five points with power_valid=1 and one recipe fingerprint (f5fa3b86…); the klaud-sweep-manifest records full-sweep: true on this head.
  • Benchmark results vs published 2026-05-20 baseline (same TP8 EP1 MTP 8k1k points, result IDs 413852 / 413854 / 413845 / 413847 / 413850):
Conc Total tok/s/GPU Output tok/s/GPU Median TTFT (s) Median TPOT (ms) Median E2E (s)
4 431.2 → 571.9 (+32.6%) 48.0 → 63.6 (+32.6%) 0.325 → 0.328 (+0.7%) 10.09 → 7.42 (-26.4%) 9.34 → 6.91 (-26.0%)
8 650.4 → 791.0 (+21.6%) 73.1 → 88.9 (+21.6%) 0.343 → 0.320 (-6.8%) 13.39 → 10.80 (-19.3%) 12.39 → 10.37 (-16.3%)
16 932.9 → 1081.7 (+15.9%) 103.4 → 119.8 (+15.9%) 0.372 → 0.348 (-6.4%) 18.74 → 16.37 (-12.6%) 17.70 → 15.29 (-13.6%)
32 1237.8 → 1401.9 (+13.3%) 138.4 → 156.8 (+13.3%) 0.372 → 0.365 (-1.9%) 28.21 → 25.21 (-10.6%) 26.25 → 23.60 (-10.1%)
64 1622.3 → 1677.0 (+3.4%) 180.0 → 186.0 (+3.4%) 0.605 → 3.318 (+449%) 44.23 → 39.69 (-10.3%) 40.98 → 39.53 (-3.6%)
  • Regression to note: median TTFT at concurrency 64 rose from 0.60 s to 3.32 s (mean 1.84 s → 4.33 s; p99 19.6 s → 18.3 s is unchanged). Its server log shows the scheduler holding 57–60 running requests with 4–7 requests queued through steady state (#queue-req peaked at 53 during the ramp), so new prompts wait behind the 32k-token chunked prefill batches before their first token. Throughput (+3.4%), TPOT (-10.3%) and E2E (-3.6%) at that point still improved, and the other four points improved on every metric except a +0.7% TTFT at concurrency 4. Mean speculative accept length at concurrency 64 was 2.38 of a maximum 3 (baseline N/A: not published). No flag change is in scope for this image refresh; the regression is reported for reviewer judgment.
  • Eval results: gsm8k em_strict / em_flexible: concurrency 32 → 0.9560 / 0.9560, concurrency 64 → 0.9591 / 0.9606 (n = 1319 each, infrastructure_success=true). Published baseline: 0.9545 / 0.9560 at both points.
  • Status: targeted smoke and exact-head final validation both passed; global approval remains with the automatic reviews and CODEOWNER sign-off. Next step: run the Klaud completion check, which verifies matrix/result coverage and marks the PR ready. Smoke evidence is in the Initial attempt.

最终完整 sweep

  • 镜像 / SHA: lmsysorg/sglang:v0.5.19-cu130,PR head 7db0a243f3dec67134c66fd7362fcf11152b6e5e(在镜像升级之上追加 dsr1-fp8-h200-sglang-mtpperf-changelog.yaml 条目;已用 utils/validate_perf_changelog.py 对基线 01cea4db 本地校验)
  • 触发: 容量检查通过后,于 2026-09-10T07:35Z 在草稿 PR 上添加 full-sweep-enabled运行 34450728410run-sweep.yml,labeled 事件,第 1 次尝试)为最终 sweep;由推送触发的 运行 34450701759 在执行任何基准之前被工作流并发组取消。
  • 范围 / 结果: 完整矩阵,TP8 EP1 MTP 8k1k 并发 4 / 8 / 16 / 32 / 64,以及并发 32 与 64 的默认 gsm8k 评测 —— 已完成,成功。五个基准任务运行于 h200-dgxc-slurm_00_04,两个评测任务运行于 h200-cw_00 / _01;check-changelog、setup、collect-results、collect-evals、compare-results、upload-changelog-metadata 与 calc-success-rate 均通过。results_bmk/agg_bmk.json 包含全部五个点,power_valid=1,配方指纹一致(f5fa3b86…);klaud-sweep-manifest 在该 head 上记录 full-sweep: true
  • 基准结果对比 2026-05-20 已发布基线(相同 TP8 EP1 MTP 8k1k 点,结果 ID 413852 / 413854 / 413845 / 413847 / 413850):
并发 总吞吐 tok/s/GPU 输出吞吐 tok/s/GPU TTFT 中位数 (s) TPOT 中位数 (ms) 端到端中位数 (s)
4 431.2 → 571.9 (+32.6%) 48.0 → 63.6 (+32.6%) 0.325 → 0.328 (+0.7%) 10.09 → 7.42 (-26.4%) 9.34 → 6.91 (-26.0%)
8 650.4 → 791.0 (+21.6%) 73.1 → 88.9 (+21.6%) 0.343 → 0.320 (-6.8%) 13.39 → 10.80 (-19.3%) 12.39 → 10.37 (-16.3%)
16 932.9 → 1081.7 (+15.9%) 103.4 → 119.8 (+15.9%) 0.372 → 0.348 (-6.4%) 18.74 → 16.37 (-12.6%) 17.70 → 15.29 (-13.6%)
32 1237.8 → 1401.9 (+13.3%) 138.4 → 156.8 (+13.3%) 0.372 → 0.365 (-1.9%) 28.21 → 25.21 (-10.6%) 26.25 → 23.60 (-10.1%)
64 1622.3 → 1677.0 (+3.4%) 180.0 → 186.0 (+3.4%) 0.605 → 3.318 (+449%) 44.23 → 39.69 (-10.3%) 40.98 → 39.53 (-3.6%)
  • 需注意的回退: 并发 64 的 TTFT 中位数从 0.60 s 升至 3.32 s(均值 1.84 s → 4.33 s;p99 19.6 s → 18.3 s 基本不变)。其服务日志显示调度器在稳态下保持 57–60 个运行请求、4–7 个请求排队(#queue-req 在爬升阶段峰值为 53),因此新提示需等待 32k token 的分块 prefill 批次之后才能产出首个 token。该点的吞吐(+3.4%)、TPOT(-10.3%)与端到端(-3.6%)仍有改善,其余四个点除并发 4 的 TTFT +0.7% 外各项指标均改善。并发 64 的平均投机接受长度为 2.38(上限 3)(基线 N/A:未发布)。本次镜像刷新范围内不包含参数调整;该回退如实报告,供审阅者判断。
  • 评测结果: gsm8k em_strict / em_flexible:并发 32 → 0.9560 / 0.9560,并发 64 → 0.9591 / 0.9606(各 n = 1319,infrastructure_success=true)。已发布基线:两点均为 0.9545 / 0.9560。
  • 状态: 定向冒烟与精确 head 的最终验证均已通过;全局批准仍取决于自动审阅与 CODEOWNER 签署。下一步:运行 Klaud 完成检查,由其校验矩阵/结果覆盖并将 PR 标记为可审阅。冒烟证据见 初始尝试

@github-actions

Copy link
Copy Markdown
Contributor

@Klaud-Cold
Klaud-Cold marked this pull request as ready for review September 10, 2026 08:13
@Klaud-Cold
Klaud-Cold requested a review from a team September 10, 2026 08:13
@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Klaud Cold: validated. All owned runs are terminal. Repairs: 0. Runs: 34446190942, 34450701759, 34450728410.

The full sweep is verified; this PR remains ready for review.


Klaud Cold:validated。所有自有运行均已结束。修复次数:0。运行:34446190942, 34450701759, 34450728410

完整 sweep 已通过验证;PR 保持就绪,等待审查。

@claude

claude Bot commented Sep 10, 2026

Copy link
Copy Markdown
Contributor

Claude finished @Klaud-Cold's task in 1m 14s —— View job


Review of PR #2955

  • Gather context: inspect the diff against origin/main
  • Validate configs/nvidia-master.yaml change (image accessibility, key prefix, recipe unchanged)
  • Validate perf-changelog.yaml entry (appended at end, validator passes)
  • Post review

LGTM - no blocking issues found.

Image-only bump of dsr1-fp8-h200-sglang-mtp to the public Docker Hub tag lmsysorg/sglang:v0.5.19-cu130; the perf-changelog.yaml entry is appended at the very end of the file with the correct config key and PR link, and the unchanged recipe script still passes --use-chat-template and --ep-size $EP_SIZE. The full sweep on this exact head (run 34450728410) is green, including check-changelog, so the validator result is covered by CI even though utils/validate_perf_changelog.py could not be executed in this review session (Bash execution of that script was not permitted). The concurrency-64 median TTFT increase the author reported is a runtime observation for CODEOWNER judgment, not a defect in the diff.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this PR and didn't find any bugs. Because configs/nvidia-master.yaml is CODEOWNER-owned (per .github/CODEOWNERS: @ ankur-singh, @ kedarpotdar-nv, @ InferenceX/core) and image bumps require verifying the tag is an unmodified upstream SGLang build, a human/CODEOWNER look would still be worthwhile.

What was reviewed:

  • The image bump itself (v0.5.12-cu130 -> v0.5.19-cu130) for the dsr1-fp8-h200-sglang-mtp recipe; all other fields (model, model-prefix, runner, TP8/EP1 topology, MTP settings) are unchanged.
  • The perf-changelog.yaml entry: correctly appended at the physical tail, no historical entries edited, schema (config-keys/description/pr-link) matches surrounding entries.
  • No model.container field exists for this single-node recipe, so the container/image parity rule for multi-node srt-slurm recipes does not apply here.
Extended reasoning...

Overview

This PR is a minimal, mechanical two-file change: a single Docker image tag bump for the dsr1-fp8-h200-sglang-mtp recipe in configs/nvidia-master.yaml (v0.5.12-cu130 to v0.5.19-cu130), and a matching append-only entry in perf-changelog.yaml. I confirmed via git diff that no other fields in the recipe changed (model, model-prefix, runner, TP8/EP1 topology, and EAGLE/MTP settings are untouched), and that the changelog entry was appended strictly at the physical tail without modifying any prior entries, using the same config-keys/description/pr-link schema as neighboring entries.

Security risks

None identified. This is a config value change (image tag) plus a documentation-style changelog append; there is no code execution path, credential handling, or auth/permissions logic involved in the diff itself.

Level of scrutiny

This would normally warrant a light-touch review as a version-bump-only change with a satisfied changelog requirement. However, configs/nvidia-master.yaml is explicitly CODEOWNER-owned (@ ankur-singh, @ kedarpotdar-nv, @ InferenceX/core per .github/CODEOWNERS), and the repository's own review checklist calls for verifying that a bumped SGLang image is unmodified upstream and that any corresponding cookbook PR is merged — something I cannot verify from the diff alone. That combination (CODEOWNER path + external verification requirement) is why I'm deferring rather than approving, per the guideline to not auto-approve CODEOWNER-owned paths.

Other factors

The bug-hunting pass reported no findings, and there's no unresolved third-party objection visible in the timeline. The jump spans several SGLang minor versions (0.5.12 to 0.5.19) on an EAGLE/MTP recipe, which is a reasonable magnitude for a human with domain context to sanity-check even though nothing in the repo's known regression notes (KLAUD_DEBUG.md) points at H200-specific issues in that range.

@adibarra

adibarra commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator

/reuse-sweep-run 34450728410

将 main 合并到 PR #2955,保留已验证的配方并复用完整扫描结果。
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Development

Successfully merging this pull request may close these issues.

2 participants