Skip to content

[Klaud Cold] Update dsr1-fp8-h200-sglang SGLang image to v0.5.19-cu130 / 将 dsr1-fp8-h200-sglang 的 SGLang 镜像更新至 v0.5.19-cu130 - #3010

Open
Klaud-Cold wants to merge 3 commits into
mainfrom
klaud/auto-0819223abfa858c3-a32566a003a59219
Open

[Klaud Cold] Update dsr1-fp8-h200-sglang SGLang image to v0.5.19-cu130 / 将 dsr1-fp8-h200-sglang 的 SGLang 镜像更新至 v0.5.19-cu130#3010
Klaud-Cold wants to merge 3 commits into
mainfrom
klaud/auto-0819223abfa858c3-a32566a003a59219

Conversation

@Klaud-Cold

@Klaud-Cold Klaud-Cold commented Sep 11, 2026

Copy link
Copy Markdown
Collaborator

Update the dsr1-fp8-h200-sglang master image from lmsysorg/sglang:v0.5.12-cu130 to the current SGLang release lmsysorg/sglang:v0.5.19-cu130 (Docker Hub index digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, build commit 0bcd822). Model, TP8 topology, 8k1k workload, concurrency grid, launch flags and evals are unchanged.

Baseline

  • Published date: 2026-05-17, old image lmsysorg/sglang:v0.5.12-cu130 (build commit 127b9e3)
  • Identity: DeepSeek-R1-0528, H200, SGLang, FP8, no speculative decoding, single node, TP8 EP1, Single-turn 8k1k, random-range-ratio dataset, concurrency 4 / 8 / 16 / 32 / 64
  • Producer: run 25980019087 attempt 2 ("Update dsr1-fp8-h200-sglang SGLang image to v0.5.12-cu130", PR #1423, head aa2df953); the same curve is still the latest published data for this family as of 2026-09-11
  • API queries: /api/v1/workflow-info?date=2026-05-17, /api/v1/benchmarks?model=DeepSeek-R1-0528&date=2026-05-17&exact=true, /api/v1/evaluations?model=DeepSeek-R1-0528&date=2026-05-17&exact=true
Conc Result ID Total tok/s/GPU Output tok/s/GPU Median TTFT (s) Median TPOT (ms) Median E2E (s)
4 409933 387.6 43.1 0.312 10.98 10.32
8 409937 607.9 68.3 0.314 13.93 13.23
16 409940 847.3 93.9 0.330 20.22 19.01
32 409935 1082.1 121.0 0.348 32.01 29.96
64 409941 1360.9 151.0 0.690 50.70 47.70
  • Evals (gsm8k, n = 1319, same producer run): concurrency 32 em_strict 0.9583 / em_flexible 0.9606 (ID 6394); concurrency 64 em_strict 0.9583 / em_flexible 0.9621 (ID 6395). The published eval rows carry disagg=true while the benchmark rows carry disagg=false for the same run; treated as an identity-metadata quirk, not a different deployment.

dsr1-fp8-h200-sglang 的主配置镜像从 lmsysorg/sglang:v0.5.12-cu130 更新至当前 SGLang 发行版 lmsysorg/sglang:v0.5.19-cu130(Docker Hub 索引摘要 sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,构建提交 0bcd822)。模型、TP8 拓扑、8k1k 工作负载、并发网格、启动参数与评测均保持不变。

基线

  • 发布日期: 2026-05-17,旧镜像 lmsysorg/sglang:v0.5.12-cu130(构建提交 127b9e3
  • 身份: DeepSeek-R1-0528,H200,SGLang,FP8,无投机解码,单节点,TP8 EP1,单轮 8k1k,random-range-ratio 数据集,并发 4 / 8 / 16 / 32 / 64
  • 生产运行: 运行 25980019087 第 2 次尝试("Update dsr1-fp8-h200-sglang SGLang image to v0.5.12-cu130",PR #1423,head aa2df953);截至 2026-09-11,该曲线仍是本家族最新的已发布数据
  • API 查询: /api/v1/workflow-info?date=2026-05-17/api/v1/benchmarks?model=DeepSeek-R1-0528&date=2026-05-17&exact=true/api/v1/evaluations?model=DeepSeek-R1-0528&date=2026-05-17&exact=true
并发 结果 ID 总吞吐 tok/s/GPU 输出吞吐 tok/s/GPU TTFT 中位数 (s) TPOT 中位数 (ms) 端到端中位数 (s)
4 409933 387.6 43.1 0.312 10.98 10.32
8 409937 607.9 68.3 0.314 13.93 13.23
16 409940 847.3 93.9 0.330 20.22 19.01
32 409935 1082.1 121.0 0.348 32.01 29.96
64 409941 1360.9 151.0 0.690 50.70 47.70
  • 评测(gsm8k,n = 1319,同一生产运行): 并发 32 em_strict 0.9583 / em_flexible 0.9606(ID 6394);并发 64 em_strict 0.9583 / em_flexible 0.9621(ID 6395)。已发布的评测行标记为 disagg=true,而同一运行的基准行为 disagg=false;视为身份元数据的差异,而非不同的部署。

🤖 Generated with Claude Code


Note

Low Risk
Config-only Docker image pin and changelog; no runtime code, auth, or workload definition changes.

Overview
Bumps the dsr1-fp8-h200-sglang master benchmark config from lmsysorg/sglang:v0.5.12-cu130 to lmsysorg/sglang:v0.5.19-cu130. DeepSeek-R1-0528 on H200 (FP8, SGLang), TP8 topology, 8k/1k scenario, concurrency sweep, launch flags, and evals are unchanged—only the container image tag moves forward.

Adds a matching perf-changelog.yaml entry for dsr1-fp8-h200-sglang documenting the image update (PR #3010).

Reviewed by Cursor Bugbot for commit f207db0. Bugbot is set up for automated code reviews on this repo. Configure here.

Update the dsr1-fp8-h200-sglang master image from lmsysorg/sglang:v0.5.12-cu130
to lmsysorg/sglang:v0.5.19-cu130. Model, TP8 topology, 8k1k workload,
concurrency grid, launch flags and evals are unchanged.

将 dsr1-fp8-h200-sglang 的主配置镜像从 lmsysorg/sglang:v0.5.12-cu130 更新至
lmsysorg/sglang:v0.5.19-cu130。模型、TP8 拓扑、8k1k 工作负载、并发网格、
启动参数与评测保持不变。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

@Klaud-Cold

Klaud-Cold commented Sep 11, 2026

Copy link
Copy Markdown
Collaborator Author

Initial attempt

  • Image / SHA: lmsysorg/sglang:v0.5.19-cu130 (Docker Hub index digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, amd64 manifest sha256:37bbbd34…) at PR head d4581fe7120df83980d6d07acb28664afb5a6709
  • Change: master image only (configs/nvidia-master.yaml, key dsr1-fp8-h200-sglang); benchmarks/single_node/fixed_seq_len/dsr1_fp8_h200.sh unchanged, no runtime patching
  • Targeted run: e2e run 34592860782 (e2e-tests.yml from main, ref=d4581fe7, test-config --trim-conc: TP8 EP1 at concurrency 4 plus the smoke gsm8k eval; dispatched 2026-09-11T11:12Z) — completed, success (1h 31m): benchmark job on h200-cw_01 (Slurm node slurm-h200-205-005), eval-only job on h200-cw_00; collect-results, collect-evals and calc-success-rate green. Server log: The server is fired up and ready to roll! after 11 min with resolved attention_backend=flashinfer, mem_fraction_static=0.82, disable_radix_cache=True, max_running_requests=256, cuda_graph_max_bs_decode=256 (from the deprecated --cuda-graph-max-bs alias, warning only), chunked_prefill_size=32768, page_size=1, moe_runner_backend=auto, fp8_gemm_runner_backend=auto, max_total_num_tokens=509930; no tracebacks.
  • Benchmark result (concurrency 4, 8k1k TP8 EP1, power_valid=1, 8-GPU avg 3418 W):
Metric Baseline v0.5.12 (2026-05-17, ID 409933) v0.5.19 smoke Delta
Total tok/s/GPU 387.6 391.9 +1.1%
Output tok/s/GPU 43.1 43.6 +1.1%
Median TTFT (s) 0.312 0.295 -5.4%
Median TPOT (ms) 10.98 10.88 -0.9%
Median E2E (s) 10.32 10.21 -1.0%
  • Eval result: gsm8k at concurrency 4 (smoke selection), em_strict 0.9591 / em_flexible 0.9598, n = 1319, infrastructure_success=true. Published baseline evals are at concurrency 32 (0.9583 / 0.9606) and 64 (0.9583 / 0.9621), so the smoke number is indicative rather than a like-for-like point.
  • Upstream source comparison: v0.5.12 @ 127b9e3 (2026-05-16) → v0.5.19 @ 0bcd822 (2026-09-04), 4893 commits. Provenance: both images carry ai.sglang.build.commit / org.opencontainers.image.revision labels equal to the annotated-tag commits above; v0.5.19-cu130 shares its digest with v0.5.19 (pushed 2026-09-04T22:50Z).
  • Coupled dependencies: CUDA base 13.0.1 → 13.0.3 (same driver>=535 requirement), cuDNN 9.13 → 9.14, FlashInfer 0.6.11.post1 → 0.6.18.
  • Flag audit: every option the script passes is still declared in v0.5.19 — --model-path, --trust-remote-code, --context-length, --mem-fraction-static, --max-running-requests, --chunked-prefill-size, --max-prefill-tokens, --disable-radix-cache, --stream-interval, --decode-log-interval, --attention-backend in server_args.py; --tensor-parallel-size / --data-parallel-size remain aliases of tp_size / dp_size. --cuda-graph-max-bs is now a deprecated alias that stores into cuda_graph_max_bs_decode (DeprecatedAliasStoreAction, warning only) and is consumed by cuda_graph_hook.py, so the decode graph cap of 256 is preserved.
  • Default / semantics deltas checked: the DeepSeek-V3 override (model_overrides/deepseek_v2.py) only auto-selects trtllm_mla on SM100 when no backend is set, so the explicit flashinfer backend on H200 (SM90) is untouched; the MoE runner auto-selection to flashinfer_trtllm is SM100-only; memory_hook.py derives mem_fraction_static only when unset, so 0.82 is kept. Kernel-level deltas on Hopper FP8: CUTLASS FP8 blockwise GEMM removed for SM90 (#30438, DeepGEMM stays the default), SM90 FP8 decode routing fix (#37018); unified radix tree default (#35081) is moot under --disable-radix-cache; compiled-kernel caches moved under SGLANG_CACHE_DIR (#32434). The MTP sibling dsr1-fp8-h200-sglang-mtp already runs this exact image and the same non-speculative flags on H200 (#2955, full sweep green 2026-09-10).
  • Decision: image-only bump; no in-scope flag changes required.
  • Next step: the perf-changelog.yaml entry is appended (validated locally against base e444727b); recheck capacity and start the final full sweep (all five concurrency points plus the default gsm8k evals at concurrency 32 and 64) with full-sweep-enabled on the draft PR.

初始尝试

  • 镜像 / SHA: lmsysorg/sglang:v0.5.19-cu130(Docker Hub 索引摘要 sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,amd64 清单 sha256:37bbbd34…),PR head d4581fe7120df83980d6d07acb28664afb5a6709
  • 改动: 仅主配置镜像(configs/nvidia-master.yaml 中的 dsr1-fp8-h200-sglang);benchmarks/single_node/fixed_seq_len/dsr1_fp8_h200.sh 未改动,无任何运行时补丁
  • 定向运行: e2e 运行 34592860782(从 main 触发 e2e-tests.ymlref=d4581fe7,test-config --trim-conc:TP8 EP1 并发 4 加冒烟 gsm8k 评测;于 2026-09-11T11:12Z 触发)—— 已完成,成功(1 小时 31 分):基准任务在 h200-cw_01(Slurm 节点 slurm-h200-205-005),仅评测任务在 h200-cw_00;collect-results、collect-evals 与 calc-success-rate 均通过。服务日志在 11 分钟后出现 The server is fired up and ready to roll!,解析后的设置为 attention_backend=flashinfermem_fraction_static=0.82disable_radix_cache=Truemax_running_requests=256cuda_graph_max_bs_decode=256(来自已弃用的 --cuda-graph-max-bs 别名,仅告警)、chunked_prefill_size=32768page_size=1moe_runner_backend=autofp8_gemm_runner_backend=automax_total_num_tokens=509930;无任何 traceback。
  • 基准结果(并发 4,8k1k TP8 EP1,power_valid=1,8 卡平均功耗 3418 W):
指标 基线 v0.5.12(2026-05-17,ID 409933) v0.5.19 冒烟 变化
总吞吐 tok/s/GPU 387.6 391.9 +1.1%
输出吞吐 tok/s/GPU 43.1 43.6 +1.1%
TTFT 中位数 (s) 0.312 0.295 -5.4%
TPOT 中位数 (ms) 10.98 10.88 -0.9%
端到端中位数 (s) 10.32 10.21 -1.0%
  • 评测结果: 并发 4 的 gsm8k(冒烟选择),em_strict 0.9591 / em_flexible 0.9598,n = 1319,infrastructure_success=true。已发布基线评测位于并发 32(0.9583 / 0.9606)与 64(0.9583 / 0.9621),因此冒烟数值仅供参考,不是同一并发点的对比。
  • 上游源码对比: v0.5.12 @ 127b9e3(2026-05-16)→ v0.5.19 @ 0bcd822(2026-09-04),共 4893 个提交。来源确认:两个镜像的 ai.sglang.build.commit / org.opencontainers.image.revision 标签均等于上述附注标签对应的提交;v0.5.19-cu130v0.5.19 共享同一摘要(2026-09-04T22:50Z 推送)。
  • 耦合依赖: CUDA 基础镜像 13.0.1 → 13.0.3(同样要求 driver>=535),cuDNN 9.13 → 9.14,FlashInfer 0.6.11.post1 → 0.6.18。
  • 参数审计: 脚本传入的全部选项在 v0.5.19 中仍有定义 —— --model-path--trust-remote-code--context-length--mem-fraction-static--max-running-requests--chunked-prefill-size--max-prefill-tokens--disable-radix-cache--stream-interval--decode-log-interval--attention-backendserver_args.py--tensor-parallel-size / --data-parallel-size 仍是 tp_size / dp_size 的别名。--cuda-graph-max-bs 现为已弃用别名,写入 cuda_graph_max_bs_decodeDeprecatedAliasStoreAction,仅告警),并由 cuda_graph_hook.py 消费,因此解码图上限 256 得以保留。
  • 已核查的默认值 / 语义变化: DeepSeek-V3 覆写(model_overrides/deepseek_v2.py)仅在 SM100 且未设置后端时自动选择 trtllm_mla,H200(SM90)上显式指定的 flashinfer 不受影响;MoE runner 自动选择 flashinfer_trtllm 仅限 SM100;memory_hook.py 仅在未设置时推导 mem_fraction_static,因此 0.82 保留。Hopper FP8 内核层面的变化:SM90 的 CUTLASS FP8 blockwise GEMM 已删除(#30438,DeepGEMM 仍为默认),SM90 FP8 解码路由修复(#37018);统一 radix 树默认开启(#35081)在 --disable-radix-cache 下无影响;编译内核缓存迁移至 SGLANG_CACHE_DIR#32434)。MTP 同类配方 dsr1-fp8-h200-sglang-mtp 已在 H200 上使用同一镜像及相同的非投机参数运行(#2955,2026-09-10 完整 sweep 通过)。
  • 决定: 仅升级镜像;无需在范围内调整参数。
  • 下一步: 已追加 perf-changelog.yaml 条目(已对基线 e444727b 本地校验);重新检查容量并在草稿 PR 上以 full-sweep-enabled 启动最终完整 sweep(全部五个并发点及并发 32 与 64 的默认 gsm8k 评测)。

Append the perf-changelog.yaml entry for the dsr1-fp8-h200-sglang
SGLang image update to v0.5.19-cu130 (PR #3010).

为 dsr1-fp8-h200-sglang 的 SGLang 镜像更新至 v0.5.19-cu130(PR #3010)
追加 perf-changelog.yaml 条目。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Sync with main so the perf-changelog.yaml tail entry for PR #3010 applies
on top of the entries merged since the candidate base; the image bump and
changelog entry are unchanged.

与 main 同步,使 PR #3010 的 perf-changelog.yaml 尾部条目位于此后合并的条目
之上;镜像更新与 changelog 条目本身不变。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@Klaud-Cold

Klaud-Cold commented Sep 11, 2026

Copy link
Copy Markdown
Collaborator Author

Final full sweep

  • Image / SHA: lmsysorg/sglang:v0.5.19-cu130 at PR head f207db02fbb7d8016d1aca9c67907b5e674dc915 (image bump + perf-changelog.yaml entry for dsr1-fp8-h200-sglang, merged with main e5f3e41c; validated locally with utils/validate_perf_changelog.py against origin/main)
  • Trigger: full-sweep-enabled applied to the draft PR at 2026-09-11T12:27Z after a passing capacity check. That label event started no sweep because the PR had become CONFLICTING: main appended perf-changelog.yaml entries after the candidate base, so GitHub skipped the pull_request workflows. Fixed by merging main into the branch, taking main's changelog bytes verbatim and re-appending this PR's entry at the tail (f207db02, pushed 12:40Z after another passing capacity check). Run 34600160889 (run-sweep.yml, synchronize event with full-sweep-enabled present) is the final sweep.
  • Scope / outcome: full matrix, TP8 EP1 8k1k at concurrency 4 / 8 / 16 / 32 / 64 plus the default gsm8k evals at concurrency 32 and 64 — completed, success (12:40Z → 13:27Z benchmarks/evals, run green at ~14:05Z after collectors). Canary (concurrency 4) on h200-cw_00, concurrency 8 on h200-cw_01, concurrency 16 / 32 / 64 on h200-dgxc-slurm_00 / _01 / _03, evals on h200-dgxc-slurm_04 (32) and _02 (64); check-changelog, setup, collect-results, collect-evals, compare-results, upload-changelog-metadata and calc-success-rate are green. results_bmk/agg_bmk.json holds all five points with power_valid=1 and one recipe fingerprint (1a0513ad…); the klaud-sweep-manifest records full-sweep: true on this head.
  • Benchmark results vs published 2026-05-17 baseline (same TP8 EP1 8k1k points, result IDs 409933 / 409937 / 409940 / 409935 / 409941):
Conc Total tok/s/GPU Output tok/s/GPU Median TTFT (s) Median TPOT (ms) Median E2E (s)
4 387.6 → 394.3 (+1.7%) 43.1 → 43.9 (+1.7%) 0.312 → 0.302 (-3.0%) 10.98 → 10.87 (-1.0%) 10.32 → 10.20 (-1.2%)
8 607.9 → 621.1 (+2.2%) 68.3 → 69.8 (+2.2%) 0.314 → 0.307 (-2.4%) 13.93 → 13.72 (-1.5%) 13.23 → 13.02 (-1.6%)
16 847.3 → 860.7 (+1.6%) 93.9 → 95.4 (+1.6%) 0.330 → 0.321 (-2.6%) 20.22 → 19.89 (-1.6%) 19.01 → 18.70 (-1.6%)
32 1082.1 → 1129.3 (+4.4%) 121.0 → 126.3 (+4.4%) 0.348 → 0.336 (-3.6%) 32.01 → 30.57 (-4.5%) 29.96 → 28.58 (-4.6%)
64 1360.9 → 1378.0 (+1.3%) 151.0 → 152.9 (+1.3%) 0.690 → 0.713 (+3.4%) 50.70 → 49.85 (-1.7%) 47.70 → 47.04 (-1.4%)
  • Regression to note: median TTFT at concurrency 64 rose from 0.690 s to 0.713 s (+3.4%); its mean TTFT fell 2.25 s → 2.17 s and p99 TTFT is unchanged (19.28 s → 19.41 s), and throughput, TPOT and E2E improved at that point. Every other point improved on every metric. Throughput gains are modest (+1.3% to +4.4%), as expected for an image-only refresh on the unchanged flashinfer / DeepGEMM Hopper path.
  • Eval results: gsm8k em_strict / em_flexible: concurrency 32 → 0.9606 / 0.9613, concurrency 64 → 0.9538 / 0.9538 (n = 1319 each, infrastructure_success=true). Published baseline: 0.9583 / 0.9606 (32) and 0.9583 / 0.9621 (64); the concurrency-64 score is 0.45 pt lower, within one standard error (≈0.55 pt).
  • Status: targeted smoke and exact-head final validation both passed; global approval remains with the automatic reviews and CODEOWNER sign-off. Next step: run the Klaud completion check, which verifies matrix/result coverage and marks the PR ready. Smoke evidence is in the Initial attempt.

最终完整 sweep

  • 镜像 / SHA: lmsysorg/sglang:v0.5.19-cu130,PR head f207db02fbb7d8016d1aca9c67907b5e674dc915(镜像升级 + dsr1-fp8-h200-sglangperf-changelog.yaml 条目,并与 main e5f3e41c 合并;已用 utils/validate_perf_changelog.pyorigin/main 本地校验)
  • 触发: 容量检查通过后,于 2026-09-11T12:27Z 在草稿 PR 上添加 full-sweep-enabled。该标签事件未启动 sweep,因为 PR 已变为 CONFLICTINGmain 在候选基线之后追加了 perf-changelog.yaml 条目,GitHub 因此跳过了 pull_request 工作流。修复方式:将 main 合并进分支,逐字采用 main 的 changelog 字节并在尾部重新追加本 PR 的条目(f207db02,于 12:40Z 在再次通过容量检查后推送)。运行 34600160889run-sweep.yml,携带 full-sweep-enabled 的 synchronize 事件)为最终 sweep。
  • 范围 / 结果: 完整矩阵,TP8 EP1 8k1k 并发 4 / 8 / 16 / 32 / 64,以及并发 32 与 64 的默认 gsm8k 评测 —— 已完成,成功(基准/评测 12:40Z → 13:27Z,收集器完成后约 14:05Z 运行变绿)。金丝雀(并发 4)在 h200-cw_00,并发 8 在 h200-cw_01,并发 16 / 32 / 64 在 h200-dgxc-slurm_00 / _01 / _03,评测在 h200-dgxc-slurm_04(32)与 _02(64);check-changelog、setup、collect-results、collect-evals、compare-results、upload-changelog-metadata 与 calc-success-rate 均通过。results_bmk/agg_bmk.json 包含全部五个点,power_valid=1,配方指纹一致(1a0513ad…);klaud-sweep-manifest 在该 head 上记录 full-sweep: true
  • 基准结果对比 2026-05-17 已发布基线(相同 TP8 EP1 8k1k 点,结果 ID 409933 / 409937 / 409940 / 409935 / 409941):
并发 总吞吐 tok/s/GPU 输出吞吐 tok/s/GPU TTFT 中位数 (s) TPOT 中位数 (ms) 端到端中位数 (s)
4 387.6 → 394.3 (+1.7%) 43.1 → 43.9 (+1.7%) 0.312 → 0.302 (-3.0%) 10.98 → 10.87 (-1.0%) 10.32 → 10.20 (-1.2%)
8 607.9 → 621.1 (+2.2%) 68.3 → 69.8 (+2.2%) 0.314 → 0.307 (-2.4%) 13.93 → 13.72 (-1.5%) 13.23 → 13.02 (-1.6%)
16 847.3 → 860.7 (+1.6%) 93.9 → 95.4 (+1.6%) 0.330 → 0.321 (-2.6%) 20.22 → 19.89 (-1.6%) 19.01 → 18.70 (-1.6%)
32 1082.1 → 1129.3 (+4.4%) 121.0 → 126.3 (+4.4%) 0.348 → 0.336 (-3.6%) 32.01 → 30.57 (-4.5%) 29.96 → 28.58 (-4.6%)
64 1360.9 → 1378.0 (+1.3%) 151.0 → 152.9 (+1.3%) 0.690 → 0.713 (+3.4%) 50.70 → 49.85 (-1.7%) 47.70 → 47.04 (-1.4%)
  • 需注意的回退: 并发 64 的 TTFT 中位数从 0.690 s 升至 0.713 s(+3.4%);该点的 TTFT 均值由 2.25 s 降至 2.17 s,p99 TTFT 基本不变(19.28 s → 19.41 s),吞吐、TPOT 与端到端均有改善。其余各点所有指标均改善。吞吐增幅较小(+1.3% 至 +4.4%),符合在未改动的 flashinfer / DeepGEMM Hopper 路径上仅刷新镜像的预期。
  • 评测结果: gsm8k em_strict / em_flexible:并发 32 → 0.9606 / 0.9613,并发 64 → 0.9538 / 0.9538(各 n = 1319,infrastructure_success=true)。已发布基线:0.9583 / 0.9606(32)与 0.9583 / 0.9621(64);并发 64 的分数低 0.45 个百分点,在一个标准误(约 0.55 个百分点)以内。
  • 状态: 定向冒烟与精确 head 的最终验证均已通过;全局批准仍取决于自动审阅与 CODEOWNER 签署。下一步:运行 Klaud 完成检查,由其校验矩阵/结果覆盖并将 PR 标记为可审阅。冒烟证据见 初始尝试

@github-actions

Copy link
Copy Markdown
Contributor

@Klaud-Cold
Klaud-Cold marked this pull request as ready for review September 11, 2026 13:29
@Klaud-Cold
Klaud-Cold requested a review from a team September 11, 2026 13:29
@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Klaud Cold: validated. All owned runs are terminal. Repairs: 0. Runs: 34592860782, 34600160889.

The full sweep is verified; this PR remains ready for review.


Klaud Cold:validated。所有自有运行均已结束。修复次数:0。运行:34592860782, 34600160889

完整 sweep 已通过验证;PR 保持就绪,等待审查。

@claude

claude Bot commented Sep 11, 2026

Copy link
Copy Markdown
Contributor

Claude finished @Klaud-Cold's task in 1m 18s —— View job


Review of PR #3010

  • Gather context (diff vs origin/main, changed files)
  • Validate master config change (public image, key prefix)
  • Validate perf-changelog.yaml entry (appended at end, correct config/PR link)
  • Post review summary

LGTM - no blocking issues found.

The diff against origin/main touches exactly two files. The image moves to lmsysorg/sglang:v0.5.19-cu130, a public Docker Hub tag already used by dsr1-fp8-h200-sglang-mtp and several qwen3.5 recipes, and the digest pinned elsewhere in the config for this tag matches the one cited in the PR body. The perf-changelog.yaml entry is appended at the very end of the file, names the right config key, and links this PR. The config key keeps the dsr1 prefix and the rest of the block is unchanged. Full sweep run 34600160889 is reported green on the exact PR head.

Note: the local utils/validate_perf_changelog.py run was not permitted in this session, so the entry was checked by inspection only. The check-changelog job in the sweep run covers the same validation.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good, straightforward image-version bump with matching changelog entry.

What was reviewed: the image field bump in configs/nvidia-master.yaml for dsr1-fp8-h200-sglang (v0.5.12-cu130 to v0.5.19-cu130), confirmed it now matches the sibling dsr1-fp8-h200-sglang-mtp entry which already used v0.5.19-cu130. Verified the perf-changelog.yaml addition is appended only at the physical end of the file with existing content untouched, and the new entry's config-key, description, and pr-link are well-formed and match the PR.

Extended reasoning...

Overview

This PR touches exactly two files: configs/nvidia-master.yaml (bumping the image field of the dsr1-fp8-h200-sglang recipe from lmsysorg/sglang:v0.5.12-cu130 to lmsysorg/sglang:v0.5.19-cu130) and perf-changelog.yaml (appending one new changelog entry documenting that bump). No other fields in the recipe (model, model-prefix, runner, precision, framework, multinode, scenarios/search-space) were changed, and the diff confirms this matches exactly what the PR description claims.

Security risks

None. This is a version string change to a container image reference in a YAML config, plus an append-only changelog entry. No code execution paths, credentials, auth, or permission logic are touched.

Level of scrutiny

Low. This is a mechanical, single-line image bump of the kind the repo's own conventions (per AGENTS.md/CONTRIBUTING.md as summarized in the review scope) treat as routine, provided the changelog is updated correctly and sibling recipes stay consistent. Both conditions are verified directly from the file contents: the sibling dsr1-fp8-h200-sglang-mtp recipe already uses v0.5.19-cu130 (so this brings the pair in line), and the changelog entry is appended cleanly at end-of-file without altering any prior bytes.

Other factors

No CODEOWNERS-restricted or security-sensitive paths are involved. The changelog's append-only invariant (new entries only at the tail, no reformatting of existing entries) is respected, which is the main structural risk for this file type. No test suite runs are needed beyond what a YAML/config-key validator would already catch, and the change is small and self-contained enough that a human review would only be duplicating this same mechanical check.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

1 participant