Skip to content

[Klaud Cold] Update qwen3.5-fp8-h100-sglang-mtp SGLang image to v0.5.19-cu130 / 将 qwen3.5-fp8-h100-sglang-mtp 的 SGLang 镜像升级至 v0.5.19-cu130 - #2952

Merged
adibarra merged 3 commits into
mainfrom
klaud/auto-3f335c30ec28319b-3d9c146c5e9ad2d6
Sep 10, 2026
Merged

[Klaud Cold] Update qwen3.5-fp8-h100-sglang-mtp SGLang image to v0.5.19-cu130 / 将 qwen3.5-fp8-h100-sglang-mtp 的 SGLang 镜像升级至 v0.5.19-cu130#2952
adibarra merged 3 commits into
mainfrom
klaud/auto-3f335c30ec28319b-3d9c146c5e9ad2d6

Conversation

@Klaud-Cold

@Klaud-Cold Klaud-Cold commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator

Update the qwen3.5-fp8-h100-sglang-mtp master image from lmsysorg/sglang:v0.5.14-cu130 to the current SGLang release lmsysorg/sglang:v0.5.19-cu130 (Docker Hub digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, tag commit sgl-project/sglang@0bcd822). The recipe script, model, TP8/EP8 topology, EAGLE/MTP settings and the 8k1k concurrency range are unchanged.

Baseline

  • Published date: 2026-07-05 (benchmarks?model=Qwen-3.5-397B-A17B&date=2026-07-05&exact=true, workflow-info?date=2026-07-05)
  • Old image: lmsysorg/sglang:v0.5.14-cu130 (digest sha256:5027e95bf6ec536856b1b52a91d1f35ff5c564ab83e8a94758a169ff09bb8df3, tag commit sgl-project/sglang@49e384c)
  • Workload / topology: single-node H100, Qwen/Qwen3.5-397B-A17B-FP8, SGLang FP8, TP8 EP8, EAGLE MTP (3 steps, top-k 1, 4 draft tokens), fixed-seq-len 8k1k (ISL 8192 / OSL 1024), concurrency 4 / 8 / 16 / 32, random dataset with chat template
  • Producer: run 28719988201 (head 2eaf828b84ebe5cac3c68fabd8a18be99fc6d924, changelog PR #2060); benchmark result IDs 432614, 432616, 432593, 432601
Conc Total tok/s/GPU Output tok/s/GPU Median TTFT (s) Median TPOT (ms) Median E2E (s)
4 686.1 76.3 0.410 5.93 5.79
8 662.5 74.5 0.623 12.45 12.51
16 1297.7 143.8 0.445 13.25 12.42
32 1618.7 181.0 0.668 21.23 20.20
  • Published eval: gsm8k at concurrency 32, em_strict 0.9697 / em_flexible 0.9598 (evaluation ID 8335, same producer run). The public evaluations feed labels this row disagg: true with 64 prefill/decode GPUs, which does not match the single-node TP8 recipe; the benchmark rows above carry the correct disagg: false identity.

qwen3.5-fp8-h100-sglang-mtp 的主配置镜像从 lmsysorg/sglang:v0.5.14-cu130 升级到当前 SGLang 发布版 lmsysorg/sglang:v0.5.19-cu130(Docker Hub 摘要 sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,标签提交 sgl-project/sglang@0bcd822)。配方脚本、模型、TP8/EP8 拓扑、EAGLE/MTP 设置以及 8k1k 并发范围均保持不变。

基线

  • 发布日期: 2026-07-05(benchmarks?model=Qwen-3.5-397B-A17B&date=2026-07-05&exact=trueworkflow-info?date=2026-07-05
  • 旧镜像: lmsysorg/sglang:v0.5.14-cu130(摘要 sha256:5027e95bf6ec536856b1b52a91d1f35ff5c564ab83e8a94758a169ff09bb8df3,标签提交 sgl-project/sglang@49e384c
  • 工作负载 / 拓扑: 单节点 H100,Qwen/Qwen3.5-397B-A17B-FP8,SGLang FP8,TP8 EP8,EAGLE MTP(3 步、top-k 1、4 个草稿 token),固定序列长度 8k1k(ISL 8192 / OSL 1024),并发 4 / 8 / 16 / 32,随机数据集并使用聊天模板
  • 数据来源: 运行 28719988201(head 2eaf828b84ebe5cac3c68fabd8a18be99fc6d924,changelog PR #2060);基准结果 ID 432614、432616、432593、432601
并发 总吞吐 tok/s/GPU 输出吞吐 tok/s/GPU TTFT 中位数 (s) TPOT 中位数 (ms) 端到端中位数 (s)
4 686.1 76.3 0.410 5.93 5.79
8 662.5 74.5 0.623 12.45 12.51
16 1297.7 143.8 0.445 13.25 12.42
32 1618.7 181.0 0.668 21.23 20.20
  • 已发布评测: 并发 32 的 gsm8k,em_strict 0.9697 / em_flexible 0.9598(评测 ID 8335,同一数据来源运行)。公共评测接口将该行标记为 disagg: true 且 64 个 prefill/decode GPU,与单节点 TP8 配方不符;上表基准行的 disagg: false 身份是正确的。

Note

Low Risk
Config-only SGLang container image version bump for a single benchmark key; no application logic or security-sensitive code changes.

Overview
Bumps the qwen3.5-fp8-h100-sglang-mtp benchmark recipe in nvidia-master.yaml from lmsysorg/sglang:v0.5.14-cu130 to lmsysorg/sglang:v0.5.19-cu130.

Adds a matching perf-changelog.yaml entry noting the digest and that the recipe script, TP8/EP8 topology, and EAGLE/MTP settings are unchanged.

Reviewed by Cursor Bugbot for commit 9cda017. Bugbot is set up for automated code reviews on this repo. Configure here.

…0.5.19-cu130

Bump the qwen3.5-fp8-h100-sglang-mtp master image from
lmsysorg/sglang:v0.5.14-cu130 to lmsysorg/sglang:v0.5.19-cu130
(Docker Hub digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,
tag commit sgl-project/sglang@0bcd822).
Recipe script, model, TP8/EP8 topology, EAGLE/MTP settings and the 8k1k
concurrency range are unchanged.

将 qwen3.5-fp8-h100-sglang-mtp 的主配置镜像从 lmsysorg/sglang:v0.5.14-cu130
升级到 lmsysorg/sglang:v0.5.19-cu130(Docker Hub 摘要
sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,
标签提交 sgl-project/sglang@0bcd822)。
配方脚本、模型、TP8/EP8 拓扑、EAGLE/MTP 设置以及 8k1k 并发范围保持不变。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

@Klaud-Cold

Klaud-Cold commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator Author

Initial attempt

  • Image / SHA: lmsysorg/sglang:v0.5.19-cu130 (digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9) at PR head 66b3f0f4a8a0604a79479903d9e652c1700ed0b1
  • Change: master image only (configs/nvidia-master.yaml, key qwen3.5-fp8-h100-sglang-mtp); benchmarks/single_node/fixed_seq_len/qwen3.5_fp8_h100_mtp.sh unchanged, no runtime patching
  • Targeted run: e2e run 34422693420 (e2e-tests.yml from main, ref=66b3f0f4, test-config --trim-conc: TP8 EP8 MTP at concurrency 4 plus the smoke gsm8k eval; dispatched 2026-09-10T00:46Z) — completed, mixed: benchmark job passed on h100-cw_01 (attempt 1), eval-only job failed on h100-cw_00
  • Benchmark result (concurrency 4, 8k1k, vs published 2026-07-05 baseline):
Metric Baseline v0.5.14 v0.5.19 Delta
Total tok/s/GPU 686.1 816.1 +18.9%
Output tok/s/GPU 76.3 90.8 +18.9%
Median TTFT (s) 0.410 0.257 -37.3%
Median TPOT (ms) 5.93 5.21 -12.1%
Median E2E (s) 5.79 4.97 -14.2%
  • Eval result: N/A — the eval-only server never became ready. First error: on rank TP2, during prefill CUDA-graph capture, Triton's CompiledKernel.__init__ raised FileNotFoundError for /root/.cache/sglang/triton/.../fused_moe_kernel.cubin (the compile-cache group metadata existed but the cubin was not yet present). This is a transient compile-cache race between the 8 TP ranks sharing the per-node SGLANG_CACHE_DIR (kernel caches moved there in v0.5.18, #32434); the identical launch command on h100-cw_01 captured all prefill graphs and served the benchmark, and no flag or model incompatibility was involved.
  • Next step: Repair 1/5 — rerun only the failed eval job (no code change).
  • Upstream source comparison: v0.5.14 @ 49e384c (2026-06-25) → v0.5.19 @ 0bcd822 (2026-09-03). Provenance: release-docker.yml tags v{version}-cu130 from a git clone --branch v{version} build (docker/Dockerfile); Docker Hub v0.5.19 and v0.5.19-cu130 share one digest, pushed 2026-09-04.
  • Coupled dependencies: torch 2.11.0 → 2.13.0, flashinfer 0.6.12 → 0.6.18, sgl-kernel 0.4.4 → 0.4.6.post1, transformers 5.8.1 → 5.12.1, CUDA base 13.0.1 → 13.0.3 (same CUDA 13.0 driver requirement as the old cu130 image).
  • Flag audit: all 23 launch flags used by the script are still defined (server_args.py plus the new arg_groups/ hooks) with unchanged choices for --quantization fp8, --attention-backend flashinfer, --kv-cache-dtype fp8_e4m3, --mamba-ssm-dtype bfloat16 and --speculative-algorithm EAGLE. --enable-flashinfer-allreduce-fusion is now deprecated and maps to --flashinfer-allreduce-fusion-backend auto with a warning (serving_hook.py); SGLANG_ENABLE_SPEC_V2 is a removal warning in both tags (spec V2 always on). The bf16 SSM-state requirements in attention_hook.py are SM100-only and do not apply to H100 (SM90), where linear-attention decode stays on Triton.
  • Behavior deltas noted: unified radix tree is the default for every configuration in v0.5.19 (#35081; moot under --disable-radix-cache); CUTLASS FP8 blockwise was deleted for SM90 in v0.5.16 (#30438), so H100 FP8 GEMMs use the remaining SM90 paths — the agentic qwen3.5-fp8-h100-sglang-agentic-mtp family already published on v0.5.16-cu130 (2026-08-07) with the same model and MTP settings. qwen3.5-fp8-b200-sglang runs v0.5.19-cu130 on main (PR WIP - SGL B200 FP8 8k1k  #2866, full sweep 34175132645 green) as SM100 evidence.
  • Decision: image-only bump; no in-scope flag changes required.

初始尝试

  • 镜像 / SHA: lmsysorg/sglang:v0.5.19-cu130(摘要 sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9),PR head 66b3f0f4a8a0604a79479903d9e652c1700ed0b1
  • 改动: 仅主配置镜像(configs/nvidia-master.yaml 中的 qwen3.5-fp8-h100-sglang-mtp);benchmarks/single_node/fixed_seq_len/qwen3.5_fp8_h100_mtp.sh 未改动,无任何运行时补丁
  • 定向运行: e2e 运行 34422693420(从 main 触发 e2e-tests.ymlref=66b3f0f4,test-config --trim-conc:TP8 EP8 MTP、并发 4,附带 gsm8k 冒烟评测;2026-09-10T00:46Z 触发)—— 已完成,结果不一:基准任务在 h100-cw_01 通过(第 1 次尝试),仅评测任务在 h100-cw_00 失败
  • 基准结果(并发 4,8k1k,对比 2026-07-05 已发布基线):
指标 基线 v0.5.14 v0.5.19 变化
总吞吐 tok/s/GPU 686.1 816.1 +18.9%
输出吞吐 tok/s/GPU 76.3 90.8 +18.9%
TTFT 中位数 (s) 0.410 0.257 -37.3%
TPOT 中位数 (ms) 5.93 5.21 -12.1%
端到端中位数 (s) 5.79 4.97 -14.2%
  • 评测结果: N/A —— 仅评测的服务端未能就绪。首个错误:TP2 rank 在 prefill CUDA graph 捕获期间,Triton 的 CompiledKernel.__init__/root/.cache/sglang/triton/.../fused_moe_kernel.cubin 抛出 FileNotFoundError(编译缓存的分组元数据已存在,但 cubin 尚未落盘)。这是 8 个 TP rank 共享节点级 SGLANG_CACHE_DIR(v0.5.18 起内核缓存迁移至此,#32434)时的瞬时编译缓存竞争;相同的启动命令在 h100-cw_01 上完整捕获了全部 prefill graph 并完成了基准测试,不涉及任何参数或模型不兼容。
  • 下一步: 修复 1/5 —— 仅重跑失败的评测任务(不改代码)。
  • 上游源码对比: v0.5.14 @ 49e384c(2026-06-25)→ v0.5.19 @ 0bcd822(2026-09-03)。来源确认:release-docker.ymlgit clone --branch v{version} 构建并打 v{version}-cu130 标签(docker/Dockerfile);Docker Hub 上 v0.5.19v0.5.19-cu130 共享同一摘要,推送于 2026-09-04。
  • 耦合依赖: torch 2.11.0 → 2.13.0,flashinfer 0.6.12 → 0.6.18,sgl-kernel 0.4.4 → 0.4.6.post1,transformers 5.8.1 → 5.12.1,CUDA 基础镜像 13.0.1 → 13.0.3(与旧 cu130 镜像相同的 CUDA 13.0 驱动要求)。
  • 参数审计: 脚本使用的全部 23 个启动参数仍然存在(server_args.py 及新的 arg_groups/ 钩子),--quantization fp8--attention-backend flashinfer--kv-cache-dtype fp8_e4m3--mamba-ssm-dtype bfloat16--speculative-algorithm EAGLE 的可选值未变。--enable-flashinfer-allreduce-fusion 已弃用,会告警并映射为 --flashinfer-allreduce-fusion-backend autoserving_hook.py);SGLANG_ENABLE_SPEC_V2 在两个标签中都只是移除告警(spec V2 始终开启)。attention_hook.py 中的 bf16 SSM 状态要求仅针对 SM100,不适用于 H100(SM90),其线性注意力解码仍使用 Triton。
  • 已注意到的行为变化: v0.5.19 对所有配置默认启用统一 radix tree(#35081;在 --disable-radix-cache 下无影响);v0.5.16 删除了 SM90 的 CUTLASS FP8 blockwise(#30438),H100 的 FP8 GEMM 走剩余的 SM90 路径 —— agentic 家族 qwen3.5-fp8-h100-sglang-agentic-mtp 已在 v0.5.16-cu130 上以相同模型和 MTP 设置发布(2026-08-07)。qwen3.5-fp8-b200-sglang 在 main 上运行 v0.5.19-cu130(PR WIP - SGL B200 FP8 8k1k  #2866,完整扫描 34175132645 全绿),可作为 SM100 佐证。
  • 决定: 仅升级镜像;无需在范围内修改参数。

@Klaud-Cold

Klaud-Cold commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator Author

Repair 1/5

  • Image / SHA: unchanged — lmsysorg/sglang:v0.5.19-cu130 at PR head 66b3f0f4a8a0604a79479903d9e652c1700ed0b1
  • Change: none in the repository. Rerun of only the failed eval-only job of the initial run (run 34422693420, attempt 2, requested 2026-09-10T01:14Z); the passing benchmark job is not repeated.
  • Reason: the Initial attempt eval server died in a transient Triton compile-cache race (FileNotFoundError for fused_moe_kernel.cubin on TP2) while the identical launch on another node served the benchmark. Nothing in the recipe or image needs changing; the rerun checks reproducibility.
  • Result: passed — eval-only job succeeded on h100-cw_00 (server ready, gsm8k on 1319 samples). Scores vs the published 2026-07-05 eval (baseline measured at concurrency 32, this smoke at concurrency 4):
gsm8k metric Baseline v0.5.14 (conc 32) v0.5.19 (conc 4) Threshold
exact_match strict 0.9697 0.9697 ≥ 0.94
exact_match flexible 0.9598 0.9606 ≥ 0.94

The Triton compile-cache race did not recur; together with the attempt-1 benchmark point this completes the targeted smoke (startup/compatibility evidence, not a full curve).

  • Next step: append the perf-changelog.yaml entry, recheck capacity and start the final full sweep (full-sweep-enabled, PR kept draft).

修复 1/5

  • 镜像 / SHA: 未变 —— lmsysorg/sglang:v0.5.19-cu130,PR head 66b3f0f4a8a0604a79479903d9e652c1700ed0b1
  • 改动: 仓库无改动。仅重跑初始运行中失败的评测任务(运行 34422693420,第 2 次尝试,2026-09-10T01:14Z 发起);已通过的基准任务不重复执行。
  • 原因: 初始尝试的评测服务端因瞬时 Triton 编译缓存竞争(TP2 上 fused_moe_kernel.cubinFileNotFoundError)退出,而另一节点上相同的启动命令完成了基准测试。配方与镜像均无需修改;重跑用于确认可复现性。
  • 结果: 通过 —— 仅评测任务在 h100-cw_00 成功(服务端就绪,gsm8k 1319 条样本)。与 2026-07-05 已发布评测对比(基线在并发 32 测得,本次冒烟在并发 4):
gsm8k 指标 基线 v0.5.14(并发 32) v0.5.19(并发 4) 阈值
exact_match strict 0.9697 0.9697 ≥ 0.94
exact_match flexible 0.9598 0.9606 ≥ 0.94

Triton 编译缓存竞争未再出现;结合第 1 次尝试的基准点,定向冒烟测试完成(属于启动/兼容性证据,不是完整曲线)。

  • 下一步: 追加 perf-changelog.yaml 条目,重新检查容量,并启动最终完整扫描(full-sweep-enabled,PR 保持草稿)。

…v0.5.19-cu130

Append the perf-changelog.yaml entry for the qwen3.5-fp8-h100-sglang-mtp
SGLang image update to lmsysorg/sglang:v0.5.19-cu130 (PR #2952).

为 qwen3.5-fp8-h100-sglang-mtp 的 SGLang 镜像升级至
lmsysorg/sglang:v0.5.19-cu130(PR #2952)追加 perf-changelog.yaml 条目。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@Klaud-Cold

Klaud-Cold commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator Author

Final full sweep

  • Image / SHA: lmsysorg/sglang:v0.5.19-cu130 at PR head 5cb93a2d98a8194fdb07e63c9dce17f5bf665131 (adds the perf-changelog.yaml entry on top of the image bump; validate_perf_changelog.py against origin/main passed locally)
  • Run: run-sweep 34429032159 (full-sweep-enabled applied 2026-09-10T02:20Z with the PR kept draft; the concurrent push-triggered run 34429017350 was superseded and cancelled by the workflow's concurrency group)
  • Scope: the complete qwen3.5-fp8-h100-sglang-mtp matrix — 8k1k TP8 EP8 MTP at concurrency 4 / 8 / 16 / 32 plus the default gsm8k eval at concurrency 32 (canary selection may run the lowest point first)
  • Result: passed — run-sweep completed success on the exact PR head (attempt 1): 4/4 benchmark points, the gsm8k eval, collect/compare/changelog-metadata jobs and the klaud-sweep-manifest (full-sweep: true) all green; every point reports power_valid: 1. Points ran on h100-cw_00 (c4) and h100-dgxc-slurm (c8/c16/c32); the eval on h100-cw_01.
  • Throughput / latency vs the published 2026-07-05 baseline (same TP8 EP8 MTP 8k1k recipe):
Conc Total tok/s/GPU (base → new) Output tok/s/GPU (base → new) Median TTFT s (base → new) Median TPOT ms (base → new) Median E2E s (base → new)
4 686.1 → 815.8 (+18.9%) 76.3 → 90.7 (+18.9%) 0.410 → 0.258 (-37.1%) 5.93 → 5.19 (-12.5%) 5.79 → 4.97 (-14.1%)
8 662.5 → 1177.7 (+77.8%) 74.5 → 132.4 (+77.8%) 0.623 → 0.288 (-53.7%) 12.45 → 7.11 (-42.9%) 12.51 → 6.89 (-44.9%)
16 1297.7 → 1630.6 (+25.6%) 143.8 → 180.6 (+25.6%) 0.445 → 0.302 (-32.1%) 13.25 → 10.38 (-21.7%) 12.42 → 9.77 (-21.4%)
32 1618.7 → 2079.1 (+28.4%) 181.0 → 232.4 (+28.4%) 0.668 → 0.427 (-36.2%) 21.23 → 16.65 (-21.6%) 20.20 → 15.65 (-22.5%)
  • Eval (gsm8k, concurrency 32, 1319 samples):
Metric Baseline v0.5.14 v0.5.19 Delta Threshold
exact_match strict 0.9697 0.9651 -0.0046 (within ±0.0047 SE) ≥ 0.94
exact_match flexible 0.9598 0.9591 -0.0007 ≥ 0.94
  • Notes: the conc 8 baseline point (662.5 tok/s/GPU, below the conc 4 value) looks like an outlier in the published curve, so its +77.8% delta overstates the typical gain; the other three points sit at +19% to +28%. No regressions observed; the small gsm8k strict decrease is inside one standard error.
  • Next step: finish verifies matrix/result coverage and marks the PR ready for the automatic reviews. Repairs used: 1/5 (an eval rerun with no code change).

最终完整扫描

  • 镜像 / SHA: lmsysorg/sglang:v0.5.19-cu130,PR head 5cb93a2d98a8194fdb07e63c9dce17f5bf665131(在镜像升级之上追加 perf-changelog.yaml 条目;本地对 origin/main 运行 validate_perf_changelog.py 通过)
  • 运行: run-sweep 34429032159(2026-09-10T02:20Z 添加 full-sweep-enabled,PR 保持草稿;由推送触发的并发运行 34429017350 已被工作流并发组取代并取消)
  • 范围: 完整的 qwen3.5-fp8-h100-sglang-mtp 矩阵 —— 8k1k TP8 EP8 MTP、并发 4 / 8 / 16 / 32,加上并发 32 的默认 gsm8k 评测(canary 选择可能先运行最低并发点)
  • 结果: 通过 —— run-sweep 在 PR 精确 head 上以 success 完成(第 1 次尝试):4/4 个基准点、gsm8k 评测、collect/compare/changelog-metadata 任务以及 klaud-sweep-manifestfull-sweep: true)全部通过;每个点均报告 power_valid: 1。基准点运行于 h100-cw_00(c4)和 h100-dgxc-slurm(c8/c16/c32);评测运行于 h100-cw_01
  • 吞吐 / 延迟对比 2026-07-05 已发布基线(同一 TP8 EP8 MTP 8k1k 配方):
并发 总吞吐 tok/s/GPU(基线 → 新) 输出吞吐 tok/s/GPU(基线 → 新) TTFT 中位数 s(基线 → 新) TPOT 中位数 ms(基线 → 新) 端到端中位数 s(基线 → 新)
4 686.1 → 815.8 (+18.9%) 76.3 → 90.7 (+18.9%) 0.410 → 0.258 (-37.1%) 5.93 → 5.19 (-12.5%) 5.79 → 4.97 (-14.1%)
8 662.5 → 1177.7 (+77.8%) 74.5 → 132.4 (+77.8%) 0.623 → 0.288 (-53.7%) 12.45 → 7.11 (-42.9%) 12.51 → 6.89 (-44.9%)
16 1297.7 → 1630.6 (+25.6%) 143.8 → 180.6 (+25.6%) 0.445 → 0.302 (-32.1%) 13.25 → 10.38 (-21.7%) 12.42 → 9.77 (-21.4%)
32 1618.7 → 2079.1 (+28.4%) 181.0 → 232.4 (+28.4%) 0.668 → 0.427 (-36.2%) 21.23 → 16.65 (-21.6%) 20.20 → 15.65 (-22.5%)
  • 评测(gsm8k,并发 32,1319 条样本):
指标 基线 v0.5.14 v0.5.19 变化 阈值
exact_match strict 0.9697 0.9651 -0.0046(在 ±0.0047 标准误内) ≥ 0.94
exact_match flexible 0.9598 0.9591 -0.0007 ≥ 0.94
  • 说明: 基线中并发 8 的点(662.5 tok/s/GPU,低于并发 4)疑似已发布曲线中的异常值,因此其 +77.8% 的变化高估了典型收益;其余三个点的提升在 +19% 至 +28% 之间。未观察到回归;gsm8k strict 的微小下降在一个标准误之内。
  • 下一步: finish 校验矩阵/结果覆盖并将 PR 标记为可评审,以启动自动评审。已使用修复次数:1/5(一次不改代码的评测重跑)。

@github-actions

Copy link
Copy Markdown
Contributor

@Klaud-Cold
Klaud-Cold marked this pull request as ready for review September 10, 2026 03:03
@Klaud-Cold
Klaud-Cold requested a review from a team September 10, 2026 03:03
@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Klaud Cold: validated. All owned runs are terminal. Repairs: 1. Runs: 34422693420, 34429017350, 34429032159.

The full sweep is verified; this PR remains ready for review.


Klaud Cold:validated。所有自有运行均已结束。修复次数:1。运行:34422693420, 34429017350, 34429032159

完整 sweep 已通过验证;PR 保持就绪,等待审查。

@claude

claude Bot commented Sep 10, 2026

Copy link
Copy Markdown
Contributor

Claude finished @Klaud-Cold's task in 1m 13s —— View job


Review: PR #2952

  • Gather context (diff vs origin/main, commits, changelog tail)
  • Verify master-config change (image public, key/prefix, rest of recipe untouched)
  • Verify perf-changelog.yaml entry (appended at end, format, PR link, no append-only misuse)
  • Run validate_perf_changelog.py against origin/main — could not execute locally (Python execution is not in this job's allowed tools); relied on CI instead, where the changelog-metadata job in run-sweep 34429032159 passed on the exact PR head 5cb93a2d9
  • Post review summary

LGTM - no blocking issues found

The diff is exactly two hunks: the image: line for qwen3.5-fp8-h100-sglang-mtp in configs/nvidia-master.yaml (public Docker Hub tag lmsysorg/sglang:v0.5.19-cu130; model, model-prefix, runner, TP8/EP8 search space and MTP scenario untouched) and a single perf-changelog.yaml entry appended at the very end of the file with the matching config key and PR link. No append-only flag or eval modifiers are involved, and the recipe script is not modified, so the full sweep on this head (4/4 points plus gsm8k, all green) covers the change.
· Branch

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good, straightforward config change.

What was reviewed: the configs/nvidia-master.yaml hunk (confirmed only the image tag for qwen3.5-fp8-h100-sglang-mtp changed, all other recipe fields untouched) and the perf-changelog.yaml addition (confirmed it is a single append at the file tail with no reordering or byte changes to prior entries). Also checked for a model.container field that would need to mirror the image bump — this recipe schema has no container field anywhere in the file, so that concern doesn't apply.

Extended reasoning...

Overview

The PR changes exactly two files: configs/nvidia-master.yaml, where the qwen3.5-fp8-h100-sglang-mtp recipe's image field is bumped from lmsysorg/sglang:v0.5.14-cu130 to lmsysorg/sglang:v0.5.19-cu130, and perf-changelog.yaml, which gets one new entry appended at the end of the file documenting the bump, referencing the target image digest, and noting the recipe script/topology/EAGLE-MTP settings are unchanged. I ran the diff directly and confirmed both hunks match this description precisely — no other lines in the recipe block changed, and the changelog addition is a clean append with no modification to existing bytes.

Security risks

None. This is a data-only configuration change (a container image tag bump and a changelog entry) with no code paths, credentials, or access-control logic involved.

Level of scrutiny

Low. This is exactly the kind of small, mechanical, self-contained change the approval guidelines call out as approvable — a version bump with no ambiguity and a matching, correctly-formed changelog entry. The repo's own review conventions (append-only changelog, image-bump-must-be-a-real-image-change) are both satisfied per my direct read of the diff.

Other factors

The repository convention note raised the question of whether a model.container field also needs updating alongside image. I checked and this recipe/config schema has no container field anywhere in the file, so that check does not apply here. No CODEOWNERS restriction is evident for these two files, and the bug-hunting pass reported zero findings.

@adibarra

adibarra commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator

/reuse-sweep-run 34429032159

将 main 合并到 PR #2952,保留已验证的配方并复用完整扫描结果。
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Development

Successfully merging this pull request may close these issues.

2 participants