Skip to content

[Klaud Cold] Update qwen3.5-fp8-h200-sglang SGLang image to v0.5.19-cu130 / 将 qwen3.5-fp8-h200-sglang 的 SGLang 镜像升级至 v0.5.19-cu130 - #2965

Closed
Klaud-Cold wants to merge 2 commits into
mainfrom
klaud/auto-5e2f62b5f064c984-3d9c146c5e9ad2d6
Closed

[Klaud Cold] Update qwen3.5-fp8-h200-sglang SGLang image to v0.5.19-cu130 / 将 qwen3.5-fp8-h200-sglang 的 SGLang 镜像升级至 v0.5.19-cu130#2965
Klaud-Cold wants to merge 2 commits into
mainfrom
klaud/auto-5e2f62b5f064c984-3d9c146c5e9ad2d6

Conversation

@Klaud-Cold

Copy link
Copy Markdown
Collaborator

Update the qwen3.5-fp8-h200-sglang master image from lmsysorg/sglang:v0.5.14-cu130 to the current SGLang release lmsysorg/sglang:v0.5.19-cu130 (Docker Hub digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, tag commit sgl-project/sglang@0bcd822). The recipe script, model, TP8/EP8 topology and the 8k1k concurrency range are unchanged.

Baseline

  • Published date: 2026-07-04 (benchmarks?model=Qwen-3.5-397B-A17B&date=2026-07-04&exact=true, workflow-info?date=2026-07-04, evaluations?model=Qwen-3.5-397B-A17B&date=2026-07-04&exact=true)
  • Old image: lmsysorg/sglang:v0.5.14-cu130 (digest sha256:5027e95bf6ec536856b1b52a91d1f35ff5c564ab83e8a94758a169ff09bb8df3, tag commit sgl-project/sglang@49e384c)
  • Workload / topology: single-node H200, Qwen/Qwen3.5-397B-A17B-FP8, SGLang FP8, TP8 EP8, no speculative decoding, fixed-seq-len 8k1k (ISL 8192 / OSL 1024), concurrency 4 / 8 / 16 / 32 / 64, random dataset
  • Producer: run 28719990565 (head e016f0c14bf2aab899de08b538d9b343069b07a5, changelog PR #2061); benchmark result IDs 432218, 432222, 432231, 432227, 432220
Conc Total tok/s/GPU Output tok/s/GPU Median TTFT (s) Median TPOT (ms) Median E2E (s)
4 486.9 54.2 0.345 8.68 8.18
8 765.2 86.0 0.356 10.97 10.56
16 1124.1 124.6 0.357 15.41 14.36
32 1522.5 170.3 0.376 22.82 21.44
64 1992.1 221.0 0.561 35.70 33.38
  • Published evals: gsm8k at concurrency 32, em_strict 0.9651 / em_flexible 0.9583 (evaluation ID 8294) and at concurrency 64, em_strict 0.9689 / em_flexible 0.9583 (evaluation ID 8293), both from the same producer run. The public evaluations feed labels these rows disagg: true with 64 prefill/decode GPUs, which does not match the single-node TP8 recipe; the benchmark rows above carry the correct disagg: false identity.

qwen3.5-fp8-h200-sglang 的主配置镜像从 lmsysorg/sglang:v0.5.14-cu130 升级到当前 SGLang 发布版 lmsysorg/sglang:v0.5.19-cu130(Docker Hub 摘要 sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,标签提交 sgl-project/sglang@0bcd822)。配方脚本、模型、TP8/EP8 拓扑以及 8k1k 并发范围均保持不变。

基线

  • 发布日期: 2026-07-04(benchmarks?model=Qwen-3.5-397B-A17B&date=2026-07-04&exact=trueworkflow-info?date=2026-07-04evaluations?model=Qwen-3.5-397B-A17B&date=2026-07-04&exact=true
  • 旧镜像: lmsysorg/sglang:v0.5.14-cu130(摘要 sha256:5027e95bf6ec536856b1b52a91d1f35ff5c564ab83e8a94758a169ff09bb8df3,标签提交 sgl-project/sglang@49e384c
  • 工作负载 / 拓扑: 单节点 H200,Qwen/Qwen3.5-397B-A17B-FP8,SGLang FP8,TP8 EP8,无投机解码,固定序列长度 8k1k(ISL 8192 / OSL 1024),并发 4 / 8 / 16 / 32 / 64,随机数据集
  • 数据来源: 运行 28719990565(head e016f0c14bf2aab899de08b538d9b343069b07a5,changelog PR #2061);基准结果 ID 432218、432222、432231、432227、432220
并发 总吞吐 tok/s/GPU 输出吞吐 tok/s/GPU TTFT 中位数 (s) TPOT 中位数 (ms) 端到端中位数 (s)
4 486.9 54.2 0.345 8.68 8.18
8 765.2 86.0 0.356 10.97 10.56
16 1124.1 124.6 0.357 15.41 14.36
32 1522.5 170.3 0.376 22.82 21.44
64 1992.1 221.0 0.561 35.70 33.38
  • 已发布评测: 并发 32 的 gsm8k,em_strict 0.9651 / em_flexible 0.9583(评测 ID 8294);并发 64 的 gsm8k,em_strict 0.9689 / em_flexible 0.9583(评测 ID 8293),均来自同一数据来源运行。公共评测接口将这些行标记为 disagg: true 且 64 个 prefill/decode GPU,与单节点 TP8 配方不符;上表基准行的 disagg: false 身份是正确的。

🤖 Generated with Claude Code

…-cu130

Bump the master image for the qwen3.5-fp8-h200-sglang family from
lmsysorg/sglang:v0.5.14-cu130 to lmsysorg/sglang:v0.5.19-cu130
(Docker Hub digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,
tag commit sgl-project/sglang@0bcd822). The recipe script, model,
TP8/EP8 topology and the 8k1k concurrency range are unchanged.

将 qwen3.5-fp8-h200-sglang 家族的主配置镜像从 lmsysorg/sglang:v0.5.14-cu130
升级到 lmsysorg/sglang:v0.5.19-cu130(Docker Hub 摘要
sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,
标签提交 sgl-project/sglang@0bcd822)。配方脚本、模型、TP8/EP8 拓扑和 8k1k
并发范围均保持不变。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

@Klaud-Cold

Klaud-Cold commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator Author

Initial attempt

  • Image / SHA: lmsysorg/sglang:v0.5.19-cu130 (digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9) at PR head 15b97450f713bde9b77872f91ecf0cbdf80b90ad
  • Change: master image only (configs/nvidia-master.yaml, key qwen3.5-fp8-h200-sglang); benchmarks/single_node/fixed_seq_len/qwen3.5_fp8_h200.sh unchanged, no runtime patching
  • Targeted run: e2e run 34477386674 (e2e-tests.yml from main, ref=15b97450, test-config --trim-conc: TP8 EP8 at concurrency 4 plus the smoke gsm8k eval; dispatched 2026-09-10T12:33Z) — completed, success: benchmark job on h200-cw_00, eval-only job on h200-cw_01, collect-results / collect-evals / calc-success-rate all green (attempt 1)
  • Benchmark result (concurrency 4, 8k1k, vs published 2026-07-04 baseline, result ID 432218):
Metric Baseline v0.5.14 v0.5.19 Delta
Total tok/s/GPU 486.9 543.0 +11.5%
Output tok/s/GPU 54.2 60.4 +11.5%
Median TTFT (s) 0.345 0.222 -35.6%
Median TPOT (ms) 8.68 7.87 -9.3%
Median E2E (s) 8.18 7.34 -10.3%

Server log: The server is fired up and ready to roll! on all ranks with resolved attention_backend=flashinfer, linear_attn_backend=triton, mem_fraction_static=0.8, disable_radix_cache=True, kv_cache_dtype=fp8_e4m3, quantization=fp8, ep_size=8; new log lines are the expected deprecation warnings for --enable-flashinfer-allreduce-fusion (mapped to flashinfer_allreduce_fusion_backend=auto) and --cuda-graph-max-bs (mapped to cuda_graph_max_bs_decode=4). power_valid=1 in the aggregate.

  • Eval result: gsm8k at concurrency 4 (smoke selection), em_strict 0.9666 / em_flexible 0.9591, n = 1319, infrastructure_success=true. Published baseline evals were 0.9651 / 0.9583 at concurrency 32 and 0.9689 / 0.9583 at concurrency 64, so this is indicative rather than a like-for-like point. The targeted smoke (startup/compatibility evidence, not a full curve) is complete.
  • Next step: the perf-changelog.yaml entry is appended (head 6d558094c11d407242970cd6ab3b3a0a19035c6b, validate_perf_changelog.py against origin/main passed); recheck capacity and start the final full sweep (full-sweep-enabled, PR kept draft).
  • Upstream source comparison: v0.5.14 @ 49e384c (2026-06-25) → v0.5.19 @ 0bcd822 (2026-09-03). Provenance: both images carry ai.sglang.build.commit / org.opencontainers.image.revision labels equal to those tag commits, and v0.5.19 and v0.5.19-cu130 share one Docker Hub digest pushed 2026-09-04 (docker/Dockerfile builds from the v${SGL_VERSION} tag).
  • Coupled dependencies: torch 2.11.0 → 2.13.0, flashinfer 0.6.12 → 0.6.18, sgl-kernel 0.4.4 → 0.4.6.post1, transformers 5.8.1 → 5.12.1, CUDA base 13.0.1 → 13.0.3 (same cuda>=13.0 / driver ≥ 535 requirement as the old cu130 image).
  • Flag audit: all 22 launch flags used by the script are still defined in v0.5.19 (server_args.py plus the new arg_groups/ hooks), with --expert-parallel-size kept as an alias of --ep-size and unchanged choices for --attention-backend flashinfer, --kv-cache-dtype fp8_e4m3, --quantization fp8 and --mamba-ssm-dtype bfloat16. --enable-flashinfer-allreduce-fusion is deprecated and maps to --flashinfer-allreduce-fusion-backend auto with a warning (serving_hook.py). The new Qwen3.5 hybrid override (model_overrides/qwen3_5.py) and the bf16 SSM-state defaults in attention_hook.py apply only to SM100 or when no attention backend is given, so the explicit flashinfer backend on H200 (SM90) keeps its path with Triton linear-attention decode.
  • Behavior deltas noted: CUTLASS FP8 blockwise GEMM was deleted for SM90 in v0.5.16 (#30438) and v0.5.19 adds an SM90 FP8 decode routing fix (#37018), so H200 FP8 GEMM kernels differ from the baseline image; unified radix tree is the default in v0.5.19 (#35081; moot under --disable-radix-cache); compiled-kernel caches moved under SGLANG_CACHE_DIR in v0.5.18 (#32434). The same image already passed a full sweep for the sibling qwen3.5-fp8-h100-sglang-mtp (#2952, Hopper SM90, same model) and the concurrency-4 smoke for dsr1-fp8-h200-sglang-mtp (#2955, same H200 fleet).
  • Decision: image-only bump; no in-scope flag changes required.

初始尝试

  • 镜像 / SHA: lmsysorg/sglang:v0.5.19-cu130(摘要 sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9),PR head 15b97450f713bde9b77872f91ecf0cbdf80b90ad
  • 改动: 仅主配置镜像(configs/nvidia-master.yaml 中的 qwen3.5-fp8-h200-sglang);benchmarks/single_node/fixed_seq_len/qwen3.5_fp8_h200.sh 未改动,无任何运行时补丁
  • 定向运行: e2e 运行 34477386674(从 main 触发 e2e-tests.ymlref=15b97450,test-config --trim-conc:TP8 EP8、并发 4,附带 gsm8k 冒烟评测;2026-09-10T12:33Z 触发)—— 已完成,成功:基准任务在 h200-cw_00,仅评测任务在 h200-cw_01,collect-results / collect-evals / calc-success-rate 全部通过(第 1 次尝试)
  • 基准结果(并发 4,8k1k,对比 2026-07-04 已发布基线,结果 ID 432218):
指标 基线 v0.5.14 v0.5.19 变化
总吞吐 tok/s/GPU 486.9 543.0 +11.5%
输出吞吐 tok/s/GPU 54.2 60.4 +11.5%
TTFT 中位数 (s) 0.345 0.222 -35.6%
TPOT 中位数 (ms) 8.68 7.87 -9.3%
端到端中位数 (s) 8.18 7.34 -10.3%

服务日志:所有 rank 均出现 The server is fired up and ready to roll!,解析后的设置为 attention_backend=flashinferlinear_attn_backend=tritonmem_fraction_static=0.8disable_radix_cache=Truekv_cache_dtype=fp8_e4m3quantization=fp8ep_size=8;新增日志仅为预期中的弃用告警:--enable-flashinfer-allreduce-fusion(映射为 flashinfer_allreduce_fusion_backend=auto)和 --cuda-graph-max-bs(映射为 cuda_graph_max_bs_decode=4)。聚合结果中 power_valid=1

  • 评测结果: 并发 4 的 gsm8k(冒烟选择),em_strict 0.9666 / em_flexible 0.9591,n = 1319,infrastructure_success=true。已发布基线评测为并发 32 的 0.9651 / 0.9583 和并发 64 的 0.9689 / 0.9583,因此该数值仅供参考,不是同一并发点的对比。定向冒烟测试(属于启动/兼容性证据,不是完整曲线)已完成。
  • 下一步: perf-changelog.yaml 条目已追加(head 6d558094c11d407242970cd6ab3b3a0a19035c6b,本地对 origin/main 运行 validate_perf_changelog.py 通过);重新检查容量并启动最终完整扫描(full-sweep-enabled,PR 保持草稿)。
  • 上游源码对比: v0.5.14 @ 49e384c(2026-06-25)→ v0.5.19 @ 0bcd822(2026-09-03)。来源确认:两个镜像的 ai.sglang.build.commit / org.opencontainers.image.revision 标签均等于对应标签提交,且 v0.5.19v0.5.19-cu130 在 Docker Hub 上共享同一摘要(2026-09-04 推送);docker/Dockerfilev${SGL_VERSION} 标签构建。
  • 耦合依赖: torch 2.11.0 → 2.13.0,flashinfer 0.6.12 → 0.6.18,sgl-kernel 0.4.4 → 0.4.6.post1,transformers 5.8.1 → 5.12.1,CUDA 基础镜像 13.0.1 → 13.0.3(与旧 cu130 镜像相同的 cuda>=13.0 / 驱动 ≥ 535 要求)。
  • 参数审计: 脚本使用的全部 22 个启动参数在 v0.5.19 中仍有定义(server_args.py 及新的 arg_groups/ 钩子),--expert-parallel-size 仍是 --ep-size 的别名,--attention-backend flashinfer--kv-cache-dtype fp8_e4m3--quantization fp8--mamba-ssm-dtype bfloat16 的可选值未变。--enable-flashinfer-allreduce-fusion 已弃用,会告警并映射为 --flashinfer-allreduce-fusion-backend autoserving_hook.py)。新增的 Qwen3.5 混合架构覆盖(model_overrides/qwen3_5.py)和 attention_hook.py 中的 bf16 SSM 状态默认值仅在 SM100 或未显式指定注意力后端时生效,因此 H200(SM90)上显式指定的 flashinfer 后端路径不变,线性注意力解码仍使用 Triton。
  • 已注意到的行为变化: v0.5.16 删除了 SM90 的 CUTLASS FP8 blockwise GEMM(#30438),v0.5.19 新增 SM90 FP8 解码路由修复(#37018),因此 H200 的 FP8 GEMM 内核与基线镜像不同;v0.5.19 默认启用统一 radix tree(#35081;在 --disable-radix-cache 下无影响);v0.5.18 起编译内核缓存迁移至 SGLANG_CACHE_DIR#32434)。同一镜像已在同模型的 Hopper(SM90)姊妹家族 qwen3.5-fp8-h100-sglang-mtp#2952)完成全量扫描,并在同一 H200 集群的 dsr1-fp8-h200-sglang-mtp#2955)通过并发 4 冒烟测试。
  • 决定: 仅升级镜像;无需在范围内修改参数。

….19-cu130

Append the perf-changelog.yaml entry for the qwen3.5-fp8-h200-sglang
SGLang image update (PR #2965), preserving all prior bytes.

为 qwen3.5-fp8-h200-sglang 的 SGLang 镜像更新(PR #2965)追加
perf-changelog.yaml 条目,保留所有历史字节。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Capacity deferral before the final full sweep

  • State: the Initial attempt targeted smoke passed (benchmark + gsm8k eval, run 34477386674, terminal success). The perf-changelog.yaml entry is committed at head 6d558094c11d407242970cd6ab3b3a0a19035c6b and validated locally. No final sweep was started; full-sweep-enabled was never applied. Repairs used: 0/5.
  • Reason: the mandatory pre-transition capacity check for the H200 target clusters of this family did not pass (two consecutive checks at 2026-09-10T14:13Z), so the final full sweep is deferred rather than queued.
  • Disposition: this PR is closed and its branch released so a later autosweep can retry the same candidate from the published data. The image itself showed no incompatibility: all launch flags resolved, the server started on both nodes, and the concurrency-4 point improved on the published baseline.

最终完整扫描前的容量推迟

  • 状态: 初始尝试的定向冒烟测试已通过(基准 + gsm8k 评测,运行 34477386674,终态 success)。perf-changelog.yaml 条目已提交至 head 6d558094c11d407242970cd6ab3b3a0a19035c6b 并在本地验证通过。未启动最终扫描;从未添加 full-sweep-enabled。已用修复次数:0/5。
  • 原因: 该家族 H200 目标集群的强制性转换前容量检查未通过(2026-09-10T14:13Z 连续两次检查),因此最终完整扫描被推迟而未排队。
  • 处置: 关闭此 PR 并释放分支,以便后续自动扫描根据已发布数据重试同一候选。镜像本身未显示不兼容:所有启动参数均已解析,服务端在两台节点上均成功启动,并发 4 的点相对已发布基线有所提升。

@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Klaud Cold: capacity-deferred. Finishing cleanup; owned child runs will be stopped and checked before closure.


Klaud Cold:capacity-deferred。正在完成清理;将先停止并确认自有子运行的状态,再关闭 PR。

@Klaud-Cold Klaud-Cold closed this Sep 10, 2026
@Klaud-Cold
Klaud-Cold deleted the klaud/auto-5e2f62b5f064c984-3d9c146c5e9ad2d6 branch September 10, 2026 14:14
@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Klaud Cold: capacity-deferred. All owned runs are terminal. Repairs: 0. Runs: 34477386674, 34487565175.

PR closed; branch deleted for retry.


Klaud Cold:capacity-deferred。所有自有运行均已结束。修复次数:0。运行:34477386674, 34487565175

PR 已关闭;分支已删除,后续可以重试。

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Development

Successfully merging this pull request may close these issues.

1 participant