Skip to content

[Klaud Cold] Update dsr1-fp4-b300-sglang SGLang image to v0.5.19-cu130 / 将 dsr1-fp4-b300-sglang 的 SGLang 镜像更新至 v0.5.19-cu130 - #2951

Closed
Klaud-Cold wants to merge 1 commit into
mainfrom
klaud/auto-023c120132f497f0-a32566a003a59219
Closed

[Klaud Cold] Update dsr1-fp4-b300-sglang SGLang image to v0.5.19-cu130 / 将 dsr1-fp4-b300-sglang 的 SGLang 镜像更新至 v0.5.19-cu130#2951
Klaud-Cold wants to merge 1 commit into
mainfrom
klaud/auto-023c120132f497f0-a32566a003a59219

Conversation

@Klaud-Cold

Copy link
Copy Markdown
Collaborator

Refresh the dsr1-fp4-b300-sglang single-node SGLang image from lmsysorg/sglang:v0.5.12-cu130 to lmsysorg/sglang:v0.5.19-cu130 (digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, SGLang commit 0bcd822 = tag v0.5.19). Model, precision, topology, workloads and benchmarks/single_node/fixed_seq_len/dsr1_fp4_b300.sh are unchanged.

Baseline

  • Published date: 2026-05-22 (public API workflow-info?date=2026-05-22 and benchmarks?model=DeepSeek-R1-0528&date=2026-05-22&exact=true, filtered to hardware b300, framework sglang, precision fp4, spec none, non-disagg, ISL/OSL 8192/1024).
  • Old image: lmsysorg/sglang:v0.5.12-cu130 (SGLang commit 127b9e3 = tag v0.5.12, CUDA 13.0.1, FlashInfer 0.6.11.post1).
  • Workload / topology: nvidia/DeepSeek-R1-0528-FP4-V2, fixed-seq-len 8k/1k single turn, TP4/EP4 conc 1–128 and TP8/EP8 conc 1–16 (13 points), runner b300 (cluster b300-dsxe), evals gsm8k at TP4 conc 64/128.
  • Producer run: all 13 points and both evals come from run 26306422380 attempt 1, head SHA 6a4fe2a1272bcd6c9df8ea6539f593088ba90437 (#1534). Note: that sweep's overall conclusion is recorded as failure (it covered six DSR1 SGLang configs); the 13 dsr1-fp4-b300-sglang points were ingested and are the currently published curve. Logical curve snapshot id 1943 is not the producer.
TP/EP conc tput/GPU (tok/s) output tput/GPU median TTFT (s) median TPOT (s) median E2EL (s)
TP4/EP4 1 355.6 39.7 0.190 0.0061 5.72
TP4/EP4 2 605.0 67.9 0.242 0.0071 6.91
TP4/EP4 4 1045.5 116.3 0.228 0.0082 7.67
TP4/EP4 8 1651.0 185.6 0.253 0.0103 9.82
TP4/EP4 16 2400.5 266.0 0.492 0.0142 13.33
TP4/EP4 32 3321.2 371.4 0.734 0.0204 19.48
TP4/EP4 64 4495.5 498.7 1.057 0.0306 29.21
TP4/EP4 128 5661.9 627.1 1.560 0.0488 46.37
TP8/EP8 1 186.5 20.8 0.132 0.0059 5.45
TP8/EP8 2 329.9 37.0 0.165 0.0065 6.34
TP8/EP8 4 580.7 64.6 0.182 0.0074 6.89
TP8/EP8 8 968.2 108.8 0.195 0.0088 8.30
TP8/EP8 16 1497.8 166.0 0.384 0.0113 10.69

Published evals (gsm8k, evaluations?date=2026-05-22, same producer run; the API labels these rows disagg: true although the recipe is aggregated single-node):

TP/EP conc gsm8k em_strict gsm8k em_flexible
TP4/EP4 64 0.9575 0.9575
TP4/EP4 128 0.9545 0.9560

dsr1-fp4-b300-sglang 单节点 SGLang 镜像从 lmsysorg/sglang:v0.5.12-cu130 更新至 lmsysorg/sglang:v0.5.19-cu130(摘要 sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,SGLang 提交 0bcd822 即标签 v0.5.19)。模型、精度、拓扑、工作负载以及 benchmarks/single_node/fixed_seq_len/dsr1_fp4_b300.sh 均保持不变。

基线

  • 发布日期: 2026-05-22(公共 API workflow-info?date=2026-05-22benchmarks?model=DeepSeek-R1-0528&date=2026-05-22&exact=true,按硬件 b300、框架 sglang、精度 fp4、无投机解码、非分离式、ISL/OSL 8192/1024 过滤)。
  • 旧镜像: lmsysorg/sglang:v0.5.12-cu130(SGLang 提交 127b9e3 即标签 v0.5.12,CUDA 13.0.1,FlashInfer 0.6.11.post1)。
  • 工作负载 / 拓扑: nvidia/DeepSeek-R1-0528-FP4-V2,fixed-seq-len 8k/1k 单轮,TP4/EP4 并发 1–128 与 TP8/EP8 并发 1–16(共 13 个点),runner b300(集群 b300-dsxe),评测为 TP4 并发 64/128 的 gsm8k。
  • 生产运行: 全部 13 个点与两项评测均来自 run 26306422380 attempt 1,head SHA 6a4fe2a1272bcd6c9df8ea6539f593088ba90437#1534)。说明:该 sweep 的整体结论记录为 failure(涵盖六个 DSR1 SGLang 配置),但这 13 个 dsr1-fp4-b300-sglang 点已被摄入并构成当前发布曲线。逻辑曲线快照 id 1943 并非生产运行。

基准数值见上方英文表格(吞吐 tok/s/GPU、输出吞吐、中位 TTFT/TPOT/E2EL)。

已发布评测(gsm8k,evaluations?date=2026-05-22,同一生产运行;API 将这些行标记为 disagg: true,但该配方实际为聚合式单节点):TP4/EP4 并发 64 em_strict 0.9575 / em_flexible 0.9575;并发 128 em_strict 0.9545 / em_flexible 0.9560。

Update the dsr1-fp4-b300-sglang master image from lmsysorg/sglang:v0.5.12-cu130
to lmsysorg/sglang:v0.5.19-cu130, pinned by manifest digest. Model, precision,
topology, workloads and the recipe script are unchanged.

将 dsr1-fp4-b300-sglang 的 SGLang 镜像从 lmsysorg/sglang:v0.5.12-cu130 更新至
lmsysorg/sglang:v0.5.19-cu130(按 manifest 摘要固定)。模型、精度、拓扑、
工作负载和配方脚本均保持不变。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

@Klaud-Cold

Klaud-Cold commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator Author

Initial attempt — failed (infrastructure: model not staged on DSXE)

  • Image / SHA: lmsysorg/sglang:v0.5.19-cu130@sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9 on PR head 77874d7a7e6856bbebd7aab928a55a98ee9d4ae6.
  • Run: targeted smoke e2e-tests run 34422578516 (test-config --config-files configs/nvidia-master.yaml --config-keys dsr1-fp4-b300-sglang --trim-conc, fail-fast, Klaud background priority). Points: TP4/EP4 conc 1 (benchmark + gsm8k eval) and TP8/EP8 conc 1. Startup/compatibility check only, not the full curve.
  • Result: failure. Jobs: TP4/EP4 c1 eval-only failure, TP8/EP8 c1 failure, TP4/EP4 c1 benchmark cancelled by fail-fast. results_bmk/agg_bmk.json and eval_results_all/agg_eval_all.json are empty ([]); no benchmark or eval ran. All jobs are terminal.
  • Benchmark / eval deltas vs. baseline: N/A (server never became ready).
  • First server error (both failed jobs, identical): sglang.launch_server died in TokenizerManagerAutoTokenizer.from_pretrained('/data/home/sa-gha-runner/models/DeepSeek-R1-0528-FP4-V2') → transformers 5.12.1 ValueError: Couldn't instantiate the backend tokenizer from one of: (1) a tokenizers library serialization file ..., i.e. no tokenizer.json was present in the model directory at load time. The HF repo nvidia/DeepSeek-R1-0528-FP4-V2 (rev 25a138f) does ship tokenizer.json, and the same checkpoint loads under the same transformers 5.12.1 pin on B200 (dsr1-fp4-b200-sglang, v0.5.16-cu130, published 2026-08-06), so this is not an engine/tokenizer-library incompatibility.
  • Diagnosis: runners/launch_b300-dsxe.sh stages DeepSeek-R1-0528-NVFP4-v2 (the Dynamo recipes' checkpoint) but not DeepSeek-R1-0528-FP4-V2, so this recipe is routed to the shared writable download dir /data/home/sa-gha-runner/models/DeepSeek-R1-0528-FP4-V2. The TP4 c1 benchmark job started first and began hf download (174 files; ~25 % fetched by 00:50 UTC). The eval job (00:50:38) and TP8 job (00:52:47) then found that directory non-empty, skipped the download per dsr1_fp4_b300.sh, launched the server against the partial checkpoint and failed on the missing tokenizer.json. Fail-fast cancelled the TP4 job at 00:54 mid-download, leaving a partial directory on shared storage, so a re-dispatch would skip the download again and fail the same way. The B300 b300 runner was remapped to cluster:b300-dsxe on 2026-09-04 (#2826); no dsr1-fp4-b300-sglang job has run on DSXE before, and the issue is independent of the image (the shipped v0.5.12 image would hit the same race). The image itself imported fine (enroot squash for the digest-pinned tag was already present) and all recipe flags were accepted by v0.5.19.
  • Upstream delta (v0.5.12 → v0.5.19): unchanged from the analysis above (old = SGLang 127b9e3, new = 0bcd822 = tag v0.5.19; CUDA 13.0.1→13.0.3, FlashInfer 0.6.11.post1→0.6.18, sgl-kernel 0.4.2.post2→0.4.6.post1, torch 2.11→2.13.0, transformers 5.6.0→5.12.1). --cuda-graph-max-bs and --enable-flashinfer-allreduce-fusion are deprecated aliases but still accepted (confirmed in the server log). Defaults changed on this path: MoE deferred finalize on for NVFP4 + flashinfer_trtllm DeepSeek-V3 family (#33618); FlashInfer MNNVL pure allreduce auto-enabled (#30700); unified radix tree default (#35081) is moot with --disable-radix-cache.
  • Next step: stop as readiness-blocked (missing staged weights on cluster:b300-dsxe). Repairs used: 0/5. The image change is not at fault; the candidate branch is released so a later autosweep can retry once DeepSeek-R1-0528-FP4-V2 is staged (or the partial writable directory is cleared) on DSXE. Fixing STAGED_MODELS/the launcher is shared-code and outside this PR's scope.

初次尝试 — 失败(基础设施:模型未在 DSXE 上预置)

  • 镜像 / SHA: lmsysorg/sglang:v0.5.19-cu130@sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,PR head 77874d7a7e6856bbebd7aab928a55a98ee9d4ae6
  • 运行: 定向冒烟 e2e-tests run 34422578516test-config --config-files configs/nvidia-master.yaml --config-keys dsr1-fp4-b300-sglang --trim-conc,fail-fast,Klaud 后台优先级)。测试点:TP4/EP4 并发 1(基准 + gsm8k 评测)与 TP8/EP8 并发 1。仅为启动/兼容性检查,并非完整曲线。
  • 结果: failure。作业:TP4/EP4 c1 eval-only failure,TP8/EP8 c1 failure,TP4/EP4 c1 基准被 fail-fast cancelledresults_bmk/agg_bmk.jsoneval_results_all/agg_eval_all.json 均为空([]),未产生任何基准或评测结果。所有作业均已终止。
  • 相对基线的基准 / 评测差异: N/A(服务端从未就绪)。
  • 首个服务端错误(两个失败作业完全一致): sglang.launch_serverTokenizerManagerAutoTokenizer.from_pretrained('/data/home/sa-gha-runner/models/DeepSeek-R1-0528-FP4-V2') 处退出,transformers 5.12.1 报 ValueError: Couldn't instantiate the backend tokenizer from one of: (1) a tokenizers library serialization file ...,即加载时模型目录中不存在 tokenizer.json。HF 仓库 nvidia/DeepSeek-R1-0528-FP4-V2(rev 25a138f)包含 tokenizer.json,且同一检查点在 B200 上以相同的 transformers 5.12.1 版本可正常加载(dsr1-fp4-b200-sglangv0.5.16-cu130,2026-08-06 发布),因此这不是引擎 / 分词器库的兼容性问题。
  • 诊断: runners/launch_b300-dsxe.sh 预置了 DeepSeek-R1-0528-NVFP4-v2(Dynamo 配方使用的检查点),但没有预置 DeepSeek-R1-0528-FP4-V2,因此该配方被路由到共享可写下载目录 /data/home/sa-gha-runner/models/DeepSeek-R1-0528-FP4-V2。TP4 c1 基准作业最先启动并开始 hf download(174 个文件;至 00:50 UTC 约完成 25%)。随后评测作业(00:50:38)与 TP8 作业(00:52:47)发现该目录非空,按 dsr1_fp4_b300.sh 逻辑跳过下载,直接用不完整的检查点启动服务端并因缺少 tokenizer.json 失败。fail-fast 于 00:54 在下载中途取消了 TP4 作业,在共享存储上留下了不完整的目录,因此重新派发会再次跳过下载并以同样方式失败。B300 的 b300 runner 于 2026-09-04 重映射到 cluster:b300-dsxe#2826);此前从未有 dsr1-fp4-b300-sglang 作业在 DSXE 上运行,且该问题与镜像无关(原 v0.5.12 镜像也会遇到同样的竞争)。镜像本身导入正常(按摘要固定的标签对应的 enroot squash 已存在),v0.5.19 也接受了配方的全部参数。
  • 上游差异(v0.5.12 → v0.5.19): 与上文分析一致(旧 = SGLang 127b9e3,新 = 0bcd822 即标签 v0.5.19;CUDA 13.0.1→13.0.3,FlashInfer 0.6.11.post1→0.6.18,sgl-kernel 0.4.2.post2→0.4.6.post1,torch 2.11→2.13.0,transformers 5.6.0→5.12.1)。--cuda-graph-max-bs--enable-flashinfer-allreduce-fusion 为弃用别名但仍被接受(服务端日志已确认)。该路径默认值变化:NVFP4 + flashinfer_trtllm 的 DeepSeek-V3 系列默认开启 MoE deferred finalize(#33618);自动启用 FlashInfer MNNVL 纯 allreduce(#30700);统一 radix 树默认开启(#35081)因 --disable-radix-cache 而不产生影响。
  • 下一步:readiness-blockedcluster:b300-dsxe 缺少预置权重)终止。已用修复次数:0/5。镜像改动本身无问题;释放候选分支,待 DeepSeek-R1-0528-FP4-V2 在 DSXE 上预置完成(或清理不完整的可写目录)后,后续自动扫描可重试。修改 STAGED_MODELS / 启动器属于共享代码,不在本 PR 范围内。

@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Stop: readiness-blocked. The targeted smoke (run 34422578516) failed because nvidia/DeepSeek-R1-0528-FP4-V2 is not staged on cluster:b300-dsxe and the concurrent on-demand download left a partial model directory; see the Initial attempt for the diagnosis. No repairs used; the image is not implicated. Closing the PR and releasing the branch so the candidate can be retried after staging.


终止:readiness-blocked。 定向冒烟(run 34422578516)失败,原因是 nvidia/DeepSeek-R1-0528-FP4-V2 未在 cluster:b300-dsxe 上预置,且并发的按需下载留下了不完整的模型目录;诊断详见初次尝试。未使用修复次数;镜像本身无问题。关闭 PR 并释放分支,以便预置完成后重试该候选。

@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Klaud Cold: readiness-blocked. Finishing cleanup; owned child runs will be stopped and checked before closure.


Klaud Cold:readiness-blocked。正在完成清理;将先停止并确认自有子运行的状态,再关闭 PR。

@Klaud-Cold Klaud-Cold closed this Sep 10, 2026
@Klaud-Cold
Klaud-Cold deleted the klaud/auto-023c120132f497f0-a32566a003a59219 branch September 10, 2026 01:00
@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Klaud Cold: readiness-blocked. All owned runs are terminal. Repairs: 0. Runs: 34422578516.

PR closed; branch deleted for retry.


Klaud Cold:readiness-blocked。所有自有运行均已结束。修复次数:0。运行:34422578516

PR 已关闭;分支已删除,后续可以重试。

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Development

Successfully merging this pull request may close these issues.

1 participant