Skip to content

[Klaud Cold] Update dsr1-fp8-mi325x-sglang SGLang ROCm image to v0.5.19-rocm700-mi30x / 将 dsr1-fp8-mi325x-sglang 的 SGLang ROCm 镜像更新至 v0.5.19-rocm700-mi30x - #2949

Merged
adibarra merged 3 commits into
mainfrom
klaud/auto-9d94bebbcc742ac8-d7346b01615deec4
Sep 10, 2026
Merged

[Klaud Cold] Update dsr1-fp8-mi325x-sglang SGLang ROCm image to v0.5.19-rocm700-mi30x / 将 dsr1-fp8-mi325x-sglang 的 SGLang ROCm 镜像更新至 v0.5.19-rocm700-mi30x#2949
adibarra merged 3 commits into
mainfrom
klaud/auto-9d94bebbcc742ac8-d7346b01615deec4

Conversation

@Klaud-Cold

@Klaud-Cold Klaud-Cold commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator

Update the dsr1-fp8-mi325x-sglang recipe image from lmsysorg/sglang:v0.5.12-rocm700-mi30x to lmsysorg/sglang:v0.5.19-rocm700-mi30x (digest sha256:590a815c128d7d5c83ef771ad768c9f8be82f64f7d0dbd98dc7b9f085b268a58). Model, TP8 topology, concurrency points, launch flags and evals are unchanged.

Baseline

  • Published date: 2026-05-18, old image lmsysorg/sglang:v0.5.12-rocm700-mi30x (SGLang tag commit 127b9e3, AITER a6bb499, base rocm/sgl-dev:rocm7-vllm-20250904)
  • Workload/topology: DeepSeek-R1-0528 FP8, MI325X, SGLang aggregated, no speculative decoding, ISL 8192 / OSL 1024, TP8, concurrency 4/8/16/32/64
  • Producer: run 26008679544 attempt 1 ("Run Sweep - Update dsr1-fp8-mi325x-sglang SGLang image to v0.5.12-rocm700-mi30x")
  • API queries: GET /api/v1/workflow-info?date=2026-05-18; GET /api/v1/benchmarks?model=DeepSeek-R1-0528&date=2026-05-18&exact=true&sequence=8k/1k filtered to mi325x / sglang / fp8 / spec none / disagg false; GET /api/v1/evaluations filtered to the producer run
  • Source links: SGLang v0.5.12 -> SGLang v0.5.19 (commit 0bcd822, AITER c16d44b, same ROCm 7.0 base image), docker/rocm.Dockerfile @ v0.5.19
conc tput/GPU (tok/s) output tput/GPU (tok/s) median TTFT (s) median TPOT (s) median E2EL (s)
4 224.98 25.03 0.703 0.0184 17.71
8 379.12 42.62 0.560 0.0215 21.14
16 616.81 68.35 0.439 0.0268 25.65
32 849.43 94.99 0.470 0.0396 37.42
64 1028.90 114.15 0.485 0.0671 62.75

Published evals (gsm8k, n_eff 1319): conc 32 em_strict 0.9530 / em_flexible 0.9553; conc 64 em_strict 0.9545 / em_flexible 0.9553. The published eval rows carry disagg: true although the recipe is aggregated; treated as a metadata quirk of the published data.


dsr1-fp8-mi325x-sglang 配方的镜像从 lmsysorg/sglang:v0.5.12-rocm700-mi30x 更新至 lmsysorg/sglang:v0.5.19-rocm700-mi30x(摘要 sha256:590a815c128d7d5c83ef771ad768c9f8be82f64f7d0dbd98dc7b9f085b268a58)。模型、TP8 拓扑、并发点、启动参数与评测均保持不变。

基线

  • 发布日期:2026-05-18,旧镜像 lmsysorg/sglang:v0.5.12-rocm700-mi30x(SGLang 标签提交 127b9e3,AITER a6bb499,基础镜像 rocm/sgl-dev:rocm7-vllm-20250904
  • 工作负载/拓扑:DeepSeek-R1-0528 FP8,MI325X,SGLang 聚合部署,无投机解码,ISL 8192 / OSL 1024,TP8,并发 4/8/16/32/64
  • 产出运行:run 26008679544 attempt 1
  • API 查询:GET /api/v1/workflow-info?date=2026-05-18GET /api/v1/benchmarks?model=DeepSeek-R1-0528&date=2026-05-18&exact=true&sequence=8k/1k,筛选 mi325x / sglang / fp8 / 无投机解码 / 非分离式;GET /api/v1/evaluations 筛选至该产出运行
  • 源码链接:SGLang v0.5.12 -> SGLang v0.5.19(提交 0bcd822,AITER c16d44b,ROCm 7.0 基础镜像相同),docker/rocm.Dockerfile @ v0.5.19
并发 每 GPU 吞吐 (tok/s) 每 GPU 输出吞吐 (tok/s) TTFT 中位数 (s) TPOT 中位数 (s) E2EL 中位数 (s)
4 224.98 25.03 0.703 0.0184 17.71
8 379.12 42.62 0.560 0.0215 21.14
16 616.81 68.35 0.439 0.0268 25.65
32 849.43 94.99 0.470 0.0396 37.42
64 1028.90 114.15 0.485 0.0671 62.75

已发布评测(gsm8k,n_eff 1319):并发 32 em_strict 0.9530 / em_flexible 0.9553;并发 64 em_strict 0.9545 / em_flexible 0.9553。已发布评测行标记为 disagg: true,而该配方为聚合部署,视为发布数据的元数据差异。

🤖 Generated with Claude Code


Note

Low Risk
Config-only container image pin for a single benchmark recipe; no runtime code or topology changes.

Overview
Bumps the dsr1-fp8-mi325x-sglang AMD master recipe to SGLang ROCm v0.5.19-rocm700-mi30x (from v0.5.12-rocm700-mi30x). Model, TP8 fixed 8k/1k sweep, concurrency range, and other recipe fields are unchanged.

Adds a perf-changelog.yaml entry for config key dsr1-fp8-mi325x-sglang that records the image update and links PR #2949.

Reviewed by Cursor Bugbot for commit 31c2396. Bugbot is set up for automated code reviews on this repo. Configure here.

…9-rocm700-mi30x

Bump lmsysorg/sglang from v0.5.12-rocm700-mi30x to v0.5.19-rocm700-mi30x
(digest sha256:590a815c128d7d5c83ef771ad768c9f8be82f64f7d0dbd98dc7b9f085b268a58)
for the DeepSeek-R1-0528 FP8 MI325X SGLang 8k1k recipe. Same ROCm 7.0
base image; model, TP8 topology, concurrencies and launch flags unchanged.

将 dsr1-fp8-mi325x-sglang 的 SGLang ROCm 镜像从 v0.5.12-rocm700-mi30x 更新至
v0.5.19-rocm700-mi30x(同一 ROCm 7.0 基础镜像)。模型、TP8 拓扑、并发点与
启动参数均保持不变。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

@Klaud-Cold

Klaud-Cold commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator Author

Initial attempt

  • Image: lmsysorg/sglang:v0.5.19-rocm700-mi30x (digest sha256:590a815c…268a58, pushed 2026-09-04)
  • Head: 36d40ac25311826918083bdb2c7befff3a09d2a2
  • Change: master-config image bump only (configs/amd-master.yaml, key dsr1-fp8-mi325x-sglang)
  • Smoke run: e2e run 34422505103, e2e-tests.yml on main with ref=36d40ac2, test-config --config-files configs/amd-master.yaml --config-keys dsr1-fp8-mi325x-sglang --trim-conc (1 point: TP8, conc 4, 8k1k, eval enabled)
  • Upstream review (v0.5.12 127b9e3 -> v0.5.19 0bcd822):
    • Same gfx942 ROCm 7.0 base image rocm/sgl-dev:rocm7-vllm-20250904 in docker/rocm.Dockerfile; AITER pin a6bb499 -> c16d44b.
    • --cuda-graph-max-bs is now a deprecated alias of --cuda-graph-max-bs-decode (server_args.py); still accepted.
    • New AITER rule in arg_groups/attention_hook.py: with --attention-backend aiter and context length > 8192, --mem-fraction-static is scaled by 0.85 unless SGLANG_AITER_HONOR_EXPLICIT_MEM_FRACTION=1. The throughput server therefore runs at an effective 0.68 instead of 0.8; left as shipped, no script change.
    • SGLANG_USE_AITER, SGLANG_AITER_MLA_PERSIST, --kv-cache-dtype fp8_e4m3, --chunked-prefill-size, --max-prefill-tokens, --num-continuous-decode-steps, --disable-radix-cache unchanged; the v0.5.19 DeepSeek-R1 deployment snippet lists the same MI325X FP8 flags.
  • Results: targeted smoke passed (run concluded success at 02:19 UTC). Benchmark job produced real output (40 requests, 294,050 input / 36,805 output tokens, 185 s window, power valid on 8/8 GPUs). Eval job (gsm8k, lm-eval, n_eff 1319) finished at 02:18 UTC with infrastructure_success: true.
point metric baseline v0.5.12 v0.5.19 smoke delta
TP8 c4 tput/GPU (tok/s) 224.98 223.37 -0.7%
TP8 c4 output tput/GPU (tok/s) 25.03 24.85 -0.7%
TP8 c4 median TTFT (s) 0.703 0.738 +5.0%
TP8 c4 median TPOT (s) 0.0184 0.0186 +0.8%
TP8 c4 median E2EL (s) 17.71 17.85 +0.8%
eval metric published baseline (conc 32 / 64) v0.5.19 smoke (conc 4)
gsm8k em_strict 0.9530 / 0.9545 0.9591
gsm8k em_flexible 0.9553 / 0.9553 0.9613

The smoke eval ran on the trimmed concurrency-4 point while the published evals are at concurrency 32 and 64, so the eval comparison is indicative only.

Server log confirms the AITER rule: mem_fraction_static resolved to 0.68 (0.8 x 0.85); KV pool 2,659,207 fp8 tokens per GPU (87 GB), far above the ~606k tokens needed at concurrency 64. Weights loaded in 224 s; --cuda-graph-max-bs deprecation warning printed only.

  • Next step: append the perf-changelog.yaml entry, recheck capacity and start the final full sweep (full-sweep-enabled). This smoke result is startup/compatibility evidence, not final validation.

初次尝试

  • 镜像:lmsysorg/sglang:v0.5.19-rocm700-mi30x(摘要 sha256:590a815c…268a58,2026-09-04 推送)
  • 提交:36d40ac25311826918083bdb2c7befff3a09d2a2
  • 改动:仅更新主配置镜像(configs/amd-master.yaml,键 dsr1-fp8-mi325x-sglang
  • 冒烟运行:e2e run 34422505103,在 main 上以 ref=36d40ac2test-config --config-files configs/amd-master.yaml --config-keys dsr1-fp8-mi325x-sglang --trim-conc 触发 e2e-tests.yml(1 个点:TP8,并发 4,8k1k,含评测)
  • 上游源码对比(v0.5.12 127b9e3 -> v0.5.19 0bcd822):
    • docker/rocm.Dockerfile 中 gfx942 ROCm 7.0 基础镜像 rocm/sgl-dev:rocm7-vllm-20250904 相同;AITER 固定提交 a6bb499 -> c16d44b
    • --cuda-graph-max-bs 现为 --cuda-graph-max-bs-decode 的已弃用别名,仍可用。
    • attention_hook.py 新增 AITER 规则:--attention-backend aiter 且上下文长度 > 8192 时,--mem-fraction-static 会乘以 0.85(除非设置 SGLANG_AITER_HONOR_EXPLICIT_MEM_FRACTION=1)。吞吐服务实际按 0.68 而非 0.8 运行;按镜像原样运行,不改脚本。
    • 其余环境变量与参数均未变化;v0.5.19 的 DeepSeek-R1 部署文档列出的 MI325X FP8 参数与本脚本一致。
  • 结果:定向冒烟通过(运行于 02:19 UTC 成功结束)。基准任务产出真实结果(40 个请求,输入 294,050 / 输出 36,805 token,185 秒窗口,8/8 GPU 功耗有效)。评测任务(gsm8k,lm-eval,n_eff 1319)于 02:18 UTC 完成,infrastructure_success: true
指标 基线 v0.5.12 v0.5.19 冒烟 差异
TP8 c4 每 GPU 吞吐 (tok/s) 224.98 223.37 -0.7%
TP8 c4 每 GPU 输出吞吐 (tok/s) 25.03 24.85 -0.7%
TP8 c4 TTFT 中位数 (s) 0.703 0.738 +5.0%
TP8 c4 TPOT 中位数 (s) 0.0184 0.0186 +0.8%
TP8 c4 E2EL 中位数 (s) 17.71 17.85 +0.8%
评测 指标 已发布基线(并发 32 / 64) v0.5.19 冒烟(并发 4)
gsm8k em_strict 0.9530 / 0.9545 0.9591
gsm8k em_flexible 0.9553 / 0.9553 0.9613

冒烟评测在裁剪后的并发 4 点上运行,而已发布评测位于并发 32 和 64,因此评测对比仅供参考。

服务日志证实 AITER 规则生效:mem_fraction_static 解析为 0.68(0.8 x 0.85);每 GPU KV 池 2,659,207 个 fp8 token(87 GB),远高于并发 64 所需的约 606k token。权重加载 224 秒;仅打印 --cuda-graph-max-bs 弃用警告。

  • 下一步:追加 perf-changelog.yaml 条目,重新检查容量并启动最终全量 sweep(full-sweep-enabled)。此冒烟结果仅为启动/兼容性证据,不构成最终验证。

…p entry

Append the perf-changelog entry for the dsr1-fp8-mi325x-sglang image update
(lmsysorg/sglang v0.5.12-rocm700-mi30x -> v0.5.19-rocm700-mi30x), linking
PR #2949. Historical entries are unchanged.

为 dsr1-fp8-mi325x-sglang 的镜像更新(lmsysorg/sglang v0.5.12-rocm700-mi30x ->
v0.5.19-rocm700-mi30x)追加 perf-changelog 条目,关联 PR #2949。历史条目保持不变。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@Klaud-Cold

Klaud-Cold commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator Author

Final full sweep

  • Head: 453890e61466ad84970a9f7001e1ba179d8678cf (image bump + perf-changelog.yaml entry for dsr1-fp8-mi325x-sglang, PR link [Klaud Cold] Update dsr1-fp8-mi325x-sglang SGLang ROCm image to v0.5.19-rocm700-mi30x / 将 dsr1-fp8-mi325x-sglang 的 SGLang ROCm 镜像更新至 v0.5.19-rocm700-mi30x #2949; validate_perf_changelog.py base 189dde4f -> head 453890e6 passed, historical bytes preserved)
  • Capacity check on mi325x passed at 02:22 UTC; label full-sweep-enabled applied to the draft PR at 02:22:39 UTC.
  • Sweep runs on this head: run 34429203486 (labeled event) and run 34429182346 (push event just before the label; expected to skip or be superseded).
  • Scope: full family, TP8 at concurrency 4/8/16/32/64 with default evals at concurrency 32 and 64 (no trimming, no eval-selection or append-only modifiers).
  • Results: final full sweep passed on the exact PR head. Run 34429203486 concluded success at 02:54 UTC with all 5 benchmark points, both default evals, collectors and compare-results green; power valid on every point; klaud-sweep-manifest records head 453890e6, full-sweep: true, recipe fingerprint f0fb10a4…c85c. The push-event run 34429182346 was cancelled by the concurrency group before any benchmark job started. The "Claude … exceeding the configured maximum of 8" annotations shown on the run belong to a separate Claude review check, not to sweep jobs.
conc tput/GPU (tok/s) base -> new output tput/GPU base -> new median TTFT (s) base -> new median TPOT (s) base -> new median E2EL (s) base -> new
4 224.98 -> 223.94 (-0.5%) 25.03 -> 24.91 (-0.5%) 0.703 -> 0.756 (+7.5%) 0.0184 -> 0.0185 (+0.6%) 17.71 -> 17.79 (+0.4%)
8 379.12 -> 385.84 (+1.8%) 42.62 -> 43.38 (+1.8%) 0.560 -> 0.450 (-19.7%) 0.0215 -> 0.0211 (-1.7%) 21.14 -> 20.94 (-0.9%)
16 616.81 -> 626.19 (+1.5%) 68.35 -> 69.39 (+1.5%) 0.439 -> 0.440 (+0.2%) 0.0268 -> 0.0264 (-1.5%) 25.65 -> 25.34 (-1.2%)
32 849.43 -> 854.62 (+0.6%) 94.99 -> 95.57 (+0.6%) 0.470 -> 0.457 (-2.6%) 0.0396 -> 0.0392 (-0.9%) 37.42 -> 36.98 (-1.2%)
64 1028.90 -> 1035.63 (+0.7%) 114.15 -> 114.89 (+0.7%) 0.485 -> 0.479 (-1.3%) 0.0671 -> 0.0665 (-0.9%) 62.75 -> 62.13 (-1.0%)

Throughput is at parity to slightly better at every point; the only regression is median TTFT at concurrency 4 (+7.5%, 0.053 s), reported without a rejection threshold.

eval conc metric published baseline v0.5.19 final sweep
gsm8k 32 em_strict / em_flexible 0.9530 / 0.9553 0.9560 / 0.9568
gsm8k 64 em_strict / em_flexible 0.9545 / 0.9553 0.9538 / 0.9553

Both evals: n_eff 1319, infrastructure_success: true, no integration error. Differences are within one standard error (about 0.0055).

  • Next step: run the finish verification, which checks matrix/result coverage and marks the PR ready for review. Targeted success, final validation and PR approval are distinct; this comment records final validation only.

最终全量 sweep

  • 提交:453890e61466ad84970a9f7001e1ba179d8678cf(镜像更新 + dsr1-fp8-mi325x-sglangperf-changelog.yaml 条目,PR 链接 [Klaud Cold] Update dsr1-fp8-mi325x-sglang SGLang ROCm image to v0.5.19-rocm700-mi30x / 将 dsr1-fp8-mi325x-sglang 的 SGLang ROCm 镜像更新至 v0.5.19-rocm700-mi30x #2949validate_perf_changelog.py 基线 189dde4f -> 头 453890e6 通过,历史字节保持不变)
  • mi325x 容量检查于 02:22 UTC 通过;02:22:39 UTC 为草稿 PR 添加 full-sweep-enabled 标签。
  • 此提交上的 sweep 运行:run 34429203486(标签事件)与 run 34429182346(标签前的推送事件;预计跳过或被取代)。
  • 范围:完整系列,TP8 并发 4/8/16/32/64,并发 32 与 64 含默认评测(不裁剪,无评测筛选或 append-only 修饰)。
  • 结果:最终全量 sweep 在 PR 精确提交上通过Run 34429203486 于 02:54 UTC 成功结束,5 个基准点、2 个默认评测、收集器与 compare-results 全部通过;每个点功耗有效;klaud-sweep-manifest 记录提交 453890e6full-sweep: true、配方指纹 f0fb10a4…c85c。推送事件的 run 34429182346 在任何基准任务开始前被并发组取消。运行页面上 "Claude … exceeding the configured maximum of 8" 的注解来自独立的 Claude 审阅检查,与 sweep 任务无关。
并发 每 GPU 吞吐 (tok/s) 基线 -> 新 每 GPU 输出吞吐 基线 -> 新 TTFT 中位数 (s) 基线 -> 新 TPOT 中位数 (s) 基线 -> 新 E2EL 中位数 (s) 基线 -> 新
4 224.98 -> 223.94 (-0.5%) 25.03 -> 24.91 (-0.5%) 0.703 -> 0.756 (+7.5%) 0.0184 -> 0.0185 (+0.6%) 17.71 -> 17.79 (+0.4%)
8 379.12 -> 385.84 (+1.8%) 42.62 -> 43.38 (+1.8%) 0.560 -> 0.450 (-19.7%) 0.0215 -> 0.0211 (-1.7%) 21.14 -> 20.94 (-0.9%)
16 616.81 -> 626.19 (+1.5%) 68.35 -> 69.39 (+1.5%) 0.439 -> 0.440 (+0.2%) 0.0268 -> 0.0264 (-1.5%) 25.65 -> 25.34 (-1.2%)
32 849.43 -> 854.62 (+0.6%) 94.99 -> 95.57 (+0.6%) 0.470 -> 0.457 (-2.6%) 0.0396 -> 0.0392 (-0.9%) 37.42 -> 36.98 (-1.2%)
64 1028.90 -> 1035.63 (+0.7%) 114.15 -> 114.89 (+0.7%) 0.485 -> 0.479 (-1.3%) 0.0671 -> 0.0665 (-0.9%) 62.75 -> 62.13 (-1.0%)

各点吞吐持平或略有提升;唯一回退是并发 4 的 TTFT 中位数(+7.5%,0.053 秒),如实报告,不设拒绝阈值。

评测 并发 指标 已发布基线 v0.5.19 最终 sweep
gsm8k 32 em_strict / em_flexible 0.9530 / 0.9553 0.9560 / 0.9568
gsm8k 64 em_strict / em_flexible 0.9545 / 0.9553 0.9538 / 0.9553

两项评测:n_eff 1319,infrastructure_success: true,无集成错误。差异均在一个标准误(约 0.0055)以内。

  • 下一步:执行 finish 校验,核对矩阵/结果覆盖并将 PR 标记为可审阅。定向成功、最终验证与 PR 批准是不同阶段;本评论仅记录最终验证。

@github-actions

Copy link
Copy Markdown
Contributor

@Klaud-Cold
Klaud-Cold marked this pull request as ready for review September 10, 2026 02:57
@Klaud-Cold
Klaud-Cold requested a review from a team September 10, 2026 02:57
@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Klaud Cold: validated. All owned runs are terminal. Repairs: 0. Runs: 34422505103, 34429182346, 34429203486.

The full sweep is verified; this PR remains ready for review.


Klaud Cold:validated。所有自有运行均已结束。修复次数:0。运行:34422505103, 34429182346, 34429203486

完整 sweep 已通过验证;PR 保持就绪,等待审查。

@claude

claude Bot commented Sep 10, 2026

Copy link
Copy Markdown
Contributor

Claude finished @Klaud-Cold's task in 1m 10s —— View job


Review of PR #2949

  • Gather context (diff, changed files)
  • Validate master config change (image accessibility, key prefix)
  • Validate perf-changelog.yaml entry (appended at end, format, no append-only misuse)
  • Run changelog validator (local run not permitted in this session; CI runs validate_perf_changelog.py, and the author reports it passed on this head)
  • Post review

LGTM - no blocking issues found

  • configs/amd-master.yaml:81: image lmsysorg/sglang:v0.5.19-rocm700-mi30x is a public Docker Hub tag; config key prefix dsr1 is valid. Only the image line changed.
  • perf-changelog.yaml:7025-7029: new entry is appended at the very end of the file, references the correct config key and PR link, and carries no append-only flag, so no isolation analysis is required.
  • Final full sweep run 34429203486 passed on the PR head.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good, straightforward config change. Checked the image tag bump in configs/amd-master.yaml for dsr1-fp8-mi325x-sglang and the matching perf-changelog.yaml entry — the description/pr-link format matches prior image-bump entries in the changelog, and the change is appended at the tail with no edits to existing bytes. This recipe is single-node (multinode: false) with no model.container field, so the srt-slurm "model.container must equal image" coupling rule in AGENTS.md doesn't apply here.

Extended reasoning...

Overview

The diff touches exactly two files: configs/amd-master.yaml (a one-line image tag bump for the dsr1-fp8-mi325x-sglang recipe, from sglang v0.5.12 to v0.5.19) and perf-changelog.yaml (a new appended entry documenting that bump). No code, launcher scripts, or benchmark logic is modified.

Security risks

None. No secrets, credentials, or executable logic are introduced. The changed lines are a container image tag string and a YAML changelog record.

Level of scrutiny

This matches the repo's own definition of a mechanical, low-risk change: a version bump with no ambiguity. I verified the perf-changelog.yaml entry is appended strictly at the tail (git diff shows only additions after the last existing entry, no prior bytes touched), satisfying the append-only/byte-sensitive requirement in AGENTS.md. I also checked the "Multi-node srt-slurm changes... For image bumps, model.container must equal image" rule; this recipe has multinode: false and no model.container field in its schema, so that coupling rule is inapplicable. The changelog description/pr-link format is consistent with numerous prior image-bump entries in the file (e.g. PRs #595, #816, #1031, #1041, #1321).

Other factors

No CODEOWNERS restriction is apparent for configs/ or perf-changelog.yaml beyond normal review. No test suite changes are needed for a pure data/config bump. The PR conversation timeline shows no outstanding CHANGES_REQUESTED review or unresolved objection from another reviewer. The bug hunter reported no findings and the exit reason was dry_streak (a valid completion state). Overall this is simple enough that a human does not need to re-verify it.

@adibarra

adibarra commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator

/reuse-sweep-run 34429203486

将 main 合并到 PR #2949,保留已验证的配方并复用完整扫描结果。
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Development

Successfully merging this pull request may close these issues.

2 participants