Skip to content

feat(agentx): retune Kimi-K3 FP4 MI355X ATOM DSpark recipe on _0907 - #2852

Merged
adibarra merged 10 commits into
mainfrom
kimik3-atom-agentx-0907
Sep 10, 2026
Merged

adibarra merged 10 commits into
mainfrom
kimik3-atom-agentx-0907

Conversation

@zejunchen-zejun

@zejunchen-zejun zejunchen-zejun commented Sep 7, 2026 •

Copy link
Copy Markdown
Collaborator

Track ROCm/ATOM#2131, the checked-in Kimi-K3 AgentX recipe, onto image kimi_k3_agentic_0907 and publish the full validated concurrency set [1, 2, 4, 8, 12, 14, 16, 32, 40, 56, 64], replacing [1, 2, 4, 8, 12, 32, 40, 56]. Concurrency 14, 16, and 64 are new points and every existing point is retuned.

The image matters for its triton pin, not only its ATOM revision. The 0903 build shipped triton 3.8.0, which costs roughly 2.1x on prefill TTFT for this workload -- concurrency 1 p50 TTFT goes from ~0.79 s to ~1.5 s -- with decode untouched, and neither the aiter revision nor the ATOM checkout nor the build flavour recovers it. 0907 is back on 3.7.x. The regression is silent, so the launcher now records the triton version and warns when it is not 3.7.x.

Serving is three bands. Concurrency 1-4 is the latency floor: GPU-resident, dcp 1, 7 draft tokens at golden AL 3.84. Concurrency 8 through 16 shard the decode KV read with dcp 8, back the paged KV with the LMCache DRAM tier, run 3 draft tokens at golden AL 3.00, and rebuild the KDA recurrent state in place with ReplaySSM. Concurrency 32 and up serve without a draft model. Both acceptance lengths come from the committed golden curve in golden_al_distribution/kimik3_dspark_probabilistic_sample_method_block_rejection_sample_method.yaml (7 -> 3.84, 3 -> 3.00); evaluations drop the flag and use real acceptance.

Full CUDA graphs are captured over the dense range [2 .. window x (1 + draft tokens)] via --level 3, --cudagraph-mode FULL, and an explicit --cudagraph-capture-sizes. The window is 2 x concurrency, and the (1 + draft) factor covers the verify step of a DSpark round, which submits one row per draft token on top of the accepted token. Concurrency 14 runs the concurrency 16 server verbatim, window pinned at 32, and only moves the client: deriving the window from 2 x concurrency there gives 28 and the server dies during graph warmup.

The CPU state-offload tier is gone. ReplaySSM rebuilds the KDA state from the in-GPU checkpoint ring on every point that keeps a draft model, so OFFLOAD_STATE and its companions and the --state-checkpoint-slots override all drop out and the whole per-rank CPU budget goes to the paged KV. LMCACHE_MAX_LOCAL_CPU_SIZE is per rank, so the aggregate TOTAL_CPU_DRAM_GB is still divided by TP as the agentic README requires: dram-utilization 0.343 lands 128 GB/rank for concurrency 8 through 40, and 0.513 lands 192 GB/rank for concurrency 56 and 64.

Two failure modes the recipe documents are now closed in the launcher rather than left to the image. The Inferact/Kimi-K3-DSpark draft is staged into the shared Hugging Face cache before the server starts, because an uncached repo id makes every rank pull the same 7 GB checkpoint at once and the log simply stops after loading the drafter -- measured at ~0.7 MB/s on one cluster, about three hours, with every GPU at 0%. And the server registers under --served-model-name $MODEL, which is the name the AgentX client asks for on the wire.

将 MI355X 上 Kimi-K3 的 ATOM AgentX 提交对齐到 ROCm/ATOM#2131 中已合入的 recipe,镜像换成 kimi_k3_agentic_0907,并发点从 [1, 2, 4, 8, 12, 32, 40, 56] 换成 recipe 完整验证过的 [1, 2, 4, 8, 12, 14, 16, 32, 40, 56, 64]。14、16、64 为新增点,其余各点逐点重调。

换镜像的关键不只是 ATOM 版本,还有 triton 的版本。0903 镜像带的是 triton
3.8.0,在该负载上 prefill TTFT 大约变差 2.1 倍——并发 1 的 p50 TTFT 从约 0.79 秒涨到约 1.5 秒,而 decode 完全不受影响;换 aiter、换 ATOM checkout、 换构建方式都救不回来。0907 回到 3.7.x。这个退化是静默的,因此脚本现在会把
triton 版本打进日志,并在不是 3.7.x 时告警。

服务分三档。并发 1-4 是时延下界:全部驻留 GPU,dcp 1,草稿 7 token 对应
golden AL 3.84。并发 8 到 16 用 dcp 8 把 decode 的 KV 读取分摊到 8 张卡, 由 LMCache DRAM 层承接分页 KV,草稿 3 token 对应 golden AL 3.00,并用 ReplaySSM 就地重建 KDA 循环状态。并发 32 及以上不加载草稿模型。两个接受长度
均取自仓库内已提交的 golden 曲线(7 -> 3.84,3 -> 3.00);评测不传该参数, 使用真实接受率。

通过 --level 3、--cudagraph-mode FULL 和显式的 --cudagraph-capture-sizes, 在 [2 .. 窗口 x (1 + 草稿 token 数)] 这个稠密区间上抓取完整 CUDA graph。 窗口取 2 x 并发,(1 + 草稿) 这个系数覆盖 DSpark 一轮中的 verify 步——它在 被接受的 token 之上,每个草稿 token 再提交一行。并发 14 完整复用并发 16 的
server 配置(窗口固定为 32),只改客户端并发:在这里按 2 x 并发推出 28,
server 会在 graph warmup 阶段直接挂掉。

CPU state offload 层被移除。凡是保留草稿模型的并发点,KDA 状态都由 ReplaySSM 从 GPU 内的 checkpoint ring 就地重建,因此 OFFLOAD_STATE 及其配套变量、 --state-checkpoint-slots 覆盖项全部去掉,每 rank 的 CPU 预算整块留给分页 KV。 LMCACHE_MAX_LOCAL_CPU_SIZE 仍是每 rank 设置,所以按 agentic README 的要求 把聚合预算 TOTAL_CPU_DRAM_GB 除以 TP:dram-utilization 0.343 在并发 8 到 40 给到每 rank 128 GB,0.513 在并发 56 和 64 给到每 rank 192 GB。

recipe 里记录的两个失败模式现在由脚本兜住,而不是依赖镜像。草稿模型
Inferact/Kimi-K3-DSpark 在起 server 之前先落进共享 HF 缓存:仓库 id 未命中 缓存时,八个 rank 会同时去拉同一份 7 GB 权重,日志在加载 drafter 之后就停住,
某个集群上实测约 0.7 MB/s、约三小时,期间所有 GPU 都是 0%。另外 server 现在 用 --served-model-name $MODEL 注册,这正是 AgentX 客户端在协议上请求的名字。


Note

Medium Risk
Large benchmark launcher and matrix changes affect published AgentX numbers and run-to-run KV sizing; backward-compatible GPU-clean helper only tightens gating where callers pass a threshold.

Overview
Retunes the MI355X Kimi-K3 FP4 ATOM AgentX recipe onto image kimi_k3_agentic_0907 and expands the published concurrency set to [1, 2, 4, 8, 12, 14, 16, 32, 40, 56, 64] (adds 14, 16, 64; retunes the rest).

The launcher is reorganized into interactive / mid / throughput bands with per-CONC server knobs: ReplaySSM on mid-band draft points, no draft from 32 up, full CUDA graphs (--level 3, explicit --cudagraph-capture-sizes scaled by draft depth), and concurrency 14 reusing the concurrency 16 server (CUDA-graph window pinned at 32). CPU KDA state offload and --state-checkpoint-slots are removed; the full per-rank LMCache budget goes to paged KV, with 128 GB/rank (dram 0.343) through 40 and 192 GB/rank (0.513) at 56/64.

Operational fixes in the script: unset leaked HTTP proxy vars from the 0907 image, wait_for_amd_gpu_clean 1 before launch (shared benchmark_lib.sh helper now accepts a VRAM % threshold), pre-cache Inferact/Kimi-K3-DSpark, --served-model-name $MODEL, explicit NUMA mapping for LMCache, and AITER_REUSE_IDENTICAL_COMM_GROUPS=1 only at 56/64. configs/amd-master.yaml and perf-changelog.yaml track the image, search space, and dram blocks.

Reviewed by Cursor Bugbot for commit 8af6126. Bugbot is set up for automated code reviews on this repo. Configure here.

@github-actions

github-actions Bot commented Sep 7, 2026

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have。

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

Track ROCm/ATOM#2131, the checked-in Kimi-K3 AgentX recipe, onto image
kimi_k3_agentic_0907 and publish the full validated concurrency set
[1, 2, 4, 8, 12, 14, 16, 32, 40, 56, 64], replacing [1, 2, 4, 8, 12, 32, 40, 56].
Concurrency 14, 16, and 64 are new points and every existing point is retuned.

Serving is three bands. Concurrency 1-4 is the latency floor: GPU-resident,
dcp 1, 7 draft tokens at golden AL 3.84. Concurrency 8 through 16 shard the
decode KV read with dcp 8, back the paged KV with the LMCache DRAM tier, run
3 draft tokens at golden AL 3.00, and rebuild the KDA recurrent state in place
with ReplaySSM. Concurrency 32 and up serve without a draft model. Both
acceptance lengths come from the committed golden curve in
golden_al_distribution/kimik3_dspark_probabilistic_sample_method_block_rejection_sample_method.yaml
(7 -> 3.84, 3 -> 3.00); evaluations drop the flag and use real acceptance.

Full CUDA graphs are captured over the dense range [2 .. window x (1 + draft
tokens)] via --level 3, --cudagraph-mode FULL, and an explicit
--cudagraph-capture-sizes. The window is 2 x concurrency, and the (1 + draft)
factor covers the verify step of a DSpark round, which submits one row per
draft token on top of the accepted token. Concurrency 14 runs the concurrency
16 server verbatim, window pinned at 32, and only moves the client: deriving
the window from 2 x concurrency there gives 28 and the server dies during
graph warmup.

The CPU state-offload tier is gone. ReplaySSM rebuilds the KDA state from the
in-GPU checkpoint ring on every point that keeps a draft model, so OFFLOAD_STATE
and its companions and the --state-checkpoint-slots override all drop out and
the whole per-rank CPU budget goes to the paged KV. LMCACHE_MAX_LOCAL_CPU_SIZE
is per rank, so the aggregate TOTAL_CPU_DRAM_GB is still divided by TP as the
agentic README requires: dram-utilization 0.343 lands 128 GB/rank for
concurrency 8 through 40, and 0.513 lands 192 GB/rank for concurrency 56 and 64.
The eight ranks are pinned across the two sockets explicitly, and
AITER_REUSE_IDENTICAL_COMM_GROUPS is set only at concurrency 56 and 64.

Two failure modes the recipe documents are now closed in the launcher rather
than left to the image. The Inferact/Kimi-K3-DSpark draft is staged into the
shared Hugging Face cache before the server starts, because an uncached repo id
makes every rank pull the same 7 GB checkpoint at once and the log simply stops
after loading the drafter -- measured at ~0.7 MB/s on one cluster, about three
hours, with every GPU at 0%. And the server registers under
--served-model-name $MODEL, which is the name the AgentX client asks for on the
wire.

将 MI355X 上 Kimi-K3 的 ATOM AgentX 提交对齐到 ROCm/ATOM#2131 中已合入的
recipe,镜像换成 kimi_k3_agentic_0907,并发点从 [1, 2, 4, 8, 12, 32, 40, 56]
换成 recipe 完整验证过的 [1, 2, 4, 8, 12, 14, 16, 32, 40, 56, 64]。14、16、64
为新增点,其余各点逐点重调。

服务分三档。并发 1-4 是时延下界:全部驻留 GPU,dcp 1,草稿 7 token 对应
golden AL 3.84。并发 8 到 16 用 dcp 8 把 decode 的 KV 读取分摊到 8 张卡,
由 LMCache DRAM 层承接分页 KV,草稿 3 token 对应 golden AL 3.00,并用
ReplaySSM 就地重建 KDA 循环状态。并发 32 及以上不加载草稿模型。两个接受长度
均取自仓库内已提交的 golden 曲线(7 -> 3.84,3 -> 3.00);评测不传该参数,
使用真实接受率。

通过 --level 3、--cudagraph-mode FULL 和显式的 --cudagraph-capture-sizes,
在 [2 .. 窗口 x (1 + 草稿 token 数)] 这个稠密区间上抓取完整 CUDA graph。
窗口取 2 x 并发,(1 + 草稿) 这个系数覆盖 DSpark 一轮中的 verify 步——它在
被接受的 token 之上,每个草稿 token 再提交一行。并发 14 完整复用并发 16 的
server 配置(窗口固定为 32),只改客户端并发:在这里按 2 x 并发推出 28,
server 会在 graph warmup 阶段直接挂掉。

CPU state offload 层被移除。凡是保留草稿模型的并发点,KDA 状态都由 ReplaySSM
从 GPU 内的 checkpoint ring 就地重建,因此 OFFLOAD_STATE 及其配套变量、
--state-checkpoint-slots 覆盖项全部去掉,每 rank 的 CPU 预算整块留给分页 KV。
LMCACHE_MAX_LOCAL_CPU_SIZE 仍是每 rank 设置,所以按 agentic README 的要求
把聚合预算 TOTAL_CPU_DRAM_GB 除以 TP:dram-utilization 0.343 在并发 8 到 40
给到每 rank 128 GB,0.513 在并发 56 和 64 给到每 rank 192 GB。八个 rank 按
显式映射绑定到两个 socket;AITER_REUSE_IDENTICAL_COMM_GROUPS 仅在并发 56
和 64 打开。

recipe 里记录的两个失败模式现在由脚本兜住,而不是依赖镜像。草稿模型
Inferact/Kimi-K3-DSpark 在起 server 之前先落进共享 HF 缓存:仓库 id 未命中
缓存时,八个 rank 会同时去拉同一份 7 GB 权重,日志在加载 drafter 之后就停住,
某个集群上实测约 0.7 MB/s、约三小时,期间所有 GPU 都是 0%。另外 server 现在
用 --served-model-name $MODEL 注册,这正是 AgentX 客户端在协议上请求的名字。

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@seungrokj seungrokj added AMD full-sweep-enabled agentx AgentX benchmarks, recipes, and infrastructure labels Sep 7, 2026
@zejunchen-zejun
zejunchen-zejun marked this pull request as ready for review September 7, 2026 09:22
@zejunchen-zejun
zejunchen-zejun requested a review from a team September 7, 2026 09:22
@claude

claude Bot commented Sep 7, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @zejunchen-zejun's task in 3m 2s —— View job


Review

  • Gather context (diff vs origin/main, read changed files)
  • Review benchmark script kimik3_fp4_mi355x_atom_mtp.sh
  • Review configs/amd-master.yaml (image accessibility, concurrency set, key format)
  • Review perf-changelog.yaml (entry appended at end, content)
  • Post review feedback

Summary: One 🟡 WARNING (inline): the PR description claims the launcher "records the triton version and warns when it is not 3.7.x", but no such check exists in the diff — since the silent triton 3.8.0 TTFT regression is the stated reason for the image swap, either add the guard or correct the description. Everything else verified clean: the perf-changelog.yaml entry is correctly appended at the end, the image is a public rocm/atom-dev tag, config conc-lists exactly match the script's case arms, concurrency 14 does run the concurrency-16 server verbatim (pinned window 32 → identical GRAPH_MAX=128 and flags), the hf download retry loop is errexit-safe, and the per-rank LMCache budgets (0.343 → 128 GB/rank, 0.513 → 192 GB/rank at TP8) are arithmetically consistent between config, script, and changelog. The --served-model-name $MODEL claim also checks out against benchmark_lib.sh:3294.

Comment thread benchmarks/single_node/agentic/kimik3_fp4_mi355x_atom_mtp.sh

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this PR and didn't find any bugs. Because it's a large, numerically dense retune of a production benchmark recipe (concurrency-band restructuring, CUDA-graph capture-size derivation, LMCache/NUMA/state-offload changes, and the matching config/changelog updates), a human look would still be worthwhile to sanity-check the tuning choices before the results are published.

What was reviewed: the three-band knob table and CUDA-graph capture-size math in kimik3_fp4_mi355x_atom_mtp.sh, the ReplaySSM/state-offload removal and LMCACHE_MAX_LOCAL_CPU_SIZE resizing, and the corresponding dram-utilization/conc-list restructuring in configs/amd-master.yaml. Checked that the perf-changelog.yaml entry is appended at the tail without altering prior bytes, and that the config's image tag matches the recipe's new kimi_k3_agentic_0907 image.

Extended reasoning...

Overview

This PR retunes a single-node MI355X AgentX benchmark recipe (Kimi-K3 FP4 via ATOM/DSpark): a new image tag, an expanded/retuned concurrency matrix, a reorganized three-band per-concurrency knob table in the launcher script, removal of the CPU state-offload tier in favor of ReplaySSM, new NUMA-pinning and AITER comm-group env vars, explicit CUDA-graph capture-size derivation, a retrying model-download step, and matching updates to configs/amd-master.yaml and perf-changelog.yaml.

Security risks

None identified. No new external inputs, no auth/crypto/permission-boundary code. The hf download retry loop and NUMA pinning are benchmark-infra concerns, not security-sensitive paths.

Level of scrutiny

Config/benchmark-tuning changes are normally low-scrutiny, but this one is unusually dense: it touches interacting numeric derivations (window = 2CONC, GRAPH_MAX = window(1+draft), per-rank CPU sizing via TOTAL_CPU_DRAM_GB/TP, dram-utilization aggregates) across two files that must stay consistent, plus a special-cased concurrency-14 point that intentionally overrides the general derivation. That combination of hand-computed constants and one deliberate exception is exactly the kind of change where a human's sanity check adds value even though the automated pass found nothing conclusive.

Other factors

I independently verified the things the bug hunter's ruled-out list flagged: the config's dram-utilization aggregates (1028 GB / 8 ≈ 128 GB/rank, 1538 GB / 8 ≈ 192 GB/rank) match the launcher's per-rank sizing comments, the conc-list splits in amd-master.yaml align with the script's case statement, and the perf-changelog entry is appended at the tail without disturbing prior bytes. No test suite covers these benchmark scripts (they run against real clusters), so review is the primary safety net here.

@github-actions

github-actions Bot commented Sep 7, 2026

Copy link
Copy Markdown
Contributor

@github-actions

github-actions Bot commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

seungrokj and others added 3 commits September 8, 2026 09:43
The kimi_k3_agentic_0907 image was built with podman on a host behind a local
squid, and the build environment leaked into the image config. HTTP_PROXY,
HTTPS_PROXY, http_proxy, and https_proxy are all baked in as
http://127.0.0.1:3128, with NO_PROXY covering only localhost. The 0821 and 0903
images carried none of them.

Nothing listens on that port inside the container on a benchmark node, so every
outbound request the launcher makes before the server starts -- the uv bootstrap
and aiperf install in install_agentic_deps, the trace-corpus fetch, and both hf
downloads -- would reach a dead proxy. Clear the four variables and keep the
loopback exemption, which leaves the AIPerf client and the /metrics scrape
unaffected either way.

kimi_k3_agentic_0907 镜像是在一台走本地 squid 的机器上用 podman 构建的,
构建环境泄漏进了镜像配置:HTTP_PROXY、HTTPS_PROXY、http_proxy、https_proxy
四个变量都被写死成 http://127.0.0.1:3128,而 NO_PROXY 只覆盖 localhost。
0821 和 0903 镜像都没有这些变量。

在 benchmark 节点上,容器内没有任何进程监听该端口,因此起 server 之前脚本发出的
所有外部请求——install_agentic_deps 里的 uv 引导与 aiperf 安装、trace 语料拉取、
以及两次 hf download——都会打到一个不存在的代理上。这里清掉这四个变量,并保留
回环地址的豁免,AIPerf 客户端与 /metrics 抓取不受影响。

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Fold seungrokj's unset in e994743 into the block above. Same four variables,
but the launcher's first outbound request is the hf download of the target
checkpoint, which runs before rocm-smi and so before that unset; clearing them
after it leaves that call pointed at the dead proxy. One unset, placed before
every network call, plus the loopback NO_PROXY exemption.

把 e994743(seungrokj)中的 unset 合并进上面的代码块。变量完全相同,但脚本的
第一次外部请求是拉取目标 checkpoint 的 hf download,它在 rocm-smi 之前执行,
也就在原先那处 unset 之前;在其之后再清理,这一次调用仍会打到不存在的代理上。
现在只保留一处 unset,置于所有网络调用之前,并保留回环地址的 NO_PROXY 豁免。

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@github-actions

github-actions Bot commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

@github-actions

github-actions Bot commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

1 similar comment
@github-actions

github-actions Bot commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

@seungrokj seungrokj left a comment •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Will review again.

@seungrokj
seungrokj self-requested a review September 8, 2026 07:49
@SemiAnalysisAI SemiAnalysisAI deleted a comment from Klaud-Cold Sep 8, 2026
@github-actions

github-actions Bot commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

@github-actions

github-actions Bot commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

1 similar comment
@github-actions

github-actions Bot commented Sep 9, 2026

Copy link
Copy Markdown
Contributor

seungrokj and others added 2 commits September 9, 2026 21:06
…rift

The default 10% GPU-clean gate admits ~28.8 GB of prior-job residual on the
288 GB parts. ATOM sizes KV from torch.cuda.mem_get_info() at launch, so that
residual is counted as used, folded into non_torch, and subtracted from
available_for_kv -- drifting the KV pool several GB between identical reruns
(non_torch 25.8 vs 31.8 GB on two c4 runs). Add an optional threshold to
wait_for_amd_gpu_clean (default unchanged) and gate the pre-server check at 1%.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Sep 9, 2026

Copy link
Copy Markdown
Contributor

@seungrokj

Copy link
Copy Markdown
Collaborator

/reuse-sweep-run

@seungrokj seungrokj left a comment •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

As a PR reviewer and CODEOWNER, I have reviewed this and have:

  • Verified that as of the moment of typing this, this is the latest version of PR_REVIEW_CHECKLIST.md
  • Verified that the general code quality meets the InferenceX standard and does not make the code quality any worse.
  • Verified that this PR has passed PR validation. Please link to GitHub Action workflow that shows this.
  • Verified that this PR passes evals. Please link to GitHub Action workflow that shows this.
  • Verified that speculative decoding PRs uses chat templates to align the AL distribution to real world
  • For agentic workloads: verified that speculative-decoding configs (EAGLE / MTP / draft models) run with simulated synthetic acceptance, with the acceptance-length value taken from the committed golden AL curve in golden_al_distribution/ for that model, thinking mode, and draft length. A submission may choose any supported draft length, but it may not substitute a different acceptance target.
  • Verified against the current MODELS.md that this PR does not submit a deprecated model, scenario, or model-scenario combination.
  • Verified that the model architecture isn't changed with benchmark hacks like using --hf-overrides to skipping indexer for every x layers on models that don't natively support this. As a general rule, we won't accept optimizations that reduces the number of model architecture FLOPs. Anything that makes that same computation run faster is fair game; FLOPs at lower precisions is fine, given that the config passes private evals. As an general north star princple, we should only use optimizations which is used in production by customers that care about accuracy
  • If an company claims that they support vLLM/SGLang as first class LLM inference engines on their hardware, I have verified that the respective vLLM submission made using upstream https://hub.docker.com/u/vllm docker repo, upstream SGLang https://hub.docker.com/u/lmsysorg docker repo. The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet as supported by vLLM/SGLang community maintainers
  • If an company claims that they support vLLM/SGLang as first class upstream in-tree LLM inference engines on their hardware, I have have verified that the respective vLLM/SGLang submission has been made before additional frameworks (TRT-LLM, ATOM, etc.). The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet.
  • Verified that every single-node vLLM/SGLang recipe in this PR is documented in the official vLLM recipes and/or the SGLang cookbook:
    • I linked the corresponding upstream PR in the vLLM recipe repo or SGLang repo and verified that it is MERGED before this InferenceX PR merges. An opened, draft, or closed-without-merge upstream PR does not satisfy this requirement. If the matching recipe was already published, I linked the published recipe/cookbook page in the additional detail section below.
  • Verified that this PR does not patch the inference engine or serving stack — the pinned image must run as shipped. This covers .patch files / git apply / patch, inline patches embedded in benchmark scripts (e.g. a python3/sed heredoc that rewrites installed engine sources before serving), in-place edits of site-packages, monkey-patching, overwriting container files, and installing forked/rebuilt engine wheels on top of the pinned image. The only exception is a patch covered by a filled-out waiver at docs/waiver/<PR_NUMBER>.md — named after the PR that introduces the patch and filed in that same PR, stating what is patched, why the unmodified upstream image cannot run this benchmark, the upstream PR/issue link, and the removal plan — which I have linked below in the additional detail section.
  • If this PR uses append-only: true, verified that it only adds generated points or recipe variants inside a selected existing config/scenario and existing same-image visual curve: every previously generated point remains present with the same recipe, no prior point is removed or rerun, and every benchmark-affecting change in the complete diff can affect only the corresponding newly appended points (never an existing point), regardless of which file contains it.
  • If any of the above criteria cannot reasonably be satisfied, I have provided additional reasoning below.

Additional detail section:

Signed: seungrokj

@Klaud-Cold

Copy link
Copy Markdown
Collaborator

✅✅✅ Verdict: PASS ✅✅✅

✅ Check 0 (CODEOWNER): PASS — @seungrokj is a named owner of configs/amd-master.yaml; the other three changed files fall under the * catch-all, which any recognized CODEOWNER satisfies.
✅ Check 1 (passing sweep on in-PR commit): PASS — run 34349642428 on in-PR commit 8a7bcceb has all 11 agentic / and all 11 agentic eval / jobs success (none skipped). Head 5a4cbc8c is a merge of main with the recipe script, benchmark_lib.sh, and the kimik3-fp4-mi355x-atom-agentic-mtp entry byte-identical to 8a7bcceb.
✅ Check 2 (evals pass): PASS — agg_eval_all.json from that run has 11/11 GSM8K results, em_strict 0.958–0.967 (n=1319 each) against the 0.90 bar, all on image rocm/atom-dev:...kimi_k3_agentic_0907, matching this PR's config.
➖ Check 3 (recipe link): N/A — ATOM-framework submission; the recipe-link requirement covers single-node vLLM/SGLang recipes only. Informational: the sign-off links the published ROCm/ATOM Kimi-K3 recipe, which matches model, MI355 TP8, fp8 KV, block-size 128, and the identical ptpc_fp8 online-quant config; the AgentX-specific args (DSpark 7/3, DCP8, LMCache, FULL cudagraph, synthetic AL) match ROCm/ATOM#2131, which is still OPEN.
✅ Check 4 (reuse command): PASS — /reuse-sweep-run posted by seungrokj (COLLABORATOR) on 2026-09-10.
✅ Check 5 (latest template): PASS — all 15 items of the current PR_REVIEW_CHECKLIST.md are present and checked.
✅ Check 6 (upstream image / engine-first): PASS — framework: atom, so the upstream-image rule does not apply; the vLLM sibling kimik3-fp4-mi355x-vllm-agentic-mtp already exists for the same model-prefix and runner: cluster:mi355x-amds.
✅ Check 7 (deprecated models): PASS — kimik3 Agentic coding (DSpark arm) is active in MODELS.md as of 2026-09-10.
✅ Check 8 (architecture hacks): PASS — no --hf-overrides/model-override args; ptpc_fp8 online quant and --state-checkpoint-interval-tokens -1 change precision/checkpoint placement, not architecture FLOPs.
✅ Check 9 (spec-decode chat template): PASS — the replay adds --apply-chat-template and the eval runs local-chat-completions --apply_chat_template.
✅ Check 10 (engine patches): PASS — no patch files, heredoc edits, site-packages rewrites, or engine wheel installs; the script only clears leaked proxy env vars and pre-stages the draft weights.
✅ Check 11 (agentic golden AL): PASS — --spec-decode-acceptance-length 3.84 at 7 draft tokens and 3.00 at 3 draft tokens match golden_al_distribution/kimik3_dspark_probabilistic_sample_method_block_rejection_sample_method.yaml; evals drop the flag and use real acceptance.
➖ Check 12 (append-only): N/A — the new perf-changelog.yaml entry does not set append-only: true.

@seungrokj

Copy link
Copy Markdown
Collaborator

@Oseltamivir @adibarra can you plz review this ?

@adibarra
adibarra merged commit 56babc8 into main Sep 10, 2026
7 of 8 checks passed
@adibarra
adibarra deleted the kimik3-atom-agentx-0907 branch September 10, 2026 15:59
wufann pushed a commit to wufann/InferenceX that referenced this pull request Sep 20, 2026
…emiAnalysisAI#2852)

* feat(agentx): retune Kimi-K3 FP4 MI355X ATOM DSpark recipe on _0907

Track ROCm/ATOM#2131, the checked-in Kimi-K3 AgentX recipe, onto image
kimi_k3_agentic_0907 and publish the full validated concurrency set
[1, 2, 4, 8, 12, 14, 16, 32, 40, 56, 64], replacing [1, 2, 4, 8, 12, 32, 40, 56].
Concurrency 14, 16, and 64 are new points and every existing point is retuned.

Serving is three bands. Concurrency 1-4 is the latency floor: GPU-resident,
dcp 1, 7 draft tokens at golden AL 3.84. Concurrency 8 through 16 shard the
decode KV read with dcp 8, back the paged KV with the LMCache DRAM tier, run
3 draft tokens at golden AL 3.00, and rebuild the KDA recurrent state in place
with ReplaySSM. Concurrency 32 and up serve without a draft model. Both
acceptance lengths come from the committed golden curve in
golden_al_distribution/kimik3_dspark_probabilistic_sample_method_block_rejection_sample_method.yaml
(7 -> 3.84, 3 -> 3.00); evaluations drop the flag and use real acceptance.

Full CUDA graphs are captured over the dense range [2 .. window x (1 + draft
tokens)] via --level 3, --cudagraph-mode FULL, and an explicit
--cudagraph-capture-sizes. The window is 2 x concurrency, and the (1 + draft)
factor covers the verify step of a DSpark round, which submits one row per
draft token on top of the accepted token. Concurrency 14 runs the concurrency
16 server verbatim, window pinned at 32, and only moves the client: deriving
the window from 2 x concurrency there gives 28 and the server dies during
graph warmup.

The CPU state-offload tier is gone. ReplaySSM rebuilds the KDA state from the
in-GPU checkpoint ring on every point that keeps a draft model, so OFFLOAD_STATE
and its companions and the --state-checkpoint-slots override all drop out and
the whole per-rank CPU budget goes to the paged KV. LMCACHE_MAX_LOCAL_CPU_SIZE
is per rank, so the aggregate TOTAL_CPU_DRAM_GB is still divided by TP as the
agentic README requires: dram-utilization 0.343 lands 128 GB/rank for
concurrency 8 through 40, and 0.513 lands 192 GB/rank for concurrency 56 and 64.
The eight ranks are pinned across the two sockets explicitly, and
AITER_REUSE_IDENTICAL_COMM_GROUPS is set only at concurrency 56 and 64.

Two failure modes the recipe documents are now closed in the launcher rather
than left to the image. The Inferact/Kimi-K3-DSpark draft is staged into the
shared Hugging Face cache before the server starts, because an uncached repo id
makes every rank pull the same 7 GB checkpoint at once and the log simply stops
after loading the drafter -- measured at ~0.7 MB/s on one cluster, about three
hours, with every GPU at 0%. And the server registers under
--served-model-name $MODEL, which is the name the AgentX client asks for on the
wire.

将 MI355X 上 Kimi-K3 的 ATOM AgentX 提交对齐到 ROCm/ATOM#2131 中已合入的
recipe,镜像换成 kimi_k3_agentic_0907,并发点从 [1, 2, 4, 8, 12, 32, 40, 56]
换成 recipe 完整验证过的 [1, 2, 4, 8, 12, 14, 16, 32, 40, 56, 64]。14、16、64
为新增点,其余各点逐点重调。

服务分三档。并发 1-4 是时延下界:全部驻留 GPU,dcp 1,草稿 7 token 对应
golden AL 3.84。并发 8 到 16 用 dcp 8 把 decode 的 KV 读取分摊到 8 张卡,
由 LMCache DRAM 层承接分页 KV,草稿 3 token 对应 golden AL 3.00,并用
ReplaySSM 就地重建 KDA 循环状态。并发 32 及以上不加载草稿模型。两个接受长度
均取自仓库内已提交的 golden 曲线(7 -> 3.84,3 -> 3.00);评测不传该参数,
使用真实接受率。

通过 --level 3、--cudagraph-mode FULL 和显式的 --cudagraph-capture-sizes,
在 [2 .. 窗口 x (1 + 草稿 token 数)] 这个稠密区间上抓取完整 CUDA graph。
窗口取 2 x 并发,(1 + 草稿) 这个系数覆盖 DSpark 一轮中的 verify 步——它在
被接受的 token 之上,每个草稿 token 再提交一行。并发 14 完整复用并发 16 的
server 配置(窗口固定为 32),只改客户端并发:在这里按 2 x 并发推出 28,
server 会在 graph warmup 阶段直接挂掉。

CPU state offload 层被移除。凡是保留草稿模型的并发点,KDA 状态都由 ReplaySSM
从 GPU 内的 checkpoint ring 就地重建,因此 OFFLOAD_STATE 及其配套变量、
--state-checkpoint-slots 覆盖项全部去掉,每 rank 的 CPU 预算整块留给分页 KV。
LMCACHE_MAX_LOCAL_CPU_SIZE 仍是每 rank 设置,所以按 agentic README 的要求
把聚合预算 TOTAL_CPU_DRAM_GB 除以 TP:dram-utilization 0.343 在并发 8 到 40
给到每 rank 128 GB,0.513 在并发 56 和 64 给到每 rank 192 GB。八个 rank 按
显式映射绑定到两个 socket;AITER_REUSE_IDENTICAL_COMM_GROUPS 仅在并发 56
和 64 打开。

recipe 里记录的两个失败模式现在由脚本兜住,而不是依赖镜像。草稿模型
Inferact/Kimi-K3-DSpark 在起 server 之前先落进共享 HF 缓存:仓库 id 未命中
缓存时,八个 rank 会同时去拉同一份 7 GB 权重,日志在加载 drafter 之后就停住,
某个集群上实测约 0.7 MB/s、约三小时,期间所有 GPU 都是 0%。另外 server 现在
用 --served-model-name $MODEL 注册,这正是 AgentX 客户端在协议上请求的名字。

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Update kimik3_fp4_mi355x_atom_mtp.sh

* fix(agentx): clear the proxy env the 0907 ATOM image bakes in

The kimi_k3_agentic_0907 image was built with podman on a host behind a local
squid, and the build environment leaked into the image config. HTTP_PROXY,
HTTPS_PROXY, http_proxy, and https_proxy are all baked in as
http://127.0.0.1:3128, with NO_PROXY covering only localhost. The 0821 and 0903
images carried none of them.

Nothing listens on that port inside the container on a benchmark node, so every
outbound request the launcher makes before the server starts -- the uv bootstrap
and aiperf install in install_agentic_deps, the trace-corpus fetch, and both hf
downloads -- would reach a dead proxy. Clear the four variables and keep the
loopback exemption, which leaves the AIPerf client and the /metrics scrape
unaffected either way.

kimi_k3_agentic_0907 镜像是在一台走本地 squid 的机器上用 podman 构建的,
构建环境泄漏进了镜像配置:HTTP_PROXY、HTTPS_PROXY、http_proxy、https_proxy
四个变量都被写死成 http://127.0.0.1:3128,而 NO_PROXY 只覆盖 localhost。
0821 和 0903 镜像都没有这些变量。

在 benchmark 节点上,容器内没有任何进程监听该端口,因此起 server 之前脚本发出的
所有外部请求——install_agentic_deps 里的 uv 引导与 aiperf 安装、trace 语料拉取、
以及两次 hf download——都会打到一个不存在的代理上。这里清掉这四个变量,并保留
回环地址的豁免,AIPerf 客户端与 /metrics 抓取不受影响。

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(agentx): move the 0907 proxy unset ahead of the model download

Fold seungrokj's unset in e994743 into the block above. Same four variables,
but the launcher's first outbound request is the hf download of the target
checkpoint, which runs before rocm-smi and so before that unset; clearing them
after it leaves that call pointed at the dead proxy. One unset, placed before
every network call, plus the loopback NO_PROXY exemption.

把 e994743(seungrokj)中的 unset 合并进上面的代码块。变量完全相同,但脚本的
第一次外部请求是拉取目标 checkpoint 的 hf download,它在 rocm-smi 之前执行,
也就在原先那处 unset 之前;在其之后再清理,这一次调用仍会打到不存在的代理上。
现在只保留一处 unset,置于所有网络调用之前,并保留回环地址的 NO_PROXY 豁免。

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(agentx): gate K3 ATOM launch on near-empty VRAM to stop KV-pool drift

The default 10% GPU-clean gate admits ~28.8 GB of prior-job residual on the
288 GB parts. ATOM sizes KV from torch.cuda.mem_get_info() at launch, so that
residual is counted as used, folded into non_torch, and subtracted from
available_for_kv -- drifting the KV pool several GB between identical reruns
(non_torch 25.8 vs 31.8 GB on two c4 runs). Add an optional threshold to
wait_for_amd_gpu_clean (default unchanged) and gate the pre-server check at 1%.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: seungrokj <144636725+seungrokj@users.noreply.github.com>
Co-authored-by: seungrokj <seungrok.jung@amd.com>
Co-authored-by: Alec Ibarra <93070681+adibarra@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

agentx AgentX benchmarks, recipes, and infrastructure AMD full-sweep-enabled

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

4 participants