feat(agentx): retune Kimi-K3 FP4 MI355X ATOM DSpark recipe on _0907 - #2852
Conversation
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
Track ROCm/ATOM#2131, the checked-in Kimi-K3 AgentX recipe, onto image kimi_k3_agentic_0907 and publish the full validated concurrency set [1, 2, 4, 8, 12, 14, 16, 32, 40, 56, 64], replacing [1, 2, 4, 8, 12, 32, 40, 56]. Concurrency 14, 16, and 64 are new points and every existing point is retuned. Serving is three bands. Concurrency 1-4 is the latency floor: GPU-resident, dcp 1, 7 draft tokens at golden AL 3.84. Concurrency 8 through 16 shard the decode KV read with dcp 8, back the paged KV with the LMCache DRAM tier, run 3 draft tokens at golden AL 3.00, and rebuild the KDA recurrent state in place with ReplaySSM. Concurrency 32 and up serve without a draft model. Both acceptance lengths come from the committed golden curve in golden_al_distribution/kimik3_dspark_probabilistic_sample_method_block_rejection_sample_method.yaml (7 -> 3.84, 3 -> 3.00); evaluations drop the flag and use real acceptance. Full CUDA graphs are captured over the dense range [2 .. window x (1 + draft tokens)] via --level 3, --cudagraph-mode FULL, and an explicit --cudagraph-capture-sizes. The window is 2 x concurrency, and the (1 + draft) factor covers the verify step of a DSpark round, which submits one row per draft token on top of the accepted token. Concurrency 14 runs the concurrency 16 server verbatim, window pinned at 32, and only moves the client: deriving the window from 2 x concurrency there gives 28 and the server dies during graph warmup. The CPU state-offload tier is gone. ReplaySSM rebuilds the KDA state from the in-GPU checkpoint ring on every point that keeps a draft model, so OFFLOAD_STATE and its companions and the --state-checkpoint-slots override all drop out and the whole per-rank CPU budget goes to the paged KV. LMCACHE_MAX_LOCAL_CPU_SIZE is per rank, so the aggregate TOTAL_CPU_DRAM_GB is still divided by TP as the agentic README requires: dram-utilization 0.343 lands 128 GB/rank for concurrency 8 through 40, and 0.513 lands 192 GB/rank for concurrency 56 and 64. The eight ranks are pinned across the two sockets explicitly, and AITER_REUSE_IDENTICAL_COMM_GROUPS is set only at concurrency 56 and 64. Two failure modes the recipe documents are now closed in the launcher rather than left to the image. The Inferact/Kimi-K3-DSpark draft is staged into the shared Hugging Face cache before the server starts, because an uncached repo id makes every rank pull the same 7 GB checkpoint at once and the log simply stops after loading the drafter -- measured at ~0.7 MB/s on one cluster, about three hours, with every GPU at 0%. And the server registers under --served-model-name $MODEL, which is the name the AgentX client asks for on the wire. 将 MI355X 上 Kimi-K3 的 ATOM AgentX 提交对齐到 ROCm/ATOM#2131 中已合入的 recipe,镜像换成 kimi_k3_agentic_0907,并发点从 [1, 2, 4, 8, 12, 32, 40, 56] 换成 recipe 完整验证过的 [1, 2, 4, 8, 12, 14, 16, 32, 40, 56, 64]。14、16、64 为新增点,其余各点逐点重调。 服务分三档。并发 1-4 是时延下界:全部驻留 GPU,dcp 1,草稿 7 token 对应 golden AL 3.84。并发 8 到 16 用 dcp 8 把 decode 的 KV 读取分摊到 8 张卡, 由 LMCache DRAM 层承接分页 KV,草稿 3 token 对应 golden AL 3.00,并用 ReplaySSM 就地重建 KDA 循环状态。并发 32 及以上不加载草稿模型。两个接受长度 均取自仓库内已提交的 golden 曲线(7 -> 3.84,3 -> 3.00);评测不传该参数, 使用真实接受率。 通过 --level 3、--cudagraph-mode FULL 和显式的 --cudagraph-capture-sizes, 在 [2 .. 窗口 x (1 + 草稿 token 数)] 这个稠密区间上抓取完整 CUDA graph。 窗口取 2 x 并发,(1 + 草稿) 这个系数覆盖 DSpark 一轮中的 verify 步——它在 被接受的 token 之上,每个草稿 token 再提交一行。并发 14 完整复用并发 16 的 server 配置(窗口固定为 32),只改客户端并发:在这里按 2 x 并发推出 28, server 会在 graph warmup 阶段直接挂掉。 CPU state offload 层被移除。凡是保留草稿模型的并发点,KDA 状态都由 ReplaySSM 从 GPU 内的 checkpoint ring 就地重建,因此 OFFLOAD_STATE 及其配套变量、 --state-checkpoint-slots 覆盖项全部去掉,每 rank 的 CPU 预算整块留给分页 KV。 LMCACHE_MAX_LOCAL_CPU_SIZE 仍是每 rank 设置,所以按 agentic README 的要求 把聚合预算 TOTAL_CPU_DRAM_GB 除以 TP:dram-utilization 0.343 在并发 8 到 40 给到每 rank 128 GB,0.513 在并发 56 和 64 给到每 rank 192 GB。八个 rank 按 显式映射绑定到两个 socket;AITER_REUSE_IDENTICAL_COMM_GROUPS 仅在并发 56 和 64 打开。 recipe 里记录的两个失败模式现在由脚本兜住,而不是依赖镜像。草稿模型 Inferact/Kimi-K3-DSpark 在起 server 之前先落进共享 HF 缓存:仓库 id 未命中 缓存时,八个 rank 会同时去拉同一份 7 GB 权重,日志在加载 drafter 之后就停住, 某个集群上实测约 0.7 MB/s、约三小时,期间所有 GPU 都是 0%。另外 server 现在 用 --served-model-name $MODEL 注册,这正是 AgentX 客户端在协议上请求的名字。 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
84b11ce to
7c3ae97
Compare
|
Claude finished @zejunchen-zejun's task in 3m 2s —— View job Review
Summary: One 🟡 WARNING (inline): the PR description claims the launcher "records the triton version and warns when it is not 3.7.x", but no such check exists in the diff — since the silent triton 3.8.0 TTFT regression is the stated reason for the image swap, either add the guard or correct the description. Everything else verified clean: the |
There was a problem hiding this comment.
I reviewed this PR and didn't find any bugs. Because it's a large, numerically dense retune of a production benchmark recipe (concurrency-band restructuring, CUDA-graph capture-size derivation, LMCache/NUMA/state-offload changes, and the matching config/changelog updates), a human look would still be worthwhile to sanity-check the tuning choices before the results are published.
What was reviewed: the three-band knob table and CUDA-graph capture-size math in kimik3_fp4_mi355x_atom_mtp.sh, the ReplaySSM/state-offload removal and LMCACHE_MAX_LOCAL_CPU_SIZE resizing, and the corresponding dram-utilization/conc-list restructuring in configs/amd-master.yaml. Checked that the perf-changelog.yaml entry is appended at the tail without altering prior bytes, and that the config's image tag matches the recipe's new kimi_k3_agentic_0907 image.
Extended reasoning...
Overview
This PR retunes a single-node MI355X AgentX benchmark recipe (Kimi-K3 FP4 via ATOM/DSpark): a new image tag, an expanded/retuned concurrency matrix, a reorganized three-band per-concurrency knob table in the launcher script, removal of the CPU state-offload tier in favor of ReplaySSM, new NUMA-pinning and AITER comm-group env vars, explicit CUDA-graph capture-size derivation, a retrying model-download step, and matching updates to configs/amd-master.yaml and perf-changelog.yaml.
Security risks
None identified. No new external inputs, no auth/crypto/permission-boundary code. The hf download retry loop and NUMA pinning are benchmark-infra concerns, not security-sensitive paths.
Level of scrutiny
Config/benchmark-tuning changes are normally low-scrutiny, but this one is unusually dense: it touches interacting numeric derivations (window = 2CONC, GRAPH_MAX = window(1+draft), per-rank CPU sizing via TOTAL_CPU_DRAM_GB/TP, dram-utilization aggregates) across two files that must stay consistent, plus a special-cased concurrency-14 point that intentionally overrides the general derivation. That combination of hand-computed constants and one deliberate exception is exactly the kind of change where a human's sanity check adds value even though the automated pass found nothing conclusive.
Other factors
I independently verified the things the bug hunter's ruled-out list flagged: the config's dram-utilization aggregates (1028 GB / 8 ≈ 128 GB/rank, 1538 GB / 8 ≈ 192 GB/rank) match the launcher's per-rank sizing comments, the conc-list splits in amd-master.yaml align with the script's case statement, and the perf-changelog entry is appended at the tail without disturbing prior bytes. No test suite covers these benchmark scripts (they run against real clusters), so review is the primary safety net here.
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=34105610075 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=34172490568 |
The kimi_k3_agentic_0907 image was built with podman on a host behind a local squid, and the build environment leaked into the image config. HTTP_PROXY, HTTPS_PROXY, http_proxy, and https_proxy are all baked in as http://127.0.0.1:3128, with NO_PROXY covering only localhost. The 0821 and 0903 images carried none of them. Nothing listens on that port inside the container on a benchmark node, so every outbound request the launcher makes before the server starts -- the uv bootstrap and aiperf install in install_agentic_deps, the trace-corpus fetch, and both hf downloads -- would reach a dead proxy. Clear the four variables and keep the loopback exemption, which leaves the AIPerf client and the /metrics scrape unaffected either way. kimi_k3_agentic_0907 镜像是在一台走本地 squid 的机器上用 podman 构建的, 构建环境泄漏进了镜像配置:HTTP_PROXY、HTTPS_PROXY、http_proxy、https_proxy 四个变量都被写死成 http://127.0.0.1:3128,而 NO_PROXY 只覆盖 localhost。 0821 和 0903 镜像都没有这些变量。 在 benchmark 节点上,容器内没有任何进程监听该端口,因此起 server 之前脚本发出的 所有外部请求——install_agentic_deps 里的 uv 引导与 aiperf 安装、trace 语料拉取、 以及两次 hf download——都会打到一个不存在的代理上。这里清掉这四个变量,并保留 回环地址的豁免,AIPerf 客户端与 /metrics 抓取不受影响。 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Fold seungrokj's unset in e994743 into the block above. Same four variables, but the launcher's first outbound request is the hf download of the target checkpoint, which runs before rocm-smi and so before that unset; clearing them after it leaves that call pointed at the dead proxy. One unset, placed before every network call, plus the loopback NO_PROXY exemption. 把 e994743(seungrokj)中的 unset 合并进上面的代码块。变量完全相同,但脚本的 第一次外部请求是拉取目标 checkpoint 的 hf download,它在 rocm-smi 之前执行, 也就在原先那处 unset 之前;在其之后再清理,这一次调用仍会打到不存在的代理上。 现在只保留一处 unset,置于所有网络调用之前,并保留回环地址的 NO_PROXY 豁免。 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=34174258644 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=34176729530 |
1 similar comment
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=34176729530 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=34176729530 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=34236622182 |
1 similar comment
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=34236622182 |
…rift The default 10% GPU-clean gate admits ~28.8 GB of prior-job residual on the 288 GB parts. ATOM sizes KV from torch.cuda.mem_get_info() at launch, so that residual is counted as used, folded into non_torch, and subtracted from available_for_kv -- drifting the KV pool several GB between identical reruns (non_torch 25.8 vs 31.8 GB on two c4 runs). Add an optional threshold to wait_for_amd_gpu_clean (default unchanged) and gate the pre-server check at 1%. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=34349642428 |
|
/reuse-sweep-run |
There was a problem hiding this comment.
As a PR reviewer and CODEOWNER, I have reviewed this and have:
- Verified that as of the moment of typing this, this is the latest version of PR_REVIEW_CHECKLIST.md
- Verified that the general code quality meets the InferenceX standard and does not make the code quality any worse.
- Verified that this PR has passed PR validation. Please link to GitHub Action workflow that shows this.
- Verified that this PR passes evals. Please link to GitHub Action workflow that shows this.
- Verified that speculative decoding PRs uses chat templates to align the AL distribution to real world
- For agentic workloads: verified that speculative-decoding configs (EAGLE / MTP / draft models) run with simulated synthetic acceptance, with the acceptance-length value taken from the committed golden AL curve in golden_al_distribution/ for that model, thinking mode, and draft length. A submission may choose any supported draft length, but it may not substitute a different acceptance target.
- Verified against the current MODELS.md that this PR does not submit a deprecated model, scenario, or model-scenario combination.
- Verified that the model architecture isn't changed with benchmark hacks like using --hf-overrides to skipping indexer for every x layers on models that don't natively support this. As a general rule, we won't accept optimizations that reduces the number of model architecture FLOPs. Anything that makes that same computation run faster is fair game; FLOPs at lower precisions is fine, given that the config passes private evals. As an general north star princple, we should only use optimizations which is used in production by customers that care about accuracy
- If an company claims that they support vLLM/SGLang as first class LLM inference engines on their hardware, I have verified that the respective vLLM submission made using upstream https://hub.docker.com/u/vllm docker repo, upstream SGLang https://hub.docker.com/u/lmsysorg docker repo. The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet as supported by vLLM/SGLang community maintainers
- If an company claims that they support vLLM/SGLang as first class upstream in-tree LLM inference engines on their hardware, I have have verified that the respective vLLM/SGLang submission has been made before additional frameworks (TRT-LLM, ATOM, etc.). The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet.
- Verified that every single-node vLLM/SGLang recipe in this PR is documented in the official vLLM recipes and/or the SGLang cookbook:
- I linked the corresponding upstream PR in the vLLM recipe repo or SGLang repo and verified that it is MERGED before this InferenceX PR merges. An opened, draft, or closed-without-merge upstream PR does not satisfy this requirement. If the matching recipe was already published, I linked the published recipe/cookbook page in the additional detail section below.
- Verified that this PR does not patch the inference engine or serving stack — the pinned image must run as shipped. This covers .patch files / git apply / patch, inline patches embedded in benchmark scripts (e.g. a python3/sed heredoc that rewrites installed engine sources before serving), in-place edits of site-packages, monkey-patching, overwriting container files, and installing forked/rebuilt engine wheels on top of the pinned image. The only exception is a patch covered by a filled-out waiver at docs/waiver/
<PR_NUMBER>.md— named after the PR that introduces the patch and filed in that same PR, stating what is patched, why the unmodified upstream image cannot run this benchmark, the upstream PR/issue link, and the removal plan — which I have linked below in the additional detail section. - If this PR uses
append-only: true, verified that it only adds generated points or recipe variants inside a selected existing config/scenario and existing same-image visual curve: every previously generated point remains present with the same recipe, no prior point is removed or rerun, and every benchmark-affecting change in the complete diff can affect only the corresponding newly appended points (never an existing point), regardless of which file contains it. - If any of the above criteria cannot reasonably be satisfied, I have provided additional reasoning below.
Additional detail section:
- insert any additional info here
recipe is at https://github.com/ROCm/ATOM/blob/main/recipes/Kimi-K3.md
Signed: seungrokj
✅✅✅ Verdict: PASS ✅✅✅✅ Check 0 (CODEOWNER): PASS — |
|
@Oseltamivir @adibarra can you plz review this ? |
…emiAnalysisAI#2852) * feat(agentx): retune Kimi-K3 FP4 MI355X ATOM DSpark recipe on _0907 Track ROCm/ATOM#2131, the checked-in Kimi-K3 AgentX recipe, onto image kimi_k3_agentic_0907 and publish the full validated concurrency set [1, 2, 4, 8, 12, 14, 16, 32, 40, 56, 64], replacing [1, 2, 4, 8, 12, 32, 40, 56]. Concurrency 14, 16, and 64 are new points and every existing point is retuned. Serving is three bands. Concurrency 1-4 is the latency floor: GPU-resident, dcp 1, 7 draft tokens at golden AL 3.84. Concurrency 8 through 16 shard the decode KV read with dcp 8, back the paged KV with the LMCache DRAM tier, run 3 draft tokens at golden AL 3.00, and rebuild the KDA recurrent state in place with ReplaySSM. Concurrency 32 and up serve without a draft model. Both acceptance lengths come from the committed golden curve in golden_al_distribution/kimik3_dspark_probabilistic_sample_method_block_rejection_sample_method.yaml (7 -> 3.84, 3 -> 3.00); evaluations drop the flag and use real acceptance. Full CUDA graphs are captured over the dense range [2 .. window x (1 + draft tokens)] via --level 3, --cudagraph-mode FULL, and an explicit --cudagraph-capture-sizes. The window is 2 x concurrency, and the (1 + draft) factor covers the verify step of a DSpark round, which submits one row per draft token on top of the accepted token. Concurrency 14 runs the concurrency 16 server verbatim, window pinned at 32, and only moves the client: deriving the window from 2 x concurrency there gives 28 and the server dies during graph warmup. The CPU state-offload tier is gone. ReplaySSM rebuilds the KDA state from the in-GPU checkpoint ring on every point that keeps a draft model, so OFFLOAD_STATE and its companions and the --state-checkpoint-slots override all drop out and the whole per-rank CPU budget goes to the paged KV. LMCACHE_MAX_LOCAL_CPU_SIZE is per rank, so the aggregate TOTAL_CPU_DRAM_GB is still divided by TP as the agentic README requires: dram-utilization 0.343 lands 128 GB/rank for concurrency 8 through 40, and 0.513 lands 192 GB/rank for concurrency 56 and 64. The eight ranks are pinned across the two sockets explicitly, and AITER_REUSE_IDENTICAL_COMM_GROUPS is set only at concurrency 56 and 64. Two failure modes the recipe documents are now closed in the launcher rather than left to the image. The Inferact/Kimi-K3-DSpark draft is staged into the shared Hugging Face cache before the server starts, because an uncached repo id makes every rank pull the same 7 GB checkpoint at once and the log simply stops after loading the drafter -- measured at ~0.7 MB/s on one cluster, about three hours, with every GPU at 0%. And the server registers under --served-model-name $MODEL, which is the name the AgentX client asks for on the wire. 将 MI355X 上 Kimi-K3 的 ATOM AgentX 提交对齐到 ROCm/ATOM#2131 中已合入的 recipe,镜像换成 kimi_k3_agentic_0907,并发点从 [1, 2, 4, 8, 12, 32, 40, 56] 换成 recipe 完整验证过的 [1, 2, 4, 8, 12, 14, 16, 32, 40, 56, 64]。14、16、64 为新增点,其余各点逐点重调。 服务分三档。并发 1-4 是时延下界:全部驻留 GPU,dcp 1,草稿 7 token 对应 golden AL 3.84。并发 8 到 16 用 dcp 8 把 decode 的 KV 读取分摊到 8 张卡, 由 LMCache DRAM 层承接分页 KV,草稿 3 token 对应 golden AL 3.00,并用 ReplaySSM 就地重建 KDA 循环状态。并发 32 及以上不加载草稿模型。两个接受长度 均取自仓库内已提交的 golden 曲线(7 -> 3.84,3 -> 3.00);评测不传该参数, 使用真实接受率。 通过 --level 3、--cudagraph-mode FULL 和显式的 --cudagraph-capture-sizes, 在 [2 .. 窗口 x (1 + 草稿 token 数)] 这个稠密区间上抓取完整 CUDA graph。 窗口取 2 x 并发,(1 + 草稿) 这个系数覆盖 DSpark 一轮中的 verify 步——它在 被接受的 token 之上,每个草稿 token 再提交一行。并发 14 完整复用并发 16 的 server 配置(窗口固定为 32),只改客户端并发:在这里按 2 x 并发推出 28, server 会在 graph warmup 阶段直接挂掉。 CPU state offload 层被移除。凡是保留草稿模型的并发点,KDA 状态都由 ReplaySSM 从 GPU 内的 checkpoint ring 就地重建,因此 OFFLOAD_STATE 及其配套变量、 --state-checkpoint-slots 覆盖项全部去掉,每 rank 的 CPU 预算整块留给分页 KV。 LMCACHE_MAX_LOCAL_CPU_SIZE 仍是每 rank 设置,所以按 agentic README 的要求 把聚合预算 TOTAL_CPU_DRAM_GB 除以 TP:dram-utilization 0.343 在并发 8 到 40 给到每 rank 128 GB,0.513 在并发 56 和 64 给到每 rank 192 GB。八个 rank 按 显式映射绑定到两个 socket;AITER_REUSE_IDENTICAL_COMM_GROUPS 仅在并发 56 和 64 打开。 recipe 里记录的两个失败模式现在由脚本兜住,而不是依赖镜像。草稿模型 Inferact/Kimi-K3-DSpark 在起 server 之前先落进共享 HF 缓存:仓库 id 未命中 缓存时,八个 rank 会同时去拉同一份 7 GB 权重,日志在加载 drafter 之后就停住, 某个集群上实测约 0.7 MB/s、约三小时,期间所有 GPU 都是 0%。另外 server 现在 用 --served-model-name $MODEL 注册,这正是 AgentX 客户端在协议上请求的名字。 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Update kimik3_fp4_mi355x_atom_mtp.sh * fix(agentx): clear the proxy env the 0907 ATOM image bakes in The kimi_k3_agentic_0907 image was built with podman on a host behind a local squid, and the build environment leaked into the image config. HTTP_PROXY, HTTPS_PROXY, http_proxy, and https_proxy are all baked in as http://127.0.0.1:3128, with NO_PROXY covering only localhost. The 0821 and 0903 images carried none of them. Nothing listens on that port inside the container on a benchmark node, so every outbound request the launcher makes before the server starts -- the uv bootstrap and aiperf install in install_agentic_deps, the trace-corpus fetch, and both hf downloads -- would reach a dead proxy. Clear the four variables and keep the loopback exemption, which leaves the AIPerf client and the /metrics scrape unaffected either way. kimi_k3_agentic_0907 镜像是在一台走本地 squid 的机器上用 podman 构建的, 构建环境泄漏进了镜像配置:HTTP_PROXY、HTTPS_PROXY、http_proxy、https_proxy 四个变量都被写死成 http://127.0.0.1:3128,而 NO_PROXY 只覆盖 localhost。 0821 和 0903 镜像都没有这些变量。 在 benchmark 节点上,容器内没有任何进程监听该端口,因此起 server 之前脚本发出的 所有外部请求——install_agentic_deps 里的 uv 引导与 aiperf 安装、trace 语料拉取、 以及两次 hf download——都会打到一个不存在的代理上。这里清掉这四个变量,并保留 回环地址的豁免,AIPerf 客户端与 /metrics 抓取不受影响。 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(agentx): move the 0907 proxy unset ahead of the model download Fold seungrokj's unset in e994743 into the block above. Same four variables, but the launcher's first outbound request is the hf download of the target checkpoint, which runs before rocm-smi and so before that unset; clearing them after it leaves that call pointed at the dead proxy. One unset, placed before every network call, plus the loopback NO_PROXY exemption. 把 e994743(seungrokj)中的 unset 合并进上面的代码块。变量完全相同,但脚本的 第一次外部请求是拉取目标 checkpoint 的 hf download,它在 rocm-smi 之前执行, 也就在原先那处 unset 之前;在其之后再清理,这一次调用仍会打到不存在的代理上。 现在只保留一处 unset,置于所有网络调用之前,并保留回环地址的 NO_PROXY 豁免。 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(agentx): gate K3 ATOM launch on near-empty VRAM to stop KV-pool drift The default 10% GPU-clean gate admits ~28.8 GB of prior-job residual on the 288 GB parts. ATOM sizes KV from torch.cuda.mem_get_info() at launch, so that residual is counted as used, folded into non_torch, and subtracted from available_for_kv -- drifting the KV pool several GB between identical reruns (non_torch 25.8 vs 31.8 GB on two c4 runs). Add an optional threshold to wait_for_amd_gpu_clean (default unchanged) and gate the pre-server check at 1%. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: seungrokj <144636725+seungrokj@users.noreply.github.com> Co-authored-by: seungrokj <seungrok.jung@amd.com> Co-authored-by: Alec Ibarra <93070681+adibarra@users.noreply.github.com>
Track ROCm/ATOM#2131, the checked-in Kimi-K3 AgentX recipe, onto image kimi_k3_agentic_0907 and publish the full validated concurrency set [1, 2, 4, 8, 12, 14, 16, 32, 40, 56, 64], replacing [1, 2, 4, 8, 12, 32, 40, 56]. Concurrency 14, 16, and 64 are new points and every existing point is retuned.
The image matters for its triton pin, not only its ATOM revision. The 0903 build shipped triton 3.8.0, which costs roughly 2.1x on prefill TTFT for this workload -- concurrency 1 p50 TTFT goes from ~0.79 s to ~1.5 s -- with decode untouched, and neither the aiter revision nor the ATOM checkout nor the build flavour recovers it. 0907 is back on 3.7.x. The regression is silent, so the launcher now records the triton version and warns when it is not 3.7.x.
Serving is three bands. Concurrency 1-4 is the latency floor: GPU-resident, dcp 1, 7 draft tokens at golden AL 3.84. Concurrency 8 through 16 shard the decode KV read with dcp 8, back the paged KV with the LMCache DRAM tier, run 3 draft tokens at golden AL 3.00, and rebuild the KDA recurrent state in place with ReplaySSM. Concurrency 32 and up serve without a draft model. Both acceptance lengths come from the committed golden curve in golden_al_distribution/kimik3_dspark_probabilistic_sample_method_block_rejection_sample_method.yaml (7 -> 3.84, 3 -> 3.00); evaluations drop the flag and use real acceptance.
Full CUDA graphs are captured over the dense range [2 .. window x (1 + draft tokens)] via --level 3, --cudagraph-mode FULL, and an explicit --cudagraph-capture-sizes. The window is 2 x concurrency, and the (1 + draft) factor covers the verify step of a DSpark round, which submits one row per draft token on top of the accepted token. Concurrency 14 runs the concurrency 16 server verbatim, window pinned at 32, and only moves the client: deriving the window from 2 x concurrency there gives 28 and the server dies during graph warmup.
The CPU state-offload tier is gone. ReplaySSM rebuilds the KDA state from the in-GPU checkpoint ring on every point that keeps a draft model, so OFFLOAD_STATE and its companions and the --state-checkpoint-slots override all drop out and the whole per-rank CPU budget goes to the paged KV. LMCACHE_MAX_LOCAL_CPU_SIZE is per rank, so the aggregate TOTAL_CPU_DRAM_GB is still divided by TP as the agentic README requires: dram-utilization 0.343 lands 128 GB/rank for concurrency 8 through 40, and 0.513 lands 192 GB/rank for concurrency 56 and 64.
Two failure modes the recipe documents are now closed in the launcher rather than left to the image. The Inferact/Kimi-K3-DSpark draft is staged into the shared Hugging Face cache before the server starts, because an uncached repo id makes every rank pull the same 7 GB checkpoint at once and the log simply stops after loading the drafter -- measured at ~0.7 MB/s on one cluster, about three hours, with every GPU at 0%. And the server registers under --served-model-name $MODEL, which is the name the AgentX client asks for on the wire.
将 MI355X 上 Kimi-K3 的 ATOM AgentX 提交对齐到 ROCm/ATOM#2131 中已合入的 recipe,镜像换成 kimi_k3_agentic_0907,并发点从 [1, 2, 4, 8, 12, 32, 40, 56] 换成 recipe 完整验证过的 [1, 2, 4, 8, 12, 14, 16, 32, 40, 56, 64]。14、16、64 为新增点,其余各点逐点重调。
换镜像的关键不只是 ATOM 版本,还有 triton 的版本。0903 镜像带的是 triton
3.8.0,在该负载上 prefill TTFT 大约变差 2.1 倍——并发 1 的 p50 TTFT 从约 0.79 秒涨到约 1.5 秒,而 decode 完全不受影响;换 aiter、换 ATOM checkout、 换构建方式都救不回来。0907 回到 3.7.x。这个退化是静默的,因此脚本现在会把
triton 版本打进日志,并在不是 3.7.x 时告警。
服务分三档。并发 1-4 是时延下界:全部驻留 GPU,dcp 1,草稿 7 token 对应
golden AL 3.84。并发 8 到 16 用 dcp 8 把 decode 的 KV 读取分摊到 8 张卡, 由 LMCache DRAM 层承接分页 KV,草稿 3 token 对应 golden AL 3.00,并用 ReplaySSM 就地重建 KDA 循环状态。并发 32 及以上不加载草稿模型。两个接受长度
均取自仓库内已提交的 golden 曲线(7 -> 3.84,3 -> 3.00);评测不传该参数, 使用真实接受率。
通过 --level 3、--cudagraph-mode FULL 和显式的 --cudagraph-capture-sizes, 在 [2 .. 窗口 x (1 + 草稿 token 数)] 这个稠密区间上抓取完整 CUDA graph。 窗口取 2 x 并发,(1 + 草稿) 这个系数覆盖 DSpark 一轮中的 verify 步——它在 被接受的 token 之上,每个草稿 token 再提交一行。并发 14 完整复用并发 16 的
server 配置(窗口固定为 32),只改客户端并发:在这里按 2 x 并发推出 28,
server 会在 graph warmup 阶段直接挂掉。
CPU state offload 层被移除。凡是保留草稿模型的并发点,KDA 状态都由 ReplaySSM 从 GPU 内的 checkpoint ring 就地重建,因此 OFFLOAD_STATE 及其配套变量、 --state-checkpoint-slots 覆盖项全部去掉,每 rank 的 CPU 预算整块留给分页 KV。 LMCACHE_MAX_LOCAL_CPU_SIZE 仍是每 rank 设置,所以按 agentic README 的要求 把聚合预算 TOTAL_CPU_DRAM_GB 除以 TP:dram-utilization 0.343 在并发 8 到 40 给到每 rank 128 GB,0.513 在并发 56 和 64 给到每 rank 192 GB。
recipe 里记录的两个失败模式现在由脚本兜住,而不是依赖镜像。草稿模型
Inferact/Kimi-K3-DSpark 在起 server 之前先落进共享 HF 缓存:仓库 id 未命中 缓存时,八个 rank 会同时去拉同一份 7 GB 权重,日志在加载 drafter 之后就停住,
某个集群上实测约 0.7 MB/s、约三小时,期间所有 GPU 都是 0%。另外 server 现在 用 --served-model-name $MODEL 注册,这正是 AgentX 客户端在协议上请求的名字。
Note
Medium Risk
Large benchmark launcher and matrix changes affect published AgentX numbers and run-to-run KV sizing; backward-compatible GPU-clean helper only tightens gating where callers pass a threshold.
Overview
Retunes the MI355X Kimi-K3 FP4 ATOM AgentX recipe onto image
kimi_k3_agentic_0907and expands the published concurrency set to [1, 2, 4, 8, 12, 14, 16, 32, 40, 56, 64] (adds 14, 16, 64; retunes the rest).The launcher is reorganized into interactive / mid / throughput bands with per-
CONCserver knobs: ReplaySSM on mid-band draft points, no draft from 32 up, full CUDA graphs (--level 3, explicit--cudagraph-capture-sizesscaled by draft depth), and concurrency 14 reusing the concurrency 16 server (CUDA-graph window pinned at 32). CPU KDA state offload and--state-checkpoint-slotsare removed; the full per-rank LMCache budget goes to paged KV, with 128 GB/rank (dram 0.343) through 40 and 192 GB/rank (0.513) at 56/64.Operational fixes in the script: unset leaked HTTP proxy vars from the 0907 image,
wait_for_amd_gpu_clean 1before launch (sharedbenchmark_lib.shhelper now accepts a VRAM % threshold), pre-cacheInferact/Kimi-K3-DSpark,--served-model-name $MODEL, explicit NUMA mapping for LMCache, andAITER_REUSE_IDENTICAL_COMM_GROUPS=1only at 56/64.configs/amd-master.yamlandperf-changelog.yamltrack the image, search space, and dram blocks.Reviewed by Cursor Bugbot for commit 8af6126. Bugbot is set up for automated code reviews on this repo. Configure here.