Repository navigation
(recipe) Kimi-K3 AgentX: FP8 prefill attention, PrefillDelayer from C16, capture from bs=1 - #2382
Merged
Merged
Conversation
…16, capture from bs=1 Adopt the configuration of the 2026-09-23 AgentX sweep (C1/4/14/16/48/56/72 on 8x MI355X): - ATOM_USE_FLYDSL_FP8_PREFILL_ATTN=1 on every band - ATOM_PREFILL_DECODE_INTERVAL=4 and ATOM_PREFILL_DELAYER_MAX_QUEUE_MS=5000 on CONC >= 16 - CUDA-graph capture sizes start at 1 (seq -s, 1) - container prerequisite: the 92 FlyDSL FP8 prefill kernels are AOT-built (ROCm/aiter#5796) and ATOM carries the merge_attn_states runtime-arg fix (#2378), with a one-line check - ReplaySSM paragraph now lists C14 among the on-bands, matching the table Against the current recipe's best runs: C48-C72 +10-15% total and +10-17% output throughput, ITL -14-18%; C1/C4 ITL p90 -12-13%; C14 TTFT p90 -19%. The delayer raises TTFT from C16 up (p50 +37-92%). The launcher's exported env and argv match the sweep's server script on all 14 bands, apart from default-equal or logging-only values. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Contributor
🏷️ CI GuideRuns automatically on every eligible PR before approval:
Heavy model tests:
|
Drop the FP8 prefill precompile prerequisite subsection and restore the section-0 intro. Revert the ReplaySSM paragraph edit. The recipe diff is now limited to ATOM_USE_FLYDSL_FP8_PREFILL_ATTN=1, the C>=16 PrefillDelayer variables and capture sizes starting at 1, plus their table rows. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This was referenced Sep 24, 2026
gbyu-amd
added a commit
that referenced
this pull request
Sep 28, 2026
…m bs=1 (#2417) * fix(ci-mesh): align Kimi-K3 AgentX cells with the recipe (#2382) Sync the K3 P/D nightly cells with recipes/Agentic-Kimi-K3.md as of d8cdd4e: - ATOM_USE_FLYDSL_FP8_PREFILL_ATTN=1 on every band - PrefillDelayer (ATOM_PREFILL_DECODE_INTERVAL=4, ATOM_PREFILL_DELAYER_MAX_QUEUE_MS=5000) from CONC 16 up - `cudagraph: auto` captures the dense range from bs=1 instead of 2 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(ci-mesh): drop PrefillDelayer env from K3 P/D cells Its benefit was measured on the aggregated recipe server; on the P/D split the decode instance never installs it and the prefill instance has no decode batch to protect, so keep the cells unchanged there. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: ganyi <ygan@amd.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seungrokj
added a commit
to SemiAnalysisAI/InferenceX
that referenced
this pull request
Sep 29, 2026
…cipe (#3407) * feat(agentx): bump Kimi-K3 FP4 MI355X ATOM image to 0924 and track recipe Track recipes/Agentic-Kimi-K3.md as retuned in ROCm/ATOM#2382: enable FlyDSL FP8 prefill attention on every band, and hold a ready prefill for four decode passes from concurrency 16 up. The published concurrency set and every other launch argument are unchanged. 将 MI355X Kimi-K3 FP4 ATOM AgentX 提交切换到 0924 镜像,并跟随 ROCm/ATOM#2382 重调后的 recipe:全部并发开启 FlyDSL FP8 prefill attention;并发 16 及以上时让就绪的 prefill 等待 4 个 decode 轮次。 已发布的并发点集合与其余启动参数保持不变。 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(agentx): point perf-changelog entry at PR 3407 将 perf-changelog 条目的 pr-link 指向 PR 3407。 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(agentx): bump Kimi-K3 FP4 MI355X ATOM image to nightly_202609251613 Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat(agentx): restore Kimi-K3 MI355X ATOM LMCache bands on srt-slurm Carry srt-slurm patch 507 so ATOM aggregate workers accept extra-kv-connectors, and restore the DCP8 bands from the legacy config and ROCm/ATOM recipes/Agentic-Kimi-K3.md: concurrency 14 and 16 with DSpark 3 and ReplaySSM, 48, 56 and 72 without a draft, all on the in-process lmcache_offload connector (128 GB/rank, 192 GB/rank at 56 and 72). The PrefillDelayer applies from concurrency 16 up. The single-node adapter now reads ATOM's decode-context-parallel-size for the DCP_SIZE check, as it does for vLLM. * chore(agentx): Kimi-K3 FP4 MI355X ATOM image 0925, FlyDSL prefill, prefill delay, DCP8 LMCache - Move kimik3-fp4-mi355x-atom-agentic-mtp to rocm/atom-dev:nightly_202609251613, tracking recipes/Agentic-Kimi-K3.md as retuned in ROCm/ATOM#2382. - Enable FlyDSL FP8 prefill attention (ATOM_USE_FLYDSL_FP8_PREFILL_ATTN=1) at every concurrency. - From conc 16 up, hold a ready prefill for four decode passes (ATOM_PREFILL_DECODE_INTERVAL=4, ATOM_PREFILL_DELAYER_MAX_QUEUE_MS=5000); conc 1, 4 and 14 unchanged. - Restore DCP8 LMCache bands on the native srt-slurm recipe: conc 14/16 (DSpark 3, ReplaySSM) and 48/56/72 (no draft) use ATOM's in-process lmcache_offload connector via roles.agg.args.extra-kv-connectors (srt-slurm patch 507), 128 GB/rank up to 48 and 192 GB/rank at 56/72. Conc 1 and 4 stay GPU-resident. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(agentx): Kimi-K3 FP4 MI355X ATOM image 0925, FlyDSL prefill, prefill delay, DCP8 LMCache - Move kimik3-fp4-mi355x-atom-agentic-mtp to rocm/atom-dev:nightly_202609251613, tracking recipes/Agentic-Kimi-K3.md as retuned in ROCm/ATOM#2382. - Enable FlyDSL FP8 prefill attention (ATOM_USE_FLYDSL_FP8_PREFILL_ATTN=1) at every concurrency. - From conc 16 up, hold a ready prefill for four decode passes (ATOM_PREFILL_DECODE_INTERVAL=4, ATOM_PREFILL_DELAYER_MAX_QUEUE_MS=5000); conc 1, 4 and 14 unchanged. - Restore DCP8 LMCache bands on the native srt-slurm recipe: conc 14/16 (DSpark 3, ReplaySSM) and 48/56/72 (no draft) use ATOM's in-process lmcache_offload connector via roles.agg.args.extra-kv-connectors (srt-slurm patch 507), 128 GB/rank up to 48 and 192 GB/rank at 56/72. Conc 1 and 4 stay GPU-resident. - Draft model precision unchanged: online_quant_config still excludes every Inferact/Kimi-K3-DSpark linear (layers.*, context_proj), so its weights and activations stay BF16, and it shares the target's FP8 KV cache (kv_cache_dtype fp8). FlyDSL FP8 prefill attention applies only to the target; the draft's block pass runs as decode attention. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: seungrokj <144636725+seungrokj@users.noreply.github.com> Co-authored-by: seungrokj <seungrok.jung@amd.com> Co-authored-by: Cameron Quilici <cameron@semianalysis.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
This updates
recipes/Agentic-Kimi-K3.mdto the configuration of a 7-point AgentX sweep on 8×MI355X (C1/4/14/16/48/56/72, 3600 s each):ATOM_USE_FLYDSL_FP8_PREFILL_ATTN=1on every band (FlyDSL FP8 prefill attention).ATOM_PREFILL_DECODE_INTERVAL=4andATOM_PREFILL_DELAYER_MAX_QUEUE_MS=5000on CONC ≥ 16 (the PrefillDelayer: 4 protected decode passes after each prefill, plus prefill coalescing bounded by a 5 s queue age).ATOM_ENABLE_PREFILL_DELAYERalready defaults to 1, so these two variables are enough to construct thePrefillDelayer. The sweep checked this on every C ≥ 16 point: each one loggedPrefillDelayer initializedonce.seq -s, 1, dense[1, …, graph_max].The recipe changes are limited to these environment variables and launch arguments, plus the matching rows in the configuration table. Nothing else in the launcher or client changes. Every other value (DCP, max-num-seqs, LMCache size, comm-group reuse, ReplaySSM, spec, AIPerf flags and env) already matched the sweep.
Dependencies
Both dependencies are merged: #2378 (2026-09-24) and ROCm/aiter#5796 (2026-09-24). The sweep ran on ATOM d9f0720 + #2378, aiter 2887d4899 + #5796, and triton 3.7.0.
Results vs the current recipe's best measured runs
Deltas are relative to the best local run of the current recipe per band (reuse A/B-aligned). Bold marks a change outside the run-to-run band measured with 5 same-config c64 runs: tok/s ±4.7%, ITL ±6.2%, TTFT ±11.7%.
¹ At C1/C4 throughput is set by the trace's pacing: both runs served nearly the same requests. Latency is the signal there.
Absolute values of the new runs:
Trade-off to be aware of
The PrefillDelayer protects decode by holding prefills, so TTFT rises wherever it is on. The delayer's own counters show the hold rate climbing with concurrency: C16 31%, C48 38%, C56 51%, C72 51%. At C48–C72 this buys +10–17% output throughput and −14–18% ITL. At C16 the throughput change is inside noise, while TTFT p50 goes from 696 ms (C14, delayer off) to 1,450 ms. So the C ≥ 16 threshold is the sweep's setting, not a tuned one; C16 may be better served without it.
Not measured with the new knobs
C2, C8, C12, C32, C40, C64 and C80 follow the same rules (FP8 prefill everywhere, PrefillDelayer from C16 up) but were not re-run in this sweep. The effects above cannot be split per knob either: the ATOM/aiter changes above, FP8 prefill, the bs=1 capture and the delayer all changed together.
How the recipe was checked
The launcher block from this PR and the sweep's server script were each executed up to the server start for all 14 bands, dumping exported env and argv, then diffed. The only remaining differences are default-equal (
ATOM_USE_FLYDSL_GATHER_KV_B_PROJ=1is the default;0.90vs0.9) or logging-only (AITER_LOG_LEVEL). The AIPerf client, including itsAIPERF_*env, is identical to what the sweep ran, apart from URLs and paths.🤖 Generated with Claude Code