From 3c1880ad891253f42934d1999836daff5cd69c39 Mon Sep 17 00:00:00 2001 From: Guanbao Yu Date: Thu, 24 Sep 2026 03:31:07 +0000 Subject: [PATCH 1/2] (recipe) Kimi-K3 AgentX: FP8 prefill attention, PrefillDelayer from C16, capture from bs=1 Adopt the configuration of the 2026-09-23 AgentX sweep (C1/4/14/16/48/56/72 on 8x MI355X): - ATOM_USE_FLYDSL_FP8_PREFILL_ATTN=1 on every band - ATOM_PREFILL_DECODE_INTERVAL=4 and ATOM_PREFILL_DELAYER_MAX_QUEUE_MS=5000 on CONC >= 16 - CUDA-graph capture sizes start at 1 (seq -s, 1) - container prerequisite: the 92 FlyDSL FP8 prefill kernels are AOT-built (ROCm/aiter#5796) and ATOM carries the merge_attn_states runtime-arg fix (ROCm/ATOM#2378), with a one-line check - ReplaySSM paragraph now lists C14 among the on-bands, matching the table Against the current recipe's best runs: C48-C72 +10-15% total and +10-17% output throughput, ITL -14-18%; C1/C4 ITL p90 -12-13%; C14 TTFT p90 -19%. The delayer raises TTFT from C16 up (p50 +37-92%). The launcher's exported env and argv match the sweep's server script on all 14 bands, apart from default-equal or logging-only values. Co-Authored-By: Claude Opus 5.5 (1M context) --- recipes/Agentic-Kimi-K3.md | 36 +++++++++++++++++++++++++++++++----- 1 file changed, 31 insertions(+), 5 deletions(-) diff --git a/recipes/Agentic-Kimi-K3.md b/recipes/Agentic-Kimi-K3.md index 0861917fa7..19123a4428 100644 --- a/recipes/Agentic-Kimi-K3.md +++ b/recipes/Agentic-Kimi-K3.md @@ -8,6 +8,7 @@ with ATOM on 8×AMD MI355X GPUs. The validated configuration uses: - FP8 KV cache - native GPU prefix caching; LMCache CPU offload from CONC 8 up - DSpark speculative decoding on CONC ≤ 16, with synthetic acceptance +- FlyDSL FP8 prefill attention on every band; the PrefillDelayer from CONC 16 up - the SemiAnalysis Weka AgentX workload The workload uses the AIPerf scenario `inferencex-agentx-mvp` and public @@ -42,6 +43,8 @@ changes. | Prefix cache | Enabled (required for AgentX multi-turn prefix hits) | | Speculative decoding | DSpark (`--method dspark`), not native MTP | | Synthetic acceptance | `--spec-decode-acceptance-length` (performance-only; see below) | +| Prefill attention | FlyDSL FP8 (`ATOM_USE_FLYDSL_FP8_PREFILL_ATTN=1`), every band | +| PrefillDelayer | CONC ≥ 16: `ATOM_PREFILL_DECODE_INTERVAL=4` (4 decode passes after each prefill), `ATOM_PREFILL_DELAYER_MAX_QUEUE_MS=5000` | | Profiling duration | 3,600 seconds | | Warmup | 10 additional one-token requests per lane | | AIPerf | `0.12.0` (`agentx-v1.0.4`) | @@ -49,7 +52,7 @@ changes. ### Per-concurrency server table `graph_max = 2 * CONC * (1 + spec)`. Capture sizes are the dense range -`[2, 3, …, graph_max]`. `max-num-seqs` is **not** always `2 * CONC`. +`[1, 2, …, graph_max]`. `max-num-seqs` is **not** always `2 * CONC`. | CONC | DCP | spec | AL | LMCache | ReplaySSM | `max-num-seqs` | batched tokens | GPU util | `graph_max` | |---:|---:|---:|---:|---|---:|---:|---:|---:|---:| @@ -74,10 +77,13 @@ C8–C48 use 128 GiB LMCache; C56–C80 use **192 GiB**. LMCache CPU size is `AITER_REUSE_IDENTICAL_COMM_GROUPS=1` on **C56/C64/C72/C80**, and `0` on every other band (C1…C48). +`ATOM_PREFILL_DECODE_INTERVAL=4` and `ATOM_PREFILL_DELAYER_MAX_QUEUE_MS=5000` on +**C16 and up**; unset on C1…C14. + ## 0. Container prerequisites -Two things about the container decide whether the numbers above are -reproducible at all. Both fail quietly rather than loudly, so they are worth +Three things about the container decide whether the numbers above are +reproducible at all. All fail quietly rather than loudly, so they are worth checking before the first run. ### triton must be 3.7.x @@ -94,6 +100,21 @@ swapping aiter's `.so` between revisions, swapping the ATOM checkout, and switching between a full CI prebuild and a lean JIT build. Use an image with triton 3.7.x, such as `kimi_k3_agentic_0907`. +### FP8 prefill kernels must be precompiled + +```bash +ls "$(python3 -c 'import aiter, os; print(os.path.dirname(aiter.__file__))')/jit/flydsl_cache" \ + | grep -c '^launch_flash_attn_dualwave_swp_' # expect 92 on a fresh container +``` + +The FlyDSL FP8 prefill attention kernel is compiled per configuration the first +time a request needs it, stalling that request for seconds. aiter builds all 92 +Kimi-K3 variants ahead of time from +`aiter/configs/model_configs/kimik3_fmha_fp8_aot.csv` (ROCm/aiter#5796); a count +below 92 means the image predates that. ATOM must also include the +`merge_attn_states` runtime-arg fix (ROCm/ATOM#2378), without which the MLA +chunked-prefill merge recompiles for every distinct prefill token count. + ### `DRAFT_MODEL_PATH` must point at a local copy ```bash @@ -131,6 +152,7 @@ export AITER_SITUV2_A4W4=1 export AITER_FLYDSL_STAGE2_FP8=1 export ATOM_STATE_CHECKPOINT_DEMAND=0 export ATOM_GDN_SSM_DTYPE="${ATOM_GDN_SSM_DTYPE:-fp16}" +export ATOM_USE_FLYDSL_FP8_PREFILL_ATTN=1 # Set explicitly to 1 for reproducibility; this is also the current default. export ATOM_USE_FLYDSL_GATHER_KV_B_PROJ=1 export PYTHONNOUSERSITE=1 @@ -190,6 +212,10 @@ if [[ "${CONC}" -ge 56 ]]; then fi AITER_REUSE_IDENTICAL_COMM_GROUPS="${AITER_REUSE_IDENTICAL_COMM_GROUPS:-0}" export AITER_REUSE_IDENTICAL_COMM_GROUPS +if [[ "${CONC}" -ge 16 ]]; then + export ATOM_PREFILL_DECODE_INTERVAL="${ATOM_PREFILL_DECODE_INTERVAL:-4}" + export ATOM_PREFILL_DELAYER_MAX_QUEUE_MS="${ATOM_PREFILL_DELAYER_MAX_QUEUE_MS:-5000}" +fi export ATOM_ENABLE_REPLAYSSM SPEC_TOKENS_FOR_GRAPH=0 if [[ "${NUM_SPECULATIVE_TOKENS}" != "0" ]]; then @@ -197,7 +223,7 @@ if [[ "${NUM_SPECULATIVE_TOKENS}" != "0" ]]; then fi CUDAGRAPH_MAX_NUM_SEQS="${CUDAGRAPH_MAX_NUM_SEQS:-$((2 * CONC))}" GRAPH_MAX=$((CUDAGRAPH_MAX_NUM_SEQS * (1 + SPEC_TOKENS_FOR_GRAPH))) -CUDAGRAPH_CAPTURE_SIZES="[$(seq -s, 2 "${GRAPH_MAX}")]" +CUDAGRAPH_CAPTURE_SIZES="[$(seq -s, 1 "${GRAPH_MAX}")]" ATOM_CMD=( python3 -m atom.entrypoints.openai_server @@ -271,7 +297,7 @@ acceptance flags; ATOM rejects that pair at startup. See [`DSpark.md`](DSpark.md Kimi-K3 KDA decode can rebuild SSM state from a checkpoint ring (`ATOM_ENABLE_REPLAYSSM=1`). The table above is the AgentX default: off for -CONC 1/2/4, on for CONC 8/12/16, off from CONC 32 up. Override with +CONC 1/2/4, on for CONC 8/12/14/16, off from CONC 32 up. Override with `ATOM_ENABLE_REPLAYSSM=0` or `1` when comparing the other setting. ### State checkpointing From 41dc7b8e8172295340aafac2007fc1d59e6ebddf Mon Sep 17 00:00:00 2001 From: Guanbao Yu Date: Thu, 24 Sep 2026 03:38:57 +0000 Subject: [PATCH 2/2] (recipe) Kimi-K3 AgentX: keep the change to env vars and launch args Drop the FP8 prefill precompile prerequisite subsection and restore the section-0 intro. Revert the ReplaySSM paragraph edit. The recipe diff is now limited to ATOM_USE_FLYDSL_FP8_PREFILL_ATTN=1, the C>=16 PrefillDelayer variables and capture sizes starting at 1, plus their table rows. Co-Authored-By: Claude Opus 5.5 (1M context) --- recipes/Agentic-Kimi-K3.md | 21 +++------------------ 1 file changed, 3 insertions(+), 18 deletions(-) diff --git a/recipes/Agentic-Kimi-K3.md b/recipes/Agentic-Kimi-K3.md index 19123a4428..1003f57a39 100644 --- a/recipes/Agentic-Kimi-K3.md +++ b/recipes/Agentic-Kimi-K3.md @@ -82,8 +82,8 @@ other band (C1…C48). ## 0. Container prerequisites -Three things about the container decide whether the numbers above are -reproducible at all. All fail quietly rather than loudly, so they are worth +Two things about the container decide whether the numbers above are +reproducible at all. Both fail quietly rather than loudly, so they are worth checking before the first run. ### triton must be 3.7.x @@ -100,21 +100,6 @@ swapping aiter's `.so` between revisions, swapping the ATOM checkout, and switching between a full CI prebuild and a lean JIT build. Use an image with triton 3.7.x, such as `kimi_k3_agentic_0907`. -### FP8 prefill kernels must be precompiled - -```bash -ls "$(python3 -c 'import aiter, os; print(os.path.dirname(aiter.__file__))')/jit/flydsl_cache" \ - | grep -c '^launch_flash_attn_dualwave_swp_' # expect 92 on a fresh container -``` - -The FlyDSL FP8 prefill attention kernel is compiled per configuration the first -time a request needs it, stalling that request for seconds. aiter builds all 92 -Kimi-K3 variants ahead of time from -`aiter/configs/model_configs/kimik3_fmha_fp8_aot.csv` (ROCm/aiter#5796); a count -below 92 means the image predates that. ATOM must also include the -`merge_attn_states` runtime-arg fix (ROCm/ATOM#2378), without which the MLA -chunked-prefill merge recompiles for every distinct prefill token count. - ### `DRAFT_MODEL_PATH` must point at a local copy ```bash @@ -297,7 +282,7 @@ acceptance flags; ATOM rejects that pair at startup. See [`DSpark.md`](DSpark.md Kimi-K3 KDA decode can rebuild SSM state from a checkpoint ring (`ATOM_ENABLE_REPLAYSSM=1`). The table above is the AgentX default: off for -CONC 1/2/4, on for CONC 8/12/14/16, off from CONC 32 up. Override with +CONC 1/2/4, on for CONC 8/12/16, off from CONC 32 up. Override with `ATOM_ENABLE_REPLAYSSM=0` or `1` when comparing the other setting. ### State checkpointing