Skip to content

(recipe) Kimi-K3 AgentX: FP8 prefill attention, PrefillDelayer from C16, capture from bs=1 - #2382

Merged
zhuyuhua-v merged 2 commits into
mainfrom
recipe/kimik3-agentx-fp8-prefill-delayer
Sep 24, 2026
Merged

zhuyuhua-v merged 2 commits into
mainfrom
recipe/kimik3-agentx-fp8-prefill-delayer

Conversation

@gbyu-amd

@gbyu-amd gbyu-amd commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

What

This updates recipes/Agentic-Kimi-K3.md to the configuration of a 7-point AgentX sweep on 8×MI355X (C1/4/14/16/48/56/72, 3600 s each):

  • ATOM_USE_FLYDSL_FP8_PREFILL_ATTN=1 on every band (FlyDSL FP8 prefill attention).
  • ATOM_PREFILL_DECODE_INTERVAL=4 and ATOM_PREFILL_DELAYER_MAX_QUEUE_MS=5000 on CONC ≥ 16 (the PrefillDelayer: 4 protected decode passes after each prefill, plus prefill coalescing bounded by a 5 s queue age). ATOM_ENABLE_PREFILL_DELAYER already defaults to 1, so these two variables are enough to construct the PrefillDelayer. The sweep checked this on every C ≥ 16 point: each one logged PrefillDelayer initialized once.
  • CUDA-graph capture sizes start at 1 instead of 2: seq -s, 1, dense [1, …, graph_max].

The recipe changes are limited to these environment variables and launch arguments, plus the matching rows in the configuration table. Nothing else in the launcher or client changes. Every other value (DCP, max-num-seqs, LMCache size, comm-group reuse, ReplaySSM, spec, AIPerf flags and env) already matched the sweep.

Dependencies

Both dependencies are merged: #2378 (2026-09-24) and ROCm/aiter#5796 (2026-09-24). The sweep ran on ATOM d9f0720 + #2378, aiter 2887d4899 + #5796, and triton 3.7.0.

Results vs the current recipe's best measured runs

Deltas are relative to the best local run of the current recipe per band (reuse A/B-aligned). Bold marks a change outside the run-to-run band measured with 5 same-config c64 runs: tok/s ±4.7%, ITL ±6.2%, TTFT ±11.7%.

CONC delayer tot tok/s out tok/s TTFT p50 TTFT p90 ITL p50 ITL p90 interactivity p90
1 – +5.6%¹ +5.3%¹ −3.9% +0.8% −8.5% −12.4% +14.3%
4 – −0.1%¹ −1.9%¹ −0.7% −6.3% −6.3% −13.3% +15.4%
14 – +4.0% +4.8% −10.3% −19.3% −2.3% −1.8% +1.9%
16 ✓ +3.9% +3.3% +91.7% +32.9% −3.2% −12.3% +13.9%
48 ✓ +9.7% +10.4% +36.8% +18.9% −13.7% −15.0% +17.6%
56 ✓ +11.9% +15.8% +61.1% +24.8% −16.4% −18.4% +22.1%
72 ✓ +14.7% +17.1% +67.3% +89.1% −16.0% −16.0% +19.0%

¹ At C1/C4 throughput is set by the trace's pacing: both runs served nearly the same requests. Latency is the signal there.

Absolute values of the new runs:

CONC tot tok/s per chip out tok/s TTFT p50 / p90 (ms) ITL p50 / p90 (ms)
1 12,981 1,623 104.5 733 / 1,805 6.28 / 6.57
4 23,074 2,884 168.0 588 / 1,339 6.43 / 8.27
14 59,352 7,419 483.1 696 / 1,638 12.07 / 18.84
16 68,033 8,504 507.5 1,450 / 2,897 13.21 / 19.68
48 102,340 12,792 680.9 1,790 / 4,214 47.12 / 62.38
56 108,821 13,603 772.1 2,050 / 4,395 52.74 / 67.07
72 117,353 14,669 837.5 2,492 / 7,496 74.43 / 100.40

Trade-off to be aware of

The PrefillDelayer protects decode by holding prefills, so TTFT rises wherever it is on. The delayer's own counters show the hold rate climbing with concurrency: C16 31%, C48 38%, C56 51%, C72 51%. At C48–C72 this buys +10–17% output throughput and −14–18% ITL. At C16 the throughput change is inside noise, while TTFT p50 goes from 696 ms (C14, delayer off) to 1,450 ms. So the C ≥ 16 threshold is the sweep's setting, not a tuned one; C16 may be better served without it.

Not measured with the new knobs

C2, C8, C12, C32, C40, C64 and C80 follow the same rules (FP8 prefill everywhere, PrefillDelayer from C16 up) but were not re-run in this sweep. The effects above cannot be split per knob either: the ATOM/aiter changes above, FP8 prefill, the bs=1 capture and the delayer all changed together.

How the recipe was checked

The launcher block from this PR and the sweep's server script were each executed up to the server start for all 14 bands, dumping exported env and argv, then diffed. The only remaining differences are default-equal (ATOM_USE_FLYDSL_GATHER_KV_B_PROJ=1 is the default; 0.90 vs 0.9) or logging-only (AITER_LOG_LEVEL). The AIPerf client, including its AIPERF_* env, is identical to what the sweep ran, apart from URLs and paths.

🤖 Generated with Claude Code

…16, capture from bs=1

Adopt the configuration of the 2026-09-23 AgentX sweep (C1/4/14/16/48/56/72
on 8x MI355X):

- ATOM_USE_FLYDSL_FP8_PREFILL_ATTN=1 on every band
- ATOM_PREFILL_DECODE_INTERVAL=4 and ATOM_PREFILL_DELAYER_MAX_QUEUE_MS=5000
  on CONC >= 16
- CUDA-graph capture sizes start at 1 (seq -s, 1)
- container prerequisite: the 92 FlyDSL FP8 prefill kernels are AOT-built
  (ROCm/aiter#5796) and ATOM carries the merge_attn_states runtime-arg fix
  (#2378), with a one-line check
- ReplaySSM paragraph now lists C14 among the on-bands, matching the table

Against the current recipe's best runs: C48-C72 +10-15% total and +10-17%
output throughput, ITL -14-18%; C1/C4 ITL p90 -12-13%; C14 TTFT p90 -19%.
The delayer raises TTFT from C16 up (p50 +37-92%).

The launcher's exported env and argv match the sweep's server script on
all 14 bands, apart from default-equal or logging-only values.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown
Contributor

🏷️ CI Guide

Runs automatically on every eligible PR before approval:

  • ✅ Pre Checkin: Black, Ruff, catalog schema validation, non-GPU unit tests

Heavy model tests:

  • ✅ Run after the PR is approved and Pre Checkin passes
  • ✅ Run immediately when an approval review is submitted
  • ✅ Can be requested before approval with labels
Label Tests
ci:full Run all heavy PR model tests: native ATOM, vLLM, and SGLang
ci:atom Run native ATOM model accuracy tests
ci:vllm Run ATOM vLLM OOT model accuracy tests
ci:sglang Run ATOM SGLang model accuracy tests

Heavy jobs are skipped when the PR is not approved and no matching ci:* label is present.
Add labels via the sidebar or gh pr edit 2382 --add-label <label>

Drop the FP8 prefill precompile prerequisite subsection and restore the
section-0 intro. Revert the ReplaySSM paragraph edit. The recipe diff is
now limited to ATOM_USE_FLYDSL_FP8_PREFILL_ATTN=1, the C>=16
PrefillDelayer variables and capture sizes starting at 1, plus their
table rows.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gbyu-amd
gbyu-amd requested a review from zhuyuhua-v September 24, 2026 06:04
@zhuyuhua-v
zhuyuhua-v merged commit d8cdd4e into main Sep 24, 2026
1 check passed
@zhuyuhua-v
zhuyuhua-v deleted the recipe/kimik3-agentx-fp8-prefill-delayer branch September 24, 2026 06:07
gbyu-amd added a commit that referenced this pull request Sep 28, 2026
…m bs=1 (#2417)

* fix(ci-mesh): align Kimi-K3 AgentX cells with the recipe (#2382)

Sync the K3 P/D nightly cells with recipes/Agentic-Kimi-K3.md as of
d8cdd4e:

- ATOM_USE_FLYDSL_FP8_PREFILL_ATTN=1 on every band
- PrefillDelayer (ATOM_PREFILL_DECODE_INTERVAL=4,
  ATOM_PREFILL_DELAYER_MAX_QUEUE_MS=5000) from CONC 16 up
- `cudagraph: auto` captures the dense range from bs=1 instead of 2

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(ci-mesh): drop PrefillDelayer env from K3 P/D cells

Its benefit was measured on the aggregated recipe server; on the P/D
split the decode instance never installs it and the prefill instance
has no decode batch to protect, so keep the cells unchanged there.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: ganyi <ygan@amd.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seungrokj added a commit to SemiAnalysisAI/InferenceX that referenced this pull request Sep 29, 2026
…cipe (#3407)

* feat(agentx): bump Kimi-K3 FP4 MI355X ATOM image to 0924 and track recipe

Track recipes/Agentic-Kimi-K3.md as retuned in ROCm/ATOM#2382:
enable FlyDSL FP8 prefill attention on every band, and hold a ready
prefill for four decode passes from concurrency 16 up. The published
concurrency set and every other launch argument are unchanged.

将 MI355X Kimi-K3 FP4 ATOM AgentX 提交切换到 0924 镜像,并跟随
ROCm/ATOM#2382 重调后的 recipe:全部并发开启 FlyDSL FP8 prefill
attention;并发 16 及以上时让就绪的 prefill 等待 4 个 decode 轮次。
已发布的并发点集合与其余启动参数保持不变。

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(agentx): point perf-changelog entry at PR 3407

将 perf-changelog 条目的 pr-link 指向 PR 3407。

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(agentx): bump Kimi-K3 FP4 MI355X ATOM image to nightly_202609251613

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(agentx): restore Kimi-K3 MI355X ATOM LMCache bands on srt-slurm

Carry srt-slurm patch 507 so ATOM aggregate workers accept
extra-kv-connectors, and restore the DCP8 bands from the legacy config
and ROCm/ATOM recipes/Agentic-Kimi-K3.md: concurrency 14 and 16 with
DSpark 3 and ReplaySSM, 48, 56 and 72 without a draft, all on the
in-process lmcache_offload connector (128 GB/rank, 192 GB/rank at 56 and
72). The PrefillDelayer applies from concurrency 16 up.

The single-node adapter now reads ATOM's decode-context-parallel-size
for the DCP_SIZE check, as it does for vLLM.

* chore(agentx): Kimi-K3 FP4 MI355X ATOM image 0925, FlyDSL prefill, prefill delay, DCP8 LMCache

- Move kimik3-fp4-mi355x-atom-agentic-mtp to rocm/atom-dev:nightly_202609251613,
  tracking recipes/Agentic-Kimi-K3.md as retuned in ROCm/ATOM#2382.
- Enable FlyDSL FP8 prefill attention (ATOM_USE_FLYDSL_FP8_PREFILL_ATTN=1) at
  every concurrency.
- From conc 16 up, hold a ready prefill for four decode passes
  (ATOM_PREFILL_DECODE_INTERVAL=4, ATOM_PREFILL_DELAYER_MAX_QUEUE_MS=5000);
  conc 1, 4 and 14 unchanged.
- Restore DCP8 LMCache bands on the native srt-slurm recipe: conc 14/16
  (DSpark 3, ReplaySSM) and 48/56/72 (no draft) use ATOM's in-process
  lmcache_offload connector via roles.agg.args.extra-kv-connectors (srt-slurm
  patch 507), 128 GB/rank up to 48 and 192 GB/rank at 56/72. Conc 1 and 4
  stay GPU-resident.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(agentx): Kimi-K3 FP4 MI355X ATOM image 0925, FlyDSL prefill, prefill delay, DCP8 LMCache

- Move kimik3-fp4-mi355x-atom-agentic-mtp to rocm/atom-dev:nightly_202609251613,
  tracking recipes/Agentic-Kimi-K3.md as retuned in ROCm/ATOM#2382.
- Enable FlyDSL FP8 prefill attention (ATOM_USE_FLYDSL_FP8_PREFILL_ATTN=1) at
  every concurrency.
- From conc 16 up, hold a ready prefill for four decode passes
  (ATOM_PREFILL_DECODE_INTERVAL=4, ATOM_PREFILL_DELAYER_MAX_QUEUE_MS=5000);
  conc 1, 4 and 14 unchanged.
- Restore DCP8 LMCache bands on the native srt-slurm recipe: conc 14/16
  (DSpark 3, ReplaySSM) and 48/56/72 (no draft) use ATOM's in-process
  lmcache_offload connector via roles.agg.args.extra-kv-connectors (srt-slurm
  patch 507), 128 GB/rank up to 48 and 192 GB/rank at 56/72. Conc 1 and 4
  stay GPU-resident.
- Draft model precision unchanged: online_quant_config still excludes every
  Inferact/Kimi-K3-DSpark linear (layers.*, context_proj), so its weights and
  activations stay BF16, and it shares the target's FP8 KV cache
  (kv_cache_dtype fp8). FlyDSL FP8 prefill attention applies only to the
  target; the draft's block pass runs as decode attention.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: seungrokj <144636725+seungrokj@users.noreply.github.com>
Co-authored-by: seungrokj <seungrok.jung@amd.com>
Co-authored-by: Cameron Quilici <cameron@semianalysis.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants