Skip to content

[AMD][MI355X] DSv4.1 Flash: enable the vLLM mono decode kernels (image TBD) / [AMD][MI355X] DSv4.1 Flash:启用 vLLM mono decode 内核(镜像待定) - #3690

Open
Fangzhou-Ai wants to merge 5 commits into
mainfrom
amd/dsv41flash-mi355x-a4w4-int4-quickreduce
Open

Fangzhou-Ai wants to merge 5 commits into
mainfrom
amd/dsv41flash-mi355x-a4w4-int4-quickreduce

Conversation

@Fangzhou-Ai

@Fangzhou-Ai Fangzhou-Ai commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator

Enables the DeepSeek-V4.1 mono decode kernels from vllm-project/vllm#60397 in the MI355X DeepSeek-V4.1-Flash vLLM DSpark AgentX recipe (dsv41flash-fp4-mi355x-vllm-agentic-dspark), to measure their end-to-end gain on the full ladder.

Changes

  • VLLM_ROCM_MONO_DECODE=1. The kernels run a decode step of up to 48 rows as two persistent launches a layer, at TP2 and TP4. With five DSpark drafts that is concurrency 1-8; larger steps take the regular path, so c16-c64 should be unchanged.
  • The image is TBD. It will be the first ROCm 10.0 nightly that contains vllm#60397; the sweep runs once that image is published and pinned here.
  • Carried from earlier revisions of this PR:
    • VLLM_ROCM_QUICK_REDUCE_QUANTIZATION=INT4 (vllm-project/recipes#1061);
    • the upstream 16384-token prefill chunk at every point, made possible by vllm#58014;
    • concurrency 1-64 on both TP arms.
  • a4w4 is no longer part of this PR.

For reviewers

  • The automatic GSM8K row is TP4 c64, where a decode step is 384 rows, so it does not exercise the mono path. The kernels' accuracy evidence is in vllm#60397: GPQA-Diamond, AIME25, RULER NIAH up to 256K and GSM8K against the same build with the kernels off, plus GPU numerics tests in vLLM CI.
  • vllm#60397 measured this recipe at c1-c8 against InferenceX's published runs: P90 interactivity 1.30-1.56x at TP2 and 1.38-1.66x at TP4. This sweep re-measures it on the pinned image.
中文

在 MI355X DeepSeek-V4.1-Flash vLLM DSpark AgentX 配置(dsv41flash-fp4-mi355x-vllm-agentic-dspark)中启用 vllm-project/vllm#60397 的 DeepSeek-V4.1 mono decode 内核,以在完整并发梯度上测量其端到端收益。

改动

  • VLLM_ROCM_MONO_DECODE=1:在 TP2 和 TP4 下,每层用两个常驻内核处理不超过 48 行的 decode 步。DSpark 有 5 个草稿 token,对应并发 1-8;更大的步走原有路径,因此 c16-c64 预期不变。
  • 镜像暂为 TBD:将使用第一个包含 vllm#60397 的 ROCm 10.0 nightly,镜像发布并固定后再运行 sweep。
  • 沿用本 PR 之前的设置:
    • VLLM_ROCM_QUICK_REDUCE_QUANTIZATION=INT4(vllm-project/recipes#1061);
    • 所有点使用上游默认的 16384 token prefill 分块(依赖 vllm#58014);
    • 两个 TP 配置的并发均为 1-64。
  • 本 PR 不再包含 a4w4。

评审注意

  • 自动 GSM8K 评测点为 TP4 c64,每个 decode 步 384 行,不经过 mono 路径。内核的精度证据见 vllm#60397:与关闭内核的同一构建对比的 GPQA-Diamond、AIME25、最长 256K 的 RULER NIAH 和 GSM8K,以及 vLLM CI 中的 GPU 数值测试。
  • vllm#60397 在 c1-c8 上将本配置与 InferenceX 已发布结果对比:P90 交互性在 TP2 提升 1.30-1.56 倍,在 TP4 提升 1.38-1.66 倍。本次 sweep 将在固定镜像上重新测量。
AI model disclosure

Prepared with Claude Opus 5.5 (claude-opus-5-5) in Claude Code: the recipe, master-config and changelog edits, the merge with main, and this description. Earlier revisions of this PR were prepared with Claude Opus 5.5 through a Cursor agent.

@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps run only when the PR appends an inferencex-e2e/perf-changelog.yaml entry and carries exactly one primary label: full-sweep-fail-fast (strongly recommended; canary plus per-matrix fail-fast), full-sweep-enabled (canary; matrix jobs continue after a failure), or non-canary-full-sweep-enabled (no canary or fail-fast). The modifiers all-evals, evals-only, and agentx-fast require a primary label. On fork PRs, a maintainer applies the label. See sweep labels and reuse.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**只有当 PR 在 inferencex-e2e/perf-changelog.yaml 末尾追加了条目,并且恰好带有一个主标签时,才会运行扫描:full-sweep-fail-fast(强烈推荐;canary 加逐矩阵 fail-fast)、full-sweep-enabled(有 canary;矩阵任务在失败后继续运行)或 non-canary-full-sweep-enabled(无 canary,也无 fail-fast)。修饰标签 all-evals、evals-only 和 agentx-fast 必须与主标签一起使用。fork PR 的标签由维护者添加。参见扫描标签与复用。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

Fangzhou-Ai added a commit that referenced this pull request Oct 3, 2026
填写 perf-changelog 中 #3690 的 PR 链接。

Co-authored-by: Cursor <cursoragent@cursor.com>
@Fangzhou-Ai Fangzhou-Ai added the full-sweep-enabled Full sweep with canary gate; matrix jobs run to completion despite failures label Oct 3, 2026

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Findings marked 🟡 are optional suggestions and need no follow-up push.

Comment on lines +65 to +66
# Native DeepSeek-V4.1 a4w4 support from vllm-project/vllm#58819.
VLLM_ROCM_USE_AITER_MOE_A4W4_DSV4: '1'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 Setting VLLM_ROCM_USE_AITER_MOE_A4W4_DSV4=1 (agentic.yaml:66) drops MoE activation precision from a8w4 to a4w4, and the recipe's own changelog says this reaches the "shared DSpark MoE layers" used by the draft (perf-changelog.yaml:9283). The old comment at line 42 confirms the prior baseline used CK a8w4; this submission-added env var lowers activations to 4-bit for those same layers when DSpark verifies, which CONTRIBUTING.md's Draft-model precision rule forbids: "Online or offline quantization of draft ... activations ... below the shipped precision, whether through a flag, environment variable" (CONTRIBUTING.md:77-80), and the baseline for comparison is the same image with no draft-related settings from the submission (CONTRIBUTING.md:40-42). …

Why this was flagged

…Matching the #3571 GSM8K baseline does not exempt this per CONTRIBUTING.md's own text, since draft-precision changes shift accuracy loss to acceptance rate, which GSM8K does not measure. Fix: either confine the a4w4 opt-in so DSpark's shared MoE path stays at the shipped baseline precision (a8w4) or obtain an explicit CODEOWNER draft-precision waiver per CONTRIBUTING.md before enabling it process-wide.

Trigger: the role env sets VLLM_ROCM_USE_AITER_MOE_A4W4_DSV4=1 (agentic.yaml:66) on a DSpark recipe where speculative-config uses method dspark (agentic.yaml:56) and DSpark verification reuses the target's MoE weights (changelog says 'shared DSpark MoE layers', perf-changelog.yaml:9283). Before this diff, moe-backend: aiter resolved to the CK a8w4 kernel (old comment, now replaced at agentic.yaml:42); after this diff the same layers run a4w4, i.e. activations drop from 8-bit to 4-bit for both target and draft tokens.

Verification: normal; acknowledged in diff. This is a genuine Draft-model precision policy concern that the base branch does not have.

Decisive facts from the diff itself:

  • agentic.yaml:66 adds VLLM_ROCM_USE_AITER_MOE_A4W4_DSV4: '1' to the role env of a DSpark (speculative-decoding) recipe (speculative-config: '{"method":"dspark",...}', agentic.yaml:56).
  • The old comment at line 42 ("The plain name lets the CK a8w4 MoE kernel win vLLM's priority list") establishes the pre-change out-of-the-box default was CK a8w4 (8-bit activations).

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@chunfangamd can we get a waiver on that? Or I can simply remove the a4w4 flag here.

Comment thread inferencex-e2e/perf-changelog.yaml Outdated
@github-actions

github-actions Bot commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

@Fangzhou-Ai

Copy link
Copy Markdown
Collaborator Author

@Fangzhou-Ai

Copy link
Copy Markdown
Collaborator Author

Closing this PR as a4w4 doesn't show significant perf gain. and a lower precision of draft model requires extra waiver.

1 similar comment
@Fangzhou-Ai

Copy link
Copy Markdown
Collaborator Author

Closing this PR as a4w4 doesn't show significant perf gain. and a lower precision of draft model requires extra waiver.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code review found no new issues

No new issues were found in this update; 1 finding from earlier reviews is still open above.

This review covers commit 144ab45, which is no longer the latest commit on this pull request; later commits are not covered by it.

Fangzhou-Ai added a commit that referenced this pull request Oct 5, 2026
perf-changelog.yaml is rebuilt as origin/main byte for byte, then the
#3690 and #3735 entries are appended at the tail. This keeps the
GLM-5.2 #3616 entry's config-keys header, which the #3690 merge commit
had folded into the preceding pr-link scalar.

perf-changelog.yaml 按字节重建为 origin/main 的内容,再在末尾追加 #3690
与 #3735 两条记录。这样可保留 GLM-5.2 #3616 条目的 config-keys 头,该头
曾被 #3690 的合并提交并入前一条的 pr-link 标量。

Co-authored-by: Cursor <cursoragent@cursor.com>
Fangzhou-Ai added a commit that referenced this pull request Oct 5, 2026
Resolve perf-changelog.yaml to origin/main; the #3690 entry is re-appended
at the tail in the next commit.

将 perf-changelog.yaml 解析为 origin/main 的版本;#3690 的记录将在下一个提交
中重新追加到文件末尾。

Co-authored-by: Cursor <cursoragent@cursor.com>
Fangzhou-Ai added a commit that referenced this pull request Oct 5, 2026
Replace --max-num-batched-tokens 16384 with 8192 on the thirteen TP2/TP4
points that still carried it; TP2 c64 was already 8192. The chunk is now
uniform across the arm, matching the B300 arm and the B200 TP4 arm from
concurrency 8 up.

A prefill chunk scheduled into the same engine step as a decode batch
stalls every decoding sequence for the length of that chunk, so the
stall scales with the chunk. On the two TP4 ep1 arms, the p90-minus-p50
ITL spread is 10.81 ms at c64 on MI355X at 16384 against 1.74 ms on B200
at 8192.

Re-append the #3690 changelog entry at the tail. Its merge commit had
folded the config-keys header of the GLM-5.2 #3616 entry into the
preceding pr-link scalar, which re-keyed #3616 and left #3690's own
entry unparseable.

将仍为 16384 的十三个 TP2/TP4 数据点的 --max-num-batched-tokens 改为 8192;
TP2 c64 原本已是 8192。该分块现在在整个臂上统一,与 B300 臂以及并发 8 及以上
的 B200 TP4 臂一致。

与 decode 批次调度进同一个引擎步的 prefill 分块,会让所有正在解码的序列停顿
该分块的时长,因此停顿时间随分块大小增长。在两个 TP4 ep1 臂上,c64 的 ITL
p90 减 p50 差值在 16384 的 MI355X 上为 10.81 ms,而在 8192 的 B200 上为 1.74 ms。

同时在文件末尾重新追加 #3690 的 changelog 记录。其合并提交曾把 GLM-5.2 #3616
条目的 config-keys 头并入前一条的 pr-link 标量,导致 #3616 被错误归键,
且 #3690 自身的记录无法被解析。

Co-authored-by: Cursor <cursoragent@cursor.com>
Fangzhou-Ai added a commit that referenced this pull request Oct 6, 2026
Merges the updated #3690 branch, so this branch again differs from it by
exactly one thing. Narrows that difference: --max-num-batched-tokens 8192
now applies only to override_tp4_c32, override_tp4_c64 and
override_tp2_c32 (override_tp2_c64 was already 8192); concurrency 1-16
returns to 16384. The 14-point A/B in run 37375212446 against baseline run
37080504692 found gains only at c32 and above (TP2 c32 ITL p90 -14.4%,
TP4 c64 -27.8%) and repeated TTFT regressions below it (TP4 c4 +29.5%,
TP2 c4 +18.8%), matching the profiled chunked-prefill step counts.

perf-changelog.yaml conflicted at its append-only tail; resolved to main's
bytes verbatim plus the #3690 and #3735 entries.

合并更新后的 #3690 分支,使本分支与其重新只差一处,并收窄这处差异:
--max-num-batched-tokens 8192 现在仅作用于 override_tp4_c32、
override_tp4_c64 与 override_tp2_c32(override_tp2_c64 本就是 8192),
并发 1-16 回到 16384。运行 37375212446 对基线运行 37080504692 的 14 点
A/B 显示收益只出现在并发 32 及以上(TP2 c32 的 ITL p90 -14.4%,TP4 c64
-27.8%),低于该并发则 TTFT 反复恶化(TP4 c4 +29.5%,TP2 c4 +18.8%),
与 profile 中分块 prefill 步的计数一致。

perf-changelog.yaml 在仅可追加的尾部产生冲突,解析为逐字节保留 main 的
内容并追加 #3690 与 #3735 两个条目。

Co-authored-by: Cursor <cursoragent@cursor.com>
@Fangzhou-Ai Fangzhou-Ai added the full-sweep-enabled Full sweep with canary gate; matrix jobs run to completion despite failures label Oct 6, 2026
@Fangzhou-Ai

Copy link
Copy Markdown
Collaborator Author

a4w4 dropped from this branch; the ladder now sweeps a8w4

Run 37424787697 on this image failed all nine TP=2 points and passed all nine TP=4 points. Root cause, reproduced locally on MI355X and bisected with a TP2/TP4 A/B on one build:

mxfp4_round_up_hidden_size_and_intermediate_size relaxes AITER_MXFP4_BF16 SiLU shards to 128-alignment (added for Kimi-K3's A16W4 SiTU kernel). The a4w4 opt-in rides the same backend and activation and inherits that relaxation, but AITER's MXFP4 kernels e8m0_shuffle the scales in groups of 8 columns and therefore stride 256 elements of K per group.

DSv4.1-Flash shards moe_intermediate_size 2304 to 1152 at TP=2: 128-aligned but not 256-aligned, so vLLM pads nothing, AITER keeps its inline-sort fast path, and that path addresses the padded stride and walks off the buffers sized on the exact width:

Memory access fault by GPU node-4 on address 0x7f84c9d33000. Reason: Unknown.
[aiter] Error in moe_sorting: CUDA error: an illegal memory access was encountered

TP=4 only escapes because its 576 -> 640 padding makes AITER reject the same path — it logs inline-sort config is unsupported (hidden/intermediate padding) 34440 times per boot, where TP=2 logs it zero times.

What changed here

VLLM_ROCM_USE_AITER_MOE_A4W4_DSV4 is unset, so the experts run a8w4 through the generic AITER CK selector. That is the only change; the ROCm 10.0 repin, INT4 quick reduce, the ladder extension to c192 and all shape overrides are untouched, so this sweep still qualifies the new nightly across the full curve. The recipe comment, both configuration-procedures docs and the changelog entry were updated to match, and origin/main was merged in to clear the conflict that was blocking the sweep from starting.

Fix status

vllm-project/vllm#60273 rounds the a4w4 intermediate to 256. It works (TP=2 boots, GSM8K 0.9629 flexible-extract over all 1319 questions) but costs +11% MoE weights at TP=2, +33% at TP=4, and moves both rows off every tuned AITER row, so it is not the fix this recipe should wait for.

The defect is AITER-side and the cheaper fix is validated locally: _mxfp4_inline_sort_unsupported already has the full fallback machinery wired and already rejects on padding, it just never checks alignment. Adding inter_dim % 256 != 0 there keeps the native 1152 shard, costs zero extra memory, keeps the tuned a4w4 FlyDSL kernels that TP=4 uses today, and scores GSM8K 0.9674 flexible-extract at TP=2 with 0 memory faults. a4w4 returns to this recipe once that lands in an image.

AI assistance was used for the root-cause investigation and this change.

@Fangzhou-Ai Fangzhou-Ai changed the title [AMD][MI355X] Enable DSv4.1 Flash a4w4 and INT4 quick reduce on the latest ROCm 10 nightly / [AMD][MI355X] 在最新 ROCm 10 nightly 上为 DSv4.1 Flash 启用 a4w4 与 INT4 quick reduce [AMD][MI355X] Qualify the latest ROCm 10 nightly for DSv4.1 Flash with INT4 quick reduce / [AMD][MI355X] 在最新 ROCm 10 nightly 上以 INT4 quick reduce 验证 DSv4.1 Flash Oct 6, 2026
Repin the MI355X DeepSeek-V4.1-Flash DSpark recipe to
nightly-rocm100-21d93d0d and retune the ladder around it.

- Leave VLLM_ROCM_USE_AITER_MOE_A4W4_DSV4 unset. The a4w4 opt-in faults
  during graph capture at TP=2; all nine TP=2 points failed and all nine
  TP=4 points passed in run 37424787697. vllm-project/vllm#60273 tracks
  the fix, so the experts stay on the generic AITER CK a8w4 selector.
- Enable VLLM_ROCM_QUICK_REDUCE_QUANTIZATION=INT4 so tensor-parallel
  all-reduces above the quick-reduce size threshold use the INT4 codec.
- Hold both TP rows at concurrency 64. Run 37510565585 measured TP=2
  c128 at 150,683 tok/s/chip against 151,896 at c64 with prompt tokens
  within 0.5%, so the extra concurrency bought real prefill rather than
  scored tokens.
- Set VLLM_SHARED_EXPERTS_STREAM_TOKEN_THRESHOLD=1024. Upstream caps the
  shared-expert overlap at 256 tokens, which predates speculative
  decoding: five DSpark drafts make a decode step CONC x 6 tokens, so
  c64 submits 384 and the overlap never engages above c42. 1024 covers
  the ladder and matches the ATOM default, and stays below the prefill
  chunk so memory profiling never takes the overlap.
- Restore the upstream 16384 prefill chunk at TP2 c64, making the chunk
  uniform across all 14 points. That point had capped it at 8192 to buy
  KV room, but the room was taken by a profiling bug rather than by the
  chunk: the ROCm paged MXFP4 indexer sizes its decode logits workspace
  as (rows, max-model-len), and with no ambient vLLM config in scope the
  row count fell back to the batched-token count, so at this recipe's 1M
  context the reservation scaled linearly with the chunk.
  vllm-project/vllm#58014 bounds rows by max-num-seqs x (1 + 5 drafts),
  and the repinned image carries that fix.

Rebased onto main: the configuration-procedure guides this branch had
edited were removed by #3789, so those edits are dropped.

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Fangzhou-Ai and others added 2 commits October 10, 2026 05:09
- Set VLLM_ROCM_MONO_DECODE=1 (vllm-project/vllm#60397): two persistent
  launches a layer for decode steps of up to 48 rows; larger steps take
  the regular path.
- Image TBD until the first ROCm 10.0 nightly with #60397 is published.
- Drop the a4w4 notes and shorten the comments and the changelog entry.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Fangzhou-Ai Fangzhou-Ai changed the title [AMD][MI355X] Qualify the latest ROCm 10 nightly for DSv4.1 Flash with INT4 quick reduce / [AMD][MI355X] 在最新 ROCm 10 nightly 上以 INT4 quick reduce 验证 DSv4.1 Flash [AMD][MI355X] DSv4.1 Flash: enable the vLLM mono decode kernels (image TBD) / [AMD][MI355X] DSv4.1 Flash:启用 vLLM mono decode 内核(镜像待定) Oct 10, 2026
@Fangzhou-Ai Fangzhou-Ai removed the full-sweep-enabled Full sweep with canary gate; matrix jobs run to completion despite failures label Oct 10, 2026
Fangzhou-Ai and others added 2 commits October 10, 2026 05:12
Keep the upstream default for simplicity.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Replace the TBD placeholder with nightly-rocm100-7d0b4e57, the first ROCm
10.0 nightly built after vllm-project/vllm#60397 merged, so the
VLLM_ROCM_MONO_DECODE=1 setting on this recipe has kernels to run.

将 TBD 占位符替换为 nightly-rocm100-7d0b4e57,这是 vllm-project/vllm#60397
合并后构建的首个 ROCm 10.0 nightly,使本 recipe 中的 VLLM_ROCM_MONO_DECODE=1
设置有对应的 kernel 可用。

Co-authored-by: Cursor <cursoragent@cursor.com>
@Fangzhou-Ai Fangzhou-Ai added the full-sweep-enabled Full sweep with canary gate; matrix jobs run to completion despite failures label Oct 10, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

full-sweep-enabled Full sweep with canary gate; matrix jobs run to completion despite failures

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

3 participants