Repository navigation
[AMD] Pack Qwen3.5 GDN input projections on ROCm - #39902
Merged
Merged
Conversation
zijiecode
requested review from
Fridge003,
Ying1123,
hnyls2002,
ispobock and
merrymercy
as code owners
September 17, 2026 04:53
yichiche
approved these changes
Sep 18, 2026
yichiche
left a comment
Collaborator
There was a problem hiding this comment.
Guard by if_hip. LGTM.
4 of 5 tasks
4 of 5 tasks
HaiShaw
approved these changes
Sep 22, 2026
This was referenced Sep 23, 2026
This was referenced Sep 24, 2026
4 of 5 tasks
2 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
The ROCm path runs
in_proj_qkvzandin_proj_baseparately. Combining them into one packed AITER GEMM reduces launch and input-quantization overhead for small batches. At large M, the packed GEMM can be less efficient than the separate projections, so this PR limits the packed path toM <= 64. Here M is the number of token rows in the current projection, not the configured concurrency.Modifications
M <= 64, quantize the input once, run one GEMM, and split the output into QKVZ and BA. Larger batches keep the separate projections.Accuracy Tests
amd/Qwen3.5-397B-A17B-MXFP4-AttnFP8-V2, 5-shot GSM8K, chat-completions, 16384-token output budget, temperature 0, and real MTP.Benchmarking and Profiling
AttnFP8-V2 on MI355X, TP4/EP1, C4, no KV offload, recv60, and MTP with 3 steps / 4 draft tokens and simulated acceptance length 3.39. Each configuration ran AgentX for 3600 seconds with 10 warmup requests per lane. The baseline uses separate projections; this PR retains the 64-token gate.
Checklist
CI States
Latest PR Test (Base): ✅ Run #35553508320
Latest PR Test (Extra): ❌ Run #35553508147
Latest PR Test (AMD ROCm 10): ❌ Run #35553508206