Skip to content

[AMD] Pack Qwen3.5 GDN input projections on ROCm - #39902

Merged
HaiShaw merged 4 commits into
sgl-project:mainfrom
zijiecode:pr/qwen35-gdn-packed-rocm
Sep 22, 2026
Merged

HaiShaw merged 4 commits into
sgl-project:mainfrom
zijiecode:pr/qwen35-gdn-packed-rocm

Conversation

@zijiecode

@zijiecode zijiecode commented Sep 17, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

The ROCm path runs in_proj_qkvz and in_proj_ba separately. Combining them into one packed AITER GEMM reduces launch and input-quantization overhead for small batches. At large M, the packed GEMM can be less efficient than the separate projections, so this PR limits the packed path to M <= 64. Here M is the number of token rows in the current projection, not the configured concurrency.

Modifications

  • Enable the ROCm packed path only for Qwen3.5 dense and MoE models. Other models sharing the GDN classes keep their existing behavior.
  • Pack the QKVZ and BA weights and their FP8 scales before CUDA Graph capture, without requantizing the checkpoint weights.
  • For supported batches with M <= 64, quantize the input once, run one GEMM, and split the output into QKVZ and BA. Larger batches keep the separate projections.
  • Preserve the guards for unsupported layouts and LoRA. FP8 packing is disabled with pipeline parallelism, and online weight updates are rejected while the derived FP8 cache is active.

Accuracy Tests

amd/Qwen3.5-397B-A17B-MXFP4-AttnFP8-V2, 5-shot GSM8K, chat-completions, 16384-token output budget, temperature 0, and real MTP.

Configuration Strict-match Flexible-extract
Baseline 97.88% 97.80%
This PR 97.95% 97.95%

Benchmarking and Profiling

AttnFP8-V2 on MI355X, TP4/EP1, C4, no KV offload, recv60, and MTP with 3 steps / 4 draft tokens and simulated acceptance length 3.39. Each configuration ran AgentX for 3600 seconds with 10 warmup requests per lane. The baseline uses separate projections; this PR retains the 64-token gate.

Metric Baseline This PR Change
Standard p90 interactivity (tok/s/user) 263.34 278.73 +5.84%
Total throughput (tok/s/chip) 6209.85 6280.55 +1.14%
Output throughput (tok/s/chip) 68.96 70.14 +1.71%
TPOT mean (ms/token) 3.304 3.111 -5.83%
TPOT p50 (ms/token) 3.150 3.000 -4.77%
TPOT p90 (ms/token) 3.797 3.588 -5.52%

Checklist

  • Format your code according to the Format code with pre-commit guide.
  • Add unit tests.
  • Update documentation.
  • Provide accuracy and speed benchmark results.
  • Follow the SGLang code style guidance.

CI States

Latest PR Test (Base): ✅ Run #35553508320
Latest PR Test (Extra): ❌ Run #35553508147
Latest PR Test (AMD ROCm 10): ❌ Run #35553508206

@yichiche yichiche added the run-ci CI: run the baseline test suite on this PR label Sep 17, 2026

@yichiche yichiche left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Guard by if_hip. LGTM.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants