Skip to content

[Perf] Pack the Qwen3.5 GDN input projection on CUDA, including from the dense wrapper - #42126

Open
SamMausberg wants to merge 2 commits into
sgl-project:mainfrom
SamMausberg:perf/qwen3-5-packed-gdn-in-proj-cuda
Open

SamMausberg wants to merge 2 commits into
sgl-project:mainfrom
SamMausberg:perf/qwen3-5-packed-gdn-in-proj-cuda

Conversation

@SamMausberg

@SamMausberg SamMausberg commented Oct 1, 2026 •

Copy link
Copy Markdown

Motivation

Qwen3_5GatedDeltaNet.finalize_fused_in_proj packs in_proj_qkvz and in_proj_ba into one weight. On CUDA, _forward_input_proj then runs one GEMM on the packed weight for up to 1,024 rows, instead of two (qwen3_5.py#L798-L812). For Qwen3.5, the packing runs only from Qwen3_5ForCausalLM.prepare_before_cuda_graph_capture (#L1758-L1767). Two things keep it from running for the dense checkpoints on CUDA:

  1. The model runner calls that hook on the top-level model (base_runner.py#L260-L264). Qwen3_5MoeForConditionalGeneration and qwen3_5_text.Qwen3_5ForCausalLM pass it on to their backbone. Qwen3_5ForConditionalGeneration does not, and it is the class the published dense checkpoints declare (Qwen/Qwen3.5-4B, for example).
  2. The backbone hook packs only under aiter. The CUDA branch of _forward_input_proj came with support qwen 3.8 flash next #37500. Today only Qwen4-Exp reaches it, and Qwen4-Exp packs at load.

The first point also affects ROCm: the packing that #39902 added for dense Qwen3.5 does not run through the dense wrapper. That part is from code reading; I have no ROCm hardware to test it on.

Modifications

  • Qwen3_5ForConditionalGeneration.prepare_before_cuda_graph_capture passes the call on to its backbone, as the MoE wrapper does.
  • The backbone hook packs on CUDA too, for the same model types (qwen3_5_text, qwen3_5_moe_text). _QWEN3_5_ROCM_PACKED_MODEL_TYPES is renamed _QWEN3_5_PACKED_MODEL_TYPES. Qwen4-Exp still packs at load. The existing checks in finalize_fused_in_proj still apply, so on CUDA only unquantized BF16 projections without bias are packed, and nothing is packed with LoRA.
  • On CUDA, finalize_fused_in_proj keeps the two GEMMs when the packed width per rank is not a multiple of 8. cuBLAS picks slower kernels when the packed BF16 output rows are not 16-byte aligned. Among the Qwen3.5 models this affects 0.8B, 2B and 27B at TP 8.
  • Three CPU test cases in test_qwen3_5_pipeline_parallel.py, described below.
The alignment check and the tests
  • The packed width per rank is the in_proj_qkvz rows plus the in_proj_ba rows. For 0.8B, 2B and 27B at TP 8 it is 1,028, 1,028 and 2,060. There the packed GEMM was 1.03, 1.33 and 1.61 times as slow as the pair (geometric mean over M; up to 2.46 times). Its in_proj_qkvz output was not bitwise equal to the separate GEMM's at any M.
  • test_dense_wrapper_packs_on_cuda checks that the hook, called on either wrapper with CUDA, packs the projection. It also checks that the packed path splits its output into the same qkvz and ba as the separate projections. On main it fails for both wrappers: the dense one has no hook to call, and the MoE one passes the call on but nothing is packed on CUDA.
  • test_packing_needs_cuda_or_aiter_and_qwen3_5 covers the negative cases: nothing is packed without CUDA or aiter, or for qwen4_exp_text. On main it fails only because the dense wrapper has no hook.
  • test_cuda_keeps_separate_gemms_for_unaligned_packed_rows uses the per-rank shapes of 0.8B at TP 8 and fails without the alignment check.

Accuracy Tests

  • Served, Qwen/Qwen3.5-4B with this change gives the same tokens and the same top-5 logprobs as main on all 160 prompts, both one request at a time and in batches of 16.
  • Per GEMM, for Qwen3.5-0.8B and 2B at TP 1-4, and for 4B and 35B-A3B at every TP, every output element of the packed GEMM is bitwise equal to the separate GEMMs' at every M.
  • For hidden size 3,072 and above (9B, 27B, 122B-A10B and 397B-A17B), the in_proj_ba part can differ from the separate GEMM by up to 0.031. That is because cuBLAS picks a different kernel for that small GEMM on its own.
Setup and results

Per GEMM, in isolation on a GH200: packed against separate, with in_proj_ba on a side stream as under CUDA graphs. The run covers the eight Qwen3.5 GDN shapes at TP 1, 2, 4 and 8 and 23 values of M from 1 to 1,024, with random weights streamed from HBM. The results above are with the alignment check. The configurations with hidden size 3,072 and above are bitwise at 3-12 of the 23 M values. Their in_proj_qkvz part is bitwise except at one M each for 9B at TP 4 and 8 and 27B at TP 4.

An earlier isolated benchmark of Qwen3.5-4B with its real weights found the same bitwise equality at power-of-two M from 1 to 1,024 (results).

The served runs used this change before the alignment check was added. The check does not apply to Qwen3.5-4B, whose packed width is 12,352 rows.

Served, Qwen/Qwen3.5-4B on one GH200, radix cache off, greedy, 160 prompts (80 MT-Bench, 80 GSM8K), 256 new tokens, main (40e1bb0) against this change:

Prompts with identical tokens Prompts with identical top-5 logprobs
one request at a time 160/160 160/160
batches of 16 160/160 160/160

As a control, main against itself in batches of 16 also gives 160/160 for both. Each pair of servers ran at --mem-fraction-static 0.25 --max-running-requests 16 with 16 mamba slots, and every batch of 16 ran together.

Speed Tests and Profiling

Served, the gain is small: +0.6% from concurrency 64, and within the spread between the two main runs at 16 and 32. At concurrency 1 the packed path is 1.2% slower, which is outside that spread. That is so even though the packed GEMM on its own was 1-2% faster at M = 1. I have not traced where the served loss at one request comes from. A lower bound of 64 rows on the CUDA branch would keep single-request decoding on the separate GEMMs; I can add it if you prefer that trade-off.

Setup and results

Per GEMM, same run as above: across the 29 packed configurations, the packed time over the separate time has a geometric mean over M of 0.87-0.99. It is 0.90-1.03 for M up to 64 and 0.78-0.97 above. The worst single points are 1.18 (9B at TP 8, M = 100 and 333) and 1.17 (27B at TP 4, M = 16). For Qwen3.5-4B at TP 1 the ratio is 0.98-0.99 at M 1-8, up to 1.03 at M 12-48, and 0.84-0.95 from M 64.

Served, main (40e1bb0) against this change, run in the order main, packed, packed, main (A-B-B-A):

  • Setup: Qwen/Qwen3.5-4B on one GH200, plain decoding, FlashInfer attention, CUDA graphs and the overlap scheduler on, --mem-fraction-static 0.80 --max-running-requests 128.
  • Workload: bench_serving with random 1,024-token inputs and 256-token outputs, 16 prompts at concurrency 1 and four times the concurrency above that.
  • Load: other processes used 1.2-1.6 CPU cores on average during each run.

Output tokens per second:

Concurrency A1 main B1 packed B2 packed A2 main Packed vs main (means) A1-A2 spread
1 274.5 271.9 271.8 275.8 -1.2% 0.5%
16 1729.3 1741.4 1723.3 1733.9 +0.0% 0.3%
32 3053.6 3099.7 3074.4 3084.4 +0.6% 1.0%
64 6642.8 6693.7 6705.7 6671.4 +0.6% 0.4%
128 8009.7 8036.2 8070.0 7993.3 +0.6% 0.2%

Checklist


CI States

Latest PR Test (Base): ❌ Run #36921850209
Latest PR Test (Extra): ❌ Run #36921849767
Latest PR Test (AMD ROCm 10): ❌ Run #36921849933

Qwen3_5ForConditionalGeneration did not forward prepare_before_cuda_graph_capture
to its backbone, so the published dense checkpoints never reached the packing
hook, and the hook packed in_proj_qkvz and in_proj_ba only under aiter. Forward
the hook as the MoE wrapper does, and pack on CUDA too, where _forward_input_proj
already runs the packed weight for up to 1,024 rows.
…are unaligned

When in_proj_qkvz and in_proj_ba together have a row count that is not a
multiple of 8 per rank, the packed BF16 output rows are not 16-byte aligned and
cuBLAS falls back to slower kernels (Qwen3.5-0.8B, 2B and 27B at TP 8: 1.03x,
1.33x and 1.61x the separate GEMMs on a GH200, up to 2.46x). Skip packing there.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant