Repository navigation
[Perf] Pack the Qwen3.5 GDN input projection on CUDA, including from the dense wrapper - #42126
Open
SamMausberg wants to merge 2 commits into
Open
SamMausberg wants to merge 2 commits into
SamMausberg wants to merge 2 commits into
Conversation
Qwen3_5ForConditionalGeneration did not forward prepare_before_cuda_graph_capture to its backbone, so the published dense checkpoints never reached the packing hook, and the hook packed in_proj_qkvz and in_proj_ba only under aiter. Forward the hook as the MoE wrapper does, and pack on CUDA too, where _forward_input_proj already runs the packed weight for up to 1,024 rows.
…are unaligned When in_proj_qkvz and in_proj_ba together have a row count that is not a multiple of 8 per rank, the packed BF16 output rows are not 16-byte aligned and cuBLAS falls back to slower kernels (Qwen3.5-0.8B, 2B and 27B at TP 8: 1.03x, 1.33x and 1.61x the separate GEMMs on a GH200, up to 2.46x). Skip packing there.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
Qwen3_5GatedDeltaNet.finalize_fused_in_projpacksin_proj_qkvzandin_proj_bainto one weight. On CUDA,_forward_input_projthen runs one GEMM on the packed weight for up to 1,024 rows, instead of two (qwen3_5.py#L798-L812). For Qwen3.5, the packing runs only fromQwen3_5ForCausalLM.prepare_before_cuda_graph_capture(#L1758-L1767). Two things keep it from running for the dense checkpoints on CUDA:Qwen3_5MoeForConditionalGenerationandqwen3_5_text.Qwen3_5ForCausalLMpass it on to their backbone.Qwen3_5ForConditionalGenerationdoes not, and it is the class the published dense checkpoints declare (Qwen/Qwen3.5-4B, for example)._forward_input_projcame with support qwen 3.8 flash next #37500. Today only Qwen4-Exp reaches it, and Qwen4-Exp packs at load.The first point also affects ROCm: the packing that #39902 added for dense Qwen3.5 does not run through the dense wrapper. That part is from code reading; I have no ROCm hardware to test it on.
Modifications
Qwen3_5ForConditionalGeneration.prepare_before_cuda_graph_capturepasses the call on to its backbone, as the MoE wrapper does.qwen3_5_text,qwen3_5_moe_text)._QWEN3_5_ROCM_PACKED_MODEL_TYPESis renamed_QWEN3_5_PACKED_MODEL_TYPES. Qwen4-Exp still packs at load. The existing checks infinalize_fused_in_projstill apply, so on CUDA only unquantized BF16 projections without bias are packed, and nothing is packed with LoRA.finalize_fused_in_projkeeps the two GEMMs when the packed width per rank is not a multiple of 8. cuBLAS picks slower kernels when the packed BF16 output rows are not 16-byte aligned. Among the Qwen3.5 models this affects 0.8B, 2B and 27B at TP 8.test_qwen3_5_pipeline_parallel.py, described below.The alignment check and the tests
in_proj_qkvzrows plus thein_proj_barows. For 0.8B, 2B and 27B at TP 8 it is 1,028, 1,028 and 2,060. There the packed GEMM was 1.03, 1.33 and 1.61 times as slow as the pair (geometric mean over M; up to 2.46 times). Itsin_proj_qkvzoutput was not bitwise equal to the separate GEMM's at any M.test_dense_wrapper_packs_on_cudachecks that the hook, called on either wrapper with CUDA, packs the projection. It also checks that the packed path splits its output into the sameqkvzandbaas the separate projections. On main it fails for both wrappers: the dense one has no hook to call, and the MoE one passes the call on but nothing is packed on CUDA.test_packing_needs_cuda_or_aiter_and_qwen3_5covers the negative cases: nothing is packed without CUDA or aiter, or forqwen4_exp_text. On main it fails only because the dense wrapper has no hook.test_cuda_keeps_separate_gemms_for_unaligned_packed_rowsuses the per-rank shapes of 0.8B at TP 8 and fails without the alignment check.Accuracy Tests
in_proj_bapart can differ from the separate GEMM by up to 0.031. That is because cuBLAS picks a different kernel for that small GEMM on its own.Setup and results
Per GEMM, in isolation on a GH200: packed against separate, with
in_proj_baon a side stream as under CUDA graphs. The run covers the eight Qwen3.5 GDN shapes at TP 1, 2, 4 and 8 and 23 values of M from 1 to 1,024, with random weights streamed from HBM. The results above are with the alignment check. The configurations with hidden size 3,072 and above are bitwise at 3-12 of the 23 M values. Theirin_proj_qkvzpart is bitwise except at one M each for 9B at TP 4 and 8 and 27B at TP 4.An earlier isolated benchmark of Qwen3.5-4B with its real weights found the same bitwise equality at power-of-two M from 1 to 1,024 (results).
The served runs used this change before the alignment check was added. The check does not apply to Qwen3.5-4B, whose packed width is 12,352 rows.
Served, Qwen/Qwen3.5-4B on one GH200, radix cache off, greedy, 160 prompts (80 MT-Bench, 80 GSM8K), 256 new tokens, main (40e1bb0) against this change:
As a control, main against itself in batches of 16 also gives 160/160 for both. Each pair of servers ran at
--mem-fraction-static 0.25 --max-running-requests 16with 16 mamba slots, and every batch of 16 ran together.Speed Tests and Profiling
Served, the gain is small: +0.6% from concurrency 64, and within the spread between the two main runs at 16 and 32. At concurrency 1 the packed path is 1.2% slower, which is outside that spread. That is so even though the packed GEMM on its own was 1-2% faster at M = 1. I have not traced where the served loss at one request comes from. A lower bound of 64 rows on the CUDA branch would keep single-request decoding on the separate GEMMs; I can add it if you prefer that trade-off.
Setup and results
Per GEMM, same run as above: across the 29 packed configurations, the packed time over the separate time has a geometric mean over M of 0.87-0.99. It is 0.90-1.03 for M up to 64 and 0.78-0.97 above. The worst single points are 1.18 (9B at TP 8, M = 100 and 333) and 1.17 (27B at TP 4, M = 16). For Qwen3.5-4B at TP 1 the ratio is 0.98-0.99 at M 1-8, up to 1.03 at M 12-48, and 0.84-0.95 from M 64.
Served, main (40e1bb0) against this change, run in the order main, packed, packed, main (A-B-B-A):
--mem-fraction-static 0.80 --max-running-requests 128.bench_servingwith random 1,024-token inputs and 256-token outputs, 16 prompts at concurrency 1 and four times the concurrency above that.Output tokens per second:
Checklist
CI States
Latest PR Test (Base): ❌ Run #36921850209
Latest PR Test (Extra): ❌ Run #36921849767
Latest PR Test (AMD ROCm 10): ❌ Run #36921849933