Repository navigation
[AMD] One-launch small-M MoE router for Qwen3.5 on gfx950 - #41133
Merged
Merged
Conversation
At 1-8 tokens the router GEMM, softmax top-k and shared-expert append were three latency-bound launches ahead of the small-M MoE kernel; one kernel now writes the same 11-slot ids and weights. Co-authored-by: Cursor <cursoragent@cursor.com>
4 of 5 tasks
chuyeh
marked this pull request as ready for review
September 28, 2026 05:54
chuyeh
requested review from
BBuf,
DarkSharpness,
HaiShaw,
HydraQYH,
celve and
yuan-luo
as code owners
September 28, 2026 05:54
…router New SGLANG_* flags belong in environ.py; the router's two DPP reductions now share one helper (bit-identical output). Co-authored-by: Cursor <cursoragent@cursor.com>
yichiche
approved these changes
Oct 1, 2026
yichiche
left a comment
Collaborator
There was a problem hiding this comment.
Guarded by _use_aiter checks plus a process-wide disable flag, this fuses the gate GEMM, softmax top-10, and shared-expert append into one launch, used only when M<=8, with a clear perf gain at low concurrency. LGTM.
HaiShaw
approved these changes
Oct 8, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
On MI355X at low concurrency, each Qwen3.5-397B MoE layer routes a verify batch of M = 4 or 8 tokens (4 MTP verify tokens × 1–2 running requests) through three launches before the small-M MoE kernel from #40204:
[M, 4096] x [512, 4096])_router_triton_kernel(softmax over 512 experts, top-10, renormalize; one wave per token doing 10 serial argmax passes)_fused_append_shared_experts_with_weights_kernel(the shared-expert gate GEMV and slot 10)All three are latency-bound at these sizes. Together they take 19.9 µs per layer in a verify step (eager trace, M = 4). Across 60 layers that is about 1.2 ms per verify step, and the MTP layer adds more in draft and draft extend. This PR replaces them with one launch.
Modifications
python/sglang/kernels/ops/moe/smallm_router_gfx950/(__init__.py,smallm_router.hip). It is built withhipccon first use and launched through ctypes, reusing [AMD] Small-M MXFP4 fused-MoE kernel for gfx950 (Qwen) #40204's_hipcc,_hip_lib,_checkand_Kernel. No build-time or wheel change.gate.weightandshared_expert_gate.weight) with fp32 accumulation. The last block to finish, elected by a self-resetting ticket counter so the kernel is graph-safe, rounds the routed logits to bf16, then runs top-10 (lowest id wins ties, as in_router_triton_kernel), softmax and renormalize. It writes slot 10 = (512,sigmoid(shared logit) * scale). The output is exactly thetopk_ids [M, 11] int32/topk_weights [M, 11] fp32thatsmallm_moe_gfx950already consumes.qwen2_moe.py: one guarded call at the top of_forward_router_experts(5 lines). The guard requires gfx95, 1 <= M <= 8, bf16 hidden 4096, gate[512, 4096]and shared gate[1, 4096]in bf16, exactly one fused shared expert, EP 1, the auto/aiter MoE runner, top-10 with renormalize, no EPLB remap or simulated routing, and not inside a piecewise CUDA graph. Everything else takes today's path. The routed-expert capture and expert-distribution recorder hooks still run.SGLANG_ROCM_SMALLM_ROUTER=0turns it off. A build or load failure disables it for the process.test/registered/amd/test_smallm_router_gfx950.py. It covers ids and weights against today'smoe_fused_gate+fused_append_shared_experts_with_weightsat 1 / 3 / 4 / 8 tokens on exactly representable inputs (many ties), fallback at 9 and 41 tokens and for unsupported dtype and shape, the off switch, and graph replay == eager. It passes onv0.5.20-rocm720-mi35x-20260923.Accuracy Tests
Router logits are rounded to bf16 before top-k, as today. Against an fp64-then-bf16 reference on real activations (10,022 token rows), this kernel agrees on 100% of expert choices. Today's path agrees on 98.3%, because its bf16 gate GEMM is not correctly rounded at M >= 2.
amd/Qwen3.5-397B-A17B-MXFP4-AttnFP8-V2(revision e17e5f0), MI355X TP4, InferenceX recipe (FP8 KV cache, EAGLE MTP3 steps / 4 draft tokens, real draft), lm-eval
local-chat-completionswith chat template, 5-shot GSM8K (InferenceXtask), thinking disabled, 16384-token output budget, temperature 0. At most 2 running requests, so every verify batch (<= 8 tokens) is inside the kernel's range.
1417345f5f)GPQA diamond (198 questions, repeat 8, simple-evals grader), same servers, thinking on, temperature 0.6 / top_p 0.95 /
top_k 20, 32768-token output budget:
1417345f5f)Same GSM8K task at TP2 (thinking disabled, at most 2 running requests):
1417345f5f)Benchmarking and Profiling
Kernel microbenchmarks
MI355X, one GPU, HIP-graph replay, 512 MB memset before each call to flush L2 / MALL, routing for one MoE layer, in µs:
In the live server the step saves more than this table suggests: today's
_router_triton_kernelruns at 11.2 µs per layer in verify and 23.6 µs in draft extend (about 7 µs in isolation), while the new kernel costs about the same everywhere.AgentX benchmark
Setup:
rocm/sgl-dev:v0.5.20-rocm720-mi35x-20260923.--kv-cache-dtype fp8_e4m3 --mem-fraction-static 0.80, EAGLE MTP 3 / 1 / 4 with simulated acceptance length 3.39, BF16 MTP layer as in the recipe.Baseline is main
1417345f5f(includes #39901, #39902 and #40204); this PR is the same commit plus this PR.Concurrency 1:
Concurrency 8, average of two full runs per configuration (the second pass in reverse order). This is a no-regression check: verify batches are 4 x running requests, so up to 32 tokens, above this kernel's M <= 8 guard; the router change is inactive once more than two requests run:
TPOT is reported as median and p90 rather than mean: in this replay a few short-output requests wait behind another request's long chunked prefill (up to 4.3 s for a 2-token answer), and a single such request moves the mean of a 3600 s run by up to 3 ms/token. Matching every request across the six concurrency-8 runs (baseline, #41133, #41134, two runs each), no request stalled in both runs of a PR without also stalling in a baseline run.
At TP2 with the official HiCache high-concurrency setup this router runs only at 8 or fewer tokens, so it is idle once more than two requests run. Concurrency 20 TPOT median is 6.94 ms versus 6.97 ms for the baseline. At concurrency 40 the median stays within about 1 ms of the baseline runs (14.7–15.3 ms), and the same tree with
SGLANG_ROCM_SMALLM_ROUTER=0lands on the same median.Decode step time
Concurrency 1, one request, same image and recipe, 15 requests per configuration over three alternating server blocks. Baseline is this PR with
SGLANG_ROCM_SMALLM_ROUTER=0:Checklist
CI States
Latest PR Test (Base): ✅ Run #36722177674
Latest PR Test (Extra): ❌ Run #36722177232
Latest PR Test (AMD ROCm 10): ❌ Run #36722177864