[GDN] Honor configured linear-attn verify backend in the kernel dispatcher - #34592
Merged
Merged
Conversation
GDNKernelDispatcher derived the verify kernel purely from whether decode or prefill selected FlashInfer, silently overriding --linear-attn-verify-backend. On SM90 the FlashInfer MTP verify path (gated_delta_rule_mtp) asserts a fp32 SSM state, so NEXTN speculative decoding could not be combined with --mamba-ssm-dtype bfloat16 at all. An explicitly configured triton verify backend now wins; the auto rule is unchanged otherwise. Qwen3.8-27B-FP8 1xH200 4096in/1024out: MTP(3/1/4) + bf16 SSM state now runs, bs=64 2092.7 -> 2143.4 tok/s (vs previous best per-shape split), bs=1 266.3, GSM8K-500 0.954 (unchanged vs fp32-state gate). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
BBuf
requested review from
Fridge003,
HaiShaw,
Qiaolin-Yu,
hebiao064,
ispobock and
merrymercy
as code owners
August 12, 2026 15:35
saturn-acc
pushed a commit
to saturn-acc/sglang
that referenced
this pull request
Aug 16, 2026
…tcher (sgl-project#34592) Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
5 tasks done
Atituiset
pushed a commit
to Atituiset/sglang
that referenced
this pull request
Sep 10, 2026
…tcher (sgl-project#34592) Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
GDNKernelDispatcherderives its verify kernel purely from whether the decode or prefill backend selected FlashInfer, silently overriding an explicitly configured--linear-attn-verify-backend. The server logs end up contradicting themselves:This matters on SM90: the FlashInfer MTP verify path (
flashinfer.gdn_decode.gated_delta_rule_mtp) asserts a fp32 SSM state, so any GDN model served with--mamba-ssm-dtype bfloat16crashes at startup as soon as NEXTN speculative decoding is enabled — and there is currently no way to force the Triton verify kernel (which handles bf16 state fine) because the flag is ignored:Modifications
GDNKernelDispatchertakes the configured verify backend; an explicittritonchoice now wins, while the existing auto rule (FlashInfer verify when the selected FlashInfer kernel supports MTP verify) is unchanged when the flag is unset.With the fix,
--speculative-algorithm NEXTN+--mamba-ssm-dtype bfloat16+--linear-attn-verify-backend tritoncompose. Measured on a GDN hybrid 27B (FP8) on 1x H200, 4096-in/1024-out: bs=64 2092.7 -> 2143.4 output tok/s vs the previous best non-composable configs, bs=1 266.3 with accept length 2.5-3.1, GSM8K-500 accuracy unchanged (0.954) vs the fp32-state gate.Checklist
🤖 Generated with Claude Code
CI States
Latest PR Test (Base): ❌ Run #31689642600
Latest PR Test (Extra): ❌ Run #31689642504