Repository navigation
[Moe] Fix flashinfer_trtllm silently dropping swiglu_limit clamped SwiGLU activation - #39920
Merged
hnyls2002 merged 1 commit intoSep 17, 2026
Merged
Conversation
Models that plumb the clamped-SwiGLU limit through swiglu_limit (GLM-5 via DeepseekV2MoE, Qwen3-Next) silently lost the clamp on the flashinfer_trtllm FP8 MoE path: it only reads gemm1_clamp_limit, which stays None for these models, so the expert activation degrades to plain SwiGLU. Fall back to swiglu_limit when gemm1_clamp_limit is unset; the TRT-LLM kernel applies it via gemm1_clamp_limit natively. GSM8K, GLM-5.3-Flash FP8, TP4/EP4 on B200: flashinfer_trtllm 92.65-92.87 -> 93.48-93.71 (triton baseline 93.56-93.63).
Dovis01
requested review from
Alisehen,
AniZpZ,
BBuf,
Edwardf0t1,
FlamingoPg,
HaiShaw,
OrangeRedeng,
b8zhong,
ch-wan and
mmangkad
as code owners
September 17, 2026 07:00
Collaborator
|
/rerun-test test_fp8_moe_runner_ownership.py test_moe_runner_extensions.py |
Collaborator
|
/tag-and-rerun-ci |
Contributor
|
Results for 🚀 |
This was referenced Sep 18, 2026
fungaren
pushed a commit
to fungaren/sglang
that referenced
this pull request
Sep 20, 2026
… dropping swiglu_limit clamped SwiGLU activation (sgl-project#39920) (sgl-project#40035) Co-authored-by: Shijin Zhang <75300765+Dovis01@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
Models with a clamped SwiGLU activation (e.g. GLM-5.3-Flash / GLM-5-Next, Qwen3-Next style) plumb the clamp limit through
FusedMoE(swiglu_limit=...)->MoeRunnerConfig.swiglu_limit. All MoE runner backends are supposed to honor it:tritonpasses it into the fused kernel,deep_gemmapplies it viasilu_and_mul_clamp/_apply_swiglu_limit, butflashinfer_trtllmsilently ignores it.The flashinfer_trtllm FP8 path only materializes
gemm1_clamp_limit(the TRT-LLM kernel parameter), which staysNonefor models that use theswiglu_limitfield. With no error and no warning, the expert activation degrades from the trainedsilu(clamp(gate, max=L)) * clamp(up, -L, L)(with L = 10.0) to a plainsilu(gate) * up, which measurably hurts accuracy.On GLM-5.3-Flash FP8 (blockwise 128x128 FP8, 288 routed experts, top-8, 1 shared expert,
swiglu_limit=10.0), GSM8K 5-shot with only--moe-runner-backendvarying:Before/after ranges do not overlap (within-group variance is ~0.2pp); after the fix flashinfer_trtllm matches the triton baseline exactly. fix #39797. Thx for finding this issue. @NolenLiang
Modifications
Single-file change in
python/sglang/srt/layers/quantization/fp8.py,Fp8MoEMethod._prepare_flashinfer_trtllm_activation_params(): whengemm1_clamp_limitis not explicitly set, fall back toswiglu_limit.Why here and not in the model code: this function is the only consumer that builds the TRT-LLM SwiGLU parameters (
gemm1_alpha/beta/clamp_limit), so the two naming pipelines (gemm1_clamp_limitused by gpt_oss / bailing_moe_v3,swiglu_limitused by DeepseekV2MoE-family models) converge at a single point;tritonconsumes both fields independently (gemm1_limitandswiglu_limit), so setting both at the model level would change behavior of unrelated backends, while this fallback is scoped to the flashinfer_trtllm preparation only. Models that already setgemm1_clamp_limitexplicitly (gpt_oss, bailing_moe_v3) are unaffected, and the same silent drop is fixed for every other model that routes the clamp throughswiglu_limit(e.g. Qwen3-Next) whenever flashinfer_trtllm is used.The TRT-LLM kernel natively supports the clamp: with
ActivationType.Swiglu,gemm1_clamp_limit=LcomputesX2 = clamp(X2, max=L),X1 = clamp(X1, -L, L),out = X2 * sigmoid(alpha * X2) * (X1 + beta)(alpha defaults to 1.0, beta to 0.0), which matches the referenceswiglu_clampedinmodels/glm5_next.pyexactly. Verified against flashinfer 0.6.18 (trtllm_fp8_block_scale_moe,Fp8QuantizationType.DeepSeekFp8).Accuracy Tests
Model:
zai-org/GLM-5.3-Flash(FP8, blockwise 128x128), 4x B200, TP=4/EP=4. Launch (only--moe-runner-backendvaries between runs):Eval (lm-evaluation-harness,
local-completions):lm_eval --model local-completions \ --model_args "model=glm-5-next-fp8,base_url=http://127.0.0.1:8000/v1/completions,num_concurrent=32,max_retries=3,tokenized_requests=False,tokenizer_backend=None" \ --tasks gsm8k --num_fewshot 5 --batch_size 1 \ --gen_kwargs temperature=0,max_gen_toks=512 --seed 1234 --log_samplesResults (exact_match, strict-match):
Fixed server sanity-checked with deterministic generation (temperature=0) on GSM8K-style prompts before running the evals.
Speed Tests and Profiling
Not affected. The change only materializes one extra per-expert fp32 tensor (
[num_local_experts], 288 bytes here) once at weight-load time; the hot kernel call is unchanged (thegemm1_clamp_limitargument was already being passed, asNone). No throughput regression expected or observed.CI States
Latest PR Test (Base): ⏳ Run #35192393525
Latest PR Test (Extra): ❌ Run #35192393195
Latest PR Test (AMD ROCm 10): ⏳ Run #35192393447