Skip to content

[Moe] Fix flashinfer_trtllm silently dropping swiglu_limit clamped SwiGLU activation - #39920

Merged
hnyls2002 merged 1 commit into
sgl-project:mainfrom
Dovis01:fix/flashinfer-trtllm-swiglu-limit
Sep 17, 2026
Merged

hnyls2002 merged 1 commit into
sgl-project:mainfrom
Dovis01:fix/flashinfer-trtllm-swiglu-limit

Conversation

@Dovis01

@Dovis01 Dovis01 commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Motivation

Models with a clamped SwiGLU activation (e.g. GLM-5.3-Flash / GLM-5-Next, Qwen3-Next style) plumb the clamp limit through FusedMoE(swiglu_limit=...) -> MoeRunnerConfig.swiglu_limit. All MoE runner backends are supposed to honor it: triton passes it into the fused kernel, deep_gemm applies it via silu_and_mul_clamp / _apply_swiglu_limit, but flashinfer_trtllm silently ignores it.

The flashinfer_trtllm FP8 path only materializes gemm1_clamp_limit (the TRT-LLM kernel parameter), which stays None for models that use the swiglu_limit field. With no error and no warning, the expert activation degrades from the trained silu(clamp(gate, max=L)) * clamp(up, -L, L) (with L = 10.0) to a plain silu(gate) * up, which measurably hurts accuracy.

On GLM-5.3-Flash FP8 (blockwise 128x128 FP8, 288 routed experts, top-8, 1 shared expert, swiglu_limit=10.0), GSM8K 5-shot with only --moe-runner-backend varying:

moe_runner_backend GSM8K strict-match (runs) mean
triton 0.9356 / 0.9363 0.9360
deep_gemm 0.9249 / 0.9287 / 0.9295 / 0.9310 / 0.9340 / 0.9348 0.9305
flashinfer_trtllm (before) 0.9265 / 0.9280 / 0.9287 0.9277
flashinfer_trtllm (this fix) 0.9348 / 0.9371 0.9360

Before/after ranges do not overlap (within-group variance is ~0.2pp); after the fix flashinfer_trtllm matches the triton baseline exactly. fix #39797. Thx for finding this issue. @NolenLiang

Modifications

Single-file change in python/sglang/srt/layers/quantization/fp8.py, Fp8MoEMethod._prepare_flashinfer_trtllm_activation_params(): when gemm1_clamp_limit is not explicitly set, fall back to swiglu_limit.

Why here and not in the model code: this function is the only consumer that builds the TRT-LLM SwiGLU parameters (gemm1_alpha/beta/clamp_limit), so the two naming pipelines (gemm1_clamp_limit used by gpt_oss / bailing_moe_v3, swiglu_limit used by DeepseekV2MoE-family models) converge at a single point; triton consumes both fields independently (gemm1_limit and swiglu_limit), so setting both at the model level would change behavior of unrelated backends, while this fallback is scoped to the flashinfer_trtllm preparation only. Models that already set gemm1_clamp_limit explicitly (gpt_oss, bailing_moe_v3) are unaffected, and the same silent drop is fixed for every other model that routes the clamp through swiglu_limit (e.g. Qwen3-Next) whenever flashinfer_trtllm is used.

The TRT-LLM kernel natively supports the clamp: with ActivationType.Swiglu, gemm1_clamp_limit=L computes X2 = clamp(X2, max=L), X1 = clamp(X1, -L, L), out = X2 * sigmoid(alpha * X2) * (X1 + beta) (alpha defaults to 1.0, beta to 0.0), which matches the reference swiglu_clamped in models/glm5_next.py exactly. Verified against flashinfer 0.6.18 (trtllm_fp8_block_scale_moe, Fp8QuantizationType.DeepSeekFp8).

Accuracy Tests

Model: zai-org/GLM-5.3-Flash (FP8, blockwise 128x128), 4x B200, TP=4/EP=4. Launch (only --moe-runner-backend varies between runs):

 python3 -m sglang.launch_server \
  --model-path zai-org/GLM-5.3-Flash \
  --tp-size 4 --ep-size 4 --mem-fraction-static 0.75 \
  --dsa-prefill-backend trtllm --dsa-decode-backend trtllm \
  --kv-cache-dtype fp8_e4m3 \
  --moe-runner-backend flashinfer_trtllm \
  --context-length 16384 --disable-radix-cache \
  --host 0.0.0.0 --port 8000 --served-model-name glm-5-next-fp8

Eval (lm-evaluation-harness, local-completions):

lm_eval --model local-completions \
  --model_args "model=glm-5-next-fp8,base_url=http://127.0.0.1:8000/v1/completions,num_concurrent=32,max_retries=3,tokenized_requests=False,tokenizer_backend=None" \
  --tasks gsm8k --num_fewshot 5 --batch_size 1 \
  --gen_kwargs temperature=0,max_gen_toks=512 --seed 1234 --log_samples

Results (exact_match, strict-match):

config run 1 run 2 run 3 mean
flashinfer_trtllm before 0.9287 0.9265 0.9280 0.9277
flashinfer_trtllm after 0.9371 0.9348 – 0.9360
triton (baseline) 0.9363 0.9356 – 0.9360

Fixed server sanity-checked with deterministic generation (temperature=0) on GSM8K-style prompts before running the evals.

Speed Tests and Profiling

Not affected. The change only materializes one extra per-expert fp32 tensor ([num_local_experts], 288 bytes here) once at weight-load time; the hot kernel call is unchanged (the gemm1_clamp_limit argument was already being passed, as None). No throughput regression expected or observed.


CI States

Latest PR Test (Base): ⏳ Run #35192393525
Latest PR Test (Extra): ❌ Run #35192393195
Latest PR Test (AMD ROCm 10): ⏳ Run #35192393447

Models that plumb the clamped-SwiGLU limit through swiglu_limit
(GLM-5 via DeepseekV2MoE, Qwen3-Next) silently lost the clamp on the
flashinfer_trtllm FP8 MoE path: it only reads gemm1_clamp_limit, which
stays None for these models, so the expert activation degrades to plain
SwiGLU. Fall back to swiglu_limit when gemm1_clamp_limit is unset; the
TRT-LLM kernel applies it via gemm1_clamp_limit natively.

GSM8K, GLM-5.3-Flash FP8, TP4/EP4 on B200:
flashinfer_trtllm 92.65-92.87 -> 93.48-93.71 (triton baseline 93.56-93.63).
@hnyls2002

Copy link
Copy Markdown
Collaborator

/rerun-test test_fp8_moe_runner_ownership.py test_moe_runner_extensions.py

@hnyls2002

Copy link
Copy Markdown
Collaborator

/tag-and-rerun-ci

@github-actions

github-actions Bot commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test test_fp8_moe_runner_ownership.py test_moe_runner_extensions.py:

🚀 ubuntu-latest (2 tests): ✅ View workflow run

cd test/ && python3 registered/unit/layers/quantization/test_fp8_moe_runner_ownership.py
cd test/ && python3 registered/unit/layers/moe/test_moe_runner_extensions.py

@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label Sep 17, 2026
@hnyls2002
hnyls2002 merged commit aebae58 into sgl-project:main Sep 17, 2026
173 of 206 checks passed
Qiaolin-Yu pushed a commit that referenced this pull request Sep 17, 2026
… dropping swiglu_limit clamped SwiGLU activation (#39920) (#40035)

Co-authored-by: Shijin Zhang <75300765+Dovis01@users.noreply.github.com>
fungaren pushed a commit to fungaren/sglang that referenced this pull request Sep 20, 2026
… dropping swiglu_limit clamped SwiGLU activation (sgl-project#39920) (sgl-project#40035)

Co-authored-by: Shijin Zhang <75300765+Dovis01@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

run-ci CI: run the baseline test suite on this PR

Projects

None yet

2 participants