Repository navigation
Conversation
sigmoid_gate_topk_renorm calls sgl-kernel-xpu's inkling_gate_topk_renorm under is_xpu() when the installed wheel has it, and otherwise keeps the existing Triton kernel. The fused kernel replaces the tl.topk sort with masked-max passes and renormalizes in log space (sgl-project/sgl-kernel-xpu#528). Also gate both nvcc JIT paths on .is_cuda: XPU satisfies torch.version.hip is None, so it took the CUDA JIT branch and asserted.
This was referenced Oct 5, 2026
jmunetong
marked this pull request as ready for review
October 6, 2026 22:47
jmunetong
requested review from
BBuf,
DarkSharpness,
HaiShaw,
HydraQYH,
celve and
yuan-luo
as code owners
October 6, 2026 22:47
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
The Inkling MoE gate epilogue, everything between the 258-column gate GEMM and the MoE dispatch, does not run on Intel XPU today, and once unblocked it lands on a path that's 18–53× slower than necessary.
This supersedes the approach in #36385, which put the XPU optimization inside the shared
_router_triton_kerneland changed CUDA/ROCm code. Following review there, the optimized kernel now lives in sgl-kernel-xpu (sgl-project/sgl-kernel-xpu#528). This PR only dispatches to it on XPU, and leaves every CUDA/ROCm code path unchanged.Modifications
2 files, +28/−1.
sigmoid_gate_topk_renorm.py, enablement. The production-shape CUDA JIT gate checked onlytorch.version.hip is None, which XPU also satisfies, so all 64 MoE layers took the CUDA JIT path and died onassert logits.is_cuda. It now also requireslogits.is_cuda. On CUDA and ROCm the added conjunct is always true, so the condition is unchanged.sigmoid_gate_topk_renorm.py, dispatch. Underis_xpu(), callssgl_kernel.inkling_gate_topk_renorm, which has the same arguments, return tuple and dtypes. It is looked up withgetattr(sgl_kernel, "inkling_gate_topk_renorm", None), so against a wheel without the op, XPU keeps the existing Triton kernel. The fast path activates once a wheel containing Add Inkling MoE gate top-k + logsigmoid renorm Triton kernel sgl-kernel-xpu#528 is pinned inpyproject_xpu.toml.inkling_common/moe.py. Same.is_cudaconjunct on the gate-GEMV condition, whose branches are nvcc JIT kernels. Again a no-op on CUDA and ROCm.moe_fused_gate.py,environ.pyand the shared_sigmoid_gate_topk_renorm_kernelare untouched.Accuracy Tests
Covered by the kernel PR, sgl-project/sgl-kernel-xpu#528. It has 21 cases against an fp64 oracle:
sigmoid / Σ sigmoidreturns NaNAgainst the existing Triton kernel at the production shape, indices are bit-identical and weights match within 2e-3.
The dispatch itself was verified at runtime by profiling which kernel ran, on Intel Arc Pro B60:
inkling_gate_topk_renorm_inkling_gate_topk_renorm_kernel_inkling_gate_topk_renorm_kernel_sigmoid_gate_topk_renorm_kernel_sigmoid_gate_topk_renorm_kernelNo sglang-side unit test is added. The XPU CI installs the pinned wheel, which lacks the op, so a test of the fast path could not exercise it until the pin moves. I'll add one with the pin bump.
Speed Tests and Profiling
Device time (profiler self-time) at the production shape (256 routed + 2 sink, k=6, fp32). Same process, interleaved, min of 5 reps × 200 iters. Intel Arc Pro B60, torch 2.15.0.dev20260912+xpu, triton 3.8.0. Measured in a single session on a shared host, so treat these as indicative.
tl.topk)sgl_kernel.inkling_gate_topk_renormChecklist
Review and Merge Process
/tag-and-rerun-ci,/tag-run-ci-label,/rerun-failed-ciCI States
Latest PR Test (Base): ❌ Run #37542803695
Latest PR Test (Extra): ❌ Run #37542803206
Latest PR Test (AMD ROCm 10): ❌ Run #37542803557