Skip to content

[XPU] Dispatch the Inkling gate epilogue to sgl-kernel-xpu - #42670

Open
jmunetong wants to merge 2 commits into
sgl-project:mainfrom
jmunetong:xpu/inkling-gate-sgl-kernel-xpu
Open

jmunetong wants to merge 2 commits into
sgl-project:mainfrom
jmunetong:xpu/inkling-gate-sgl-kernel-xpu

Conversation

@jmunetong

@jmunetong jmunetong commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Motivation

The Inkling MoE gate epilogue, everything between the 258-column gate GEMM and the MoE dispatch, does not run on Intel XPU today, and once unblocked it lands on a path that's 18–53× slower than necessary.

This supersedes the approach in #36385, which put the XPU optimization inside the shared _router_triton_kernel and changed CUDA/ROCm code. Following review there, the optimized kernel now lives in sgl-kernel-xpu (sgl-project/sgl-kernel-xpu#528). This PR only dispatches to it on XPU, and leaves every CUDA/ROCm code path unchanged.

Modifications

2 files, +28/−1.

  • sigmoid_gate_topk_renorm.py, enablement. The production-shape CUDA JIT gate checked only torch.version.hip is None, which XPU also satisfies, so all 64 MoE layers took the CUDA JIT path and died on assert logits.is_cuda. It now also requires logits.is_cuda. On CUDA and ROCm the added conjunct is always true, so the condition is unchanged.
  • sigmoid_gate_topk_renorm.py, dispatch. Under is_xpu(), calls sgl_kernel.inkling_gate_topk_renorm, which has the same arguments, return tuple and dtypes. It is looked up with getattr(sgl_kernel, "inkling_gate_topk_renorm", None), so against a wheel without the op, XPU keeps the existing Triton kernel. The fast path activates once a wheel containing Add Inkling MoE gate top-k + logsigmoid renorm Triton kernel sgl-kernel-xpu#528 is pinned in pyproject_xpu.toml.
  • inkling_common/moe.py. Same .is_cuda conjunct on the gate-GEMV condition, whose branches are nvcc JIT kernels. Again a no-op on CUDA and ROCm.

moe_fused_gate.py, environ.py and the shared _sigmoid_gate_topk_renorm_kernel are untouched.

Accuracy Tests

Covered by the kernel PR, sgl-project/sgl-kernel-xpu#528. It has 21 cases against an fp64 oracle:

  • the production shape across T ∈ {1…4096}, contiguous and strided
  • ties, a dominating shared sink, and all-underflow logits, where the existing kernel's sigmoid / Σ sigmoid returns NaN
  • packed output

Against the existing Triton kernel at the production shape, indices are bit-identical and weights match within 2e-3.

The dispatch itself was verified at runtime by profiling which kernel ran, on Intel Arc Pro B60:

installed sgl-kernel-xpu plain packed
with inkling_gate_topk_renorm _inkling_gate_topk_renorm_kernel _inkling_gate_topk_renorm_kernel
without (current pinned wheel) _sigmoid_gate_topk_renorm_kernel _sigmoid_gate_topk_renorm_kernel

No sglang-side unit test is added. The XPU CI installs the pinned wheel, which lacks the op, so a test of the fast path could not exercise it until the pin moves. I'll add one with the pin bump.

Speed Tests and Profiling

Device time (profiler self-time) at the production shape (256 routed + 2 sink, k=6, fp32). Same process, interleaved, min of 5 reps × 200 iters. Intel Arc Pro B60, torch 2.15.0.dev20260912+xpu, triton 3.8.0. Measured in a single session on a shared host, so treat these as indicative.

µs T=1 T=8 T=64 T=512 T=4096 T=8192
existing Triton kernel (tl.topk) 101.94 117.13 123.58 223.89 639.23 1231.98
sgl_kernel.inkling_gate_topk_renorm 4.95 6.24 6.96 4.22 20.57 37.09
speedup 20.6× 18.8× 17.8× 53.0× 31.1× 33.2×

Checklist

Review and Merge Process

  1. Ping Merge Oncalls to start the process. See the PR Merge Process.
  2. Get approvals from CODEOWNERS and other reviewers.
  3. Trigger CI tests with comments or contact authorized users to do so.
    • Common commands include /tag-and-rerun-ci, /tag-run-ci-label, /rerun-failed-ci
  4. After green CI and required approvals, ask Merge Oncalls or people with Write permission to merge the PR.

CI States

Latest PR Test (Base): ❌ Run #37542803695
Latest PR Test (Extra): ❌ Run #37542803206
Latest PR Test (AMD ROCm 10): ❌ Run #37542803557

sigmoid_gate_topk_renorm calls sgl-kernel-xpu's inkling_gate_topk_renorm
under is_xpu() when the installed wheel has it, and otherwise keeps the
existing Triton kernel. The fused kernel replaces the tl.topk sort with
masked-max passes and renormalizes in log space (sgl-project/sgl-kernel-xpu#528).

Also gate both nvcc JIT paths on .is_cuda: XPU satisfies
torch.version.hip is None, so it took the CUDA JIT branch and asserted.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant