Repository navigation
Conversation
chrikrah
left a comment
There was a problem hiding this comment.
#37037 is open on this file and fixes the same NameError. It landed 2026-08-29, twelve days before
this, moves the selection into a _get_env() with a CPU fallback, gives CPUEnv a real body with
os.cpu_count() rather than an empty dict, and adds test/registered/unit/test_check_env.py. Neither
pull request mentions the other. Whichever goes first, the second is a rebase rather than a rewrite,
since your XPU branch and its CPU branch sit in the same if chain.
Both are also stuck on the same thing, and it is not the code. pr-gate exits on the label:
PR Labels: []
...
Missing required label 'run-ci'. See https://docs.sglang.io/developer_guide/contribution_guide.html#how-to-trigger-ci-tests
/tag-and-rerun-ci adds it, from any account in .github/CI_PERMISSIONS.json. Neither of us is among
those 251, so somebody who is has to say it before any suite starts. lint is the one check that did
run, and it passed.
One confirmation from the outside: on origin/main a CPU-only host raises your exact
NameError: name 'env' is not defined at check_env.py:594, with no XPU present. is_xpu() returns
False there and no other branch matches, so the unbound env is not Intel-specific. is_cpu() is not
the branch for it either: srt/utils/common.py:218 returns False unless SGLANG_USE_CPU_ENGINE=1.
chrikrah
left a comment
There was a problem hiding this comment.
Correcting myself on the overlap: #37037 is not the CPU half of this file. It adds its own
XPUEnv(BaseEnv) with torch.xpu.is_available(), device_count() and get_device_name(), and
if is_xpu(): return XPUEnv() inside _get_env(). So the two collide on the class name and on the whole
dispatch block, not only at the tail, and whichever goes second is a rewrite rather than a rebase.
On the ordering: #37037 was opened 2026-08-29, covers CPU as well as XPU and carries
test/registered/unit/test_check_env.py. This one is the fuller XPU report, 314 lines against 114, with
the package lists and the SYCL runtime detail. #29568 proposed the CPU fallback before either and its own
author closed it, so no maintainer has turned that idea down.
On this host the new else is the branch that runs, so the crash becomes the RuntimeError. Better, and
still no output for the Environment field the bug template requires.
For the label: @Kangyan-Zhou merged the MPS and MUSA additions to this same chain, #20753 and #16959, and
is listed in .github/CI_PERMISSIONS.json. /tag-and-rerun-ci here and on #37037 would start the suites.
|
@airMeng on the duplicate question, with today's state. #39571 is closed unmerged, so it is not the overlap. #37037 is still open and is, on this file: it adds its own #37037 already ships I read the state today and have not run either branch. |
Motivation
python3 -m sglang.check_envcrashes on Intel XPU machines:The dispatch block at the bottom of
check_env.pyonly handles CUDA, HIP, NPU,MUSA and MPS. There is no XPU branch and no
else, so on an XPU host everycondition is false,
envis never bound, and the script dies with aNameErrorthat gives the user no clue what went wrong.
Modifications
1. Fix the
NameErrorAdded an
is_xpu()branch to the dispatch chain, plus a finalelsethat raisesa descriptive
RuntimeErrorinstead of falling through to an unbound name on anyfuture unsupported platform.
2. New
XPUEnv(BaseEnv)Reports XPU-native information rather than CUDA-shaped fields:
torch.xpu.get_device_properties, grouped by devicename): device id, architecture, vendor, SYCL platform, compute-runtime version,
backend version, UUID, total memory, compute units, EU count, Xe-core (subslice)
count, max work-group size, sub-group sizes, SLM size, memory bus width/clock.
fp16, fp64, atomic64, bf16_conversions, XMX (dpas), XMX tf32, 2d_block_io.torch.version.xpu(SYCL version), AOT arch list andgencode flags, bf16/tf32 support. An arch-list mismatch is the most common cause
of XPU-only kernel failures, so surfacing it here saves a round-trip.
ONEAPI_ROOT/CMPLR_ROOT,icpx --version,sycl-ls, and the kernel-mode driver in use (xevsi915) with its version.sgl_kernelhealth: whether the extension imports and how many ops itregistered — an importable-but-opless build is the usual symptom of an arch
mismatch.
ZE_AFFINITY_MASK,ZE_FLAT_DEVICE_HIERARCHY,ONEAPI_DEVICE_SELECTOR,SYCL_CACHE_*,UR_L0_*,CCL_*,TORCH_LLM_ALLREDUCE,SGLANG_ATTENTION_BACKEND, etc.xpu-smi topology -mandxpu-smi discovery(PCI BDF, DRM device,SOC UUID) — the Intel analogue of
nvidia-smi topo -m.3. XPU-appropriate package list
Drops CUDA-only entries that are always
Module Not Foundon XPU(
flashinfer_python,flashinfer_cubin,flashinfer_jit_cache,triton,sglang-kernel) and adds the XPU stack:sgl-kernel/sgl-kernel-xpu,triton-xpu,intel-sycl-rt,intel-cmplr-*,intel-opencl-rt,intel-pti,dpcpp-cpp-rt,umf,tcmlib, the oneMKL SYCL libs,oneccl, andimpi-rt.sgl-kernel-xpuis aliased to thesgl-kerneldistribution, since the XPUkernels ship under that distribution name.
All external commands (
icpx,sycl-ls,xpu-smi) are wrapped so a missingbinary degrades to
Not Availableor an omitted section rather than failing thewhole report.
No behavioural change for CUDA/HIP/NPU/MUSA/MPS — those paths are untouched.
Accuracy
Verified on 2 x Intel Arc Pro B70 and 8 x Intel Arc Pro B60. Sample output:
CI States
Latest PR Test (Base): ❌ Run #34936881811
Latest PR Test (Extra): ❌ Run #34936881789
Latest PR Test (AMD ROCm 10): ❌ Run #34936881782