Repository navigation
Conversation
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
kpjeeja
force-pushed
the
origin/jeeja/check_env_fix
branch
from
September 15, 2026 08:45
a015202 to
cc7d043
Compare
airMeng
reviewed
Sep 17, 2026
| """Environment checker for Intel XPU""" | ||
|
|
||
| EXTRA_PACKAGE_LIST = [ | ||
| "intel-extension-for-pytorch", |
|
|
||
| EXTRA_PACKAGE_LIST = [ | ||
| "intel-extension-for-pytorch", | ||
| "pytorch-triton-xpu", |
Collaborator
There was a problem hiding this comment.
pytorch-triton-xpu has been replaced by triton-xpu
check_env picked its env class from a chain with no XPU case and no else, so on an Intel GPU host `python3 -m sglang.check_env` died with `NameError: name 'env' is not defined` -- exactly when a bug reporter needs it. XPUEnv reports the devices, driver version, SYCL build and xpu-smi topology. Selection keys on the torch build (torch.version.xpu) rather than is_xpu(), which also requires a visible device: an XPU that is present but unusable -- driver mismatch, no /dev/dri access -- is what a bug report is most often about, and it has to print the XPU stack instead of falling through to "none detected". CPUEnv reports the CPU engine that docker/xeon.Dockerfile enables (SGLANG_USE_CPU_ENGINE, CPU model, the AVX512/AMX flags its kernels need), and UnknownEnv is the fallback for a host with no accelerator sglang recognizes. BaseEnv now copies PACKAGE_LIST, so a subclass appending its extras no longer mutates the module global for every env built afterwards. The dispatch chain moves out of __main__ into select_env() so it can be tested; test/registered/unit/test_check_env.py covers the fallback, the dispatch priority, the unusable-XPU path, device grouping and the CPU reporter.
kpjeeja
force-pushed
the
origin/jeeja/check_env_fix
branch
from
September 17, 2026 09:50
cc7d043 to
862bf7f
Compare
Collaborator
Contributor
Author
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
python3 -m sglang.check_envcan crash when the environment does not match one of the accelerator backends handled by the existing dispatch logic. The dispatch chain in__main__had no XPU branch and no fallbackelsebranch, soenvcould remain undefined:This is particularly problematic when collecting environment information for a bug report. For example, an Intel GPU/XPU environment, a CPU-engine environment, or a host where the accelerator stack is installed but currently unusable can produce a traceback instead of an environment report.
This change makes environment selection explicit and ensures that
check_envalways produces a report, including for XPU, CPU-engine, and otherwise unsupported environments.Modifications
python/sglang/check_env.pyselect_env()— moves environment selection out ofif __name__ == "__main__"into a module-level function that always returns aBaseEnv. This makes the selection logic importable and unit-testable.__main__is reduced toselect_env().check_env(). The selection order is CUDA → HIP → NPU → MUSA → MPS → XPU → CPU → unknown; existing backends retain their current ordering and behavior.XPUEnv— adds XPU-specific environment reporting, including XPU availability, devices grouped by name and memory, Level-Zero driver versions, the XPU build version reported bytorch.version.xpu, andxpu-smi topology -m. It also reports relevant packages:intel-extension-for-pytorch,pytorch-triton-xpu,intel-sycl-rt, andoneccl. Thexpu-smiprobe uses a 15-second timeout so that a hung Level-Zero driver does not block the environment report indefinitely.is_xpu_v2()— detects an XPU-enabled PyTorch build usingtorch.version.xpu is not None, rather than relying onis_xpu(), which also depends on a usable/visible XPU device. This allowscheck_envto report an XPU environment even when the XPU device is currently unavailable, instead of incorrectly falling through to the generic fallback.CPUEnv— reportsSGLANG_USE_CPU_ENGINE,platform.machine(), the CPU model, and the availability ofavx512f,avx512_bf16,amx_bf16, andamx_int8as reported bylscpu. Iflscpuis unavailable, only the CPU ISA information is omitted.UnknownEnv— provides a fallback for environments where no supported accelerator or CPU-engine configuration is detected. It reportsAccelerator: none detectedtogether with the standard Python, PyTorch, package, and ulimit information.BaseEnv.__init__— copiesPACKAGE_LISTintoself.package_listinstead of aliasing the module-level list. Previously, a subclass extendingself.package_listcould modify the globalPACKAGE_LIST, causing packages from one backend to leak into subsequently created environment objects in the same process.PACKAGE_LIST— addstorch_memory_saver.platformimport fromMPSEnv.get_info()to the module-level imports becauseCPUEnvalso uses it.test/registered/unit/hardware_backend/test_check_env.py(new,register_cpu_ci(est_time=5, suite="base-a-test-cpu"))select_env()returnsUnknownEnvinstead of raising when no detector matches, and that the fallback produces a complete report.XPUEnvreportsXPU available: Falsewhen no usable XPU device is available, without attempting device-property queries.xpu-smiexecution degrades gracefully to an empty topology rather than aborting the environment report.CPUEnvusing capturedlscpuoutput for both a Xeon configuration with AVX512/AMX support and an EPYC configuration without AVX512/AMX, as well as the missing-lscpucase.No runtime or model-execution code is modified; this change is limited to the
check_envdiagnostic entry point and its unit tests.Accuracy Tests
Not applicable — this change does not modify kernels, model execution, or inference behavior.
Speed Tests and Profiling
Not applicable — this change only affects the diagnostic CLI.
Checklist
CI States
Latest PR Test (Base): ❌ Run #35207328925
Latest PR Test (Extra): ❌ Run #35207328657
Latest PR Test (AMD ROCm 10): ❌ Run #35207328975