Skip to content

Add XPU and CPU branches to check_env so it stops raising NameError - #39571

Closed
kpjeeja wants to merge 1 commit into
sgl-project:mainfrom
kpjeeja:origin/jeeja/check_env_fix
Closed

kpjeeja wants to merge 1 commit into
sgl-project:mainfrom
kpjeeja:origin/jeeja/check_env_fix

Conversation

@kpjeeja

@kpjeeja kpjeeja commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

Motivation

python3 -m sglang.check_env can crash when the environment does not match one of the accelerator backends handled by the existing dispatch logic. The dispatch chain in __main__ had no XPU branch and no fallback else branch, so env could remain undefined:

$ python3 -m sglang.check_env
Traceback (most recent call last):
  File "/usr/lib/python3.12/runpy.py", line 198, in _run_module_as_main
    return _run_code(code, main_globals, None,
  ...
  File ".../python/sglang/check_env.py", line 591, in <module>
    env.check_env()
    ^^^
NameError: name 'env' is not defined

This is particularly problematic when collecting environment information for a bug report. For example, an Intel GPU/XPU environment, a CPU-engine environment, or a host where the accelerator stack is installed but currently unusable can produce a traceback instead of an environment report.

This change makes environment selection explicit and ensures that check_env always produces a report, including for XPU, CPU-engine, and otherwise unsupported environments.

Modifications

python/sglang/check_env.py

  • select_env() — moves environment selection out of if __name__ == "__main__" into a module-level function that always returns a BaseEnv. This makes the selection logic importable and unit-testable. __main__ is reduced to select_env().check_env(). The selection order is CUDA → HIP → NPU → MUSA → MPS → XPU → CPU → unknown; existing backends retain their current ordering and behavior.
  • XPUEnv — adds XPU-specific environment reporting, including XPU availability, devices grouped by name and memory, Level-Zero driver versions, the XPU build version reported by torch.version.xpu, and xpu-smi topology -m. It also reports relevant packages: intel-extension-for-pytorch, pytorch-triton-xpu, intel-sycl-rt, and oneccl. The xpu-smi probe uses a 15-second timeout so that a hung Level-Zero driver does not block the environment report indefinitely.
  • is_xpu_v2() — detects an XPU-enabled PyTorch build using torch.version.xpu is not None, rather than relying on is_xpu(), which also depends on a usable/visible XPU device. This allows check_env to report an XPU environment even when the XPU device is currently unavailable, instead of incorrectly falling through to the generic fallback.
  • CPUEnv — reports SGLANG_USE_CPU_ENGINE, platform.machine(), the CPU model, and the availability of avx512f, avx512_bf16, amx_bf16, and amx_int8 as reported by lscpu. If lscpu is unavailable, only the CPU ISA information is omitted.
  • UnknownEnv — provides a fallback for environments where no supported accelerator or CPU-engine configuration is detected. It reports Accelerator: none detected together with the standard Python, PyTorch, package, and ulimit information.
  • BaseEnv.__init__ — copies PACKAGE_LIST into self.package_list instead of aliasing the module-level list. Previously, a subclass extending self.package_list could modify the global PACKAGE_LIST, causing packages from one backend to leak into subsequently created environment objects in the same process.
  • PACKAGE_LIST — adds torch_memory_saver.
  • Moves the platform import from MPSEnv.get_info() to the module-level imports because CPUEnv also uses it.

test/registered/unit/hardware_backend/test_check_env.py (new, register_cpu_ci(est_time=5, suite="base-a-test-cpu"))

  • Verifies that select_env() returns UnknownEnv instead of raising when no detector matches, and that the fallback produces a complete report.
  • Verifies dispatch priority: an XPU-enabled PyTorch build is selected before the CPU fallback; a CUDA build takes precedence over later probes; and a CPU-engine environment is not reported as an unsupported/unknown environment.
  • Verifies that XPUEnv reports XPU available: False when no usable XPU device is available, without attempting device-property queries.
  • Verifies device grouping: identical devices are collapsed into a single entry, while different device configurations are reported separately, with driver information included.
  • Verifies that missing, timed-out, or non-zero xpu-smi execution degrades gracefully to an empty topology rather than aborting the environment report.
  • Verifies CPUEnv using captured lscpu output for both a Xeon configuration with AVX512/AMX support and an EPYC configuration without AVX512/AMX, as well as the missing-lscpu case.

No runtime or model-execution code is modified; this change is limited to the check_env diagnostic entry point and its unit tests.

Accuracy Tests

Not applicable — this change does not modify kernels, model execution, or inference behavior.

Speed Tests and Profiling

Not applicable — this change only affects the diagnostic CLI.

Checklist


CI States

Latest PR Test (Base): ❌ Run #35207328925
Latest PR Test (Extra): ❌ Run #35207328657
Latest PR Test (AMD ROCm 10): ❌ Run #35207328975

@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@kpjeeja
kpjeeja force-pushed the origin/jeeja/check_env_fix branch from a015202 to cc7d043 Compare September 15, 2026 08:45
Comment thread python/sglang/check_env.py Outdated
"""Environment checker for Intel XPU"""

EXTRA_PACKAGE_LIST = [
"intel-extension-for-pytorch",

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ipex is in maintain mode

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

fixed

Comment thread python/sglang/check_env.py Outdated

EXTRA_PACKAGE_LIST = [
"intel-extension-for-pytorch",
"pytorch-triton-xpu",

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

pytorch-triton-xpu has been replaced by triton-xpu

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

fixed

check_env picked its env class from a chain with no XPU case and no else, so on
an Intel GPU host `python3 -m sglang.check_env` died with
`NameError: name 'env' is not defined` -- exactly when a bug reporter needs it.

XPUEnv reports the devices, driver version, SYCL build and xpu-smi topology.
Selection keys on the torch build (torch.version.xpu) rather than is_xpu(),
which also requires a visible device: an XPU that is present but unusable --
driver mismatch, no /dev/dri access -- is what a bug report is most often about,
and it has to print the XPU stack instead of falling through to "none detected".

CPUEnv reports the CPU engine that docker/xeon.Dockerfile enables
(SGLANG_USE_CPU_ENGINE, CPU model, the AVX512/AMX flags its kernels need), and
UnknownEnv is the fallback for a host with no accelerator sglang recognizes.
BaseEnv now copies PACKAGE_LIST, so a subclass appending its extras no longer
mutates the module global for every env built afterwards.

The dispatch chain moves out of __main__ into select_env() so it can be tested;
test/registered/unit/test_check_env.py covers the fallback, the dispatch
priority, the unusable-XPU path, device grouping and the CPU reporter.
@kpjeeja
kpjeeja force-pushed the origin/jeeja/check_env_fix branch from cc7d043 to 862bf7f Compare September 17, 2026 09:50
@airMeng

airMeng commented Sep 17, 2026

Copy link
Copy Markdown
Collaborator

@kpjeeja what is the difference between this PR and #38873? can you merge as one PR?

@kpjeeja

kpjeeja commented Sep 21, 2026

Copy link
Copy Markdown
Contributor Author

@kpjeeja what is the difference between this PR and #38873? can you merge as one PR?

yes, this is Duplicate PR

@kpjeeja kpjeeja closed this Sep 21, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants