Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
57 commits
Select commit Hold shift + click to select a range
ce194e1
Studio: keep a Windows update from leaving a CPU-only PyTorch
danielhanchen Aug 27, 2026
6770055
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Aug 27, 2026
9d543e4
Studio: unset the published torch flavor portably
danielhanchen Aug 27, 2026
1bf632b
Studio: close three gaps in the Windows torch flavor invariant
danielhanchen Aug 27, 2026
66f58e4
Studio: enforce the published ROCm flavor too, by delegation
danielhanchen Aug 27, 2026
3b1c37e
Studio: do not skip the flavor invariant on an explicitly pinned GPU …
danielhanchen Aug 27, 2026
65e1492
Studio: let an explicit torch index pin outrank the recorded flavor
danielhanchen Aug 27, 2026
b600fe0
Studio: keep the Windows ARM64 no-torchaudio exception in the flavor …
danielhanchen Aug 27, 2026
88a8eb4
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Aug 27, 2026
145afd2
Studio: verify the delegated ROCm repair, and enforce a pinned CPU index
danielhanchen Aug 27, 2026
e116d30
Studio: verify the repaired family, and re-settle what the torch rele…
danielhanchen Aug 27, 2026
15fd973
Studio: stop the post-repair resync from undoing the repair
danielhanchen Aug 27, 2026
f8fca1a
Studio: re-pin torchao when the CUDA major moves, not just the release
danielhanchen Aug 27, 2026
cde627a
Studio: report the GPUs the OS sees, not just the ones PyTorch opened
danielhanchen Aug 27, 2026
22c5b51
Studio: four corrections to the GPU mismatch reporting
danielhanchen Aug 27, 2026
5a3b7fa
Studio: inventory Linux AMD cards, and stop three false mismatch verd…
danielhanchen Aug 27, 2026
f7376fa
Studio: record the flavor off Windows, and keep the GPU verdict current
danielhanchen Aug 27, 2026
cd5fd6c
Studio: inventory Intel GPUs, keep unknown distinct, and stay off the…
danielhanchen Aug 27, 2026
32927d4
Studio: corroborate Windows adapters, exclude plain Intel iGPUs, re-d…
danielhanchen Aug 27, 2026
fb404a9
Studio: XPU recovery and masks, one-to-one adapter corroboration, qui…
danielhanchen Aug 30, 2026
59b7952
Studio: the XPU triton swap now has a third repair point on Windows
danielhanchen Aug 30, 2026
d3f0460
Studio: recovery really re-detects, the health path never blocks, and…
danielhanchen Aug 30, 2026
4229100
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Aug 30, 2026
b7692fc
Studio: the health path measures torch off-thread, and an unimportabl…
danielhanchen Aug 30, 2026
83dfc3f
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Aug 30, 2026
eb367b2
Studio: per-vendor masks, untagged GPU builds off disk, and a broken …
danielhanchen Aug 30, 2026
a8153ee
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Aug 30, 2026
4b9f659
Studio: CUDA_VISIBLE_DEVICES masks AMD too, and no request path retri…
danielhanchen Aug 30, 2026
695cac4
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Aug 30, 2026
f412709
Studio: an AMD card the installers decline is not a broken install
danielhanchen Aug 30, 2026
7457db4
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Aug 30, 2026
848c913
Studio: an absent nvidia-smi is an answer, and the masks are compared…
danielhanchen Aug 30, 2026
35897c8
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Aug 30, 2026
6bbc047
Studio: release the re-detection guard, read the verdict once, and st…
danielhanchen Aug 30, 2026
f5ba713
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Aug 30, 2026
2cce0fb
Studio: keep the hardware module loadable without the rest of the pac…
danielhanchen Aug 30, 2026
2a3a0d4
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Aug 30, 2026
d829c83
Studio: five more gaps in the Windows flavor invariant
danielhanchen Aug 30, 2026
54e77af
Studio: ARM64 exception on the delegated ROCm repair, and a broken to…
danielhanchen Aug 30, 2026
8c44df0
Studio: probe fallbacks the installer already has, and an automatic C…
danielhanchen Aug 30, 2026
71ba7bb
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Aug 30, 2026
a91bf50
Studio: per-vendor probe accounting, exact adapter matches, and a der…
danielhanchen Aug 30, 2026
2213bf5
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Aug 30, 2026
2e38923
Studio: every supported AMD arch, HIP precedence over the CUDA alias,…
danielhanchen Aug 30, 2026
ac0070a
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Aug 30, 2026
25c6d47
Studio: keep a stated backend's provenance, count an untagged XPU run…
danielhanchen Aug 30, 2026
a542d80
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Aug 30, 2026
368a190
Reduce comment volume across the branch
danielhanchen Aug 31, 2026
8c785cb
Reduce comment volume across the branch
danielhanchen Aug 31, 2026
1b5085b
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Aug 31, 2026
89ec8e4
Studio: evaluate the sidebar poll decision, do not just match its text
danielhanchen Aug 31, 2026
24aff32
Studio: the unimportable-torch test passed for the wrong reason below…
danielhanchen Aug 31, 2026
fc1e500
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Aug 31, 2026
175ceae
Studio: fixes from review round 2
danielhanchen Aug 31, 2026
6fd8a77
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Aug 31, 2026
8bd0865
Merge commit '6fd8a77a1b8753cd2b7b3664c5c15fdd3e4457e2' into pr-9858-…
danielhanchen Aug 31, 2026
c4c725c
Add staging CI workflows for unslothai/unsloth#9858
danielhanchen Aug 31, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
31 changes: 31 additions & 0 deletions .github/workflows/staging-9858-macos-15.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,31 @@
name: "staging-9858 macos-15"
on:
push:
branches: ["pr-9858-xplat-ci"]
paths-ignore:
- '**/*.md'
workflow_dispatch:
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: true
permissions:
contents: read
defaults:
run:
shell: bash
jobs:
test:
runs-on: macos-15
timeout-minutes: 30
env:
UNSLOTH_COMPILE_DISABLE: '1'
UNSLOTH_IS_PRESENT: '1'
steps:
- uses: actions/checkout@v4
with:
persist-credentials: false
- uses: actions/setup-python@v5
with:
python-version: '3.12'
- run: python -m pip install -r studio/backend/requirements/studio.txt pytest pytest-asyncio pytest-timeout python-multipart
- run: rc=0; PYTHONPATH=studio/backend python -m pytest studio/backend/tests/test_torch_cpu_build_on_nvidia_host.py tests/python/test_install_python_stack.py tests/studio/install/test_cuda_repair.py tests/studio/install/test_gpu_detection_followups.py tests/studio/install/test_pr5940_followups.py tests/studio/install/test_rocm_arch_table_parity.py tests/studio/install/test_rocm_support.py tests/studio/install/test_windows_torch_flavor_invariant.py tests/studio/test_xpu_triton_swap.py -q --tb=short -k 'not llama_cpp_load_progress_live and not TestGpuAutoSelection and not TestPreSpawnGpuResolution and not TestPerGpuFitGuardAllCounts and not TestTransformersIntrospection and not test_returns_cuda_when_cuda_available and not test_calls_cuda_cache_when_cuda' > pytest_out.txt 2>&1 || rc=$?; cat pytest_out.txt; if [ "$rc" = "5" ]; then echo "no tests ran (deps absent on this runner)"; rc=0; fi; if [ "$rc" = "2" ] && grep -qE "No module named .(torch|unsloth_zoo|transformers)." pytest_out.txt && ! grep -qE "^(FAILED|ERROR) " pytest_out.txt; then echo "collection needs a dep this runner does not ship"; rc=0; fi; if [ "$rc" = "1" ] && ! grep -qE "^ERROR " pytest_out.txt; then tot=$(grep -cE "^FAILED " pytest_out.txt); dep=$(grep -cE "^FAILED .* No module named .(torch|unsloth_zoo|transformers).$" pytest_out.txt); if [ "$tot" -gt 0 ] && [ "$tot" = "$dep" ]; then echo "only tests needing a dep this runner does not ship failed"; rc=0; fi; fi; exit "$rc"
31 changes: 31 additions & 0 deletions .github/workflows/staging-9858-ubuntu-latest.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,31 @@
name: "staging-9858 ubuntu-latest"
on:
push:
branches: ["pr-9858-xplat-ci"]
paths-ignore:
- '**/*.md'
workflow_dispatch:
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: true
permissions:
contents: read
defaults:
run:
shell: bash
jobs:
test:
runs-on: ubuntu-latest
timeout-minutes: 30
env:
UNSLOTH_COMPILE_DISABLE: '1'
UNSLOTH_IS_PRESENT: '1'
steps:
- uses: actions/checkout@v4
with:
persist-credentials: false
- uses: actions/setup-python@v5
with:
python-version: '3.12'
- run: python -m pip install -r studio/backend/requirements/studio.txt pytest pytest-asyncio pytest-timeout python-multipart
- run: rc=0; PYTHONPATH=studio/backend python -m pytest studio/backend/tests/test_torch_cpu_build_on_nvidia_host.py tests/python/test_install_python_stack.py tests/studio/install/test_cuda_repair.py tests/studio/install/test_gpu_detection_followups.py tests/studio/install/test_pr5940_followups.py tests/studio/install/test_rocm_arch_table_parity.py tests/studio/install/test_rocm_support.py tests/studio/install/test_windows_torch_flavor_invariant.py tests/studio/test_xpu_triton_swap.py -q --tb=short -k 'not llama_cpp_load_progress_live and not TestGpuAutoSelection and not TestPreSpawnGpuResolution and not TestPerGpuFitGuardAllCounts and not TestTransformersIntrospection and not test_returns_cuda_when_cuda_available and not test_calls_cuda_cache_when_cuda' > pytest_out.txt 2>&1 || rc=$?; cat pytest_out.txt; if [ "$rc" = "5" ]; then echo "no tests ran (deps absent on this runner)"; rc=0; fi; if [ "$rc" = "2" ] && grep -qE "No module named .(torch|unsloth_zoo|transformers)." pytest_out.txt && ! grep -qE "^(FAILED|ERROR) " pytest_out.txt; then echo "collection needs a dep this runner does not ship"; rc=0; fi; if [ "$rc" = "1" ] && ! grep -qE "^ERROR " pytest_out.txt; then tot=$(grep -cE "^FAILED " pytest_out.txt); dep=$(grep -cE "^FAILED .* No module named .(torch|unsloth_zoo|transformers).$" pytest_out.txt); if [ "$tot" -gt 0 ] && [ "$tot" = "$dep" ]; then echo "only tests needing a dep this runner does not ship failed"; rc=0; fi; fi; exit "$rc"
31 changes: 31 additions & 0 deletions .github/workflows/staging-9858-windows-latest.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,31 @@
name: "staging-9858 windows-latest"
on:
push:
branches: ["pr-9858-xplat-ci"]
paths-ignore:
- '**/*.md'
workflow_dispatch:
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: true
permissions:
contents: read
defaults:
run:
shell: bash
jobs:
test:
runs-on: windows-latest
timeout-minutes: 30
env:
UNSLOTH_COMPILE_DISABLE: '1'
UNSLOTH_IS_PRESENT: '1'
steps:
- uses: actions/checkout@v4
with:
persist-credentials: false
- uses: actions/setup-python@v5
with:
python-version: '3.12'
- run: python -m pip install -r studio/backend/requirements/studio.txt pytest pytest-asyncio pytest-timeout python-multipart
- run: rc=0; PYTHONPATH=studio/backend python -m pytest studio/backend/tests/test_torch_cpu_build_on_nvidia_host.py tests/python/test_install_python_stack.py tests/studio/install/test_cuda_repair.py tests/studio/install/test_gpu_detection_followups.py tests/studio/install/test_pr5940_followups.py tests/studio/install/test_rocm_arch_table_parity.py tests/studio/install/test_rocm_support.py tests/studio/install/test_windows_torch_flavor_invariant.py tests/studio/test_xpu_triton_swap.py -q --tb=short -k 'not llama_cpp_load_progress_live and not TestGpuAutoSelection and not TestPreSpawnGpuResolution and not TestPerGpuFitGuardAllCounts and not TestTransformersIntrospection and not test_returns_cuda_when_cuda_available and not test_calls_cuda_cache_when_cuda' > pytest_out.txt 2>&1 || rc=$?; cat pytest_out.txt; if [ "$rc" = "5" ]; then echo "no tests ran (deps absent on this runner)"; rc=0; fi; if [ "$rc" = "2" ] && grep -qE "No module named .(torch|unsloth_zoo|transformers)." pytest_out.txt && ! grep -qE "^(FAILED|ERROR) " pytest_out.txt; then echo "collection needs a dep this runner does not ship"; rc=0; fi; if [ "$rc" = "1" ] && ! grep -qE "^ERROR " pytest_out.txt; then tot=$(grep -cE "^FAILED " pytest_out.txt); dep=$(grep -cE "^FAILED .* No module named .(torch|unsloth_zoo|transformers).$" pytest_out.txt); if [ "$tot" -gt 0 ] && [ "$tot" = "$dep" ]; then echo "only tests needing a dep this runner does not ship failed"; rc=0; fi; fi; exit "$rc"
18 changes: 18 additions & 0 deletions install.sh
Original file line number Diff line number Diff line change
Expand Up @@ -4559,6 +4559,15 @@ while [ -n "$_torch_index_leaf" ] && [ "${_torch_index_leaf%/}" != "$_torch_inde
done
_torch_index_leaf="${_torch_index_leaf##*/}"
_torch_index_leaf=$(printf '%s' "$_torch_index_leaf" | tr '[:upper:]' '[:lower:]')
# Whether the caller had already STATED a backend before the assignment below overwrites it.
# setup.sh documents UNSLOTH_TORCH_BACKEND=cpu as the way to keep a deliberate CPU install,
# and on a GPU-less host the resolved value is cpu too, so without this the manifest cannot
# tell a stated choice from the automatic answer.
if [ -n "${UNSLOTH_TORCH_BACKEND:-}" ]; then
_torch_backend_was_stated=true
else
_torch_backend_was_stated=false
fi
case "$_torch_index_leaf" in
rocm*|gfx*) export UNSLOTH_TORCH_BACKEND="rocm" ;;
cpu) export UNSLOTH_TORCH_BACKEND="cpu" ;;
Expand All @@ -4568,6 +4577,15 @@ case "$_torch_index_leaf" in
*) unset UNSLOTH_TORCH_BACKEND ;;
esac

# Derived from the index this script RESOLVED, which on a GPU-less machine is "cpu" whether
# or not anyone asked. Without the marker every ordinary Linux CPU install is recorded as a
# deliberate choice, and a machine that later gains a GPU is never offered the repair.
if [ -n "${UNSLOTH_TORCH_BACKEND:-}" ] && [ "$_torch_backend_was_stated" != true ]; then
export UNSLOTH_TORCH_BACKEND_SOURCE="resolved"
else
unset UNSLOTH_TORCH_BACKEND_SOURCE
fi

# Whether TORCH_INDEX_URL names an actual pip ROCm family (rocm<digit>* / gfx*), gating the
# ROCm-only side effects below (AMD bitsandbytes, ROCm-torch repair). Digit-gated so a leaf
# merely STARTING with "rocm" isn't force-repaired from the wrong path.
Expand Down
52 changes: 52 additions & 0 deletions studio/backend/core/inference/llama_cpp.py
Original file line number Diff line number Diff line change
Expand Up @@ -10610,6 +10610,10 @@ def _probe_or_none():
# only when an axis is actually quantized.
_tensor_quant_kv_unsupported_binaries: set[tuple[str, int]] = set()

# Binary dirs already reported by _warn_missing_windows_cuda_runtime. The env is rebuilt
# for every launch and every --list-devices probe, so one line per binary is enough.
_missing_cuda_runtime_warned: set[str] = set()

@classmethod
def _binary_key(cls, binary: Optional[str]) -> Optional[tuple[str, int]]:
"""(path, mtime_ns); ns mtime re-probes a same-second binary swap."""
Expand Down Expand Up @@ -10733,6 +10737,47 @@ def _add(path: Path) -> None:
_add(site_packages / "torch" / "lib")
return out

@classmethod
def _warn_missing_windows_cuda_runtime(cls, binary_dir: str, path_dirs: list[str]) -> None:
"""Say so when a CUDA llama-server has no cudart to load. Diagnostic only.

The CUDA prebuilt links ``cudart64_*.dll`` / ``cublas64_*.dll`` and takes them
from the managed venv -- ``torch/lib`` or the ``nvidia/*`` wheels, per
_windows_pip_nvidia_dll_dirs. A 2.11.0+cpu torch ships neither, so ggml cannot
load its CUDA backend and ``llama-server.exe --list-devices`` prints
``Available devices: (none)`` while UNSLOTH_PREBUILT_INFO.json still says
``backend cuda`` (#8473, HF discussion 87). Today that is entirely silent.

Changes nothing about the launch: the process still starts, still falls back to
CPU, and a custom build with the DLLs somewhere else is not second-guessed.
Never raises -- a diagnostic must not be able to stop a load.
"""
try:
if binary_dir in cls._missing_cuda_runtime_warned:
return
# Same identification _installed_ggml_backends uses: the official prebuilts are
# single-backend, so the ggml CUDA lib beside llama-server IS the build.
ggml_cuda = os.path.join(binary_dir, "ggml-cuda.dll")
if not os.path.isfile(ggml_cuda):
return
for directory in path_dirs:
try:
names = os.listdir(directory)
except OSError:
continue
if any(name.lower().startswith("cudart64_") for name in names):
return
cls._missing_cuda_runtime_warned.add(binary_dir)
logger.warning(
"llama.cpp is the CUDA build (%s) but no cudart64_*.dll was found on its "
"DLL search path. The CUDA ggml backend will not load and llama-server "
"will report no devices. This is what a CPU-only PyTorch in the managed "
"environment looks like; repair the installation to restore GPU support.",
ggml_cuda,
)
except Exception as e:
logger.debug(f"CUDA runtime DLL diagnostic failed: {e}")

@staticmethod
def _build_windows_path_dirs(binary_dir: str, prefix: str, cuda_path: str) -> list[str]:
"""Ordered PATH entries prepended so llama-server.exe resolves cudart /
Expand Down Expand Up @@ -10771,6 +10816,13 @@ def _llama_server_env_for_binary(
)
existing_path = env.get("PATH", "")
env["PATH"] = ";".join(path_dirs) + ";" + existing_path
# Warn against the FULL search path, inherited entries included: a hand-installed CUDA
# toolkit puts cudart64_*.dll on PATH without the venv or CUDA_PATH knowing, and warning
# on the prepended directories alone told working custom setups to repair a fine install.
LlamaCppBackend._warn_missing_windows_cuda_runtime(
binary_dir,
path_dirs + [d for d in existing_path.split(";") if d],
)

# ROCm: the prebuilt bundles rocblas.dll but NOT the Tensile
# kernel files (rocblas/library/*.dat + *.hsaco); the DLL searches
Expand Down
19 changes: 12 additions & 7 deletions studio/backend/main.py
Original file line number Diff line number Diff line change
Expand Up @@ -1561,11 +1561,15 @@ def _hardware_snapshot() -> Optional[tuple[bool, Optional[str], Optional[str]]]:
generation = _hw_module.DETECTION_GENERATION
device = _hw_module.DEVICE
chat_only = bool(_hw_module.CHAT_ONLY)
reason = getattr(_hw_module, "CHAT_ONLY_REASON", None)
# Inside the guarded read, with the reason it belongs to. Read after it, a forced
# re-detect starting in between would pair this reply's reason with a detail from
# a different pass, or with none at all.
detail = getattr(_hw_module, "CHAT_ONLY_DETAIL", None)
# Refreshed, not the frozen global: the three inventory-sensitive verdicts can change
# after startup (an eGPU attached, a driver that finished restarting). Reason and detail
# come back together, or a forced re-detect starting in between would pair this reply's
# reason with a detail from a different pass.
try:
reason, detail = _hw_module.current_chat_only_verdict()
except Exception:
reason = getattr(_hw_module, "CHAT_ONLY_REASON", None)
detail = getattr(_hw_module, "CHAT_ONLY_DETAIL", None)
if (
device is not None
and _hw_module.DETECTION_COMPLETE.is_set()
Expand Down Expand Up @@ -2057,8 +2061,9 @@ def _get_cached_system_gpu_info(logger) -> tuple[dict[str, Any], dict[str, Any]]
logger.debug(f"Could not resolve gpu_ids support: {e}")
llama_uses_vulkan = False
gpu_ids_supported = True
# Preserve backend/index metadata from the visibility probe: a CPU training host can expose
# a Vulkan inference GPU, and the UI must label it Vulkan, not the top-level CPU backend.
# The spread also carries `physical_devices` and `mismatch`: GPUs the OS sees that this PyTorch
# cannot open (#8473). They stay their own fields, because `devices` below is the runtime-usable
# list that model fit budgets against and the training device picker pins from.
gpu_info = {
**visibility_info,
"available": visibility_info.get("available", False),
Expand Down
25 changes: 25 additions & 0 deletions studio/backend/tests/conftest.py
Original file line number Diff line number Diff line change
Expand Up @@ -942,3 +942,28 @@ def real_prequant_safe_globals(monkeypatch):
monkeypatch.setattr(pq, "_SAFE_GLOBALS_REGISTERED", None)
monkeypatch.setattr(pq, "_RESOLVED_SAFE_GLOBALS", set())
return resolver


@pytest.fixture(autouse = True)
def _no_carried_over_hardware_measurements():
"""Both hardware caches start empty for every test, as they do in a fresh process.

The torch build snapshot and the physical GPU inventory are module globals with a
60 second TTL, so one test's host -- a suite that makes `import torch` fail, say --
would otherwise answer for every test that ran within a minute of it. Cleared
afterwards as well, so a test that warms one deliberately does not leak either.
"""
from utils.hardware import hardware as _hw

def _clear():
# Under the locks: a non-blocking read hands the refresh to a daemon thread that holds
# these while it writes, so clearing without waiting lets a previous test's REAL host land
# in the cache a moment later. Torch lock FIRST, then the inventory lock, because that is
# the order the background refresh takes them in.
with _hw._torch_build_snapshot_lock, _hw._physical_gpu_inventory_lock:
_hw._torch_build_snapshot_cache = None
_hw._physical_gpu_inventory_cache = None

_clear()
yield
_clear()
Loading
Loading