Studio macOS: non-blocking startup, drop obsolete macOS pin, MLX self-heal - #208
danielhanchen wants to merge 8 commits into
Conversation
…X self-heal - Defer the llama.cpp capability + freshness probes off the FastAPI lifespan critical path onto a daemon thread so they no longer gate Application startup complete (cold/slow GitHub freshness check was adding tens of seconds on macOS). Opt out with UNSLOTH_DISABLE_UPDATE_CHECK=1. - Remove the obsolete upstream macOS prebuilt pin (b9415 / pinned_macos_release_tag); macOS now routes to the unslothai/llama.cpp fork's own Mac builds, and the Mach-O minos preflight remains the backstop. - Report chat_only_reason in /api/health and show why Train/Export are greyed out in the sidebar; add a background MLX self-heal that reinstalls mlx/mlx-lm/mlx-vlm by name on Apple Silicon when it is missing, then re-detects (opt out UNSLOTH_DISABLE_MLX_AUTOREPAIR=1). - Guard load_model_defaults against a None model id. - Tests for all of the above.
… macOS fixes Sweep the upstream workflows and add five throwaway staging workflows: deterministic pytest (macos-14 + ubuntu-latest), real install+boot+MLX on macos-14, startup-time bisection across pre-unslothai#5528 / main / fix, and an MLX self-heal e2e. Not for merge.
There was a problem hiding this comment.
Code Review
This pull request introduces several improvements to the startup process and platform support. Key changes include moving the llama.cpp startup probes off the critical path to a background daemon thread to prevent startup delays, implementing an automatic background self-heal/reinstall mechanism for MLX on Apple Silicon when it is missing, and exposing a chat_only_reason via the health check API so the frontend can display helpful tooltips explaining why training/exporting is disabled. Additionally, the obsolete macOS upstream pin logic has been removed, and a guard was added to prevent errors when loading model defaults with a null or empty model ID. The review feedback highlights three important improvements: invalidating Python's import caches using importlib.invalidate_caches() after the background MLX installation to ensure the package is immediately importable, managing the lifecycle of the pip/uv installation subprocess with subprocess.Popen and an atexit handler to prevent orphaned background processes, and internationalizing the new user-facing tooltip strings in the sidebar using the t() translation function.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
| def mlx_available() -> bool: | ||
| try: | ||
| import mlx.core # noqa: F401 | ||
|
|
||
| return True | ||
| except Exception: | ||
| return False |
There was a problem hiding this comment.
In Python, if a package is installed via a subprocess while the parent process is already running, the import system's finders will not automatically detect the new package because their directory caches are not updated. To ensure import mlx.core succeeds after a successful background installation, you must call importlib.invalidate_caches() before attempting the import.
| def mlx_available() -> bool: | |
| try: | |
| import mlx.core # noqa: F401 | |
| return True | |
| except Exception: | |
| return False | |
| def mlx_available() -> bool: | |
| try: | |
| import importlib | |
| importlib.invalidate_caches() | |
| import mlx.core # noqa: F401 | |
| return True | |
| except Exception: | |
| return False |
| def attempt_mlx_repair(*, timeout: int = _REPAIR_TIMEOUT_S) -> bool: | ||
| """Install mlx + mlx-lm + mlx-vlm by name into the running venv. Best-effort; | ||
| returns True iff `import mlx.core` works afterwards.""" | ||
| cmd = _pip_install_cmd("--upgrade", *MLX_PACKAGES) | ||
| logger.info("MLX self-heal: installing %s", ", ".join(MLX_PACKAGES)) | ||
| try: | ||
| result = subprocess.run( | ||
| cmd, | ||
| stdout = subprocess.PIPE, | ||
| stderr = subprocess.STDOUT, | ||
| text = True, | ||
| timeout = timeout, | ||
| ) | ||
| except subprocess.TimeoutExpired: | ||
| logger.warning("MLX self-heal timed out after %ss; staying chat-only", timeout) | ||
| return False | ||
| except Exception as exc: # pragma: no cover - environment dependent | ||
| logger.warning("MLX self-heal could not start: %s", exc) | ||
| return False | ||
| if result.returncode != 0: | ||
| tail = (result.stdout or "")[-2000:] | ||
| logger.warning("MLX self-heal failed (staying chat-only):\n%s", tail) | ||
| return False | ||
| return mlx_available() |
There was a problem hiding this comment.
Spawning pip or uv via subprocess.run in a daemon thread can lead to orphaned background processes if the main server process exits or is terminated. The daemon thread is abruptly killed, but the child pip process continues running in the background, holding a lock on the virtual environment (.venv). To prevent this, manage the subprocess lifecycle using subprocess.Popen and register an atexit handler to cleanly terminate the process on exit.
import atexit
_active_proc: subprocess.Popen[str] | None = None
_proc_lock = threading.Lock()
def attempt_mlx_repair(*, timeout: int = _REPAIR_TIMEOUT_S) -> bool:
"""Install mlx + mlx-lm + mlx-vlm by name into the running venv. Best-effort;
returns True iff import mlx.core works afterwards."""
global _active_proc
cmd = _pip_install_cmd("--upgrade", *MLX_PACKAGES)
logger.info("MLX self-heal: installing %s", ", ".join(MLX_PACKAGES))
try:
with _proc_lock:
_active_proc = subprocess.Popen(
cmd,
stdout = subprocess.PIPE,
stderr = subprocess.STDOUT,
text = True,
)
try:
stdout, _ = _active_proc.communicate(timeout = timeout)
returncode = _active_proc.returncode
except subprocess.TimeoutExpired:
with _proc_lock:
if _active_proc:
_active_proc.kill()
stdout, _ = _active_proc.communicate()
logger.warning("MLX self-heal timed out after %ss; staying chat-only", timeout)
return False
finally:
with _proc_lock:
_active_proc = None
except Exception as exc: # pragma: no cover - environment dependent
logger.warning("MLX self-heal could not start: %s", exc)
return False
if returncode != 0:
tail = (stdout or "")[-2000:]
logger.warning("MLX self-heal failed (staying chat-only):\n%s", tail)
return False
return mlx_available()
def terminate_mlx_repair() -> None:
"""Terminate the active MLX repair subprocess if running."""
global _active_proc
with _proc_lock:
if _active_proc:
logger.info("Terminating background MLX self-heal process...")
_active_proc.terminate()
try:
_active_proc.wait(timeout = 5)
except subprocess.TimeoutExpired:
_active_proc.kill()
_active_proc = None
atexit.register(terminate_mlx_repair)| const trainExportDisabledHint: string | undefined = !chatOnly | ||
| ? undefined | ||
| : chatOnlyReason === "mlx_unavailable" | ||
| ? "Training needs MLX. Run `unsloth studio update` to enable Train and Export." | ||
| : chatOnlyReason === "intel_mac" | ||
| ? "Training needs Apple Silicon or a GPU. Intel Macs are chat-only." | ||
| : chatOnlyReason === "no_gpu" | ||
| ? "Training needs an NVIDIA or AMD GPU." | ||
| : undefined; |
There was a problem hiding this comment.
These user-facing tooltip strings are hardcoded in English. To support internationalization (i18n) and adhere to the project's standards, use the t() translation function with appropriate keys.
| const trainExportDisabledHint: string | undefined = !chatOnly | |
| ? undefined | |
| : chatOnlyReason === "mlx_unavailable" | |
| ? "Training needs MLX. Run `unsloth studio update` to enable Train and Export." | |
| : chatOnlyReason === "intel_mac" | |
| ? "Training needs Apple Silicon or a GPU. Intel Macs are chat-only." | |
| : chatOnlyReason === "no_gpu" | |
| ? "Training needs an NVIDIA or AMD GPU." | |
| : undefined; | |
| const trainExportDisabledHint: string | undefined = !chatOnly | |
| ? undefined | |
| : chatOnlyReason === "mlx_unavailable" | |
| ? t("shell.navigation.trainDisabledHint.mlx_unavailable") | |
| : chatOnlyReason === "intel_mac" | |
| ? t("shell.navigation.trainDisabledHint.intel_mac") | |
| : chatOnlyReason === "no_gpu" | |
| ? t("shell.navigation.trainDisabledHint.no_gpu") | |
| : undefined; |
… macos-14 Adapt studio-mac-ui-smoke (install + GGUF load + chat inference via playwright_chat_ui.py) and mlx-ci (run_real_mlx_smoke.py: train gemma-3-270m + save lora/merged_16bit/gguf + reload) to run against this branch on real Apple Silicon.
_install_llama_cpp_macos builds into ~/.unsloth/llama.cpp, so the gguf reload locator must check LLAMA_CPP_DEFAULT_DIR, not only ./llama.cpp.
llama-cli timed out at 300s in the gguf reload step: a freshly cmake-built Metal binary can stall on first-run shader init and some builds block on inherited stdin despite -no-cnv. Run the load+generate smoke on CPU (-ngl 0, sub-second for 270M) with stdin closed.
Validate the exported GGUF by parsing it with gguf.GGUFReader (proves it is a loadable model) and make the llama-cli generation best-effort: a freshly built binary can be slow on first run, so a timeout is a soft skip rather than a job failure. Pairs with the zoo Release-build fix that speeds up llama-cli.
|
Staging run finished; closing. Staging PRs exist to run CI on a spare queue and are never merged. |
Throwaway staging PR to validate the macOS Studio fixes on a real Apple Silicon (macos-14) runner. Not for merge.
What this validates
Workflows