Skip to content

Studio macOS: non-blocking startup, drop obsolete macOS pin, MLX self-heal - #208

Closed
danielhanchen wants to merge 8 commits into
macos-studio-fixes-basefrom
macos-studio-fixes-validation
Closed

danielhanchen wants to merge 8 commits into
macos-studio-fixes-basefrom
macos-studio-fixes-validation

Conversation

@danielhanchen

Copy link
Copy Markdown
Owner

Throwaway staging PR to validate the macOS Studio fixes on a real Apple Silicon (macos-14) runner. Not for merge.

What this validates

  1. Slow startup on macOS. The llama.cpp capability + freshness probes (added in Studio: warn when llama.cpp prebuilt is too old for MTP unslothai/unsloth#5528/Studio: warn when llama.cpp prebuilt is at least 3 days behind unslothai/unsloth#5529, 2026-05-18) ran inline in the FastAPI lifespan, so a cold or slow GitHub freshness check blocked "Application startup complete". They now run on a daemon thread. Opt out with UNSLOTH_DISABLE_UPDATE_CHECK=1.
  2. Obsolete macOS prebuilt pin. pinned_macos_release_tag / b9415 only fired on the upstream ggml-org path, but published_repo_for_host routes every macOS host to the unslothai/llama.cpp fork (since Source llama.cpp prebuilts from unslothai/llama.cpp (CUDA, ROCm, macOS) unslothai/unsloth#5963), so it was dead code. Removed; the Mach-O minos preflight remains the backstop.
  3. Train and Export greyed out on macOS. Both are gated by CHAT_ONLY, which is false only when Apple Silicon and import mlx.core both hold. mlx is only pulled transitively via unsloth-zoo and the resolver can silently drop it. Added a background self-heal that reinstalls mlx/mlx-lm/mlx-vlm by name and re-detects (opt out UNSLOTH_DISABLE_MLX_AUTOREPAIR=1), plus chat_only_reason in /api/health and a sidebar hint.
  4. Harmless None.lower() guard in load_model_defaults.

Workflows

  • pytest (macos-14) and pytest (ubuntu-latest): deterministic backend tests, including the non-blocking-startup invariant and the MLX self-heal logic.
  • install + boot + MLX (macos-14): real install.sh with torch+mlx, boot timing, resolved versions, chat_only + chat_only_reason.
  • startup bisection (macos-14): boot timing across pre-Studio: warn when llama.cpp prebuilt is too old for MTP unslothai/unsloth#5528 / current main / this fix.
  • MLX self-heal e2e (macos-14): uninstall mlx, boot, assert it self-heals and chat_only flips false.

…X self-heal

- Defer the llama.cpp capability + freshness probes off the FastAPI lifespan
  critical path onto a daemon thread so they no longer gate Application startup
  complete (cold/slow GitHub freshness check was adding tens of seconds on macOS).
  Opt out with UNSLOTH_DISABLE_UPDATE_CHECK=1.
- Remove the obsolete upstream macOS prebuilt pin (b9415 / pinned_macos_release_tag);
  macOS now routes to the unslothai/llama.cpp fork's own Mac builds, and the Mach-O
  minos preflight remains the backstop.
- Report chat_only_reason in /api/health and show why Train/Export are greyed out in
  the sidebar; add a background MLX self-heal that reinstalls mlx/mlx-lm/mlx-vlm by
  name on Apple Silicon when it is missing, then re-detects (opt out
  UNSLOTH_DISABLE_MLX_AUTOREPAIR=1).
- Guard load_model_defaults against a None model id.
- Tests for all of the above.
… macOS fixes

Sweep the upstream workflows and add five throwaway staging workflows:
deterministic pytest (macos-14 + ubuntu-latest), real install+boot+MLX on macos-14,
startup-time bisection across pre-unslothai#5528 / main / fix, and an MLX self-heal e2e.
Not for merge.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces several improvements to the startup process and platform support. Key changes include moving the llama.cpp startup probes off the critical path to a background daemon thread to prevent startup delays, implementing an automatic background self-heal/reinstall mechanism for MLX on Apple Silicon when it is missing, and exposing a chat_only_reason via the health check API so the frontend can display helpful tooltips explaining why training/exporting is disabled. Additionally, the obsolete macOS upstream pin logic has been removed, and a guard was added to prevent errors when loading model defaults with a null or empty model ID. The review feedback highlights three important improvements: invalidating Python's import caches using importlib.invalidate_caches() after the background MLX installation to ensure the package is immediately importable, managing the lifecycle of the pip/uv installation subprocess with subprocess.Popen and an atexit handler to prevent orphaned background processes, and internationalizing the new user-facing tooltip strings in the sidebar using the t() translation function.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment on lines +46 to +52
def mlx_available() -> bool:
try:
import mlx.core # noqa: F401

return True
except Exception:
return False

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

In Python, if a package is installed via a subprocess while the parent process is already running, the import system's finders will not automatically detect the new package because their directory caches are not updated. To ensure import mlx.core succeeds after a successful background installation, you must call importlib.invalidate_caches() before attempting the import.

Suggested change
def mlx_available() -> bool:
try:
import mlx.core # noqa: F401
return True
except Exception:
return False
def mlx_available() -> bool:
try:
import importlib
importlib.invalidate_caches()
import mlx.core # noqa: F401
return True
except Exception:
return False

Comment on lines +63 to +86
def attempt_mlx_repair(*, timeout: int = _REPAIR_TIMEOUT_S) -> bool:
"""Install mlx + mlx-lm + mlx-vlm by name into the running venv. Best-effort;
returns True iff `import mlx.core` works afterwards."""
cmd = _pip_install_cmd("--upgrade", *MLX_PACKAGES)
logger.info("MLX self-heal: installing %s", ", ".join(MLX_PACKAGES))
try:
result = subprocess.run(
cmd,
stdout = subprocess.PIPE,
stderr = subprocess.STDOUT,
text = True,
timeout = timeout,
)
except subprocess.TimeoutExpired:
logger.warning("MLX self-heal timed out after %ss; staying chat-only", timeout)
return False
except Exception as exc: # pragma: no cover - environment dependent
logger.warning("MLX self-heal could not start: %s", exc)
return False
if result.returncode != 0:
tail = (result.stdout or "")[-2000:]
logger.warning("MLX self-heal failed (staying chat-only):\n%s", tail)
return False
return mlx_available()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Spawning pip or uv via subprocess.run in a daemon thread can lead to orphaned background processes if the main server process exits or is terminated. The daemon thread is abruptly killed, but the child pip process continues running in the background, holding a lock on the virtual environment (.venv). To prevent this, manage the subprocess lifecycle using subprocess.Popen and register an atexit handler to cleanly terminate the process on exit.

import atexit

_active_proc: subprocess.Popen[str] | None = None
_proc_lock = threading.Lock()


def attempt_mlx_repair(*, timeout: int = _REPAIR_TIMEOUT_S) -> bool:
    """Install mlx + mlx-lm + mlx-vlm by name into the running venv. Best-effort;
    returns True iff import mlx.core works afterwards."""
    global _active_proc
    cmd = _pip_install_cmd("--upgrade", *MLX_PACKAGES)
    logger.info("MLX self-heal: installing %s", ", ".join(MLX_PACKAGES))
    try:
        with _proc_lock:
            _active_proc = subprocess.Popen(
                cmd,
                stdout = subprocess.PIPE,
                stderr = subprocess.STDOUT,
                text = True,
            )
        try:
            stdout, _ = _active_proc.communicate(timeout = timeout)
            returncode = _active_proc.returncode
        except subprocess.TimeoutExpired:
            with _proc_lock:
                if _active_proc:
                    _active_proc.kill()
                    stdout, _ = _active_proc.communicate()
            logger.warning("MLX self-heal timed out after %ss; staying chat-only", timeout)
            return False
        finally:
            with _proc_lock:
                _active_proc = None
    except Exception as exc:  # pragma: no cover - environment dependent
        logger.warning("MLX self-heal could not start: %s", exc)
        return False

    if returncode != 0:
        tail = (stdout or "")[-2000:]
        logger.warning("MLX self-heal failed (staying chat-only):\n%s", tail)
        return False
    return mlx_available()


def terminate_mlx_repair() -> None:
    """Terminate the active MLX repair subprocess if running."""
    global _active_proc
    with _proc_lock:
        if _active_proc:
            logger.info("Terminating background MLX self-heal process...")
            _active_proc.terminate()
            try:
                _active_proc.wait(timeout = 5)
            except subprocess.TimeoutExpired:
                _active_proc.kill()
            _active_proc = None


atexit.register(terminate_mlx_repair)

Comment on lines +278 to +286
const trainExportDisabledHint: string | undefined = !chatOnly
? undefined
: chatOnlyReason === "mlx_unavailable"
? "Training needs MLX. Run `unsloth studio update` to enable Train and Export."
: chatOnlyReason === "intel_mac"
? "Training needs Apple Silicon or a GPU. Intel Macs are chat-only."
: chatOnlyReason === "no_gpu"
? "Training needs an NVIDIA or AMD GPU."
: undefined;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

These user-facing tooltip strings are hardcoded in English. To support internationalization (i18n) and adhere to the project's standards, use the t() translation function with appropriate keys.

Suggested change
const trainExportDisabledHint: string | undefined = !chatOnly
? undefined
: chatOnlyReason === "mlx_unavailable"
? "Training needs MLX. Run `unsloth studio update` to enable Train and Export."
: chatOnlyReason === "intel_mac"
? "Training needs Apple Silicon or a GPU. Intel Macs are chat-only."
: chatOnlyReason === "no_gpu"
? "Training needs an NVIDIA or AMD GPU."
: undefined;
const trainExportDisabledHint: string | undefined = !chatOnly
? undefined
: chatOnlyReason === "mlx_unavailable"
? t("shell.navigation.trainDisabledHint.mlx_unavailable")
: chatOnlyReason === "intel_mac"
? t("shell.navigation.trainDisabledHint.intel_mac")
: chatOnlyReason === "no_gpu"
? t("shell.navigation.trainDisabledHint.no_gpu")
: undefined;

… macos-14

Adapt studio-mac-ui-smoke (install + GGUF load + chat inference via playwright_chat_ui.py)
and mlx-ci (run_real_mlx_smoke.py: train gemma-3-270m + save lora/merged_16bit/gguf + reload)
to run against this branch on real Apple Silicon.
_install_llama_cpp_macos builds into ~/.unsloth/llama.cpp, so the gguf
reload locator must check LLAMA_CPP_DEFAULT_DIR, not only ./llama.cpp.
llama-cli timed out at 300s in the gguf reload step: a freshly cmake-built
Metal binary can stall on first-run shader init and some builds block on
inherited stdin despite -no-cnv. Run the load+generate smoke on CPU (-ngl 0,
sub-second for 270M) with stdin closed.
Validate the exported GGUF by parsing it with gguf.GGUFReader (proves it is a
loadable model) and make the llama-cli generation best-effort: a freshly built
binary can be slow on first run, so a timeout is a soft skip rather than a job
failure. Pairs with the zoo Release-build fix that speeds up llama-cli.
@danielhanchen

Copy link
Copy Markdown
Owner Author

Staging run finished; closing. Staging PRs exist to run CI on a spare queue and are never merged.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants