Skip to content

Studio: in-app Update llama.cpp button to install the latest prebuilt - #6097

Merged
danielhanchen merged 7 commits into
mainfrom
studio/llama-cpp-update-button
Jun 10, 2026
Merged

danielhanchen merged 7 commits into
mainfrom
studio/llama-cpp-update-button

Conversation

@danielhanchen

Copy link
Copy Markdown
Member

Summary

Adds an in-app "Update llama.cpp" affordance to Unsloth Studio. When the installed llama.cpp prebuilt is behind the latest published release, Studio shows a small, non-invasive banner with an "Update llama.cpp" button. Clicking it downloads the latest prebuilt and replaces the current one in the background, then reports when the new build is ready.

This is the action half of the existing freshness detection from #5529 (which only warned that the prebuilt was behind). Detection is unchanged; this adds the one-click update.

What it does

  • Backend utils/llama_cpp_update.py: get_update_status() reuses llama_cpp_freshness.check_prebuilt_freshness and adds update_available (installed tag != latest tag). start_update() runs on a daemon thread: unload the running model, run install_llama_prebuilt.py --llama-tag latest, reset caches, re-read the install marker. Fails open: if the installer cannot be located it returns a clear error and never disturbs inference.
  • Backend routes/llama.py: GET /api/llama/update-status (optional ?force_refresh) and POST /api/llama/update, registered at /api/llama.
  • Frontend hooks/use-llama-update-check.ts + components/llama-update-banner.tsx: a non-invasive bottom-right banner that appears shortly after load when an update is available, fades on its own after about 10 seconds, and re-surfaces roughly hourly as a gentle reminder. The button POSTs the swap and polls the job to completion. Mounted in app/provider.tsx for both the web and desktop shells, gated off on the same routes as the existing web update banner.

Testing

  • 6 hermetic backend tests (tests/test_llama_cpp_update.py): no-marker, update-available, up-to-date, refuse-without-marker, happy-path (installer invoked, job success), installer-failure. All pass.
  • End-to-end on an NVIDIA B200: fresh-installed older prebuilts (b9334 and b9493) via the installer, confirmed update_available, applied the update, and verified the install marker advanced to the latest tag and the new llama-server runs. Both cases pass, including a cuda12 to cuda13 line change.
  • The banner flow (appear, auto-hide, re-show, click to update) was exercised against a live Studio build.

Notes

  • Pairs with the prebuilt installer (studio/install_llama_prebuilt.py); the update simply re-runs it for the latest tag, so selection and validation are unchanged.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces in-app update capabilities for llama.cpp prebuilt binaries, adding backend endpoints to check update status and trigger background updates, along with a frontend banner and polling hook to notify users. The review feedback highlights a bug where polling is not initiated if an update is already running on page load. Additionally, it points out a layering violation and potential circular dependency caused by importing the inference backend inside the utility module, suggesting instead to pass the backend instance directly from the route handler to safely unload the model before updating.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

I am having trouble creating individual review comments. Click here to see my feedback.

studio/frontend/src/hooks/use-llama-update-check.ts (118-146)

high

Bug: Polling is not started if an update is already running on page load

If the page is loaded or refreshed while an update is already in progress (e.g., started in another tab or from a previous session), surfaceIfAvailable correctly detects next.job.state === 'running' and sets applying to true. However, the polling interval (pollTimer.current) is only ever started inside the apply() function. As a result, the banner will display the 'Updating llama.cpp...' status indefinitely without ever polling the backend to detect when the update completes.

Suggestion

Add a useEffect hook to automatically spin up the polling interval if we detect an active background job on mount or status change.

  useEffect(() => {
    if (!enabled) return;
    let canceled = false;

    const firstTimer = setTimeout(() => {
      fetchStatus().then((s) => {
        if (!canceled) surfaceIfAvailable(s);
      });
    }, FIRST_CHECK_DELAY_MS);

    const reminder = setInterval(() => {
      fetchStatus().then((s) => {
        if (!canceled) surfaceIfAvailable(s);
      });
    }, REMINDER_INTERVAL_MS);

    return () => {
      canceled = true;
      clearTimeout(firstTimer);
      clearInterval(reminder);
      clearHideTimer();
      if (pollTimer.current) {
        clearInterval(pollTimer.current);
        pollTimer.current = null;
      }
    };
  }, [enabled, surfaceIfAvailable, clearHideTimer]);

  // Automatically poll to completion if an update is already running on load
  useEffect(() => {
    if (!enabled) return;
    if (status?.job.state === "running" && !pollTimer.current) {
      setApplying(true);
      setVisible(true);
      clearHideTimer();
      pollTimer.current = setInterval(async () => {
        const s = await fetchStatus();
        if (!s) return;
        setStatus(s);
        if (s.job.state !== "running") {
          if (pollTimer.current) clearInterval(pollTimer.current);
          pollTimer.current = null;
          setApplying(false);
          if (s.job.state === "success") {
            setVisible(false);
          } else {
            armAutoHide();
          }
        }
      }, JOB_POLL_INTERVAL_MS);
    }
  }, [enabled, status?.job.state, armAutoHide, clearHideTimer]);

  const dismiss = useCallback(() => {
    clearHideTimer();
    setVisible(false);
  }, [clearHideTimer]);

studio/backend/utils/llama_cpp_update.py (150-163)

medium

Layering Violation & Circular Dependency

Importing get_llama_cpp_backend from routes.inference inside a utility module (utils.llama_cpp_update) introduces a layering violation and a potential circular dependency chain (since routes typically import from utilities).

Furthermore, calling get_llama_cpp_backend() inside a background thread without request context can fail if the dependency expects a FastAPI Request or state object. If it fails, the running model won't be unloaded, which will cause the update to fail on Windows because the active llama-server.exe binary will be locked.

Suggestion

Pass the backend instance directly from the route handler to start_update(), falling back to the deferred import only if not provided.

def _run_update(
    install_dir: Path,
    repo: str,
    from_tag: Optional[str],
    script: Path,
    backend: Optional[object] = None,
) -> None:
    """Worker: unload the running model, run the installer, refresh caches."""
    model_was_loaded = False
    try:
        # Release the binary so the atomic swap can replace files in use.
        try:
            if backend is None:
                from routes.inference import get_llama_cpp_backend

                backend = get_llama_cpp_backend()

            if backend is not None:
                model_was_loaded = bool(getattr(backend, "is_loaded", False))
                if model_was_loaded:
                    backend.unload_model()
        except Exception as exc:
            logger.debug("llama update: unload before install failed", error = str(exc))

studio/backend/utils/llama_cpp_update.py (217-266)

medium

Support passing backend to start_update

Update start_update to accept an optional backend instance and pass it to the background worker thread _run_update.

def start_update(backend: Optional[object] = None) -> dict:
    """Kick off a background update. Idempotent: a second call while one is
    running returns the in-flight job rather than starting another."""
    binary = _find_binary()
    install_dir = _install_dir_for(binary)
    marker = read_install_marker(binary)
    if install_dir is None or not marker:
        return {
            "started": False,
            "reason": "no_prebuilt_marker",
            "message": (
                "This llama.cpp install was not provisioned from an Unsloth "
                "prebuilt (source build or custom path); in-app update is "
                "unavailable."
            ),
            "job": get_update_status()["job"],
        }
    script = _installer_script()
    if script is None:
        return {
            "started": False,
            "reason": "installer_missing",
            "message": "install_llama_prebuilt.py could not be located.",
            "job": get_update_status()["job"],
        }
    repo = marker.get("published_repo") or DEFAULT_PUBLISHED_REPO
    from_tag = marker.get("tag") or marker.get("release_tag")

    with _job_lock:
        if _job["state"] == _JOB_RUNNING:
            return {"started": False, "reason": "already_running", "job": dict(_job)}
        _job.update(
            state = _JOB_RUNNING,
            message = "Downloading and installing the latest llama.cpp prebuilt...",
            from_tag = from_tag,
            to_tag = None,
            error = None,
            started_at = _utcnow(),
            finished_at = None,
        )
        job_snapshot = dict(_job)

    thread = threading.Thread(
        target = _run_update,
        args = (install_dir, repo, from_tag, script, backend),
        name = "llama-cpp-update",
        daemon = True,
    )
    thread.start()
    return {"started": True, "reason": None, "job": job_snapshot}

studio/backend/routes/llama.py (73-77)

medium

Pass backend to start_update

Pass the active llama_cpp_backend instance to start_update to ensure it can be safely unloaded before the update starts.

@router.post("/update", response_model = LlamaUpdateActionResponse)
async def llama_update(
    current_subject: str = Depends(get_current_subject),
) -> LlamaUpdateActionResponse:
    try:
        from routes.inference import get_llama_cpp_backend
        backend = get_llama_cpp_backend()
    except Exception:
        backend = None
    return LlamaUpdateActionResponse(**start_update(backend = backend))

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 9e07a79be9

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

setVisible(true);
clearHideTimer();
try {
const res = await authFetch("/api/llama/update", { method: "POST" });

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Handle refused update responses

When /api/llama/update returns HTTP 200 with started: false (for example installer_missing, or a marker disappearing between the status check and the click), this hook ignores the response body and starts polling for only success or error. The backend leaves the job in idle for those refusal cases, so the promise never resolves and the banner stays stuck in "Updating..." until the component unmounts; parse the action response and surface reason/message instead of entering the completion poll when no job was started.

Useful? React with 👍 / 👎.

Comment on lines +159 to +161
model_was_loaded = bool(getattr(backend, "is_loaded", False))
if model_was_loaded:
backend.unload_model()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Unload active llama-server processes before replacing binaries

This only unloads when is_loaded is true, but LlamaCppBackend also has an active-but-not-healthy/loading state where _process exists and is_loaded is false. If a user clicks update while a GGUF is still starting (or has become unhealthy), the installer runs with the existing llama-server process still alive; on Windows that can keep the executable locked and on other platforms it leaves a server running from the old binary during the swap. Check the backend's active process state (or call unload_model() whenever a process exists) before invoking the installer.

Useful? React with 👍 / 👎.

Comment on lines +308 to +310
<LlamaUpdateBanner
enabled={showApp && !HIDDEN_TITLEBAR_SIDEBAR_ROUTES.has(pathname)}
/>

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Mount the llama updater for macOS desktop sessions

In the Tauri branch this banner is added only after the shouldUseCustomWindowTitlebar() guard; when that guard is false, provider.tsx returns content earlier. I checked window-titlebar.tsx, where that guard explicitly returns false on macOS, so macOS desktop users never mount the new llama.cpp update banner and therefore cannot use the in-app update button even when the backend reports an available prebuilt update.

Useful? React with 👍 / 👎.

Comment on lines +103 to +108
if (next.job.state === "running") {
// A swap is in progress (e.g. started in another tab) -- keep it up.
setApplying(true);
setVisible(true);
clearHideTimer();
return;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Clear applying state for externally started jobs

When this tab observes job.state === "running" from a job started elsewhere (the comment even calls out another tab), it sets applying and keeps the banner visible but never starts the 3s completion poll. Later hourly status checks with success, error, or no update do not clear applying, so the banner remains stuck in the disabled "Updating..." state with no dismiss button; start polling when a running job is detected or clear applying when the job is no longer running.

Useful? React with 👍 / 👎.

except Exception as exc:
logger.debug("llama update: unload before install failed", error = str(exc))

cmd = [

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve ROCm installer overrides during updates

The setup scripts pass --has-rocm/--rocm-gfx specifically because install_llama_prebuilt.py can miss ROCm or the gfx target on amd-smi-only/name-inferred hosts; this in-app path reconstructs the installer command without those overrides. On an existing HIP prebuilt installed via those setup-detected flags, clicking Update can make the helper plan a non-HIP asset or fail instead of refreshing the same ROCm bundle, so the update should carry the marker/runtime ROCm selection (or rerun the setup detector) into the command.

Useful? React with 👍 / 👎.

except Exception as exc:
logger.debug("llama update: unload before install failed", error = str(exc))

cmd = [

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Use the same simple install policy as setup

This command omits --simple-policy, while both setup scripts include it when installing prebuilts. For installs whose marker points at ggml-org/llama.cpp (macOS, Windows, and CPU-only Linux), the helper's default policy looks for Unsloth-style manifest/checksum assets rather than directly scanning upstream release assets, so the in-app update can fail with no installable release even though setup.sh/setup.ps1 would successfully update the same installation.

Useful? React with 👍 / 👎.

The previous change re-exported llama_router from routes/__init__.py and
imported it via the `from routes import (...)` aggregator. That tripped two CI
gates:

- Backend tests: test_desktop_auth builds a hermetic fake `routes` module to
  import main.py in isolation, and that fake did not carry llama_router, so
  `from routes import (..., llama_router, ...)` failed.
- Source lint (import-hoist): a newly added re-export in routes/__init__.py is
  only referenced via __all__ (a string), which the checker counts as an unused
  hoisted import.

Import llama_router the same way settings_router is already imported -- directly
from its submodule (`from routes.llama import router as llama_router`) -- and
drop the routes/__init__.py re-export. Extend the test_desktop_auth fake to stub
routes.llama exactly like routes.settings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: f2a0fef06a

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

repo,
]
logger.info("llama update: installing", cmd = " ".join(cmd))
proc = subprocess.run(

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Block GGUF loads while the update is installing

While this long installer subprocess is running, the inference load path is still free to start a new llama-server because load_model() only serializes on its own _serial_load_lock, which this updater never takes. If a user clicks Update and then loads a GGUF before the download/validation/swap finishes, that new process can keep serving from the old executable (or lock it on Windows), so the job can report success while Studio is still running the pre-update binary. Put the backend into a maintenance/update state or coordinate with the llama backend lock for the whole install window.

Useful? React with 👍 / 👎.

shimmyshimmer and others added 3 commits June 10, 2026 16:19
Resolve conflicts from the #5963 installer rework:
- main.py: keep both the routes.llama and hub.* router imports/registrations
- provider.tsx: keep main's DownloadManagerPanel/TooltipProvider alongside the
  LlamaUpdateBanner, and mount the banner on the macOS desktop (native titlebar)
  path so macOS users get the in-app update
- test_desktop_auth.py: union the routes.llama and routes.prompts stub modules
…ller

Re-run install_llama_prebuilt.py the way setup.sh/setup.ps1 now do, and harden
the update flow against the edge cases raised in review.

Installer invocation (utils/llama_cpp_update.py):
- Forward the AMD gfx target the same way setup does. The install marker's asset
  name carries the gfx family for lemonade HIP bundles, so derive --rocm-gfx
  (or --has-rocm for the fork ROCm-version bundles) from it; without this, the
  update could mis-plan a non-HIP asset on amd-smi-only / name-inferred hosts.
- Do not pass --cpu-fallback: setup only uses it for the arm64 source-build
  rescue, and it force-drops GPU detection; re-running into the same install-dir
  and published-repo already reproduces the right CPU bundle.
- No --simple-policy (removed from setup by #5963; the default policy scans the
  published repo).

Update safety:
- Put the backend into a maintenance state for the install window: set a flag
  under the serial load lock and unload whenever a server process exists
  (is_active, not just is_loaded). load_model() rejects fast while the flag is
  set so no server starts from a half-swapped binary; the flag clears in a
  finally. Fails open when the backend is unavailable.
- Frontend hook: handle a 200 response that did not start a job (no marker /
  installer missing) by surfacing the reason instead of polling forever, and
  track an externally started job to completion so the banner does not stick on
  Updating. Shared poll helper used by both paths.

Tests: rocm-gfx / has-rocm derivation, ggml-org CPU marker has no --cpu-fallback,
no --simple-policy, backward-compatible marker without an asset field, refusal
paths, and the maintenance-flag set/clear plus fail-open coverage.
@danielhanchen

Copy link
Copy Markdown
Member Author

Rebased on main to pick up #5963 (the prebuilt installer now sourced from unslothai/llama.cpp) and reworked the update path to invoke install_llama_prebuilt.py the same way setup.sh / setup.ps1 do, plus the review fixes below.

Installer invocation now matches the published-binaries flow:

  • Forwards --published-repo from the marker, and for ROCm installs forwards --rocm-gfx (derived from the marker asset's gfx family, e.g. app-...-rocm-gfx110X...) or --has-rocm for the fork ROCm-version bundles, so an amd-smi-only or name-inferred host refreshes the same HIP bundle instead of mis-planning a non-HIP asset.
  • Does not pass --cpu-fallback: re-running into the same install dir and published repo reproduces the correct CPU bundle, and --cpu-fallback force-drops GPU detection (it is only setup's arm64 source-build rescue).
  • Does not pass --simple-policy: Source llama.cpp prebuilts from unslothai/llama.cpp (CUDA, ROCm, macOS) #5963 removed it from setup, and the installer's default policy now scans the published repo.

Review comments:

  • Unload active server before replacing binaries: now unloads whenever a server process exists (is_active), not only when fully loaded, so the loading / unhealthy window (and the Windows executable lock) is covered.
  • Block GGUF loads while the update is installing: the updater now sets a maintenance flag under the serial load lock and unloads, and load_model() rejects fast while the flag is set, so no server starts from a half-swapped binary. The flag clears in a finally, and the coordination fails open when the backend is unavailable.
  • Handle refused update responses: a 200 with started: false (no marker / installer missing) now surfaces the reason instead of entering the completion poll, so the banner no longer sticks on "Updating...".
  • Clear applying state for externally started jobs: a job observed as running (started elsewhere) is now tracked to completion via a shared poll, so the banner clears when it finishes.
  • Mount the updater for macOS desktop sessions: the banner is now mounted on the native-title-bar path too, so macOS desktop gets the in-app update.
  • Preserve ROCm installer overrides during updates: covered by the --rocm-gfx / --has-rocm forwarding above.
  • Use the same simple install policy as setup: no longer applicable. Source llama.cpp prebuilts from unslothai/llama.cpp (CUDA, ROCm, macOS) #5963 removed --simple-policy from both setup scripts, so the update intentionally omits it.

Hermetic backend tests cover the installer-argument construction, the refusal paths, and the maintenance-flag set/clear plus fail-open; the frontend type-check passes.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 6aa97ae7ef

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

"--published-repo",
repo,
]
cmd.extend(_rocm_install_args(asset))

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve CPU-fallback installs during updates

For Linux arm64 GPU machines where setup fell back to the ggml-org CPU prebuilt, the installed marker is still treated as supported here, but this reconstructed command never forwards the original --cpu-fallback choice. That flag is what makes install_llama_prebuilt.py clear GPU detection via _apply_host_overrides(... force_cpu=True); without it, the upstream resolver has no arm64+NVIDIA CPU branch, so clicking the in-app update fails instead of refreshing the existing CPU bundle.

Useful? React with 👍 / 👎.

Comment on lines +130 to +132
update_available = bool(
freshness.get("has_marker") and installed and latest and installed != latest
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Compare against the installable macOS fallback tag

On pre-macOS-26 machines with an older ggml-org/llama.cpp marker, install_llama_prebuilt.py deliberately resolves latest to the pinned b9415 fallback via pinned_macos_release_tag, but this status check compares the installed tag to GitHub's raw latest release. After a successful update/no-op to b9415, update_available remains true forever, so the hourly banner keeps prompting for an update that the installer will never apply; compute the same host-aware install target before surfacing the banner.

Useful? React with 👍 / 👎.

# (a live process also locks the exe on Windows during the swap).
if getattr(backend, "is_active", False):
model_was_active = True
backend.unload_model()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Avoid dropping the active model on failed updates

When a GGUF model is active and the installer later fails before replacing anything (for example a network, checksum, or no-compatible-asset error), this unload has already killed the user's running llama-server and the error path only records _JOB_ERROR; it never restores the model. Since the installer stages and validates before the final swap, delay the unload until the swap window or preserve/reload the prior model so a failed update does not unnecessarily interrupt the session.

Useful? React with 👍 / 👎.

@danielhanchen
danielhanchen merged commit dab0b77 into main Jun 10, 2026
31 of 34 checks passed
@danielhanchen
danielhanchen deleted the studio/llama-cpp-update-button branch June 10, 2026 17:04
danielhanchen added a commit that referenced this pull request Jun 10, 2026
…missed (#6162)

Follow-up to #6097. The update banner appeared 8s after load and auto-hid after
about 10s. Show it about 1s after a newer prebuilt is detected and keep it up
until the user dismisses it (click outside or the X) or runs the update; it
stays during an in-progress update so the progress is visible.

Co-authored-by: danielhanchen <michaelhan2050@gmail.com>
danielhanchen added a commit that referenced this pull request Jul 20, 2026
…back #7213 (#7228)

* test(studio): add e2e test for cpu-fallback overriding vulkan

* feat(studio): add UNSLOTH_LLAMA_CPP_BACKEND env var

* feat(studio): add UNSLOTH_LLAMA_CPP_BACKEND env var

* Preserve UNSLOTH_LLAMA_CPP_BACKEND=cpu across llama.cpp updates for PR #7228

The in-app updater rebuilt the installer command without --cpu-fallback and
only re-asserted Vulkan, so accepting a llama.cpp update after forcing CPU on
an Intel iGPU host re-ran host detection and routed back to the crashing Vulkan
bundle (#7213). Record install_kind in the prebuilt marker and re-assert
--cpu-fallback on update when the installed bundle is CPU.

Also make setup.sh's UNSLOTH_LLAMA_CPP_BACKEND check case-insensitive to match
setup.ps1, and add tests for the updater CPU preservation and the setup.sh flag
plumbing.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Trim and validate UNSLOTH_LLAMA_CPP_BACKEND, warn on unknown values for PR #7228

Trim surrounding whitespace and lowercase the value in both setup.sh and
setup.ps1, so values like ' cpu ' or 'CPU' still force the CPU-only prebuilt.
An unrecognized value (e.g. 'gpu') now prints a warning instead of silently
falling back to auto. Extend test_setup_llama_cpp_backend.py to cover both
scripts, including trimmed, empty and unknown values.

* Preserve arm64 CPU installs on update and honor CPU override in Windows prune for PR #7228

The update-path CPU preservation only matched install_kind ending in -cpu, so
arm64 CPU bundles (linux-arm64, windows-arm64) were re-routed to a GPU or source
build on update. Match the full set of CPU-only kinds instead.

Persisting install_kind also activated the previously inert Windows
mismatch-prune in setup.ps1: on a GPU host with UNSLOTH_LLAMA_CPP_BACKEND=cpu it
saw the windows-cpu marker as mismatched and deleted it every rerun. Normalize
the override once and make CPU expected so a deliberate CPU install is kept.
Extend the tests to cover both.

* Document legacy llama.cpp markers keep heal-to-GPU on update for PR #7228

Legacy prebuilt markers written before install_kind was persisted intentionally
do not force --cpu-fallback on update: the in-app updater lets them re-resolve
(heal to a GPU bundle) per the existing behavior from #6097, and only markers
that explicitly record a CPU install_kind are pinned to CPU. Add a comment and a
regression case documenting the boundary.

* Tighten llama.cpp CPU-fallback comments for PR #7228

* Fix Windows install-prune to keep valid Intel/fallback bundles for PR #7228

Persisting install_kind activated the setup.ps1 mismatch-prune, whose
expectedKinds was incomplete: the non-NVIDIA/non-AMD branch omitted
windows-vulkan (the Intel auto-route) and the GPU branches omitted the
windows-cpu/windows-arm64 fallback the installer uses when a GPU prebuilt is
missing. That made every setup rerun delete and re-download a valid Intel Vulkan
(or CPU-fallback) install. List all kinds the installer can produce per host so
only a bundle the host cannot run is pruned. Cover the full matrix in tests.

* Persist force_cpu marker flag so only forced CPU installs re-assert on update for PR #7228

* Add --force-cpu for deliberate CPU installs and warn on macOS for PR #7228

* Record force_cpu when reusing a matching CPU bundle for PR #7228

* Accept force_cpu keyword in installer test validator fakes for PR #7228

---------

Co-authored-by: danielhanchen <unslothai@gmail.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
VectorCipher added a commit to VectorCipher/unsloth that referenced this pull request Jul 20, 2026
…back unslothai#7213 (unslothai#7228)

* test(studio): add e2e test for cpu-fallback overriding vulkan

* feat(studio): add UNSLOTH_LLAMA_CPP_BACKEND env var

* feat(studio): add UNSLOTH_LLAMA_CPP_BACKEND env var

* Preserve UNSLOTH_LLAMA_CPP_BACKEND=cpu across llama.cpp updates for PR unslothai#7228

The in-app updater rebuilt the installer command without --cpu-fallback and
only re-asserted Vulkan, so accepting a llama.cpp update after forcing CPU on
an Intel iGPU host re-ran host detection and routed back to the crashing Vulkan
bundle (unslothai#7213). Record install_kind in the prebuilt marker and re-assert
--cpu-fallback on update when the installed bundle is CPU.

Also make setup.sh's UNSLOTH_LLAMA_CPP_BACKEND check case-insensitive to match
setup.ps1, and add tests for the updater CPU preservation and the setup.sh flag
plumbing.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Trim and validate UNSLOTH_LLAMA_CPP_BACKEND, warn on unknown values for PR unslothai#7228

Trim surrounding whitespace and lowercase the value in both setup.sh and
setup.ps1, so values like ' cpu ' or 'CPU' still force the CPU-only prebuilt.
An unrecognized value (e.g. 'gpu') now prints a warning instead of silently
falling back to auto. Extend test_setup_llama_cpp_backend.py to cover both
scripts, including trimmed, empty and unknown values.

* Preserve arm64 CPU installs on update and honor CPU override in Windows prune for PR unslothai#7228

The update-path CPU preservation only matched install_kind ending in -cpu, so
arm64 CPU bundles (linux-arm64, windows-arm64) were re-routed to a GPU or source
build on update. Match the full set of CPU-only kinds instead.

Persisting install_kind also activated the previously inert Windows
mismatch-prune in setup.ps1: on a GPU host with UNSLOTH_LLAMA_CPP_BACKEND=cpu it
saw the windows-cpu marker as mismatched and deleted it every rerun. Normalize
the override once and make CPU expected so a deliberate CPU install is kept.
Extend the tests to cover both.

* Document legacy llama.cpp markers keep heal-to-GPU on update for PR unslothai#7228

Legacy prebuilt markers written before install_kind was persisted intentionally
do not force --cpu-fallback on update: the in-app updater lets them re-resolve
(heal to a GPU bundle) per the existing behavior from unslothai#6097, and only markers
that explicitly record a CPU install_kind are pinned to CPU. Add a comment and a
regression case documenting the boundary.

* Tighten llama.cpp CPU-fallback comments for PR unslothai#7228

* Fix Windows install-prune to keep valid Intel/fallback bundles for PR unslothai#7228

Persisting install_kind activated the setup.ps1 mismatch-prune, whose
expectedKinds was incomplete: the non-NVIDIA/non-AMD branch omitted
windows-vulkan (the Intel auto-route) and the GPU branches omitted the
windows-cpu/windows-arm64 fallback the installer uses when a GPU prebuilt is
missing. That made every setup rerun delete and re-download a valid Intel Vulkan
(or CPU-fallback) install. List all kinds the installer can produce per host so
only a bundle the host cannot run is pruned. Cover the full matrix in tests.

* Persist force_cpu marker flag so only forced CPU installs re-assert on update for PR unslothai#7228

* Add --force-cpu for deliberate CPU installs and warn on macOS for PR unslothai#7228

* Record force_cpu when reusing a matching CPU bundle for PR unslothai#7228

* Accept force_cpu keyword in installer test validator fakes for PR unslothai#7228

---------

Co-authored-by: danielhanchen <unslothai@gmail.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants