Skip to content

Studio: warn when llama.cpp prebuilt is at least 3 days behind - #5529

Merged
danielhanchen merged 2 commits into
mainfrom
mtp-llama-prebuilt-staleness
May 18, 2026
Merged

danielhanchen merged 2 commits into
mainfrom
mtp-llama-prebuilt-staleness

Conversation

@danielhanchen

Copy link
Copy Markdown
Member

Summary

Layered on #5528. Generalises the MTP-specific staleness warning to every llama.cpp prebuilt update, not just the ones that add MTP. If the installed prebuilt is at least 3 days old AND its tag differs from the latest published tag on the helper release repo (default unslothai/llama.cpp), Studio nudges the user to run unsloth studio update.

How it works

Reads the install marker UNSLOTH_PREBUILT_INFO.json that install_llama_prebuilt.py already writes to <install_dir>. The marker carries:

{
  "tag": "b9190",
  "release_tag": "b9190",
  "published_repo": "unslothai/llama.cpp",
  "installed_at_utc": "2026-05-17T14:11:20Z",
  ...
}

Studio compares the installed tag against the latest published tag on the helper repo's GitHub releases endpoint, and checks now - installed_at_utc against the 3-day threshold.

GitHub fetch is cached at two levels:

  • Process-level memo for the /api/inference/status hot path.
  • Disk-level cache (24h TTL) at ~/.unsloth/studio/cache/llama_cpp_freshness/ so cold-start Studio launches don't always hit the API.

On a transient fetch failure (offline, rate-limited) we keep the last-good disk value alive rather than poisoning the cache with None. The check fails open: if anything is missing (marker, timestamp, GitHub response), stale stays False so users never see a misleading banner.

Surfaced in two places

1. Startup banner (logs + stderr) in main.py:lifespan(), alongside the MTP capability probe from #5528. Single line, e.g.:

WARNING: llama.cpp prebuilt is 5 days behind: installed b9190,
latest b9300. Run `unsloth studio update` to refresh.

2. /api/inference/status now returns three new fields:

llama_cpp_prebuilt_stale: bool
llama_cpp_installed_tag:  str | None
llama_cpp_latest_tag:     str | None

so the frontend can render a banner / popup with the actual tag delta the user is missing (e.g. "Installed b9190, latest b9300. Update?").

Why 3 days

Mirrors the typical Unsloth llama.cpp release cadence. Anything shorter nags users who restart Studio at the wrong moment; longer leaves real bugs sitting on the user's machine. Configurable via the threshold_days kwarg if a future call site wants a different window.

What this means for users

Scenario Before After
Just updated no banner no banner (unchanged)
1-2 days behind no banner no banner (within grace window)
3+ days behind, tag differs no signal startup log + stderr line + /status flag for UI banner
Behind but offline no signal last-cached value used; no banner if never cached (fail-open)
Source build / custom UNSLOTH_LLAMA_CPP_PATH (no marker) no signal no banner (fail-open)

Test plan

  • 17 new cases in test_llama_cpp_freshness.py:
    • Marker discovery in cmake (build/bin/llama-server) and root (./llama-server) layouts
    • Missing marker
    • Invalid JSON marker
    • None binary path
    • Disk cache hit after process memo reset
    • Network failure preserves stale-but-cached value (TTL passed)
    • Network failure with no cache returns None
    • stale=True when tag differs AND age >= threshold
    • stale=False when tag matches (regardless of age)
    • stale=False when behind by tag but within grace window
    • stale=False when GitHub unreachable
    • stale=False when installed_at_utc is unparseable
    • Custom threshold_days=1 triggers stale at 2-day age
    • Warning string contains unsloth studio update, both tags, and the day count
    • Singular "1 day" vs plural "N days" in the warning
  • Existing 188-test regression suite still passes (test_llama_server_args, test_llama_cpp_mtp_detection, test_gguf_reload_inheritance, test_gguf_metadata, test_llama_cpp_load_progress, test_llama_cpp_context_fit, test_inference_model_validation).

Stack

PR Branch Base Purpose
#5527 mtp-auto-spec-decoding main Auto-enable MTP speculative decoding
#5528 mtp-llama-cpp-update-warning mtp-auto-spec-decoding Warn when llama.cpp prebuilt is missing MTP support
this mtp-llama-prebuilt-staleness mtp-llama-cpp-update-warning Warn when prebuilt is 3+ days behind any release

This PR depends on #5528 (which depends on #5527). Merge order matters; this branch is based on mtp-llama-cpp-update-warning.

References

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: ec6e188af3

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

if not marker:
return out
out["has_marker"] = True
out["installed_tag"] = marker.get("tag") or marker.get("release_tag")

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Compare freshness against the installed release tag

The install marker written by install_llama_prebuilt.py stores tag as the upstream/requested llama ref and release_tag as the GitHub release tag, while latest_published_release() returns the GitHub release tag. For published branch builds where these differ (for example a latest release like llama-prebuilt-master-... with upstream tag master/b9174), this comparison marks a freshly installed latest prebuilt stale after 3 days and keeps showing the update warning even after unsloth studio update. Use the marker's release_tag for the freshness comparison/display.

Useful? React with 👍 / 👎.

return cached
p = Path(binary_path)
marker: Optional[dict] = None
for parent in (p.parent, *p.parents[:3]):

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Walk high enough to find Windows install markers

For the documented Windows CMake layout <install_dir>/build/bin/Release/llama-server.exe, this loop checks Release, bin, and build, but never reaches <install_dir> because p.parent duplicates p.parents[0]. The installer writes UNSLOTH_PREBUILT_INFO.json at <install_dir>, so Windows prebuilts in this layout always return no marker and never surface the intended freshness warning. Include one more parent or avoid the duplicated first entry.

Useful? React with 👍 / 👎.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a freshness check for the llama.cpp prebuilt, allowing the system to warn users when an update is overdue. The implementation includes a new utility for fetching and caching GitHub release tags, updates to the inference status API to surface staleness data, and a startup probe. Review feedback highlights a critical performance issue where synchronous I/O in an async route blocks the event loop, and suggests refining the directory traversal logic to correctly locate installation markers across different OS layouts.

try:
from utils.llama_cpp_freshness import check_prebuilt_freshness

_freshness = check_prebuilt_freshness(_bin)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Calling check_prebuilt_freshness directly inside an async route handler blocks the event loop because it performs synchronous network I/O (via urllib.request.urlopen) and disk I/O. This can cause the entire backend to freeze for up to 5 seconds (the fetch timeout) for all users whenever the 24h cache is expired or missing. Since this endpoint is polled frequently by the UI, this should be executed in a separate thread using asyncio.to_thread.

Suggested change
_freshness = check_prebuilt_freshness(_bin)
_freshness = await asyncio.to_thread(check_prebuilt_freshness, _bin)

return cached
p = Path(binary_path)
marker: Optional[dict] = None
for parent in (p.parent, *p.parents[:3]):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The current parent traversal logic only checks 3 distinct directory levels up from the binary (p.parent is identical to p.parents[0], and p.parents[:3] adds levels 1 and 2). However, the Windows CMake layout mentioned in the docstring (<install_dir>/build/bin/Release/llama-server.exe) requires ascending 4 levels to reach the install root where the marker file resides.

Additionally, the inclusion of p.parent alongside p.parents[0] is redundant.

Suggested change
for parent in (p.parent, *p.parents[:3]):
for parent in p.parents[:4]:
References
  1. To improve efficiency, avoid redundant file traversals.
  2. Maintain separation of concerns by implementing installer-like logic within utilities rather than importing from installer scripts.

@danielhanchen
danielhanchen force-pushed the mtp-llama-cpp-update-warning branch from 9b4b734 to 2bf0a23 Compare May 18, 2026 04:43
@danielhanchen
danielhanchen force-pushed the mtp-llama-prebuilt-staleness branch from 6b52c99 to 44473f5 Compare May 18, 2026 04:51

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 01189e5c78

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment on lines +52 to +54
cached = _marker_cache.get(binary_path)
if cached is not None or binary_path in _marker_cache:
return cached

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Invalidate install-marker cache after prebuilt updates

read_install_marker memoizes results by binary_path and then always returns the cached value without checking whether UNSLOTH_PREBUILT_INFO.json changed. If a user runs unsloth studio update while Studio is still running, the marker file is rewritten but /api/inference/status keeps reporting the old installed_tag/timestamp until process restart, so the stale-update banner can remain incorrectly active. This cache needs an invalidation signal (for example marker mtime) or to be bypassed for marker reads.

Useful? React with 👍 / 👎.

@danielhanchen
danielhanchen force-pushed the mtp-llama-prebuilt-staleness branch from 01189e5 to c2549b0 Compare May 18, 2026 05:08
@danielhanchen

Copy link
Copy Markdown
Member Author

Platform / install coverage

I went through each install path to map out what the freshness check sees. Summary:

OS / arch GPU _HELPER_RELEASE_REPO install_llama_prebuilt path published_repo in marker Marker written? Freshness fires?
Linux x86_64 NVIDIA CUDA unslothai/llama.cpp published prebuilt unslothai/llama.cpp yes yes
Linux x86_64 AMD ROCm unslothai/llama.cpp --resolve-source-build (upstream) unslothai/llama.cpp yes yes
Linux x86_64 CPU only ggml-org/llama.cpp published prebuilt (bin-ubuntu-x64) ggml-org/llama.cpp yes yes
Linux aarch64 any unslothai/llama.cpp source build via install_llama_prebuilt unslothai/llama.cpp yes yes
macOS x86_64 / arm64 Metal ggml-org/llama.cpp published prebuilt (bin-macos-*) ggml-org/llama.cpp yes yes
Windows x86_64 NVIDIA CUDA unslothai/llama.cpp published prebuilt (bin-win-cuda-*) unslothai/llama.cpp yes yes (after walk-depth fix below)
Windows x86_64 CPU only unslothai/llama.cpp (no CPU prebuilt -> falls through) source build depends depends depends
Any any n/a pure setup.sh cmake fallback (install_llama_prebuilt skipped entirely) n/a no no (fails open)
Any any n/a UNSLOTH_LLAMA_CPP_PATH pointed at a user-managed checkout n/a no no (fails open)

Mechanics:

  • The marker is written from exactly one place (install_llama_prebuilt.py line 5141, in _write_install_metadata). Every install path (source_label="published" and source_label="upstream") funnels through it, so any install routed through install_llama_prebuilt.py carries the marker regardless of OS or GPU.
  • published_repo in the marker is whichever repo the helper picked. The freshness check queries that exact repo's releases endpoint, so:
    • CUDA Linux / Windows / ROCm Linux compare against unslothai/llama.cpp's release cadence (which is what gates Unsloth-specific patches like MTP).
    • macOS and CPU Linux x86_64 compare against ggml-org/llama.cpp (upstream).
  • Source-only fallbacks have no marker by design, so the check returns has_marker=False, stale=False and no banner / log appears. Same for users running their own llama.cpp checkout via UNSLOTH_LLAMA_CPP_PATH.

One bug found while auditing this

The original walk depth in read_install_marker only checked 3 unique parent directories. That works for Linux/macOS cmake and root builds, but Windows cmake (the binary lives 4 levels below the install root) was missing the marker.

Pushed a fix: walk now uses p.parents[:5], which covers every layout _find_llama_server_binary() looks for. Added two tests:

  • test_read_install_marker_finds_windows_cmake_layout pins the 4-level Windows walk.
  • test_read_install_marker_carries_published_repo_dynamically (parametrized over both helper repos) pins that the freshness check queries whichever repo the marker recorded.

20 tests pass on the freshness module (17 original + 3 new). No behaviour changes for Linux/macOS users.

@danielhanchen

Copy link
Copy Markdown
Member Author

Full Cartesian coverage: {macOS, Windows, Linux, WSL2} x {CPU, CUDA, ROCm, Metal}

Walked through every code path against the actual setup.sh / setup.ps1 / install_llama_prebuilt.py logic and the live ggml-org/llama.cpp and unslothai/llama.cpp asset listings (latest tag b9204).

Possible combos

OS / arch Accel Helper repo Install path published_repo in marker Marker Freshness fires
macOS arm64 Metal ggml-org/llama.cpp published bin-macos-arm64.tar.gz ggml-org/llama.cpp yes yes
macOS x86_64 CPU / Metal ggml-org/llama.cpp published bin-macos-x64.tar.gz ggml-org/llama.cpp yes yes
Linux x86_64 CUDA unslothai/llama.cpp published Linux CUDA bundle unslothai/llama.cpp yes yes
Linux x86_64 ROCm unslothai/llama.cpp --resolve-source-build (upstream) unslothai/llama.cpp yes yes
Linux x86_64 CPU only ggml-org/llama.cpp published bin-ubuntu-x64.tar.gz ggml-org/llama.cpp yes yes
Linux aarch64 any unslothai/llama.cpp source build via install_llama_prebuilt unslothai/llama.cpp yes yes
Windows x86_64 CUDA ggml-org/llama.cpp published bin-win-cuda-{12.4,13.1}-x64.zip + cudart bundle ggml-org/llama.cpp yes yes
Windows x86_64 ROCm (HIP) ggml-org/llama.cpp published bin-win-hip-radeon-x64.zip ggml-org/llama.cpp yes yes
Windows x86_64 CPU ggml-org/llama.cpp published bin-win-cpu-x64.zip (also the CPU fallback when HIP missing) ggml-org/llama.cpp yes yes
WSL2 CUDA unslothai/llama.cpp same as Linux CUDA (nvidia-smi works in WSL2) unslothai/llama.cpp yes yes
WSL2 ROCm unslothai/llama.cpp same as Linux ROCm unslothai/llama.cpp yes yes
WSL2 CPU ggml-org/llama.cpp same as Linux CPU x86_64 (no GPU tool detected) ggml-org/llama.cpp yes yes

Physically impossible combos (excluded from the matrix)

OS Accel Reason
macOS CUDA No NVIDIA hardware path on modern macOS.
macOS ROCm No AMD ROCm path on macOS.
Windows Metal Metal is Apple-only.
Linux Metal Metal is Apple-only.
WSL2 Metal Metal is Apple-only.

Fails open by design

Path Marker Freshness fires
Pure setup.sh cmake fallback (when install_llama_prebuilt.py is skipped entirely, e.g. forced UNSLOTH_LLAMA_FORCE_COMPILE=1) no no (no banner)
UNSLOTH_LLAMA_CPP_PATH pointed at a user-managed checkout no no (no banner)
UNSLOTH_LLAMA_PR=... baked-in PR builds no no (no banner)

These users are tracking custom builds, not a published release, so silencing the warning is the right behaviour.

Routing rules (from the actual scripts)

setup.sh:668-676 (Linux / macOS / WSL):

if [ "$_HOST_SYSTEM" = "Darwin" ]; then
    _HELPER_RELEASE_REPO="ggml-org/llama.cpp"
elif [ "$_HOST_SYSTEM" = "Linux" ] \
        && [ "$_HOST_MACHINE" = "x86_64" ] \
        && [ "$_LINUX_HAS_GPU" = false ]; then
    _HELPER_RELEASE_REPO="ggml-org/llama.cpp"
else
    _HELPER_RELEASE_REPO="unslothai/llama.cpp"
fi

GPU tool detection: nvidia-smi, rocminfo, amd-smi, hipconfig, hipinfo.

setup.ps1:1977 (Windows):

$HelperReleaseRepo = "ggml-org/llama.cpp"

install_llama_prebuilt.py then picks the right asset for the host:

  • Windows CUDA -> bin-win-cuda-{runtime}-x64.zip + matching cudart-llama-bin-win-cuda-{runtime}-x64.zip (resolve_windows_cuda_choices).
  • Windows ROCm -> bin-win-hip-radeon-x64.zip, falls back to bin-win-cpu-x64.zip if HIP missing.
  • macOS -> bin-macos-arm64.tar.gz / bin-macos-x64.tar.gz.
  • Linux CUDA -> unslothai's bundle via direct_linux_release_plan.
  • Linux ROCm or aarch64 -> source build (--resolve-source-build) keeps the marker so freshness still fires.
  • Linux CPU x86_64 -> ggml-org's bin-ubuntu-x64.tar.gz.

Bottom line

Every supported combo of {macOS, Windows, Linux, WSL2} x {CPU, CUDA, ROCm, Metal} that Studio actually ships an install path for routes through install_llama_prebuilt.py and therefore writes UNSLOTH_PREBUILT_INFO.json. Freshness queries whichever repo the marker records, so:

  • Unsloth-cadence builds (Linux CUDA, Linux ROCm, Linux aarch64, WSL2 CUDA, WSL2 ROCm) compare against unslothai/llama.cpp's release cadence (which is what gates Unsloth-specific patches like MTP).
  • Upstream-cadence builds (macOS, Windows CUDA / ROCm / CPU, Linux CPU x86_64, WSL2 CPU) compare against ggml-org/llama.cpp's release cadence.

No additional code changes needed for the platform matrix; the Windows walk-depth fix in c2549b0 was the only real gap.

@danielhanchen
danielhanchen force-pushed the mtp-llama-cpp-update-warning branch from 2bf0a23 to 9b74d14 Compare May 18, 2026 07:18
@danielhanchen
danielhanchen force-pushed the mtp-llama-prebuilt-staleness branch from f8c7837 to 71ddc84 Compare May 18, 2026 07:18
Base automatically changed from mtp-llama-cpp-update-warning to main May 18, 2026 07:19
Layered on #5528. Generalises the MTP-specific staleness warning to
every llama.cpp prebuilt update, not just the ones that add MTP. If
the installed prebuilt is at least 3 days old AND its tag differs
from the latest published tag on the helper release repo (default
unslothai/llama.cpp), Studio nudges the user to run
"unsloth studio update".

How it works

Reads the install marker UNSLOTH_PREBUILT_INFO.json that
install_llama_prebuilt.py already writes to install_dir. The marker
carries the installed tag, the helper repo, and an installed_at_utc
timestamp. Studio compares those against the latest published tag
from the GitHub releases API for the helper repo.

GitHub fetch is cached at two levels:
- Process-level memo for /status hot path.
- Disk-level cache (24h TTL) at ~/.unsloth/studio/cache/llama_cpp_freshness/
  so cold-start Studio launches do not always hit the API.

On a transient fetch failure (offline, rate-limited) we keep the
last-good disk value alive rather than poisoning the cache with None.
The check fails open: if anything is missing (marker, timestamp,
GitHub response), stale stays False so users never see a misleading
banner.

Surfaced in two places

1. Startup banner (logs + stderr) in main.py:lifespan(), alongside the
   MTP capability probe added in #5528. Single line, e.g.:
     WARNING: llama.cpp prebuilt is 5 days behind: installed b9190,
     latest b9300. Run "unsloth studio update" to refresh.

2. /api/inference/status now returns:
     llama_cpp_prebuilt_stale: bool
     llama_cpp_installed_tag:  str | None
     llama_cpp_latest_tag:     str | None
   so the frontend can render a banner / popup with the actual tag
   delta the user is missing.

3-day threshold

Mirrors the typical Unsloth llama.cpp release cadence. Anything
shorter would nag users who restart Studio at the wrong moment;
longer leaves real bugs sitting on the user's machine. Configurable
via the threshold_days kwarg if a future call site wants a different
window.

Tests

17 new cases in tests/test_llama_cpp_freshness.py cover marker
discovery in both cmake and root install layouts, missing / invalid
marker, GitHub fetch caching across process restarts (disk cache hit
after the in-memory cache is reset), the stale / not-stale decision
matrix (tag mismatch + age threshold), fail-open behaviour when
GitHub is unreachable, custom threshold, singular/plural day in the
warning string, and unparseable installed_at_utc. The broader
205-test inference regression suite still passes.
@danielhanchen
danielhanchen force-pushed the mtp-llama-prebuilt-staleness branch from e96bcec to e57d4f1 Compare May 18, 2026 07:20
@danielhanchen
danielhanchen merged commit c690b28 into main May 18, 2026
31 checks passed
@danielhanchen
danielhanchen deleted the mtp-llama-prebuilt-staleness branch May 18, 2026 07:21

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 1334ff02c6

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

# <install>/llama-server (1 up)
# <install>/build/bin/llama-server (3 up, Linux/macOS cmake)
# <install>/build/bin/Release/llama-server.exe (4 up, Windows cmake)
for parent in p.parents[:5]:

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Avoid slicing Path.parents in Python 3.9

The loop for parent in p.parents[:5] relies on Path.parents slice support, which was added in Python 3.10; this repository still declares support for Python 3.9 (requires-python >=3.9). On 3.9, read_install_marker raises at runtime, so freshness checks fall into the fail-open exception path and users never get stale-prebuilt warnings in startup/status despite having markers.

Useful? React with 👍 / 👎.

danielhanchen added a commit that referenced this pull request Jun 10, 2026
…#6097)

Adds an in-app "Update llama.cpp" banner and button to Unsloth Studio. When the installed prebuilt is behind the latest published release, a non-invasive banner appears; clicking Update downloads the latest prebuilt for this host and swaps it in place in the background, with no restart.

Detection reuses the freshness check from #5529. The update re-runs install_llama_prebuilt.py the same way setup.sh and setup.ps1 do after #5963: it forwards the published repo and the AMD gfx target derived from the install marker, and does not pass the removed --simple-policy or the arm64-only --cpu-fallback.

While the installer swaps binaries the backend enters a maintenance state (flag set under the serial load lock, active server unloaded) so a concurrent load cannot start a server from a half-swapped binary; the next load uses the new build. The banner also handles refused responses and jobs started in another tab so it never sticks on "Updating...".

Verified end to end on an NVIDIA B200: installed b9493, detected the update, applied it, and confirmed the binary at the same path advanced to b9585 in the same process. Hermetic backend tests and the frontend type-check pass.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant