Studio: warn when llama.cpp prebuilt is at least 3 days behind - #5529
Conversation
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: ec6e188af3
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| if not marker: | ||
| return out | ||
| out["has_marker"] = True | ||
| out["installed_tag"] = marker.get("tag") or marker.get("release_tag") |
There was a problem hiding this comment.
Compare freshness against the installed release tag
The install marker written by install_llama_prebuilt.py stores tag as the upstream/requested llama ref and release_tag as the GitHub release tag, while latest_published_release() returns the GitHub release tag. For published branch builds where these differ (for example a latest release like llama-prebuilt-master-... with upstream tag master/b9174), this comparison marks a freshly installed latest prebuilt stale after 3 days and keeps showing the update warning even after unsloth studio update. Use the marker's release_tag for the freshness comparison/display.
Useful? React with 👍 / 👎.
| return cached | ||
| p = Path(binary_path) | ||
| marker: Optional[dict] = None | ||
| for parent in (p.parent, *p.parents[:3]): |
There was a problem hiding this comment.
Walk high enough to find Windows install markers
For the documented Windows CMake layout <install_dir>/build/bin/Release/llama-server.exe, this loop checks Release, bin, and build, but never reaches <install_dir> because p.parent duplicates p.parents[0]. The installer writes UNSLOTH_PREBUILT_INFO.json at <install_dir>, so Windows prebuilts in this layout always return no marker and never surface the intended freshness warning. Include one more parent or avoid the duplicated first entry.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Code Review
This pull request introduces a freshness check for the llama.cpp prebuilt, allowing the system to warn users when an update is overdue. The implementation includes a new utility for fetching and caching GitHub release tags, updates to the inference status API to surface staleness data, and a startup probe. Review feedback highlights a critical performance issue where synchronous I/O in an async route blocks the event loop, and suggests refining the directory traversal logic to correctly locate installation markers across different OS layouts.
| try: | ||
| from utils.llama_cpp_freshness import check_prebuilt_freshness | ||
|
|
||
| _freshness = check_prebuilt_freshness(_bin) |
There was a problem hiding this comment.
Calling check_prebuilt_freshness directly inside an async route handler blocks the event loop because it performs synchronous network I/O (via urllib.request.urlopen) and disk I/O. This can cause the entire backend to freeze for up to 5 seconds (the fetch timeout) for all users whenever the 24h cache is expired or missing. Since this endpoint is polled frequently by the UI, this should be executed in a separate thread using asyncio.to_thread.
| _freshness = check_prebuilt_freshness(_bin) | |
| _freshness = await asyncio.to_thread(check_prebuilt_freshness, _bin) |
| return cached | ||
| p = Path(binary_path) | ||
| marker: Optional[dict] = None | ||
| for parent in (p.parent, *p.parents[:3]): |
There was a problem hiding this comment.
The current parent traversal logic only checks 3 distinct directory levels up from the binary (p.parent is identical to p.parents[0], and p.parents[:3] adds levels 1 and 2). However, the Windows CMake layout mentioned in the docstring (<install_dir>/build/bin/Release/llama-server.exe) requires ascending 4 levels to reach the install root where the marker file resides.
Additionally, the inclusion of p.parent alongside p.parents[0] is redundant.
| for parent in (p.parent, *p.parents[:3]): | |
| for parent in p.parents[:4]: |
References
- To improve efficiency, avoid redundant file traversals.
- Maintain separation of concerns by implementing installer-like logic within utilities rather than importing from installer scripts.
9b4b734 to
2bf0a23
Compare
6b52c99 to
44473f5
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 01189e5c78
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| cached = _marker_cache.get(binary_path) | ||
| if cached is not None or binary_path in _marker_cache: | ||
| return cached |
There was a problem hiding this comment.
Invalidate install-marker cache after prebuilt updates
read_install_marker memoizes results by binary_path and then always returns the cached value without checking whether UNSLOTH_PREBUILT_INFO.json changed. If a user runs unsloth studio update while Studio is still running, the marker file is rewritten but /api/inference/status keeps reporting the old installed_tag/timestamp until process restart, so the stale-update banner can remain incorrectly active. This cache needs an invalidation signal (for example marker mtime) or to be bypassed for marker reads.
Useful? React with 👍 / 👎.
01189e5 to
c2549b0
Compare
Platform / install coverageI went through each install path to map out what the freshness check sees. Summary:
Mechanics:
One bug found while auditing thisThe original walk depth in Pushed a fix: walk now uses
20 tests pass on the freshness module (17 original + 3 new). No behaviour changes for Linux/macOS users. |
Full Cartesian coverage: {macOS, Windows, Linux, WSL2} x {CPU, CUDA, ROCm, Metal}Walked through every code path against the actual Possible combos
Physically impossible combos (excluded from the matrix)
Fails open by design
These users are tracking custom builds, not a published release, so silencing the warning is the right behaviour. Routing rules (from the actual scripts)
if [ "$_HOST_SYSTEM" = "Darwin" ]; then
_HELPER_RELEASE_REPO="ggml-org/llama.cpp"
elif [ "$_HOST_SYSTEM" = "Linux" ] \
&& [ "$_HOST_MACHINE" = "x86_64" ] \
&& [ "$_LINUX_HAS_GPU" = false ]; then
_HELPER_RELEASE_REPO="ggml-org/llama.cpp"
else
_HELPER_RELEASE_REPO="unslothai/llama.cpp"
fiGPU tool detection:
$HelperReleaseRepo = "ggml-org/llama.cpp"
Bottom lineEvery supported combo of {macOS, Windows, Linux, WSL2} x {CPU, CUDA, ROCm, Metal} that Studio actually ships an install path for routes through
No additional code changes needed for the platform matrix; the Windows walk-depth fix in |
2bf0a23 to
9b74d14
Compare
f8c7837 to
71ddc84
Compare
Layered on #5528. Generalises the MTP-specific staleness warning to every llama.cpp prebuilt update, not just the ones that add MTP. If the installed prebuilt is at least 3 days old AND its tag differs from the latest published tag on the helper release repo (default unslothai/llama.cpp), Studio nudges the user to run "unsloth studio update". How it works Reads the install marker UNSLOTH_PREBUILT_INFO.json that install_llama_prebuilt.py already writes to install_dir. The marker carries the installed tag, the helper repo, and an installed_at_utc timestamp. Studio compares those against the latest published tag from the GitHub releases API for the helper repo. GitHub fetch is cached at two levels: - Process-level memo for /status hot path. - Disk-level cache (24h TTL) at ~/.unsloth/studio/cache/llama_cpp_freshness/ so cold-start Studio launches do not always hit the API. On a transient fetch failure (offline, rate-limited) we keep the last-good disk value alive rather than poisoning the cache with None. The check fails open: if anything is missing (marker, timestamp, GitHub response), stale stays False so users never see a misleading banner. Surfaced in two places 1. Startup banner (logs + stderr) in main.py:lifespan(), alongside the MTP capability probe added in #5528. Single line, e.g.: WARNING: llama.cpp prebuilt is 5 days behind: installed b9190, latest b9300. Run "unsloth studio update" to refresh. 2. /api/inference/status now returns: llama_cpp_prebuilt_stale: bool llama_cpp_installed_tag: str | None llama_cpp_latest_tag: str | None so the frontend can render a banner / popup with the actual tag delta the user is missing. 3-day threshold Mirrors the typical Unsloth llama.cpp release cadence. Anything shorter would nag users who restart Studio at the wrong moment; longer leaves real bugs sitting on the user's machine. Configurable via the threshold_days kwarg if a future call site wants a different window. Tests 17 new cases in tests/test_llama_cpp_freshness.py cover marker discovery in both cmake and root install layouts, missing / invalid marker, GitHub fetch caching across process restarts (disk cache hit after the in-memory cache is reset), the stale / not-stale decision matrix (tag mismatch + age threshold), fail-open behaviour when GitHub is unreachable, custom threshold, singular/plural day in the warning string, and unparseable installed_at_utc. The broader 205-test inference regression suite still passes.
e96bcec to
e57d4f1
Compare
for more information, see https://pre-commit.ci
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 1334ff02c6
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| # <install>/llama-server (1 up) | ||
| # <install>/build/bin/llama-server (3 up, Linux/macOS cmake) | ||
| # <install>/build/bin/Release/llama-server.exe (4 up, Windows cmake) | ||
| for parent in p.parents[:5]: |
There was a problem hiding this comment.
Avoid slicing Path.parents in Python 3.9
The loop for parent in p.parents[:5] relies on Path.parents slice support, which was added in Python 3.10; this repository still declares support for Python 3.9 (requires-python >=3.9). On 3.9, read_install_marker raises at runtime, so freshness checks fall into the fail-open exception path and users never get stale-prebuilt warnings in startup/status despite having markers.
Useful? React with 👍 / 👎.
…#6097) Adds an in-app "Update llama.cpp" banner and button to Unsloth Studio. When the installed prebuilt is behind the latest published release, a non-invasive banner appears; clicking Update downloads the latest prebuilt for this host and swaps it in place in the background, with no restart. Detection reuses the freshness check from #5529. The update re-runs install_llama_prebuilt.py the same way setup.sh and setup.ps1 do after #5963: it forwards the published repo and the AMD gfx target derived from the install marker, and does not pass the removed --simple-policy or the arm64-only --cpu-fallback. While the installer swaps binaries the backend enters a maintenance state (flag set under the serial load lock, active server unloaded) so a concurrent load cannot start a server from a half-swapped binary; the next load uses the new build. The banner also handles refused responses and jobs started in another tab so it never sticks on "Updating...". Verified end to end on an NVIDIA B200: installed b9493, detected the update, applied it, and confirmed the binary at the same path advanced to b9585 in the same process. Hermetic backend tests and the frontend type-check pass.
Summary
Layered on #5528. Generalises the MTP-specific staleness warning to every
llama.cppprebuilt update, not just the ones that add MTP. If the installed prebuilt is at least 3 days old AND its tag differs from the latest published tag on the helper release repo (defaultunslothai/llama.cpp), Studio nudges the user to rununsloth studio update.How it works
Reads the install marker
UNSLOTH_PREBUILT_INFO.jsonthatinstall_llama_prebuilt.pyalready writes to<install_dir>. The marker carries:{ "tag": "b9190", "release_tag": "b9190", "published_repo": "unslothai/llama.cpp", "installed_at_utc": "2026-05-17T14:11:20Z", ... }Studio compares the installed
tagagainst the latest published tag on the helper repo's GitHub releases endpoint, and checksnow - installed_at_utcagainst the 3-day threshold.GitHub fetch is cached at two levels:
/api/inference/statushot path.~/.unsloth/studio/cache/llama_cpp_freshness/so cold-start Studio launches don't always hit the API.On a transient fetch failure (offline, rate-limited) we keep the last-good disk value alive rather than poisoning the cache with
None. The check fails open: if anything is missing (marker, timestamp, GitHub response),stalestaysFalseso users never see a misleading banner.Surfaced in two places
1. Startup banner (logs + stderr) in
main.py:lifespan(), alongside the MTP capability probe from #5528. Single line, e.g.:2.
/api/inference/statusnow returns three new fields:so the frontend can render a banner / popup with the actual tag delta the user is missing (e.g. "Installed b9190, latest b9300. Update?").
Why 3 days
Mirrors the typical Unsloth
llama.cpprelease cadence. Anything shorter nags users who restart Studio at the wrong moment; longer leaves real bugs sitting on the user's machine. Configurable via thethreshold_dayskwarg if a future call site wants a different window.What this means for users
/statusflag for UI bannerUNSLOTH_LLAMA_CPP_PATH(no marker)Test plan
test_llama_cpp_freshness.py:build/bin/llama-server) and root (./llama-server) layoutsNonebinary pathNonestale=Truewhen tag differs AND age >= thresholdstale=Falsewhen tag matches (regardless of age)stale=Falsewhen behind by tag but within grace windowstale=Falsewhen GitHub unreachablestale=Falsewheninstalled_at_utcis unparseablethreshold_days=1triggers stale at 2-day ageunsloth studio update, both tags, and the day counttest_llama_server_args,test_llama_cpp_mtp_detection,test_gguf_reload_inheritance,test_gguf_metadata,test_llama_cpp_load_progress,test_llama_cpp_context_fit,test_inference_model_validation).Stack
mtp-auto-spec-decodingmainmtp-llama-cpp-update-warningmtp-auto-spec-decodingmtp-llama-prebuilt-stalenessmtp-llama-cpp-update-warningThis PR depends on #5528 (which depends on #5527). Merge order matters; this branch is based on
mtp-llama-cpp-update-warning.References
install_llama_prebuilt.pyinstall marker schema (this repo)