Skip to content

Studio: warn when llama.cpp prebuilt is too old for MTP - #5528

Merged
danielhanchen merged 2 commits into
mainfrom
mtp-llama-cpp-update-warning
May 18, 2026
Merged

danielhanchen merged 2 commits into
mainfrom
mtp-llama-cpp-update-warning

Conversation

@danielhanchen

Copy link
Copy Markdown
Member

Summary

Layered on #5527. Now that Studio auto-enables MTP speculative decoding for MTP GGUFs, users running an outdated llama-server prebuilt would silently miss the speedup (or worse, get an unknown-flag error on load). This adds a one-shot capability probe so we can:

  1. Log a warning at startup with a clear next-step (unsloth studio update).
  2. Skip the auto-emit gracefully on the load hot path when MTP isn't available, so MTP GGUFs still load (just without spec decoding).
  3. Expose a llama_cpp_supports_mtp flag on /api/inference/status so the frontend can render a banner / popup.

How the probe works

LlamaCppBackend.probe_server_capabilities() runs llama-server --help once and looks for MTP in the --spec-type enum line.

  • Result is cached at the class level keyed on (binary_path, mtime). One subprocess call the first time, instant thereafter.
  • An unsloth studio update replaces the binary -> mtime changes -> cache key changes -> next call re-probes automatically. No process restart needed.
  • Recognises both upstream naming variants:
    • draft-mtp from the original llama.cpp PR #22673
    • mtp from later upstream commits that dropped the draft- prefix
  • The spec block uses whichever token the binary accepts, so Studio's emitted flag adapts to the user's actual prebuilt.

Surfaced in three places

1. Startup banner (logs + stderr) in main.py:lifespan():

WARNING: llama.cpp prebuilt is missing MTP support
(--spec-type mtp / draft-mtp). Run `unsloth studio update` to refresh
it. MTP GGUFs will load without speculative decoding.

Both structlog.warning(...) (structured logs) and print(..., flush=True) to stderr so it's visible in the terminal where the user launched unsloth studio.

2. Load-time fallback in load_model's spec block:

When the user loads an MTP GGUF and the probe says no, we log a warning and skip the auto-emit instead of letting llama-server fail with an unknown-flag error. The model still loads, just without speculative decoding.

3. /api/inference/status now returns llama_cpp_supports_mtp: bool. Frontend can render a persistent banner / dismissible popup pointing at unsloth studio update. Default True so existing clients that don't read the field aren't affected.

What this means for users

Scenario Before After
Updated llama.cpp + MTP GGUF MTP works MTP works (unchanged)
Outdated llama.cpp + MTP GGUF llama-server errors out on unknown flag warning logged, MTP GGUF loads without spec decoding, frontend can banner
Outdated llama.cpp + non-MTP GGUF works, no warning works, no warning (unchanged)
Outdated llama.cpp launched silent startup log + stderr line pointing at unsloth studio update

Test plan

  • 6 new probe cases in test_llama_cpp_mtp_detection.py:
    • draft-mtp detection (older naming)
    • mtp detection (renamed)
    • pre-MTP build (only ngram variants in --spec-type)
    • missing binary path
    • cache hits on identical (path, mtime)
    • cache invalidates when mtime changes (covers unsloth studio update)
  • Existing 38 MTP detection cases still pass.
  • Broader 188-test regression suite (test_llama_server_args, test_gguf_reload_inheritance, test_gguf_metadata, test_llama_cpp_load_progress, test_llama_cpp_context_fit, test_inference_model_validation) still green.

Notes

  • The probe is ~10ms for the subprocess call on first invocation, instant on subsequent calls (cached). Runs once at startup and once per first /status call after a binary update.
  • Probe failures (binary missing, subprocess error, timeout) are non-fatal: /status returns llama_cpp_supports_mtp: True (conservative default that hides the banner rather than showing a false-positive warning).
  • This PR depends on Studio: auto-enable MTP speculative decoding for MTP GGUFs #5527 (auto-MTP). Merge order matters: this branch is based on mtp-auto-spec-decoding.

References

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 9b4b734f44

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

"llama.cpp prebuilt. Continuing without "
"MTP speculative decoding."
)
self._speculative_type = None

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve the loaded-state marker after MTP fallback

When an MTP GGUF is loaded with an outdated llama-server, this fallback starts the model successfully but records self._speculative_type as None. The duplicate-load checks still normalize the same MTP request to draft-mtp (_already_in_target_state auto-promotes MTP models, and the route-level settings check compares the requested spec against the backend), so every subsequent identical /load is treated as a mismatch and needlessly tears down/restarts llama-server instead of returning already_loaded. This only shows up on the exact environment this change targets: MTP GGUFs with a prebuilt that lacks MTP support.

Useful? React with 👍 / 👎.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a capability probe for llama-server to detect MTP speculative decoding support, enabling version-specific token usage and user warnings for outdated binaries. Feedback suggests improving the subprocess call for Windows compatibility, making the help-text regex more robust, and refining error handling to distinguish between probe failures and confirmed lack of support to avoid false-positive warnings.

Comment on lines +914 to +920
result = subprocess.run(
[bin_path, "--help"],
capture_output = True,
text = True,
timeout = 10,
check = False,
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The subprocess.run call for the capability probe should include the environment and hidden window flags for consistency with other subprocess calls in this backend. This ensures that the probe can find its dependencies (like CUDA libraries) and doesn't cause a console window to flash on Windows.

Suggested change
result = subprocess.run(
[bin_path, "--help"],
capture_output = True,
text = True,
timeout = 10,
check = False,
)
result = subprocess.run(
[bin_path, "--help"],
capture_output = True,
text = True,
timeout = 10,
check = False,
env = child_env_without_native_path_secret(),
**_windows_hidden_subprocess_kwargs(),
)

break
if "draft-mtp" in spec_line:
mtp_token = "draft-mtp"
elif re.search(r"[|,\[]mtp[|,\]]", spec_line):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The regex for detecting mtp support should be more robust to handle different help output formats, such as those using curly braces {} or spaces around separators (e.g., --spec-type {none, mtp}).

Suggested change
elif re.search(r"[|,\[]mtp[|,\]]", spec_line):
elif re.search(r"(?:^|[|,\[\s{])mtp(?:$|[|,\]\s}])", spec_line):

Comment on lines +912 to +940
mtp_token: Optional[str] = None
try:
result = subprocess.run(
[bin_path, "--help"],
capture_output = True,
text = True,
timeout = 10,
check = False,
)
help_text = (result.stdout or "") + "\n" + (result.stderr or "")
# PR #22673 names the spec type ``draft-mtp``; later
# upstream commits rename it to ``mtp``. Recognise both.
spec_line = ""
for line in help_text.splitlines():
if "--spec-type" in line:
spec_line = line
break
if "draft-mtp" in spec_line:
mtp_token = "draft-mtp"
elif re.search(r"[|,\[]mtp[|,\]]", spec_line):
mtp_token = "mtp"
except (OSError, subprocess.SubprocessError) as exc:
logger.debug(f"llama-server --help probe failed: {exc}")

info = {
"found": True,
"mtp_token": mtp_token,
"supports_mtp": mtp_token is not None,
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

When the capability probe fails (e.g., due to a subprocess error or timeout), it is better to report supports_mtp as None (unknown) rather than False. This allows callers to distinguish between "probed and confirmed missing" and "probe failed", enabling a more conservative approach to showing warnings in the UI. Additionally, ensure variables are initialized within the try/except blocks rather than before the try block to avoid unnecessary code.

        try:
            result = subprocess.run(
                [bin_path, "--help"],
                capture_output = True,
                text = True,
                timeout = 10,
                check = False,
                env = child_env_without_native_path_secret(),
                **_windows_hidden_subprocess_kwargs(),
            )
            help_text = (result.stdout or "") + "\n" + (result.stderr or "")
            # PR #22673 names the spec type ``draft-mtp``; later
            # upstream commits rename it to ``mtp``. Recognise both.
            spec_line = ""
            for line in help_text.splitlines():
                if "--spec-type" in line:
                    spec_line = line
                    break
            if "draft-mtp" in spec_line:
                mtp_token = "draft-mtp"
            elif re.search(r"(?:^|[|,\[\s{])mtp(?:$|[|,\]\s}])", spec_line):
                mtp_token = "mtp"
            else:
                mtp_token = None
            supports_mtp = mtp_token is not None
        except (OSError, subprocess.SubprocessError) as exc:
            logger.debug(f"llama-server --help probe failed: {exc}")
            mtp_token = None
            supports_mtp = None

        info = {
            "found": True,
            "mtp_token": mtp_token,
            "supports_mtp": supports_mtp,
        }
References
  1. Avoid initializing variables to an empty value before a try block if the variable is guaranteed to be bound on all execution paths.

Comment thread studio/backend/main.py

_caps = LlamaCppBackend.probe_server_capabilities()
app.state.llama_cpp_capabilities = _caps
if _caps.get("found") and not _caps.get("supports_mtp"):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Update the startup check to only warn if MTP support is explicitly confirmed as missing (False), avoiding false positive warnings if the probe failed (None).

Suggested change
if _caps.get("found") and not _caps.get("supports_mtp"):
if _caps.get("found") and _caps.get("supports_mtp") is False:

# `unsloth studio update` is needed for MTP support.
try:
_caps = type(llama_backend).probe_server_capabilities()
_supports_mtp = bool(_caps.get("supports_mtp", False))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The logic for determining MTP support should handle the case where the probe failed (None) by defaulting to True. This aligns with the "conservative" approach mentioned in the comment, hiding the banner if we aren't certain that support is missing.

Suggested change
_supports_mtp = bool(_caps.get("supports_mtp", False))
_supports_mtp = _caps.get("supports_mtp")
if _supports_mtp is None:
_supports_mtp = True # be conservative: hide banner on probe failure

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 2bf0a23256

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

# Probe binary; fail gracefully on outdated prebuilts.
# Use whichever token the binary advertises
# (older: draft-mtp; renamed upstream: mtp).
caps = self.probe_server_capabilities(binary)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Accept explicit mtp speculative requests

When an API client explicitly sends speculative_type="mtp" for a binary whose probe advertises mtp, normalized_spec is "mtp", so this draft-mtp branch is skipped and the later _valid_spec_types set also rejects it, leaving speculative decoding disabled. The new token adaptation therefore only works for auto-promoted default/null requests, not for explicit mtp requests that use the token this probe now detects.

Useful? React with 👍 / 👎.

Base automatically changed from mtp-auto-spec-decoding to main May 18, 2026 07:15
Layered on #5527. Adds a one-shot llama-server --help capability probe
so users get a clear signal when their prebuilt is missing MTP support,
plus a graceful fallback if they load an MTP GGUF against an outdated
binary.

What's surfaced:

1. Startup log + stderr line in main.py:lifespan() if MTP isn't
   advertised:
     WARNING: llama.cpp prebuilt is missing MTP support
     (--spec-type mtp / draft-mtp). Run `unsloth studio update` to
     refresh it. MTP GGUFs will load without speculative decoding.
2. Load-time graceful fallback in load_model's spec block: skip the
   auto-emit and log a clear warning instead of letting llama-server
   fail with an unknown-flag error.
3. /api/inference/status now returns llama_cpp_supports_mtp: bool so
   the frontend can show a banner / popup.

Probe internals:

- Class-level cache keyed on (binary_path, mtime). One subprocess call
  the first time, instant thereafter. Touching the binary (e.g. via
  `unsloth studio update`) invalidates the cache automatically because
  the mtime changes, so the new build is picked up without restarting
  the server.
- Recognises both upstream naming forms: the original draft-mtp from
  llama.cpp PR #22673 and the renamed mtp variant in later commits.
- Spec block uses whichever token the binary accepts so we emit the
  right value regardless of which release the user has.

Tests:

- 6 new cases in test_llama_cpp_mtp_detection.py covering each probe
  variant (draft-mtp, renamed mtp, pre-MTP build, missing binary,
  mtime-based cache invalidation).
- Existing 38 MTP detection cases still pass; broader 188-test
  regression suite (server args, reload inheritance, gguf metadata,
  load progress, context fit, model validation) still green.
@danielhanchen
danielhanchen force-pushed the mtp-llama-cpp-update-warning branch from 2bf0a23 to 9b74d14 Compare May 18, 2026 07:18
@danielhanchen
danielhanchen merged commit fc04809 into main May 18, 2026
31 checks passed
@danielhanchen
danielhanchen deleted the mtp-llama-cpp-update-warning branch May 18, 2026 07:19

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 9ed7a4836c

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

# MTP capability probe (cached). Drives the UI update banner.
try:
_caps = type(llama_backend).probe_server_capabilities()
_supports_mtp = bool(_caps.get("supports_mtp", False))

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Treat missing llama-server probe as unknown, not unsupported

get_status currently maps probe output to bool(_caps.get("supports_mtp", False)), so when probe_server_capabilities cannot find a binary and returns {found: False, supports_mtp: False}, the API reports llama_cpp_supports_mtp=false. This causes a false “update llama.cpp” signal in environments where llama-server is simply absent/uninitialized (not necessarily outdated), despite the rest of the code trying to fail-open on probe failures. Gate this on found (or default to True when found is false) to avoid incorrect banner behavior.

Useful? React with 👍 / 👎.

pull Bot pushed a commit to ShinnChow/unsloth that referenced this pull request May 18, 2026
…thai#5529)

* Studio: warn when llama.cpp prebuilt is at least 3 days behind

Layered on unslothai#5528. Generalises the MTP-specific staleness warning to
every llama.cpp prebuilt update, not just the ones that add MTP. If
the installed prebuilt is at least 3 days old AND its tag differs
from the latest published tag on the helper release repo (default
unslothai/llama.cpp), Studio nudges the user to run
"unsloth studio update".

How it works

Reads the install marker UNSLOTH_PREBUILT_INFO.json that
install_llama_prebuilt.py already writes to install_dir. The marker
carries the installed tag, the helper repo, and an installed_at_utc
timestamp. Studio compares those against the latest published tag
from the GitHub releases API for the helper repo.

GitHub fetch is cached at two levels:
- Process-level memo for /status hot path.
- Disk-level cache (24h TTL) at ~/.unsloth/studio/cache/llama_cpp_freshness/
  so cold-start Studio launches do not always hit the API.

On a transient fetch failure (offline, rate-limited) we keep the
last-good disk value alive rather than poisoning the cache with None.
The check fails open: if anything is missing (marker, timestamp,
GitHub response), stale stays False so users never see a misleading
banner.

Surfaced in two places

1. Startup banner (logs + stderr) in main.py:lifespan(), alongside the
   MTP capability probe added in unslothai#5528. Single line, e.g.:
     WARNING: llama.cpp prebuilt is 5 days behind: installed b9190,
     latest b9300. Run "unsloth studio update" to refresh.

2. /api/inference/status now returns:
     llama_cpp_prebuilt_stale: bool
     llama_cpp_installed_tag:  str | None
     llama_cpp_latest_tag:     str | None
   so the frontend can render a banner / popup with the actual tag
   delta the user is missing.

3-day threshold

Mirrors the typical Unsloth llama.cpp release cadence. Anything
shorter would nag users who restart Studio at the wrong moment;
longer leaves real bugs sitting on the user's machine. Configurable
via the threshold_days kwarg if a future call site wants a different
window.

Tests

17 new cases in tests/test_llama_cpp_freshness.py cover marker
discovery in both cmake and root install layouts, missing / invalid
marker, GitHub fetch caching across process restarts (disk cache hit
after the in-memory cache is reset), the stale / not-stale decision
matrix (tag mismatch + age threshold), fail-open behaviour when
GitHub is unreachable, custom threshold, singular/plural day in the
warning string, and unparseable installed_at_utc. The broader
205-test inference regression suite still passes.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant