Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
69 commits
Select commit Hold shift + click to select a range
b965a28
Studio: say which model is missing instead of "No model loaded"
danielhanchen Jul 26, 2026
fed31af
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Jul 26, 2026
88f1e5e
Studio: page the API monitor, show model load/unload, pin the example…
danielhanchen Jul 26, 2026
41b3e3a
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Jul 26, 2026
d80cb3a
Studio: optionally download a model named in an OpenAI API request
danielhanchen Jul 26, 2026
d61ca7b
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Jul 26, 2026
658f106
Studio: add an Unload button to the API monitor
danielhanchen Jul 26, 2026
0104baa
Studio: keep the API monitor Unload button visible when idle
danielhanchen Jul 26, 2026
2f7bb58
Studio: never answer a named model with a different one
danielhanchen Jul 26, 2026
a4aa2cf
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Jul 26, 2026
3f51a73
Studio: use a simpler prompt in the API usage examples
danielhanchen Jul 26, 2026
723528d
Studio: only refuse a model reference meant for this server
danielhanchen Jul 26, 2026
d76f56d
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Jul 26, 2026
f913d33
Studio: scope the auto-download 404 cache to the caller's credentials
danielhanchen Jul 26, 2026
014e1fb
Studio: tighten the comments added by this branch
danielhanchen Jul 26, 2026
21cc4fc
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Jul 26, 2026
9eb5a7a
Studio: keep API auto-download off the server's Hugging Face identity
danielhanchen Jul 26, 2026
68723b4
Studio: stop treating a namespace as what decides model intent
danielhanchen Jul 26, 2026
86f772d
Studio: tighten the comments added since the last pass
danielhanchen Jul 26, 2026
00f1c93
Studio: match a resident model through its resolver alias
danielhanchen Jul 26, 2026
f180d58
Studio: shorten the comments added in the last pass
danielhanchen Jul 26, 2026
28b7c56
Merge remote-tracking branch 'origin/main' into r7454
danielhanchen Jul 26, 2026
300881e
Studio: keep the FLA fast-path tests hermetic across transformers ver…
danielhanchen Jul 27, 2026
46232e4
Studio: keep the /v1 admission check off the model-scanning path
danielhanchen Jul 27, 2026
738c695
Merge remote-tracking branch 'origin/main' into r7454
danielhanchen Jul 27, 2026
a551910
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Jul 27, 2026
27fdb9c
Merge remote-tracking branch 'origin/main' into r7454
danielhanchen Jul 27, 2026
7db4683
Merge branch 'fix/openai-model-not-found-error' of https://github.com…
danielhanchen Jul 27, 2026
aba8858
Studio: fix the admission hook's cold, stale and contended index paths
danielhanchen Jul 27, 2026
345bbc0
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Jul 27, 2026
5791f7b
Studio: make /v1/models and the admission hook agree on what is local
danielhanchen Jul 27, 2026
76dad06
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Jul 27, 2026
d7d113a
Studio: tighten the comments this branch adds
danielhanchen Jul 27, 2026
26b2d4a
Merge branch 'fix/openai-model-not-found-error' of https://github.com…
danielhanchen Jul 27, 2026
03e4364
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Jul 27, 2026
7be429b
Studio: four admission and catalog fixes from review
danielhanchen Jul 27, 2026
d2954e6
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Jul 27, 2026
7a96e57
Studio: make the Hub error fixture carry a status on both hub majors
danielhanchen Jul 27, 2026
03ac326
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Jul 27, 2026
da452bf
Studio: invalidate on every download, resolve bare tags, keep polling
danielhanchen Jul 27, 2026
b629605
Merge branch 'fix/openai-model-not-found-error' of https://github.com…
danielhanchen Jul 27, 2026
2b7be7e
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Jul 27, 2026
7c4918d
Studio: hold the download slot while it is in use, and keep quants to…
danielhanchen Jul 27, 2026
a901e7d
Merge branch 'fix/openai-model-not-found-error' of https://github.com…
danielhanchen Jul 27, 2026
3fbd462
Studio: keep what the resolver already knew when a download lands
danielhanchen Jul 27, 2026
d8b001f
Studio: match the quant, not just the directory, and default-select b…
danielhanchen Jul 27, 2026
9616211
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Jul 27, 2026
a260c96
Studio: one quant preference, and stop trusting a stale checkpoint
danielhanchen Jul 27, 2026
c84d0bb
Merge branch 'fix/openai-model-not-found-error' of https://github.com…
danielhanchen Jul 27, 2026
667edcc
Studio: fix the Windows path compare, and advertise a label the worke…
danielhanchen Jul 27, 2026
4e09d9d
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Jul 27, 2026
c5a06cb
Studio: a stored checkpoint needs catalog evidence, not just the swit…
danielhanchen Jul 27, 2026
2a4c67b
Merge branch 'fix/openai-model-not-found-error' of https://github.com…
danielhanchen Jul 27, 2026
278d459
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Jul 27, 2026
a59b2ab
Studio: normalize the quote style pre-commit would have rewritten
danielhanchen Jul 27, 2026
907ecd5
Merge branch 'fix/openai-model-not-found-error' of https://github.com…
danielhanchen Jul 27, 2026
20ed0b0
Studio: cover the model that just landed, and pin the quant the catal…
danielhanchen Jul 27, 2026
19c572c
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Jul 27, 2026
68199da
Studio: apply three rules everywhere they belong, not only where repo…
danielhanchen Jul 27, 2026
3ddc1b6
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Jul 27, 2026
ff58fdd
Merge remote-tracking branch 'origin/main' into fix/openai-model-not-…
danielhanchen Jul 27, 2026
b9e8fa3
Merge branch 'fix/openai-model-not-found-error' of https://github.com…
danielhanchen Jul 27, 2026
ce50721
Studio: probe before refusing busy, and scan once when the index is cold
danielhanchen Jul 27, 2026
e32b486
[pre-commit.ci] auto fixes from pre-commit.com hooks
pre-commit-ci[bot] Jul 27, 2026
a52f028
Studio: an unfinished scan is not absence, and a decided refusal is n…
danielhanchen Jul 27, 2026
091806d
Studio: keep the asyncio.timeout fallback tests runnable on Python 3.10
danielhanchen Jul 27, 2026
c24ed12
Studio: decide GGUF residency, servability and variant keys by one ru…
danielhanchen Jul 27, 2026
0684926
Merge remote-tracking branch 'origin/main' into fix/openai-model-not-…
danielhanchen Jul 27, 2026
1d82c3d
Studio: bound the Hub admission probes and stop guessing at nested mo…
danielhanchen Jul 27, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
18 changes: 17 additions & 1 deletion studio/backend/auth/authentication.py
Original file line number Diff line number Diff line change
Expand Up @@ -164,6 +164,22 @@ async def get_current_subject_allow_password_change(
)


# The literal the examples ship with; pasting one unedited is likelier than a revoked key.
API_KEY_PLACEHOLDER = f"{API_KEY_PREFIX}YOUR_KEY"


def _invalid_api_key_detail(token: str) -> str:
"""Why the key failed. Only the unedited example placeholder is called out;
every real key still gets one indistinguishable message, so this reveals
nothing about which keys exist."""
if token == API_KEY_PLACEHOLDER:
return (
"This is the placeholder key from the example. Create an API key in "
f"Unsloth Studio under Settings > API and use it in place of {API_KEY_PLACEHOLDER}."
)
return "Invalid or expired API key"


async def _get_current_subject(
credentials: HTTPAuthorizationCredentials, *, allow_password_change: bool
) -> str:
Expand All @@ -176,7 +192,7 @@ async def _get_current_subject(
if username is None:
raise HTTPException(
status_code = status.HTTP_401_UNAUTHORIZED,
detail = "Invalid or expired API key",
detail = _invalid_api_key_detail(token),
)
return username

Expand Down
127 changes: 115 additions & 12 deletions studio/backend/core/inference/api_monitor.py
Original file line number Diff line number Diff line change
Expand Up @@ -52,6 +52,13 @@ class ApiMonitorEntry:
total_tokens: Optional[int] = None
total_tokens_authoritative: bool = False
error: Optional[str] = None
# "request" (HTTP call) or "lifecycle" (model load/unload: event/reason, not a prompt; shared).
kind: str = "request"
event: Optional[str] = None
reason: Optional[str] = None
shared: bool = False
# 0-100 for a running download row; None when not applicable.
progress: Optional[float] = None

def snapshot(self, *, include_details: bool = True) -> dict[str, Any]:
duration_ms = None
Expand Down Expand Up @@ -85,6 +92,10 @@ def snapshot(self, *, include_details: bool = True) -> dict[str, Any]:
"completion_tokens": self.completion_tokens,
"total_tokens": self.total_tokens,
"error": self.error,
"kind": self.kind,
"event": self.event,
"reason": self.reason,
"progress": self.progress,
}
if include_details:
payload["prompt"] = self.prompt
Expand Down Expand Up @@ -127,6 +138,75 @@ def start(
self._trim_terminal_locked()
return entry.id

def record_lifecycle(
self,
*,
event: str,
model: str,
reason: Optional[str] = None,
running: bool = False,
) -> str:
"""Record a model load/unload alongside the request traffic that caused it.

``running=True`` opens the row (a load in progress) and the caller closes
it with the usual :meth:`finish` / :meth:`fail`; an unload is terminal on
arrival. Rows are shared, so every subject sees them, and share the same
retention budget as requests.
"""
now = time.time()
entry = ApiMonitorEntry(
id = f"apievt_{uuid.uuid4().hex[:12]}",
endpoint = f"model.{event}",
method = "",
model = model or "default",
prompt = "",
status = "running" if running else "completed",
started_at = now,
updated_at = now,
started_monotonic = time.monotonic(),
finished_at = None if running else now,
finished_monotonic = None if running else time.monotonic(),
kind = "lifecycle",
event = event,
reason = reason,
shared = True,
)
with self._lock:
self._entries.appendleft(entry)
self._trim_terminal_locked()
return entry.id

def relabel(self, entry_id: Optional[str], model: str) -> None:
"""Rename an open lifecycle row once the load resolves its real id (the
caller only has the load path up front, which may be an HF snapshot dir)."""
if not entry_id or not model:
return
with self._lock:
entry = self._find_locked(entry_id)
if entry is not None:
entry.model = model
entry.updated_at = time.time()

def set_progress(self, entry_id: Optional[str], progress: Optional[float]) -> None:
"""Update an open download row's percentage (clamped to 0-100)."""
if not entry_id or progress is None:
return
with self._lock:
entry = self._find_locked(entry_id)
if entry is not None and entry.status == "running":
entry.progress = min(100.0, max(0.0, float(progress)))
entry.updated_at = time.time()

def discard(self, entry_id: Optional[str]) -> None:
"""Drop a row that turned out not to be an event (a load that was already
satisfied, so nothing was actually loaded)."""
if not entry_id:
return
with self._lock:
entry = self._find_locked(entry_id)
if entry is not None:
self._entries.remove(entry)

def append_reply(self, entry_id: Optional[str], text: str) -> None:
if not entry_id or not text:
return
Expand Down Expand Up @@ -212,6 +292,19 @@ def finish(
self._entries.appendleft(entry)
self._trim_terminal_locked()

def fail_open(self, entry_id: Optional[str], error: str) -> None:
"""Fail only a still-open row. Unlike :meth:`fail` this never touches an
entry that already finished, so a catch-all in a ``finally`` cannot stamp
an error onto a request that in fact succeeded."""
if not entry_id:
return
with self._lock:
entry = self._find_locked(entry_id)
if entry is None or entry.finished_at is not None:
return
# Same lock as the check, so a finish() cannot land in between.
self._fail_locked(entry, error)

def fail(self, entry_id: Optional[str], error: str) -> None:
if not entry_id:
return
Expand All @@ -224,15 +317,18 @@ def fail(self, entry_id: Optional[str], error: str) -> None:
if error:
entry.error = _trim(error, 1000)
return
now = time.time()
entry.status = "error"
entry.error = _trim(error, 1000)
entry.updated_at = now
entry.finished_at = now
entry.finished_monotonic = time.monotonic()
self._entries.remove(entry)
self._entries.appendleft(entry)
self._trim_terminal_locked()
self._fail_locked(entry, error)

def _fail_locked(self, entry: ApiMonitorEntry, error: str) -> None:
now = time.time()
entry.status = "error"
entry.error = _trim(error, 1000)
entry.updated_at = now
entry.finished_at = now
entry.finished_monotonic = time.monotonic()
self._entries.remove(entry)
self._entries.appendleft(entry)
self._trim_terminal_locked()

def snapshot(
self,
Expand All @@ -244,7 +340,7 @@ def snapshot(
return [
entry.snapshot(include_details = include_details)
for entry in self._entries
if subject is None or entry.subject == subject
if self._visible(entry, subject)
]

def get(
Expand All @@ -257,22 +353,29 @@ def get(
entry = self._find_locked(entry_id)
if entry is None:
return None
if subject is not None and entry.subject != subject:
if not self._visible(entry, subject):
return None
return entry.snapshot(include_details = True)

def active_count(self, *, subject: Optional[str] = None) -> int:
# Lifecycle rows show as "running" while loading but are not in-flight API requests.
with self._lock:
return sum(
1
for entry in self._entries
if entry.status == "running" and (subject is None or entry.subject == subject)
if entry.status == "running"
and entry.kind != "lifecycle"
and (subject is None or entry.subject == subject)
)

def clear(self) -> None:
with self._lock:
self._entries.clear()

@staticmethod
def _visible(entry: ApiMonitorEntry, subject: Optional[str]) -> bool:
return subject is None or entry.subject == subject or entry.shared

def _find_locked(self, entry_id: str) -> Optional[ApiMonitorEntry]:
for entry in self._entries:
if entry.id == entry_id:
Expand Down
18 changes: 18 additions & 0 deletions studio/backend/core/inference/llama_keepwarm.py
Original file line number Diff line number Diff line change
Expand Up @@ -345,6 +345,22 @@ def _loaded_identity(backend):
return (backend.model_identifier, getattr(backend, "hf_variant", None), advertised)


def _note_idle_unload_event(freed) -> None:
"""Record an idle auto-unload in the API monitor, using the advertised repo id
from the stash so the row never shows the on-disk load path. Best-effort."""
try:
from core.inference.api_monitor import api_monitor
from core.inference.model_ids import public_model_id

identifier, variant, advertised = (list(freed) + [None, None, None])[:3]
label = public_model_id(advertised or identifier) or "model"
if variant and ":" not in label:
label = f"{label}:{variant}"
api_monitor.record_lifecycle(event = "unload", model = label, reason = "idle")
except Exception as exc:
logger.debug("idle unload monitor event failed: %s", exc)


async def idle_unload_loop(poll_seconds: float = 15.0) -> None:
"""Unload the loaded GGUF once idle past the configured TTL. Inert when off."""
from utils.openai_auto_switch_settings import (
Expand Down Expand Up @@ -407,6 +423,8 @@ async def idle_unload_loop(poll_seconds: float = 15.0) -> None:
elif manifest:
_delete_resume_files(manifest)
logger.info("Idle auto-unload: freed GGUF after %ss idle", ttl)
# An idle unload stashes for reload and skips note_model_unloaded.
_note_idle_unload_event(freed)
seen_model = None
except Exception as exc:
logger.debug("idle_unload_loop iteration failed: %s", exc)
Loading
Loading