[staging CI] unslothai/unsloth#6979 - #82
Closed
danielhanchen wants to merge 711 commits into
Closed
danielhanchen wants to merge 711 commits into
danielhanchen wants to merge 711 commits into
Conversation
…#6541) * Route dense NemotronH models to the transformers 5.10 tier Dense NemotronH models (e.g. unsloth/NVIDIA-Nemotron-3-Nano-4B) describe their layer stack with a hybrid_override_pattern that includes '-' (MLP) layers. transformers only learned to parse that ('-' -> 'mlp' in pattern_mapping, 'mlp' in valid_types and MIXER_TYPES) in 5.10; on 5.3/5.5 the config raises KeyError: '-'. The model also ships auto_map remote code, so training and inference that approve trust_remote_code load fine, but a native (TRC=False) load such as export hits the built-in parser and fails with 'Failed to load checkpoint: -'. Detect dense NemotronH from config.json (a '-' in hybrid_override_pattern, or 'mlp' in an expanded layers_block_type) and route it to the 5.10 tier, where the model loads natively without remote code. Pure-MoE NemotronH configs are unaffected and keep their existing tier. Covers both the local config.json and the remote HF-id paths, and adds tests for the detector and the resulting tier selection. * Tighten _nemotron_h_needs_mlp_support docstring * Detect dense NemotronH in nested, cached, and resolved-away configs Three gaps could still route a dense NemotronH (MLP '-' layers) to a tier below 5.10 and hit KeyError: '-': - VL wrappers (e.g. NemotronH_Nano_VL_V2) keep the dense language model under llm_config/text_config; the detector only checked the top-level model_type. Recurse into nested language configs. - Offline or blocked config fetches returned None for an already-downloaded repo. Read config.json from the HF hub cache before any network. - A local checkpoint resolves to its base before tiering, so an offline/private base discarded the local config that revealed the dense pattern. Prefer the higher tier of the resolved base and the original path. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Harden NemotronH tier detection follow-ups Address review of the nested/cached/resolved-away detection: - The local re-check ran the full tier detector on the original path, so a bare LoRA adapter under e.g. /runs/gemma-4-x/llama-lora could upgrade a default base via directory-name substrings. Gate the re-check on a real local config.json so it reads metadata, not path names. - The HF hub cache was read before any network, so an online tier check could serve stale config.json after the repo changed upstream. Consult the cache only offline or after a failed fetch. - Reading the cache imported huggingface_hub during tier detection, which runs before a sidecar venv is activated and could pin the default-env hub into sys.modules. Resolve the cache path with stdlib only. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Trim comments to be more succinct * Select newest hub-cache snapshot by mtime and retry transient config fetches The HF cache fallback in tier detection picked the lexicographically-first snapshot when refs/main was absent (commit-pinned downloads), which can be an older SHA than the Hub would load. Sort snapshots by mtime instead. A transient online fetch failure cached the hub-cache fallback under the normal (model_name, token) key, so a long-lived worker kept serving stale metadata even after connectivity recovered. Return the fallback without memoizing it so the next call retries the network. * Harden config.json tier detection against auth failures and transient blips - _load_config_json: a 401/403/404 from the raw Hub request is a definitive access answer, not an outage. Return None instead of falling back to the HF hub cache, so an unauthenticated or wrong-token request can never read another caller's cached private metadata. - _check_config_needs_510/550: only memoize the derived tier when the underlying config read was definitive (local file, offline cache, or a completed fetch). A transient fetch fallback is no longer pinned, so the tier is re-evaluated once connectivity returns instead of staying stuck on the lower tier. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Tighten comments in tier-detection auth/cache paths --------- Co-authored-by: Daniel Han <michaelhan2050@gmail.com> Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
…ds (unslothai#6535) * Auto-install SSM kernels (causal-conv1d, mamba-ssm) for inference loads Mamba/SSM hybrids (Nemotron-H/Nano, Falcon-H1, Granite-4.0-H, ...) lazily import mamba_ssm / causal_conv1d during from_pretrained, so loading them for chat failed with 'mamba-ssm is required by the Mamba model but cannot be imported'. The training worker already wheel-first installs these before a fine-tune; the inference worker did not. Add utils/ssm_runtime.ensure_ssm_runtime and call it from the inference load path so the same models load for inference. Training worker is untouched; a drift test keeps the shared detection and pinned versions in lockstep. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * ssm_runtime: invalidate import caches, skip MLX, cover LoRA base - Invalidate importlib finder caches in _is_importable and after a successful wheel install, so a kernel installed earlier in this same process is actually importable when the modeling code lazy-imports it during from_pretrained. - Skip the SSM kernel install entirely on the MLX (Apple Silicon) load path: these are CUDA/ROCm Torch kernels with no MLX use and no macOS prebuilt wheel, so the source build would fail before the MLX backend loads the model. - For LoRA loads, also run detection over the resolved base model, since an adapter id like 'me/my-lora' won't match the SSM heuristics but its SSM base (Nemotron-H, ...) is what needs the kernels. Adds tests for cache invalidation and the MLX-skip / LoRA-base worker wiring. * Tighten SSM autoinstall comments * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * ssm_runtime: verify wheel imports, HIP-aware source build, build heartbeat Address review feedback: - Verify a prebuilt wheel actually imports before trusting it; a CUDA/ABI-mismatched wheel now falls back to a source build instead of returning success and failing later with the cryptic lazy-import error. - HIP-aware source build: require hipcc on ROCm, inject clang --gcc-install-dir, and use the 1800s timeout, mirroring the training worker (ROCm has no prebuilt wheel). - Emit a status heartbeat every 60s during the source build so a long (ROCm) build does not trip the orchestrator's 300s inactivity timeout. Tests cover the wheel-not-importable fallback and the missing-hipcc ROCm bail. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Make causal-conv1d best-effort and harden the SSM source build - causal-conv1d is a fast path: models that merely want it (Qwen3-Next, LFM2) fall back to torch, so a failed install must not reject an otherwise loadable chat model on Windows/CPU/macOS or an ABI without a wheel. Only a true SSM model's mamba-ssm requirement stays fatal, matching the training worker which treats causal-conv1d as best-effort. - The source build is reached only when not importable, including a wheel that installed but failed to import; add --reinstall/--force-reinstall so it replaces the broken install instead of no-opping as already satisfied. - Add --no-cache to the ROCm uv source build to avoid reusing stale artifacts from a partial HIP build, mirroring the training worker. * Address review: install SSM kernels before transformers, harden import + Windows Codex: - Install the SSM kernels before importing transformers. run_inference_process imported core.inference.inference (which imports unsloth/transformers) before the load, and a sidecar transformers can evaluate its optional-backend gates against the import state; installing causal_conv1d/mamba_ssm afterwards left those gates unsatisfied and a Nemotron/Falcon/Granite load still failed with "mamba-ssm is required". The initial model's kernels are now installed in run_inference_process before the ML import, via a shared _ensure_ssm_kernels helper; _handle_load keeps calling it (idempotent) for a LoRA's base and for later in-process loads. - _is_importable now treats any import failure as "not importable", not only ImportError. An ABI-incompatible native kernel (undefined symbol after a torch/CUDA upgrade) raises OSError/RuntimeError; letting those escape reported ssm_runtime_install_failed instead of falling back to reinstall/source build. - Skip causal-conv1d on Windows (no prebuilt wheel), mirroring the training worker. A causal-conv1d-only model (Qwen3-Next/LFM2) no longer drops a chat load into a multi-minute untimed source build; it uses the torch fallback. mamba-ssm is still attempted for true SSM hybrids. Tests: test_ssm_runtime.py +5 (broken-kernel exceptions read as not-importable; causal-conv1d skipped on win32 while mamba-ssm still installs). 36 passed. * Trim comments to be more succinct * Run security gates before installing SSM kernels The SSM kernel auto-install is name-based (model_is_ssm is a substring match, no config fetch), so a model id merely containing an SSM substring triggered a native-package install (possibly a slow source build) before the malware and remote-code consent gates ran. Extract those gates into _run_security_gates and call it before the kernel install in both the pre-import path of run_inference_process and in _handle_load, so a blocked or nonexistent model is refused before any build. The gates are metadata-only and do not import transformers, so they are safe to run before the pre-import install. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Resolve remote LoRA bases before importing transformers _resolve_base_model only reads a local adapter_config.json, so a remote LoRA adapter whose own id has no SSM substring but whose base is a Nemotron/Falcon/ Granite model had its base discovered only by ModelConfig in _handle_load, after transformers was imported and its optional-backend availability snapshotted, so the SSM kernel install there was too late. Add _remote_lora_base, a metadata-only adapter_config.json fetch (no huggingface_hub / transformers import), and use it in the pre-import path so the base is gated and its kernels pre-installed. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Gate only loaded roots, tier on the resolved base, read offline LoRA cache Three follow-ups to the pre-import resolution: - The security gate reused the SSM target list, which for a local full fine-tune includes the config.json-recorded base. That base is never loaded, so scanning it could falsely block a safe local checkpoint. Gate only the model plus a genuine LoRA base (matching _handle_load's mc.is_lora), separate from the broader SSM-install list. - Tier activation ran on the raw adapter id, so a remote LoRA whose base needs a sidecar transformers version imported the default and failed. Resolve the base once up front and activate on it. - _remote_lora_base bailed on offline before checking the hub cache, missing a cached adapter's base. Read the cached adapter_config.json when offline or when the fetch fails. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Keep the pre-import gate transformers-free; harden remote LoRA resolution The pre-import security gate called security_load_subdirs, which imports model_config and thus transformers, snapshotting optional-backend availability before the SSM kernels are installed and defeating the ordering. Add compute_subdirs to _run_security_gates and pass False in the preflight so it scans from the root only (transformers-free); _handle_load still runs the authoritative gate with full subdir scoping after the import. _remote_lora_base now skips existing local relative paths (is_local_path) so a checkpoint like outputs/run1 is never treated as a Hub repo, and distinguishes a definitive 404 (not a LoRA -> None) from transient/offline failures (read the cache), so a repo that is now a full model no longer resolves a stale cached base. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Probe a real model id for SSM kernels; respect HF_ENDPOINT model_is_ssm is a substring match, so an arbitrary name could false-match and force a mamba-ssm install that fails the load for a non-SSM model: - a LoRA adapter id like user/falcon-h1-lora (the SSM-relevant code is the base's); - a local checkpoint under an SSM-named parent dir, e.g. /runs/falcon-h1/llama-ckpt. Add ssm_probe_identifier, which resolves the base (or a bare local checkpoint's basename) and feed that to ensure_ssm_runtime from both the pre-import path and _handle_load, so detection runs against a real model id, never an adapter id or parent folders. _remote_lora_base now honors HF_ENDPOINT so enterprise/mirror deployments resolve the adapter base instead of always hitting huggingface.co. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Tighten comments in the pre-import SSM gate/install path --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: Daniel Han <michaelhan2050@gmail.com>
The chat and sidebar scroll-fade overlays ended their gradient at the `transparent` keyword, which is transparent black. Safari 27 Beta (Liquid Glass) interpolates an opaque colour to transparent black through a grey midtone, so the fades render as two solid grey bands (top of the chat and above the composer). Fade each gradient to the theme colour at zero alpha instead, so every step keeps the same hue and no grey can appear. Fixes unslothai#6457. Co-authored-by: wasimysaid <wasimysdev@gmail.com>
…i#6475) * Studio: add Export to GGUF button on finished training runs A completed run's 'Current Run' tab greys out, so it was unclear how to export it: GGUF export lives on the separate Export page and there was no link to it from a run. Add an 'Export to GGUF' button to the run progress card (shown for completed/stopped runs) that deep-links to the Export page with that run preselected via a new ?run= search param. The Export page reads the param, selects the run, defaults to GGUF, and picks the run's main checkpoint. No retraining is required to export a finished run. * Studio: fix export deep-link checkpoint preselect and edge cases Address review feedback on the Export to GGUF deep link: - Move the main-checkpoint auto-select effect after the model-change reset effect so it runs last; previously the reset cleared the checkpoint back to null in the same commit, leaving the field empty on a deep link. - Reset the applied-run ref when the ?run= param clears (e.g. navigating to /export via the sidebar) so a later manual reselect of the same run is not treated as a deep link. - Trim trailing slashes before taking the run output-dir basename so a path like /outputs/run/ still yields a name (the button no longer disappears). * Studio: hide Export to GGUF on runs superseded by a resume A stopped run whose output_dir was later reused by a resumed run is marked resumed_later by the backend; its on-disk contents no longer match the older run's metrics. Since the export deep link selects by output-dir basename, showing the button on such a run would export the newer continuation instead of the run being viewed. Carry resumed_later into the view data and hide the button when set. --------- Co-authored-by: danielhanchen <michaelhan2050@gmail.com>
…ai#6551) * Studio: persistent per-user trust_remote_code approval cache The consent gate pins each approval to a content fingerprint (sha256 over every repo .py), but nothing was persisted, so the dialog reappeared on every fresh load of the same unchanged repo. This adds an on-disk, per-user approval cache that lets the gate skip the dialog when the same user reloads the same code, while keeping the safety guarantees intact. Two-tier validation, both must hold or the user is re-prompted: - Commit SHA (cheap, one HfApi.model_info().sha, no download): a match means a byte-identical tree to the approved revision, so the scan/download is skipped. - Content fingerprint (authoritative): used whenever the SHA is unavailable (local path / offline) and always recomputed on a SHA miss. A new or edited .py changes both the SHA and the fingerprint, so it is caught in every mode. Safety: - Keyed per subject; one user's approval never auto-runs code for another. - CRITICAL is never stored or honored (guarded on both write and read), so a hand-edited store cannot smuggle in an auto-approval. - The malware (HF unsafe-file) gate stays unconditional. - Fail-safe: a corrupt store, an unresolvable SHA, or any error degrades to "ask again", never to "auto-approve". UNSLOTH_TRC_APPROVAL_CACHE_DISABLE=1 turns the cache off entirely. New module utils/security/remote_code_approvals.py holds the store (studio_root()/security/remote_code_approvals.json, atomic write, 0600, RLock) plus the SHA resolvers. Recording happens at the single gate chokepoint when the caller supplies the matching fingerprint, so subject is just threaded through inference/training/export (orchestrators, routes, workers). The scan endpoint returns already_approved so the frontend can skip the dialog on a cache hit. Tests: new tests/test_trc_approval_cache.py covers cache miss, SHA-match skip, SHA-moved re-scan, new-file re-consent, CRITICAL never cached (write + forged read), disable flag, subject isolation, combined adapter+base key, corrupt store, and no-subject bypass. Full security suite: 101 passed. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Address review: make the approval cache skip only the prompt, never the scan Codex found that the SHA "no-scan" fast path could run untrusted code without re-consent. Removed it; the gate now always re-scans and the cache only seeds the authoritative fingerprint check, so it can skip the dialog but never the scan. - CRITICAL is hard-blocked on every load (the scan always runs), so a hand-edited store that downgrades a CRITICAL repo's severity can no longer auto-run it (P2: do not trust editable severity for SHA approvals). - The fingerprint covers external auto_map repos, so changed third-party code always re-prompts even when the primary commit SHA is unchanged; there is no longer a SHA path that bypasses the fingerprint (P1: external auto_map repos). - resolve_commit_sha is resolved fresh on every call (no memoization), so a repo whose default branch moves after approval re-prompts instead of reusing a stale cached SHA (P1: revalidate mutable Hub SHAs). The SHA is now only a conservative secondary gate: a fresh resolvable SHA must match the approved revision, else the seed is withheld; a None (local/offline) falls back to the fingerprint. - Approvals record the scanner ruleset version (SCAN_RULES_VERSION); the gate ignores approvals from an older ruleset so reclassified bytes are re-scanned and re-shown instead of silently auto-approved (P2: invalidate on scan-policy change). Tests: test_trc_approval_cache.py rewritten around the prompt-skip semantics (unchanged repo still scans; SHA move / changed code / scanner-version bump / disable flag all re-prompt; forged downgraded severity still blocks CRITICAL). 105 passed with test_consent_gate.py. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Trim comments to be more succinct * Keep run-owner subject out of persisted config; serialize approval writes Threading subject (the run owner's username / API-key id) into the training config meant _sanitize_db_config persisted it into config_json, which training-history GET returns to any authenticated user, leaking who started a run in multi-user installs. Filter subject alongside the token fields; the worker still receives it from the live config. The approval store's RLock only guards one process, but approvals are recorded from separate inference/export/training subprocesses, so concurrent writers could clobber each other on os.replace and drop an approval (re-prompt). Hold a best-effort cross-process file lock around the read-modify-write. * Fail safe on a malformed approval store A store with the right version but a non-dict shape (e.g. a hand-edited "subjects": []) passed _load()'s check, then lookup chained .get() on a list and raised, breaking every remote-code load until the file was removed. Validate that subjects is a dict in _load(), and tolerate a non-dict per-subject entry in lookup/record/forget, so a corrupt store fails safe (re-prompt) instead. * Keep subject out of the MLX W&B run config _run_mlx_training uploads the whole training config to W&B minus a sensitive set that only listed hf_token/wandb_token/s3_config, so the authenticated subject (username / API-key id) was sent to W&B as run config even though DB history already strips it. Add subject to the W&B-sensitive filter, mirroring training._sanitize_db_config. * Tighten the W&B subject-filter comment --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
…nslothai#6560) studio.backend.run.__main__ adds "--secure" argument via argparse.BooleanOptionalAction, which automatically creates negative --no-secure, that is with **NO** prefix, instead of **NOT**.
…train() (unslothai#6511) * Reset torch.compile cache poisoned by a stray forward before trainer.train() A manual forward / forward+backward run under model.train() before trainer.train() (for example a pre-train grad-norm probe like out = model(**batch); out.loss.backward()) silently poisons training when torch.compile is enabled. The stray training-mode pass is the first one in the process, so it compiles and caches the model forward and, via AOTAutograd, its backward graph in a one-off context that does not match the real training loop. When trainer.train() reuses that cached graph the gradients come out NaN/Inf, the loss never moves, and the run looks like it trains but never learns. Observed on gpt-oss-20b (loss frozen at ~4.25, grad_norm NaN from step 1) with both use_gradient_checkpointing="unsloth" and =True. It does not reproduce when the probe runs under torch.no_grad(), nor with UNSLOTH_COMPILE_DISABLE=1, and a single torch._dynamo.reset() before training fully cures it (loss 4.29 -> 0.0002, identical to a run with no probe). Resetting the gradient-checkpointing buffers, zero_grad, empty_cache, or for_training does not help, confirming the corruption lives in the torch._dynamo / torch.compile cache. get_peft_model now attaches a one-shot forward pre-hook that records whether a forward ran before train(). prepare_for_training_mode checks it at the start of train() and, if a pre-train forward was seen and torch.compile is enabled, calls torch._dynamo.reset() (plus a pristine gradient-checkpoint reset and zero_grad) and warns once. On the normal path (no pre-train forward) it is a strict no-op: no dynamo reset, no recompilation, identical loss curve. * Ignore no-grad pre-train probes and detect probes across the wrapper chain A no-grad forward (with torch.no_grad(): model(**batch)) builds no AOTAutograd backward graph, so it cannot poison the compiled training graph. Gate the marker on torch.is_grad_enabled() so such probes no longer trigger a needless dynamo reset, recompile and warning on an otherwise clean run. Also walk the model wrapper chain (PeftModel / DDP / base model) when resetting so a probe that ran on a different wrapper than self.model is still detected, and tear down every detector hook in the chain. Re-installing the detector is now idempotent and only re-registers when a prior hook was already removed. * Walk DDP/FSDP .module when scanning for the pre-train marker The chain walk followed only .model and .base_model, so a probe that fired on the model below a DDP/FSDP wrapper (which exposes it via .module) left the marker undetected and the poisoned compile cache un-reset. Add .module to the walk. * Install pre-train detector on the full-finetuning path too get_peft_model returns early when UNSLOTH_ENABLE_FULL_FINETUNING=1, before the detector was installed, so full-finetuning runs (which still use torch.compile) did not drop a graph cache poisoned by a stray pre-train forward. Install the detector before both full-finetuning early returns (FastLlamaModel and FastBaseModel). The detector is idempotent, so this never stacks duplicate hooks when get_peft_model is also called on a LoRA model. * torch.compile stray-forward reset: tighten comments (no code change) * Wire stray-forward compile-cache reset into SFT path and PEFT pass-through The pre-train forward detector is installed for plain LoRA/vision models in get_peft_model, but only RL trainers ran the reset via prepare_for_training_mode. A grad-enabled probe before SFTTrainer.train() therefore left the poisoned Dynamo cache in place and the detector hook running on every training forward. - trainer.py: wrap SFTTrainer.train to run _unsloth_reset_stray_compile_cache, which both drops the poisoned cache and tears down the detector hook. For UnslothSFTTrainer the later prepare_for_training_mode assignment supersedes it. - llama.py: arm the detector before the 'Already have LoRA adapters' early return so pre-wrapped PEFT models keep the reset capability. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Preserve detector evidence on reinstall + wire reset into plain Trainer path P2 (_utils.py): _unsloth_install_pretrain_detector cleared marker['seen'] before the live-hook early return, so a re-entrant get_peft_model/patch_peft_model after a grad-enabled probe erased the recorded poisoning while leaving the hook installed, and train() then skipped the Dynamo reset. Only reset seen when (re)installing a fresh hook; keep it when a live hook is already recording. P2 (llama.py): the detector is armed for every LoRA model, but only TRL SFT/RL train wrappers consumed it. Inject _unsloth_reset_stray_compile_cache(self) at the start of the generated _fast_inner_training_loop so a bare transformers.Trainer.train() also drops a poisoned cache and tears down the hook. Idempotent with the TRL-wrapper reset. * Make _unsloth_reset_stray_compile_cache an importable module-level helper The reset was only defined inside the RLTrainer_replacement template string, so 'from unsloth.models.rl import _unsloth_reset_stray_compile_cache' raised ImportError (swallowed) on the SFT auto-packing wrapper and the injected plain-Trainer loop - both paths kept the poisoned Dynamo cache and the dangling detector hook. Move the canonical implementation to unsloth.models._utils (next to the detector, exported in __all__). The RL trainer template now imports it (no-op fallback if the import ever fails), and trainer.py / llama.py import it from _utils too, so every training entry point actually runs the reset. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: danielhanchen <michaelhan2050@gmail.com> Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
…nslothai#6396) * Studio: detect transformers 5.3.0 tier from config.json for local checkpoints A local safetensors folder whose config.json did not match the Gemma4 (510/550) architecture signals short-circuited get_transformers_tier() to "default" (transformers 4.57.x), never reaching the name-substring check that routes Qwen3.5 to the 5.3.0 sidecar. So a local Qwen3.5 checkpoint (model_type "qwen3_5", needs transformers >= 5.2.0) loaded with 4.57.x and failed with "does not support Qwen3.5". The same model as a remote HF id worked, because it has no local config.json to trigger the short-circuit. Detect the 5.3.0 tier from config.json (model_type "qwen3_5" / architecture Qwen3_5ForCausalLM) in the local-config branch, mirroring the existing Gemma4 510/550 handling. This is a positive config signal, so it fixes local Qwen3.5 without weakening the directory-name false-positive guard (a llama checkpoint under a "gemma-4-12b-*" parent still resolves to default). Adds tests for the config-based 530 detection and local-folder tier resolution. * Studio: suppress false warning when config.json parse fails for sidecar-tier models * Studio: generalize local-checkpoint tier detection for all 5.3.0 families Expands the config.json-based tier detection to cover all known 5.3.0-tier model families (Qwen3 MoE, GLM-4.7-Flash, LFM2.5-VL) and adds a _name_or_path fallback so renamed local checkpoints with unrecognised model_type values still route correctly via the HF ID embedded in their config.json. - Expand _TRANSFORMERS_530_ARCHITECTURES / _MODEL_TYPES with verified entries from Qwen3MoeForCausalLM, Glm4MoeLiteForCausalLM, Lfm2VlForConditionalGeneration, and Qwen3_5ForConditionalGeneration (confirmed from local Qwen3.5-2B config.json) - Extract _tier_from_name() helper, deduplicating the fast-substring logic used by both the remote-path branch and the new config _name_or_path fallback - In the local-config branch: after architecture checks, resolve the tier from cfg._name_or_path / cfg.model_name before returning "default", preserving the existing directory-name false-positive guard - 79 tests passing * Studio: match 510/550 style for 530 config sets (no inline comments) * Studio: use _resolve_base_model instead of reinlining _name_or_path lookup * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: recurse into get_transformers_tier for resolved base model (Gemini suggestion) * Studio: use _tier_from_name in local-config fallback to avoid network probes Using get_transformers_tier(resolved) on the _name_or_path fallback would trigger up to 3 network fetches (config.json + tokenizer_config.json, 10s each) for every ordinary checkpoint whose _name_or_path is a plain HF ID like meta-llama/Llama-3-8B. The fallback's purpose is name-based detection on the resolved HF ID, _tier_from_name covers all known cases without I/O. * Studio: add _check_config_needs_530 to slow HF-ID fallback path Private or renamed HF repos whose model IDs lack a 5.3 substring were silently routed to the default tier. _check_config_needs_530 mirrors the existing 510/550 pattern: fetches config.json once, caches the result, and is called after the 550 check in the slow path. Includes 5 unit tests. * Studio: guard _tier_from_name fallback against local-path false positives When _name_or_path in config.json is an absolute path to the same checkpoint passed as a relative path, the textual resolved != model_name check passes and _tier_from_name would scan the directory path for substrings. Split the fallback: local directories recurse into get_transformers_tier (config check, no network I/O); HF Hub IDs use _tier_from_name (name-based, no network). * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: separator-norm aliases, model_name/_name_or_path fallback, tests - _norm_separators(): collapse _ . whitespace to - so underscore/dot model ID variants (Qwen3_5, Qwen3_Next) match the canonical substring list - _tier_from_name(): apply norm to both name and each substring so aliases resolve without duplicating the substring lists - _resolve_base_model(): try model_name then _name_or_path separately so a self-referential Unsloth model_name doesn't hide the useful HF ID in _name_or_path - Gate get_base_model_from_lora on adapter_cfg_path.is_file() to avoid eagerly importing transformers before the sidecar venv is on sys.path - 17 new tests covering _norm_separators, separator-insensitive _tier_from_name, and the model_name/_name_or_path fallback * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: only pre-resolve LoRA adapters in activation callers activate_transformers_for_subprocess and ensure_transformers_version were pre-resolving all local checkpoints via _resolve_base_model before calling get_transformers_tier. After the model_name/_name_or_path fix, a full checkpoint with a private/offline _name_or_path and no tier substring would resolve to that HF ID, which can't be probed, bypassing the local config.json model_type check entirely. Gate pre-resolution on adapter_config.json so full checkpoints go straight to get_transformers_tier, which reads config.json directly. LoRA adapters still pre-resolve as before. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: fix Qwen3.5 MoE/Qwen3.6 tier detection and dot-version false positives - Add Qwen3.5 MoE (qwen3_5_moe / Qwen3_5MoeForConditionalGeneration) and Qwen3-Next to the 5.3.0 config sets, so renamed local checkpoints route to the sidecar instead of default transformers - Let a 510/550 name match override a 530 config match, so Qwen3.6 (which reuses qwen3_5 / qwen3_5_moe config ids) still routes to the 5.5.0 sidecar - Stop normalizing version dots to hyphens so size names like Qwen3-5B and Qwen3-6B are not promoted to a 5.x sidecar; underscore aliases still match - Skip name matching for resolved values that look like stale local paths * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: close remaining codex P2s: adapter-only LoRA + 530-override path-hint guard - adapter_model-only LoRA: add import-light _is_lora_adapter_dir/_has_adapter_weights and gate activation/export pre-resolve on them, so LoRA dirs with adapter_model*.safetensors but no adapter_config.json still resolve to their base model (via _resolve_base_model's new unsloth_<model>_<ts> directory-name parse) instead of tiering off the adapter folder. - 530 override: only treat a resolved value as a name hint when it is a real Hub id; a stale/renamed local path in model_name/_name_or_path can no longer flip a correct 530 config to 550. Current folder basename still allowed. Added 7 regression tests; suite at 116 passing. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: address review feedback on tier detection - Add Qwen3.5 text-tower model types (qwen3_5_text / qwen3_5_moe_text) to the 5.3.0 config set so text-only configs with stripped architectures still route to the sidecar - Apply the Qwen3.6 name override on the remote slow path too, so a renamed or private repo whose config reuses qwen3_5 ids but names Qwen3.6 in _name_or_path selects 5.5.0 instead of 5.3.0 - Treat an existing local path (or empty value) as a path, not a Hub id, in _looks_like_hf_id so a real local checkpoint folder is not name matched - Guard _resolve_base_model against non-string config values and compare paths by realpath so relative or absolute self references resolve correctly - Keep the LoRA adapter is_file check inside the OSError guard * Studio: harden tier detection against malformed configs and bad paths - _config_matches_tier no longer raises TypeError when a malformed config.json carries a non-string model_type (e.g. a list) or non-list architectures; it fails open to no-match - guard the model_name-derived is_file/is_dir probes with _safe_is_file / _safe_is_dir so a pathological or over-long path (e.g. a Windows long path) fails open to the default tier instead of raising OSError No routing changes for any valid model; purely defensive. Verified by a cross-platform simulation (POSIX + NT path semantics) and a before/after tier matrix that is unchanged for all previously supported models. * Studio: trim verbose comments in tier detection Shorten/remove over-long comments and docstrings, mainly on internal helpers, without changing behavior. Verified code-only via comment_tools.py check; suite unchanged at 128 passing. --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: Lee Jackson <130007945+Imagineer99@users.noreply.github.com> Co-authored-by: Daniel Han <danielhanchen@gmail.com>
…hai#6484) * Studio: show tool-call progress for large GGUF tool arguments The GGUF agentic tool loop only surfaced an early provisional tool card for render_html, so any other tool (python, terminal, ...) was invisible in the UI while its arguments streamed. For a large argument such as a full HTML or code file this left the chat sitting on "Generating..." with zero progress for tens of seconds while the model was clearly working. Generalize the provisional tool_start to any enabled tool once its streamed arguments grow past a threshold (render_html still surfaces immediately, small-argument tools are unchanged). The provisional and the real tool_start share the tool_call_id so the frontend reconciles them into one card. Close the provisional on no-op, denial, parallel-drop, post-loop, and on stream errors so a card can never spin forever, surface each parallel call, and skip the early card while a human confirmation gate is active. Apply the same confirmation-gate guard to the safetensors agentic loop. Additional hardening: - Only emit a provisional card once a real, non-empty tool_call_id is known. llama.cpp can stream a tool call with an empty id, and a card keyed by "" cannot reconcile with the real tool_start (the frontend mints its own id per event), so it would dangle. - On a connection drop or other mid-iteration failure, close the dangling provisional card with an error result instead of an empty success so the UI renders it as failed rather than completed. - Mirror the provisional cleanup in the safetensors loop: close a provisional render_html card if the model generator raises mid-stream or the controller turns the call into an internal no-op. Adds regression tests for the empty-id guard, the error-result on a dropped connection, and the safetensors mid-stream exception cleanup. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: wasimysaid <wasimysdev@gmail.com>
…nslothai#6568) * Studio: accept --not-secure as a back-compat alias for --no-secure PR unslothai#6560 renamed the negative secure flag from --not-secure to --no-secure to match argparse.BooleanOptionalAction. Re-add --not-secure as a hidden, deprecated alias at both CLI layers so existing scripts and muscle memory keep working, while --no-secure stays the documented spelling. - studio/backend/run.py: extract the CLI parser into _build_arg_parser() so the flag wiring is unit-testable, and register --not-secure as a hidden store_false alias for --no-secure. Last flag wins, matching BooleanOptionalAction semantics. - unsloth_cli/commands/studio.py: add a hidden --not-secure option to `unsloth studio` and `unsloth studio run`; it forces secure off and forwards the canonical --no-secure to the backend. - Tests at both layers for the alias and its polarity. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: address review on --not-secure alias - run.py: use argparse.SUPPRESS for the --not-secure default so the alias never contributes a namespace default (the canonical --secure owns it). - studio.py: resolve --not-secure last-wins from argv via _resolve_secure() so `--not-secure --secure` keeps secure on, matching the backend's BooleanOptionalAction and how --secure/--no-secure already behave. - Add a CLI last-wins test covering both flag orders. --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
…thai#6569) * Add regression tests for the stray-forward compile-cache reset Follow-up to unslothai#6511, which fixed the bug but whose squash merge did not include the tests. These cover the two issues that fix addressed, under the GPU-free tests/conftest.py harness: - _unsloth_reset_stray_compile_cache is an exported module-level symbol in unsloth.models._utils (it previously lived only inside the RL trainer template string, so every non-RL import silently no-op'd) - _unsloth_install_pretrain_detector keeps a recorded "seen" forward on an idempotent reinstall with a live hook, and only resets it after teardown - only a grad-enabled pre-train forward marks the cache poisoned - the reset warns and clears seen when a stray forward was seen, tears the hook down even on the clean path, and walks the .model/.base_model/.module wrapper chain to reach a nested marker * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Pin UNSLOTH_COMPILE_DISABLE in the warn-path reset tests The reset only warns and resets Dynamo when UNSLOTH_COMPILE_DISABLE != "1". A GPU-free CI env that sets it to "1" would make the warn assertion in test_reset_clears_seen_and_warns_when_a_stray_forward_was_seen flaky. monkeypatch it to "0" in both warn-path tests so the warn / no-warn assertions are deterministic and test the seen flag, not the env. --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
…d wait) (unslothai#6566) * Installer: respect a declined Studio auto-start and keep Ctrl+C shutdown logs ordered The `curl | sh` Studio auto-start prompt had two issues on Linux/macOS/WSL (install.sh). install.ps1 already gates on input redirection, so Windows is unaffected. 1. Typing n, or any closed/EOF /dev/tty, still launched Studio. The read fallbacks defaulted to "y" (read failure, and the no-tty branch), so any answer other than a cleanly delivered y/n line auto-started a blocking foreground server. Default those to "n"; a real Enter still counts as yes via ${_reply:-y}. 2. On Ctrl+C the shell prompt printed in the middle of Studio's shutdown logs. The non-interactive installer shell took the default SIGINT action and died before the child finished its graceful shutdown, so the prompt raced ahead of "All subprocesses cleaned up". trap '' INT in the installer shell so it waits for Studio's own graceful shutdown. * Studio: wait for the uvicorn thread before the terminal returns on Ctrl+C Builds on unslothai#6565 by @Imagineer99. The studio server runs uvicorn in a daemon thread, so on Ctrl+C the process could return to the shell while that thread was still writing its shutdown logs, interleaving them with the prompt. Retain the uvicorn thread and join it (flushing stdout/stderr) before terminal entrypoints return, from run.py's main shutdown path and the CLI shutdown paths. Refinements over unslothai#6565: - Bound the join at 5s (_SERVER_SHUTDOWN_JOIN_TIMEOUT, matching the existing _graceful_shutdown subprocess timeouts) so a stalled uvicorn shutdown cannot hang the terminal; the timeout warning branch is now reachable. - Restore SIG_DFL for SIGINT/SIGTERM at the start of the signal handler so a second Ctrl+C force-quits, and drop the redundant in-handler wait (the post-loop wait already covers the signal path). Co-authored-by: Lee Jackson <130007945+Imagineer99@users.noreply.github.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Address review: keep child Ctrl+C working and restore SIGBREAK - install.sh: run studio in a subshell that resets INT to default (trap - INT; exec ...) so the foreground child does not inherit the installer shell's ignored SIGINT, which would otherwise swallow the studio process's own Ctrl+C and graceful shutdown. - run.py: also restore SIGBREAK to SIG_DFL in the signal handler so a second Ctrl+Break force-quits on Windows, matching SIGINT/SIGTERM. * install.sh: capture studio exit with || under set -e so the migration hint still prints * Trim shutdown-fix comments to be terser (comments only, no code change) * Dedup CLI shutdown-wait into finally blocks (review follow-up) --------- Co-authored-by: Lee Jackson <130007945+Imagineer99@users.noreply.github.com> Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
* Add HF dataset streaming mode to Studio * Added default value for datasetStreaming in training-config-store.ts * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Handle None max_steps for streaming validation * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * studio: fast-fail streaming validation and guard incompatible modes Reject dataset_streaming at the API boundary when hf_dataset is empty, the dataset is vision/audio, or max_steps is not set. Probe eval split with get_dataset_split_names before the streaming load so typos fail immediately instead of mid-training. Guard column_names=None after map on iterables. Hide the UI toggle for non-text configurations and clear the stale flag when config becomes incompatible. * studio: add streaming dataset tests, iterable helper, and streaming template/format support (WIP) Work-in-progress on top of feat/studio-dataset-streaming-mode (PR unslothai#4946): - new test_training_streaming.py and iterable.py dataset helper - streaming support in chat_templates.py and format_conversion.py - additional streaming guards in trainer.py / models / routes - frontend streaming wiring in params-section and training-config-store Committed to preserve uncommitted work before merging latest main. * studio: fix review-team findings for streaming + main merge BLOCKER: streaming + raw-text/CPT crashed on len(IterableDataset). Guard it in the start route (reject format_type=="raw" or training_type=="Continued Pretraining") and in isStreamingSupported (datasetFormat !== "raw"). Also: - models/training.py: validate hf_dataset/subset/split (charset+length, block ..//); cap dataset slice indices (le=1e9); note validator ordering - chat_templates.py: guard _apply_custom_mapping .map() for streaming - trainer.py: warn when packing+streaming - training-config-store.ts: persist-migration bump to v11 (standalone datasetStreaming backfill); add isVisionModel to NON_PERSISTED; toast on silent streamingCompatiblePatch mutations in the 4 indirect setters - tests: route rejections (max_steps, raw/cpt), slice cap, unsafe hf_dataset * studio: enable raw-text/CPT dataset streaming + streaming UX polish - raw_text: keep the lazy filter but skip len()-based row counting for IterableDatasets so raw-text / CPT can stream; guard the eval-size log - routes/trainer: drop the raw/CPT streaming block; add a defensive not-streaming guard on the eval auto-split (train_test_split) - dataset-section: streaming toggle is visible-but-disabled and lists the exact unmet requirement(s) in its tooltip; block embedding models - training-start-overlay: show "streaming (no full download)" instead of a stuck download bar for streaming runs - trim the streaming test suite to the high-value cases * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * studio: address streaming review (MLX/embedding guards, sliced eval split, rehydrate timing) - routes: reject dataset_streaming for embedding training and on Apple Silicon (MLX); both loaders materialize the full dataset instead of streaming - trainer: validate the base eval split name so streaming eval accepts HF slice syntax such as "validation[:1000]" - training-config-store: defer the onRehydrateStorage setState to a microtask so it doesn't hit the store's TDZ during synchronous hydration - test: streaming start rejects embedding models * studio: harden HF dataset streaming (column_names, split slicing, empty/eval bounds, gating) Address a deeper streaming review: - raw_text: resolve_column_names() guards IterableDataset.column_names=None (from_generator / unresolved features) so raw-text and CPT streaming no longer raise TypeError before training - models/routes: reject HF slice syntax in train_split/eval_split when streaming (load_dataset(streaming=True) raises "Bad split"); reject mixed sources (local/S3) and embedding/MLX streaming at the API, not just in the UI - trainer: an empty post-slice/filter stream fails preflight with a clear message; streaming eval is capped (STREAMING_EVAL_MAX_SAMPLES) so each eval terminates; the manual-slice shortcut falls back to a regular load when train_split is sliced - format_conversion: streaming conversions preflight the first mapped row so format errors surface before training, not mid-iteration - frontend: block streaming on Apple Silicon; clear datasetStreaming when a dataset is detected as image/audio at start * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * studio: fix CI for streaming PR (lint blocker + no-torch sandbox + preflight test) - trainer.py: drop unused `IterableDataset` import (hoist safety-net blocker). - test_training_streaming.py: only select real classes (isinstance type) when locating the trainer class, so a MagicMock-stubbed global is never passed to object.__new__ (fixes TypeError on the Python 3.10-3.13 jobs). - no-torch import sandboxes (test_e2e_no_torch_sandbox.py, test_studio_import_no_torch.py): teach the chat_templates/format_conversion exec stubs and the full-import-chain copy list about the new `.iterable` module so the AFTER/runtime cases import without torch again. --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: Roland Tannous <115670425+rolandtannous@users.noreply.github.com> Co-authored-by: Roland Tannous <rolandtannous@gravityq.ai> Co-authored-by: Etherll <61019402+Etherll@users.noreply.github.com>
* studio: let users change their password from Settings The only day-to-day way to change credentials was the destructive console command 'unsloth studio reset-password' (it deletes auth.db); the in-app change-password page is the forced first-login flow and bounces non-forced users to /login. Add a Change password control to Settings > General > Account: a small dialog that takes the current and new password and calls the existing POST /api/auth/change-password, then stores the rotated tokens it returns. Username changes remain out of scope. The dialog uses authFetch, so an expired access token is refreshed and the request retried instead of failing with a spurious expired-token error for a user who left Studio open past the token lifetime. The row is hidden in the Tauri desktop app, which authenticates via desktop auto-auth with a generated secret: there is no user-entered password to change there, and changing it would clear the desktop secret. * studio: harden settings password change * studio: harden settings password dialog UX --------- Co-authored-by: wasimysaid <wasimysdev@gmail.com>
…ai#6550) * Resolve the transformers tier by probing AutoConfig instead of guessing When the only signal is a 5.x tokenizer class, get_transformers_tier guessed the lowest 5.x sidecar (530). That misroutes models whose built-in config parser needs a higher tier: dense NemotronH ships a 5.x tokenizer but its '-' (MLP) layer only transformers 5.10 can parse, so 5.3/5.5 raise KeyError '-'. The config.json transformers_version field records the saving version, not the minimum to load, so it cannot drive routing either. Replace the weak tokenizer->530 guesses (local and remote) with a probe: parse config.json with the built-in parser (trust_remote_code=False) in each sidecar, escalating 530->550->510, and pick the first that succeeds. This generalizes to any architecture without hardcoded lists. Strong signals stay fast paths (no subprocess); the probe runs only when the tier is otherwise ambiguous and is cached by (model, commit sha). It never executes repo code, never downloads weights, never raises, and falls back to the legacy 530 guess on a transient/auth/offline failure or when no sidecar is available. UNSLOTH_DISABLE_TIER_PROBE restores the old behavior. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Address review: tier probe fallbacks and cross-platform robustness Codex: - Never escalate to 510 on uncertainty. When every sidecar was probed and none parsed with the built-in parser, the model is a remote-code / custom model_type that loads via its own code; keep the legacy 530 route instead of jumping to 510 (which would change the behavior of models that worked on the 5.3 stack). - Only cache the 530 fallback when the result is conclusive (every tier actually probed). If a sidecar was missing/uninstallable the environment is incomplete, so return 530 uncached and retry on the next call. - Do not pin the tier cache under an unknown revision: _resolve_commit_sha no longer memoizes a None sha (a transient Hub failure is retried), and _probe_tier only caches a tier when the commit sha is known. Gemini: - Wrap Path.exists() in the sha resolver in try/except OSError (a remote repo id can raise WinError 123 on Windows). - Probe script writes the error to sys.stderr.buffer as UTF-8 bytes so a non-ASCII message cannot itself raise UnicodeEncodeError under cp1252. - subprocess.run decodes stderr with errors="replace" to avoid UnicodeDecodeError on non-UTF-8 consoles. Tests: 72 passed (added partial-sidecar uncached, sha-unresolved not cached, all-failed stays 530 + cached, sha resolver retries None / handles OSError). * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Address review round 2: authenticate tier checks, stop memoizing local sigs Codex: - Thread hf_token through _check_config_needs_510/550 and _check_tokenizer_config_needs_v5 (and the underlying raw fetches). Previously a gated/private model whose only 5.x signal is tokenizer_config.json never reached the authenticated probe: the unauthenticated raw fetch failed and cached False, so the model fell through to the default 4.x tier. The per-check caches are now keyed by (model, token) so an unauthenticated miss cannot poison a later authed read, mirroring _load_config_json. - _resolve_commit_sha no longer memoizes a local directory signature. A local signature is mutable (size/mtime of config/tokenizer), so a reused/overwritten checkpoint path would otherwise keep selecting the previous tier; it is now recomputed every call. Only the immutable remote commit sha is memoized. Tests: 75 passed (added token-cache isolation + auth header, local signature not memoized, token threaded into all checks/probe). * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Address review round 3: reach activation with the token, drop SHA tier cache Codex round 3: - Thread hf_token into the activation path that actually selects a sidecar. The token-aware tier checks added last round were unreachable: activate_transformers_for_subprocess called get_transformers_tier without a token, and the inference/training/export workers passed only the model name even though they hold a request-scoped hf_token. activate_transformers_for_subprocess now takes hf_token and the three workers forward config["hf_token"], so a gated/private model whose only 5.x signal is an authenticated config/tokenizer is routed to the right sidecar instead of falling to default 4.x. - Stop importing huggingface_hub during tier detection. _probe_tier no longer resolves a commit sha, so it never pulls huggingface_hub into the worker before the sidecar venv is prepended to sys.path (activation only prepends, never purges), which would otherwise pin the default-env hub over the sidecar's pinned huggingface_hub==1.8.0. - The tier cache is now keyed by model_name for the process lifetime (a model's required tier is a property of its architecture; cleared on restart). This drops the mutable-SHA memo that masked remote revision changes and the mutable local-signature memo, removing _resolve_commit_sha / _local_dir_signature / _probe_sha_cache entirely. - Do not cache a probe success that depended on a skipped lower tier: if a lower sidecar was unavailable, the lowest valid tier may change once it installs, so the result is returned uncached and re-probed next call. Tests: 73 passed (probe imports no hub; success uncached when a lower tier is skipped; activation forwards the token). * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Trim comments to be more succinct * Re-probe overwritten local checkpoints and authenticate the probe child The AutoConfig tier probe cached its result under the bare model_name, so a local checkpoint overwritten in place (same path, new config.json) kept serving the stale sidecar. Fold a cheap config.json signature (size + mtime) into the cache key for local paths; remote ids stay name-keyed so no huggingface_hub import lands before the sidecar is activated. The probe relies on the implicit HF_TOKEN env, so an inherited HF_HUB_DISABLE_IMPLICIT_TOKEN=1 left it unauthenticated and a gated repo 401ed into the 530 fail-safe. Clear that flag in the child env when a token is set. * Keep tier probes off the log-only path and probe new 5.x archs default-first - get_transformers_tier gains probe=True/False. needs_transformers_5 (a coarse 4-vs-5 boolean used only for a spawn log and a vision-check branch) now passes probe=False, so a parent/log-only caller never spawns sidecar probes. The real activation path keeps probe=True and resolves the exact tier in the worker. - A config.json saved by transformers 5.x but matched by no fast path is now probed default-first: _probe_tier gains include_default + floor, prepending the ambient 4.57.x tier to the escalation. A model that still parses on the default is left on it (no mis-route onto a sidecar); only a config the default parser cannot read escalates to the lowest 5.x tier that parses. The transformers_version field is a cheap 'worth probing' hint only, read from the already-fetched config (no extra network); ordinary 4.x configs never probe. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Separate probe cache by mode and keep version-field 5.x visible to needs_transformers_5 - _probe_tier cache was keyed only by config.json signature, so a default-first probe that returned 'default' could be handed back to a later tokenizer/known-5.x caller (floor=530), leaving a model with a 5.x-only tokenizer on transformers 4.x. Key the cache by probe mode (floor + include_default); the legacy 530 mode keeps the bare key. - The version-field 5.x detection is a cheap config read, not a probe, so run it even when probe=False: a standard-tokenizer model whose only signal is transformers_version >= 5 now classifies as 5.x via needs_transformers_5 (returns '530' without spawning a probe), so the vision-routing fallback uses the 5.x subprocess instead of failing the default parser and marking it non-vision. The real activation path still probes default-first and may resolve 'default'. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Don't treat local checkpoints as Hub ids, and fix stale activation test double - _load_config_json / _check_tokenizer_config_needs_v5: a local checkpoint dir whose config.json / tokenizer_config.json is not yet present was being fetched from the Hub as if the path were a repo id, and the 404 miss was cached. A later call after the file is written (in-progress checkpoint) then served the stale miss, so a TokenizersBackend checkpoint fell through to the default tier. Skip the Hub fetch for local dirs and do not cache the miss, so the file is read once it appears. - test_activate_transformers_version_or_warn_*: the worker now threads hf_token into _activate_transformers_version (model_name, hf_token); update the one-arg test doubles to the real two-arg signature so the silent-success path stays silent. * Tighten comments in the AutoConfig probe and tier-selection paths * Address review: canonical probe cache key and reuse _token_cache_key - _probe_cache_key resolves config.json to its absolute realpath before keying, so a relative path or a changed cwd can't collide with or miss a prior probe result. Remote ids still fall back to the name (stat raises, caught). - _cached_config_json reuses _token_cache_key instead of re-hashing the token inline, keeping the (model, token) key derivation in one place. --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
) * Studio: honor custom HF_HOME for model download and load _setup_cache_env always derived HF_HUB_CACHE and HF_XET_CACHE from XDG_CACHE_HOME / ~/.cache, ignoring a user-set HF_HOME. Because it sets HF_HUB_CACHE explicitly and that variable takes precedence over HF_HOME in huggingface_hub, the hub cache was pinned to the standard location: a model already present under a custom HF_HOME was detected but then re-downloaded from scratch on load. Seed HF_HUB_CACHE and HF_XET_CACHE from HF_HOME when the user set it (HF's own default is $HF_HOME/hub and $HF_HOME/xet), and honor the legacy HUGGINGFACE_HUB_CACHE alias. The hub download workers call snapshot_download without a cache_dir for both the Xet and HTTP-fallback paths, so they follow HF_HUB_CACHE; fixing it here unifies detection and both transports on one root. Explicit HF_HUB_CACHE / HF_XET_CACHE stay untouched. Adds tests for the custom-HF_HOME, default, explicit-override, and legacy-alias cases. Fixes unslothai#5182. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: do not crash startup when a custom HF_HOME is not writable Seeding HF_HUB_CACHE/HF_XET_CACHE from HF_HOME means _setup_cache_env now mkdir's under a user-controlled path. A non-writable or not-yet-mounted HF_HOME (typo, offline drive) would raise and crash startup, where the old code silently fell back. Make the mkdir best-effort; the env var is still set, so HF reports a clear error at download time. Adds a regression test. * Studio: strip blank HF_HOME and isolate cache-env tests Address review: a whitespace-only HF_HOME no longer derives " /hub"; strip it and fall back to the default (matches studio_root). Tests set UNSLOTH_STUDIO_HOME to a tmp dir so _setup_cache_env's UV/VLLM mkdirs do not touch the real ~/.unsloth/studio. Adds a whitespace regression test. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
* Studio Playwright: snooze update banner before sending The llama.cpp update banner is a fixed bottom-right toast (z-9998). When an update is available it overlaps the composer's Send button and its subtree intercepts the click, so send_and_wait times out (flaky; surfaces on the Windows studio UI smoke, passes otherwise). Snooze the banner if it is showing before each send, then wait for it to detach. * Also snooze the web update banner before sending The web update banner (web-update-banner, z-9999) is a fixed bottom-right toast like the llama.cpp one and can overlap the Send button too. Loop over both banners and snooze whichever is showing.
* Studio: hide RAG embedder from the On Device list The bge-small-en-v1.5 RAG embedder (and other infra models) were already hidden from Discover but still showed up in the On Device browse list, cluttering the user's downloaded models. They are now filtered out of On Device the same way, while a search that matches still reveals the row so the user can confirm it is already downloaded. * Studio: also check path/title when hiding infra models from On Device isHiddenModelId only saw row.id and row.repoId, but local inventory rows can have a null repoId and an id that is a hash rather than the file path/name, so the llama.cpp validation probe (stories260K.gguf) could slip into the On Device list. Pass the local row's path and title too, mirroring the backend's _is_hidden_model(m.id, m.path). Addresses review feedback from gemini-code-assist on PR unslothai#6572. * Studio: exclude infra models from On Device count and dataset list The On Device hidden-model filter was applied to datasets too, so a dataset whose id/title/path contained an infra needle (bge-small-en-v1.5, stories260k.gguf) was wrongly hidden. Bypass the filter for datasets, the same way Discover and the format filter already do. The On Device header count and the Cache/Local stat pills still used the unfiltered row counts, so a fresh install with only the bge embedder cached read 1 over an empty list. Count visible (non-infra) rows instead, keeping full counts for datasets. * Studio: count search-revealed infra rows in the On Device tally The visible-row counts excluded every hidden row unconditionally, but the On Device list reveals a hidden row when the search query matches it. So with only the bge embedder cached and a "bge" search, the list showed one row while the header and Cache stat stayed 0. Reuse isVisibleInventoryRow for the counts so a query-revealed row is counted, keeping them in step with the list. --------- Co-authored-by: Daniel Han <danielhanchen@gmail.com>
The macOS-arm studio venv still installs anyio 4.14.0 despite the constraints.txt cap from unslothai#6546. mlx-vlm / mlx-lm pull anyio>=4.14, which conflicts with the anyio<4.14.0 constraint; a uv -c constraint loses that conflict so 4.14.0 gets installed, reintroducing the cancel-scope RuntimeError on Python 3.13 (unslothai#6483). UV_OVERRIDE is already applied on macOS-arm via overrides-darwin-arm64.txt and a uv override wins the conflict, so cap anyio there too. macOS-arm now resolves anyio 4.13.0.
unslothai#6578) * Restore sys.modules in test_pre_import_gate_is_transformers_free The test pops transformers and utils.models.model_config from sys.modules to assert the pre-import security gate does not re-import them, but never put them back. A later importer then rebound a fresh utils.models.model_config, so tests that had captured the original instance missed their patches and hit the real path: test_vision_cache patches _is_vision_model_uncached on the original module, but is_vision_model (still bound to that original) ran the real network lookup instead. This produced 17 spurious failures whenever test_ssm_runtime ran before test_vision_cache in the same process. Snapshot the removed modules and restore the original objects in a finally, so the assertions still run against a clean slate while later tests see the same module instances they captured at import time. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
…lothai#6573) * Studio: keep model downloads running across navigation and loads Downloads started from the chat model selector were tied to the staged pick lifecycle, so they were cancelled in cases where Hub downloads keep going. This makes the chat download flow behave like the Hub. - Leaving the chat route or switching thread/project/new chat now detaches the staging UI but keeps the in-flight transfer running in the global download manager (new keepDownload option on abandonStagedModel). - Staging a second pick no longer cancels the previous pick's download, so multiple models/variants can download at once. - Picking a model to download while another model is loading now starts the download in the background instead of refusing, since a download is independent of a load. * Studio: also background-download remote GGUF quants while a model loads isDownloadableHubRepo (wantManagerDownload) excludes GGUF sources, so an uncached remote GGUF quant picked from the chat selector while another model was loading fell through to the 'Another model is already loading' toast instead of downloading in the background. Treat an uncached remote hub GGUF as a background download too, matching the staged-pick download path. Addresses review feedback from gemini-code-assist and codex on PR unslothai#6573. * Studio: only toast a background download once it actually starts The chat background-download path (used when a model is already loading) fired the "Downloading in the background" toast unconditionally, but requestStart can return without starting a job: a cross-transport partial records a conflict that is only resolvable from the Hub download card, and a busy sibling variant returns after its own toast. So the user could be told a download started when none did, with no way to resolve the conflict from chat. requestStart now reports an outcome (started/conflict/busy/error). The chat path only shows the success toast on an actual start and points the user to the Hub when a transport conflict needs resolving. The Hub card surface keeps its existing behavior (it renders the conflict resolver, so it ignores the outcome). * Studio: report background-download outcome from real job state The chat background-download toast trusted requestStart's optimistic "started", but a start can no-op without throwing: startJob finalizes the job as "error" when the backend refuses or fails apiStart, its peer guard skips a fresh start, and hasActiveOrPendingStart trips on a snapshot, peer variant, or pending preflight that is not this request. So the user could be told a download started when none did. Derive the outcome from the actual job state of the exact key (running/cancelling = started, otherwise error/busy), so the toast only fires for a transfer that is really live. Also guard against re-downloading the model that is already loading: the /load flow downloads before it sets the checkpoint, and that fetch is not a download-manager job, so picking the same id+variant again would start a second transfer against the same cache. Detect that pick and surface a "this model is already loading" toast instead. --------- Co-authored-by: Daniel Han <danielhanchen@gmail.com>
…or, not a 4.14 cancel-scope bug) (unslothai#6579) * Studio: correct the anyio<4.14 pin rationale (mixed-install ImportError) The pin comments said "anyio 4.14+ breaks cancel scope on Python 3.13", but a clean anyio 4.14.0 works on 3.13 (cancel scopes, Event, and the asyncio backend import all pass). The actual failure is a half-resolved install: anyio 4.14 added TaskHandle, imported by __init__.py and _backends/_asyncio from _core/_tasks. When a stale 4.13 _core/_tasks (no TaskHandle) sits under 4.14's importers, the import raises ImportError and 500s the server. Correct the rationale; the <4.14 pin still stands as the way to keep one consistent anyio version. * Clarify the anyio override comment (mixed-install ImportError, not a 4.14 cancel-scope bug)
…ass) (unslothai#6548) * Use UTF-8 for Python code-execution subprocess I/O Studio's code-execution tool already tells the child to emit UTF-8 (PYTHONIOENCODING=utf-8 in _build_safe_env), but _python_exec writes the temp script and decodes the subprocess pipe with the OS default codec. On Windows (cp1252), non-ASCII in model-written code or its output -- arrows, CJK, emoji -- raises UnicodeEncodeError / UnicodeDecodeError and breaks execution. Complete the UTF-8 wiring in core/inference/tools.py: - write the temp script with encoding="utf-8" - decode _python_exec stdout as utf-8, errors="replace" - set PYTHONIOENCODING=utf-8 in _build_bypass_env too (matches _build_safe_env, so the bypass path's child also emits utf-8) The child is python with PYTHONIOENCODING=utf-8, so it emits UTF-8 regardless of the console code page and the decode is always correct. Shell execution via cmd.exe has a separate console-code-page story and is left to a follow-up. Refs unslothai#6489 * Scope Python exec UTF-8 env to Python tool * Make bash bypass test robust to a host-set PYTHONIOENCODING for PR unslothai#6548 Bypass mode preserves benign host env vars, so a host-set PYTHONIOENCODING was inherited into the bash bypass env and tripped the new assertion even though _bash_exec never adds it. Clear it in the test so the assertion checks _bash_exec, not the runner environment. --------- Co-authored-by: Lee Jackson <130007945+Imagineer99@users.noreply.github.com> Co-authored-by: Daniel Han <danielhanchen@gmail.com>
- Discover defaults to the whole Hub instead of the unsloth org; an explicit Unsloth choice is still remembered - Discover models placeholder reads Search all models to match - Give the Unsloth/All scope pill a min width so it stays readable
…ish (unslothai#6592) Polish for the in-chat model picker popover and its guided-tour step. - Search box placeholder reads Search Unsloth models, matching the Unsloth-only listing. - Search Hub button shows a Search all models tooltip on hover. - Floating Eject pill moves 1px lower so it sits closer to the bottom edge. - Results list max height trimmed by 1px (21rem to 335px) from the bottom only. - Chat guided tour Two tabs step updated to describe Unsloth-scoped search plus Search Hub for all of Hugging Face.
…lothai#6597) * Studio: refresh chat tour for the redesigned model picker - Pick a model step describes the Recommended and On Device tabs instead of the old Hub and Fine-tuned split - Find a model step (was Two tabs) covers Unsloth search vs Search Hub, the format and sort filters, and the OOM tag - Settings step now anchors to the run settings panel on the right. The old anchor sat on the open settings button, which unmounts when settings opens, so the tooltip lost its target and drifted left * Add guided tour step for the composer + menu
…pSeek thinking not streaming with a pill on) (unslothai#6947)
* Speed up Studio startup path * Studio: recheck managed binary executability on preflight cache hit and ignore stale unauthenticated platform fetches Preflight: a matching capability cache fingerprint no longer skips the runnability check when the managed binary's executable bit was cleared (size and mtime unchanged, since chmod bumps ctime not mtime). The cache fast path now confirms the binary is still executable, otherwise it falls back to the CLI help probe so preflight reports Stale and can repair, instead of returning Ready and failing later at backend start. Adds a regression test. Frontend: now that first render is no longer gated on fetchDeviceType, the initial unauthenticated health call can resolve after an authenticated platform fetch. Guard the store so a late unauthenticated or failed non-forced response cannot overwrite an already authoritative device type, tunnel URL, or secure flag. Forced refreshes and the first unauthenticated load are unaffected. * Studio: use access(X_OK) for the preflight cache executability guard A mode bitmask treats any execute bit as launchable, but the executable bits can be set only for another owner or group, or be denied by an ACL, so the current user could still hit PermissionDenied at launch and the cached fast path would wrongly return Ready. access(X_OK) checks real executability for the calling user, so an ownership or permission change correctly falls back to the CLI help probe and the Stale repair path. * Studio: ignore any stale non-forced platform fetch once authoritative Extend the platform store guard so a non-forced health response never overwrites an already authoritative result, not only unauthenticated ones. With a saved token the post-render non-forced request can be authenticated but older than a later forced refresh that already picked up the tunnel URL and secure flag; if that earlier request resolves last it would null those fields. Now any non-forced response is dropped once the store holds a server-reported platform. Forced refreshes and the first authoritative write are unaffected. * Studio: run the managed CLI help probe before trusting the preflight cache Restore running the managed CLI help probe before returning Ready from the desktop capability cache, so a managed install whose venv interpreter or a runtime dependency is broken (while path, size, mtime, and markers are unchanged) is reported Stale for repair rather than proceeding to a backend start that cannot spawn. The capability cache still skips the heavier desktop-capabilities probe on a hit, so a warm cache runs one probe instead of two. Removes the executable-access shortcut, which the help probe now subsumes. --------- Co-authored-by: Daniel Han <danielhanchen@gmail.com>
* Polish assistant message actions menu Use the circle question mark (HelpCircleIcon) for the "See response details" action instead of the file-database icon, and lowercase the "Export as markdown" label. * Align response details sheet icon
* Move New badge to System settings tab Show the "New" badge on the System tab and drop it from Connections. * Stabilize refresh revocation UI test * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
* fix: force Unsloth provider selection for opencode * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * opencode: pin the model without clobbering the user's disabled providers The session overlay wrote disabled_providers unconditionally and the inline OPENCODE_CONFIG_CONTENT set disabled_providers to an empty list. Since that inline layer outranks the user's global and project config and opencode replaces the array rather than merging it, every provider the user had disabled was silently re-enabled for the session. Only strip 'unsloth' from an existing disable list, and drop disabled_providers from the inline config. Also insert --model only on a bare launch: it is a global flag for the TUI, so placing it before a passthrough subcommand (serve/run) breaks arg parsing; a subcommand takes the model from the pinned config instead. Parse the printed OPENCODE_CONFIG_CONTENT with shlex.split in the test so it round-trips under POSIX shell quoting. * Re-enable a globally disabled opencode unsloth provider for the session A fresh OPENCODE_CONFIG overlay omits disabled_providers, and opencode replaces that array across config layers only when a higher layer sets the key, so a user's global disabled_providers of ['unsloth', ...] survived the merge and left the session provider disabled even though the overlay defines provider.unsloth and pins the model. Consult the user's global opencode config (XDG_CONFIG_HOME/opencode, or %APPDATA%/opencode on Windows) when the overlay has no list of its own, and when the effective list disables unsloth write it back to the overlay minus unsloth. The provider loads while the user's other disabled providers stay disabled. Best-effort read: a missing or unparseable global config is a no-op. * Override opencode disabled_providers in the inline layer; keep model flag for TUI flags Re-enabling a disabled unsloth provider now rides in the inline OPENCODE_CONFIG_CONTENT layer instead of the session overlay. The overlay sits below a project opencode.json, which could re-disable the provider; the inline layer outranks both global and project configs and is recomputed each run, so no-launch reruns never reuse a stale generated list. The effective disabled list is read from the project config if the repo sets one, else the global config, across config.json/opencode.json/opencode.jsonc (JSONC tolerated), and written back minus unsloth only when unsloth is disabled. Also keep the pinned --model when the opencode passthrough starts with a top-level TUI flag such as --dir or --continue; only a real subcommand (serve/run/...) takes the model from config, so a leading '-' now still gets --model injected. * Discover the opencode project config by walking up from the cwd opencode finds a project config by searching ancestor directories, not just the cwd. Walk from the cwd up to the filesystem root and use the nearest directory that sets disabled_providers, so running unsloth start opencode from a subdirectory of a repo whose root config disables unsloth still gets the inline override. * Only inject opencode --model on a bare launch; rely on the inline model pin Injecting --model whenever the passthrough started with a flag could place it before a subcommand (e.g. opencode --print-logs serve), which opencode can misparse. --model is unnecessary for any passthrough because the inline OPENCODE_CONFIG_CONTENT pins the model in the highest-priority layer, so the session model is forced without the flag. Restrict --model to the bare launch and pass any other invocation through untouched. * Register the session provider under a dedicated OpenCode id Selecting the Unsloth model reliably required the wrapper to re-enable a user-disabled unsloth provider, which meant reconstructing OpenCode's full disabled_providers resolution (global, OPENCODE_CONFIG overlay, project config discovered via --dir or an ancestor walk, .opencode directories, OPENCODE_CONFIG_DIR, config.json/opencode.json/opencode.jsonc precedence, and {env:} variable substitution) and overriding it in the inline layer. That is unbounded and cannot be kept correct. Register the session provider under a dedicated id (unsloth-studio) instead. A user's disabled_providers list would never target it, so the session model is always selectable and the overlay no longer reads or writes disabled_providers at all: the user's own disables, in whatever config layer, are left exactly as they are. This removes the JSONC parser, the config-directory scan, and the ancestor/global resolution helpers, and the tests that exercised them. * Scope the opencode session to the Studio provider opencode filters every provider, including a config-defined custom one, through its enabled_providers allowlist and disabled_providers denylist, and pinning the model does not bypass that gate (a filtered provider resolves to a not-found error). The provider arrays are also replaced, not merged, across config layers. So a user with an enabled_providers allowlist that omits the session provider would still have the Studio model filtered out. Set enabled_providers to just the session provider and clear disabled_providers in the inline OPENCODE_CONFIG_CONTENT overlay (the highest-priority layer, which replaces these arrays). This guarantees the Studio model loads regardless of the user's provider filters, without reading or reconstructing their multi-layer config. It is session-only: the overlay lives in the env for this launch and never touches the user's config files, so their normal opencode is unchanged and only this session is limited to the Studio provider. Also drop the redundant --model on --no-launch so the printed command stays append-safe for drivers that append a subcommand (the inline pin forces the model), and parse both POSIX and PowerShell no-launch output in the opencode tests so they are not shell-specific. * Pin opencode small_model to the session provider The session allowlists only the Studio provider, but opencode's separate small_model (used for lightweight tasks) could still point at another provider from the user or project config; under the allowlist that provider is filtered, so the lightweight task would resolve a not-found error mid-session even with the main model pinned. Pin small_model to the session model in the same inline overlay so every model use stays on the enabled provider. The session serves one model, so it is the only valid target, and this stays session-only like the rest of the overlay. --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: Daniel Han <danielhanchen@gmail.com> Co-authored-by: Wasim Yousef Said <wasimysdev@gmail.com>
* fix: use Windows Hermes installer from unsloth start * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Skip the Hermes setup wizard during unattended start-install unsloth start hermes auto-installs Hermes and then writes its own session-scoped Hermes config. The install commands, as written, drop into the installer's interactive setup wizard (hermes setup), which prompts for global API keys and model choice and points the user at a different global provider than the one Unsloth just configured, blocking the launch. Pass the installer's skip flag on both platforms: the PowerShell scriptblock form with -SkipSetup, and bash -s -- --skip-setup for the piped POSIX installer. * Refresh PATH from the registry after a Windows agent install A Windows installer persists the agent's directory to the User/Machine PATH in the registry and updates only its own process, so the current process keeps a stale PATH until it restarts (the installers print 'restart your terminal'). The post-install shutil.which then misses the just-installed agent and unsloth start fails with 'installed but isn't on PATH yet', forcing a re-run in a new shell. Merge the registry PATH hives back into the process before re-resolving so a freshly installed agent launches in the same invocation. No-op off Windows and on any read error; only ever augments PATH. * Fix/adjust PATH refresh for PR unslothai#6903 --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: danielhanchen <danielhanchen@gmail.com> Co-authored-by: wasimysaid <112766706+wasimysaid@users.noreply.github.com> Co-authored-by: Wasim Yousef Said <wasimysdev@gmail.com>
…slothai#6851) * Studio: heal DiffusionGemma tool calls into structured tool_calls * Fall back to supports_tools for backends without the passthrough capability * Route DiffusionGemma client tools through passthrough when enable_tools is on * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Drop orphaned strip_tool_call_markup import after syncing with main * Tighten supports_tool_passthrough comment * Re-run CI on current main --------- Co-authored-by: danielhanchen <unslothai@gmail.com> Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: Daniel Han <danielhanchen@gmail.com>
…unslothai#6900) * fix: handle case-variant GGUF cache hits for unsloth start * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * gguf cache: keep split shards co-located and isolate cache tests properly When a cached main shard was reused from an older snapshot, the extra shards were resolved independently and could come from a different snapshot dir (or a fresh download into the current ref), leaving llama.cpp unable to load a multi-shard GGUF whose pieces are split across directories. Only reuse a cached main shard when every sibling shard sits in the same snapshot; otherwise fetch the whole set together so they stay co-located. Also patch huggingface_hub.constants.HF_HUB_CACHE (not just the HF_HUB_CACHE env var) in the two cache tests that seeded a temp cache: the snapshot lookup reads the module constant, so the env-only override let the real cache leak in and skip an asserted download. * Do not let a companion-only cache snapshot shadow real GGUF variants When listing GGUF variants from the local HF cache, a newer snapshot may contain only a companion file (for example a vision projector fetched on demand) while the actual quant files live in an older snapshot. The prior scan returned the first snapshot whose vision flag was set, yielding an empty variant list and hiding the real quants. Keep scanning older snapshots for actual variants and carry the vision flag across snapshots. Also record the disk-space fallback variant's size in expected_sizes so the later cache-reuse probe can size-verify the fallback main shard instead of only checking for its existence. * Propagate cached repo casing to companions and preflight split co-location Two fixes to the case-variant GGUF cache reuse: - Resolve the requested repo id to its cached canonical casing once in load_model, up front, and pass it to the main GGUF and its companions (mmproj / MTP drafter). Previously only _download_gguf resolved the casing internally, so a case-variant request loaded the main file from the canonical cache dir while the companions kept the requested casing and missed the cached vision projector / drafter offline. Extracted the resolution into a shared _resolve_repo_id_casing helper. - Apply the split-shard co-location check in the disk-space preflight. When a split GGUF's shards are cached across different snapshots the whole set is refetched later, so counting them as cached made the preflight read 0 bytes to download, skip the smaller-variant fallback, and then fail the full download on a low-disk machine. * Reuse a co-located split GGUF snapshot and fix split fallback size probe - When reusing a cached split GGUF, scan snapshots for one that holds the whole set co-located instead of taking the newest snapshot's first shard. A newer snapshot with only the first shard no longer shadows an older complete snapshot, so an already-cached split model is reused rather than refetched (which would fail offline). - The disk-space fallback records its size in expected_sizes only for a single-file fallback. _find_smallest_fitting_variant returns the whole variant size, so using it as the first shard's expected size rejected a valid cached first shard of a split fallback and forced a re-download. * Scan for a complete split snapshot in the preflight; require a loaded catalog hit - The disk-space preflight now uses the same co-located snapshot scan as the download path (_cached_colocated_split_main) instead of the newest-snapshot probe, so a newer snapshot holding only the first shard no longer masks an older complete one and trips the smaller-variant fallback for a fully cached split model. - _resolve_model only attaches to a /v1/models entry that is actually loaded (loaded != False). /v1/models also lists cached-but-unloaded catalog entries, and matching one by case skipped /api/inference/load and left the agent pointed at a model that is not resident. * Restrict cross-snapshot GGUF cache reuse to offline Reusing a same-name blob from an older or case-variant snapshot bypasses the Hub revision/etag check, so a repo that updates a GGUF in place could serve stale weights online. Gate the cross-snapshot and case-variant reuse (both the disk-space preflight accounting and the download path) on HF_HUB_OFFLINE. Online, hf_hub_download fetches the current revision and resumes a partial download, so the reuse is unnecessary there; offline it remains the resilience fallback. Marked the two reuse regression tests as the offline scenarios they represent and added an online test asserting a fresh fetch. * Harden offline cache reuse and hub-id detection Three follow-ups on the case-variant GGUF cache path: - Honor every truthy HF_HUB_OFFLINE spelling (1/true/yes/on), not just "1", when gating the cross-snapshot and case-variant cache reuse. With HF_HUB_OFFLINE=true the Hub calls are already offline, so the reuse must trigger or the cached GGUF fails to load; route both the preflight accounting and the download path through the same offline parse the rest of the backend uses. - Resolve mmproj/MTP companions from the actual cached snapshot when offline. resolve_cached_repo_id_case can keep a partial lower-case spelling when any dir exists under the requested casing, so an hf_hub_download on that casing misses the canonical companion; scan every case-variant snapshot and return the cached path. - Restrict the case-insensitive model-id match to syntactically valid hub ids (a single namespace/name over the HF charset). A server-side relative path such as models/Llama/Foo.gguf is no longer treated as a hub id, so it cannot casefold-match a differently cased path on a case-sensitive filesystem. This is host independent, unlike the local-existence probe which cannot see a server path. * Only casefold-match model ids against a loopback Studio A two-segment string like Models/Foo is indistinguishable from a hub id, and the local Path.exists() probe in _is_hub_model_id cannot see a path that exists only on a remote Studio host. So against a remote server, casefolding could attach to a distinct server-side path (Models/Foo vs models/foo) on a case-sensitive filesystem. Gate the case-insensitive match on is_loopback_url(base): only a local Studio, where the existence probe is authoritative, casefolds. For a remote Studio the match is exact and a case-mismatched request falls through to /api/inference/load, whose already-loaded dedup resolves it correctly. --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: Daniel Han <danielhanchen@gmail.com> Co-authored-by: Wasim Yousef Said <wasimysdev@gmail.com>
…ows (unslothai#6382) (unslothai#6928) * Studio: show Hugging Face address on hover for Hub and online model rows The model selector already shows an on-disk path tooltip on local rows, but Hub and online rows showed only the bare repo id, and nothing at all when there was no VRAM estimate. Add an optional hubUrl prop and a hubRepoUrl helper that mirrors localPathTooltip, and surface huggingface.co/<repo_id> on hover for the Discover, search, and downloaded Hub rows. Local and VRAM tooltips are unchanged; the VRAM tooltip now also appends the address line. Closes unslothai#6382 * Studio: use a 700ms hover delay before the model-row tooltip Give the model-row hover tooltip (the Hugging Face address, plus the VRAM and local-path lines it shares) a 700ms open delay instead of showing it instantly, so it does not flash while sweeping the mouse down the list. * Fix/adjust GGUF tooltips for PR unslothai#6928 --------- Co-authored-by: wasimysaid <112766706+wasimysaid@users.noreply.github.com> Co-authored-by: Wasim Yousef Said <wasimysdev@gmail.com>
…nslothai#6957) * Studio: fix link, currency and indentation edge cases in LaTeX rendering Follow-up to unslothai#6914. Three fixes to studio/frontend/src/lib/latex.ts: - Skip reference-link definition URLs ([id]: url) during delimiter conversion, so escaped parens in such URLs are not rewritten as math. - Preserve the opener line's indentation when emitting a display $$ block, so a \[...\] inside a list item stays part of the list. - Stop a currency amount from pairing with a converted span's opening $, which swallowed the price into math (for example $5 + x \(y\)). * Exclude GFM footnote definitions from the reference-URL skip A footnote definition like [^1]: \(x\) had its body treated as a link destination, so leading math was left literal. Skip [^...] labels. * Merge overlapping link destination regions A reference-def token can nest inline-link spans (for example [1]: http://h/[a](b)/foo\(x\)), so the combined spans could overlap and isInRegion's binary search missed the outer one, rewriting the URL. Merge overlapping spans before the search. * Guard lineStart when the display opener is at index 0 Behavior is unchanged (lastIndexOf clamps a negative fromIndex to 0), but the explicit guard avoids relying on that implicit clamp. * Scope to indentation and currency fixes Drop the reference-link URL protection added earlier. It guards a case models effectively never emit (escaped parens in a reference-style URL), and approximating CommonMark reference definitions with a regex needs open-ended special-casing. Keep the two high-value fixes: preserve display math indentation (including multi-line bodies) inside a list item, and stop a currency amount from pairing with a converted span's opening dollar sign.
* feat(studio): route CLI trainer to MLX backend * fix(studio): harden MLX trainer routing * fix(studio): harden MLX trainer adapter routing * test(studio): assert MLX CLI activation order * fix(studio): address MLX CLI review feedback * feat(cli): support MLX in legacy script * fix(cli): adapt MLX tokenizer for raw text * fix(cli): omit unsupported MLX eval batch arg * fix(cli): feed raw text to MLX trainer * Fix CLI MLX routing and Python 3.9 annotations Route the MLX backend through create_mlx_trainer_adapter so the torch-free Apple Silicon path never imports trainer.py (torch/unsloth/trl). Replace from __future__ import annotations with typing.Optional/Union so the CLI annotations stay Python 3.9 compatible without the unused-import lint hit. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Strip return_tensors from MLX raw-text tokenizer proxy On a torch-free MLX install, RawTextDataLoader calls the tokenizer with return_tensors='pt'; the callable proxy forwarded that to the HF tokenizer, which tried to build torch tensors and failed before training. Drop return_tensors so the MLX path returns plain token ids. * Tighten CLI MLX-backend comments --------- Co-authored-by: Daniel Han <danielhanchen@gmail.com> Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
* feat(cli): detect MLX distributed launch context * feat(mlx): wire distributed inference backend * feat(cli): broadcast MLX distributed chat turns * fix(cli): wait indefinitely for distributed chat turns * fix(cli): report MLX distributed load errors cleanly * fix(mlx): route distributed vlm through loader * fix(cli): detect inline MLX host JSON * fix(studio): harden distributed object sharing * fix(studio): select JACCL distributed backend * fix(cli): abort distributed error paths * Distinguish real stream errors from model text via GenStreamError in distributed CLI * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Fail loud when MLX distributed init returns a singleton group The worker only reaches this block when distributed was explicitly requested. A singleton (size 1) group means the launch failed to form a real group (MLX built without distributed support, or an invalid launch env/hostfile); silently continuing leaves nonzero ranks looping forever on share_distributed_object. Raise instead so the surrounding handler returns a clear load error. * Tighten MLX distributed inference comments --------- Co-authored-by: Daniel Han <danielhanchen@gmail.com> Co-authored-by: danielhanchen <unslothai@gmail.com> Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
…_loss for TRL >= 1.7.0 (unslothai#6904) * Fix PEFT replacement for TRL >= 1.7.0, add missing compute_aux_loss for TRL 1.7.0 * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Fix GRPO for TRL >= 1.7.0: PEFT ref-adapter removal and return arity rl.py: for trl >= 1.7.0, scope the PEFT removal regex to the ref-adapter block only by anchoring the end on ref_param.data.copy_(param.data), so it no longer also deletes the following gradient-checkpointing enable_input_require_grads() block. Neutralize TRL 1.7.0's `if _is_quantized_model:` bf16 cast the same way the existing is_loaded_in_4bit cast is handled. rl_replacements.py: initialize _extra_moe_kwargs before use (it was referenced before assignment whenever compute_aux_loss was passed) and only request output_router_logits when the aux loss is actually wanted. rl_replacements.py: _get_per_token_logps_and_entropies now returns a 3-tuple (logps, entropies, aux_loss) for trl >= 1.7.0 and a 2-tuple for older TRL, matching how every TRL call site unpacks the result. Without this, TRL 1.7.x _generate_and_score_completions unpacks 3 values from a 2-tuple and raises "not enough values to unpack (expected 3, got 2)". * Return zero aux_loss placeholder and drop inference-mode aux collection * GRPO TRL >= 1.7.0: reject router aux-loss opt-in at init; drop zero aux placeholder Unsloth's optimized GRPO forward cannot compute the MoE router auxiliary loss. Previously an explicit opt-in (router_aux_loss_coef > 0) returned a fabricated zero, silently training without the requested load-balancing penalty. Now reject it at trainer init with a clear NotImplementedError, and return None (not zero) for the aux slot of TRL's 3-tuple. Default stays off (coef 0), so the common path is unaffected. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * GRPO hidden-states fallback: free ModelOutput before chunked log-softmax The old/ref logprob fallback binds the full ModelOutput (which holds every layer's hidden_states when output_hidden_states=True) and kept it alive across chunked_hidden_states_selective_log_softmax, an avoidable OOM on large models. Extract logits then del outputs in both the text and VLM branches. * Version-compat CI: proactively catch TRL GRPO breakage The existing TRL canary is a static symbol/source grep: it verifies symbols exist but is blind to structural changes (TRL 1.7.0's 2->3-tuple per-token-logps return arity and restructured PEFT ref-adapter block, which the fix in this PR addresses, both slipped past it because the methods still existed). Two additions: - test_trl_grpo_pinned_symbols.py: extend TRL_TAGS to 1.5/1.6/1.7 and pin the exact source-string contracts the rl.py / rl_replacements.py transforms depend on for TRL >= 1.7.0 (PEFT elif ref-adapter block + enable_input_require_grads survival, if _is_quantized_model, aux_loss_enabled anchor, compute_aux_loss arity). A future TRL change fails on main a few days before the PyPI release. - test_trl_grpo_fake_run.py + a version-compat-ci job: fake-CUDA run that drives the real GRPO/SFT/DPO source-transform patchers against latest + main TRL on a CPU-only runner (no training) and asserts the generated Unsloth trainer still satisfies the transform contracts. Catches behavioral regressions the grep cannot see. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fake-run test: use a normal Version import for the aux gate * version-compat CI: fix fake-run job gate + torch-absent collection - Drop the invalid job-level matrix if (matrix is not available in jobs.<id>.if -> 'Unrecognized named-value: matrix' fails the whole workflow). Use a single job that runs vs TRL latest always and re-runs vs TRL main only on schedule/dispatch via a step-level github.event_name guard. Validated with actionlint. - Module-level skip the fake-run test when torch is absent so daily-fresh-fetch (pytest-only, collects tests/version_compat/) does not crash on the top-level spoof import. * fake-run test: do not skip on import failure unsloth/trl are installed in the grpo-fake-run job, so a failing import is the import-time drift this canary must catch. Keep only the not-installed find_spec skips; let a real import error fail the test. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * GRPO arity gate: regex downgrade + fail loud + CI coverage The TRL < 1.7.0 per-token-logps return downgrade was an exact-string replace anchored on the full return line incl. its comment, so a reformat (e.g. pre-commit) could silently no-op it and ship a 3-tuple to older TRL. Switch to a regex tolerant of comment/whitespace drift, and raise if the anchor stops matching (re.subn count != 1) instead of failing silently. Add a monkeypatched trl_version unit test asserting both arities, since CI only installs TRL >= 1.7.0 and never exercised the downgrade otherwise. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * fake-run: give SFT/DPO a real contract, not just ast-parse The SFT/DPO fake patch runs only checked the generated trainer parses. Also assert the shared QLoRA _is_quantized_model bf16 cast is neutralized (TRL 1.7's spelling, present in both sft_trainer and dpo_trainer), so a structural TRL change to that block is caught for SFT/DPO too, not just GRPO. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * GRPO PEFT ref-adapter removal: lower gate to the TRL 1.4.0 floor The elif is_peft_model(model) and args.beta != 0.0: ref-adapter block was introduced in TRL 1.4.0 and is unchanged through 1.7.x, but the removal was gated at >= 1.7.0, so for 1.4 <= TRL < 1.7 the transform fell through to the 0.27 branch (which matches the older if is_peft_available()... form) and silently no-oped: a PEFT + beta != 0 GRPO run then computed the KL reference from the copied ref adapter instead of the base model. Lower the gate to 1.4.0 and keep the 1.7.0-only router aux-loss fail-fast nested. Widen the pinned-symbol contract test to run from 1.4.0 so the covered versions are actually exercised. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: Daniel Han <danielhanchen@gmail.com>
…#6965) * version-compat CI: fake CPU training runs for SFT/GRPO/DPO Adds a runtime layer on top of the patch-run canary: actually runs trainer.train() for a couple of steps on a CPU-only runner under the CUDA spoof, wrapping a plain tiny HF model in the Unsloth-patched trainer. Exercises the real train() loop (collation, generation, the injected _get_per_token_logps_and_entropies, loss, backward, optimizer) so a TRL or transformers change that breaks the loop at runtime -- not just the source structure -- surfaces here. No GPU, no meaningful numerics. Needs a chain of small CPU shims (eager torch.compile, dynamo suppress, cuda tensor-alloc redirect to CPU, model.for_training/for_inference equivalents) documented inline. Does not exercise Unsloth's Triton/GPU kernels (CPU can't). * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * cpu fake-train: force adamw_torch + disable dynamo for CPU runner On a real CPU-build torch runner (GitHub CI) two things bit that a CUDA-build torch with GPUs hidden masked locally: - The default optimizer is adamw_8bit (bitsandbytes), whose is_on_gpu() check dies on CPU tensors. Force optim=adamw_torch in all three configs. - import unsloth reinstalls the real torch.compile over the eager passthrough, so the GRPO hot path (chunked_selective_log_softmax) actually compiles and inductor picks the spoofed CUDA device, crashing on device props (gcnArchName). Re-apply the eager passthrough after import and flip torch._dynamo.config.disable so every @torch.compile runs eager at call time. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * cpu fake-train: write checkpoints under pytest tmp_path Use pytest's tmp_path for each trainer's output_dir instead of a hardcoded relative temp/ci_* path, so a local pytest run does not leave untracked dirs in the repo tree and the tests are CWD-independent. * version-compat CI: disable dynamo at process level for the fake-run job Set TORCHDYNAMO_DISABLE / TORCH_COMPILE_DISABLE in the fake-run step env so dynamo/inductor is off before conftest.py's early import unsloth, not only via the per-test runtime shim. Defense in depth on the GPU-less runner: the GRPO hot path never compiles regardless of when its functions were decorated. --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
* fix: launch OpenClaw local TUI by default * Fix/adjust OpenClaw launch paths for PR unslothai#6937 * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Default OpenClaw to the local TUI only on a bare invocation The first-arg startswith('-') branch rewrote passthrough globals into a broken command: OpenClaw's grammar is openclaw [--dev] [--profile <name>] <command>, so 'unsloth start openclaw --profile test' became 'openclaw tui --local --profile test', but tui does not accept --profile (or --dev), so the invocation failed. A leading '--flag value' is ambiguous between a global (--profile test) and a tui option (--message hi), so it cannot be reinterpreted safely. Default to the local TUI only when no passthrough args are given, and forward everything else verbatim so OpenClaw parses it under its own grammar. The bare-launch default (the point of this change) is preserved; explicit subcommands and global flags pass through. --------- Co-authored-by: wasimysaid <112766706+wasimysaid@users.noreply.github.com> Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: danielhanchen <danielhanchen@gmail.com> Co-authored-by: Wasim Yousef Said <wasimysdev@gmail.com>
…i#6909) * feat: detect installed coding agent CLIs in Studio settings The API-keys panel only ever showed the "claude" flavor of the `unsloth start` command, so anyone using Codex, OpenCode, OpenClaw, Hermes, or Pi had to manually rewrite the copied command by hand. Add a backend check that looks for each agent's CLI binary on PATH (shutil.which, mirroring the pattern already used elsewhere in studio/backend/utils) and expose it as GET /api/settings/coding-agents. The API-keys panel now renders a picker for all six supported agents, marks the ones it finds installed, and defaults to one of those instead of always falling back to claude. Includes unit tests for the detection helper. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * address review feedback on coding-agent detection Three fixes from PR review: - detect_installed_coding_agents now treats a PATH lookup failure as "not installed" instead of letting it bubble up and break the settings endpoint; added a regression test for it. - CodingAgentsResponse.agents is now typed as an immutable tuple instead of a list built from one, matching CODING_AGENTS itself. - Fixed a race in the API-keys panel: picking an agent while the installed-CLI check is still in flight could get silently overwritten once that check resolved. A ref now tracks whether the user has made a manual choice, so the auto-detected default only applies before that happens. * Address Codex feedback: GGUF gating and remote-detection scope - codex refuses to launch against a non-GGUF (transformers-backed) model (unsloth_cli's _require_gguf_for_codex), so auto-defaulting to it produced a copy-pasteable command that fails immediately whenever the loaded model isn't GGUF. Add useActiveModelIsGguf() (looks up the active checkpoint in the chat runtime store) and a correction effect that steers the auto-pick away from codex unless the loaded model qualifies, without ever touching a choice the user made by hand. - Detection runs via shutil.which on the Studio backend host, which isn't the same machine as the browser in a tunnel/remote session. Reword the 'installed'/'detected' copy to say so explicitly when the tunnel URL is in use, instead of implying the check ran on the viewer's own device. * Rework auto-default per review: loopback gating + inline GGUF check Replaces the previous approach with the exact shape discussed on the PR: - Export isLoopbackHost/normalizeHost from agent-command.ts. The detection endpoint runs shutil.which on the Studio backend, which only describes the browser's own machine when the base this panel targets resolves to loopback. For a LAN or tunnel/remote base, gate the whole thing off -- don't mark anything as "detected" and don't let it drive the default -- instead of just relabeling the copy. - Drop the separate GGUF-correction effect and useActiveModelIsGguf hook. Read useChatRuntimeStore.getState().activeGgufVariant inline inside the existing detection effect's .then() (so it doesn't need to sit in the effect's deps), and pick the first detected agent that isn't codex unless the loaded model is GGUF, leaving the existing default untouched when no compatible agent is detected. Verified both branches (loopback vs LAN/tunnel base, gguf vs non-gguf, manual pick preserved, no-compatible-agent fallback) with a standalone port of the .then() logic. * Address latest Codex findings: stale detection, model swap, cache - Clear detectedAgents (and skip the network call entirely) when the panel leaves a loopback base, instead of leaving a previous loopback detection result marked 'installed' for a command that now targets a LAN/tunnel/ remote host. - Add a separate, network-free correction effect keyed on the live activeGgufVariant: if codex was auto-picked while a GGUF model was loaded and the user then switches to a transformers-backed model while this panel stays mounted, steer away from codex instead of leaving a command that unsloth_cli's _require_gguf_for_codex will now reject. Never touches a manual pick. - Drop coding-agents.ts's module-lifetime cache. Installed-CLI detection is environment state, not a persisted setting, so a stale positive/negative from before the user installed something (or reopened the tab) is worse than one extra cheap local API call per mount; keep only the in-flight de-dupe for concurrent callers. Verified the correction-effect logic (gguf->non-gguf swap with/without a fallback, still-gguf no-op, manual pick never overridden) with a standalone port of the effect. * Make the codex/GGUF auto-pick symmetric in both directions The correction effect only steered away from codex when the model stopped being GGUF; it never steered back toward codex if the model became GGUF *after* a non-GGUF-gated fallback had already picked something else (e.g. codex is the only detected CLI, a transformers model is loaded so the selection correctly falls back to the claude default, then the user loads a GGUF model while the panel stays mounted -- codex never gets reconsidered). Consolidate into one effect that re-derives the preferred detected agent from scratch whenever detectedAgents or activeGgufVariant changes, in either direction, instead of only reacting to the codex-specific downgrade case. The fetch effect now only populates detectedAgents/availableAgents; this effect is the single source of truth for what gets auto-picked from that list. Never overrides a manual choice. Verified both transition directions plus the manual-pick-survives and initial-detection cases with a standalone port of the derivation logic. * Reset the auto-pick to the default when it stops being trustworthy Two more real gaps from the latest Codex pass on d988f52: - The unified derivation effect only handled the case where a *different* detected agent could take over. If codex was the only detected agent and auto-picked while a GGUF model was loaded, then the model stopped being GGUF, 'preferred' came back undefined and the effect silently left the selection on codex -- exactly the command unsloth_cli's _require_gguf_for_codex now rejects. Fall back to DEFAULT_AGENT in that case instead of leaving it untouched. - Leaving a loopback base cleared detectedAgents (so the 'installed' badges correctly disappear) but left whatever agent had been auto-picked from that now-stale, server-side-only detection still selected. Reset to DEFAULT_AGENT there too, unless the user picked by hand. Introduces a shared DEFAULT_AGENT constant instead of repeating the "claude" literal at each reset site. Verified all five cases (both new resets, both manual-pick-survives variants, and the existing multi-detected-agent fallback still preferring another compatible agent over resetting) with a standalone port of the effects. * Derive GGUF-ness from the actual loaded state, not just the variant string activeGgufVariant only covers an HF-repo GGUF pick (a specific quant variant string). A direct local .gguf file -- custom folder, LM Studio, or drag-drop -- is just as much a GGUF the codex preflight (unsloth_cli's _require_gguf_for_codex) would accept, but it never has a "variant" to report, so it read as non-GGUF here even though /api/inference/status correctly reports is_gguf: true for it. That mismatch could leave a Codex-only install not auto-selected, or reset an auto-picked Codex, for a model that actually supports it. Combined activeGgufVariant with activeNativePathToken (covers the drag-drop/picked-file case) and ggufContextLength (only ever populated when the backend last reported is_gguf: true for the active model, see applyActiveModelStatusToStore) so all three paths a model can be GGUF through are covered, matching the same is_gguf-or-equivalent check hasGgufSource already applies to a staged pick elsewhere in this codebase. * Clear stale native-path token on a non-GGUF status refresh When a native (drag-dropped or picked) GGUF was loaded and the backend later switches to a transformers model outside the UI load path, refresh() adopts the new /api/inference/status via setCheckpoint and applyActiveModelStatusToStore. Those reset activeGgufVariant and ggufContextLength but never clear activeNativePathToken, so the isGguf OR stays true after the switch and a Codex-only detection auto-selects unsloth start codex for a non-GGUF model its preflight rejects. Drop activeNativePathToken in applyActiveModelStatusToStore whenever the status is non-GGUF. A real GGUF load reports is_gguf: true, so its token is preserved (the load path owns it); only a non-GGUF status clears it. * Add the AGPL-3.0 header to the new studio contract test * Fix/adjust agent detection for PR unslothai#6909 * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: danielhanchen <danielhanchen@gmail.com> Co-authored-by: wasimysaid <112766706+wasimysaid@users.noreply.github.com>
…he 5.x sidecar (unslothai#6968) * Studio: keep transformers off sys.modules until the training worker activates the sidecar The training worker (core/training/worker.py:run_training_process) decides the per-worker Xet env flip during preflight by importing utils/hf_xet_fallback.py, which eagerly imported unsloth_zoo at module load. unsloth_zoo's __init__ imports transformers, so the default transformers 4.57.x was cached in sys.modules before activate_transformers_for_subprocess prepended the 5.x sidecar to sys.path. Since activation only edits sys.path, the already cached module won, and 5.x models failed to load their tokenizer or config: - Qwen3.5 / GLM-4.7 (tokenizer_class TokenizersBackend): "Tokenizer class TokenizersBackend does not exist or is not currently imported." - gemma-4: "... is not supported yet in transformers==4.57.6." Fix: load the shared unsloth_zoo backend lazily (only when a heavy download helper is first used, which is after activation). child_should_disable_xet and the DEFAULT_* constants are defined locally so importing the shim stays light. The download wrappers, the DownloadStallError class, start_watchdog and get_hf_download_state resolve the shared backend on first use, and the degraded no-unsloth_zoo fallback is preserved. Tests: - test_hf_xet_fallback.py: existing suite kept green via the restored _shared_* seam; the GPU-init retry test now triggers the lazy load explicitly; new guard asserts importing child_should_disable_xet does not import transformers/unsloth_zoo. - test_training_worker_import_discipline.py: new invariant test that the worker preflight imports leave transformers unimported, so this class of regression cannot return silently. Runs in studio-backend-ci (CPU only, no network/GPU/weights). * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: CPU-only guard that activation switches transformers to the model's sidecar version Adds test_worker_activates_correct_transformers.py: runs the real worker preflight (from utils.hf_xet_fallback import child_should_disable_xet) plus the real tier detection and activate_transformers_for_subprocess for a transformers-5.x model (Qwen3.5, tier 530), then asserts the in-process transformers actually switched to the 5.x sidecar. A stale pre-activation import leaves 4.57.x pinned and fails the assertion, which is exactly the TokenizersBackend regression (unslothai#6951). Self-contained CUDA spoof (mirrors tests/_zoo_aggressive_cuda_spoof.py) forces unsloth_zoo down its full, transformers-importing init path on a GPU-less runner; without it unsloth_zoo degrades and never preloads transformers, masking the bug. A one-line stub sidecar stands in for the 5.x venv, so no GPU, network, weights, or real sidecar are needed. Passes on this fix, fails on buggy main. * Studio: load the repo's canonical CUDA spoof in the correct-version guard Load tests/_zoo_aggressive_cuda_spoof.py (the committed spoof the consolidated CI already relies on) as the single source of truth so the guard matches CI and stays robust on a CPU-only torch wheel, where a partial hand-rolled spoof could miss a torch.cuda call and let the unsloth_zoo import raise (masking the bug). Falls back to a minimal inline spoof for a standalone studio checkout. Verified: passes on this fix, fails on buggy main, and the fallback path passes when the spoof file is absent. * Studio: declare the lazily-resolved xet names so ruff F822 stays green DownloadStallError, start_watchdog and get_hf_download_state are provided via the module __getattr__ (PEP 562), so ruff F822 flagged them as undefined names in __all__ and the Source-lint / pre-commit checks went red. Add annotation-only declarations (no value bound, so __getattr__ still resolves them lazily to the shared unsloth_zoo backend) to mark them defined for the linter while keeping F822 active for the rest of __all__. * Studio: tighten comments on the sidecar-activation fix and its tests * Studio: mirror the new MLX-dispatch preflight import in the import-discipline guard The worker preflight now also runs 'from core.training.training import is_apple_silicon_training_platform, should_use_mlx_training_backend' before it activates the transformers sidecar. Add that import (guarded) to the guard's preflight snippet so the invariant test stays a faithful mirror: a future change that makes core.training.training pull transformers/unsloth_zoo eagerly would then be caught too. Verified clean on the current tree (no leak). --------- Co-authored-by: danielhanchen <unslothai@gmail.com> Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: Lee Jackson <130007945+Imagineer99@users.noreply.github.com>
…othai#6311) * Studio: source CPU llama.cpp prebuilts from the unslothai fork * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: reject unknown Linux CPU arches and keep ROCm-tooling hosts off the CPU prebuilt * Studio: extend the resolve-prebuilt ROCm-tooling guard to Windows * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Studio: let ROCm-SDK-only CPU hosts take the fork CPU prebuilt * Studio: accept windows-arm64 prebuilt kind and refresh stale fork-routing comments * Studio: correct stale fork-routing comments and --resolve-prebuilt help * Refresh stale ggml-org routing comments --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: Daniel Han <danielhanchen@gmail.com>
…ring (unslothai#6946) (unslothai#6953) * fix(studio/hub): apply repo_id length limit per segment, not whole string is_valid_repo_id() applied the 96-char limit to the full "namespace/repo_name" string, so a repo with a valid (<=96 char) name but a long combined id was falsely rejected. Match huggingface_hub.validate_repo_id by checking the length per segment instead. Fixes unslothai#6946. * Fix long repo id state filenames * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: Etherll <61019402+Etherll@users.noreply.github.com> Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
…d of leaving them frozen (unslothai#6936) * models: auto-target per-expert Linear MoE experts for LoRA (gpt-oss 4bit) MoE checkpoints whose experts are stored as per-expert nn.Linear ModuleLists could not receive expert LoRA. gpt-oss bnb-4bit is the canonical case: its experts live at mlp.experts.gate_up_projs.<i> and mlp.experts.down_projs.<i> as per-expert Linear4bit modules, not a fused nn.Parameter. The target_parameters path only handles the fused nn.Parameter layout, and the plain gate_proj/up_proj/down_proj leaf names do not match the per-expert indices, so get_peft_model attached LoRA to attention only and left every expert frozen (0 of 1536 on gpt-oss-20b) even though the grouped bnb-4bit training forward exists. Add get_moe_target_modules, the module-LoRA counterpart of get_moe_target_parameters: it detects per-expert Linear ModuleLists under an experts container and returns their suffix target_modules names (gate_up_projs.<i> / down_projs.<i>). get_peft_model in both llama.py and vision.py extends target_modules with these, handling the explicit leaf-list form and the regex form (auto / all-linear / scoped). It is gated on the same MLP-in-scope condition as the parameter path, so an attention-only request still skips the experts. Also gate get_moe_target_parameters on the fused parameter actually existing, so a per-expert-Linear layout no longer produces a dead target_parameters path or a misleading "Enabling LoRA on MoE parameters" line; those experts are handled through target_modules instead. Validated on gpt-oss-20b-unsloth-bnb-4bit (transformers 5.5.0): experts attach (1536 modules, trainable 0.036 percent to 1.65 percent) across the default, None and all-linear paths; training memorizes and the LoRA adapter reproduces exactly after a cold reload in a fresh process. No regression: fused-parameter MoEs (Qwen3-30B-A3B-4bit), non-MoE models, and attention-only requests are unaffected (get_moe_target_modules returns an empty list). Merging these per-expert adapters into a merged_16bit checkpoint is handled by a companion unsloth-zoo change (saving_utils folds each per-expert delta into the fused gate_up_proj / down_proj tensor). With both, the LoRA adapter and the merged_16bit checkpoint reload the trained behavior identically. * models: scope per-expert MoE targets, keep repeat get_peft_model idempotent, warn on old zoo Address review of the per-expert Linear MoE targeting: - Scope get_moe_target_modules to the requested projection leaves (gate/up map to the gate_up ModuleList, down maps to the down ModuleList), so a narrowed request such as target_modules=["down_proj"] no longer also trains gate_up_projs, matching get_moe_target_parameters. - Detect experts through a PEFT-wrapped base_layer as well, and recompute the auto-added expert targets in the llama.py existing-adapter check, so a repeat get_peft_model call with the same arguments stays idempotent instead of raising on the saved expert targets. - Warn when the installed unsloth_zoo cannot fold these per-expert experts into a merged_16bit checkpoint (older releases keep the fused gate_up_proj / down_proj tensors and drop the per-expert deltas), so the expert LoRA is not silently lost on save_pretrained_merged; the fold lands in unsloth-zoo #885. The LoRA adapter itself is unaffected. * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
|
|
||
| function writeStoredLocale(locale: Locale): void { | ||
| try { | ||
| globalThis.localStorage?.setItem(LOCALE_STORAGE_KEY, locale); |
| profile_text += f"model_context_window = {int(window)}\n" | ||
| profile = home / f"{_CODEX_PROFILE}.config.toml" | ||
| if not profile.exists() or profile.read_text(encoding = "utf-8") != profile_text: | ||
| profile.write_text(profile_text, encoding = "utf-8") |
| existing = config.read_text(encoding = "utf-8") if config.exists() else "" | ||
| merged = _merge_codex_config(existing, base) | ||
| if merged != existing: | ||
| config.write_text(merged, encoding = "utf-8") |
| pytest.skip( | ||
| f"Server failed to start within 30 seconds. Output:\n{server_output}" | ||
| ) | ||
| server_output = stdout.decode(errors = "replace") + stderr.decode(errors = "replace") |
| yield "data: [DONE]\n\n" | ||
|
|
||
| return StreamingResponse( | ||
| gen(), |
Closes unslothai#6961) (unslothai#6970) * fix: Remove moot has_blackwell_gpu() function Fixes unslothai#6961. This function skipped flash-attn on Blackwell GPUs because no prebuilt wheel existed; Dao-AILab now ships one and url_exists() already gates resolution. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix: use torchao 0.17.0 for Blackwell Fixes unslothai#6961. Torchao 0.16.0's cpp extensions are built against CUDA 12, so on a CUDA-13 torch (cu130 / Blackwell) they fail to load with "libcudart.so.12: cannot open shared object file". Select 0.17.0 there instead: its cpp targets torch 2.11, so it is skipped cleanly rather than crashing. CUDA-12 / ROCm / CPU torch 2.10 keeps 0.16.0 and its working kernels. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Condense torchao version-selection comments (no behavior change) * Support torch 2.11 in the Studio installer via the torch2.10 prebuilt wheels Map torch 2.11 to the torch2.10 prebuilt wheels for flash-attn, causal-conv1d, and mamba through wheel_utils.prebuilt_wheel_torch_mm, applied in direct_wheel_url (filename) and flash_attn_wheel_url (version). Those torch2.10 CUDA wheels load and pass each project's own test suite on torch 2.11 (verified on B200), so a torch 2.11 environment gets the prebuilt accelerators instead of skipping or building from source. Raise _CUDA_TORCH_PKG_SPEC to <2.12.0 (torchvision <0.27.0, torchaudio <2.12.0) so the CUDA torch repair path can install torch 2.11, where torchao 0.17's cpp kernels load cleanly. Add tests for the mapping. * Keep has_blackwell_gpu as a False stub for future arch gating * Restore has_blackwell_gpu as a return-False probe kept for future arch gating Keep the nvidia-smi compute_cap detection and its two call sites, but short-circuit with return False at the top so flash-attn is no longer skipped on Blackwell (sm_100+ now has prebuilt wheels and url_exists gates resolution). Drop the early return to re-enable arch-based detection later. --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Daniel Han <danielhanchen@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Disposable CI run for unslothai#6979. Do not merge; closed after CI.