Skip to content

Studio: redesign Select model dropdown to match Hub design - #6364

Merged
danielhanchen merged 164 commits into
mainfrom
studio-model-selector-hub-design
Jun 22, 2026
Merged

danielhanchen merged 164 commits into
mainfrom
studio-model-selector-hub-design

Conversation

@shimmyshimmer

@shimmyshimmer shimmyshimmer commented Jun 16, 2026 •

Copy link
Copy Markdown
Member

What

Reworks the chat Select model dropdown so it is easier to scan and lines up with the Hub page's design language. The dropdown now carries the same detail you get on the Hub, with per-tab search, per-section sorting, a format filter and clearer model metadata.

Rows

  • Split owner / name so the model name reads as the primary text and the uploader is de-emphasized.
  • Added an audio capability badge, inferred from the model's HF tags and pipeline tag with a repo-name fallback. Vision and reasoning badges were dropped to keep the rows uncluttered.
  • Added a parameter chip and a DotTag format pill: GGUF, MLX and Safetensors repos each get a coloured tag, adapters their own. Safetensors rows now show param, the Safetensors pill and size so their meta matches GGUF and MLX.
  • The currently loaded model shows a green Loaded marker.
  • Already-downloaded models surface in the Unsloth list with a download-icon badge, and their downloaded quants carry the same delete action as the On Device tab.
  • Over-budget rows keep full-contrast text; the OOM badge alone signals the fit.

Search and sections

  • The Hub tab has three sections: Unsloth, On Device and Custom, selected by a segmented pill toggle that reuses the Hub's .hub-tab-toggle styling (the hub.css selectors are extended to also match .unsloth-model-selector-menu).
  • Search is scoped to the active tab. The Unsloth tab searches the Unsloth Hugging Face listing only; On Device filters downloaded and LM Studio models by name; Custom filters custom-folder models. Each has its own empty state.
  • A Search Hub button sits beside the search bar and opens the Hub Discover page to browse the full catalogue. It matches the section tabs (rounded, no border, soft shadow) and darkens on hover.
  • Models under the local models directory (./models) now show in a Local models group on On Device so they stay selectable, with the same format, search and chat-only GGUF rules as the other local groups.
  • Each section has its own sort dropdown, inline to the right of the tabs at a fixed width:
    • Unsloth: Trending (default), Recommended, Most likes, Downloads, Recent.
    • On Device and Custom: Recent, Downloaded, Size, Name.
  • Recent orders by when you last loaded a model; Downloaded by the file's download date, tracked separately.
  • The On Device group is labelled Unsloth when other providers (LM Studio / Ollama) also appear.
  • A format filter (All / GGUF / MLX / Safetensors) sits next to the sort, applied across every sort.

Recommended

  • The Unsloth section lists Unsloth models live from Hugging Face, sorted by the dropdown, defaulting to Trending.
  • The list pages correctly as you scroll: the observer re-attaches after each loaded page, so a heavily filtered list keeps filling instead of spinning with nothing new.
  • Recommended keeps the most recently created GGUF and MLX models that fit the device budget (0.7 * GPU + 0.7 * RAM); the other sorts list everything with the existing fit badges.
  • GGUF rows that exceed the device get the same OOM badge as safetensors. Mobile GGUF builds are excluded.
  • GGUF repos rarely expose safetensors metadata, so we request the gguf expand field from Hugging Face and read gguf.total for the param count. That sizes repos with no <n>B token in the name (Kimi, MiniMax) so they get a param chip and an OOM badge instead of slipping through unsized.
  • In chat-only mode the Unsloth list keeps only runnable GGUF/MLX formats across every sort.
  • Image and video diffusion models (text-to-image, text-to-video and similar tasks, plus known families like LTX) are dropped from the listing since they cannot run in chat.
  • The redundant unsloth/ prefix is dropped on these rows.

Accessibility and polish

  • The pill toggle uses a WAI-ARIA roving tabindex with Arrow Left/Right navigation; ArrowDown still drops focus into the list.
  • The popover is sized to the tab cluster (550px) so the left and right padding match, with the format and sort dropdowns one gap-2 from the tabs. The two dropdown menus match the Projects activity Select (tight padding, 14px radius). The Search Hub button matches the dropdown width and lines up with the Trending dropdown on the right, and the list scrollbar sits inside the box.
  • The GGUF and MLX format pills use a smaller dot and a tighter dot-to-label gap.
  • The empty On Device state names the active format filter, e.g. No downloaded MLX models yet.
  • Once the list scrolls, its top edge fades under the pinned search / tab bar, matching the sidebar.
  • The search box reads Search models.

Fine-tuned models

  • The Fine-tuned tab is folded into On Device as its own section. A train shortcut on the On Device header jumps to it, and a folder shortcut jumps to Custom Folders. Both headers always show so the shortcuts always have a target.

Load settings and staging

  • Brought up to date with main, which now carries the deferred-load staging flow from Studio: add 'Load on selection' toggle to configure load options before loading #6348. This branch uses that single staging path rather than the old pre-load popup, which is removed.
  • The gear on a downloaded quant stages the model and opens the Run settings sidebar with a Load model button, so context length, KV cache, speculative decoding and tensor parallelism can be set before the one load.
  • The Load on selection preference lives in Settings, Chat (not the popover). On: Unsloth auto-picks the best settings and loads on selection. Off: picking a model opens Run settings to customize first.

Select model settings (Settings, Chat)

  • Load on selection (default on).
  • Expand quantizations (default off): expand every On Device GGUF model's quantizations by default, or keep them behind a click. With it on, clicking a model collapses or re-expands just that model; the collapse state is in memory only, so it resets on reload and when the setting is toggled.
  • Show all quantizations (default on): On Device only. List every quant including ones not downloaded, or downloaded only. Recommended and other browse lists always show every quant.

Both options affect On Device models only, and the copy says so.

Settings layout

  • The Select model settings group (the three options above) sits in the Chat tab, above the Chat menu group.
  • The New badge moves from API keys to the Chat tab.
  • The chat menu item is relabelled Chat with Files (RAG).
  • General drops a duplicate version section and moves the llama.cpp update notification above Helper LLM.

Picker polish

  • The folder browser keeps its list mounted and dims it while refetching, so toggling Show hidden or changing folders no longer flashes.
  • Picker tooltips drop the hover grace area, so moving between the train, folder and gear icons switches the tooltip at once.
  • On Device GGUF rows drop the redundant Quantizations subheading. The Vision badge moves onto the model name, and the repo size is dropped since each quant already shows its own.
  • Eject loaded model sits in a centered footer below the scrollable list, so it stays in view no matter how far the list is scrolled.

Notes

  • On Device GGUF rows show the Vision badge on the model name; Recommended keeps it inline with the variant list.
  • Unrelated tidy-up: removed the em dashes from the Projects export and import labels.
  • Keyboard roving focus is unchanged otherwise. The pill toggle keeps role="tab".
  • The Unsloth section fetches from Hugging Face, like the search box already does, so it needs a network connection. On Device and Custom stay fully local.
  • Backend: adds read-only helper endpoints for the new UI (native context length, a /kv-cache-estimate route, and a quant GGUF resolver) and gates the recommended-folder chips on folders that actually hold a downloaded model. No existing route's contract changes and model loading, training and hardware detection are untouched.

Testing

  • tsc -b type checks clean and vite build succeeds; i18n:check passes parity.
  • Merged with main and adopted its staging (Studio: add 'Load on selection' toggle to configure load options before loading #6348) and Hub as canonical, so the staging store, runtime hook, sidebar and gguf metadata reader match main with no duplicate path. The branch is mergeable.
  • Isolated logic simulations (per-tab search scoping, local ./models surfacing, GGUF sizing and OOM, the chat-only filter, format detection, row metadata) and a cross-browser Playwright harness (Chromium, Firefox, WebKit) for the rendering, roving keyboard nav, scroll fade and localStorage behaviour all pass.

Make the chat Select model picker easier to scan by reusing the Hub
on-device card's visual language.

- Rows now split owner/name, add a param chip, a DotTag format pill,
  a tabular size, and a Loaded marker on the active model.
- Hub models / Fine-tuned tabs reuse the Hub's exact .hub-tab-toggle
  styling (selectors extended in hub.css to the selector menu).
- Add a Downloaded / Recommended / Custom section toggle on the Hub
  tab to filter the list.
- Widen the popover and nudge the scrollbar toward the edge.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request refactors the model selector UI by replacing the standard tabs component with a custom segmented pill toggle (PillTabs) and introducing a new section-based navigation system for model picking. It also improves the row presentation for models by parsing metadata to display formatted tags, parameter counts, and owner information. The review feedback highlights an opportunity to improve the accessibility of the new PillTabs component by implementing WAI-ARIA keyboard navigation patterns, such as arrow key support and proper focus management.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment on lines +246 to +262
{tabs.map((tab) => (
<button
key={tab.value}
type="button"
role="tab"
aria-selected={value === tab.value}
onClick={() => onValueChange(tab.value)}
className={cn(
"relative z-10 inline-flex h-9 min-w-0 flex-1 items-center justify-center rounded-full px-3 text-[12.5px] transition-colors",
value === tab.value
? "text-foreground"
: "text-muted-foreground hover:text-foreground",
)}
>
{tab.label}
</button>
))}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

To comply with WAI-ARIA tablist keyboard interaction patterns, only the active tab should be focusable via tabIndex={0} (with inactive tabs having tabIndex={-1}), and users should be able to navigate between tabs using the ArrowLeft and ArrowRight keys. Adding an onKeyDown handler and setting the appropriate tabIndex will significantly improve the keyboard accessibility of the segmented pill toggle.

      {tabs.map((tab, index) => (
        <button
          key={tab.value}
          type="button"
          role="tab"
          aria-selected={value === tab.value}
          tabIndex={value === tab.value ? 0 : -1}
          onClick={() => onValueChange(tab.value)}
          onKeyDown={(e) => {
            if (e.key === "ArrowRight" || e.key === "ArrowLeft") {
              e.preventDefault();
              const nextIndex = (index + (e.key === "ArrowRight" ? 1 : -1) + tabs.length) % tabs.length;
              onValueChange(tabs[nextIndex].value);
              (e.currentTarget.parentElement?.querySelectorAll('button[role="tab"]')[nextIndex] as HTMLElement)?.focus();
            }
          }}
          className={cn(
            "relative z-10 inline-flex h-9 min-w-0 flex-1 items-center justify-center rounded-full px-3 text-[12.5px] transition-colors",
            value === tab.value
              ? "text-foreground"
              : "text-muted-foreground hover:text-foreground",
          )}
        >
          {tab.label}
        </button>
      ))}

shimmyshimmer added 2 commits June 16, 2026 04:08
Put Downloaded / Recommended / Custom under the search bar in their own
row so Hub models / Fine-tuned no longer wrap. The section toggle uses a
smaller font and sizes each tab to its label instead of equal widths.
Move splitRepoLabel, classifyMetaToken, and parseMetaTokens out of
pickers.tsx into row-meta.ts. No behaviour change; keeps the presentation
logic free of React/DOM deps so it is easy to test in isolation.
@shimmyshimmer

Copy link
Copy Markdown
Member Author

Validated this end to end with isolated simulations before merging.

Pure logic (52 cases, against the extracted row-meta module)
Owner/name split including Windows backslash paths and empty / leading / trailing / double slashes, meta token classification, and the size vs param disambiguation (16 GB vs 4B, ~4.2 GB, TBD not matching TB, MoE 8x7B). All pass.

Cross browser rendering
A harness built from the real compiled CSS, driven by Playwright across Chromium (Chrome and Edge), Firefox, and WebKit (Safari engine), in light and dark mode. 63 assertions per run, all pass on every engine:

  • the broadened :is(.hub-page, .unsloth-model-selector-menu) selector and color-mix() resolve
  • sliding pill geometry lands on the active tab for both 2 and 3 tabs
  • section tabs size to their label (not equal width)
  • owner and name render flush, no overflow at 520px

Static
tsc -b and vite build clean, i18n parity passes, and the diff has no OS specific APIs so Linux, macOS, and Windows render identically.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: b612d3ccc8

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

// Section gating, ignored while searching. Downloaded stays visible while
// searching so a searched-for downloaded model doesn't disappear.
const showDownloaded = showHfSection || section === "downloaded";
const showCustom = !showHfSection && section === "custom";

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Keep custom-folder matches visible while searching

When the user is on the new Custom section and types in the search box, showHfSection becomes true, which forces showCustom to false and removes all custom-folder rows from both the DOM and the roving option list. This means custom local models cannot be searched from the Custom tab; the picker instead switches to downloaded/HF results even if the query matches a custom model, despite the UI presenting the search as spanning the Hub sections.

Useful? React with 👍 / 👎.

Comment on lines +320 to +322
const [activeTab, setActiveTab] = useState<string>(() =>
chatOnly ? chatOnlyTabsDefault : studioTabsDefault,
);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Recompute the source tab when the selection changes

Because activeTab is initialized only once, the selector can stay on hub even after the selected value is later found in loraModels or externalModels (both are populated asynchronously in the chat page). In that case reopening the dropdown for a persisted fine-tuned or connected model hides the currently loaded row until the user manually switches tabs, so the default should be derived again when the selected model/source list changes or when the popover opens.

Useful? React with 👍 / 👎.

const effectiveTab = tabs.some((tab) => tab.value === activeTab)
? activeTab
: tabs[0].value;
const [hubSection, setHubSection] = useState<HubSection>("downloaded");

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Select the Hub section that contains the loaded model

When the current checkpoint is a Hub model that is not in Downloaded (for example a recommended/search result before the cached lists catch up, or a restored selection that only exists in models), hubSection always starts as downloaded, so reopening the selector shows the empty/downloaded view and the new Loaded marker is absent until the user manually switches sections. The new section state needs the same selected-model derivation as the source tab.

Useful? React with 👍 / 👎.

shimmyshimmer added 3 commits June 16, 2026 04:56
Size the Hub models / Fine-tuned tabs to their labels (with side
padding) like the section toggle, instead of stretching full width. Add
a leading download, star, and folder icon to Downloaded, Recommended,
and Custom.
The popover is a flex column, so the fit toggle stretched full width;
add w-fit/self-start so it sizes to its content. Also hide the
Fine-tuned tab when there are no fine-tuned models, defaulting to Hub
models.
Local models (LM Studio, Ollama, custom folders) carry source "local"
and already show in the Hub tab's Downloaded / Custom sections, so
exclude them from the Fine-tuned tab and from its visibility count.
Extract the tab rules into source-tabs.ts.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 9113cd7f8e

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

// Ollama, custom folders) carry source "local" and live in the Hub tab's
// Downloaded / Custom sections instead.
const fineTunedModels = useMemo(
() => loraModels.filter((model) => isFineTunedSource(model.source)),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Keep local model sources selectable

Filtering every source: "local" entry out of the Fine-tuned tab hides local models that the chat page still passes through loraModels (including models_dir entries and LM Studio entries outside chat-only mode). The Hub picker does not add those back: its local refresh only stores lmstudio and custom sources, and LM Studio rows are only rendered for chatOnly, so users with models under the local models directory can no longer select them from the dropdown after this change.

Useful? React with 👍 / 👎.

shimmyshimmer added 2 commits June 16, 2026 05:30
Show LM Studio and other local provider models in the Downloaded
section in all modes (was chat-only). Put Recommended first and make it
the default section. Add a little more space below the search bar.
Replace the static Recommended list (and its collapse chevron) with a
sort dropdown over Unsloth's own models: Recommended, Trending, Most
likes, Downloads, Recently updated. Recommended shows recently uploaded
GGUF/MLX models that fit the device (hidden if they do not); the other
sorts list all Unsloth models, badged but never hidden. Adds a sort
option to useHfModelSearch and a pure recommended-fit helper.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 4b9dd55011

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

const rows = recommendedSearch.results.filter(
(r) => !downloadedSet.has(r.id.toLowerCase()),
);
if (recommendedSort !== "recommended") return rows;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Filter unsupported chat-only recommendations

In chat-only/macOS mode, switching the Recommended list to Downloads/Likes/Trending takes this early return and skips the runnable-format filter. useHfModelSearch returns all unsloth repos for those sorts, so regular safetensors/checkpoint repos that the rest of the picker still hides in chat-only become selectable and are sent to onSelect as Hub models, leaving chat-only users trying to load unsupported models; keep the chat-only GGUF/MLX filter even for alternate sorts.

Useful? React with 👍 / 👎.

recommendedRows.map((r) => {
const id = r.id;
const info = recommendedMeta.get(id);
const isG = isKnownGgufRepo(id);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Honor GGUF tags in recommended rows

For a Recommended search result where Hugging Face marks r.isGguf true but the repo name does not match the -GGUF suffix, the row passes the filter via isRunnableRecommendedFormat(r.id, r.isGguf) but this check ignores that hint. Such rows render and click as plain Hub checkpoints instead of opening GgufVariantExpander, so users cannot choose/download a quantized file for tagged-only GGUF repos; carry r.isGguf into the isG decision.

Useful? React with 👍 / 👎.

obs.disconnect();
};
}, [recommendedSentinel, hasMoreRecommended, recommendedPage, scrollRef]);
}, [recommendedSentinel, recommendedSearch.hasMore, recommendedSearch.fetchMore, scrollRef]);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Re-arm recommended infinite scroll after each page

When the Recommended tab has more results after the first fetchMore, the observer has already disconnected and this effect does not rerun because none of its dependencies change while hasMore stays true. That means the Recommended list can load the initial page plus one extra page, then further scrolling never fetches the remaining Hugging Face results; include a changing page/result/loading value in the dependencies or avoid permanently disconnecting.

Useful? React with 👍 / 👎.

shimmyshimmer added 7 commits June 16, 2026 06:58
…issing

GGUF and MLX repos rarely expose safetensors metadata, so a large model
with no size could pass the Recommended fit check because unknown size was
treated as fitting. Parse the parameter count from the repo id, including
the Gemma E series, and hide anything we still cannot size.
Thread tags and the pipeline tag through the model search results and add a
pure helper that infers vision, reasoning and audio plus the architecture
family, falling back to repo-name keywords when tags are absent.
Give each model row more detail and make the Hub sections easier to scan:

- Show vision, reasoning and audio badges plus the architecture family tag
  on each row, alongside the params, format and size.
- Drop the redundant unsloth/ prefix on the Recommended rows.
- Rename the Recommended section tab to Unsloth and enlarge the section tabs.
- Move the sort dropdown inline to the right of the tabs at a fixed width.
- Add Recent, Size and Downloaded sorting to the Downloaded and Custom tabs.
- Remove the header icons, pad the subheadings, and grow the list height.
- Recommended now lists the most recently created Unsloth repos.
- Narrow the sort dropdown, remove its border, and truncate long labels.
- Tighten the gap between the section tab icons and their labels.
- Remove the architecture family tag from rows since it repeats the name.
Move the segmented pill toggle out of the model selector into its own file so
the Hub picker can reuse it for a format filter without duplicating the markup.
- Re-attach the scroll observer on each loaded page so a filtered Recommended
  list keeps paging until the viewport fills instead of spinning forever with
  nothing new appearing.
- Add an All / GGUF / MLX / Safetensors toggle on the Unsloth listing that
  filters every sort.
…ce, and fade the scroll edge

Sort: default the Recommended view to Trending and add a Name option to
the On Device / Custom sort. Recent now orders by last load time while
Downloaded orders by file date, tracked in localStorage (model-usage.ts).

Formats: show the format filter on all three tabs (Unsloth, On Device,
Custom), exclude mobile GGUF builds from Recommended, and flag GGUF rows
that exceed the device with the same OOM badge as safetensors.

Polish: download-icon badge on already-downloaded Recommended rows, the
hugeicons view stroke-rounded vision badge, Search all models placeholder,
matched popover padding, and a top-edge mask fade once the list scrolls.
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Repo admins can enable using credits for code reviews in their settings.

Repos with no <n>B token in the name (Kimi, MiniMax) had no param count
and so never showed an OOM badge. Request the gguf expand field from
Hugging Face and read gguf.total, so those repos get a param chip and an
OOM badge when they exceed the device budget.

Keep the row name full contrast when over budget (the OOM badge already
signals the fit), shorten the format and sort dropdowns, narrow the
popover, and rename Recently updated to Recent and All formats to All.
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Repo admins can enable using credits for code reviews in their settings.

shimmyshimmer added 2 commits June 16, 2026 18:50
Add WAI-ARIA roving tabindex and Arrow Left/Right navigation to the pill
toggle so only the active tab is in the tab order. Keep the chat-only
GGUF/MLX filter for every Recommended sort, not just Recommended, so
chat-only users do not see unrunnable checkpoints under Trending. Feed
both listings' GGUF hints into repo detection so a tag-only GGUF in
Recommended expands variants instead of loading as a checkpoint.
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Repo admins can enable using credits for code reviews in their settings.

@shimmyshimmer

Copy link
Copy Markdown
Member Author

Went through the review feedback. Dispositions below.

Fixed in 0952e12

Already addressed on the current head

  • Re-arm recommended infinite scroll (Jupyter notebook: No module named 'unsloth' #1272): the observer effect already re-runs on recommendedSearch.results.length, so it re-attaches after each page. That review was on an earlier commit.
  • Recompute the source tab on selection change (LoRA adapters applied twice #289): the popover content is not force-mounted, so it remounts on open and re-derives the active tab from the current value. No stale tab on reopen.

Tracking separately

  • Custom matches while searching (how to train a model by dpo? #1497): correct that the Custom section is not folded into search results yet. Keeping it out of this change since it touches the roving focus order; will follow up.
  • Hub section for the loaded model ("Model is not a full model or PEFT model" despite loadable on another machine #327): the section no longer defaults to Downloaded (it is Unsloth now), so the original case is gone. Deriving the exact section from the selection depends on the async cached lists, leaving for a follow-up.
  • Local model sources in Fine-tuned (Slow Kaggle Performance (2x T4) #260): need to confirm against the chat page's loraModels wiring before changing, since the Hub tab is meant to own local models now. Verifying before touching it.

Validated the three fixes with isolated logic and cross-browser (Chromium, Firefox, WebKit) simulations: roving focus and Arrow nav, the chat-only filter across sorts, and the GGUF-hint detection. tsc -b, vite build and i18n parity are clean. Also updated the branch with main.

Search is now per section. The Unsloth tab searches the Unsloth HF
listing only, On Device filters downloaded and LM Studio models by name,
and Custom filters custom-folder models, each with its own empty state.

MLX repos get an MLX pill mirroring the GGUF tag. Downloaded quants in
the Unsloth and search lists get the same delete action as On Device.

Also: revert the model name to normal weight, narrow the popover to
558px so the format and sort dropdowns sit one gap-2 from the tabs,
tighten the dropdown menus to match the Projects activity Select, and
make the empty On Device state name the active format filter.
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Repo admins can enable using credits for code reviews in their settings.

Models under the local models directory (source models_dir) flow in as local
models but were dropped from every list: filtered out of Fine-tuned and never
re-added by the Hub picker, which kept only LM Studio and custom-folder
sources. Capture them in the local refresh and render a Local models group on
the On Device tab, with the same format, search, and chat-only GGUF rules as
the other local groups.
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.

@danielhanchen

Copy link
Copy Markdown
Member

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 35c8ed46df

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +1276 to +1278
}).catch(() => {
const store = useChatRuntimeStore.getState();
if (!store.pendingSelection) store.resetModelSettingsToLoaded();

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Clear failed auto-load stages

When an autoLoad pick finishes downloading but the subsequent selectModel call fails (for example OOM, gated repo, or validation error), selectModel only clears pendingSelection on success, so this catch usually sees the hidden auto-load stage still present and does nothing. That leaves an orphaned staged model/edited load knobs in the store even though the sheet was never opened; opening Run settings later shows the stale staged pick and those settings can leak into later loads. Clear or abandon the stage when the pending selection still matches this failed auto-load pick.

Useful? React with 👍 / 👎.

Comment on lines +2085 to +2089
const downloadedEmpty =
visibleCachedGguf.length === 0 &&
visibleCachedModelRows.length === 0 &&
sortedLmStudio.length === 0 &&
sortedLocalDir.length === 0;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3 Badge Count fine-tuned rows in the On Device empty state

When the On Device tab has only fine-tuned models (or a search matches only fine-tuned rows), downloadedEmpty is still true because it ignores fineTunedRows. The empty-state block then renders “No downloaded models yet” or “No matching models on device” immediately above the actual Fine-tuned section, so users see a false empty message despite having selectable on-device models.

Useful? React with 👍 / 👎.

// Per-model pre-load inference settings, persisted in localStorage so the load
// dialog can offer "Remember settings for <model>".

const KEY = "unsloth_load_settings";

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3 Badge Include remembered load settings in preference reset

This new localStorage key is not included in GeneralTab’s PREFS_KEYS, so using “Reset all local preferences” leaves remembered per-model context/KV/speculative/tensor-parallel settings behind. After a reset, staging the same model can silently reapply those old load settings, which contradicts the reset action’s promise to clear local preferences.

Useful? React with 👍 / 👎.

Comment thread studio/backend/routes/models.py Outdated
if entry.is_dir():
if not entry.name.startswith("."):
queue.append(entry)
elif entry.name.lower().endswith((".gguf", ".safetensors")):

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3 Badge Recognize PyTorch bin weights in folder recommendations

The new recommended-folder gate only treats .gguf and .safetensors as generic weights, but the local model scanners still support PyTorch .bin checkpoints such as pytorch_model.bin/model*.bin. When a well-known directory contains only those valid .bin weights, /recommended-folders now omits it, so the one-click folder chip disappears even though the folder has a loadable local model. Include the supported .bin weight names in this check.

Useful? React with 👍 / 👎.

- chat-page auto-load: selectModel only clears pendingSelection on success, so a
  failed auto-load left the hidden stage (and its edited load knobs) behind.
  Abandon the stage when it still matches the failed pick.
- model picker: count fine-tuned rows in the On Device empty check so a
  fine-tuned-only tab no longer shows a false 'No models on device' message
  above the Fine-tuned section.
- general settings: add the remembered per-model load settings key to PREFS_KEYS
  so 'Reset all local preferences' actually clears it.
- recommended-folders: recognize PyTorch .bin weights (gated by the scanner's
  weight-name prefixes) so a .bin-only model folder still earns a chip; add tests.
@danielhanchen

Copy link
Copy Markdown
Member

Thanks, all four of these were valid and are addressed in 7b71d64:

  • Clear failed auto-load stages (chat-page): the catch now abandons the stage when pendingSelection still matches the failed auto-load pick (it previously only reverted when the stage was already cleared, so an OOM/gated/validation failure left the orphaned stage and its load knobs behind).
  • Count fine-tuned rows in the On Device empty state (pickers): downloadedEmpty now also checks fineTunedRows, so a fine-tuned-only tab no longer renders "No models on device" above the Fine-tuned section.
  • Include remembered load settings in preference reset (general settings): added unsloth_load_settings to PREFS_KEYS so "Reset all local preferences" clears the per-model remembered settings.
  • Recognize PyTorch bin weights in folder recommendations (recommended-folders): _dir_has_downloaded_model now accepts .bin weights gated by the same name prefixes the scanner uses (pytorch_model/model/adapter_model/consolidated), so a .bin-only folder still earns a chip while tokenizer.bin and friends do not. Added tests.

Frontend typecheck/build/i18n and the backend folder-recommendation tests pass.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.

@danielhanchen

Copy link
Copy Markdown
Member

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 7b71d64d98

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread studio/backend/routes/models.py Outdated
Comment on lines +352 to +353
has_non_gguf = any(path.glob("*.safetensors")) or any(path.glob("*.bin"))
return None if has_non_gguf else "gguf"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Ignore tokenizer .bin files when classifying GGUF dirs

When a suffixless GGUF directory also contains a non-weight .bin file such as tokenizer.bin, this treats it as having non-GGUF weights and returns None instead of "gguf". The frontend relies on this hint for LM Studio/custom folders whose names do not end in -GGUF; without it, those rows are labeled/filtered as plain local checkpoints and are sent through the wrong load path instead of expanding GGUF variants. The .bin check should mirror the weight-name gating used elsewhere rather than counting every .bin companion file.

Useful? React with 👍 / 👎.

Comment on lines +40 to +43
export const CHAT_EXPAND_QUANTIZATIONS_KEY =
"unsloth_chat_expand_quantizations";
export const CHAT_SHOW_ALL_QUANTIZATIONS_KEY =
"unsloth_chat_show_all_quantizations";

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Add selector keys to preference reset

These new selector settings are persisted to localStorage, but the Reset all local preferences path only clears keys listed in GeneralTab's PREFS_KEYS, and these two keys are not included there. After a user toggles Expand quantizations or Show all quantizations, using Reset all local preferences leaves the model selector behavior changed instead of returning to defaults.

Useful? React with 👍 / 👎.

Comment on lines +1639 to +1640
(!hasGgufSource(selection) && !wantManagerDownload) ||
(store.loadOnSelection && selection.isDownloaded)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Avoid staging non-GGUF picks behind GGUF controls

When Load on selection is off, an uncached non-GGUF Hub repo now falls through this branch to stageModel so the settings sheet opens during the manager download. If a GGUF model is currently loaded, ChatSettingsPanel still derives its GGUF controls from the loaded model, so the staged safetensors/MLX pick shows stale context/KV/speculative knobs and can snapshot those values on Load; the staged pick should suppress loaded-model GGUF controls or avoid opening that sheet path for non-GGUF downloads.

Useful? React with 👍 / 👎.

Comment on lines +360 to +362
// Connected when an external model is active, else On Device when the
// user has downloads, else their last section.
setHubSection(wantsConnectedDefault ? "connected" : defaultHubSection());

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Open On Device for active fine-tuned models

After moving fine-tuned models into the Hub section, reopening the selector while a fine-tuned model is active no longer routes to where that active row lives. With no persisted On Device choice (or on a fresh profile), this falls back to Recommended, so the currently loaded fine-tuned model and its delete/reload controls are hidden until the user manually switches to On Device; the old dedicated LoRA tab did this routing via studioTabsDefault.

Useful? React with 👍 / 👎.

Comment on lines +1721 to +1723
return sortedCachedGguf.filter((c) =>
normalizeForSearch(c.repo_id).includes(q),
);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3 Badge Keep format filtering active during On Device search

When the user types a search in On Device, this branch filters cached GGUF rows only by text and stops applying formatFilter. As a result, selecting Safetensors (or MLX) and then searching still shows matching GGUF downloads, while the no-query path correctly hides them; include matchesFormatFilter in the search branch too so the format dropdown remains consistent.

Useful? React with 👍 / 👎.

…nce reset

Follow-up to the codex review on the model_format/recommended-folder paths:

- _dir_model_format and _scan_models_dir treated any .bin (incl. tokenizer.bin)
  as a non-GGUF weight, so a suffixless GGUF folder shipping a companion .bin was
  misclassified as a plain checkpoint and routed through the wrong load path.
  Factor the scanner's weight-name gating into shared _is_weight_bin /
  _has_non_gguf_weights helpers and use them everywhere (also in
  _dir_has_downloaded_model).
- PREFS_KEYS was missing the new 'Select model settings' keys (load on selection,
  expand/show-all quantizations), so 'Reset all local preferences' left them set.
- On Device cached search dropped the active format filter while a query was
  typed; keep matchesFormatFilter applied so the format dropdown stays consistent.

Adds tests for the tokenizer.bin vs weight-.bin classification.
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.

@danielhanchen

Copy link
Copy Markdown
Member

Triaged this round; three were addressed in b1abe33 and two are flagged for the staging/selector work.

Fixed

  • Ignore tokenizer .bin when classifying GGUF dirs (_dir_model_format, _scan_models_dir): both treated any .bin (including tokenizer.bin) as a non-GGUF weight, so a suffixless GGUF folder with a companion .bin was misread as a plain checkpoint and routed through the wrong load path. Factored the scanner's weight-name gating into shared _is_weight_bin / _has_non_gguf_weights helpers and used them in all three weight checks (including _dir_has_downloaded_model). Added tests for tokenizer.bin vs a real weight .bin.
  • Add selector keys to preference reset: PREFS_KEYS was missing the new "Select model settings" keys, so "Reset all local preferences" left them set. Added unsloth_chat_load_on_selection, unsloth_chat_expand_quantizations, and unsloth_chat_show_all_quantizations.
  • Keep format filtering active during On Device search: visibleCachedGguf/visibleCachedModels applied matchesFormatFilter in the no-query branch but dropped it while searching, so Safetensors-selected still showed GGUF downloads once you typed. The search branch now applies the filter too.

Flagged for the staging/selector work (not changed)

These are real but sit inside the actively-developed deferred-load and selector-redesign code and are UX/design calls, so leaving them to the author:

  • Non-GGUF picks staged behind GGUF controls (chat-page -> ChatSettingsPanel): when a GGUF model is loaded and an uncached non-GGUF Hub repo is staged, the panel still derives its GGUF controls (context/KV/speculative) from the loaded model, so those stale knobs show and can be snapshotted on Load. Wants either suppressing the GGUF controls for a non-GGUF staged pick or not opening that sheet path for non-GGUF downloads.
  • Open On Device for an active fine-tuned model (model-selector.tsx): with fine-tuned models now in the Hub section, reopening the selector while a fine-tuned model is active (no persisted section, or a fresh profile) falls back to Recommended, hiding the active row and its delete/reload controls until the user switches to On Device. Wants the default-section routing to detect an active fine-tuned model.

Frontend typecheck/build/i18n and the backend model_format/recommended-folder tests pass.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.

@danielhanchen

Copy link
Copy Markdown
Member

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 73a3f62890

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +104 to +108
if (
latest?.autoLoad &&
(latest.ggufVariant ?? null) === (variant ?? null)
) {
onAutoLoadRef.current?.(latest);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Validate the completed job before auto-loading

When an auto-load download is replaced by another staged pick with the same quantization, this completion handler can act on the new pendingSelection even though the callback belongs to the old repo: abandonStaged() cancels asynchronously and immediately stages the next pick, while the old listener remains until React re-renders. If the old job completes in that window, the variant-only check passes and onAutoLoad tries to load the new repo before its own download has finished. Please also verify the repo/model id associated with the completed job before auto-loading the latest pending selection.

Useful? React with 👍 / 👎.

Comment thread studio/backend/routes/models.py Outdated
Comment on lines +921 to +922
if m.is_file():
return True

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Require blobs before recommending Ollama folders

When an Ollama directory has manifest files but its blobs/ directory is empty or missing the referenced layer blobs (for example after a failed or pruned download), this returns true as soon as it sees any manifest file. _scan_ollama_dir only emits a selectable model after resolving manifest layers to existing blob files, so /recommended-folders can still show a recommended chip that opens to no models—the empty-scaffold regression this helper is intended to avoid. Please verify that at least one referenced blob exists before returning true.

Useful? React with 👍 / 👎.

@@ -473,9 +656,10 @@ function GgufVariantExpander({
ggufVariant: quant,
isDownloaded: isLocalPath ? true : downloaded,
expectedBytes: sizeBytes,
contextLength: nativeContext,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Start downloads even when context is already known

When a repo already has one GGUF variant downloaded, /gguf-variants returns a context_length, so this now attaches contextLength to clicks for every variant, including ones with downloaded === false. The staging effect exits on pendingHasContext before calling startDownloadRef, so choosing an undownloaded quant from that partially cached repo (especially with Load on selection enabled, where the sheet stays closed) never starts the download and never auto-loads. Only seed the context for already-local/downloaded picks, or make pendingHasContext skip only the metadata fetch, not the download.

Useful? React with 👍 / 👎.

if (!gpuData?.available || !gpuData.devices?.length) {
// No discrete GPU (e.g. Mac): still surface system RAM so memory math
// (unified memory) has a budget to work with.
const info: GpuInfo = { ...DEFAULT_GPU, systemRamAvailableGb: ramAvailableGb };

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Apply RAM budget when no GPU is present

When /api/system reports no discrete GPU (for example Mac/unified-memory hosts), this branch now returns a nonzero systemRamAvailableGb while leaving available: false, but the GGUF variant fit path still gates on gpu.available/gpuGb and returns fits before considering system RAM. As a result those hosts see every quantization as fitting (and no OOM/tight ordering) even though the Recommended filter uses the RAM budget. Please either pass/use systemRamAvailableGb for no-GPU hosts or adjust the fit classifier so the new RAM value is honored.

Useful? React with 👍 / 👎.

…on no-GPU hosts

- recommended-folders: only count an Ollama dir once its manifest resolves to an
  on-disk model blob, so a failed/pruned pull no longer surfaces an empty chip
- GGUF variant click: only seed the staged contextLength for already-downloaded
  picks, so choosing an undownloaded quant from a partially cached repo still
  starts its download (the staging effect short-circuits on a known context)
- device fit: classify GGUF variants against the system-RAM budget on no-GPU /
  unified-memory hosts instead of reporting everything as fits, and pass
  systemRamGb to every variant expander regardless of gpu.available
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.

@danielhanchen

Copy link
Copy Markdown
Member

Pushed 5ad8b38 addressing the Codex review (#6364 (review)).

Fixed

  • Require blobs before recommending Ollama folders (backend/routes/models.py). _dir_has_downloaded_model no longer returns true on the mere presence of a manifest file. It now parses each manifest, finds the application/vnd.ollama.image.model layer, and confirms the referenced blob exists on disk (blobs/<digest with ":" -> "-">), mirroring _scan_ollama_dir. A failed or pruned pull (manifest left behind, blob gone) no longer shows a Recommended chip that opens to an empty picker. Added test_ollama_manifest_without_blob_is_false and made test_ollama_with_manifest_is_true use a realistic manifest.
  • Start downloads even when context is already known (model-selector/pickers.tsx). handleVariantClick now only seeds contextLength for already-local/downloaded picks (isLocalPath || downloaded === true). The staging effect short-circuits on a known contextLength (pendingHasContext) before calling startDownload, so attaching it to an undownloaded quant from a partially cached repo was skipping the download entirely (and never auto-loading with Load on selection enabled). Undownloaded quants now download first; the header context fills in on completion as before.
  • Apply RAM budget when no GPU is present (model-selector/pickers.tsx). getGgufFit no longer returns fits for everything when gpuGb is absent. It short-circuits only when there is no budget at all (totalBudgetGb <= 0); on no-GPU / unified-memory hosts the tier collapses to fit-or-oom against the system-RAM budget. effectiveRecommended is gated on totalBudgetGb for the same reason, and systemRamGb is now passed to every variant expander unconditionally rather than gated on gpu.available, so the RAM value actually reaches the classifier on those hosts.

Flagged (not fixed here)

  • Validate the completed job before auto-loading (use-staged-model-preparation.ts). This is a real but narrow race in the staging flow. The fix needs the completion callback to know which repo finished, but useRepoDownload's onComplete only forwards (variant, bytes) (see download-manager/use-repo-download.ts:76), so a robust check requires threading the completed job's repo id through the shared JobListeners.onComplete signature in the download manager. That is your active area, so I left it for you. Happy to make that signature change if you want me to take it.

Scope confirmation (re: unsloth-only)

The chat model dropdown and its Recommended section are unsloth-only: the picker calls the search hook with ownerScope: "unsloth" and the curated id list filters on the unsloth/ prefix, so a typed query never pulls in other owners there. The export / training / onboarding base-model pickers also default to unsloth with no query, and only expand to the full Hub once you type a search (ownerScope: query ? "all" : "unsloth"), which was the earlier approved fix so people can train/export non-unsloth bases. If you would rather lock those to unsloth-only too, it is a one-line change per picker; say the word and I will flip them.

tsc -b, vite build, i18n:check, and the backend tests are green.

@danielhanchen

Copy link
Copy Markdown
Member

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 5ad8b38c7c

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

pinUnslothFirst: true,
keepUnsupportedTags: true,
accessToken,
enabled: online,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Gate Hub search to the active tab

When the picker is on the On Device or Connected section, typing in the search box still enables this Hugging Face search (and the empty recommendedSearch below is also enabled), even though those results are discarded unless section === "recommended". In the On Device flow this produces unnecessary network requests/spinner state for every local-only search and breaks the intended offline/local behavior; gate the HF hooks on the Recommended section or pass an empty/inactive query for other sections.

Useful? React with 👍 / 👎.

expectedBytes: selection.expectedBytes,
nativePathToken: selection.nativePathToken,
isGguf: selection.isGguf,
isHubRepo: wantManagerDownload || undefined,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Avoid staging non-GGUF picks as GGUF settings

When Load on selection is off and an uncached non-GGUF Hub model is selected, wantManagerDownload stages it here, but the settings sheet still computes isGguf as isLoadedGguf || pendingIsGguf. If a GGUF model is currently loaded, the staged safetensors/MLX pick opens Run settings with GGUF-only context/KV/speculative controls from the loaded model; either do not route non-GGUF picks through this staged GGUF settings path or make the sheet prefer the pending model's type while a pending selection exists.

Useful? React with 👍 / 👎.

type="button"
onClick={(e) => {
e.stopPropagation();
useChatRuntimeStore.getState().stageModel({

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Scope the GGUF settings gear to the main runtime

When this action is rendered inside the Compare pane model selectors (they reuse ModelSelector via GeneralCompareHeader), clicking the gear for a downloaded quant still calls the single chat runtime stageModel here and opens/stages the global Run settings instead of configuring that compare pane. In Compare mode this leaves the selected left/right model unchanged and mutates the main chat pending model; hide this action outside the main selector or route it through a pane-scoped handler.

Useful? React with 👍 / 👎.

sortLocalModels(
localDirModels.filter(
(m) =>
(!chatOnly || localModelIsGguf(m)) &&

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Keep local MLX models visible on Mac

In chat-only Mac builds this predicate keeps only GGUF entries from ./models, so a local MLX model is removed before the MLX format filter can match it, even though the new selector now treats MLX as runnable/recommendable on Mac. Users who place or download an MLX model under the local models directory cannot reselect it from On Device; include an MLX check (or a backend format hint) alongside localModelIsGguf.

Useful? React with 👍 / 👎.

…, keep local MLX on Mac

- model picker: only run the Hub search hooks on the Recommended section. On
  Device / Connected render local data, so typing there no longer fires HF
  requests or a spinner and the local/offline flow is preserved
- chat settings: when a pick is staged, decide the GGUF-only controls from the
  staged model's type, not the currently loaded model's. A staged non-GGUF Hub
  repo no longer inherits a loaded GGUF's context/KV/speculative controls
- On Device: keep local MLX builds in ./models selectable on Mac (chat-only ran
  GGUF/MLX only, but the filter dropped MLX before the format toggle)
@danielhanchen

Copy link
Copy Markdown
Member

Pushed a4e57f0 for the Codex review (#6364 (review)).

Fixed

  • Gate Hub search to the active tab (model-selector/pickers.tsx). The two useHubModelSearch hooks now run only when section === "recommended". The On Device and Connected sections render local data (their consumers, recommendedRows / recommendedIds / hfIds, are all recommended-only), so typing on those tabs no longer fires HF requests or spins, and the local/offline flow is preserved. On Device GGUF detection uses modelGgufIds / localModelIsGguf, not the Hub hooks, so it is unaffected.
  • Avoid staging non-GGUF picks as GGUF settings (chat/chat-settings-sheet.tsx). isGguf now prefers the staged model's type while a pending selection exists (pendingSelection != null ? pendingIsGguf : isLoadedGguf). A staged safetensors/MLX pick no longer inherits a loaded GGUF model's context/KV/speculative controls. This aligns the flag with the rest of the sheet, which already keys context off pendingIsGguf.
  • Keep local MLX models visible on Mac (model-selector/pickers.tsx). The ./models chat-only filter said "GGUF/MLX only" but only checked GGUF, dropping MLX before the format toggle. Added a localModelIsMlx name-hint helper and allow it through on Mac (isMac && localModelIsMlx(m)), matching how isRecommendableFormat treats MLX as runnable there. MLX stays hidden on non-Mac hosts since it cannot run.

Flagged (not fixed here)

  • Scope the GGUF settings gear to the main runtime (model-load-settings-action.tsx). Real issue: in Compare panes GeneralCompareHeader reuses ModelSelector, and the gear calls the global useChatRuntimeStore.stageModel directly, so clicking it in a pane mutates the main chat's pending model instead of the pane's selection. The clean fix (hide the gear outside the main selector, or route it through a pane-scoped handler) means threading a new prop through ModelSelector -> ModelSelectorContent -> HubModelPicker -> the 8 GgufVariantExpander sites, plus a Compare-UX call on whether panes should get their own load-settings affordance at all. That is your selector/Compare area, so I left the design choice to you. Happy to wire the "hide in compare" prop if you want it.

tsc -b, vite build, and lint (no new problems vs base: 31 before and after) are green.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.

@danielhanchen
danielhanchen merged commit 44d6727 into main Jun 22, 2026
45 of 48 checks passed
@danielhanchen
danielhanchen deleted the studio-model-selector-hub-design branch June 22, 2026 11:33
shimmyshimmer added a commit that referenced this pull request Jun 23, 2026
PR #6364 added a branch that overrode the unsloth owner avatar with the
bundled circle-logo-small.png sticker. Revert it so unsloth uploads use
the live Hugging Face org avatar again, falling back to the colored
initial tile. Upstream re-uploads are unaffected since they still
resolve through provider logos.

Co-authored-by: Unsloth <michaelhan@Michaels-MacBook-Pro.local>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants