Skip to content

Studio: choose the model and quant in the API usage examples - #10313

Closed
NilayYadav wants to merge 21 commits into
unslothai:mainfrom
NilayYadav:fix-usage-examples-model-picker
Closed

NilayYadav wants to merge 21 commits into
unslothai:mainfrom
NilayYadav:fix-usage-examples-model-picker

Conversation

@NilayYadav

Copy link
Copy Markdown
Collaborator

The usage examples in Settings > API only appear once a model is loaded, and there is no way to see them for a different model or quant.

This adds a model dropdown and a quant dropdown at the top of the Usage examples panel. The model list is every chat model downloaded on this server, with the loaded one marked. Picking one rewrites all the snippets. The choice is remembered like the other panel settings.

If the picked model is downloaded but not loaded and "Switch model by request" is off, a short warning explains that the request will not work yet. This keeps the guarantee from #7454 that a copied snippet does not fail.

When nothing is downloaded, the snippet names the same example model the Agents tab uses, with a note that this server does not have it yet, so the request shape is still visible.

To offer quants, /v1/models now lists every quant a repo has on disk in a new quants field. The existing quant field is unchanged. Strings are added to all locales.

@NilayYadav

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 4, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-04T23:57:39.210400Z b4a60ae Manual request
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 5b104aebb5

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +142 to +143
if (keylessOnly) {
return { servable: false, blockedBy: "keyless" };

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Account for keyless idle-reload when reporting servability

When idle auto-unload has freed a model, a credential-less request naming that exact stashed model is allowed to restore it: _maybe_auto_switch_model explicitly retains the idle-reload path for keyless callers and validates the requested ID and quant against the stash. The catalog nevertheless marks the model unloaded, so this unconditional branch reports it as blocked and tells the user to use an API key or load it manually even though the generated keyless request is runnable. Include the idle-unload state in this verdict, at least for the model/quant eligible for restoration.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not taking this one. tests/studio/test_usage_examples_model_source_contract.py::test_idle_unload_does_not_guess_the_stashed_checkpoint asserts that neither this resolver nor the panel hook may read idleReload or idleUnloadActive, because the idle stash is process-wide while the browser checkpoint is not. Reading the idle state here to soften the verdict is the guess that test exists to forbid, and it would be wrong whenever the stash holds a different model than this browser last loaded. The catalog's loaded flag stays the only evidence this component has.

@NilayYadav

NilayYadav commented Sep 4, 2026 •

Copy link
Copy Markdown
Collaborator Author

Review + GitHub Actions evidence

Pinned head: bae21304f65bef298c119e8d3532822323449811, merge base e8a88546ddef77ac6efc3d034ba4a142e7f61179. Merged with current main (971eb36f4); the one conflict, in the contract test upstream #10113 also reflowed, is resolved.

Backend A/B on Actions

Both branches ran the same workflow and the same probe against _openai_catalog_objects(), differing only in the implementation under test. A repo with UD-Q4_K_XL and Q8_0 on disk:

  • Merge base — failing run:
    ROW org/Foo = {'id': 'org/Foo', 'quant': 'UD-Q4_K_XL'} → FAIL: quants: MISSING. /v1/models advertises one quant, so a client cannot offer Q8_0.
  • PR head — passing run:
    ROW org/Foo = {'id': 'org/Foo', 'quant': 'UD-Q4_K_XL', 'quants': ['UD-Q4_K_XL', 'Q8_0']} → PASS.

UI before/after

Chromium at 1180×900, driving the repo's own smoke-settings.html (the real SettingsDialog, no backend) with one fixed /v1/models payload on both sides, from two checkouts pinned to the exact SHAs above.

fact merge base head
comboboxes in the panel 0 2 (Model, Quantization)
model the snippet names unsloth/Qwen3-30B-A3B-GGUF:UD-Q4_K_XL same
after picking gemma-3-27b-it unsloth/Qwen3-30B-A3B-GGUF:UD-Q4_K_XL (no control) unsloth/gemma-3-27b-it-GGUF:UD-Q4_K_XL

Control: the About tab, which this PR does not touch, rendered identically across the two installs — 1 differing pixel out of 1,062,000 at a max channel delta of 1, i.e. antialiasing noise rather than a content change.

PR 10313 before/after

Regression found and fixed during review

The first revision replaced one either/or with three independent conditionals, leaving no branch for model === null && placeholder === false. That is the default state after downloading a model without loading it (DEFAULT_OPENAI_AUTO_SWITCH_ENABLED = False): the panel rendered no snippet and no explanation, where the merge base had shown guidance.

downloaded, none resident, switching off snippet_len explanation shown
merge base 0 yes
first revision 0 no
head 293 yes, plus the not-loaded warning

Also fixed in review: tests/studio/test_usage_examples_model_source_contract.py went 17 passed → 7 failed because the resolver moved to lib/example-model.ts while the static contract still read usage-examples.tsx; the assertions now follow the code. A keyless caller is refused a switch server-side, so that warning no longer tells them to toggle a setting that would not help. The Agents tab and this panel now read the shipped example model from one module instead of restating the literal.

Local: 6901 frontend tests pass, tsc --noEmit clean, i18n:check:strict clean, backend catalog suites 763 passed. Two failures in test_openai_catalog.py on my machine are pre-existing — identical at the merge base, caused by a real model on disk leaking into the scan.

Codex

Eleven rounds. Every P1 and every P2 that reproduced is fixed and covered by a test that fails at the previous commit; five suggestions were declined with reasons in-thread:

  • reading the idle-unload state to soften the not-loaded verdict — test_idle_unload_does_not_guess_the_stashed_checkpoint exists to forbid exactly that guess, since the idle stash is process-wide and the browser checkpoint is not;
  • disabling copy for any unloaded pick — that is the feature; the amber line already states the precondition, and copy is disabled only for the placeholder, which names a model this server has never held;
  • a colon-bearing catalog id — by_id keys on info.model_id or public_model_id(info.id), both bare repo ids;
  • publishing GGUF alternatives for a quantless resident row — that changes what a bare id resolves to, and belongs in its own PR;
  • a second copy's quants when two copies disagree — deliberately withheld, since collect_local_models orders rows by updated_at while the resolver's alias index walks roots in a fixed order, so neither copy is reliably the one resolve_local_gguf reaches.

The later rounds circled one narrow area — two catalog rows describing a single repo — that no scan has been shown to produce. The handling there is now the conservative one: when two copies disagree, GET /v1/models advertises no quants at all rather than a pin that could 404.

chatgpt-codex-connector[bot]

This comment was marked as resolved.

chatgpt-codex-connector[bot]

This comment was marked as resolved.

chatgpt-codex-connector[bot]

This comment was marked as resolved.

chatgpt-codex-connector[bot]

This comment was marked as resolved.

@unslothai unslothai deleted a comment from chatgpt-codex-connector Bot Sep 4, 2026
@unslothai unslothai deleted a comment from chatgpt-codex-connector Bot Sep 4, 2026
@unslothai unslothai deleted a comment from chatgpt-codex-connector Bot Sep 4, 2026
@unslothai unslothai deleted a comment from chatgpt-codex-connector Bot Sep 4, 2026
@unslothai unslothai deleted a comment from chatgpt-codex-connector Bot Sep 4, 2026
@unslothai unslothai deleted a comment from chatgpt-codex-connector Bot Sep 4, 2026
@unslothai unslothai deleted a comment from chatgpt-codex-connector Bot Sep 4, 2026
@unslothai unslothai deleted a comment from chatgpt-codex-connector Bot Sep 4, 2026
@unslothai unslothai deleted a comment from chatgpt-codex-connector Bot Sep 4, 2026
@unslothai unslothai deleted a comment from chatgpt-codex-connector Bot Sep 4, 2026
chatgpt-codex-connector[bot]

This comment was marked as resolved.

@unslothai unslothai deleted a comment from chatgpt-codex-connector Bot Sep 4, 2026
@unslothai unslothai deleted a comment from chatgpt-codex-connector Bot Sep 4, 2026
@unslothai unslothai deleted a comment from chatgpt-codex-connector Bot Sep 4, 2026
@NilayYadav

Copy link
Copy Markdown
Collaborator Author

@codex review

1 similar comment
@NilayYadav

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 2cd9504a5b

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +838 to +842
const handlePickModel = (id: string) => {
// Picking the model the server serves on its own means "follow it" again.
setPickedModel(
followed !== null && sameBaseModelId(id, followed) ? null : id,
);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Make the follow-current reset reachable

When a persisted pick becomes the currently served model—especially when it is the only catalog entry—triggerModel already equals that option's id. Radix Select does not call onValueChange when the controlled value remains unchanged, so selecting the visible model cannot reach this branch and clear apiExampleModel; subsequent model loads therefore leave the examples pinned to the old choice instead of resuming follow-current behavior. Provide a distinct “follow loaded model” option or another explicit reset control.

Useful? React with 👍 / 👎.

Comment on lines +25312 to +25315
# Compare the way `_quant_list` dedupes, or two copies that hold the same
# files under different spellings read as a disagreement and lose their list.
found_key = tuple(q.lower() for q in found)
seen = quants_seen.get(key)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Compare duplicate quant inventories without order sensitivity

When two physical copies of the same model ID contain the same quant labels but their scans return the non-default labels in a different order, these tuples compare unequal and the code publishes quants: []. Every explicit quant pin is valid against either copy in this scenario, so the picker unnecessarily loses all alternatives even though the change promises to expose every on-disk quant. Compare a normalized set of labels for ambiguity while retaining one scan's order for display.

Useful? React with 👍 / 👎.

Comment on lines +966 to +970
{example.option && example.option.quants.length > 0 ? (
<Select
value={splitPinnedQuant(model ?? "").quant ?? ""}
onValueChange={handlePickQuant}
disabled={example.option.quants.length === 1}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Keep a sole unselected quant selectable

When case-variant catalog rows are merged, a loaded row with no quant can acquire a single quants entry from the other spelling, leaving model bare and the quant Select's value empty. This unconditional length-one disable then prevents selecting that only GGUF quant, so the newly added picker cannot switch from the resident non-GGUF model to the available GGUF variant. Disable a one-item selector only when that item is already the selected quant.

Useful? React with 👍 / 👎.

@NilayYadav

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 565e3f5d68

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +90 to +91
if (existing.quants.length === 0 && m.quants?.length) {
existing.quants = m.quants;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Merge vouched quants into loaded fallback rows

When /v1/models returns case-variant rows, loaded rows arrive first and may contain only the singular resident quant, while the differently cased scan row carries the full quants inventory. The singular fallback makes existing.quants.length nonzero, so this branch ignores the plural list; for example, loaded Org/Foo with quant: "Q8_0" plus scanned org/Foo with quants: ["Q8_0", "Q4_K_M"] exposes only Q8_0 in the picker. Merge the vouched plural inventory with the resident quant rather than adopting it only when the existing list is empty.

Useful? React with 👍 / 👎.

Comment thread studio/frontend/src/i18n/locales/en.ts Outdated
Comment on lines +1593 to +1594
usageNoModel:
"Load or download a model to see runnable examples. This server has no model to name yet.",
"Nothing is downloaded yet, so this example names a model this server does not have. Download one from the Hub and the example will name it.",

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3 Badge Describe the absence of a chat model accurately

This message is shown when the filtered chat-model options are empty, not when the server has no downloads. A server containing only downloaded image, audio, unsupported, or currently unservable models therefore tells the user that nothing is downloaded even though those models remain visible elsewhere in Studio. Phrase this as no compatible chat model being available, and make the equivalent correction in the other updated locales.

Useful? React with 👍 / 👎.

@NilayYadav

Copy link
Copy Markdown
Collaborator Author

@codex review

1 similar comment
@NilayYadav

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Swish!

Reviewed commit: 0f2b5cb0a9

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@NilayYadav

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Keep them coming!

Reviewed commit: b4a60aeefd

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@Imagineer99

Copy link
Copy Markdown
Member

The model and quant picker solves a real usability gap, but a saved selection with no available quants can crash the API settings panel or incorrectly show “not loaded.” Normalize the missing quant with pinned ?? null before the residency check and add regression coverage before merging.

Deterministic A/B evidence confirms four crashes and six false warnings on Linux, Windows and macOS. The one-line correction passes all 80 cases twice, plus all 21 existing resolver tests on each OS.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants