Studio: choose the model and quant in the API usage examples - #10313
NilayYadav wants to merge 21 commits into
Conversation
… moved resolver in the contract test
|
@codex review |
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
for more information, see https://pre-commit.ci
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 5b104aebb5
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| if (keylessOnly) { | ||
| return { servable: false, blockedBy: "keyless" }; |
There was a problem hiding this comment.
Account for keyless idle-reload when reporting servability
When idle auto-unload has freed a model, a credential-less request naming that exact stashed model is allowed to restore it: _maybe_auto_switch_model explicitly retains the idle-reload path for keyless callers and validates the requested ID and quant against the stash. The catalog nevertheless marks the model unloaded, so this unconditional branch reports it as blocked and tells the user to use an API key or load it manually even though the generated keyless request is runnable. Include the idle-unload state in this verdict, at least for the model/quant eligible for restoration.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Not taking this one. tests/studio/test_usage_examples_model_source_contract.py::test_idle_unload_does_not_guess_the_stashed_checkpoint asserts that neither this resolver nor the panel hook may read idleReload or idleUnloadActive, because the idle stash is process-wide while the browser checkpoint is not. Reading the idle state here to soften the verdict is the guess that test exists to forbid, and it would be wrong whenever the stash holds a different model than this browser last loaded. The catalog's loaded flag stays the only evidence this component has.
Review + GitHub Actions evidencePinned head: Backend A/B on ActionsBoth branches ran the same workflow and the same probe against
UI before/afterChromium at 1180×900, driving the repo's own
Control: the About tab, which this PR does not touch, rendered identically across the two installs — 1 differing pixel out of 1,062,000 at a max channel delta of 1, i.e. antialiasing noise rather than a content change. Regression found and fixed during reviewThe first revision replaced one either/or with three independent conditionals, leaving no branch for
Also fixed in review: Local: 6901 frontend tests pass, CodexEleven rounds. Every P1 and every P2 that reproduced is fixed and covered by a test that fails at the previous commit; five suggestions were declined with reasons in-thread:
The later rounds circled one narrow area — two catalog rows describing a single repo — that no scan has been shown to produce. The handling there is now the conservative one: when two copies disagree, |
… section that owns it
|
@codex review |
1 similar comment
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 2cd9504a5b
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| const handlePickModel = (id: string) => { | ||
| // Picking the model the server serves on its own means "follow it" again. | ||
| setPickedModel( | ||
| followed !== null && sameBaseModelId(id, followed) ? null : id, | ||
| ); |
There was a problem hiding this comment.
Make the follow-current reset reachable
When a persisted pick becomes the currently served model—especially when it is the only catalog entry—triggerModel already equals that option's id. Radix Select does not call onValueChange when the controlled value remains unchanged, so selecting the visible model cannot reach this branch and clear apiExampleModel; subsequent model loads therefore leave the examples pinned to the old choice instead of resuming follow-current behavior. Provide a distinct “follow loaded model” option or another explicit reset control.
Useful? React with 👍 / 👎.
| # Compare the way `_quant_list` dedupes, or two copies that hold the same | ||
| # files under different spellings read as a disagreement and lose their list. | ||
| found_key = tuple(q.lower() for q in found) | ||
| seen = quants_seen.get(key) |
There was a problem hiding this comment.
Compare duplicate quant inventories without order sensitivity
When two physical copies of the same model ID contain the same quant labels but their scans return the non-default labels in a different order, these tuples compare unequal and the code publishes quants: []. Every explicit quant pin is valid against either copy in this scenario, so the picker unnecessarily loses all alternatives even though the change promises to expose every on-disk quant. Compare a normalized set of labels for ambiguity while retaining one scan's order for display.
Useful? React with 👍 / 👎.
| {example.option && example.option.quants.length > 0 ? ( | ||
| <Select | ||
| value={splitPinnedQuant(model ?? "").quant ?? ""} | ||
| onValueChange={handlePickQuant} | ||
| disabled={example.option.quants.length === 1} |
There was a problem hiding this comment.
Keep a sole unselected quant selectable
When case-variant catalog rows are merged, a loaded row with no quant can acquire a single quants entry from the other spelling, leaving model bare and the quant Select's value empty. This unconditional length-one disable then prevents selecting that only GGUF quant, so the newly added picker cannot switch from the resident non-GGUF model to the available GGUF variant. Disable a one-item selector only when that item is already the selected quant.
Useful? React with 👍 / 👎.
for more information, see https://pre-commit.ci
|
@codex review |
…nd drop the dead quant control
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 565e3f5d68
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| if (existing.quants.length === 0 && m.quants?.length) { | ||
| existing.quants = m.quants; |
There was a problem hiding this comment.
Merge vouched quants into loaded fallback rows
When /v1/models returns case-variant rows, loaded rows arrive first and may contain only the singular resident quant, while the differently cased scan row carries the full quants inventory. The singular fallback makes existing.quants.length nonzero, so this branch ignores the plural list; for example, loaded Org/Foo with quant: "Q8_0" plus scanned org/Foo with quants: ["Q8_0", "Q4_K_M"] exposes only Q8_0 in the picker. Merge the vouched plural inventory with the resident quant rather than adopting it only when the existing list is empty.
Useful? React with 👍 / 👎.
| usageNoModel: | ||
| "Load or download a model to see runnable examples. This server has no model to name yet.", | ||
| "Nothing is downloaded yet, so this example names a model this server does not have. Download one from the Hub and the example will name it.", |
There was a problem hiding this comment.
Describe the absence of a chat model accurately
This message is shown when the filtered chat-model options are empty, not when the server has no downloads. A server containing only downloaded image, audio, unsupported, or currently unservable models therefore tells the user that nothing is downloaded even though those models remain visible elsewhere in Studio. Phrase this as no compatible chat model being available, and make the equivalent correction in the other updated locales.
Useful? React with 👍 / 👎.
|
@codex review |
1 similar comment
|
@codex review |
|
Codex Review: Didn't find any major issues. Swish! Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
…ry on-disk quant offered
|
@codex review |
|
Codex Review: Didn't find any major issues. Keep them coming! Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
|
The model and quant picker solves a real usability gap, but a saved selection with no available quants can crash the API settings panel or incorrectly show “not loaded.” Normalize the missing quant with Deterministic A/B evidence confirms four crashes and six false warnings on Linux, Windows and macOS. The one-line correction passes all 80 cases twice, plus all 21 existing resolver tests on each OS. |



The usage examples in Settings > API only appear once a model is loaded, and there is no way to see them for a different model or quant.
This adds a model dropdown and a quant dropdown at the top of the Usage examples panel. The model list is every chat model downloaded on this server, with the loaded one marked. Picking one rewrites all the snippets. The choice is remembered like the other panel settings.
If the picked model is downloaded but not loaded and "Switch model by request" is off, a short warning explains that the request will not work yet. This keeps the guarantee from #7454 that a copied snippet does not fail.
When nothing is downloaded, the snippet names the same example model the Agents tab uses, with a note that this server does not have it yet, so the request shape is still visible.
To offer quants, /v1/models now lists every quant a repo has on disk in a new quants field. The existing quant field is unchanged. Strings are added to all locales.