Skip to content

fix(gemini): derive AI Studio picker controls from model metadata - #111754

Open
jkorzeniak wants to merge 8 commits into
NousResearch:mainfrom
jkorzeniak:fix/google-ai-studio-picker
Open

jkorzeniak wants to merge 8 commits into
NousResearch:mainfrom
jkorzeniak:fix/google-ai-studio-picker

Conversation

@jkorzeniak

@jkorzeniak jkorzeniak commented Sep 15, 2026 •

Copy link
Copy Markdown

What changes

Google AI Studio now lists models from Google's native paginated models[].name catalog with header authentication. A successful listing, including an empty one, replaces curated IDs; failed or partial discovery retains fallback behavior. Malformed individual entries are skipped, and dedicated Computer Use routes are excluded from the generic agent picker.

Thinking controls come from cached models.dev reasoning_options. Desktop and the request builder use the same metadata for supported levels, bounded token budgets, Dynamic, and Off. Unsupported explicit choices are rejected before mutation; inherited settings are normalized. Unknown metadata leaves the override unset instead of advertising an unverified ladder. Unset, Dynamic, and Off remain distinct.

This is scoped to AI Studio; it does not include the separate Antigravity integration or change Vertex's thinking policy. Provider validation is opt-in.

Compatibility with current main

Updated through upstream e6bb65aa2fc895224dcdcfc3802fa26d87e2134f, preserving the existing branch history.

  • Retains the shared clamp_reasoning_config, provider capability metadata, OAuth extension hooks, and account-usage hook.
  • Keeps gateway-reported effort translation, including Ultra→Max labels and route tooltips.
  • Applies authoritative catalog semantics in the shared picker/setup merge. Setup must not resurrect curated IDs after a successful empty listing.

Validation

Native Windows, Python 3.11, Node 22.23.2; isolated checkouts with their own dependencies.

  • After the merge to fe32647090: 259 Python tests across 22 files via scripts/run_tests.sh; 35 Desktop tests across 4 files; full Desktop TypeScript checks passed.
  • After the update through 7c4d2a812e and setup regression fix: 39 Python tests across 4 files passed, including native catalog controls, authoritative setup behavior, and external-process provider initialization.
  • After merging current main at e6bb65aa2f: 50 Python tests across 5 files passed, covering Gemini controls, authoritative setup catalogs, OpenCode retired-model filtering, and generated contracts. Ruff and whitespace checks passed on the two conflict resolutions.
  • Combined with feat(desktop): configurable model/reasoning picker with explicit catalog refresh #112351 on the earlier shared base: 111 Python tests across 5 files and 44 Desktop tests across 4 files passed, including six local integration cases for four picker styles and Gemini budgets. Integration-only tests are not added to either PR.
  • Ruff on conflict resolutions and diff whitespace checks passed. Generated gateway contracts were regenerated and checked.

No new live inference calls or full repository suite were run for this update. Prior live observations are not a benchmark of this revision.

Focused reproduction:

scripts/run_tests.sh tests/agent/test_gemini_catalog_reasoning.py tests/tui_gateway/test_gemini_reasoning_controls.py tests/hermes_cli/test_provider_live_curated_merge.py tests/hermes_cli/test_setup_provider_catalog.py tests/tui_gateway/contracts/test_generated.py
npm run typecheck --workspace apps/desktop
npm run test --workspace apps/desktop -- src/app/shell/model-edit-submenu.test.tsx src/app/shell/model-catalog-menu.test.tsx src/lib/reasoning-effort.test.ts src/app/chat/composer/reasoning-pill.test.tsx

Limits and related work

models.dev metadata may be incomplete or stale; discovery does not make generation probes, and catalog membership does not guarantee account access or quota. Fast is not advertised for AI Studio. Validated reasoning edits during a running turn remain guarded.

Discovery overlaps with #42693 and #62267; retired-ID cleanup with #109896. Earlier closed proposals #85246 and #90801 addressed metadata-driven effort menus. This PR connects AI Studio discovery, controls, validation and request serialization.

AI assistance: OpenAI Codex assisted with investigation, implementation, tests, and this description. Human input defined scope and interaction requirements. Automated checks do not imply exhaustive human review.

@alt-glitch alt-glitch added type/bug Something isn't working P3 Low — cosmetic, nice to have comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint comp/desktop Electron desktop app (apps/desktop/*) comp/tui Terminal UI (ui-tui/ + tui_gateway/) comp/cli CLI entry point, hermes_cli/, setup wizard comp/plugins Plugin system and bundled plugins provider/gemini Google Gemini (AI Studio, Cloud Code) labels Sep 15, 2026
@jkorzeniak
jkorzeniak marked this pull request as ready for review September 15, 2026 10:10
@kyssta-exe

Copy link
Copy Markdown

Summary

AI Studio-only change: native paginated Gemini model discovery replaces the inherited OpenAI-shape reader, and the effort picker is driven by cached models.dev reasoning_options instead of a generic ladder. Picker, inventory, gateway contracts, and request builders share one thinking-control description, so unsupported levels and a misleading Off switch disappear.

What changed

  • New agent/gemini_model_catalog.py: native generativelanguage.googleapis.com discovery with pagination, fail-closed to fallback; filters to text-output function-calling models.
  • New agent/gemini_catalog_reasoning.py: describe_thinking_control / supported_efforts / build_thinking_config mapping effort, budget:*, off, Dynamic, and unset to thinkingLevel/thinkingBudget.
  • agent/models_dev.py, transports, providers, inventory, gateway contracts, Desktop reasoning UI/i18n updated to carry declared controls.

Strengths

  • Root-causes two real bugs (wrong discovery shape; picker/protocol sharing no control model) with structured data instead of another exception table.
  • Conservative failure semantics: partial/failed discovery never becomes authoritative; unknown metadata leaves overrides unset.

Findings

  • agent/gemini_model_catalog.py (fetch_models): a single malformed entry (name missing/non-str) returns None for the entire catalog — consider skipping the entry instead of discarding all good ones.
  • build_thinking_config: enabled is False emits thinkingBudget: 0, yet the PR notes Lite 3.5 rejects Off (400) — confirm the can_disable_reasoning gate always screens this path.
  • Budget clamp max(1, option["min"]) silently rewrites a declared min: 0 — surface the declared bound or comment why 0 is never valid.

Verdict

Needs minor polish (single-entry poisons catalog; verify Off gating end-to-end).

Reviewed using Hermes-Agent

A-061 review follow-up: preserve valid paginated discovery results, document budget sentinels, and cover Off validation through inventory, gateway and request builders.

Copy link
Copy Markdown
Author

Thanks for the review, @kyssta-exe. Addressed in 1dca4f60c8:

  1. Malformed catalog entries: they are now skipped, including invalid names and non-list generation methods. Valid models on later pages remain available; a failed page still invalidates discovery. The new regression cases reproduced six failures before the fix and pass afterward.
  2. Off gating: the existing resolver checks the same declared capabilities used by the picker. Explicit unsupported Off is rejected before mutation or persistence; inherited Off is normalized to an unset level/budget override before serialization. Added coverage for the gateway, native request builder, OpenAI-compatible serialization, and both profile and legacy transport paths. No change to the Off policy was needed.
  3. Minimum budget: clarified the code comments and guide. The numeric input represents positive Thinking-on budgets; zero belongs to the separately gated Off control, and -1 to Dynamic. Raw metadata and positive model-specific minima are preserved. Tests cover source minima of -1, 0 and 512, with and without toggle support, and ensure budget:0 cannot bypass the Off gate.

Validation: 280 Python tests passed across 8 files, including 23 catalog/reasoning and 10 gateway tests; Ruff and whitespace checks passed. These were offline checks, with no new live inference calls. The PR description also includes this follow-up.

The follow-up code and this response were prepared with OpenAI Codex.

A-061 audit follow-up: leave budgets unchanged for providers that do not opt into validation, hide auto in TUI status, and report catalog pagination exhaustion without sensitive data.

Copy link
Copy Markdown
Author

A further audit with Opus running in Hermes identified a TUI regression and an unintended extension of budget validation to other providers. Both were reproduced locally and addressed in 3840f233ef.

  • TUI status label: the gateway now reports an unset effort as auto, but the TUI previously rendered that value literally. effortLabel now hides it. The StatusRule component test checks that auto renders like an unset effort, while explicit effort and Fast remain visible. It failed before the fix and passes afterward.
  • Budget validation scope: the budget:N branch now sits behind validate_reasoning_selection. Providers that do not opt in retain their original settings, for both explicit and inherited selections. This avoids changing their behavior as part of an AI Studio fix. All six regression cases failed before the change and pass afterward.
  • Pagination diagnostics: exhausting the page limit now logs a fallback warning without credentials, page tokens, or response contents. Discovery still rejects a partial catalog.

Dashboard chat embeds the real TUI, so it receives the same label fix. Its separate React reasoning picker reads saved configuration rather than the changed session-info field.

Validation for this follow-up: 236 Python tests across 16 files and 61 TUI tests across 3 files passed, along with TUI typecheck/build, Ruff, ESLint, Prettier and whitespace checks. Six Dashboard helper tests also passed with a minimal Node configuration; the full Dashboard suite was not run because the standard configuration could not load a missing local dependency. No new live inference calls were made.

The PR description has been updated with these findings and validation limits. The fixes and this comment were prepared with OpenAI Codex.

…picker

# Conflicts:
#	hermes_cli/models.py
#	tests/hermes_cli/test_provider_live_curated_merge.py

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint comp/cli CLI entry point, hermes_cli/, setup wizard comp/desktop Electron desktop app (apps/desktop/*) comp/plugins Plugin system and bundled plugins comp/tui Terminal UI (ui-tui/ + tui_gateway/) P3 Low — cosmetic, nice to have provider/gemini Google Gemini (AI Studio, Cloud Code) type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants