fix(providers): select and verify the requested gemini-web model/mode before answering (#13381) - #13919
Merged
diegosouzapw merged 1 commit intoSep 17, 2026
Conversation
… before answering (#13381) Root cause: GeminiWebExecutor.execute() opened the identical fixed https://gemini.google.com/app URL and ran the identical Playwright interaction sequence for every advertised gweb/<model> id. `model` was read only AFTER the response was captured, purely to stamp the OpenAI-shaped response — never to influence what was actually clicked/typed, so two different advertised models produced byte-identical automation and the response `model` field was a caller-supplied label, not an observed fact. Fix (owner decision, Option B): a new model -> Gemini UI mode map (open-sse/executors/gemini-web/modeSelection.ts) drives an in-browser selection step before anything is typed — try the mode control, read back the active-mode indicator, and only proceed on a confirmed match. An unconfirmed model, or a requested Extended Thinking control that cannot be confirmed (#13381 follow-up comment), fails closed with 400 unsupported_control_for_provider instead of silently running the account default under the requested label. The selectors involved are UNVALIDATED (no live Gemini account from this checkout) — see the PR's "Selector set is UNVALIDATED" section and the required live smoke. Regression test: tests/unit/issue-13381-gemini-web-model-selection.test.ts
5 tasks done
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
… before answering (diegosouzapw#13381) (diegosouzapw#13919) Root cause: GeminiWebExecutor.execute() opened the identical fixed https://gemini.google.com/app URL and ran the identical Playwright interaction sequence for every advertised gweb/<model> id. `model` was read only AFTER the response was captured, purely to stamp the OpenAI-shaped response — never to influence what was actually clicked/typed, so two different advertised models produced byte-identical automation and the response `model` field was a caller-supplied label, not an observed fact. Fix (owner decision, Option B): a new model -> Gemini UI mode map (open-sse/executors/gemini-web/modeSelection.ts) drives an in-browser selection step before anything is typed — try the mode control, read back the active-mode indicator, and only proceed on a confirmed match. An unconfirmed model, or a requested Extended Thinking control that cannot be confirmed (diegosouzapw#13381 follow-up comment), fails closed with 400 unsupported_control_for_provider instead of silently running the account default under the requested label. The selectors involved are UNVALIDATED (no live Gemini account from this checkout) — see the PR's "Selector set is UNVALIDATED" section and the required live smoke. Regression test: tests/unit/issue-13381-gemini-web-model-selection.test.ts
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Refs #13381
Selector set is UNVALIDATED
This PR implements Option B from the plan-file (owner decision, 2026-09-15): real model/mode
selection in the Gemini Web UI, instead of collapsing the catalog to one honest default. It is
built from a checkout with no live Gemini account and no way to inspect Google's current DOM.
Every CSS selector used to locate the mode-switcher and Extended Thinking controls in
open-sse/executors/gemini-web/modeSelection.tsis a best-effort guess, not a value confirmedagainst a real page — and it is explicitly documented as such in that file.
That is fine by design. The executor never trusts an unconfirmed guess: it tries the control,
reads back the active-mode indicator, and only proceeds on a confirmed match. If a selector is
wrong — the likely outcome on the real Gemini UI until the live smoke below is run and these
values are corrected — the lookup fails, the read-back does not confirm the requested mode, and
the caller gets a clear
400 unsupported_control_for_providerinstead of a response silentlylabeled with a model/mode that never actually ran. Even if every selector currently misses, this
PR still ships the value the issue asked for: no more silently mislabeled responses. Today, in
practice, every request for a non-default model or for Extended Thinking will most likely fail
closed with 400 until the selectors are corrected against a live account (see the Live check
section — this PR is
hold-vpsgated on exactly that).Root cause (short)
GeminiWebExecutor.execute()opened the identical fixed URL(
https://gemini.google.com/app) and ran the identical Playwright interaction sequence(type → Enter) for every advertised
gweb/<model>id. Themodelfield was read onlyafter the response was already captured, purely to stamp the OpenAI-shaped response —
never to influence what was actually clicked/typed. Two structurally different advertised
models therefore produced byte-identical upstream behavior, and the response
modelfieldwas a caller-supplied label, not an observed fact (#13381). The follow-up comment
(2026-09-11) noted the same problem applies to Extended Thinking: eligible accounts expose a
real toggle, but OmniRoute never detected, selected, or verified it.
Fix
open-sse/executors/gemini-web/modeSelection.ts: a model → Gemini UI mode mapfor the three advertised
gweb/<model>ids, plus a descriptor for the Extended Thinkingcontrol.
gemini-3.1-prois treated as the modegemini.google.com/appalready opens to(no interaction attempted — this is also why every pre-existing gemini-web test, which all
happen to use
gemini-3.1-pro, needed zero changes). Every other model, and ExtendedThinking, go through
selectGeminiUiMode(): locate the toggle → click it → locate theactive-mode indicator → read back its text → confirm it matches the expected pattern.
Any failure at any step (
control_not_found,indicator_not_found,indicator_mismatch,or
unknown_model) returns{ confirmed: false, reason }— never a silent pass.open-sse/executors/gemini-web.ts:modelIdis now resolved before the browserlaunches (previously only after the response). Right after the page loads and before
anything is typed, the executor calls
selectGeminiModel()and, whenreasoning_effortasks for real thinking,selectGeminiExtendedThinking(). An unconfirmedresult returns
400 unsupported_control_for_provider(reusingGEMINI_WEB_UNSUPPORTED_CONTROL_CODEfrom./gemini-web/capabilities.ts, the same codefix(providers): gemini-web silently ignores reasoning_effort and tool_choice=required #9356 established) and the browser is closed via the existing
finallyblock — therequest never reaches the prompt editor.
reasoning_efforthandling changed at the executor level (see "Existing tests aligned"):before this PR it was a blanket pre-browser rejection for any effort above "minimal",
regardless of account. That is the same shape of dishonesty fix(providers): gemini-web advertises stale model IDs but never selects the requested model in Gemini UI #13381 reports — claiming
Extended Thinking can never be honored when eligible accounts genuinely expose it. It is
now attempted via the same detect-and-verify-and-fail-closed pattern as model
selection, and only rejected when the attempt cannot confirm the control.
tool_choiceforcing is unchanged: no UI control for it exists on any account, so it stillrejects before the credential check and before Playwright launches.
Regression test (path + RED → GREEN)
tests/unit/issue-13381-gemini-web-model-selection.test.ts(new, 8 cases) — a fake,controllable Playwright page proves: two different advertised models now produce different
automation traces (the plan-file's repro, inverted to assert the CORRECT behavior); a
confirmed read-back proceeds to the prompt; a missing control, a mismatched read-back, and an
unknown model id all fail closed with 400
unsupported_control_for_providerand never reachthe prompt editor; the same holds for Extended Thinking;
reasoning_effort: "none"/"minimal"never triggers the Extended Thinking control at all.
RED (against the pre-fix
gemini-web.ts,modeSelection.tspresent but unused):GREEN (against the fixed
gemini-web.ts):Gates run
npx eslint --suppressions-location config/quality/eslint-suppressions.json <changed files>→ 0 errorsnpm run check:open-sse-typecheck→ OK, 0 pre-existing errorsnode scripts/check/check-file-size.mjs→ 3 pre-existing violations, none in files this PR touches (base-red, 🔴 Release branch not green: release/v3.8.51 #13866)node scripts/check/check-complexity.mjs→ OK — 2839 violations (baseline 3218)node scripts/check/check-cognitive-complexity.mjs→ OK — 1281 violations (baseline 1437)node scripts/check/check-test-discovery.mjs→ OK, new test file discoveredDATA_DIR=$(mktemp -d) npm run check:provider-consistency→ OK, unchanged (catalog not touched)node --import tsx/esm --teston the new suite +tests/unit/gemini-web-capabilities-9356.test.ts+ every othergemini-web*/issue-*gemini-web*test file (53 + 13 + 8 = 74 tests) → all greenExisting tests aligned
tests/unit/gemini-web-capabilities-9356.test.ts: two executor-level tests updated to matchthe
reasoning_effortbehavior change described above."executor returns 400 for reasoning_effort=high before launching a browser"→ renamed andrewritten: it now launches a mocked browser with no Extended Thinking control present, and
asserts the SAME 400 /
unsupported_control_for_provider/ no-stack-trace outcome, reachedvia the new in-browser detect-and-fail-closed path instead of a blanket static reject.
"the capability guard runs ahead of the credential check"→ replaced itsreasoning_effortcase with a
tool_choice: "required"case (the control this property still genuinely appliesto post-fix(providers): gemini-web advertises stale model IDs but never selects the requested model in Gemini UI #13381:
tool_choiceforcing never needs a browser to know it cannot be honored).run()helper's default model changed from"gemini-3.6-flash"(unmapped — would nowfail closed for an unrelated reason,
unknown_model) to"gemini-3.1-pro"(the default mode,no interaction attempted), so the two Extended Thinking-specific tests exercise exactly the
control under test.
checkGeminiWebUnsupportedControlsdirectly (the pure function is untouched) or exercise
tool_choice, which is unaffected.gemini-web*test file needed a change: every existing test that callsGeminiWebExecutor.execute()already usedmodel: "gemini-3.1-pro", which is the default modeand requires zero UI interaction — confirmed by running all of them unmodified (see Gates run).
tests/unit/gemini-web.test.ts:167-178(the 3-model catalog list) is unchanged — this PRdoes not touch the registry/catalog (Option B, not Option A).
Live check (required before merge)
This PR is
hold-vpslabeled and must not be merged until the following is run against a realGemini account and the results (pass/fail per model, plus the corrected selectors if any guess
missed) are recorded here or in a follow-up commit:
(
POST /v1/chat/completionswithmodel: "gweb/<id>", e.g. "What is 2+2?"):gweb/gemini-3.1-pro— expect a normal200response; no UI interaction should be visible(Playwright trace/screenshot should show no mode-switcher click).
gweb/gemini-3.7-flash— expect the mode-switcher control to be clicked and the account'sactive-mode indicator to read a Flash-identifying string. PASS:
200with the answer,and the indicator/screenshot confirms Flash was actually active when the prompt was
submitted. FAIL:
400 unsupported_control_for_provider(selector needs correcting), or a200where the indicator does NOT show Flash (this would mean the read-back check itself isunreliable and must be fixed, not loosened).
gweb/gemini-3.1-flash-lite— same as above, expecting a Flash-Lite-identifying indicator.reasoning_effort: "high"againstgweb/gemini-3.1-pro. PASS: the Deep Think / ExtendedThinking control is clicked, the indicator confirms it is active, and the response proceeds.
FAIL:
400(selector needs correcting) or a200with no confirmation that the toggle wasactually engaged.
reasoning_effortomitted), confirm ordinaryrequests are completely unaffected (no selector lookups attempted, matching the
reasoning_effort: "none"/"minimal"unit test).GEMINI_WEB_MODEL_MODES/GEMINI_WEB_EXTENDED_THINKING_MODEinopen-sse/executors/gemini-web/modeSelection.tswith the selectors that actually matched, andremove the "UNVALIDATED" framing from the header comment once confirmed.
Not covered here
gemini-3.1-pro/gemini-3.7-flash/gemini-3.1-flash-lite) isunchanged — the issue also flagged these names as possibly stale vs. the reporter's live UI
(
gemini-3.8-flash/gemini-3.5-flash-lite); a catalog rename is a separate, narrower change(see the closed feat(providers): refresh Gemini Web model catalog #12598) and was out of scope for the owner's Option B decision, which is about
selection correctness, not renaming.
"Selector set is UNVALIDATED" and "Live check" above) — this is why the issue stays open
(
Refs, notCloses) until the live smoke passes.