Skip to content

fix(providers): select and verify the requested gemini-web model/mode before answering (#13381) - #13919

Merged
diegosouzapw merged 1 commit into
release/v3.8.51from
fix/13381-gemini-web-model-selection
Sep 17, 2026
Merged

diegosouzapw merged 1 commit into
release/v3.8.51from
fix/13381-gemini-web-model-selection

Conversation

@diegosouzapw

Copy link
Copy Markdown
Owner

Refs #13381

⚠️ base-red inherited: #13866

Selector set is UNVALIDATED

This PR implements Option B from the plan-file (owner decision, 2026-09-15): real model/mode
selection in the Gemini Web UI, instead of collapsing the catalog to one honest default. It is
built from a checkout with no live Gemini account and no way to inspect Google's current DOM.
Every CSS selector used to locate the mode-switcher and Extended Thinking controls in
open-sse/executors/gemini-web/modeSelection.ts is a best-effort guess, not a value confirmed
against a real page — and it is explicitly documented as such in that file.

That is fine by design. The executor never trusts an unconfirmed guess: it tries the control,
reads back the active-mode indicator, and only proceeds on a confirmed match. If a selector is
wrong — the likely outcome on the real Gemini UI until the live smoke below is run and these
values are corrected — the lookup fails, the read-back does not confirm the requested mode, and
the caller gets a clear 400 unsupported_control_for_provider instead of a response silently
labeled with a model/mode that never actually ran. Even if every selector currently misses, this
PR still ships the value the issue asked for
: no more silently mislabeled responses. Today, in
practice, every request for a non-default model or for Extended Thinking will most likely fail
closed with 400 until the selectors are corrected against a live account (see the Live check
section — this PR is hold-vps gated on exactly that).

Root cause (short)

GeminiWebExecutor.execute() opened the identical fixed URL
(https://gemini.google.com/app) and ran the identical Playwright interaction sequence
(type → Enter) for every advertised gweb/<model> id. The model field was read only
after the response was already captured, purely to stamp the OpenAI-shaped response —
never to influence what was actually clicked/typed. Two structurally different advertised
models therefore produced byte-identical upstream behavior, and the response model field
was a caller-supplied label, not an observed fact (#13381). The follow-up comment
(2026-09-11) noted the same problem applies to Extended Thinking: eligible accounts expose a
real toggle, but OmniRoute never detected, selected, or verified it.

Fix

  • New module open-sse/executors/gemini-web/modeSelection.ts: a model → Gemini UI mode map
    for the three advertised gweb/<model> ids, plus a descriptor for the Extended Thinking
    control. gemini-3.1-pro is treated as the mode gemini.google.com/app already opens to
    (no interaction attempted — this is also why every pre-existing gemini-web test, which all
    happen to use gemini-3.1-pro, needed zero changes). Every other model, and Extended
    Thinking, go through selectGeminiUiMode(): locate the toggle → click it → locate the
    active-mode indicator → read back its text → confirm it matches the expected pattern.
    Any failure at any step (control_not_found, indicator_not_found, indicator_mismatch,
    or unknown_model) returns { confirmed: false, reason } — never a silent pass.
  • open-sse/executors/gemini-web.ts: modelId is now resolved before the browser
    launches (previously only after the response). Right after the page loads and before
    anything is typed, the executor calls selectGeminiModel() and, when
    reasoning_effort asks for real thinking, selectGeminiExtendedThinking(). An unconfirmed
    result returns 400 unsupported_control_for_provider (reusing
    GEMINI_WEB_UNSUPPORTED_CONTROL_CODE from ./gemini-web/capabilities.ts, the same code
    fix(providers): gemini-web silently ignores reasoning_effort and tool_choice=required #9356 established) and the browser is closed via the existing finally block — the
    request never reaches the prompt editor.
  • reasoning_effort handling changed at the executor level (see "Existing tests aligned"):
    before this PR it was a blanket pre-browser rejection for any effort above "minimal",
    regardless of account. That is the same shape of dishonesty fix(providers): gemini-web advertises stale model IDs but never selects the requested model in Gemini UI #13381 reports — claiming
    Extended Thinking can never be honored when eligible accounts genuinely expose it. It is
    now attempted via the same detect-and-verify-and-fail-closed pattern as model
    selection, and only rejected when the attempt cannot confirm the control.
    tool_choice forcing is unchanged: no UI control for it exists on any account, so it still
    rejects before the credential check and before Playwright launches.

Regression test (path + RED → GREEN)

tests/unit/issue-13381-gemini-web-model-selection.test.ts (new, 8 cases) — a fake,
controllable Playwright page proves: two different advertised models now produce different
automation traces (the plan-file's repro, inverted to assert the CORRECT behavior); a
confirmed read-back proceeds to the prompt; a missing control, a mismatched read-back, and an
unknown model id all fail closed with 400 unsupported_control_for_provider and never reach
the prompt editor; the same holds for Extended Thinking; reasoning_effort: "none"/"minimal"
never triggers the Extended Thinking control at all.

RED (against the pre-fix gemini-web.ts, modeSelection.ts present but unused):

✖ #13381: distinct advertised gemini-web model IDs drive distinguishable Playwright automation …
  AssertionError [ERR_ASSERTION]: expected the automation trace to DIFFER between two distinct
  advertised models … — but it is identical
✖ #13381: a confirmed model-mode read-back lets the request proceed to the prompt
✖ #13381: a missing model-mode control fails closed with 400, not a silent 200
  502 !== 400
✖ #13381: a mismatched read-back fails closed with 400, not a silent success
  502 !== 400
✖ #13381: an advertised-but-unmapped model id fails closed instead of running the account default
  502 !== 400
✖ #13381: Extended Thinking is attempted and, once confirmed, the request proceeds
  400 !== 502
✖ #13381: Extended Thinking fails closed with 400 when the control cannot be confirmed
  message did not match /Extended Thinking/
ℹ tests 8, pass 1, fail 7

GREEN (against the fixed gemini-web.ts):

✔ #13381: distinct advertised gemini-web model IDs drive distinguishable Playwright automation …
✔ #13381: a confirmed model-mode read-back lets the request proceed to the prompt
✔ #13381: a missing model-mode control fails closed with 400, not a silent 200
✔ #13381: a mismatched read-back fails closed with 400, not a silent success
✔ #13381: an advertised-but-unmapped model id fails closed instead of running the account default
✔ #13381: Extended Thinking is attempted and, once confirmed, the request proceeds
✔ #13381: Extended Thinking fails closed with 400 when the control cannot be confirmed
✔ #13381: reasoning_effort none/minimal never triggers the Extended Thinking control at all
ℹ tests 8, pass 8, fail 0

Gates run

  • npx eslint --suppressions-location config/quality/eslint-suppressions.json <changed files> → 0 errors
  • npm run check:open-sse-typecheck → OK, 0 pre-existing errors
  • node scripts/check/check-file-size.mjs → 3 pre-existing violations, none in files this PR touches (base-red, 🔴 Release branch not green: release/v3.8.51 #13866)
  • node scripts/check/check-complexity.mjs → OK — 2839 violations (baseline 3218)
  • node scripts/check/check-cognitive-complexity.mjs → OK — 1281 violations (baseline 1437)
  • node scripts/check/check-test-discovery.mjs → OK, new test file discovered
  • DATA_DIR=$(mktemp -d) npm run check:provider-consistency → OK, unchanged (catalog not touched)
  • node --import tsx/esm --test on the new suite + tests/unit/gemini-web-capabilities-9356.test.ts + every other gemini-web*/issue-*gemini-web* test file (53 + 13 + 8 = 74 tests) → all green

Existing tests aligned

  • tests/unit/gemini-web-capabilities-9356.test.ts: two executor-level tests updated to match
    the reasoning_effort behavior change described above.
    • "executor returns 400 for reasoning_effort=high before launching a browser" → renamed and
      rewritten: it now launches a mocked browser with no Extended Thinking control present, and
      asserts the SAME 400 / unsupported_control_for_provider / no-stack-trace outcome, reached
      via the new in-browser detect-and-fail-closed path instead of a blanket static reject.
    • "the capability guard runs ahead of the credential check" → replaced its reasoning_effort
      case with a tool_choice: "required" case (the control this property still genuinely applies
      to post-fix(providers): gemini-web advertises stale model IDs but never selects the requested model in Gemini UI #13381: tool_choice forcing never needs a browser to know it cannot be honored).
    • The run() helper's default model changed from "gemini-3.6-flash" (unmapped — would now
      fail closed for an unrelated reason, unknown_model) to "gemini-3.1-pro" (the default mode,
      no interaction attempted), so the two Extended Thinking-specific tests exercise exactly the
      control under test.
    • All other tests in this file are unchanged — they call checkGeminiWebUnsupportedControls
      directly (the pure function is untouched) or exercise tool_choice, which is unaffected.
  • No other gemini-web* test file needed a change: every existing test that calls
    GeminiWebExecutor.execute() already used model: "gemini-3.1-pro", which is the default mode
    and requires zero UI interaction — confirmed by running all of them unmodified (see Gates run).
  • tests/unit/gemini-web.test.ts:167-178 (the 3-model catalog list) is unchanged — this PR
    does not touch the registry/catalog (Option B, not Option A).

Live check (required before merge)

This PR is hold-vps labeled and must not be merged until the following is run against a real
Gemini account and the results (pass/fail per model, plus the corrected selectors if any guess
missed) are recorded here or in a follow-up commit:

  1. Configure one valid Gemini Web cookie connection on a build running this branch.
  2. For each of the three advertised models, send one short request
    (POST /v1/chat/completions with model: "gweb/<id>", e.g. "What is 2+2?"):
    • gweb/gemini-3.1-pro — expect a normal 200 response; no UI interaction should be visible
      (Playwright trace/screenshot should show no mode-switcher click).
    • gweb/gemini-3.7-flash — expect the mode-switcher control to be clicked and the account's
      active-mode indicator to read a Flash-identifying string. PASS: 200 with the answer,
      and the indicator/screenshot confirms Flash was actually active when the prompt was
      submitted. FAIL: 400 unsupported_control_for_provider (selector needs correcting), or a
      200 where the indicator does NOT show Flash (this would mean the read-back check itself is
      unreliable and must be fixed, not loosened).
    • gweb/gemini-3.1-flash-lite — same as above, expecting a Flash-Lite-identifying indicator.
  3. On an account known to expose Extended Thinking, send one request with
    reasoning_effort: "high" against gweb/gemini-3.1-pro. PASS: the Deep Think / Extended
    Thinking control is clicked, the indicator confirms it is active, and the response proceeds.
    FAIL: 400 (selector needs correcting) or a 200 with no confirmation that the toggle was
    actually engaged.
  4. On an account WITHOUT Extended Thinking (or with reasoning_effort omitted), confirm ordinary
    requests are completely unaffected (no selector lookups attempted, matching the
    reasoning_effort: "none"/"minimal" unit test).
  5. Update GEMINI_WEB_MODEL_MODES / GEMINI_WEB_EXTENDED_THINKING_MODE in
    open-sse/executors/gemini-web/modeSelection.ts with the selectors that actually matched, and
    remove the "UNVALIDATED" framing from the header comment once confirmed.

Not covered here

  • The catalog itself (gemini-3.1-pro / gemini-3.7-flash / gemini-3.1-flash-lite) is
    unchanged — the issue also flagged these names as possibly stale vs. the reporter's live UI
    (gemini-3.8-flash / gemini-3.5-flash-lite); a catalog rename is a separate, narrower change
    (see the closed feat(providers): refresh Gemini Web model catalog #12598) and was out of scope for the owner's Option B decision, which is about
    selection correctness, not renaming.
  • The mode-selection/Extended Thinking selectors are unverified against a live account (see
    "Selector set is UNVALIDATED" and "Live check" above) — this is why the issue stays open
    (Refs, not Closes) until the live smoke passes.

… before answering (#13381)

Root cause: GeminiWebExecutor.execute() opened the identical fixed
https://gemini.google.com/app URL and ran the identical Playwright
interaction sequence for every advertised gweb/<model> id. `model` was
read only AFTER the response was captured, purely to stamp the
OpenAI-shaped response — never to influence what was actually
clicked/typed, so two different advertised models produced
byte-identical automation and the response `model` field was a
caller-supplied label, not an observed fact.

Fix (owner decision, Option B): a new model -> Gemini UI mode map
(open-sse/executors/gemini-web/modeSelection.ts) drives an in-browser
selection step before anything is typed — try the mode control, read
back the active-mode indicator, and only proceed on a confirmed match.
An unconfirmed model, or a requested Extended Thinking control that
cannot be confirmed (#13381 follow-up comment), fails closed with 400
unsupported_control_for_provider instead of silently running the
account default under the requested label. The selectors involved are
UNVALIDATED (no live Gemini account from this checkout) — see the PR's
"Selector set is UNVALIDATED" section and the required live smoke.

Regression test: tests/unit/issue-13381-gemini-web-model-selection.test.ts
@diegosouzapw diegosouzapw added the hold-vps PR verde, merge aguardando validação live (release-drain) label Sep 16, 2026
@diegosouzapw
diegosouzapw merged commit ceafa55 into release/v3.8.51 Sep 17, 2026
17 of 21 checks passed
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
… before answering (diegosouzapw#13381) (diegosouzapw#13919)

Root cause: GeminiWebExecutor.execute() opened the identical fixed
https://gemini.google.com/app URL and ran the identical Playwright
interaction sequence for every advertised gweb/<model> id. `model` was
read only AFTER the response was captured, purely to stamp the
OpenAI-shaped response — never to influence what was actually
clicked/typed, so two different advertised models produced
byte-identical automation and the response `model` field was a
caller-supplied label, not an observed fact.

Fix (owner decision, Option B): a new model -> Gemini UI mode map
(open-sse/executors/gemini-web/modeSelection.ts) drives an in-browser
selection step before anything is typed — try the mode control, read
back the active-mode indicator, and only proceed on a confirmed match.
An unconfirmed model, or a requested Extended Thinking control that
cannot be confirmed (diegosouzapw#13381 follow-up comment), fails closed with 400
unsupported_control_for_provider instead of silently running the
account default under the requested label. The selectors involved are
UNVALIDATED (no live Gemini account from this checkout) — see the PR's
"Selector set is UNVALIDATED" section and the required live smoke.

Regression test: tests/unit/issue-13381-gemini-web-model-selection.test.ts
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

hold-vps PR verde, merge aguardando validação live (release-drain)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant