fix: resolve Gemini 400 error and support 2M context models - #643
Gustavo-Falci wants to merge 5 commits into
Conversation
Remove `store: false` payload parameter for Gemini in openaiShim.ts. The strict validation of the Gemini API rejected requests containing this unknown parameter, resulting in an immediate 400 Bad Request error. Add missing Gemini model entries (gemini-3.1-pro-preview, gemini-3-flash-preview) at 2,000,000 context window and 65,536 max output tokens. Without these entries, the models fell back to the conservative 8k default, which triggered an immediate and infinite auto-compact loop despite their actual 2M token capacity.
Remove `store: false` payload parameter for Gemini in openaiShim.ts. The strict validation of the Gemini API rejected requests containing this unknown parameter, resulting in an immediate 400 Bad Request error.
There was a problem hiding this comment.
Pull request overview
Fixes Gemini OpenAI-compatible request failures and updates model limits so Gemini 3.x preview models can use their intended large context windows without triggering compaction misbehavior.
Changes:
- Removed the unsupported
storefield from the OpenAI chat-completions request body inopenaiShim. - Added Gemini 3.x preview models to OpenAI-compatible context window and max-output-token lookup tables.
- Removed
storefrom the Codex/responsespayload.
Reviewed changes
Copilot reviewed 3 out of 3 changed files in this pull request and generated 5 comments.
| File | Description |
|---|---|
src/utils/model/openaiContextWindows.ts |
Adds context window and max output token mappings for gemini-3-flash-preview and gemini-3.1-pro-preview. |
src/services/api/openaiShim.ts |
Removes store from the chat-completions request payload to satisfy strict-schema providers (Gemini). |
src/services/api/codexShim.ts |
Removes store from the Codex /responses payload. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| 'gemini-2.0-flash': 1_048_576, | ||
| 'gemini-2.5-pro': 1_048_576, | ||
| 'gemini-2.5-flash': 1_048_576, | ||
| 'gemini-3-flash-preview': 2_000_000, | ||
| 'gemini-3.1-pro-preview': 2_000_000, |
| 'gemini-2.0-flash': 8_192, | ||
| 'gemini-2.5-pro': 65_536, | ||
| 'gemini-2.5-flash': 65_536, | ||
| 'gemini-3-flash-preview': 65_536, | ||
| 'gemini-3.1-pro-preview': 65_536, |
| 'gemini-3-flash-preview': 2_000_000, | ||
| 'gemini-3.1-pro-preview': 2_000_000, |
| const body: Record<string, unknown> = { | ||
| model: request.resolvedModel, | ||
| messages: openaiMessages, | ||
| stream: params.stream ?? false, | ||
| store: false, | ||
| } |
| content: [{ type: 'input_text', text: '' }], | ||
| }, | ||
| ], | ||
| store: false, | ||
| stream: true, | ||
| } |
Vasanthdev2004
left a comment
There was a problem hiding this comment.
Review: PR #643 — Fix Gemini 400 error and support 2M context models
Reviewed on head b55f36d. CI green ✅. kevincodex1 approved. 3 files, +4/-2.
✅ store: false removal — correct
The store parameter is OpenAI-specific (controls whether conversations are saved in ChatGPT). Gemini's strict JSON schema validation rejects unknown fields with 400 Bad Request. Removing it from the chat-completions body is the right fix — simpler than adding provider-specific stripping logic, and store defaults to false on OpenAI anyway.
The leftover delete body.store for Mistral in openaiShim.ts:1230 is now a no-op (since store is never set) but harmless — not worth touching in this PR.
The /responses endpoint fallback (line 1375) still has store: false, but that code path is GitHub Copilot only (isGithub), which is an OpenAI-compatible endpoint that accepts the parameter. Not a concern for Gemini.
The codexShim.ts removal is consistent — Codex's /responses body also shouldn't send store: false to non-OpenAI providers.
✅ Context window entries — correct values
gemini-3-flash-preview and gemini-3.1-pro-preview at 2,000,000 context / 65,536 max output match Google's published specs.
🔧 Missing: provider-qualified "google/" entries
The existing table has "google/gemini-2.0-flash": 1_048_576 and "google/gemini-2.5-pro": 1_048_576, but the new entries lack the "google/" qualified versions:
google/gemini-3-flash-preview→ 2,000,000 / 65,536google/gemini-3.1-pro-preview→ 2,000,000 / 65,536
Without these, users selecting Gemini models via the google/ prefix will still hit the 128k fallback and the infinite auto-compact loop (same root cause as Issue #635 that PR #636 fixes). This is the same class of bug the PR is trying to fix — it should be complete.
🟡 Nit: PR description
The PR description only mentions the openaiShim.ts change, but the diff also removes store: false from codexShim.ts. Minor, but worth mentioning for completeness.
🟡 Nit: test coverage
No test coverage for the new model entries. Not blocking — this is a data table update and the values match published specs — but adding a couple of assertions in context.test.ts for gemini-3-flash-preview / gemini-3.1-pro-preview (and their google/ qualified forms) would prevent regressions.
Verdict: Needs changes 🔧
One blocker: add "google/gemini-3-flash-preview" and "google/gemini-3.1-pro-preview" entries to both OPENAI_CONTEXT_WINDOWS and OPENAI_MAX_OUTPUT_TOKENS. Without them, provider-qualified model selection still falls back to 128k, which is the exact bug this PR is fixing.
Add explicit test cases for `gemini-3.1-pro-preview` and `gemini-3-flash-preview` to verify that their context window (2,000,000) and max output tokens (65,536) are correctly resolved. Also includes assertions for their `google/` prefixed variants to ensure OpenRouter compatibility is maintained and protected against regressions.
There was a problem hiding this comment.
Pull request overview
Updates the OpenAI-compatible shims and model-limit tables to work with newer Gemini 3.x models by avoiding schema-invalid request fields and by correctly configuring very large context windows/output caps to prevent auto-compact loops.
Changes:
- Remove the
store: falsefield from the OpenAI shim chat-completions request payload (fixes Gemini strict JSON 400s). - Add explicit 2,000,000-context + 65,536-max-output mappings for Gemini 3.x (native and
google/-prefixed) in the OpenAI-compatible limits tables. - Add a focused unit test covering Gemini 3.x context/output cap resolution.
Reviewed changes
Copilot reviewed 4 out of 4 changed files in this pull request and generated 3 comments.
| File | Description |
|---|---|
| src/utils/model/openaiContextWindows.ts | Adds Gemini 3.x context window + max output token mappings for native and OpenRouter-style model IDs. |
| src/utils/context.test.ts | Adds unit coverage for Gemini 3.x model cap resolution paths. |
| src/services/api/openaiShim.ts | Removes store: false from chat-completions payload to avoid Gemini request rejection. |
| src/services/api/codexShim.ts | Removes store: false from /responses payload construction. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| 'google/gemini-2.5-pro': 1_048_576, | ||
| 'google/gemini-2.0-flash': 1_048_576, | ||
| 'google/gemini-2.5-pro': 1_048_576, | ||
| 'google/gemini-3-flash-preview': 2_000_000, |
| // gemini-3-flash-preview | ||
| expect(getContextWindowForModel('gemini-3-flash-preview')).toBe(2_000_000) | ||
| expect(getModelMaxOutputTokens('gemini-3-flash-preview')).toEqual({ | ||
| default: 65_536, | ||
| upperLimit: 65_536, | ||
| }) | ||
|
|
||
| // google/ prefixed versions | ||
| expect(getContextWindowForModel('google/gemini-3.1-pro-preview')).toBe(2_000_000) | ||
| expect(getContextWindowForModel('google/gemini-3-flash-preview')).toBe(2_000_000) | ||
| expect(getModelMaxOutputTokens('google/gemini-3.1-pro-preview')).toEqual({ |
| }, | ||
| ], | ||
| store: false, | ||
| stream: true, |
There was a problem hiding this comment.
Thanks for the thorough review, @kevincodex1 @Vasanthdev2004 ! I've just pushed a new commit addressing all your points:
Blocker resolved: Added the google/gemini-3-flash-preview and google/gemini-3.1-pro-preview entries to both OPENAI_CONTEXT_WINDOWS and OPENAI_MAX_OUTPUT_TOKENS. OpenRouter users will now correctly get the 2M context window instead of the 128k fallback.
Test coverage added (Nit): Added explicit assertions in context.test.ts for both the native and google/ prefixed variants to prevent future regressions.
PR Description updated (Nit): Updated the summary to explicitly mention the codexShim.ts cleanup, along with the new prefixes and tests.
Everything should be green and good to go. Let me know if you need anything else!
|
@Vasanthdev2004 please have a look when you can bro |
Vasanthdev2004
left a comment
There was a problem hiding this comment.
Re-review: PR #643 — Fix Gemini 400 error and support 2M context models (head db9a975)
CI green ✅. 4 files, +40/-6. Previously requested changes on head b55f36d.
Blocker fix check
1. Missing google/ prefixed context window entries ✅ Fixed
Both google/gemini-3-flash-preview and google/gemini-3.1-pro-preview now have 2,000,000 context entries in OPENAI_CONTEXT_WINDOWS.
2. Missing google/ prefixed max_output_tokens entries ✅ Fixed
Both google/ prefixed entries now have 65,536 max output in OPENAI_MAX_OUTPUT_TOKENS.
New additions (since last review)
3. Test coverage ✅
New test in context.test.ts verifies:
gemini-3.1-pro-preview: 2M context, 65,536 output ✅gemini-3-flash-preview: 2M context, 65,536 output ✅google/gemini-3.1-pro-preview: 2M context ✅google/gemini-3-flash-preview: 2M context ✅google/gemini-3.1-pro-preview: 65,536 output ✅
Good — tests cover both bare and google/ prefixed forms, which was exactly the bug pattern (users selecting via prefixed form hit fallback without these entries).
Verdict: Approve-ready ✅
Both blockers resolved. store: false removal is correct, context window/output entries are complete for both naming forms, and tests verify the fix.
gnanam1990
left a comment
There was a problem hiding this comment.
Thanks for tightening up the context-window coverage here. The added tests for the Gemini 3.x model limits are helpful, but on the current diff the shim path still appears to retain store: false, so the PR description seems to overstate the full 400-error fix. I’d be more comfortable approving this once the actual request-shape fix is fully reflected in the code and the description matches the branch precisely.
- Strip store field from request body for local providers (Ollama, vLLM) that reject unknown JSON fields with 400 errors - Add Gemini 3.x model context windows and output token limits (gemini-3-flash-preview, gemini-3.1-pro-preview, google/ OpenRouter variants) - Preserve reasoning_content on assistant tool-call message replays for providers that require it (Kimi k2.5, DeepSeek reasoner) - Use conservative max_output_tokens fallback (4096/16384) for unknown 3P models to prevent vLLM/Ollama 400 errors from exceeding max_model_len Consolidates fixes from: #258, #268, #237, #643, #666, #677 Co-authored-by: auriti <auriti@users.noreply.github.com> Co-authored-by: Gustavo-Falci <Gustavo-Falci@users.noreply.github.com> Co-authored-by: lttlin <lttlin@users.noreply.github.com> Co-authored-by: Durannd <Durannd@users.noreply.github.com>
gnanam1990
left a comment
There was a problem hiding this comment.
Two separate changes bundled: (a) new Gemini 3.x context entries — tests included, LGTM. (b) removing store: false from openaiShim and codexShim — this does fix the Gemini 400, but it also silently flips OpenAI requests back to server-side conversation storage (OpenAI defaults store to true), which is a privacy-surface change for every OpenAI user. Could you gate the omission on provider — e.g. skip store: false when the base URL is Gemini's, keep it for OpenAI? Happy to re-review once scoped. Thanks!
jatmn
left a comment
There was a problem hiding this comment.
Findings
-
[P1] Keep
store: falsefor OpenAI-compatible requests that accept it
src/services/api/openaiShim.ts:1191
The Gemini 400 fix removesstore: falsefrom the shared OpenAI-compatible chat-completions body, so native OpenAI requests no longer opt out of server-side storage. OpenAI currently stores Chat Completions by default for new accounts, and the Responses API also defaultsstoretotrue, so this silently changes the privacy behavior for every OpenAI user while fixing Gemini. Please scope the omission to strict-schema providers like Gemini/Mistral, or otherwise keepstore: falsefor OpenAI endpoints that support it. -
[P1] Preserve the Codex Responses storage opt-out
src/services/api/codexShim.ts:494
The same regression exists on the Codex/responsespath: removingstore: falsemeans generated responses are stored by default unless the account has Zero Data Retention. This path is privacy-sensitive and had an explicit opt-out before this PR, so the Gemini compatibility fix should not drop it for providers/endpoints that accept the field.
|
Closing as abandoned |
Summary
what changed:
store: falseparameter from the request payload for Gemini models in bothsrc/services/api/openaiShim.tsandsrc/services/api/codexShim.ts.gemini-3.1-pro-previewandgemini-3-flash-preview(along with theirgoogle/prefixed variants for OpenRouter compatibility) toOPENAI_CONTEXT_WINDOWS(2,000,000) andOPENAI_MAX_OUTPUT_TOKENS(65,536) insrc/utils/model/openaiContextWindows.ts.src/utils/context.test.ts.why it changed:
storeparameter caused the API to immediately reject the request with a400 Bad Requesterror. This was fixed across both shims for consistency.google/prefixed models prevents this fallback.Impact
Testing
bun run buildbun run smokecontext.test.tsto assert that context window and max output tokens resolve correctly for both native andgoogle/prefixed 3.x models. Manually tested the CLI to verify the 400 error is gone and the 2M context registers properly.Notes
geminiprovider ->gemini-3.1-pro-preview,gemini-3-flash-preview, and their OpenRoutergoogle/variants.gemini-*inopenaiContextWindows.tsin the future to automatically grant 1M/2M context to newer model iterations without requiring a manual PR update.