feat: Gemini provider + OpenAI context window guard - #15
Merged
Merged
Conversation
Member
|
please fix conflicts @gnanam1990 |
Collaborator
Author
|
i will check @kevincodex1 |
gnanam1990
force-pushed
the
feat/gemini-provider
branch
from
April 1, 2026 11:43
658e51e to
0578543
Compare
Adds Google Gemini as a first-class provider using Gemini's OpenAI-compatible endpoint, supporting gemini-2.0-flash, gemini-2.5-pro, and gemini-2.0-flash-lite across all three model tiers (opus/sonnet/haiku). - Add 'gemini' to APIProvider type with CLAUDE_CODE_USE_GEMINI env detection - Map all 11 model configs to appropriate Gemini models per tier - Route Gemini through existing OpenAI shim (generativelanguage.googleapis.com) - Support GEMINI_API_KEY and GOOGLE_API_KEY for authentication - Fix model display name to show actual Gemini model instead of Claude fallback - Add Gemini support to provider-launch, provider-bootstrap, system-check scripts - Add dev:gemini npm script for local development Bootstrap: bun run profile:init -- --provider gemini --api-key <key> Launch: bun run dev:gemini Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Without this fix, getContextWindowForModel() returns 200k for all OpenAI
models (the Claude default), causing two problems:
1. Auto-compact/warnings trigger at wrong thresholds (200k instead of 128k)
2. getModelMaxOutputTokens() returns 32k causing 400 errors from APIs that
cap output tokens lower (gpt-4o supports max 16384)
Fix:
- Add openaiContextWindows.ts with known context window sizes and max output
token limits for 30+ OpenAI-compatible models (OpenAI, DeepSeek, Groq,
Mistral, Ollama, LM Studio)
- Hook into getContextWindowForModel() so correct input limits are used
- Hook into getModelMaxOutputTokens() so correct output limits are sent,
preventing 400 "max_tokens is too large" errors
All existing warning, blocking, and auto-compact infrastructure works
automatically once the correct limits are returned.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
gnanam1990
force-pushed
the
feat/gemini-provider
branch
from
April 1, 2026 12:12
0578543 to
4ca94b2
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Two features on this branch:
1. Native Gemini Provider (
CLAUDE_CODE_USE_GEMINI=1)Adds Google Gemini as a first-class provider using Gemini's OpenAI-compatible endpoint — no OpenRouter needed, just a free API key from https://aistudio.google.com/apikey.
gemini-2.5-pro-preview-03-25, sonnet →gemini-2.0-flash, haiku →gemini-2.0-flash-litegenerativelanguage.googleapis.com/v1beta/openai)GEMINI_API_KEYandGOOGLE_API_KEYprovider-launch,provider-bootstrap,system-checksupportdev:geminiscript2. Context Window Guard for OpenAI-compatible Models
Fixes two bugs when using
CLAUDE_CODE_USE_OPENAI=1:Bug 1 — Wrong thresholds:
getContextWindowForModel()returned 200k (Claude default) for all OpenAI models. Forgpt-4o(128k context), auto-compact and warnings never fired — users hit hard API errors without warning.Bug 2 — 400 on every message:
getModelMaxOutputTokens()returned 32k default.gpt-4ocaps at 16,384 output tokens, so every request failed:Fix: New
src/utils/model/openaiContextWindows.ts— lookup table of context window sizes and max output token limits for 30+ models (OpenAI, DeepSeek, Groq, Mistral, Ollama, LM Studio). Hooked intogetContextWindowForModel()andgetModelMaxOutputTokens(). All existing warning, blocking, and auto-compact infrastructure works automatically.Tested:
bun run dev:openaiwithgpt-4o— clean response, no 400 error.