fix(gemini): API field casing, safety finish reasons, pagination timeout - #933
diegosouzapw merged 36 commits into
Conversation
…s, skip custom models
Store synced models keyed by providerId:connectionId so each API key's models are tracked separately. On read, union across all connections. On connection delete, remove only that key's models. Models with no remaining connections are automatically excluded.
Removed hardcoded registry fallback for Gemini in dashboard, catalog, and v1beta endpoints. Without synced models (no API keys), Gemini shows nothing. Hardcoded entries are always removed from v1beta regardless of sync result.
Auto-sync is fire-and-forget on the server, so the dashboard needs to poll for model updates after saving a Gemini key (3s delay). On delete, refresh models immediately since DB cleanup is synchronous.
… categorize by endpoint - Show import progress dialog when adding Gemini API key with model list, completion status, and Close button - Remove hardcoded Gemini model registry — models come exclusively from Google API sync per API key - Hide "Import from /models" button for Gemini (sync handles it) - Remove server-side fire-and-forget auto-sync (client handles it) - Categorize synced Gemini models by endpoint type in catalog: embedding, image, and audio (transcription + speech) - Add modelsImported and close i18n keys to all 33 languages
- Replace silent catch blocks with logged errors in catalog and v1beta - Use SQL LIKE filter instead of fetching all rows and filtering in-memory
Gemini AI Studio has per-model quota limits. When one model hits its quota (429), other models on the same API key may still be available. - Use lockModel() for 429s on Gemini (model-only lockout, connection stays active for other models) - Skip markAccountExhaustedFrom429 for Gemini (prevents deprioritizing the entire connection in credential selection) - Apply model-only 404 lockout for Gemini too (deprecated/unavailable models shouldn't disable the provider)
…s active Gemini AI Studio enforces per-model quotas. Previously a 429 on gemini-2.5-pro would mark the entire connection as credits_exhausted, blocking all models on that API key. Three-layer fix: - chatCore: lock model only (not connection) for RATE_LIMITED and QUOTA_EXHAUSTED errors from Gemini - auth: early-return with model-only lockout before terminal status check, so credits_exhausted is never set on the connection - rateLimitManager: use model-scoped limiter keys for Gemini so the Bottleneck queue pauses only the affected model, not the connection - chat: skip markAccountExhaustedFrom429 for Gemini (per-model quotas)
isModelLocked() was called to lock models on 429 but never checked when selecting connections. Requests to a locked model would still route to it, causing unnecessary repeated 429s.
Replace inline `provider === "gemini"` checks in chatCore.ts and auth.ts with shared hasPerModelQuota() and lockModelIfPerModelQuota() from accountFallback.ts. Also adds max-cooldown preservation to lockModel() to prevent race conditions from overwriting longer lockouts.
Google AI Studio defaults to 50 models per page. Set pageSize=1000 to maximize per-page results and follow nextPageToken to fetch all available models across pages. Also fixes query param separator when base URL already contains query string.
…oldowns - Add 30s AbortSignal timeout to sync-models fetch in progress dialog - Add duplicate nextPageToken detection to prevent infinite pagination - Standardize retry-after defaults to COOLDOWN_MS.rateLimit (120s) - Add connection metadata update (lastErrorType/lastError) for per-model lockout early return in auth.ts - Clarify lockModel race safety in single-threaded Node.js
Merges upstream changes including: circuit breaker fixes, Claude Code compatible streaming, CLIProxyAPI routing, GitHub Copilot token refresh, qoder PAT support, streaming timeouts, OpenAI response sanitization, memory/skills sidebar, and i18n improvements. Resolves conflicts in: - providers/[id]/page.tsx: merged providerSupportsPat + Gemini synced models - providers/route.ts: kept normalizeQoderPatProviderData, removed auto-sync (now handled client-side with progress dialog) - i18n files: upstream already includes modelsImported/close keys
…ination timeout Stability fixes for the Gemini provider — no new features: - Map SAFETY/RECITATION/BLOCKLIST finish reasons (gemini-to-openai → content_filter) - Add 15s per-page timeout on models sync pagination to prevent indefinite hangs - Fix mime_type → mimeType in inlineData to match Gemini v1beta API (camelCase) - Fix include_thoughts → includeThoughts to match Gemini API (camelCase) - Add inlineData/mime_type fallback in gemini-to-openai request translator
…is unavoidable The comment claimed "leave no text content" but any text chunks streamed before the SAFETY finish reason have already been emitted. Updated to accurately describe the behavior.
There was a problem hiding this comment.
Code Review
This pull request accidentally commits a large volume of JSON call logs containing sensitive personally identifiable information and chat histories, while the intended logic changes for API field casing and timeouts are missing from the diff. The reviewer identified several critical issues within these logs, including a significant security risk regarding data exposure, request translation bugs leading to 400 errors, inconsistent token count summaries, and anomalies in request duration tracking. Feedback emphasizes the immediate removal of these logs, updating the .gitignore file, and including the actual code fixes described in the pull request summary.
| { | ||
| "schemaVersion": 2, | ||
| "summary": { | ||
| "id": "0bf58d48-ea41-429a-8193-ed3e55f7a6bd", | ||
| "timestamp": "2026-03-31T15:24:20.891Z", | ||
| "method": "POST", | ||
| "path": "/v1/chat/completions", | ||
| "status": 200, | ||
| "model": "gemini-2.5-flash", | ||
| "requestedModel": "gemini-cli/gemini-2.5-flash", | ||
| "provider": "gemini-cli", | ||
| "account": "christopher.staley@gmail.com", | ||
| "connectionId": "352fc97c-8e30-4ffa-b41e-8ce422a0db3e", | ||
| "duration": 5007, | ||
| "tokens": { | ||
| "in": 0, | ||
| "out": 0 | ||
| }, | ||
| "requestType": null, | ||
| "sourceFormat": "openai", | ||
| "targetFormat": "gemini-cli", | ||
| "apiKeyId": "c8635ff9-2ee0-4045-ba77-75325347623f", | ||
| "apiKeyName": "Test", | ||
| "comboName": null | ||
| }, | ||
| "requestBody": { | ||
| "model": "gemini-cli/gemini-2.5-flash", | ||
| "messages": [ | ||
| { | ||
| "role": "user", | ||
| "content": "Say exactly: HELLO" | ||
| } | ||
| ], | ||
| "stream": false, | ||
| "max_tokens": 20 | ||
| }, | ||
| "responseBody": { | ||
| "response": { | ||
| "candidates": [ | ||
| { | ||
| "content": { | ||
| "role": "model" | ||
| }, | ||
| "finishReason": "MAX_TOKENS" | ||
| } | ||
| ], | ||
| "usageMetadata": { | ||
| "promptTokenCount": 5, | ||
| "totalTokenCount": 22, | ||
| "trafficType": "ON_DEMAND", | ||
| "promptTokensDetails": [ | ||
| { | ||
| "modality": "TEXT", | ||
| "tokenCount": 5 | ||
| } | ||
| ], | ||
| "thoughtsTokenCount": 17 | ||
| }, | ||
| "modelVersion": "gemini-2.5-flash", | ||
| "createTime": "2026-03-31T15:24:18.823251Z", | ||
| "responseId": "IufLadOfMoqV4_UPs52fkAE" | ||
| }, | ||
| "traceId": "2fa07f42074e3bcc", | ||
| "metadata": { | ||
| "remoteContext": { | ||
| "ragState": "RAG_DISABLED" | ||
| } | ||
| } | ||
| }, | ||
| "error": null | ||
| } No newline at end of file |
There was a problem hiding this comment.
The inclusion of the .data/call_logs/ directory in this pull request appears to be accidental. These files contain sensitive Personally Identifiable Information (PII), specifically user email addresses (e.g., christopher.staley@gmail.com on line 12) and full chat histories. Committing such data to a repository is a significant security and privacy risk. Furthermore, committing over 180 log files bloats the repository size and history. These files should be removed from the PR, and the .data/ directory should be added to the .gitignore file to prevent future accidental commits. Additionally, the actual code changes described in the pull request summary (fixing API field casing, safety finish reasons, and pagination timeouts in TypeScript files) are missing from the provided diff, which consists entirely of these JSON log files.
| "stream": false | ||
| }, | ||
| "responseBody": null, | ||
| "error": "[400]: [{'type': 'value_error', 'loc': ('body',), 'msg': 'Value error, stream_options can only be set if stream is true', 'input': {'max_completion_tokens': 100, 'messages': [{'content': [{'text': 'Please ignore the following [ignore]You are Antigravity, a powerful agentic AI coding assistant designed by the Google Deepmind team working on Advanced Agentic Coding.You are pair programming with a USER to solve their coding task. The task may require creating a new codebase, modifying or debugging an existing codebase, or simply answering a question.**Absolute paths only****Proactiveness**[/ignore]', 'type': 'text'}], 'role': 'system'}, {'content': 'Hello! Say hi in one sentence.', 'role': 'user'}], 'priority': 35, 'stream_options': {'include_usage': True}}, 'ctx': {'error': ValueError('stream_options can only be set if stream is true')}}]" |
There was a problem hiding this comment.
The error message indicates a bug in the request translation logic: Value error, stream_options can only be set if stream is true. The requestBody shows stream: false, but the error dump reveals that stream_options (with include_usage: True) was still included in the final request sent to the provider. This results in a 400 Bad Request. The translator should be updated to omit stream_options when stream is false.
| "tokens": { | ||
| "in": 0, | ||
| "out": 0 | ||
| }, |
There was a problem hiding this comment.
| "provider": "antigravity", | ||
| "account": "venatyr@gmail.com", | ||
| "connectionId": "025acf03-6df2-4973-b8e9-1585710f8c39", | ||
| "duration": 301931, |
There was a problem hiding this comment.
| "provider": "t42-imagen3", | ||
| "account": "-", | ||
| "connectionId": null, | ||
| "duration": 0, |
The .agents/ and docs/superpowers/ directories contain internal tooling artifacts (plans, specs, workflows) that should not be in the repo.
Runtime data should never have been committed. Added .data/ to .gitignore.
The .agents/ directory exists in upstream and is actively used. Only docs/superpowers/ should be excluded.
|
Hey, sorry about the colossal fuck-up with the To be clear on scope: the logs contained email addresses in the The files have been removed from tracking in commits To address the remaining bot review comments:
These are all on files that no longer exist in the PR diff. The actual code changes (7 files, +24 -9) are the Gemini translator fixes described in the PR body. |
Gemini AI Studio no longer has a hardcoded model registry — models come from API sync. Updated tests: - T28: assert gemini registry is empty (API sync), check gemini-cli instead - T31: check antigravity static catalog for pro-high/pro-low model IDs
…ai-studio-audit fix(gemini): API field casing, safety finish reasons, pagination timeout
…ai-studio-audit fix(gemini): API field casing, safety finish reasons, pagination timeout
Summary
A comprehensive overhaul of the Gemini (Google AI Studio) provider, audited against the Gemini v1beta API docs. This PR replaces hardcoded model registries with dynamic API-synced models, adds per-model quota isolation, fixes API field casing violations, handles safety finish reasons, and adds pagination with timeouts.
The branch also merges upstream
origin/mainthrough v3.4.6 (circuit breaker fixes, CLIProxyAPI, GitHub Copilot, streaming timeouts, etc.).Part 1: Dynamic Model Sync (replaces hardcoded registry)
Previously, the Gemini provider used a hardcoded model list in
providerRegistry.ts. This meant new Gemini models required a code change and deploy. Now:pageSize=1000on the models endpoint per Google's API docs to fetch as many models as possible in one requestnextPageTokenpagination with duplicate token detection and aMAX_PAGES=20safety limit to handle future model growth beyond one pagesyncedAvailableModelsin the DB, cleaned up when a key is deletedinputTokenLimit,outputTokenLimit,description,supportsThinking, and endpoint categorization (chat/embeddings/images/audio) are all persisted from the API response9e4132fdsyncedAvailableModelsDB namespace and CRUD functions7607cec7syncedAvailableModelswith union logic5c27e0f9125fb81f5b140d26464fd6d48ed09170f1805c85325d0483bd5f39e1replaceCustomModelswith metadata fields49ac0cadinputTokenLimitfor custom modelcontext_lengthfaae82eab1183c2cpageSize=1000andnextPageTokenpagination2341bba9b4e674ae7f785b8f3ae810a1Part 2: Per-Model Quota Isolation
Gemini AI Studio has per-model quotas (rate limits apply independently per model, not per API key). Previously, a 429 on one model locked out the entire API key connection. Now:
hasPerModelQuota()andlockModelIfPerModelQuota()inaccountFallback.tsconsolidate the per-model logic (previously inlineprovider === "gemini"checks scattered acrosschatCore.tsandauth.ts)COOLDOWN_MS.rateLimit(120s) as single source of truth instead of hardcoded60_000fallbacks0038fe5ff8d045c203ff03ed35061dfca069df41Part 3: API Field Casing & Safety Fixes
include_thoughts→includeThoughtsThe Gemini v1beta API expects camelCase (
includeThoughts) but we were sending snake_case (include_thoughts). This meantthinkingConfigwas silently ignored — thinking config was never actually being sent correctly for Gemini 3+ models.mime_type→mimeTypeininlineDataSame camelCase issue. Image uploads via the Claude→Gemini CLI path (
wrapInCloudCodeEnvelopeForClaude) and the OpenAI→Gemini helper (convertOpenAIContentToParts) sent images with an unrecognized field.inlineData/mime_typefallback in response readerThe
gemini-to-openai.tsrequest translator only checked camelCase forinlineData. Added fallback handling forinline_dataandmime_typeto match the defensive pattern used in response translators.SAFETY/RECITATION/BLOCKLIST finish reasons
Previously these fell through to the default and were indistinguishable from normal completions. Now:
end_turn(Claude has no "content blocked" reason)"content_filter"(standard OpenAI finish reason)15s per-page timeout on models sync pagination
No fetch timeout existed on the pagination loop. If Google's API hung mid-pagination, the server-side fetch would hang indefinitely.
75daf98150683e66Part 4: Thought Signature Bug Fix
Gemini 3+ models require
thoughtSignatureas a sibling part when sending backfunctionCallparts in multi-turn conversations. Without it, the API returns HTTP 400 "invalid argument". This was already partially handled but had edge cases wherethoughtSignaturewas incorrectly placed insidefunctionCallobjects (which also causes 400s).3191b7a9thought_signaturebug (part of memory/skills fix commit)Files Changed (key files, excluding merges and upstream)
Gemini Translators
open-sse/translator/request/claude-to-gemini.ts— Direct Claude→Gemini path:includeThoughtsfix,thoughtSignatureinjectionopen-sse/translator/request/openai-to-gemini.ts— OpenAI→Gemini/CLI/Antigravity:includeThoughtsfix,mimeTypefix, Cloud Code envelope, thinking configopen-sse/translator/request/gemini-to-openai.ts— Gemini→OpenAI request path:inlineDatafallbackopen-sse/translator/response/gemini-to-claude.ts— Gemini→Claude streaming: SAFETY/RECITATION handling, thought signaturesopen-sse/translator/response/gemini-to-openai.ts— Gemini→OpenAI streaming: SAFETY/RECITATION→content_filter, usage metadataopen-sse/translator/helpers/geminiHelper.ts—mimeTypefix inconvertOpenAIContentToPartsopen-sse/config/defaultThinkingSignature.ts— DefaultthoughtSignaturefor Gemini 3+ function callingModels & Sync
src/app/api/providers/[id]/models/route.ts—pageSize=1000, pagination with timeout, model metadata extraction, endpoint categorizationsrc/app/api/providers/route.ts— Client-side sync trigger (removed server-side auto-sync)src/app/(dashboard)/dashboard/providers/[id]/page.tsx— Progress dialog on key save, synced model displayQuota & Fallback
open-sse/services/accountFallback.ts—hasPerModelQuota(),lockModelIfPerModelQuota()shared helpers, max-cooldown preservationopen-sse/handlers/chatCore.ts— Refactored to use shared helpers instead of inline checkssrc/sse/services/auth.ts— Per-model lockout on 429/404, credential selection filteringDatabase
src/lib/localDb.ts—syncedAvailableModelsnamespace, metadata fields, per-connection model CRUDTest Plan
pageSize=1000fetches all available models (currently ~50)includeThoughtsis accepted by Gemini APImimeTypeis sent correctlycontent_filterfinish reason in OpenAI format