Repository navigation
fix(google_genai): forward response schema and tool parameters through the generateContent adapter - #42067
Conversation
…h the generateContent adapter
|
I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".
|
|
| @@ -582,14 +650,6 @@ def translate_completion_to_generate_content( | |||
| ), | |||
There was a problem hiding this comment.
Response field removed Both response modes drop
text, breaking consumers without the required compatibility flag; updated tests now mask the regression by requiring its absence
Rule Used: What: avoid backwards-incompatible changes without user-controlled flags Why: This breaks current behaviour for users using existing functionality Example of BAD: this PR (#22164) introduced run_post_custom... (source)
There was a problem hiding this comment.
The Gemini API defines no top-level text field; the SDK computes .text client-side from candidates parts, so nothing consumes it
There was a problem hiding this comment.
You're right. The top-level text field is not part of the Gemini generateContent response contract, and the google-genai SDK derives .text from candidates[].content.parts[]. Removing this adapter-only convenience field therefore does not break the supported Gemini response shape or SDK behavior. I’m withdrawing this finding.
| return None | ||
| return { | ||
| "type": "json_schema", | ||
| "json_schema": {"name": "response", "schema": _strip_gemini_only_schema_keys(schema)}, |
There was a problem hiding this comment.
Schema enforcement weakened Google response schemas reach the provider without
strict: true, so downstream responses may violate the caller's requested shape
| "json_schema": {"name": "response", "schema": _strip_gemini_only_schema_keys(schema)}, | |
| "json_schema": { | |
| "name": "response", | |
| "schema": _strip_gemini_only_schema_keys(schema), | |
| "strict": True, | |
| }, |
There was a problem hiding this comment.
OpenAI strict mode rejects schemas lacking additionalProperties false or every property required, which Gemini dialect schemas never carry; strict would break working calls
There was a problem hiding this comment.
You're right. OpenAI strict structured outputs require additionalProperties: false and every object property to be listed in required (including nested objects). Gemini schemas do not necessarily include those constraints, so adding strict: true here could turn currently valid requests into 400s. Since this adapter does not rewrite schemas into OpenAI's strict subset, strict should remain unset. I'm withdrawing this comment; no change is needed.
| schema: Final = _validated( | ||
| _JSON_OBJECT_SCHEMA, normalize_json_schema_types(_first_present(fields, _RESPONSE_SCHEMA_KEYS)) | ||
| ) | ||
| if schema is None or schema.get("type") != "object": |
There was a problem hiding this comment.
OpenAI requires an object root, so wrapping needs an unwrap on both response paths; non-object roots keep merge-base behavior and are tracked separately
There was a problem hiding this comment.
Understood. Since OpenAI requires an object-root response_format, forwarding array/string/enum roots would require wrapping them and then unwrapping the model output on both the non-streaming and streaming response paths, which is outside this focused fix. Keeping those roots on merge-base behavior and tracking them separately is reasonable; I’m resolving this comment.
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 5eb967d. Configure here.
TLDR
Problem this solves:
generateContenton a non-Gemini deployment ignoredresponse_schemaandresponse_mime_typefunctionDeclarations[].parametersnever reached the model, onlyparametersJsonSchemadidtextfield next tocandidatesHow it solves it:
response_schemabecomes ajson_schemaresponse_formaton the deploymentpropertyOrderingkeys are stripped and types lowercased so Anthropic and OpenAI accept itresponse_format(e.g.openai/gpt-4)parametersis read whenparametersJsonSchemais absent, and a declaration that is not a JSON object is dropped as before instead of being forwardedtextis dropped from non-streaming and streaming responsesUser Flow
Before: a developer whose app sends Gemini-native
generateContentrequests to a deployment backed by Claude (or OpenAI) gets markdown prose instead of the JSON they asked for, and the model never sees the tools' parametersPOST https://litellm-domain/v1beta/models/claude-behind-native-route:generateContent(raw REST or the google-genai SDK withbase_urlpointed at the gateway) withtools[0].functionDeclarations[].parametersdeclaringpark_idanddate,generationConfig.response_mime_type: "application/json"and aresponse_schemacandidates[0].content.parts[0].textis markdown prose (**EPCOT in a nutshell**...), the SDK'sresponse.parsedisNone, and the body carries a non-standard top-leveltextfield repeating the prose next tocandidatesandusageMetadatatoolConfig.functionCallingConfig.mode: "ANY"to force apark_hours_lookupcall, and the returnedfunctionCall.argsare guessed names (park,date) because the declaredpark_idnever reached the modelPOST https://litellm-domain/v1beta/models/claude-behind-native-route:streamGenerateContent?alt=sse, and every SSE chunk is prose and carries the same extra top-leveltextAfter: the same requests come back as JSON matching the schema, tool calls use the declared parameter names, and the body is the standard Gemini shape
POST https://litellm-domain/v1beta/models/claude-behind-native-route:generateContentwith the same tools andgenerationConfigcandidates[0].content.parts[0].textis a JSON document matchingresponse_schema, the SDK'sresponse.parsedis populated, and the body carries onlycandidatesandusageMetadata, no top-leveltextpark_hours_lookupwithargskeyedpark_idanddate, exactly as declared:streamGenerateContent?alt=ssechunks concatenate to that same JSON document and carry no top-leveltextRelevant issues
Supersedes #37568 and #35921, which each carried the
parametershalf of this fixAffected release
Linear ticket
Resolves LIT-7160
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Both legs boot the proxy the same way, with 2 uvicorn workers, DB-less, against the real Anthropic, OpenAI, and Gemini APIs. Before runs the merge base f49fd22 on port 28639, After runs this PR's tip on port 56558. Long
textvalues in the outputs are cut at[...]proxy_config.yamlpayload.json(two tools declared withparameters, a Pydantic-styleresponse_schemawith$defs){ "contents": [{"role": "user", "parts": [{"text": "Summarize EPCOT for a family with two kids ages 6 and 9. Mention one must-do and one tip. Do not call tools."}]}], "tools": [{"functionDeclarations": [ { "name": "park_hours_lookup", "description": "Look up park hours for a given date and park.", "parameters": { "type": "object", "properties": { "park_id": {"type": "string", "description": "Stable park identifier, e.g. epcot"}, "date": {"type": "string", "description": "ISO-8601 date, e.g. 2026-05-01"} }, "required": ["park_id", "date"] } }, { "name": "dining_availability_hint", "description": "Returns a short hint about dining availability patterns (not a booking).", "parameters": { "type": "object", "properties": { "party_size": {"nullable": true, "type": "integer"}, "meal_period": {"enum": ["breakfast", "lunch", "dinner", "snack"], "type": "string"}, "preferences": {"items": {"type": "string"}, "nullable": true, "type": "array"} }, "required": ["meal_period"] } } ]}], "generationConfig": { "response_mime_type": "application/json", "response_schema": { "$defs": { "Highlight": { "type": "object", "title": "Highlight", "description": "A single bullet highlight.", "properties": { "title": {"type": "string", "title": "Title"}, "detail": {"anyOf": [{"type": "string"}, {"type": "null"}], "default": null, "title": "Detail"} }, "required": ["title"] } }, "type": "object", "title": "ParkTipResponse", "description": "Structured park tip for testing tools + JSON schema.", "properties": { "park_name": {"type": "string", "title": "Park Name"}, "summary": {"type": "string", "title": "Summary"}, "highlights": {"type": "array", "title": "Highlights", "items": {"$ref": "#/$defs/Highlight"}}, "confidence": {"type": "string", "enum": ["low", "medium", "high"], "title": "Confidence"} }, "required": ["park_name", "summary", "highlights", "confidence"] } } }payload_tools_off.jsonispayload.jsonplus"toolConfig": {"functionCallingConfig": {"mode": "NONE"}}payload_tool_call.json{ "contents": [{"role": "user", "parts": [{"text": "What are EPCOT's park hours on 2026-10-01? Call park_hours_lookup."}]}], "tools": [{"functionDeclarations": [{ "name": "park_hours_lookup", "description": "Look up park hours for a given date and park.", "parameters": { "type": "object", "properties": { "park_id": {"type": "string", "description": "Stable park identifier, e.g. epcot"}, "date": {"type": "string", "description": "ISO-8601 date, e.g. 2026-05-01"} }, "required": ["park_id", "date"] } }]}], "toolConfig": {"functionCallingConfig": {"mode": "ANY"}} }payload_array_root.json{ "contents": [{"role": "user", "parts": [{"text": "List three EPCOT attractions for kids ages 6 and 9."}]}], "generationConfig": {"responseMimeType": "application/json", "responseSchema": {"type": "ARRAY", "items": {"type": "STRING"}}} }payload_mime_only.json{ "contents": [{"role": "user", "parts": [{"text": "Name one EPCOT attraction for kids ages 6 and 9."}]}], "generationConfig": {"responseMimeType": "application/json"} }sdk_leg.py(google-genai 1.37.0, the client pointed at the proxy)Before (f49fd22)
Structured output over REST, Claude deployment
textnext tocandidatesandusageMetadataForced tool call over REST, Claude deployment
parkbecause the declaredpark_idnever reached itgoogle-genai SDK, Claude deployment
response.parsedisNoneand the text is proseStreaming over REST, Claude deployment
textkey and the chunks add up to proseStructured output over REST, OpenAI deployment
texton OpenAIControl, Gemini deployment on the same route
Leftover path, schema on a deployment that rejects
response_format(openai/gpt-4)Leftover path, array-root schema on the OpenAI deployment
Leftover path, JSON mime type with no schema on the OpenAI deployment
google-genai SDK, OpenAI deployment
response.parsedisNoneand the text is proseAfter (5eb967d)
Structured output over REST, Claude deployment
candidatesandusageMetadataForced tool call over REST, Claude deployment
park_idanddategoogle-genai SDK, Claude deployment
response.parsedis a populatedParkTipResponseStreaming over REST, Claude deployment
textand the chunks add up to the JSON documentStructured output over REST, OpenAI deployment
Control, Gemini deployment on the same route
Leftover path, schema on a deployment that rejects
response_format(openai/gpt-4)response_formatLeftover path, array-root schema on the OpenAI deployment
Leftover path, JSON mime type with no schema on the OpenAI deployment
google-genai SDK, OpenAI deployment
response.parsedis a populatedParkTipResponseType
🐛 Bug Fix
Caveats (if any)
Low
nullable: truepasses through and Anthropic ignores it, so a required nullable field cannot come backnullthere. Left as is: rewriting it toanyOfwithnullis a Gemini-to-JSON-Schema translation layer this fix does not need, and OpenAI honors the keyapplication/jsonmaps;text/x.enumand other mime types leaveresponse_formatunset, which is the merge-base behavior for themmock_responseon this route still returns the top-leveltext(litellm/google_genai/main.py). Left as is: nothing asserts on it, no real request reaches it, and a commit now resets both bot verdicts plus the per-commit QA and live risk legs for a mock-only fieldintegration-extensions,logging_testing,integration-cost) are red here and on everymainpipeline since 2026-09-19 23:48Z (anmcp2.x import, a GCS pub/sub logging test, and a Fireworks cache-read price); none touches this adapter. Fixing them here would mean mergingmainin or patching unrelated tests, which resets the bot verdicts and the per-commit QA for checks the merge gate does not readproxy_e2e_anthropic_messages_testsjob is red on the two Bedrocktest_all_beta_headerscases ("invalid beta flag"), which failed on everymainpipeline until fix(test): unbreak the integration-cost and proxy_e2e_anthropic_messages CircleCI jobs on main #42048 (7966f50, LIT-8149) landed after this branch's merge base; the latestmainpipeline passes them and neither case touches this adapterDesign decisions
json_schemaname is the fixedresponseandstrictis not set, so a schema passes without every property being required and withoutadditionalProperties: falseresponseJsonSchemawins overresponseSchemawhen both are set, since it is the caller's JSON Schema as written; the snake_case spellings are accepted because the google-genai SDK sends them. The picked schema is validated as a JSON object with Pydantic before it is forwarded, and anything else (a config object from the Python SDK, a schema that is not a mapping) falls through to the pre-fix prose behavior instead of failing the requestpropertyOrderingandproperty_orderingare stripped, at every nesting level, because Anthropic rejects them and OpenAI ignores them;nullable,format,example,title, anddefaultpass through because both providers accept them, so no Gemini-to-JSON-Schema rewrite is addedresponse_formatas supported, soopenai/gpt-4keeps returning prose instead of a 400; a provider that cannot be resolved fails open and sends the schemaparametersJsonSchemawins overparameters, a null value falls through to the other key, and a value that is not a JSON object (an int, a string) is dropped as the merge base did, since forwarding it turned a 200 into a 500 in the live drivetextis removed rather than kept, since the Gemini API has no such field, nothing in the repo reads it, and the SDK computes.textclient-sideLive PR risk
Base f49fd22 and tip 5eb967d each ran as a real proxy with Claude, OpenAI, and Gemini deployments behind the native route, twelve REST cells plus the
litellm.agenerate_contentSDK entrypoint, with the outbound provider request captured from the proxy log and diffed per cellparametersis not a JSON object (5,"") from a 200 into a 500, since OpenAI rejected the body and the Anthropic transformation raised; the tip validates the declaration as a JSON object first and drops anything else exactly as the base didtextis gone from every response and stream chunk (nothing in litellm, the dashboard, or the docs reads it, and the SDK computes.textclient-side), and the provider now receivestools[].parameters/input_schemaandtext.format/output_format, which is the fixOBJECT,STRING,propertyOrdering,nullable) are accepted by both providers and the forced call comes back with the declaredpark_id; a non-object, null, or 5 KB stringresponseSchema, a non-stringresponseMimeType, and conflicting camelCase and snake_case keys (camelCase wins) all keep the base behaviorlitellm/utils.py(function setup, wrapped in try/except), the v3 parallel request limiter (token estimation, fails open), the/v1betaendpoints, andlitellm.generate_content; the legacy suites that reach it (tests/proxy_unit_tests/test_google_endpoint_routing.py,test_google_gemini_proxy_request.py,tests/llm_translation/test_openai.py::test_openai_via_gemini_streaming_bridge) pass at the tip;mainmoved 81 commits since the merge base, none on the adapter or its neighbors, and the merge is cleantext/x.enumand other non-JSON mime types, and a deployment reached through a wildcard or model group aliasFinal Attestation