Skip to content

fix(google_genai): forward response schema and tool parameters through the generateContent adapter - #42067

Merged
mateo-berri merged 3 commits into
mainfrom
litellm_genai_adapter_response_schema_tool_params
Sep 20, 2026
Merged

mateo-berri merged 3 commits into
mainfrom
litellm_genai_adapter_response_schema_tool_params

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 20, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Native generateContent on a non-Gemini deployment ignored response_schema and response_mime_type
  • Tool functionDeclarations[].parameters never reached the model, only parametersJsonSchema did
  • Responses carried a non-standard top-level text field next to candidates

How it solves it:

  • An object-root response_schema becomes a json_schema response_format on the deployment
  • Gemini-only propertyOrdering keys are stripped and types lowercased so Anthropic and OpenAI accept it
  • The schema is skipped on deployments that reject response_format (e.g. openai/gpt-4)
  • parameters is read when parametersJsonSchema is absent, and a declaration that is not a JSON object is dropped as before instead of being forwarded
  • The top-level text is dropped from non-streaming and streaming responses

User Flow

Before: a developer whose app sends Gemini-native generateContent requests to a deployment backed by Claude (or OpenAI) gets markdown prose instead of the JSON they asked for, and the model never sees the tools' parameters

  1. They send POST https://litellm-domain/v1beta/models/claude-behind-native-route:generateContent (raw REST or the google-genai SDK with base_url pointed at the gateway) with tools[0].functionDeclarations[].parameters declaring park_id and date, generationConfig.response_mime_type: "application/json" and a response_schema
  2. They get a 200 whose candidates[0].content.parts[0].text is markdown prose (**EPCOT in a nutshell**...), the SDK's response.parsed is None, and the body carries a non-standard top-level text field repeating the prose next to candidates and usageMetadata
  3. They send the same request with toolConfig.functionCallingConfig.mode: "ANY" to force a park_hours_lookup call, and the returned functionCall.args are guessed names (park, date) because the declared park_id never reached the model
  4. They send the same body to POST https://litellm-domain/v1beta/models/claude-behind-native-route:streamGenerateContent?alt=sse, and every SSE chunk is prose and carries the same extra top-level text

After: the same requests come back as JSON matching the schema, tool calls use the declared parameter names, and the body is the standard Gemini shape

  1. They send the same POST https://litellm-domain/v1beta/models/claude-behind-native-route:generateContent with the same tools and generationConfig
  2. They get a 200 whose candidates[0].content.parts[0].text is a JSON document matching response_schema, the SDK's response.parsed is populated, and the body carries only candidates and usageMetadata, no top-level text
  3. The forced tool call comes back as park_hours_lookup with args keyed park_id and date, exactly as declared
  4. The :streamGenerateContent?alt=sse chunks concatenate to that same JSON document and carry no top-level text

Relevant issues

Supersedes #37568 and #35921, which each carried the parameters half of this fix

Affected release

Linear ticket

Resolves LIT-7160

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Both legs boot the proxy the same way, with 2 uvicorn workers, DB-less, against the real Anthropic, OpenAI, and Gemini APIs. Before runs the merge base f49fd22 on port 28639, After runs this PR's tip on port 56558. Long text values in the outputs are cut at [...]

python litellm/proxy/proxy_cli.py --config proxy_config.yaml --port <port> --num_workers 2 --detailed_debug

proxy_config.yaml

model_list:
  - model_name: claude-behind-native-route
    litellm_params:
      model: anthropic/claude-opus-5
      api_key: os.environ/ANTHROPIC_API_KEY
  - model_name: openai-behind-native-route
    litellm_params:
      model: openai/gpt-5.6
      api_key: os.environ/OPENAI_API_KEY
  - model_name: openai-legacy-behind-native-route
    litellm_params:
      model: openai/gpt-4
      api_key: os.environ/OPENAI_API_KEY
  - model_name: gemini-control
    litellm_params:
      model: gemini/gemini-3.8-flash
      api_key: os.environ/GEMINI_API_KEY
general_settings:
  master_key: sk-1234

payload.json (two tools declared with parameters, a Pydantic-style response_schema with $defs)

{
  "contents": [{"role": "user", "parts": [{"text": "Summarize EPCOT for a family with two kids ages 6 and 9. Mention one must-do and one tip. Do not call tools."}]}],
  "tools": [{"functionDeclarations": [
    {
      "name": "park_hours_lookup",
      "description": "Look up park hours for a given date and park.",
      "parameters": {
        "type": "object",
        "properties": {
          "park_id": {"type": "string", "description": "Stable park identifier, e.g. epcot"},
          "date": {"type": "string", "description": "ISO-8601 date, e.g. 2026-05-01"}
        },
        "required": ["park_id", "date"]
      }
    },
    {
      "name": "dining_availability_hint",
      "description": "Returns a short hint about dining availability patterns (not a booking).",
      "parameters": {
        "type": "object",
        "properties": {
          "party_size": {"nullable": true, "type": "integer"},
          "meal_period": {"enum": ["breakfast", "lunch", "dinner", "snack"], "type": "string"},
          "preferences": {"items": {"type": "string"}, "nullable": true, "type": "array"}
        },
        "required": ["meal_period"]
      }
    }
  ]}],
  "generationConfig": {
    "response_mime_type": "application/json",
    "response_schema": {
      "$defs": {
        "Highlight": {
          "type": "object",
          "title": "Highlight",
          "description": "A single bullet highlight.",
          "properties": {
            "title": {"type": "string", "title": "Title"},
            "detail": {"anyOf": [{"type": "string"}, {"type": "null"}], "default": null, "title": "Detail"}
          },
          "required": ["title"]
        }
      },
      "type": "object",
      "title": "ParkTipResponse",
      "description": "Structured park tip for testing tools + JSON schema.",
      "properties": {
        "park_name": {"type": "string", "title": "Park Name"},
        "summary": {"type": "string", "title": "Summary"},
        "highlights": {"type": "array", "title": "Highlights", "items": {"$ref": "#/$defs/Highlight"}},
        "confidence": {"type": "string", "enum": ["low", "medium", "high"], "title": "Confidence"}
      },
      "required": ["park_name", "summary", "highlights", "confidence"]
    }
  }
}

payload_tools_off.json is payload.json plus "toolConfig": {"functionCallingConfig": {"mode": "NONE"}}

payload_tool_call.json

{
  "contents": [{"role": "user", "parts": [{"text": "What are EPCOT's park hours on 2026-10-01? Call park_hours_lookup."}]}],
  "tools": [{"functionDeclarations": [{
    "name": "park_hours_lookup",
    "description": "Look up park hours for a given date and park.",
    "parameters": {
      "type": "object",
      "properties": {
        "park_id": {"type": "string", "description": "Stable park identifier, e.g. epcot"},
        "date": {"type": "string", "description": "ISO-8601 date, e.g. 2026-05-01"}
      },
      "required": ["park_id", "date"]
    }
  }]}],
  "toolConfig": {"functionCallingConfig": {"mode": "ANY"}}
}

payload_array_root.json

{
  "contents": [{"role": "user", "parts": [{"text": "List three EPCOT attractions for kids ages 6 and 9."}]}],
  "generationConfig": {"responseMimeType": "application/json", "responseSchema": {"type": "ARRAY", "items": {"type": "STRING"}}}
}

payload_mime_only.json

{
  "contents": [{"role": "user", "parts": [{"text": "Name one EPCOT attraction for kids ages 6 and 9."}]}],
  "generationConfig": {"responseMimeType": "application/json"}
}

sdk_leg.py (google-genai 1.37.0, the client pointed at the proxy)

import json
import sys

from google import genai
from google.genai import types
from pydantic import BaseModel


class Highlight(BaseModel):
    title: str
    detail: str | None = None


class ParkTipResponse(BaseModel):
    park_name: str
    summary: str
    highlights: list[Highlight]
    confidence: str


park_hours_lookup = types.FunctionDeclaration(
    name="park_hours_lookup",
    description="Look up park hours for a given date and park.",
    parameters=types.Schema(
        type="OBJECT",
        properties={
            "park_id": types.Schema(type="STRING", description="Stable park identifier, e.g. epcot"),
            "date": types.Schema(type="STRING", description="ISO-8601 date, e.g. 2026-05-01"),
        },
        required=["park_id", "date"],
    ),
)

client = genai.Client(api_key="sk-1234", http_options={"base_url": sys.argv[1]})
response = client.models.generate_content(
    model=sys.argv[2],
    contents="Summarize EPCOT for a family with two kids ages 6 and 9. Mention one must-do and one tip. Do not call tools.",
    config=types.GenerateContentConfig(
        response_mime_type="application/json",
        response_schema=ParkTipResponse,
        tools=[types.Tool(function_declarations=[park_hours_lookup])],
    ),
)
print("parsed:", response.parsed)
print("text starts with:", repr((response.text or "")[:200]))
try:
    json.loads(response.text or "")
    print("text parses as JSON: True")
except Exception as exc:
    print("text parses as JSON: False ->", exc)

Before (f49fd22)

Structured output over REST, Claude deployment

  1. Run
curl -s http://127.0.0.1:28639/v1beta/models/claude-behind-native-route:generateContent -H 'x-goog-api-key: sk-1234' -H 'Content-Type: application/json' --data-binary @payload.json | jq '{top_level_keys: keys, text: .candidates[0].content.parts[0].text}'
  1. 200 with markdown prose, and a top-level text next to candidates and usageMetadata
{
  "top_level_keys": [
    "candidates",
    "text",
    "usageMetadata"
  ],
  "text": "**EPCOT in a nutshell (kids 6 & 9):**\n\nEPCOT is split into two halves with very different vibes. The front is **World Celebration/Discovery/Nature** — space, science, and nature rides plus the big Spaceship Earth \"golf ball.\" The back half is **World Showcase**, a ring of 11 country pavilions around a lagoon with food, shops, street performers, and a few rides tucked inside.\n\nFor your ages, the sweet [...]
}

Forced tool call over REST, Claude deployment

  1. Run
curl -s http://127.0.0.1:28639/v1beta/models/claude-behind-native-route:generateContent -H 'x-goog-api-key: sk-1234' -H 'Content-Type: application/json' --data-binary @payload_tool_call.json | jq -c '.candidates[0].content.parts[] | select(.functionCall) | .functionCall'
  1. The model guessed park because the declared park_id never reached it
{"name":"park_hours_lookup","args":{"park":"EPCOT","date":"2026-10-01"}}

google-genai SDK, Claude deployment

  1. Run
python sdk_leg.py http://127.0.0.1:28639 claude-behind-native-route
  1. response.parsed is None and the text is prose
parsed: None
text starts with: '**EPCOT at a glance (ages 6 & 9)**\n\nEPCOT is Walt Disney World\'s "edutainment" park, split into two very different halves. The front half (World Celebration, Discovery, and Nature) is space-, science-'
text parses as JSON: False -> Expecting value: line 1 column 1 (char 0)

Streaming over REST, Claude deployment

  1. Run
curl -sN 'http://127.0.0.1:28639/v1beta/models/claude-behind-native-route:streamGenerateContent?alt=sse' -H 'x-goog-api-key: sk-1234' -H 'Content-Type: application/json' --data-binary @payload.json | grep '^data: ' | sed 's/^data: //' | jq -sc '{chunk_keys: (map(keys) | add | unique), text: (map(.candidates[0].content.parts[0].text // "") | join(""))}'
  1. Every chunk carries the extra text key and the chunks add up to prose
{"chunk_keys":["candidates","text","usageMetadata"],"text":"**EPCOT in a nutshell** — It's Disney World's \"edutainment\" park, split into two halves: **World Celebration/Discovery/Nature** (space, sea, and imagination-themed attractions) and **World Showcase** (11 country pavilions circling a lagoon). It's less ride-dense than Magic Kingdom, but there's plenty for a 6- and 9-year-old, and the pace is calmer.\n\n**Go [...]

Structured output over REST, OpenAI deployment

  1. Run
curl -s http://127.0.0.1:28639/v1beta/models/openai-behind-native-route:generateContent -H 'x-goog-api-key: sk-1234' -H 'Content-Type: application/json' --data-binary @payload.json | jq '{top_level_keys: keys, text: .candidates[0].content.parts[0].text}'
  1. Same prose and extra text on OpenAI
{
  "top_level_keys": [
    "candidates",
    "text",
    "usageMetadata"
  ],
  "text": "EPCOT blends Disney attractions with science, culture, and food from around the world. Kids ages 6 and 9 will likely enjoy rides such as Remy’s Ratatouille Adventure, Frozen Ever After, and The Seas with Nemo & Friends, plus exploring the interactive World Showcase pavilions.\n\n**Must-do:** Ride **Remy’s Ratatouille Adventure**—a fun, family-friendly 3D adventure.\n\n**Tip:** EPCOT involves lots of walkin [...]
}

Control, Gemini deployment on the same route

  1. Run
curl -s http://127.0.0.1:28639/v1beta/models/gemini-control:generateContent -H 'x-goog-api-key: sk-1234' -H 'Content-Type: application/json' --data-binary @payload_tools_off.json | jq '{top_level_keys: keys, text: .candidates[0].content.parts[0].text}'
  1. Gemini itself returns JSON and the standard Gemini keys
{
  "top_level_keys": [
    "candidates",
    "modelVersion",
    "responseId",
    "usageMetadata"
  ],
  "text": "{\"park_name\":\"EPCOT\",\"summary\":\"EPCOT offers a fantastic blend of innovative rides, interactive play areas, and global culture that keeps kids ages 6 and 9 fully engaged throughout the day.\",\"highlights\":[{\"title\":\"Must-Do Attraction\",\"detail\":\"Experience Remy's Ratatouille Adventure in the France Pavilion, a trackless 4D ride that delights elementary-aged kids.\"},{\"title\":\"Helpful Tip [...]
}

Leftover path, schema on a deployment that rejects response_format (openai/gpt-4)

  1. Run
curl -s -o /tmp/out.json -w 'HTTP %{http_code}\n' http://127.0.0.1:28639/v1beta/models/openai-legacy-behind-native-route:generateContent -H 'x-goog-api-key: sk-1234' -H 'Content-Type: application/json' --data-binary @payload_tools_off.json; jq '{top_level_keys: keys, text: .candidates[0].content.parts[0].text}' /tmp/out.json
  1. 200 with prose
HTTP 200
{
  "top_level_keys": [
    "candidates",
    "text",
    "usageMetadata"
  ],
  "text": "EPCOT, also known as the Experimental Prototype Community of Tomorrow, is the perfect fusion of science, culture, and entertainment. At approximately twice the size of the Magic Kingdom, this Walt Disney World theme park is split into two major areas: Future World and World Showcase.\n\nIn Future World, your family will experience technological innovations and get a chance to explore the ocean's depth, the [...]
}

Leftover path, array-root schema on the OpenAI deployment

  1. Run
curl -s -o /tmp/out.json -w 'HTTP %{http_code}\n' http://127.0.0.1:28639/v1beta/models/openai-behind-native-route:generateContent -H 'x-goog-api-key: sk-1234' -H 'Content-Type: application/json' --data-binary @payload_array_root.json; jq '{top_level_keys: keys, text: .candidates[0].content.parts[0].text}' /tmp/out.json
  1. 200 with prose
HTTP 200
{
  "top_level_keys": [
    "candidates",
    "text",
    "usageMetadata"
  ],
  "text": "- **Remy’s Ratatouille Adventure** – A fun, trackless 3D ride through Gusteau’s kitchen; no height requirement.\n- **Frozen Ever After** – A gentle boat ride featuring Anna, Elsa, and songs from *Frozen*; no height requirement.\n- **Soarin’ Around the World** – A simulated hang-gliding adventure over famous landmarks; minimum height is **40 inches (102 cm)**."
}

Leftover path, JSON mime type with no schema on the OpenAI deployment

  1. Run
curl -s -o /tmp/out.json -w 'HTTP %{http_code}\n' http://127.0.0.1:28639/v1beta/models/openai-behind-native-route:generateContent -H 'x-goog-api-key: sk-1234' -H 'Content-Type: application/json' --data-binary @payload_mime_only.json; jq '{top_level_keys: keys, text: .candidates[0].content.parts[0].text}' /tmp/out.json
  1. 200 with prose
HTTP 200
{
  "top_level_keys": [
    "candidates",
    "text",
    "usageMetadata"
  ],
  "text": "**Remy’s Ratatouille Adventure** in the France Pavilion—family-friendly, immersive, and has no height requirement."
}

google-genai SDK, OpenAI deployment

  1. Run
python sdk_leg.py http://127.0.0.1:28639 openai-behind-native-route
  1. response.parsed is None and the text is prose
parsed: None
text starts with: 'EPCOT blends family-friendly rides, Disney characters, hands-on exhibits, and food and culture from around the world.\n\n- **Must-do:** *Remy’s Ratatouille Adventure*—a fun, immersive ride the whole fam'
text parses as JSON: False -> Expecting value: line 1 column 1 (char 0)

After (5eb967d)

Structured output over REST, Claude deployment

  1. Run
curl -s http://localhost:56558/v1beta/models/claude-behind-native-route:generateContent -H 'x-goog-api-key: sk-1234' -H 'Content-Type: application/json' --data-binary @payload.json | jq '{top_level_keys: keys, text: .candidates[0].content.parts[0].text}'
  1. 200 with a JSON document matching the schema, and only candidates and usageMetadata
{
  "top_level_keys": [
    "candidates",
    "usageMetadata"
  ],
  "text": "{\"park_name\":\"EPCOT\",\"summary\":\"EPCOT splits into two moods that work well for a 6- and 9-year-old: the future-focused World Celebration/Discovery/Nature side with big rides and hands-on science, and World Showcase, a walkable loop of 11 country pavilions with snacks, street performers, and character spots. It's a lot of walking and less thrill-dense than other parks, so plan a relaxed pace, plenty  [...]
}

Forced tool call over REST, Claude deployment

  1. Run
curl -s http://localhost:56558/v1beta/models/claude-behind-native-route:generateContent -H 'x-goog-api-key: sk-1234' -H 'Content-Type: application/json' --data-binary @payload_tool_call.json | jq -c '.candidates[0].content.parts[] | select(.functionCall) | .functionCall'
  1. The call uses the declared park_id and date
{"name":"park_hours_lookup","args":{"park_id":"epcot","date":"2026-10-01"}}

google-genai SDK, Claude deployment

  1. Run
python sdk_leg.py http://localhost:56558 claude-behind-native-route
  1. response.parsed is a populated ParkTipResponse
parsed: park_name='EPCOT' summary="EPCOT is Walt Disney World's celebration of innovation and world cultures, split into a future-focused front half (World Celebration, World Discovery, World Nature) and the World Showcase, a ring of 11 country pavilions around a lagoon. For a 6- and 9-year-old, it's a great mix: a few thrill-lite attractions, lots of hands-on play, characters from Frozen, Nemo, and Moana, and plenty [...]
text starts with: '{"park_name":"EPCOT","summary":"EPCOT is Walt Disney World\'s celebration of innovation and world cultures, split into a future-focused front half (World Celebration, World Discovery, World Nature) and'
text parses as JSON: True

Streaming over REST, Claude deployment

  1. Run
curl -sN 'http://localhost:56558/v1beta/models/claude-behind-native-route:streamGenerateContent?alt=sse' -H 'x-goog-api-key: sk-1234' -H 'Content-Type: application/json' --data-binary @payload.json | grep '^data: ' | sed 's/^data: //' | jq -sc '{chunk_keys: (map(keys) | add | unique), text: (map(.candidates[0].content.parts[0].text // "") | join(""))}'
  1. No chunk carries text and the chunks add up to the JSON document
{"chunk_keys":["candidates","usageMetadata"],"text":"{\"park_name\":\"EPCOT\",\"summary\":\"EPCOT pairs future-focused attractions in World Celebration/Discovery/Nature with a walk-around-the-world tour of 11 country pavilions in World Showcase. For a 6- and 9-year-old, the front half of the park delivers the rides and hands-on play, while World Showcase works best as a slower afternoon of snacks, street performers,  [...]

Structured output over REST, OpenAI deployment

  1. Run
curl -s http://localhost:56558/v1beta/models/openai-behind-native-route:generateContent -H 'x-goog-api-key: sk-1234' -H 'Content-Type: application/json' --data-binary @payload.json | jq '{top_level_keys: keys, text: .candidates[0].content.parts[0].text}'
  1. Same JSON document and standard keys on OpenAI
{
  "top_level_keys": [
    "candidates",
    "usageMetadata"
  ],
  "text": "{\"park_name\":\"EPCOT\",\"summary\":\"EPCOT blends family-friendly rides, interactive exhibits, global food, and character experiences. Kids ages 6 and 9 will likely enjoy exploring World Celebration, World Nature, and the country pavilions around World Showcase.\",\"highlights\":[{\"title\":\"Must-do\",\"detail\":\"Ride Remy’s Ratatouille Adventure, a playful, trackless 4D attraction that shrinks your fa [...]
}

Control, Gemini deployment on the same route

  1. Run
curl -s http://localhost:56558/v1beta/models/gemini-control:generateContent -H 'x-goog-api-key: sk-1234' -H 'Content-Type: application/json' --data-binary @payload_tools_off.json | jq '{top_level_keys: keys, text: .candidates[0].content.parts[0].text}'
  1. Unchanged, Gemini never goes through this adapter
{
  "top_level_keys": [
    "candidates",
    "modelVersion",
    "responseId",
    "usageMetadata"
  ],
  "text": "{\"park_name\":\"EPCOT\",\"summary\":\"EPCOT offers an exciting mix of futuristic discovery and global culture with plenty of interactive experiences tailored for 6- and 9-year-olds.\",\"highlights\":[{\"title\":\"Must-Do Attraction\",\"detail\":\"Remy's Ratatouille Adventure in the France Pavilion is a trackless 4D experience beloved by kids in this age bracket.\"},{\"title\":\"Top Tip\",\"detail\":\"Take [...]
}

Leftover path, schema on a deployment that rejects response_format (openai/gpt-4)

  1. Run
curl -s -o /tmp/out.json -w 'HTTP %{http_code}\n' http://localhost:56558/v1beta/models/openai-legacy-behind-native-route:generateContent -H 'x-goog-api-key: sk-1234' -H 'Content-Type: application/json' --data-binary @payload_tools_off.json; jq '{top_level_keys: keys, text: .candidates[0].content.parts[0].text}' /tmp/out.json
  1. Still 200 with prose, the schema is not sent to a model that rejects response_format
HTTP 200
{
  "top_level_keys": [
    "candidates",
    "usageMetadata"
  ],
  "text": "EPCOT, short for Experimental Prototype Community of Tomorrow, is a unique theme park in the Walt Disney World Resort that is all about discovering, amazement, and inspiration. It's divided into two main sections, Future World and the World Showcase. \n\nFuture World is teeming with thrilling attractions based on innovative and futuristic concepts. The Spaceship Earth is an iconic ride that takes guests on [...]
}

Leftover path, array-root schema on the OpenAI deployment

  1. Run
curl -s -o /tmp/out.json -w 'HTTP %{http_code}\n' http://localhost:56558/v1beta/models/openai-behind-native-route:generateContent -H 'x-goog-api-key: sk-1234' -H 'Content-Type: application/json' --data-binary @payload_array_root.json; jq '{top_level_keys: keys, text: .candidates[0].content.parts[0].text}' /tmp/out.json
  1. Still 200 with prose, a non-object root schema is not forwarded
HTTP 200
{
  "top_level_keys": [
    "candidates",
    "usageMetadata"
  ],
  "text": "- **Remy’s Ratatouille Adventure** – A fun, trackless 3D ride through Gusteau’s restaurant.\n- **Frozen Ever After** – A gentle boat ride featuring songs and characters from *Frozen*.\n- **Spaceship Earth** – A slow-moving journey through the history of human communication inside EPCOT’s iconic sphere."
}

Leftover path, JSON mime type with no schema on the OpenAI deployment

  1. Run
curl -s -o /tmp/out.json -w 'HTTP %{http_code}\n' http://localhost:56558/v1beta/models/openai-behind-native-route:generateContent -H 'x-goog-api-key: sk-1234' -H 'Content-Type: application/json' --data-binary @payload_mime_only.json; jq '{top_level_keys: keys, text: .candidates[0].content.parts[0].text}' /tmp/out.json
  1. Still 200 with prose, a mime type without a schema is not forwarded
HTTP 200
{
  "top_level_keys": [
    "candidates",
    "usageMetadata"
  ],
  "text": "**Remy’s Ratatouille Adventure** — a fun, trackless 3D ride with no height requirement, great for both ages 6 and 9."
}

google-genai SDK, OpenAI deployment

  1. Run
python sdk_leg.py http://localhost:56558 openai-behind-native-route
  1. response.parsed is a populated ParkTipResponse
parsed: park_name='EPCOT' summary='EPCOT blends family-friendly attractions with hands-on science, imaginative worlds, and international pavilions. It’s a strong choice for ages 6 and 9, with rides and activities that appeal to both kids and adults.' highlights=[Highlight(title='Must-do', detail='Ride Remy’s Ratatouille Adventure, a playful, trackless 3D experience that shrinks your family to Remy’s size.'), Highligh [...]
text starts with: '{"park_name":"EPCOT","summary":"EPCOT blends family-friendly attractions with hands-on science, imaginative worlds, and international pavilions. It’s a strong choice for ages 6 and 9, with rides and a'
text parses as JSON: True

Type

🐛 Bug Fix

Caveats (if any)

Low

  • A schema the provider rejects for another reason now returns the provider's 400, where before it was prose. Left as is: the earlier 200 ignored the schema the caller asked for, and the live probes show Anthropic and OpenAI accept every Gemini Schema keyword except property ordering, which is stripped
  • nullable: true passes through and Anthropic ignores it, so a required nullable field cannot come back null there. Left as is: rewriting it to anyOf with null is a Gemini-to-JSON-Schema translation layer this fix does not need, and OpenAI honors the key
  • Only application/json maps; text/x.enum and other mime types leave response_format unset, which is the merge-base behavior for them
  • A JSON mime type with no schema still yields prose on non-Gemini deployments (follow-up ticket LIT-8222)
  • Non-object root schemas (arrays, enums) are still not forwarded (follow-up ticket LIT-8223)
  • mock_response on this route still returns the top-level text (litellm/google_genai/main.py). Left as is: nothing asserts on it, no real request reaches it, and a commit now resets both bot verdicts plus the per-commit QA and live risk legs for a mock-only field
  • Three non-required CircleCI jobs (integration-extensions, logging_testing, integration-cost) are red here and on every main pipeline since 2026-09-19 23:48Z (an mcp 2.x import, a GCS pub/sub logging test, and a Fireworks cache-read price); none touches this adapter. Fixing them here would mean merging main in or patching unrelated tests, which resets the bot verdicts and the per-commit QA for checks the merge gate does not read
  • The non-required proxy_e2e_anthropic_messages_tests job is red on the two Bedrock test_all_beta_headers cases ("invalid beta flag"), which failed on every main pipeline until fix(test): unbreak the integration-cost and proxy_e2e_anthropic_messages CircleCI jobs on main #42048 (7966f50, LIT-8149) landed after this branch's merge base; the latest main pipeline passes them and neither case touches this adapter

Design decisions

  1. Only object-root schemas are forwarded, since OpenAI rejects a non-object root; wrapping the root in an object was rejected because it would change the JSON the caller gets back
  2. The json_schema name is the fixed response and strict is not set, so a schema passes without every property being required and without additionalProperties: false
  3. responseJsonSchema wins over responseSchema when both are set, since it is the caller's JSON Schema as written; the snake_case spellings are accepted because the google-genai SDK sends them. The picked schema is validated as a JSON object with Pydantic before it is forwarded, and anything else (a config object from the Python SDK, a schema that is not a mapping) falls through to the pre-fix prose behavior instead of failing the request
  4. Only propertyOrdering and property_ordering are stripped, at every nesting level, because Anthropic rejects them and OpenAI ignores them; nullable, format, example, title, and default pass through because both providers accept them, so no Gemini-to-JSON-Schema rewrite is added
  5. The schema is only sent when the deployment's provider lists response_format as supported, so openai/gpt-4 keeps returning prose instead of a 400; a provider that cannot be resolved fails open and sends the schema
  6. Tool parameters keep the existing lowercasing only; parametersJsonSchema wins over parameters, a null value falls through to the other key, and a value that is not a JSON object (an int, a string) is dropped as the merge base did, since forwarding it turned a 200 into a 500 in the live drive
  7. The top-level text is removed rather than kept, since the Gemini API has no such field, nothing in the repo reads it, and the SDK computes .text client-side

Live PR risk

Base f49fd22 and tip 5eb967d each ran as a real proxy with Claude, OpenAI, and Gemini deployments behind the native route, twelve REST cells plus the litellm.agenerate_content SDK entrypoint, with the outbound provider request captured from the proxy log and diffed per cell

  • Breaking: none at the tip. Every cell that returned 200 at the base returns 200 at the tip, one provider request each side. The first tip (47ebfa1) turned a tool declaration whose parameters is not a JSON object (5, "") from a 200 into a 500, since OpenAI rejected the body and the Anthropic transformation raised; the tip validates the declaration as a JSON object first and drops anything else exactly as the base did
  • Backward incompatible: the top-level text is gone from every response and stream chunk (nothing in litellm, the dashboard, or the docs reads it, and the SDK computes .text client-side), and the provider now receives tools[].parameters / input_schema and text.format / output_format, which is the fix
  • Regression risk: the outbound diff per cell contains only those two differences. Gemini-dialect declarations (OBJECT, STRING, propertyOrdering, nullable) are accepted by both providers and the forced call comes back with the declared park_id; a non-object, null, or 5 KB string responseSchema, a non-string responseMimeType, and conflicting camelCase and snake_case keys (camelCase wins) all keep the base behavior
  • Dependency graph: the adapter's callers outside its module are litellm/utils.py (function setup, wrapped in try/except), the v3 parallel request limiter (token estimation, fails open), the /v1beta endpoints, and litellm.generate_content; the legacy suites that reach it (tests/proxy_unit_tests/test_google_endpoint_routing.py, test_google_gemini_proxy_request.py, tests/llm_translation/test_openai.py::test_openai_via_gemini_streaming_bridge) pass at the tip; main moved 81 commits since the merge base, none on the adapter or its neighbors, and the merge is clean
  • Not verified: deployments other than OpenAI, Anthropic, and Gemini behind the native route (Bedrock, Vertex non-Gemini, Azure), text/x.enum and other non-JSON mime types, and a deployment reached through a wildcard or model group alias

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR
  • 5eb967d passes /live-pr-risk

@devin-ai-integration

devin-ai-integration Bot commented Sep 20, 2026 •

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@codspeed

codspeed Bot commented Sep 20, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_genai_adapter_response_schema_tool_params (5eb967d) with main (5417abd)

Open in CodSpeed

@greptile-apps

greptile-apps Bot commented Sep 20, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The PR appears safe to merge; no outstanding correctness, security, or repository-rule violations remain.

Findings

  1. P1 Response field removed ▶
  2. P1 Schema enforcement weakened ▶
  3. P1 Response schemas silently dropped ▶

Summary

This PR improves the Google GenAI generateContent adapter for non-Gemini deployments.

  • Converts supported object-root response schemas into OpenAI-compatible response_format values while removing Gemini-only ordering fields.
  • Forwards both parametersJsonSchema and parameters tool declarations, with the latest revision dropping malformed non-object schemas.
  • Removes the non-standard top-level text field from streaming and non-streaming responses.
  • Adds focused unit and adapter-level coverage for schema forwarding, tool parameters, unsupported models, and response shape.

Reviews (3) · Last reviewed commit: "fix(google_genai): drop non-object tool ..."

@@ -582,14 +650,6 @@ def translate_completion_to_generate_content(
),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Response field removed Both response modes drop text, breaking consumers without the required compatibility flag; updated tests now mask the regression by requiring its absence

Rule Used: What: avoid backwards-incompatible changes without user-controlled flags Why: This breaks current behaviour for users using existing functionality Example of BAD: this PR (#22164) introduced run_post_custom... (source)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The Gemini API defines no top-level text field; the SDK computes .text client-side from candidates parts, so nothing consumes it

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You're right. The top-level text field is not part of the Gemini generateContent response contract, and the google-genai SDK derives .text from candidates[].content.parts[]. Removing this adapter-only convenience field therefore does not break the supported Gemini response shape or SDK behavior. I’m withdrawing this finding.

return None
return {
"type": "json_schema",
"json_schema": {"name": "response", "schema": _strip_gemini_only_schema_keys(schema)},

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Schema enforcement weakened Google response schemas reach the provider without strict: true, so downstream responses may violate the caller's requested shape

Suggested change
"json_schema": {"name": "response", "schema": _strip_gemini_only_schema_keys(schema)},
"json_schema": {
"name": "response",
"schema": _strip_gemini_only_schema_keys(schema),
"strict": True,
},

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

OpenAI strict mode rejects schemas lacking additionalProperties false or every property required, which Gemini dialect schemas never carry; strict would break working calls

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You're right. OpenAI strict structured outputs require additionalProperties: false and every object property to be listed in required (including nested objects). Gemini schemas do not necessarily include those constraints, so adding strict: true here could turn currently valid requests into 400s. Since this adapter does not rewrite schemas into OpenAI's strict subset, strict should remain unset. I'm withdrawing this comment; no change is needed.

schema: Final = _validated(
_JSON_OBJECT_SCHEMA, normalize_json_schema_types(_first_present(fields, _RESPONSE_SCHEMA_KEYS))
)
if schema is None or schema.get("type") != "object":

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Response schemas silently dropped Array, string, and enum schemas return no response_format, leaving downstream generation unconstrained despite the caller's structured-output request

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

OpenAI requires an object root, so wrapping needs an unwrap on both response paths; non-object roots keep merge-base behavior and are tracked separately

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Understood. Since OpenAI requires an object-root response_format, forwarding array/string/enum roots would require wrapping them and then unwrapping the model output on both the non-streaming and streaming response paths, which is outside this focused fix. Keeping those roots on merge-base behavior and tracking them separately is reasonable; I’m resolving this comment.

@codecov

codecov Bot commented Sep 20, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@mateo-berri mateo-berri added run-ci and removed run-ci labels Sep 20, 2026
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri mateo-berri added run-ci and removed run-ci labels Sep 20, 2026
@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

@mateo-berri

Copy link
Copy Markdown
Contributor

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 5eb967d. Configure here.

@mateo-berri mateo-berri left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@mateo-berri
mateo-berri merged commit b444709 into main Sep 20, 2026
145 of 150 checks passed
@mateo-berri
mateo-berri deleted the litellm_genai_adapter_response_schema_tool_params branch September 20, 2026 02:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant