Managed batches fixes for vertex - #22464
Conversation
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
Greptile SummaryThis PR fixes several issues with Vertex AI managed batches and file operations: corrects the
Confidence Score: 3/5
|
| Filename | Overview |
|---|---|
| litellm/files/main.py | Added vertex_ai to afile_retrieve Literal type, but the sync file_retrieve it delegates to was not updated — will cause runtime failures for vertex_ai file retrieval. |
| litellm/llms/vertex_ai/files/transformation.py | Implements GCS file retrieve/delete/content operations. Delete response reconstruction omits the bucket name from the returned file ID, producing an incorrect gs:// URI. |
| litellm/llms/vertex_ai/batches/transformation.py | Correctly fixes GcsSource(uris=...) to pass a list instead of a bare string, matching the updated type definition. |
| litellm/types/llms/vertex_ai.py | Type fix: GcsSource.uris changed from str to List[str] to match the Vertex AI API spec. |
| litellm/batches/batch_utils.py | Simplified batch cost calculation by extracting usageMetadata directly instead of using VertexGeminiConfig transformation. Clean, correct approach using batch_cost_calculator. |
| litellm/llms/vertex_ai/batches/handler.py | Adds try/except around HTTP POST to log error body on HTTPStatusError. Minor style issue with redundant hasattr check and f-string in logger. |
| enterprise/litellm_enterprise/proxy/hooks/managed_files.py | Adds fallback to check litellm_metadata.model_info when top-level model_info is absent — fixes batch file ID resolution for managed files. |
Flowchart
%%{init: {'theme': 'neutral'}}%%
flowchart TD
A[afile_retrieve vertex_ai] --> B[file_retrieve sync]
B -->|"❌ Missing vertex_ai in Literal"| C[Routing Failure]
D[Batch Create Request] --> E[VertexAIBatchTransformation]
E --> F["GcsSource(uris=[file_id])"]
F --> G[Vertex AI API POST]
G -->|HTTPStatusError| H[Log error body + re-raise]
G -->|200 OK| I[Transform to LiteLLMBatch]
J[Batch Retrieve] --> K[Get output file content]
K --> L[calculate_vertex_ai_batch_cost_and_usage]
L --> M[Extract usageMetadata per line]
M --> N[batch_cost_calculator]
N --> O[Aggregate cost + Usage]
P[File Delete Request] --> Q[_parse_gcs_uri]
Q --> R["GCS Storage API DELETE /b/{bucket}/o/{object}"]
R --> S[transform_delete_file_response]
S -->|"❌ Missing bucket in gs:// URI"| T[Incorrect file ID returned]
Last reviewed commit: b16397a
| async def afile_retrieve( | ||
| file_id: str, | ||
| custom_llm_provider: Literal["openai", "azure", "gemini", "hosted_vllm", "manus"] = "openai", | ||
| custom_llm_provider: Literal["openai", "azure", "gemini", "vertex_ai", "hosted_vllm", "manus"] = "openai", |
There was a problem hiding this comment.
Sync file_retrieve missing vertex_ai provider
afile_retrieve now accepts "vertex_ai" (and "gemini"), but it delegates to the synchronous file_retrieve at line 337 which still only accepts Literal["openai", "azure", "hosted_vllm", "manus"]. Calling afile_retrieve(file_id, custom_llm_provider="vertex_ai") will hit the sync function with a provider value it doesn't recognise, causing a type mismatch and likely a routing failure at runtime.
The Literal type on file_retrieve (line 339) needs to be updated to include "gemini" and "vertex_ai" as well.
| file_id = "deleted" | ||
| if hasattr(raw_response, "request") and raw_response.request: | ||
| url = str(raw_response.request.url) | ||
| if "/o/" in url: | ||
| import urllib.parse | ||
| encoded_name = url.split("/o/")[-1].split("?")[0] | ||
| file_id = f"gs://{urllib.parse.unquote(encoded_name)}" |
There was a problem hiding this comment.
Reconstructed delete file ID is missing the bucket name
The URL format is https://storage.googleapis.com/storage/v1/b/{bucket}/o/{encoded_object}. When this code does url.split("/o/")[-1], it extracts only the object path, omitting the bucket name. So for a file originally at gs://my-bucket/path/to/file.jsonl, the returned ID would be gs://path/to/file.jsonl instead of gs://my-bucket/path/to/file.jsonl.
The bucket name should also be extracted from the URL (between /b/ and /o/) and prepended. For example:
| file_id = "deleted" | |
| if hasattr(raw_response, "request") and raw_response.request: | |
| url = str(raw_response.request.url) | |
| if "/o/" in url: | |
| import urllib.parse | |
| encoded_name = url.split("/o/")[-1].split("?")[0] | |
| file_id = f"gs://{urllib.parse.unquote(encoded_name)}" | |
| file_id = "deleted" | |
| if hasattr(raw_response, "request") and raw_response.request: | |
| url = str(raw_response.request.url) | |
| if "/o/" in url: | |
| import urllib.parse | |
| bucket_part = url.split("/b/")[-1].split("/o/")[0] | |
| encoded_name = url.split("/o/")[-1].split("?")[0] | |
| file_id = f"gs://{bucket_part}/{urllib.parse.unquote(encoded_name)}" |
| error_body = e.response.text if hasattr(e, 'response') else "N/A" | ||
| litellm.verbose_logger.error( | ||
| f"Vertex AI batch create failed: status={e.response.status_code}, body={error_body[:1000]}" | ||
| ) |
There was a problem hiding this comment.
Redundant hasattr check and unsafe f-string in logger call
httpx.HTTPStatusError always has a .response attribute (it's a constructor parameter), so hasattr(e, 'response') is always True — this guard is misleading. More importantly, using an f-string in the logger.error() call means the string is always formatted, even when the error log level is disabled. Prefer %-style formatting:
| error_body = e.response.text if hasattr(e, 'response') else "N/A" | |
| litellm.verbose_logger.error( | |
| f"Vertex AI batch create failed: status={e.response.status_code}, body={error_body[:1000]}" | |
| ) | |
| error_body = e.response.text | |
| litellm.verbose_logger.error( | |
| "Vertex AI batch create failed: status=%s, body=%s", | |
| e.response.status_code, error_body[:1000], | |
| ) |
|
Automated patch bundle from next-100 unresolved backlog expansion.\nGenerated due limited direct branch-write access; please apply/cherry-pick minimal edits below.\n\n## PR #22464 — Unresolved thread summary
Minimal patch proposals
|
Sameerlite
left a comment
There was a problem hiding this comment.
LGTM, just make the small change i have asked
| async def afile_retrieve( | ||
| file_id: str, | ||
| custom_llm_provider: Literal["openai", "azure", "gemini", "hosted_vllm", "manus"] = "openai", | ||
| custom_llm_provider: Literal["openai", "azure", "gemini", "vertex_ai", "hosted_vllm", "manus"] = "openai", |
Managed batches - Address PR bot comments from #22464
* add explicit caching to litellm proxy for gemini models via injection * fix: add missing `supports_function_calling` for deepinfra models All 55 deepinfra models that had `supports_tool_choice: true` were missing the `supports_function_calling` flag, causing `litellm.supports_function_calling()` to incorrectly return False. Fixes #22619 Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * Managed batches - Address PR bot comments from #22464 * feat(togetherai): add support for TogetherAI Qwen3.5-397B-A17B model * Agent Tracing - support context_id based trace id propogation + nested llm calls (#22626) * style(ui/): distinguish agent calls from llm calls on ui * feat: initial grouping working * feat: set stable contextid for a2a calls - allows for easily passing to downstream llm/mcp calls * feat(a2a_endpoints.py): fix tracing to avoid recreating logging objects for the same call allows stable trace id usage * fix(guardrail_endpoints): handle string ui_type values in _build_field_dict _build_field_dict unconditionally called .value on ui_type, which crashes for guardrail configs that use plain strings (e.g. BlockCodeExecutionGuardrailConfigModel uses "multiselect" and "percentage"). Now checks with hasattr before calling .value. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: propagate trace/session id from headers in MCP server calls Cherry-picked mcp_server/server.py fixes from 6feb9ba: adds get_chain_id_from_headers to extract x-litellm-trace-id / x-litellm-session-id from raw headers, and uses it in call_tool and list_tools to keep spend logs and tracing consistent with A2A. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> * [Feat] UI - Add Open in New Tab on leftnav Bar (#22731) * Add minimal dev_config.yaml for proxy development Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com> * feat(ui): wrap left nav items in <a> tags for open-in-new-tab support Nav items are now rendered as <a> elements with proper href attributes, enabling right-click → 'Open in new tab', Ctrl/Cmd+click, and middle-click to open any sidebar page in a new browser tab. Normal clicks continue to use SPA navigation (no full page reload). Applied to both leftnav.tsx (query-param routing) and Sidebar2.tsx (Next.js file-based routing). Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com> --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com> * [Feat] Add Tool Policies for AI Gateway (#22732) * fix: fix ui render * fix: fix minor bugs * refactor: use prisma functions instead of raw sql (safer) * fix(add-new-tiles-to-tool-policies): allow developer to see what's available * feat: ensure tool allowlist runs correctly for tool names + mcp's * refactor: more ui improvements * feat: working key tool blocking * feat(tools): show tool logs * refactor: backend code improvements * refactor: improve log viewer for tools * fix: address PR review feedback for tool access control - Add missing blocked_tools column to root schema.prisma (schema drift) - Invalidate ToolPolicyRegistry after policy mutations so changes take effect immediately - Remove dead code: unused get_effective_policies, get_tool_policies_cached, and helpers Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: race condition in permission resolution and remove duplicate allowlist check - Use atomic update_many with object_permission_id=None to prevent concurrent requests from creating orphaned permission rows and losing tool blocks - Remove duplicate allowed_tools enforcement from guardrail (already enforced in auth layer via check_tools_allowlist) - Move inline uuid import to module level Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * update to account for userAgent * UI - Add ToolDetails * input/output policy * LiteLLM_PolicyAttachmentTable * LiteLLM_PolicyAttachmentTable * fix: add _enqueue_tool_registry_upsert * fix: tool mgmt endpoints * tool mgmt endpoints * Update tests/test_litellm/proxy/db/test_tool_registry_writer.py Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> * Update tests/test_litellm/proxy/db/test_tool_registry_writer.py Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> * Update tests/test_litellm/proxy/db/test_tool_registry_writer.py Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> * fix: sync root schema.prisma and fix test_tool_registry_writer for input/output policy - Migrate root schema.prisma LiteLLM_ToolTable from call_policy to input_policy/output_policy, add missing user_agent and last_used_at columns (now consistent with litellm/proxy/schema.prisma and litellm-proxy-extras) - Fix SpendLogToolIndex comment across all three schema files - Fix all call_policy references in test_tool_registry_writer.py: swapped update_tool_policy arguments, wrong get_tools_by_names return type assertions, _mock_tool_row setting call_policy instead of input_policy Addresses Greptile review feedback on PR #22732. Made-with: Cursor --------- Co-authored-by: Krrish Dholakia <krrishdholakia@gmail.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> * feat(proxy): add key_alias, key_hash, requested_model DD APM span tags (#22710) * feat(proxy): add key_alias, key_hash, requested_model tags to DD APM spans * refactor(proxy): consolidate DD APM tag helpers into DDSpanTagger class * refactor(proxy): move DDSpanTagger to its own file litellm/proxy/dd_span_tagger.py --------- Co-authored-by: liweiguang <codingpunk@gmail.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> Co-authored-by: Ephrim Stanley <ephrim.stanley@point72.com> Co-authored-by: Varad Khonde <varadkhonde@gmail.com> Co-authored-by: Krish Dholakia <krrishdholakia@gmail.com> Co-authored-by: Sameer Kankute <sameer@berri.ai> Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com> Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com> Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
…es-feb27 Managed batches fixes for vertex
…es-mar3 Managed batches - Address PR bot comments from BerriAI#22464
…I#21881) * add explicit caching to litellm proxy for gemini models via injection * fix: add missing `supports_function_calling` for deepinfra models All 55 deepinfra models that had `supports_tool_choice: true` were missing the `supports_function_calling` flag, causing `litellm.supports_function_calling()` to incorrectly return False. Fixes BerriAI#22619 * Managed batches - Address PR bot comments from BerriAI#22464 * feat(togetherai): add support for TogetherAI Qwen3.5-397B-A17B model * Agent Tracing - support context_id based trace id propogation + nested llm calls (BerriAI#22626) * style(ui/): distinguish agent calls from llm calls on ui * feat: initial grouping working * feat: set stable contextid for a2a calls - allows for easily passing to downstream llm/mcp calls * feat(a2a_endpoints.py): fix tracing to avoid recreating logging objects for the same call allows stable trace id usage * fix(guardrail_endpoints): handle string ui_type values in _build_field_dict _build_field_dict unconditionally called .value on ui_type, which crashes for guardrail configs that use plain strings (e.g. BlockCodeExecutionGuardrailConfigModel uses "multiselect" and "percentage"). Now checks with hasattr before calling .value. * fix: propagate trace/session id from headers in MCP server calls Cherry-picked mcp_server/server.py fixes from 6feb9ba: adds get_chain_id_from_headers to extract x-litellm-trace-id / x-litellm-session-id from raw headers, and uses it in call_tool and list_tools to keep spend logs and tracing consistent with A2A. --------- * [Feat] UI - Add Open in New Tab on leftnav Bar (BerriAI#22731) * Add minimal dev_config.yaml for proxy development Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com> * feat(ui): wrap left nav items in <a> tags for open-in-new-tab support Nav items are now rendered as <a> elements with proper href attributes, enabling right-click → 'Open in new tab', Ctrl/Cmd+click, and middle-click to open any sidebar page in a new browser tab. Normal clicks continue to use SPA navigation (no full page reload). Applied to both leftnav.tsx (query-param routing) and Sidebar2.tsx (Next.js file-based routing). Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com> --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com> * [Feat] Add Tool Policies for AI Gateway (BerriAI#22732) * fix: fix ui render * fix: fix minor bugs * refactor: use prisma functions instead of raw sql (safer) * fix(add-new-tiles-to-tool-policies): allow developer to see what's available * feat: ensure tool allowlist runs correctly for tool names + mcp's * refactor: more ui improvements * feat: working key tool blocking * feat(tools): show tool logs * refactor: backend code improvements * refactor: improve log viewer for tools * fix: address PR review feedback for tool access control - Add missing blocked_tools column to root schema.prisma (schema drift) - Invalidate ToolPolicyRegistry after policy mutations so changes take effect immediately - Remove dead code: unused get_effective_policies, get_tool_policies_cached, and helpers * fix: race condition in permission resolution and remove duplicate allowlist check - Use atomic update_many with object_permission_id=None to prevent concurrent requests from creating orphaned permission rows and losing tool blocks - Remove duplicate allowed_tools enforcement from guardrail (already enforced in auth layer via check_tools_allowlist) - Move inline uuid import to module level * update to account for userAgent * UI - Add ToolDetails * input/output policy * LiteLLM_PolicyAttachmentTable * LiteLLM_PolicyAttachmentTable * fix: add _enqueue_tool_registry_upsert * fix: tool mgmt endpoints * tool mgmt endpoints * Update tests/test_litellm/proxy/db/test_tool_registry_writer.py Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> * Update tests/test_litellm/proxy/db/test_tool_registry_writer.py Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> * Update tests/test_litellm/proxy/db/test_tool_registry_writer.py Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> * fix: sync root schema.prisma and fix test_tool_registry_writer for input/output policy - Migrate root schema.prisma LiteLLM_ToolTable from call_policy to input_policy/output_policy, add missing user_agent and last_used_at columns (now consistent with litellm/proxy/schema.prisma and litellm-proxy-extras) - Fix SpendLogToolIndex comment across all three schema files - Fix all call_policy references in test_tool_registry_writer.py: swapped update_tool_policy arguments, wrong get_tools_by_names return type assertions, _mock_tool_row setting call_policy instead of input_policy Addresses Greptile review feedback on PR BerriAI#22732. Made-with: Cursor --------- Co-authored-by: Krrish Dholakia <krrishdholakia@gmail.com> Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> * feat(proxy): add key_alias, key_hash, requested_model DD APM span tags (BerriAI#22710) * feat(proxy): add key_alias, key_hash, requested_model tags to DD APM spans * refactor(proxy): consolidate DD APM tag helpers into DDSpanTagger class * refactor(proxy): move DDSpanTagger to its own file litellm/proxy/dd_span_tagger.py --------- Co-authored-by: liweiguang <codingpunk@gmail.com> Co-authored-by: Ephrim Stanley <ephrim.stanley@point72.com> Co-authored-by: Varad Khonde <varadkhonde@gmail.com> Co-authored-by: Krish Dholakia <krrishdholakia@gmail.com> Co-authored-by: Sameer Kankute <sameer@berri.ai> Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com> Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com> Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
Relevant issues
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
tests/litellm/directory, Adding at least 1 test is a hard requirement - see detailsmake test-unit@greptileaiand received a Confidence Score of at least 4/5 before requesting a maintainer reviewCI (LiteLLM team)
Branch creation CI run
Link:
CI run for the last commit
Link:
Merge / cherry-pick CI run
Links:
Type
🆕 New Feature
🐛 Bug Fix
🧹 Refactoring
📖 Documentation
🚄 Infrastructure
✅ Test
Changes